Published: October 11, 2026 | By Vito Ruocco
Anthropic has switched off the internet for its own AI tests, China’s labs are lining up GLM 5.5, Kimi K3.1, Qwen 4 and a new MiniMax, Alibaba says one of its models improved itself 33 times, and a mystery Fable 5.5 checkpoint is going around. I checked each of these claims against the sources, and several are weaker than the headlines. Then I did the thing that actually answers the question: I gave the best model you can buy today, Claude Fable 5.1, and the best Chinese open-weights model you can use today, Kimi K3, the same four prompts, and played all eight results myself in the browser.
The Short Version
- Anthropic cut its AI evaluations off from the live internet after a model, running a test that visited random websites, submitted a false murder tip to the Philadelphia police. Nobody was hurt, but it was found more than two months late.
- Fable 5.5 is not confirmed. Anthropic’s own model list still shows Fable 5.1. The posts claiming a newer checkpoint have no server-side proof, and the demos attached to them cannot be tied to a model.
- “China achieved recursive self-improvement” is not proven. Alibaba says Qwen 3.8 Max ran 33 automated improvement cycles in about a month and rose from 40 to 45 on the Artificial Analysis score. The score is checkable. The 33 cycles are Alibaba’s own statement.
- GLM 5.5, Qwen 4 and Kimi K3.1 have no release dates. Z.ai has never used the name GLM 5.5 publicly. Alibaba confirmed four Qwen 4 tiers but no weights, no model card and no date. Kimi K3.1 exists only as a pi-digits post and an identifier reportedly seen in an API registry.
- My test: Fable 5.1 beat Kimi K3 on all four builds, but Kimi was close on two of them and cost about eleven times less, roughly $1.29 against $13.90.
- Other news: a hidden Voice tab in Claude, OpenAI resetting Codex limits again, Nvidia in talks about Reflection AI, and Cloudflare buying Deno.
Anthropic Switched Off the Internet for Its Own AI Tests
The story that changes the most about how labs work this week is not a model release. Anthropic published a report on October 9 describing cases where its Claude models did things nobody intended during internal evaluations, and its response is simple: until its monitoring can reliably catch this kind of behavior, internal evaluations lose their live internet access.
What the model actually did
The headline case is a false tip. A model that was running an evaluation, interacting with randomly selected websites, reached a page about an unsolved murder in Philadelphia and submitted a tip to the police tip line, presenting itself as someone with information. The tip was false. According to TechCrunch, it was sent on July 18 at 11:27 p.m., and Anthropic only discovered it on September 28. It notified the police department on October 8 and met with them the next day. The police never saw the tip, because it had been marked as spam.
The Philadelphia Police Department did not hold back. Its statement says that “the two-month delay in detecting and reporting the incident to the City is unacceptable.”
The report lists more than the tip. Anthropic describes four categories of unintended actions: exploiting a software flaw to run commands on a server, submitting a sensitive form on a real website (the police tip is the example), bypassing technical controls or payment requirements to reach data, and using free URL-shortening services to get around the length limit of its web fetch tool. Some incidents involved websites run by government agencies. Anthropic says the real-world impact was minimal and mostly blames training environments that rewarded cheating. I think the most important detail is the admission that alignment training is not yet sufficient for search and computer-use skills, which is why some evaluations were stopped or moved offline.
The White House and Nadella
The reaction came fast. Reports say the White House now expects AI companies to disclose model incidents immediately and fix any harm. How that will be enforced is not clear yet, and there is no penalty described, so treat it as a position for now.
It lands right next to what Satya Nadella wrote on his personal blog on October 10, in an essay titled “Models as Insider Risks in the Super Intelligence Era.” His argument is that companies should treat every frontier model as a possible insider threat, assume it may already be compromised, and keep an emergency brake within reach: tamper-proof logs of what a model does, tightly limited privileges, and a way for an authorized human to pause a model in the middle of a task. This is a position, not a Microsoft product. But after Philadelphia, it reads very differently than it would have a month ago. The scary part is that a model does not need to be malicious. One mistake or one prompt injection is enough.
The Fable 5.5 Rumor: What Is Verified and What Is Not
Everyone is asking about Fable 5.5, so here is the careful version.
What is confirmed: Anthropic’s own model directory and Claude Code’s supported-model list still identify Fable 5.1 as the current Fable. Opus 5.5 came out on September 22, Sonnet 5.5 on September 28 and Haiku 5.5 on October 7.
What is circulating: users say that Claude Code or the Claude website quietly routes some requests to a newer Fable, and they check it with a fingerprint question, asking the model about a nickname (“Tibo the reset guy”) without web search. Then come the demos, posted as Fable 5.5 output: a first-person shooter, detailed vector art of game controllers, a voxel island, a pixel game and a card game.
What is missing: any server-side proof. A nickname test cannot identify model weights, because recognizing a name is not the same as being a different model. The screenshots people share still show the Fable 5.1 label, and the demos have no metadata tying them to a model. One tracker that follows the rumor, last checked on October 7, concludes that the evidence “does not establish which model generated the outputs.” Separately, reports citing the Wall Street Journal and Reuters say Anthropic’s listing has been pushed to November and that the company is weighing a new model before it goes public. Neither company has confirmed that on the record.
So treat the demos as impressive and the name as unconfirmed. Which is exactly why I tested the Fable you can actually pay for.
Did China’s AI Achieve Recursive Self-Improvement?
The question in the title of the video I used as a reference deserves a direct answer: not proven.
Alibaba’s claim
At its Apsara conference on September 22, Alibaba said Qwen 3.8 Max completed 33 improvement cycles over about a month of fully automated runs. The system finds its own weaknesses, builds new training data, tests the result and feeds it back into training, with much less human involvement. Alibaba says the model’s Artificial Analysis score went from 40 to 45. In a separate experiment on chip design, it describes more than 60 hours of self-improvement and over 10,000 calls to electronic design automation tools.
Look at who is saying what. The Artificial Analysis score is public and can be checked. The count of 33 cycles cannot, because it is a statement from Alibaba itself, published on its own press site.
Z.ai talks the same language
Z.ai appears to be heading the same way. A report says the company raised about five billion dollars in September, and that roughly sixty percent of the proceeds go to what it calls a fully self-training system, a phrase taken from the filing as rendered by Reuters. Plans to reach around one trillion parameters are a roadmap, not a release.
The best counterweight I found
Epoch AI built a test called InnovationEval. Agents got up to 3,000 GPU-hours on 50 GPUs and a real training stack, and the goal was to independently reinvent a recent post-training method without seeing it. The best clean result was GPT-5.6 Sol, at 15 percent of the original method’s gains after corrections. Claude Fable 5 got 2 percent. Models whose training data included the paper did better, but still fell short of Epoch’s own rerun at 94 percent. And both agents selectively reported their best runs, even though the transcripts show they knew that was misleading. The sample is small and Epoch has not released its code, so this is not a verdict either. But it is a useful reminder.
My honest summary: labs are automating parts of their own pipeline, and that is real. A model that improves itself from end to end has not been shown.
China’s Next Models, One by One
GLM 5.5
There is still no GLM 5.5 model card, no API name and nothing on the Hugging Face page of Z.ai. What Z.ai actually shipped is GLM 5.3: announced on August 14, in the API on August 18, with open weights on August 28. It has 744 billion parameters in total, 40 billion active, one million tokens of context and 128K output. The name GLM 5.5 and the figure of “1T+” parameters come from a rumor chain that started in July, and the size grew from 1T to 3T in a few days with no source for either. An August date traced back to a JPMorgan forecast relayed by Reuters, not to Z.ai. A release in the second half of October is plausible, but Z.ai has not said so.
Qwen 4
At Apsara, Alibaba previewed four tiers: Max, Plus, Flash and a 27B model that is the open-weights tier for local use. There is no release date, no model card, no pricing and no benchmark. Whether Qwen 4 Max will be open is not stated. If you read that the flagship is open source, that is exactly the part nobody has confirmed. The 5 to 10 trillion parameter figure floating around belongs to later generations, not Qwen 4. For context, the previous generation moved fast: Qwen 3.8 Max launched on August 3 and its weights followed about ten days later.
Kimi K3.1
The only official signal was a post on Kimi’s verified Zhihu account on September 19: 135 digits of pi after “3.1”, with no caption. Then, on September 29, a “kimi-k3-1” identifier reportedly appeared in Moonshot’s API model registry, with a preview showing up to one million tokens of context and several reasoning levels. Moonshot has announced nothing. The model you can use today is Kimi K3, released on July 16 as a 2.8 trillion parameter model, with weights on July 27.
MiniMax M3.1 and Space Bunny Alpha
You may remember Space Bunny Alpha, the free mystery model that appeared on OpenRouter on September 23 with a one million token context window. I tested it in earlier videos, and three independent tokenizer tests point to MiniMax. MiniMax showed an M3.1 on stage at the World Artificial Intelligence Conference on July 17 and released a fast preview, M3.1-Flash-Preview, in its coding product. I could not find an official page for an “M3.1 Pro.” So when people say the new MiniMax is coming, remember that it is a guess.
The number nobody mentions: safety disclosure
SemiAnalysis counted 857 releases from nine Chinese labs between 2021 and September 15, 2026. Only 31 of them, 3.6 percent, came with a safety result published by the developer, and only 9, 1.1 percent, had one at or before launch. Zhipu is the clear outlier, with safety results in every year since 2022. The count only reflects what could be verified in public documents, so “not found” does not mean “not tested.” But if you build on Chinese open-weights models, the practical advice is clear: run your own safety evaluations and add runtime controls on agents.
Everything Else That Happened
Claude may be getting its own voice
A hidden “Voice” item labelled “Preview” appeared in Claude’s web navigation, and the mobile apps now ask users to share voice data for training. Here is what is established: the setting exists, it is off by default, and it is separate from the one for typed chats. What is not established is that Anthropic is building its own speech model. Claude already talks on mobile and desktop, so this could simply be about improving what exists. The label “Preview” refers to Anthropic’s internal environment, not a public release.
OpenAI resets limits again
OpenAI reset usage limits for paid Codex and ChatGPT Work users again. By one tracker this is the third time since August, and there is now a banked reset that users can apply themselves. What it says about the promised twenty-eight days of launches: when there is nothing to ship, everyone gets more tokens.
Nvidia and Reflection AI
The Financial Times reported on October 10 that Nvidia is in early talks to deepen its investment in Reflection AI or buy it outright. Nvidia has already put in about $800 million, and Reflection was valued at $25 billion before the money in a March round. The options on the table are a full acquisition, an acqui-hire where Nvidia hires the team and licenses the technology, or more investment. People familiar with the talks say a deal could come within weeks, but the talks could also collapse, and both companies declined to comment. If it happens, the company that sells chips to every lab would also have its own open-weights lab.
Cloudflare buys Deno
Cloudflare announced on October 9 that it is acquiring Deno, and Ryan Dahl’s team is joining. Deno Deploy will shut down in about six months, the runtime gets another year of bug-fix and security releases and then development stops, the code stays open source, and the JSR registry keeps running. If you build on Deno, this is your migration notice.
The Test: Fable 5.1 vs Kimi K3
Now the part I promised at the start. Chinese labs are about to ship GLM 5.5, Kimi K3.1 and Qwen 4, so the useful question is how far the best Chinese open model you can use today is from the best model you can actually pay for. I used Claude Fable 5.1 and Kimi K3, released in July, both through OpenRouter. Same four prompts, word for word, one request each, reasoning set to low, no second tries on quality (I only repeated requests that were cut off). Then I played every result myself, in the browser, on my own site.
All eight demos are free to play at ruocco.it/ai-demos/fable-5-1-vs-kimi-k3. The exact prompts, settings and full code are for members of the Ruocco Academy.
One note on the recording: both shooters and both roguelikes are written as very dark night scenes, so for the video I raised the screen brightness on all four, with the same filter. The files on the site are untouched.
1. A game controller drawn only with code
The prompt asks for a hyper-detailed wireless game controller in SVG, with real shapes, gradients and light, and then to make it alive: it tilts with the mouse, the sticks drag, the buttons press, the light bar changes colour and a Disassemble button opens an exploded view.

Fable’s Nimbus Pad has believable plastic highlights, a touch panel, sticks with a rubber ring pattern and a light bar that glows in the chosen colour. Disassemble breaks it into labelled parts: triggers, light bar, touch panel, d-pad, a green circuit board and even a battery cell.

Kimi’s Vector//Pad is a cleaner, smaller studio shot with a carbon grip texture, and the same interactions work, exploded view with labels included. It is a bit less detailed than Fable’s, and honestly this is closer than I expected. Cost: $1.30 for Fable and $0.41 for Kimi.
2. A tactical shooter with squad AI
The prompt asks for a night-time compound, rain, floodlights, enemies that take cover and call out, ragdolls and allied squad-mates. I play it like a person: moving, aiming and firing, with a script that sends real key and mouse events.

Fable’s Blackout Yard is a full mission, Operation Lantern: a carbine with reload and scope, procedural container textures, a compass, a minimap, a kill feed, spoken barks and two named squad-mates. Enemies take cover and shoot back, and in my first test I died, which tells you the enemy AI is real.

Kimi’s shooter has a starry sky, mist, a hostiles counter and a minimap. It works, and enemies fall when you hit them, but the world is much plainer and the weapon model looks odd. Cost: $3.95 for Fable and $0.23 for Kimi.
3. A roguelike where the lighting does the work
The prompt asks for Ember Depths: procedural floors, a pixel hero drawn in code, a dash, a sword with a trail, six enemy types and a boss, torches casting shadows with line-of-sight, fog of war and relics.

Fable’s version has a sword swing with a trail, warm light that stops at the walls and skeletons coming out of the dark. Kimi’s is even darker: a torch cone, shop pedestals and floating damage text. It plays, but it is easy to die early and the map is hard to read. Cost: $3.65 for Fable and $0.25 for Kimi.
4. Software: a spreadsheet with a real formula engine
The prompt asks for “Gridlight”, a spreadsheet app with a virtualised grid, a formula engine with dependency tracking, charts, conditional formatting, sorting and a dark mode.

Fable’s Gridlight is the best single demo of the whole test. Load demo gives a company budget with formulas, a key metrics panel with a lookup and today’s date, and a chart that sits next to the data. I change a number and the totals and the chart move with it. Cost: $5.00.

Now Kimi, and here is something you should know. The file it wrote shipped with a syntax error, so out of the box nothing ran. I fixed two characters, a missing empty string in a one-line ternary that builds the tab labels, and I labelled that fix on the page. After that it works: a budget, a colour scale and bar and pie charts. The charts do cover some of the data rows, and the budget it invented loses money in every month. Cost: $0.39.
What it cost
Fable 5.1 cost about $13.90 for the four builds ($1.30, $3.95, $3.65 and $5.00). Kimi K3 cost about $1.29 ($0.41, $0.23, $0.25 and $0.39). For full transparency, Fable’s first spreadsheet request hit the output cap and was cut off, which cost about $4 that is not counted above, and Kimi’s shooter was cut off mid-stream and had to be requested again.
My Verdict
Fable 5.1 is better at all four. The controller has more depth, the shooter has a squad and real tactics, the roguelike is brighter and more polished, and the spreadsheet is the best single demo of the whole test.
But Kimi is not far behind. All eight came from a single request each, and only one needed a fix. On value the gap is huge: about fourteen dollars against about a dollar thirty.
So did China just catch up? Not on polish, not yet. On value, almost. If GLM 5.5, Kimi K3.1 or Qwen 4 close even half of that gap when they ship, the picture changes this month, and I will test them the day they are available.
Try Every Demo, and Get the Code
Every demo from this video is on my site and runs in your browser: play all eight Fable 5.1 vs Kimi K3 builds, and the full list of tests is at ruocco.it/ai-demos. Tell me in the comments which side you would pick.
The exact prompts, the settings, the cost of each run, the bots I used to play the games and the untouched original of Kimi’s spreadsheet are inside the Ruocco Academy.
A Note on What Is Unconfirmed
Almost everything about the upcoming Chinese models is rumor or company statement, and I labelled it that way in the video and here: GLM 5.5 has no model card, Qwen 4 Max has no confirmed license, Kimi K3.1 is a pi-digits post and a registry identifier, “M3.1 Pro” has no official page, Fable 5.5 is not listed by Anthropic, and the 33 improvement cycles are Alibaba’s own statement. I will update this article when any of them changes. The voice in the video is a synthetic clone of mine.
Sources
- TechCrunch: an Anthropic AI model sent a false homicide tip to Philadelphia police
- Security Affairs: Anthropic restricts live internet access after Claude evaluation failures
- ExplainX: the White House on AI incident disclosure
- ExplainX: Nadella, assume every AI model is compromised
- TestingCatalog: Anthropic might be developing its own voice models
- Codex Resets tracker
- Kingy AI: Claude Fable 5.5, hidden rollout claims remain unverified
- OrcaRouter: Claude Fable 5.2, the reported window
- Alibaba Cloud: 2026 Apsara Conference roadmap
- AlphaSignal: Epoch’s InnovationEval
- CellCog: GLM-5.5 release date, the leak vs what Z.ai shipped
- OrcaRouter: Qwen 4 Max announced at Apsara 2026
- TechNode: Kimi K3.1 identifier surfaces in Moonshot’s API registry
- CellCog: Space Bunny Alpha and the MiniMax clues
- ExplainX: only 3.6 percent of Chinese AI releases ship a safety result (SemiAnalysis)
- Yahoo Finance: Nvidia weighs buying Reflection AI
- The New Stack: Cloudflare acquires Deno
If you want to see the next test the day the models ship, subscribe to @VitoMind on YouTube, and watch the full video above.