Published September 22, 2026 — by Vito Ruocco
Introduction: The Frontier Just Got an Open-Source King
When a Chinese tech giant like Xiaomi claims to have built the strongest open-source model in the world, you pay attention. When they put a live cost counter — $3.47 million — right on the landing page, you start to believe it. And when a seasoned AI tester like yours truly spends an afternoon clicking through every single demo on the release page instead of reading benchmark tables, you know this is something different.
Xiaomi has just released MiMo V2.6, and the company is not mincing words: this is the most capable open-source model on the planet, at least according to the independent Artificial Analysis Intelligence Index v4.3. After spending significant time with every demo, every benchmark, and every interactive element on the release page, I can tell you this: the hype is real — and remarkably, it’s all accessible right now.
In this article, I’ll take you through everything MiMo V2.6 offers: the raw numbers, the live-trained reinforcement learning pipeline, the jaw-dropping demos — from playable city builders to robot arm simulations to formal mathematical proofs — and what it all means for the AI landscape in late 2026.
🎥 Watch the full video walkthrough on my channel:
What’s in the Box: Pro, Flash, and Ultra Speed
MiMo V2.6 ships as two natively multimodal models, plus a speed-optimized serving mode:
- MiMo V2.6 Pro — The flagship. Their most capable model, designed to push the frontier on every metric that matters.
- MiMo V2.6 Flash — The balance of intelligence, speed, and cost. Ideal for production workloads where latency matters.
- Pro Ultra Speed — The same Pro checkpoint, served up to 20 times faster at the same quality. This is not a separate model — just an optimized inference pipeline that will matter enormously for real-time applications.
The entire release is framed as a step on what Xiaomi calls the RSI Path — Recursive Self-Improvement through scalable reinforcement learning on tasks that can be verified. The model keeps pushing its own frontier through exploration and feedback. It’s a philosophy, not just a release.
The Numbers: Benchmarks That Tell the Real Story
Let’s get the benchmarks out of the way — because the truly impressive part comes after them.
On Encoding, MiMo V2.6 Pro scores 71.9 on DPS 1.1. For context, Claude Opus 5 and GPT-6 Astra sit at 74. So it’s two points behind the absolute frontier — and it’s open. On Program Bench, Pro scores 26.5 against Opus at 37 — a real gap that Xiaomi acknowledges honestly. On their in-house MiMo Code Bench, it ranks second only to Opus.
The General Intelligence section is where things get interesting. On GDP Value — an aggregated rating from Artificial Analysis on real professional work — Pro scores 163, right behind Claude Fable 5.1 and Opus. On Automation Bench, it scores 53.1, which is actually above Opus and GPT-6 Astra. The weak spot is Terminal Bench 4.0 at 34.9 against 59.6 for Astra — and I appreciate that Xiaomi left this on the page instead of hiding it. Honesty in AI benchmarks is rare, and it deserves respect.
On Visual and Cyber benchmarks, MiMo V2.6 sits in the middle of the pack on visual coding, but on Cyber Gym both models are at the top — 94 and 95.1 — well ahead of everything else listed.
The headline claim comes from the Artificial Analysis Intelligence Index v4.3: MiMo V2.6 Pro scores 46.32, putting it ahead of KIMI 3.8 and maintaining its position as the highest-scoring open-source model. You can see it in the bar chart on the release page — the vermillion bar with the arrow, sitting comfortably among the closed-frontier models. The chart that really matters, though, is the second one: Intelligence vs. Cost per Task on a log scale. Pro sits right on the Pareto line — the attractive quadrant where you get frontier intelligence at a fraction of the price. The reason is simple: Xiaomi kept the exact same API prices as V2 and just made the model smarter.
Live RL Training: Streaming the Frontier in Real Time
This is the section where Xiaomi did something nobody else does. They ran the final reinforcement learning stage live, in public, and streaming. In under six days, Flash and Pro each completed 30 RL steps over roughly 750,000 projects. The cost: approximately $850,000 for Flash and $2.62 million for Pro.
The results are visible in real training curves. The pass rate on the training task went up 25% and 12% respectively. On DSWE — which was held out as a validation set — Pro went from 50.4 to 72.6. You can hover over any step and read the value. DSW4 Pro starts at 58.4 at step one and ends at 72.6 at step 30. Automation bench goes from 46.5 to 55. Visual coding from 68.5 to 74.9. The lower band shows total tokens per task, and you can see the model getting more verbose on coding tasks as it improves while staying flat on general workflows.
Xiaomi scaled along three axes: bigger batches on a fully asynchronous setup (1,568 samples per update at up to 1 million context), more tasks and environments across coding, agentic, visual, and cyber domains, and more compute — comparing answers inside a group to get a sharper reward signal. They froze the router to stop training drift and built a whole defense against reward hacking. And they say they’re open-sourcing all of it: the technical report, the training environments, and the RL code.
The live streaming page even has a notices log, and honestly, this is my favorite thing on the whole site. “Pro run restarted at step 17 because of a GPU out of memory error from expert load imbalance — network issue between the training cluster and the grader — removed the cyber dataset because they saw bad patterns in the rollout logs.” That’s what training a frontier model actually looks like, and they left it all up.
From Vibe Coding to Vibe World
This section of the release page is titled “From Vibe Coding to Vibe World”, and the idea is that the model’s coding ability has generalized. Give MiMo V2.6 a goal, a harness, and time, and it builds, tests, and ships. V2.6 adds 3D spatial reasoning, multimodal perception, and a computer-use agent. Natural language no longer stops at software — it extends into interactive worlds.
Game Development
Given an image, a video, or a text prompt, MiMo V2.6 splits the request across multiple agents that build a 3D scene, write interaction logic, and visually verify the result. The demos are spectacular. A full city builder — Isla Presidente in the colonial era — complete with treasury, citizens, workers, tourists, approval rating, a directive panel, a mini-map, and a build bar with a dozen buildings. That’s a fully tropical-style game, generated. Then there’s a third-person open-world demo: a character on horseback riding through tall grass with mountains in the distance, wind moving through the field, complete with a Unity engine watermark — meaning this was built inside a real game engine, not just a browser canvas. And an assassin-style climbing game with a character on top of a tower in a desert city, camera orbiting around to show the entire city laid out below.
3D Modeling Inside Blender
MiMo V2.6 works inside Blender from a text description or reference image, producing actual 3D assets you could animate, 3D print, or drop into a game. The demos include a red sports car with racing stripes turning in the viewport with proper materials and reflections, a three-masted sailing ship with rigging and sails on a display stand, a ginkgo tree with thousands of tiny leaves and a proper trunk, and an architectural visualization — a modern house at night with warm interior lights, a pool, and full camera control. This is the kind of output an architectural studio charges real money for.
Embodied Simulations
This is where it gets deeply technical. A Franka Panda robot arm receives camera feeds directly, with the model reasoning about and controlling the arm in a closed loop. The task is on screen: “Use the gripper to place the red block on the red pad and the blue block on the blue pad.” On the left, you can literally read the model’s reasoning as it executes — checking that the block is on the pad from the wrist camera. On the right, the action output shows position deltas and gripper state. A second task involves picking up a black bowl between a plate and a ramekin and placing it on the plate, with the reasoning showing the model considering the rim of the bowl and where the fingers should close. Outcome: task success. Yes, it’s a simulation rather than a real arm, and the page is honest about that — but the loop is real.
Visual Design, Presentations, and Music
MiMo V2.6 claims that one instruction becomes a complete front-end or slide deck with structured layout, real components, interaction, and animation. The model is fluent with Figma and with image and video generation tools. And the caption under the gallery says something revealing: “MiMo has its own taste” — a dig at every model that produces the same pale purple gradient landing page.
And these are not screenshots. They are live HTML pages loaded in iframes. A creative studio called Salazar with an orange planet morphing in the hero section — fully scrollable. A creative technology studio called Mantis with a WebGL 3D object rotating in the hero, a fake system readout with coordinates and a frame counter, and text that decrypts itself as you scroll. A climate journalism site about Central Asia with an editorial layout, serif headlines, numbered section navigation, photo captions, and language switches for English, Russian, and Kazakh. A design director’s portfolio with a live-looking airport dashboard showing on-time percentages, open exceptions, SLA risk, and a data table. Twelve sites in total, each with its own typography, its own color system, and none looking like a template.
Presentation design: three decks from one-line briefs. An internal engineering training called “How an LLM Actually Learns” — 14 slides covering parameters, embeddings, attention, backpropagation, all laid out with diagrams and a consistent dark theme. A research note on the critical minerals supercycle (in Chinese) with a completely different editorial look. And a design review proposing three routes for a company’s first annual report. Three briefs, three styles.
Video clips: MiMo handles visual design, shot sequencing, music, and beat-synced editing end-to-end. Educational content about Fourier composition — a wave on the left, its frequency components on the right, with the animation building the idea step by step. The narration is generated with MiMo V2 TTS and synced to the visuals.
Music composition: Pro was asked to compose an orchestral piece for about ten instruments. It wrote the score, then converted it to MIDI itself. The piece, called “Night Road,” features strings, oboe, clarinet, trumpet, trombone, French horn, glockenspiel, and timpani. Two piano pieces in A minor — one slow called “Fading Light” and one moderate called “Quiet Soliloquy” — complete the musical demo suite.
Scientific Research: Catching Forever Chemicals and Proving Chaos
The part that surprised me most was the research section. Xiaomi’s material scientists asked Pro to design a brand-new metal-organic framework capable of capturing PFAS — the “forever chemicals” that contaminate water supplies worldwide. The model searched the literature and patents, proposed hypotheses, checked novelty, then ran dry experiments. It called open-source computational tools, set up the simulation environment itself, and calculated binding strengths. The result is a rotatable 3D model of the framework with the PFAS molecule caught inside — drawn in vermillion. Three candidates (A50, B50 with a CF3 biphenyl ligand, and C50), with absorption energy updating as you switch between them: -2.78 electron volts for the first one. These geometries come from the model’s own optimization runs.
The second research demo is pure mathematics: the model helped formalize the full main theorem of Li and Yorke’s 1975 paper — “Period Three Implies Chaos” — in Lean 4. Over 6,000 lines of Lean code, verified by the kernel. No unfinished placeholders. And the model had no Lean-specific post-training. The diagram shows the logistic map bifurcating into chaos with the period-3 window clearly marked. The full Lean code is downloadable from the page.
Availability and Pricing
Both models are live today. You can access them through AI Studio, MiMo Code, MiMo Desktop, the MiMo API platform, and OpenRouter. The desktop app — previously in early access — is now in its first official release with Pro and Flash built in, pitching itself as an all-in-one agent for slides, docs, sheets, and file organization with smart scheduling.
One important note: the app is not yet available in South Korea, the UK, or EU member states. So if you’re in Europe like me, the desktop app isn’t available yet — but the API and OpenRouter are fully accessible.
Pricing is where the benchmarks really matter:
- Flash: $0.14/million input tokens, $0.28/million output
- Pro: $0.435/input, $0.87/output
- Ultra Speed (20x faster): $0.0435/input, $0.087/output
- Cache hits: Nearly free at $0.0033/million on Flash; cache writes are free for now
Compare that to what you pay for a frontier closed model and you’ll understand the Pareto chart. On OpenRouter, Xiaomi has eight models listed, and the three new ones are already there — Pro, Flash, and Pro Ultra Speed, all with a 1-million-token context window. The listing tells you Pro is a model of over a trillion parameters and Flash is a 309-billion-parameter mixture of experts. Nothing to sign up for — just switch the model name.
And on Hugging Face, the collection has the RL checkpoints for both Pro and Flash, plus a distilled 9-billion-parameter model based on Qwen that you can run locally. The weights are actually open. This is not an API-only release.
The Honest Verdict: Where MiMo V2.6 Shines and Where It Doesn’t
The appendix table on the release page gives the honest view. Pro is best or near-best on the agentic coding rails. It’s competitive on GDP value. And it’s clearly behind on the security exploitation benchmarks — Exploit Bench at 47.9 against 100 for GPT-6 Astra, and Terminal Bench 4.0 at 34.9. If your work is long terminal sessions, the frontier closed models are still ahead. For everything else, this is a fraction of the price.
What convinced me was not the table. It was clicking through the page. A playable city builder. A Blender car. A robot arm whose reasoning you can read. Twelve websites that each look designed by a different human. A formal proof in Lean. And a live cost counter from the training run that included GPU-out-of-memory errors and network issues — the honest messiness of real frontier research.
MiMo V2.6 is, right now, the best open-source model on the independent index. And the demos prove it — not through marketing claims, but through actual, interactive, downloadable, verifiable work.
🎥 Watch my full walkthrough of every demo on the MiMo V2.6 release page:
All links to every demo are in the video description. If you want more content where I actually test these models instead of reading press releases, subscribe and let me know which demo you want me to rebuild from scratch.
Article by Vito Ruocco. Video and original testing by Vito Ruocco. Cover image generated with AI assistance.