Xiaomi's MiMo-V2.6 Targets Agents, Cybersecurity and 3D Worlds in Fully Open-Source Release
Key Takeaways
- •MiMo-V2.6-Pro claimed the top rating among open-source models on the Artificial Analysis Intelligence Index with a score of 46.32, surpassing Kimi K3 and Qwen3.8 Max.
- •The series keeps API pricing at its V2.5 predecessor's levels, extending the intelligence-versus-cost frontier without raising prices.
- •On the DeepSWE v1.1 software engineering benchmark, MiMo-V2.6-Flash jumped to 67.9 from MiMo-V2.5-Pro's 19.0, while Pro scored 71.9, trailing DeepSeek V4.1 Flash and the Claude and GPT leaders.
- •The live-streamed reinforcement learning run finished in under six days across roughly 750,000 trajectories, costing about $0.85 million for Flash and $2.62 million for Pro, with held-out DeepSWE gains of roughly 17 and 14 points respectively.
- •In research applications, MiMo-V2.6-Pro helped formalize the Li-Yorke theorem in Lean 4 with more than 6,000 lines of kernel-verified code and contributed to screening novel MOF materials for PFAS adsorption.

Xiaomi has released and fully open-sourced MiMo-V2.6, the newest generation of its natively omnimodal AI models, casting the launch as a key step toward scaling reinforcement learning (RL) on verifiable, complex tasks. The lineup consists of two core models — MiMo-V2.6-Pro, which the company describes as its most capable model to date, and the more cost-efficient MiMo-V2.6-Flash — joined by a Pro-UltraSpeed variant that delivers output speeds up to 20 times faster.
On the Artificial Analysis Intelligence Index (v4.3, September 2026), MiMo-V2.6-Pro posted a score of 46.32, overtaking Kimi K3 and Qwen3.8 Max to claim the top rating among open-source models. The series also keeps the API pricing of its V2.5 predecessor, pushing the intelligence-versus-cost Pareto frontier outward at unchanged cost. It is part of a broader pattern in which open and closed frontier models are scored on the same public benchmarks, keeping the open-source tier in direct, like-for-like comparison with closed systems.
The models benchmark competitively across disciplines. On DeepSWE v1.1, a long-horizon software engineering test, MiMo-V2.6-Pro scored 71.9 — behind DeepSeek V4.1 Flash (74.2) and both Claude Opus 5 and GPT 6 Astra (74.0 each) — while MiMo-V2.6-Flash reached 67.9, a dramatic jump from MiMo-V2.5-Pro's 19.0. In general agentic workflows, the series led Automation Bench v1.0.6 among frontier models with the sole exception of DeepSeek V4.1 Flash, and it recorded scores of 94.0 and 95.1 on the cyber-focused CyberGym benchmark.
Introducing Xiaomi MiMo-V2.6 — Pro & Flash. Frontier intelligence, all the modalities, built in public. Two omnimodal models, advancing through scaled reinforcement learning Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks Pro scores… pic.twitter.com/oqfYPC00uK
— Xiaomi MiMo (@XiaomiMiMo) September 21, 2026
Reinforcement Learning at Scale, Streamed in Public
The technical backbone of the release is a scaled RL training run that Xiaomi broadcast live. In under six days, each model completed 30 RL steps spanning roughly 750,000 trajectories, at costs of approximately $0.85 million for Flash and $2.62 million for Pro. Average pass rates on training tasks rose by 25% and 12% in relative terms, respectively, and gains on held-out benchmarks were substantial: DeepSWE v1.1 scores climbed by roughly 17 points for Flash (48.8 to 65.68) and 14 points for Pro (58.4 to 72.57). According to the company, the RL process proved sample-efficient and generalized beyond the training distribution.
Scaling efforts followed three axes: larger batches on a fully asynchronous architecture (1,568 samples per update, context lengths of up to 1 million, and 3.5–3.7 billion tokens per step); a multi-task suite covering coding, general agents, visual and cyber domains; and increased grader compute that uses relative comparisons to produce more precise reward signals. The team also froze the router to curb training drift and deployed a defense against reward hacking combining reward design, adversarial evaluation, anomaly detection and cross-verification.
Beyond benchmarks, Xiaomi showcases applied capabilities under what it calls “Vibe World,” extending natural-language programming from software into interactive 3D worlds, game development, Blender-based 3D modeling, closed-loop robotic arm control, frontend and presentation design, and end-to-end video and music production. In research settings, MiMo-V2.6-Pro contributed to the computational screening of novel MOF materials for PFAS adsorption and helped formalize the Li–Yorke theorem in Lean 4, generating more than 6,000 lines of kernel-verified code without any Lean-specific post-training. The Lean 4 result illustrates the release's stated focus on verifiable tasks, with the proof assistant's kernel — not a human reviewer — confirming the generated output.
The models can be accessed through the MiMo API Platform, AI Studio, MiMo Desktop and OpenRouter, with pricing unchanged from V2.5. Xiaomi is also open-sourcing the full technical report, training environments and RL code to enable reproduction and further research. Independent reproductions of the six-day run and third-party checks of the reported scores are the immediate follow-ons to watch as the release circulates.
Source: Metaverse Post