Xiaomi Opens Invite-Only Beta for MiMo Desktop, Becoming First China-Based Lab to Offer Computer Use
Key Takeaways
- •Xiaomi launched an invite-only beta of MiMo Desktop, becoming the first China-based lab to offer computer-use capabilities to users via its flagship models.
- •MiMo Desktop accepts multi-format inputs such as spreadsheets, images, video, and PDFs, and delivers editable outputs including documents, presentations, 3D models, and software projects.
- •The overseas version provides full computer control, reading screen content and operating keyboard and mouse across applications, with built-in result verification and a Record & Replay function.
- •The system uses smart scheduling to route tasks between standard and flagship models based on complexity and cost, with cache optimization achieving hit rates of up to 99% within a session.
- •Xiaomi's groundwork includes the open-source MiMo-VL-7B, which set a then-record 56.1 score on the OSWorld-G GUI grounding benchmark, and its MiMo-V2-Pro flagship now ranks in the same tier as Claude 4.5 Sonnet, GPT-5.2, and Gemini 3.0 Pro on agent benchmarks.

Xiaomi has launched invitation-only testing for MiMo Desktop, a desktop AI application designed to execute complete workflows rather than generate responses in a chat interface. With this release, the Chinese technology company becomes the first China-based laboratory to offer computer-use capabilities to users through its flagship models. The move places Xiaomi in a category where Western labs have moved first: Anthropic introduced computer use as a standard API capability for Claude in late 2024, and OpenAI, Google, and others have since shipped browser- and desktop-operating agents, making Xiaomi's entry notable as the first from a China-based lab.
Real-world work rarely starts with a well-defined prompt. It typically involves spreadsheets, images, video, PDF documents, audio recordings, and compressed files that together form the context of a task. MiMo Desktop accepts these multi-format inputs directly, with no prior organization or format conversion required. Users describe their objectives in natural language, and the system then interprets the materials, decomposes the task, invokes the appropriate tools, and delivers editable outputs — including documents, spreadsheets, presentations, web pages, audio, video, 3D models, and software engineering projects.
Two design decisions set the product apart. First, the result preview is not a static rendering but a fully interactive deliverable that incorporates component structure, interaction logic, and data visualization. Users can operate and revise it within the session, whether the output is a data dashboard, a game prototype, or a presentation. Second, modifications are made by selection: users highlight a chart, paragraph, or region, describe the desired change, and only that portion is regenerated. Every revision is preserved in a version history, allowing users to compare drafts and roll back to earlier states.
The system also features "Smart" scheduling. Instead of requiring users to pick a model manually, MiMo Desktop evaluates task type, complexity, and cost requirements, routing routine work to standard models for speed and complex, multi-step tasks to flagship models for higher output quality. Large-scale tasks can be distributed across multiple agent sessions that collaborate while keeping independent memory and workspaces. Xiaomi reports that cache optimization achieves hit rates of up to 99% within a session on long tasks, reducing redundant computation and controlling operational costs.
Browser and Computer Control as Core Capabilities
The second dimension of the release concerns control. MiMo Desktop operates the browser as both an information source and an execution environment: it opens pages, retrieves information, completes forms, extracts assets, and imports findings directly into the task. When producing web-based outputs, the system can also verify deliverables through the browser during the workflow.
The most important feature is full computer control, available in the overseas version. MiMo Desktop reads screen content and operates the keyboard and mouse across applications — opening files, verifying data, and transferring information between programs — with built-in result verification that can trigger corrections or halt execution when necessary. For stable, repetitive processes, a Record & Replay function lets users demonstrate a workflow once and have the system re-execute it through natural language instructions.
This reflects a deliberate architectural approach: the desktop is treated as the operating environment, and the model functions as an agent within it, rather than as a language interface layered on top of existing applications.
A Year in the Making: How Xiaomi Built the Technical Case for Computer Use
Computer-use capability is not primarily an interface problem but a technical one, and Xiaomi has been developing the underlying foundation for more than a year. The groundwork was laid in mid-2025 with MiMo-VL-7B, the company's open-source vision-language model, which achieved a then-record score of 56.1 on OSWorld-G — the standard benchmark for GUI grounding — outperforming models built specifically for the task, such as UI-TARS. This capability was developed through a four-stage training pipeline covering mobile, web, and desktop interfaces, along with a substantial corpus of Chinese GUI data and a training task that infers intermediate actions from before-and-after screenshots — precisely the perceptual foundation required for a computer-use agent.
The model lineage has continued to advance. By early 2026, Xiaomi's flagship MiMo-V2-Pro ranked in the same tier as Claude 4.5 Sonnet, GPT-5.2, and Gemini 3.0 Pro on coding agent, general agent, and tool-use benchmarks. This gives the Desktop client a top-tier model for complex tasks while the routing layer maintains cost efficiency for routine operations. Xiaomi also brings more than two decades of experience in consumer hardware and software ecosystems, which is reflected in the product's focus on the practical conditions of a user's desktop environment rather than on benchmark performance alone.
For observers of the AI sector, the release signals a broader shift in evaluation criteria: the frontier is moving from reasoning benchmarks toward operational competence — agents that complete forms, verify their own output, and stay within defined compute budgets. Xiaomi's combination of GUI-grounding models, a frontier-tier flagship, and a computer-use interface now in beta places the company among the leading contenders in the emerging category of autonomous desktop agents. What to watch next is how the invite-only beta performs on real desktops outside controlled benchmarks — computer-use agents from other labs have shown the format's potential while also demonstrating reliability challenges on multi-step workflows — and whether Xiaomi opens the beta to broader access beyond the overseas version's full computer control.
Source: Metaverse Post