Xiaomi is publicly streaming selected telemetry from two unfinished MiMo V2.6 RL runs, Pro and Flash: roughly 2 billion tokens per step across 1,568 prompts × 16 rollouts. The notable technical signal is that Xiaomi says it is scaling not only model compute, but also multi task agent environments and grader side com...
Published byEdited with GPT-5.6 TerraImages generated with GPT Image 2
Research answer

Create a landscape editorial hero image for this Studio Global article: What is Xiaomi’s public MiMo-V2.6 reinforcement-learning post-training livestream at mimo.xiaomi.com/rl/, what does it reveal about the para. Article summary: Xiaomi’s MiMo-V2.6 page is an unusual public dashboard for *ongoing* RL post-training, not a release of a finished model. It exposes selected trainer-log telemetry for parallel unreleased Pro and Flash runs—reward and be. Topic tags: general, education, general web, user generated, academic. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, water
Xiaomi’s MiMo team has made an unusual part of model development public: a dashboard for the reinforcement-learning post-training of two unreleased models, MiMo-V2.6-Pro and MiMo-V2.6-Flash. Rather than publishing only a finished benchmark table, the page exposes selected training telemetry such as reward curves, rollout counts, step timing, running compute costs, evaluation results, and infrastructure events. 1
8
The key takeaway is not that the dashboard proves a final model result. It is that Xiaomi is treating large-scale agentic RL operations—rollouts, environments, verification, grading, updates, failures, and spending—as something worth observing in public.
Reports on the public page describe parallel Pro and Flash runs that began on September 15, 2026. The visible dashboard tracks progress while the models remain in training; Xiaomi has not announced final specifications, pricing, or public availability for either model. 1
8
The disclosed batch configuration is substantial: 1,568 prompts with 16 rollouts per prompt, or about 25,000 rollouts, and approximately 2 billion tokens per RL step. Xiaomi says the system operates fully asynchronously. 24
That matters because this is not presented as a simple preference-optimization loop over static text pairs. Luo Fuli, who leads the MiMo team, described three scaling dimensions:
In practical terms, the public description implies a pipeline in which models generate actions or solutions, execute them in task environments, receive verifier or rubric-derived feedback, and then update from that feedback. At this scale, generating rollouts and running evaluators can be as operationally important as the optimizer itself.
A fully asynchronous RL system allows rollout generation, environment execution, scoring, and model updates to proceed without every component waiting for the slowest task. That design is especially relevant for mixed agent workloads, where some tasks may finish quickly while others require longer tool use or more expensive grading.
Xiaomi’s stated focus on grader compute is also meaningful. In agentic tasks, a final outcome may depend on a sequence of actions rather than one response. Assigning reward across that sequence—what Xiaomi calls agentic in-group credit assignment—becomes part of the training system rather than an afterthought. 24
The dashboard therefore offers an operational lesson: scaling RL for agents is not simply a matter of adding training GPUs. It also requires enough environment, rollout, and verifier capacity to keep the training loop supplied with useful feedback.
Early reporting based on the public dashboard put combined spending at more than $1.08 million over roughly 36 hours after Pro began, or about $30,000 per hour on average across the two runs. 4 Other snapshots reported approximately $830,000 for Pro and $360,000 for Flash, alongside a Pro restart associated with node VRAM and a Flash restart after a dataset error at step 15.
9
Those figures are striking, but they need careful interpretation:
Still, the counter makes one point difficult to miss: Xiaomi is representing this as a costly, infrastructure-heavy online RL effort, not a lightweight post-training experiment.
The dashboard reportedly includes reward curves and mid-training coding evaluations. One reported snapshot put DeepSWE v1.1 at 63.72 for Pro and 60.77 for Flash. 9
These are useful diagnostics, not final capability claims. A rising reward curve can be consistent with better task completion, but it may also reflect changes in task mix, variance, grader behavior, reward optimization, or benchmark exposure. Likewise, a mid-run benchmark score is not a final model result unless the evaluation protocol, sampling settings, contamination controls, and final checkpoint-selection process are documented and independently checked.
The right reading is: the stream lets observers watch training signals evolve. It does not by itself show how well a finished MiMo-V2.6 model will generalize outside the displayed evaluation setup.
Luo Fuli said the team spent nearly half a year investigating a single question: how far reinforcement learning can scale. She characterized MiMo-V2.6 as still in the middle of its RL run and said Xiaomi would open-source details "piece by piece" over coming weeks. 24
That is a meaningful commitment to further technical disclosure, but it should not be read as a complete reproducibility release. The statement does not, by itself, promise all training data, prompts, grader weights, environment implementations, optimizer state, every checkpoint, or full cluster logs.
MiMo-V2.6’s transparency is about process visibility. Xiaomi is exposing selected operational telemetry from an ongoing language-and-agent post-training run.
Xiaomi-Robotics-U0 was a different kind of project and a different kind of openness. The robotics model is described as a 38-billion-parameter autoregressive world foundation model for embodied intelligence, unifying tasks including text-to-image generation, image editing, embodied scene generation, embodied transfer, and embodied video generation. 25
34 Xiaomi’s public repository describes released model and framework materials for that work.
32
The distinction is straightforward:
Both can be valuable. Neither automatically supplies everything needed to reproduce the full result from scratch.
The MiMo stream is more revealing than a conventional final benchmark announcement, but several important uncertainties remain.
A public dashboard can be near-real-time while still containing delayed updates, backfilled values, corrected data, pauses, restarts, or manual interventions. Its existence is evidence of unusual disclosure, not cryptographic proof of uninterrupted continuity.
Public counts of prompts, rollouts, tokens, and costs do not prove the absence of undisclosed teacher-model outputs, proprietary data, filtered human data, hidden benchmark exposure, or other off-dashboard inputs.
Without a fully published evaluation harness, frozen test sets, sampling settings, repeated seeds, and independent replication, neither reward curves nor interim DeepSWE results can settle the final general capability of the models.
The dashboard can show progress and incidents without revealing the full task mixture, reward weighting, grader reliability, hardware inventory, exact cost methodology, data provenance, or final checkpoint-selection criteria.
Xiaomi’s MiMo-V2.6 dashboard is a rare public observability experiment for large-scale agentic RL. Its disclosed configuration—roughly 2 billion tokens per step, 1,568 prompts, 16 asynchronous rollouts per prompt, mixed environments, and substantial grader-side work—shows that the company is framing RL scaling as a systems problem as much as a model-training problem. 24
The reported spend, reward curves, benchmark snapshots, and visible failures make the process more inspectable than a post hoc launch post. But they remain selected, self-reported telemetry from unfinished runs. The dashboard is strong evidence of operational transparency—not independent proof of continuous live training, complete training provenance, or final model quality.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
Xiaomi is publicly streaming selected telemetry from two unfinished MiMo V2.6 RL runs, Pro and Flash: roughly 2 billion tokens per step across 1,568 prompts × 16 rollouts.
Xiaomi is publicly streaming selected telemetry from two unfinished MiMo V2.6 RL runs, Pro and Flash: roughly 2 billion tokens per step across 1,568 prompts × 16 rollouts. The notable technical signal is that Xiaomi says it is scaling not only model compute, but also multi task agent environments and grader side computation, using test case and rubric based rewards.
The dashboard is unusually transparent about operational progress and some failures, but public curves cannot establish uninterrupted live continuity, clean data provenance, or independently replicated capability.