MiniMax's model family has grown rapidly in just over one year. The following table shows how the planned M3 Pro compares to its predecessors:
| Model | Release Date | Total Parameters | Active Parameters (MoE) | Key Features |
|---|---|---|---|---|
| MiniMax-M1 | June 2025 | 456B | 45.9B | First open-weight hybrid-attention reasoning model; 1M context; Apache 2.0 BD |
| MiniMax-M3 | June 1, 2026 | ~428B | ~23B | Frontier coding (59.0% SWE-bench Pro), 1M-token context, native image/video/computer-use multimodality WX |
| M3 Pro (planned) | Q3 2026 (target) | 2.7T | TBD (likely MoE) | Potentially the largest open-weight model ever; largest by a Chinese firm RT |
The jump from M3's ~428B total parameters to 2.7 trillion is a ~6.3x increase. If MiniMax continues its Mixture-of-Experts (MoE) approach, the active parameter count per forward pass could remain manageable—perhaps in the tens to hundreds of billions—while the total parameter footprint becomes unprecedented for an open-weight release MR.
DeepSeek V4 Pro, released April 23, 2026, is the current open-weight benchmark leader in pure coding performance. Here is how the two compare, and where the planned 2.7T model would fit:
Crucially, the new model's name may differ at launch. While internal sources refer to "M3 Pro," Reuters' report does not use that codename, and the final name is unconfirmed RFF.
The M3 Pro is the leading edge of a sustained push by Chinese AI labs to release powerful open-weight models. This wave includes:
This reflects a deliberate strategy: Chinese labs release open-weight models under permissive licenses (Apache 2.0, MIT) that allow commercial use and modification. This approach bypasses export controls on hardware by making the models themselves freely available to global developers, and it positions openness as a competitive advantage against Western labs like OpenAI, Anthropic, and Google that keep frontier models closed GRV.
The planned 2.7T model, if realized, would make several statements:
The trend signals that China is no longer just a hardware-constrained follower—it is a leader in open-weight model releases, using openness as a competitive moat against Western closed-model leaders GR.
Deploying a 2.7 trillion-parameter model at scale poses extreme challenges, even before considering training costs:
That said, MiniMax's M3 already uses sparse attention mechanisms to achieve a 1M-token context window efficiently W. The company may employ similar architectural innovations for the 2.7T model, potentially making deployment less daunting than the raw parameter count suggests.
Chinese open-weight models have gained significant traction globally for several practical reasons:
The M3 Pro, if it ships on schedule and as planned, would be the strongest statement yet in this direction: a Chinese open-weight model that is not merely competitive on benchmarks or cost, but the largest model of any kind, open or closed, that the world has ever seen.