AMD 喺 7 月 24 日公佈 Instella MoE,一個 16B 總參數、每 Token 只激活 2.8B 參數嘅混合專家模型,用自家 Instinct MI300X 同 MI325X GPU 由頭 train 起,仲用 ROCm 軟件棧 新架構包括 Gated Multi head Latent Attention(Gated MLA)同 FarSkip Collective,效率勁過傳統 MoE,預訓練階段訓練速度快咗 12.7% 完全公開:權重、訓練設定、中間檢查點、推論代碼全部喺 Hugging Face 攞得,用 Research RAIL 授權,畀學術同研究用

Create a landscape editorial hero image for this Studio Global article: Search & fact-check with cited sources for What are the key details of AMD's July 24 release of Instella-MoE, a fully open 16-billion-parame. Article summary: Here are the key verified details of AMD's Instella-MoE release, based primarily on the official AMD ROCm blog post (July 24, 2026) and corroborating news sources.. Topic tags: general, academic, documentation, general web, user generated. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fake numbers, clickbait thumbnails
On July 24, 2026, AMD announced Instella-MoE, a fully open 16-billion-parameter Mixture-of-Experts (MoE) language model trained entirely from scratch on AMD Instinct MI300X and MI325X GPUs using the ROCm software stack . The model activates only 2.8 billion parameters per token, making it highly efficient while delivering competitive benchmark performance. This release is a strategic milestone, demonstrating that AMD's hardware and open-source software stack can support large-scale AI development from pre-training through reinforcement learning, rivaling the dominant Nvidia CUDA ecosystem
.
Instella-MoE employs a shared-plus-routed expert configuration with 2 always-active shared experts and 6 routed experts selected from a pool of 64 per token . The model consists of 27 decoder layers, each with a hidden size of 2,048
. Two architectural innovations stand out:
The model also uses a Multi-Token Prediction (MTP) objective during its pre-training and mid-training phases .
After a dedicated long-context extension training stage, Instella-MoE supports a 64K-token context .
AMD published the complete training pipeline, which consists of six sequential stages :
The Instella-MoE-16B-A3B base model achieved an average score of 76.7 on standard benchmarks, positioning it among the top fully open models at its scale . It reportedly outperformed SmolLM3-3B and OLMo-3-7B, and competed with larger models despite activating only 2.8 billion parameters per token
.
In line with its commitment to open AI, AMD released all training artifacts openly :
All artifacts are available on Hugging Face under the amd/Instella-MoE-16B-A3B namespace, licensed under a Research RAIL (Responsible AI License) for academic and research purposes .
Instella-MoE is more than a model release — it is a direct proof point that AMD's ROCm software stack and Instinct GPU hardware can support frontier-scale AI development, from pre-training through RL, using fully open infrastructure. By delivering competitive benchmarks against leading open-source models trained on Nvidia hardware, AMD positions its ecosystem as a viable alternative to Nvidia's CUDA for large-scale AI research and development .
Studio Global AI
Use this topic as a starting point for a fresh source-backed answer, then compare citations before you share it.
AMD 喺 7 月 24 日公佈 Instella MoE,一個 16B 總參數、每 Token 只激活 2.8B 參數嘅混合專家模型,用自家 Instinct MI300X 同 MI325X GPU 由頭 train 起,仲用 ROCm 軟件棧
AMD 喺 7 月 24 日公佈 Instella MoE,一個 16B 總參數、每 Token 只激活 2.8B 參數嘅混合專家模型,用自家 Instinct MI300X 同 MI325X GPU 由頭 train 起,仲用 ROCm 軟件棧 新架構包括 Gated Multi head Latent Attention(Gated MLA)同 FarSkip Collective,效率勁過傳統 MoE,預訓練階段訓練速度快咗 12.7%
完全公開:權重、訓練設定、中間檢查點、推論代碼全部喺 Hugging Face 攞得,用 Research RAIL 授權,畀學術同研究用