Xing4.0 29B A4B is China Telecom AI’s open weight agent focused model: it has about 29B total parameters, activates roughly 4B per token and supports a native 256K token context. It routes each token to four of 64 experts alongside one shared expert; its weights are available on Hugging Face under Apache 2.0.
Published byEdited with GPT-6 SolImages generated with GPT Image 2
Research answer

Create a landscape editorial hero image for this Studio Global article: What is China Telecom AI’s open-sourced Xing4.0-29B-A4B agentic Mixture-of-Experts model, and what are its parameter count, active experts,. Article summary: Xing4.0-29B-A4B is China Telecom AI’s open-weight Mixture-of-Experts language model for coding and other multi-step agent tasks. It has about 29 billion parameters in total, but activates about 4 billion for each token—n. Topic tags: general, general web. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fake numbers, click
Xing4.0-29B-A4B is China Telecom AI’s open-weight Mixture-of-Experts (MoE) language model aimed at coding and other multi-step agent tasks. Its name highlights the distinction: the model has about 29 billion parameters in total, while roughly 4 billion are active for each token. That does not mean a deployment needs to store only 4 billion parameters. 1
10
The model has 64 routed experts plus one shared expert. For each token, it selects four routed experts; the shared expert also participates. Its native context window is 256K tokens, described as extensible to 512K. The larger figure is an extension claim, not the native window. 1
2
Xing4.0-29B-A4B combines multi-head latent attention (MLA), manifold-constrained hyper-connections (mHC) and multi-token prediction (MTP). The reported design uses MLA to reduce attention-cache demands and mHC to help stabilize training, alongside MTP for multi-token generation. 1
5
The model was trained on Huawei Ascend hardware using the MindSpore ecosystem. Its developers report that MoE communication tuning, selective recomputation, graph-operator fusion and custom Ascend C operators increased training throughput by approximately 96% over an out-of-the-box baseline. That is a reported training-system improvement—not a claim that users will see 96% faster inference. 5
13
The weights are available on Hugging Face as XingChen-AGI/Xing4.0-29B-A4B under the Apache 2.0 license. The model materials show Transformers and vLLM usage; reported inference and deployment support also includes SGLang and KTransformers. For fine-tuning, the developers list LLaMA-Factory and MindFormers. Agent-framework adaptations are reported for OpenCode, Claude Code, OpenClaw and Hermes. 12
13
8
Reported results include 75.00 on SWE-bench Verified, 57.50 on Terminal-Bench 2.1 and 93.52 on SuperCLUE’s agent evaluation. These figures come from model materials and coverage of the release; they should not be treated as independently reproduced results. 6
19
8
China Telecom also says low-bit quantization and memory optimization can bring GPU memory use to about 15 GB; other coverage describes the figure for a 4-bit version. It is a deployment claim, not a guarantee for every setup. Quantization choice, runtime configuration and context length matter when assessing whether the model will fit on a particular GPU. 6
8
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
Xing4.0 29B A4B is China Telecom AI’s open weight agent focused model: it has about 29B total parameters, activates roughly 4B per token and supports a native 256K token context.
Xing4.0 29B A4B is China Telecom AI’s open weight agent focused model: it has about 29B total parameters, activates roughly 4B per token and supports a native 256K token context. It routes each token to four of 64 experts alongside one shared expert; its weights are available on Hugging Face under Apache 2.0.