On September 1, 2026, iFLYTEK subsidiary Ciyuan Xinghuo open sourced Spark X2.5 4B and X2.5 1.7B, compact edge models with a claimed native context window of up to 1 million tokens. The models use hybrid attention and are positioned for agent tasks, coding, math, instruction following, and local document reasoning—w...
Published byEdited with GPT-5.6 TerraImages generated with GPT Image 2
Research answer

Create a landscape editorial hero image for this Studio Global article: What did iFLYTEK’s wholly owned subsidiary launch and open-source on September 1, 2026 with Spark X2.5-4B and Spark X2.5-1.7B—including thei. Article summary: On September 1, 2026, iFLYTEK’s wholly owned subsidiary Ciyuan Xinghuo launched and open-sourced the edge-deployable Spark/Xinghuo X2.5-4B and X2.5-1.7B models. The company claims they are the first edge models with a na. Topic tags: general, general web. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fake numbers, clic
iFLYTEK’s wholly owned subsidiary Ciyuan Xinghuo launched and open-sourced two edge-oriented language models—Spark (Xinghuo) X2.5-4B and Spark X2.5-1.7B—on September 1, 2026. Their standout specification is a native context window of up to one million tokens on local devices. iFLYTEK describes the pair as the first edge models in this category to offer that capability; that leadership claim should be treated as a company claim rather than an independently validated result. 2
The 4B- and 1.7B-parameter models use a hybrid-attention architecture and were tuned around agent or tool use, programming, mathematical reasoning, and instruction following. iFLYTEK says the models perform strongly against similarly sized open models, but the announcement does not provide independent comparative results. 2
Their intended deployment targets include local and edge environments such as smart hardware, vehicles, robots, and Internet-of-Things devices—settings where cloud connectivity, latency, or data handling can be constraints. 3
Long-document workflows often rely on chunking: a manual, policy archive, or codebase is broken into smaller passages and retrieved one at a time. That can be useful, but it can also omit a rule, exception, or dependency found in another section.
iFLYTEK’s example is an after-sales manual. With the full manual available in the prompt, the model can reason across provisions on returns, fault handling, exceptions, logistics, data erasure, and responsibility for costs. The company frames this as an alternative to losing cross-chapter context when a document is split into isolated chunks. 2
A large window is not, by itself, proof of reliable reasoning or factual accuracy. For practical deployments, teams still need to test retrieval quality, response latency, hardware memory requirements, and whether the model correctly follows the applicable rules across a full document.
iFLYTEK positions Spark X2.5 for local productivity assistance, though it has not published a detailed office-automation feature list or a dedicated office benchmark in the release material. 2
For other use cases, the company reports:
According to iFLYTEK, both models were trained entirely on domestic computing platforms using roughly 20 trillion diverse pretraining tokens, then further trained with supervised fine-tuning and reinforcement learning. 2
The announced compatibility list includes NVIDIA, Huawei, Hygon, and Hemo hardware; vLLM, SGLang, and llama.cpp inference frameworks; Ollama and LM Studio deployment paths; and LLaMA-Factory for incremental fine-tuning. Actual speed and memory use will depend on model choice, quantization, context length, and the target hardware. 2
iFLYTEK said it released the model weights, source repositories, and deployment documentation, with availability through Hugging Face and GitHub. It also announced associated APIs on its Starry MaaS platform with limited-time free access. The announcement did not specify the offer’s expiry date or quota, so developers should verify current access terms before planning around the promotion. 2
6
Spark X2.5’s differentiator is not simply that it is open source or small enough for edge use. It is the combination of compact 1.7B and 4B models with a vendor-claimed million-token native context window. If that capability holds up in deployment, it could be particularly useful for private local analysis of full manuals, technical archives, or codebases where cross-document context matters. The key next step for adopters is hands-on evaluation: test the models on representative long-context tasks and on the exact device stack where they will run.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
On September 1, 2026, iFLYTEK subsidiary Ciyuan Xinghuo open sourced Spark X2.5 4B and X2.5 1.7B, compact edge models with a claimed native context window of up to 1 million tokens.
On September 1, 2026, iFLYTEK subsidiary Ciyuan Xinghuo open sourced Spark X2.5 4B and X2.5 1.7B, compact edge models with a claimed native context window of up to 1 million tokens. The models use hybrid attention and are positioned for agent tasks, coding, math, instruction following, and local document reasoning—without necessarily splitting a large manual or codebase into separate retrieval ch...
Weights, code repositories, and deployment documentation were released, with stated support for major inference stacks and a limited time API offer whose quota and end date were not specified.