ZGCM 1 7B is a 7.39B parameter, 256K context model released for mathematical reasoning and tool assisted search. Its value as a research baseline lies in the combination of thinking and direct response modes, a documented training recipe, and released data and intermediate checkpoints.
Published byEdited with GPT-6 SolImages generated with GPT Image 2
Research answer

Create a landscape editorial hero image for this Studio Global article: What is ZGCM-1-7B, who released it and under what license, and what makes it a useful open baseline for mathematical reasoning and tool-assi. Article summary: ZGCM-1-7B is a 7.39-billion-parameter dense language model trained from scratch for mathematical reasoning and tool-assisted search. It was released by researchers at Zhongguancun Academy and the Zhongguancun Institute o. Topic tags: general, academic, general web, user generated. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, char
ZGCM-1-7B is a dense language model trained from scratch to combine mathematical reasoning with active, tool-assisted information gathering. Researchers at Zhongguancun Academy and the Zhongguancun Institute of Artificial Intelligence released it as a model that others can inspect and test, rather than only compare through benchmark scores. 2
4
The model has 7.39 billion parameters. Its model listing identifies the release as MIT-licensed; the CC BY 4.0 label on the research paper applies to the paper and should not be mistaken for the model’s license. Researchers should check the terms attached to each resource they use. 2
7
1
ZGCM-1-7B is a dense, decoder-only Transformer with a stated 256K-token context window. It supports both a deliberate thinking mode and a direct-response mode in the same model, making it possible to study different response styles without changing models. 2
4
Its attention design interleaves gated sliding-window layers with full-attention layers; one model description specifies 27 of the former and five of the latter. The authors report that the hybrid design cuts KV-cache use by 6.4× and increases throughput at 256K context by 3.94× relative to their full-attention comparison. Those are comparison-specific system results, not a promise of the same speedup on every deployment. 4
10
The published recipe runs from pretraining through progressive long-context mid-training and general-agentic supervised fine-tuning. It combines hybrid attention with a stable FP8 Muon optimizer, extends the curriculum across 16K, 64K, and 256K stages, and reformulates interaction traces as Markov decision process transitions. The reported mid-training curriculum covers 600 billion tokens. 1
2
10
Alongside the final model, the project has published a dataset and intermediate checkpoints; reporting on the release also describes training code, recipes, and logs. That gives researchers more of the training path to examine than final weights alone. A checkpoint labeled for the 16K data stage, for example, documents its training context and cumulative mid-training tokens. 2
3
6
9
The team describes researcher-directed AI agents assisting with development tasks, including data curation and cluster operations. That is a description of its workflow—not evidence that agents independently designed or validated the model, or a measurement of how much they contributed to its scores. 2
8
The reported evaluations give ZGCM-1-7B 97.13% on MATH-500, 75.00% on AIME 2026, 63.09% on WebWalkerQA, and 19.43% on BrowseComp. Results are not uniformly ahead of comparison models: the published comparisons also show losses on AIME 2024 and AIME 2025. Separately, the authors report reaching the same pretraining loss in roughly 4.2× less time than their AdamW/BF16 baseline. These figures depend on the stated benchmarks, evaluation setup, and system baselines; they do not establish equivalent performance in every search workflow. 7
10
The model listing provides a Transformers text-generation example for zgcagi/ZGCM-1-7B that uses trust_remote_code=True. Review the repository’s custom code before enabling that option. The provided material does not establish a minimum GPU, memory requirement, or package-version requirement for running the final model, so those should not be inferred from the benchmark or checkpoint descriptions. 2
19
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
ZGCM 1 7B is a 7.39B parameter, 256K context model released for mathematical reasoning and tool assisted search.
ZGCM 1 7B is a 7.39B parameter, 256K context model released for mathematical reasoning and tool assisted search. Its value as a research baseline lies in the combination of thinking and direct response modes, a documented training recipe, and released data and intermediate checkpoints.