China Telecom’s Hangzhou branch, ZTE and Zhonghao Xinying deployed a 1,024 chip Chana TPU cluster rated at 400 PFLOPS. The project is described as China Telecom’s first large scale domestic TPU cluster deployment and East China’s first domestic TPU thousand card cluster; a second phase is planned to exceed 10,000 PF...
Research answer

Create a landscape editorial hero image for this Studio Global article: How did China Telecom’s Hangzhou branch, ZTE, and Zhonghao Xinying deploy China’s first domestic 1,024-chip Chana TPU cluster delivering 400. Article summary: China Telecom’s Hangzhou branch, ZTE, and Zhonghao Xinying built the first phase as a production-oriented, fully domestic TPU stack rather than merely installing accelerators: 1,024 first-generation Chana TPUs were integ. Topic tags: general, academic, general web, user generated. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, char
The Hangzhou deployment is significant less for its 1,024-chip headline than for what was integrated around those chips. China Telecom’s Hangzhou branch, ZTE and Zhonghao Xinying combined first-generation Chana TPUs with the Taize computing platform, servers, cluster networking and scheduling into a 400-PFLOPS system for large-model training, AI inference and scientific computing. Reporting describes it as China Telecom’s first large-scale domestic-TPU cluster deployment and East China’s first domestic TPU thousand-card cluster. 17
20
Phase one uses 1,024 Chana TPUs on Zhonghao Xinying’s Taize platform. The partners describe the system as domestically built across the accelerator, server hardware, cluster networking and compute scheduling layers. 17
18
That distinction matters. A large AI cluster is not simply a rack of accelerators: usable performance depends on how servers, network fabric, model software and the scheduler work together. The project’s stated purpose is to turn the TPU into an operable computing service rather than a standalone chip demonstration. 18
Large-language-model inference has two different stages:
These stages have different resource profiles. Separating them lets an inference platform allocate and schedule capacity for prompt processing and token generation independently, instead of making one device pool serve both workloads under a single schedule. Production implementations also require cluster orchestration and network configuration, as shown by deployment guidance for Prefill–Decode-separated SGLang services. 2
Hangzhou Telecom said that months of software-and-hardware tuning—covering model versions, operators, server configurations, network bandwidth and scheduling—improved token-output efficiency by more than 10× compared with the start of the year. It also reported a fall in cost per million tokens from roughly ¥100 to below ¥10. 9
11
Those figures should be read as operator-reported deployment results, not universal TPU benchmarks. The available reports do not disclose the model mix, precision, batch sizes, traffic pattern, utilization rate or cost-accounting method. That makes it impossible to independently reproduce or compare the claimed savings with another inference platform.
The cluster links three capabilities that are usually evaluated separately:
For a domestic AI chip, this kind of operator environment is useful because it shifts the test from peak-chip specifications to service delivery: scheduling, model adaptation, network behavior and ongoing operations all affect whether customers receive predictable inference capacity. The project’s architecture explicitly includes compute resources, scheduling, model adaptation, operations management and industry applications. 18
A TPU is a tensor-oriented AI accelerator. Its promise is not peak FLOPS alone, but efficient execution of the matrix-heavy workloads common in large-model training and inference. At cluster scale, however, accelerator performance is inseparable from interconnect, memory behavior, collective communication, scheduling, reliability and model-software compatibility.
That is why a domestic alternative must prove more than that its silicon works. Models, inference engines, operators, quantization approaches and customer workflows need to run reliably on the new stack. The supplied reporting supports the Hangzhou deployment and its planned expansion, but it does not establish broad parity with mainstream GPU software ecosystems across models and frameworks.
The partners plan a second phase using Zhonghao Xinying’s second-generation Xuyu TPU, with total planned compute exceeding 10,000 PFLOPS. 17
19
The company reports that Xuyu delivers 896 TFLOPS of mixed-precision floating-point performance and 1,792 TOPS for 8-bit inference, with a 600 W rated power draw. It also claims Xuyu offers roughly three times the performance of Chana and 50% lower power consumption than conventional chips at comparable performance. These are company-reported specifications. 19
20
At the system level, Xuyu is designed for supernodes with up to 2,048 chips connected through all-optical interconnect, according to the company’s public presentation. 20
The project demonstrates an effort to assemble a domestic AI-computing stack from chip through platform: accelerator, instruction set, servers, networking, scheduling and service operations. 18
19 Its most meaningful test will be whether that stack can sustain diverse production workloads with compatible software, predictable performance and transparent economics.
The 400-PFLOPS deployment and the planned Xuyu expansion are widely reported. The more eye-catching claims—more than 10× higher token-output efficiency and sub-¥10 cost per million tokens—are plausible goals for a better-utilized inference system, but remain claims from the operator and project participants until benchmark methodology and workload details are disclosed. 9
11
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
China Telecom’s Hangzhou branch, ZTE and Zhonghao Xinying deployed a 1,024 chip Chana TPU cluster rated at 400 PFLOPS.
China Telecom’s Hangzhou branch, ZTE and Zhonghao Xinying deployed a 1,024 chip Chana TPU cluster rated at 400 PFLOPS. The project is described as China Telecom’s first large scale domestic TPU cluster deployment and East China’s first domestic TPU thousand card cluster; a second phase is planned to exceed 10,000 PFLOPS with Zhonghao...