China Telecom Debuts Xing4.0 Coding Agent Model
China Telecom has released Xing4.0-29B-A4B, an open-source mixture-of-experts model that allows developers to run a powerful coding agent locally with minimal hardware requirements.

China Telecom's AI team has launched Xing4.0-29B-A4B, a sparse mixture-of-experts model designed for local coding agents, repository analysis, and long-context research. Released under an Apache 2.0 license, the model features 29 billion total parameters but activates only 4 billion parameters per token. This architectural choice significantly reduces the computational power required for inference while keeping the full parameter set available in system or GPU memory. The model was trained end-to-end on Huawei Ascend 910C NPUs using the MindSpore framework, marking a milestone for this hardware scale.
Structurally, Xing4.0 consists of 40 layers with a hidden size of 3,584. It routes each token through four of its 64 specialized experts, alongside one shared expert. To handle massive codebases, the model features a native context window of 256,000 tokens, which can be extended up to 512,000 tokens. This long-context capability is made practical through Multi-head Latent Attention, which compresses the key-value cache to prevent memory usage from spiraling out of control during long-context operations.
In performance evaluations, Xing4.0 achieved a score of 75.0 on SWE-bench Verified and 57.5 on Terminal-Bench 2.1, demonstrating strong capabilities in resolving real-world software engineering issues. To facilitate local deployment, the model has been converted to the GGUF format, with available quantizations ranging from a highly compressed 9.9 GB (IQ2_M) version to a full 62.5 GB (F16) version.
For software developers and AI practitioners, this release lowers the barrier to running advanced coding assistants on consumer-grade hardware. By utilizing GGUF quantizations, developers can run a highly capable 29-billion-parameter class model on local machines with as little as 19 GB of memory. This enables private, offline repository analysis and agentic workflows without relying on expensive cloud APIs or high-end enterprise GPUs.
This is our own summary of reporting by AlphaSignal



