Home / News EN / Huawei Accelerates AI Chips: Ascend 960DT Arrives Q1 2027

Huawei Accelerates AI Chips: Ascend 960DT Arrives Q1 2027

Ascend 960DT launch. Huawei has moved the availability of the Ascend 960DT accelerator to Q1 2027 and presented a roadmap that includes 960PR, 970, and 980 solutions. Meanwhile, details were provided on Atlas 960E SuperPoD and OceanStor M900.

Ascend 960DT launch

The Chinese company has accelerated the development of its artificial intelligence infrastructure, moving the launch of the Ascend 960DT to Q1 2027. Initially scheduled for Q3 next year, Huawei announced this acceleration during Huawei Connect 2026.

Ascend 960DT launch: why it matters

A Huawei spokesperson stated: “The Ascend 960DT should be ready in Q1 2027. The Ascend 960 chips are arriving ahead of schedule, doubling performance and advancing year over year.” The Ascend 960DT will be the successor to the Ascend 950 and will introduce a new SIMD/SIMT architecture with support for various numerical formats, including FP32, HF32, FP16, BF16, FP8, MXFP8, HiF8, MXFP4, and HiF4. The “H” formats consider both precision and dynamic range simultaneously, aiming to represent information that would normally require two floating-point formats with a single format.

In terms of performance, Huawei indicates up to 2 PFLOPS in FP8 and 4 PFLOPS in FP4. The chip is expected to integrate 288 GB of HBM memory with a bandwidth of 9.6 TB/s, an increase of 2.4 times compared to the Ascend 950 generation. Internal interconnection will reach 2.2 TB/s.

The Ascend 960PR is also scheduled for Q3 2027, a variant primarily intended for inference, with performance up to 8 PFLOPS FP4, 192 GB of HBM with 2.4 TB/s bandwidth, and 2.2 TB/s interconnection.

The 2028 roadmap includes the Ascend 970, credited with 3.6 PFLOPS FP8 and 14 PFLOPS FP4, 288 GB of HBM at 14.4 TB/s, and 4.4 TB/s interconnection. For 2029, the Ascend 980 is planned, for which Huawei preliminarily indicates 7.2 PFLOPS FP8, 28 PFLOPS FP4, 384 GB of HBM at 38.4 TB/s, and 8 TB/s interconnection. The specifications for the 980 are still preliminary and may undergo changes before launch.

The most ambitious project presented is the Atlas 960E SuperPoD system, scheduled for Q3 2027. It is a fully liquid-cooled platform that uses Near-Package Optics (NPO) solutions through Hi-ONE optical modules, with a declared transmission speed of 7.2 Tbps per module. The SuperPoD integrates 4,096 NPUs, unified addressed memory, and UnifiedBus interconnection.

What changes and what are the effects

Huawei indicates overall performance up to 8 Exaflops FP8 and 16 Exaflops FP4, with a total unified memory of 256 TB. Other corporate communications refer to approximately 1 PB of HBM, a difference linked to how system memory and fabric memory are accounted for. The platform is expected to use around 5,500 Hi-ONE modules, reducing power consumption by over 550 kW compared to configurations based on traditional optical links.

Compared to the Atlas 950 SuperPoD, Huawei indicates up to 2.3 times the performance in training and 2.5 times in inference for models with 10 trillion parameters, with a reduction in inference latency of up to 70%. Multiple SuperPods can be connected together reaching 512,000 accelerators and, in multi-track configurations, up to one million NPUs.

The Peerium Computing Architecture uses UnifiedBus to connect processors, memory, storage, and network. Huawei is also addressing the issue of KV cache growth, particularly relevant with models featuring increasingly long contexts. The OceanStor M900 solution leverages the Lingqu network to create a petabyte-scale global cache hierarchy, extending available capacity from memory and VRAM down to SSDs.

A single cluster can reach 64 PB of capacity, while the amount of KV cache available for each NPU moves, according to Huawei, from the order of gigabytes to that of terabytes. The system combines CPU, network, and storage controller to allow the NPU to access SSDs directly without protocol conversions and without passing through the CPU. Huawei declares a reduction in access latency from milliseconds to approximately 60 microseconds (about 90%) and an aggregate bandwidth up to 40 TB/s.

Source and further reading on Ascend 960DT launch: original article.

* Content created with the assistance of artificial intelligence systems.