The configuration shown on the floor connected 1,024 Ascend 950DT NPUs across 16 cabinets — a fraction of the full system, which Huawei says links up to 8,192 processors across 160 cabinets covering roughly 1,000 square meters, with multiple SuperPoDs able to interconnect into a SuperCluster scaling beyond 500,000 Ascend NPUs. Huawei's self-reported figures put the 8,192-chip configuration at 8 exaflops of FP8 performance, 16 exaflops at FP4, and 16.3 petabytes per second of interconnect bandwidth, using a proprietary all-optical protocol called UnifiedBus that functions more like a shared memory fabric than a conventional network.
None of these figures have been independently verified by any third-party benchmarking organization, and technical analysts have explicitly flagged that cluster-level performance claims of this kind are difficult to validate externally. But the architecture answers a specific constraint: Huawei cannot manufacture chips that match Nvidia's per-unit performance, so its approach scales through sheer interconnect density instead — UnifiedBus lets thousands of NPUs share a unified memory pool, which becomes essential when training models with hundreds of trillions of parameters.
The debut lands as Huawei also plans to enter South Korea's AI accelerator market in Q4 2026, pitching aggressive pricing against Nvidia's H20, and as DeepSeek has said that once Ascend 950 supernodes reach mass production later this year, pricing on its DeepSeek-V4-Pro model will drop substantially — a preview of how domestic Chinese compute capacity may reshape model economics regardless of whether Huawei's raw performance claims hold up.