Sixteen cabinets, 1,024 NPUs — that's what was actually standing on the floor in Shanghai

The 2026 World AI Conference opened on July 17 at the Shanghai World Expo Exhibition Center and ran through July 19. Huawei used it to do something it had not done before: put the Atlas 950 SuperPoD (昇腾950超节点) on a show floor as physical, assembled hardware rather than a slide. Huawei's Chinese newsroom and IT之家 both date the debut to July 17. Seoul Economic Daily and Huawei Central date it to July 16, because that's when the media preview ran, the day before the doors opened. Both are right — preview on the 16th, official exhibition on the 17th — and if you see the two dates fighting in your feed, that's why.

Here's the deal, and it's the part almost every aggregated summary got wrong: the machine on the floor was not the machine in the headlines. Per Huawei's own release, the exhibited system was 1,024 Ascend NPUs across 16 cabinets — which lines up exactly with the "64 NPUs per cabinet" figure Huawei published at MWC Barcelona in March, and with what on-site observers counted. The official spec for that 1,024-card configuration is 1 EFLOPS FP8, 2 EFLOPS FP4, a 256TB globally unified memory address space, and 3μs round-trip latency between NPUs.

A lot of coverage attached those exact numbers — 1 EFLOPS / 2 EFLOPS / 256TB — to the phrase "when scaled to 8,192 chips." That attribution is simply wrong, and it badly undersells the claim while also misplacing it. The official full-scale numbers for an 8,192-NPU Atlas 950 SuperPoD are an order of magnitude up: 8 EFLOPS FP8, 16 EFLOPS FP4, 1,152TB of total memory, and 16 (16.3) PB/s of interconnect bandwidth. And those figures did not debut in Shanghai this month. They were unveiled on September 18, 2025, at Huawei Connect 2025, in a keynote by rotating chairman Eric Xu (Xu Zhijun). What WAIC added was metal, not math.

The same correction applies to the comparison everyone quoted. "6.7x the compute and 15x the memory of Nvidia's NVL144" is a claim about the 8,192-NPU full build, not about the 1,024-card unit that was exhibited — and Huawei's original framing included two more multipliers alongside those: 56.8x the NPU count and 62x the interconnect bandwidth. One more fix while we're here: the widely repeated line that Atlas 950 SuperPoD specs "were first unveiled at MWC Barcelona in March" is false. First unveiling was Huawei Connect in September 2025. MWC Barcelona on March 2, 2026 — presented by Seaway Zhang, president of Huawei's computing product line — was the international market launch, aimed at buyers outside China. Get that order backwards and you lose the single most important fact about this product: it has been in public development for ten months, and it still hasn't shipped.

The cast — a company with no shareholders, a chairman who admits he's behind, and a bus named Lingqu

Start with Huawei itself, because its ownership structure explains its behavior. Founded by Ren Zhengfei in 1987, headquartered in Shenzhen, employee-owned with no outside shareholders and no listing. That matters more than it sounds. A company with no quarterly stock price to defend can publish a four-year silicon roadmap and grind against it without having to explain a bad quarter to anybody. Per the 2025 annual report (released March 31, 2026, audited by KPMG), revenue was CNY 880.9 billion — roughly $124 billion — with net profit of CNY 68 billion and R&D spending of CNY 192.3 billion, or 21.8% of revenue. Cumulative R&D over the past decade: CNY 1.382 trillion. Rotating chair Sabrina Meng framed it as "we are moving toward a future that is full of uncertainty, so we have to remain true to our strategy and maintain strategic focus." Worth noting the same reporting cycle wasn't uniformly good news — CNBC reported that Huawei's cloud revenue declined in 2025.

Eric Xu is the voice of this product line, and the interesting thing about him is that he says two apparently contradictory things and both are load-bearing. At Huawei Connect 2025 he went maximally confident: "We are confident that, for the next few years, the Atlas 950 SuperPoD will remain the world's most powerful SuperPoD," and on the cluster tier, "The Atlas 950 SuperCluster will unequivocally be the world's most powerful computing cluster." But in the WAIC context, per Seoul Economic Daily, he conceded the opposite on the axis that matters most to engineers: Huawei is behind Nvidia at the single-chip level and won't catch up quickly — but it is confident at the supernode and cluster level.

Put those together and you have the entire Chinese AI-silicon strategy in one sentence. If you cannot win per chip, change the unit of competition to the system. Export controls took away EUV, which caps per-die performance. They did not take away the ability to bolt thousands of weaker dies together — that's a packaging, optics, and protocol problem, and those are problems Huawei has spent thirty years solving in telecom.

Which brings us to the third character, which isn't a person. Lingqu (灵衢), also called UnifiedBus, version 2.0. This is the actual product. Huawei published the technical specification in September 2025, and the stated design goal is blunt: make 10,000+ NPUs behave like a single machine. The claimed properties are TB/s-class bandwidth, 2.1μs latency, optical-path fault detection at the 100ns level, and all-optical links beyond 200 meters. That last one is the strategically important number. If you can hold coherent, low-latency connectivity across 200+ meters of fiber, you can scatter cabinets across a data center floor and still address them as one memory space. The exhibited 1,024-card unit hitting 3μs RTT is the evidence that this layer works in silicon, not just in a whitepaper.

And the chip inside is the Ascend 950DT, which sits in the middle of a published four-generation ladder: 950PR (prefill and recommendation workloads, 1 PFLOPS FP8 / 2 PFLOPS MXFP4, Huawei's own "HiBL 1.0" HBM, Q1 2026) → 950DT (decode and training, in-house "HiZQ 2.0" HBM at 144GB and 4TB/s, 2TB/s of per-chip interconnect, Q4 2026) → Ascend 960 (Q4 2027, roughly 2x the 950) → Ascend 970 (Q4 2028, 4 PFLOPS FP8). The detail to circle is that Huawei is branding its own HBM. High-bandwidth memory is the sharpest edge of the export-control regime, and Huawei is putting product names on it. Whether that represents genuine independence or a label on a domestic partner's output is the question the entire roadmap rests on — more on that below.

What actually happened — separating the exhibit from the announcement

Stated precisely, the new fact created by WAIC 2026 is exactly one fact: a pre-commercial supernode exists as assembled hardware, was built out to 1,024 cards, and was shown running. The specifications had been public for ten months. But do not treat "the hardware exists" as trivial. The graveyard of AI accelerator projects is full of impressive specs that never became assembled racks, and wiring 1,024 NPUs across 16 cabinets into a single addressable memory space is a real integration milestone even at one-eighth scale.

The full-scale build is worth visualizing physically, because the numbers stop being abstract. 160 cabinets — 128 compute plus 32 networking — occupying roughly 1,000 square meters. That's about four tennis courts of floor space for one computer. Huawei's claimed throughput at that scale is 4.91 million tokens/second for training (17x its own prior-generation Atlas 900 A3) and 19.6 million tokens/second for inference (26.5x). The stated purpose is unambiguous: training trillion-parameter models, and high-concurrency inference at data-center scale. Commercial availability target: Q4 2026, which has not happened yet.

Item Exhibited at WAIC 2026 8,192-NPU full build (announced spec)
When / where Jul 16 preview, Jul 17 exhibition — WAIC, Shanghai Sep 18, 2025 — Huawei Connect 2025, Eric Xu keynote
NPU count 1,024 8,192 (Ascend 950DT)
Cabinets 16 (64 NPUs each) 160 (128 compute + 32 networking), ~1,000 m²
FP8 1 EFLOPS 8 EFLOPS
FP4 2 EFLOPS 16 EFLOPS
Memory 256TB globally unified address space 1,152TB total
Interconnect TB-class NPU-to-NPU, 3μs RTT 16 (16.3) PB/s
Protocol Lingqu / UnifiedBus 2.0 Lingqu / UnifiedBus 2.0
Claim vs NVL144 6.7x compute · 15x memory · 56.8x NPUs · 62x bandwidth
Availability Demo / exhibition stage Q4 2026 target
Third-party validation None None

Look at the memory row and you'll notice the arithmetic doesn't close. If 8,192 NPUs give you 1,152TB, that's ~144GB per NPU, so 1,024 cards should yield about 144TB. The 950DT's HBM spec is indeed 144GB, and 1,024 × 144GB ≈ 147TB. Yet the exhibited configuration is listed at 256TB — nearly double. The most defensible reading is to take Huawei's wording literally: the 256TB is a "globally unified memory address space," not a sum of HBM. It almost certainly folds in host memory or pooled DRAM mapped into the same addressing scheme. Which means 1,152TB and 256TB are not the same kind of number and must not be compared side by side. A great deal of the coverage did exactly that.

The SuperPoD also isn't the top of Huawei's stated stack. Above it sits the Atlas 950 SuperCluster — 64 SuperPoDs, 520,000+ Ascend 950DT chips, 524 EFLOPS FP8, 1 ZFLOPS FP4, also targeted at Q4 2026 — and above that the Atlas 960 SuperCluster at 1 million+ NPUs and 2 ZFLOPS FP8 for Q4 2027. Huawei claims the Atlas 950 SuperCluster carries 2.5x the compute units and 1.3x the compute of xAI's Colossus. Same caveat throughout: these are vendor figures for systems that do not exist yet.

Two more things Huawei put on the WAIC floor deserve attention because they're about deployed product rather than roadmap. First: the previous generation, the Ascend 384 supernode (CloudMatrix 384), has shipped in more than 750 commercial deployments across internet, telecom, finance, education, healthcare, transportation and manufacturing, and Huawei described it as the only domestic supernode to have trained a SOTA model. Second, on software: the CANN open-source community now reports 67 projects, over 12.44 million lines of code, and 3,500+ monthly active developers, alongside 3,000 partners, 7,000 solutions and 2,000 enterprise customers. And on July 18, the Atlas 950 SuperPoD won the SAIL Award, WAIC's top honor.

Who gets what out of this

Huawei buys credibility and time. It cannot win the per-chip fight and its own chairman said so. But if the unit of competition becomes "how many accelerators can you make behave as one machine," then a per-die deficit gets reframed as a problem you solve with more dies and better optics. Showing metal converts that argument from a claim into an artifact, and the SAIL Award stamps state-level endorsement on it. For an enterprise procurement committee, "a spec that will exist" and "a system I watched run in a hall in Shanghai" are not remotely the same document.

China's industrial base gets a procurement alternative, which is the actual point. The backdrop explains the timing. When Nvidia's H20 was effectively barred from China in April 2025, Nvidia took a $5.5 billion write-down; policy has since swung between loosening and re-tightening, and the follow-on B30A was also blocked from Chinese sale. With per-chip supply constrained by Washington, the workaround China chose was system-level integration — supernodes. Chinese media are now calling 2026 the "year one of Chinese supernodes," and WAIC 2026, a state-level event whose opening involved President Xi Jinping, was built to showcase exactly that.

There's an export play too, and it's more deliberate than it looks. At MWC Barcelona in March, Huawei didn't just show the Atlas 950 — it showed the TaiShan 950 SuperPoD for general-purpose computing and the Atlas 850E, an air-cooled system scaling from 8 to 1,024 NPUs. That 850E is the tell. An air-cooled box starting at eight accelerators isn't aimed at hyperscalers; it's aimed at mid-size enterprises and telco edge sites in markets that either can't buy Nvidia or would rather not. Huawei is building the bottom rungs of a ladder, not just the top.

Nvidia arguably benefits too, in a perverse way. Jensen Huang has argued for years that export controls accelerate Chinese self-sufficiency rather than preventing it, and a photograph of assembled Atlas 950 cabinets is the single best exhibit for that argument in front of US policymakers. That doesn't make it a gift from Huawei — both sides simply find the same image useful.

The party that got nothing yet is the customer. Commercial availability is a Q4 2026 target, and the Ascend 950DT silicon inside it is scheduled for the same quarter. A chip and the system built from it debuting simultaneously is a schedule with no slack, and schedules with no slack in semiconductors slip. What's proven today is existence. How many you can buy is a completely different question, and nobody has answered it.

Precedent — NVL72 proved the category, Ponte Vecchio and Graphcore proved the failure modes

The category itself is validated, and Nvidia validated it. GB200 NVL72, announced at GTC in March 2024 and shipping in volume through 2025, established the "sell an entire rack as one GPU" paradigm and drove a substantial share of Nvidia's data-center revenue surge. So the supernode isn't a Huawei invention or a Chinese workaround — it's the shape the frontier converged on, and Huawei is pushing the same shape to an extreme that its constraints require. Going from 72 GPUs per rack to 8,192 per pod is a difference of magnitude, not of direction.

Huawei has its own success precedent, and it's the strongest card in the deck: CloudMatrix / Ascend 384, announced in 2025, now at 750+ commercial deployments as of July 2026. That number is what separates Huawei from a vendor announcing vapor. This is a company that shipped the previous generation into production at real customers. The honest caveat is that 384 to 8,192 is a 21x jump, and supernode difficulty scales worse than linearly — cable and optical module counts, failure probability across a coherent domain, and synchronization overhead all grow together.

Now the failure precedents, and the first one is uncomfortably close. Intel's Ponte Vecchio and the Aurora supercomputer were announced in 2019 with enormous fanfare, slipped repeatedly, and only delivered in 2023–24. The successor, Falcon Shores, was removed from the commercial lineup in January 2025. Intel had money, fabs, and talent, and still spent four years in the gap between the announced spec and the delivered machine. Atlas 950 currently occupies precisely that gap: spec published, hardware demonstrated, product not shipped.

The second failure precedent kills a different way. Graphcore positioned IPU-POD as the alternative to Nvidia, and the hardware was genuinely interesting. It failed to win demand and sold to SoftBank in July 2024 for roughly $400 million, against a prior valuation around $2.8 billion. What killed it was not FLOPS — it was software and ecosystem gravity. Nobody could justify porting off CUDA. This is exactly why Huawei bothered to put CANN community statistics on a hardware show floor: it knows precisely where Graphcore died. But it's also fair to note that 3,500 monthly active developers is a different order of magnitude from CUDA's ecosystem, and that gap is not closed by a press release.

How the competition answers

Nvidia's response is already scheduled. The Vera Rubin NVL144 is due in the second half of 2026: 3.6 EFLOPS FP4 for inference, 1.2 EFLOPS FP8 for training, roughly 3.3x GB300 NVL72. The Rubin GPU is a two-reticle design at 50 PFLOPS FP4 with 288GB of HBM4, paired with an 88-core Vera CPU, and 75TB of fast memory per rack. So when Huawei says "6.7x," the target is a product that hasn't launched, and the comparison sets one rack (NVL144) against 160 cabinets (Atlas 950). That is not a like-for-like benchmark by any reading. Normalize for floor space, power, cooling and total cost of ownership and most of the multiple evaporates. Stated fairly, Huawei's claim mostly reduces to "our system is bigger," which is true and is not the same as "our system is more efficient."

The more interesting counterplay came from inside China. WAIC 2026's computing hall was the largest gathering of non-Nvidia AI infrastructure ever assembled, and Alibaba Cloud was the standout. Its Zhenwu M890 chip inside the Panjiu AL128 SuperPoD — 128 cards per cabinet, a cable-free orthogonal architecture, Alibaba's own ALink System protocol, Pb/s-class bandwidth and sub-100ns latency — was named one of the exhibition's treasures. More consequentially, Alibaba began selling SuperPoD-class compute on its public cloud for the first time, and disclosed cumulative deliveries of over 560,000 Zhenwu chips.

Sit with that last number. While Huawei describes an 8,192-NPU system arriving in Q4, Alibaba has already put 560,000 of its own accelerators into service — and because it owns the cloud they plug into, it has no procurement problem to solve at all. Add Enflame, Sugon (Dengfeng 8000), ZTE's OEX supernode, Kunlunxin, Inspur's REX81 SuperPoD 4096, and Moore Threads' 256-GPU fully-connected system (Xinhua, July 20), and the Chinese AI silicon market is not a two-way Nvidia-versus-Huawei fight. It's eight-plus domestic camps competing for the same substitution demand. Huawei's most dangerous rival may be in Hangzhou, not Santa Clara.

Expect Nvidia to play three cards. Deepen software lock-in — keep raising the migration cost across CUDA, NCCL and the inference stack. Lobby on policy — press the "controls build Huawei" argument to widen what it's allowed to sell into China. And demand third-party benchmarks. That third one stings most, because 8 EFLOPS, 6.7x, 15x and 62x are all vendor-reported figures with no MLPerf or equivalent independent validation. Self-reported numbers can carry headlines for a while. They cannot carry a procurement decision indefinitely.

So what actually changes

If you build AI infrastructure — nothing changes this quarter, but the direction is legible. The supernode is moving from "a rack as one GPU" to "a data center floor as one machine." That's why Lingqu 2.0 advertises 200m+ all-optical reach and microsecond RTT, and it's why Alibaba's cable-free orthogonal design attacks the same problem from a different angle. The concept worth internalizing is the globally unified memory address space: once model parameters reach trillion scale, the size and access latency of the addressable memory domain start determining the bottleneck more than your partitioning strategy does. If you're evaluating the Huawei stack specifically, do your own diligence on CANN maturity and on PyTorch, vLLM and SGLang support levels — real-world throughput is decided there, not by peak FLOPS.

If you're an investor — separate what was demonstrated from what was asserted. Demonstrated: an Atlas 950 SuperPoD exists as hardware at 1,024 cards, was exhibited in Shanghai, won the SAIL Award, and the prior-generation Ascend 384 has shipped 750+ deployments. Asserted by Huawei with no independent validation: 8 EFLOPS FP8, 16 EFLOPS FP4, 1,152TB, the 6.7x/15x/56.8x/62x comparisons, and the Q4 2026 availability date. And the real risk isn't performance — it's supply. SMIC's advanced-node yield is undisclosed and some analysts estimate around 40%; SMIC plans to roughly double capacity to about 70,000 wafers per month during 2026, but the absence of EUV is a structural ceiling. The harder bottleneck is memory: SemiAnalysis has argued in "Huawei Ascend Production Ramp: HBM is The Bottleneck" that CXMT's 2026 HBM output of roughly 2 million stacks translates to only about 250,000–300,000 Ascend 910C-equivalent packages. How many 8,192-NPU pods can physically be built is the ending of this story, and it will be written by CXMT and SMIC, not by Huawei's marketing.

If you follow policy — read this as an interim scorecard on export controls, and it cuts both ways. The controls worked: no EUV, HBM as the binding constraint, and Huawei's own chairman conceding a per-chip deficit. The controls also produced an adaptation: compensate for per-die inefficiency with volume and system integration, and absorb the TCO penalty by deploying where power is relatively cheap. But that adaptation has a geography. Selling a 160-cabinet, 1,000-square-meter system into markets with expensive electricity and expensive real estate is an entirely different proposition — and notably, Huawei has not published power consumption or cooling costs for this configuration. Until those numbers exist, every argument about export competitiveness is floating.

If you're a general reader — this doesn't touch your week. But it connects to something that eventually will. If Chinese companies can reliably train large models on domestic infrastructure, the flow of Chinese open-weight models that anyone worldwide can download and run for free continues — and over the past few years that flow has meaningfully shaped what independent developers and small companies can afford to build. Sixteen cabinets in a Shanghai exhibition hall look remote from anything you use. They're connected to the question of where the free models you might use in three years get trained.

🥄 Three Things You're Probably Wondering

— So what does this mean for me? Nothing directly. Indirectly, if Chinese firms can keep training frontier-scale models on domestic hardware, the supply of freely downloadable open-weight models keeps flowing — which is one of the main forces holding down what AI capability costs for everyone else. That's a multi-year link, not a this-week one.

— Is this actually ahead of Nvidia now? Too early to say, and the framing is off. The 6.7x claim compares one NVL144 rack to 160 cabinets of Atlas 950, and NVL144 hasn't shipped yet. Every headline figure here is Huawei-reported with no MLPerf-style independent validation. Eric Xu himself conceded Huawei trails Nvidia at the single-chip level and won't close that gap quickly.

— Does this mean China has solved its AI chip problem? Not yet. The binding constraint is manufacturing, not design. SMIC is running advanced nodes without EUV at an undisclosed yield, and analysts estimate CXMT's 2026 HBM output covers only around 250,000–300,000 Ascend-class packages. How many of these pods can be built is what settles the question, and that number has not been disclosed by anyone.

Sources

Numbers are as of announcement and may change.