跳到正文
北京时间
原文
Epoch AI:Gradient Updates· Venkat Somala·· 11 天前精选AI 评分65

Epoch AI 分析:华为 AI 芯片到 2030 年仍落后 Nvidia 约 4 年

How far behind Nvidia is Huawei?

AI 导读

Epoch AI 报告估算,2026 年华为总 AI 算力产出约为 Nvidia 的 1/25,Ascend 950 性能约为 B300 的 1/7、2022 年 H100 的一半。报告认为即便华为路线图全部兑现并计入走私 HBM 情景,到 2030 年其在芯片性能和算力产出上仍落后 Nvidia 约 4 年,出口管制是主要制约因素。

推荐理由

报告用量化估算比较双方芯片性能、产量和 HBM 供给,读者可以据此理解出口管制下差距的成因与走向。

正文 · 原文

This is a summary of a more detailed report, available on our website.


The US-China AI competition spans everything from models and data centers to chips and the manufacturing equipment used to make them. Chips and manufacturing equipment form the foundation of that stack and determine how much compute each side can field.

This is why semiconductor export controls have been a centerpiece of US AI policy, and why so much depends on how China’s homegrown chip industry stacks up against America’s.

Huawei, China’s leading AI chip designer, laid out its chip roadmap last year and showcased an accelerated version last week at its annual conference. The roadmap is ambitious: it plans to pack more compute into each chip, connect far more chips into a single high-speed system, and mature the software that determines how much of that hardware is actually used.

But Huawei currently lags Nvidia on both per-chip performance and the number of chips produced. Multiply the two, and Huawei will produce roughly 25× less total compute than Nvidia in 2026. US export controls constrain its ability to improve either factor, while Nvidia continues to push the frontier.

With Nvidia making strides, can Huawei close the gap?

After crunching the numbers, we think almost certainly not. Between now and 2030, Huawei will likely remain around four years behind in both chip performance and compute production.

Huawei’s starting position in 2026

In 2026, Huawei is significantly behind Nvidia. To understand how far Huawei needs to go to catch up, we can compare Huawei and Nvidia on the total compute each produces, which is the number of chips each makes multiplied by the performance of those chips. We convert all performance measurements to the equivalent number of Nvidia H100 chips, called H100-equivalents (H100e).

Performance: Huawei’s flagship AI chip for 2026, the Ascend 950, delivers roughly 7× less compute throughput than Nvidia’s B300 and half that of Nvidia’s 2022 H100. In other words, Huawei’s latest and most powerful GPU delivers only half the performance of a two-generation-old Nvidia chip that first started shipping in 2022, suggesting that Huawei’s chip performance currently trails Nvidia’s by four years.1

Volume: For 2026, we estimate that Huawei will produce 1.5 million chips — one-fourth of Nvidia’s roughly 6 million.

Total compute: With one-fourth the chips and each chip one-seventh the performance, Huawei will produce roughly 25× less compute than Nvidia in 2026.

Huawei’s roadmap to improve performance by 2030

Closing that gap would require major progress in both individual chips and the systems connecting them — and Huawei has laid out an ambitious roadmap to improve both.

Chip performance

Huawei starts out at a disadvantage on chip performance because Huawei’s fabrication partner SMIC can’t pack transistors as densely as Nvidia’s partner TSMC can. Transistors are the building blocks of a chip’s performance, so greater transistor density allows engineers to pack more compute onto one chip.

US export controls restrict sales to China of the advanced lithography machines needed to produce the densest chips, and China’s efforts to produce these machines domestically likely won’t materialize until after 2030.

Huawei’s roadmap targets a 7x performance improvement over two Ascend generations. Since the Ascend chips can’t scale transistor density quickly enough, it must gain performance improvements from other methods.

Bigger chips: First, Huawei can make its chips more powerful by making them bigger and packing each with more or larger logic dies (the parts that perform the computation), but the trade-off is higher costs and more defects. Meanwhile, Nvidia is also expanding its packages, possibly with fewer defects given its more mature process.

Chip Architecture: Huawei’s updated Ascend roadmap optimizes performance for lower precision formats, likely to match what frontier workloads are moving towards. This architecture update will allow the Ascend chips to perform more calculations with the same amount of silicon, significantly improving the Ascend’s performance running frontier AI workloads.

Vertical stacking: Huawei’s big, long-term bet is to stack logic dies vertically to fit far more transistors within a given footprint than SMIC’s process can print in a single layer. Stacking also shortens communication time within the chip, which allows the chip to run faster with less power. However, this technique, which Huawei calls LogicFolding, won’t reach the Ascend line until 2030. Nvidia’s Feynman generation, expected in 2028, will reportedly stack at least two layers of its superior logic dies, two years before the first Ascend attempts to do so. This would compound the 2× advantage from transistor density into a 4× advantage.

Faster systems

Like all major chip designers, Huawei is focused on optimizing the entire computing system, rather than just the individual chip.

A decade ago, frontier models could be trained on a handful of GPUs. Today, training frontier models requires tens of thousands of AI chips. At that scale, performance depends heavily on how quickly chips can communicate and how effectively software can coordinate work across them. Huawei is pursuing both levers.

Larger systems: Huawei plans to connect more chips inside a single domain. A domain is a group of chips wired together tightly so that they behave more like one giant system, which reduces the time spent waiting on communication from other parts of the system.

Nvidia’s domain focuses on having extremely fast links within one rack of 72 chips. Inside the rack, any chip can talk to any other chip at full speed. Since Nvidia’s superior chips can keep communication-intensive work within fewer chips, it doesn’t need to compromise on communication speed for a larger domain size.

Huawei goes the other way, compensating for weaker chips with larger domains; its planned “SuperPoDs” will connect 8,192 chips. However, maintaining high-bandwidth simultaneous communication from thousands of chips comes with unfeasibly steep trade-offs. Instead, Huawei implements a communication hierarchy where closer chips have more bandwidth, and more distant chips have to go through several hops. Whether this will perform well depends on Huawei’s software getting better at keeping communication local. Unfortunately, it’s hard to say how well the SuperPoD will perform in practice because Huawei has not published key performance metrics.

Better software: The software that coordinates workloads across chips can move performance by more than 10×, and it’s currently one of Nvidia’s strongest levers. Huawei’s equivalent of Nvidia’s CUDA is called CANN, which was first released in 2018 and makes working with Ascend chips difficult. When DeepSeek tried to train its R2 model on Ascend chips, Huawei sent its own engineers in and still couldn’t complete a training run, prompting DeepSeek to revert to Nvidia. However, if Huawei can capture the feedback loop from having Chinese labs use its chips instead of Nvidia’s, the software gap could close fast. For now, though, that possibility is still nascent: CANN remains years behind CUDA, and while Chinese labs are increasingly using domestic chips, they still reach for Nvidia hardware when they can get it.

Huawei will remain three to four years behind in chip performance

On paper, Nvidia’s B300 delivers roughly 7× more compute performance than Huawei’s Ascend 950 series. As Nvidia moves from its Blackwell chips to Rubin and Rubin Ultra, the gap between each company’s best chip will widen to roughly 9× by 2027. The first Ascend chip expected to come close to the B300 in peak performance is the 970, slated to ship in late 2028. Given that the B300 arrived in 2025, that would put Huawei roughly three to four years behind Nvidia at the chip level even out to 2029.

Can Huawei make up the difference in volume?

Suppose Huawei delivers on all of this: the Ascend line hits its roadmap targets, the SuperPoDs perform well when scaled to 8,000 chips, and CANN matures. Even then, its chip performance would still trail Nvidia’s by years, since Nvidia is pulling the same levers from a more advanced base.

If Huawei cannot match Nvidia’s performance, whether chip-for-chip or system-for-system, could it compensate by producing more chips?

Using domestic HBM: Huawei is currently bottlenecked on high-bandwidth memory (HBM), as China’s leading memory manufacturer, CXMT, and the smaller XMC are only beginning to ramp up HBM production, and their supply cannot yet meet the demand from Chinese chip designers.

Across 2026–2028, we estimate that Nvidia could produce tens of millions of H100-equivalents a year, rising to nearly 100 million by 2028. Huawei, using CXMT and XMC as its sole HBM sources, could increase production to around 1.5 million H100e a year by 2028, roughly 1.5% of Nvidia’s output.

By 2030, China’s domestic HBM production is expected to ramp substantially — optimistically by 18× its 2026 production. However, our back-of-the-envelope calculation suggests that Huawei’s optimistic 2030 compute production will be about what Nvidia will produce this year. In other words, Huawei’s compute output will continue to trail Nvidia’s by roughly 4 years.

Including smuggled HBM: In 2025, the gap was filled by illegally acquired components and a stockpile of more than 10 million foreign HBM stacks that Chinese firms built up before the December 2024 export controls took effect.2 That stockpile is expected to run out in 2026, forcing Chinese AI chip volumes to fall back to what CXMT can produce and what China can smuggle.

So could more smuggling be the solution? Suppose Huawei smuggled ten times its domestic HBM in 2028 — an implausibly large flow that would cost over $40 billion even at 2025 HBM prices. Under this scenario, Huawei’s output would go from 1.5 million H100e to roughly 16 million, still only about 16% of Nvidia’s production that year.

Greater HBM supply, whether smuggled or domestic, moves Huawei’s output but does not close the gap.

Export controls are biting

Huawei began 2026 producing roughly 25× less AI compute than Nvidia. Its chips are roughly three to four years behind, and the supply chain that produces them cannot match Nvidia’s.

Between now and 2030, Huawei will likely remain around four years behind in both chip performance and compute production. Export controls drive much of this disadvantage, hitting both volume and performance. While not stopping Huawei’s progress entirely, export controls have made it far harder, slower, and more expensive for Huawei to compete.

Meanwhile, Nvidia is compounding its lead. Nvidia’s HBM suppliers — SK Hynix, Samsung, and Micron — will spend hundreds of billions of dollars through 2030 building fabrication plants. Nvidia’s Feynman is likely to deliver a significant performance boost over Rubin, and by 2030 Nvidia could even start shipping Feynman’s successor.

Huawei is one of the world’s most capable engineering organizations, and it may eventually catch up — by outperforming Nvidia on vertical stacking, improving its software dramatically enough to close the gap with CUDA, gaining access to cutting-edge lithography machines, and significantly scaling production. But overcoming these disadvantages will take time, and the available information strongly suggests it will not happen by 2030.

This is a summary of a more detailed report, available on our website.

1

Nvidia’s real lead is likely even bigger than this comparison suggests, because Huawei’s CANN software stack is far less mature than Nvidia's CUDA, so less of each Ascend chip’s potential gets used.

2

Much of the Ascend production to date has relied on foreign components. A teardown reportedly found that the Ascend 910C parts use TSMC logic and Samsung or SK Hynix HBM rather than fully domestic silicon. Even if Huawei hits the performance targets on its future Ascend roadmap, it remains an open question whether China's domestic supply chain can provide the key components at the scale needed to produce millions of Ascend units.

来源:Epoch AI:Gradient Updates · epochai.substack.com