跳到正文
北京时间
原文
Epoch AI:研究、数据与评测· Venkat Somala·· 2026-08-13精选AI 评分61

Epoch AI 估算 AI 芯片性价比每年增长约 49%

The performance per dollar of AI chips purchased each quarter has grown by an average of 49% per year

AI 导读

Epoch AI 估算 2023 年 Q1 至 2025 年 Q4 每季度购买的 AI 芯片综合性价比,从约 5.6×10^11 升至约 1.4×10^12 bit-operations per second per dollar(2025 年不变美元),指数拟合平均每年增长约 49%(90% CI:36%–66%),倍增时间 1.7 年。

推荐理由

Epoch AI 用自建的芯片销量和价格数据量化了每季度购买 AI 芯片的性价比增长,并说明了 Nvidia 高利润率对均价的影响。

正文

Learn more about this graph

We estimate the aggregate performance per dollar of the AI chips purchased in each quarter. Performance per dollar rose from about 5.6 × 1011 bit-operations per second per dollar in the first quarter of 2023 to about 1.4 × 1012 by the end of 2025, in constant 2025 dollars. Fitting an exponential trend gives an average growth rate of about 49% per year (90% CI: 36% to 66% per year).

The gains are largely due to chips getting exponentially more powerful with each new generation. While recent chips have become more expensive, they have grown more powerful at an even faster rate. In 2025 dollars, NVIDIA’s GB300 costs about 5.5 times the P100’s 2016 launch price, yet it delivers roughly 200 times the performance, making it about 37 times more cost effective.

While chips are available with performance per dollar as high as 3.6 × 1012 bit-operations per second per dollar (Google’s TPU v6e), the average across each quarter’s purchases is lower. Most AI hardware spending currently goes towards Nvidia’s GPUs, which command a relatively high margin and thus have a lower price-perf than some custom ASICs.

Data

Our analysis draws data from several Epoch datasets:

  • Chip sales: Quarterly units of each accelerator, from our AI Chip Sales Hub. We begin the series in 2023 because coverage before then is sparse, and end it in Q4 2025 because sales estimates for later quarters are not yet complete. Note that the Chip Sales hub, and thus this analysis, models the amount of units and compute sold rather than deployed each quarter.

  • Performance: We use each chip’s Total Processing Performance (TPP), a precision-normalized throughput metric, equal to peak operations per second multiplied by the bit width of that number format.

  • Price: We use one representative purchase price per chip. These are the purchase prices rather than the rental prices, and they cover the accelerator card or module rather than a full server. Most of these chips have no single MSRP, so we estimate what buyers were paying around the time each chip started shipping in volume. For Nvidia, AMD, and Huawei we use reported transaction prices, executive statements, and order values, and take a geometric mean when we only have a range.

    Google and Amazon did not sell their chips directly, so we use what it costs them to procure the chips from their manufacturing partners, based on BOM modeling plus the partner’s margin (see our Chip Sales Hub methodology).

    For the H100/H200, whose price we model as slightly declining over its long sales run, we use year-specific prices: $30,000 in 2023, $28,000 in 2024, and $26,000 in 2025. Other chips keep one price across all quarters, so we do not capture any later price cuts or spot prices. You can find the price data here.

The analysis covers 24 chips with both sales and price data: Nvidia A100, A800, H100/H200, H800, H20, GB200, and GB300; AMD Instinct MI250X, MI300A, MI300X, MI308X, MI325X, MI350X, and MI355X; Google TPU v4, v4i, v5e, v5p, v6e, and v7; Huawei Ascend 910B and 910C; Amazon Trainium2; and Cambricon Siyuan 590. The H100/H200 are combined in the sales data, and we use H100 specs and prices for the combined line. For the Blackwell generation we use the GB200 and GB300 per-GPU specifications.

For the plot, we group the A800, the H800, AMD’s Instinct GPUs, Huawei’s Ascend chips, Cambricon’s Siyuan 590, and all Google TPUs except the v6e into a combined “Other” category.

Analysis

For each quarter (Q), performance per dollar across all chips sold is:

\[\text{perf}/\$(Q) = \frac{\sum_i \text{TPP}_i \times \text{units}_i(Q)}{\sum_i \text{price}_i \times \text{units}_i(Q)}\]

where the unit term is the number of chip i sold within quarter Q, and TPP is its Total Processing Performance. The output is a dollar-weighted average of each chip’s performance per dollar, weighted by its share of that quarter’s spending. Since Nvidia accounted for most spending from 2023-2025, the average is pulled toward its higher-priced hardware.

Spending is adjusted for inflation using the consumer price index. Each quarter’s new spending is converted into Q4 2025 dollars at that quarter’s index value, so all figures are in constant 2025 dollars. There was no CPI published for October 2025, due to the government shutdown, so the Q4 2025 anchor averages November and December.

A log-linear fit over the 12 quarters from Q1 2023 to Q4 2025 gives an annual growth rate of about 49%, a doubling time of 1.7 years. The 90% confidence interval, from bootstrap resampling, is roughly 36% to 66% per year, or doubling times between 1.4 and 2.3 years. In nominal dollars, the growth rate is 45% per year. The growth comes in spurts: price-performance was nearly flat at 6% per year from 2023 to mid-2024, and then grew to roughly double per year as Blackwell-generation chips took over spending. Chips bought in 2024 averaged 23% better performance per dollar than 2023 purchases, and 2025 purchases averaged 91% better than 2024’s.

Assumptions and limitations

Our Chip Sales estimates, and thus the volumes used in this analysis, are chips that are sold rather than deployed. This is a price-performance measure of the hardware that has shipped, not of active computing capacity. Active capacity may differ both because old chips have depreciated, and because newly-sold chips have not yet become operational.

We use peak theoretical performance as given by each chip’s spec sheet, converted to peak bit-operations per second. Real-world performance depends heavily on software support for the hardware, the model architecture, the workload shape, the rack-scale system design, the networking topology, and memory and interconnect bandwidth. Realized performance per dollar on any given workload can differ substantially from on-paper specs.

Prices vary across vendors and designers. We price each chip at its cost to the primary owner of that compute. Nvidia, AMD, and Huawei primarily sell accelerators to external customers, so we use external sale prices, which include the designer’s margin. Google’s and Amazon’s accelerators are built primarily to serve their own internal workloads and to be rented to customers through their cloud platforms. Both companies design their chips in-house and procure them from manufacturing partners, paying production cost plus the margin of their manufacturing partners, which is typically lower than Nvidia margins. Their prices in this analysis are therefore closer to procurement cost than to a market price, which flatters TPU and Trainium price-performance relative to what an external buyer would likely pay. The margin Google and Amazon charge when renting this hardware to customers through their cloud platforms is not counted here. Google and Broadcom have recently begun selling TPUs directly to external customers, but those sales do not yet appear in our data.

Transaction prices vary. We use a single representative price per chip, while actual sale prices vary by customer, purchase volume, and timing, and are especially uncertain for chips sold into restricted markets.

The series covers 24 accelerators with price data. Coverage is broad but not complete. The analysis excludes Amazon’s first-generation Trainium and Inferentia, for which we lack a price estimate (less than half a percent of tracked compute sold), and does not cover accelerators our Chip Sales Hub does not track, including Intel’s Gaudi, Groq, Cerebras, Microsoft’s Maia, Meta’s MTIA, Tesla’s Dojo, etc.

Download this data

Performance per dollar of AI chips bought each quarter (2025 USD)

Explore this data

AI Chip Sales

AI Chip Sales

AI chip sales data.

来源:Epoch AI:研究、数据与评测 · epoch.ai