Token Pricing: How China and the US Reprice, and the New Tiers
This board answers four things: are token prices still falling, who is raising and who is cutting, which tier is moving, and where the cost per unit of intelligence is heading.
Read from this board: List prices stopped falling in the second half of 2025 while the cost per unit of intelligence keeps collapsing; since September 22 the flagship tier has held while the mainstream tier cut on a new generation, so the tight-compute signal now lives in the flagship tier and in Chinese vendors' price increases.
Cheapest US flagship
$4
Sep 2026|blended, per million tokens
DeepSeek mainstream price
$1.98
Sep 2026|peak-hour price
Cheapest way to a score of 40
$0.237
Sep 2026|AA intelligence index, blended price
Chinese vendors' token share
77.6% / 22.4%
2026-09-20|tokens / spend, Vercel gateway 7-day avg
How to read this board
Core view: From 2023 to mid-2025 token prices deflated steeply. The turn came in the second half of 2025: new flagship generations stopped getting cheaper. When compute costs stop falling, model vendors stop passing savings on, and the end of token deflation is price evidence of a tight market. From September 22, 2026 it has to be read in two layers: the $20 flagship tier did not move, while three vendors cut their mainstream tier on the same day, so the mainstream tier is back to fighting for volume.
- The US rule is that a model's list price almost never changes in its lifetime; repricing happens by launching a new model. Counting same-model changes alone makes the US vendors look like they only cut, because every increase comes through a new generation.
- The Chinese rule is launch at a promotional price, revert when it expires, then adjust around events; DeepSeek raised the same mainstream model three times in two years.
- List price and cost per unit of intelligence are two different things: list prices have stopped falling while the lowest cost at equal ability keeps collapsing, so a smaller application-layer bill may just mean a cheaper model of equal ability.
- List price is not task cost: a chatty model can cost more even when it lists cheaper, and the same model at different effort levels can differ five-fold in task cost.
Term: Blended price = (3 × input + 1 × output) ÷ 4, per million tokens at list. It is a list price, not a transaction price, and excludes cache, batch and enterprise discounts. The mainstream tier means models at $3 to $5.6; the flagship tier means $10 to $20.
1All pricing events: same-model repricing plus generational repricing
Dark bars are repricing of the same model; light bars are the price gap between a new generation and the one before. Looking at same-model changes alone makes the US vendors look like they only cut, because their increases all come through new generations. From July 2025 to September 2026 the US vendors raised 13 times and cut 19; 12 of the 13 raises came from OpenAI's and Google's mainstream and mid tiers. China raised 12 times and cut 4, a more consistent direction. The three generational cuts on September 22 were the first time the mainstream tier moved down together.
2Launch-price ladders: US versus China
Each line is the launch-month list price of successive generations in the same product tier. US: OpenAI's mainstream line climbed from the GPT-5 low to $11.25, then GPT-6 Sol dropped to $4 on September 22, the first move down since August 2025; Google Flash went from $0.175 to $3 and then halved to $1.5; Anthropic went the other way, cutting Opus to $10 and then $8 while opening a $20 Fable tier above it; xAI cut with each generation to win share. China's three flagship lines only went up: Kimi five-fold, GLM more than double, Qwen up by half before a generational cut.
US Models: Launch Price by Generation, Same Tier (USD per Million Tokens, log)#
How to read this chart
Each line is the launch-month list price of successive generations in the same product tier, not same-model repricing. Diamonds mark $20 premium SKUs.
Source: Published model pricing and the Artificial Analysis index; FinSight compilation · Updated 2026-09-30
China Flagship Lines: Launch Price by Generation (USD per Million Tokens, log)#
How to read this chart
Same y-range as the US chart for comparison. Kimi K2 to K3 went from $1 to $6; GLM 4.6 to 5.2 from $0.96 to $2.15; Qwen Max 3.8 fell back to $3 at its generation change.
Source: Published model pricing and the Artificial Analysis index; FinSight compilation · Updated 2026-09-30
3DeepSeek's mainstream model: three increases in two years
The mainstream chain runs from deepseek-chat through V4-Pro. Dashed lines are pricing events: February 2025 promotion expiry back to list (+173%), September V3.1 output up 50% with the night discount removed, December V3.2 cut 64%, June 2026 the V4-Pro generation, July double pricing at peak hours, August 16 five times at peak and 2.5 times off-peak. Once September list prices reflected it, the blended price went from $0.544 to $1.98, a new high for the chain. Volume did not fall after the increase, it doubled; the driver was supply that could not keep up, a price clearing a tight market rather than weak demand.
DeepSeek Main Models: Blended Price and Pricing Events (USD per Million Tokens, log)#
How to read this chart
The main line runs deepseek-chat from V2 through V3.2 to V4-Pro. Dashed lines are pricing events: the February 2025 promotion ending, the September increase and end of night discounts, the December cut, the June 2026 generation change, double pricing at peak hours from July, and another increase in August.
Source: Published model pricing and the Artificial Analysis index; FinSight compilation · Updated 2026-09-30
4Flagship tiers, US versus China: what each dollar buys
One row per model; x is the blended price, the bracket is the Artificial Analysis intelligence index. The blue band is the mainstream tier at $3 to $5.6, the orange band the flagship tier at $10 to $20. As of September 30: the two $20 flagships, Fable 5.1 and Astra, score below Opus 5.5 at $8 and Sonnet 5.5 at $4, the first time the flagship tier is not where the top scores sit; four new generations entered the mainstream tier within a month; DeepSeek V4-Pro at its $1.98 peak price scores only 36.
China vs US Flagship Price Bands (AA Index v4.3, Snapshot Sep 30 2026, Pricing as of Sep 22)#
How to read this chart
One model per row; x is blended price (log), the bracket is the intelligence score. The blue band is the workhorse band at $3 to $5.6, the orange band the flagship band at $10 to $20.
Source: Published model pricing and the Artificial Analysis index; FinSight compilation · Updated 2026-09-30
5Paying for speed: fast channels became a standard price axis
Left: the premium a fast or peak-hour channel charges over the standard price for the same model, mostly converging around 2x. Right: who started selling a fast channel and when, from xAI's first in April 2025 to Astra shipping with one at launch in September 2026, one vendor to six in 18 months. Token pricing has split into several dimensions: intelligence, speed, time of day and data rights.
"Paying for Speed": What the Fast Lane Costs (Multiple of List)#
How to read this chart
The fast lane's price as a multiple of the standard price for the same model. Only OpenAI publishes the speed gain (2x to 2.5x speed at 2x price, same intelligence).
Source: Published model pricing and the Artificial Analysis index; FinSight compilation · Updated 2026-09-30
"Paying for Speed" Spreads: Who Started Selling a Fast Lane, and When#
How to read this chart
X is when each fast lane first appeared; y is its multiple of the standard price.
Source: Published model pricing and the Artificial Analysis index; FinSight compilation · Updated 2026-10-01
6Price of intelligence: list prices held, cost per unit of intelligence still collapsing
Left: one dot per model, colour by launch period, dashed lines are each period's price frontier; on September 30 the top of the frontier sat outside the $20 tier for the first time, with Opus 5.5 buying the highest score at $8. Right: the lowest cost of reaching an intelligence score of 30, 40 or 50; the 40 line fell from $10 to $0.24 in four months, the 50 line from $10 to $4 in two. A falling application-layer bill may just mean a switch to a cheaper model of equal ability, not falling usage. Note that the cheap end of the frontier is mostly Chinese and open models; US enterprises still buy mainly from the big four.
Intelligence Index vs Blended Price (AA Index v4.3, Snapshot Sep 30 2026)#
How to read this chart
One dot per model (intelligence score 20 or above, 246 models). Color is launch period; dashed lines are the cumulative price-performance frontier for each period (list-price basis).
Source: Published model pricing and the Artificial Analysis index; FinSight compilation · Updated 2026-10-01
Cheapest Way to Buy the Same Capability, Over Time (log)#
How to read this chart
The lowest blended price on record to reach intelligence scores of 30, 40, 50 and 60, by launch date at current prices. The 60-plus line: $20 in June, $8 in July, $3 in August, $2 in September.
Source: Published model pricing and the Artificial Analysis index; FinSight compilation · Updated 2026-09-30
7The chattiness tax: list price and task cost can come apart
Left: how many tokens each model burns on the same benchmark; Grok 4.7 lists at $3 but talks the most, so its task cost is high, while GPT-6.1 Sol reaches a high score with the fewest tokens. Right: CursorBench 4.0, an agentic coding benchmark; each line is one model swept across effort levels; Opus 5.5 at high effort scores almost the same as its maximum for a third of the money, the new value sweet spot; task cost and list price can differ by 17 times.
How Chatty Each Model Is: Tokens Burned Running the Same Benchmark (Implied, Millions)#
How to read this chart
Implied consumption = task cost ÷ list price, for the 17 models scoring 40 or above. Red is China, blue is the US.
Source: Published model pricing and the Artificial Analysis index; FinSight compilation · Updated 2026-09-30
CursorBench 4.0: Agentic Coding Score × Cost per Task (Effort Sweep, Snapshot Sep 29 2026)#
How to read this chart
Cursor's official benchmark, built from real multi-file, loosely specified tasks. Version 4.0 (Sep 10 2026) added long-horizon tasks, so scores are not comparable with 3.2: the top score fell from 73% to 52% because the tasks got harder, not because models got worse. X is actual cost per task (log), y is score; dots of one color on one line are the same model swept across effort settings from low to high. Chinese models and the previous generation are not on the 4.0 board.
Source: Cursor public benchmarks; FinSight compilation · Updated 2026-09-30
8Real traffic: Chinese vendors' volume and money after the price increase
A direction gauge, not a market-wide meter: Vercel AI Gateway open data, a sample of overseas app developers that excludes direct enterprise contracts and traffic inside China, so read the direction, not the level. Chinese vendors combined: their share of tokens rose to 77% in the six weeks after the August 16 price increase and held there, their share of spend rose two weeks running to 29%, and their share of requests sits at 44%. Volume held, price went up, money came in: demand looks more price-insensitive, not less.
China Models Combined Share: Tokens vs Dollars vs Requests (7-Day Average)#
How to read this chart
China's models combined. Five weeks after the price increase token share rose from 53% to 78%, almost all from DeepSeek's new release; dollar share went from 15% to 22% and request share is 44%. Remember these are shares, and the sample is Vercel's overseas developers.
Source: FinSight compilation and estimates · Updated 2026-09-30
9Extension: flagship and small-model chains
Left: blended price of each vendor's flagship relay across generations; steep deflation from 2023 to mid-2025, and from the second half of 2025 new flagship generations stopped getting cheaper. Right: the small-model line is the cost floor for high-volume, low-value tokens and decides whether AI features fit inside free products; small models are rising too, Haiku 4.5 costs four times 3 Haiku, and only OpenAI keeps pushing small-model prices down.
Flagship Lines by Vendor: Blended Price (USD per Million Tokens, log)#
How to read this chart
Each line hands off from one flagship generation to the next, so it jumps at each generation change; hover to see the model of the month. List prices exclude caching and batch discounts.
Source: Published model pricing and the Artificial Analysis index; FinSight compilation · Updated 2026-09-30
Small-Model Lines: Blended Price (log)#
Source: FinSight compilation and estimates · Updated 2026-09-30
10Detail: same-model repricing events, generally available models
From the LiteLLM monthly price history; the size is the month-on-month change in blended price. Chinese vendors' time-limited promotions and off-peak discounts mostly do not reach list prices, so their actual count of changes is understated.
Detail: Same-Model Repricing on GA Releases#
| Camp | When | Model | Change |
|---|---|---|---|
| 中國 | 2025/02 | deepseek-chat(V3 促銷到期) | +173% |
| 中國 | 2025/09 | deepseek-chat(V3.1、取消夜間優惠) | +83% |
| 中國 | 2025/12 | deepseek-chat(V3.2) | -64% |
| 中國 | 2025/09 | deepseek-reasoner(V3.1 統一定價) | -9% |
| 中國 | 2025/12 | deepseek-reasoner(V3.2) | -64% |
| 中國 | 2026/09 | deepseek-v4-pro(8/16 漲價,尖峰價口徑) | +264% |
| 中國 | 2026/09 | deepseek-v4-flash(同上) | +277% |
| 美系 | 2023/11 | claude-2 | -27% |
| 美系 | 2023/11 | claude-instant-1.2 | -90% |
| 美系 | 2024/07 | gemini-1.5-pro(預覽價轉正式價) | +900% |
| 美系 | 2024/09 | gemini-1.5-flash | -75% |
| 美系 | 2024/11 | gpt-4o | -42% |
| 美系 | 2025/01 | o1-mini | -63% |
| 美系 | 2025/02 | claude-3-5-haiku | -20% |
| 美系 | 2025/06 | o3 | -80% |
| 美系 | 2025/09 | gpt-3.5-turbo | -54% |
| 美系 | 2026/03 | gemini-1.5-flash | -57% |
| 美系 | 2026/08 | gpt-5.6-luna | -80% |
| 美系 | 2026/08 | gpt-5.6-terra | -20% |
| 美系 | 2026/09 | gpt-5.6-sol(8/21 起限時 3 個月,美系首見促銷價) | -29% |
Source: Published model pricing and the Artificial Analysis index; FinSight compilation · Updated 2026-10-01
