LLM price barometer
Every list-price move measured between 15–19 Aug, biggest first. Percentages are on the blended 30/70 input/output price. Data snapshot: 2026-08-19.
Got more expensive (6)
- +328.4% DeepSeek V4 Pro 0813 (DeepSeek) — blended list price $0.74 to $3.17 per 1M tokens; input +203.4%, output +355.2%. Large move: worth confirming against the vendor's own rate card before you act on it.
- +109.1% GLM 5.2 (Zhipu) — blended list price $1.16 to $2.42 per 1M tokens; input +109.1%, output +109.1%. Large move: worth confirming against the vendor's own rate card before you act on it.
- +75.4% Kimi K2.6 (Moonshot) — blended list price $1.76 to $3.08 per 1M tokens; input +75.4%, output +75.4%.
- +25.8% DeepSeek V4 Flash 0423 (DeepSeek) — blended list price $0.11 to $0.14 per 1M tokens; input +25.7%, output +25.8%.
- +4.7% DeepSeek V3.1 Terminus (DeepSeek) — blended list price $0.75 to $0.78 per 1M tokens; input 0.0%, output +5.3%.
- +4.4% Qwen3 30B A3B (Alibaba) — blended list price $0.39 to $0.40 per 1M tokens; input +8.3%, output +4.0%.
Got cheaper (7)
- -44.8% Qwen3.6 27B (Alibaba) — blended list price $2.70 to $1.49 per 1M tokens; input -50.0%, output -44.4%.
- -21.1% Kimi K2.5 (Moonshot) — blended list price $2.17 to $1.71 per 1M tokens; input -21.1%, output -21.1%.
- -20.0% Nemotron 3.5 Lightning (Nvidia) — blended list price $0.20 to $0.16 per 1M tokens; input -20.0%, output -20.0%.
- -18.0% Gemma 4 26B A4B (Google) — blended list price $0.32 to $0.26 per 1M tokens; input -41.7%, output -15.0%.
- -13.2% Qwen3.5-122B-A10B (Alibaba) — blended list price $1.77 to $1.53 per 1M tokens; input -10.3%, output -13.3%.
- -9.1% GLM 4.6 (Zhipu) — blended list price $1.71 to $1.55 per 1M tokens; input -9.1%, output -9.1%.
- -1.1% Gemma 4 31B (Google) — blended list price $0.27 to $0.27 per 1M tokens; input -10.0%, output 0.0%.
How these figures are calculated
- The window is real: it runs from the newest snapshot back to the oldest one within the preceding seven days, and those two dates are the ones named above. While the archive is young the window is shorter than a week.
- Prices are blended 30/70 input to output, the same figure the Efficiency score ranks on.
- List prices only. Batch and cache discounts are conditional on the customer implementing them.
- Corrections are excluded rather than annotated: when a published rate was wrong at the source and later fixed, the difference is a data correction, not a price move.
- A missing, zero or negative baseline yields no entry. New listings, renames and delistings get no percentage rather than an invented one.
The full cost-per-request formulas are in the cost per request methodology guide. For what to do about a price move, see the Cloud FinOps skill on GitHub.