FinOps guides on LLM and cloud pricing
Practical FinOps guides on LLM and cloud pricing: batch and cache economics, model selection, and what the published rate cards leave out. Every figure is recomputed from the catalogue at build time. Data snapshot: 2026-09-29.
- ARM against x86: what Graviton, Ampere and Axion save
The 20% ARM discount is a discount against Intel. Measured at matched vCPU and memory across 4 clouds, it halves against AMD. Current prices. - Batch API economics: when the -50% applies
Batch APIs cut LLM token prices roughly in half, but only for async workloads and only where a batch rate is published. Worked examples, current prices. - Context windows: what a million tokens costs to fill
A 1M context window is a spec-sheet capability and an invoice line. What one full prompt costs across the catalogue, and what re-reading it costs. - The cost of reasoning: what thinking models charge for output
Reasoning tokens are billed as output tokens, where prompt caching cannot reach them. What a thinking budget costs across the catalogue, at current prices. - How we compute cost per request
The exact formulas behind the list-price and after-discounts columns: token profiles, cache and batch rules, what the numbers exclude and where they mislead. - Open weights against proprietary APIs: the cost gap in numbers
Open-licensed models undercut proprietary APIs at matched quality across most of the range, but the open catalogue stops below the top tier. Live prices. - Prompt caching: the discount nobody budgets for
Cache reads bill at a fraction of the input price, but the saving that reaches the bill is far smaller than the headline. Break-even by use case, live prices. - Quality per dollar: the models that beat their price class
A model's price class comes from its rate card, not its quality. Measured against Arena scores, most of the catalogue has a cheaper equal. Current prices. - Bedrock, Vertex, Azure or direct: what the channel changes
The same model is sold by its maker and 3 hyperscalers. Across the catalogue the token rate is usually identical; the regions and commitment terms are not.