How we compute cost per request
A specification, not an estimate. These are the formulas behind every cost-per-request figure on this site.
Every model row carries two cost figures: Cost / 1K req (list price) and Cost / 1K req (after discounts). The table prints them per 1,000 requests so a column of them can be compared without counting zeros. The formulas and worked examples below are per single request, so multiply by 1,000 to land on the table's figure. They answer different questions, and the gap between them is the only part of the number you control. This page gives both formulas in full and the assumptions behind them. More usefully, it lists the things they leave out on purpose, because that is where a budget built on them goes wrong.
The list-price formula
Cost per request at the published rate card, with no optimisation assumed:
cost = (inputTokens / 1,000,000) x inputPrice
+ (outputTokens / 1,000,000) x outputPrice
That is the whole of it. The token counts come from the use-case profile you pick in the scenario bar; the prices come from the catalogue. Nothing else enters.
Worked, on Claude Sonnet 5 (Anthropic, $2.00/M in, $10.00/M out) at the Support Ticket profile:
| Component | Tokens | Rate /1M | Cost |
|---|---|---|---|
| Input | 1,500 | $2.00 | $0.0030 |
| Output | 500 | $10.00 | $0.0050 |
| Cost per request (list price) | $0.0080 |
The after-discounts formula
Same request, priced as if you had implemented the two optimisations the provider publishes a rate for. Four rules, applied in order:
- Batch substitution. If the use case is asynchronous and the model publishes both batch rates, the batch input and output rates replace the list rates outright (where that discount is real).
- Cache read on the hit-rate share. The profile's cache hit rate is the fraction of input tokens billed at the published cache-read rate; the remainder bills at whatever rate rule 1 left in place (what that discount is worth).
- Cache write, amortised. A token can only be read from the cache after it has been written there, and providers that publish a write rate charge more for it than for ordinary input. The formula assumes one token written for every 5 read. A written token is one of the cache misses, already billed at the input rate by rule 2. What is added is the premium: the published write rate minus that input rate, never below zero. Under batch the write rate is taken as published, because no provider publishes a batch rate for it. A model that publishes a read rate but no write rate is assumed to write at the ordinary input rate: premium zero.
- Fall back, never invent. A model that publishes no batch rate keeps list prices for rule 1. A model that publishes no cache-read rate has its cached share billed at the ordinary input rate, which makes the discount zero and leaves no write to charge. And if the write premium outweighs the read saving, the input side is billed uncached, since caching that does not pay is caching you would not switch on: the figure is never above list.
writePremium = max(0, cacheWriteRate - inputRate) (0 when no write rate is published)
writtenShare = min(hitRate / 5, 1 - hitRate)
effectiveInput = cacheReadRate x hitRate
+ inputRate x (1 - hitRate)
+ writePremium x writtenShare
effectiveInput = min(effectiveInput, inputRate)
cost = (inputTokens / 1,000,000) x effectiveInput
+ (outputTokens / 1,000,000) x outputRate
The 5 is a modelling assumption, not a published figure: how often a cached prefix is re-read before it must be rewritten is a property of your traffic. In a conversation, each turn re-reads the earlier turns and writes its own, which works out to (n - 1) / 2 reads per write over n turns, so 5 corresponds to a 11-turn session. A prompt shared by all requests under steady traffic does far better, and the same prompt at low volume does worse, because entries expire between requests. 45 of 292 priced models publish a write rate.
Rule 4 is the one that matters when reading the table. The after-discounts column is never lower than what the provider will honour, so where a model publishes nothing, the two figures are identical and the table prints a dash in place of the second one. A dash means "nothing published", not "no savings available".
The same request, now at the Support Ticket profile's 60% cache hit rate (this profile is synchronous, so no batch rate applies):
| Component | Tokens | Rate /1M | Cost |
|---|---|---|---|
| Input, cache hit (60%) | 900 | $0.20 | $0.000180 |
| Input, cache miss | 600 | $2.00 | $0.0012 |
| Cache write premium, amortised (1 write per 5 reads) | 180 | $0.50 | $0.000090 |
| Output | 500 | $10.00 | $0.0050 |
| Cost per request (after discounts) | $0.0065 |
Claude Sonnet 5 publishes a cache-write rate of $2.50/M, 1.25x its input rate, so the written tokens cost the difference on top of the input rate their cache miss already paid.
A 19.1% reduction, entirely from prompt caching. Note where it comes from: output is untouched, so the saving is bounded by the input share of the request before caching is even discussed.
An asynchronous profile exercises both rules. At Meeting Summary (10,000 in, 1,200 out, 10% cache hit, batch-eligible) the batch rates replace list on both sides first, and the cache rate then applies to 10% of input:
| Component | Tokens | Rate /1M | Cost |
|---|---|---|---|
| Input, cache hit (10%) | 1,000 | $0.20 | $0.000200 |
| Input, cache miss (batch rate) | 9,000 | $1.00 | $0.0090 |
| Cache write premium, amortised (1 write per 5 reads) | 200 | $1.50 | $0.000300 |
| Output (batch rate) | 1,200 | $5.00 | $0.0060 |
| Cost per request (after discounts) | $0.015 |
$0.032 at list, $0.015 after discounts: 51.6% off. The monthly figure is this number multiplied by your volume and nothing else. At 100,000 requests a month, $3.2K becomes $1.6K.
Do not read 51.6% as typical. This example was chosen because it publishes both a cache and a batch rate, which only 70 of 292 priced models do. The median model on this profile saves far less, for the reason set out in the next section.
How much the after-discounts column is worth
Measured across the 292 models carrying both prices, counting only those that receive a discount for that profile. Knowledge Q&A moves furthest at 21.0%; Marketing Content moves least at 3.9%.
The ordering rewards cache hit rate rather than batch eligibility, which is worth pausing on. Batch is the larger discount, typically half off both sides. But it can only apply to the 70 models publishing a batch rate, so on an asynchronous profile the median model still gets nothing from it. A high cache hit rate pays out on every model publishing a cache rate. That is why the Meeting Summary example above saved 51.6% while the median model on the same profile saves 6.3%. One model publishes batch rates; most do not.
| Use case | Tokens in / out | Cache hit | Batch | Models discounted | Median saving |
|---|---|---|---|---|---|
| Support Ticket | 1,500 / 500 | 60% | No | 197 of 292 | 20.3% |
| Knowledge Q&A | 2,000 / 800 | 70% | No | 197 of 292 | 21.0% |
| Meeting Summary | 10,000 / 1,200 | 10% | Yes | 203 of 292 | 6.3% |
| Marketing Content | 2,500 / 1,800 | 20% | No | 197 of 292 | 3.9% |
| Coding Task | 3,000 / 2,000 | 50% | No | 197 of 292 | 10.4% |
| Invoice Processing | 1,500 / 600 | 30% | Yes | 226 of 292 | 11.2% |
| Call Summary | 2,000 / 700 | 10% | Yes | 203 of 292 | 4.4% |
| Agent Workflow | 6,000 / 3,000 | 70% | No | 197 of 292 | 18.0% |
These hit rates are modelling assumptions, not measurements of your traffic. They are the single biggest lever on the after-discounts column, so treat that column as a shape rather than a quote.
Which models can be discounted at all
Of the 292 priced models, 207 publish at least one discount rate. The other 85 publish neither and therefore have an after-discounts cost identical to list. That is a disclosure gap, not a pricing fact: a provider may well run a batch endpoint it has not published a rate card for. Where the table shows a dash, the correct reading is that we have nothing to quote.
What these numbers deliberately exclude
Every exclusion below is a decision, not an oversight. Each one would have made the figure look better and less true.
- Your negotiated rates. These are public list prices. Committed spend, marketplace agreements and enterprise terms all sit outside them, and they are usually the largest single difference between this page and your invoice.
- Your real cache behaviour. The write is charged at one per 5 reads and the hit rates assume steady traffic. At 10,000 requests a month a request arrives about every 4 minutes against a typical 5-minute cache lifetime, so entries expire between calls, writes multiply and the after-discounts figure is optimistic. Longer-lived cache tiers, sold at a higher write rate, are not modelled at all.
- Retries and failures. A request that errors after generating output is billed. Timeouts, tool-call loops and validation retries are real line items that no token profile models.
- Reasoning tokens. On models that think before answering, the reasoning trace bills as output and can dwarf the visible answer. The profiles here count the answer.
- The system prompt, past the profile. Input tokens are the profile's assumption, not your prompt. A long system prompt, a tool schema or a RAG context window can multiply the input side.
- Everything that is not tokens. Fine-tuning, embeddings, storage, egress, the compute you run around the model, and the engineering time to implement the caching this page credits you for.
- Throughput and rate limits. Two models at the same cost per request are not equivalent if one needs three times the concurrency to hold your latency target.
Using the number honestly
- Compare with it; do not budget with it. Ranking models against one another is what a common set of assumptions is good for. A committed monthly figure needs your own token counts.
- Measure your input-output split before trusting any saving. Caching cannot touch output, so an output-heavy workload has a low ceiling however good the hit rate.
- Treat an identical column pair as missing data. Check the provider's own documentation before concluding a model has no discount path.
- Re-run after any model switch. Discount coverage varies far more between providers than list prices do, so a migration can change the optimisation case even when the headline rates look similar.
For a team early in its FinOps practice the first useful step is not choosing a cheaper model, it is splitting the bill into input, output, cache reads and cache writes. Almost every decision on this page becomes obvious once that split exists, and almost none of them are answerable without it.
The FinOps reasoning behind this guide comes from OptimNow's open-source practice library: github.com/OptimNow/cloud-finops-skills.
Questions this guide answers
What is the difference between list price and after-discounts cost per request?
List price is the request's input and output tokens at the published rate card, with nothing else. After discounts swaps in the batch rate when the workload is asynchronous and a batch rate is published, bills the cached share of input at the cache-read rate, and adds the cache-write premium. On the 2026-09-30 OptimToken catalogue snapshot, 207 of the 292 priced models publish at least one discount rate. For the other 85 the two figures are identical.
Which use case saves the most after discounts?
On the 2026-09-30 OptimToken catalogue snapshot, Knowledge Q&A saves the most: a median 21.0% across the models that receive a discount on that profile. Marketing Content saves the least, at 3.9%. Cache hit rate decides the order, not batch eligibility, because batch only applies to models that publish a batch rate.
What does a dash in the after-discounts column mean?
It means the model publishes no cache-read or batch rate, so there is nothing to quote. It does not mean no saving exists: a provider may run a batch endpoint it has not published a rate for. On the 2026-09-30 OptimToken catalogue snapshot that applies to 85 of the 292 priced models.
Can I budget with the after-discounts cost per request?
Use it to compare models, not to commit a budget. It assumes modelled cache hit rates and one cache write per 5 reads, and it leaves out negotiated rates, retries, reasoning tokens and everything that is not tokens. On the 2026-09-30 OptimToken catalogue snapshot, the worked example, Claude Sonnet 5, drops from $0.032 to $0.015 per Meeting Summary request, 51.6% off. That is not typical: only 70 of the 292 priced models publish a batch rate.
Your real cost also depends on caching and batching, negotiated discounts, marketplace agreements, and the contracts you already have. OptimNow audits exactly that.
Get a FinOps reviewData snapshot: 2026-09-30 · Compare all models · Cloud compute pricing · All guides