OptimToken All guides

How we compute cost per request

A specification, not an estimate. These are the formulas behind every cost-per-request figure on this site.

Every model row carries two cost figures: Cost / 1K req (list price) and Cost / 1K req (after discounts). The table prints them per 1,000 requests so a column of them can be compared without counting zeros. The formulas and worked examples below are per single request, so multiply by 1,000 to land on the table's figure. They answer different questions, and the gap between them is the only part of the number you control. This page gives both formulas in full and the assumptions behind them. More usefully, it lists the things they leave out on purpose, because that is where a budget built on them goes wrong.

The list-price formula

Cost per request at the published rate card, with no optimisation assumed:

cost = (inputTokens / 1,000,000) x inputPrice
     + (outputTokens / 1,000,000) x outputPrice

That is the whole of it. The token counts come from the use-case profile you pick in the scenario bar; the prices come from the catalogue. Nothing else enters.

Worked, on Claude Sonnet 5 (Anthropic, $2.00/M in, $10.00/M out) at the Support Ticket profile:

ComponentTokensRate /1MCost
Input1,500$2.00$0.0030
Output500$10.00$0.0050
Cost per request (list price)$0.0080

The after-discounts formula

Same request, priced as if you had implemented the two optimisations the provider publishes a rate for. Four rules, applied in order:

  1. Batch substitution. If the use case is asynchronous and the model publishes both batch rates, the batch input and output rates replace the list rates outright (where that discount is real).
  2. Cache read on the hit-rate share. The profile's cache hit rate is the fraction of input tokens billed at the published cache-read rate; the remainder bills at whatever rate rule 1 left in place (what that discount is worth).
  3. Cache write, amortised. A token can only be read from the cache after it has been written there, and providers that publish a write rate charge more for it than for ordinary input. The formula assumes one token written for every 5 read. A written token is one of the cache misses, already billed at the input rate by rule 2. What is added is the premium: the published write rate minus that input rate, never below zero. Under batch the write rate is taken as published, because no provider publishes a batch rate for it. A model that publishes a read rate but no write rate is assumed to write at the ordinary input rate: premium zero.
  4. Fall back, never invent. A model that publishes no batch rate keeps list prices for rule 1. A model that publishes no cache-read rate has its cached share billed at the ordinary input rate, which makes the discount zero and leaves no write to charge. And if the write premium outweighs the read saving, the input side is billed uncached, since caching that does not pay is caching you would not switch on: the figure is never above list.
writePremium   = max(0, cacheWriteRate - inputRate)   (0 when no write rate is published)
writtenShare   = min(hitRate / 5, 1 - hitRate)

effectiveInput = cacheReadRate x hitRate
               + inputRate x (1 - hitRate)
               + writePremium x writtenShare
effectiveInput = min(effectiveInput, inputRate)

cost = (inputTokens / 1,000,000) x effectiveInput
     + (outputTokens / 1,000,000) x outputRate

The 5 is a modelling assumption, not a published figure: how often a cached prefix is re-read before it must be rewritten is a property of your traffic. In a conversation, each turn re-reads the earlier turns and writes its own, which works out to (n - 1) / 2 reads per write over n turns, so 5 corresponds to a 11-turn session. A prompt shared by all requests under steady traffic does far better, and the same prompt at low volume does worse, because entries expire between requests. 45 of 292 priced models publish a write rate.

Rule 4 is the one that matters when reading the table. The after-discounts column is never lower than what the provider will honour, so where a model publishes nothing, the two figures are identical and the table prints a dash in place of the second one. A dash means "nothing published", not "no savings available".

The same request, now at the Support Ticket profile's 60% cache hit rate (this profile is synchronous, so no batch rate applies):

ComponentTokensRate /1MCost
Input, cache hit (60%)900$0.20$0.000180
Input, cache miss600$2.00$0.0012
Cache write premium, amortised (1 write per 5 reads)180$0.50$0.000090
Output500$10.00$0.0050
Cost per request (after discounts)$0.0065

Claude Sonnet 5 publishes a cache-write rate of $2.50/M, 1.25x its input rate, so the written tokens cost the difference on top of the input rate their cache miss already paid.

A 19.1% reduction, entirely from prompt caching. Note where it comes from: output is untouched, so the saving is bounded by the input share of the request before caching is even discussed.

An asynchronous profile exercises both rules. At Meeting Summary (10,000 in, 1,200 out, 10% cache hit, batch-eligible) the batch rates replace list on both sides first, and the cache rate then applies to 10% of input:

ComponentTokensRate /1MCost
Input, cache hit (10%)1,000$0.20$0.000200
Input, cache miss (batch rate)9,000$1.00$0.0090
Cache write premium, amortised (1 write per 5 reads)200$1.50$0.000300
Output (batch rate)1,200$5.00$0.0060
Cost per request (after discounts)$0.015

$0.032 at list, $0.015 after discounts: 51.6% off. The monthly figure is this number multiplied by your volume and nothing else. At 100,000 requests a month, $3.2K becomes $1.6K.

Do not read 51.6% as typical. This example was chosen because it publishes both a cache and a batch rate, which only 70 of 292 priced models do. The median model on this profile saves far less, for the reason set out in the next section.

How much the after-discounts column is worth

Median discount the after-discounts column applies, by use case Median discount the after-discounts column applies, by use case Knowledge Q&A 21.0% Support Ticket 20.3% Agent Workflow 18.0% Invoice Processing 11.2% Coding Task 10.4% Meeting Summary 6.3% Call Summary 4.4% Marketing Content 3.9%

Measured across the 292 models carrying both prices, counting only those that receive a discount for that profile. Knowledge Q&A moves furthest at 21.0%; Marketing Content moves least at 3.9%.

The ordering rewards cache hit rate rather than batch eligibility, which is worth pausing on. Batch is the larger discount, typically half off both sides. But it can only apply to the 70 models publishing a batch rate, so on an asynchronous profile the median model still gets nothing from it. A high cache hit rate pays out on every model publishing a cache rate. That is why the Meeting Summary example above saved 51.6% while the median model on the same profile saves 6.3%. One model publishes batch rates; most do not.

Use caseTokens in / outCache hitBatchModels discountedMedian saving
Support Ticket1,500 / 50060%No197 of 29220.3%
Knowledge Q&A2,000 / 80070%No197 of 29221.0%
Meeting Summary10,000 / 1,20010%Yes203 of 2926.3%
Marketing Content2,500 / 1,80020%No197 of 2923.9%
Coding Task3,000 / 2,00050%No197 of 29210.4%
Invoice Processing1,500 / 60030%Yes226 of 29211.2%
Call Summary2,000 / 70010%Yes203 of 2924.4%
Agent Workflow6,000 / 3,00070%No197 of 29218.0%

These hit rates are modelling assumptions, not measurements of your traffic. They are the single biggest lever on the after-discounts column, so treat that column as a shape rather than a quote.

Which models can be discounted at all

Published discount rates across 292 priced models Published discount rates across 292 priced models Cache and batch rates 64 Cache read only 137 Batch only 6 Neither published 85

Of the 292 priced models, 207 publish at least one discount rate. The other 85 publish neither and therefore have an after-discounts cost identical to list. That is a disclosure gap, not a pricing fact: a provider may well run a batch endpoint it has not published a rate card for. Where the table shows a dash, the correct reading is that we have nothing to quote.

What these numbers deliberately exclude

Every exclusion below is a decision, not an oversight. Each one would have made the figure look better and less true.

Using the number honestly

For a team early in its FinOps practice the first useful step is not choosing a cheaper model, it is splitting the bill into input, output, cache reads and cache writes. Almost every decision on this page becomes obvious once that split exists, and almost none of them are answerable without it.

The FinOps reasoning behind this guide comes from OptimNow's open-source practice library: github.com/OptimNow/cloud-finops-skills.

Questions this guide answers

What is the difference between list price and after-discounts cost per request?

List price is the request's input and output tokens at the published rate card, with nothing else. After discounts swaps in the batch rate when the workload is asynchronous and a batch rate is published, bills the cached share of input at the cache-read rate, and adds the cache-write premium. On the 2026-09-30 OptimToken catalogue snapshot, 207 of the 292 priced models publish at least one discount rate. For the other 85 the two figures are identical.

Which use case saves the most after discounts?

On the 2026-09-30 OptimToken catalogue snapshot, Knowledge Q&A saves the most: a median 21.0% across the models that receive a discount on that profile. Marketing Content saves the least, at 3.9%. Cache hit rate decides the order, not batch eligibility, because batch only applies to models that publish a batch rate.

What does a dash in the after-discounts column mean?

It means the model publishes no cache-read or batch rate, so there is nothing to quote. It does not mean no saving exists: a provider may run a batch endpoint it has not published a rate for. On the 2026-09-30 OptimToken catalogue snapshot that applies to 85 of the 292 priced models.

Can I budget with the after-discounts cost per request?

Use it to compare models, not to commit a budget. It assumes modelled cache hit rates and one cache write per 5 reads, and it leaves out negotiated rates, retries, reasoning tokens and everything that is not tokens. On the 2026-09-30 OptimToken catalogue snapshot, the worked example, Claude Sonnet 5, drops from $0.032 to $0.015 per Meeting Summary request, 51.6% off. That is not typical: only 70 of the 292 priced models publish a batch rate.

These are list prices.

Your real cost also depends on caching and batching, negotiated discounts, marketplace agreements, and the contracts you already have. OptimNow audits exactly that.

Get a FinOps review

Data snapshot: 2026-09-30 · Compare all models · Cloud compute pricing · All guides