OptimToken All guides

Batch API economics: when the -50% applies

A FinOps guide by OptimNow · prices from the live OptimToken catalogue

Most large providers now sell a batch tier: submit a file of requests, get the results back asynchronously, pay roughly half the token price. Prompt caching cuts a larger percentage, but only off input tokens. Batch is the discount that applies to output as well, which is why it moves a bill that generation dominates. What it costs instead is latency, and it is not offered on every model. This guide shows where the −50% is real, with numbers from current prices.

What batch pricing is

A batch API is not a faster way to send many requests: it is a slower one. You upload a job (typically a JSONL file of requests), the provider schedules it on spare capacity, and commits to returning results within a window, usually 24 hours. In exchange, both input and output tokens are billed at a discount, typically 50%. The provider wins because batch jobs smooth out their GPU utilization; you win if, and only if, nobody is waiting on the answer.

Which workloads qualify

The test is simple: if nobody needs the result before tomorrow, batch it. That covers a lot of real enterprise volume:

What never qualifies: anything interactive. Support chat, coding assistants, agent loops that chain on each other's output: a 24-hour ceiling is meaningless there. Of the 8 use-case presets on the pricing table, 3 are marked batch-eligible for this reason: meeting summary, invoice processing and call summary.

Who publishes batch prices

Here the −50% story gets narrower. Today, 67 of 288 models on OptimToken publish batch rates. A batch tier is not a property of the model, it is a property of the vendor's billing: it clusters in a handful of catalogues and is absent from the rest.

Models publishing a batch rate, by provider Models publishing a batch rate, by provider OpenAI 34 / 50 Alibaba 0 / 41 Google 11 / 23 Zhipu 2 / 16 Mistral 6 / 16 DeepSeek 0 / 15 Anthropic 12 / 13 MiniMax 0 / 8 Moonshot 1 / 7 xAI 1 / 7
The 10 largest providers by catalogue size. The pale bar is every model the provider lists, the dark bar is the subset publishing a batch rate.

OpenAI, Google, Mistral and Anthropic price at least 30% of their own catalogue at a batch rate. Zhipu (2 of 16), Moonshot (1 of 7) and xAI (1 of 7) publish one or two isolated rows. That is a model decision rather than a billing policy, and should not be read as batch support. Alibaba, DeepSeek and MiniMax publish none at all, across 64 catalogued models between them. For those vendors the strategy is a low list price rather than a discount tier, which is a reasonable answer to the same problem and a very different one to plan around. If your shortlist sits entirely in that last group, batch is not a lever you have.

When a batch rate exists, it is almost always half

The rate itself turns out to be the predictable part. Of the 67 models publishing a full batch tier, 63 price it at exactly 50% off a 30/70 blend of input and output. The median is 50%. Unlike prompt-cache rates, which spread across the whole range, batch pricing has converged on one number.

Batch discount against list price (n=67 publishing a rate) Batch discount against list price (n=67 publishing a rate) Exactly 50% off 63 Partial discount 2 No discount 0 Batch priced above list 0
Discount measured on a 30/70 input-output blend, the same weighting the pricing table ranks on. Counts the 67 models that publish a batch rate; the 221 that publish none are the subject of the chart above.

The exceptions are worth knowing before you write a migration off the headline. The check is one subtraction against the list price, and it belongs in the migration plan rather than after it.

Worked example: invoice processing at 100K requests/month

Using the Invoice Processing preset (1,500 input + 600 output tokens per request, an async workload), at 100,000 requests per month:

ModelCost/req (list)Cost/req (batch)Monthly (list)Monthly (batch)Batch saves
Claude Sonnet 5 $0.0090 $0.0045 $900 $450 $450 (−50%)
GPT-5.5 $0.025 $0.013 $2.5K $1.3K $1.3K (−50%)
DeepSeek V4 Flash 0731 $0.000223 not published $22 $22 $0 (no batch rate)

Two lessons. First, when batch rates exist, the saving is real money at volume: no code changes to the prompts, just a different submission path. Second, a cheap model at list price can still undercut an expensive model at batch price. DeepSeek V4 Flash at list ($0.02/M in, $0.32/M out) costs a fraction of Claude Sonnet 5's batch rate for this extraction-style workload. Batch discounts are an optimization within a model choice, not a substitute for making the model choice well. Compare the two side by side before assuming the discount settles it.

The caveats that bite in production

When not to bother

Skip batch when the workload is interactive, or when a deadline sits closer than the SLA window. Skip it too when the monthly spend is small. An engineer-day of pipeline work does not pay back on a $40/month job, and at that scale a budget model at list price is usually the better lever. And never batch-migrate a workload before checking whether a cheaper model would beat the discount outright. The pricing table sorts by optimized cost per use case for exactly that question.

Figures on this page come from the 2026-09-28 price snapshot and rebuild nightly.

The FinOps reasoning behind this guide comes from OptimNow's open-source practice library: github.com/OptimNow/cloud-finops-skills. If you want your own async workloads sized against the batch and cache rates you are entitled to, OptimNow runs that review.

Questions this guide answers

Does the batch API always give 50% off?

Only where a batch rate is published, and then almost always. On the 2026-09-28 OptimToken catalogue snapshot, 63 of the 67 models with a full batch tier price it at exactly 50% off a 30/70 blend of input and output. The median discount is 50%.

How many LLMs have a batch API price?

On the 2026-09-28 OptimToken catalogue snapshot, 67 of the 288 models publish both a batch input and a batch output rate. Batch pricing clusters by vendor: OpenAI, Google, Mistral and Anthropic price at least 30% of their own catalogue at a batch rate, while Alibaba, DeepSeek and MiniMax publish none.

Which workloads can use batch pricing?

Any workload where nobody waits on the answer: invoice and document processing, call and meeting summaries, evaluations and backfills. Interactive work such as support chat, coding assistants and agent loops never qualifies. On the 2026-09-28 OptimToken catalogue snapshot, 3 of the 8 use-case presets on the pricing table are marked batch-eligible: meeting summary, invoice processing and call summary.

Is a batch discount worth more than switching to a cheaper model?

Often not. On the 2026-09-28 OptimToken catalogue snapshot, an invoice-processing request costs $0.0045 on Claude Sonnet 5 at its batch rate and $0.000223 on DeepSeek V4 Flash 0731 at list price. The list-price model is 20.1x cheaper with no discount applied. Batch is an optimisation within a model choice, not a substitute for making the model choice well.

These are list prices.

Your real cost also depends on caching and batching, negotiated discounts, marketplace agreements, and the contracts you already have. OptimNow audits exactly that.

Get a FinOps review

Data snapshot: 2026-09-28 · Compare all models · Cloud compute pricing · All guides