OptimToken All guides

LLM API pricing statistics: cheapest and median

A FinOps guide by OptimNow · prices from the live OptimToken catalogue

This page puts numbers on LLM API pricing: how many models carry a token price, which are the cheapest and the dearest, what the median model costs, and how wide each price tier is. Every figure is recomputed from the OptimToken catalogue when it refreshes, once a day. Each sentence that states one carries the date of the snapshot it describes, which is 2026-10-04 for this version.

These are statistics of the prices OptimToken quotes: one host's rate per model, in USD per 1M tokens, at list price. What that covers and what it leaves out is set out under How these figures are calculated.

Key figures on 2026-10-04

How many models, and who makes them

Price tierModelsShareModel makers
All tiers285100%42
Frontier3011%5
Mid-tier11340%24
Budget14250%34

As of 2026-10-04, the model makers with the most token-priced models are OpenAI (49), Alibaba (41) and Google (19).

The catalogue also lists 7 image models. An image model's rate is not comparable with a text model's token price, so none enters the figures on this page.

What a million tokens costs

As of 2026-10-04, the median price per 1M tokens by price tier is: Frontier, $5.00 input and $27.50 output (30 models); Mid-tier, $1.10 input and $4.25 output (113 models); Budget, $0.15 input and $0.50 output (142 models). As of 2026-10-04, the median token-priced model is quoted 4.0x as much for output tokens as for input tokens.

Price tierModelsMedian inputMedian outputMedian blendedInput rangeOutput range
All tiers285$0.40$2.00$1.47$0.003 to $150.00$0.03 to $600.00
Frontier30$5.00$27.50$20.75$2.50 to $150.00$15.00 to $600.00
Mid-tier113$1.10$4.25$3.35$0.003 to $4.35$2.00 to $14.00
Budget142$0.15$0.50$0.40$0.015 to $1.00$0.03 to $1.95

The blended price (30% input, 70% output) is the mix the pricing table ranks on. A tier is a price band, assigned from the quoted price alone. It is not a quality grade: the cheapest Frontier model is the lowest price in the top band, whatever it scores. Which models beat their price band is the subject of the quality per dollar guide.

Token-priced models by blended price per 1M tokens (n=285) Token-priced models by blended price per 1M tokens (n=285) $0.01 to $0.10 9 (3%) $0.10 to $1 110 (39%) $1 to $10 132 (46%) $10 to $100 30 (11%) $100 to $1,000 4 (1%)
Snapshot of 2026-10-04. One band per power of ten of the blended price (30% input, 70% output); a band includes its lower bound. The highlighted band holds the median model.

As of 2026-10-04, 119 of 285 token-priced models (42%) are quoted below $1 per 1M tokens on the blended price.

The cheapest and the dearest models

As of 2026-10-04, the 3 cheapest token-priced LLM APIs on the blended price (30% input, 70% output) are Mistral Nemo (Mistral) at $0.027, Ling 3.0 Flash VL (inclusionai) at $0.049 and Ling 3.0 Flash (inclusionai) at $0.05 per 1M tokens.

ModelModel makerInput, per 1MOutput, per 1MBlendedQuoted host
Mistral NemoMistral$0.019$0.03$0.027DeepInfra (fp8 build)
Ling 3.0 Flash VLinclusionai$0.021$0.062$0.049Novita
Ling 3.0 Flashinclusionai$0.021$0.063$0.05Novita
Llama 3.1 8B InstructMeta$0.05$0.08$0.071Groq
Mistral Small 3Mistral$0.05$0.08$0.071DeepInfra (fp8 build)

The price quoted for Mistral Nemo on 2026-10-04 is DeepInfra's rate for a quantized fp8 build, one of 6 hosts selling the model. Mistral, the model maker, charges $0.15 per 1M input tokens and $0.15 per 1M output tokens.

Each figure is the price OptimToken quotes for the model, which is one host's rate: another host can charge more or less for the same model. Each model's page lists every host selling it; the quoted hosts above are from the 2026-10-04 read. Where a model runs, and what the channel changes, is covered in a guide of its own.

As of 2026-10-04, the cheapest model in each price tier, ranked on the blended price and quoted per 1M tokens, is: Frontier, GPT-5.4 (OpenAI) at $2.50 input and $15.00 output; Mid-tier, Seed 1.6 (ByteDance) at $0.25 input and $2.00 output, level with 3 other models; Budget, Mistral Nemo (Mistral) at $0.019 input and $0.03 output.

The cheapest model of each tier, on each of the 3 prices:

TierOn input, per 1MOn output, per 1MOn the blended price
All tiersDeepSeek V4.1 Flash
$0.003
Mistral Nemo
$0.03
Mistral Nemo
$0.027
FrontierGPT-5.4
$2.50
GPT-5.4
$15.00 (6 others at this price)
GPT-5.4
$11.25
Mid-tierDeepSeek V4.1 Flash
$0.003
Seed 1.6
$2.00 (7 others at this price)
Seed 1.6
$1.47 (3 others at this price)
BudgetDeepSeek V4 Flash 0731
$0.015
Mistral Nemo
$0.03
Mistral Nemo
$0.027

As of 2026-10-04, the lowest input price, $0.003 per 1M tokens, is quoted for DeepSeek V4.1 Flash (DeepSeek), whose output is quoted at $2.40 per 1M tokens: 153 of the 285 models are cheaper on the blended price.

And the dearest:

TierOn input, per 1MOn output, per 1MOn the blended price
All tierso1-pro
$150.00
o1-pro
$600.00
o1-pro
$465.00
Frontiero1-pro
$150.00
o1-pro
$600.00
o1-pro
$465.00
Mid-tierMiMo-V2.6-Pro-UltraSpeed
$4.35
GPT-5.2
$14.00 (3 others at this price)
GPT-5.2
$10.32 (3 others at this price)
BudgetSonar
$1.00
Qwen3.6 Plus
$1.95
Morph V3 Large
$1.60

Open source, open weights and proprietary

Openness is read from the licence, and it is a separate axis from the price tier. Open source means published weights under an OSI-approved licence with no field-of-use limit. Open weights means published weights under a licence that restricts something. Proprietary means API only. A model with no licence on record is counted on its own, as unknown, and assumed to be neither.

OpennessModelsShareMedian inputMedian outputMedian blended
Open source3512%$0.28$1.25$0.90
Open weights7827%$0.20$1.00$0.77
Proprietary12544%$1.25$5.00$3.80
Unknown4716%$0.30$1.20$0.97

As of 2026-10-04, the median open-weights model is quoted below the median proprietary model, which costs 5.0x as much on the blended price.

Each of those medians is taken on the price OptimToken quotes, which is one host's rate per model. On the 2026-10-04 read of host prices, among the models sold by more than one host, the dearest host charged a median 1.8x the cheapest across 24 open-source models, 1.9x the cheapest across 37 open-weights models, 1.00x the cheapest across 69 proprietary models and 1.4x the cheapest across 6 models with no established licence.

Token-priced models by openness and price tier (n=285) Token-priced models by openness and price tier (n=285) Snapshot of 2026-10-04. An outlined cell is an empty one. Frontier Mid-tier Budget Open source 1223 Open weights 2949 Proprietary 276137 Unknown 31133
Rows are openness groups, columns are price tiers. The accent marks a group with no model in the Frontier tier.

As of 2026-10-04, none of the 30 Frontier models is open source or open weights: 27 are proprietary and 3 carry no established licence.

These medians compare groups of different models, at whatever quality each group holds. The comparison at matched quality, and what a licence buys beyond the price, is in Open weights against proprietary APIs.

Batch and cache rates

As of 2026-10-04, a batch rate is published for 67 of 285 token-priced models (24%) and a prompt-cache read rate for 195 (68%); 61 carry both and 84 carry neither. A cache write rate, the price of storing a prompt in the cache before it can be read back, is published for 47 of them.

When each discount applies, and what it is worth on a real workload, is covered in Batch API economics and Prompt caching.

What moved

From 27 Sep to 4 Oct 2026: 4 price changes (3 up, 1 down) among 287 models compared. 18 other moves are flagged as a change of quoted host or an outlier listing.

Since tracking began on 2026-08-15, the price OptimToken quotes has changed at least once for 56 of the 312 models tracked: 357 changes in all, up to 2026-10-04. Of those changes, 152 are established as a change of quoted host, a test that needs both days' host prices, archived from 2026-08-30. For the other 256 models, the quoted price has not changed since the first day in the archive.

A price change, here as on the barometer, is a move in the price we quote that is not flagged as a change of quoted host or as an outlier listing. The price barometer lists every move of the latest window with its thresholds, and each completed week keeps a page of its own.

How these figures are calculated

How to cite this page

The dataset behind this page is published under Creative Commons Attribution 4.0. Quote a figure with its date, because the figures change:

OptimToken by OptimNow, LLM API pricing statistics, snapshot of 2026-10-04, https://optimtoken.optimnow.io/guides/llm-pricing-statistics

The same catalogue is available as JSON at /api/llm-models and through the MCP server.

The FinOps reasoning behind this guide comes from OptimNow's open-source practice library: github.com/OptimNow/cloud-finops-skills. If you want these list prices set against what your own contracts charge, OptimNow runs that review.

Questions this guide answers

What is the cheapest LLM API?

As of 2026-10-04, the cheapest of 285 token-priced LLM APIs in the OptimToken catalogue is Mistral Nemo (Mistral), quoted at $0.019 per 1M input tokens and $0.03 per 1M output tokens, which is $0.027 on the blended price (30% input, 70% output). The price quoted for Mistral Nemo on 2026-10-04 is DeepInfra's rate for a quantized fp8 build, one of 6 hosts selling the model. Mistral, the model maker, charges $0.15 per 1M input tokens and $0.15 per 1M output tokens. As of 2026-10-04, the cheapest model in each price tier, ranked on the blended price and quoted per 1M tokens, is: Frontier, GPT-5.4 (OpenAI) at $2.50 input and $15.00 output; Mid-tier, Seed 1.6 (ByteDance) at $0.25 input and $2.00 output, level with 3 other models; Budget, Mistral Nemo (Mistral) at $0.019 input and $0.03 output.

How much do LLM APIs cost per million tokens?

As of 2026-10-04, the median price across 285 token-priced LLM APIs in the OptimToken catalogue is $0.40 per 1M input tokens and $2.00 per 1M output tokens, and the median blended price (30% input, 70% output) is $1.47. As of 2026-10-04, quoted prices run from $0.003 to $150.00 per 1M input tokens and from $0.03 to $600.00 per 1M output tokens. As of 2026-10-04, the median price per 1M tokens by price tier is: Frontier, $5.00 input and $27.50 output (30 models); Mid-tier, $1.10 input and $4.25 output (113 models); Budget, $0.15 input and $0.50 output (142 models).

Are open-weight models cheaper than proprietary ones?

As of 2026-10-04, the median open-weights model is quoted below the median proprietary model, which costs 5.0x as much on the blended price. As of 2026-10-04, the median blended price (30% input, 70% output) per 1M tokens is $0.90 for the 35 open-source models, $0.77 for the 78 open-weights models, $3.80 for the 125 proprietary models and $0.97 for the 47 models with no established licence. Each of those medians is taken on the price OptimToken quotes, which is one host's rate per model. On the 2026-10-04 read of host prices, among the models sold by more than one host, the dearest host charged a median 1.8x the cheapest across 24 open-source models, 1.9x the cheapest across 37 open-weights models, 1.00x the cheapest across 69 proprietary models and 1.4x the cheapest across 6 models with no established licence.

What is the most expensive LLM API?

As of 2026-10-04, the most expensive of 285 token-priced LLM APIs in the OptimToken catalogue is o1-pro (OpenAI), quoted at $150.00 per 1M input tokens and $600.00 per 1M output tokens, which is $465.00 on the blended price (30% input, 70% output).

How many LLM APIs publish batch and prompt-cache prices?

As of 2026-10-04, a batch rate is published for 67 of 285 token-priced models (24%) and a prompt-cache read rate for 195 (68%); 61 carry both and 84 carry neither. A cache write rate, the price of storing a prompt in the cache before it can be read back, is published for 47 of them.

How often do LLM API prices change?

Since tracking began on 2026-08-15, the price OptimToken quotes has changed at least once for 56 of the 312 models tracked: 357 changes in all, up to 2026-10-04. Of those changes, 152 are established as a change of quoted host, a test that needs both days' host prices, archived from 2026-08-30. From 27 Sep to 4 Oct 2026: 4 price changes (3 up, 1 down) among 287 models compared. 18 other moves are flagged as a change of quoted host or an outlier listing.

These are list prices.

Your real cost also depends on caching and batching, negotiated discounts, marketplace agreements, and the contracts you already have. OptimNow audits exactly that.

Get a FinOps review

Data snapshot: 2026-10-04 · Compare all models · Cloud compute pricing · All guides