LLM API pricing statistics: cheapest and median
A FinOps guide by OptimNow · prices from the live OptimToken catalogue
This page puts numbers on LLM API pricing: how many models carry a token price, which are the cheapest and the dearest, what the median model costs, and how wide each price tier is. Every figure is recomputed from the OptimToken catalogue when it refreshes, once a day. Each sentence that states one carries the date of the snapshot it describes, which is 2026-10-04 for this version.
These are statistics of the prices OptimToken quotes: one host's rate per model, in USD per 1M tokens, at list price. What that covers and what it leaves out is set out under How these figures are calculated.
Key figures on 2026-10-04
- As of 2026-10-04, the OptimToken catalogue quotes a token price for 285 LLM APIs from 42 model makers: 30 Frontier, 113 Mid-tier and 142 Budget.
- As of 2026-10-04, the cheapest of 285 token-priced LLM APIs in the OptimToken catalogue is Mistral Nemo (Mistral), quoted at $0.019 per 1M input tokens and $0.03 per 1M output tokens, which is $0.027 on the blended price (30% input, 70% output).
- As of 2026-10-04, the most expensive of 285 token-priced LLM APIs in the OptimToken catalogue is o1-pro (OpenAI), quoted at $150.00 per 1M input tokens and $600.00 per 1M output tokens, which is $465.00 on the blended price (30% input, 70% output).
- As of 2026-10-04, the median price across 285 token-priced LLM APIs in the OptimToken catalogue is $0.40 per 1M input tokens and $2.00 per 1M output tokens, and the median blended price (30% input, 70% output) is $1.47.
- As of 2026-10-04, quoted prices run from $0.003 to $150.00 per 1M input tokens and from $0.03 to $600.00 per 1M output tokens.
- As of 2026-10-04, the median blended price (30% input, 70% output) per 1M tokens is $0.90 for the 35 open-source models, $0.77 for the 78 open-weights models, $3.80 for the 125 proprietary models and $0.97 for the 47 models with no established licence.
- As of 2026-10-04, a batch rate is published for 67 of 285 token-priced models (24%) and a prompt-cache read rate for 195 (68%); 61 carry both and 84 carry neither.
How many models, and who makes them
| Price tier | Models | Share | Model makers |
|---|---|---|---|
| All tiers | 285 | 100% | 42 |
| Frontier | 30 | 11% | 5 |
| Mid-tier | 113 | 40% | 24 |
| Budget | 142 | 50% | 34 |
As of 2026-10-04, the model makers with the most token-priced models are OpenAI (49), Alibaba (41) and Google (19).
The catalogue also lists 7 image models. An image model's rate is not comparable with a text model's token price, so none enters the figures on this page.
What a million tokens costs
As of 2026-10-04, the median price per 1M tokens by price tier is: Frontier, $5.00 input and $27.50 output (30 models); Mid-tier, $1.10 input and $4.25 output (113 models); Budget, $0.15 input and $0.50 output (142 models). As of 2026-10-04, the median token-priced model is quoted 4.0x as much for output tokens as for input tokens.
| Price tier | Models | Median input | Median output | Median blended | Input range | Output range |
|---|---|---|---|---|---|---|
| All tiers | 285 | $0.40 | $2.00 | $1.47 | $0.003 to $150.00 | $0.03 to $600.00 |
| Frontier | 30 | $5.00 | $27.50 | $20.75 | $2.50 to $150.00 | $15.00 to $600.00 |
| Mid-tier | 113 | $1.10 | $4.25 | $3.35 | $0.003 to $4.35 | $2.00 to $14.00 |
| Budget | 142 | $0.15 | $0.50 | $0.40 | $0.015 to $1.00 | $0.03 to $1.95 |
The blended price (30% input, 70% output) is the mix the pricing table ranks on. A tier is a price band, assigned from the quoted price alone. It is not a quality grade: the cheapest Frontier model is the lowest price in the top band, whatever it scores. Which models beat their price band is the subject of the quality per dollar guide.
As of 2026-10-04, 119 of 285 token-priced models (42%) are quoted below $1 per 1M tokens on the blended price.
The cheapest and the dearest models
As of 2026-10-04, the 3 cheapest token-priced LLM APIs on the blended price (30% input, 70% output) are Mistral Nemo (Mistral) at $0.027, Ling 3.0 Flash VL (inclusionai) at $0.049 and Ling 3.0 Flash (inclusionai) at $0.05 per 1M tokens.
| Model | Model maker | Input, per 1M | Output, per 1M | Blended | Quoted host |
|---|---|---|---|---|---|
| Mistral Nemo | Mistral | $0.019 | $0.03 | $0.027 | DeepInfra (fp8 build) |
| Ling 3.0 Flash VL | inclusionai | $0.021 | $0.062 | $0.049 | Novita |
| Ling 3.0 Flash | inclusionai | $0.021 | $0.063 | $0.05 | Novita |
| Llama 3.1 8B Instruct | Meta | $0.05 | $0.08 | $0.071 | Groq |
| Mistral Small 3 | Mistral | $0.05 | $0.08 | $0.071 | DeepInfra (fp8 build) |
The price quoted for Mistral Nemo on 2026-10-04 is DeepInfra's rate for a quantized fp8 build, one of 6 hosts selling the model. Mistral, the model maker, charges $0.15 per 1M input tokens and $0.15 per 1M output tokens.
Each figure is the price OptimToken quotes for the model, which is one host's rate: another host can charge more or less for the same model. Each model's page lists every host selling it; the quoted hosts above are from the 2026-10-04 read. Where a model runs, and what the channel changes, is covered in a guide of its own.
As of 2026-10-04, the cheapest model in each price tier, ranked on the blended price and quoted per 1M tokens, is: Frontier, GPT-5.4 (OpenAI) at $2.50 input and $15.00 output; Mid-tier, Seed 1.6 (ByteDance) at $0.25 input and $2.00 output, level with 3 other models; Budget, Mistral Nemo (Mistral) at $0.019 input and $0.03 output.
The cheapest model of each tier, on each of the 3 prices:
| Tier | On input, per 1M | On output, per 1M | On the blended price |
|---|---|---|---|
| All tiers | DeepSeek V4.1 Flash $0.003 | Mistral Nemo $0.03 | Mistral Nemo $0.027 |
| Frontier | GPT-5.4 $2.50 | GPT-5.4 $15.00 (6 others at this price) | GPT-5.4 $11.25 |
| Mid-tier | DeepSeek V4.1 Flash $0.003 | Seed 1.6 $2.00 (7 others at this price) | Seed 1.6 $1.47 (3 others at this price) |
| Budget | DeepSeek V4 Flash 0731 $0.015 | Mistral Nemo $0.03 | Mistral Nemo $0.027 |
As of 2026-10-04, the lowest input price, $0.003 per 1M tokens, is quoted for DeepSeek V4.1 Flash (DeepSeek), whose output is quoted at $2.40 per 1M tokens: 153 of the 285 models are cheaper on the blended price.
And the dearest:
| Tier | On input, per 1M | On output, per 1M | On the blended price |
|---|---|---|---|
| All tiers | o1-pro $150.00 | o1-pro $600.00 | o1-pro $465.00 |
| Frontier | o1-pro $150.00 | o1-pro $600.00 | o1-pro $465.00 |
| Mid-tier | MiMo-V2.6-Pro-UltraSpeed $4.35 | GPT-5.2 $14.00 (3 others at this price) | GPT-5.2 $10.32 (3 others at this price) |
| Budget | Sonar $1.00 | Qwen3.6 Plus $1.95 | Morph V3 Large $1.60 |
Open source, open weights and proprietary
Openness is read from the licence, and it is a separate axis from the price tier. Open source means published weights under an OSI-approved licence with no field-of-use limit. Open weights means published weights under a licence that restricts something. Proprietary means API only. A model with no licence on record is counted on its own, as unknown, and assumed to be neither.
| Openness | Models | Share | Median input | Median output | Median blended |
|---|---|---|---|---|---|
| Open source | 35 | 12% | $0.28 | $1.25 | $0.90 |
| Open weights | 78 | 27% | $0.20 | $1.00 | $0.77 |
| Proprietary | 125 | 44% | $1.25 | $5.00 | $3.80 |
| Unknown | 47 | 16% | $0.30 | $1.20 | $0.97 |
As of 2026-10-04, the median open-weights model is quoted below the median proprietary model, which costs 5.0x as much on the blended price.
Each of those medians is taken on the price OptimToken quotes, which is one host's rate per model. On the 2026-10-04 read of host prices, among the models sold by more than one host, the dearest host charged a median 1.8x the cheapest across 24 open-source models, 1.9x the cheapest across 37 open-weights models, 1.00x the cheapest across 69 proprietary models and 1.4x the cheapest across 6 models with no established licence.
As of 2026-10-04, none of the 30 Frontier models is open source or open weights: 27 are proprietary and 3 carry no established licence.
These medians compare groups of different models, at whatever quality each group holds. The comparison at matched quality, and what a licence buys beyond the price, is in Open weights against proprietary APIs.
Batch and cache rates
As of 2026-10-04, a batch rate is published for 67 of 285 token-priced models (24%) and a prompt-cache read rate for 195 (68%); 61 carry both and 84 carry neither. A cache write rate, the price of storing a prompt in the cache before it can be read back, is published for 47 of them.
When each discount applies, and what it is worth on a real workload, is covered in Batch API economics and Prompt caching.
What moved
From 27 Sep to 4 Oct 2026: 4 price changes (3 up, 1 down) among 287 models compared. 18 other moves are flagged as a change of quoted host or an outlier listing.
Since tracking began on 2026-08-15, the price OptimToken quotes has changed at least once for 56 of the 312 models tracked: 357 changes in all, up to 2026-10-04. Of those changes, 152 are established as a change of quoted host, a test that needs both days' host prices, archived from 2026-08-30. For the other 256 models, the quoted price has not changed since the first day in the archive.
A price change, here as on the barometer, is a move in the price we quote that is not flagged as a change of quoted host or as an outlier listing. The price barometer lists every move of the latest window with its thresholds, and each completed week keeps a page of its own.
How these figures are calculated
- Source. Prices are list prices in USD per 1M tokens, read once a day from OpenRouter's public model list. Negotiated rates, committed-use discounts and provisioned capacity are not visible in them.
- One price per model. OptimToken quotes one price per model: the rate of one host, the one OpenRouter uses as its reference for the model. Many models are sold by several hosts at different rates, so these are statistics of quoted prices, not of every host's price. On the 2026-10-04 read of host prices, 136 of the 285 token-priced models with host data were sold by more than one host, and for 35 of those the dearest host charged at least twice the cheapest.
- Population. The statistics cover the 285 models in the Frontier, Mid-tier and Budget tiers that carry a price on both input and output. Image models are left out.
- Median. The middle value once the models are ranked on a price. With an even number of models it is the mean of the 2 middle values. The input, output and blended medians are computed separately, so the blended median is not the blend of the other 2.
- Blended price. One figure per model, weighted 30% input, 70% output. Your own workload has its own token mix: the cost per request methodology prices 8 of them.
- Model makers. Counted from the maker named in the catalogue, ignoring capitalisation.
- Refresh. The page is rebuilt from the catalogue once a day. A figure quoted from it is a figure of its date.
How to cite this page
The dataset behind this page is published under Creative Commons Attribution 4.0. Quote a figure with its date, because the figures change:
OptimToken by OptimNow, LLM API pricing statistics, snapshot of 2026-10-04, https://optimtoken.optimnow.io/guides/llm-pricing-statistics
The same catalogue is available as JSON at /api/llm-models and through the MCP server.
The FinOps reasoning behind this guide comes from OptimNow's open-source practice library: github.com/OptimNow/cloud-finops-skills. If you want these list prices set against what your own contracts charge, OptimNow runs that review.
Questions this guide answers
What is the cheapest LLM API?
As of 2026-10-04, the cheapest of 285 token-priced LLM APIs in the OptimToken catalogue is Mistral Nemo (Mistral), quoted at $0.019 per 1M input tokens and $0.03 per 1M output tokens, which is $0.027 on the blended price (30% input, 70% output). The price quoted for Mistral Nemo on 2026-10-04 is DeepInfra's rate for a quantized fp8 build, one of 6 hosts selling the model. Mistral, the model maker, charges $0.15 per 1M input tokens and $0.15 per 1M output tokens. As of 2026-10-04, the cheapest model in each price tier, ranked on the blended price and quoted per 1M tokens, is: Frontier, GPT-5.4 (OpenAI) at $2.50 input and $15.00 output; Mid-tier, Seed 1.6 (ByteDance) at $0.25 input and $2.00 output, level with 3 other models; Budget, Mistral Nemo (Mistral) at $0.019 input and $0.03 output.
How much do LLM APIs cost per million tokens?
As of 2026-10-04, the median price across 285 token-priced LLM APIs in the OptimToken catalogue is $0.40 per 1M input tokens and $2.00 per 1M output tokens, and the median blended price (30% input, 70% output) is $1.47. As of 2026-10-04, quoted prices run from $0.003 to $150.00 per 1M input tokens and from $0.03 to $600.00 per 1M output tokens. As of 2026-10-04, the median price per 1M tokens by price tier is: Frontier, $5.00 input and $27.50 output (30 models); Mid-tier, $1.10 input and $4.25 output (113 models); Budget, $0.15 input and $0.50 output (142 models).
Are open-weight models cheaper than proprietary ones?
As of 2026-10-04, the median open-weights model is quoted below the median proprietary model, which costs 5.0x as much on the blended price. As of 2026-10-04, the median blended price (30% input, 70% output) per 1M tokens is $0.90 for the 35 open-source models, $0.77 for the 78 open-weights models, $3.80 for the 125 proprietary models and $0.97 for the 47 models with no established licence. Each of those medians is taken on the price OptimToken quotes, which is one host's rate per model. On the 2026-10-04 read of host prices, among the models sold by more than one host, the dearest host charged a median 1.8x the cheapest across 24 open-source models, 1.9x the cheapest across 37 open-weights models, 1.00x the cheapest across 69 proprietary models and 1.4x the cheapest across 6 models with no established licence.
What is the most expensive LLM API?
As of 2026-10-04, the most expensive of 285 token-priced LLM APIs in the OptimToken catalogue is o1-pro (OpenAI), quoted at $150.00 per 1M input tokens and $600.00 per 1M output tokens, which is $465.00 on the blended price (30% input, 70% output).
How many LLM APIs publish batch and prompt-cache prices?
As of 2026-10-04, a batch rate is published for 67 of 285 token-priced models (24%) and a prompt-cache read rate for 195 (68%); 61 carry both and 84 carry neither. A cache write rate, the price of storing a prompt in the cache before it can be read back, is published for 47 of them.
How often do LLM API prices change?
Since tracking began on 2026-08-15, the price OptimToken quotes has changed at least once for 56 of the 312 models tracked: 357 changes in all, up to 2026-10-04. Of those changes, 152 are established as a change of quoted host, a test that needs both days' host prices, archived from 2026-08-30. From 27 Sep to 4 Oct 2026: 4 price changes (3 up, 1 down) among 287 models compared. 18 other moves are flagged as a change of quoted host or an outlier listing.
Your real cost also depends on caching and batching, negotiated discounts, marketplace agreements, and the contracts you already have. OptimNow audits exactly that.
Get a FinOps reviewData snapshot: 2026-10-04 · Compare all models · Cloud compute pricing · All guides