OptimToken All guides

Quality per dollar: the models that beat their price class

A FinOps guide by OptimNow · prices from the live OptimToken catalogue

OptimToken files every model into a price class: Frontier, Mid-tier or Budget. The label comes from the published rate card and from nothing else. Quality is measured separately, by the LMArena preference score. This guide puts the two side by side and asks which models deliver quality their price does not reflect, and which charge a premium the scores cannot justify. Of the 265 models in the catalogue, 71 carry both a price and an Arena score. Everything below is computed from those 71.

A price class is not a quality tier

Frontier spans 1324 to 1507 on Arena. Budget spans 1211 to 1469. One picture shows the problem: the classes overlap across most of their range.

Arena score range by price class Arena score range by price class 71 ranked models. A tick marks each class's median. 1324 1469 Frontier (n=16) 1324 1507 1473 Mid-tier (n=34) 1335 1490 1450 Budget (n=21) 1211 1469 1390 LMArena score
Price class comes from published price alone (see the category rules on the pricing table). Arena scores are a separate input, last updated 2026-09-01.

44 of the 71 ranked models sit in the band where Frontier and Budget overlap (1324 to 1469 on Arena). Inside it, blended list prices run from $0.14 to $57.00 per 1M tokens: a 417x spread across models the benchmark cannot tell apart. On average, paying more still buys quality, since the class medians rank in the expected order. A shortlist is not made of averages. For an individual model, the class label predicts the bill and little else.

What the efficiency score measures

The pricing table's Efficiency score joins the two axes into one number. Each ranked model is placed twice, by Arena score and by blended list price, and the score is 60% of the quality rank plus 40% of the price rank. The weighting leans toward quality, so a cheap model with a weak score cannot climb on price alone. GLM 5.3 Flash tops it today at 82, on Arena 1469 and $0.20 per 1M blended. Both halves are relative, so the ranking moves as prices do.

Highest efficiency scores (top 10 of 71 ranked models) Highest efficiency scores (top 10 of 71 ranked models) GLM 5.3 Flash (Budget) 82 Gemini 3.7 Flash (Mid-tier) 78 Grok 4.20 (Mid-tier) 74 Gemini 3.6 Flash (Mid-tier) 73 GLM 5.3 (Mid-tier) 72 GLM 5.2 (Mid-tier) 70 Grok 4.20 Multi-Agent (Mid-tier) 69 Qwen3.7 Max (Mid-tier) 66 DeepSeek V4 Pro 0423 (Budget) 66 GPT-5.6 Luna (Budget) 66
Score = 60% quality percentile + 40% price percentile, ranked among the 71 models carrying an Arena score.

The substitution test

The same question in cash terms: for each ranked model, is there a cheaper model with an equal or better Arena score? Substitutes are drawn only from models the vendor still sells. A deprecated model can be on your bill; it should not be on your shortlist, and OptimToken now marks vendor-announced deprecations in the pricing table.

The answer is yes for 59 of the 71 ranked models, at a median premium of 9.3x on the blended price. Priced on a workload's own token mix, the individual premiums move; the worked examples below are costed on one such mix. One caveat carries more weight than the aggregate: DeepSeek V4 Flash 0423 ($0.14 per 1M blended) is the cheapest answer in 31 of those 59 cases. This is a single model, licensed Apache 2.0, sitting far below the price of its measured quality, not a broad menu of substitutes. Rule it out for licensing or data-residency reasons and most of that headroom goes with it.

Worked example: the three largest premiums

The three largest premiums among models still sold today, costed on the Support Ticket preset (1,500 input + 500 output tokens per request) at 100,000 requests per month. The premium column is the ratio of the two monthly costs before rounding:

Paying forCheaper equalMonthlyMonthly (alternative)Premium
Claude Opus 4.5
Arena 1469
GLM 5.3 Flash
Arena 1469
$2.0K $24 84.2x
Claude Sonnet 4.5
Arena 1455
GLM 5.3 Flash
Arena 1469
$1.2K $24 50.5x
GPT-5.2
Arena 1436
DeepSeek V4 Flash 0423
Arena 1436
$963 $20 47.8x

These figures are the cost of inference only: tokens at the public list price. They include no cache or batch discounts, no negotiated rates, and nothing around the API call. Self-hosting an open-weights model is a different cost structure entirely, GPUs instead of tokens, and the open-weights guide prices that gap.

These are substitution candidates, not recommendations. Arena measures general preference on open-ended prompts, not your workload. What a gap this size does support is an evaluation: at these prices, two days of testing on your own traffic pays for itself many times over.

Where the premium is real

The test cuts both ways. 12 of the 59 ranked models their vendor still sells have no cheaper equal anywhere in the catalogue. For those, the price buys measured quality nothing else offers for less. Claude Fable 5 sits at the top, at Arena 1507 and $38.00 per 1M blended. Its efficiency score is only 61, because the score rewards value, not capability. If an evaluation says a workload needs that level, a low efficiency score is no argument against it: the substitution does not exist. Compare the two ends of the range to see what the money buys.

Old models keep their launch price

Among the 21 models paying 20x or more, 14 were released at least 12 months before this snapshot: 67% of the group, against 38% across the whole ranked population. The pattern in this catalogue is consistent: vendors rarely reprice an old model downward. They publish a new one and leave the old rate card in place, and several of the largest premiums sit on models whose vendor has already announced retirement.

The useful habit is an inventory, not a negotiation. List which model each workload pins, and check it against the catalogue quarterly. A workload pinned to a 2-year-old model pays an old price for old quality, and neither half shows up in the invoice.

What these numbers do not say

The pricing table's FinOps Friendly badge applies a stricter version of this test; 13 of the 71 ranked models hold it today. Sort the pricing table by Efficiency for the full ranking, and check the cost columns against your own token mix before acting on any of it. Figures on this page come from the 2026-09-06 price snapshot and rebuild nightly.

The FinOps reasoning behind this guide comes from OptimNow's open-source practice library: github.com/OptimNow/cloud-finops-skills.

These are list prices.

Your real cost also depends on caching and batching, negotiated discounts, marketplace agreements, and the contracts you already have. OptimNow audits exactly that.

Get a FinOps review

Data snapshot: 2026-09-06 · Compare all models · Cloud compute pricing · All guides