Quality per dollar: the models that beat their price class
A FinOps guide by OptimNow · prices from the live OptimToken catalogue
OptimToken files every model into a price class: Frontier, Mid-tier or Budget. The label comes from the published rate card and from nothing else. Quality is measured separately, by the LMArena preference score. This guide puts the two side by side and asks which models deliver quality their price does not reflect, and which charge a premium the scores cannot justify. Of the 265 models in the catalogue, 71 carry both a price and an Arena score. Everything below is computed from those 71.
A price class is not a quality tier
Frontier spans 1324 to 1507 on Arena. Budget spans 1211 to 1469. One picture shows the problem: the classes overlap across most of their range.
44 of the 71 ranked models sit in the band where Frontier and Budget overlap (1324 to 1469 on Arena). Inside it, blended list prices run from $0.14 to $57.00 per 1M tokens: a 417x spread across models the benchmark cannot tell apart. On average, paying more still buys quality, since the class medians rank in the expected order. A shortlist is not made of averages. For an individual model, the class label predicts the bill and little else.
What the efficiency score measures
The pricing table's Efficiency score joins the two axes into one number. Each ranked model is placed twice, by Arena score and by blended list price, and the score is 60% of the quality rank plus 40% of the price rank. The weighting leans toward quality, so a cheap model with a weak score cannot climb on price alone. GLM 5.3 Flash tops it today at 82, on Arena 1469 and $0.20 per 1M blended. Both halves are relative, so the ranking moves as prices do.
The substitution test
The same question in cash terms: for each ranked model, is there a cheaper model with an equal or better Arena score? Substitutes are drawn only from models the vendor still sells. A deprecated model can be on your bill; it should not be on your shortlist, and OptimToken now marks vendor-announced deprecations in the pricing table.
The answer is yes for 59 of the 71 ranked models, at a median premium of 9.3x on the blended price. Priced on a workload's own token mix, the individual premiums move; the worked examples below are costed on one such mix. One caveat carries more weight than the aggregate: DeepSeek V4 Flash 0423 ($0.14 per 1M blended) is the cheapest answer in 31 of those 59 cases. This is a single model, licensed Apache 2.0, sitting far below the price of its measured quality, not a broad menu of substitutes. Rule it out for licensing or data-residency reasons and most of that headroom goes with it.
Worked example: the three largest premiums
The three largest premiums among models still sold today, costed on the Support Ticket preset (1,500 input + 500 output tokens per request) at 100,000 requests per month. The premium column is the ratio of the two monthly costs before rounding:
| Paying for | Cheaper equal | Monthly | Monthly (alternative) | Premium |
|---|---|---|---|---|
| Claude Opus 4.5 Arena 1469 |
GLM 5.3 Flash Arena 1469 |
$2.0K | $24 | 84.2x |
| Claude Sonnet 4.5 Arena 1455 |
GLM 5.3 Flash Arena 1469 |
$1.2K | $24 | 50.5x |
| GPT-5.2 Arena 1436 |
DeepSeek V4 Flash 0423 Arena 1436 |
$963 | $20 | 47.8x |
These figures are the cost of inference only: tokens at the public list price. They include no cache or batch discounts, no negotiated rates, and nothing around the API call. Self-hosting an open-weights model is a different cost structure entirely, GPUs instead of tokens, and the open-weights guide prices that gap.
These are substitution candidates, not recommendations. Arena measures general preference on open-ended prompts, not your workload. What a gap this size does support is an evaluation: at these prices, two days of testing on your own traffic pays for itself many times over.
Where the premium is real
The test cuts both ways. 12 of the 59 ranked models their vendor still sells have no cheaper equal anywhere in the catalogue. For those, the price buys measured quality nothing else offers for less. Claude Fable 5 sits at the top, at Arena 1507 and $38.00 per 1M blended. Its efficiency score is only 61, because the score rewards value, not capability. If an evaluation says a workload needs that level, a low efficiency score is no argument against it: the substitution does not exist. Compare the two ends of the range to see what the money buys.
Old models keep their launch price
Among the 21 models paying 20x or more, 14 were released at least 12 months before this snapshot: 67% of the group, against 38% across the whole ranked population. The pattern in this catalogue is consistent: vendors rarely reprice an old model downward. They publish a new one and leave the old rate card in place, and several of the largest premiums sit on models whose vendor has already announced retirement.
The useful habit is an inventory, not a negotiation. List which model each workload pins, and check it against the catalogue quarterly. A workload pinned to a 2-year-old model pays an old price for old quality, and neither half shows up in the invoice.
What these numbers do not say
- Arena is one benchmark, and a general one. It says nothing about latency, tool calling, structured output or languages other than English. Treat a gap under roughly 20 points as noise.
- Coverage is partial. 71 of 265 catalogued models carry an Arena score. A model missing here is unranked, not bad.
- The price axis is a 30/70 blend. A workload with a different token mix reorders it, which is why the table's cost columns run on use-case profiles instead.
- Switching costs money. Prompt rework, re-evaluation and a second vendor relationship are real. The gaps above are large enough to survive that; a 1.5x gap usually is not.
The pricing table's FinOps Friendly badge applies a stricter version of this test; 13 of the 71 ranked models hold it today. Sort the pricing table by Efficiency for the full ranking, and check the cost columns against your own token mix before acting on any of it. Figures on this page come from the 2026-09-06 price snapshot and rebuild nightly.
The FinOps reasoning behind this guide comes from OptimNow's open-source practice library: github.com/OptimNow/cloud-finops-skills.
Your real cost also depends on caching and batching, negotiated discounts, marketplace agreements, and the contracts you already have. OptimNow audits exactly that.
Get a FinOps reviewData snapshot: 2026-09-06 · Compare all models · Cloud compute pricing · All guides