Open weights against proprietary APIs: the cost gap in numbers
A FinOps guide by OptimNow · prices from the live OptimToken catalogue
The argument for open-licensed models usually arrives as a price argument: publish the weights, let anyone serve them, watch the rate drop. The numbers support that, up to a point, and the point is sharper than the argument suggests. This guide measures the gap three ways: across every model, at matched quality, and at the ceiling where the open models run out.
Openness is not a price tier
Two things about a model are easy to run together. One is what it costs: Frontier, Mid-tier or Budget, read from the rate card. The other is whose hardware can run it, read from the licence. Open source means OSI-approved terms with no field-of-use limit. Open weights means published weights under a licence that restricts something. Proprietary means API only, and Unknown means no licence is recorded. Of 265 priced models today, 32 are open source, 76 are open weights, 121 are proprietary and 36 carry no established licence. The unknown group stays unknown rather than guessed.
Read straight off, the proprietary median of $3.41 sits 4.3x above the open median of $0.80. That is the number most often quoted, and on its own it proves little. The two groups are not the same models at different prices; they are different models. Proprietary catalogues carry the frontier tier, plus every superseded flagship still on sale behind it. The open groups sit almost entirely in the middle of the market.
The quality-matched comparison, and the ceiling
Arena scores give a common yardstick, on 71 of 265 priced models (27 open, 44 proprietary), scored as of 2026-09-01. Splitting the ranked models into score bands shows why the median comparison misleads. It also shows the finding that matters more than any price ratio on this page.
The best-ranked open model is GLM 5.3 at 1484 (Open source, Apache 2.0). Above that score sit 7 proprietary models and no open ones: Claude Fable 5 (1507), Claude Opus 4.6 (1497), Claude Opus 4.7 (1494), Claude Opus 5 (1492), among others. For a workload that needs quality above 1484 on this benchmark, openness is not a lever. No price buys an open model there, because none exists.
Below the ceiling, the gap is large
Inside the overlap the comparison becomes fair, and it is worth running per model rather than per group. For each ranked proprietary model, take every open model with an equal or better Arena score, then the cheapest of those. 37 of 44 proprietary models have such a match. 34 of those 37 matches are cheaper, at a median of 24.9x.
One caveat belongs next to that figure. DeepSeek V4 Flash 0423 is the cheapest qualifying substitute in 18 of the 37 matches, and 2 providers account for every substitute in the set (Zhipu, DeepSeek). The gap sits in a small number of aggressively priced models, not across the open catalogue. Read it as "a handful of open models are priced far below their measured rank", not as "open licensing is cheaper". It is also fragile: those rates are set by whoever serves the weights, and they move.
How fragile is measurable. Each price above is one seller's rate, and most open models are sold by several sellers at once. Repeat the matched test pricing every open model at its cheapest seller and the median gap is 24.9x. Price them at their dearest and it is 7.5x, on the same models and the same day. The proprietary side barely moves under the same treatment, because there is usually only one place to buy it. So treat the 24.9x as the middle of a range you land in by choosing a seller, not as a rate card. Which seller a price belongs to, and what else the choice changes, is a guide of its own.
What crossing the ceiling costs
At the top of the open range the rule reverses. One proprietary model is both better ranked and cheaper than the best open model in the catalogue: Gemini 3.7 Flash (1490, $2.85 blended), against GLM 5.3 at 1484 and $3.50. A buyer picking the strongest open model available is paying more for less measured quality than a mid-tier proprietary model on the same page.
What that costs at volume
The Knowledge Q&A preset (2,000 input + 800 output tokens per request), at 1,000,000 requests per month:
| Model | Openness | ELO | List in / out per 1M | Cost/req | Monthly at 1M req |
|---|---|---|---|---|---|
| GLM 5.3 | Open source | 1484 | $1.40 / $4.40 | $0.0063 | $6.3K |
| Gemini 3.7 Flash | Proprietary | 1490 | $0.75 / $3.75 | $0.0045 | $4.5K |
| Claude Opus 4 | Proprietary | 1414 | $15.00 / $75.00 | $0.090 | $90.0K |
| DeepSeek V4 Flash 0423 | Open source | 1436 | $0.08 / $0.16 | $0.000290 | $290 |
The first two rows are the ceiling trade: GLM 5.3 is the best open model on offer, Gemini 3.7 Flash the cheapest thing that outranks it. Compare them side by side before the licence decides it. The last two are the widest matched gap in the catalogue, Claude Opus 4 against DeepSeek V4 Flash 0423 at 417x. 8 of the 10 widest gaps sit on proprietary models released at least 12 months before this snapshot. Read a multiple that size as a signal about a model your workloads have outgrown, not as the going rate for proprietary inference.
The licence is not the price, it is the exit
Every price on this page is a hosted API rate, open models included. Publishing weights does not make inference free. It makes the serving market competitive, and someone still runs the GPUs. What the licence buys is optionality: the right to move the same model onto your own hardware if the rate moves, the vendor withdraws it, or data residency rules out the API.
"Competitive" is measurable, and it is the sharpest difference between the two groups on this page. Across the models sold by more than one seller, an open model's dearest seller charges a median 1.7x its cheapest (n=61). For proprietary models the same figure is 1.00x (n=64): you pay what the vendor charges, wherever you buy it. That is the licence showing up in the price before anyone self-hosts anything. One side of this comparison is a negotiation and the other is a rate card. It also cuts the other way: 38 of the open models priced here have no first-party API at all. Every rate quoted for them is a reseller's, and there is no vendor list price to fall back on if that reseller withdraws.
That right is not uniform, which is why open source and open weights are kept apart. 10 of the 27 ranked open models carry OSI-approved terms (Apache 2.0 or MIT); 17 carry a licence that restricts something: an acceptable-use clause, a monthly-active-user threshold, or a non-commercial term. Both are self-hostable. Only one is unconditionally deployable in a commercial product, and the check is a licence read, not a category glance.
Pricing the exit is a separate exercise, and the per-token gap is a poor guide to it. Self-hosting swaps a variable per-token bill for a fixed per-hour one. GPUs bill whether or not they serve traffic, storage is extra, and so is the operations work that TCO calculators leave out: driver and inference-stack upgrades, capacity sizing, preemption handling, model revalidation on every update, on-call. OptimNow's practice reference puts the crossover for a dedicated deployment at roughly 200M to 500M tokens per day, depending on model size and quantisation, with 0.5 to 2 dedicated ML-Ops FTE. Below those volumes, and below FinOps Run maturity, the managed API is usually the cheaper answer even when the per-token rate is higher.
What to do with this
- Check the ceiling before the price. If the workload needs quality above 1484, the open question is closed and the decision is which proprietary model.
- Run the substitution per workload, not per catalogue. The median 24.9x is a distribution, not a rate. Your shortlist may sit in the part of it where the gap is small or negative.
- Name the seller before you quote the price. An open model's rate is a range across sellers, not a number: the median gap between the cheapest and dearest is 1.7x. A business case built on the cheapest rate is built on one supplier's pricing decision.
- Read the licence before the roadmap. Open weights and open source are different commercial positions, and the difference only appears in production.
- Treat these as list prices, not spend. Arena is one general benchmark; treat a gap under roughly 20 points as noise.
Figures on this page come from the 2026-09-06 price snapshot and rebuild nightly. Seller prices are from the 2026-09-06 read. Filter the pricing table by Openness in the Business view to run this against your own shortlist.
The FinOps reasoning behind this guide comes from OptimNow's open-source practice library: github.com/OptimNow/cloud-finops-skills. If you want the self-hosted question priced against 90 days of your own usage data, OptimNow runs that review.
Your real cost also depends on caching and batching, negotiated discounts, marketplace agreements, and the contracts you already have. OptimNow audits exactly that.
Get a FinOps reviewData snapshot: 2026-09-06 · Compare all models · Cloud compute pricing · All guides