Vision, audio, agents: what each capability adds to the bill
A FinOps guide by OptimNow · prices from the live OptimToken catalogue
A product team says the feature needs to read screenshots. Someone then asks what vision costs, and the catalogue appears to answer: models that accept images carry a median blended price 2.88x the models that do not. That number is real and it is the wrong one to budget against. It compares vendors, not capabilities. This guide separates the two on 285 priced models. The cost that remains is the size of the shortlist each requirement leaves.
Three capabilities worth measuring, one worth skipping
OptimToken reads capability tags from upstream model metadata. Image input and audio come from the model's declared input and output modalities. Tool use comes from the parameters the model accepts. Those three are facts about the model, so they can be priced.
The Code tag cannot, and this guide leaves it out. OptimToken assigns it partly from a price threshold, so asking what Code adds to the bill would return the threshold. Quoting a code premium from this catalogue would be circular. Saying so beats dropping it quietly.
The catalogue-wide premium is a mix effect
Capabilities cluster with vendors. A house whose entire line-up is multimodal and priced for enterprises puts every one of its models on the with-vision side. A house selling small text models puts every one of its models on the other. The gap between the two sides is the gap between those two houses.
Compare inside one vendor's catalogue instead and the premium shrinks. Of the providers listing at least 2 models on each side, the median ratio between their vision models and their text-only models is 1.67x, against 2.88x across the catalogue. In 2 of 8 the ratio sits below 1, meaning the vendor's vision models are the cheaper half of its own line-up.
The spread across vendors is wide, from 0.57x to 2.23x, so image input is not free everywhere. The vendor you buy from decides more of the bill than the capability does.
Counter-example: dropping a capability can raise the bill
The usual FinOps instinct is to treat model selection as rightsizing, and to stop paying for what the feature does not use. On this catalogue that instinct misfires at 2 of the 8 vendors with enough models to compare. At Zhipu the median vision model costs $0.986 per 1M blended tokens against $1.720 for its text-only models, so removing image input from the requirement moves the shortlist to the more expensive half.
At Google, the cheapest model accepting images is Gemma 3 4B at $0.085 per 1M blended tokens. Drop the image requirement and the cheapest thing on the same shelf is Gemma 2 27B at $0.650, a 7.65x gap in the wrong direction. Compare them side by side. The vendor set the price by the model's position in the line-up. The capability came with it.
What a requirement costs instead: the shortlist
If the rate card does not price capability, something else does. Every requirement removes models from consideration. A thinner shortlist costs you the competition that was holding the price down.
Read as a shortlist, the four requirements are not comparable at all. Tool use leaves 253 of 285 models, 89% of the catalogue. The cheapest model carrying it is Mistral Nemo at $0.027 per 1M blended tokens, the cheapest model in the catalogue outright. Requiring tool use costs nothing at the floor. The budget option already has it.
Worked example: the audio requirement
Audio behaves differently. It is the requirement worth planning around. Across the catalogue, 24 of 285 models accept or produce it, 8% of the total. The catalogue-wide price ratio is only 1.54x, so on a rate-card reading audio looks close to free. The cheapest model that handles audio is MiMo-V2.6-Flash at $0.238 per 1M blended tokens, against $0.027 for Mistral Nemo with no capability requirement at all. Adding audio to the specification multiplies the floor by 8.91x.
Concentration compounds it. Google publishes 13 of the 24 audio models, 54% of the shortlist, spread across 7 providers in total. A requirement met by one vendor on most of the shortlist is a negotiating position as much as a technical choice, and it belongs in the risk register next to the cost estimate. Where audio is a minority of requests, a router is the cheaper structure. Send the audio requests to a model that handles them and leave the rest on the catalogue floor. Paying the audio floor on every call buys audio for calls that carry none.
How to use this when choosing a model
These steps assume a team that already tags requests by type, the Walk stage in FinOps maturity terms. A team still reading one monthly invoice has nothing to route yet, and request-level tagging comes first.
- Specify capabilities per request type rather than per feature. One multimodal model for a workload that is 5% images sets the floor for the other 95%.
- Check the vendor's own ladder before paying for a capability. Within a line-up the premium ran a median 1.67x for image input, and went the other way at 2 of 8 vendors.
- Count the shortlist, then look at its concentration. A requirement that leaves 24 models, most from one vendor, is a different risk than one that leaves 253.
- Set the quality bar before the capability list. The bar belongs to the product owner; the model choice is the cheapest path that clears it. Capability tags say what a model accepts. They say nothing about how well it does the job.
These are published list prices for the model's own API, on a 30/70 input-output blend (the formula is here). They do not include what a capability costs in tokens once it is used. An image is billed as input tokens when it reaches the model, and a long screenshot is a long prompt. Price the payload with that formula; this guide prices the shortlist only.
Figures on this page come from the 2026-10-04 price snapshot and rebuild nightly.
The FinOps reasoning behind this guide comes from OptimNow's open-source practice library: github.com/OptimNow/cloud-finops-skills. If you want your own model shortlist sized against the capabilities your workload uses, OptimNow runs that review.
Questions this guide answers
Do models with image input cost more?
Across the catalogue yes, inside one vendor's line-up mostly not. On the 2026-10-04 OptimToken catalogue snapshot, models accepting images carry a median blended price 2.88x that of models that do not. Comparing within each provider that lists at least 2 models on each side, the median ratio falls to 1.67x, and at 2 of 8 providers the vision models are the cheaper half. The catalogue-wide gap measures the difference between vendors, not the price of the capability.
How many models support audio?
On the 2026-10-04 OptimToken catalogue snapshot, 24 of 285 priced text models accept or produce audio, 8% of the catalogue. The cheapest is MiMo-V2.6-Flash at $0.238 per 1M blended tokens, against $0.027 for Mistral Nemo with no capability requirement, so requiring audio multiplies the price floor by 8.91x.
Does requiring tool use narrow the choice of model?
Barely. On the 2026-10-04 OptimToken catalogue snapshot, 253 of 285 priced models accept tool-use parameters, 89% of the catalogue, and the cheapest of them (Mistral Nemo at $0.027 per 1M blended tokens) is also the cheapest model in the catalogue outright. Requiring tool use costs nothing at the floor.
What does a capability requirement cost?
The shortlist rather than the rate. Each requirement removes models from consideration, and the floor rises with it. On the 2026-10-04 OptimToken catalogue snapshot: tool use leaves 253 of 285 models and a floor 1.00x the catalogue floor; a reasoning control leaves 211 of 285 models and a floor 1.85x the catalogue floor; image input leaves 174 of 285 models and a floor 1.85x the catalogue floor; audio leaves 24 of 285 models and a floor 8.91x the catalogue floor.
Your real cost also depends on caching and batching, negotiated discounts, marketplace agreements, and the contracts you already have. OptimNow audits exactly that.
Get a FinOps reviewData snapshot: 2026-10-04 · Compare all models · Cloud compute pricing · All guides