OptimToken All guides

Bedrock, Vertex, Azure or direct: what the channel changes

A FinOps guide by OptimNow · prices from the live OptimToken catalogue

Most models on a shortlist are sold in more than one place. Claude is on Anthropic's own API, on AWS Bedrock and on Google Vertex. GPT-5.6 is on OpenAI's API, on Azure AI Foundry and, since 2026, on Bedrock. This guide calls each of those places a channel. Procurement usually starts by asking which channel is cheapest. Across the models the hub tracks, that question has a dull answer, and the differences worth negotiating are elsewhere.

How many channels sell a model

The hub keeps a hand-checked list of where each model is sold. It covers 60 of the 265 models in the catalogue. 22 are sold through 4 channels, 22 through 3, 16 through 2. The list says where a model can be bought, and nothing about the price there.

Channels carrying each model (n=60 with a checked entry) Channels carrying each model (n=60 with a checked entry) OpenRouter 60 / 60 Direct API 54 / 60 AWS Bedrock 28 / 60 Azure AI Foundry 25 / 60 Google Vertex 19 / 60
The pale bar is all 60 models, the dark bar is the ones that channel carries. Every model is on OpenRouter, because that is where the hub reads its prices.

6 of the 60 have no API from their maker. Llama 4 Maverick, Llama 4 Scout, Nova Premier 1.0, Nova Micro 1.0 and 2 more are sold only through someone else. Meta publishes Llama weights and runs no service of its own, and Amazon sells Nova through Bedrock alone, so for those a cloud is the only shop.

The price usually does not move

Every night the hub reads the price each seller lists for a model: the maker, the three clouds, and any reseller. For 27 models both the maker and at least one cloud publish a price, which gives 61 comparisons. 53 of them are identical, input and output, to the cent.

Take Claude Fable 5, sold through 4 channels and priced by 5 sellers on 2026-09-06:

SellerInput /1MOutput /1MBlended 30/70Against the maker
Claude Platform on AWS $10.00 $50.00 $38.00 same
Azure $10.00 $50.00 $38.00 same
Amazon Bedrock $10.00 $50.00 $38.00 same
Google $10.00 $50.00 $38.00 same
Anthropic (the model's maker) $10.00 $50.00 $38.00 same

Moving that workload from the maker to a cloud saves nothing on the rate. The clouds resell at the maker's list price, and that holds for most proprietary models in the catalogue.

The 8 comparisons that do differ deserve a read before anyone concludes that a cloud marks the model up. 5 sit at exactly 10% above the maker's rate. AWS documents a 10% premium on regional endpoints for recent Claude models, against its global ones. So the likeliest reading is an endpoint choice, not a markup. The hub cannot see which endpoint each row prices, so treat that as a lead to check. The other 3 gaps are wider, and a gap that wide usually means two different products. Sellers list a cheaper asynchronous tier and a dearer priority tier next to the standard one. Check that both rows are the same tier before you act on the gap.

Where the seller is the whole price decision

That flat price is a feature of proprietary models. Split the 60 by licence and the two halves behave differently. The 46 proprietary models have a median of 2 sellers and a median gap of 1.00x between the cheapest and the dearest. The 14 open-weights and open-source models have a median of 4 sellers and a median gap of 1.66x. One maker setting a price is a rate card. Many operators serving the same downloadable weights is a market, and a market has a spread.

DeepSeek V3.2 is the widest case: 15 sellers, from $0.28 per 1M tokens at GMICloud to $4.05 at SambaNova. That is 14.5x for the same weights, and here the seller decides the bill. The cheap end is not free of catches, because sellers differ on quantization, context limit and throughput, so test the cheapest one on your own prompts before you budget on it. Compare the two examples side by side.

Command A is the opposite case: sold through 4 channels, and exactly one of them publishes a price the hub can read. A channel can carry a model without telling you what it costs there. Four logos on a slide are not four quotes.

What does move: where the model runs

A model bought through a cloud runs only where that cloud serves it, and the map differs by model. On Bedrock, Claude Haiku 4.5 is served in 12 regions and R1 in 3. 2 of the 17 models on the checked list have no European region at all: DeepSeek V3.2 and R1.

Where each model runs on AWS Bedrock Where each model runs on AWS Bedrock Regions serving the model, by geography. An empty cell means none. North America Europe Asia Pacific Claude Haiku 4.5 Claude Haiku 4.5: us-east-1, us-east-2, us-west-23Claude Haiku 4.5: eu-central-1, eu-north-1, eu-south-1, eu-south-2, eu-west-1, eu-west-36Claude Haiku 4.5: ap-northeast-1, ap-northeast-3, ap-southeast-13 Claude Opus 4.7 Claude Opus 4.7: us-east-1, us-east-2, us-west-23Claude Opus 4.7: eu-central-1, eu-north-1, eu-south-1, eu-south-2, eu-west-1, eu-west-36Claude Opus 4.7: ap-northeast-1, ap-northeast-3, ap-southeast-13 Claude Opus 4.8 Claude Opus 4.8: us-east-1, us-east-2, us-west-23Claude Opus 4.8: eu-central-1, eu-north-1, eu-south-1, eu-south-2, eu-west-1, eu-west-36Claude Opus 4.8: ap-northeast-1, ap-northeast-3, ap-southeast-13 Claude Sonnet 4.5 Claude Sonnet 4.5: us-east-1, us-east-2, us-west-23Claude Sonnet 4.5: eu-central-1, eu-north-1, eu-south-1, eu-south-2, eu-west-1, eu-west-36Claude Sonnet 4.5: ap-northeast-1, ap-northeast-3, ap-southeast-13 Claude Sonnet 4.6 Claude Sonnet 4.6: us-east-1, us-east-2, us-west-23Claude Sonnet 4.6: eu-central-1, eu-north-1, eu-south-1, eu-south-2, eu-west-1, eu-west-36Claude Sonnet 4.6: ap-northeast-1, ap-northeast-3, ap-southeast-13 Claude 3 Haiku Claude 3 Haiku: us-east-1, us-east-2, us-west-23Claude 3 Haiku: eu-central-1, eu-west-1, eu-west-33Claude 3 Haiku: ap-northeast-1, ap-northeast-2, ap-south-1, ap-southeast-1, ap-southeast-25 Claude Opus 4.5 Claude Opus 4.5: us-east-1, us-east-2, us-west-23Claude Opus 4.5: eu-central-1, eu-north-1, eu-south-1, eu-south-2, eu-west-1, eu-west-36Claude Opus 4.5: ap-northeast-1, ap-southeast-12 Claude Opus 4.6 Claude Opus 4.6: us-east-1, us-east-2, us-west-23Claude Opus 4.6: eu-central-1, eu-north-1, eu-south-1, eu-south-2, eu-west-1, eu-west-36Claude Opus 4.6: ap-northeast-1, ap-southeast-12 Claude Opus 5 Claude Opus 5: us-east-1, us-east-2, us-west-23Claude Opus 5: eu-central-1, eu-north-1, eu-south-1, eu-south-2, eu-west-1, eu-west-36Claude Opus 5: ap-northeast-1, ap-southeast-12 Claude Sonnet 5 Claude Sonnet 5: us-east-1, us-east-2, us-west-23Claude Sonnet 5: eu-central-1, eu-north-1, eu-south-1, eu-south-2, eu-west-1, eu-west-36Claude Sonnet 5: ap-northeast-1, ap-southeast-12 Claude Fable 5 Claude Fable 5: us-east-1, us-east-2, us-west-23Claude Fable 5: eu-central-1, eu-west-12Claude Fable 5: ap-northeast-1, ap-southeast-12 GPT-5.6 Luna GPT-5.6 Luna: us-east-1, us-east-2, us-west-23GPT-5.6 Luna: eu-central-1, eu-west-12GPT-5.6 Luna: ap-northeast-1, ap-southeast-12 GPT-5.6 Sol GPT-5.6 Sol: us-east-1, us-east-2, us-west-23GPT-5.6 Sol: eu-central-1, eu-west-12GPT-5.6 Sol: ap-northeast-1, ap-southeast-12 GPT-5.6 Terra GPT-5.6 Terra: us-east-1, us-east-2, us-west-23GPT-5.6 Terra: eu-central-1, eu-west-12GPT-5.6 Terra: ap-northeast-1, ap-southeast-12 Mistral Large Mistral Large: us-east-1, us-east-2, us-west-23Mistral Large: eu-west-11Mistral Large: ap-northeast-11 DeepSeek V3.2 DeepSeek V3.2: us-east-1, us-east-2, us-west-23DeepSeek V3.2: not served in EuropeDeepSeek V3.2: ap-northeast-11 R1 R1: us-east-1, us-east-2, us-west-23R1: not served in EuropeR1: not served in Asia Pacific
From the 2026-08-20 sweep of 18 Bedrock regions. Bedrock is the only channel that publishes this per model, so it is the one shown. Global inference profiles are not counted: they route anywhere, which is not an answer to a residency question.

A model can win the benchmark and fit the budget and still be unusable for a workload that must keep its data in Europe. Teams usually find this out at deployment, after the shortlist has closed. Check the map while there are still alternatives on it.

What really differs: the commitment terms

Same price, different regions. The third difference appears when the workload is large enough to reserve capacity. The three clouds sell reserved inference on terms that do not line up, and this is where a channel choice becomes hard to undo:

AWS BedrockGoogle VertexAzure AI Foundry
Commitment scopeLocked to one modelOne publisher, any of its modelsPooled units, any model
Upgrade mid-termNoWithin the publisherYes
Overflow past reserved capacityBuild the failover yourselfPay-as-you-go by defaultOpt-in spillover per deployment
Cost attributionIAM principal, inference profiles, project tagsBilling labels, BigQuery exportResource tags, cost analysis

Read the top row against a model roadmap that moves every few months. A commitment locked to one model is stranded the day its successor ships, and the two prices you compared at the start were identical anyway. This is the FinOps Foundation's GenAI capacity framing, as distilled in OptimNow's practice library: commitment terms, not token rates, are where a channel decision gets expensive. Size any reservation to average load, with a spillover path for the peaks.

Running the check on your own shortlist

Figures on this page come from the 2026-09-06 catalogue snapshot and the 2026-09-06 seller prices, and rebuild nightly. The pricing table filters by Available on, and every model page lists its sellers with the price each one publishes.

The FinOps reasoning behind this guide comes from OptimNow's open-source practice library: github.com/OptimNow/cloud-finops-skills. If you want a channel and commitment review run against your own inference usage, OptimNow does that work.

These are list prices.

Your real cost also depends on caching and batching, negotiated discounts, marketplace agreements, and the contracts you already have. OptimNow audits exactly that.

Get a FinOps review

Data snapshot: 2026-09-06 · Compare all models · Cloud compute pricing · All guides