Llama 3.1 8B Instruct API Pricing
Meta · Budget · Llama 3.x · released 2024-07. Prices in USD per 1 million tokens. Data snapshot: 2026-09-11.
- Input: $0.05 per 1M tokens
- Output: $0.08 per 1M tokens
- Batch API (input/output): not published
- Prompt-cache read: $0.03
- Context window: 131K
- Arena ELO: 1211 (as of 2026-09-02)
Cost per request by business use case
List price vs optimized (prompt caching on the cached input share, batch API for async workloads, only where the provider publishes those rates).
| Use case | Tokens (in/out) | Cost/req (list) | Cost/req (optimized) | Monthly @100K req |
|---|---|---|---|---|
| Support Ticket | 1,500 / 500 | $0.000115 | $0.000093 | $9 |
| Knowledge Q&A | 2,000 / 800 | $0.000164 | $0.000129 | $13 |
| Meeting Summary | 10,000 / 1,200 | $0.000596 | $0.000571 | $57 |
| Marketing Content | 2,500 / 1,800 | $0.000269 | $0.000257 | $26 |
| Coding Task | 3,000 / 2,000 | $0.000310 | $0.000273 | $27 |
| Invoice Processing | 1,500 / 600 | $0.000123 | $0.000112 | $11 |
| Call Summary | 2,000 / 700 | $0.000156 | $0.000151 | $15 |
| Agent Workflow | 6,000 / 3,000 | $0.000540 | $0.000435 | $44 |
Llama 3.1 8B Instruct pricing by host
A host is a company that runs the model and sells it by the token — the lab that made it, a cloud platform like Amazon Bedrock, Azure or Vertex, or an independent provider. One model, several sellers, and often several prices. This model is sold by 5 different hosts. The price quoted above is the one we publish so that models stay comparable; it is not necessarily the cheapest way to buy this model. List prices as published on 2026-09-11 via OpenRouter — negotiated rates, committed-use discounts and provisioned capacity are not shown. These hosts do not all serve the model at the same numeric precision, so the cheapest is not a like-for-like substitute for the dearest.
| Host | Input $/1M | Output $/1M |
|---|---|---|
| DeepInfra — fp8 | $0.02 | $0.04 |
| Novita — fp8 | $0.02 | $0.05 |
| Groq (the price shown above) | $0.05 | $0.08 |
| CoreWeave — bf16 | $0.22 | $0.22 |
| Cloudflare — fp8 | $0.15 | $0.29 |
Models comparable to Llama 3.1 8B Instruct
The nearest Budget models on Arena ELO, so a substitution is compared on quality as well as price.
- Phi 4 (Microsoft) — $0.07 in / $0.14 out per 1M, Arena ELO 1256
- Claude 3 Haiku (Anthropic) — $0.25 in / $1.25 out per 1M, Arena ELO 1261
- Gemma 2 27B (Google) — $0.65 in / $0.65 out per 1M, Arena ELO 1289
- Llama 3.1 70B Instruct (Meta) — $0.40 in / $0.40 out per 1M, Arena ELO 1293
Common questions about Llama 3.1 8B Instruct pricing
- How much does Llama 3.1 8B Instruct cost?
- Llama 3.1 8B Instruct costs $0.05 per million input tokens and $0.08 per million output tokens on the API.
- Does Llama 3.1 8B Instruct offer batch or cached pricing?
- Prompt-cache reads cost $0.03/M input tokens. No batch pricing is published.
- What would Llama 3.1 8B Instruct cost per month?
- For a support-ticket workload (1,500 input + 500 output tokens per request) at 100K requests/month, Llama 3.1 8B Instruct costs about $12 at list prices.
These are list prices: real cost also depends on caching, batching, negotiated discounts, marketplace agreements, and existing contracts. Get a FinOps review from OptimNow or compare all models.
Build the business case for Llama 3.1 8B Instruct in the AI ROI Calculator: add your business value on top of these costs to get ROI, break-even and payback.
All models · Compare side by side · Cloud compute pricing · Price barometer · FinOps guides · API and MCP server