Llama 3.3 70B Instruct meta-llama/llama-3.3-70b-instruct
Served by 22 providers. Cheapest: DeepInfra at $0.155 blended /1M tokens. The most expensive listing costs 7.9× as much for the same model.
| Provider | Type | Input /1M | Output /1M | Cached input | Blended | Context | Notes | Source |
|---|---|---|---|---|---|---|---|---|
| DeepInfra meta-llama/Llama-3.3-70B-Instruct-Turbo |
Inference hosts | $0.10 | $0.32 | — | $0.155 | 131k | fp8 |
✓ src src2 2026-09-18 |
| OpenRouter meta-llama/llama-3.3-70b-instruct |
LLM routers | $0.10 | $0.32 | — | $0.155 | 131k | api modality text->text |
✓ src 2026-09-18 |
| Inference.net llama-3.3-70b-instruct |
Inference hosts | $0.13 | $0.40 | — | $0.198 | 131k | https://inference.net/models shows the same figures |
✓ src 2026-09-18 |
| Requesty nebius/meta-llama/Llama-3.3-70B-Instruct |
LLM routers | $0.13 | $0.40 | — | $0.198 | 128k | fp8 upstream route nebius; geo eu; 2 other upstream routes listed (input $0.23-$0.39/M); API list price; Requesty's 5% markup on model cost (pricing page) is not included |
✓ src 2026-09-18 |
| Novita AI meta-llama/llama-3.3-70b-instruct |
Inference hosts | $0.135 | $0.40 | — | $0.201 | 12k |
✓ src src2 2026-09-18 |
|
| AkashML meta-llama/Llama-3.3-70B-Instruct |
Inference hosts | $0.20 | $0.52 | $0.10 | $0.28 | — |
✓ src src2 2026-09-18 |
|
| Parasail | Inference hosts | $0.22 | $0.50 | $0.11 | $0.29 | — | fp8 |
✓ src 2026-09-18 |
| GMI Cloud Inference Engine meta-llama/Llama-3.3-70B-Instruct |
Inference hosts | $0.25 | $0.75 | — | $0.375 | — | from GMI console public pricing API (console is client-rendered; docs say prices live in console Model Hub) |
✓ src src2 2026-09-18 |
| Hyperbolic meta-llama/Llama-3.3-70B-Instruct |
Inference hosts | $0.40 | $0.40 | — | $0.40 | 131k | Docs list a single "$0.40/M tokens" price (no input/output split). Inference docs section is marked hidden/noindex; availability should be re-verified. |
◐ src 2026-09-18 |
| Featherless meta-llama/Llama-3.3-70B-Instruct |
Inference hosts | $0.65 | $0.75 | $0.52 | $0.675 | 33k | Developer-plan per-token rate; model_class llama33-70b; also usable flat-rate on $25/mo Chat plan (32K ctx, human chat only) |
✓ src src2 src3 2026-09-18 |
| CoreWeave meta-llama/Llama-3.3-70B-Instruct |
Inference hosts | $0.71 | $0.71 | — | $0.71 | 128k | Context from docs model table (rounded, e.g. "262k"). |
✓ src src2 2026-09-18 |
| Microsoft Foundry (Azure AI Foundry / Azure OpenAI) Llama-3.3-70B-Instruct |
Hyperscaler AI platforms | $0.71 | $0.71 | — | $0.71 | — | Global Standard, Azure Retail Prices API meter 'Llama 3.3 70B Inp glbl Tokens' (eastus); API meter is per 1K tokens, multiplied by 1000 (unit conversion only); Data Zone/Regional $0.781 in / $0.781 out |
✓ src src2 2026-09-18 |
| Amazon Bedrock meta.llama3-3-70b-instruct-v1:0 |
Hyperscaler AI platforms | $0.72 | $0.72 | — | $0.72 | 128k | batch −50% |
✓ src src2 2026-09-18 |
| Google Vertex AI (Gemini Enterprise Agent Platform) | Hyperscaler AI platforms | $0.72 | $0.72 | — | $0.72 | — | batch −50% Managed MaaS API. Batch $0.36/$0.36. |
✓ src 2026-09-18 |
| SambaNova Cloud (SambaCloud) Meta-Llama-3.3-70B-Instruct |
Inference hosts | $0.60 | $1.20 | — | $0.75 | 128k | Production model |
✓ src src2 2026-09-18 |
| IBM watsonx.ai llama-3-3-70b-instruct |
Hyperscaler AI platforms | $0.753 | $0.753 | — | $0.753 | — | Page shows a single rate "USD 0.7526" per 1M tokens; IBM states input and completion tokens are charged at the same rate (RU metric). |
✓ src src2 2026-09-18 |
| OVHcloud AI Endpoints Meta-Llama-3_3-70B-Instruct |
Inference hosts | $0.769 | $0.769 | — | $0.769 | — | fp8 list price EUR 0.67/M in, EUR 0.67/M out; context shown on catalog as "131K" (rounded); converted at ECB EURUSD 1.1481 2026-09-17 (latest ECB reference rate) |
✓ src 2026-09-18 |
| Cloudflare Workers AI @cf/meta/llama-3.3-70b-instruct-fp8-fast |
Inference hosts | $0.293 | $2.25 | — | $0.783 | — | fp8 Billed in neurons ($0.011/1k neurons); USD per-M-token equivalent as published by Cloudflare |
✓ src 2026-09-18 |
| Scaleway llama-3.3-70b-instruct |
Inference hosts | $1.03 | $1.03 | — | $1.03 | — | batch −50% list price EUR 0.9/M in, EUR 0.9/M out; Paris region; prices before tax; Batches API -50%; converted at ECB EURUSD 1.1481 2026-09-17 (latest ECB reference rate) |
✓ src 2026-09-18 |
| Together AI meta-llama/Llama-3.3-70B-Instruct-Turbo |
Inference hosts | $1.04 | $1.04 | — | $1.04 | 131k | fp8 |
✓ src 2026-09-18 |
| Venice.ai API llama-3.3-70b |
Inference hosts | $0.70 | $2.80 | — | $1.22 | 128k | fp8 |
✓ src src2 2026-09-18 |
| Groq llama-3.3-70b-versatile |
Inference hosts | — | — | — | — | 131k | batch −50% Listed as Enterprise / Contact Sales (no public per-token price). OpenRouter endpoint still shows Groq at $0.59/$0.79. |
◐ src 2026-09-18 |
✓ from the provider's own pricing page · ◐ partly confirmed there · 2° third-party source. Router prices may exclude the router's own fees; see each router's page.