Groq
- Categories
- Inference hosts
- Owned by
- Groq, Inc. (private)
- HQ
- US
- Website
- https://groq.com
- Pricing page
- https://console.groq.com/docs/models
- Fee model
- per_token
- Fees
- Per-token on LPU hardware; Batch API 50% discount; cached input tokens 50% discount (discount does not stack with batch)
- Free tier
- Free plan with rate limits (Developer plan limits listed per model)
- OpenAI-compatible API
- yes
- Notes
- groq.com/pricing is client-rendered and returned no prices; prices taken from console docs model table. Llama 3.1 8B, Llama 3.3 70B and MiniMax M2.7 now listed as Enterprise / Contact Sales.
- Sources
- https://console.groq.com/docs/models https://console.groq.com/docs/batch https://console.groq.com/docs/prompt-caching collected 2026-09-18
LLM prices (7 models, USD per 1M tokens)
| Model | Input | Output | Cached input | Context | Source |
|---|---|---|---|---|---|
| gpt-oss-120b openai/gpt-oss-120b |
$0.15 | $0.60 | — | 131k | ✓ src |
| gpt-oss-20b openai/gpt-oss-20b |
$0.075 | $0.30 | — | 131k | ✓ src |
| gpt-oss-safeguard-20b openai/gpt-oss-safeguard-20b |
$0.075 | $0.30 | — | 131k | ✓ src |
| Llama 3.1 8B Instruct llama-3.1-8b-instant |
— | — | — | 131k | ◐ src |
| Llama 3.3 70B Instruct llama-3.3-70b-versatile |
— | — | — | 131k | ◐ src |
| MiniMax M2.7 minimaxai/minimax-m2.7 |
— | — | — | 197k | ◐ src |
| Qwen3.8-27B qwen/qwen3.8-27b |
$0.80 | $4 | — | 131k | ✓ src |