Together AI
- Categories
- GPU clouds, Inference hosts
- Owned by
- Together Computer, Inc. (private)
- HQ
- US
- Website
- https://www.together.ai
- Pricing page
- https://www.together.ai/pricing
- Offers
- GPU clusters, inference API, fine-tuning
- Fee model
- per_token+per_second_gpu
- Fees
- Serverless per-token; Dedicated Inference per GPU-hour (HGX H100 $5.49 on-demand, promo $3.99 until 2026-09-30; HGX B200 $8.99); Provisioned Throughput sold in PTUs; Batch API priced separately
- Free tier
- no
- OpenAI-compatible API
- yes
- Notes
- Serverless prices per 1M tokens; docs table gives context length and quantization per model. Also sells GPU clusters (H100 $3.99/h on-demand).
- Sources
- https://www.together.ai/pricing https://docs.together.ai/docs/serverless-models collected 2026-09-18
LLM prices (27 models, USD per 1M tokens)
| Model | Input | Output | Cached input | Context | Source |
|---|---|---|---|---|---|
| Cogito v2.1 671B | $1.25 | $1.25 | — | — | ◐ src |
| DeepSeek V4 Flash 0731 deepseek-ai/DeepSeek-V4-Flash-0731 |
$0.14 | $0.28 | $0.03 | 1.05M | ✓ src |
| DeepSeek V4 Pro 0813 deepseek-ai/DeepSeek-V4-Pro-0813 |
$1.32 | $3.96 | $0.13 | 1.05M | ✓ src |
| DeepSeek V4.1 Flash deepseek-ai/DeepSeek-V4.1-Flash |
$0.30 | $1.20 | $0.006 | 1M | ✓ src |
| Gemma 4 31B | $0.39 | $0.97 | — | — | ◐ src |
| GLM-5.2 zai-org/GLM-5.2 |
$1.40 | $4.40 | $0.26 | 1.05M | ✓ src |
| GLM-5.3 zai-org/GLM-5.3 |
$1.40 | $4.40 | $0.26 | 1.05M | ✓ src |
| GLM-5.3 Flash zai-org/GLM-5.3-Flash |
$0.15 | $0.50 | $0.03 | 1.05M | ✓ src |
| gpt-oss-120b openai/gpt-oss-120b |
$0.15 | $0.60 | — | 131k | ✓ src |
| Inkling thinkingmachines/Inkling |
$1 | $4.05 | $0.17 | 524k | ✓ src |
| Kimi K3 moonshotai/Kimi-K3 |
$3 | $15 | $0.30 | 1.05M | ✓ src |
| Llama 3 8B Instruct Lite | $0.14 | $0.14 | — | — | ◐ src |
| Llama 3.3 70B Instruct meta-llama/Llama-3.3-70B-Instruct-Turbo |
$1.04 | $1.04 | — | 131k | ✓ src |
| MiniMax M2.7 | $0.30 | $1.20 | $0.06 | — | ◐ src |
| MiniMax M3 MiniMaxAI/MiniMax-M3 |
$0.30 | $1.20 | $0.06 | 524k | ✓ src |
| Muse Glimmer 30B meta-models/Muse-Glimmer-30B |
$0.35 | $1.50 | $0.04 | 131k | ✓ src |
| Qwen2.5 7B Instruct Turbo | $0.30 | $0.30 | — | — | ◐ src |
| Qwen3 235B A22B 2507 | $0.20 | $0.60 | — | — | ◐ src |
| Qwen3.5 9B Qwen/Qwen3.5-9B |
$0.17 | $0.25 | — | 262k | ✓ src |
| Qwen3.5-397B-A17B | $0.60 | $3.60 | $0.35 | — | ◐ src |
| Qwen3.6 Plus Qwen/Qwen3.6-Plus |
$0.50 | $3 | — | 1M | ✓ src |
| Qwen3.7 Max Qwen/Qwen3.7-Max |
$2.50 | $7.50 | $0.50 | — | ◐ src src2 |
| Qwen3.7 Plus Qwen/Qwen3.7-Plus |
$0.32 | $1.28 | — | 1M | ✓ src |
| Qwen3.8 Flash Qwen/Qwen3.8-Flash |
$0.15 | $0.47 | — | 1M | ✓ src |
| Qwen3.8-2.4T-A95B Qwen/Qwen3.8-2.4T-A95B |
$2 | $6 | $0.50 | — | ◐ src src2 |
| Rnj-1 Instruct | $0.15 | $0.15 | — | — | ◐ src |
| Ternary Bonsai 27B Prism-ML/Ternary-Bonsai-27B |
$0 | $0 | — | 262k | ✓ src |
GPU rates
| Provider | GPU | GPUs | VRAM/GPU | Billing | Per GPU-hr | Instance /hr | Region | Source |
|---|---|---|---|---|---|---|---|---|
| Together AI | H100 SXM | 1 | 80 GB | spot | $1.99 | $1.99 | ✓ src ⓘ | |
| Together AI | H200 | 1 | 141 GB | spot | $2.99 | $2.99 | ✓ src ⓘ | |
| Together AI | H100 SXM | 1 | 80 GB | on demand | $3.99 | $3.99 | ✓ src ⓘ | |
| Together AI | B200 | 1 | 180 GB | spot | $4.09 | $4.09 | ✓ src ⓘ | |
| Together AI | H200 | 1 | 141 GB | on demand | $5.99 | $5.99 | ✓ src ⓘ | |
| Together AI | B200 | 1 | 180 GB | on demand | $8.19 | $8.19 | ✓ src ⓘ |