Venice.ai API
- Categories
- Inference hosts
- Owned by
- Venice.ai (private)
- Website
- https://venice.ai
- Pricing page
- https://docs.venice.ai/overview/pricing
- Fee model
- per_token
- Fees
- Per-token USD pricing; also payable with DIEM (1 Diem = $1/day of compute). Some models offered in E2EE/TEE variants at different prices.
- Free tier
- no
- OpenAI-compatible API
- yes
- Bring your own key
- no
- Notes
- Privacy-focused: models tagged Private (no retention) or Anonymized (proxied to upstream, e.g. Claude/Gemini/GPT). Also resells closed models. HQ country not stated on the pages used.
- Sources
- https://docs.venice.ai/overview/pricing https://api.venice.ai/api/v1/models?type=text collected 2026-09-18
LLM prices (33 models, USD per 1M tokens)
| Model | Input | Output | Cached input | Context | Source |
|---|---|---|---|---|---|
| DeepSeek V3.2 deepseek-v3.2 |
$0.33 | $0.48 | $0.16 | 160k | ✓ src src2 |
| DeepSeek V4 Flash 0423 deepseek-v4-flash |
$0.138 | $0.275 | $0.028 | 1M | ✓ src src2 |
| DeepSeek V4 Flash 0731 deepseek-v4-flash-0731 |
$0.175 | $0.35 | $0.035 | 1M | ✓ src src2 |
| DeepSeek V4 Pro 0423 deepseek-v4-pro |
$1.65 | $3.30 | $0.33 | 1M | ✓ src src2 |
| Gemma 3 27B IT google-gemma-3-27b-it |
$0.12 | $0.20 | — | 198k | ✓ src src2 |
| Gemma 4 26B A4B google-gemma-4-26b-a4b-it |
$0.13 | $0.40 | $0.05 | 256k | ✓ src src2 |
| Gemma 4 31B google-gemma-4-31b-it |
$0.12 | $0.36 | $0.09 | 256k | ✓ src src2 |
| GLM 4.6 zai-org-glm-4.6 |
$0.43 | $1.75 | $0.08 | 198k | ✓ src src2 |
| GLM 4.7 zai-org-glm-4.7 |
$0.55 | $2.65 | $0.11 | 198k | ✓ src src2 |
| GLM 4.7 Flash zai-org-glm-4.7-flash |
$0.06 | $0.40 | $0.01 | 128k | ✓ src src2 |
| GLM 5 zai-org-glm-5 |
$1 | $3.20 | $0.20 | 198k | ✓ src src2 |
| GLM 5.1 zai-org-glm-5-1 |
$1.54 | $4.84 | $0.286 | 200k | ✓ src src2 |
| GLM-5.2 zai-org-glm-5-2 |
$1.40 | $4.40 | $0.26 | 1M | ✓ src src2 |
| GLM-5.3 z-ai-glm-5-3 |
$1.75 | $5.50 | $0.325 | 1M | ✓ src src2 |
| GLM-5.3 Flash z-ai-glm-5-3-flash |
$0.15 | $0.50 | $0.03 | 1.05M | ✓ src src2 |
| gpt-oss-120b openai-gpt-oss-120b |
$0.07 | $0.30 | — | 128k | ✓ src src2 |
| Hermes 3 Llama 3.1 405b hermes-3-llama-3.1-405b |
$1.10 | $3 | — | 128k | ✓ src src2 |
| Kimi K2.5 kimi-k2-5 |
$0.56 | $3.50 | $0.22 | 256k | ✓ src src2 |
| Kimi K2.6 kimi-k2-6 |
$0.75 | $3.50 | $0.16 | 256k | ✓ src src2 |
| Kimi K3 kimi-k3 |
$3.75 | $18.75 | $0.375 | 1M | ✓ src src2 |
| Llama 3.3 70B Instruct llama-3.3-70b |
$0.70 | $2.80 | — | 128k | ✓ src src2 |
| MiniMax M2.5 minimax-m25 |
$0.27 | $0.95 | $0.03 | 198k | ✓ src src2 |
| MiniMax M2.7 minimax-m27 |
$0.375 | $1.50 | $0.069 | 198k | ✓ src src2 |
| Mistral Small 3.2 24B mistral-small-3-2-24b-instruct |
$0.094 | $0.25 | — | 256k | ✓ src src2 |
| Qwen3 235B A22B 2507 qwen3-235b-a22b-instruct-2507 |
$0.15 | $0.75 | — | 128k | ✓ src src2 |
| Qwen3 Coder 480B A35B qwen3-coder-480b-a35b-instruct-turbo |
$0.35 | $1.50 | $0.04 | 256k | ✓ src src2 |
| Qwen3 Next 80B A3B qwen3-next-80b |
$0.35 | $1.90 | — | 256k | ✓ src src2 |
| Qwen3-235B-A22B-Thinking-2507 qwen3-235b-a22b-thinking-2507 |
$0.45 | $3.50 | — | 128k | ✓ src src2 |
| Qwen3.5-397B-A17B qwen3-5-397b-a17b |
$0.75 | $4.50 | — | 128k | ✓ src src2 |
| Qwen3.6 27B qwen3-6-27b |
$0.325 | $3.25 | — | 256k | ✓ src src2 |
| Qwen3.6 35B A3B qwen3-6-35b-a3b |
$0.10 | $1 | — | 256k | ✓ src src2 |
| Qwen3.8-2.4T-A95B qwen-3-8-2-4t-a95b |
$2.50 | $7.50 | $0.312 | 262k | ✓ src src2 |
| Qwen3.8-27B qwen-3-8-27b |
$0.45 | $3.20 | — | 262k | ✓ src src2 |