GMI Cloud Inference Engine
- Categories
- Inference hosts
- Owned by
- GMI Cloud (private)
- HQ
- US
- Website
- https://www.gmicloud.ai
- Pricing page
- https://www.gmicloud.ai/model-library
- Fee model
- per_token+per_second_gpu
- Fees
- Serverless MaaS per 1M tokens, frequent percentage discounts off list; dedicated endpoints and GPUs by the hour (H100 from $2.00/GPU-h, H200 from $2.60, B200 from $4.00, GB200 from $8.00)
- Free tier
- no
- OpenAI-compatible API
- yes
- Bring your own key
- no
- Notes
- Marketing pages are client-rendered; docs say current prices live in the console Model Hub. Offer prices read from the console's unauthenticated pricing API (vendor-served); figures match OpenRouter's GMICloud endpoints. Also resells closed models (Claude, GPT, Gemini) — not recorded.
- Sources
- https://console.gmicloud.ai/api/v1/billing/model_prices https://www.gmicloud.ai/pricing https://docs.gmicloud.ai/inference-engine/billing/price collected 2026-09-18
LLM prices (32 models, USD per 1M tokens)
| Model | Input | Output | Cached input | Context | Source |
|---|---|---|---|---|---|
| DeepSeek V3.1 deepseek-ai/DeepSeek-V3.1 |
$0.27 | $1 | — | — | ✓ src src2 |
| DeepSeek V3.2 deepseek-ai/DeepSeek-V3.2 |
$0.29 | $0.43 | — | — | ✓ src src2 |
| DeepSeek V4 Flash 0423 deepseek-ai/DeepSeek-V4-Flash |
$0.14 | $0.28 | — | — | ✓ src src2 |
| DeepSeek V4 Pro 0423 deepseek-ai/DeepSeek-V4-Pro |
$1.74 | $3.48 | — | — | ✓ src src2 |
| Gemma 3 27B IT google/gemma-3-27b-it |
$0.09 | $0.16 | — | — | ✓ src src2 |
| Gemma 4 26B A4B google/gemma-4-26b-a4b-it |
$0.13 | $0.40 | — | — | ✓ src src2 |
| Gemma 4 31B google/gemma-4-31b-it |
$0.14 | $0.40 | — | — | ✓ src src2 |
| GLM 4.6 zai-org/GLM-4.6 |
$0.60 | $2 | — | — | ✓ src src2 |
| GLM 4.7 zai-org/GLM-4.7-FP8 |
$0.60 | $2.20 | — | — | ✓ src src2 |
| GLM 5 zai-org/GLM-5-FP8 |
$1 | $3.20 | — | — | ✓ src src2 |
| GLM 5.1 zai-org/GLM-5.1-FP8 |
$1.40 | $4.40 | — | — | ✓ src src2 |
| GLM-5.3 zai-org/GLM-5.3 |
$1.40 | $4.40 | — | — | ✓ src src2 |
| gpt-oss-120b openai/gpt-oss-120b |
$0.05 | $0.25 | — | — | ✓ src src2 |
| gpt-oss-20b openai/gpt-oss-20b |
$0.04 | $0.15 | — | — | ✓ src src2 |
| Kimi K2 0905 moonshotai/Kimi-K2-Instruct-0905 |
$0.57 | $2.29 | — | — | ✓ src src2 |
| Kimi K2.5 moonshotai/Kimi-K2.5 |
$0.60 | $3 | — | — | ✓ src src2 |
| Kimi K2.6 moonshotai/Kimi-K2.6 |
$0.95 | $4 | — | — | ✓ src src2 |
| Kimi K3 moonshotai/kimi-k3 |
$3 | $15 | — | — | ✓ src src2 |
| Llama 3.3 70B Instruct meta-llama/Llama-3.3-70B-Instruct |
$0.25 | $0.75 | — | — | ✓ src src2 |
| Llama 4 Maverick 17B Instruct meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8 |
$0.25 | $0.80 | — | — | ✓ src src2 |
| Llama 4 Scout 17B Instruct meta-llama/Llama-4-Scout-17B-16E-Instruct |
$0.08 | $0.50 | — | — | ✓ src src2 |
| MiniMax M2.5 MiniMaxAI/MiniMax-M2.5 |
$0.30 | $1.20 | — | — | ✓ src src2 |
| MiniMax M2.7 MiniMaxAI/MiniMax-M2.7 |
$0.30 | $1.20 | — | — | ✓ src src2 |
| MiniMax M3 MiniMaxAI/MiniMax-M3 |
$0.30 | $1.20 | — | — | ✓ src src2 |
| NVIDIA Nemotron 3 Ultra (Preview) nvidia/nemotron-3-ultra-550b-a55b |
$0.80 | $2.60 | — | — | ✓ src src2 |
| Qwen3 235B A22B 2507 Qwen/Qwen3-235B-A22B-Instruct-2507-FP8 |
$0.35 | $1.40 | — | — | ✓ src src2 |
| Qwen3 32B Qwen/Qwen3-32B-FP8 |
$0.10 | $0.60 | — | — | ✓ src src2 |
| Qwen3 Coder 480B A35B Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8 |
$0.90 | $4.50 | — | — | ✓ src src2 |
| Qwen3 Next 80B A3B Qwen/Qwen3-Next-80B-A3B-Instruct |
$0.15 | $1.50 | — | — | ✓ src src2 |
| Qwen3.5-397B-A17B Qwen/Qwen3.5-397B-A17B |
$0.60 | $3.60 | — | — | ✓ src src2 |
| Qwen3.8-27B Qwen/Qwen3.8-27B |
$0.45 | $3.20 | — | — | ✓ src src2 |
| R1 0528 deepseek-ai/DeepSeek-R1-0528 |
$0.57 | $2.29 | — | — | ✓ src src2 |