Nscale Serverless Inference
- Categories
- Inference hosts
- Owned by
- Nscale Global Holdings (private)
- HQ
- GB
- Website
- https://www.nscale.com/product/serverless
- Pricing page
- https://www.nscale.com/product/serverless
- Fee model
- per_token
- Fees
- Billed per 1M tokens (input and output); images per megapixel
- Free tier
- $5 free credit for new accounts
- OpenAI-compatible API
- yes
- Bring your own key
- no
- Notes
- Per-model prices are not on any public page (console / authenticated /v1/models only; docs pricing page access-restricted). Offer prices taken from Hugging Face Inference Providers router, which passes provider prices through → secondary.
- Sources
- https://www.nscale.com/product/serverless https://docs.nscale.com/docs/ai-services/models https://router.huggingface.co/v1/models collected 2026-09-18
LLM prices (11 models, USD per 1M tokens)
| Model | Input | Output | Cached input | Context | Source |
|---|---|---|---|---|---|
| DeepSeek R1 Distill Qwen 14B deepseek-ai/DeepSeek-R1-Distill-Qwen-14B |
$0.20 | $0.20 | — | 131k | 2° src |
| gpt-oss-120b openai/gpt-oss-120b |
$0.10 | $0.40 | — | 131k | 2° src |
| gpt-oss-20b openai/gpt-oss-20b |
$0.05 | $0.20 | — | 131k | 2° src |
| Llama 3.1 8B Instruct meta-llama/Llama-3.1-8B-Instruct |
$0.06 | $0.06 | — | 131k | 2° src |
| Llama 4 Scout 17B Instruct meta-llama/Llama-4-Scout-17B-16E-Instruct |
$0.09 | $0.29 | — | 890k | 2° src |
| qwen2.5-coder-32b-instruct Qwen/Qwen2.5-Coder-32B-Instruct |
$0.06 | $0.20 | — | 131k | 2° src |
| Qwen3 14B Qwen/Qwen3-14B |
$0.07 | $0.20 | — | 41k | 2° src |
| Qwen3 235B A22B Qwen/Qwen3-235B-A22B |
$0.20 | $0.60 | — | 32k | 2° src |
| Qwen3 235B A22B 2507 Qwen/Qwen3-235B-A22B-Instruct-2507 |
$0.20 | $0.60 | — | 33k | 2° src |
| Qwen3 32B Qwen/Qwen3-32B |
$0.08 | $0.25 | — | 41k | 2° src |
| Qwen3 8B Qwen/Qwen3-8B |
$0.07 | $0.18 | — | 41k | 2° src |