GLM-5.3 Flash z-ai/glm-5.3-flash
Served by 19 providers. Cheapest: Inference.net at $0.138 blended /1M tokens. The most expensive listing costs 1.7× as much for the same model.
| Provider | Type | Input /1M | Output /1M | Cached input | Blended | Context | Notes | Source |
|---|---|---|---|---|---|---|---|---|
| Inference.net glm-5.3-flash |
Inference hosts | $0.09 | $0.28 | $0.02 | $0.138 | 1.05M | https://inference.net/models shows the same figures |
✓ src 2026-09-18 |
| OpenRouter z-ai/glm-5.3-flash |
LLM routers | $0.09 | $0.30 | $0.018 | $0.143 | 1.05M | api modality text+image+video->text |
✓ src 2026-09-18 |
| Atlas Cloud zai-org/glm-5.3-flash |
Inference hosts | $0.15 | $0.50 | $0.03 | $0.237 | 1.05M | fp8 |
✓ src 2026-09-18 |
| Baseten | Inference hosts | $0.15 | $0.50 | $0.03 | $0.237 | — | Model API id not shown on pricing page. |
◐ src 2026-09-18 |
| Cloudflare Workers AI @cf/zai-org/glm-5.3-flash |
Inference hosts | $0.15 | $0.50 | $0.03 | $0.237 | — | Billed in neurons ($0.011/1k neurons); USD per-M-token equivalent as published by Cloudflare; requires Workers Paid or AI Gateway credits |
✓ src 2026-09-18 |
| CoreWeave zai-org/GLM-5.3-Flash |
Inference hosts | $0.15 | $0.50 | $0.05 | $0.237 | 1.05M | Context from docs model table (rounded, e.g. "262k"). |
✓ src src2 2026-09-18 |
| Crusoe | Inference hosts | $0.15 | $0.50 | $0.03 | $0.237 | — |
✓ src 2026-09-18 |
|
| DeepInfra zai-org/GLM-5.3-Flash |
Inference hosts | $0.15 | $0.50 | — | $0.237 | 1.05M | fp4 Price columns are DeepInfra's list price; its models API flags a 50% promotional discount on this model (no end date given), so billed price is lower. |
✓ src src2 2026-09-18 |
| DigitalOcean glm-5.3-flash |
Inference hosts | $0.15 | $0.50 | $0.03 | $0.237 | 1.05M |
✓ src src2 2026-09-18 |
|
| Featherless zai-org/GLM-5.3-Flash |
Inference hosts | $0.15 | $0.50 | $0.03 | $0.237 | 262k | Developer-plan per-token rate; model_class glm5-next-321b3; also usable flat-rate on $25/mo Chat plan (32K ctx, human chat only) |
✓ src src2 src3 2026-09-18 |
| Fireworks AI accounts/fireworks/models/glm-5p3-flash |
Inference hosts | $0.15 | $0.50 | $0.03 | $0.237 | — | batch −50% Priority $0.1875/$0.0375/$0.625 |
✓ src 2026-09-18 |
| FriendliAI zai-org/GLM-5.3-Flash |
Inference hosts | $0.15 | $0.50 | $0.03 | $0.237 | 1.05M | context from pricing page embedded model data |
✓ src 2026-09-18 |
| Novita AI zai-org/glm-5.3-flash |
Inference hosts | $0.15 | $0.50 | $0.03 | $0.237 | 1.05M |
✓ src src2 2026-09-18 |
|
| Parasail | Inference hosts | $0.15 | $0.50 | $0.03 | $0.237 | — |
✓ src 2026-09-18 |
|
| Requesty zai/glm-5.3-flash |
LLM routers | $0.15 | $0.50 | $0.03 | $0.237 | 1M | upstream route zai; geo sg; 7 other upstream routes listed (input $0.15-$0.2/M); API list price; Requesty's 5% markup on model cost (pricing page) is not included |
✓ src 2026-09-18 |
| SiliconFlow | Inference hosts | $0.15 | $0.50 | $0.03 | $0.237 | — | context shown on pricing page as "1049K" |
✓ src 2026-09-18 |
| Together AI zai-org/GLM-5.3-Flash |
Inference hosts | $0.15 | $0.50 | $0.03 | $0.237 | 1.05M | fp8 |
✓ src 2026-09-18 |
| Venice.ai API z-ai-glm-5-3-flash |
Inference hosts | $0.15 | $0.50 | $0.03 | $0.237 | 1.05M |
✓ src src2 2026-09-18 |
|
| Z.ai (Zhipu GLM API, international) glm-5.3-flash |
Model labs | $0.15 | $0.50 | $0.03 | $0.237 | 1M | Cached-input storage limited-time free. |
✓ src 2026-09-18 |
✓ from the provider's own pricing page · ◐ partly confirmed there · 2° third-party source. Router prices may exclude the router's own fees; see each router's page.