GLM-5.3 FlashX
Z.AIUpdated October 2026. GLM-5.3 FlashX by Z.AI: $0.37/M cache-miss input, $1.25/M output tokens. Cached input: $0.075/M. 1.0M context, 128K max output. Function Calling & JSON Mode. Free calculator + compare 40+ models.
Input Price
$0.37
cache miss / 1M tokens
Cached Input
$0.075
per 1M tokens
Output Price
$1.25
per 1M tokens
Context Window
1.0M
tokens
Specifications
| Provider | Z.AI |
| Model ID | glm-5.3-flashx |
| Input Price | $0.37 / 1M cache-miss tokens |
| Cached Input Price | $0.075 / 1M tokens |
| Output Price | $1.25 / 1M tokens |
| Context Window | 1.0M tokens |
| Max Output | 128K tokens |
| Capabilities | textfunction_callingstructured_output |
| Release Date | 2026-09 |
| Pricing Source | Official Z.AI pricing |
| Price Verified | 2026-10-02 · GLM-5.3, GLM-5.3 Flash, and GLM-5.3 FlashX now have published pay-as-you-go API prices. |
| Notes | Faster GLM-5.3 API route. Cached-input storage is free for a limited time. |
Monthly Cost Estimates
Estimated monthly costs based on different daily usage levels (assuming 50% input / 50% output split). Input estimates use cache-miss pricing, so cache-heavy workloads can be lower.
| Daily Tokens | Monthly Cost | Annual Cost |
|---|---|---|
| 10K | $0.243 | $2.92 |
| 50K | $1.22 | $14.58 |
| 100K | $2.43 | $29.16 |
| 500K | $12.15 | $145.80 |
| 1.0M | $24.30 | $291.60 |
About GLM-5.3 FlashX
GLM-5.3 FlashX is a large language model by Z.AI. It features a 1.0M token context window with up to 128K tokens of output per request. The model supports 3 capabilities: text, function_calling, structured_output.
At $0.37 per million cache-miss input tokens and $1.25 per million output tokens, GLM-5.3 FlashX is positioned as a cost-effective option in the Z.AI lineup. Repeated prefix input can be charged at $0.075 per million cached tokens. Use our Token Counter to estimate how many tokens your prompts use, and our Pricing Calculator to compare costs across all models.
GLM-5.3 FlashX Key Details
- Pricing: $0.37/M cache-miss input tokens, $0.075/M cached input tokens, $1.25/M output tokens
- Context window: 1.0M tokens — one of the largest available
- Max output: 128K tokens per response
- Capabilities: text, function_calling, structured_output
- Highlights: Faster GLM-5.3 API route. Cached-input storage is free for a limited time.
- Released: 2026-09
Other Z.AI Models
Similar Price Range
Related Tools
FAQ
How much does GLM-5.3 FlashX cost?
GLM-5.3 FlashX costs $0.37 per million cache-miss input tokens and $1.25 per million output tokens. Cached input costs $0.075 per million tokens. For a typical workload of 100K input tokens/day and 50K output tokens/day, expect approximately $2.99/month before cache-hit savings.
What is GLM-5.3 FlashX's context window?
GLM-5.3 FlashX supports a context window of 1.0M tokens. This means your combined input prompt and output response can be up to 1.0M tokens. The maximum output per response is 128K tokens.
Is GLM-5.3 FlashX good for my use case?
GLM-5.3 FlashX supports text, function_calling, structured_output. As a budget-friendly model, it works well for high-volume tasks like classification, summarization, and simple generation. Use our Pricing Calculator to compare with alternatives.