DevTk.AI

GLM-5.3 FlashX

Z.AI

Updated October 2026. GLM-5.3 FlashX by Z.AI: $0.37/M cache-miss input, $1.25/M output tokens. Cached input: $0.075/M. 1.0M context, 128K max output. Function Calling & JSON Mode. Free calculator + compare 40+ models.

Input Price

$0.37

cache miss / 1M tokens

Cached Input

$0.075

per 1M tokens

Output Price

$1.25

per 1M tokens

Context Window

1.0M

tokens

Specifications

ProviderZ.AI
Model IDglm-5.3-flashx
Input Price$0.37 / 1M cache-miss tokens
Cached Input Price$0.075 / 1M tokens
Output Price$1.25 / 1M tokens
Context Window1.0M tokens
Max Output128K tokens
Capabilities
textfunction_callingstructured_output
Release Date2026-09
Pricing SourceOfficial Z.AI pricing
Price Verified2026-10-02 · GLM-5.3, GLM-5.3 Flash, and GLM-5.3 FlashX now have published pay-as-you-go API prices.
NotesFaster GLM-5.3 API route. Cached-input storage is free for a limited time.

Monthly Cost Estimates

Estimated monthly costs based on different daily usage levels (assuming 50% input / 50% output split). Input estimates use cache-miss pricing, so cache-heavy workloads can be lower.

Daily TokensMonthly CostAnnual Cost
10K $0.243 $2.92
50K $1.22 $14.58
100K $2.43 $29.16
500K $12.15 $145.80
1.0M $24.30 $291.60

About GLM-5.3 FlashX

GLM-5.3 FlashX is a large language model by Z.AI. It features a 1.0M token context window with up to 128K tokens of output per request. The model supports 3 capabilities: text, function_calling, structured_output.

At $0.37 per million cache-miss input tokens and $1.25 per million output tokens, GLM-5.3 FlashX is positioned as a cost-effective option in the Z.AI lineup. Repeated prefix input can be charged at $0.075 per million cached tokens. Use our Token Counter to estimate how many tokens your prompts use, and our Pricing Calculator to compare costs across all models.

GLM-5.3 FlashX Key Details

  • Pricing: $0.37/M cache-miss input tokens, $0.075/M cached input tokens, $1.25/M output tokens
  • Context window: 1.0M tokens — one of the largest available
  • Max output: 128K tokens per response
  • Capabilities: text, function_calling, structured_output
  • Highlights: Faster GLM-5.3 API route. Cached-input storage is free for a limited time.
  • Released: 2026-09

Other Z.AI Models

Similar Price Range

Related Tools

FAQ

How much does GLM-5.3 FlashX cost?

GLM-5.3 FlashX costs $0.37 per million cache-miss input tokens and $1.25 per million output tokens. Cached input costs $0.075 per million tokens. For a typical workload of 100K input tokens/day and 50K output tokens/day, expect approximately $2.99/month before cache-hit savings.

What is GLM-5.3 FlashX's context window?

GLM-5.3 FlashX supports a context window of 1.0M tokens. This means your combined input prompt and output response can be up to 1.0M tokens. The maximum output per response is 128K tokens.

Is GLM-5.3 FlashX good for my use case?

GLM-5.3 FlashX supports text, function_calling, structured_output. As a budget-friendly model, it works well for high-volume tasks like classification, summarization, and simple generation. Use our Pricing Calculator to compare with alternatives.