DevTk.AI

GLM-5.3 Flash

Z.AI

Updated October 2026. GLM-5.3 Flash by Z.AI: $0.15/M cache-miss input, $0.5/M output tokens. Cached input: $0.03/M. 1.0M context, 128K max output. Function Calling & JSON Mode. Free calculator + compare 40+ models.

Input Price

$0.15

cache miss / 1M tokens

Cached Input

$0.03

per 1M tokens

Output Price

$0.5

per 1M tokens

Context Window

1.0M

tokens

Specifications

ProviderZ.AI
Model IDglm-5.3-flash
Input Price$0.15 / 1M cache-miss tokens
Cached Input Price$0.03 / 1M tokens
Output Price$0.5 / 1M tokens
Context Window1.0M tokens
Max Output128K tokens
Capabilities
textfunction_callingstructured_output
Release Date2026-09
Pricing SourceOfficial Z.AI pricing
Price Verified2026-10-02 · GLM-5.3, GLM-5.3 Flash, and GLM-5.3 FlashX now have published pay-as-you-go API prices.
NotesLowest-cost GLM-5.3 route. Cached-input storage is free for a limited time.

Monthly Cost Estimates

Estimated monthly costs based on different daily usage levels (assuming 50% input / 50% output split). Input estimates use cache-miss pricing, so cache-heavy workloads can be lower.

Daily TokensMonthly CostAnnual Cost
10K $0.0975 $1.17
50K $0.4875 $5.85
100K $0.975 $11.70
500K $4.88 $58.50
1.0M $9.75 $117.00

About GLM-5.3 Flash

GLM-5.3 Flash is a large language model by Z.AI. It features a 1.0M token context window with up to 128K tokens of output per request. The model supports 3 capabilities: text, function_calling, structured_output.

At $0.15 per million cache-miss input tokens and $0.5 per million output tokens, GLM-5.3 Flash is positioned as a cost-effective option in the Z.AI lineup. Repeated prefix input can be charged at $0.03 per million cached tokens. Use our Token Counter to estimate how many tokens your prompts use, and our Pricing Calculator to compare costs across all models.

GLM-5.3 Flash Key Details

  • Pricing: $0.15/M cache-miss input tokens, $0.03/M cached input tokens, $0.5/M output tokens
  • Context window: 1.0M tokens — one of the largest available
  • Max output: 128K tokens per response
  • Capabilities: text, function_calling, structured_output
  • Highlights: Lowest-cost GLM-5.3 route. Cached-input storage is free for a limited time.
  • Released: 2026-09

Other Z.AI Models

Similar Price Range

Related Tools

FAQ

How much does GLM-5.3 Flash cost?

GLM-5.3 Flash costs $0.15 per million cache-miss input tokens and $0.5 per million output tokens. Cached input costs $0.03 per million tokens. For a typical workload of 100K input tokens/day and 50K output tokens/day, expect approximately $1.20/month before cache-hit savings.

What is GLM-5.3 Flash's context window?

GLM-5.3 Flash supports a context window of 1.0M tokens. This means your combined input prompt and output response can be up to 1.0M tokens. The maximum output per response is 128K tokens.

Is GLM-5.3 Flash good for my use case?

GLM-5.3 Flash supports text, function_calling, structured_output. As a budget-friendly model, it works well for high-volume tasks like classification, summarization, and simple generation. Use our Pricing Calculator to compare with alternatives.