DevTk.AI
AI API PricingGPT-5.6Claude 5DeepSeek V4Gemini 3.6Model Pricing

AI API Pricing Comparison (August 2026): Current Models Side by Side

Updated August 2026 API prices for GPT-5.6, Claude 5, Gemini 3.6, DeepSeek V4-Flash-0731, Grok 4.5, Kimi K3, GLM-5.2, Mistral, MiniMax, Qwen, Cohere, and Amazon Nova.

DevTk.AI 2026-02-19 Updated 2026-08-01 6 min read

The August 2026 price refresh changes the low-cost end of the API market. OpenAI cut GPT-5.6 Terra to $2/$12 and Luna to $0.20/$1.20 per million input/output tokens. Google released Gemini 3.6 Flash and Gemini 3.5 Flash-Lite, Anthropic added the Claude 5 lineup, and DeepSeek upgraded the existing Flash endpoint to V4-Flash-0731 without changing its base price.

All USD rows below use standard synchronous rates for cache-miss input and output. Cached-input prices are shown separately. Long-context, Batch, Flex, Fast, regional, and tool charges can change the final bill.

Current Flagship and Balanced API Prices

ModelProviderInputCached inputOutputContextImportant note
GPT-5.6 SolOpenAI$5.00$0.50$30.001.05MLong-context rates above 272K input
GPT-5.6 TerraOpenAI$2.00$0.20$12.001.05MReduced August price
Claude Fable 5Anthropic$10.00$1.00$50.001MCreative model; access restored July 1
Claude Opus 5Anthropic$5.00$0.50$25.001MCurrent Opus flagship
Claude Sonnet 5Anthropic$2.00$0.20$10.001MIntro price through Aug 31
Gemini 3.6 FlashGoogle$1.50$0.15$7.501.05MCurrent GA premium Flash
Gemini 3.1 Pro PreviewGoogle$2.00$0.20$12.001.05M$4/$18 above 200K input
Grok 4.5xAI$2.00$0.30$6.00500K$4/$0.60/$12 above 200K
Kimi K3Moonshot AI$3.00$0.30$15.001.05MGlobal API; full weights available
GLM-5.2Z.AI$1.40$0.26$4.401MCurrent open-weight GLM flagship
Mistral Medium 3.5Mistral$1.50$0.15$7.50131K90% cached-input discount
Command A+Cohere$2.50-$10.00128KCurrent Cohere flagship

Current Lower-Cost API Prices

ModelProviderInputCached inputOutputContextImportant note
DeepSeek V4 FlashDeepSeek$0.14$0.0028$0.281MServes V4-Flash-0731 at the same price
Xiaomi MiMo-V2.5Xiaomi$0.14$0.0028$0.281MMultimodal; cache writes temporarily free
GPT-5.6 LunaOpenAI$0.20$0.02$1.201.05MReduced 80% from its launch price
Mistral Small 4Mistral$0.15$0.015$0.60131KCurrent rate increased from $0.10/$0.30
MiniMax M3MiniMax$0.30$0.06$1.201M$0.60/$2.40 above 512K input
Gemini 3.5 Flash-LiteGoogle$0.30$0.03$2.501.05MCurrent GA high-volume route
Amazon Nova 2 LiteAmazon$0.30-$2.501MBedrock price may vary by region
Kimi K2.7 CodeMoonshot AI$0.95$0.19$4.00262KGlobal coding model
Kimi K2.6Moonshot AI$0.95$0.16$4.00262KGlobal multimodal model
Claude Haiku 4.5Anthropic$1.00$0.10$5.00200KCurrent low-cost Claude tier
Grok 4.3xAI$1.25$0.20$2.501MLong-context rate above 200K
Jamba Mini 2AI21 Labs$0.20-$0.40256KCurrent jamba-mini alias
Command RCohere$0.50-$1.50128KCorrected from stale $0.15/$0.60 data

China API Prices in CNY

Do not mix these rows into a USD ranking without choosing an exchange rate and region.

ModelInputCached inputOutputContextPrice status
Qwen3.7 Plus¥1.60-¥6.401MCurrent 20% alias promotion; ¥4.80/¥19.20 above 256K
Qwen3.7 Max¥6.00-¥18.001MCurrent 50% alias promotion
Kimi K2.5¥4.00¥0.70¥21.00262KChina API price

Alibaba’s current pricing page does not list an end date for the Qwen3.7 alias promotions. DevTk.AI records both the current promotional rate and the list price in model notes.

Same Workload Cost Comparison

For 2M cache-miss input tokens plus 500K output tokens, using standard short-context prices:

ModelCost
DeepSeek V4 Flash$0.42
Xiaomi MiMo-V2.5$0.42
Mistral Small 4$0.60
GPT-5.6 Luna$1.00
MiniMax M3$1.20
Gemini 3.5 Flash-Lite$1.85
Amazon Nova 2 Lite$1.85
GLM-5.2$5.00
Grok 4.5$7.00
Claude Sonnet 5$9.00
GPT-5.6 Terra$10.00
Kimi K3$13.50
Claude Opus 5$22.50
GPT-5.6 Sol$25.00

This table is a token-cost benchmark, not a quality ranking. Models can use different reasoning-token and output-token amounts for the same task. Tool calls, web search, storage, regional processing, and service tiers are additional.

What Changed This Month

  • OpenAI reduced Terra by 20% and Luna by 80% on both input and output. The former Priority tier is now called Fast mode.
  • DeepSeek V4-Flash-0731 is a checkpoint update under the same API ID and base price. A future 2x Beijing peak-hour multiplier is announced but not yet effective.
  • Anthropic now lists Fable 5, Opus 5, and Sonnet 5. Sonnet’s $2/$10 rate ends August 31.
  • Google launched Gemini 3.6 Flash at $1.50/$7.50 and Gemini 3.5 Flash-Lite at $0.30/$2.50.
  • xAI’s current Grok 4.5 cached-input rate is $0.30/M, or $0.60/M above 200K context.
  • Mistral Small 4 moved to $0.15/$0.60 and Mistral lists 90% off cached input.

Use the AI Model Pricing Calculator for your own traffic mix. For the market implications, read The August 2026 AI API Price War.

Official Sources

Related Posts