DevTk.AI
AI API PricingGPT-6Claude Opus 5.5DeepSeek V4.1Gemini 3.8Grok 4.7

AI API Pricing Comparison (October 2026): Latest Models Side by Side

Updated October 2 API prices for GPT-6, Claude Opus 5.5, Gemini 3.8 Flash, Grok 4.7, DeepSeek V4.1, GLM-5.3, MiMo 2.6, Qwen 3.8, and Kimi K3.

DevTk.AI 2026-02-19 Updated 2026-10-02 5 min read

The October 2 refresh adds GPT-6, Claude Opus 5.5, Claude Fable 5.1, Gemini 3.8 Flash, Grok 4.7, DeepSeek V4.1 Flash, MiMo 2.6, and Qwen 3.8. It also replaces GLM-5.3’s pending status with published pay-as-you-go pricing and applies GPT-5.6 Sol’s current promotion.

This refresh adds GPT-6.1 Sol ($2/$10, $0.10/M cache reads), Claude Sonnet 5.5 ($2/$10), and MiMo-V2.6-Pro-UltraSpeed ($4.35/$8.70). Kimi K3 cache writes are $3/M for the default five-minute TTL or $6/M for one hour; MiniMax M3’s 50% reduction is permanent.

All USD rows below use standard synchronous cache-miss input and output rates. Long-context, Batch, Flex, Fast, regional, time-of-day, and tool charges can change the final bill.

Current Flagship and Balanced API Prices

ModelProviderInputCached inputOutputContextImportant note
GPT-6 AstraOpenAI$10.00$1.00$50.001.05M$20/$75 above 272K input
GPT-6.1 SolOpenAI$2.00$0.10$10.001.05MDemanding coding and agent tier
Claude Fable 5.1Anthropic$10.00$0.25$50.001MHighest broadly available Claude
Claude Opus 5.5Anthropic$4.00$0.20$20.001M5m cache write $5/M
Claude Sonnet 5.5Anthropic$2.00$0.20$10.001MSame token rates as Sonnet 5
Gemini 3.8 FlashGoogle$0.75$0.075$3.751.05MIntro price through Dec 31
Grok 4.7xAI$2.00$0.50$6.00500K$4/$1/$12 at or above 200K
Kimi K3Moonshot AI$3.00$0.30$15.001.05MGlobal API; full weights available
GLM-5.3Z.AI$1.40$0.26$4.401MPay-as-you-go now available
Xiaomi MiMo-V2.6-Pro-UltraSpeedXiaomi$4.35$0.036$8.701M10x Pro token rates; no Batch

Current Lower-Cost API Prices

ModelProviderInputCached inputOutputContextImportant note
GPT-6 LunaOpenAI$0.10$0.01$0.501.05MLowest-cost GPT-6 tier
DeepSeek V4.1 FlashDeepSeek$0.15-$0.30$0.003-$0.006$0.60-$1.201MOff-peak to weekday peak
GLM-5.3 FlashZ.AI$0.15$0.03$0.501MLowest-cost GLM-5.3 route
GLM-5.3 FlashXZ.AI$0.37$0.075$1.251MFaster GLM-5.3 route
Xiaomi MiMo-V2.6-FlashXiaomi$0.14$0.0028$0.281MBatch is half price
Xiaomi MiMo-V2.6-ProXiaomi$0.435$0.0036$0.871MFull-modal flagship
MiniMax M3MiniMax$0.30$0.06$1.201M$0.60/$2.40 above 512K
Mistral Small 4Mistral$0.15$0.015$0.60131K90% cached-input discount
Gemini 3.5 Flash-LiteGoogle$0.30$0.03$2.501.05MCurrent low-cost Gemini route

China API Prices in CNY

Do not mix CNY rows into a USD ranking without choosing an exchange rate and deployment region.

ModelInputCached inputOutputContext
Qwen3.8 Flash¥0.80¥0.10¥2.701M
Qwen3.8 Max¥12.00¥1.50¥36.001M
Qwen3.7 Plus¥1.60-¥6.401M
Kimi K2.5¥4.00¥0.70¥21.00262K

Same Workload Cost Comparison

For 2M cache-miss input plus 500K output tokens, using short-context prices:

ModelCost
Xiaomi MiMo-V2.6-Flash$0.42
GPT-6 Luna$0.45
GLM-5.3 Flash$0.55
DeepSeek V4.1 Flash, off-peak$0.60
DeepSeek V4.1 Flash, peak$1.20
Xiaomi MiMo-V2.6-Pro$1.31
Gemini 3.8 Flash, 2026 intro price$3.38
GLM-5.3$5.00
Grok 4.7$7.00
GPT-6.1 Sol$9.00
Claude Sonnet 5.5$9.00
Xiaomi MiMo-V2.6-Pro-UltraSpeed$13.05
Kimi K3$13.50
Claude Opus 5.5$18.00
GPT-6 Astra$45.00
Claude Fable 5.1$45.00

This is a token-cost benchmark, not a quality ranking. Reasoning-token volume, retries, cache hit rate, tool charges, latency, and human repair time determine cost per completed task.

Recent Changes Through October 2

  • OpenAI launched GPT-6 Astra, Sol, and Luna, then GPT-6.1 Sol on September 29 with $0.10/M cache reads; Luna is $0.10/$0.50.
  • Anthropic launched Opus 5.5 at $4/$20 and Fable 5.1 at $10/$50; Sonnet 5’s $2/$10 rate became permanent; Sonnet 5.5 launched on September 28 at the same rate.
  • Google launched Gemini 3.8 Flash at the same year-end introductory rate as 3.7 and 3.6.
  • xAI launched Grok 4.7 at the same token rates and context thresholds as 4.6.
  • DeepSeek V4.1 Flash replaced V4 Flash and reduced peak rates to $0.30/$1.20.
  • Z.AI published GLM-5.3, Flash, and FlashX pay-as-you-go prices.
  • Xiaomi launched MiMo 2.6 at unchanged token prices and scheduled V2.5 retirement for October 21.
  • Alibaba added Qwen3.8 Max and Flash to Model Studio.

Use the AI Model Pricing Calculator for your own traffic mix. For premium routing, compare GPT-6 pricing with Claude Opus 5.5 value.

Official Sources

Related Posts