AI API Pricing Comparison (October 2026): Latest Models Side by Side
Updated October 2 API prices for GPT-6, Claude Opus 5.5, Gemini 3.8 Flash, Grok 4.7, DeepSeek V4.1, GLM-5.3, MiMo 2.6, Qwen 3.8, and Kimi K3.
The October 2 refresh adds GPT-6, Claude Opus 5.5, Claude Fable 5.1, Gemini 3.8 Flash, Grok 4.7, DeepSeek V4.1 Flash, MiMo 2.6, and Qwen 3.8. It also replaces GLM-5.3’s pending status with published pay-as-you-go pricing and applies GPT-5.6 Sol’s current promotion.
This refresh adds GPT-6.1 Sol ($2/$10, $0.10/M cache reads), Claude Sonnet 5.5 ($2/$10), and MiMo-V2.6-Pro-UltraSpeed ($4.35/$8.70). Kimi K3 cache writes are $3/M for the default five-minute TTL or $6/M for one hour; MiniMax M3’s 50% reduction is permanent.
All USD rows below use standard synchronous cache-miss input and output rates. Long-context, Batch, Flex, Fast, regional, time-of-day, and tool charges can change the final bill.
Current Flagship and Balanced API Prices
| Model | Provider | Input | Cached input | Output | Context | Important note |
|---|---|---|---|---|---|---|
| GPT-6 Astra | OpenAI | $10.00 | $1.00 | $50.00 | 1.05M | $20/$75 above 272K input |
| GPT-6.1 Sol | OpenAI | $2.00 | $0.10 | $10.00 | 1.05M | Demanding coding and agent tier |
| Claude Fable 5.1 | Anthropic | $10.00 | $0.25 | $50.00 | 1M | Highest broadly available Claude |
| Claude Opus 5.5 | Anthropic | $4.00 | $0.20 | $20.00 | 1M | 5m cache write $5/M |
| Claude Sonnet 5.5 | Anthropic | $2.00 | $0.20 | $10.00 | 1M | Same token rates as Sonnet 5 |
| Gemini 3.8 Flash | $0.75 | $0.075 | $3.75 | 1.05M | Intro price through Dec 31 | |
| Grok 4.7 | xAI | $2.00 | $0.50 | $6.00 | 500K | $4/$1/$12 at or above 200K |
| Kimi K3 | Moonshot AI | $3.00 | $0.30 | $15.00 | 1.05M | Global API; full weights available |
| GLM-5.3 | Z.AI | $1.40 | $0.26 | $4.40 | 1M | Pay-as-you-go now available |
| Xiaomi MiMo-V2.6-Pro-UltraSpeed | Xiaomi | $4.35 | $0.036 | $8.70 | 1M | 10x Pro token rates; no Batch |
Current Lower-Cost API Prices
| Model | Provider | Input | Cached input | Output | Context | Important note |
|---|---|---|---|---|---|---|
| GPT-6 Luna | OpenAI | $0.10 | $0.01 | $0.50 | 1.05M | Lowest-cost GPT-6 tier |
| DeepSeek V4.1 Flash | DeepSeek | $0.15-$0.30 | $0.003-$0.006 | $0.60-$1.20 | 1M | Off-peak to weekday peak |
| GLM-5.3 Flash | Z.AI | $0.15 | $0.03 | $0.50 | 1M | Lowest-cost GLM-5.3 route |
| GLM-5.3 FlashX | Z.AI | $0.37 | $0.075 | $1.25 | 1M | Faster GLM-5.3 route |
| Xiaomi MiMo-V2.6-Flash | Xiaomi | $0.14 | $0.0028 | $0.28 | 1M | Batch is half price |
| Xiaomi MiMo-V2.6-Pro | Xiaomi | $0.435 | $0.0036 | $0.87 | 1M | Full-modal flagship |
| MiniMax M3 | MiniMax | $0.30 | $0.06 | $1.20 | 1M | $0.60/$2.40 above 512K |
| Mistral Small 4 | Mistral | $0.15 | $0.015 | $0.60 | 131K | 90% cached-input discount |
| Gemini 3.5 Flash-Lite | $0.30 | $0.03 | $2.50 | 1.05M | Current low-cost Gemini route |
China API Prices in CNY
Do not mix CNY rows into a USD ranking without choosing an exchange rate and deployment region.
| Model | Input | Cached input | Output | Context |
|---|---|---|---|---|
| Qwen3.8 Flash | ¥0.80 | ¥0.10 | ¥2.70 | 1M |
| Qwen3.8 Max | ¥12.00 | ¥1.50 | ¥36.00 | 1M |
| Qwen3.7 Plus | ¥1.60 | - | ¥6.40 | 1M |
| Kimi K2.5 | ¥4.00 | ¥0.70 | ¥21.00 | 262K |
Same Workload Cost Comparison
For 2M cache-miss input plus 500K output tokens, using short-context prices:
| Model | Cost |
|---|---|
| Xiaomi MiMo-V2.6-Flash | $0.42 |
| GPT-6 Luna | $0.45 |
| GLM-5.3 Flash | $0.55 |
| DeepSeek V4.1 Flash, off-peak | $0.60 |
| DeepSeek V4.1 Flash, peak | $1.20 |
| Xiaomi MiMo-V2.6-Pro | $1.31 |
| Gemini 3.8 Flash, 2026 intro price | $3.38 |
| GLM-5.3 | $5.00 |
| Grok 4.7 | $7.00 |
| GPT-6.1 Sol | $9.00 |
| Claude Sonnet 5.5 | $9.00 |
| Xiaomi MiMo-V2.6-Pro-UltraSpeed | $13.05 |
| Kimi K3 | $13.50 |
| Claude Opus 5.5 | $18.00 |
| GPT-6 Astra | $45.00 |
| Claude Fable 5.1 | $45.00 |
This is a token-cost benchmark, not a quality ranking. Reasoning-token volume, retries, cache hit rate, tool charges, latency, and human repair time determine cost per completed task.
Recent Changes Through October 2
- OpenAI launched GPT-6 Astra, Sol, and Luna, then GPT-6.1 Sol on September 29 with $0.10/M cache reads; Luna is $0.10/$0.50.
- Anthropic launched Opus 5.5 at $4/$20 and Fable 5.1 at $10/$50; Sonnet 5’s $2/$10 rate became permanent; Sonnet 5.5 launched on September 28 at the same rate.
- Google launched Gemini 3.8 Flash at the same year-end introductory rate as 3.7 and 3.6.
- xAI launched Grok 4.7 at the same token rates and context thresholds as 4.6.
- DeepSeek V4.1 Flash replaced V4 Flash and reduced peak rates to $0.30/$1.20.
- Z.AI published GLM-5.3, Flash, and FlashX pay-as-you-go prices.
- Xiaomi launched MiMo 2.6 at unchanged token prices and scheduled V2.5 retirement for October 21.
- Alibaba added Qwen3.8 Max and Flash to Model Studio.
Use the AI Model Pricing Calculator for your own traffic mix. For premium routing, compare GPT-6 pricing with Claude Opus 5.5 value.