AI API Pricing Comparison (August 2026): Current Models Side by Side
Updated August 2026 API prices for GPT-5.6, Claude 5, Gemini 3.6, DeepSeek V4-Flash-0731, Grok 4.5, Kimi K3, GLM-5.2, Mistral, MiniMax, Qwen, Cohere, and Amazon Nova.
The August 2026 price refresh changes the low-cost end of the API market. OpenAI cut GPT-5.6 Terra to $2/$12 and Luna to $0.20/$1.20 per million input/output tokens. Google released Gemini 3.6 Flash and Gemini 3.5 Flash-Lite, Anthropic added the Claude 5 lineup, and DeepSeek upgraded the existing Flash endpoint to V4-Flash-0731 without changing its base price.
All USD rows below use standard synchronous rates for cache-miss input and output. Cached-input prices are shown separately. Long-context, Batch, Flex, Fast, regional, and tool charges can change the final bill.
Current Flagship and Balanced API Prices
| Model | Provider | Input | Cached input | Output | Context | Important note |
|---|---|---|---|---|---|---|
| GPT-5.6 Sol | OpenAI | $5.00 | $0.50 | $30.00 | 1.05M | Long-context rates above 272K input |
| GPT-5.6 Terra | OpenAI | $2.00 | $0.20 | $12.00 | 1.05M | Reduced August price |
| Claude Fable 5 | Anthropic | $10.00 | $1.00 | $50.00 | 1M | Creative model; access restored July 1 |
| Claude Opus 5 | Anthropic | $5.00 | $0.50 | $25.00 | 1M | Current Opus flagship |
| Claude Sonnet 5 | Anthropic | $2.00 | $0.20 | $10.00 | 1M | Intro price through Aug 31 |
| Gemini 3.6 Flash | $1.50 | $0.15 | $7.50 | 1.05M | Current GA premium Flash | |
| Gemini 3.1 Pro Preview | $2.00 | $0.20 | $12.00 | 1.05M | $4/$18 above 200K input | |
| Grok 4.5 | xAI | $2.00 | $0.30 | $6.00 | 500K | $4/$0.60/$12 above 200K |
| Kimi K3 | Moonshot AI | $3.00 | $0.30 | $15.00 | 1.05M | Global API; full weights available |
| GLM-5.2 | Z.AI | $1.40 | $0.26 | $4.40 | 1M | Current open-weight GLM flagship |
| Mistral Medium 3.5 | Mistral | $1.50 | $0.15 | $7.50 | 131K | 90% cached-input discount |
| Command A+ | Cohere | $2.50 | - | $10.00 | 128K | Current Cohere flagship |
Current Lower-Cost API Prices
| Model | Provider | Input | Cached input | Output | Context | Important note |
|---|---|---|---|---|---|---|
| DeepSeek V4 Flash | DeepSeek | $0.14 | $0.0028 | $0.28 | 1M | Serves V4-Flash-0731 at the same price |
| Xiaomi MiMo-V2.5 | Xiaomi | $0.14 | $0.0028 | $0.28 | 1M | Multimodal; cache writes temporarily free |
| GPT-5.6 Luna | OpenAI | $0.20 | $0.02 | $1.20 | 1.05M | Reduced 80% from its launch price |
| Mistral Small 4 | Mistral | $0.15 | $0.015 | $0.60 | 131K | Current rate increased from $0.10/$0.30 |
| MiniMax M3 | MiniMax | $0.30 | $0.06 | $1.20 | 1M | $0.60/$2.40 above 512K input |
| Gemini 3.5 Flash-Lite | $0.30 | $0.03 | $2.50 | 1.05M | Current GA high-volume route | |
| Amazon Nova 2 Lite | Amazon | $0.30 | - | $2.50 | 1M | Bedrock price may vary by region |
| Kimi K2.7 Code | Moonshot AI | $0.95 | $0.19 | $4.00 | 262K | Global coding model |
| Kimi K2.6 | Moonshot AI | $0.95 | $0.16 | $4.00 | 262K | Global multimodal model |
| Claude Haiku 4.5 | Anthropic | $1.00 | $0.10 | $5.00 | 200K | Current low-cost Claude tier |
| Grok 4.3 | xAI | $1.25 | $0.20 | $2.50 | 1M | Long-context rate above 200K |
| Jamba Mini 2 | AI21 Labs | $0.20 | - | $0.40 | 256K | Current jamba-mini alias |
| Command R | Cohere | $0.50 | - | $1.50 | 128K | Corrected from stale $0.15/$0.60 data |
China API Prices in CNY
Do not mix these rows into a USD ranking without choosing an exchange rate and region.
| Model | Input | Cached input | Output | Context | Price status |
|---|---|---|---|---|---|
| Qwen3.7 Plus | ¥1.60 | - | ¥6.40 | 1M | Current 20% alias promotion; ¥4.80/¥19.20 above 256K |
| Qwen3.7 Max | ¥6.00 | - | ¥18.00 | 1M | Current 50% alias promotion |
| Kimi K2.5 | ¥4.00 | ¥0.70 | ¥21.00 | 262K | China API price |
Alibaba’s current pricing page does not list an end date for the Qwen3.7 alias promotions. DevTk.AI records both the current promotional rate and the list price in model notes.
Same Workload Cost Comparison
For 2M cache-miss input tokens plus 500K output tokens, using standard short-context prices:
| Model | Cost |
|---|---|
| DeepSeek V4 Flash | $0.42 |
| Xiaomi MiMo-V2.5 | $0.42 |
| Mistral Small 4 | $0.60 |
| GPT-5.6 Luna | $1.00 |
| MiniMax M3 | $1.20 |
| Gemini 3.5 Flash-Lite | $1.85 |
| Amazon Nova 2 Lite | $1.85 |
| GLM-5.2 | $5.00 |
| Grok 4.5 | $7.00 |
| Claude Sonnet 5 | $9.00 |
| GPT-5.6 Terra | $10.00 |
| Kimi K3 | $13.50 |
| Claude Opus 5 | $22.50 |
| GPT-5.6 Sol | $25.00 |
This table is a token-cost benchmark, not a quality ranking. Models can use different reasoning-token and output-token amounts for the same task. Tool calls, web search, storage, regional processing, and service tiers are additional.
What Changed This Month
- OpenAI reduced Terra by 20% and Luna by 80% on both input and output. The former Priority tier is now called Fast mode.
- DeepSeek V4-Flash-0731 is a checkpoint update under the same API ID and base price. A future 2x Beijing peak-hour multiplier is announced but not yet effective.
- Anthropic now lists Fable 5, Opus 5, and Sonnet 5. Sonnet’s $2/$10 rate ends August 31.
- Google launched Gemini 3.6 Flash at $1.50/$7.50 and Gemini 3.5 Flash-Lite at $0.30/$2.50.
- xAI’s current Grok 4.5 cached-input rate is $0.30/M, or $0.60/M above 200K context.
- Mistral Small 4 moved to $0.15/$0.60 and Mistral lists 90% off cached input.
Use the AI Model Pricing Calculator for your own traffic mix. For the market implications, read The August 2026 AI API Price War.
Official Sources
Related Posts
The August 2026 AI API Price War: OpenAI Cuts, DeepSeek V4-0731, and the New Cost Floor
2026-08-01
Kimi K3 vs GLM-5.2 vs DeepSeek V4: Price and Agent Routing
2026-07-19
Chinese AI Models in 2026: Kimi K3, GLM-5.2, DeepSeek V4, MiniMax, and Qwen
2026-06-14