The August 2026 AI API Price War: Google Discounts as DeepSeek Raises Prices
Google halved Gemini 3.7/3.6 Flash through year-end while DeepSeek raised V4 rates and xAI launched Grok 4.6. Here is the new API price-war map.
The August API market is no longer a simple race to the bottom. OpenAI opened the month with major GPT-5.6 Terra and Luna reductions. Google then launched Gemini 3.7 Flash at a year-end introductory rate and applied the same discount to 3.6 Flash. DeepSeek moved the other way: its new peak/off-peak V4 price schedule took effect on August 17 and materially raised both discounted and maximum rates from the previous flat price.
The result is a price war built around temporary discounts, time-based billing, context thresholds, and task success, not one universal cheapest model.
What Changed
| Provider and model | Current change | Effective price |
|---|---|---|
| OpenAI GPT-5.6 Luna | 80% permanent reduction from launch | $0.20 input / $1.20 output |
| Google Gemini 3.7 Flash | New GA model with year-end introduction | $0.75 / $3.75 through Dec 31 |
| Google Gemini 3.6 Flash | Existing model receives same introduction | $0.75 / $3.75 through Dec 31 |
| DeepSeek V4 Flash | New time-of-day rates | $0.22/$0.66 off-peak; $0.44/$1.32 peak |
| xAI Grok 4.6 | New model; cache is higher than 4.5 | $2/$0.50 cached/$6 below 200K |
| Z.AI GLM-5.3 | New model, but pay-as-you-go price pending | Coding Plan available; API coming soon |
Prices are USD per million input/output tokens. DeepSeek rows show cache-miss input and output. Grok’s middle value is cached input.
Google Is Buying Year-End Adoption
Gemini 3.7 Flash and 3.6 Flash currently cost $0.75/M input, $0.075/M cached input, and $3.75/M output. Google has already scheduled those values to double on January 1, 2027.
This is an aggressive current price, especially because both models provide roughly 1M context, native text/image/video/audio/PDF input, and Batch/Flex support. The catch is forecast risk: a December run rate cannot be copied into a January budget.
DeepSeek Is Monetizing Time
DeepSeek’s old V4 Flash price of $0.14/M cache-miss input and $0.28/M output is gone. The new Flash price is $0.22/$0.66 off-peak and $0.44/$1.32 peak. Peak hours are 01:00-04:00 and 06:00-10:00 UTC.
The off-peak discount creates a new operational advantage for queued work, but it does not preserve the former price. Even off-peak output is more than twice the old output rate.
DeepSeek V4 Pro now ranges from $0.66/$1.98 off-peak to $1.32/$3.96 peak for cache-miss input/output. Pro has moved out of the default budget tier and should be justified by better completion rates.
xAI Holds Token Prices but Raises Cache
Grok 4.6 uses the same $2/M input and $6/M output as Grok 4.5 below 200K. Cached input rose from $0.30 to $0.50, a 66.7% increase. At 200K prompt tokens, all request tokens switch to $4 input, $1 cached input, and $12 output.
xAI’s strategy is capability-led rather than price-led. The migration case rests on xhigh reasoning and improved agent/coding results, while cache-heavy teams face a higher bill.
A Normalized Workload
For 2M cache-miss input plus 500K output tokens across short requests:
| Model | Cost |
|---|---|
| Xiaomi MiMo-V2.5 | $0.42 |
| Mistral Small 4 | $0.60 |
| DeepSeek V4 Flash, off-peak | $0.77 |
| GPT-5.6 Luna | $1.00 |
| MiniMax M3 | $1.20 |
| DeepSeek V4 Flash, peak | $1.54 |
| Gemini 3.5 Flash-Lite | $1.85 |
| Gemini 3.7 Flash, current introduction | $3.38 |
| GLM-5.2 | $5.00 |
| Grok 4.6 | $7.00 |
| Kimi K3 | $13.50 |
This is not a quality ranking. Reasoning-token volume, retries, tool charges, context limits, multimodal preprocessing, and human repair time determine cost per completed task.
Who Wins the Price War?
- Lowest ordinary token bill: MiMo-V2.5 and Mistral Small 4 in this example.
- Scheduled long-context text: DeepSeek V4 Flash remains strong off-peak.
- Low-cost OpenAI route: GPT-5.6 Luna is now competitive with peak DeepSeek traffic.
- Multimodal agents through year-end: Gemini 3.7 Flash combines low introductory pricing with broad inputs and Google tools.
- Capability-led escalation: Grok 4.6, Kimi K3, GLM-5.2, and premium GPT/Claude models must win on completion rather than unit cost.
What Developers Should Do
Store pricing with effective dates and conditions. DeepSeek needs a time-of-day field; Grok needs a 200K threshold; Gemini needs a January 1 scheduled price; temporary Claude or Qwen rates need expiry monitoring.
Then evaluate cost per successful task rather than cost per million tokens. A useful dashboard includes cache hit rate, reasoning/output tokens, retry count, p95 latency, tool fees, and human repair minutes.
Use the AI Pricing Calculator and current API price comparison for the live table. For DeepSeek specifically, read Is DeepSeek Still Good Value?.
Official sources checked August 17, 2026: Google pricing, DeepSeek pricing, xAI pricing, Z.AI pricing, and OpenAI pricing.
Related Posts
Is DeepSeek Still Good Value After the August 2026 Price Increase?
2026-08-17
AI API Pricing Comparison (August 2026): Latest Models Side by Side
2026-02-19
Grok 4.6 vs Gemini 3.7 Flash: API Price, Coding, and the 200K Cost Cliff
2026-08-17