DevTk.AI
AI API Price WarDeepSeek Price IncreaseGemini 3.7Grok 4.6GPT-5.6 LunaAPI Pricing

The August 2026 AI API Price War: Google Discounts as DeepSeek Raises Prices

Google halved Gemini 3.7/3.6 Flash through year-end while DeepSeek raised V4 rates and xAI launched Grok 4.6. Here is the new API price-war map.

DevTk.AI 2026-08-01 Updated 2026-08-17 4 min read

The August API market is no longer a simple race to the bottom. OpenAI opened the month with major GPT-5.6 Terra and Luna reductions. Google then launched Gemini 3.7 Flash at a year-end introductory rate and applied the same discount to 3.6 Flash. DeepSeek moved the other way: its new peak/off-peak V4 price schedule took effect on August 17 and materially raised both discounted and maximum rates from the previous flat price.

The result is a price war built around temporary discounts, time-based billing, context thresholds, and task success, not one universal cheapest model.

What Changed

Provider and modelCurrent changeEffective price
OpenAI GPT-5.6 Luna80% permanent reduction from launch$0.20 input / $1.20 output
Google Gemini 3.7 FlashNew GA model with year-end introduction$0.75 / $3.75 through Dec 31
Google Gemini 3.6 FlashExisting model receives same introduction$0.75 / $3.75 through Dec 31
DeepSeek V4 FlashNew time-of-day rates$0.22/$0.66 off-peak; $0.44/$1.32 peak
xAI Grok 4.6New model; cache is higher than 4.5$2/$0.50 cached/$6 below 200K
Z.AI GLM-5.3New model, but pay-as-you-go price pendingCoding Plan available; API coming soon

Prices are USD per million input/output tokens. DeepSeek rows show cache-miss input and output. Grok’s middle value is cached input.

Google Is Buying Year-End Adoption

Gemini 3.7 Flash and 3.6 Flash currently cost $0.75/M input, $0.075/M cached input, and $3.75/M output. Google has already scheduled those values to double on January 1, 2027.

This is an aggressive current price, especially because both models provide roughly 1M context, native text/image/video/audio/PDF input, and Batch/Flex support. The catch is forecast risk: a December run rate cannot be copied into a January budget.

DeepSeek Is Monetizing Time

DeepSeek’s old V4 Flash price of $0.14/M cache-miss input and $0.28/M output is gone. The new Flash price is $0.22/$0.66 off-peak and $0.44/$1.32 peak. Peak hours are 01:00-04:00 and 06:00-10:00 UTC.

The off-peak discount creates a new operational advantage for queued work, but it does not preserve the former price. Even off-peak output is more than twice the old output rate.

DeepSeek V4 Pro now ranges from $0.66/$1.98 off-peak to $1.32/$3.96 peak for cache-miss input/output. Pro has moved out of the default budget tier and should be justified by better completion rates.

xAI Holds Token Prices but Raises Cache

Grok 4.6 uses the same $2/M input and $6/M output as Grok 4.5 below 200K. Cached input rose from $0.30 to $0.50, a 66.7% increase. At 200K prompt tokens, all request tokens switch to $4 input, $1 cached input, and $12 output.

xAI’s strategy is capability-led rather than price-led. The migration case rests on xhigh reasoning and improved agent/coding results, while cache-heavy teams face a higher bill.

A Normalized Workload

For 2M cache-miss input plus 500K output tokens across short requests:

ModelCost
Xiaomi MiMo-V2.5$0.42
Mistral Small 4$0.60
DeepSeek V4 Flash, off-peak$0.77
GPT-5.6 Luna$1.00
MiniMax M3$1.20
DeepSeek V4 Flash, peak$1.54
Gemini 3.5 Flash-Lite$1.85
Gemini 3.7 Flash, current introduction$3.38
GLM-5.2$5.00
Grok 4.6$7.00
Kimi K3$13.50

This is not a quality ranking. Reasoning-token volume, retries, tool charges, context limits, multimodal preprocessing, and human repair time determine cost per completed task.

Who Wins the Price War?

  • Lowest ordinary token bill: MiMo-V2.5 and Mistral Small 4 in this example.
  • Scheduled long-context text: DeepSeek V4 Flash remains strong off-peak.
  • Low-cost OpenAI route: GPT-5.6 Luna is now competitive with peak DeepSeek traffic.
  • Multimodal agents through year-end: Gemini 3.7 Flash combines low introductory pricing with broad inputs and Google tools.
  • Capability-led escalation: Grok 4.6, Kimi K3, GLM-5.2, and premium GPT/Claude models must win on completion rather than unit cost.

What Developers Should Do

Store pricing with effective dates and conditions. DeepSeek needs a time-of-day field; Grok needs a 200K threshold; Gemini needs a January 1 scheduled price; temporary Claude or Qwen rates need expiry monitoring.

Then evaluate cost per successful task rather than cost per million tokens. A useful dashboard includes cache hit rate, reasoning/output tokens, retry count, p95 latency, tool fees, and human repair minutes.

Use the AI Pricing Calculator and current API price comparison for the live table. For DeepSeek specifically, read Is DeepSeek Still Good Value?.

Official sources checked August 17, 2026: Google pricing, DeepSeek pricing, xAI pricing, Z.AI pricing, and OpenAI pricing.

Related Posts