DeepSeek API Pricing 2026: V4.1 Flash and Pro Peak vs Off-Peak Rates
Current DeepSeek V4.1 Flash and V4 Pro API prices, weekday peak and off-peak token rates, cache costs, model IDs, vision support, and cost examples.
DeepSeek replaced V4 Flash with DeepSeek-V4.1-Flash on September 10, 2026 and reduced its token prices. The API still uses time-of-day pricing. Peak hours are 01:00-04:00 and 06:00-10:00 UTC on weekdays, excluding Chinese public holidays; every other hour receives a 50% off-peak discount.
The current model IDs are deepseek-flash and deepseek-v4-pro. Both have a 1M-token context window, 384K maximum output, thinking and non-thinking modes, JSON output, tool calls, Responses API, and OpenAI- and Anthropic-compatible endpoints. V4.1 Flash also adds image input; V4 Pro remains text-only.
Current DeepSeek Prices
USD per million tokens:
| Model and billing period | Cached input | Cache-miss input | Output |
|---|---|---|---|
deepseek-flash (V4.1 Flash), off-peak | $0.003 | $0.15 | $0.60 |
deepseek-flash (V4.1 Flash), peak | $0.006 | $0.30 | $1.20 |
deepseek-v4-pro, off-peak | $0.022 | $0.66 | $1.98 |
deepseek-v4-pro, peak | $0.044 | $1.32 | $3.96 |
DevTk.AI’s calculator uses the conservative peak rate. For workloads that can be scheduled, calculate a second forecast with the off-peak rate.
Peak Schedule
| Billing period | UTC | Beijing time |
|---|---|---|
| Peak window 1 | 01:00-04:00 weekdays | 09:00-12:00 weekdays |
| Peak window 2 | 06:00-10:00 weekdays | 14:00-18:00 weekdays |
| Off-peak | All other hours and holidays | All other hours and Chinese public holidays |
The 2x statement compares the current peak and off-peak prices. The September V4.1 Flash release lowered both tiers from the August V4 Flash rates: peak cache-miss input fell from $0.44 to $0.30 and peak output from $1.32 to $1.20.
Current Versions and IDs
| API ID | Served version | Context | Max output |
|---|---|---|---|
deepseek-flash | DeepSeek-V4.1-Flash | 1M | 384K |
deepseek-v4-pro | DeepSeek-V4-Pro-0813 | 1M | 384K |
Legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp are still accepted, but their retired models now route to V4.1 Flash and are billed at the current Flash price. Use deepseek-flash for new integrations.
Cost Formula and Cache Fields
cost =
prompt_cache_hit_tokens * cached_input_price / 1,000,000
+ prompt_cache_miss_tokens * cache_miss_input_price / 1,000,000
+ completion_tokens * output_price / 1,000,000
DeepSeek exposes the relevant usage counters:
usage.prompt_cache_hit_tokens
usage.prompt_cache_miss_tokens
Caching is best effort. Stable system prompts, repository rules, tool schemas, and conversation prefixes have the best chance of reuse. Monitor the actual cache-hit ratio rather than budgeting for a perfect cache.
Cost Examples
For 2M cache-miss input plus 500K output tokens across multiple requests:
| Model and period | Cost |
|---|---|
| V4.1 Flash, off-peak | $0.60 |
| V4.1 Flash, peak | $1.20 |
| V4 Pro, off-peak | $2.31 |
| V4 Pro, peak | $4.62 |
With 85% of input served from cache, the same workload becomes approximately $0.350 off-peak or $0.700 peak on V4.1 Flash. The result still depends on actual cache behavior and output-token use.
OpenAI-Compatible Example
from openai import OpenAI
client = OpenAI(
api_key="your-deepseek-api-key",
base_url="https://api.deepseek.com"
)
response = client.chat.completions.create(
model="deepseek-flash",
messages=[
{"role": "system", "content": "You are a careful coding assistant."},
{"role": "user", "content": "Explain this repository structure."}
],
)
print(response.choices[0].message.content)
For Anthropic-format coding agents:
export ANTHROPIC_BASE_URL=https://api.deepseek.com/anthropic
export ANTHROPIC_AUTH_TOKEN=your-deepseek-api-key
Is DeepSeek Still Cheap?
V4.1 Flash is again one of the lowest-priced managed long-context models, especially off-peak and with repeated prefixes. For the 2M-input/500K-output example, it costs $0.60 off-peak and $1.20 at peak. GPT-6 Luna is still cheaper at $0.45 for the same short-context token mix, while V4.1 Flash offers vision, a 1M context window, and up to 384K output.
V4 Pro should be an escalation route. Its value depends on reducing failed tool loops, retries, or human repair enough to offset the higher price.
Read the updated DeepSeek value analysis and use the AI Pricing Calculator for your workload.
Official sources checked September 25, 2026: DeepSeek pricing, DeepSeek changelog, context caching, and Anthropic API compatibility.