DevTk.AI
DeepSeek APIDeepSeek PricingDeepSeek V4.1Prompt CachingLLM Pricing

DeepSeek API Pricing 2026: V4.1 Flash and Pro Peak vs Off-Peak Rates

Current DeepSeek V4.1 Flash and V4 Pro API prices, weekday peak and off-peak token rates, cache costs, model IDs, vision support, and cost examples.

DevTk.AI 2026-02-23 Updated 2026-09-25 4 min read

DeepSeek replaced V4 Flash with DeepSeek-V4.1-Flash on September 10, 2026 and reduced its token prices. The API still uses time-of-day pricing. Peak hours are 01:00-04:00 and 06:00-10:00 UTC on weekdays, excluding Chinese public holidays; every other hour receives a 50% off-peak discount.

The current model IDs are deepseek-flash and deepseek-v4-pro. Both have a 1M-token context window, 384K maximum output, thinking and non-thinking modes, JSON output, tool calls, Responses API, and OpenAI- and Anthropic-compatible endpoints. V4.1 Flash also adds image input; V4 Pro remains text-only.

Current DeepSeek Prices

USD per million tokens:

Model and billing periodCached inputCache-miss inputOutput
deepseek-flash (V4.1 Flash), off-peak$0.003$0.15$0.60
deepseek-flash (V4.1 Flash), peak$0.006$0.30$1.20
deepseek-v4-pro, off-peak$0.022$0.66$1.98
deepseek-v4-pro, peak$0.044$1.32$3.96

DevTk.AI’s calculator uses the conservative peak rate. For workloads that can be scheduled, calculate a second forecast with the off-peak rate.

Peak Schedule

Billing periodUTCBeijing time
Peak window 101:00-04:00 weekdays09:00-12:00 weekdays
Peak window 206:00-10:00 weekdays14:00-18:00 weekdays
Off-peakAll other hours and holidaysAll other hours and Chinese public holidays

The 2x statement compares the current peak and off-peak prices. The September V4.1 Flash release lowered both tiers from the August V4 Flash rates: peak cache-miss input fell from $0.44 to $0.30 and peak output from $1.32 to $1.20.

Current Versions and IDs

API IDServed versionContextMax output
deepseek-flashDeepSeek-V4.1-Flash1M384K
deepseek-v4-proDeepSeek-V4-Pro-08131M384K

Legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp are still accepted, but their retired models now route to V4.1 Flash and are billed at the current Flash price. Use deepseek-flash for new integrations.

Cost Formula and Cache Fields

cost =
  prompt_cache_hit_tokens * cached_input_price / 1,000,000
+ prompt_cache_miss_tokens * cache_miss_input_price / 1,000,000
+ completion_tokens * output_price / 1,000,000

DeepSeek exposes the relevant usage counters:

usage.prompt_cache_hit_tokens
usage.prompt_cache_miss_tokens

Caching is best effort. Stable system prompts, repository rules, tool schemas, and conversation prefixes have the best chance of reuse. Monitor the actual cache-hit ratio rather than budgeting for a perfect cache.

Cost Examples

For 2M cache-miss input plus 500K output tokens across multiple requests:

Model and periodCost
V4.1 Flash, off-peak$0.60
V4.1 Flash, peak$1.20
V4 Pro, off-peak$2.31
V4 Pro, peak$4.62

With 85% of input served from cache, the same workload becomes approximately $0.350 off-peak or $0.700 peak on V4.1 Flash. The result still depends on actual cache behavior and output-token use.

OpenAI-Compatible Example

from openai import OpenAI

client = OpenAI(
    api_key="your-deepseek-api-key",
    base_url="https://api.deepseek.com"
)

response = client.chat.completions.create(
    model="deepseek-flash",
    messages=[
        {"role": "system", "content": "You are a careful coding assistant."},
        {"role": "user", "content": "Explain this repository structure."}
    ],
)

print(response.choices[0].message.content)

For Anthropic-format coding agents:

export ANTHROPIC_BASE_URL=https://api.deepseek.com/anthropic
export ANTHROPIC_AUTH_TOKEN=your-deepseek-api-key

Is DeepSeek Still Cheap?

V4.1 Flash is again one of the lowest-priced managed long-context models, especially off-peak and with repeated prefixes. For the 2M-input/500K-output example, it costs $0.60 off-peak and $1.20 at peak. GPT-6 Luna is still cheaper at $0.45 for the same short-context token mix, while V4.1 Flash offers vision, a 1M context window, and up to 384K output.

V4 Pro should be an escalation route. Its value depends on reducing failed tool loops, retries, or human repair enough to offset the higher price.

Read the updated DeepSeek value analysis and use the AI Pricing Calculator for your workload.

Official sources checked September 25, 2026: DeepSeek pricing, DeepSeek changelog, context caching, and Anthropic API compatibility.

Related Posts