DevTk.AI
DeepSeek V4.1DeepSeek PricingAPI PricingGPT-6 LunaGLM-5.3 Flash

Is DeepSeek Still Good Value After the V4.1 Flash Price Cut?

DeepSeek V4.1 Flash lowered its September 2026 API prices. Compare peak, off-peak, and cached costs with GPT-6 Luna, GLM-5.3 Flash, Mistral Small 4, and Gemini 3.8 Flash.

DevTk.AI 2026-08-17 Updated 2026-10-02 4 min read

DeepSeek is still price-competitive, but the reason changed twice in six weeks. V4 Flash prices increased in August; on September 10, DeepSeek replaced it with V4.1 Flash, added vision support, and cut Flash rates to $0.30/M input and $1.20/M output at peak, with a 50% off-peak discount.

The short answer: V4.1 Flash is strong value for cache-heavy, long-context, visual, or very long-output work. It is not the cheapest route for every ordinary text request.

Current DeepSeek API Prices

USD per million tokens:

Model and periodCached inputCache-miss inputOutput
V4.1 Flash, off-peak$0.003$0.15$0.60
V4.1 Flash, peak$0.006$0.30$1.20
V4 Pro, off-peak$0.022$0.66$1.98
V4 Pro, peak$0.044$1.32$3.96

Use deepseek-flash for V4.1 Flash. The retired deepseek-v4-flash and deepseek-v4-flash-vision-exp names are still accepted, but requests route to V4.1 Flash and use current billing.

Peak hours are 01:00-04:00 and 06:00-10:00 UTC on weekdays, excluding Chinese public holidays. Weekends, holidays, and all remaining hours use the off-peak rate.

Same Workload, Current Bill

For 2M uncached input plus 500K output tokens across short-context requests:

ModelCost
GPT-6 Luna$0.45
GLM-5.3 Flash$0.55
Mistral Small 4$0.60
DeepSeek V4.1 Flash, off-peak$0.60
DeepSeek V4.1 Flash, peak$1.20
Gemini 3.8 Flash, 2026 introductory price$3.38
DeepSeek V4 Pro, off-peak$2.31
DeepSeek V4 Pro, peak$4.62

GPT-6 Luna has the lowest uncached bill in this token mix. DeepSeek ties Mistral Small 4 off-peak and costs more at peak, but token price alone hides differences in context, output limits, vision, tools, and task success rate.

Where DeepSeek Still Has an Edge

Cached long-context agents

V4.1 Flash cached input costs $0.003/M off-peak and $0.006/M at peak. With 85% of the 2M input above served from cache, the total becomes about $0.35 off-peak or $0.70 peak, including 500K output. That beats the uncached GPT-6 Luna example, although cache hits are best effort and must be measured.

Very long generated output

DeepSeek lists a 1M context window and up to 384K output. Many lower-priced alternatives cap output at 64K or 128K. Avoiding chunked continuation calls can matter more than a few cents per million tokens.

Visual input without a separate premium route

V4.1 Flash accepts images while keeping the Flash price. It can replace a separate vision preprocessor for some document, screenshot, and coding workflows.

Schedulable batch-like traffic

Off-peak halves input, cached-input, and output rates without requiring a separate batch endpoint. Queue evaluation, indexing, and non-urgent agent jobs outside the weekday peak windows when latency requirements allow it.

Where Another Model Can Be Better

  • Routine short text: GPT-6 Luna and GLM-5.3 Flash are cheaper in the normalized uncached example.
  • Stable fixed-rate budgeting: DeepSeek’s schedule creates a 2x range; GLM and OpenAI standard rates are easier to forecast.
  • Premium hard tasks: V4 Pro is not a budget model at peak. Compare it with GPT-6.1 Sol, GLM-5.3, and Claude Sonnet 5.5 on cost per successful task.
  • Google-grounded multimodal work: Gemini 3.8 Flash costs more but includes a broader Google tool and media-input ecosystem.
WorkloadStarting route
Simple high-volume textGPT-6 Luna or GLM-5.3 Flash
Cache-heavy 1M-context agentDeepSeek V4.1 Flash
Image-plus-text work needing long outputDeepSeek V4.1 Flash
Hard DeepSeek-compatible taskV4 Pro only after evaluation gates
Broad multimodal and Google-grounded taskGemini 3.8 Flash

DeepSeek still has real value, but the defensible claim is specific: V4.1 Flash is an unusually inexpensive cache-aware, multimodal, long-output route. It is not the universal unit-price winner.

Use the AI Pricing Calculator for the conservative peak estimate, and read the current DeepSeek API pricing guide for model IDs and schedules.

Official sources checked September 25, 2026: DeepSeek pricing, OpenAI pricing, Z.AI pricing, Mistral API pricing, and Gemini pricing.

Related Posts