AI Model Pricing Calculator
Compare costs across 40+ AI models, including DeepSeek V4 cache-hit pricing. Estimate your monthly spend instantly.
Workload presets
Cache-hit share is applied only to models with published cached-input pricing.
100,000 tokens/day = 3,000,000 tokens/month
Activates provider long-context rates at published thresholds
50,000 tokens/day = 1,500,000 tokens/month
0% of input tokens use cached pricing when available
| Model↑↓ | Provider | Input $/1M↑↓ | Output $/1M↑↓ | Context Window | Monthly Cost↑ |
|---|---|---|---|---|---|
Amazon Nova MicroCheapest | Amazon | $0.035 | $0.14 | 128,000 | $0.3150 |
Amazon Nova Lite | Amazon | $0.06 | $0.24 | 300,000 | $0.5400 |
GPT-5 Nano | OpenAI | $0.05 cached $0.005 | $0.4 | 400,000 | $0.7500 |
Xiaomi MiMo-V2.5 | Xiaomi MiMo | $0.14 cached $0.0028 | $0.28 | 1,000,000 | $0.8400 |
Gemini 2.5 Flash-Lite | $0.1 cached $0.01 | $0.4 | 1,000,000 | $0.9000 | |
Jamba Mini 2 | AI21 Labs | $0.2 | $0.4 | 256,000 | $1.20 |
GPT-4o mini | OpenAI | $0.15 cached $0.075 | $0.6 | 128,000 | $1.35 |
Mistral Small 4 | Mistral | $0.15 cached $0.015 | $0.6 | 131,072 | $1.35 |
GPT-5.6 Luna | OpenAI | $0.2 cached $0.02 | $1.20 | 1,050,000 | $2.40 |
GPT-5.4 Nano | OpenAI | $0.2 cached $0.02 | $1.25 | 400,000 | $2.48 |
Xiaomi MiMo-V2.5-Pro | Xiaomi MiMo | $0.435 cached $0.0036 | $0.87 | 1,000,000 | $2.61 |
MiniMax M3 | MiniMax | $0.3 cached $0.06 | $1.20 | 1,000,000 | $2.70 |
Gemini 3.1 Flash-Lite | $0.25 cached $0.025 | $1.50 | 1,000,000 | $3.00 | |
DeepSeek V4 Flash Peak Off-peak: $0.22 input / $0.007 cached / $0.66 output | DeepSeek | $0.44 cached $0.014 | $1.32 | 1,000,000 | $3.30 |
GPT-5 Mini | OpenAI | $0.25 cached $0.025 | $2.00 | 400,000 | $3.75 |
Mistral Large 3 | Mistral | $0.5 cached $0.05 | $1.50 | 128,000 | $3.75 |
Command R | Cohere | $0.5 | $1.50 | 128,000 | $3.75 |
Gemini 3.5 Flash-Lite | $0.3 cached $0.03 | $2.50 | 1,048,576 | $4.65 | |
Gemini 2.5 Flash | $0.3 cached $0.03 | $2.50 | 1,000,000 | $4.65 | |
Amazon Nova 2 Lite | Amazon | $0.3 | $2.50 | 1,000,000 | $4.65 |
Grok Build 0.1 | xAI | $1.00 cached $0.2 | $2.00 | 256,000 | $6.00 |
Amazon Nova Pro | Amazon | $0.8 | $3.20 | 300,000 | $7.20 |
Grok 4.3 | xAI | $1.25 cached $0.2 | $2.50 | 1,000,000 | $7.50 |
Grok 4.20 Reasoning | xAI | $1.25 cached $0.2 | $2.50 | 1,000,000 | $7.50 |
Grok 4.20 Non-Reasoning | xAI | $1.25 cached $0.2 | $2.50 | 1,000,000 | $7.50 |
Grok 4.20 Multi-Agent | xAI | $1.25 cached $0.2 | $2.50 | 1,000,000 | $7.50 |
Gemini 3.7 Flash Introductory through 2026-12-31 Standard from 2027-01-01: $1.50 input / $0.15 cached / $7.50 output | $0.75 cached $0.075 | $3.75 | 1,048,576 | $7.88 | |
Gemini 3.6 Flash Introductory through 2026-12-31 Standard from 2027-01-01: $1.50 input / $0.15 cached / $7.50 output | $0.75 cached $0.075 | $3.75 | 1,048,576 | $7.88 | |
Kimi K2.6 | Moonshot AI | $0.95 cached $0.16 | $4.00 | 262,144 | $8.85 |
Kimi K2.7 Code | Moonshot AI | $0.95 cached $0.19 | $4.00 | 262,144 | $8.85 |
GPT-5.4 Mini | OpenAI | $0.75 cached $0.075 | $4.50 | 400,000 | $9.00 |
GLM-5 Turbo | Z.AI | $1.20 cached $0.24 | $4.00 | 200,000 | $9.60 |
GLM-5V Turbo | Z.AI | $1.20 cached $0.24 | $4.00 | 200,000 | $9.60 |
DeepSeek V4 Pro Peak Off-peak: $0.66 input / $0.022 cached / $1.98 output | DeepSeek | $1.32 cached $0.044 | $3.96 | 1,000,000 | $9.90 |
Claude Haiku 4.5 | Anthropic | $1.00 cached $0.1 | $5.00 | 200,000 | $10.50 |
GLM-5.2 | Z.AI | $1.40 cached $0.26 | $4.40 | 1,000,000 | $10.80 |
GLM-5.1 | Z.AI | $1.40 cached $0.26 | $4.40 | 200,000 | $10.80 |
Grok 4.6 | xAI | $2.00 cached $0.5 | $6.00 | 500,000 | $15.00 |
Grok 4.5 | xAI | $2.00 cached $0.3 | $6.00 | 500,000 | $15.00 |
Mistral Medium 3.5 | Mistral | $1.50 cached $0.15 | $7.50 | 131,072 | $15.75 |
o3 | OpenAI | $2.00 cached $0.5 | $8.00 | 200,000 | $18.00 |
Gemini 3.5 Flash | $1.50 cached $0.15 | $9.00 | 1,048,576 | $18.00 | |
Jamba Large 1.7 | AI21 Labs | $2.00 | $8.00 | 256,000 | $18.00 |
GPT-5 | OpenAI | $1.25 cached $0.125 | $10.00 | 400,000 | $18.75 |
Gemini 2.5 Pro | $1.25 cached $0.125 | $10.00 | 1,048,576 | $18.75 | |
Claude Sonnet 5 | Anthropic | $2.00 cached $0.2 | $10.00 | 1,000,000 | $21.00 |
Command A+ | Cohere | $2.50 | $10.00 | 128,000 | $22.50 |
Command R+ | Cohere | $2.50 | $10.00 | 128,000 | $22.50 |
GPT-5.6 Terra | OpenAI | $2.00 cached $0.2 | $12.00 | 1,050,000 | $24.00 |
Gemini 3.1 Pro Preview | $2.00 cached $0.2 | $12.00 | 1,048,576 | $24.00 | |
GPT-5.3-Codex | OpenAI | $1.75 cached $0.175 | $14.00 | 400,000 | $26.25 |
GPT-5.4 | OpenAI | $2.50 cached $0.25 | $15.00 | 1,050,000 | $30.00 |
Claude Sonnet 4.6 | Anthropic | $3.00 cached $0.3 | $15.00 | 1,000,000 | $31.50 |
Kimi K3 | Moonshot AI | $3.00 cached $0.3 | $15.00 | 1,048,576 | $31.50 |
Claude Opus 5 | Anthropic | $5.00 cached $0.5 | $25.00 | 1,000,000 | $52.50 |
Claude Opus 4.8 | Anthropic | $5.00 cached $0.5 | $25.00 | 1,000,000 | $52.50 |
Claude Opus 4.6 | Anthropic | $5.00 cached $0.5 | $25.00 | 1,000,000 | $52.50 |
GPT-5.6 Sol | OpenAI | $5.00 cached $0.5 | $30.00 | 1,050,000 | $60.00 |
GPT-5.5 | OpenAI | $5.00 cached $0.5 | $30.00 | 1,050,000 | $60.00 |
Claude Fable 5 | Anthropic | $10.00 cached $1.00 | $50.00 | 1,000,000 | $105.00 |
o3-pro | OpenAI | $20.00 | $80.00 | 200,000 | $180.00 |
GPT-5.5 Pro | OpenAI | $30.00 | $180.00 | 1,050,000 | $360.00 |
How to Use This Tool
- Enter your estimated daily input tokens (the text you send to the AI) and daily output tokens (the AI's response length).
- Use the provider filter to narrow results to specific providers like OpenAI, Anthropic, Google, or others.
- Sort by monthly cost, input price, or output price to find the most cost-effective model for your use case.
- Click on a provider name to visit their official pricing page and sign up for API access.
- Compare multiple models side-by-side to find the best price-to-performance ratio for your specific workload.
Understanding AI API Pricing in 2026
AI API pricing is based on tokens processed, with separate rates for input tokens (your prompts) and output tokens (the model's responses). Prices are typically quoted per million tokens. Current examples include GPT-5.6 Luna at $0.20/$1.20 and Claude Sonnet 5 at its August introductory rate of $2/$10.
The AI pricing landscape has become increasingly competitive in 2026. DeepSeek V4 Flash is especially aggressive for cache-heavy agent workloads, while premium models like GPT-5.6 Sol and Claude Opus 5 offer stronger reasoning at higher price points. The right choice depends on task complexity, cache-hit rate, and output length.
Several factors beyond per-token pricing affect your total cost: Batch API discounts, prompt caching or provider-side context caching, and context window usage. DeepSeek V4, Anthropic, OpenAI, and Xiaomi MiMo all expose cached-input economics in different ways, so a naive cache-miss estimate can overstate real agent bills.
For cost optimization, consider these strategies: Use lower-cost models such as DeepSeek V4 Flash or GPT-5.6 Luna for simple tasks. Use prompt caching for system prompts that don't change. Batch non-urgent requests for 50% savings. Monitor token usage with observability tools like Helicone or Langfuse.
Last updated: August 2026
FAQ
How is the monthly cost calculated?
Monthly cost = (daily input tokens × input price per token × 30) + (daily output tokens × output price per token × 30). Prices are based on the latest published API pricing.
How often are prices updated?
We update pricing data regularly. The last update date is shown on the page. AI model prices change frequently, so always verify with the provider's official pricing page.
Which model is cheapest?
It depends on your use case. For cache-heavy agent traffic, DeepSeek V4 Flash can be extremely cheap. For simple tasks, GPT-4o-mini and Claude Haiku offer strong value. For complex reasoning, larger models like GPT-5, Claude Sonnet, or DeepSeek V4 Pro may be more cost-effective despite higher per-token costs.
What is Batch API pricing?
Several providers offer Batch API pricing at ~50% discount for requests that don't need real-time responses. OpenAI's Batch API, Anthropic's Message Batches, and Google's batch endpoints all offer significant savings for bulk processing like document analysis, data extraction, or content generation jobs.
How does prompt caching reduce costs?
Prompt caching and context caching store repeated prompt prefixes on the provider side. DeepSeek V4 Flash is a clear example: cache-hit input is $0.007/M off-peak and $0.014/M at peak, versus $0.22/M and $0.44/M for cache misses. This is especially valuable for coding agents that repeatedly send repository rules, system prompts, and stable context.
Which model offers the best value in 2026?
It depends on your use case. DeepSeek V4 Flash is strong for cache-heavy agent workloads and high-volume text tasks. Smaller models are better for simple extraction, while GPT-5, Claude Sonnet, or DeepSeek V4 Pro can be worth the extra cost for harder reasoning and coding.
Related Blog Posts
Official K3 pricing, OpenAI-compatible setup, automatic caching, multimodal input, and output limits.
DeepSeek V4-Flash-0731 API Pricing 2026Current V4 Flash and Pro pricing, cache-hit math, peak-hour status, and agent cost examples.
Gemini 3.6 Flash and 3.5 Flash-Lite API PricingGoogle's current Gemini prices, caching, Batch/Flex, and 1M-context model choices.
GPT-5.6 API Pricing: Sol, Terra, Luna, Caching, and Long ContextOfficial GPT-5.6 standard and long-context pricing with cache-write and cache-read costs.
AI API Pricing Comparison: August 2026Current side-by-side pricing comparison of the major AI API providers.
How to Reduce AI API Costs: 10 Proven StrategiesPractical tips to cut your AI API spend by 50-90%.