DevTk.AI

AI Model Pricing Directory

Browse pricing for 40+ AI models including GPT-5, Claude, Gemini, DeepSeek, Llama & more. Compare input/output costs, context windows, and capabilities.

Last updated: 2026-08-17 · 65 models from 14 providers

3 provider sources need re-verification: Meta, Meta (via providers), Alibaba Cloud.

65

Models

14

Providers

$0.035

Cheapest Input/M

1.1M

Max Context

Need to estimate costs?

Use our interactive Pricing Calculator to compare models side-by-side and estimate monthly costs.

Open Pricing Calculator

OpenAI

15 models Verified 2026-08-01
Model Input / 1M Output / 1M Context
GPT-5.6 Sol

GPT-5.6 flagship for complex professional and agentic work. The gpt-5.6 alias routes to Sol. Explicit cache writes cost 1.25x uncached input; requests above 272K input tokens use long-context pricing.

$5.00 $30.00 1.1M
GPT-5.6 Terra

Balanced GPT-5.6 production model at OpenAI's reduced August 2026 price. Explicit cache writes cost 1.25x uncached input; requests above 272K input tokens use long-context pricing.

$2.00 $12.00 1.1M
GPT-5.6 Luna

High-volume GPT-5.6 routing model at OpenAI's reduced August 2026 price, with the same 1.05M context window and 128K maximum output as Sol and Terra.

$0.2 $1.20 1.1M
GPT-5.5

Previous OpenAI frontier model for complex reasoning, coding, and professional work. GPT-5.6 Sol is the current flagship; prompts above 272K input tokens use higher long-context pricing.

$5.00 $30.00 1.1M
GPT-5.5 Pro

Higher-compute GPT-5.5 variant for the hardest professional tasks. No cached-input discount is listed on the official model page.

$30.00 $180.00 1.1M
GPT-5.4

Legacy OpenAI model for coding and professional work. GPT-5.6 Terra now uses the same standard $2.50/$15 price point.

$2.50 $15.00 1.1M
GPT-5 $1.25 $10.00 400K
GPT-5 Mini $0.25 $2.00 400K
GPT-5 Nano $0.05 $0.4 400K
GPT-4o mini $0.15 $0.6 128K
o3-pro

Highest reasoning capability for elite tasks.

$20.00 $80.00 200K
o3

Standard reasoning model.

$2.00 $8.00 200K
GPT-5.4 Mini

Current lower-cost GPT-5.4 production model.

$0.75 $4.50 400K
GPT-5.4 Nano

Current high-volume GPT-5.4 routing model.

$0.2 $1.25 400K
GPT-5.3-Codex

Current dedicated Codex API model for long-horizon agentic coding.

$1.75 $14.00 400K

Anthropic

7 models Verified 2026-08-01
Model Input / 1M Output / 1M Context
Claude Fable 5

Claude model for fiction and long-form creative work. Global API access was restored on 2026-07-01. The cache-write price shown is for a 5-minute TTL; 1-hour writes cost $20/M.

$10.00 $50.00 1.0M
Claude Opus 5

Current Claude Opus flagship for complex coding, agent, and knowledge-work tasks. The cache-write price shown is for a 5-minute TTL; 1-hour writes cost $10/M.

$5.00 $25.00 1.0M
Claude Sonnet 5

Introductory pricing through 2026-08-31. On 2026-09-01, standard pricing becomes $3/M input, $0.30/M cache reads, $3.75/M 5-minute cache writes, and $15/M output.

$2.00 $10.00 1.0M
Claude Opus 4.8

Previous-generation Opus model. Full 1M context remains available at standard pricing.

$5.00 $25.00 1.0M
Claude Opus 4.6

Previous-generation Opus model. Full 1M context is available at standard pricing.

$5.00 $25.00 1.0M
Claude Sonnet 4.6

Previous-generation balanced Claude model. Full 1M context remains available at standard pricing.

$3.00 $15.00 1.0M
Claude Haiku 4.5

Fastest Claude model. 200K context, 64K max output.

$1.00 $5.00 200K

Google

9 models Verified 2026-08-17
Model Input / 1M Output / 1M Context
Gemini 3.7 Flash

Current GA Gemini Flash workhorse for coding, multimodal reasoning, and multi-step agents. The 2026 introductory rate also applies to Batch and Flex at half the listed synchronous rate.

$0.75 $3.75 1.0M
Gemini 3.6 Flash

Previous GA Gemini Flash model. Google applies the same introductory 2026 price as Gemini 3.7 Flash; the standard rate returns on 2027-01-01.

$0.75 $3.75 1.0M
Gemini 3.5 Flash-Lite

Current GA low-cost Gemini model for high-throughput multimodal workloads.

$0.3 $2.50 1.0M
Gemini 3.5 Flash

Stable Gemini 3.5 Flash model for agentic loops, coding cycles, long-horizon tasks, search grounding, Batch API, Flex, context caching, and multimodal inputs.

$1.50 $9.00 1.0M
Gemini 3.1 Pro Preview

Preview Gemini Pro model. Inputs above 200K tokens use higher long-context pricing.

$2.00 $12.00 1.0M
Gemini 3.1 Flash-Lite

Fastest Gemini model. Optimized for high-throughput, multimodal tasks.

$0.25 $1.50 1.0M
Gemini 2.5 Pro

Long context >200K: $2.50 input, $15.00 output per 1M.

$1.25 $10.00 1.0M
Gemini 2.5 Flash $0.3 $2.50 1.0M
Gemini 2.5 Flash-Lite $0.1 $0.4 1.0M
Model Input / 1M Output / 1M Context
Grok 4.6

Current xAI frontier reasoning model with low, medium, high, and xhigh effort. xAI publishes no text-output limit. Prompts at or above 200K tokens bill all tokens at long-context rates; Priority Processing costs 2x standard rates.

$2.00 $6.00 500K
Grok 4.5

Previous xAI frontier model for coding, agentic tasks, and knowledge work. xAI does not publish a separate maximum-output limit; prompts at or above 200K tokens use long-context pricing.

$2.00 $6.00 500K
Grok 4.3

Lower-cost general-purpose xAI API model with configurable reasoning. xAI does not publish a separate maximum-output limit; requests above 200K tokens use long-context pricing.

$1.25 $2.50 1.0M
Grok 4.20 Reasoning

Pinned Grok 4.20 reasoning endpoint. xAI does not publish a separate maximum-output limit; requests above 200K tokens use long-context pricing.

$1.25 $2.50 1.0M
Grok 4.20 Non-Reasoning

Pinned Grok 4.20 endpoint without reasoning tokens. xAI does not publish a separate maximum-output limit; requests above 200K tokens use long-context pricing.

$1.25 $2.50 1.0M
Grok 4.20 Multi-Agent

Beta multi-agent endpoint for parallel research with 4-agent and 16-agent configurations. xAI does not publish a separate maximum-output limit.

$1.25 $2.50 1.0M
Grok Build 0.1

Public-beta xAI model optimized for coding and software-building workflows. Requests above 200K tokens use long-context pricing; xAI does not publish a separate maximum-output limit.

$1.00 $2.00 256K

Mistral

3 models Verified 2026-08-01
Model Input / 1M Output / 1M Context
Mistral Large 3 $0.5 $1.50 128K
Mistral Medium 3.5

Current balanced Mistral production model.

$1.50 $7.50 131K
Mistral Small 4

Current low-cost Mistral model for high-volume production routing. Mistral lists a 90% discount for cached input tokens.

$0.15 $0.6 131K

DeepSeek

2 models Verified 2026-08-17
Model Input / 1M Output / 1M Context
DeepSeek V4 Flash

Serves DeepSeek-V4-Flash-0731. New time-of-day pricing took effect on 2026-08-16 at 16:00 UTC. Peak hours are 01:00-04:00 and 06:00-10:00 UTC; off-peak rates are 50% lower. The retired deepseek-chat and deepseek-reasoner aliases were discontinued on 2026-07-24.

$0.44 $1.32 1.0M
DeepSeek V4 Pro

Serves DeepSeek-V4-Pro-0813. New time-of-day pricing took effect on 2026-08-16 at 16:00 UTC. Peak hours are 01:00-04:00 and 06:00-10:00 UTC; off-peak rates are 50% lower.

$1.32 $3.96 1.0M

Xiaomi MiMo

2 models Verified 2026-08-01
Model Input / 1M Output / 1M Context
Xiaomi MiMo-V2.5-Pro

Open-sourced under MIT. New pay-as-you-go pricing took effect on 2026-05-27 00:00 Beijing time. Domestic pricing is ¥3.00/M cache-miss input, ¥0.025/M cached input, and ¥6.00/M output. Cache writing is currently free for a limited time.

$0.435 $0.87 1.0M
Xiaomi MiMo-V2.5

Native full-modal model with text, image, video, and audio understanding. New pay-as-you-go pricing took effect on 2026-05-27 00:00 Beijing time. Domestic pricing is ¥1.00/M cache-miss input, ¥0.02/M cached input, and ¥2.00/M output.

$0.14 $0.28 1.0M

MiniMax

1 models Verified 2026-08-01
Model Input / 1M Output / 1M Context
MiniMax M3

Latest MiniMax frontier coding and agent model. Supports Adaptive Thinking, tool use, and native image/video input. Inputs above 512K tokens use long-context pricing and currently require limited-access availability.

$0.3 $1.20 1.0M

Z.AI

4 models Verified 2026-08-17
Model Input / 1M Output / 1M Context
GLM-5.2

Latest GLM model with published pay-as-you-go USD API pricing. GLM-5.3 is newer and available in Coding Plan, but its standard API access and token pricing are still coming soon.

$1.40 $4.40 1.0M
GLM-5.1

Previous GLM flagship for long-horizon coding and agent tasks. It remains available, while GLM-5.2 adds a 1M-token context window.

$1.40 $4.40 200K
GLM-5 Turbo

Current faster GLM model optimized for tool use and agent workflows.

$1.20 $4.00 200K
GLM-5V Turbo

Current multimodal GLM agent model for text, image, and video input.

$1.20 $4.00 200K

Moonshot AI

4 models Verified 2026-08-01
Model Input / 1M Output / 1M Context
Kimi K2.5

China API pricing in CNY. Current K2.5 multimodal model with text, image, and video input.

¥4.00 ¥21.00 262K
Kimi K2.6

Global API pricing in USD for the general Kimi multimodal model.

$0.95 $4.00 262K
Kimi K2.7 Code

Global API pricing in USD. Current Kimi coding model with thinking-only operation.

$0.95 $4.00 262K
Kimi K3

Always-reasoning K3 API with automatic context caching and max reasoning effort. Input plus output cannot exceed 1,048,576 tokens; the full model weights are available.

$3.00 $15.00 1.0M

Alibaba

2 models Verified 2026-08-01
Model Input / 1M Output / 1M Context
Qwen3.7 Plus

Current Qwen3.7 Plus alias promotional price in China, listed in CNY. List price is ¥2/¥8 per million tokens below 256K and ¥6/¥24 above 256K; Alibaba lists no promotion end date.

¥1.60 ¥6.40 1.0M
Qwen3.7 Max

Current Qwen3.7 Max alias promotional price in China, listed in CNY. List price is ¥12/M input and ¥36/M output; Alibaba lists no promotion end date.

¥6.00 ¥18.00 1.0M

Cohere

3 models Verified 2026-08-01
Model Input / 1M Output / 1M Context
Command A+

Current Cohere flagship for agentic, multilingual, RAG, and enterprise workloads.

$2.50 $10.00 128K
Command R+

Enterprise RAG-optimized with citation grounding.

$2.50 $10.00 128K
Command R $0.5 $1.50 128K

AI21 Labs

2 models Verified 2026-08-01
Model Input / 1M Output / 1M Context
Jamba Large 1.7

Current Jamba Large API alias with a hybrid SSM-Transformer architecture and 256K context. The stable site slug is retained for existing links.

$2.00 $8.00 256K
Jamba Mini 2

Current Jamba Mini API alias with 256K context. The stable site slug is retained for existing links.

$0.2 $0.4 256K

Amazon

4 models Verified 2026-08-01
Model Input / 1M Output / 1M Context
Amazon Nova 2 Lite

Current Amazon Nova 2 model on Bedrock for multimodal reasoning, tool use, and long-context workloads.

$0.3 $2.50 1.0M
Amazon Nova Pro

Via AWS Bedrock. 300K context.

$0.8 $3.20 300K
Amazon Nova Lite $0.06 $0.24 300K
Amazon Nova Micro

Text-only. Lowest cost option on Bedrock.

$0.035 $0.14 128K

Related Resources