AI 模型定价目录
浏览 40+ AI 模型定价,包括 GPT-5、Claude、Gemini、DeepSeek、Llama 等。对比输入/输出价格、上下文窗口和能力。
最后更新:2026-08-17 · 来自 14 家提供商的 65 个模型
3 家提供商需要重新核验:Meta、Meta (via providers)、Alibaba Cloud。
65
模型
14
提供商
$0.035
最低输入/M
1.1M
最大上下文
需要估算成本?
使用我们的定价计算器,交互式对比各模型并估算每月费用。
| 模型 | 输入 / 百万 | 输出 / 百万 | 上下文 |
|---|---|---|---|
| GPT-5.6 Sol GPT-5.6 flagship for complex professional and agentic work. The gpt-5.6 alias routes to Sol. Explicit cache writes cost 1.25x uncached input; requests above 272K input tokens use long-context pricing. | $5.00 | $30.00 | 1.1M |
| GPT-5.6 Terra Balanced GPT-5.6 production model at OpenAI's reduced August 2026 price. Explicit cache writes cost 1.25x uncached input; requests above 272K input tokens use long-context pricing. | $2.00 | $12.00 | 1.1M |
| GPT-5.6 Luna High-volume GPT-5.6 routing model at OpenAI's reduced August 2026 price, with the same 1.05M context window and 128K maximum output as Sol and Terra. | $0.2 | $1.20 | 1.1M |
| GPT-5.5 Previous OpenAI frontier model for complex reasoning, coding, and professional work. GPT-5.6 Sol is the current flagship; prompts above 272K input tokens use higher long-context pricing. | $5.00 | $30.00 | 1.1M |
| GPT-5.5 Pro Higher-compute GPT-5.5 variant for the hardest professional tasks. No cached-input discount is listed on the official model page. | $30.00 | $180.00 | 1.1M |
| GPT-5.4 Legacy OpenAI model for coding and professional work. GPT-5.6 Terra now uses the same standard $2.50/$15 price point. | $2.50 | $15.00 | 1.1M |
| GPT-5 | $1.25 | $10.00 | 400K |
| GPT-5 Mini | $0.25 | $2.00 | 400K |
| GPT-5 Nano | $0.05 | $0.4 | 400K |
| GPT-4o mini | $0.15 | $0.6 | 128K |
| o3-pro Highest reasoning capability for elite tasks. | $20.00 | $80.00 | 200K |
| o3 Standard reasoning model. | $2.00 | $8.00 | 200K |
| GPT-5.4 Mini Current lower-cost GPT-5.4 production model. | $0.75 | $4.50 | 400K |
| GPT-5.4 Nano Current high-volume GPT-5.4 routing model. | $0.2 | $1.25 | 400K |
| GPT-5.3-Codex Current dedicated Codex API model for long-horizon agentic coding. | $1.75 | $14.00 | 400K |
| 模型 | 输入 / 百万 | 输出 / 百万 | 上下文 |
|---|---|---|---|
| Claude Fable 5 Claude model for fiction and long-form creative work. Global API access was restored on 2026-07-01. The cache-write price shown is for a 5-minute TTL; 1-hour writes cost $20/M. | $10.00 | $50.00 | 1.0M |
| Claude Opus 5 Current Claude Opus flagship for complex coding, agent, and knowledge-work tasks. The cache-write price shown is for a 5-minute TTL; 1-hour writes cost $10/M. | $5.00 | $25.00 | 1.0M |
| Claude Sonnet 5 Introductory pricing through 2026-08-31. On 2026-09-01, standard pricing becomes $3/M input, $0.30/M cache reads, $3.75/M 5-minute cache writes, and $15/M output. | $2.00 | $10.00 | 1.0M |
| Claude Opus 4.8 Previous-generation Opus model. Full 1M context remains available at standard pricing. | $5.00 | $25.00 | 1.0M |
| Claude Opus 4.6 Previous-generation Opus model. Full 1M context is available at standard pricing. | $5.00 | $25.00 | 1.0M |
| Claude Sonnet 4.6 Previous-generation balanced Claude model. Full 1M context remains available at standard pricing. | $3.00 | $15.00 | 1.0M |
| Claude Haiku 4.5 Fastest Claude model. 200K context, 64K max output. | $1.00 | $5.00 | 200K |
| 模型 | 输入 / 百万 | 输出 / 百万 | 上下文 |
|---|---|---|---|
| Gemini 3.7 Flash Current GA Gemini Flash workhorse for coding, multimodal reasoning, and multi-step agents. The 2026 introductory rate also applies to Batch and Flex at half the listed synchronous rate. | $0.75 | $3.75 | 1.0M |
| Gemini 3.6 Flash Previous GA Gemini Flash model. Google applies the same introductory 2026 price as Gemini 3.7 Flash; the standard rate returns on 2027-01-01. | $0.75 | $3.75 | 1.0M |
| Gemini 3.5 Flash-Lite Current GA low-cost Gemini model for high-throughput multimodal workloads. | $0.3 | $2.50 | 1.0M |
| Gemini 3.5 Flash Stable Gemini 3.5 Flash model for agentic loops, coding cycles, long-horizon tasks, search grounding, Batch API, Flex, context caching, and multimodal inputs. | $1.50 | $9.00 | 1.0M |
| Gemini 3.1 Pro Preview Preview Gemini Pro model. Inputs above 200K tokens use higher long-context pricing. | $2.00 | $12.00 | 1.0M |
| Gemini 3.1 Flash-Lite Fastest Gemini model. Optimized for high-throughput, multimodal tasks. | $0.25 | $1.50 | 1.0M |
| Gemini 2.5 Pro Long context >200K: $2.50 input, $15.00 output per 1M. | $1.25 | $10.00 | 1.0M |
| Gemini 2.5 Flash | $0.3 | $2.50 | 1.0M |
| Gemini 2.5 Flash-Lite | $0.1 | $0.4 | 1.0M |
| 模型 | 输入 / 百万 | 输出 / 百万 | 上下文 |
|---|---|---|---|
| Grok 4.6 Current xAI frontier reasoning model with low, medium, high, and xhigh effort. xAI publishes no text-output limit. Prompts at or above 200K tokens bill all tokens at long-context rates; Priority Processing costs 2x standard rates. | $2.00 | $6.00 | 500K |
| Grok 4.5 Previous xAI frontier model for coding, agentic tasks, and knowledge work. xAI does not publish a separate maximum-output limit; prompts at or above 200K tokens use long-context pricing. | $2.00 | $6.00 | 500K |
| Grok 4.3 Lower-cost general-purpose xAI API model with configurable reasoning. xAI does not publish a separate maximum-output limit; requests above 200K tokens use long-context pricing. | $1.25 | $2.50 | 1.0M |
| Grok 4.20 Reasoning Pinned Grok 4.20 reasoning endpoint. xAI does not publish a separate maximum-output limit; requests above 200K tokens use long-context pricing. | $1.25 | $2.50 | 1.0M |
| Grok 4.20 Non-Reasoning Pinned Grok 4.20 endpoint without reasoning tokens. xAI does not publish a separate maximum-output limit; requests above 200K tokens use long-context pricing. | $1.25 | $2.50 | 1.0M |
| Grok 4.20 Multi-Agent Beta multi-agent endpoint for parallel research with 4-agent and 16-agent configurations. xAI does not publish a separate maximum-output limit. | $1.25 | $2.50 | 1.0M |
| Grok Build 0.1 Public-beta xAI model optimized for coding and software-building workflows. Requests above 200K tokens use long-context pricing; xAI does not publish a separate maximum-output limit. | $1.00 | $2.00 | 256K |
| 模型 | 输入 / 百万 | 输出 / 百万 | 上下文 |
|---|---|---|---|
| Mistral Large 3 | $0.5 | $1.50 | 128K |
| Mistral Medium 3.5 Current balanced Mistral production model. | $1.50 | $7.50 | 131K |
| Mistral Small 4 Current low-cost Mistral model for high-volume production routing. Mistral lists a 90% discount for cached input tokens. | $0.15 | $0.6 | 131K |
| 模型 | 输入 / 百万 | 输出 / 百万 | 上下文 |
|---|---|---|---|
| DeepSeek V4 Flash Serves DeepSeek-V4-Flash-0731. New time-of-day pricing took effect on 2026-08-16 at 16:00 UTC. Peak hours are 01:00-04:00 and 06:00-10:00 UTC; off-peak rates are 50% lower. The retired deepseek-chat and deepseek-reasoner aliases were discontinued on 2026-07-24. | $0.44 | $1.32 | 1.0M |
| DeepSeek V4 Pro Serves DeepSeek-V4-Pro-0813. New time-of-day pricing took effect on 2026-08-16 at 16:00 UTC. Peak hours are 01:00-04:00 and 06:00-10:00 UTC; off-peak rates are 50% lower. | $1.32 | $3.96 | 1.0M |
| 模型 | 输入 / 百万 | 输出 / 百万 | 上下文 |
|---|---|---|---|
| Xiaomi MiMo-V2.5-Pro Open-sourced under MIT. New pay-as-you-go pricing took effect on 2026-05-27 00:00 Beijing time. Domestic pricing is ¥3.00/M cache-miss input, ¥0.025/M cached input, and ¥6.00/M output. Cache writing is currently free for a limited time. | $0.435 | $0.87 | 1.0M |
| Xiaomi MiMo-V2.5 Native full-modal model with text, image, video, and audio understanding. New pay-as-you-go pricing took effect on 2026-05-27 00:00 Beijing time. Domestic pricing is ¥1.00/M cache-miss input, ¥0.02/M cached input, and ¥2.00/M output. | $0.14 | $0.28 | 1.0M |
| 模型 | 输入 / 百万 | 输出 / 百万 | 上下文 |
|---|---|---|---|
| MiniMax M3 Latest MiniMax frontier coding and agent model. Supports Adaptive Thinking, tool use, and native image/video input. Inputs above 512K tokens use long-context pricing and currently require limited-access availability. | $0.3 | $1.20 | 1.0M |
| 模型 | 输入 / 百万 | 输出 / 百万 | 上下文 |
|---|---|---|---|
| GLM-5.2 Latest GLM model with published pay-as-you-go USD API pricing. GLM-5.3 is newer and available in Coding Plan, but its standard API access and token pricing are still coming soon. | $1.40 | $4.40 | 1.0M |
| GLM-5.1 Previous GLM flagship for long-horizon coding and agent tasks. It remains available, while GLM-5.2 adds a 1M-token context window. | $1.40 | $4.40 | 200K |
| GLM-5 Turbo Current faster GLM model optimized for tool use and agent workflows. | $1.20 | $4.00 | 200K |
| GLM-5V Turbo Current multimodal GLM agent model for text, image, and video input. | $1.20 | $4.00 | 200K |
| 模型 | 输入 / 百万 | 输出 / 百万 | 上下文 |
|---|---|---|---|
| Kimi K2.5 China API pricing in CNY. Current K2.5 multimodal model with text, image, and video input. | ¥4.00 | ¥21.00 | 262K |
| Kimi K2.6 Global API pricing in USD for the general Kimi multimodal model. | $0.95 | $4.00 | 262K |
| Kimi K2.7 Code Global API pricing in USD. Current Kimi coding model with thinking-only operation. | $0.95 | $4.00 | 262K |
| Kimi K3 Always-reasoning K3 API with automatic context caching and max reasoning effort. Input plus output cannot exceed 1,048,576 tokens; the full model weights are available. | $3.00 | $15.00 | 1.0M |
| 模型 | 输入 / 百万 | 输出 / 百万 | 上下文 |
|---|---|---|---|
| Qwen3.7 Plus Current Qwen3.7 Plus alias promotional price in China, listed in CNY. List price is ¥2/¥8 per million tokens below 256K and ¥6/¥24 above 256K; Alibaba lists no promotion end date. | ¥1.60 | ¥6.40 | 1.0M |
| Qwen3.7 Max Current Qwen3.7 Max alias promotional price in China, listed in CNY. List price is ¥12/M input and ¥36/M output; Alibaba lists no promotion end date. | ¥6.00 | ¥18.00 | 1.0M |
| 模型 | 输入 / 百万 | 输出 / 百万 | 上下文 |
|---|---|---|---|
| Command A+ Current Cohere flagship for agentic, multilingual, RAG, and enterprise workloads. | $2.50 | $10.00 | 128K |
| Command R+ Enterprise RAG-optimized with citation grounding. | $2.50 | $10.00 | 128K |
| Command R | $0.5 | $1.50 | 128K |
| 模型 | 输入 / 百万 | 输出 / 百万 | 上下文 |
|---|---|---|---|
| Jamba Large 1.7 Current Jamba Large API alias with a hybrid SSM-Transformer architecture and 256K context. The stable site slug is retained for existing links. | $2.00 | $8.00 | 256K |
| Jamba Mini 2 Current Jamba Mini API alias with 256K context. The stable site slug is retained for existing links. | $0.2 | $0.4 | 256K |
| 模型 | 输入 / 百万 | 输出 / 百万 | 上下文 |
|---|---|---|---|
| Amazon Nova 2 Lite Current Amazon Nova 2 model on Bedrock for multimodal reasoning, tool use, and long-context workloads. | $0.3 | $2.50 | 1.0M |
| Amazon Nova Pro Via AWS Bedrock. 300K context. | $0.8 | $3.20 | 300K |
| Amazon Nova Lite | $0.06 | $0.24 | 300K |
| Amazon Nova Micro Text-only. Lowest cost option on Bedrock. | $0.035 | $0.14 | 128K |