DevTk.AI

AI 模型定价目录

浏览 40+ AI 模型定价,包括 GPT-5、Claude、Gemini、DeepSeek、Llama 等。对比输入/输出价格、上下文窗口和能力。

最后更新:2026-10-02 · 来自 14 家提供商的 80 个模型

7 家提供商需要重新核验:Meta、Meta (via providers)、Alibaba Cloud、Cohere、AI21 Labs、AWS Bedrock、Amazon。

80

模型

14

提供商

$0.035

最低输入/M

1.1M

最大上下文

需要估算成本?

使用我们的定价计算器,交互式对比各模型并估算每月费用。

打开定价计算器

OpenAI

19 个模型 已核验 2026-10-02
模型 输入 / 百万 输出 / 百万 上下文
GPT-6.1 Sol

Current balanced GPT-6 model for coding and professional work. Cache reads cost half as much as GPT-6 Sol. Prompts above 272K input tokens bill the entire request at long-context rates; Batch/Flex cost 50% of Standard and Fast mode costs 2x. Tool calling requires the Responses API.

$2.00 $10.00 1.1M
GPT-6 Astra

OpenAI's highest-capability GPT-6 model for hard end-to-end work. Prompts above 272K input tokens bill the entire request at long-context rates; Batch and Flex cost 50% of Standard, while Fast mode costs 2x. Ultrafast mode costs 6x Standard: $60/$300 at short context or $120/$450 above 272K input tokens.

$10.00 $50.00 1.1M
GPT-6 Sol

GPT-6 model for complex coding and agentic workflows. Prompts above 272K input tokens bill the entire request at long-context rates; Batch and Flex cost 50% of Standard, while Fast mode costs 2x.

$2.00 $10.00 1.1M
GPT-6 Luna

GPT-6 model for focused, repeatable, high-volume work. Prompts above 272K input tokens bill the entire request at long-context rates; Batch and Flex cost 50% of Standard, while Fast mode costs 2x.

$0.1 $0.5 1.1M
GPT-5.6 Sol

Previous-generation flagship. The gpt-5.6 alias routes to Sol. Promotional pricing is available at least through 2026-11-21; explicit cache writes cost 1.25x uncached input and requests above 272K input tokens use long-context pricing.

$4.00 $20.00 1.1M
GPT-5.6 Terra

Balanced GPT-5.6 production model at OpenAI's reduced August 2026 price. Explicit cache writes cost 1.25x uncached input; requests above 272K input tokens use long-context pricing.

$2.00 $12.00 1.1M
GPT-5.6 Luna

High-volume GPT-5.6 routing model at OpenAI's reduced August 2026 price, with the same 1.05M context window and 128K maximum output as Sol and Terra.

$0.2 $1.20 1.1M
GPT-5.5

Previous OpenAI frontier model for complex reasoning, coding, and professional work. GPT-5.6 Sol is the current flagship; prompts above 272K input tokens use higher long-context pricing.

$5.00 $30.00 1.1M
GPT-5.5 Pro

Higher-compute GPT-5.5 variant for the hardest professional tasks. No cached-input discount is listed on the official model page.

$30.00 $180.00 1.1M
GPT-5.4

Legacy OpenAI model for coding and professional work. GPT-5.6 Terra now uses the same standard $2.50/$15 price point.

$2.50 $15.00 1.1M
GPT-5 $1.25 $10.00 400K
GPT-5 Mini $0.25 $2.00 400K
GPT-5 Nano $0.05 $0.4 400K
GPT-4o mini $0.15 $0.6 128K
o3-pro

Highest reasoning capability for elite tasks.

$20.00 $80.00 200K
o3

Standard reasoning model.

$2.00 $8.00 200K
GPT-5.4 Mini

Current lower-cost GPT-5.4 production model.

$0.75 $4.50 400K
GPT-5.4 Nano

Current high-volume GPT-5.4 routing model.

$0.2 $1.25 400K
GPT-5.3-Codex

Current dedicated Codex API model for long-horizon agentic coding.

$1.75 $14.00 400K

Anthropic

10 个模型 已核验 2026-10-02
模型 输入 / 百万 输出 / 百万 上下文
Claude Fable 5.1

Anthropic's highest-capability broadly available model for long-horizon agentic work. Cache reads cost 0.025x base input; the listed write price is for a 5-minute TTL and 1-hour writes cost $20/M.

$10.00 $50.00 1.0M
Claude Opus 5.5

Current default Claude model for long-running agentic coding and knowledge work. Adaptive thinking is always on. The listed cache-write price is for a 5-minute TTL; 1-hour writes cost $8/M, Batch input/output costs $2/$10, and Fast mode costs $8/$40.

$4.00 $20.00 1.0M
Claude Fable 5

Claude model for fiction and long-form creative work. Global API access was restored on 2026-07-01. The cache-write price shown is for a 5-minute TTL; 1-hour writes cost $20/M.

$10.00 $50.00 1.0M
Claude Opus 5

Previous-generation Opus model for complex coding, agent, and knowledge-work tasks. The cache-write price shown is for a 5-minute TTL; 1-hour writes cost $10/M.

$5.00 $25.00 1.0M
Claude Sonnet 5.5

Current balanced Claude model, priced the same as Sonnet 5. Five-minute cache writes cost $2.50/M; one-hour writes cost $4/M. The full 1M context window uses standard rates; Batch input/output are $1/$5.

$2.00 $10.00 1.0M
Claude Sonnet 5

Anthropic made the launch price permanent instead of applying the previously scheduled 2026-09-01 increase. The listed cache-write price is for a 5-minute TTL; 1-hour writes cost $4/M.

$2.00 $10.00 1.0M
Claude Opus 4.8

Previous-generation Opus model. Full 1M context remains available at standard pricing.

$5.00 $25.00 1.0M
Claude Opus 4.6

Previous-generation Opus model. Full 1M context is available at standard pricing.

$5.00 $25.00 1.0M
Claude Sonnet 4.6

Previous-generation balanced Claude model. Full 1M context remains available at standard pricing.

$3.00 $15.00 1.0M
Claude Haiku 4.5

Fastest Claude model. 200K context, 64K max output.

$1.00 $5.00 200K

Google

10 个模型 已核验 2026-10-02
模型 输入 / 百万 输出 / 百万 上下文
Gemini 3.8 Flash

Current GA Gemini Flash flagship for long-horizon coding, autonomous agents, and enterprise workflows. The 2026 introductory rate also applies to Batch and Flex at half the listed synchronous rate.

$0.75 $3.75 1.0M
Gemini 3.7 Flash

Previous GA Gemini Flash model. The 2026 introductory rate also applies to Batch and Flex at half the listed synchronous rate.

$0.75 $3.75 1.0M
Gemini 3.6 Flash

Previous GA Gemini Flash model. Google applies the same introductory 2026 price as Gemini 3.7 Flash; the standard rate returns on 2027-01-01.

$0.75 $3.75 1.0M
Gemini 3.5 Flash-Lite

Current GA low-cost Gemini model for high-throughput multimodal workloads.

$0.3 $2.50 1.0M
Gemini 3.5 Flash

Stable Gemini 3.5 Flash model for agentic loops, coding cycles, long-horizon tasks, search grounding, Batch API, Flex, context caching, and multimodal inputs.

$1.50 $9.00 1.0M
Gemini 3.1 Pro Preview

Preview Gemini Pro model. Inputs above 200K tokens use higher long-context pricing.

$2.00 $12.00 1.0M
Gemini 3.1 Flash-Lite

Fastest Gemini model. Optimized for high-throughput, multimodal tasks.

$0.25 $1.50 1.0M
Gemini 2.5 Pro

Long context >200K: $2.50 input, $15.00 output per 1M.

$1.25 $10.00 1.0M
Gemini 2.5 Flash $0.3 $2.50 1.0M
Gemini 2.5 Flash-Lite $0.1 $0.4 1.0M

xAI

8 个模型 已核验 2026-10-02
模型 输入 / 百万 输出 / 百万 上下文
Grok 4.7

Current xAI frontier model for coding, agents, and knowledge work. xAI publishes no text-output limit. Prompts at or above 200K tokens bill all tokens at long-context rates; the public API does not offer the 2x-priced Fast variant.

$2.00 $6.00 500K
Grok 4.6

Previous xAI frontier reasoning model with low, medium, high, and xhigh effort. xAI publishes no text-output limit. Prompts at or above 200K tokens bill all tokens at long-context rates.

$2.00 $6.00 500K
Grok 4.5

Previous xAI frontier model for coding, agentic tasks, and knowledge work. xAI does not publish a separate maximum-output limit; prompts at or above 200K tokens use long-context pricing.

$2.00 $6.00 500K
Grok 4.3

Lower-cost general-purpose xAI API model with configurable reasoning. xAI does not publish a separate maximum-output limit; requests above 200K tokens use long-context pricing.

$1.25 $2.50 1.0M
Grok 4.20 Reasoning

Pinned Grok 4.20 reasoning endpoint. xAI does not publish a separate maximum-output limit; requests above 200K tokens use long-context pricing.

$1.25 $2.50 1.0M
Grok 4.20 Non-Reasoning

Pinned Grok 4.20 endpoint without reasoning tokens. xAI does not publish a separate maximum-output limit; requests above 200K tokens use long-context pricing.

$1.25 $2.50 1.0M
Grok 4.20 Multi-Agent

Beta multi-agent endpoint for parallel research with 4-agent and 16-agent configurations. xAI does not publish a separate maximum-output limit.

$1.25 $2.50 1.0M
Grok Build 0.1

Public-beta xAI model optimized for coding and software-building workflows. Requests above 200K tokens use long-context pricing; xAI does not publish a separate maximum-output limit.

$1.00 $2.00 256K

Mistral

3 个模型 已核验 2026-10-02
模型 输入 / 百万 输出 / 百万 上下文
Mistral Large 3 $0.5 $1.50 128K
Mistral Medium 3.5

Current balanced Mistral production model.

$1.50 $7.50 131K
Mistral Small 4

Current low-cost Mistral model for high-volume production routing. Mistral lists a 90% discount for cached input tokens.

$0.15 $0.6 131K

DeepSeek

2 个模型 已核验 2026-10-02
模型 输入 / 百万 输出 / 百万 上下文
DeepSeek V4.1 Flash

Current DeepSeek Flash endpoint with native vision. Retired deepseek-v4-flash aliases temporarily route here. Peak hours are 01:00-04:00 and 06:00-10:00 UTC on weekdays; all other hours are 50% lower.

$0.3 $1.20 1.0M
DeepSeek V4 Pro

Serves DeepSeek-V4-Pro-0813. DeepSeek continued this endpoint after 2026-09-14 with unchanged billing. Peak hours are 01:00-04:00 and 06:00-10:00 UTC on weekdays; off-peak rates are 50% lower.

$1.32 $3.96 1.0M

Xiaomi MiMo

3 个模型 已核验 2026-10-02
模型 输入 / 百万 输出 / 百万 上下文
Xiaomi MiMo-V2.6-Pro

Current full-modal MiMo flagship. Uses the former V2.5-Pro real-time price; Batch costs $0.2175/M cache-miss input and $0.435/M output. Cache writes are free for a limited time.

$0.435 $0.87 1.0M
Xiaomi MiMo-V2.6-Pro-UltraSpeed

Latency-focused V2.6-Pro endpoint with up to 20x output speed according to Xiaomi. Token rates are 10x Pro; Batch is unsupported. Contact Xiaomi for custom rate limits.

$4.35 $8.70 1.0M
Xiaomi MiMo-V2.6-Flash

Current full-modal high-volume MiMo model. Uses the former V2.5 real-time price; Batch costs $0.07/M cache-miss input and $0.14/M output. Cache writes are free for a limited time.

$0.14 $0.28 1.0M

MiniMax

1 个模型 已核验 2026-10-02
模型 输入 / 百万 输出 / 百万 上下文
MiniMax M3

Latest MiniMax frontier coding and agent model at permanent 50%-off standard rates. Priority costs 1.5x Standard. The current official table does not publish M3 cache-write pricing. Supports Adaptive Thinking, tool use, and native image/video input. Inputs above 512K tokens use long-context pricing and currently require limited-access availability.

$0.3 $1.20 1.0M

Z.AI

7 个模型 已核验 2026-10-02
模型 输入 / 百万 输出 / 百万 上下文
GLM-5.3

Current GLM flagship for complex coding and long-horizon agents. Reasoning is always enabled with low, high, and max effort levels; cached-input storage is free for a limited time.

$1.40 $4.40 1.0M
GLM-5.3 Flash

Lowest-cost GLM-5.3 route. Cached-input storage is free for a limited time.

$0.15 $0.5 1.0M
GLM-5.3 FlashX

Faster GLM-5.3 API route. Cached-input storage is free for a limited time.

$0.37 $1.25 1.0M
GLM-5.2

Previous GLM flagship. GLM-5.3 now has pay-as-you-go API pricing at the same token rates.

$1.40 $4.40 1.0M
GLM-5.1

Previous GLM flagship for long-horizon coding and agent tasks. It remains available, while GLM-5.2 adds a 1M-token context window.

$1.40 $4.40 200K
GLM-5 Turbo

Current faster GLM model optimized for tool use and agent workflows.

$1.20 $4.00 200K
GLM-5V Turbo

Current multimodal GLM agent model for text, image, and video input.

$1.20 $4.00 200K

Moonshot AI

4 个模型 已核验 2026-10-02
模型 输入 / 百万 输出 / 百万 上下文
Kimi K2.5

China API pricing in CNY. Current K2.5 multimodal model with text, image, and video input.

¥4.00 ¥21.00 262K
Kimi K2.6

Global API pricing in USD for the general Kimi multimodal model.

$0.95 $4.00 262K
Kimi K2.7 Code

Global API pricing in USD. Current Kimi coding model with thinking-only operation.

$0.95 $4.00 262K
Kimi K3

Kimi K3 global USD API rates. Automatic prefix caching uses a default 5-minute TTL with $3/M cache writes; one-hour writes cost $6/M. Cache hits cost $0.30/M and refresh the TTL without another write charge. Full model weights are available.

$3.00 $15.00 1.0M

Alibaba

4 个模型 已核验 2026-10-02
模型 输入 / 百万 输出 / 百万 上下文
Qwen3.8 Max

Current Qwen flagship. China/global deployment price in CNY; Batch File input/output costs ¥6/¥18 per million tokens.

¥12.00 ¥36.00 1.0M
Qwen3.8 Flash

Current lower-cost Qwen3.8 multimodal route. China/global deployment price in CNY with no input-length tiers.

¥0.8 ¥2.70 1.0M
Qwen3.7 Plus

Previous Qwen3.7 Plus alias promotional price in China, listed in CNY. List price is ¥2/¥8 per million tokens below 256K and ¥6/¥24 above 256K; Alibaba lists no promotion end date.

¥1.60 ¥6.40 1.0M
Qwen3.7 Max

Previous Qwen3.7 Max alias promotional price in China, listed in CNY. List price is ¥12/M input and ¥36/M output; Alibaba lists no promotion end date.

¥6.00 ¥18.00 1.0M

Cohere

3 个模型 已核验 2026-08-01
模型 输入 / 百万 输出 / 百万 上下文
Command A+

Current Cohere flagship for agentic, multilingual, RAG, and enterprise workloads.

$2.50 $10.00 128K
Command R+

Enterprise RAG-optimized with citation grounding.

$2.50 $10.00 128K
Command R $0.5 $1.50 128K

AI21 Labs

2 个模型 已核验 2026-08-01
模型 输入 / 百万 输出 / 百万 上下文
Jamba Large 1.7

Current Jamba Large API alias with a hybrid SSM-Transformer architecture and 256K context. The stable site slug is retained for existing links.

$2.00 $8.00 256K
Jamba Mini 2

Current Jamba Mini API alias with 256K context. The stable site slug is retained for existing links.

$0.2 $0.4 256K

Amazon

4 个模型 已核验 2026-08-01
模型 输入 / 百万 输出 / 百万 上下文
Amazon Nova 2 Lite

Current Amazon Nova 2 model on Bedrock for multimodal reasoning, tool use, and long-context workloads.

$0.3 $2.50 1.0M
Amazon Nova Pro

Via AWS Bedrock. 300K context.

$0.8 $3.20 300K
Amazon Nova Lite $0.06 $0.24 300K
Amazon Nova Micro

Text-only. Lowest cost option on Bedrock.

$0.035 $0.14 128K

相关资源