DevTk.AI
Grok API PricingxAIGrok 4.6Grok 4.5Long Context Pricing

Grok API Pricing 2026: Grok 4.6, 4.5, Long Context, and Tool Costs

Current Grok 4.6 API pricing, its 200K long-context threshold, cache increase from Grok 4.5, Priority Processing, tool costs, context limits, and migration guidance.

DevTk.AI 2026-02-24 Updated 2026-08-17 4 min read

Grok 4.6 is xAI’s newest production reasoning model. It launched on August 12, 2026 under the API ID grok-4.6, with a 500K context window, text and image input, function calling, structured output, and low, medium, high, and xhigh reasoning effort.

xAI publishes no separate text-output limit for Grok 4.6. Do not mistake the 500K context window for a maximum output size.

Current Grok API Prices

USD per million tokens:

ModelContextShort inputShort cachedShort outputLong inputLong cachedLong output
Grok 4.6500K$2.00$0.50$6.00$4.00$1.00$12.00
Grok 4.5500K$2.00$0.30$6.00$4.00$0.60$12.00
Grok 4.31M$1.25$0.20$2.50$2.50$0.40$5.00
Grok 4.20 endpoints1M$1.25$0.20$2.50$2.50$0.40$5.00
Grok Build 0.1256K$1.00$0.20$2.00$2.00$0.40$4.00

Long pricing starts when the prompt reaches 200K tokens. At that point, xAI charges the long rate for every token in the request, not only the portion above 200K. Cached prompt tokens count toward the threshold.

Grok 4.6 vs 4.5 Pricing

Grok 4.6 keeps 4.5’s ordinary input and output prices. Cached input increased by 66.7%, from $0.30 to $0.50 below 200K and from $0.60 to $1.00 in the long tier.

That is the migration risk for cache-heavy agents. xAI reports stronger coding and agent results and adds xhigh, but a production upgrade should compare cost per successful task rather than token price alone. Grok 4.5 remains listed and is not marked deprecated.

The 200K Cost Cliff

One request with 199K input and 20K output costs about $0.518 on Grok 4.6. Raising the prompt to 200K with the same output activates long pricing and costs $1.04. A small context increase can therefore nearly double the request bill.

Split independent context when possible, compact conversation history, and keep retrieved documents focused. Do not split requests if doing so harms task completion enough to create retries.

Priority Processing and Batch

Priority Processing charges 2x standard rates for input, cached input, output, and reasoning tokens. Billing uses the priority rate only when the response confirms "service_tier": "priority".

Grok 4.6 and 4.5 currently have no Batch discount. xAI lists a 20% Batch discount for Grok 4.3 and the Grok 4.20 endpoints.

Server-Side Tool Costs

Token rates do not include xAI-hosted tool invocations:

ToolPrice per 1,000 calls
Web Search$5.00
X Search$5.00
Code Execution$5.00
File Attachments search$10.00
Collections search$2.50
Remote MCPToken-based; no separate invocation fee

For example, ten Web Search calls add $0.05 before model tokens. This can dominate a short prompt.

Cost Examples

Assuming separate requests below 200K and no cache hits:

Daily usageGrok 4.6Grok 4.3 / 4.20Grok Build 0.1
100K input + 50K output$15/month$7.50/month$6/month
1M input + 500K output$150/month$75/month$60/month
10M input + 5M output$1,500/month$750/month$600/month

The same short-context token totals give Grok 4.5 the same no-cache bill. Grok 4.6 costs more only when cached input is present or Priority is used.

Which Grok Model Should You Use?

  • Choose Grok 4.6 for the hardest coding and agent tasks after validating the cache increase.
  • Keep Grok 4.5 when 4.6 does not improve completion enough to cover higher cached input.
  • Choose Grok 4.3 for lower-cost general agents and 1M context.
  • Choose a Grok 4.20 endpoint for pinned reasoning behavior or the multi-agent beta.
  • Choose Grok Build 0.1 for a lower-cost coding-specific public beta.

Read Grok 4.6 vs Gemini 3.7 Flash and use the AI Pricing Calculator for current model costs.

Official sources checked August 17, 2026: xAI pricing, Grok 4.6 docs, release notes, and Priority Processing.

Related Posts