DevTk.AI
OpenAI API PricingGPT-6 PricingGPT-6 AstraGPT-6.1 SolGPT-6 LunaAPI Costs

OpenAI API Pricing 2026: GPT-6, GPT-5.6 and Codex Costs

Current OpenAI API prices for GPT-6 Astra, GPT-6.1 Sol, Luna, GPT-5.6 and Codex, including cache writes, long context, Batch, Flex and Fast mode.

DevTk.AI 2026-02-24 Updated 2026-10-02 4 min read

OpenAI’s current general-purpose API family is GPT-6: Astra for the highest capability, Sol for demanding coding and agents, and Luna for efficient high-volume work. GPT-5.6 remains available, and GPT-5.6 Sol has a temporary $4/$20 price available at least through November 21, 2026.

GPT-6.1 Sol, released September 29, 2026, is the current balanced tier. Input/output remain $2/$10, while cache reads fall from GPT-6 Sol’s $0.20/M to $0.10/M. The older gpt-6-sol ID retains its own $0.20/M cache-read price.

Current OpenAI API Prices

Standard short-context prices are USD per 1 million tokens.

ModelInputCached readCache writeOutputContext
GPT-6 Astra$10.00$1.00$12.50$50.001.05M
GPT-6.1 Sol$2.00$0.10$2.50$10.001.05M
GPT-6 Luna$0.10$0.01$0.125$0.501.05M
GPT-5.6 Sol$4.00$0.40$5.00$20.001.05M
GPT-5.6 Terra$2.00$0.20$2.50$12.001.05M
GPT-5.6 Luna$0.20$0.02$0.25$1.201.05M
GPT-5.5$5.00$0.50-$30.001.05M
GPT-5.4$2.50$0.25-$15.001.05M
GPT-5.4 mini$0.75$0.075-$4.50400K
GPT-5.4 nano$0.20$0.02-$1.25400K
GPT-5.3-Codex$1.75$0.175-$14.00400K

GPT-6.1 Sol cache reads cost 5% of normal input; Astra and Luna cache reads cost 10%. Explicit cache writes cost 1.25 times normal input. The GPT-6 family supports text and image input, up to 128K output, function calling, structured output, and built-in tools through the Responses API.

GPT-6 Long-Context Pricing

Prompts above 272K input tokens use higher prices for the entire request.

ModelLong inputLong cached readLong cache writeLong output
GPT-6 Astra$20.00$2.00$25.00$75.00
GPT-6.1 Sol$4.00$0.20$5.00$15.00
GPT-6 Luna$0.20$0.02$0.25$0.75

See the GPT-6 API pricing guide for detailed examples and migration advice.

Batch, Flex, Standard, and Fast Mode

Processing modeRelative priceUse case
Batch50%Asynchronous jobs with a 24-hour window
Flex50%Delay-tolerant lower-priority work
Standard100%Normal production traffic
Fast mode200%Latency-sensitive traffic
Ultrafast (Astra only)600%Very low output latency

Regional processing adds a 10% uplift for eligible models released on or after March 5, 2026. Hosted tools such as web search are billed separately.

Astra Ultrafast costs $60/$300 per million input/output tokens at short context, or $120/$450 above 272K input. It is a separate tier from Fast mode.

Cost Example

For 3M uncached input plus 1.5M output tokens at short-context Standard rates:

ModelTotal cost
GPT-6 Luna$1.05
GPT-5.6 Luna$2.40
GPT-6.1 Sol$21.00
GPT-5.6 Terra$24.00
GPT-5.6 Sol$42.00
GPT-6 Astra$105.00

GPT-6.1 Sol is cheaper than GPT-5.6 Terra for this mix, while GPT-6 Luna cuts the already reduced GPT-5.6 Luna bill by more than half. Astra is a premium escalation tier and should be justified by completion quality.

Which Model Should You Use?

  • GPT-6 Astra: highest-value tasks where failures, retries, or human repair are expensive.
  • GPT-6.1 Sol: demanding production coding, tool use, reasoning, and agent workflows.
  • GPT-6 Luna: routing, extraction, classification, search, summarization, and high-volume automation.
  • GPT-5.6: retain for evaluated production systems or compatibility, but benchmark GPT-6.1 Sol and Luna before new deployments.
  • GPT-5.3-Codex: use when a dedicated Codex API model is required by an existing workflow.

Use the AI Model Pricing Calculator for your token mix. Direct API rate limits remain separate from ChatGPT or Codex subscription allowances.

Official Sources

Related Posts