GPT-6 API Pricing Guide 2026: Astra, GPT-6.1 Sol & Luna Token Costs
Current GPT-6 API prices for Astra, GPT-6.1 Sol, and Luna, including cached input, cache writes, long-context rates, Batch, Flex, Fast mode, and cost examples.
OpenAI’s GPT-6 API family has three tiers: GPT-6 Astra for the hardest end-to-end work, GPT-6.1 Sol for demanding coding and agent workflows, and GPT-6 Luna for focused high-volume tasks. All three provide a 1.05M-token context window, up to 128K output tokens, text and image input, and text output.
GPT-6.1 Sol, released September 29, 2026, is the current balanced tier. Input/output remain $2/$10, while cache reads fall from GPT-6 Sol’s $0.20/M to $0.10/M. The older gpt-6-sol ID retains its own $0.20/M cache-read price.
GPT-6 Standard Pricing
Prices are USD per 1 million tokens for requests with no more than 272K input tokens.
| Model | API ID | Input | Cached input | Cache write | Output |
|---|---|---|---|---|---|
| GPT-6 Astra | gpt-6-astra | $10.00 | $1.00 | $12.50 | $50.00 |
| GPT-6.1 Sol | gpt-6.1-sol | $2.00 | $0.10 | $2.50 | $10.00 |
| GPT-6 Luna | gpt-6-luna | $0.10 | $0.01 | $0.125 | $0.50 |
Cache reads cost 5% of normal input on GPT-6.1 Sol and 10% on Astra and Luna. Explicit cache writes cost 1.25 times normal input. Tool calls, regional processing, and hosted tools can add separate charges.
Long-Context Pricing Above 272K
When a prompt contains more than 272K input tokens, OpenAI applies the higher rates to the entire request, not only the tokens above the threshold.
| Model | Long input | Long cached input | Long cache write | Long output |
|---|---|---|---|---|
| GPT-6 Astra | $20.00 | $2.00 | $25.00 | $75.00 |
| GPT-6.1 Sol | $4.00 | $0.20 | $5.00 | $15.00 |
| GPT-6 Luna | $0.20 | $0.02 | $0.25 | $0.75 |
For large repositories and document sets, crossing 272K doubles input and cache rates and raises output by 50%. Context trimming and retrieval can therefore matter more than choosing a cheaper model.
Batch, Flex, Standard, and Fast Mode
| Processing mode | Relative price | Typical use |
|---|---|---|
| Batch | 50% of Standard | Asynchronous jobs with a 24-hour window |
| Flex | 50% of Standard | Delay-tolerant, lower-priority traffic |
| Standard | 100% | Normal interactive production traffic |
| Fast mode | 200% of Standard | Latency-sensitive workloads |
| Ultrafast (Astra only) | 600% | Very low output latency |
At short context, GPT-6.1 Sol costs $1/$5 in Batch or Flex and $4/$20 in Fast mode. GPT-6 Luna falls to $0.05/$0.25 in Batch or Flex.
Astra Ultrafast costs $60/$300 per million input/output tokens at short context, or $120/$450 above 272K input. It is a separate tier from Fast mode.
Cost Example
For 2M uncached input tokens plus 500K output tokens at short-context Standard rates:
| Model | Input cost | Output cost | Total |
|---|---|---|---|
| GPT-6 Astra | $20.00 | $25.00 | $45.00 |
| GPT-6.1 Sol | $4.00 | $5.00 | $9.00 |
| GPT-6 Luna | $0.20 | $0.25 | $0.45 |
This is a token-volume comparison, not a quality ranking. Astra costs 5 times Sol and 100 times Luna for this mix, so it must save enough retries, output tokens, tool calls, or human repair time to justify the difference.
GPT-6 vs GPT-5.6 Pricing
GPT-5.6 Sol currently has a promotional rate of $4/$20, available at least through November 21, 2026. GPT-6.1 Sol costs $2/$10, exactly half that rate. GPT-6 Luna is also half the input price of GPT-5.6 Luna and cuts output from $1.20 to $0.50.
GPT-6 Astra is a different premium tier at $10/$50. Do not treat it as the default migration target for every GPT-5.6 workload. Start by testing GPT-6.1 Sol at the same effective reasoning effort, then escalate only when your evals show a meaningful completion-rate gain.
Which GPT-6 Model Should You Use?
- GPT-6 Astra: high-value computer use, difficult research, complex software engineering, cybersecurity, and professional workflows where failure is expensive.
- GPT-6.1 Sol: the practical default for demanding coding, agents, tool use, and production reasoning.
- GPT-6 Luna: extraction, classification, routing, repository search, summarization, and repeatable high-volume work.
Use the AI Model Pricing Calculator for your traffic mix, and compare the premium tier with Claude Opus 5.5’s value analysis.