GPT-6 Sol and Luna: half the price, and what else changed
OpenAI's two new GPT-6 models cost about half as much to run as the versions they replace. Here is what each is for, what the published results do and do not show, and how to choose between Sol, Luna and Astra.
By Monolith

Most of the AI work inside a business falls into two piles. One pile holds careful, multi-step jobs: fixing a website, reconciling a month of invoices, drafting a proposal from a messy brief. The other holds high-volume routine work: summarizing support emails, pulling fields out of forms, answering the same product question for the hundredth time. The two piles need different things from a model, and they carry very different bills.
OpenAI's release on September 22, 2026 is aimed squarely at that split. GPT-6 Sol and GPT-6 Luna are two new models trained with methods similar to GPT-6 Astra, the flagship OpenAI launched earlier in September. Astra remains the company's most capable model; Sol and Luna bring much of its progress to faster, cheaper models, and OpenAI has cut their API prices in half compared with GPT-5.6's promotional rates. 1
Three GPT-6 models, three kinds of work
OpenAI now offers three GPT-6 models. Astra is the one to choose, in OpenAI's words, when you want the best results and an uncompromising experience. Sol is built for complex coding and agentic workflows, meaning jobs where the model plans and carries out several steps with tools rather than answering a single question. Luna is described as the company's most efficient model for focused, high-volume tasks. 123
| Model | Input | Cached input | Output | Knowledge cutoff |
|---|---|---|---|---|
| GPT-6 Sol | $2.00 | $0.20 | $10.00 | April 20, 2026 |
| GPT-6 Luna | $0.10 | $0.01 | $0.50 | May 18, 2026 |
Input is the material you send; output is what the model writes back, including its working-out on harder problems. Cached input is material the system has already seen in a recent request, such as a long set of instructions reused on every call, and it is billed at a tenth of the normal input rate. The knowledge cutoff is the point after which the model knows nothing unless you give it current material or connect a search tool. Luna's is about a month later than Sol's, but for anything recent, both need to be told. 23
Astra, the top of the family, has its own guide covering computer use and its higher price.
What the price cut actually means
| Model | Input: GPT-5.6 to GPT-6 | Output: GPT-5.6 to GPT-6 | Change |
|---|---|---|---|
| Sol | $4.00 to $2.00 | $20.00 to $10.00 | 50% lower on both |
| Luna | $0.20 to $0.10 | $1.20 to $0.50 | 50% lower input, 58% lower output |
Read the comparison carefully. OpenAI measures the saving against GPT-5.6 promotional pricing, and Sol's $4 and $20 rates were already a promotion when GPT-5.6 launched. Luna's output price falls further than the headline suggests, from $1.20 to $0.50, which is a 58% cut. 1
Price per token is only half of a bill. The other half is how many tokens a job takes. Artificial Analysis, an independent firm that runs the same tests across many models, measured the cost of running its full test suite and found GPT-6 Sol at $1.06 per task against $1.99 for GPT-5.6 Sol, and GPT-6 Luna at $0.07 against $0.18 for GPT-5.6 Luna. Both new models actually wrote slightly more per task, about 31,000 output tokens for Sol against 29,000 and 51,000 for Luna against 41,000, so the savings come from the lower price rather than from shorter answers. 4
A few billing rules are worth knowing before a large project. A prompt longer than 272,000 tokens is billed at twice the input rate and 1.5 times the output rate for the whole request. Writing material into the cache for reuse costs 1.25 times the normal input rate, and each later read of it costs a tenth. Batch and Flex processing, for work that can wait, costs half the standard rate, and the faster Fast mode costs twice as much. 23
What OpenAI's own results show
OpenAI's headline comparison uses AutomationBench, a test in which an AI agent completes end-to-end business workflows using 47 tools across sales, marketing, operations, customer support, finance and HR. That is close to the kind of work many businesses hope to automate, which makes it the most relevant of OpenAI's published tests for this audience. 1
| Model (reasoning effort) | Score | Cost per task |
|---|---|---|
| GPT-6 Sol (xhigh) | 33.2% | $0.27 |
| Claude Fable 5.1 with Opus 5 fallback (max) | 31.4% | More than 8.9 times Sol |
| GPT-6 Astra (low) | 30.3% | 3.9 times Sol |
| Claude Opus 5 (max) | 26.9% | 11.1 times Sol |
The pattern is clear: on this test, Sol at its second-highest reasoning setting edges past the Claude models in OpenAI's comparison at a small fraction of their cost per task. The absolute numbers are also worth sitting with. The best result here completes about a third of these workflows, so an automated process built on any of these models still needs checks and a person to handle what the model cannot finish. OpenAI also reports that Luna at high effort improved on its predecessor by 5.4 points while costing 58% less per task. 1
OpenAI published three more comparisons worth knowing, each with a condition attached. On DeepSWE v1.1, a test of long software engineering tasks in real codebases, Sol at maximum effort scored 68.8%, within 1.1 points of the best Claude Fable 5 result at about 80% lower cost per task; OpenAI notes it used Fable 5 because Fable 5.1 scores were not available. On OSWorld 2.0, which tests an AI operating a computer through everyday and professional workflows, Sol at xhigh scored 60.5% against 60.3% for Claude Opus 5 at medium, again at about 80% lower cost. And on OpenAI's internal factuality test, Sol made about half as many mistakes as GPT-5.6 Sol. That last test uses real conversations in which users had already flagged an error, which OpenAI says is not representative of typical use. 1
What independent testing found
Artificial Analysis tested both models at maximum effort on launch day. Its summary is more measured than the launch messaging: prices are about half, but on its broad Intelligence Index and its Coding Agent Index, the new models score level with GPT-5.6, with progress on some tests and regressions on others. 4
| Measure | GPT-6 Sol | GPT-5.6 Sol | GPT-6 Luna | GPT-5.6 Luna |
|---|---|---|---|---|
| Cost per task, full test suite | $1.06 | $1.99 | $0.07 | $0.18 |
| Output tokens per task | 31,000 | 29,000 | 51,000 | 41,000 |
| Coding Agent Index (Codex) | 57 | 55 | 41 | 43 |
| AutomationBench-AA | 62% | 60% | 53% | 50% |
| Hallucination rate (lower is better) | 60% | 92% | 77% | 93% |
| Accuracy on the same test | 54% | 59% | 44% | 43% |
Two findings matter most for a business. The first is how the hallucination rate fell. A hallucination is a confident answer that is wrong. On Artificial Analysis's knowledge test, Sol's rate dropped from 92% to 60% largely because it now declines to answer more often: it attempted 83% of questions against 99% before, which cut wrong answers by about a quarter but also lowered accuracy by five points. In practice, expect more replies along the lines of I do not know, and build your workflow to route those to a person rather than treating them as failures. 4
The second is a regression in the kind of work agencies care about. On GDPval-AA, a test of professional deliverables across 44 occupations, Sol dropped about 100 Elo points and Luna about 75. Elo is a head-to-head rating borrowed from chess, where a higher number means the model's work wins more often when compared directly with others. Luna also slipped about 45 points on AA-Briefcase, a test of multi-week office projects, while Sol held level. After inspecting hundreds of outputs, Artificial Analysis traced the drops to weaker presentation and deliverables that left out required elements. If your team uses GPT-5.6 Sol for client-facing documents, compare the two on a real brief before switching. 4
Coding is split. Sol gained two points on the Coding Agent Index, while Luna lost two and scored lower than its predecessor on DeepSWE v1.1 in Artificial Analysis's run. OpenAI reported a higher DeepSWE score for Luna in its own setup; the two used different test harnesses, which is why a single public score should never settle a decision. 14

Which model for which job
| Job | Start with | Why | Check yourself |
|---|---|---|---|
| Summarizing, sorting or extracting from large volumes | Luna | Lowest price, and the job OpenAI designed it for | A sample of outputs each week, and how often it declines to answer |
| Multi-step business workflows across apps | Sol | Best cost for the score on OpenAI's AutomationBench | Every step that sends, pays or publishes, before it runs |
| Coding and code review | Sol | Gains in both OpenAI's and independent coding results | Tests, a human review and whether the change is ready to merge |
| Client-facing documents and decks | Compare Sol with your current model | Independent tests found weaker presentation | Structure, missing sections and whether it follows the brief |
| The hardest, highest-stakes work | Astra | OpenAI says it remains the best across the board | Cost per accepted result, not cost per token |
The table points to a practical habit: route work by type rather than choosing one model for everything. A support team might send ticket summaries to Luna, escalate billing disputes that need several systems to Sol, and reserve Astra for the rare case that justifies its price. The savings come from matching the model to the job, and the quality comes from checking the result where it matters.
Where you can use them
In ChatGPT, both models are rolling out in the Work and Codex areas for Plus, Pro, Business, Enterprise and Edu users, gradually over launch day. Free and Go users can try Luna in the desktop app. OpenAI says the models are not yet available in ChatGPT's regular Chat view. Developers can call them through the API as gpt-6-sol and gpt-6-luna. 1
One detail for anyone comparing ChatGPT with published scores: OpenAI ran its evaluations in its research environment or through the API, and notes that production ChatGPT can behave slightly differently because of its own system instructions and tools. 1
Specifications for the people wiring it up
| Specification | GPT-6 Sol | GPT-6 Luna |
|---|---|---|
| API model ID | gpt-6-sol | gpt-6-luna |
| Context window | 1,050,000 tokens | 1,050,000 tokens |
| Maximum output | 128,000 tokens | 128,000 tokens |
| Input and output | Text and image in; text out | Text and image in; text out |
| Reasoning effort | none, low, medium (default), high, xhigh, max | none, low, medium (default), high, xhigh, max |
| Long prompts | Above 272,000 input tokens: 2x input, 1.5x output | Above 272,000 input tokens: 2x input, 1.5x output |
| Cheaper processing | Batch and Flex at half price; Fast at double | Batch and Flex at half price; Fast at double |
Reasoning effort sets how much the model thinks before answering. Both models accept none, which skips the thinking step for quick, simple replies, and default to medium. OpenAI's documentation adds that function calling through the older Chat Completions interface only works with effort set to none; for tools and reasoning together, use the Responses interface. The context window is the total amount of material one request can hold, which is large enough for a long contract or a small codebase, but capacity is not a promise that every detail will be used well. 23
OpenAI also improved caching for agents and long conversations. The system now reuses more context by default, a new dashboard shows how much input is being cached, and developers can raise or lower reasoning effort, or switch tools on and off, partway through a conversation without losing the cache discount. For an agent that runs the same long instructions all day, those changes can matter as much as the headline price. 1

Safety results, and what they measure
OpenAI says Sol and Luna carry over Astra's alignment work, and reports lower rates of misleading claims about their own coding work. In its coding deception test, where tasks are deliberately chosen to tempt a model into dishonesty and effort is set to maximum, Sol was caught in 1.3% of answers against 10.4% for GPT-5.6 Sol, and Luna in 2.8% against 9.5%. Astra measured 0.5%. OpenAI stresses that these are stress tests and that deception is much rarer in typical use. For a business, the practical reading is that the models are less likely to claim work is done when it is not, which reduces, but does not remove, the need to check. 1
Monolith's take: GPT-6 Sol and Luna make capable AI markedly cheaper to run, and Sol is a strong choice for multi-step workflows and coding. Independent testing suggests they are not smarter than GPT-5.6 overall, and weaker on polished client deliverables. Route routine volume to Luna, agentic work to Sol, and test document quality on a real brief before switching anything client-facing.
Is GPT-6 Sol better than GPT-5.6 Sol?
Should we use GPT-6 Sol or GPT-6 Luna?
Can I use GPT-6 Sol and Luna in ChatGPT?
Anthropic released Claude Opus 5.5 the same day. See how it compares on price and long coding jobs.
Read for this feature. The numbers match the markers in the text.
- OpenAI: Introducing GPT-6 Sol and Luna, September 22, 2026openai.com
- OpenAI API documentation: GPT-6 Sol; checked September 22, 2026developers.openai.com
- OpenAI API documentation: GPT-6 Luna; checked September 22, 2026developers.openai.com
- Artificial Analysis: GPT-6 Sol and Luna evaluation results, posted on X, September 22, 2026x.com
Planning which parts of your work to automate, and with which model? Talk it through with us.
Want to apply this to a project? Here is the related service.
25 AI instructions to try on everyday work
Prompts are instructions you give an AI tool. This PDF includes 25 examples for inquiries, marketing and admin. Adapt them to your task and check the results.