InsightsPublished 12 min read

GPT-6 Sol and Luna: half the price, and what else changed

OpenAI's two new GPT-6 models cost about half as much to run as the versions they replace. Here is what each is for, what the published results do and do not show, and how to choose between Sol, Luna and Astra.

By Monolith

Two colleagues at a shared desk, one lit by a sunset in the left windows, the other by evening light with a crescent moon in the right windows
Two colleagues working at a shared desk as the sun sets in one window and the moon rises in another.

Most of the AI work inside a business falls into two piles. One pile holds careful, multi-step jobs: fixing a website, reconciling a month of invoices, drafting a proposal from a messy brief. The other holds high-volume routine work: summarizing support emails, pulling fields out of forms, answering the same product question for the hundredth time. The two piles need different things from a model, and they carry very different bills.

OpenAI's release on September 22, 2026 is aimed squarely at that split. GPT-6 Sol and GPT-6 Luna are two new models trained with methods similar to GPT-6 Astra, the flagship OpenAI launched earlier in September. Astra remains the company's most capable model; Sol and Luna bring much of its progress to faster, cheaper models, and OpenAI has cut their API prices in half compared with GPT-5.6's promotional rates. 1

Three GPT-6 models, three kinds of work

OpenAI now offers three GPT-6 models. Astra is the one to choose, in OpenAI's words, when you want the best results and an uncompromising experience. Sol is built for complex coding and agentic workflows, meaning jobs where the model plans and carries out several steps with tools rather than answering a single question. Luna is described as the company's most efficient model for focused, high-volume tasks. 123

ModelInputCached inputOutputKnowledge cutoff
GPT-6 Sol$2.00$0.20$10.00April 20, 2026
GPT-6 Luna$0.10$0.01$0.50May 18, 2026
OpenAI API documentation, checked September 22, 2026. Prices in USD per million tokens for standard-length requests. A token is a small piece of text, roughly three quarters of a word. 23

Input is the material you send; output is what the model writes back, including its working-out on harder problems. Cached input is material the system has already seen in a recent request, such as a long set of instructions reused on every call, and it is billed at a tenth of the normal input rate. The knowledge cutoff is the point after which the model knows nothing unless you give it current material or connect a search tool. Luna's is about a month later than Sol's, but for anything recent, both need to be told. 23

Astra, the top of the family, has its own guide covering computer use and its higher price.

What the price cut actually means

ModelInput: GPT-5.6 to GPT-6Output: GPT-5.6 to GPT-6Change
Sol$4.00 to $2.00$20.00 to $10.0050% lower on both
Luna$0.20 to $0.10$1.20 to $0.5050% lower input, 58% lower output
OpenAI's published price change, USD per million tokens. The GPT-5.6 Sol figures were themselves a promotional rate. 1

Read the comparison carefully. OpenAI measures the saving against GPT-5.6 promotional pricing, and Sol's $4 and $20 rates were already a promotion when GPT-5.6 launched. Luna's output price falls further than the headline suggests, from $1.20 to $0.50, which is a 58% cut. 1

Price per token is only half of a bill. The other half is how many tokens a job takes. Artificial Analysis, an independent firm that runs the same tests across many models, measured the cost of running its full test suite and found GPT-6 Sol at $1.06 per task against $1.99 for GPT-5.6 Sol, and GPT-6 Luna at $0.07 against $0.18 for GPT-5.6 Luna. Both new models actually wrote slightly more per task, about 31,000 output tokens for Sol against 29,000 and 51,000 for Luna against 41,000, so the savings come from the lower price rather than from shorter answers. 4

A few billing rules are worth knowing before a large project. A prompt longer than 272,000 tokens is billed at twice the input rate and 1.5 times the output rate for the whole request. Writing material into the cache for reuse costs 1.25 times the normal input rate, and each later read of it costs a tenth. Batch and Flex processing, for work that can wait, costs half the standard rate, and the faster Fast mode costs twice as much. 23

What OpenAI's own results show

OpenAI's headline comparison uses AutomationBench, a test in which an AI agent completes end-to-end business workflows using 47 tools across sales, marketing, operations, customer support, finance and HR. That is close to the kind of work many businesses hope to automate, which makes it the most relevant of OpenAI's published tests for this audience. 1

Model (reasoning effort)ScoreCost per task
GPT-6 Sol (xhigh)33.2%$0.27
Claude Fable 5.1 with Opus 5 fallback (max)31.4%More than 8.9 times Sol
GPT-6 Astra (low)30.3%3.9 times Sol
Claude Opus 5 (max)26.9%11.1 times Sol
OpenAI's AutomationBench 1.0.6 results, published September 22, 2026. Higher scores are better; cost is per task relative to GPT-6 Sol. Competitor scores come from public reports. OpenAI notes the Fable 5.1 cost is understated because it omits Opus 5 fallbacks, which ran on about 40% of tasks. 1

The pattern is clear: on this test, Sol at its second-highest reasoning setting edges past the Claude models in OpenAI's comparison at a small fraction of their cost per task. The absolute numbers are also worth sitting with. The best result here completes about a third of these workflows, so an automated process built on any of these models still needs checks and a person to handle what the model cannot finish. OpenAI also reports that Luna at high effort improved on its predecessor by 5.4 points while costing 58% less per task. 1

OpenAI published three more comparisons worth knowing, each with a condition attached. On DeepSWE v1.1, a test of long software engineering tasks in real codebases, Sol at maximum effort scored 68.8%, within 1.1 points of the best Claude Fable 5 result at about 80% lower cost per task; OpenAI notes it used Fable 5 because Fable 5.1 scores were not available. On OSWorld 2.0, which tests an AI operating a computer through everyday and professional workflows, Sol at xhigh scored 60.5% against 60.3% for Claude Opus 5 at medium, again at about 80% lower cost. And on OpenAI's internal factuality test, Sol made about half as many mistakes as GPT-5.6 Sol. That last test uses real conversations in which users had already flagged an error, which OpenAI says is not representative of typical use. 1

What independent testing found

Artificial Analysis tested both models at maximum effort on launch day. Its summary is more measured than the launch messaging: prices are about half, but on its broad Intelligence Index and its Coding Agent Index, the new models score level with GPT-5.6, with progress on some tests and regressions on others. 4

MeasureGPT-6 SolGPT-5.6 SolGPT-6 LunaGPT-5.6 Luna
Cost per task, full test suite$1.06$1.99$0.07$0.18
Output tokens per task31,00029,00051,00041,000
Coding Agent Index (Codex)57554143
AutomationBench-AA62%60%53%50%
Hallucination rate (lower is better)60%92%77%93%
Accuracy on the same test54%59%44%43%
Artificial Analysis results, posted September 22, 2026. All four models at maximum reasoning effort. The Coding Agent Index row uses OpenAI's Codex tool; the other rows use Artificial Analysis's own setup, so compare along each row, not between rows. 4

Two findings matter most for a business. The first is how the hallucination rate fell. A hallucination is a confident answer that is wrong. On Artificial Analysis's knowledge test, Sol's rate dropped from 92% to 60% largely because it now declines to answer more often: it attempted 83% of questions against 99% before, which cut wrong answers by about a quarter but also lowered accuracy by five points. In practice, expect more replies along the lines of I do not know, and build your workflow to route those to a person rather than treating them as failures. 4

The second is a regression in the kind of work agencies care about. On GDPval-AA, a test of professional deliverables across 44 occupations, Sol dropped about 100 Elo points and Luna about 75. Elo is a head-to-head rating borrowed from chess, where a higher number means the model's work wins more often when compared directly with others. Luna also slipped about 45 points on AA-Briefcase, a test of multi-week office projects, while Sol held level. After inspecting hundreds of outputs, Artificial Analysis traced the drops to weaker presentation and deliverables that left out required elements. If your team uses GPT-5.6 Sol for client-facing documents, compare the two on a real brief before switching. 4

Coding is split. Sol gained two points on the Coding Agent Index, while Luna lost two and scored lower than its predecessor on DeepSWE v1.1 in Artificial Analysis's run. OpenAI reported a higher DeepSWE score for Luna in its own setup; the two used different test harnesses, which is why a single public score should never settle a decision. 14

An operations coordinator working through routine customer requests on a laptop in a bright small office, with a colleague on the phone behind her
An operations coordinator working through a steady stream of routine customer requests.

Which model for which job

JobStart withWhyCheck yourself
Summarizing, sorting or extracting from large volumesLunaLowest price, and the job OpenAI designed it forA sample of outputs each week, and how often it declines to answer
Multi-step business workflows across appsSolBest cost for the score on OpenAI's AutomationBenchEvery step that sends, pays or publishes, before it runs
Coding and code reviewSolGains in both OpenAI's and independent coding resultsTests, a human review and whether the change is ready to merge
Client-facing documents and decksCompare Sol with your current modelIndependent tests found weaker presentationStructure, missing sections and whether it follows the brief
The hardest, highest-stakes workAstraOpenAI says it remains the best across the boardCost per accepted result, not cost per token
Monolith's suggested starting points based on the published results above. These are proposed trials, not measured outcomes for any particular business.

The table points to a practical habit: route work by type rather than choosing one model for everything. A support team might send ticket summaries to Luna, escalate billing disputes that need several systems to Sol, and reserve Astra for the rare case that justifies its price. The savings come from matching the model to the job, and the quality comes from checking the result where it matters.

Where you can use them

In ChatGPT, both models are rolling out in the Work and Codex areas for Plus, Pro, Business, Enterprise and Edu users, gradually over launch day. Free and Go users can try Luna in the desktop app. OpenAI says the models are not yet available in ChatGPT's regular Chat view. Developers can call them through the API as gpt-6-sol and gpt-6-luna. 1

One detail for anyone comparing ChatGPT with published scores: OpenAI ran its evaluations in its research environment or through the API, and notes that production ChatGPT can behave slightly differently because of its own system instructions and tools. 1

Specifications for the people wiring it up

SpecificationGPT-6 SolGPT-6 Luna
API model IDgpt-6-solgpt-6-luna
Context window1,050,000 tokens1,050,000 tokens
Maximum output128,000 tokens128,000 tokens
Input and outputText and image in; text outText and image in; text out
Reasoning effortnone, low, medium (default), high, xhigh, maxnone, low, medium (default), high, xhigh, max
Long promptsAbove 272,000 input tokens: 2x input, 1.5x outputAbove 272,000 input tokens: 2x input, 1.5x output
Cheaper processingBatch and Flex at half price; Fast at doubleBatch and Flex at half price; Fast at double
OpenAI API documentation for each model, checked September 22, 2026. 23

Reasoning effort sets how much the model thinks before answering. Both models accept none, which skips the thinking step for quick, simple replies, and default to medium. OpenAI's documentation adds that function calling through the older Chat Completions interface only works with effort set to none; for tools and reasoning together, use the Responses interface. The context window is the total amount of material one request can hold, which is large enough for a long contract or a small codebase, but capacity is not a promise that every detail will be used well. 23

OpenAI also improved caching for agents and long conversations. The system now reuses more context by default, a new dashboard shows how much input is being cached, and developers can raise or lower reasoning effort, or switch tools on and off, partway through a conversation without losing the cache discount. For an agent that runs the same long instructions all day, those changes can matter as much as the headline price. 1

A finance lead pointing at a wall display showing an automated workflow while an operations manager watches in a meeting room
A finance lead and an operations manager watching an automated month-end workflow run.

Safety results, and what they measure

OpenAI says Sol and Luna carry over Astra's alignment work, and reports lower rates of misleading claims about their own coding work. In its coding deception test, where tasks are deliberately chosen to tempt a model into dishonesty and effort is set to maximum, Sol was caught in 1.3% of answers against 10.4% for GPT-5.6 Sol, and Luna in 2.8% against 9.5%. Astra measured 0.5%. OpenAI stresses that these are stress tests and that deception is much rarer in typical use. For a business, the practical reading is that the models are less likely to claim work is done when it is not, which reduces, but does not remove, the need to check. 1

The short version

Monolith's take: GPT-6 Sol and Luna make capable AI markedly cheaper to run, and Sol is a strong choice for multi-step workflows and coding. Independent testing suggests they are not smarter than GPT-5.6 overall, and weaker on polished client deliverables. Route routine volume to Luna, agentic work to Sol, and test document quality on a real brief before switching anything client-facing.

Questions, answered
Is GPT-6 Sol better than GPT-5.6 Sol?
It costs about half as much per task and improves on several tests, including business workflows and coding. Independent testing found its overall intelligence level with GPT-5.6 Sol and a drop in the quality of professional deliverables, so test document work before switching. 14
Should we use GPT-6 Sol or GPT-6 Luna?
Use Luna for high-volume tasks with a clear goal, such as summarizing or extracting information, and Sol for complex, multi-step work and coding. Reserve GPT-6 Astra for the hardest work where the best result justifies a higher price. 123
Can I use GPT-6 Sol and Luna in ChatGPT?
Yes, in ChatGPT Work and Codex on Plus, Pro, Business, Enterprise and Edu plans, rolling out gradually. Free and Go users can try Luna in the desktop app. OpenAI says the models are not yet in the regular Chat view. 1

Anthropic released Claude Opus 5.5 the same day. See how it compares on price and long coding jobs.

Sources

Read for this feature. The numbers match the markers in the text.

  1. OpenAI: Introducing GPT-6 Sol and Luna, September 22, 2026openai.com
  2. OpenAI API documentation: GPT-6 Sol; checked September 22, 2026developers.openai.com
  3. OpenAI API documentation: GPT-6 Luna; checked September 22, 2026developers.openai.com
  4. Artificial Analysis: GPT-6 Sol and Luna evaluation results, posted on X, September 22, 2026x.com

Planning which parts of your work to automate, and with which model? Talk it through with us.

Where this leads

Want to apply this to a project? Here is the related service.

Free download

25 AI instructions to try on everyday work

Prompts are instructions you give an AI tool. This PDF includes 25 examples for inquiries, marketing and admin. Adapt them to your task and check the results.

Enter your email to access the PDF. This form does not sign you up for a marketing sequence.