NewsStocksGPT-6 Sol and Luna Launch With 50% API Price Cuts as Independent Tests Show Mixed Performance

GPT-6 Sol and Luna Launch With 50% API Price Cuts as Independent Tests Show Mixed Performance

Author: Metaverse Post·

Key Takeaways

  • GPT-6 Sol and GPT-6 Luna launch with API prices 50% below GPT-5.6 promotional rates, set at $2 and $0.10 per million input tokens respectively, which OpenAI attributes to caching and inference infrastructure gains.
  • OpenAI's benchmarks show Sol outperforming Claude Opus 5 on AutomationBench (33.2% versus 26.9%) at roughly 9% of its per-task cost and matching it on OSWorld 2.0 at about one-fifth the cost.
  • Independent testing by Artificial Analysis confirmed steep per-task cost declines—Sol from $1.99 to $1.06 and Luna from $0.18 to $0.07—but found capability scores roughly level, with knowledge-work regressions of about 100 Elo points for Sol and 75 for Luna on GDPval-AA v2.1.
  • Sol's hallucination rate on the AA-Omniscience benchmark fell from 92% to 60%, though part of the gain came from the model declining to answer more questions, which lowered its accuracy by 5 points.
  • Both models are available in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users and through the API as gpt-6-sol and gpt-6-luna, with Luna additionally accessible to Free and Go users via the desktop app.
GPT-6 Sol and Luna Launch With 50% API Price Cuts as Independent Tests Show Mixed Performance

OpenAI has released GPT-6 Sol and GPT-6 Luna, expanding its GPT-6 model family alongside the flagship GPT-6 Astra, which was introduced earlier this month. The company says the new models deliver performance approaching Astra’s level in professional work, factuality, coding, and computer use, while reducing API prices by 50% compared with the promotional pricing for the previous GPT-5.6 generation.

OpenAI attributed the lower prices to improvements in caching and inference infrastructure, saying the resulting savings are being passed directly to users. GPT-6 Sol now costs $2 per million input tokens and $10 per million output tokens, down from $4 and $20, respectively. GPT-6 Luna costs $0.10 per million input tokens and $0.50 per million output tokens, compared with its previous prices of $0.20 and $1.20. Because API billing scales with token volume, per-token rates are the primary cost lever for developers, and halving them compounds across the large token volumes that agent workloads can generate.

OpenAI describes Astra as its highest-capability model for the most demanding projects, while Sol and Luna are intended to make advanced AI more practical for higher-volume, everyday workloads.

GPT-6 Sol and Luna just landed in Astra’s orbit. Both launch today with API prices 50% lower than GPT-5.6. Build with Sol. Scale with Luna. To production and beyond. pic.twitter.com/ZCEFp4JdjV — OpenAI Developers (@OpenAIDevs) September 22, 2026

GPT-6 Sol and Luna just landed in Astra’s orbit. Both launch today with API prices 50% lower than GPT-5.6. Build with Sol. Scale with Luna. To production and beyond. pic.twitter.com/ZCEFp4JdjV

Benchmark results and technical improvements

On AutomationBench, which evaluates agents across 47 business tools in sales, marketing, operations, support, finance, and human resources, GPT-6 Sol at xhigh effort scored 33.2% at a cost of $0.27 per task. That compared with 26.9% for Claude Opus 5 at maximum effort, while Sol’s cost per task was approximately 9% of Claude Opus 5’s. Cost-per-task comparisons of this kind have become a standard feature of frontier model releases, sitting alongside headline benchmark scores.

On Agents’ Last Exam, Sol at maximum effort achieved a score of 56.4%, exceeding Claude Opus 5’s highest score at approximately 60% lower cost. In coding evaluations, GPT-6 Sol scored 68.8% on DeepSWE v1.1, within 1.1 percentage points of Claude Fable 5’s best result at approximately 80% lower cost. Luna scored 66.6%, comparable to Opus 5 and Fable 5 at medium effort, at 93–96% lower cost. On OSWorld 2.0, Sol matched Claude Opus 5, scoring 60.5% versus 60.3%, at about one-fifth of the cost.

OpenAI also reported that Sol’s factual error rate was cut in half relative to its predecessor in an internal evaluation of de-identified conversations in which users had flagged mistakes. At higher effort, Luna matched the reliability of GPT-5.6 Sol at roughly one-hundredth of the cost. Alignment evaluations showed both models improving over their predecessors, including lower rates of deception in deliberately challenging coding scenarios.

OpenAI also upgraded prompt caching to increase default cache hit rates. The system provides a 90% discount on cached input-token reads and includes a monitoring dashboard, diagnostic tools, and controls that allow developers to adjust reasoning effort and tool availability without invalidating cached context. GitHub reported that these changes reduced the share of prompt tokens requiring fresh processing by more than 50% across billions of requests.

GPT-6 Sol and Luna became available in ChatGPT Work and Codex for Plus Pro, Business, Enterprise, and Edu users. Luna is also available to Free and Go users through the desktop app. Both models are offered through the API as gpt-6-sol and gpt-6-luna, with a gradual rollout throughout the day.

Independent analysis finds cost gains and mixed performance

Third-party testing by Artificial Analysis broadly supported OpenAI’s efficiency claims but offered a more mixed assessment of the models’ capabilities. Using its Intelligence Index, the firm found that GPT-6 Sol at maximum effort cost approximately $1.06 per task, compared with $1.99 for its predecessor. Luna’s cost fell from $0.18 to $0.07 per task. Artificial Analysis said both models had reached the cost-efficiency Pareto frontier.

The firm said the savings came entirely from the price reduction, because both models used slightly more output tokens per task than their predecessors.

GPT-6 Sol and Luna push the cost efficiency frontier by halving cost relative to GPT-5.6 Sol and Luna. Intelligence Index and Coding Agent Index scores remain level with GPT-5.6, with progress in some evaluations and regressions in others Pricing is approximately half that of… pic.twitter.com/eubnxPVsyN — Artificial Analysis (@ArtificialAnlys) September 22, 2026

GPT-6 Sol and Luna push the cost efficiency frontier by halving cost relative to GPT-5.6 Sol and Luna. Intelligence Index and Coding Agent Index scores remain level with GPT-5.6, with progress in some evaluations and regressions in others Pricing is approximately half that of… pic.twitter.com/eubnxPVsyN

Performance results showed both improvements and regressions. Sol’s Coding Agent Index score increased by 2 points to 57, with gains on Terminal-Bench 4.0 and SWE-Atlas-QnA. Luna’s score fell by 2 points, with declines on SWE-Atlas-QnA and DeepSWE v1.1. The divergent picture relative to OpenAI’s launch-day figures partly reflects differences in benchmark suites, effort levels, and scoring methodologies between the two sets of results.

On the AA-Omniscience benchmark, Sol’s hallucination rate declined from 92% to 60%. Artificial Analysis noted that the improvement was partly attributable to the model declining to answer more questions, which reduced its accuracy by 5 points. That distinction—gains from stronger factual grounding versus gains from answering fewer questions—is a recurring nuance in hallucination measurement.

The largest regressions appeared in knowledge-work evaluations. Sol fell by approximately 100 Elo points on GDPval-AA v2.1, while Luna declined by about 75 points. Manual inspection attributed the lower scores to shorter deliverables that omitted required rubric elements and to reduced presentation quality. Read alongside OpenAI’s figures, the independent results frame this launch primarily as a cost story: pricing fell across the board, while capability movement varied by benchmark and task category.