Advancing the price-performance frontier with GPT-5.6
Quick Answer
OpenAI's GPT-5.6 models, Luna and Terra, now offer up to 80% and 20% price reductions, respectively, enhancing efficiency and performance for businesses.
Quick Take
Luna achieves nearly 99% lower costs per task compared to Fable 5, while Fast mode in the API boosts speeds by up to 2.5x for GPT-5.6 Sol.
Key Points
- GPT-5.6 Luna now costs 80% less, Terra 20% less, enhancing affordability.
- Luna delivers performance at 6 cents per task, nearly nine times faster than previous models.
- Fast mode in the API offers 2.5x speed increase for GPT-5.6 Sol at double the price.
- Businesses can optimize workflows by balancing intelligence and cost with Luna and Sol.
- New pricing allows for economical high-volume AI applications across various sectors.
DeepSignal Analysis
What happened
OpenAI has introduced GPT-5.6 models, Luna and Terra, with significant price reductions of 80% and 20%, respectively. Luna is designed for high-volume tasks, offering nearly 99% lower costs per task compared to Fable 5. Additionally, a new Fast mode in the API for GPT-5.6 Sol increases processing speeds by up to 2.5 times.
Key evidence
- GPT-5.6 Luna will now cost 80% less, while Terra will cost 20% less, enhancing affordability for businesses.
- Luna achieves nearly 99% lower costs per task compared to Fable 5, making it a cost-effective option for high-volume work.
- Fast mode for GPT-5.6 Sol delivers speeds up to 2.5 times faster than Standard processing, improving efficiency in API usage.
Why it matters
These updates reflect OpenAI's ongoing efforts to enhance the efficiency and affordability of AI technologies. By reducing costs and improving processing speeds, businesses can leverage AI more effectively, potentially leading to broader adoption across various sectors. The ability to optimize workflows with different models allows companies to tailor their AI usage to specific needs, which could drive innovation and productivity.
📖 Reader Mode
~5 min readYesterday, we shared how GPT‑5.6 helped make itself more efficient to run. Today, we’re passing those gains on to customers with lower prices for GPT‑5.6 Luna(opens in a new window) and Terra(opens in a new window) and faster performance with GPT‑5.6 Sol in the API. Together, these updates help customers get more from every dollar they invest in AI and move faster when time matters.
Starting today, GPT‑5.6 Luna, our fastest and most affordable model, will cost 80% less, while GPT‑5.6 Terra, our balanced model for everyday work, will cost 20% less. These lower prices for Luna and Terra are also reflected in how usage is counted against paid subscriptions when using Codex and ChatGPT Work. Luna gives businesses a far more cost-effective way to handle high-volume work at very high levels of quality. It can use tools and complete multi-step workflows, making a broader range of AI applications practical to run at scale.
Making advanced intelligence more abundant and affordable is central to OpenAI’s mission to ensure AGI benefits all of humanity. These changes put that commitment into practice. They reflect years of improvements in how our models are built, served, and put to work.
We’re also introducing Fast mode in the API, which replaces our Priority Processing offering. For GPT‑5.6 Sol, Fast mode now delivers up to 2.5× faster speeds than Standard processing at twice the price, with no change in intelligence. Fast mode is backward compatible: requests tagged priority will automatically use Fast mode.
1 of 6
Matching intelligence to the outcome
Using AI efficiently begins with the outcome. The stakes, cost of error, urgency, and scale determine the right balance of intelligence, speed, reliability, and cost. That balance can change from one step of a workflow to the next.
GPT‑5.6 gives businesses much more room to optimize that equation. Luna delivers performance comparable to models that were frontier-class a year ago at roughly 6 cents on the dollar per task, and at nearly nine times the speed. On professional work, as measured by Agents’ Last Exam, Luna outperforms Fable 5 at an estimated cost per task nearly 99% lower.
In practice, businesses can define the outcome and quality standard they need, then use evaluations to determine where additional intelligence materially improves the result and where faster, lower-cost processing can deliver the same quality. A coding workflow, for example, might use Sol to resolve uncertainty and define the plan, then use Luna to implement well-specified changes, write and run tests, and evaluate the results. Another workflow may call for a different balance.
The GPT‑5.6 family expands the range of those choices. Businesses can apply the maximum useful intelligence at every stage while paying the right price for the value it creates.
Delivering that flexibility starts with making every layer behind the models more efficient.
How we advance the efficiency frontier
Our efficiency edge comes from improving the models, the inference systems that run them, and the agentic harness that connects them to tools and context. GPT‑5.6 models take a more direct path through work. Better routing keeps hardware productive, optimized production software generates tokens more efficiently, and smarter context management helps agents avoid repeating completed work. Together, these improvements let us complete more useful work with the same compute, reducing the time, tokens, and cost required for each result.
GPT‑5.6 Sol is increasingly helping us find and deliver the next round of gains. Within a human-led process, Sol autonomously rewrote and optimized production kernels, designed and ran hundreds of experiments to improve token generation, and monitored training, intervening when problems arose. The kernel work helped reduce the end-to-end cost of serving the model by 20%, while its experiments increased token-generation efficiency by more than 15%. This work continues, creating a tighter feedback loop: as our models improve and are able to work more autonomously, our ability to improve efficiencies accelerates. Read more about the engineering behind GPT‑5.6.
A compute strategy built for scale
Meeting demand for abundant intelligence requires both more compute and more productive compute. We are building a resilient infrastructure portfolio and matching each workload to the systems best suited to run it. That approach supports both ends of the price-performance curve. At the lower-cost end, the new Luna and Terra prices make high-volume work economical at much greater scale. At the frontier end, Fast mode gives API customers faster access to Sol when response time is important.
Enterprises can move more AI into everyday operations without sacrificing speed on their most consequential work. Large-scale document analysis, customer-interaction classification, and routine implementation can become economical to run broadly, while complex Sol workloads can move faster when the premium is justified.
The gains can compound. Within a human-led process, more capable models help our technical team find the next generation of improvements, shortening the path to better performance and lower costs. Our strategy remains focused on advancing both capability and efficiency so each generation of intelligence can accomplish more work at a lower cost.
Availability and pricing
GPT‑5.6 Terra and Luna remain available in ChatGPT Work, Codex, and the OpenAI API. In ChatGPT Work and Codex, Free and Go users can access Terra, while Plus, Pro, Business, and Enterprise users can choose Terra and Luna.
Starting July 30, API pricing is $2 per million input tokens and $12 per million output tokens for Terra, and $0.20 per million input tokens and $1.20 per million output tokens for Luna. Sol pricing remains unchanged. ChatGPT and Codex subscription prices and quota budgets remain unchanged, while Terra and Luna usage now consumes fewer credits. Pricing changes will begin rolling out in AWS later today.
Fast mode for GPT‑5.6 Sol replaces Priority Processing in the API and aligns with /fast in Codex. Existing API requests tagged priority will continue to work. View complete API pricing details.
— Originally published at openai.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from OpenAI Blog
See more →Scientific computing in the age of agentic AI
AI agents are transforming scientific computing by streamlining software development, enabling researchers to focus on discovery. Projects using Codex and Claude Code report accelerated development and improved maintenance, though challenges in validating AI outputs remain. Long-term stewardship of research software is crucial to ensure reliability and reproducibility.