OpenAI cut GPT-5.6 Luna 80% and Terra 20% on 30 July 2026: the LLM price war is here

OpenAI cut GPT-5.6 Luna 80% and Terra 20% on 30 July 2026. The new prices and what they mean for model spend.

Read time
12 min
Word count
1.8K
Sections
11
FAQs
8
Share
Descending glowing bar-chart columns over a dark data network grid
OpenAI's 30 July cut resets the cheap-tier math.
On this page · 11 sections
  1. What OpenAI actually changed on 30 July 2026
  2. GPT-5.6 pricing before and after the cut
  3. Why now: the price war OpenAI is answering
  4. Where GPT-5.6 now sits against the frontier
  5. What the cut changes for your model spend
  6. A worked example: what the cut saves on a real workload
  7. What enterprise buyers should do now
  8. India-specific considerations
  9. FAQ
  10. How eCorpIT can help
  11. References

Summary. On 30 July 2026 OpenAI cut the price of GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20%. Luna now costs $0.20 per million input tokens and $1.20 per million output tokens; Terra costs $2 and $12. GPT-5.6 Sol kept its $5 and $30 rate but gained a Fast mode that runs up to 2.5 times quicker than standard at twice the price, with no change in intelligence. OpenAI says the cuts came from efficiency work that reduced the end-to-end cost of serving the models by about 20%. This is the move we anticipated in June, when the Wall Street Journal reported OpenAI was weighing cuts as Anthropic's annualised revenue climbed toward a reported $47 billion and Anthropic passed OpenAI on business adoption. The rumour is now a price change you can act on. The short version: Luna at $0.20 input is now one of the cheapest capable models available, which makes tiered routing more valuable than ever, and the right response is to re-run your own cost math rather than assume your current model choice still wins.

What OpenAI actually changed on 30 July 2026

The specifics matter, because two of the three GPT-5.6 tiers changed and one did not. Per OpenAI's developer announcement, Luna, the fastest and cheapest tier, dropped 80% to $0.20 per million input tokens and $1.20 per million output tokens. Terra, the balanced everyday tier, dropped 20% to $2 input and $12 output. Sol, the frontier reasoning tier, held at $5 input and $30 output.

Sol instead gained speed. Its new Fast mode delivers up to 2.5 times faster responses than standard processing at twice the price, with no change in intelligence, and requests tagged as priority automatically use it. So Sol buyers now choose between standard latency at the standard price and roughly 2.5 times faster output at double the cost.

The cuts were not a margin giveaway. OpenAI attributes them to engineering: kernel-level work reduced the end-to-end cost of serving the models by about 20%, and further experiments increased token-generation efficiency by more than 15%. That framing matters for planning, because efficiency-driven cuts tend to hold, where promotional cuts often reverse. The GPT-5.6 series itself, Sol, Terra, and Luna, launched on 9 July 2026, so this reprice came three weeks after general availability, per OpenAI's GPT-5.6 pricing page.

GPT-5.6 pricing before and after the cut

GPT-5.6 tier Input ($/M) before Input ($/M) after Output ($/M) before Output ($/M) after Change
Luna (fast, cheapest) ~$1.00 $0.20 ~$6.00 $1.20 -80%
Terra (balanced default) ~$2.50 $2.00 ~$15.00 $12.00 -20%
Sol (frontier reasoning) $5.00 $5.00 $30.00 $30.00 No cut; new Fast mode (2.5x speed, 2x price)

The "before" figures for Luna and Terra are implied by OpenAI's stated percentage cuts against the current published prices; the "after" figures and percentages are what OpenAI announced. The takeaway is the shape of the change, not decimal precision: Luna became five times cheaper, Terra came down modestly, and Sol traded a price cut for a speed option.

Why now: the price war OpenAI is answering

This cut did not happen in isolation. Through the first half of 2026, Anthropic went from challenger to the company setting the pace. Its annualised revenue rose from roughly $1 billion in January 2025 to a reported $47 billion by May 2026, and the Ramp AI Index for May 2026 put Anthropic ahead of OpenAI on business adoption for the first time, at 34.4% against 32.3%. Anthropic then launched Claude Fable 5 and Mythos 5 on 9 June 2026 at $10 and $50 per million tokens, with Fable 5 posting 80.3% on SWE-Bench Pro. That combination of aggressive pricing and coding-benchmark strength is what pushed enterprise buyers to reconsider defaults.

OpenAI CEO Sam Altman had already signalled the direction, saying publicly that "costs had become a huge issue" for business buyers and that the company would find ways to deliver more value for less spend. Both companies are also moving toward public listings, and each filed confidentially for an initial public offering during this period. Cutting the price of the two cheaper GPT-5.6 tiers is a direct answer to that pressure: it defends the high-volume, cost-sensitive workloads most exposed to a cheaper competitor, while keeping the frontier Sol tier priced for margin.

Where GPT-5.6 now sits against the frontier

Model Input ($/M) Output ($/M) Position
GPT-5.6 Luna $0.20 $1.20 Cheapest OpenAI tier after the cut
GPT-5.6 Terra $2.00 $12.00 Balanced OpenAI default
GPT-5.6 Sol $5.00 $30.00 OpenAI frontier reasoning
Claude Fable 5 $10.00 $50.00 Anthropic frontier (launched June 2026)
Claude Sonnet 4.5 $3.00 $15.00 Anthropic production default
Claude Haiku 4.5 $0.80 $4.00 Anthropic cheap tier

Anthropic prices are as reported around their mid-2026 availability and should be checked against Anthropic's current rate card before you commit, since this market reprices often. Read against that field, the cut lands hardest at the bottom. Luna at $0.20 input and $1.20 output is now cheaper than Claude Haiku 4.5 on both sides, which changes the default for high-volume, simple traffic. Terra at $2 and $12 sits below Claude Sonnet 4.5 on input and above it on output, so the choice there is genuinely workload-dependent. Sol stays the premium option and now competes on speed rather than price.

What the cut changes for your model spend

The instinct after a price cut is to switch to the cheaper model everywhere. That is usually wrong. The better response is to re-run the routing math, because the cut changes where the break-even points sit, not which architecture wins.

Luna became a serious cheap-tier default. At $0.20 input, sending the 60% to 80% of production traffic that is classification, extraction, short answers, and routing decisions to Luna now costs a fraction of what a frontier call costs. If you were using a competitor's cheap tier for that traffic, Luna is worth a benchmark on your own prompts. This is the same tiered-routing logic covered in the hybrid model routing framework: match each request to the cheapest model that clears your quality bar, and let a gateway enforce it.

Terra's smaller cut rewards measurement over assumption. A 20% reduction is real but not transformative, so whether Terra or Claude Sonnet 4.5 is cheaper for your workload depends on your input-to-output ratio. Output-heavy workloads, such as long generations, feel Terra's $12 output rate more than input-heavy retrieval-augmented workloads do. Model this against your actual token mix rather than the headline number, using the same discipline as our GPT-5.6 inference cost analysis.

Sol's Fast mode is a latency lever, not a cost lever. Paying double for up to 2.5 times faster output only makes sense where response time drives revenue or user retention, such as interactive coding or live agents. For batch and background work, standard Sol is the cheaper and correct choice. Deciding which Sol mode a workload deserves is now part of tier selection across GPT-5.6.

A worked example: what the cut saves on a real workload

Numbers make the point better than percentages. Take an application handling 10 million requests a month, averaging 500 input tokens and 300 output tokens per request, with tiered routing sending 70% of traffic to Luna and 30% to Terra. Using OpenAI's post-cut prices, and the pre-cut prices implied by the stated reductions, the monthly bill looks like this.

Tier and share Before the cut After the cut
Luna (70%, 7M requests) $16,100 $3,220
Terra (30%, 3M requests) $17,250 $13,800
Total per month $33,350 $17,020

The Luna line does the heavy lifting. Its 3,500 million input tokens fall from about $3,500 to $700, and its 2,100 million output tokens fall from about $12,600 to $2,520, because the 80% cut applies to the tier carrying most of the volume. Terra's smaller 20% cut trims its line from $17,250 to $13,800. The workload's total bill drops roughly 49%, from $33,350 to $17,020 a month, which is about ₹27.7 lakh falling to ₹14.1 lakh at a rate near ₹83 to the dollar.

Two lessons sit inside that table. First, the savings concentrate where the cheap, high-volume tier runs, so a routing design that pushes simple traffic to Luna captures most of the benefit. Second, output tokens dominate the bill on both tiers, so prompt and response design that trims output length moves the number as much as the price cut did. Run this same calculation on your own token mix before you change anything.

What enterprise buyers should do now

Three concrete steps follow from this cut. First, re-benchmark your cheap tier: if Luna at $0.20 beats what you run today on your own prompts, the migration usually pays for itself quickly on high-volume paths. Second, avoid long lock-ins at the moment: this market is repricing on a monthly cadence, and Anthropic is likely to respond, so quarter-by-quarter terms or explicit benefit-of-future-cuts clauses protect you. Third, instrument before you optimise: you cannot route to the cheapest capable model without measuring cost per request, so a gateway or spend tracker is the prerequisite, not an afterthought, and it is the highest-use first step in any LLM cost programme.

The larger pattern holds across the frontier model comparison: capability is converging while price competition intensifies, so the durable advantage is not picking one model forever but building a system that can move workloads to whichever tier wins this quarter.

India-specific considerations

For Indian teams, the Luna cut is unusually valuable. Cost-sensitive engineering teams here already design AI-heavy workflows, and an 80% reduction on the cheap tier flows straight to unit economics for products that run many small model calls per user. A workload that sends a high share of traffic to Luna at $0.20 input now carries a materially lower rupee cost than it did in early July.

Two cautions apply. Token counts differ by language: Hindi, Tamil, Bengali, and other non-Latin scripts tokenise less efficiently than English, so the same task costs more tokens, and you should benchmark on your actual multilingual prompts rather than trusting an English-based estimate. And where prompts carry personal data, the Digital Personal Data Protection Act 2023 governs how that data is processed, so a cheaper model does not remove the need to design data handling correctly; applications should be designed aligned with DPDP requirements regardless of which provider is cheapest this month.

FAQ

How eCorpIT can help

eCorpIT helps teams turn a moving price war into a controlled AI bill. Our senior engineering teams benchmark models on your own production prompts, design tiered routing so simple traffic lands on cheap tiers like GPT-5.6 Luna while hard tasks reach the right frontier model, and instrument cost per request so you can see the impact of every reprice. If you want to re-model your LLM spend after this cut or plan a migration, see our LLM migration and cost-optimization service and reach us through our contact page.

References

  1. OpenAI Developer Community, "Announcing a major price drop for 5.6 Terra and Luna and Fast mode for 5.6-Sol": community.openai.com
  1. OpenAI, "Advancing the price-performance frontier with GPT-5.6": openai.com
  1. OpenAI API, "Pricing": developers.openai.com/api/docs/pricing
  1. CNBC, "OpenAI cuts prices for two of its GPT-5.6 AI models as companies grow sensitive to costs" (30 July 2026): cnbc.com
  1. OpenAI Help Center, "A preview of GPT-5.6 Sol, Terra, and Luna": help.openai.com
  1. CNBC, "OpenAI mulls slashing prices as it competes with Anthropic for users: WSJ" (11 June 2026): cnbc.com
  1. Anthropic, "Claude Fable 5 and Mythos 5": anthropic.com
  1. VentureBeat, "Anthropic finally beat OpenAI in business AI adoption": venturebeat.com
  1. SaaStr, "Anthropic Just Passed OpenAI in Revenue": saastr.com
  1. Reuters via Yahoo Finance, "OpenAI considers drastic price cuts, anticipating war for users with Anthropic": finance.yahoo.com

_Last updated: 31 July 2026._

Frequently asked

Quick answers.

01 How much did OpenAI cut GPT-5.6 prices on 30 July 2026?
OpenAI cut GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20%. Luna now costs $0.20 per million input tokens and $1.20 per million output tokens, and Terra costs $2 input and $12 output. Sol kept its $5 and $30 pricing but added a faster processing mode.
02 Did GPT-5.6 Sol get cheaper?
No. Sol held its price at $5 per million input tokens and $30 per million output tokens. Instead it gained a Fast mode that runs up to 2.5 times faster than standard processing at twice the price, with no change in the model's intelligence. Requests tagged as priority automatically use Fast mode.
03 Why did OpenAI cut GPT-5.6 prices?
OpenAI attributes the cuts to efficiency work that reduced the end-to-end cost of serving the models by about 20%, with token-generation efficiency up more than 15%. The competitive backdrop matters too: Anthropic's revenue and enterprise adoption had surged through mid-2026, and both firms are preparing for public listings.
04 Is Luna now the cheapest capable model?
At $0.20 per million input tokens and $1.20 output, Luna is among the cheapest models from a major lab, and after the cut it undercuts Anthropic's Claude Haiku 4.5 on both input and output. Whether it is right for your workload depends on quality on your own prompts, so benchmark before switching production traffic.
05 Should I switch all my traffic to Luna?
No. The cut changes break-even points, not the value of tiered routing. Send simple, high-volume tasks to Luna, keep harder reasoning on Terra or Sol, and route each request to the cheapest model that clears your quality bar. Switching everything to the cheapest tier usually degrades quality on the tasks that need a stronger model.
06 How does GPT-5.6 Terra compare with Claude Sonnet 4.5 now?
After the 20% cut, Terra costs $2 input and $12 output, against Claude Sonnet 4.5 at a reported $3 and $15. Terra is cheaper on input and, depending on your output volume, can be cheaper overall. Because the outcome depends on your token mix, model it on your real workload rather than the headline rates.
07 Will Anthropic respond to this cut?
Very likely. The 2026 price war has been bilateral, with each lab responding to the other, and both are heading toward public listings where growth and margin narratives matter. Expect pricing or capability moves from Anthropic, so avoid long lock-ins and keep your architecture able to shift workloads between providers.
08 What is the safest way to plan around frequent price changes?
Instrument cost per request, route to the cheapest capable model through a gateway, and keep contracts short. Because prices are moving monthly, a system that can switch workloads between tiers and providers protects you better than a single fixed model choice. Re-run your cost math each quarter as new rates land.

About the author

Manu Shukla

Founder & Director

Founder of eCorpIT. Hands-on engineer leading senior-only delivery for AI apps, custom software, and cloud systems for global clients.

Subscribe

One engineering note a week. No fluff, no spam.

Senior-architect playbooks on AI agents, mobile apps, cloud, security, data, and marketing — delivered every Wednesday.

Past the reading

Read enough. Let's build something.

A senior architect responds in 24 working hours with scope, indicative cost, and a timeline. NDA before any technical conversation.