OpenAI Cuts GPT-5.6 Luna and Terra Prices, Adds Fast Mode for Sol

CodexView original changelog

OpenAI cut API pricing on two of its three GPT-5.6 models roughly three weeks after their general availability, dropping GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20%, while flagship GPT-5.6 Sol pricing stays unchanged. OpenAI also introduced Fast mode for Sol in the API, delivering up to 2.5x the throughput of Standard processing at 2x the price, replacing the previous Priority Processing tier. Because Codex and ChatGPT Work meter usage against these same per-token rates, the price cuts translate directly into more usage headroom for Codex users at existing subscription tiers, with quota budgets otherwise unchanged.

Key Takeaways

  • GPT-5.6 Luna drops 80% in price β€” from $1.00/$6.00 to $0.20/$1.20 per million input/output tokens, the steepest cut in the announcement.
  • GPT-5.6 Terra falls 20% β€” from $2.50/$15.00 to $2.00/$12.00 per million tokens, while flagship Sol pricing is untouched.
  • Fast mode replaces Priority Processing for Sol in the API, offering up to 2.5x throughput at roughly 2x the standard price.
  • Codex and ChatGPT Work usage goes further β€” OpenAI confirmed Luna/Terra usage now consumes less quota, with no subscription price or quota-budget changes.
  • The cuts were driven by GPT-5.6 Sol optimizing its own inference stack, rewriting production GPU kernels and its speculative-decoding draft model to cut serving costs.
  • The reductions land roughly three weeks after GPT-5.6's July 9 general availability, signaling OpenAI is moving quickly on cost competitiveness versus rival model providers.

Luna and Terra Get Major Price Cuts

Three weeks after launching the GPT-5.6 model family, OpenAI announced steep API price reductions for its two lower-cost tiers. GPT-5.6 Luna, the fastest and cheapest model in the lineup, drops 80% in price β€” from $1.00 / $6.00 per million input/output tokens down to $0.20 / $1.20. GPT-5.6 Terra, the balanced everyday model, falls 20% β€” from $2.50 / $15.00 down to $2.00 / $12.00 per million tokens. Pricing for the flagship GPT-5.6 Sol model remains unchanged at $5.00 / $30.00 per million tokens.

OpenAI attributed the reductions to efficiency work carried out during GPT-5.6's own development cycle. According to an engineering post published a day earlier, Sol was tasked with optimizing its own serving infrastructure: it rewrote production GPU inference kernels, cutting end-to-end serving costs by roughly 20%, and redesigned the model's speculative-decoding draft model, improving token-generation efficiency by more than 15%. OpenAI framed the price cuts as a direct pass-through of those efficiency gains rather than a promotional discount.

Fast Mode Replaces Priority Processing for Sol

Alongside the price cuts, OpenAI introduced Fast mode for GPT-5.6 Sol in the API. Fast mode delivers up to 2.5x the throughput of Standard processing, priced at roughly double the standard rate ($10 / $60 per million input/output tokens). It replaces OpenAI's earlier Priority Processing offering, giving API customers a straightforward way to trade cost for lower latency on Sol without any change in model intelligence or output quality.

What This Means for Codex and ChatGPT Work

Because Codex and ChatGPT Work bill usage against the same underlying per-token model rates, the Luna and Terra price cuts translate directly into cheaper usage for Codex users. OpenAI confirmed that subscription prices and quota budgets for Codex and ChatGPT Work are not changing, but that Terra and Luna usage now consumes less of a user's quota than before β€” effectively stretching existing plans further for tasks routed to the cheaper tiers. For teams running high-volume or long-horizon agentic workloads on Luna or Terra inside Codex, this is a meaningful reduction in effective cost per task with no action required on their part.