⚡ Uncle Cat AI Radar
IndustryModels

OpenAI cuts GPT-5.6 Luna price 80% and Terra 20%

OpenAI slashed API pricing across its GPT-5.6 line, dropping Luna to $0.20 per million input tokens and adding a paid Fast mode for Sol.

OpenAI announced a sharp reduction in API pricing for its GPT-5.6 family on Thursday, cutting the cost of the entry-tier Luna model by 80 percent and the mid-tier Terra by 20 percent, effective immediately.

What changed

Luna now bills at $0.20 per million input tokens and $1.20 per million output tokens, down from roughly $1.00 and $6.00. Terra moves to $2.00 and $12.00 per million input and output tokens respectively. OpenAI also introduced a Fast processing mode for GPT-5.6 Sol in the API, offering up to 2.5 times the throughput of standard processing at twice the standard price — a rare instance of the company selling latency as a separate product tier rather than folding it into the base rate.

The company attributed the reduction to efficiency work spanning three layers: the model itself, the inference stack, and the agentic harness. It said GPT-5.6 Sol was used to optimise GPU-level software, trimming its own serving costs by about 20 percent, while speculative decoding improvements lifted token generation throughput by more than 15 percent. OpenAI framed the combined effect in workload terms, saying a task that cost roughly a dollar on comparable models a year ago now runs at about six cents on Luna and completes substantially faster.

OpenAI simultaneously moved the automatic code-review feature in the ChatGPT app and the Codex CLI from GPT-5.4 to GPT-5.6 Luna, saying the model swap combined with the new price should cut auto-review costs sharply. Cognition updated its FrontierCode 1.1 leaderboard within hours to reflect the new economics, placing the GPT-5.6 series on the cost-performance Pareto curve. The Arena leaderboard also repriced the models the same day.

Why it matters

Inference pricing has become the main competitive front between US and Chinese labs, with open-weight Chinese models from Zhipu, Moonshot and DeepSeek undercutting Western APIs by an order of magnitude. An 80 percent cut on a general-purpose tier compresses that gap directly and resets the build-versus-buy calculation for agentic products, where token consumption scales with task length rather than user count. For developers running long-horizon agents, the change is closer to a structural shift in unit economics than a routine discount.

Sources