
OpenAI just cut the API price of GPT-5.6 Luna by 80%.
Luna moved from $1 per million input tokens and $6 per million output tokens to $0.20 and $1.20. GPT-5.6 Terra fell 20%, from $2.50/$15 to $2/$12. Sol did not get cheaper, although OpenAI added an API Fast mode that can reach 2.5× Standard speed at twice the price.
The paid ChatGPT and Codex subscription prices also did not change. Instead, Luna and Terra now consume fewer credits inside those products. OpenAI says the cut came from improvements across the models, inference stack, and agent harness: the same work now takes fewer tokens and less compute, and some of that saving is being passed on to customers.
That is already a notable update three weeks after the GPT-5.6 launch. But the more important story is what it does to the rest of the market.
The middle of the market just collapsed
At $0.20/$1.20, Luna is no longer merely a cheaper frontier-family model. It has dropped into a price range where every “good enough, reasonably priced” model has to defend its reason for existing.

The direct token comparison is stark. Gemini 3.6 Flash costs $1.50/$7.50. Claude Sonnet 5 is currently $2/$10 under its introductory pricing. The global Z.ai API lists GLM-5.2 at $1.40/$4.40. These products are not interchangeable, but Luna now starts from a radically lower cost base.
The capability gap is also smaller than the price gap suggests. On the Artificial Analysis Intelligence Index, Luna at max reasoning and GLM-5.2 at max both score around 51, while Gemini 3.6 Flash sits around 50. Using Luna’s new token rates against the benchmark’s measured token mix puts its weighted task cost at roughly $0.06, compared with about $0.27 for GLM-5.2 and close to $0.50 for Gemini.
The six-cent figure is a weighted benchmark workload, not a universal per-task bill. Gemini still handles more input types and Claude has different agent strengths, but neither changes the central result: Luna now offers comparable benchmark intelligence from a dramatically lower cost base.
Still, the direction is hard to miss. Similar benchmark intelligence is now available at a very different price.
Why GLM-5.2 is in the awkward position
For readers outside China, GLM-5.2 needs a little context. Zhipu AI operates a global service under the Z.ai brand and a China-facing service under BigModel. GLM-5.2 itself is a strong open-weight, text-only model built around long-horizon coding and a 1M-token context window. Its global API price is $1.40/$4.40; the China API lists ¥8/¥28.
GLM’s problem is not that it suddenly became a bad model. The problem is that its old pitch combined three things: strong coding, open weights, and a low operating price. Luna now attacks the third point while adding image input.
The subscription story makes the timing worse. On the China-facing BigModel platform, the monthly GLM Coding Plan moved from ¥49/¥149/¥469 to ¥118/¥538/¥1,078 for Lite, Pro, and Max. These are Chinese yuan prices—renminbi, or RMB—not U.S. dollars. Put another way, Lite rose by about 2.4×, Pro by 3.6×, and Max by 2.3×. BigModel did make the new system more transparent by publishing five-hour and weekly point pools, but the headline for Chinese developers is still a steep price increase.
Z.ai’s global plans now cost $18, $72, and $160 a month. They are coding subscriptions rather than general API credits, but that detail does not rescue the pricing story. In both China and the global market, GLM has moved toward premium-tool pricing at exactly the moment Luna has moved downmarket.
That is the squeeze. GLM-5.2 remains a capable open model, but buyers no longer get a simple bargain: its API costs more than Luna, its coding subscription costs much more than it used to, and the GLM brand does not yet command the same flagship pull as the top closed models. Customers now have to choose GLM because they specifically want GLM—not because it is the obvious value option. It does retain one practical edge: its open-weight, less restrictive deployment can be valuable in cybersecurity work, as Hugging Face showed when it used GLM-5.2 for incident forensics after hosted models rejected real exploit data.
Not every Chinese model is hurt in the same way
Kimi K3 is fighting a different battle. It is aimed directly at GPT-5.6 Sol and Claude Fable 5, not at Luna. Its official API is $3 for cache-miss input and $15 for output, with a 1M context window, native visual understanding, and open weights.
I have not used K3 myself, but the community praise around it is unusually consistent: it is excellent at front-end generation, surprisingly strong at building playable games, open-weight, and one of the few coding models repeatedly praised for visual taste. Those four strengths give it a concrete identity. K3 is not merely claiming flagship benchmark numbers; it is producing interfaces and games that look intentionally designed.
That clearer ambition protects K3 from Luna’s price cut. If a developer wants the cheapest high-volume model, Luna and DeepSeek are the natural comparison. If they want an open, multimodal, aesthetically stronger flagship that can stand alongside Sol and Fable 5, K3 has a coherent reason to exist. GLM-5.2 currently sits between those stories—too expensive to own the cost floor, but not differentiated enough to own the flagship ceiling.
DeepSeek owns the opposite end. DeepSeek V4 Flash is $0.14/$0.28. V4 Pro charges $0.435 for cache-miss input, $0.003625 for cache hits, and $0.87 for output. Luna can be cheaper on ordinary input and may deliver more intelligence per task at some reasoning levels, but DeepSeek’s cache economics and output price remain extremely difficult to beat.
So the Chinese model market is not one block:
- Kimi K3 protects the high end with flagship performance and native multimodality.
- DeepSeek protects the cost floor with aggressive output and cache pricing.
- GLM-5.2 is caught between them, just as Luna’s price moves down into its territory.
That is why I think GLM-5.2 feels this cut most sharply. OpenAI did not merely make Luna cheaper. It removed much of the space where a capable mid-priced model could win by being “almost frontier, but more affordable.”
The broader lesson is the same one DeepSeek has been pushing for a long time: capability matters, but cost is itself a capability. Once models are good enough to complete the work, the next competition is over how cheaply and reliably they can keep doing it.
Sources: OpenAI’s price-cut announcement, GPT-5.6 Luna model page, Gemini 3.6 Flash, Claude Sonnet 5, Z.ai API pricing, Z.ai Coding Plan, BigModel pricing, GLM-5.2, DeepSeek API pricing, Kimi K3, Hugging Face’s July 2026 security incident disclosure, and Artificial Analysis.