If anyone is still spending the most money to call OpenAI’s strongest model today, OpenAI would instead advise them to switch to a different one. On July 30, OpenAI released a price adjustment announcement: GPT-5.6 Luna dropped 80% in price, and Terra dropped 20%. Seeing this news, it’s easy to focus on the price war starting in Silicon Valley. However, if you carefully read the official technical documentation and API call guidelines, you will realize that the truly noteworthy move is not the price numbers themselves.
This is essentially the first time OpenAI has told users that many tasks do not actually require the strongest model. The official gave a very specific suggestion: for a complex task, first use GPT-5.6 Sol to complete requirements analysis and solution design, then hand it over to Luna to execute, write code, and run tests. The most expensive model is responsible for thinking, and the cheapest model is responsible for doing the work. Two years ago, this strategy would have been equivalent to commercial self-denial. After all, in the past, the entire AI industry desperately told the market that its own model was the smartest. Yet today, OpenAI comes out and says there is no need to always buy the most expensive one. This is far more important than the price cuts.
1.The Unspoken Understanding of Silicon Valley's Two Giants
Let’s look at what has happened in the past two weeks. On July 30, OpenAI adjusted prices: Luna dropped 80%, Terra dropped 20%. Sol did not drop in price, but added Fast mode, which can be up to 2.5 times faster than the standard mode at double the price, with the same intelligence level. Top-tier models remain. What really begins to scale is the low-end and mid-range. A week earlier, on July 24, Anthropic did almost the same thing. Claude Opus 5 was released, priced at $5 per million input tokens and $25 per million output tokens, exactly half of Fable 5. More than performance breakthroughs, Anthropic emphasized its economy: at only half the price, you can obtain frontier reasoning capabilities infinitely close to Fable 5. Just a month ago, Fable 5 was Anthropic's flagship—strongest reasoning, longest context, highest price. Yet a month later, Anthropic itself found a half-price alternative for its flagship. If only one company did this, it could be seen as product adjustment; when two do it almost simultaneously, it is no coincidence. Both are actively reducing the importance of flagship models.
2.Flagships for Show, Volume Models for Money
I have been thinking: why now? The answer is not complicated. In the past, the greatest value of flagship models was not making money, but proving technological leadership. After GPT-4 came out, OpenAI's valuation kept rising. Every time Claude updated, Anthropic redefined its technical position. Flagship models bear brand value. But the real places where enterprises spend money are not there. An enterprise runs millions of API calls every day: customer service, search, approvals, code generation, agent execution—these high-frequency tasks account for the vast majority of token consumption. What enterprise buyers care about most is not Benchmark No. 1, but cost per task, stability, and ROI. When call scale expands to tens of millions per day, the slight intelligence advantage of flagship models is instantly erased by huge computing costs. OpenAI's actions and statements mark a turning point: top-tier flagships no longer bear the burden of making money.
3.AI Enters the "Mass-Market Model" Era
The auto industry has long experienced this. Twenty years ago, the BMW 7 Series defined BMW's height, the S-Class upheld Mercedes-Benz's luxury, and the A8 established Audi's flagship image; flagship models determined brand ceiling. But the real volume and profit base always came from the BMW 3 Series, Mercedes-Benz C-Class, and Audi A4. Later, it became more obvious: Model S proved Tesla could build cars; what really made it a global automaker was Model 3 and Model Y. Flagships prove capability; mass-market models provide scale. AI is now entering this stage. Sol and Fable will continue to exist—they are responsible for pushing technical limits and refreshing benchmarks. What truly bears commercialization will increasingly become Luna, Terra, and Opus. OpenAI even publicly wrote a recommended workflow this time: Sol is responsible for planning, Luna for execution. This is no longer a single model; OpenAI is designing a division of labor among models. What enterprises buy in the future will not be a model, but a system of models. What really determines cost is not the chief architect, but the construction crew working every day.
4.Models Begin to Optimize Models
One detail is more interesting to me than the price cut itself. OpenAI mentioned in its technical note that this price reduction was not only because it purchased more GPUs or due to scale. The real reason is that models are beginning to participate in optimizing models. Specifically, under the guidance of human engineers, Sol autonomously rewrote and optimized the underlying production kernel. It designed hundreds of experiments by itself, improved token generation efficiency, and even participated in monitoring the model training process, intervening directly when problems were found. As a result, end-to-end operating costs dropped 20%, and token generation efficiency increased 15%.
This information can be easily overlooked, but its significance is great. In the past, efficiency improvements relied on engineers. After a model went live, people gradually optimized inference frameworks, CUDA, cache strategies, and scheduling algorithms; squeezing out a dozen percentage points of efficiency improvement over a year was already good. Now, technological evolution has entered a new path: models begin to take over engineering optimization of underlying code and compute scheduling, iterating around the clock. This is a continuously accelerating loop. The smarter the model, the stronger its optimization capability; the faster optimization, the faster costs drop; the lower costs, the greater call volume; the greater call volume, the more data generated to continue training. Looking back at OpenAI's price changes over the past two and a half years: GPT-4's initial input price was $30 per million tokens, GPT-4o dropped to $5, and GPT-4o mini reached $0.15. Luna is now priced close to the cheap range of the early mini, yet its overall intelligence level far exceeds the expensive GPT-4 from two years ago. In just over two years, prices have fallen by nearly two orders of magnitude. If models continue to participate in optimizing themselves, this curve will most likely continue downward. What is truly frightening is not that prices dropped 80% today, but that cost reduction has begun to have a self-driving capability.
5.From "Who Is Smartest" to "Who Is Most Worth It"
I increasingly feel that when people discuss AI competition today, they sometimes still apply the framework of the previous phase. For example, people often still ask: who is the smartest? GPT, Claude, Gemini, DeepSeek—whenever a new model is released, the media immediately looks at the leaderboards. Whoever ranks first wins. However, the new moves by OpenAI and Anthropic break this pattern; they send a new signal to the market: the appeal of a single performance champion is fading. OpenAI mentioned a sentence in the announcement that I find particularly critical: it suggested that developers match different models based on task importance, cost of error, urgency, and scale. Notice that it is no longer discussing models, but tasks. In the past, an enterprise deploying AI mostly had only one choice. Starting today, it is more like building an organization: the most important tasks use Sol, daily execution goes to Luna, and perhaps even lighter models for simpler tasks in the future. This is very similar to cloud computing back then: no one places all data on the most expensive SSD; hot data on SSD, ordinary data on HDD, cold data on object storage. What people discuss has never been which hard drive is fastest, but how to build the entire system most cost-effectively. DeepSeek actually sensed this direction earlier than Silicon Valley. Over the past six months, it has hardly emphasized that it is necessarily the smartest in the world; instead, it repeatedly said a few words: cheap, fast enough, good enough. After cache hits, every million tokens are so cheap they are almost negligible. It has consistently bet on one thing: what enterprises need is not a world champion, but the deployment option with the best overall ROI. This is not a short-term price war; the two Silicon Valley giants’ follow-up now verifies the inevitability of this business path.
6.What OpenAI Really Wants to Sell Is Not the Model
By now, OpenAI's strategy is clear: what it really wants to sell is not models, but call volume. In the past, model vendors relied on high technological premiums for high gross margins; now, the operating logic has shifted to exchanging extremely low thresholds for super-large-scale traffic ecosystems. Microsoft really made money not because Windows was expensive, but because all computers installed Windows. AWS made money not because of high per-server profit, but because countless applications run on it every day around the world. Platform revenue has always depended on penetration. This explains why Sol's price has not moved: Sol carries the brand—it proves OpenAI is still the company with the highest technical ceiling. Luna is the revenue engine. OpenAI hopes developers will form a new default habit: use Luna for writing agents, for running workflows, and for batch execution; only call Sol when encountering truly difficult problems. Once this default is established, subsequent competition will become very difficult. Migrating the underlying model for an enterprise means retesting, validating, and adapting the entire workflow; migration costs will rise higher and higher. The real moat has begun to shift from capability leadership to ecosystem stickiness.
7.Flagship Models No Longer Determine Direction
Taking a longer view, the AI industry is undergoing a very typical inflection point of economies of scale. When the unit cost of computing drops to extremely low levels, the market's total demand for tokens does not decrease as unit prices fall—it explodes exponentially. Flagship models no longer determine direction, because the center of gravity of technological evolution has shifted from exploring the ceiling of intelligence to industrializing cost reduction in computing. When API call costs become negligible, the form and boundaries of models will gradually blur. Enterprise developers will no longer focus on the cost of each token, but will seamlessly embed AI into every business process. The most profound aspect of this shift is that the collapse of API prices is raising the migration cost of the entire software engineering stack. Once enterprise workflows, agent scheduling networks, and automated data pipelines are all built on combinations of low-cost models, the underlying model provider locks in the computing pipelines for the next decade. Over the past few years, large model companies sold intelligence. Starting today, they are beginning to sell efficiency. These are two completely different stories.
In the past, we used to analogize AI development to consumer electronics, expecting new record-breaking flagship products from time to time. But perhaps the true victorious form of AI is not to be a sensational product for a moment. When the steam engine was first invented, people marveled at its productivity; yet today, electricity flows everywhere, driving the operation of civilization, and no one specifically discusses it anymore. When humanity no longer fanatically discusses which flagship model has refreshed which IQ benchmark, AI silently embeds itself into every system and every instruction, becoming the water and electricity that are called by default behind all automated processes. Only then will its true era have just begun.



