
Large Models Are a Pharmaceutical Business with a Ten-Month Patent Period
In the recent speech by DeepSeek's Liang Wenfeng that has been circulating, there is an interesting data point that happens to intersect with my recent thoughts. Therefore, I will write down some predictions about the entire large model industry for future falsification and correction of my thinking.
He said that DeepSeek's API pricing standard is: buy a batch of equipment, and recover the cost in ten months. That is, the profit earned in ten months can cover the hardware investment. By the way, a ten-month payback period and a six-fold profit are the same thing. Because a six-fold profit means that the GPU is depreciated over sixty months, or five years.
In a previous article, I proposed the $/Watt paradigm and calculated that NVIDIA takes about $35/W from the industry chain based on $/W. Using this scale and the ten-month payback benchmark, DeepSeek's current overall $/W performance is $35 ÷ 10 × 12 ≈ $42/W·yr. I mentioned earlier that Anthropic is above $70 (and may even reach $200, so it can afford to rent X.ai's $50/W data center), and the silicon profitability of cloud vendors is between $20 and $30. Although DeepSeek cannot afford to rent a data center like X.ai's, its profitability already exceeds that of cloud vendors. For DeepSeek's investors, especially in China, this is already a good business.
Why Ten Months
Let's return to the number ten months that Liang Wenfeng mentioned. In his speech, he described it as DeepSeek's own decision. But what I want to prove here is that this is actually caused by the entire industry structure and is not controlled by DeepSeek. It is determined by the current pace of evolution of large models.
Currently, the general rhythm of the entire industry is that regardless of the size of the model's niche, a model will inevitably have a cheaper and more capable substitute appear ten months after its release. This substitute could be from a competitor or from the company itself. Because even if others don't come to revolutionize it, the company itself will actively do so.
Looking at DeepSeek's own example is clear enough. V3.2-Exp was released on September 29, 2025, and on the same day, API prices were reduced by more than 50%, with the output price dropping to 3 yuan per million tokens. The interface for the previous generation V3.1-Terminus was only temporarily retained until October 15—two weeks. On December 1, the official version of V3.2 was launched, and by mid-2026, V4 was already in service.
In this, if we want to conduct attribution analysis, the reasons are of course multifaceted: architectural improvements, new training data and training methods, or even purely the result of scaling laws—larger models are now more efficient. But attribution analysis is not very important here; what is important is that this is a structural result, not the decision of a single company.
Look at OpenAI's models and Anthropic's models; they follow roughly the same rhythm. Every ten months, the model loses its "halo." Although GPT-4 has been on the API for over three years, the time it could command a premium was much shorter (prices dropped after eight months). A model can serve traffic for a long time, but its price will continue to decline after ten months until it is retired.
What exactly is this "ten months"? In fact, it can be characterized as the model's "pricing power"—after ten months, you can no longer charge at the price set at launch. This is due to competition and the industry chain; if you don't lower the price, the model will no longer be attractive. So ten months appears to be a choice, but in reality, it is not a choice; it is a constraint derived from the industry rhythm. The payback period of your model must be shorter than the premium window. Otherwise, by the time your equipment hasn't paid for itself, the pricing power is already gone, and the investment is not viable.

how price collapse
The Isomorphism of the Pricing Power Window and Moore's Law
The length of the pricing power for large models spans two different structures. Let's first discuss its determinants.
As mentioned earlier, this pricing power window is not determined by a single manufacturer. Structurally, it's like how Moore's Law is not a decision by Intel to double integration every eighteen months. It is driven by the entire industry. This is contrary to Liang Wenfeng's narrative. So some readers will surely ask, what facts support your view?
GPT-4 was launched in March 2023 at $30/$60; in November of the same year at DevDay, GPT-4 Turbo was launched at $10/$30, with input costs reduced by three times and output by half. Eight months.
Gemini 1.5 Pro was released at I/O in May 2024, and the price adjustment effective October 1 saw input costs drop by 64% and output by 52%. Four and a half months.
Claude 3 Opus was launched in March 2024 at $15/$75; on June 20, Claude 3.5 Sonnet was released at $3/$15, surpassing Opus on most benchmarks. Three and a half months. The Opus 4 series repeated the same thing: $15/$75 in May 2025, and on November 24, Opus 4.5 dropped to $5/$25, a 67% reduction. Six months.
o3 is the most extreme case. Launched in April 2025, on June 10, the official price was directly reduced by 80% to $2/$8. Two months.
Secondly, the history of Moore's Law provides a lesson with the same structure.
Moore's Law was never a physical law; it was a timetable jointly adhered to by the industry. Companies that couldn't keep up were left behind by history and the industry chain.
The most direct example is DEC's Alpha. This was one of the most powerful microprocessors in the 1990s, architecturally领先, but DEC did not have the ability to sustain continuous process investment with the revenue from a single company's products. Alpha's technological advantages could not be translated into iteration speed, and it exited by the late 1990s. Also disappearing during the same period were Motorola's 88000, Sun's SPARC in the general market, and MIPS's position in workstations—not because their designs were poor, but because they couldn't keep up with the density improvements per unit of time.
Even more remarkable is that the industry chain will eventually actively homogenize the pace. In the later stages of Moore's Law, the industry's approach was to turn "doubling every eighteen months" from an observation into a coordination mechanism: the International Technology Roadmap for Semiconductors (ITRS) was jointly developed by the industry starting in the late 1990s, writing down node timelines, equipment specifications, and material requirements years in advance, with all manufacturers and equipment suppliers scheduling production accordingly. This was not a decision by any single company; it was the entire chain fixing the pace, because no one wanted to be faster or slower than others—being faster meant no supporting ecosystem, being slower meant elimination.
The current position of large models is very similar to this. There is no ITRS, but there is a functionally equivalent thing: public benchmarks, open weights, and competitor pricing visible to everyone. How long after a company's model release it must lower prices doesn't need to be written in any document, but everyone knows that number is not controlled by them.
The Isomorphism of the Pricing Power Window and Drug Patents
What is the structural equivalent of this ten-month drop in pricing power? I see it as isomorphic to the expiration of a drug patent, the emergence of generics, and the collapse of prices. Therefore, any large model company, regardless of its name, can be structurally viewed as a pharmaceutical business with a patent period of only ten months. This is an extremely high-risk business—a failure in new drug development might mean there is no next window.
Of course, everyone's goal is ultimately to create that panacea, called AGI. On the path to that universal drug, intermediate products are rapidly iterated and taken offline on a ten-month cycle.
Looking at the industry chain, it's clear that Liang Wenfeng does not have the freedom to extend ten months to twenty months, even if he wanted to. He himself explained this principle: those who want 5% will lose to those willing to take 1%, and those willing to take 1% will lose to those willing to take 0.1%. Everyone is running desperately on a Red Queen Race. So focus is correct; sometimes there isn't even enough oxygen to reach the next station, let alone to stop and admire the flowers and the moon.

Based on these structural judgments, I write down two falsifiable viewpoints.
First, in the large model game, many current players will fade away in disappointment. China may be left with two leading players chasing the United States, while the remaining players will transform and no longer take large models as their main business.
Currently, large model companies essentially subsidize training with inference profits. If the model's "patent" period is too short, the R&D invested in training the model will not yield returns.
Many model manufacturers' financing is far from sufficient for even one failure, and they also lack enough computing resources to simultaneously do both training and inference well. Training and inference compete for the same GPUs, creating an internal zero-sum game between "making drugs" and "making money with existing drugs." Every GPU hour used to serve old models is an hour not used to train new models. This is also the reason I mentioned earlier that training and inference may ultimately be completely separated across different manufacturers.
Back then, when OpenAI retired GPT-4o and cut Sora, the reasons given were to free up computing power and resources. So the reason for retirement was not the disappearance of demand, but the need to reallocate computing resources. Therefore, using the product logic of "I have users, so I'll continue" to view the large model business is wrong: having users means you need to focus even more and not be distracted. OpenAI learned this lesson through bloody experience.
If leading companies dare not waste resources, smaller companies must be even more cautious. Perhaps a few leading companies can do both well for a period of time, having both advanced models and API products. Historically, this is the Intel model: vertical integration, node leadership, betting that integration beats division of labor. Intel's integration was indeed optimal for twenty years—until it could no longer afford process investment with its own product revenue.
Second, Chinese model companies do not need to catch up with the United States to catch up with the United States.
A popular narrative is when Chinese models will catch up with the United States. In fact, from a structural perspective, this is completely unimportant. The pricing power window has only two Nash equilibria: being one generation ahead, or being simultaneous. So Chinese models will not progressively compress this window: ten months, eight months, six months, three months, etc. Instead, as long as a certain compression crosses the current Nash equilibrium, it will be sufficient.
Let's do a thought experiment. Suppose OpenAI now spends a lot of cost to train the next-generation model, and Chinese players make a model that is slightly inferior but ten times cheaper—this is Kimi's direction of large parameters with low activation. Then, OpenAI's pricing power patent period is indirectly compressed to six months, or even three months. In this way, within the patent period, the probability of recovering training and hardware costs is reduced. As the leading player, it needs more time to recover costs. Structurally, the motivation to immediately train the next model automatically weakens to the point where everyone starts together. So once we observe that Chinese models reduce the pricing power window of frontier laboratory models to three months or less, the entire industry will automatically slide to that Nash equilibrium where everyone starts together.
Liang Wenfeng said DeepSeek is one to two years behind, using one-twentieth of the computing power, with the goal of narrowing the gap to three to six months. Note that he didn't say to catch up, and it's not necessary. This is the power of structure. This is also the real role of open weights: not to grab market share, but to shorten the opponent's patent cycle.
Agreeing with Liang Wenfeng, I believe that as long as Chinese laboratories have enough NVIDIA cards, we will see such convergence within one to two years. This has nothing to do with whether Huawei's cards are used to achieve such a structural breakthrough. I would even say that for model manufacturers, first doing their best to achieve such a breakthrough and then switching to Huawei's cards is the simplest path, which can be completed in one year, rather than using Huawei's cards first, causing training difficulties, falling behind by two to three years, and then trying to catch up—at which point catching up becomes even more difficult, and they might be permanently locked in a generation behind.
In the end, the entire large model industry will rapidly evolve under a new structure. This has nothing to do with the AGI narrative; it is determined by the interest structure of the entire industry. Under this structural force, all narratives cannot bypass a structural fact: all large model manufacturers are pharmaceutical factories with a very short patent period.