
Review of Kimi K3: Small-Town Youth Strikes Capitalists

The Three Carriages of Chinese AI
For a long time, Chinese AI models could be roughly divided into three routes.
The first route is actively making models smaller. For example, minimax-m3, kimi k2.5, hy3, and other models that actively limit activated parameters to around 20B. Whether 20B activated parameters count as small is itself a subtle question, but in engineering terms, we usually consider it small.
Smaller activated parameters are often a hard economic constraint. Maximizing performance under economic constraints is a typical "small model" approach.
The second route is DeepSeek, and after DeepSeek's emergence, other vendors with similar technology stacks and model sizes. The characteristic of the DeepSeek route is not only that the technology is fully open source, but more importantly, DeepSeek-V3/R1 was a relatively large model from its inception, benchmarking against OpenAI's top-tier models.
With the explosion of R1, DeepSeek supplemented a large amount of economically optimized technology stacks. Therefore, the most typical feature of the DeepSeek route is "high performance while ensuring inference economy." GLM-5 is a systematic inheritor and practitioner of this route, and GLM-5.2 is one of the better-known achievements along this route.
The third is kimi. Kimi's technology stack has always been very independent, with a distant relationship to DeepSeek. The K2.5~K2.7 series absorbed a significant portion of DeepSeek's technology stack, while the new K3 series reverted to using its own proprietary technology.
K3 pushes total parameters to 2.8T, with estimated activated parameters of 50B, using entirely its own technology stack. The scaling law of size has taken effect, but at the same time, it has caused a significant decline in the economy of kimi k3, pushing it into the price range of Sonnet.
Although still very cheap compared to Claude, given that kimi's inference cluster is located in China and considering the limitations on the number of cards, kimi is likely not prioritizing economy as the primary requirement when making K3.
In other words, K3 is a product of an independent technology stack, extremely large size, and a willingness to pursue high model performance at all costs.
Compared to models on the DeepSeek route that balance performance and economy, K3's goal from the start was not to create a better R1, but to benchmark against the most cutting-edge ultra-large models from OpenAI and Anthropic.
Why is K3 So Large?
Describing Kimi as an independent technology route that has never had any relation to DeepSeek is not accurate.
The Kimi K2 paper makes it very clear: K2 uses a super-sparse MoE and MLA architecture similar to DeepSeek-V3. Both have 61 layers, both have shared experts, and each token activates 8 routing experts; on this basis, K2 increases the number of experts to 384, halves the attention heads, and uses its own MuonClip, data recipes, agentic post-training, and independent infrastructure to create its own fork.
Therefore, K2 is more like a Kimi branch on the DeepSeek-V3 chassis, rather than a different species built from scratch.
What is truly noteworthy is that the base size of K2.5 to K2.7 has basically not expanded further: total parameters remain around 1T, activated parameters about 32B, still 384 experts selecting 8.
In other words, Kimi has proven that it can squeeze the same chassis for many generations through post-training and agent systems without expanding the base.
But K3 is different. K3 is therefore not an ordinary version upgrade. It represents Kimi actively ending the phase of "continuing post-training on an existing base" and returning to scaling pre-training.

K3 heavily uses its own proprietary technologies that were previously scattered, such as its own training base, its own attention mechanism, and its own expert routing structure. At this point, the kinship between K3 and DeepSeek has become relatively distant.
K3 simultaneously aims to prove two things: one, that kimi can define its own technology paradigm; two, that kimi is qualified to use new technology to train a completely new ultra-large model, benchmarking against OpenAI and Anthropic.
As it turns out, K3 succeeded.
From Public Utility to Privileged Monopoly
Before DeepSeek, China did have usable models, but they were barely passable.
The emergence of DeepSeek for the first time brought a high-performance model available in China, becoming a public utility for a generation of Chinese AI. The core proposition of public utility is Adoption: how to broadly promote "useful" capabilities—on an already strained inference base?
Therefore, DeepSeek not only undertook the function of performance breakthroughs but also bore a significant responsibility for economic optimization.
From V3 to V4, DeepSeek has always maintained API prices close to charity, backed by a mature technology stack optimized for inference.
At this stage, the problem facing Chinese AI was: first, meet the performance needs of public utility; second, reduce the computational cost of public utility.
For Chinese vendors other than DeepSeek, the tasks suddenly became two: first, performance must catch up to R1, otherwise the model has no competitive qualification; second, price, throughput, and deployment efficiency must be as close to DeepSeek as possible, otherwise adoption simply won't take off.
GLM's phased stack switching is the clearest microcosm of this industrial pressure.
In this sense, DeepSeek brought Chinese models into the Adoption phase. The core issue of this phase is: how to let more people use sufficiently powerful intelligence, and how to handle the ensuing surge in call volume.
But a frontier lab will sooner or later have to answer another question: if we only ever pursue making things cheaper to catch up with others, where will the higher-performance models come from?
All major internet companies made a choice: find ways to bypass the ban and import from the US. For example, the famous Alibaba's Singapore entity, allowing users to remotely log into Singapore machines to run Claude Code.
And then it was blocked. If Uncle Sam hadn't quickly followed up with the 5.6-Sol cut, Dario could probably have posted a few more articles stepping on the Chinese.
Through extensive engineering optimization, China has gradually achieved the foundation of AI public utility. But without larger, higher-performance frontier models, it will always be one beat behind, falling into a long-term disadvantageous "privilege competition."
The Sword Dance Aimed at Pei Gong
K3's 2.8T is Moonshot's answer to this question.
K2.5 to K2.7 can still be understood as improving cost-effectiveness on a 1T-A32B base through continued pre-training, RL, tool environments, and agent systems.
But K3 directly pushes total capacity to 2.8T, accepts deployment requirements of 64 or more accelerator cards, and raises API prices from K2.7 Code's $0.95/$4 to $3/$15. Although this price is only comparable to Sonnet, it is still very expensive.
The most critical part is the 64 accelerator cards. Where can you find so many 64-card nodes in China's inference clusters?
K3 may still want to maintain economy, but it's almost impossible for K3 to be economical. A 2.8T giant model is a huge challenge even for local inference companies in the US. Therefore, K3's intention is clearly not for broader adoption, but to set the standard for open-source models that can be widely adopted on the more cutting-edge battlefield.
Compared to most PR articles that claim to "crush US models," Kimi's official wording is more moderate: K3 as a whole still lags behind Fable 5 and GPT-5.6 Sol.
Of course, viewed from another angle, it's also radical: it's only slightly behind the top-tier models of OpenAI and Anthropic.
Once it enters this tier, the first thing K3 shakes is Anthropic's most important commercial narrative over the past year.
Lao A's Eternal Movement Story
Anthropic's methodology has always been very clear: we create an irreplaceable model. Once users discover it can do work that other models cannot, they are willing to pay a premium for higher performance.
Fable 5 is the most extreme product of this methodology. It targets the most difficult knowledge work, coding, and complex asynchronous tasks that can last for days, with API prices reaching $10 per million tokens for input and $50 for output, twice the regular price of Opus 4.8.
Anthropic is not selling more expensive tokens, but a new task domain: projects that previously could not be reliably completed can now be handed over to Fable. In its scenarios, Fable has irreplaceable value.
As long as this irreplaceability exists, the price premium is entirely reasonable. After all, high-performance models replace high-knowledge human labor, which is much more expensive than the model.
After Fable's release, the market briefly formed an illusion that OpenAI could no longer catch up to Anthropic. SemiAnalysis's super bullish projection even predicted that Anthropic's ARR could reach $300 billion by the end of 2027, corresponding to a $6 trillion enterprise value at 20 times ARR.
This prediction also relied on a whole set of assumptions: API accounting for 75%-85% of revenue, high API gross margins, extremely high net revenue retention, and continued penetration of Claude Code.
Standing at the time point of July 18, it seemed like SemiAnalysis was hallucinating. But back on July 8, quite a few people thought it made a lot of sense.
Reality arranged a highly dramatic timeline for this narrative. The projection of a $6 trillion valuation was widely circulated on July 8, and GPT-5.6 Sol was officially released on July 9.
Lao A's Eternal Movement Story lasted less than a month before being overturned. In between, there were also a few articles criticizing Chinese model distillation.
Who Killed the Robin?
After GLM-5.2 was released, the author wrote an article systematically analyzing the significance of GLM-5.2, which was widely hailed by the masses as a "masterpiece."
Considering that Zhipu has never paid the author for promotion, the author will still maintain the status of an independent researcher (laughs).
However, GLM-5.2 has indeed shaken Anthropic's foundation, but at the time people didn't realize the seriousness of the matter.
GLM-5.2 performs better than Opus 4.8 on most tasks, second only to Opus 4.7. You read that right: it's better than 4.8 while being second only to 4.7, which makes one suspect how much water Anthropic diluted in 4.8.
At the same time, GLM-5.2's price is only one-fifth of Opus, and it is not subject to the rate limits of Anthropic's subscription plans. The emergence of GLM-5.2 has already significantly threatened the pricing anchor of Anthropic's main model segment.
Thus, the market wanted to price Anthropic's "more cutting-edge models," leading to SemiAnalysis's 6 trillion valuation god-like prediction.
Then, Sol directly collided with Fable on the highest capability tier. Anthropic subsequently extended Fable's subscription quota multiple times and eventually announced it would become a permanent benefit of Max and Team Premium.

And then K3 came out. Its overall capability lags behind Fable and Sol, and only behind Fable and Sol. At the same time, K3 is open source, deployable on any inference cluster.
Kids, advanced AI models really do grow in the fields.
Asymmetric Competition
In the past, the logic of US export controls on advanced computing chips to China was very straightforward: by restricting upstream computing power, slow down China's training of frontier models and building of its own semiconductor ecosystem. National security, military, and surveillance uses were the public legal reasons, while maintaining US technological leadership was the concurrent industrial goal.
Of course, the bigger problem is actually that the US itself doesn't have enough advanced chips. The US is not only worried about China obtaining computing power, but also does not want Chinese companies to participate in bidding, squeezing the ability of US labs and data centers to buy cards.
This set of controls certainly increased the difficulty for China to train frontier models, but it also created a subtle backlash: it deprived Chinese AI companies of the ability to earn excess profits, while turning this disadvantage into a weaponized tool in the game.
Since you won't let me make excess profits, and I don't want to be locked out of technology, why not just give up profits after making the model?
Moreover, open-source models are not picky about graphics cards. Chinese GPUs can deploy them, and US computing clusters can too. Fireworks, Together AI, and a bunch of NeoClouds all have plenty of GPUs.
Thus, the China-US AI competition has produced an extremely abstract combination:
Chinese models + US and global inference clusters VS US closed-source AI labs
Chinese labs bear the expensive cost of model development, exchanging open weights for adoption and frontier status; US cloud vendors and inference platforms use local GPUs to handle calls, pricing at cost, and making good money.
The ones that truly lose profits are the US frontier labs that rely on closed weights and capability scarcity to charge exorbitant prices.
For Chinese labs, open weights sacrifice the potential API rental income that is limited by their own inference capacity.
For US closed-source labs, what is sacrificed is the real excess profits built on abundant local computing power and global capability scarcity.
Therefore, the US faces not a simple choice of "continue sanctions or lift sales," but an impossible trinity:
This is a typical asymmetric competition. Washington loves to talk about asymmetric competition, and now the term has come back to Washington.
Pledge Allegiance to Micron!
Pushing this chain one step further upstream touches the entire AI CapEx narrative.
In the past, the high inference gross margins of closed-source frontier models were the most important buffer in the AI CapEx chain. As long as model capability was sufficiently scarce, APIs and subscriptions could maintain high prices; labs and cloud vendors with high gross margins could tolerate more expensive GPUs, faster depreciation, and more aggressive data center investments; upstream in turn continued to expand production and maintain high pricing.
This chain can be simply written as:
Model capability scarcity → High API prices → High downstream gross margins → Ability to bear GPU premiums → CapEx expansion → Upstream sustained prosperity.
When multiple platforms can deploy the same weights, inference prices will gradually shift from "what the frontier lab is willing to charge" to "how much it costs for an efficient cluster to run a successful task."
Open weights do not change the fact that computing power costs money; they cancel the "model tax" attached to computing power consumption. Fireworks does not have to bear the pre-training cost of K3 for Moonshot, nor does it have to pay the full model rent per token upstream to a closed-source model.
It still has to bear GPU, inference engine, quantization, caching, networking, financing, SLA, and enterprise sales costs, but it only needs to quote based on actual computing costs, utilization, service costs, and competitive profit. This is typical cost-based pricing logic, and for industries that rely on prosperity, cost-based pricing is disastrous for the upstream.
This will reversely change the capital return calculation for every GPU.
In the past, service providers could put expensive hardware costs into high-priced model APIs and then pass them on to end customers.
When inference platforms compete around the same open weights, it becomes harder for any single entity to continue passing on this premium. Thus, GPU procurement requires higher utilization, shorter payback periods, lower financing costs, and more certain customer contracts.
In the short term, downstream buyers are more fragmented, which does not weaken the bargaining power of computing power producers (especially NVDA). But in the long run, the systematic decline in downstream gross margins will inevitably lead to poorer pricing flexibility upstream, posing an uncontrollable and huge risk to the entire AI CapEx narrative.
Too Big to Fail
At this point, K3's 2.8T has taken on two seemingly opposite meanings.
On the one hand, K3 is not a very economical model, pushing the model into the Sonnet price range and the deployment range of 64+ cards.
On the other hand, it is precisely this "not prioritizing economy first" that gives K3 the qualification to enter the highest capability group, shake Anthropic's model rent, restructure the profits of China-US inference, and influence the AI CapEx narrative.
In the past, many companies liked to talk about edge computing, small sizes, and "good enough." These routes are certainly important; they determine how many devices, users, and industries the existing intelligence can spread to.
But the word "good enough" is very boring; it only applies to markets where the capability boundary is already fixed.
Frontier AI faces a constantly evolving set of tasks: just yesterday the model learned single-file programming, today it needs to handle an entire codebase; just yesterday it learned tool calling, today the next generation needs to work continuously for hours or even days.
Moderately sized models answer "how to be used by more people," while larger models compete for "who has the right to issue oracles."
A large model that is already strong but expensive can later be made faster and cheaper through sparsification, quantization, caching, distillation, and inference engineering, and can also serve as a teacher model for an entire product line of small models.
A small model that never learned a certain capability from the start will find it difficult to compensate by extending thinking or deployment optimization out of thin air.
This asymmetry makes "going big won't lose" a rough but often effective frontier methodology.
In 2024, OpenAI released o1, shifting the focus of scaling from larger pre-training bases to training-time reinforcement learning and test-time thinking computation.
o1 itself is a smaller-sized reasoning model. It proved that "smaller base + more thinking" can achieve remarkable progress in math, code, and formal reasoning, thereby starting a cycle of making models smaller.
The smaller the base, the more it relies on test-time compute to compensate; the more successful the reasoning techniques, the less incentive the organization has to expand the base. Thus, OpenAI's models began to rapidly deteriorate.
In 2025, Anthropic went in the opposite direction. Anthropic insisted on making larger and more expensive models, then turning capabilities into products through Claude Code.
User loyalty to Claude is not because it's cheap, but because on the most difficult real software engineering tasks, they trust Opus more. By the end of 2025, Anthropic claimed that Claude Code had captured more than half of the AI coding market, with enterprise market share rising from 24% to 40%. Revenue also surpassed OpenAI.
OpenAI then realized again that inference efficiency and routing cannot replace the highest capability base. GPT-5 re-established a complete ladder of main, thinking, mini, and Pro; by GPT-5.6, it explicitly set three model sizes—Sol, Terra, Luna—and placed the large Sol back in the flagship position for the most difficult work.
Thus, Codex users quickly surpassed 10 million, forcing Anthropic to add Fable back to the subscription plan.
Now, K3 is released, second only to Sol and Fable, even sparking discussions about whether US frontier AI is still doing well.
When you are truly willing to compete for first place, the result is likely not too bad. This is Too Big to Fail.
Conclusion: Small-Town Youth Strikes Capitalists
China's AI model companies are often described as small-town youth: limited resources, unfavorable location, even worse international reputation, barely enough computing power to support inference, just hoping not to fall behind in training.
At the same time, advanced AI labs in the US exhibit a capitalist demeanor: discussing anthropology and safety, machines of loving grace, whether the future world belongs to freedom or unfreedom, whether AI belongs to the earth or the sky.
With such a dramatic gap, it's surprising that a scene of small-town youth striking capitalists can unfold. No matter how you explain it, it's enough to make you laugh out loud.
——Actually, no need to rush. The good show is still to come.