Reproducing the "DeepSeek Moment"? Wall Street Agrees: Kimi K3 Actually Strengthens Compute Demand
- Core Thesis: The release of Kimi K3 as the world's largest open-source model triggered market panic over weakening compute demand. However, Wall Street investment banks generally believe that K3 will accelerate, rather than reduce, AI infrastructure demand, consistent with the "Jevons paradox."
- Key Elements:
- Kimi K3 features 2.8 trillion parameters, a 1M token context window, and native multimodal capabilities. It has topped authoritative programming benchmarks, with performance rivaling top-tier closed-source models.
- K3's pricing (USD 15 per million output tokens) is lower than Claude Fable 5 and Opus 4.8, but higher than DeepSeek and others, positioning it to offer near-frontier capabilities at a lower cost.
- Investment banks like Nomura and Citigroup point out that more efficient models will stimulate greater application and token processing demand, ultimately driving up consumption of compute, memory, and storage.
- Large-scale deployment of K3 requires supernode clusters, and its inference side has higher demand for HBM and server memory (e.g., DDR5, eSSD) compared to other frontier models.
- OpenRouter data shows that the global share of token usage for Chinese AI models has surged from less than 2% a year ago to over 45%, indicating accelerating ecosystem penetration.
Original author: Long Yue
Original source: Wall Street CN
The market panicked over Kimi K3 as "DeepSeek Moment 2.0," but this time, Wall Street's assessment is entirely different.
Late on July 16th, Moonshot AI launched Kimi K3 in Shanghai. This open-source model with 2.8 trillion parameters scored 57 on the Artificial Analysis Intelligence Index, ranking third to fourth globally, on par with Anthropic's Claude Opus 4.8 and OpenAI's GPT-5.5. More critically, on the Frontend Code Arena leaderboard created by UC Berkeley, K3 topped the charts with a score of 1679, surpassing Claude Fable 5 and GPT-5.6 Sol, becoming the first open-source model to outperform all closed-source foreign models on a prominent coding benchmark.
On July 17th, the US semiconductor sector experienced a notable decline. The market's conditioned reflex is easy to understand – in early 2025, the release of DeepSeek R1 triggered a sharp sell-off in computing stocks. The logic was: Chinese models are getting stronger. Do US AI companies really need to spend so much on computing power? If Chinese models can approach frontier capabilities at lower costs, should the demand for NVIDIA, HBM, servers, and networking equipment be reassessed?
However, according to reports from the ZhuiFeng trading desk, the latest research reports from investment banks including UBS, Nomura, BofA Securities, and Citigroup suggest: Kimi K3 is not the end of computing demand, but an accelerator.
Kimi K3 and DeepSeek R1 represent different types of shocks. R1 showed the market "efficiency"; K3 highlights "scale." With 2.8 trillion parameters, a 1M token context window, always-on reasoning, native multimodality, and a MoE architecture, these features are not an asset-light story. They will collectively increase the pressure on inference, memory, networking, and storage.

How Strong is Kimi K3?
Kimi K3 was released by Moonshot AI on July 16, 2026, with the full model weights scheduled to open on July 27th. It is an open-weight large model with 2.8 trillion parameters, hailed by many institutions as the largest open-weight LLM currently available.
Core configurations include three key points:
First, a 1M token context window. The model can handle longer texts, larger codebases, and more complex enterprise documents and research tasks.
Second, always-on reasoning. It's not just simple Q&A, but designed for long-chain reasoning and Agent tasks.
Third, native vision capabilities. K3 processes not only text but also handles multimodal tasks involving video, images, game development, frontend design, and CAD.
Architecturally, K3 utilizes Kimi Delta Attention, Attention Residuals, and Stable LatentMoE. For the MoE part, it activates 16 experts per token out of 896 total experts. Moonshot AI states that compared to Kimi K2, the overall scaling efficiency has improved approximately 2.5 times.
This explains why K3 is not a "cheaper K2." Nomura's compiled pricing shows K3's input price is $3 per million tokens, cache hit input is $0.30, and output is $15 per million tokens. According to Artificial Analysis metrics, K3's cost per task is approximately $0.94. This price is lower than Claude Fable 5's ~$2.75 and Claude Opus 4.8's ~$1.80, similar to GPT-5.6 Sol's $1.04, but significantly higher than GLM-5.2's $0.32-$0.47, and far above DeepSeek V4 Pro's $0.04.
Thus, K3's positioning is not to be the cheapest, but to offer capabilities approaching frontier models at a lower price point.

Four Investment Banks Reach a Consensus: This is Not Demand Destruction
Addressing market fears of a "DeepSeek Moment" impact, Nomura's Asia Pacific Technology team analyst Duan Bing wrote: "We believe competition and innovation in the global large model market will not cease. As we get closer to Artificial General Intelligence (AGI), the application of generative AI across consumer and enterprise sectors will continue to expand. Frontier AI labs and hyperscale cloud platforms are likely to maintain investment at this stage to sustain their competitive positions – as the scaling laws continue, we interpret this competition as positive for the AI infrastructure value chain."
Citigroup semiconductor analyst Peter Lee, in a July 19th report, titled it directly "Another Jevons Paradox." What does the Jevons Paradox mean? Simply put: Increased efficiency of coal steam engines led to greater coal consumption because more people could afford it and more use cases emerged. The same applies to AI models – when high-quality models become cheaper, developers and enterprises deploy more applications and process more tokens, ultimately leading to higher overall compute consumption.
Peter Lee believes that even if K3 is widely adopted, demand for general memory such as server DDR5 and eSSD will still increase. This is because while K3's inference efficiency is comparable to other frontier models, its KV cache footprint grows with longer contexts, placing greater, not lesser, demands on memory.
BofA Securities semiconductor analyst Vivek Arya was even more direct in his July 17th report. He argued that the response from US frontier AI labs "is not less compute, but more." If Chinese open-source models continue to approach the frontier, OpenAI, Anthropic, and Google must maintain differentiation through larger-scale training, heavier inference, and faster iteration. Arya also highlighted a background detail often overlooked: media reports suggest Google's Gemini 3.5 Pro has been delayed by months, with coding performance not meeting internal targets, and its frontier leadership is "becoming increasingly difficult to defend."
UBS analyst Timo Arcuri's team noted in a July 20th report that parallels between K3 and DeepSeek R1 do exist, but K3 is more about scale – it is the world's largest open-source model with 2.8 trillion parameters and a 1 million token context window. The analysts emphasized that open-source models typically require more memory than closed-source frontier models because of longer context windows, and even with quantization, the absolute KV cache demand continues to grow, making open-source model deployment more dependent on HBM and storage.

Who Truly Benefits in This Competition?
Storage: The most directly benefited sector. UBS estimates show that the cumulative free cash flow (FCF) for the storage and memory segment by 2028 is expected to reach approximately 30% of their market cap, the highest among all sub-sectors – with Micron (MU) alone at 47%. Both Citigroup and Nomura maintain Buy ratings on Samsung Electronics, citing the extremely tight global memory supply situation. Citigroup's Peter Lee pointed out that the inference side memory requirements for Kimi K3 are comparable to other frontier models, and the expansion of KV cache volumes will directly boost demand for server DDR5 and enterprise SSDs (eSSD). He specifically noted that large-scale deployment of Kimi K3 requires "supernode" cluster configurations with over 64 GPUs.

Computing Infrastructure: TSMC and NVIDIA are primary beneficiaries. Whether it's the continued validity of training-side scaling laws or the growth in inference-side token demand, both ultimately point to increased demand for advanced process chips. Nomura reiterates Buy ratings on TSMC, ASE, and MediaTek. NVIDIA has publicly stated that the performance-per-watt for inference on modern MoE models, such as on the GB300 NVL72, is up to 25 times greater than the previous Hopper architecture. Models like K3 are natural beneficiaries of NVIDIA's latest hardware.
Networking: Supernode trend creates structural opportunities. Kimi K3 requires supernode clusters. Domestic Chinese computing power, constrained by export controls on high-end chips, relies more on supernode architectures to compensate for the single-card performance gap. This drives demand for networking layer suppliers like optical modules and optical chips. Nomura favors Zhongji Innolight and Suzhou Xuchuang (likely referring to subsidiaries or partners related to InnoLight).
Cloud Platforms: Benefiting from ecosystem aggregation effects. Cloud platforms hosting multiple frontier open-source models have stronger pricing power and are not dependent on a single closed-source model vendor. Nomura favors Alibaba (BABA) as a core player in China's AI cloud ecosystem, as well as data center operators like GDS Holdings (GDS) and VNET Group (VNET).
How Fast is the Global Penetration of Chinese AI Models?
This might be the most easily underestimated data point in the entire narrative.
According to statistics from the open API gateway OpenRouter, the proportion of global developer traffic using tokens from Chinese AI models was less than 2% a year ago, but has now surpassed 45%. BofA Merrill Lynch's data confirms the acceleration of overall AI penetration: approximately 55% of US enterprises have already subscribed to AI models, platforms, or tools, with Anthropic's enterprise adoption rate at 42% and OpenAI's at 40%. Top-tier AI consumers (the top 1% of enterprise users) are spending $4,833 per employee per month on AI.
The market is differentiating. On one hand, Chinese open-source models represented by DeepSeek and Kimi K3 cover the economy and mid-to-high-end value markets. On the other hand, top-tier US frontier models are focusing on more complex workloads (like scientific computing) to maintain their technological and pricing premiums. Nomura's assessment is that leading large model players on both sides of the Pacific will benefit – provided they can consistently stay on the leading edge of the technology curve.


What K3 Truly Changes is the Pace of Competition, Not a Single Company's Story
Following the K3 release, the market's initial reaction was to compare it with DeepSeek R1. This comparison is useful, but shouldn't stop at "are we going to sell off AI hardware again."
DeepSeek forced the market to reassess training efficiency. K3 shows the market something else: open-source models can also push scale, long context, Agent capabilities, and multimodality into the frontier domain.
This will compel frontier US labs to continue investing and allow Chinese models to expand further within the global developer ecosystem. Top-tier closed-source models retain their technological and pricing premiums, open-source models cover more price points and deployment scenarios, cloud vendors provide model distribution and enterprise implementation, and the hardware chain bears the burden of training and inference.
In the short term, trading may be volatile due to the "DeepSeek memory." In the medium term, as long as token usage continues to grow, and long-context and Agent-based applications continue to proliferate, computing power, HBM, storage, networking, and IDC remain unavoidable cost items.
This is why multiple institutions arrived at similar conclusions following K3: stronger open-source models are not the endpoint of AI infrastructure demand; instead, they might be the entry point for the next wave of demand diffusion.
However, BofA Merrill Lynch also explicitly highlighted a tail risk: "If the pace of efficiency gains outpaces the growth in workloads, we might see some pullback in infrastructure buildout." In other words, if models become increasingly cheaper to run but usage does not expand proportionally, the growth narrative for computing demand becomes less compelling.


