Reproducing the "DeepSeek moment"? Wall Street Says: Kimi K3 Actually Strengthens Computing Power Demand
- Core Thesis: The release of Kimi K3 as the world's largest open-source model triggered market panic over weakening computing power demand. However, Wall Street investment banks generally believe that K3 will accelerate, rather than reduce, AI infrastructure demand, aligning with the "Jevons paradox."
- Key Elements:
- Kimi K3 features 2.8 trillion parameters, a 1M token context window, and native multimodal capabilities. It has topped authoritative programming benchmarks, with performance rivaling top-tier closed-source models.
- K3's pricing ($15 per million output tokens) is lower than Claude Fable 5 and Opus 4.8 but higher than DeepSeek, positioning itself as providing near-frontier capabilities at a lower cost.
- Investment banks like Nomura and Citigroup point out that more efficient models will stimulate greater application and token processing demand, ultimately driving up consumption of computing power, memory, and storage.
- Large-scale deployment of K3 requires supernode clusters, and its inference side demands higher consumption of HBM and server memory (e.g., DDR5, eSSD) compared to other frontier models.
- OpenRouter data shows that the global share of token usage by Chinese AI models has surged from less than 2% a year ago to over 45%, indicating an acceleration in ecosystem penetration.
Original Author: Long Yue
Original Source: Wall Street News
The market panicked over Kimi K3 as if it were "DeepSeek Moment 2.0," but this time, Wall Street's assessment is starkly different.
Late on the night of July 16, Moonshot AI launched Kimi K3 in Shanghai. This open-source model with 2.8 trillion parameters scored 57 on the Artificial Analysis intelligence index, ranking third to fourth globally, on par with Anthropic's Claude Opus 4.8 and OpenAI's GPT-5.5. More critically, on the Frontend Code Arena leaderboard created by UC Berkeley, K3 topped the charts with a score of 1679, surpassing Claude Fable 5 and GPT-5.6 Sol, becoming the first open-source model to outperform all closed-source foreign models on an authoritative coding benchmark.
On July 17, the U.S. semiconductor sector experienced a notable decline. The market's conditioned reflex is understandable—at the beginning of 2025, the release of DeepSeek R1 triggered a sharp sell-off in compute stocks. The logic was: Chinese models are getting stronger, so do U.S. AI companies really need to spend so much on compute? If Chinese models can approach frontier capabilities at lower costs, should the demand for NVIDIA, HBM, servers, and network equipment be reassessed?
However, according to Zhuifeng Trading Desk sources, the latest research reports from investment banks including UBS, Nomura, BofA Merrill Lynch, and Citigroup argue that: Kimi K3 is not the end of compute demand, but an accelerator.
Kimi K3 and DeepSeek R1 are not the same kind of shock. R1 showed the market more about "efficiency"; K3 highlights "scale." With 2.8 trillion parameters, a 1M token context window, always-on reasoning, native multimodality, and an MoE architecture, these features are not a story of asset-light models. They will collectively raise the pressure on reasoning, memory, networking, and storage.

How Powerful is Kimi K3?
Kimi K3 was released by Moonshot AI on July 16, 2026, with the full model weights scheduled to open on July 27. It is an open-source weight model with 2.8 trillion parameters, hailed by multiple institutions as the largest open-source weight LLM to date.
Core configurations include three key points:
First, a 1M token context window. The model can process longer texts, larger codebases, and more complex corporate documents and research tasks.
Second, always-on reasoning. It is not just simple Q&A, but targets long-chain reasoning and Agent tasks.
Third, native vision capabilities. K3 handles not only text but also multimodal tasks such as video, images, game development, frontend design, and CAD.
Architecturally, K3 utilizes Kimi Delta Attention, Attention Residuals, and Stable LatentMoE. The MoE portion activates 16 experts per token out of 896 experts. Moonshot AI states that overall scaling efficiency has improved approximately 2.5 times compared to Kimi K2.
This explains why K3 is not a "cheaper K2." Pricing compiled by Nomura shows that K3's input price is $3 per million tokens, cached-hit input is $0.30, and output is $15 per million tokens; according to Artificial Analysis metrics, K3's cost per task is approximately $0.94. This price is lower than Claude Fable 5's ~$2.75 and Claude Opus 4.8's ~$1.80, close to GPT-5.6 Sol's $1.04, but significantly higher than GLM-5.2's $0.32–$0.47, and far above DeepSeek V4 Pro's $0.04.
So, K3's positioning is not about being the cheapest, but about approaching frontier model capabilities at a lower price.

Four Investment Banks Consensus: This is Not Demand Weakening
In response to market fears of a "DeepSeek moment" shock, Duan Bing, an analyst in Nomura's Asia-Pacific Technology team, wrote: "We believe that competition and innovation in the global large language model market will not stop. As we get closer to Artificial General Intelligence (AGI), the application of generative AI on both consumer and enterprise sides will continue to expand. Frontier AI labs and hyperscale cloud platforms are likely to continue investing at this stage to maintain competitive positions—as scaling laws continue, we interpret this competition as a positive signal for the AI infrastructure value chain."
In a report on July 19, Citigroup semiconductor analyst Peter Lee directly titled it "Another Jevons Paradox." What does the Jevons Paradox mean? Simply put: increased efficiency of coal steam engines led to greater coal consumption because more people could afford it and it was usable in more scenarios. The same applies to AI models—when high-quality models become cheaper, developers and enterprises deploy more applications and process more tokens, ultimately increasing compute consumption.
Peter Lee believes that even if K3 is widely adopted, demand for general-purpose memory like server DDR5 and eSSD will still increase. The reason is that while K3's inference efficiency is comparable to other frontier models, its KV cache footprint will expand as context grows, placing greater, not lesser, pressure on memory.
BofA Securities semiconductor analyst Vivek Arya was even more direct in his July 17 report. He argues that the response of U.S. frontier AI labs will be "not less compute, but more." If Chinese open-source models continue to close the gap, OpenAI, Anthropic, and Google must maintain differentiation through larger-scale training, heavier inference, and faster iteration. Arya also highlighted an easily overlooked context: media reports indicate that Google's Gemini 3.5 Pro has been delayed by several months and its coding performance has missed internal targets, making its frontier leadership position "increasingly difficult to defend."
In a July 20 report, UBS analyst Timo Arcuri's team noted that parallels between K3 and DeepSeek R1 do exist, but K3 is more about scale—it is the world's largest open-source model with 2.8 trillion parameters and a 1 million token context window. The analyst emphasized that open-source models typically consume more memory than closed-source frontier models because their context windows are longer, and KV cache demand, even after quantization, continues to grow in absolute terms, making open-source model deployment more reliant on HBM and storage.

Who Truly Benefits in This Competition?
Storage: The most direct beneficiary sector. UBS estimates that the cumulative free cash flow (FCF) of the storage and memory sector through 2028 is expected to reach approximately 30% of its current market capitalization, the highest among all sub-sectors—with Micron (MU) alone accounting for 47%. Both Citigroup and Nomura maintain buy ratings on Samsung Electronics, citing an extremely tight supply situation in the global memory market. Citigroup analyst Peter Lee pointed out that the inference-side memory demand of Kimi K3 is no less than that of other frontier models, and the expansion of KV cache volume will directly drive up demand for server DDR5 and enterprise solid-state drives (eSSD). He specifically noted: large-scale deployment of Kimi K3 requires "supernode" cluster configurations with over 64 GPUs.

Compute Infrastructure: TSMC and NVIDIA are the primary beneficiaries. Whether it's the continued validity of scaling laws on the training side or the growth of token demand on the inference side, both ultimately point towards increased demand for advanced process chips. Nomura reiterates buy ratings on TSMC, ASE, MediaTek, and others. NVIDIA has publicly stated that the performance-per-watt of modern MoE (Mixture of Experts) model inference on the GB300 NVL72 can be up to 25 times better than the previous generation Hopper architecture. Models like K3 are naturally beneficiaries of NVIDIA's latest hardware.
Networking: The supernode trend creates structural opportunities. Kimi K3 requires supernode clusters. Domestic Chinese compute power, constrained by export controls on high-end chips, relies more on supernode architectures to compensate for single-card performance gaps, driving demand for network layer suppliers like optical modules and optical chips. Nomura is bullish on Zhongji Innolight and Suzhou Innolight.
Cloud Platforms: Benefit from ecosystem agglomeration effects. Cloud platforms hosting multiple frontier open-source models have stronger bargaining power and are not dependent on a single closed-source model supplier. Nomura favors Alibaba (BABA) as the core of China's AI cloud ecosystem, along with data center operators like GDS Holdings and VNET.
How Fast is the Global Penetration of Chinese AI Models?
This might be the most easily underestimated data point in the entire narrative.
According to statistics from the open API gateway OpenRouter, the proportion of token usage from Chinese AI models as a share of global developer traffic has surged from less than 2% a year ago to over 45% today. Data from BofA Merrill Lynch confirms the acceleration of overall AI penetration: approximately 55% of U.S. enterprises have now subscribed to AI models, platforms, or tools, with Anthropic's enterprise adoption rate at 42% and OpenAI's at 40%. Top-tier AI consumers (the top 1% of corporate users) spend an average of $4,833 per employee per month on AI.
The market is bifurcating. On one side are Chinese open-source models represented by DeepSeek and Kimi K3, covering the low-cost and mid-to-high-end value-for-money markets. On the other side, top-tier U.S. frontier models are focusing on more complex workloads like scientific computing, maintaining technological and pricing premiums. Nomura's assessment is that leading LLM players on both sides of the China-U.S. divide will benefit—provided they can continuously stay at the forefront of the technology curve.


K3 Truly Changes the Pace of Competition, Not Just One Company's Story
After the K3 launch, the market's first reaction was still to compare it with DeepSeek R1. This comparison is useful, but it shouldn't stop at the "Is it going to crush AI hardware again?" layer.
DeepSeek made the market reassess training efficiency. K3 showed the market another thing: open-source models can also push scale, long context, Agents, and multimodality to the frontier.
This will force U.S. frontier labs to continue investing, and it will also allow Chinese models to continue expanding within the global developer ecosystem. Top-tier closed-source models retain technological and pricing premiums, open-source models cover more price points and deployment scenarios, cloud providers facilitate model distribution and enterprise implementation, and the hardware chain bears the pressure of training and inference.
In the short term, trading may fluctuate due to the "DeepSeek memory." In the medium term, as long as token usage continues to grow and long-context and Agents continue to proliferate, compute, HBM, storage, networking, and IDC will remain unavoidable cost items.
This is also why multiple institutions reached similar conclusions after K3: stronger open-source models are not the end of AI infrastructure demand; instead, they may be the entry point for the next wave of demand expansion.
However, BofA Merrill Lynch also explicitly left a tail risk: "If the pace of efficiency gains exceeds the growth of workloads, we might see some pullback in infrastructure buildout." In other words, if models become cheaper and cheaper, but usage does not expand proportionally, the growth narrative for compute demand could be compromised.


