SemiAnalysis: Kimi K3's KDA Mechanism Boosts Attention Efficiency, But Will Require More GPU, HBM, DRAM, and Networking, Not Less
- Core Thesis: Moonshot AI's Kimi K3 model employs a linear attention mechanism, leading the market to worry it might weaken demand for high-end AI hardware like Nvidia's. However, analysis suggests that the sheer scale of K3's over 2.8 trillion parameters and its inference architecture demands may instead reinforce the need for high-performance GPUs, HBM, and high-speed interconnect devices, while long-term, lower costs driven by this efficiency will stimulate the deployment of more applications.
- Key Factors:
- K3's parameter count exceeds 2.8 trillion, requiring over 1.5TB of HBM for model weights. The KV cache also requires substantial offloading to non-HBM storage, meaning HBM demand remains undiminished.
- Efficient inference for K3 requires a large-scale expansion domain architecture comprising at least 64 chips, highly compatible with rack-scale AI systems like Nvidia's GB200/GB300 NVL72.
- K3 employs a WideEP optimization strategy, distributing 896 experts across multiple GPUs. This increases the inter-chip data exchange requirements for high-speed network interconnects, such as copper backplanes.
- Market dissent exists, noting that Huawei's Ascend 950 SuperPod, with its 64-chip configuration, can also meet K3's expansion domain needs, making Nvidia not the sole beneficiary.
- SemiAnalysis cites Jevons Paradox, arguing that linear attention reduces inference costs. This will drive a massive expansion of AI applications, consequently stimulating long-term growth in demand for hardware like GPUs.
- The truly crucial variable is that if leading AI companies (like OpenAI) massively adopt linear attention, the hardware demand structure for long-context inference could undergo significant changes.
Original Author: Li Jia
Original Source: Wall Street CN
The large model Kimi K3 from Dark Side of the Moon (Moonshot AI), which employs a linear attention mechanism, has sparked market concerns that demand for Nvidia, HBM, and networking equipment might be weakened.
However, semiconductor research firm SemiAnalysis has recently offered a starkly different assessment: K3's massive parameter scale and inference architecture requirements will likely not diminish demand for high-end AI hardware; instead, they may further strengthen demand for Nvidia's high-end GPUs, HBM, and high-speed interconnect devices.
SemiAnalysis points out that K3 has over 2.8 trillion parameters, with model weights requiring more than 1.5TB of HBM capacity. Even under relatively limited concurrent user scenarios, the KV cache still necessitates extensive offloading to CPU DDR5 memory and NVMe storage, leaving no significant surplus in HBM space.
More importantly, Dark Side of the Moon previously indicated that efficient inference deployment for K3 requires a large-scale expansion domain architecture consisting of at least 64 chips. This hardware requirement closely aligns with the design direction of rack-scale AI systems like Nvidia's GB200/GB300 NVL72.
SemiAnalysis believes the market's interpretation of linear attention as "reducing GPU demand" is misguided. The real impact may be quite the opposite: More efficient model architectures lower the cost of AI inference, driving broader application adoption, which in turn stimulates long-term demand for GPUs, HBM, DRAM, and network infrastructure.

Demand for Nvidia Chips Remains Robust Despite Linear Attention Iteration
The market's concerns primarily stem from the Kimi Delta Attention (KDA) mechanism adopted by Kimi K3.
Compared to the traditional Transformer attention mechanism, KDA can significantly reduce the data transfer requirements for the KV cache, potentially alleviating network bandwidth pressure by up to approximately 10x. This change reminded some investors of the concerns about AI hardware demand that followed the release of DeepSeek R1, leading to fears that improved model efficiency could reduce reliance on high-performance computing hardware.
However, SemiAnalysis argues that this view overlooks another core requirement for large model inference: the computational and interconnect pressure stemming from model parameter scale.
K3 has over 2.8 trillion parameters, meaning its model weights alone require deployment on large-scale distributed computing systems. Additionally, K3 employs a Wide Expert Parallelism (WideEP) strategy, distributing 896 expert modules across multiple GPUs so that each GPU only handles a subset of expert weights, thereby improving computational utilization.
Nevertheless, WideEP introduces new challenges: Frequent data exchange between experts requires more powerful network interconnect capabilities. SemiAnalysis notes that the copper backplane interconnect architecture used by GB200/GB300 NVL72 offers 18 times more intra-rack bandwidth than traditional DGX B200 systems, making it highly suitable for such large-scale expert parallel inference tasks.
In other words, the KV cache communication savings from KDA may be partially offset by the weight exchange demands of WideEP. The overall pressure on AI infrastructure does not decrease significantly.


The 64-Chip Expansion Domain Does Not Benefit Nvidia Exclusively
However, dissenting voices exist in the market. An informed source, GDP (@bookwormengr), points out that the "64-chip expansion domain" mentioned by Dark Side of the Moon does not necessarily imply the Nvidia NVL72 solution. Huawei's Ascend 950 SuperPod also features a 64-chip configuration and possesses Unified Bus (UB) memory expansion capabilities similar to NVLink.
From an architectural perspective, the Ascend 950 SuperPod can scale across 16 racks to support up to 1024 NPUs, making it equally competitive in meeting large-scale model inference demands. Therefore, while the increased hardware demand from K3 does not mean Nvidia will be the sole beneficiary, the trend towards stronger demand for high-end AI interconnect systems remains clear.

Jevons' Paradox: AI Efficiency Gains May Drive Hardware Demand Growth
SemiAnalysis further invokes Jevons' Paradox to explain the trends in AI infrastructure.
This theory posits that when a technology improves resource utilization efficiency and lowers unit costs, demand often does not decrease; instead, it may grow due to expanded application scope. Applied to the AI domain, the reduction in inference costs facilitated by linear attention mechanisms could encourage more enterprises to deploy AI applications, further expanding the scale of global AI inference. This ultimately drives demand growth for GPUs, HBM, DRAM, and high-speed networking equipment.
However, GDP remains cautious about this viewpoint. He acknowledges the long-term logic of Jevons' Paradox but notes that KDA has practical significance for optimizing state storage in long-context tasks. Even with massive model weight sizes, the actual memory pressure might be lower than market intuition suggests, due to the use of 4-bit quantization and highly sparse design.
He believes that the truly important variable to watch is whether the AI companies with the highest global inference demand—OpenAI, Anthropic, and Google DeepMind—have already adopted, or will adopt in the future, linear attention schemes similar to KDA or DeepSeek CSA/HCA. If leading AI companies widely adopt such architectures, the demand for memory and interconnect resources in long-context inference could see a significant decline. This would be a crucial factor shaping the future demand structure for AI hardware.
Market Focus Shifts: Can AI Demand Growth Offset Architectural Efficiency Gains?
Overall, the emergence of Kimi K3 does not simply point towards a decline in AI hardware demand. For Nvidia, the real competitive focus may shift from "single-card performance" to "system-level capabilities"—including high-speed interconnects, rack-scale expansion, and large-scale inference optimization.
As model scales continue to grow, even with ongoing optimization of attention mechanisms, AI infrastructure will still face pressure from increasing parameter counts, expert parallelism, data exchange, and inference throughput demands.
The core question the market needs to focus on in the future is not whether linear attention reduces the consumption of individual resources, but whether the new demand generated by the expanding scale of AI applications can consistently outpace the resource savings brought about by architectural efficiency improvements.


