BTC
ETH
HTX
SOL
BNB
ดูตลาด
简中
繁中
English
日本語
한국어
ภาษาไทย
Tiếng Việt

SemiAnalysis: Kimi K3's KDA Mechanism Improves Attention Efficiency, But Will Require More GPUs, HBM, DRAM, and Networking, Not Less

星球君的朋友们
Odaily资深作者
2026-07-20 02:31
บทความนี้มีประมาณ 2204 คำ การอ่านทั้งหมดใช้เวลาประมาณ 4 นาที
While Kimi K3's linear attention has sparked short-term concerns about hardware demand, SemiAnalysis points out that its massive scale of over 2.8 trillion parameters and inference architecture will instead intensify the demand for high-end GPUs, HBM, and high-speed interconnects.
สรุปโดย AI
ขยาย
  • Core Thesis: Moonshot AI's Kimi K3 model employs a linear attention mechanism, raising market concerns that it could weaken demand for high-end AI hardware like Nvidia's. However, analysis suggests that the sheer scale of K3's over 2.8 trillion parameters and its inference architecture requirements may instead strengthen demand for high-performance GPUs, HBM, and high-speed interconnect devices, while long-term cost reductions could spur broader application adoption.
  • Key Factors:
    1. K3's parameter count exceeds 2.8 trillion, with model weights requiring over 1.5TB of HBM. Additionally, the KV cache still requires significant offloading to non-HBM storage, meaning HBM demand remains undiminished.
    2. Efficient inference for K3 necessitates a large-scale domain architecture composed of at least 64 chips, highly compatible with rack-scale AI systems like Nvidia's GB200/GB300 NVL72.
    3. K3 employs a WideEP optimization strategy, distributing 896 experts across multiple GPUs, thereby increasing the demand for high-speed network interconnects (e.g., copper backplanes) for inter-expert data transfer.
    4. Diverse market opinions exist, noting that Huawei's Ascend 950 SuperPod with 64 chips can also meet K3's extended domain requirements, meaning Nvidia is not the sole beneficiary.
    5. Referencing Jevons Paradox, SemiAnalysis argues that lower inference costs from linear attention will drive massive expansion in AI applications, thereby stimulating long-term growth in demand for hardware like GPUs.
    6. The truly critical variable is that if leading AI companies (e.g., OpenAI) adopt linear attention at scale, the hardware demand structure for long-context inference could undergo significant changes.

Original Author: Li Jia

Original Source: Wall Street News

The large language model Kimi K3, developed by Moonshot AI, adopts a linear attention mechanism, sparking market concerns that demand for NVIDIA, HBM, and networking equipment could be weakened.

However, semiconductor research firm SemiAnalysis recently offered a starkly different assessment: The massive parameter size and inference architecture requirements of K3 will not weaken demand for high-end AI hardware; instead, they may further strengthen the need for NVIDIA's high-end GPUs, HBM, and high-speed interconnect equipment.

SemiAnalysis points out that K3 has over 2.8 trillion parameters, with model weight capacity exceeding 1.5TB of HBM. Even in scenarios with relatively limited concurrent users, the KV cache still requires significant offloading to CPU DDR5 memory and NVMe storage, leaving no obvious surplus in HBM capacity.

More importantly, Moonshot AI has previously indicated that efficient inference deployment of K3 requires a large-scale scaling domain architecture consisting of at least 64 chips. This hardware requirement strongly aligns with the design direction of rack-scale AI systems like NVIDIA's GB200/GB300 NVL72.

SemiAnalysis believes the market misinterpretation, which viewed linear attention as "reducing GPU demand," is flawed. The real impact might be quite the opposite: More efficient model architectures lower the cost of AI inference, driving more applications to deployment, and thereby stimulating long-term demand for GPUs, HBM, DRAM, and network infrastructure.

Demand for NVIDIA Chips Remains Robust Despite Linear Attention Iterations

Market concerns primarily stem from the Kimi Delta Attention (KDA) mechanism employed by Kimi K3. Compared to the traditional Transformer attention mechanism, KDA can significantly reduce the data transfer requirements of the KV cache, potentially alleviating network bandwidth pressure by up to approximately 10 times. This change reminded some investors of the market's concerns over AI hardware demand following the release of DeepSeek R1, leading to the belief that improved model efficiency could reduce dependence on high-end computing hardware.

However, SemiAnalysis argues that this assessment overlooks another core requirement of large model inference: the computational and interconnection pressure driven by parameter scale.

K3 possesses over 2.8 trillion parameters, meaning its model weights inherently rely on large-scale distributed computing systems for deployment. Furthermore, K3 adopts Wide Expert Parallelism (WideEP) optimization strategy, distributing its 896 expert modules across multiple GPUs so that each GPU only handles a portion of the expert weights, thereby improving compute utilization.

Nevertheless, WideEP also introduces new challenges: Frequent data exchange between experts requires more powerful network interconnection capabilities. SemiAnalysis notes that the copper backplane interconnect architecture of GB200/GB300 NVL72 offers 18 times the intra-rack bandwidth compared to traditional DGX B200 systems, making it highly suitable for such large-scale expert parallel inference tasks.

In other words, the reduction in KV cache communication demand achieved by KDA may be partially offset by the weight exchange requirements inherent in WideEP, resulting in no significant overall decrease in AI infrastructure pressure.

NVIDIA Not the Sole Beneficiary of the 64-Chip Scaling Domain

However, there are dissenting voices in the market. An informed source, GDP (@bookwormengr), points out that the "64-chip scaling domain" mentioned by Moonshot AI does not necessarily refer to NVIDIA's NVL72 solution. Huawei's Ascend 950 SuperPod also features a 64-chip configuration and possesses unified bus (UB) memory expansion capabilities similar to NVLink.

From an architectural perspective, the Ascend 950 SuperPod can scale to 1024 NPUs across 16 racks, making it competitive for meeting the inference needs of large-scale models. Therefore, while the increased hardware demand driven by K3 does not mean NVIDIA will be the sole beneficiary, the trend towards strengthened demand for high-end AI interconnect systems remains clear.

Jevons' Paradox: AI Efficiency Gains Could Fuel Hardware Demand Growth

SemiAnalysis further invokes Jevons' Paradox to interpret the AI infrastructure trend. This theory posits that when a technology improves resource utilization efficiency and reduces unit costs, demand often does not decline but may actually grow due to expanded application scope. Applied to AI, reducing inference costs via linear attention mechanisms could drive more enterprises to deploy AI applications, further expanding the global scale of AI inference, ultimately boosting demand for GPUs, HBM, DRAM, and high-speed networking equipment.

However, GDP remains cautious. While acknowledging the long-term logic of Jevons' Paradox, he points out the practical significance of KDA in optimizing state storage for long-context tasks. Even with enormous model weight size, the actual memory pressure might be lower than market intuition suggests, due to the use of 4-bit quantization and highly sparse design.

He believes that the truly critical variable to watch is whether the AI companies with the largest global inference demand—OpenAI, Anthropic, and Google DeepMind—have already adopted or will adopt linear attention mechanisms like KDA or DeepSeek CSA/HCA. If these leading AI companies broadly adopt such architectures, the demand for memory and interconnect resources in long-context inference could see a significant decline, becoming a major factor influencing the future structure of AI hardware demand.

Market Focus Shifts: Can AI Demand Growth Outpace Architectural Efficiency Gains?

Overall, the emergence of Kimi K3 does not simply point towards declining demand for AI hardware. For NVIDIA, the true competitive battleground may be shifting from "single-card performance" to "system-level capabilities," including high-speed interconnect, rack-scale expansion, and large-scale inference optimization.

As model scales continue to grow, even with ongoing optimization of attention mechanisms, AI infrastructure will continue to face pressure from increasing parameter counts, expert parallelism, data exchange, and inference throughput.

The core question the market needs to focus on in the future is not whether linear attention reduces consumption of a single resource, but whether the incremental demand generated by the expansion of AI application scale can consistently surpass the resource savings brought by architectural efficiency improvements.

AI
ยินดีต้อนรับเข้าร่วมชุมชนทางการของ Odaily
กลุ่มสมาชิก
https://t.me/Odaily_News
กลุ่มสนทนา
https://t.me/Odaily_GoldenApe
บัญชีทางการ
https://twitter.com/OdailyChina
กลุ่มสนทนา
https://t.me/Odaily_CryptoPunk
ค้นหา
สารบัญบทความ
ดาวน์โหลดแอพ Odaily พลาเน็ตเดลี่
ให้คนบางกลุ่มเข้าใจ Web3.0 ก่อน
IOS
Android