BTC
ETH
HTX
SOL
BNB
시장 동향 보기
简中
繁中
English
日本語
한국어
ภาษาไทย
Tiếng Việt

Marvell Unveils AI "Memory Disaggregation" Architecture to Solve Agentic AI Inference Bandwidth Bottlenecks

2026-08-04 13:17

Odaily News – Marvell Technology has announced a new generation of memory solution portfolio for AI infrastructure, covering server-level AI storage, rack-level CXL memory expansion and pooling, and multi-cabinet optical interconnect shared memory, aimed at addressing the growing memory capacity and bandwidth bottlenecks in Agentic AI inference.

Marvell stated that as AI models scale up, context windows lengthen, and KV Cache demand grows, the traditional tightly coupled compute-memory architecture is limiting AI inference efficiency. Through memory disaggregation, memory resources can scale more independently from compute resources, improving GPU utilization and reducing data movement latency. The products announced this time include:

Bravera SC6 PCIe 6.0 SSD Controller: Designed for AI inference storage scenarios, it helps cloud service providers migrate more KV Cache to high-performance SSDs, improving infrastructure efficiency. The product adopts a multi-vendor NAND-compatible architecture and is expected to begin sampling in Q4 2026.

Marvell Structera X Memory Expansion Solution: Based on CXL technology, it supports rack-level memory expansion and resource pooling, helping data centers share and allocate memory resources more flexibly while reducing AI infrastructure costs.

Marvell Photonic Fabric Optical Interconnect Memory Solution: Builds a cross-cabinet shared memory architecture through optical interconnect technology, supporting up to 32TB of warm KV Cache offloading and helping AI inference clusters improve throughput.

Marvell stated that the Photonic Fabric solution can achieve up to 2 to 3 times token throughput improvement within existing data center space and power constraints, supporting larger-scale models and longer-context AI applications.

Marvell executive Will Chu stated that AI infrastructure is shifting from single-server architectures to systems where compute, memory, and connectivity operate collaboratively. In the future, memory needs to scale more independently to improve resource utilization and token efficiency.

As demand for AI agents and large model inference continues to grow, memory capacity, bandwidth, and data transfer efficiency are becoming the new focus of competition in AI infrastructure, following compute power.