Marvell Launches AI "Memory Disaggregation" Architecture to Address Agentic AI Inference Bandwidth Bottlenecks
Odaily News: Marvell Technology has announced a new generation of memory solution portfolio for AI infrastructure, covering server-level AI storage, rack-level CXL memory expansion and pooling, as well as multi-cabinet optically interconnected shared memory. The solutions aim to address the growing memory capacity and bandwidth bottlenecks in Agentic AI inference.
Marvell stated that as AI models scale up, context windows lengthen, and KV Cache demand grows, traditional tightly coupled compute-memory architectures are limiting AI inference efficiency. Through memory disaggregation, memory resources can be scaled independently of compute resources, improving GPU utilization and reducing data movement latency. The products announced include:
Bravera SC6 PCIe 6.0 SSD Controller: Designed for AI inference storage scenarios, it helps cloud service providers migrate more KV Cache to high-performance SSDs, improving infrastructure efficiency. The product adopts a multi-vendor NAND-compatible architecture and is expected to begin sampling in Q4 2026.
Marvell Structera X Memory Expansion Solution: Based on CXL technology, it supports rack-level memory expansion and resource pooling, enabling data centers to share and allocate memory resources more flexibly while reducing AI infrastructure costs.
Marvell Photonic Fabric Optically Interconnected Memory Solution: Builds a multi-cabinet shared memory architecture through optical interconnect technology, supporting offloading of up to 32TB of warm KV Cache and helping AI inference clusters improve throughput.
Marvell stated that the Photonic Fabric solution can deliver up to 2-3x token throughput improvement within existing data center space and power constraints, supporting larger-scale models and longer-context AI applications.
Marvell executive Will Chu stated that AI infrastructure is shifting from single-server architectures to systems where compute, memory, and connectivity operate collaboratively. Moving forward, memory will need to scale more independently to improve resource utilization and token efficiency.
As demand for AI agents and large model inference continues to grow, memory capacity, bandwidth, and data transfer efficiency are becoming the new focal points of competition in AI infrastructure, following compute power.
