Marvell’s new portfolio tackles AI’s memory bottleneck instead of adding more compute

Home Semiconductor News Marvell’s new portfolio tackles AI’s memory bottleneck instead of adding more compute
Marvell Technology

Marvell’s four-tier KV cache architecture spans SSDs, CXL pooling, and optical fabric

In sum – what we know:

  • A four-tier cache – Marvell’s architecture layers HBM, CXL memory, optical shared memory, and PCIe 6.0 SSDs to keep hot tokens close and cold cache cheap.
  • Optical shared memory – The Photonic Fabric pools up to 32TB of warm KV cache across racks at ~30 meters, with claimed 2–3× inference throughput gains.
  • The neutral play – Unlike Nvidia’s NVLink and AMD’s Infinity Fabric, Marvell’s portfolio is built to plug into any accelerator ecosystem without picking a side.

Marvell Technology has unveiled a new AI memory and storage infrastructure portfolio at the “Future of Memory and Storage” (FMS) event. The portfolio spans CXL-based memory expansion, PCIe 6.0 SSD controllers, and an optical shared memory fabric — all aimed at hyperscalers and cloud providers running large language models and the newer wave of agentic AI applications.

Marvell argues the core bottleneck for AI infrastructure is shifting away from insufficient compute and toward slow data transmission, slow retrieval, and limited memory — particularly for the key-value (KV) caches that balloon as context windows grow. By decoupling memory capacity and bandwidth from compute scaling, the company says operators can improve token throughput, efficiency, and total cost of ownership without simply buying more accelerators.

Bravera SC6 PCIe 6.0 SSD controller

The Bravera SC6 is a new enterprise SSD controller built around a specific job — moving KV cache off accelerators and onto SSDs. That matters because high-bandwidth memory on GPUs and XPUs is expensive and perpetually scarce, and a growing share of it is being consumed by inference state rather than active computation. Offloading colder cache to flash frees that HBM up for the work it’s actually good at.

Marvell says the SC6 doubles the performance of the widely deployed Bravera SC5, with PCIe 6.0 support providing the throughput and latency headroom to make SSD-resident KV cache viable for large-model serving. Whether flash can genuinely keep inference pipelines fed will depend on real-world benchmarks, which Marvell hasn’t published yet — but the architectural logic is sound, and the KV cache problem is real. The Bravera SC6 is expected to begin sampling at the end of this year.

Structera CXL memory portfolio and pooling

Structera is the more mature part of the story. The portfolio spans three lines — Structera X for CXL memory expansion, Structera A for near-memory acceleration, and Structera S for CXL switching and rack-level memory pooling — with inline hardware compression built into the X and A devices to stretch capacity further. The Structera X platform, developed alongside hyperscaler partners, is where CXL is now crossing from evaluation into real deployment. One practical draw is that Structera lets operators extend the life of existing DDR4 investments while still supporting demanding AI workloads, which is a meaningful TCO lever when memory budgets are ballooning.

At FMS, Marvell is extending Structera X for rack-scale memory expansion and pooling, pushing the architecture toward PCIe 6.0 and CXL 3.x connectivity and more advanced multi-host memory sharing. The bigger idea is memory pooling — letting multiple servers and XPUs draw from shared memory resources rather than being confined to whatever DRAM sits on their own boards. (Marvell’s dedicated CXL switch for this, the Structera S 30260, got its live debut back at OFC in the spring.) In other words, memory stops being a restricted per-node resource and becomes a flexible rack-level one.

That shift is particularly relevant for agentic AI, where workloads jump across systems and need fast access to shared state. An agent that hands off tasks between nodes doesn’t want to rebuild its context from scratch every time.

Photonic Fabric and optical shared memory

Photonic Fabric is the most novel piece of the portfolio, and arguably the most differentiated. It creates a new shared-memory tier using optical interconnects, enabling pod-level memory sharing across multiple XPUs and racks at distances up to roughly 30 meters — well beyond what electrical interconnects can manage. Marvell says the fabric supports up to 32 TB of “warm” KV cache offloaded from local memory into the shared tier.

The headline claim is 2–3× higher token throughput for AI inference, achieved within existing data center power envelopes and footprints. That’s a Marvell number, not an independently validated one, and it deserves the usual scrutiny until customer deployments back it up. But the underlying appeal is clear enough. Scaling memory this way sidesteps the constant GPU upgrade cycle and the ever-larger, power-hungrier accelerator clusters that hyperscalers are currently locked into.

It’s also where Marvell diverges most sharply from Nvidia’s NVLink and AMD’s Infinity Fabric, both of which rely on electrical interconnects and memory that lives in or near the box. An optical shared memory fabric plays a different game — trading some proximity for reach and pooling flexibility.

Octeon DPU and multi-tier KV cache management

Rounding out the portfolio is the Octeon DPU, which offloads storage and networking tasks from CPUs and handles high-bandwidth, low-latency data movement between compute nodes and storage systems. That plumbing matters more than it sounds when KV cache and inference state are scattered across SSDs, CXL pools, and optical fabric rather than sitting in one place.

Taken together, the pieces form a four-tier KV cache architecture:

  • Tier 1: HBM and local DRAM on the GPU or accelerator — fastest access, but limited and expensive.
  • Tier 2: CXL-attached memory expansion via Structera, for larger caches at still-low latency.
  • Tier 3: Photonic Fabric shared memory across racks, holding up to 32 TB of warm KV cache.
  • Tier 4: PCIe 6.0 SSDs via Bravera SC6, for colder cache and inference state.

The goal is to keep the hottest tokens and context close to compute while giving everything else a cheaper, scalable reservoir. It’s the same tiering logic storage has used for decades, applied to inference state.

Competitive positioning

Strategically, Marvell is positioning itself as the neutral infrastructure supplier. Nvidia and AMD each want to sell you their own vertically integrated interconnect and memory hierarchy; Marvell’s portfolio is designed to plug into Nvidia, AMD, or custom XPU ecosystems without picking a side. In a market where hyperscalers are increasingly building their own accelerators, that neutrality has real value.

It also completes a narrative Marvell has been assembling for over a year. The company’s earlier 2nm custom SRAM work targeted memory bandwidth inside the chip; this portfolio extends the story off-chip, across servers, racks, and pods. Whether the 2–3× throughput claims hold up in production remains to be seen, but the direction is hard to argue with. The industry has spent years fixated on GPU counts. Marvell is betting the next fight is over where the data lives — and how fast it can get to the compute.

What you need to know in 5 minutes

Join 37,000+ professionals receiving the AI Infrastructure Daily Newsletter

This field is for validation purposes and should be left unchanged.

This website uses cookies to improve your experience. We'll assume you're ok with this, but you can opt-out if you wish. Accept Read More