Skip to main content

Marvell CXL Switch Enables 48TB AI Memory Pools

·2229 words·11 mins
Marvell CXL AI Infrastructure AI Memory PCIe 6.0 KV Cache Data Centers Hyperscalers
Table of Contents

Marvell CXL Switch Enables 48TB AI Memory Pools

Marvell is positioning memory as a distinct infrastructure layer for next-generation AI systems.

At FMS 2026, the company expanded its AI memory portfolio across three levels: PCIe 6.0 SSD controllers for server storage, CXL-based memory expansion and pooling for racks, and photonic shared memory for larger-scale systems. The strategy targets one of the most important bottlenecks emerging in agentic AI inference: KV cache capacity and bandwidth.

The portfolio includes the Bravera SC6 PCIe 6.0 SSD controller, the Structera CXL memory family, and Photonic Fabric shared-memory technology. Among the announcements, Structera is particularly significant because Marvell says its CXL 3.x switching technology can connect up to 16 or 32 hosts and create shared memory pools reaching 48TB.

Marvell also claims that Structera-based GPU memory pooling can increase inference throughput by up to 4.8x while reducing time-to-first-token (TTFT) by 82.7% in its benchmark configurations. These figures are vendor-published results, so independent production-scale validation remains important.

🧠 Marvell Builds a Three-Tier AI Memory Architecture
#

Marvell’s FMS 2026 portfolio is structured around the idea that AI memory should scale independently from compute.

The three major layers are:

  1. Server level: Bravera SC6 PCIe 6.0 SSD controllers.
  2. Rack level: Structera CXL memory expansion and pooling.
  3. Cabinet level: Photonic Fabric shared memory across XPUs and racks.

This approach reflects a broader change in AI infrastructure. As models become larger and inference sessions maintain longer contexts, memory requirements increasingly become a limiting factor even when sufficient compute resources are available.

For agentic AI, the problem is particularly pronounced because KV caches can remain resident for extended periods and consume substantial amounts of high-bandwidth memory.

Marvell’s strategy is therefore to create additional memory tiers that can be attached, pooled, and accessed independently of local CPU or GPU memory.

Bravera SC6 Targets PCIe 6.0 AI Storage
#

The Bravera SC6 is Marvell’s next-generation SSD controller designed for PCIe Gen 6 and NVMe 2.2.

According to Marvell, the controller provides approximately twice the performance of the PCIe 5.0-generation Bravera SC5. It supports NAND from multiple vendors and delivers data-transfer rates of up to 3,600 MT/s across 16 channels.

The controller integrates 15 Arm CPU cores, including 12 Cortex-R82 cores, and is expected to begin sampling in Q4 2026.

Its vendor-neutral NAND architecture is particularly relevant to hyperscalers, which generally prefer flexibility in memory and storage sourcing rather than architectures tied to a single NAND supplier.

Photonic Fabric Extends Shared Memory Across Racks
#

At the larger system level, Marvell’s Photonic Fabric technology creates a shared memory layer across multiple XPUs and racks separated by distances of up to 50 meters.

Marvell says the technology can provide up to 32TB of warm KV-cache capacity and increase token throughput by approximately 2x to 3x within existing power envelopes.

This represents a different approach from simply adding more local HBM. Instead of forcing every accelerator to carry the maximum amount of expensive high-bandwidth memory, shared memory allows infrastructure designers to create additional capacity tiers that can be dynamically accessed as workloads require.

🔗 Structera Moves CXL From Evaluation Toward Deployment
#

The most consequential part of Marvell’s announcement is arguably the Structera family.

Marvell says the Structera X 2404 and 2504 platforms have already begun shipping to hyperscale cloud providers. If sustained, this would indicate that CXL memory expansion is moving beyond technology evaluation toward real infrastructure deployment.

Structera X Supports DDR4 and DDR5 Expansion
#

The Structera X family provides CXL-attached memory using conventional DDR memory resources.

The Structera X 2404 uses DDR4 and provides a mechanism for hyperscalers to reuse existing memory infrastructure rather than immediately migrating every system to higher-cost DDR5.

The Structera X 2504 targets DDR5-based deployments where additional memory bandwidth justifies the newer memory technology.

Both platforms provide four DDR channels per controller and communicate with host systems through standard CXL interfaces.

This separation between compute and memory creates a more flexible infrastructure model. Operators can increase memory capacity without replacing the host CPU platform or redesigning the entire server architecture.

GPU Memory Pooling Shows Large Potential Gains
#

Marvell reports that Structera-based GPU memory pooling achieved:

  • Up to 4.8x higher inference throughput
  • Up to 82.7% lower time-to-first-token
  • Expanded memory capacity beyond local GPU memory
  • More efficient utilization of existing accelerator resources

These results should be interpreted as vendor benchmark data rather than evidence of universal production gains. The company has not publicly disclosed the specific hyperscalers involved, deployment scale, or complete workload configurations behind the results.

Nevertheless, the magnitude of the reported improvement illustrates why CXL memory pooling is attracting attention as AI inference becomes increasingly memory-bound.

📦 CXL Could Become a Standard Memory Expansion Layer
#

The economics of CXL become more interesting as server memory requirements continue to rise.

Each Structera X controller provides four DDR channels. The DDR5-based 2504 can support up to eight DIMMs per controller, while the DDR4-based 2404 can support up to three DIMMs per channel.

Depending on DIMM density and configuration, a single expansion card can approach 1TB of memory capacity.

A fully populated server requiring several terabytes of additional memory could therefore use multiple CXL expansion controllers connected to the same CPU socket.

Multi-Terabyte Memory Requirements Are Becoming Normal
#

Large language models with trillion-parameter-scale architectures and long context windows can produce working sets measured in multiple terabytes.

At an illustrative 8TB memory requirement per socket, a server could require several CXL controllers depending on DIMM density and configuration.

This creates a potential scaling model in which CXL controllers become a standard server connectivity component, similar to network adapters.

The underlying economic relationship is straightforward: as the number of servers increases, demand for memory expansion controllers increases alongside them.

The difference is that CXL allows memory capacity to scale more independently from CPU and GPU configurations.

Structera A Adds Near-Memory Compute
#

Marvell’s portfolio extends beyond passive memory expansion.

The Structera A 2504 incorporates 16 Arm Neoverse V2 cores running at up to 3.2 GHz. It supports up to 4TB of memory and up to 200 GB/s of bandwidth for workloads such as deep-learning recommendation models and vector search.

This architecture places compute closer to expanded memory, reducing the need to move every operation back through the primary CPU or GPU.

The concept becomes increasingly relevant for workloads where data movement, rather than arithmetic throughput, is the dominant bottleneck.

Structera S Creates a Large Shared Memory Pool
#

The Structera S 30260 represents the highest-capacity CXL component in the announced portfolio.

The CXL 3.x switch can connect 16 or 32 hosts and create a shared memory pool of up to 48TB.

Marvell specifies total bandwidth of up to 4TB/s, with unidirectional round-trip latency below 460 nanoseconds.

This creates a fundamentally different memory hierarchy:

  • Local HBM provides the highest bandwidth and lowest latency.
  • Local DRAM provides general-purpose system memory.
  • CXL-attached memory provides expanded capacity.
  • Shared CXL pools provide capacity that can be accessed across multiple hosts.

The value proposition is therefore not simply “more memory.” It is the ability to allocate memory capacity independently of individual CPU or GPU sockets.

🏗️ Hyperscaler Economics Could Accelerate CXL Adoption
#

One of CXL’s strongest advantages is its potential to improve memory utilization.

Traditional server architectures often provision memory for peak requirements, leaving substantial capacity underutilized during normal operation.

CXL makes it possible to separate memory capacity from compute resources and dynamically allocate additional memory where workloads require it.

The DDR4-based Structera X 2404 is particularly interesting from this perspective because it can reuse existing DDR4 memory assets.

Rather than discarding older memory infrastructure when servers are upgraded, hyperscalers can potentially convert that capacity into externally attached memory pools.

This creates an economic incentive independent of raw bandwidth improvements.

CXL Becomes More Attractive as AI Inference Shifts Toward Memory
#

Training workloads traditionally dominate discussions around accelerator throughput and HBM bandwidth.

Inference introduces a different optimization problem.

Long-context and agentic workloads can maintain large KV caches while simultaneously serving many concurrent users. Increasing accelerator utilization therefore depends not only on compute throughput but also on keeping the required state available at an acceptable latency and cost.

This creates an opportunity for CXL to act as an intermediate memory tier between expensive local accelerator memory and slower storage.

The resulting architecture resembles a hierarchical memory system rather than a GPU-centric design.

⚔️ Marvell Faces Competition From CXL and GPU Ecosystems
#

Marvell’s broad portfolio gives it exposure across multiple memory infrastructure layers, but the competitive environment is already crowded.

Astera Labs operates in the CXL memory controller and connectivity market, while XConn develops CXL switching and hub technologies. Montage Technology also provides CXL memory expansion solutions.

Memory manufacturers including Samsung and SK hynix are simultaneously developing CXL memory products and ecosystem partnerships.

Marvell’s interoperability across Intel Xeon, AMD EPYC, and Arm platforms is therefore strategically important. A vendor-neutral architecture allows hyperscalers to deploy the technology across heterogeneous CPU environments rather than locking memory infrastructure to one processor vendor.

NVIDIA Represents a Different Architectural Model
#

The larger competitive question involves NVIDIA’s tightly integrated GPU memory architecture.

NVIDIA keeps much of the AI memory hierarchy close to the accelerator through HBM, NVLink, and Grace-connected memory.

This architecture offers extremely high bandwidth and tightly optimized communication, but it also creates a more vertically integrated system.

Marvell is pursuing the opposite direction: make memory a more open, composable infrastructure layer that can be shared across processors and accelerators.

Photonic Fabric represents the company’s attempt to extend that concept across larger physical domains.

The outcome may depend on whether hyperscalers prefer open, composable memory infrastructure or continue adopting increasingly integrated accelerator platforms.

💡 The Real Opportunity Is KV Cache Offloading
#

The most important long-term application for Marvell’s architecture may not be conventional server memory expansion.

It is KV cache offloading for AI inference.

As context windows expand, the KV cache can become one of the largest consumers of accelerator memory. Keeping every cache locally in HBM is expensive and can force operators to provision more accelerator memory than average workloads actually require.

A hierarchical architecture could instead keep hot KV data close to the accelerator while moving less frequently accessed cache segments into CXL or photonic shared memory.

This approach could improve memory utilization without requiring every GPU to carry maximum-capacity HBM.

The effectiveness of this model will ultimately depend on software maturity, cache-management policies, latency tolerance, and the ability of AI frameworks to exploit heterogeneous memory transparently.

🔬 Key Technical Constraints to Watch
#

Despite the potential, several factors could limit CXL’s expansion into AI infrastructure.

Latency Versus Local DRAM and HBM
#

CXL-attached memory is inherently farther from the processor than local DRAM, while shared memory introduces additional switching and fabric latency.

For latency-sensitive workloads, additional capacity is useful only if software can intelligently determine which data belongs in which memory tier.

CXL Software Maturity
#

Hardware deployment alone does not create a usable memory pool.

Operating systems, hypervisors, schedulers, AI frameworks, and memory-management software must understand CXL topology and intelligently allocate workloads across local and pooled memory.

Software maturity will therefore be a major determinant of real-world adoption.

Custom Silicon From Hyperscalers
#

Large cloud providers increasingly design custom accelerators, networking chips, and memory subsystems.

If hyperscalers decide to integrate CXL functionality directly into custom infrastructure, merchant silicon vendors could face margin and differentiation pressure.

Marvell’s advantage is the breadth of its portfolio, but that advantage must translate into measurable deployment and total-cost benefits.

📊 What to Watch Through 2027
#

Several developments will reveal whether Marvell’s AI memory strategy can move from product announcements to sustained infrastructure demand:

  • Structera X deployment disclosures: Public confirmation of volume hyperscaler deployments would strengthen Marvell’s claim that CXL is moving from evaluation to production.
  • Bravera SC6 sampling: Q4 2026 sampling and successful qualification across multiple NAND suppliers will test Marvell’s PCIe 6.0 storage strategy.
  • Production benchmark validation: Independent measurements of the reported 4.8x inference-throughput improvement and 82.7% TTFT reduction will be critical.
  • NVIDIA’s KV-cache architecture: Future rack-scale systems will reveal whether shared memory remains an open infrastructure layer or becomes increasingly tied to proprietary GPU ecosystems.
  • CXL 3.x competition: Products from Astera Labs, XConn, and other vendors will determine how quickly shared-memory fabrics become standardized infrastructure.
  • AI software support: Framework-level support for tiered memory and KV-cache placement may ultimately matter more than raw hardware capacity.

🚀 Conclusion
#

Marvell’s FMS 2026 announcements point toward a broader architectural shift in AI infrastructure: memory is becoming a composable resource rather than a fixed property of each compute device.

The Structera S 30260’s ability to connect up to 32 hosts and create a 48TB shared CXL memory pool illustrates how far this model can scale. Meanwhile, the Bravera SC6 addresses high-speed storage, and Photonic Fabric extends memory sharing across multiple XPUs and racks.

The strategic significance goes beyond the individual products. If agentic AI inference continues to increase KV-cache requirements faster than local HBM capacity can economically scale, hyperscalers will need additional memory tiers.

CXL provides a standards-based path toward that architecture.

The remaining question is whether its latency, software complexity, and ecosystem maturity can compete with tightly integrated GPU memory systems. Marvell has assembled one of the broadest merchant-silicon portfolios targeting this transition. The next phase will be determined not by specifications alone, but by whether hyperscalers deploy these technologies at production scale and achieve measurable improvements in inference economics.

Related

Meta Vistara: Reusing DDR4 Memory with CXL at Hyperscale
·1683 words·8 mins
Meta CXL Data Centers Memory Tiering DDR4 Linux Kernel AI Infrastructure Server Hardware
Eaton and the AI Data Center Power Bottleneck: The Hidden Infrastructure Layer
·835 words·4 mins
AI Infrastructure Data Centers Eaton Power Systems Electrical Engineering Hyperscalers UPS Grid Architecture Semiconductors Ecosystem 800VDC
From Optical Interconnects to Optical Computing: The Photonics Era
·1024 words·5 mins
Photonics Optical Computing CPO AI Infrastructure Data Centers CXL Interconnects Semiconductors High-Performance Computing AI Scaling