CXL Workloads: KV Cache, Memory Pooling, and Tiering
🔎 Introduction #
In our previous article, we explored Compute Express Link (CXL), including its relationship with PCIe and DDR memory, the three protocols (CXL.io, CXL.cache, and CXL.mem), device Types 1–3, and the CXL memory module product lines offered by major memory manufacturers.
Understanding the protocol is only the first step. The more practical question is: Which real-world workloads benefit from CXL, and how can its additional memory capacity justify the associated latency cost?
This question is particularly relevant to large language model (LLM) inference, cloud virtualization, and data-intensive server applications. These workloads often encounter memory-capacity constraints before exhausting their available compute resources.
CXL addresses this problem by extending memory capacity beyond local DRAM and enabling more flexible allocation of memory resources. Depending on the hardware and platform architecture, CXL can also support memory sharing or pooling across hosts.
However, CXL memory is not a drop-in replacement for local DDR or GPU high-bandwidth memory (HBM). Its latency and bandwidth characteristics differ, and cross-host shared-memory implementations may require explicit synchronization and consistency management.
The central design trade-off can be summarized in two properties.
Advantage: Capacity and flexibility. CXL can expand the memory address space available to a host and, with suitable pooling infrastructure, allow memory capacity to be allocated more dynamically across systems.
Cost: Additional latency. Depending on the device, topology, and access path, CXL memory may exhibit approximately 170–300 ns of access latency, or roughly two to three times the latency of local DRAM in representative configurations. Actual results vary by platform, and bandwidth can also be substantially lower than that of local memory.
These properties are two sides of the same trade-off. The important question is not whether CXL is faster than DDR, but whether the additional capacity and improved resource utilization outweigh the cost of accessing slower memory.
Three application scenarios illustrate where CXL can be useful:
- KV cache offloading for LLM inference.
- Memory pooling for virtual machine consolidation.
- Hot/warm/cold memory tiering.
🧠 Use Case 1: KV Cache Offloading for LLM Inference #
One promising application of CXL in AI infrastructure is providing additional capacity for LLM key-value (KV) caches.
Why KV Cache Becomes a Bottleneck #
During autoregressive inference, an LLM generates Key and Value tensors for the tokens it processes. These tensors are retained and reused during subsequent decoding steps, avoiding repeated computation for previously processed context.
The KV cache grows approximately linearly with sequence length, the number of concurrent requests, the number of model layers, and the number of KV heads.
For a conventional transformer implementation, its approximate size can be expressed as:
$$ \text{KV Cache Size} = 2 \times B \times S \times L \times H_{kv} \times D_h \times P $$Where:
- \(B\) is the batch size or number of sequences.
- \(S\) is the sequence length in tokens.
- \(L\) is the number of transformer layers.
- \(H_{kv}\) is the number of key-value heads.
- \(D_h\) is the dimension of each head.
- \(P\) is the number of bytes per stored element.
- The leading factor of 2 accounts for the Key and Value tensors.
This formula assumes a conventional dense KV cache with a uniform representation across layers. Actual implementations may differ because of sliding-window attention, latent attention architectures, quantization, block allocation, and other memory-saving techniques.
Consider an illustrative deployment of Llama 3.1 405B using FP8 model weights and an FP16 KV cache. Under the conventional architecture assumptions above, the cache requires approximately 0.5 MiB per token.
A single request containing 100,000 tokens can therefore consume around 48 GiB of KV cache. Supporting 128 requests of similar length would require approximately 6 TiB of cache capacity.
These figures are illustrative rather than universal. The precise requirement depends on the model configuration, context length, batch size, cache representation, and allocation overhead.
The underlying issue remains the same: GPU HBM is expensive and capacity-constrained. As the KV cache grows, it competes with model weights, activations, and other inference buffers for GPU memory. Every additional byte assigned to the cache can reduce the number of requests the system can serve concurrently.
Limitations of Existing Offloading Strategies #
To relieve HBM pressure, inference systems increasingly move KV cache data into host DRAM or NVMe SSDs.
These approaches extend available capacity, but they introduce different costs.
Host DRAM offloading provides substantially more capacity than GPU HBM in many systems, but moving cache blocks between the CPU and GPU generally requires data transfers over PCIe or another supported interconnect. DMA operations, staging buffers, and transfer scheduling can add overhead.
NVMe SSD offloading provides much larger storage capacity at lower cost per byte, but block I/O introduces substantially higher latency than memory access. Moving cache data between the SSD and GPU can become expensive when the data must be retrieved on a latency-sensitive path.
Both approaches require the runtime to track cache locations, decide which blocks to retain or evict, and coordinate data movement.
The key challenge is therefore not simply finding additional storage capacity. It is placing data in a memory tier that offers the best compromise between capacity, access latency, transfer overhead, and cost.
What CXL Adds to the Memory Hierarchy #
CXL Type 3 memory devices can expose additional memory capacity through the host’s physical address space. When the platform supports and enables the relevant memory mappings, the CPU can access this memory using ordinary load and store instructions instead of submitting block I/O requests.
This creates several opportunities for KV cache management.
Capacity expansion: CXL can provide a large memory tier beyond local DRAM or GPU HBM, depending on the system architecture.
Byte-addressable access: CPU software can access mapped CXL memory at memory granularity, making it suitable for data structures and cache-management techniques that do not naturally fit block-oriented storage.
An intermediate performance tier: CXL memory generally offers a lower-latency access path than SSD storage, while providing additional capacity beyond local DRAM.
CXL does not eliminate the cost of moving KV cache data to and from GPU HBM. GPU accessibility, DMA paths, cache placement, and data-movement scheduling remain important design considerations.
Furthermore, CXL shared memory across hosts should not automatically be treated as a fully coherent shared address space. Depending on the platform, cross-host visibility and synchronization may require software-managed protocols.
Case Study: SK hynix TraCT #
TraCT is a research system that uses CXL shared memory for disaggregated LLM serving. It is designed to address the cost of transferring KV tensors between prefill and decode workers and to support a rack-wide prefix-aware KV cache.
Instead of relying exclusively on network-based transfers, TraCT uses CXL shared memory as a common transfer and caching layer. Its implementation integrates with NVIDIA Dynamo and uses CXL load/store and DMA operations to move KV blocks.
This architecture supports two important use cases:
- Prefill-to-decode KV transfer: KV tensors generated during prompt processing can be made available to separate decode workers through shared memory.
- Rack-level prefix caching: Reusable prefixes can be retained in a shared cache, reducing repeated processing when requests contain previously seen context.
The design also addresses challenges associated with non-coherent shared memory, including synchronization, data visibility, and shared data structures.
In the reported evaluation, TraCT improved average time to first token (TTFT) by up to 9.8 times and peak throughput by up to 1.6 times compared with the evaluated RDMA and DRAM-based caching baselines.
These are workload-specific experimental results, not guarantees for arbitrary production deployments. Nevertheless, they demonstrate how CXL can improve inference efficiency when KV transfer, prefix reuse, or network contention becomes a major bottleneck.
The broader lesson is that CXL can serve as more than a passive memory expansion device. With suitable hardware and software, it can become part of the communication and caching architecture for distributed inference.
🏢 Use Case 2: Memory Pooling and Virtual Machine Consolidation #
A second major CXL workload is memory pooling for cloud servers. Unlike KV cache offloading, which focuses on moving application data between memory tiers, memory pooling primarily targets resource utilization and infrastructure cost.
The Stranded Memory Problem #
Traditional cloud servers are provisioned with a fixed combination of CPU cores and local DRAM.
However, different workloads consume these resources at different rates. A CPU-intensive workload may use most of its allocated cores while leaving a large amount of memory idle. Conversely, a memory-intensive workload may exhaust its DRAM allocation while leaving CPU capacity underutilized.
This mismatch creates the stranded memory problem: one resource becomes a bottleneck while another remains available.
The issue is particularly important in virtualized data centers, where many virtual machines share a physical server. If CPU and memory allocations cannot be adjusted independently enough to reflect workload demand, the operator may need to provision additional servers despite having unused resources elsewhere.
Overprovisioning local DRAM to accommodate peak demand can further increase capital expenditure and power consumption.
How CXL Memory Pooling Works #
CXL 2.0 introduced support for memory pooling architectures that allow physical memory capacity to be allocated more flexibly among connected hosts.
A pooling system can use multi-logical-device (MLD) memory devices and a Fabric Manager (FM) to manage memory allocation according to the capabilities of the hardware and platform.
The important distinction is between dynamic allocation and simultaneous sharing.
In a conventional exclusive-allocation model, a memory resource or allocated region belongs to one host at a time. The system can reassign capacity when demand changes, but it does not necessarily permit multiple hosts to access the same memory region simultaneously.
Exclusive assignment simplifies ownership and isolation and avoids some of the consistency problems associated with concurrent access. True cross-host sharing is a separate capability that requires suitable hardware support and software coordination.
With memory pooling, servers no longer need to provision all their potential peak memory demand as local DRAM. They can use a smaller local-memory configuration and draw additional capacity from a shared pool when workloads require it.
Because different servers rarely reach their peak demand at precisely the same time, pooled memory can improve aggregate utilization.
The architecture also makes it possible to schedule CPU and memory resources more independently, giving the infrastructure scheduler additional flexibility when placing and sizing virtual machines.
Case Study: Microsoft Azure Pond #
Microsoft Research’s Pond examines CXL-based memory pooling for cloud platforms.
The research combines deployment traces from 100 production clusters spanning 75 days with performance measurements from 158 workloads. The evaluation uses a simulated CXL latency environment to investigate the trade-off between local-memory performance and overall DRAM utilization.
Under a reported configuration with memory latency at 2.22 times that of local DRAM, a 16-slot shared pool, and a target performance degradation of 5%, the end-to-end simulation reduced required DRAM capacity by 7%.
The research also reports that this DRAM-cost reduction translates into approximately a 3.5% reduction in total cloud-server cost under its modeled assumptions.
These results illustrate a fundamental characteristic of memory pooling: modest changes to the memory-access path can be worthwhile when they substantially improve resource utilization across a large infrastructure fleet.
The value does not come from making memory access faster. It comes from providing enough additional capacity and allocation flexibility to avoid unnecessary overprovisioning.
For cloud providers, the resulting benefits can include improved server utilization, lower memory procurement costs, and more flexible VM placement.
🗂️ Use Case 3: Hot/Warm/Cold Memory Tiering #
Memory pooling addresses where capacity is allocated. Memory tiering addresses where individual data pages should reside.
Rather than treating CXL memory as a replacement for local DRAM, a tiered architecture uses CXL as an additional memory level with different latency, bandwidth, and capacity characteristics.
Why Memory Tiering Is Necessary #
CXL memory is generally slower than directly attached DDR memory. Placing every allocation in CXL memory can therefore degrade applications that frequently access data on the critical path.
A tiered architecture instead classifies data according to its access characteristics.
| Memory tier | Typical data | Placement strategy |
|---|---|---|
| Hot | Frequently accessed data, latency-sensitive working sets | Keep in local DDR or GPU HBM, depending on the workload |
| Warm | Reusable data with moderate access frequency | Consider CXL placement when capacity or resource utilization justifies the latency |
| Cold | Infrequently accessed data or capacity-dominated working sets | Favor CXL memory when slower access remains acceptable |
The exact classification depends on the application and its access patterns. A dataset that is cold during one workload phase can become hot during another.
From the CPU’s perspective, CXL memory can appear as a slower NUMA memory node, allowing the operating system to use NUMA-aware allocation and migration mechanisms to manage placement.
The goal is not to eliminate access to slower memory. It is to keep latency-sensitive data close to the processor while placing less critical data in a larger, potentially less expensive tier.
Automated Tiering vs. Explicit Allocation #
There are two broad approaches to managing memory placement.
Operating-system-managed tiering #
For general-purpose workloads, automated page placement can reduce the need for application-specific memory management.
Linux can use memory-access monitoring and NUMA mechanisms to classify data, adjust page-reclamation priorities, and move pages between memory tiers. The OS can make placement decisions based on observed access behavior rather than requiring every application to identify hot and cold data explicitly.
One relevant mechanism is DAMON_LRU_SORT, which uses the Linux Data Access MONitor (DAMON) framework to identify regions with different access patterns and adjust their relative priority in the least recently used (LRU) lists.
The Linux kernel documentation describes a default hotness threshold of 50% access frequency and a coldness threshold of 120 seconds without access for this mechanism. These are configurable policy parameters, not universal definitions of hot and cold memory across Linux.
DAMON_LRU_SORT changes page-reclamation priorities; it does not by itself implement every aspect of a CXL tiering policy. Actual movement between memory nodes depends on the additional placement and migration mechanisms configured on the system.
Reference: Linux kernel documentation for DAMON-based LRU-list sorting.
Application-aware placement #
For workloads with predictable access patterns, explicit memory placement can provide more control.
LLM inference is a useful example because the serving runtime often has detailed knowledge of KV cache ownership, request lifecycle, prefix reuse, and the separation between prefill and decode.
An application-aware policy might keep frequently accessed KV blocks in GPU HBM, place reusable prefixes in a CXL-backed shared cache, and retain less frequently accessed data in lower-cost memory tiers.
This strategy can avoid relying exclusively on generic page-access heuristics. It also allows the runtime to distinguish between data that is frequently accessed by the CPU and data whose actual access path involves GPU DMA or device-specific memory operations.
The trade-off is complexity: the application must understand its data lifecycle and coordinate placement with the operating system and hardware.
Avoiding Page Thrashing #
Tiering systems must prevent repeated migration between memory nodes.
Suppose a page is moved from local DDR to CXL because it appears cold. If the application accesses the page shortly afterward, the page may need to return to local DRAM. A poorly tuned placement policy can repeatedly move the same page between tiers, wasting bandwidth and CPU cycles.
Transparent Page Placement (TPP) and related tiering mechanisms can use access information and subsequent fault behavior to guide migration decisions.
A common strategy is to promote a page only after sufficient evidence that it is frequently accessed. Requiring repeated evidence of hotness helps prevent a single access from triggering an unnecessary migration to local DRAM.
Conversely, pages that remain inactive for long periods can be demoted to CXL when local-memory pressure justifies the change.
The effectiveness of these mechanisms depends on the sampling interval, migration overhead, application access patterns, and the relative latency and bandwidth of the memory tiers.
Case Study: Meta Vistara #
Meta’s Vistara project demonstrates a production-oriented application of CXL memory expansion and tiering.
Rather than relying only on newly purchased memory devices, Vistara uses a custom CXL memory-expansion controller to connect reusable DDR4 memory from retired servers to newer systems designed for DDR5.
A reported server configuration combines:
- 768 GB of local DDR5 memory.
- 256 GB of CXL-attached DDR4 memory.
- 1 TB of total memory capacity.
The host system exposes the additional CXL memory as a CPU-less NUMA node. Linux memory-management mechanisms then help keep frequently accessed pages in local DDR5 while placing colder pages in the additional DDR4 tier.
The reported software stack includes Transparent Page Placement (TPP) and Transparent Memory Offloading (TMO), combining reactive page placement with proactive memory offloading.
This design allows a single server to retain latency-sensitive data in faster local memory while using recycled memory to expand total capacity.
Meta’s reported experiments show meaningful benefits in selected workloads. In continuous integration and development environments, enabling CXL increased the number of tasks or virtual machines supported per server by 33%. For machine-learning workloads, the reported configuration reduced the required server count by 25% while increasing throughput by 4%.
These findings demonstrate how memory tiering and hardware reuse can improve data-center economics. The results remain dependent on workload characteristics, the tiering policy, and the hardware configuration; they should not be interpreted as universal performance gains.
Vistara also illustrates that CXL is not exclusively a mechanism for connecting newly manufactured memory. It can provide a standardized high-level interface for integrating different memory technologies into a system, provided the controller and software stack support the required configurations.
⚖️ Which Workloads Are Best Suited to CXL? #
The three use cases reveal a common pattern: CXL is particularly useful when memory capacity or resource utilization is a more significant constraint than raw access latency.
Three workload characteristics are especially relevant.
Capacity Is the Primary Bottleneck #
CXL is attractive when a workload cannot scale because its local memory capacity is exhausted, even though processing resources remain available.
Examples include:
- LLM serving environments limited by KV cache capacity.
- Cloud servers whose local DRAM is overprovisioned for peak demand.
- Large in-memory working sets that exceed economically practical local-memory configurations.
In these cases, accepting slower access to some portion of the working set may be preferable to reducing concurrency, provisioning additional servers, or repeatedly evicting useful data.
The Workload Can Tolerate Higher Latency #
CXL is most beneficial when a substantial part of the data can reside outside the fastest memory tier without violating the application’s latency requirements.
Cold data is an obvious candidate, but the principle is broader. Some warm data can also be placed in CXL if its access frequency is low enough or if the system can overlap data movement with computation.
Conversely, data accessed on every latency-critical operation should generally remain in the fastest practical tier.
For LLM inference, this means that actively used KV blocks and model data on the critical attention path should remain in GPU HBM whenever possible. Reusable prefixes, cache blocks that are temporarily inactive, and other capacity-dominated data may be suitable for offloading, depending on the execution path.
Access Patterns Can Be Managed Effectively #
The value of CXL increases when the system can identify which data should remain local and which can safely move to a slower tier.
Predictable application behavior allows explicit placement policies, while general-purpose workloads can benefit from OS-managed monitoring and migration.
In both cases, the placement mechanism must account for the cost of moving data. An overly aggressive policy may perform more migration than the avoided latency or capacity pressure justifies.
The best design therefore combines capacity planning, access-pattern analysis, and workload-specific performance measurements.
🔗 CXL and HBF: Complementary Approaches to Memory Expansion #
CXL and High Bandwidth Flash (HBF) address different parts of the memory-capacity problem.
In AI systems, HBM provides the high-bandwidth memory used directly by GPUs. HBF is intended to provide another high-capacity tier closer to the accelerator-side memory hierarchy, while CXL expands memory capacity on the CPU side and can support broader resource sharing.
The two technologies differ in physical placement, access mechanisms, bandwidth, and latency. However, they follow a similar architectural principle: introduce a larger memory tier beneath the fastest, most expensive memory rather than attempting to replace that tier entirely.
For GPU-side workloads, an additional accelerator-oriented tier can help expand available capacity. For CPU-side workloads and server infrastructure, CXL can expose additional memory capacity, facilitate memory pooling, and support OS- or application-managed tiering.
These approaches are complementary rather than mutually exclusive.
The right memory architecture depends on where the capacity bottleneck occurs, which processor accesses the data, how frequently the data is reused, and how much latency the application can tolerate.
✅ Conclusion #
CXL’s value becomes clearer when it is evaluated against concrete workload requirements rather than as a standalone interconnect technology.
The three application scenarios demonstrate its potential:
- KV cache offloading: CXL can provide additional capacity for LLM inference, support reusable prefix caching, and reduce the cost of transferring KV tensors between prefill and decode workers.
- Memory pooling: CXL can help cloud providers reduce stranded memory capacity by allocating shared resources more flexibly across servers and virtual machines.
- Memory tiering: CXL can hold warm and cold data outside local DRAM while allowing operating systems and applications to keep latency-sensitive working sets in faster memory.
The common conditions are straightforward: capacity must be a significant bottleneck, the workload must tolerate the additional latency for at least part of its data, and the software stack must manage placement and movement effectively.
CXL does not eliminate the need for fast memory. Instead, it expands the available hierarchy beneath expensive local DRAM and GPU HBM, making more capacity accessible and improving the flexibility with which memory resources can be used.
Ultimately, the success of a CXL deployment depends less on the protocol alone than on the interaction between hardware topology, memory allocation, workload behavior, and software placement policies. When these elements are designed together, CXL can turn otherwise stranded capacity into a useful resource for AI serving and large-scale cloud infrastructure.