AMD Zen 6 IBS Memory Profiler: Linux Memory Tiering
AMD is introducing a dedicated hardware memory profiler with its Zen 6 architecture, giving the Linux kernel a more direct way to identify frequently accessed memory pages and optimize their placement across memory tiers.
The new IBS Memory Profiler, presented by AMD engineer Bharata Bhasker Rao at the Linux Plumbers Conference 2026 in Prague, extends AMD’s existing Instruction-Based Sampling (IBS) framework with a lightweight facility dedicated to data memory access sampling.
Unlike conventional profiling approaches that may depend on page-table inspection or memory-access faults to infer page activity, the new hardware can report both virtual and physical addresses, identify the source of memory accesses, and provide information relevant to NUMA placement. These signals can help the operating system distinguish frequently accessed pages in slower memory from pages that can remain in a lower-cost, higher-capacity tier.
AMD is developing integration with the Linux pghot subsystem, a proposed framework for collecting memory hotness information from different sources and coordinating page promotion. The target use case is particularly relevant to servers combining local DRAM with Compute Express Link (CXL) memory expansion.
For background on the underlying architecture, see KAD’s guide to CXL memory, system interconnects, and tiered memory management. The new profiler addresses an important software challenge in that architecture: determining which pages deserve access to the fastest memory tier.
🔬 How the IBS Memory Profiler Works #
Instruction-Based Sampling is an existing AMD processor performance-monitoring mechanism. Traditional IBS provides instruction-fetch and execution-operation sampling, helping developers investigate instruction behavior, execution bottlenecks, and performance characteristics.
Zen 6 adds a separate IBS Memory Profiler designed specifically for data memory accesses. Its independent hardware implementation enables memory-focused sampling without relying exclusively on the primary IBS facility used by other profiling workloads.
Memory Access Information #
According to AMD’s technical documentation and the Linux Plumbers Conference presentation, the profiler can collect several categories of information for sampled accesses:
- Instruction linear address: Identifies the instruction associated with the memory operation.
- Data linear address: Records the virtual address of the accessed data.
- Data physical address: Identifies the physical memory address, allowing the kernel to associate the access with a physical page.
- Memory source information: Distinguishes sources such as cache, local DRAM, and external memory, including lower-tier memory accessed through CXL.
- NUMA-related information: Helps track memory placement and relationships between the processor performing an access and the node containing the accessed data.
The physical address is especially valuable for memory management. The kernel can identify the corresponding physical page frame without having to infer the page solely from an instruction address or depend on an additional address-translation step within the profiling path.
This makes the profiler useful for workloads in which page-placement decisions depend on actual memory-access behavior rather than aggregate CPU utilization.
AMD documents the feature in its AMD64 Zen 6 Instruction-Based Sampling Extensions and Features specification.
Software-Adjustable Sampling #
Collecting information for every memory access would impose unnecessary overhead on most systems. The profiler therefore uses statistical sampling to balance observation coverage against runtime cost.
Software can configure the sampling period, controlling how frequently hardware records memory accesses. A shorter period generally provides more observations, but also increases the amount of profiling data that must be processed.
A longer period reduces the sampling rate, which can be useful when minimizing overhead is more important than capturing every part of a workload’s access pattern.
The hardware also supports filtering mechanisms that reduce the volume of irrelevant observations. These include filtering for L3 cache misses, applying load-latency thresholds, and excluding accesses associated with selected instruction-address classes.
With appropriate filtering, the profiler can focus on accesses more likely to matter for memory placement instead of processing large numbers of cache hits that do not benefit from migration to another memory tier.
A Dedicated Sampling Facility #
The Zen 6 IBS Memory Profiler operates independently of the primary IBS instance and has its own interrupt mechanism.
This separation allows the kernel to use the memory profiler for page-placement decisions without treating it as a replacement for general-purpose performance-monitoring tools.
The implementation described in AMD’s Linux patches also includes per-CPU sample storage, deferred processing, and CPU-hotplug handling. These components help move sampled data into kernel memory-management infrastructure while managing processor lifecycle events.
The feature remains specialized, however. It does not replace instruction-fetch profiling, general-purpose performance tracing, or application-level instrumentation. The initial implementation also has limitations, including the absence of hardware virtualization support and IBS-style buffering functionality.
🧠 Why Hardware-Assisted Memory Tiering Matters #
Modern servers increasingly combine memory technologies with different capacity, bandwidth, latency, and cost characteristics.
Local DRAM typically provides the fastest general-purpose CPU memory access, while CXL-connected memory can expand capacity beyond what is practical with directly attached DRAM alone.
The trade-off is that accesses to a slower memory tier can increase application latency. The operating system therefore needs to distinguish pages that benefit from local DRAM from pages that can safely remain in an expanded memory pool.
The Hot-Page Placement Problem #
A memory page’s importance depends on how frequently and how recently it is accessed, as well as the performance cost of those accesses.
A page accessed repeatedly during a critical application loop may benefit significantly from placement in local DRAM. A page that is rarely accessed may consume valuable fast-memory capacity without contributing much to performance.
A tiered memory policy generally attempts to:
- Identify frequently accessed pages residing in slower memory.
- Promote those pages to a faster tier when sufficient capacity is available.
- Demote less active pages when local memory pressure justifies the move.
- Avoid excessive migration that wastes bandwidth or causes repeated movement between tiers.
The difficult part is obtaining sufficiently accurate access information without creating excessive overhead.
Traditional Linux NUMA balancing can use hinting faults and other mechanisms to discover where memory is accessed. These approaches provide useful information, but the detection process can introduce additional work through page-table operations, faults, or memory scans.
A hardware memory profiler offers another source of evidence, allowing the kernel to observe data accesses directly and feed the resulting information into a broader memory-management policy.
Why Physical Addresses Are Important #
Linux memory management ultimately needs to make decisions about physical pages, not just virtual addresses used by application instructions.
A sampled physical address can help the kernel identify the page being accessed and determine whether that page resides in local DRAM or a slower memory node.
Combined with source information and NUMA topology, this allows the memory-management subsystem to distinguish accesses that indicate potential tiering opportunities from accesses that are already being served efficiently.
For CXL-based memory systems, the distinction is particularly important because the operating system must balance the larger capacity of the external memory tier against the latency-sensitive requirements of active workloads.
KAD’s guide to installing and configuring CXL memory expansion in Linux servers provides additional background on the hardware and operating-system configuration needed before tiering policies can take advantage of this information.
🐧 Integrating the Profiler with Linux pghot
#
AMD is developing support for feeding IBS Memory Profiler samples into pghot, a proposed Linux subsystem for hot-page tracking and promotion.
Rather than implementing a separate migration policy for every available source of access information, pghot aims to provide a shared framework that can accept memory hotness signals from multiple producers.
These sources may include NUMA hint faults, page-table-based tracking, and hardware-assisted sampling mechanisms such as AMD IBS. Other platforms could supply corresponding information through their own supported hardware facilities.
Separating Detection from Migration #
The central design principle behind pghot is to separate the collection of access information from the decision to move pages between memory tiers.
The framework can collect observations, determine which pages are hot, and coordinate promotion through a common mechanism. A dedicated promotion policy can then manage migration rate limits and other decisions without requiring each information source to implement its own complete migration engine.
AMD’s IBS Memory Profiler acts as one of these information sources. It supplies hardware-derived access observations through the pghot_record_access() interface.
This architecture offers several potential benefits:
Multiple detection mechanisms: The subsystem can consume information from different hardware and software sources.
Consistent promotion policy: Page migration decisions can be coordinated through a shared mechanism instead of being duplicated across individual profiling implementations.
Reduced dependence on hinting faults: Hardware samples can provide page-access evidence without requiring every observation to originate from a NUMA hint fault.
Better adaptability to tiered memory: Hardware signals can help the kernel identify pages that are good candidates for promotion from CXL-connected memory to local DRAM.
The framework is intended to be vendor-agnostic at the policy level, even though the IBS Memory Profiler driver itself is specific to compatible AMD Zen 6 processors.
Status of the Linux Implementation #
As of October 10, 2026, the IBS Memory Profiler driver and its pghot integration remain under development rather than forming a generally available, merged mainline Linux feature.
AMD engineer Bharata Bhasker Rao posted an RFC patch series to the Linux kernel mailing lists on September 24, 2026. The series separates the x86-specific profiler driver from the wider pghot development work and documents the hardware interface, sample processing, and integration path.
The proposal is available in the Linux kernel mailing-list archive. The Linux Plumbers Conference session also provides context on the design challenges involved in making hardware-provided page-hotness information available through a common kernel interface.
The current development status is important for platform planning. Availability depends not only on the processor’s hardware capability, but also on kernel support, configuration options, subsystem integration, and the eventual acceptance and maturity of the relevant patches.
📈 Initial Benchmarks and Performance Implications #
AMD’s early testing indicates that the profiler can provide useful memory-placement information with relatively low overhead.
The development report includes measurements from a 256-CPU AMD EPYC Venice server configured with three memory nodes: local DRAM as the upper tier and a CPU-less CXL node as the lower tier.
The benchmark comparison evaluated a baseline without tiering, existing NUMA-balancing-based promotion, and hardware-hint-driven promotion using IBS Memory Profiler samples through pghot.
Reported Workload Results #
The early patch report includes the following normalized performance results, where the no-tiering configuration serves as the 1.00× baseline.
| Workload | NUMA-balancing-based tiering | IBS hardware-hint-driven promotion |
|---|---|---|
| NAS BT | 2.34× | 2.15× |
| Graph500 BFS | 2.34× | 3.17× |
llama.cpp decoding |
1.19× | 1.22× |
| Redis with memtier | 1.04× | 1.00× |
| Pointer-chasing test, sampling period 10,000 | 2.24× | 1.49× |
| Pointer-chasing test, sampling period 5,008 | 2.24× | 3.48× |
| Memory microbenchmark | 2.55× | 2.22× |
These figures are taken from the developer’s preliminary patch report. They are workload-specific comparisons, not universal speedup guarantees.
The results illustrate several important points.
Hardware hints can recover much of the benefit of existing tiering. In the NAS BT test, hardware-hint-driven promotion achieved 2.15× performance relative to the no-tiering baseline, compared with 2.34× for the NUMA-balancing-based approach.
Some workloads can benefit more from hardware sampling. Graph500 BFS reached 3.17× with hardware hints, compared with 2.34× using the existing tiering configuration.
Sampling coverage matters. For the pointer-chasing workload, the shorter sampling period improved the reported speedup from 1.49× to 3.48×. This reflects the importance of capturing enough accesses from the active working set to make effective migration decisions.
Not every workload improves. Redis with memtier remained approximately at the no-tiering baseline under the tested hardware-hint settings. The patch report attributes this result to incomplete coverage of the hot set compared with the existing promotion mechanism.
The report also notes that most configurations were tested with a single run, while llama.cpp and the microbenchmark used averages of three runs. The results should therefore be treated as preliminary evidence rather than a comprehensive independent benchmark.
The broader lesson is that the quality of a hardware sampling source depends on more than its ability to report addresses. Sampling period, filtering, workload access patterns, promotion policy, memory topology, and migration overhead all influence the final performance result.
⚙️ Engineering Challenges and Limitations #
Although dedicated memory sampling provides more direct access information, it does not automatically solve every problem associated with memory tiering.
Sampling Coverage Versus Overhead #
A longer sampling period reduces profiling activity but risks missing important portions of a workload’s hot working set.
A shorter period can provide more observations, but increases sample volume and the processing required to convert that information into useful page-placement decisions.
The optimal sampling configuration therefore depends on workload behavior, the available CPU resources, and the performance requirements of the memory-management policy.
Migration Can Cost More Than the Benefit #
Promoting a page from CXL memory to local DRAM consumes memory bandwidth and requires the kernel to coordinate the migration.
If a page is accessed only once or remains inactive after promotion, the migration may provide little benefit. Excessive movement can also create contention and displace other pages that are more valuable in the faster tier.
A successful implementation must combine access sampling with promotion thresholds, aging, rate limiting, and policies that account for available memory capacity.
Platform and Kernel Compatibility #
The profiler depends on Zen 6 hardware support and compatible kernel driver code. Existing systems cannot obtain the same hardware-derived information simply by enabling a generic Linux profiling option.
The proposed pghot integration also depends on related kernel infrastructure. As the patch series evolves, configuration options, interfaces, sampling controls, and migration behavior may change.
Developers and operators evaluating the feature should track the patch series and verify the exact kernel revision and configuration required for their target hardware.
NUMA Topology and Workload Sensitivity #
The value of hardware-assisted sampling depends on the structure of the memory hierarchy.
A system with local DDR5 and a CXL expansion node has different migration costs from a multi-socket system in which some DRAM is remote but still directly attached to another processor.
Similarly, the access characteristics of a large graph workload may differ substantially from those of a database, language-model inference engine, or pointer-intensive application.
Hardware profiling provides better evidence for making placement decisions, but the policy must still be tuned and evaluated against the real workload.
✅ Conclusion #
AMD Zen 6’s IBS Memory Profiler introduces a dedicated hardware facility for observing data memory accesses and supplying Linux with finer-grained information about page hotness and memory placement.
Its ability to report virtual and physical addresses, memory-source information, and NUMA-related details makes it particularly relevant to systems combining local DRAM with CXL memory expansion.
Through the proposed pghot integration, the profiler can provide hardware-derived access information to a common page-promotion framework instead of requiring each sampling mechanism to implement its own complete migration policy.
Early results from AMD’s EPYC Venice testing show promising behavior, including low-overhead sampling and competitive performance across several tiered-memory workloads. They also demonstrate that sampling coverage and workload characteristics can determine whether hardware-assisted promotion outperforms existing approaches.
Three conclusions stand out:
- Dedicated hardware sampling improves visibility: The kernel can obtain direct information about actual memory accesses and their physical locations.
- The memory-management policy remains essential: Sampling must be combined with effective promotion, aging, and migration controls.
- Mainline availability is still developing: The driver and
pghotintegration require further kernel development, review, and validation.
As server architectures adopt increasingly heterogeneous memory configurations, hardware-software co-design will become more important. Zen 6’s IBS Memory Profiler is a step toward a memory-management model in which CPU hardware supplies precise access evidence and the operating system uses it to place data more intelligently across the memory hierarchy.