↓ Skip to main content

NVIDIA Vera vs. EPYC and Xeon: Why 176 Cores Win at HPC

NVIDIA Vera vs. EPYC and Xeon: Why 176 Cores Win at HPC

NVIDIA’s Vera CPU is challenging a long-standing assumption in high-performance computing (HPC): that more CPU cores and higher clock frequencies necessarily translate into better application performance.

In a series of benchmarks published by Phoronix on October 6, 2026, a dual-socket NVIDIA Vera server was compared with flagship AMD EPYC and Intel Xeon systems across computational fluid dynamics, weather forecasting, numerical computing, and other demanding workloads. Despite having only 176 physical cores compared with 256 cores in the competing dual-socket EPYC 9755 and Xeon 6980P systems, Vera delivered competitive results and led in several memory-intensive workloads.

The explanation lies in the combination of NVIDIA’s custom Olympus CPU cores, high-bandwidth LPDDR5X memory, and a system architecture designed to keep processor cores supplied with data. Vera’s memory subsystem offers up to 1.2 TB/s of bandwidth per CPU, substantially exceeding the bandwidth available from the DDR5-6400 configuration used by the tested AMD systems.

The findings also reveal an important distinction in server CPU evaluation: core count, clock frequency, and peak throughput do not independently determine real-world performance. Memory bandwidth, workload characteristics, software configuration, and power efficiency can be equally important.

For additional background, KAD’s article on NVIDIA Vera and its entry into the server CPU market examines the processor’s broader architectural goals. Its analysis of 2026 data-center CPU trends and architecture shifts provides further context on the growing importance of memory bandwidth and heterogeneous computing.

🧪 Test Platforms and Configurations
#

The benchmark compares four dual-socket server configurations: NVIDIA Vera, AMD EPYC 9575F, AMD EPYC 9755, and Intel Xeon 6980P.

The systems differ in CPU architecture, memory technology, operating-system configuration, and maximum clock frequency. These differences are important when interpreting the results because the benchmark measures complete platforms rather than isolated CPU cores.

Hardware Comparison
#

Platform CPU configuration Total cores Total threads Memory configuration
NVIDIA Vera 2P 2 × 88-core Vera 176 352 16 × 96 GB Samsung LPDDR5-9600
AMD EPYC 9575F 2P 2 × 64-core EPYC 9575F 128 256 24 × 64 GB Samsung DDR5-6400
AMD EPYC 9755 2P 2 × 128-core EPYC 9755 256 512 24 × 64 GB Samsung DDR5-6400
Intel Xeon 6980P 2P 2 × 128-core Xeon 6980P 256 512 24 × 64 GB Micron MRDIMM-8800

Each configuration provides approximately 1.5 TB of installed memory. Vera’s main distinction is therefore not simply memory capacity, but the bandwidth and power characteristics of its LPDDR5X-based memory subsystem.

The EPYC 9575F represents a high-frequency alternative with fewer cores. The EPYC 9755 and Xeon 6980P represent higher-core-count platforms intended to maximize throughput across heavily threaded server workloads.

Software and Compiler Environment
#

Phoronix used the following software configurations:

Component NVIDIA Vera AMD EPYC and Intel Xeon
Operating system Customized Ubuntu 24.04 LTS Ubuntu 26.04 LTS
Linux kernel Patched Linux 6.17-based kernel Distribution-provided kernel
Kernel page size 64 KB Native configuration
Compiler GCC 15.3 GCC 15-series compiler
Benchmark settings Stock compiler flags reported for the tests Modern default environment

The Vera system used a customized Linux environment aligned with NVIDIA’s recommended software configuration for the platform. The x86 systems used their respective Ubuntu environments and compiler defaults.

Consequently, the results represent performance under the tested software stacks. They should not be interpreted as perfectly controlled measurements of CPU microarchitecture alone. Kernel configuration, compiler versions, optimization behavior, and memory subsystems can all influence HPC performance.

For the full methodology and individual benchmark graphs, refer to the original Phoronix NVIDIA Vera HPC benchmark report.

🚀 Why Vera Excels in Memory-Bound HPC Workloads
#

The defining characteristic of Vera is its high-bandwidth memory subsystem.

HPC applications frequently operate on large datasets and repeatedly access arrays, grids, sparse matrices, and other memory-intensive structures. When a workload cannot fully utilize the processor’s arithmetic units because it is waiting for data, adding more cores or increasing clock frequency may deliver diminishing returns.

In these situations, memory bandwidth determines how effectively the CPU can execute the workload.

LPDDR5X Memory Bandwidth
#

NVIDIA Vera supports up to 1.2 TB/s of memory bandwidth per CPU through its LPDDR5X memory subsystem. A dual-socket system combines two processors, each with its own local memory resources.

This design provides substantial bandwidth for applications involving large working sets, repeated memory accesses, and bandwidth-intensive numerical operations.

The tested systems use different memory technologies:

  • NVIDIA Vera: LPDDR5-9600 memory, as identified in Phoronix’s test configuration.
  • AMD EPYC: DDR5-6400 memory.
  • Intel Xeon: MRDIMM-8800 memory.

The higher transfer rate of Vera’s memory contributes to its advantages in several benchmarks. However, transfer rate alone does not determine system bandwidth; channel count, memory-controller design, access patterns, and NUMA placement also matter.

LPDDR5X is particularly attractive in this context because it combines high bandwidth with comparatively low memory power consumption. It allows Vera to supply data to its cores at a rate that traditional server memory configurations may struggle to match without specialized high-bandwidth modules.

Olympus CPU Architecture
#

Vera uses NVIDIA-designed Olympus CPU cores rather than conventional Arm Neoverse cores. Its architecture is intended to sustain high instructions per cycle, particularly under demanding workloads involving irregular memory accesses and complex software execution.

NVIDIA describes Olympus as incorporating a wide instruction front end, advanced branch prediction, deep out-of-order execution, and memory-prefetching mechanisms designed to reduce stalls.

These features matter in HPC because real-world performance depends on how effectively a core converts available clock cycles into completed work.

A processor with fewer cores can compete against a higher-core-count alternative when its cores execute useful instructions more efficiently and receive data quickly enough to avoid unnecessary stalls.

The combination of Olympus cores and high-bandwidth memory gives Vera a credible path to competitive HPC performance without matching the core count of the largest EPYC and Xeon configurations.

🌦️ WRF and HPCG: Memory Bandwidth Delivers Results
#

Two of the most revealing tests in the Phoronix suite are WRF and HPCG. Both help demonstrate why memory behavior can be more important than core count alone.

WRF: Weather Research and Forecasting
#

The Weather Research and Forecasting (WRF) model simulates atmospheric processes using computationally demanding numerical methods.

WRF performance depends on several factors, including floating-point execution, memory bandwidth, memory latency, and how efficiently the workload is distributed across CPU cores.

In Phoronix’s test, the dual-socket Vera system delivered the fastest WRF performance recorded in that particular benchmark evaluation. It also led in performance per watt.

The result is especially notable because the Vera system has fewer cores than the dual EPYC 9755 and Xeon 6980P systems.

Phoronix identified Vera’s LPDDR5-9600 memory as an important contributor, particularly compared with the EPYC systems configured with DDR5-6400. The results suggest that the WRF workload benefited substantially from the additional memory bandwidth.

This does not establish that Vera will lead in every WRF configuration. Grid resolution, compiler optimizations, MPI communication, thread placement, and memory locality can alter the relative performance of competing systems.

Nevertheless, the measured result shows that a high-bandwidth memory architecture can deliver a meaningful advantage in conventional scientific computing.

HPCG: High Performance Conjugate Gradients
#

HPCG evaluates performance using computational and memory-access patterns associated with iterative linear-system solvers. It is generally more representative of memory-sensitive numerical operations than a benchmark designed primarily to maximize dense arithmetic throughput.

Vera performed strongly in the HPCG evaluation, reinforcing the importance of its memory subsystem.

In memory-intensive tests, a CPU may be limited by the rate at which it can retrieve and update data rather than by the number of arithmetic operations it can theoretically execute per second.

HPCG therefore provides another example of why the ranking of server CPUs can differ significantly from their ranking by core count or peak clock frequency.

Future comparisons with next-generation EPYC processors using higher-speed MRDIMM memory and subsequent Intel Xeon platforms will help establish whether Vera retains its advantage when competing systems gain additional memory bandwidth.

🧮 OpenFOAM: Performance Depends on Mesh Size
#

OpenFOAM is a computational fluid dynamics (CFD) application used to simulate fluid flow and related physical processes. Its performance depends on the numerical algorithms, mesh size, memory-access characteristics, and parallel execution strategy.

The Phoronix benchmark examined multiple mesh sizes, revealing that the best-performing CPU depended on the workload configuration.

Small Mesh
#

In the small-mesh drivaerFastback test, Vera delivered competitive mesh-processing performance and finished ahead of the EPYC 9575F and Xeon 6980P in overall execution time.

The higher-core-count EPYC 9755 remained ahead in total execution time. On a per-core basis, Vera ranked behind the high-frequency EPYC 9575F.

Vera nevertheless delivered the best performance per watt in this particular test.

This result illustrates the difference between per-core performance and aggregate throughput. A CPU with fewer cores can perform well on a particular workload without necessarily beating a processor with more cores across every execution-time measurement.

Medium Mesh
#

With the medium mesh, Vera maintained strong mesh-processing performance and ranked first in the mesh-time measurement.

The Intel Xeon 6980P secured the fastest overall execution time, while Vera followed closely behind.

Vera also performed strongly on a per-core basis and retained a performance-per-watt advantage in the reported test.

The change in the overall winner demonstrates that scaling a numerical application can alter how its work maps to competing processors. The best result depends not only on raw memory bandwidth, but also on how the workload distributes computation, data movement, and synchronization.

Large Mesh
#

The large-mesh OpenFOAM workload was particularly favorable to Vera.

Phoronix reported that the dual-socket Vera system established a substantial lead in overall performance, with LPDDR5-9600 memory playing an important role in sustaining throughput.

Large computational meshes increase the volume of working data and can intensify pressure on the memory subsystem. When memory bandwidth becomes a major bottleneck, Vera’s architecture can make better use of its available cores than a system with more cores but a less favorable bandwidth-to-core ratio.

The result is not evidence that core count is irrelevant. Rather, it shows that additional cores are most useful when the memory system and workload can keep them productively occupied.

📐 LIBXSMM, QuantLib, and QMCPACK: Different Workloads, Different Winners
#

The benchmark suite also included dense and sparse matrix operations, quantitative finance calculations, and quantum Monte Carlo simulations.

These tests help distinguish Vera’s architectural strengths from a generalized claim that it is faster in every HPC workload.

LIBXSMM Matrix Operations
#

LIBXSMM provides highly optimized kernels for matrix multiplication and other numerical operations. Performance depends on matrix dimensions, computational intensity, memory reuse, and the efficiency of vectorized arithmetic.

Vera performed well in several of the tested dense and sparse matrix scenarios, benefiting from its memory bandwidth and per-core execution capability.

However, the EPYC 9575F was faster in the largest GEMM configurations tested.

Large dense matrix multiplication can have a different performance profile from memory-bound stencil computations or irregular numerical algorithms. When data reuse and arithmetic throughput dominate, higher core frequencies and specialized optimization behavior may be decisive.

LIBXSMM therefore demonstrates that Vera’s strengths are workload-dependent rather than universal.

QuantLib
#

QuantLib is a quantitative finance library that includes computationally intensive numerical routines.

In the reported benchmark, Vera outperformed the Xeon 6980P and performed strongly against the EPYC 9755.

More importantly, Vera delivered competitive results with 176 cores despite the 256-core configurations of the EPYC 9755 and Xeon 6980P.

Phoronix also observed that Vera achieved this result at a reported peak frequency of approximately 3.3 GHz, compared with up to 3.9 GHz for the Xeon 6980P and 5.0 GHz for the EPYC 9575F.

The result supports the argument that core efficiency and application characteristics can matter as much as clock frequency and core count.

QMCPACK
#

QMCPACK is a quantum Monte Carlo simulation package used in computational science.

In the tested scenarios, Vera competed closely with the EPYC 9575F for the leading position.

This result further reinforces the importance of evaluating individual application characteristics. Different scientific codes place different demands on floating-point execution, memory access, and parallel scaling, so no single hardware configuration should be expected to dominate every test.

⚡ Core Count, Clock Frequency, and Power Consumption
#

The benchmark challenges a simplistic comparison based on processor specifications alone.

Vera has fewer cores than the EPYC 9755 and Xeon 6980P, but it delivers competitive results in several tests. Its advantages become clearer when considering its memory architecture, per-core performance, and power consumption together.

Core Count and Threading
#

The dual-socket configurations provide the following resources:

Platform Physical cores Hardware threads
NVIDIA Vera 2P 176 352
AMD EPYC 9575F 2P 128 256
AMD EPYC 9755 2P 256 512
Intel Xeon 6980P 2P 256 512

Vera provides 176 physical cores, compared with 256 in each of the two highest-core-count x86 configurations.

NVIDIA Spatial Multithreading allows each Vera core to execute two hardware threads while partitioning certain core resources to improve execution predictability. Hardware-thread counts should not be interpreted as a direct measure of performance equivalence across different CPU architectures.

The benchmark demonstrates that Vera can compensate for its lower core count in selected workloads through a combination of efficient core execution and high memory bandwidth.

Clock Frequencies
#

Phoronix recorded the following peak frequencies during the HPC testing:

Platform Reported peak frequency
NVIDIA Vera 3.2–3.3 GHz
Intel Xeon 6980P Up to 3.9 GHz
AMD EPYC 9755 Up to 4.3 GHz
AMD EPYC 9575F Up to 5.0 GHz

Vera’s competitive performance at lower clock frequencies points to the importance of effective instructions per cycle and memory availability.

However, peak frequency is only one part of the performance equation. Different architectures perform different amounts of useful work per cycle, and benchmark scaling can depend on the distribution of serial and parallel operations.

Consequently, a processor operating at a lower frequency may still complete a specific workload faster if it executes the relevant instructions more efficiently or encounters fewer memory stalls.

CPU Power Consumption
#

Across the benchmark suite, Phoronix reported the following power figures for the dual-socket systems:

Measurement NVIDIA Vera 2P
Average CPU power consumption 842 W
Peak measured CPU power consumption 1,044 W

The average power consumption was comparable to that of the dual EPYC 9755 system under the tested conditions. The Intel Xeon 6980P system recorded the highest power consumption among the three competing platforms in the overall comparison.

Vera also achieved strong performance-per-watt results in workloads such as WRF and several OpenFOAM scenarios.

These figures represent the power measurements reported for the CPU platforms during the tests, not a universal estimate of total server power consumption. Actual data-center energy costs must also account for memory, networking, cooling, power-conversion losses, and workload utilization.

📊 What the Results Tell HPC Operators
#

The benchmark highlights several considerations for organizations choosing CPUs for scientific computing and other high-performance workloads.

Memory Bandwidth Can Outweigh Core Count
#

A high-core-count CPU does not necessarily deliver higher application performance when its memory subsystem cannot supply data quickly enough to keep the cores busy.

Vera’s LPDDR5X subsystem provides substantial bandwidth, helping it perform well in WRF, HPCG, and selected OpenFOAM configurations.

For HPC workloads with large working sets or intensive memory traffic, evaluating memory bandwidth per core can be more useful than comparing total core counts alone.

Performance Must Be Evaluated Per Workload
#

Vera led in some benchmarks but did not win every test.

The EPYC 9755 remained ahead in the small-mesh OpenFOAM total execution-time result. The Xeon 6980P led the medium-mesh OpenFOAM execution-time comparison, while the EPYC 9575F won in the largest LIBXSMM GEMM configurations.

These differences reinforce the need to benchmark actual production applications with representative datasets rather than relying on theoretical throughput or isolated microbenchmarks.

System Configuration Matters
#

The systems in this comparison differed in memory technology, operating-system configuration, and hardware design.

Vera’s performance advantage cannot be attributed exclusively to its CPU architecture. Its high-bandwidth memory subsystem is an integral part of the platform and contributes directly to the observed results.

Similarly, comparisons with future EPYC and Xeon generations will need to account for their updated memory technologies and architectural improvements.

Performance per Watt Is a Critical Metric
#

For large HPC installations, the cost of completing a workload includes both hardware acquisition and energy consumption.

A system that completes a simulation faster while using less power can improve throughput per rack and reduce operating costs. Conversely, a configuration with more cores may be preferable if it delivers better aggregate throughput for a workload that scales efficiently across those cores.

The relevant metric is therefore not simply watts or execution time in isolation, but the energy and total resource cost required to complete useful work.

🔮 What Comes Next for Server CPU Competition?
#

NVIDIA Vera demonstrates that Arm-based server processors can compete directly with established x86 platforms in demanding scientific workloads.

Its performance also highlights the importance of treating the CPU, memory subsystem, and software stack as an integrated platform. As memory capacity and bandwidth increase, CPU architecture alone becomes a less reliable predictor of application performance.

Future comparisons will be particularly interesting as AMD EPYC Venice and next-generation Intel Xeon platforms introduce newer CPU architectures and higher-bandwidth memory configurations.

For a broader view of how these developments fit into the industry, KAD’s data-center CPU architecture analysis examines the wider competitive landscape across Intel, AMD, and NVIDIA.

The next generation of HPC benchmarks will help determine whether Vera’s current advantages persist when competing platforms close the memory-bandwidth gap.

✅ Conclusion
#

Phoronix’s dual-socket CPU comparison shows why NVIDIA Vera can compete with AMD EPYC and Intel Xeon platforms despite having fewer physical cores.

The combination of 176 Olympus cores, Spatial Multithreading, and high-bandwidth LPDDR5X memory enables Vera to deliver leading results in several memory-intensive HPC workloads. WRF, HPCG, and selected OpenFOAM tests highlight the benefits of sustaining high data throughput, while LIBXSMM and other workloads demonstrate that performance remains application-dependent.

Three conclusions stand out:

  • Memory bandwidth matters: Vera’s LPDDR5X subsystem helps sustain core utilization in memory-intensive scientific workloads.
  • Fewer cores can still deliver competitive performance: Effective core execution and system-level balance can compensate for a lower core count.
  • No platform wins every workload: Architecture, memory technology, software configuration, and workload characteristics determine the final ranking.

For HPC operators, the practical takeaway is to evaluate complete server platforms against representative production workloads. Core count, clock frequency, memory bandwidth, performance per watt, and software compatibility should all inform purchasing decisions.

Vera’s results do not establish that Arm is universally faster than x86. They demonstrate something more useful: when CPU architecture and memory bandwidth are balanced around the workload, a system with fewer cores can compete effectively against substantially larger conventional server CPUs.

Related