↓ Skip to main content

Arista SU-144 Challenges NVLink With Open Ethernet Scale-Up

Arista SU-144 Challenges NVLink With Open Ethernet Scale-Up

AI clusters increasingly depend on two distinct networking layers: Scale-Up and Scale-Out. Scale-Up links accelerators inside a tightly coupled compute domain, where bandwidth, latency, flow control, and collective communication efficiency are critical. Scale-Out connects larger numbers of compute nodes or separate compute domains.

Arista is now pushing Ethernet deeper into the Scale-Up layer with its Etherlink SU-144 architecture, based on Ethernet for Scale-Up Networking (ESUN). According to Arista’s published design specifications, the architecture can connect up to 144 accelerators in a single-hop topology and scale to as many as 1,024 accelerators through cross-rack expansion.

The significance goes beyond a single product architecture. SU-144 highlights a broader industry debate over whether AI Scale-Up networking should remain dominated by tightly integrated proprietary interconnects such as NVIDIA NVLink or evolve toward open Ethernet-based standards capable of supporting multiple accelerator vendors.

🧩 Scale-Up vs. Scale-Out: Two Different Networking Domains
#

The distinction between Scale-Up and Scale-Out is essential when evaluating AI cluster architectures.

Scale-Up focuses on communication between accelerators inside a tightly coupled compute domain. It prioritizes extremely high bandwidth, low latency, efficient collective operations, and tightly controlled traffic behavior. Depending on the architecture, Scale-Up can extend beyond a single rack and use switches to connect accelerators across multiple racks.

Scale-Out, by contrast, connects multiple compute nodes or larger compute domains. Traditional deployments commonly use InfiniBand or Ethernet, with RoCEv2 widely adopted for Ethernet-based AI networking.

The major technologies competing in or around the Scale-Up space include:

  • NVIDIA NVLink/NVSwitch
  • AMD Infinity Fabric
  • Google ICI
  • UALink
  • ESUN
  • Emerging Ethernet-based approaches associated with the Ultra Ethernet Consortium (UEC)

The Scale-Out ecosystem has its own stack. UEC is developing technologies such as Ultra Ethernet Transport (UET), while NVIDIA’s Spectrum-X provides an Ethernet platform specifically optimized for AI networking.

The boundaries are therefore architectural rather than simply physical. A Scale-Up fabric can span racks, while Scale-Out can operate across a much larger cluster.

⚙️ ESUN Simplifies the Network Stack for Scale-Up
#

One of the more notable ESUN design choices is its use of Layer 2 forwarding combined with a compact header.

Removing the IP header reduces protocol overhead, but information normally carried by that header cannot simply disappear. Congestion notification, class-of-service information, load-balancing information, and hop limits still need to be represented.

ESUN therefore introduces a 4-byte ESUN Header containing fields for these functions:

  • EH-ECN for explicit congestion notification
  • EH-CoS for class-of-service handling
  • A 16-bit Flow Label for load balancing
  • A 4-bit TTL field to prevent forwarding loops

The 4-byte figure refers specifically to the ESUN header itself rather than the total header overhead of the complete packet.

ESUN also incorporates mechanisms defined by the UEC, including Credit-Based Flow Control (CBFC) and Link-Level Retry (LLR).

CBFC regulates transmission according to available receiver-buffer credits, helping prevent buffer exhaustion. LLR, meanwhile, allows transmission errors to be recovered locally at the link level rather than requiring every error to trigger an end-to-end recovery operation.

These mechanisms address important problems in high-performance fabrics, but they should not automatically be interpreted as a guarantee of low tail latency under heavy congestion. Actual latency behavior depends on the complete endpoint, switching, buffering, congestion-control, and software implementation.

ESUN 1.0 should therefore be viewed as a foundational networking specification rather than proof that every implementation will behave identically. Product-level support, endpoint transport implementations, and cross-vendor interoperability still need to be established through documentation and real-world testing.

🚀 NVLink’s Advantage Comes From Full-System Co-Design #

The strongest counterpoint to Ethernet Scale-Up remains NVIDIA’s NVLink ecosystem.

According to NVIDIA’s July 2026 technical disclosures, sixth-generation NVLink, introduced with the Vera Rubin platform, delivers up to 3.6 TB/s of bidirectional bandwidth per GPU. A Vera Rubin NVL72 configuration reaches approximately 260 TB/s of aggregate rack-level bandwidth, while its integrated networking capabilities provide up to 130 TFLOPS of rack-level FP8 in-network compute for operations such as All-Reduce and Broadcast.

These numbers illustrate why comparing NVLink with a conventional Ethernet switch purely on port bandwidth can be misleading.

NVLink is deeply co-designed with the accelerator architecture, switching silicon, collective-communication mechanisms, and software stack. The resulting performance comes from coordination across the entire platform rather than from the network device alone.

That integration represents one of NVIDIA’s major advantages in tightly coupled training systems.

At the same time, the continued emergence of non-NVIDIA accelerators creates an opening for Ethernet-based Scale-Up solutions.

Interestingly, NVIDIA itself is among the founding members of ESUN. The OCP ESUN initiative includes 12 founding members, including NVIDIA, AMD, Arista, Broadcom, Meta, and Microsoft.

This makes the market more complicated than a simple open-versus-proprietary contest.

NVIDIA continues to develop its proprietary NVLink platform while participating in open standards. Its NVLink Fusion strategy also provides semi-custom ASICs and CPUs with a path into NVIDIA rack-scale AI infrastructure, with companies such as d-Matrix and MediaTek announcing adoption.

In other words, major industry players are increasingly hedging across multiple interconnect strategies rather than committing exclusively to one ecosystem.

🏗️ SU-144 Targets Rack-Scale and Cross-Rack Scale-Up
#

Arista’s SU-144 architecture is designed around three physical configurations.

The first is an orthogonal chassis architecture capable of directly connecting up to 144 XPUs, with liquid-cooling support for systems ranging from approximately 100 kW to 400 kW.

The second is a cabled backplane architecture, while the third supports cross-rack expansion, allowing the architecture to scale toward 1,024 accelerators.

At the software level, SU-144 supports Arista’s EOS as well as open-source SONiC, reinforcing its positioning around a software-defined and relatively open networking model.

Arista also cites 1.6 Pb/s and 6.5 Pb/s figures in its optical-connectivity roadmap. These figures should not be interpreted as the bandwidth of a currently shipping SU-144 system. The approximately 1.6 Pb/s figure corresponds to current OSFP-based designs, while future XPO and CPO/NPO approaches are intended to enable substantially higher optical connectivity, potentially reaching approximately 6.5 Pb/s.

There are two important limitations to the current positioning.

First, SU-144 is not a complete full-rack appliance SKU. Arista has stated that its price lists do not include complete rack SKUs. Instead, Arista supplies the networking equipment and software deployed inside liquid-cooled racks, while system integration partners provide the complete rack solution.

Second, independent evaluation remains limited. Arista has not yet published sufficient large-scale production deployment data or comprehensive, directly comparable workload benchmarks to establish how SU-144 performs against competing Scale-Up architectures under identical conditions.

Consequently, questions around sustained performance, reliability, congestion behavior, operational complexity, and total cost of ownership remain open.

🌐 Custom Accelerators Create the Biggest Opening for Ethernet
#

The strongest argument for Ethernet Scale-Up is not necessarily that it will immediately replace NVLink in NVIDIA’s largest training systems.

Instead, the opportunity comes from accelerator diversification.

Cloud providers and hyperscalers are increasingly developing custom accelerators, while AMD and other vendors continue expanding their presence in AI compute. A networking architecture that can accommodate multiple accelerator families can reduce dependence on a single vertically integrated ecosystem.

Ethernet has several structural advantages in this environment. It benefits from a large established networking supply chain, broad switch and optical-component availability, and a substantial pool of networking software and engineering expertise.

However, those advantages do not automatically make Ethernet interoperable with every accelerator. Endpoint implementations, drivers, collective-communication libraries, congestion-control mechanisms, and accelerator software stacks all have to support the underlying fabric.

The strategies of major hyperscalers illustrate this point.

Meta is a co-lead of the ESUN specification while also deploying AMD accelerators for AI workloads. Its participation in ESUN and UALink provides a path toward more vendor-neutral infrastructure for heterogeneous compute environments.

Microsoft is another ESUN co-lead and is simultaneously developing its own Maia accelerator family. Ethernet-based Scale-Up standards therefore align with its broader interest in building infrastructure that is not exclusively dependent on a single accelerator vendor.

Google takes another approach, developing its TPU-specific ICI interconnect alongside its broader data-center networking architecture. Its network is divided into communication domains, with ICI handling tightly coupled accelerator communication, Virgo serving as a Scale-Out fabric between Pods, and Jupiter supporting north-south traffic.

This illustrates an important principle: hyperscale networks are increasingly designed around communication domains and workload requirements, rather than being divided simply according to physical rack boundaries.

🔄 Open and Proprietary Fabrics Are Likely to Coexist
#

The competitive landscape therefore does not fit neatly into an “Ethernet versus NVLink” narrative.

NVIDIA’s NVLink remains difficult to displace in systems where the GPU, switch fabric, collective communication, and software stack are designed together. For large NVIDIA training clusters, this degree of co-design provides a substantial performance and optimization advantage.

Ethernet Scale-Up is approaching the problem from a different direction.

Its primary value proposition is ecosystem breadth: standard networking components, multiple switch vendors, open software options, and the possibility of supporting heterogeneous accelerator environments.

ESUN, UALink, and UEC are consequently addressing different but overlapping parts of the broader AI networking stack. Their practical value will depend not only on specifications but also on endpoint adoption, software support, interoperability testing, deployment scale, and economics.

The key question is therefore not whether open Ethernet will simply “beat” proprietary interconnects.

It is whether the AI industry will continue moving toward a heterogeneous compute model in which the flexibility of an open Scale-Up fabric becomes more valuable than the peak performance of a vertically integrated one.

📊 Conclusion: Scale-Up Networking Is Moving Toward Layered Coexistence
#

Today, NVIDIA-based AI clusters generally rely on NVLink/NVSwitch for Scale-Up and InfiniBand or Ethernet for Scale-Out. That architecture remains highly effective because NVIDIA controls much of the hardware and software stack end to end.

But the rise of custom accelerators and alternative platforms is changing the economics of the Scale-Up market.

Arista’s SU-144 demonstrates how Ethernet vendors are attempting to move beyond traditional Scale-Out networking and directly address rack-scale accelerator communication. ESUN provides the standards framework, while technologies such as CBFC and LLR attempt to provide the flow-control and reliability characteristics required by tightly coupled AI workloads.

In the short term, large NVIDIA training systems are likely to retain NVLink because its system-level co-design advantages cannot be reproduced simply by replacing the network switch.

Over the medium term, however, growing deployments of AMD and custom accelerators should increase demand for open, multi-vendor Scale-Up fabrics.

The likely outcome is not the elimination of proprietary interconnects. Instead, AI infrastructure is heading toward a layered coexistence model, where proprietary fabrics remain dominant in tightly optimized accelerator platforms while open Ethernet-based technologies expand wherever heterogeneous compute, vendor flexibility, and infrastructure reuse become higher priorities.

Related