Skip to main content

Alibaba Cloud Reveals AI Infrastructure Trends for the Agentic AI Era

·1103 words·6 mins
Alibaba Cloud Agentic-Ai AI Infrastructure CXL UALink Optical Interconnect Memory Architecture Open Compute
Table of Contents

Alibaba Cloud Reveals AI Infrastructure Trends for the Agentic AI Era

At the 2026 Open Compute Technology Conference in Beijing, Alibaba Cloud presented its vision for the next generation of AI infrastructure, emphasizing that the rapid emergence of Agentic AI is fundamentally reshaping system architecture rather than merely increasing compute demand.

Under the conference theme “Boundless Intelligent Computing: Open, Diverse, Scalable,” discussions focused on GW-scale AI computing centers, AI-native infrastructure, high-speed networking, and compute-energy optimization. During multiple technical sessions, Alibaba Cloud highlighted four major infrastructure trends that will define future AI platforms: collaborative heterogeneous computing, memory-centric architectures, optical Scale-Up interconnects, and AI-enabled firmware.

Collectively, these trends illustrate a shift from GPU-centric acceleration toward tightly integrated systems where CPUs, GPUs, memory, storage, networking, and firmware cooperate as a unified computing platform.

πŸ–₯️ Trend 1: Agentic AI Is Driving a New Computing Paradigm
#

According to Alibaba Cloud, AI infrastructure has already experienced two major architectural transitions:

  • Traditional machine learning relied primarily on CPU-centric computing.
  • Deep learning shifted workloads toward GPU-centric acceleration.
  • Agentic AI introduces a third paradigm built around CPU-GPU collaboration.

As AI agents evolve beyond simple prompt-response interactions into autonomous workflows driven by Loop Engineering, the CPU assumes a much larger role in orchestrating task execution, scheduling tools, handling networking, and coordinating memory.

At the same time, technologies such as vCPU oversubscription significantly increase CPU utilization, while growing volumes of LLM inference requests place additional pressure on GPU clusters.

Chen Jian, Senior Director of Server R&D at Alibaba Cloud, described modern AI infrastructure as a collaborative ecosystem rather than isolated GPU workloads. Future performance depends on deep coordination between:

  • CPUs
  • GPUs
  • Multi-tier memory
  • High-speed interconnects
  • System software

This holistic approach reflects the increasing complexity of large-scale Agentic AI deployments.

🧠 Trend 2: Memory Architecture Becomes the Primary Bottleneck
#

Alibaba Cloud identified the rapidly expanding inference context of AI agents as one of the industry’s most significant infrastructure challenges.

Unlike conventional chatbot interactions, Agentic AI frequently performs:

  • Multi-turn conversations
  • Tool invocations
  • External knowledge retrieval
  • Workflow execution

Each operation expands the inference context, often increasing context size by more than three times.

Meanwhile, KV Cache capacity requirements continue to grow dramatically. As context windows expand from 64K tokens toward 10 million tokens, memory demands increase from hundreds of gigabytes to multiple terabytes.

The resulting bottleneck is no longer GPU compute aloneβ€”it is the ability of compute, memory, and storage resources to operate as a unified architecture.

Panjiu UMX Memory-Compute Architecture
#

To address this challenge, Alibaba Cloud introduced Panjiu UMX (Unified Memory/Storage Extension).

UMX is designed to extend memory capacity across heterogeneous systems using:

  • UALink for GPU memory expansion
  • CXL for CPU memory expansion

Unlike traditional storage access mechanisms, UMX preserves native Load/Store memory semantics, allowing compute nodes to directly access heterogeneous storage resources including:

  • DRAM
  • Storage-Class Memory (SCM)
  • SSDs

These devices are connected through the ALink Scale-Up Fabric, providing nanosecond-scale latency while enabling a tiered memory hierarchy that scales from on-chip registers to hundreds of terabytes.

The objective is to eliminate the memory wall limiting long-context AI inference.

In-House Hardware and Software Integration
#

Alibaba Cloud also emphasized that UMX is the result of full-stack internal development spanning:

  • Interconnect silicon
  • Memory technologies
  • Switching hardware
  • System software

Its self-developed CXL Switch delivers:

  • 32 GT/s transmission bandwidth
  • Sub-100 ns switching latency

According to Alibaba Cloud, this architecture reduces large language model first-token latency by 89.6%.

The underlying research on three-tier memory disaggregation also received the SIGMOD 2025 Industrial Track Best Paper Award, highlighting its technical significance.

🌐 Trend 3: Scale-Up Networks Are Transitioning Toward Optical Interconnects
#

Another major architectural transition involves GPU interconnect technology.

As large AI models continue scaling, communication bandwidth between accelerators becomes just as important as raw compute performance.

Alibaba Cloud noted that increasing GPU compute capacity effectively requires proportional increases in Scale-Up interconnect bandwidth.

However, as electrical signaling advances toward 224G and eventually 448G, traditional copper interconnects face increasingly severe distance limitations.

This makes optical interconnect technology a practical necessity.

NPO Instead of Immediate CPO Adoption
#

Rather than adopting Co-Packaged Optics (CPO) immediately, Alibaba Cloud currently favors Near-Packaged Optics (NPO).

Compared with CPO, the NPO approach offers several practical advantages:

  • Better fault isolation
  • Easier component replacement
  • Greater supply-chain flexibility
  • Decoupled optical and switching design

Alibaba Cloud views optical and copper technologies as complementary rather than competing solutions.

Open Standards Drive Ecosystem Growth
#

Alibaba Cloud also stressed that open standards are becoming increasingly important for Scale-Up networking.

The company participates in the development of:

  • UALink
  • OIF NPO specifications
  • Industry-wide interoperability standards

Building on these efforts, Alibaba Cloud introduced SNPO, its next-generation optical module architecture designed specifically for Scale-Up AI servers.

Key capabilities include:

  • NPO and NPC electro-optical compatibility
  • Copper-optical hybrid deployment
  • Single-mode and multi-mode fiber support
  • 3.2T and 6.4T module configurations

According to Alibaba Cloud, SNPO is the industry’s first commercially deployed solution in 2026 supporting both NPO and NPC compatibility while targeting ultra-dense super-node server architectures.

πŸ€– Trend 4: AI Intelligence Extends Into Server Firmware
#

Alibaba Cloud’s final trend moves AI intelligence below the operating system and into server firmware.

Managing ultra-large AI clusters requires increasingly sophisticated operational tooling, making conventional command-line management less efficient.

To simplify infrastructure management, Alibaba Cloud introduced an AI-enabled BMC (Baseboard Management Controller) platform for its Panjiu servers.

Natural Language Infrastructure Management
#

The new firmware architecture integrates:

  • D-Bus communication
  • Alibaba Cloud’s self-developed BMC-CLI
  • OpenBMC

Together, these components allow administrators to manage servers using natural language instead of memorizing command syntax.

The system supports both:

  • Cloud-hosted language models
  • Local AI model deployment

By embedding AI directly into firmware, Alibaba Cloud extends intelligent operations and maintenance from cloud management platforms down to the hardware layer itself.

πŸ† Building Native Infrastructure for Agentic AI
#

Alibaba Cloud concluded that future AI infrastructure will no longer revolve around GPU performance alone.

Instead, competitive advantage will increasingly depend on deep integration across:

  • CPUs
  • GPUs
  • Multi-tier memory
  • High-speed Scale-Up networking
  • Storage
  • Firmware intelligence

The company has already invested heavily in open ecosystem development as a board member of both the CXL Consortium and UALink Consortium.

Its recent initiatives include:

  • The industry’s first CXL memory-pooled super-node server
  • The UALink-compatible ALink ecosystem
  • The Panjiu AL128 super-node server
  • Full-stack hardware-software co-design spanning chips, interconnects, servers, storage, and firmware

As Agentic AI continues driving exponential growth in model complexity and inference workloads, Alibaba Cloud’s roadmap reflects a broader industry transition toward tightly integrated, collaborative computing platforms. Rather than treating CPUs, GPUs, memory, networking, and firmware as independent components, the next generation of AI infrastructure will increasingly optimize them as a unified system capable of supporting autonomous, long-context AI applications at hyperscale.

Related

Beluga: CXL-Based KV Cache Architecture Cuts TTFT by 89.6%
·690 words·4 mins
CXL Memory Architecture Alibaba Cloud LLM GPU
Samsung CXL Memory Modules and HBM3E Drive AI Scalability
·733 words·4 mins
Samsung CXL HBM3e Memory Architecture AI Infrastructure Data Center DRAM Composable Infrastructure
Arm CEO: AI CPU Demand Is 'Off the Charts' as Agentic AI Reshapes Data Centers
·1196 words·6 mins
ARM AI Infrastructure CPU Agentic-Ai Data Centers Semiconductors Cloud Computing Neoverse