Skip to main content

Cerebras CS-4 AI Rack Delivers 750 PFLOPS with WSE-3 Turbo

·1247 words·6 mins
Cerebras CS-4 WSE-3 Turbo AI Infrastructure AI Accelerators Wafer-Scale Computing Generative AI HPC
Table of Contents

Cerebras CS-4 AI Rack Delivers 750 PFLOPS with WSE-3 Turbo

Cerebras has introduced the CS-4, a new rack-scale AI computing platform powered by its latest WSE-3 Turbo (WSE-3T) wafer-scale processor.

The company claims a dramatic performance advantage over conventional GPU-based systems for certain AI generation workloads. In one comparison, a task requiring approximately 30 seconds on a GPU rack can reportedly be completed by the CS-4 in roughly 1 second.

The system combines three WSE-3T processors, delivering up to 750 PFLOPS of AI compute per rack, while substantially increasing memory bandwidth and reducing system-level latency.

Initial CS-4 shipments are expected to begin this quarter, positioning Cerebras for a more direct challenge to large-scale GPU infrastructure used for generative AI.

πŸš€ WSE-3 Turbo Preserves Cerebras’ Wafer-Scale Advantage
#

At the heart of the CS-4 is the WSE-3 Turbo, an upgraded version of Cerebras’ existing WSE-3 wafer-scale processor.

The processor retains Cerebras’ defining approach: instead of dividing AI compute across thousands of conventional accelerator chips, the company places an enormous number of compute resources onto a single wafer-scale device.

Reported WSE-3T specifications include:

Specification WSE-3 Turbo
Transistors 4 trillion
AI cores 900,000
Silicon area 46,225 mmΒ²
On-chip SRAM 44 GB
AI compute 125 PFLOPS
Sparse AI compute Up to 250 PFLOPS
Memory bandwidth 43.2 PB/s
Reported latency 2 ms

These figures preserve the fundamental architectural advantage of wafer-scale computing: extremely high compute density combined with massive local memory bandwidth.

Memory bandwidth remains a defining advantage
#

The WSE-3T’s reported 43.2 PB/s of memory bandwidth is particularly important for AI workloads.

Large language models frequently encounter memory bandwidth and communication bottlenecks rather than purely arithmetic limitations.

By placing 44 GB of SRAM directly on the processor and providing enormous internal bandwidth, Cerebras can reduce the need to repeatedly move data between separate accelerator chips and external memory subsystems.

The company also reports that system latency has improved from approximately 5 ms to 2 ms, further increasing the responsiveness of inference workloads.

⚑ CS-4 Targets GPU Rack-Level AI Performance
#

The CS-4 is built around Cerebras’ new Nexus platform architecture.

Rather than treating the rack as a collection of independent accelerator nodes, the platform uses a modular three-layer design covering:

  • Compute
  • Power
  • I/O

This architecture is intended to reduce bottlenecks between the wafer-scale processors, power infrastructure, and external networking components.

Backpack power architecture moves conversion closer to the wafer
#

One of the more unusual aspects of CS-4 is its pluggable backpack design.

The architecture places power-conversion modules close to the wafer, reducing the distance electrical power needs to travel before reaching the processor.

Cerebras claims this approach can substantially reduce power losses and provide approximately twice as much power to the chip compared with its previous design.

For wafer-scale processors operating at extremely high power levels, power delivery becomes a fundamental system-design problem rather than a peripheral engineering detail.

πŸ“Š CS-4 Delivers 750 PFLOPS Per Rack
#

Each CS-4 rack contains three WSE-3T processors.

With each processor delivering up to 250 PFLOPS under sparse workloads, the system can reach a reported 750 PFLOPS of total AI compute.

The company also claims approximately a 10x improvement in throughput per watt compared with the previous-generation CS-3.

This is particularly important as AI infrastructure operators increasingly evaluate systems using performance-per-watt rather than peak compute alone.

A higher-performance accelerator that requires proportionally less power can reduce:

  • Data-center electricity consumption
  • Cooling requirements
  • Rack-level power constraints
  • Operating expenditure
  • Infrastructure density limitations

However, actual efficiency will depend heavily on workload characteristics, model architecture, utilization, and software optimization.

🧠 GPT-OSS 120B Test Shows Extreme Inference Throughput
#

Cerebras reports that a single CS-4 rack can achieve more than 4,400 tokens per second per user when running the GPT-OSS 120B model.

The company also presents a more dramatic comparison.

For an identical generation workload, Cerebras claims a CS-4 can complete the task in approximately 1 second, while a conventional GPU rack requires around 30 seconds.

If reproduced under equivalent workload and system conditions, that would represent a substantial difference in inference latency and throughput.

The comparison needs workload context
#

Raw tokens-per-second figures should not be interpreted as universal performance measurements.

AI inference performance depends on several variables, including:

  • Batch size
  • Context length
  • Quantization
  • Sequence length
  • Model architecture
  • Memory utilization
  • Number of concurrent users
  • Sampling configuration
  • Software stack
  • Interconnect topology

The most meaningful comparison is therefore between systems running the same model, workload, precision, and serving configuration.

Nevertheless, the reported result demonstrates the type of workload Cerebras’ wafer-scale architecture is designed to target: large-model inference where communication and memory movement can dominate execution time.

🌐 CS-4 Can Scale Beyond a Single Rack
#

Cerebras is not positioning the CS-4 solely as a standalone inference appliance.

The platform can be connected into larger clusters through its interconnect architecture, allowing multiple wafer-scale systems to work together on extremely large models.

Cerebras claims the architecture can support frontier models exceeding 50 trillion parameters.

This is important because wafer-scale computing does not eliminate the need for distributed infrastructure.

Instead, the architecture attempts to minimize communication overhead within each accelerator while providing mechanisms to scale across multiple systems when the model exceeds the resources of a single wafer.

🀝 Cerebras and AMD Target the Broader AI Infrastructure Market
#

Cerebras’ primary competitive target remains large-scale GPU infrastructure, particularly systems built around NVIDIA accelerators.

However, the company is also expanding its ecosystem through a partnership with AMD.

The two companies plan to combine Cerebras’ wafer-scale rack systems with AMD’s Helios AI rack architecture.

This could create a complementary infrastructure model in which different accelerator technologies are optimized for different stages or characteristics of AI workloads.

Rather than replacing GPUs universally, Cerebras can potentially position wafer-scale processors as specialized infrastructure for workloads where extremely high bandwidth and low communication overhead provide a meaningful advantage.

πŸ”§ Wafer-Scale Computing Takes a Different Approach to AI Scaling
#

The CS-4 highlights a fundamental architectural difference between Cerebras and conventional GPU platforms.

Traditional AI infrastructure scales by connecting large numbers of individual accelerator chips through high-speed interconnects.

Cerebras instead attempts to maximize the amount of computation and memory bandwidth available within a single enormous processor.

This approach can reduce the frequency of communication between separate accelerator devices and potentially improve efficiency for tightly coupled workloads.

The trade-off is that wafer-scale processors require highly specialized manufacturing, packaging, power delivery, cooling, and software infrastructure.

Cerebras therefore faces a different engineering challenge from conventional GPU vendors: the company must make the entire system architecture work together as a tightly integrated computing platform.

🎯 CS-4 Raises the Stakes in AI Accelerator Competition
#

The launch of CS-4 demonstrates that the AI accelerator market is continuing to diversify beyond conventional GPU architectures.

With the WSE-3T, Cerebras is pushing wafer-scale computing to a new performance level:

  • 4 trillion transistors
  • 900,000 AI cores
  • 44 GB on-chip SRAM
  • 43.2 PB/s memory bandwidth
  • 125 PFLOPS per processor
  • Up to 750 PFLOPS per CS-4 rack
  • More than 4,400 tokens/s per user on GPT-OSS 120B

The reported 30-second-versus-1-second generation comparison is particularly attention-grabbing, although independent testing will be necessary to determine how broadly that advantage applies across real-world workloads.

With initial shipments expected this quarter and a growing partnership with AMD, Cerebras is clearly positioning the CS-4 as more than a specialized research platform.

Its larger ambition is to establish wafer-scale computing as a serious alternative for large-scale AI inference and training infrastructureβ€”and to challenge the assumption that adding more conventional GPUs is always the most efficient path to scaling AI.

Related

Why CXL Is Losing the AI Accelerator Interconnect Race
·1341 words·7 mins
CXL AI Accelerators PCIe NVLink SerDes HBM AI Infrastructure Interconnects
Do We Still Need GPUs? How AI-Accelerated CPUs Could Change HPC
·1840 words·9 mins
AI GPUs CPUs HPC AI Accelerators HBM LLM Supercomputing
Samsung PM1763 PCIe Gen6 SSD Enters Enterprise Production
·711 words·4 mins
Samsung PM1763 PCIe Gen6 Enterprise SSD NVMe AI Infrastructure HPC Data Center