Google’s 10th-Gen TPU May Add AMD CPU Cores for AI Workloads
Google is reportedly working with AMD on its 10th-generation Tensor Processing Unit (TPU), potentially marking a significant change in how the company’s custom AI accelerators are designed.
The next-generation TPU is expected to combine specialized accelerator hardware with general-purpose CPU cores inside the same package, targeting workloads that require substantially more CPU-side processing than conventional AI training.
The reported collaboration would also represent AMD’s first deep involvement in the development of Google’s custom AI ASICs, breaking from the design model used across the previous nine TPU generations.
🤝 Google Reportedly Brings AMD Into TPU Development #
According to reports, Google is collaborating with AMD on its 10th-generation TPU. The partnership would represent a major shift for Google’s custom accelerator program, which has historically relied on Broadcom for TPU chip design across its first nine generations.
Google has developed extensive internal expertise in AI accelerator architecture, making the reported involvement of AMD particularly notable. Rather than simply supplying conventional components, AMD is reportedly contributing more deeply to the underlying architecture and implementation of the next TPU generation.
The move could reflect Google’s changing hardware requirements as AI workloads increasingly extend beyond conventional matrix-heavy model training.
🧠 In-Package CPU Cores Target Reinforcement Learning #
One of the most significant reported changes is the addition of general-purpose CPU cores directly within the TPU package.
Traditional LLM training can place the majority of computational demand on specialized accelerators. However, reinforcement learning, inference pipelines, and emerging agentic AI workloads can introduce substantially greater CPU-side requirements.
These workloads frequently involve orchestration, environment interaction, control logic, data processing, and other operations that are less efficiently handled by dedicated matrix-processing hardware.
Integrating CPU resources alongside the TPU could therefore reduce data movement between separate processors while allowing CPU-intensive operations to execute closer to the accelerator.
Google’s Existing TPU Designs Already Use CPU Resources #
Google has already explored tightly coupled CPU and TPU configurations.
Its 8i inference TPU reportedly pairs one internally developed Axion CPU with every two TPUs. Earlier, the 7th-generation training TPU paired one Intel Xeon Emerald Rapids processor with every four TPUs.
The reported 10th-generation design could take this concept further by bringing the CPU directly into the accelerator package rather than relying on discrete host processors.
Such integration could provide tighter communication between general-purpose compute and AI acceleration while potentially improving system-level efficiency for increasingly heterogeneous workloads.
⚙️ AMD’s Packaging and CPU IP Strengthen the Case #
AMD has several technologies and product architectures that align closely with the reported requirements.
Its advanced packaging capabilities, 3D SoIC technology, and extensive x86 CPU IP portfolio provide the foundation for combining general-purpose processors with specialized accelerators in tightly integrated packages.
AMD has already demonstrated this architectural approach with products such as the Instinct MI300A, which combines Zen CPU cores and CDNA GPU compute resources within a unified accelerator package.
The experience gained from these hybrid CPU-accelerator architectures could make AMD a natural technology partner for a TPU design that requires both high-throughput AI compute and substantial general-purpose processing.
🚀 TPU Architecture Could Shift Toward Agentic AI #
The reported design direction highlights a broader change in AI hardware requirements.
As AI systems evolve from conventional model training toward inference-heavy, reinforcement-learning, and agentic workloads, accelerator performance alone is increasingly insufficient. The surrounding CPU, memory, interconnect, and orchestration infrastructure can become critical determinants of overall system efficiency.
A TPU with integrated CPU resources could allow Google to optimize the complete compute stack around these workloads instead of treating the accelerator and host CPU as separate components.
If the reports are accurate, Google’s 10th-generation TPU could therefore represent more than another accelerator-generation upgrade. It may signal a broader architectural transition toward heterogeneous AI compute, where specialized accelerators and general-purpose CPU resources are designed together around the requirements of next-generation AI systems.
Specific architectural details and the exact scope of AMD’s involvement remain subject to official confirmation.