Google releases 8th-generation TPU chip TPU 8t, targeting large-scale pre-training and dense embedding workloads

AIMPACT message: On April 27 (UTC+8), Google Cloud released its eighth-generation TPU chip, TPU 8t, optimized for large-scale pretraining and embedding-intensive workloads. The chip uses a 3D torus network topology, and a single supercomputing cluster can scale to 9,600 chips. Key innovations include a SparseCore dedicated accelerator that handles sparse memory access patterns for embedding lookups; overlapping and balanced scaling of the VPU/MXU to improve FLOPs utilization; and native FP4 support, which doubles MXU throughput while maintaining large-model accuracy. The Virgo network architecture provides up to 4x data center network bandwidth, using high-radix switches to achieve a flat two-tier non-blocking topology. With JAX and Pathways, a single training cluster can scale to more than 1 million TPU chips; the Virgo network can connect more than 134,000 TPU 8t chips, providing up to 47 Pb/s of non-blocking bidirectional bandwidth. TPU 8i is optimized for real-time inference and post-training, and both systems integrate Arm-based Axion CPUs to eliminate host bottlenecks. These systems are key components of Google Cloud AI supercomputers. (Source: InFoQ)
View Original
This page may contain third-party content, which is provided for information purposes only (not representations/warranties) and should not be considered as an endorsement of its views by Gate, nor as financial or professional advice. See Disclaimer for details.
  • Reward
  • Comment
  • Repost
  • Share
Comment
Add a comment
Add a comment
No comments
  • Pinned