Google is looking to embed the Gemini architecture into chips, achieving up to a 10x improvement in inference energy efficiency

robot
Abstract generation in progress
ME AI news, according to Beating monitoring, Google has been reported to be developing the Frozen v2 AI inference chip. It will directly harden parts of Gemini’s architecture into the chip, reducing computation and data transfer when running models. Internally, it’s expected that the number of tokens it can process per watt will reach 6 to 10 times that of the latest TPU. Google plans to deploy it as early as 2028. Google hopes to use it to ease the shortage of compute power. The gap has already forced Google Cloud to reject some external customer orders. Frozen v2 will coexist in parallel with TPU. TPU can run different models, while Frozen v2 trades flexibility for efficiency. This is also an architectural bet. If Gemini later makes major changes to its underlying design, this batch of chips may not be usable. (Source: BlockBeats)
View Original
This page may contain third-party content, which is provided for information purposes only (not representations/warranties) and should not be considered as an endorsement of its views by Gate, nor as financial or professional advice. See Disclaimer for details.
  • Reward
  • Comment
  • Repost
  • Share
Comment
Add a comment
Add a comment
No comments
  • Pinned