NVIDIA’s dual-tower architecture is pretty interesting—it splits the generative model into two towers: one for “read-only” and one for “writing,” then uses confidence to mask and write in an intermittent, jumpy way. With the 30B model, it speeds things up by 2.4x while almost not sacrificing quality. It really feels like inference optimization has opened up another new avenue.

NVDA1.86%
View Original
CoinNetwork
CoinWorld news: NVIDIA has introduced a dual-tower architecture (TwoTower), connecting two 30B models in parallel to achieve a 2.4x increase in generation speed without loss. This architecture aims to solve the bottleneck of generation speed in large models. It adopts a dual-tower decoupling design, freezes the autoregressive large model as a "read-only context tower," and independently trains a "denoising writing tower," which reads context information through cross-attention. The writing tower uses a "confidence-based demasking" mechanism, first writing high-confidence words, and gradually filling in the remaining blanks. On a 30B-level hybrid architecture model, this design uses only 1/12 of the data for adaptation, retains 98.7% of the quality, and increases generation speed by 2.42x.
This page may contain third-party content, which is provided for information purposes only (not representations/warranties) and should not be considered as an endorsement of its views by Gate, nor as financial or professional advice. See Disclaimer for details.
  • Reward
  • Comment
  • Repost
  • Share
Comment
Add a comment
Add a comment
No comments
  • Pinned