MoE architecture for video generation, with 3 billion parameters but only activating 300 million, this wave of efficiency optimization indeed hits the pain points of embodied intelligence.

View Original
CoinNetwork
Coin World News, Ant LingBot announced the open-source release of LingBot-Video, the world's first video generation foundation model for embodied intelligence. The model is based on a mixture-of-experts (MoE) architecture, designed to improve video generation efficiency. LingBot-Video has a total of 3 billion parameters, with only about 300 million parameters activated during generation. Compared to traditional dense architectures of similar parameter scale, inference efficiency is improved by approximately 3 times. This design enables the model to achieve visual expressiveness from large-scale parameters while being more suitable for the efficient inference requirements of embodied intelligence.
This page may contain third-party content, which is provided for information purposes only (not representations/warranties) and should not be considered as an endorsement of its views by Gate, nor as financial or professional advice. See Disclaimer for details.
  • Reward
  • Comment
  • Repost
  • Share
Comment
Add a comment
Add a comment
No comments
  • Pinned