LongCat-2.0 has been open-sourced. Its 30T tokens pre-training and the full disclosure of training stability solutions at the ten-thousand-card scale are now all public—another step forward in the infrastructure for homegrown large models.

View Original
CoinNetwork
CoinWorld news, today, Meituan officially released the new generation trillion-parameter large model LongCat-2.0, which will be open-sourced. LongCat-2.0's pre-training data scale exceeds 30t tokens, covering Chinese, English, multilingual, code, and other types of data. Facing hardware failures, communication anomalies, memory pressure, and numerical fluctuations in ten-thousand-card-level training, the LongCat team has overcome the challenges of domestic computing power training from three aspects: stability, correctness, and efficiency. In terms of stability, through hccl exception handling, elastic scaling of cards, and automatic fault recovery, the monthly average daily failure rate is reduced by more than 70%. In terms of correctness, through self-developed deterministic operators, bitwise consistency verification, and parameter detection, the reliability of training results is ensured, while the calculation accuracy of key modules is improved and the reduce logic is optimized based on practice.
This page may contain third-party content, which is provided for information purposes only (not representations/warranties) and should not be considered as an endorsement of its views by Gate, nor as financial or professional advice. See Disclaimer for details.
  • Reward
  • Comment
  • Repost
  • Share
Comment
Add a comment
Add a comment
No comments
  • Pinned