This open-source model scene really has a bit of the vibe of a computer-city assembled machine.



Baseten didn’t re-train a complete multimodal model; instead, it took the MoonViT vision encoder from Kimi K2.6 and plugged it into GLM-5.2, which originally only processed text. The 744B core of GLM wasn’t touched, the entire vision tower was frozen, and what was truly trained was only the PatchMerger with about 49.5 million parameters in the middle.

The result reached about 55% on MMMU-Pro, and the official said it’s close to Claude 4.5 Haiku. In other words, to add vision, speech, or other capabilities to an open-source model in the future, you may really not need to burn training costs from scratch every time—just pick existing parts and hook them up.

But don’t just look at “49.5 million parameters is cheap.” These weights are about 466GB. Running full 1 million-context requires 49.5M200 cards, and even at 256K it still needs 4. The R&D barrier is indeed down, but the deployment bills are far from friendly.

So I think this may not be all good news for model companies. As capabilities become easier to piece together, the ones that can keep charging rent for real may still be $NVDA, inference platforms, and cloud hosting providers. In the future, people may not only download models—you might also need to learn how to assemble models first 🤣
#Ourbit 不只 Crypto,全球熱門資產一站交易。
NVDA-1.58%
View Original
This page may contain third-party content, which is provided for information purposes only (not representations/warranties) and should not be considered as an endorsement of its views by Gate, nor as financial or professional advice. See Disclaimer for details.
  • Reward
  • Comment
  • Repost
  • Share
Comment
Add a comment
Add a comment
No comments
  • Pinned