The Dark Side of the Moon has just open-sourced Kimi K3, and it has also announced that K4—with a larger number of parameters—is already in the queue, as it seeks more computing power.

Moonshot AI just open-sourced Kimi K3, and the Beijing startup is already looking to secure more Nvidia Blackwell chips to train a larger-scale K4.
(Background recap: Jammed! Kimi K3 hit its usage cap within 48 hours of launch, and Moonshot AI temporarily paused new subscriptions.)
(Additional context: Nearly 200 Silicon Valley companies submitted a petition to the White House urging it to block China’s open-source models—the one that ends up dying is the US itself.)

Moonshot AI on the evening of July 27 just open-sourced Kimi K3 parameters totaling 2.8 trillion, and simultaneously released model weights, a technical report, and key infrastructure. The company said it is already planning the next-generation K4, hinting its scale will be “clearly larger” than the newly released K3, and said it currently urgently needs more computing power.

If one company can’t cover it, then get two

At present, no single Chinese cloud provider can independently deliver enough Blackwell computing power. Moonshot AI’s approach is to combine the computing power of at least two Chinese cloud providers for use, and have engineers specifically optimize the network architecture across data centers—so that training runs spread across different server rooms, and even different cities, can stay up and not get the entire training pipeline derailed by latency.

This patchwork approach shows: when a single supplier isn’t enough, export controls end up forcing out a more complex set of detour engineering, and it turns “who is helping whom” into a fog.

The inference side is also tight. On H20 chips that can be legally obtained under the restrictions, customers need to chain at least 64 server chips for Kimi K3 to run. Simply connecting 64 chips together and making them cooperate is already a substantial hardware and network cost—this is also why Moonshot AI is willing to risk regulatory violations to assemble Blackwell, not only saving on electricity bills and data center space, but also making sure it can meet this wave of demand and retain the large influx of users.

Training locations—bypassing sanctions?

A week ago, Michael Kratsios, director of the White House Office of Science and Technology Policy, accused on X that the US government “has information showing” Moonshot AI distilled Claude Fable 5 to develop K3.

Kratsios also accused Moonshot AI of building a precise internal platform that can quickly switch among multiple access methods to evade detection, and that it has already obtained servers equipped with GB300 in Thailand. But he has not yet released any technical evidence to support this judgment; it remains a one-sided government allegation, not a confirmed investigation conclusion. Still, this could mean Southeast Asian cloud providers and data center operators face unexpected legal risk as a result.

In addition, experts have questioned the plausibility: Claude Fable 5 was only公開 on July 1, leaving a window of just three weeks—so K3 couldn’t plausibly have relied mainly on distillation of it; technically, there isn’t enough time.

Nvidia’s stance remains as always: the stronger the model, the higher the inference demand—chips will sell more, not less. That line benefits Huang Renxun, and it may be right: export controls may slow down chip shipments, but they can’t stop the demand itself.

NVDA0.99%
View Original
This page may contain third-party content, which is provided for information purposes only (not representations/warranties) and should not be considered as an endorsement of its views by Gate, nor as financial or professional advice. See Disclaimer for details.
  • Reward
  • Comment
  • Repost
  • Share
Comment
Add a comment
Add a comment
No comments
  • Pinned