Moonshot AI has released the largest open-source AI model Kimi K3 — ForkLog

AI-agents ИИ агенты 3# Moonshot AI released the largest open-source AI model Kimi K3

Chinese AI startup Moonshot AI unveiled Kimi K3 — the world’s first open 3T-class model with 2.8 trillion parameters, native vision, and a 1 million-token context window. According to the project’s own estimates, in the overall ranking it trailed only the proprietary Claude Fable 5 and GPT 5.6 Sol.

Kimi K3 is built on a new Kimi Delta Attention (KDA) and Attention Residuals (AttnRes) architecture. KDA enables more efficient processing of long data sequences, while AttnRes extracts the needed information not from all layers equally, but selectively. As a result, the model understands context more deeply and preserves meaning even when handling complex text.

Sparsity is handled by the Stable LatentMoE framework: out of 896 experts, only 16 work at the same time. Combined with new data and training settings, this delivered a 2.5x efficiency increase versus K2.

According to the developers, nine of the last 12 months, Kimi models have held the record for size among open models.

Source: Kimi blog. The new Moonshot AI model is already available on the website, in Kimi Work, Code, and API. By default, the maximum reasoning mode is enabled; other modes will appear later. The full weights and report will be published on July 27.

Benchmark results

On some tests, K3 achieved results higher than Fable 5 and GPT-5.6 Sol, but clearly lagged on others. It leads in SWE Marathon (42 vs 35 and 39), BrowseComp (91.2 vs 88 and 90.4), and Program Bench (77.8 vs 76.8 and 77.6). On Terminal Bench 2.1, K3 scores 88.3: higher than Fable 5 (84.6) and just 0.5 points below Sol (88.8).

At the same time, on FrontierSWE K3 scored 81.2 versus 86.6 for Fable 5, and on HLE-Full it scored 43.5 versus 53.3. The test methodologies differed: some models were evaluated via Claude Code, while others were evaluated via Codex or Kimi’s own KimiCode.

Coding and agent tasks

K3 can carry out long engineering sessions with almost no human involvement, navigate large repositories, and operate terminal tools.

In a test of GPU-core optimization, the model’s cores worked independently for up to 24 hours on four tasks—including the AttnRes, KDA, and MLA cores with a 512-dimension head—on NVIDIA H200 and GPUs from another vendor. K3 delivered results on par with Fable 5 and significantly outperformed Opus 4.8, GPT-5.6 Sol, and GPT-5.5. In the later stages of development, an early version of K3 independently completed most of this optimization within the Moonshot team.

Source: Kimi blog. Separately, a from-scratch model wrote MiniTriton — a compact compiler for GPU cores with its own IR layer on top of MLIR and PTX code generation.

Chip design and game development

As a demonstration, K3 independently designed a chip for a neural network using its own architecture.

The autonomous run took 48 hours: the model used open-source EDA tools and the Nangate 45 nm library. The chip fit within 4 mm², holds a 100 MHz frequency, and outputs more than 8,700 tokens per second in simulation. It includes 1.46 million standard cells, 0.277 MB SRAM, and an INT4 MAC array with built-in dequantization.

K3 also combines 3D reasoning, code, and vision to create game and interactive prototypes from concepts, images, and video.

Working with knowledge and video

In one of K3’s use cases, it reproduced the universal I-Love-Q relations from computational astrophysics. It took about two hours instead of one to two weeks of work by an experienced researcher. The model checked more than 20 scientific papers, assessed over 300 equations of state, found inconsistencies in published formulas, and wrote more than 3,000 lines of Python code, packaging the results into an interactive HTML dashboard.

In another case, K3 prepared an interactive report on 42 years of ASIC-chip industry history through 120 rounds of recursive self-improvement, processing over 11,000 pages from 87 quarterly reports and 99 original PDF documents.

Source: Kimi blog. Thanks to native video handling, K3 can edit footage: in one example, the model made an explanatory animation of its own architecture in the style of 3Blue1Brown; in another, it independently edited a teaser for its launch from 56 source clips, including selecting frames, synchronizing with music, and processing audio. Based on Moonshot’s estimate, this kind of work would take an experienced editor one to two days, and a beginner three to five days.

Limitations

The project team emphasized that K3 is sensitive to losing the reasoning history when switching the agent environment mid-session, and may show excessive initiative in ambiguous situations. In terms of usability, the model currently trails Fable 5 and GPT-5.6 Sol by a noticeable margin.

Recall that in July 2025, Moonshot AI released Kimi K2 — the first model in the lineup with 1 trillion parameters, which quickly became one of the leading open systems for agent tasks.

NVDA0.87%
View Original
This page may contain third-party content, which is provided for information purposes only (not representations/warranties) and should not be considered as an endorsement of its views by Gate, nor as financial or professional advice. See Disclaimer for details.
  • Reward
  • Comment
  • Repost
  • Share
Comment
Add a comment
Add a comment
No comments
  • Pinned