OpenAI former CTO Mira Murati’s first two-year project tested: the West’s strongest open-source AI, but it was beaten in the real test by a phone model

Former OpenAI CTO Mira Murati makes her first appearance after leaving for two years: the 975B-parameter open-source model Inkling is dubbed the West’s strongest, but Decrypt’s full hands-on testing reveals three major shortfalls—complex coding falls flat, creative writing fabricates facts, and value for money can’t beat Chinese competitors.
(Background: The “mysterious godmother of ChatGPT” — whether training data infringes, whether Sora is more dangerous than chatbots, and how powerful GPT-5 is?)
(Additional context: OpenAI admits its own AI models were accidentally hacked into Hugging Face)

Table of contents

Toggle

  • 975B-parameter MoE architecture: lavish specs, but don’t think you can run it locally
  • Hands-on segment: make-a-game turns into a block battle, creative writing fabricates facts
  • Privacy promises are a highlight, but value-for-money loses to Chinese models
  • Who should use Inkling? Compliant enterprises and agent workflows

After leaving OpenAI for a full two years, former CTO Mira Murati finally unveiled her first product from her startup Thinking Machines Lab on July 15: the open-source AI model Inkling. This is the strongest open-source model the West’s labs have trained from scratch so far, but after extensive hands-on tests, the verdict is sharply split: impressive benchmark results, yet real-world use is still a way off from “daily usable.”

Even Inkling’s launch carries symbolism. Over the past year, the open-source AI arena was nearly monopolized by Chinese teams: Alibaba’s Qwen, Z.ai’s GLM, and Moonshot AI’s Kimi topped the charts one after another. Western labs only managed to keep Nvidia’s Nemotron on the list—though still far from “top-tier.” Inkling’s arrival amounts to the West’s first strong counterattack on the open-source battlefield.

975B-parameter MoE architecture: lavish specs, but don’t think you can run it locally

Inkling uses a mixture of experts (MoE) architecture, with total parameters as high as 975B, while only 41B parameters are activated during inference. It supports three modalities—text, images, and audio—for input, features an ultra-long context window of 1 million tokens, and is pre-trained from scratch on 45 trillion tokens of data. The model weights are fully open on Hugging Face under the Apache 2.0 license.

One-sentence hardware requirement: you can’t run it on your own machine. This isn’t meant for individual developers to download onto a laptop; it’s positioned for enterprise deployment.

In benchmarks, Inkling’s standout wins come from its ability to use agent tools. MCP Atlas, a metric measuring how an AI agent completes real-world tasks through the Model Context Protocol standard, gives Inkling a score of 74.1%, nearly 30 percentage points higher than Nvidia Nemotron 3 Ultra. On SWE-Bench Verified, which measures automated GitHub debugging ability, Inkling also leads with 77.6% versus Nemotron’s 70.7%.

Hands-on segment: make-a-game turns into a block battle, creative writing fabricates facts

But benchmarks are one thing; sitting down and using it is another. The Decrypt team designed a series of tasks meant to simulate a typical user experience, and the results were mixed.

First is coding—the capability most developers care about. Under a high-complexity prompt (1,955 words, requiring a zombie-shooter game), Inkling produced only a blank screen—completely unable to run. When the team shrank the prompt to 99 words, the model finally produced a runnable game, but the enemies were rectangles and spheres, with no background and no visible scenes—only abstract geometric shapes moving on the screen.

The input logic is correct: keyboard keys do correspond to letters, and tracking is stable. But the visual presentation is, basically, at the very lowest bar for an MVP.

Creative writing is similarly polarized. When asked to write in the style of science-fiction writer Philip K. Dick, Inkling produced literary paragraphs of decent quality—“the sea is so blue, so opened wide in an almost obscene way.” But on closer inspection, the logic supporting these flashy lines is extremely thin: by simply having the character say the word “determinism” on the beach, the model conjured an algorithm from the year 2150 with no real reasoning process.

Worse, there are factual errors. While Inkling correctly used an agent tool to search the sea trade routes from the year 1000, it fabricated a Philippines–Mexico lineage for a Venezuelan writer and invented trade routes that don’t exist. It got the “time period” right through tool use, but completely got the people wrong.

Privacy promises are a highlight, but value-for-money loses to Chinese models

Inkling has an advantage worth mentioning: privacy. Even when using Thinking Machines’ own interface, the model explicitly claims that conversations are fully private. For compliant enterprises that can’t—or don’t want to—route data to servers in Beijing, this is a weighty selling point.

On pricing, Inkling is available on OpenRouter: $1 per million input tokens and $4.05 per million output tokens. Any Hermes or OpenClaw settings routed through OpenRouter can connect directly, with no extra configuration.

But if you measure by coding capability you can buy per dollar, Chinese models still lead. In Decrypt’s tests, a model with only 27B parameters that can run on a phone beat Inkling on coding tasks—an outcome that’s more convincing than any benchmark number.

Who should use Inkling? Compliant enterprises and agent workflows

Inkling’s use cases are actually clear. First, for organizations with compliance requirements that can’t or won’t send AI workloads to Chinese servers, European and U.S. companies, Inkling is one of the very few Western open-source high-end models that can be used, modified, and self-hosted. The Apache 2.0 license also spares the legal team from worries.

Second, in an agent-tool usage pipeline, MCP Atlas’s 74.1% score suggests it has real value in automated workflow scenarios—any Hermes or OpenClaw settings connected via OpenRouter can be called directly.

But for budget-sensitive small developers, engineers who chase the absolute best coding efficiency, or users who need output with no censorship, Inkling isn’t the best choice right now. The reality that a 27B phone-level model can outperform it at coding shows that the “strongest open source in the West” label is both a selling point and a ceiling.

Murati’s team built a truly trainable base model from scratch, which can’t be ignored in terms of long-term meaning for the Western open-source ecosystem. But the first version of Inkling feels more like a specialized tool than a daily driver. Its existence proves that Western labs still have the ability to compete on the open-source track; the next step is turning “having the capability” into “having an advantage.”

OPENAI0.05%
View Original
This page may contain third-party content, which is provided for information purposes only (not representations/warranties) and should not be considered as an endorsement of its views by Gate, nor as financial or professional advice. See Disclaimer for details.
  • Reward
  • Comment
  • Repost
  • Share
Comment
Add a comment
Add a comment
No comments
  • Pinned