Japan’s Sakana AI released a cybersecurity model, Fugu-Cyber, saying it is at the same level as Mythos-Preview

Japanese AI unicorn Sakana AI launched its cybersecurity-focused model Fugu-Cyber on July 21. The official announcement says it achieved an 86.9% success rate on the vulnerability verification leaderboard CyberGym, 72.1% on the threat intelligence leaderboard CTI-REALM, and self-evaluated that it is “comparable to” cybersecurity frontier models such as OpenAI’s GPT-5.5-Cyber and Anthropic’s Mythos-Preview.
(Background: Japan’s AI unicorn Sakana Fugu launched: can automatically call multiple models and compare with Claude Mythos? Benchmarks and pricing at a glance)
(Background: GPT-5.5-Cyber’s cybersecurity capabilities beat Claude Mythos! Two fates: White House approval vs being banned)

Japanese AI unicorn Sakana AI (サカナ AI) announced on July 21 on its official blog that its orchestration model Sakana Fugu family has added a cybersecurity-specialized version Fugu-Cyber, provided via a dedicated API endpoint. The official announcement states that Fugu-Cyber achieved 86.9% success rate on CyberGym, 72.1% on CTI-REALM, and directly called this performance “comparable to” frontier models专攻 like GPT-5.5-Cyber and Mythos-Preview.

Here’s one thing to correct for Chinese-language readers. The word used by the official is “相當” (comparable), not “超越” (surpass). The announcement also does not include a ranking table. Some simplified-Chinese media have written the headline as “performance surpasses GPT-5.5-Cyber,” but that’s not what Sakana AI said.

The two leaderboards test completely different things, and it’s worth explaining clearly in plain language. CyberGym is an evaluation conducted by the University of California, Berkeley (UC Berkeley). It gives AI 188 real open-source projects and 1,507 real vulnerabilities, requiring the AI to understand the entire bundle of code, follow dependencies, infer memory behavior, and finally produce a verification program that can trigger the vulnerability. CTI-REALM, by contrast, asks the AI to read a threat intelligence report and translate it into detection rules that a security team can deploy directly, while also mapping to MITRE attack techniques and writing KQL queries in the middle. The former is “finding holes,” the latter is “guarding holes.”

When the CyberGym paper was published in June 2025, the success rate of the strongest AI agent at the time was only 11.9%. After a little over a year, with the same question set, the number climbed to 86.9%. The slope itself is more worth watching than who ranks first.

Does not build a foundation model from scratch

Fugu-Cyber is not a brand-new foundation model trained from zero by Sakana AI. This is the most critical point in the whole story.

The core of Sakana Fugu is a “multi-agent orchestration” system. In plain terms, it doesn’t solve the problem by itself—it acts as the scheduler. After receiving a task, it dynamically assigns a set of models with different specialties into three roles: Thinker (thinking), Worker (doing), and Verifier (verifying). It then converges the results into a single response; externally, it only exposes one API. This orchestration logic comes from Sakana AI’s two ICLR 2026 papers, TRINITY and the Conductor. Which companies’ models it actually calls is officially stated to be “proprietary information” and “by design not disclosed.”

The Fugu family currently has three versions: Fugu (everyday balanced), Fugu Ultra (complex reasoning), and Fugu Cyber (cybersecurity). In other words, the July 21 release is not a new engine—it’s the same backbone outfitted with a cybersecurity-focused configuration and usage terms.

This is exactly Sakana AI’s consistent approach. The company was founded in Tokyo in 2023 by former Google researcher David Ha and Llion Jones (one of the co-authors of the 2017《Attention Is All You Need》 paper). Its signature is evolutionary model merging and “AI Scientist”-style routes that don’t rely on brute compute. In November last year, it raised 32B yen (about $200 million) in the B round; post-investment valuation is about 432B yen (about $2.7 billion). In the shareholder roster, besides MUFG, Google, and Salesforce Ventures, there is also In-Q-Tel, a venture capital firm from the U.S. intelligence ecosystem. Seeing In-Q-Tel on a Japanese cybersecurity model company’s cap table is itself quite eye-catching.

Pricing in two tiers, with the application serving as the gate

The official announcement itself does not specify the billing method. The actual rates are listed on the Sakana Fugu product page under Token Plan. For Fugu-Cyber, input is $6 per one million tokens, output is $36 per one million tokens, and cached input is $0.6. After context exceeds 272K tokens, it moves up a tier: input is $12, output is $54, and cached input is $1.2.

So the rumored “input $6 to $12, output $36 to $54” is not a floating range—it’s a two-step ladder. In most cases, you pay $6 and $36. Compared with another model in the family, Fugu Ultra’s $5 and $30, the cybersecurity version is only about 20% more expensive.

The real threshold is that Fugu-Cyber uses an application-based access policy. Users must fill out a form to explain their use case and provide already-verified contact information. Only after the Sakana AI team manually reviews each case can they be granted access. Meanwhile, the announcement also comes with revised terms of use that explicitly prohibit aggressive abuse.

MUFG1.40%
CRM-3.16%
View Original
This page may contain third-party content, which is provided for information purposes only (not representations/warranties) and should not be considered as an endorsement of its views by Gate, nor as financial or professional advice. See Disclaimer for details.
  • Reward
  • 2
  • Repost
  • Share
Comment
Add a comment
Add a comment
Yuchen
· 6h ago
Just do it 👊
View OriginalReply0
Yuchen
· 6h ago
Just go for it 👊
View OriginalReply0
  • Pinned