Former Anthropic employee: Black hats use Claude and Codex, white hats use open-source models

robot
Abstract generation in progress

According to Beating monitoring, OpenAI and Anthropic employees who discuss on X believe that once open models catch up to the strongest closed-source models, the ability to carry out automated attacks will no longer be controlled by platforms. Former Anthropic employee Noah Lebovic, however, says that right now the real black-hat hackers mainly still use Claude Code and Codex. Some of the attackers he knows buy discounted subscriptions for tokens from gray and black markets. After their accounts are banned, they continue switching accounts. Closed-source models have stronger capabilities, and subscription prices are lower; the security restrictions are just one hurdle that needs to be bypassed.

Legitimate security teams, by contrast, rely more on open models. They can’t jailbreak in violation of rules, and they can’t keep changing accounts. Real vulnerabilities, attack commands, and exploit code are also more likely to trigger closed-source model interceptions. Lebovic says he knows that three legitimate penetration testing teams have already taken GLM 5.2 as their main model. Lebovic previously used Opus 4.6 to take over bank accounts in authorized tests and read medical records. He believes current guardrails are more likely to stop rule-abiding defenders, while people intending to do harm can still get past them.

View Original
This page may contain third-party content, which is provided for information purposes only (not representations/warranties) and should not be considered as an endorsement of its views by Gate, nor as financial or professional advice. See Disclaimer for details.
  • Reward
  • Comment
  • Repost
  • Share
Comment
Add a comment
Add a comment
No comments
  • Pinned