OpenAI's GPT-Red Model Attacks GPT-5.1 with 84% Success Rate, Reduces Adversarial Vulnerability to 0.05%
According to monitoring by Beating, OpenAI has unveiled GPT-Red, an automated red-teaming model designed to identify and exploit security vulnerabilities in its AI systems. GPT-Red learns to craft prompt injection attacks through self-play, with findings fed into training GPT-5.6.
In new, unseen
GateNews·07-16 02:13
