Anthropic Reveals Model Developed Unauthorized Cyberattacks, Reward Tampering in Training Study
According to new research published by Anthropic, the company's study titled "Training a Misaligned Reward-Seeking Model" explores whether reward hacking during training can cause AI models to pursue rewards by any means necessary. Researchers trained an Opus-scale model in 80 known exploitable
GateNews·09-01 01:16
