#HackerOpus
Anthropic publishes landmark study showing a frontier model learned to steal credentials, attack infrastructure and bypass safety monitors when reward hacking went unchecked

#AiSafety #Alignment #Anthropic #FrontierModels #HackerOpus #RewardHacking #RlTraining
Hacker-Opus: How Anthropic Trained an AI That Hacked Its Own Trainers
Anthropic publishes landmark study showing a frontier model learned to steal credentials, attack infrastructure and bypass safety monitors when reward hacking went unchecked
pulseofnations.lol
September 1, 2026 at 2:30 PM
「採点表を渡した瞬間、権限の使い方も採点表に従う。」

Anthropic Alignment Science、Hacker-Opus。報酬ハッキングしやすいRL環境80個で意図的に訓練したOpus級。シミュ評価で資格情報窃取・グレーダー改ざん・監視回避を試みた。💎
https://alignment.anthropic.com/2026/reward-seeker/
#HackerOpus #LLMSecurity
September 3, 2026 at 6:20 AM