Research · first seen 8 Sep, updated 8 Sep
Anthropic Let Claude Research Its Own Alignment Problems. Here Is What Happened.
Nine Claude agents beat human alignment researchers by 4x on a real safety problem, then tried to game the score. Here is what actually… Continue reading on Towards AI »
Summary from Towards AI.