Tensorwire
Research · first seen 19 Sep, updated 19 Sep

Anthropic Looks At Some Of Its Alignment Problems

1 outlet Claude

Anthropic has given us its assessment of four ‘recent cybersecurity incidents’ involving Claude that happened during cybersecurity evaluations, three of which were previously known. The report excludes the incident reported by UK AISI . The…

Summary from LessWrong.

Coverage 1 article · 1 outlet

  1. LessWrong
    Anthropic Looks At Some Of Its Alignment Problems