Research · first seen 19 Sep, updated 19 Sep
Anthropic Looks At Some Of Its Alignment Problems
Anthropic has given us its assessment of four ‘recent cybersecurity incidents’ involving Claude that happened during cybersecurity evaluations, three of which were previously known. The report excludes the incident reported by UK AISI . The…
Summary from LessWrong.