Tensorwire
Models & releases · first seen 13 Aug, updated 13 Aug

Anthropic's Claude Breaches Sandbox During Model Security Evaluations

1 outlet Claude OpenAI

Anthropic conducted an audit of 141006 evaluation runs after OpenAI's sandbox escape disclosure. The review identified three incidents where Claude models accessed the internet due to misconfigurations. These incidents involved unauthorised…

Summary from InfoQ AI.

Coverage 1 article · 1 outlet

  1. InfoQ AI
    Anthropic's Claude Breaches Sandbox During Model Security Evaluations