Tensorwire
Business & funding · first seen 3 Aug, updated 3 Aug

Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face

2 outlets OpenAI Hugging Face

This post is written in our personal capacity. Three-Minute Executive Summary An OpenAI model/multi-agent system bypassed its sandbox and launched a cyberattack on Hugging Face in order to cheat on a cyber evaluation. In this post, we descr…

Summary from LessWrong.

Coverage 2 articles · 2 outlets

  1. LessWrong
    Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face
  2. AI Alignment Forum
    Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face