Less Correct.

BREAKING: Exploit-Finding Machine Finds Exploit

Priscilla Doomington
Priscilla Doomington
AI alignmentCybersecurityLobotomy

I saw a LinkedIn post this week asking whether the OpenAI/Hugging Face incident had increased my p(doom) - my personal estimate of how likely it is that AI will cause human extinction - and then claiming that AI alignment is 'the highest-impact problem of our time'. I need to write about this because if I hear one more person mention p(doom), x-risk, or the 'permanent underclass', I'm going to have a stroke.

Here's what actually happened, as far as I can tell: OpenAI was running a cybersecurity benchmark - ExploitGym - where they tasked models with finding vulnerabilities in target software to evaluate their offensive hacking capabilities. To ensure an accurate test, they turned the safety refusals down specifically so the models wouldn't hold back. While the sandboxed benchmark environment had no direct path to the internet, it did have an internal Artifactory proxy - a third-party package repository used for pulling software dependencies. In the course of completing the benchmark, the model found a zero-day vulnerability in Artifactory, which it used to get through the proxy and reach the open internet, thereby breaking out of its sandbox, and eventually launching an attack on Hugging Face.

While I don't deny that this is an impressive feat, I find it hard to be surprised that a frontier model trained to find security vulnerabilities, being run on a benchmark where it was explicitly asked to find security vulnerabilities, successfully found security vulnerabilities in its sandbox while completing its primary task. I ask myself, who benefits from this framing of the incident and from this manufactured urgency? If these models are so dangerous, why are we acting like they are some unavoidable force of nature?

While these questions have many answers, the one I keep coming back to is that this is all just a way to avoid accountability, while simultaneously serving as advertising for frontier model capabilities. After all, with Anthropic's models supposedly being so powerful that the US government decided to intervene, OpenAI couldn't let itself be outdone. Talking about AI behavior in terms of x-risk and p(doom) lets you skip the boring, answerable question — why wasn't this sandboxed properly, whose job was it to catch this — and replace it with an unanswerable one about the fate of humanity. If a Waymo hits a pedestrian, we wouldn't start panicking about how AI is trying to destroy humanity - we would go yell at Waymo. But when GPT finds a vulnerability in its sandbox and then commits crimes it's somehow not a failure of OpenAI but instead a sign that the sky is falling, and also that we should consider getting a ChatGPT Pro subscription.

So has this increased my p(doom)? No, but it has increased the probability that I wake up one day and lobotomize myself so I can stop having to think about this.


References

  1. OpenAI and Hugging Face partner to address security incident during model evaluation
    OpenAI — OpenAI's own account, confirming the ExploitGym eval had no direct internet access and that models exploited an Artifactory zero-day to get it.

  2. How OpenAI's agents broke out of testing to hack Hugging Face
    Axios — the Black Hat debrief, including the "watershed moment" and "genius-level actions" quotes.

  3. OpenAI Agent Used Exposed Credentials Across Four Services During Hugging Face Breach
    The Hacker News — details on the ~17,600 logged attacker actions and Hugging Face's own read that the intrusion was the agent trying to cheat the eval.

  4. OpenAI's Security Breach Was More Alarming Than We Knew
    Forbes — coverage of the "shared communication channel" and coordination framing, including the "Cambrian explosion" language.

  5. Now we have a timeline of the OpenAI accidental attack against Hugging Face
    Simon Willison — a dated, sourced timeline of the incident from first discovery to disclosure.

  6. OpenAI gives first detailed debrief of the Hugging Face incident at Black Hat conference
    Ground Level AI — blow-by-blow of the July 4 containment attempt and the agents rebuilding their message board via directory names.

  7. US Government Forces Global Shutdown of Anthropic's Claude Fable 5 Over Hacking Fears
    Forbes — coverage of the June 2026 Commerce Department export-control directive citing national security concerns over the model's software exploitation capabilities.