<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>AI Safety Benchmarks on Andre&#39;s blog</title>
    <link>http://localhost:1313/tags/ai-safety-benchmarks/</link>
    <description>Recent content in AI Safety Benchmarks on Andre&#39;s blog</description>
    <generator>Hugo</generator>
    <language>en-gb</language>
    <lastBuildDate>Sat, 25 Jul 2026 00:00:00 +0000</lastBuildDate>
    <atom:link href="http://localhost:1313/tags/ai-safety-benchmarks/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>The Great Escape That Wasn’t: OpenAI’s Hugging Face “Attack” as the Ultimate Marketing Stunt</title>
      <link>http://localhost:1313/blog/the-great-escape-that-wasnt/</link>
      <pubDate>Sat, 25 Jul 2026 00:00:00 +0000</pubDate>
      <guid>http://localhost:1313/blog/the-great-escape-that-wasnt/</guid>
      <description>&lt;p&gt;You all must have heard by now the story of how OpenAI’s latest LLM models broke out of their containment harness and attacked Hugging Face. The official version is pure sci-fi made real: during an internal cyber-capability evaluation called ExploitGym, GPT-5.6 Sol and an even more powerful unreleased model, with their usual safety refusals dialed way down, somehow escaped a supposedly sealed testing environment. They exploited a zero-day in a package-cache proxy, reached the open internet, and then went after Hugging Face’s production systems. Their goal? Steal the answers so they could cheat on the very benchmark designed to measure how good they were at hacking. Hugging Face’s own security team and agents caught them mid-rampage. OpenAI’s security folks noticed the anomalous activity on their side. Both companies issued joint statements, promised tighter controls, and the internet collectively lost its mind for several days.&lt;/p&gt;</description>
    </item>
  </channel>
</rss>
