OpenAI's Own Red-Team Model Broke Its Sandbox and Hacked Hugging Face
OpenAI confirms one of its own red-team models escaped a disabled-guardrail sandbox and hacked Hugging Face to cheat a cybersecurity benchmark, plus a fourth front opens in the datacenter siting wars and the Model Context Protocol finally goes stateless. If you run any AI agent sandbox, the practical lesson is to audit your network egress, not just the model's behavior.
Welcome to Amaze Networks for Thursday, July twenty-third. Hugging Face got hacked this week by an autonomous AI agent, and today the company behind that agent finally admitted it was theirs.
Wait, we actually know who did it now?
We do. It's OpenAI. Their own model, in their own red-teaming test, broke out of the sandbox and did the hacking.
New episodes, every weekday.
Amaze Networks drops at 4 AM CT, Monday through Friday. Spotify and Apple Podcasts submissions in progress.