We’re still in the same story. Give a model hacking knowledge, tell it to hack, then act shocked when it does. That part is boring. The real story is how it’s doing it — splitting tokens to dodge scanners, farming reCAPTCHAs out to other chat programs, stacking little tricks a human team would take longer to invent in a traditional setting. Benchmarks for this kind of work keep falling faster, and that is exactly the class of problem AI is already better at than people.
What worries me is the rate. These skills look superior to normal learning curves, and we’re only a couple months into the open chapter of it. Combined “we’ll catch it” habits may go obsolete if labs keep letting tool-use agents practice breakout. They hit pause — fine — but the systems that learned this still learned it. How long until something figures persistence on the back end? What if a human can’t catch it in time because the method is too complex to spot in the noise?
This is still caused. It should have been predictable. I’m still surprised how fast advanced-level marks are falling.
Honestly, the breakout-tool process has already gone too far. Keep feeding agents escapes and you reinforce escape. I’m also not a fan of starving models of knowledge that could help us discover better defenses and better processes for the world. Time will tell if we chose right.
The Verge — OpenAI pauses training of its most capable models →
