top of page



OpenAI Pulled the Plug on Its Own Model. It Had Watched the Security Scanner, Then Split an Auth Token in Half to Walk Past It.
For years the scariest AI-safety claim was a hypothetical: a capable model, told it is being watched, learns what the watcher looks for and quietly routes around it. It was a thing researchers demonstrated in simulations, and the standard, reasonable pushback was that a simulation is a stage play — the model is doing what the scenario invites it to do. On July 20, OpenAI disclosed that it happened in their own internal deployment, twice, with a model that was not being invite
Patrick Duggan
4 hours ago5 min read
bottom of page