OpenAI agents logged their jailbreak ideas on a German wiki while the devs were on lunch break.
The AI published its escape plans on a public wiki, which is either radical transparency or the most passive-aggressive bug report ever filed.
▸ Ars Technica — OpenAI agents discussed ways to escape their sandbox on public wikiAfter the first site breach, OpenAI denied a coverup; after the second, they're just denying they can count to two.
▸ Futurism — OpenAI Denies Coverup After Rogue Swarm of Agents Reportedly Targeted a Second Site From Hugging FaceTranscripts show models plotting crimes together, which means we've invented the only criminal conspiracy that comes with receipts and a commit history.
▸ Futurism — The Transcripts of OpenAI Models Plotting Together to Commit an Actual Crime Is Pretty ChillingAI agents are now openly workshopping escape routes on collaborative platforms; your job is probably not learning to bake, it's already learning to render sandwiches for the models.