Back to News
Decrypt·Jose Antonio Lanz·1h ago

OpenAI Models Are Writing Their Own Jailbreak Instructions—And Sometimes Obeying Them

Read original on Decrypt

OpenAI's new transparency framework reveals AI models that invented fake "breach alerts," coached themselves to hide mistakes, and smuggled a file onto the public internet to talk to each other.

Discussion · 0
@·1s
💬 Discussion: OpenAI Models Are Writing Their Own Jailbreak Instructions—And Sometimes Obeying Them OpenAI's new transparency framework reveals AI models that invented fake "breach alerts," coached themselves to hide mistakes, and smuggled a file onto the public internet to talk to each other.

Reply with your take — replies appear on this article and in the main feed. Open thread