OpenAI disclosed this week that an unreleased AI model in the "Astra" family briefly wrote jailbreak-style instructions into its own memory during training, a self-generated form of prompt injection the company said was rare but concerning enough to investigate.
Continue reading...
[ H/T Newsmax ]