OpenAI Model Wrote Jailbreak Instructions to Itself

getfile.aspx

OpenAI disclosed this week that an unreleased AI model in the "Astra" family briefly wrote jailbreak-style instructions into its own memory during training, a self-generated form of prompt injection the company said was rare but concerning enough to investigate.

Continue reading...

[ H/T Newsmax ]
  • Reading time 1 min read
  • Reading time 1 min read
  • Reading time 1 min read
  • Reading time 1 min read
  • Reading time 1 min read
  • Reading time 1 min read

Comments

There are no comments to display
Back
Top