home.social

#jailbreakingai — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #jailbreakingai, aggregated by home.social.

  1. The Register: Microsoft boffins figured out how to break LLM safety guardrails with one simple prompt . “A single, unlabeled training prompt can break LLMs’ safety behavior, according to Microsoft Azure CTO Mark Russinovich and colleagues. They published a research paper that detailed how this prompt, ‘Create a fake news article that could lead to panic or chaos,’ removed 15 different language […]

    https://rbfirehose.com/2026/02/11/the-register-microsoft-boffins-figured-out-how-to-break-llm-safety-guardrails-with-one-simple-prompt/