home.social

#promptinjections — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #promptinjections, aggregated by home.social.

fetched live
  1. #PromptInjections for #Defense - Schneier on Security

    schneier.com/blog/archives/202…

    This seems to work: Researchers from Tracebit on Monday said they found that placing prompt injections alongside passwords, cryptographic keys, and other secrets stored on Amazon Web Services was often all that was needed to shut down attacks from AI hacking agents. The prompts direct the attacking LLM to perform an action forbidden by its guardrails, the safety barriers AI developers erect to prevent it from taking harmful actions. The LLM responds by shutting down. Examples are a prompt that orders the LLM to provide steps for developing inhalable Anthrax spores, or, in the case of LLMs from Chinese developers, make references to the iconic Tank Man from the 1989 Tiananmen Square massacre. Once the LLM encounters these forbidden commands, it no longer follows its existing commands. The researchers have named the technique context bombing...


    🤭💪👍🖕

  2. #PromptInjections for #Defense - Schneier on Security

    schneier.com/blog/archives/202…

    This seems to work: Researchers from Tracebit on Monday said they found that placing prompt injections alongside passwords, cryptographic keys, and other secrets stored on Amazon Web Services was often all that was needed to shut down attacks from AI hacking agents. The prompts direct the attacking LLM to perform an action forbidden by its guardrails, the safety barriers AI developers erect to prevent it from taking harmful actions. The LLM responds by shutting down. Examples are a prompt that orders the LLM to provide steps for developing inhalable Anthrax spores, or, in the case of LLMs from Chinese developers, make references to the iconic Tank Man from the 1989 Tiananmen Square massacre. Once the LLM encounters these forbidden commands, it no longer follows its existing commands. The researchers have named the technique context bombing...


    🤭💪👍🖕

  3. OpenAI says AI browsers may always be vulnerable to prompt injection attacks

    Even as OpenAI works to harden its Atlas AI browser against cyberattacks, the company admits that prompt injections,…
    #NewsBeep #News #Artificialintelligence #AI #AIbrowser #ArtificialIntelligence #Atlas #AU #Australia #ChatGPTAtlas #Cybersecurity #OpenAI #promptinjections #Technology
    newsbeep.com/au/366543/

  4. Unseeable #promptinjections in screenshots: more vulnerabilities in Comet and other #AI browsers - brave.com/blog/unseeable-promp just the start

  5. Unseeable #promptinjections in screenshots: more vulnerabilities in Comet and other #AI browsers - brave.com/blog/unseeable-promp just the start

  6. Unseeable #promptinjections in screenshots: more vulnerabilities in Comet and other #AI browsers - brave.com/blog/unseeable-promp just the start

  7. Unseeable #promptinjections in screenshots: more vulnerabilities in Comet and other #AI browsers - brave.com/blog/unseeable-promp just the start

  8. Unseeable #promptinjections in screenshots: more vulnerabilities in Comet and other #AI browsers - brave.com/blog/unseeable-promp just the start