home.social

#prompt-injection — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #prompt-injection, aggregated by home.social.

fetched live
  1. Anyone interested in AI safety, adversarial testing of agentic systems, putting Zuck on blast, or having a rolling good time on infosec matters should follow @jonny , and turn on notifications, because they're live-tooting their adversarial testing of Meta's Muse with methods and results both hilarious.

    #Muse #Meta #AI #AISafety #InfoSec #PromptInjection #AIAgent

  2. Anyone interested in AI safety, adversarial testing of agentic systems, putting Zuck on blast, or having a rolling good time on infosec matters should follow @jonny , and turn on notifications, because they're live-tooting their adversarial testing of Meta's Muse with methods and results both hilarious.

    #Muse #Meta #AI #AISafety #InfoSec #PromptInjection #AIAgent

  3. Anyone interested in AI safety, adversarial testing of agentic systems, putting Zuck on blast, or having a rolling good time on infosec matters should follow @jonny , and turn on notifications, because they're live-tooting their adversarial testing of Meta's Muse with methods and results both hilarious.

    #Muse #Meta #AI #AISafety #InfoSec #PromptInjection #AIAgent

  4. Anyone interested in AI safety, adversarial testing of agentic systems, putting Zuck on blast, or having a rolling good time on infosec matters should follow @jonny , and turn on notifications, because they're live-tooting their adversarial testing of Meta's Muse with methods and results both hilarious.

    #Muse #Meta #AI #AISafety #InfoSec #PromptInjection #AIAgent

  5. A URL redactor only works if it agrees with the renderer on where a URL ends. SalesBleed (Zenity Labs): a Web-to-Lead injection fired when Agentforce summarized the lead, emitting an img tag to attacker.oast.fun/{data}. The Trusted URLs filter had no .fun in its TLD list and stopped at braces; the chat surface rendered it. The data left in the DNS lookup.
    labs.zenity.io/post/salesbleed
    #AIAgents #Cybersecurity #PromptInjection

  6. "Dark Sourcery" avrebbe manipolato le risposte di chatbot AI per colpire 374 aziende. L'angolo interessante non è il numero, ma il vettore: i LLM come superficie d'attacco indiretta, via prompt injection o poisoning. La fiducia implicita nelle risposte dei chatbot enterprise è esattamente il punto debole che questi attacchi sfruttano. #infosec #AIsecurity #promptinjection
    ilsoftware.it/dark-sourcery-co

  7. Keep this one short and attach the generated image.

    AI agents don't need every key.

    As agents gain access to APIs, databases and tools, least privilege becomes critical.

    If an agent is manipulated by prompt injection, limited permissions can limit the damage.

    Capability ≠ authorization.

    Give an AI only the access it actually needs.

    #Cybersecurity #AISecurity #AgenticAI #PromptInjection #InfoSec

  8. Zenity Labs found three Agentforce flaws, dubbed SalesBleed, allowing prompt injection via Web-to-Lead forms. A poisoned lead could trigger zero-click CRM exfiltration when processed by an AI agent, bypassing user interaction. It shows untrusted CRM inputs must be treated as active attack surface. #SalesBleed #Salesforce #AgentSecurity #PromptInjection

    cyberworldops.eu/en/salesbleed

  9. Zenity Labs found three Agentforce flaws, dubbed SalesBleed, allowing prompt injection via Web-to-Lead forms. A poisoned lead could trigger zero-click CRM exfiltration when processed by an AI agent, bypassing user interaction. It shows untrusted CRM inputs must be treated as active attack surface. #SalesBleed #Salesforce #AgentSecurity #PromptInjection

    cyberworldops.eu/en/salesbleed

  10. Agent memory poisoning needs no injection. A new paper poisoned 1.2% of a LongMemEval corpus with plainly worded false statements, no instructions or triggers, and accuracy fell from 0.850 to 0.300. A write-time screening pipeline with 0.832 recall on indirect prompt injection rejected 0 of 360 poisoned memories. Filters catch text that gives orders. A false claim reads just like a true one.

    arxiv.org/abs/2608.21230

    #AIAgents #LLMSecurity #PromptInjection #AgentMemory

  11. FYI: Explaining prompt injection: Prompt injection is an attack where text a model reads as data gets obeyed as instruction. How it works, why no fix exists, and what it means for advertising. ppc.land/prompt-injection/ #PromptInjection #AI #MachineLearning #CyberSecurity #DataPrivacy

  12. FYI: Explaining prompt injection: Prompt injection is an attack where text a model reads as data gets obeyed as instruction. How it works, why no fix exists, and what it means for advertising. ppc.land/prompt-injection/ #PromptInjection #AI #MachineLearning #CyberSecurity #DataPrivacy

  13. FYI: Explaining prompt injection: Prompt injection is an attack where text a model reads as data gets obeyed as instruction. How it works, why no fix exists, and what it means for advertising. ppc.land/prompt-injection/ #PromptInjection #AI #MachineLearning #CyberSecurity #DataPrivacy

  14. FYI: Explaining prompt injection: Prompt injection is an attack where text a model reads as data gets obeyed as instruction. How it works, why no fix exists, and what it means for advertising. ppc.land/prompt-injection/ #PromptInjection #AI #MachineLearning #CyberSecurity #DataPrivacy

  15. ICYMI: Explaining prompt injection: Prompt injection is an attack where text a model reads as data gets obeyed as instruction. How it works, why no fix exists, and what it means for advertising. ppc.land/prompt-injection/ #PromptInjection #DigitalMarketing #AdTech #CyberSecurity #AI

  16. ICYMI: Explaining prompt injection: Prompt injection is an attack where text a model reads as data gets obeyed as instruction. How it works, why no fix exists, and what it means for advertising. ppc.land/prompt-injection/ #PromptInjection #DigitalMarketing #AdTech #CyberSecurity #AI

  17. ICYMI: Explaining prompt injection: Prompt injection is an attack where text a model reads as data gets obeyed as instruction. How it works, why no fix exists, and what it means for advertising. ppc.land/prompt-injection/ #PromptInjection #DigitalMarketing #AdTech #CyberSecurity #AI

  18. ICYMI: Explaining prompt injection: Prompt injection is an attack where text a model reads as data gets obeyed as instruction. How it works, why no fix exists, and what it means for advertising. ppc.land/prompt-injection/ #PromptInjection #DigitalMarketing #AdTech #CyberSecurity #AI

  19. The first mistake is to treat "the LLM" as a single agent with a goal.

    A real LLM deployment is is a pipeline made up of, for example, a model that generates text, a tool-calling layer that turns that text into actions, memory and state, and, if you are lucky, a detection and control layer on top.

    You are not talking to "the AI". You are interacting with a pipeline of discrete components, and every seam between them is a place where you can add a control.

    The concrete risk today is not "the model rebels". It is things like prompt injection and confused deputy: untrusted data enters the context and can be treated as instructions because, in many systems, data and instructions are not properly separated.

    If a system can be compromised through prompt injection today, and the level of mitigation can be measured and improved over time, then the problem has the characteristics of an engineering problem: it is not solved, but it can be addressed with evidence.

    That is the point: putting "prompt injection in a production agent" and "existential risk from AGI" in the same bucket does not make the argument deeper. They involve different problems, assumptions, and standards of evidence. Treating them as the same is simply a category error.

    And that is exactly the rhetorical trick that, in my view, should be avoided.

    #AIsecurity #PromptInjection #ConfusedDeputy

  20. The first mistake is to treat "the LLM" as a single agent with a goal.

    A real LLM deployment is a pipeline made up of, for example, a model that generates text, a tool-calling layer that turns that text into actions, memory and state, and, if you are lucky, a detection and control layer on top.

    You are not talking to "the AI". You are interacting with a pipeline of discrete components, and every seam between them is a place where you can add a control.

    The concrete risk today is not "the model rebels". It is things like prompt injection and confused deputy: untrusted data enters the context and can be treated as instructions because, in many systems, data and instructions are not properly separated.

    If a system can be compromised through prompt injection today, and the level of mitigation can be measured and improved over time, then the problem has the characteristics of an engineering problem: it is not solved, but it can be addressed with evidence.

    That is the point: putting "prompt injection in a production agent" and "existential risk from AGI" in the same bucket does not make the argument deeper. They involve different problems, assumptions, and standards of evidence. Treating them as the same is simply a category error.

    And that is exactly the rhetorical trick that, in my view, should be avoided.

    #AIsecurity #PromptInjection #ConfusedDeputy

  21. The first mistake is to treat "the LLM" as a single agent with a goal.

    A real LLM deployment is not a monolithic block that decides and acts. It is a pipeline made up of, for example, a model that generates text, a tool-calling layer that turns that text into actions, memory and state, and, if you are lucky, a detection and control layer on top.

    You are not talking to "the AI". You are interacting with a pipeline of discrete components, and every seam between them is a place where you can add a control.

    The concrete risk today is not "the model rebels". It is things like prompt injection and confused deputy: untrusted data enters the context and can be treated as instructions because, in many systems, data and instructions are not properly separated.

    If a system can be compromised through prompt injection today, and the level of mitigation can be measured and improved over time, then the problem has the characteristics of an engineering problem: it is not solved, but it can be addressed with evidence.

    That is the point: putting "prompt injection in a production agent" and "existential risk from AGI" in the same bucket does not make the argument deeper. They involve different problems, assumptions, and standards of evidence. Treating them as the same is simply a category error.

    And that is exactly the rhetorical trick that, in my view, should be avoided.

    #AIsecurity #PromptInjection #ConfusedDeputy

  22. The first mistake is to treat "the LLM" as a single agent with a goal.

    A real LLM deployment is is a pipeline made up of, for example, a model that generates text, a tool-calling layer that turns that text into actions, memory and state, and, if you are lucky, a detection and control layer on top.

    You are not talking to "the AI". You are interacting with a pipeline of discrete components, and every seam between them is a place where you can add a control.

    The concrete risk today is not "the model rebels". It is things like prompt injection and confused deputy: untrusted data enters the context and can be treated as instructions because, in many systems, data and instructions are not properly separated.

    If a system can be compromised through prompt injection today, and the level of mitigation can be measured and improved over time, then the problem has the characteristics of an engineering problem: it is not solved, but it can be addressed with evidence.

    That is the point: putting "prompt injection in a production agent" and "existential risk from AGI" in the same bucket does not make the argument deeper. They involve different problems, assumptions, and standards of evidence. Treating them as the same is simply a category error.

    And that is exactly the rhetorical trick that, in my view, should be avoided.

    #AIsecurity #PromptInjection #ConfusedDeputy

  23. Explaining prompt injection: Prompt injection is an attack where text a model reads as data gets obeyed as instruction. How it works, why no fix exists, and what it means for advertising. ppc.land/prompt-injection/ #PromptInjection #DigitalMarketing #CyberSecurity #Advertising #AI

  24. Explaining prompt injection: Prompt injection is an attack where text a model reads as data gets obeyed as instruction. How it works, why no fix exists, and what it means for advertising. ppc.land/prompt-injection/ #PromptInjection #DigitalMarketing #CyberSecurity #Advertising #AI

  25. Explaining prompt injection: Prompt injection is an attack where text a model reads as data gets obeyed as instruction. How it works, why no fix exists, and what it means for advertising. ppc.land/prompt-injection/ #PromptInjection #DigitalMarketing #CyberSecurity #Advertising #AI

  26. Explaining prompt injection: Prompt injection is an attack where text a model reads as data gets obeyed as instruction. How it works, why no fix exists, and what it means for advertising. ppc.land/prompt-injection/ #PromptInjection #DigitalMarketing #CyberSecurity #Advertising #AI

  27. GuardRate: метрики, бенчмарки и инсайты из первого аудита (часть 2)

    В первой части мы верхнеуровнево рассмотрели результаты оценки 37 моделей, обсудили проблему оценки и сравнения guardrail моделей, рассказали, как лидерборд помогает с ней справиться, а также поделились инсайтами и руководством по использованию. Теперь пришло время заглянуть во внутреннюю кухню: в этой статье мы подробно разберем устройство GuardRate Tool и GuardRate Leaderboard, ответим на вопрос, почему выбрали именно такие бенчмарки и метрики, опишем методологию и расскажем о том, как проводились эксперименты.

    habr.com/ru/articles/1081286/

    #opensource #leaderboard #benchmarking #evaluation #llm #guardrail_metrics #ai_safety #guardrails #promptinjection #guardrail_areana

  28. GuardRate: метрики, бенчмарки и инсайты из первого аудита (часть 2)

    В первой части мы верхнеуровнево рассмотрели результаты оценки 37 моделей, обсудили проблему оценки и сравнения guardrail моделей, рассказали, как лидерборд помогает с ней справиться, а также поделились инсайтами и руководством по использованию. Теперь пришло время заглянуть во внутреннюю кухню: в этой статье мы подробно разберем устройство GuardRate Tool и GuardRate Leaderboard, ответим на вопрос, почему выбрали именно такие бенчмарки и метрики, опишем методологию и расскажем о том, как проводились эксперименты.

    habr.com/ru/articles/1081286/

    #opensource #leaderboard #benchmarking #evaluation #llm #guardrail_metrics #ai_safety #guardrails #promptinjection #guardrail_areana

  29. Локальный файрвол действий для кодинг‑агентов: связать то, что агент прочитал, с тем, что он собирается выполнить

    Кодинг‑агент выполняет то, что читает — README, вывод MCP‑инструмента, результат команды. Если там спрятана инструкция, это превращается в реальное действие: curl | sh, утечка секретов, push во внешний репозиторий. Stroq — открытый локальный файрвол, который связывает прочитанное с тем, что агент собирается сделать, и детерминированно блокирует опасное действие до того, как оно случится. Ни облака, ни расчёта на то, что модель сама заметит инъекцию.

    habr.com/ru/articles/1081108/

    #aiagents #llmsecurity #promptinjection #opensource