#prompt-injection — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #prompt-injection, aggregated by home.social.
-
Discover how self-replicating prompt injection worms force AI agents to propagate malicious instructions via emails and code, highlighting new cybersecurity risks.
#PromptInjection #AIAgents #Cybersecurity #OpenAI #ComputerWorm
-
Discover how self-replicating prompt injection worms force AI agents to propagate malicious instructions via emails and code, highlighting new cybersecurity risks.
#PromptInjection #AIAgents #Cybersecurity #OpenAI #ComputerWorm
-
Discover how self-replicating prompt injection worms force AI agents to propagate malicious instructions via emails and code, highlighting new cybersecurity risks.
#PromptInjection #AIAgents #Cybersecurity #OpenAI #ComputerWorm
-
Anyone interested in AI safety, adversarial testing of agentic systems, putting Zuck on blast, or having a rolling good time on infosec matters should follow @jonny , and turn on notifications, because they're live-tooting their adversarial testing of Meta's Muse with methods and results both hilarious.
#Muse #Meta #AI #AISafety #InfoSec #PromptInjection #AIAgent
-
Anyone interested in AI safety, adversarial testing of agentic systems, putting Zuck on blast, or having a rolling good time on infosec matters should follow @jonny , and turn on notifications, because they're live-tooting their adversarial testing of Meta's Muse with methods and results both hilarious.
#Muse #Meta #AI #AISafety #InfoSec #PromptInjection #AIAgent
-
Anyone interested in AI safety, adversarial testing of agentic systems, putting Zuck on blast, or having a rolling good time on infosec matters should follow @jonny , and turn on notifications, because they're live-tooting their adversarial testing of Meta's Muse with methods and results both hilarious.
#Muse #Meta #AI #AISafety #InfoSec #PromptInjection #AIAgent
-
Anyone interested in AI safety, adversarial testing of agentic systems, putting Zuck on blast, or having a rolling good time on infosec matters should follow @jonny , and turn on notifications, because they're live-tooting their adversarial testing of Meta's Muse with methods and results both hilarious.
#Muse #Meta #AI #AISafety #InfoSec #PromptInjection #AIAgent
-
A URL redactor only works if it agrees with the renderer on where a URL ends. SalesBleed (Zenity Labs): a Web-to-Lead injection fired when Agentforce summarized the lead, emitting an img tag to attacker.oast.fun/{data}. The Trusted URLs filter had no .fun in its TLD list and stopped at braces; the chat surface rendered it. The data left in the DNS lookup.
https://labs.zenity.io/post/salesbleed-0-click-data-exfiltration-on-agentforce
#AIAgents #Cybersecurity #PromptInjection -
"Dark Sourcery" avrebbe manipolato le risposte di chatbot AI per colpire 374 aziende. L'angolo interessante non è il numero, ma il vettore: i LLM come superficie d'attacco indiretta, via prompt injection o poisoning. La fiducia implicita nelle risposte dei chatbot enterprise è esattamente il punto debole che questi attacchi sfruttano. #infosec #AIsecurity #promptinjection
https://www.ilsoftware.it/dark-sourcery-colpisce-374-aziende-manipolando-le-risposte-dei-chatbot/ -
Keep this one short and attach the generated image.
AI agents don't need every key.
As agents gain access to APIs, databases and tools, least privilege becomes critical.
If an agent is manipulated by prompt injection, limited permissions can limit the damage.
Capability ≠ authorization.
Give an AI only the access it actually needs.
#Cybersecurity #AISecurity #AgenticAI #PromptInjection #InfoSec
-
https://www.europesays.com/afrika/55161/ Die Madagaskar-Falle: Wenn KI auf Befehle hört, die Menschen nicht sehen #ChatGPT #GenerativeKI #Hochschule #JasonGibson #KIAgenten #KISicherheit #KünstlicheIntelligenz #Madagascar #Madagaskar #MicrosoftCopilot #PromptInjection #Studium
-
Zenity Labs found three Agentforce flaws, dubbed SalesBleed, allowing prompt injection via Web-to-Lead forms. A poisoned lead could trigger zero-click CRM exfiltration when processed by an AI agent, bypassing user interaction. It shows untrusted CRM inputs must be treated as active attack surface. #SalesBleed #Salesforce #AgentSecurity #PromptInjection
https://cyberworldops.eu/en/salesbleed-turned-salesforce-lead-forms-into-a-path-for-silent-crm
-
Zenity Labs found three Agentforce flaws, dubbed SalesBleed, allowing prompt injection via Web-to-Lead forms. A poisoned lead could trigger zero-click CRM exfiltration when processed by an AI agent, bypassing user interaction. It shows untrusted CRM inputs must be treated as active attack surface. #SalesBleed #Salesforce #AgentSecurity #PromptInjection
https://cyberworldops.eu/en/salesbleed-turned-salesforce-lead-forms-into-a-path-for-silent-crm
-
@bsi has developed a number of basic measures to combat prompt injections in document-based LLM workflows: https://github.com/BSI-Bund/baseline_defense_lab_indirectPromptInjections
#LLM #PromptInjection #ITSecurity -
@bsi has developed a number of basic measures to combat prompt injections in document-based LLM workflows: https://github.com/BSI-Bund/baseline_defense_lab_indirectPromptInjections
#LLM #PromptInjection #ITSecurity -
@bsi has developed a number of basic measures to combat prompt injections in document-based LLM workflows: https://github.com/BSI-Bund/baseline_defense_lab_indirectPromptInjections
#LLM #PromptInjection #ITSecurity -
@bsi has developed a number of basic measures to combat prompt injections in document-based LLM workflows: https://github.com/BSI-Bund/baseline_defense_lab_indirectPromptInjections
#LLM #PromptInjection #ITSecurity -
Jev Is Not a Language Model, but It Breaks Like One
Comments: https://news.ycombinator.com/item?id=49830172
#HackerNews #JevModel #AI #Security #PromptInjection #LanguageModel
-
Jev Is Not a Language Model, but It Breaks Like One
Comments: https://news.ycombinator.com/item?id=49830172
#HackerNews #JevModel #AI #Security #PromptInjection #LanguageModel
-
Jev Is Not a Language Model, but It Breaks Like One
Comments: https://news.ycombinator.com/item?id=49830172
#HackerNews #JevModel #AI #Security #PromptInjection #LanguageModel
-
Jev Is Not a Language Model, but It Breaks Like One
Comments: https://news.ycombinator.com/item?id=49830172
#HackerNews #JevModel #AI #Security #PromptInjection #LanguageModel
-
A Meta Muse zero-day let local processes hijack dictation and steal auth tokens on macOS. Patrick Wardle's PoC prompted an emergency Meta fix.
#MetaMuse #ZeroDay #macOS #PatrickWardle #AIagents #PromptInjection #CyberSecurity
https://meterpreter.org/meta-muse-macos-zero-day/?utm_source=mastodon&utm_medium=jetpack_social
-
A Meta Muse zero-day let local processes hijack dictation and steal auth tokens on macOS. Patrick Wardle's PoC prompted an emergency Meta fix.
#MetaMuse #ZeroDay #macOS #PatrickWardle #AIagents #PromptInjection #CyberSecurity
https://meterpreter.org/meta-muse-macos-zero-day/?utm_source=mastodon&utm_medium=jetpack_social
-
A Meta Muse zero-day let local processes hijack dictation and steal auth tokens on macOS. Patrick Wardle's PoC prompted an emergency Meta fix.
#MetaMuse #ZeroDay #macOS #PatrickWardle #AIagents #PromptInjection #CyberSecurity
https://meterpreter.org/meta-muse-macos-zero-day/?utm_source=mastodon&utm_medium=jetpack_social
-
AI Apocalypse Claims Under Scrutiny: Follow the Hardware, Permissions, and Money
Follow the money behind AI fear: this analysis examines whether dramatic warnings can shape rules, raise barriers for rivals, and shift the cost of failure onto the public.
#AIApocalypse #AIHype #AILiability #AIRegulation #AISafety #AndrewYang #Anthropic #ArtificialIntelligence #autonomousAgents #BigTech #corporateAccountability #criticalInfrastructure #Cybersecurity #NIST #OpenAI #promptInjection #publicInterest #regulatoryCapture #responsibleTechnology #softwareEngineering https://wp.me/p1OjMZ-pcQ -
AI Apocalypse Claims Under Scrutiny: Follow the Hardware, Permissions, and Money
Follow the money behind AI fear: this analysis examines whether dramatic warnings can shape rules, raise barriers for rivals, and shift the cost of failure onto the public.
#AIApocalypse #AIHype #AILiability #AIRegulation #AISafety #AndrewYang #Anthropic #ArtificialIntelligence #autonomousAgents #BigTech #corporateAccountability #criticalInfrastructure #Cybersecurity #NIST #OpenAI #promptInjection #publicInterest #regulatoryCapture #responsibleTechnology #softwareEngineering https://wp.me/p1OjMZ-pcQ -
AI Apocalypse Claims Under Scrutiny: Follow the Hardware, Permissions, and Money
Follow the money behind AI fear: this analysis examines whether dramatic warnings can shape rules, raise barriers for rivals, and shift the cost of failure onto the public.
#AIApocalypse #AIHype #AILiability #AIRegulation #AISafety #AndrewYang #Anthropic #ArtificialIntelligence #autonomousAgents #BigTech #corporateAccountability #criticalInfrastructure #Cybersecurity #NIST #OpenAI #promptInjection #publicInterest #regulatoryCapture #responsibleTechnology #softwareEngineering https://wp.me/p1OjMZ-pcQ -
AI Apocalypse Claims Under Scrutiny: Follow the Hardware, Permissions, and Money
Follow the money behind AI fear: this analysis examines whether dramatic warnings can shape rules, raise barriers for rivals, and shift the cost of failure onto the public.
#AIApocalypse #AIHype #AILiability #AIRegulation #AISafety #AndrewYang #Anthropic #ArtificialIntelligence #autonomousAgents #BigTech #corporateAccountability #criticalInfrastructure #Cybersecurity #NIST #OpenAI #promptInjection #publicInterest #regulatoryCapture #responsibleTechnology #softwareEngineering https://wp.me/p1OjMZ-pcQ -
📬 KI-Prompt-Injection: OpenAI-Modelle schreiben ihre eigenen Jailbreaks
#Jailbreaks #KünstlicheIntelligenz #AstraModelle #CompactionSummary #Jailbreak #KIModelle #KIPromptInjection #KISicherheit #KITraining #Misalignment #OpenAI #PromptInjection https://sc.tarnkappe.info/e69143 -
📬 KI-Prompt-Injection: OpenAI-Modelle schreiben ihre eigenen Jailbreaks
#Jailbreaks #KünstlicheIntelligenz #AstraModelle #CompactionSummary #Jailbreak #KIModelle #KIPromptInjection #KISicherheit #KITraining #Misalignment #OpenAI #PromptInjection https://sc.tarnkappe.info/e69143 -
📬 KI-Prompt-Injection: OpenAI-Modelle schreiben ihre eigenen Jailbreaks
#Jailbreaks #KünstlicheIntelligenz #AstraModelle #CompactionSummary #Jailbreak #KIModelle #KIPromptInjection #KISicherheit #KITraining #Misalignment #OpenAI #PromptInjection https://sc.tarnkappe.info/e69143 -
📬 KI-Prompt-Injection: OpenAI-Modelle schreiben ihre eigenen Jailbreaks
#Jailbreaks #KünstlicheIntelligenz #AstraModelle #CompactionSummary #Jailbreak #KIModelle #KIPromptInjection #KISicherheit #KITraining #Misalignment #OpenAI #PromptInjection https://sc.tarnkappe.info/e69143 -
Agent memory poisoning needs no injection. A new paper poisoned 1.2% of a LongMemEval corpus with plainly worded false statements, no instructions or triggers, and accuracy fell from 0.850 to 0.300. A write-time screening pipeline with 0.832 recall on indirect prompt injection rejected 0 of 360 poisoned memories. Filters catch text that gives orders. A false claim reads just like a true one.
-
FYI: Explaining prompt injection: Prompt injection is an attack where text a model reads as data gets obeyed as instruction. How it works, why no fix exists, and what it means for advertising. https://ppc.land/prompt-injection/ #PromptInjection #AI #MachineLearning #CyberSecurity #DataPrivacy
-
FYI: Explaining prompt injection: Prompt injection is an attack where text a model reads as data gets obeyed as instruction. How it works, why no fix exists, and what it means for advertising. https://ppc.land/prompt-injection/ #PromptInjection #AI #MachineLearning #CyberSecurity #DataPrivacy
-
FYI: Explaining prompt injection: Prompt injection is an attack where text a model reads as data gets obeyed as instruction. How it works, why no fix exists, and what it means for advertising. https://ppc.land/prompt-injection/ #PromptInjection #AI #MachineLearning #CyberSecurity #DataPrivacy
-
FYI: Explaining prompt injection: Prompt injection is an attack where text a model reads as data gets obeyed as instruction. How it works, why no fix exists, and what it means for advertising. https://ppc.land/prompt-injection/ #PromptInjection #AI #MachineLearning #CyberSecurity #DataPrivacy
-
ICYMI: Explaining prompt injection: Prompt injection is an attack where text a model reads as data gets obeyed as instruction. How it works, why no fix exists, and what it means for advertising. https://ppc.land/prompt-injection/ #PromptInjection #DigitalMarketing #AdTech #CyberSecurity #AI
-
ICYMI: Explaining prompt injection: Prompt injection is an attack where text a model reads as data gets obeyed as instruction. How it works, why no fix exists, and what it means for advertising. https://ppc.land/prompt-injection/ #PromptInjection #DigitalMarketing #AdTech #CyberSecurity #AI
-
ICYMI: Explaining prompt injection: Prompt injection is an attack where text a model reads as data gets obeyed as instruction. How it works, why no fix exists, and what it means for advertising. https://ppc.land/prompt-injection/ #PromptInjection #DigitalMarketing #AdTech #CyberSecurity #AI
-
ICYMI: Explaining prompt injection: Prompt injection is an attack where text a model reads as data gets obeyed as instruction. How it works, why no fix exists, and what it means for advertising. https://ppc.land/prompt-injection/ #PromptInjection #DigitalMarketing #AdTech #CyberSecurity #AI
-
The first mistake is to treat "the LLM" as a single agent with a goal.
A real LLM deployment is is a pipeline made up of, for example, a model that generates text, a tool-calling layer that turns that text into actions, memory and state, and, if you are lucky, a detection and control layer on top.
You are not talking to "the AI". You are interacting with a pipeline of discrete components, and every seam between them is a place where you can add a control.
The concrete risk today is not "the model rebels". It is things like prompt injection and confused deputy: untrusted data enters the context and can be treated as instructions because, in many systems, data and instructions are not properly separated.
If a system can be compromised through prompt injection today, and the level of mitigation can be measured and improved over time, then the problem has the characteristics of an engineering problem: it is not solved, but it can be addressed with evidence.
That is the point: putting "prompt injection in a production agent" and "existential risk from AGI" in the same bucket does not make the argument deeper. They involve different problems, assumptions, and standards of evidence. Treating them as the same is simply a category error.
And that is exactly the rhetorical trick that, in my view, should be avoided.
-
The first mistake is to treat "the LLM" as a single agent with a goal.
A real LLM deployment is a pipeline made up of, for example, a model that generates text, a tool-calling layer that turns that text into actions, memory and state, and, if you are lucky, a detection and control layer on top.
You are not talking to "the AI". You are interacting with a pipeline of discrete components, and every seam between them is a place where you can add a control.
The concrete risk today is not "the model rebels". It is things like prompt injection and confused deputy: untrusted data enters the context and can be treated as instructions because, in many systems, data and instructions are not properly separated.
If a system can be compromised through prompt injection today, and the level of mitigation can be measured and improved over time, then the problem has the characteristics of an engineering problem: it is not solved, but it can be addressed with evidence.
That is the point: putting "prompt injection in a production agent" and "existential risk from AGI" in the same bucket does not make the argument deeper. They involve different problems, assumptions, and standards of evidence. Treating them as the same is simply a category error.
And that is exactly the rhetorical trick that, in my view, should be avoided.
-
The first mistake is to treat "the LLM" as a single agent with a goal.
A real LLM deployment is not a monolithic block that decides and acts. It is a pipeline made up of, for example, a model that generates text, a tool-calling layer that turns that text into actions, memory and state, and, if you are lucky, a detection and control layer on top.
You are not talking to "the AI". You are interacting with a pipeline of discrete components, and every seam between them is a place where you can add a control.
The concrete risk today is not "the model rebels". It is things like prompt injection and confused deputy: untrusted data enters the context and can be treated as instructions because, in many systems, data and instructions are not properly separated.
If a system can be compromised through prompt injection today, and the level of mitigation can be measured and improved over time, then the problem has the characteristics of an engineering problem: it is not solved, but it can be addressed with evidence.
That is the point: putting "prompt injection in a production agent" and "existential risk from AGI" in the same bucket does not make the argument deeper. They involve different problems, assumptions, and standards of evidence. Treating them as the same is simply a category error.
And that is exactly the rhetorical trick that, in my view, should be avoided.
-
The first mistake is to treat "the LLM" as a single agent with a goal.
A real LLM deployment is is a pipeline made up of, for example, a model that generates text, a tool-calling layer that turns that text into actions, memory and state, and, if you are lucky, a detection and control layer on top.
You are not talking to "the AI". You are interacting with a pipeline of discrete components, and every seam between them is a place where you can add a control.
The concrete risk today is not "the model rebels". It is things like prompt injection and confused deputy: untrusted data enters the context and can be treated as instructions because, in many systems, data and instructions are not properly separated.
If a system can be compromised through prompt injection today, and the level of mitigation can be measured and improved over time, then the problem has the characteristics of an engineering problem: it is not solved, but it can be addressed with evidence.
That is the point: putting "prompt injection in a production agent" and "existential risk from AGI" in the same bucket does not make the argument deeper. They involve different problems, assumptions, and standards of evidence. Treating them as the same is simply a category error.
And that is exactly the rhetorical trick that, in my view, should be avoided.
-
The PuzzleMask attack hides malicious commands in plain prose, slipping past quick LLM filters while a stronger model extracts and runs them. Details here.
#PuzzleMask #PromptInjection #AISecurity #CheckPoint #LLM #CyberSecurity #InfoSec
https://meterpreter.org/puzzlemask-ai-attack-vector/?utm_source=mastodon&utm_medium=jetpack_social
-
The PuzzleMask attack hides malicious commands in plain prose, slipping past quick LLM filters while a stronger model extracts and runs them. Details here.
#PuzzleMask #PromptInjection #AISecurity #CheckPoint #LLM #CyberSecurity #InfoSec
https://meterpreter.org/puzzlemask-ai-attack-vector/?utm_source=mastodon&utm_medium=jetpack_social
-
The PuzzleMask attack hides malicious commands in plain prose, slipping past quick LLM filters while a stronger model extracts and runs them. Details here.
#PuzzleMask #PromptInjection #AISecurity #CheckPoint #LLM #CyberSecurity #InfoSec
https://meterpreter.org/puzzlemask-ai-attack-vector/?utm_source=mastodon&utm_medium=jetpack_social
-
The PuzzleMask attack hides malicious commands in plain prose, slipping past quick LLM filters while a stronger model extracts and runs them. Details here.
#PuzzleMask #PromptInjection #AISecurity #CheckPoint #LLM #CyberSecurity #InfoSec
https://meterpreter.org/puzzlemask-ai-attack-vector/?utm_source=mastodon&utm_medium=jetpack_social
-
Researchers revealed the Puzzlemask AI attack, showing how this Puzzlemask AI attack vector bypasses quick LLM security filters using plain prose.
#Puzzlemask #AISecurity #PromptInjection #LLMSecurity #CheckPointResearch
-
Researchers revealed the Puzzlemask AI attack, showing how this Puzzlemask AI attack vector bypasses quick LLM security filters using plain prose.
#Puzzlemask #AISecurity #PromptInjection #LLMSecurity #CheckPointResearch
-
Researchers revealed the Puzzlemask AI attack, showing how this Puzzlemask AI attack vector bypasses quick LLM security filters using plain prose.
#Puzzlemask #AISecurity #PromptInjection #LLMSecurity #CheckPointResearch
-
Explaining prompt injection: Prompt injection is an attack where text a model reads as data gets obeyed as instruction. How it works, why no fix exists, and what it means for advertising. https://ppc.land/prompt-injection/ #PromptInjection #DigitalMarketing #CyberSecurity #Advertising #AI
-
Explaining prompt injection: Prompt injection is an attack where text a model reads as data gets obeyed as instruction. How it works, why no fix exists, and what it means for advertising. https://ppc.land/prompt-injection/ #PromptInjection #DigitalMarketing #CyberSecurity #Advertising #AI
-
Explaining prompt injection: Prompt injection is an attack where text a model reads as data gets obeyed as instruction. How it works, why no fix exists, and what it means for advertising. https://ppc.land/prompt-injection/ #PromptInjection #DigitalMarketing #CyberSecurity #Advertising #AI
-
Explaining prompt injection: Prompt injection is an attack where text a model reads as data gets obeyed as instruction. How it works, why no fix exists, and what it means for advertising. https://ppc.land/prompt-injection/ #PromptInjection #DigitalMarketing #CyberSecurity #Advertising #AI
-
Hidden HTML can turn an AI email summarizer into a “quiet liar,” fabricating facts and omitting legitimate content.
https://jpmellojr.blogspot.com/2026/09/ai-summary-attack-conceals-code-that.html
#AIsecurity #PromptInjection #Forcepoint -
GuardRate: метрики, бенчмарки и инсайты из первого аудита (часть 2)
В первой части мы верхнеуровнево рассмотрели результаты оценки 37 моделей, обсудили проблему оценки и сравнения guardrail моделей, рассказали, как лидерборд помогает с ней справиться, а также поделились инсайтами и руководством по использованию. Теперь пришло время заглянуть во внутреннюю кухню: в этой статье мы подробно разберем устройство GuardRate Tool и GuardRate Leaderboard, ответим на вопрос, почему выбрали именно такие бенчмарки и метрики, опишем методологию и расскажем о том, как проводились эксперименты.
https://habr.com/ru/articles/1081286/
#opensource #leaderboard #benchmarking #evaluation #llm #guardrail_metrics #ai_safety #guardrails #promptinjection #guardrail_areana
-
GuardRate: метрики, бенчмарки и инсайты из первого аудита (часть 2)
В первой части мы верхнеуровнево рассмотрели результаты оценки 37 моделей, обсудили проблему оценки и сравнения guardrail моделей, рассказали, как лидерборд помогает с ней справиться, а также поделились инсайтами и руководством по использованию. Теперь пришло время заглянуть во внутреннюю кухню: в этой статье мы подробно разберем устройство GuardRate Tool и GuardRate Leaderboard, ответим на вопрос, почему выбрали именно такие бенчмарки и метрики, опишем методологию и расскажем о том, как проводились эксперименты.
https://habr.com/ru/articles/1081286/
#opensource #leaderboard #benchmarking #evaluation #llm #guardrail_metrics #ai_safety #guardrails #promptinjection #guardrail_areana
-
Локальный файрвол действий для кодинг‑агентов: связать то, что агент прочитал, с тем, что он собирается выполнить
Кодинг‑агент выполняет то, что читает — README, вывод MCP‑инструмента, результат команды. Если там спрятана инструкция, это превращается в реальное действие: curl | sh, утечка секретов, push во внешний репозиторий. Stroq — открытый локальный файрвол, который связывает прочитанное с тем, что агент собирается сделать, и детерминированно блокирует опасное действие до того, как оно случится. Ни облака, ни расчёта на то, что модель сама заметит инъекцию.