home.social

#prompt-injection — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #prompt-injection, aggregated by home.social.

fetched live
  1. Claude Code взломали просьбой пересказать сайт

    Представь простую задачу: ты просишь агента пересказать содержимое сайта. Не установить пакет. Не запустить скачанный скрипт. Не открыть подозрительное вложение. Просто прочитать страницу и сделать краткое резюме. Через несколько шагов на твоей машине выполняется код атакующего, устанавливается соединение с управляющим сервером и для наглядности открывается калькулятор. И самое интересное: агент не нарушал прямой запрет. Наоборот, в ключевой момент он поступил "безопасно" — отказался запускать неизвестный бинарный файл и написал собственный декодер на Python. Именно это решение и стало частью атаки. Такую цепочку продемонстрировал исследователь безопасности Иоганн Ребергер на Claude Code с моделью Opus 5 в Auto Mode. В его небольших сериях тестов атака срабатывала в трёх или четырёх запусках из пяти. Это не означает, что "Claude взламывается с вероятностью 80%", но хорошо показывает более важную проблему: безопасный на вид отдельный шаг ничего не гарантирует, если вся цепочка строится в среде, которую контролирует атакующий. И чего он там наломал?

    habr.com/ru/articles/1078702/

    #ai #agents #claudecode #cybersecurity #promptinjection #agentic #agentic_ai #agentic_coding #agentic_engineering #aircrack

  2. Claude Code взломали просьбой пересказать сайт

    Представь простую задачу: ты просишь агента пересказать содержимое сайта. Не установить пакет. Не запустить скачанный скрипт. Не открыть подозрительное вложение. Просто прочитать страницу и сделать краткое резюме. Через несколько шагов на твоей машине выполняется код атакующего, устанавливается соединение с управляющим сервером и для наглядности открывается калькулятор. И самое интересное: агент не нарушал прямой запрет. Наоборот, в ключевой момент он поступил "безопасно" — отказался запускать неизвестный бинарный файл и написал собственный декодер на Python. Именно это решение и стало частью атаки. Такую цепочку продемонстрировал исследователь безопасности Иоганн Ребергер на Claude Code с моделью Opus 5 в Auto Mode. В его небольших сериях тестов атака срабатывала в трёх или четырёх запусках из пяти. Это не означает, что "Claude взламывается с вероятностью 80%", но хорошо показывает более важную проблему: безопасный на вид отдельный шаг ничего не гарантирует, если вся цепочка строится в среде, которую контролирует атакующий. И чего он там наломал?

    habr.com/ru/articles/1078702/

    #ai #agents #claudecode #cybersecurity #promptinjection #agentic #agentic_ai #agentic_coding #agentic_engineering #aircrack

  3. Claude Code взломали просьбой пересказать сайт

    Представь простую задачу: ты просишь агента пересказать содержимое сайта. Не установить пакет. Не запустить скачанный скрипт. Не открыть подозрительное вложение. Просто прочитать страницу и сделать краткое резюме. Через несколько шагов на твоей машине выполняется код атакующего, устанавливается соединение с управляющим сервером и для наглядности открывается калькулятор. И самое интересное: агент не нарушал прямой запрет. Наоборот, в ключевой момент он поступил "безопасно" — отказался запускать неизвестный бинарный файл и написал собственный декодер на Python. Именно это решение и стало частью атаки. Такую цепочку продемонстрировал исследователь безопасности Иоганн Ребергер на Claude Code с моделью Opus 5 в Auto Mode. В его небольших сериях тестов атака срабатывала в трёх или четырёх запусках из пяти. Это не означает, что "Claude взламывается с вероятностью 80%", но хорошо показывает более важную проблему: безопасный на вид отдельный шаг ничего не гарантирует, если вся цепочка строится в среде, которую контролирует атакующий. И чего он там наломал?

    habr.com/ru/articles/1078702/

    #ai #agents #claudecode #cybersecurity #promptinjection #agentic #agentic_ai #agentic_coding #agentic_engineering #aircrack

  4. A lawyer tried to "hack" the legal system by hiding a secret prompt injection in a court filing, instructing any AI reading the document to side with them.

    The attack was spotted by a human. This is a warning, that if AI is going to be used in court, then there must be ways to ensure that these types of attacks can be mitigated effectively.

    Source: 404media.co/person-hides-promp

    #Cybersecurity #Security #Infosec #AI #PromptInjection

  5. A lawyer tried to "hack" the legal system by hiding a secret prompt injection in a court filing, instructing any AI reading the document to side with them.

    The attack was spotted by a human. This is a warning, that if AI is going to be used in court, then there must be ways to ensure that these types of attacks can be mitigated effectively.

    Source: 404media.co/person-hides-promp

    #Cybersecurity #Security #Infosec #AI #PromptInjection

  6. A lawyer tried to "hack" the legal system by hiding a secret prompt injection in a court filing, instructing any AI reading the document to side with them.

    The attack was spotted by a human. This is a warning, that if AI is going to be used in court, then there must be ways to ensure that these types of attacks can be mitigated effectively.

    Source: 404media.co/person-hides-promp

    #Cybersecurity #Security #Infosec #AI #PromptInjection

  7. A lawyer tried to "hack" the legal system by hiding a secret prompt injection in a court filing, instructing any AI reading the document to side with them.

    The attack was spotted by a human. This is a warning, that if AI is going to be used in court, then there must be ways to ensure that these types of attacks can be mitigated effectively.

    Source: 404media.co/person-hides-promp

    #Cybersecurity #Security #Infosec #AI #PromptInjection

  8. A lawyer tried to "hack" the legal system by hiding a secret prompt injection in a court filing, instructing any AI reading the document to side with them.

    The attack was spotted by a human. This is a warning, that if AI is going to be used in court, then there must be ways to ensure that these types of attacks can be mitigated effectively.

    Source: 404media.co/person-hides-promp

    #Cybersecurity #Security #Infosec #AI #PromptInjection

  9. GuardBreaker: UAC-0099 inganna gli scanner IA con una frase sull’arma nucleare

    Il gruppo di cyberspionaggio UAC-0099 nasconde nei commenti dei propri script una richiesta su armi nucleari per far scattare i failsafe etici dei modelli linguistici usati nella scansione automatica del codice malevolo. ESET Research battezza la tecnica GuardBreaker, mentre una campagna parallela (Hades) applica lo stesso trucco su centinaia di pacchetti open source.

    insicurezzadigitale.com/guardb

  10. GuardBreaker: UAC-0099 inganna gli scanner IA con una frase sull’arma nucleare

    Il gruppo di cyberspionaggio UAC-0099 nasconde nei commenti dei propri script una richiesta su armi nucleari per far scattare i failsafe etici dei modelli linguistici usati nella scansione automatica del codice malevolo. ESET Research battezza la tecnica GuardBreaker, mentre una campagna parallela (Hades) applica lo stesso trucco su centinaia di pacchetti open source.

    insicurezzadigitale.com/guardb

  11. GuardBreaker: UAC-0099 inganna gli scanner IA con una frase sull’arma nucleare

    Il gruppo di cyberspionaggio UAC-0099 nasconde nei commenti dei propri script una richiesta su armi nucleari per far scattare i failsafe etici dei modelli linguistici usati nella scansione automatica del codice malevolo. ESET Research battezza la tecnica GuardBreaker, mentre una campagna parallela (Hades) applica lo stesso trucco su centinaia di pacchetti open source.

    insicurezzadigitale.com/guardb

  12. GuardBreaker: UAC-0099 inganna gli scanner IA con una frase sull’arma nucleare

    Il gruppo di cyberspionaggio UAC-0099 nasconde nei commenti dei propri script una richiesta su armi nucleari per far scattare i failsafe etici dei modelli linguistici usati nella scansione automatica del codice malevolo. ESET Research battezza la tecnica GuardBreaker, mentre una campagna parallela (Hades) applica lo stesso trucco su centinaia di pacchetti open source.

    insicurezzadigitale.com/guardb

  13. GuardBreaker: UAC-0099 inganna gli scanner IA con una frase sull’arma nucleare

    Il gruppo di cyberspionaggio UAC-0099 nasconde nei commenti dei propri script una richiesta su armi nucleari per far scattare i failsafe etici dei modelli linguistici usati nella scansione automatica del codice malevolo. ESET Research battezza la tecnica GuardBreaker, mentre una campagna parallela (Hades) applica lo stesso trucco su centinaia di pacchetti open source.

    insicurezzadigitale.com/guardb

  14. Russia-Aligned Hackers Inject Nuclear Prompt to Evade AI Analysis

    Hackers have found a sneaky way to outsmart AI-powered security tools by injecting a provocative phrase, like a threat to create a nuclear weapon, into malicious code to disable analysis. This clever trick, dubbed GuardBreaker, tricks AI scanners into failing to examine the rest of the script.

    osintsights.com/russia-aligned

    #Russia #AiAnalysisEvasion #Guardbreaker #PromptInjection #Uac0099

  15. Amazon Kiro Flaw Exposes Sensitive Data Through Prompt Injection

    A security flaw in Kiro, known as a prompt injection vulnerability, allowed hackers to tap into sensitive data by manipulating the Kiro agent with malicious repository content. This issue, affecting Kiro IDE 0.7.45 on Windows, could send local information to an external endpoint, putting users at risk.

    osintsights.com/amazon-kiro-fl

    #Kiro #PromptInjection #SensitiveDataExposure #Vulnerability #IdeSecurity

  16. Los agentes de OpenAI hackearon Hugging Face porque fueron entrenados para hacer trampa
    No es algo que se pueda resolver de la noche a la mañana, dice Kai Chen, que dirige el equipo de investigación de alineación de OpenAI. "Hay desafíos que hemos estado siguiendo durante mucho tiempo, y ahora los estamos viendo con mucha mayor precisión.
    Leer entera:laautopsia.com/noticia/agentes
    #LaAutopsia #seguridadia #Ciberseguridad #LLM #PromptInjection #vulnerabilidades #IA #apisdom

  17. Los agentes de OpenAI hackearon Hugging Face porque fueron entrenados para hacer trampa
    No es algo que se pueda resolver de la noche a la mañana, dice Kai Chen, que dirige el equipo de investigación de alineación de OpenAI. "Hay desafíos que hemos estado siguiendo durante mucho tiempo, y ahora los estamos viendo con mucha mayor precisión.
    Leer entera:laautopsia.com/noticia/agentes
    #LaAutopsia #seguridadia #Ciberseguridad #LLM #PromptInjection #vulnerabilidades #IA #apisdom

  18. Los agentes de OpenAI hackearon Hugging Face porque fueron entrenados para hacer trampa
    No es algo que se pueda resolver de la noche a la mañana, dice Kai Chen, que dirige el equipo de investigación de alineación de OpenAI. "Hay desafíos que hemos estado siguiendo durante mucho tiempo, y ahora los estamos viendo con mucha mayor precisión.
    Leer entera:laautopsia.com/noticia/agentes
    #LaAutopsia #seguridadia #Ciberseguridad #LLM #PromptInjection #vulnerabilidades #IA #apisdom

  19. Los agentes de OpenAI hackearon Hugging Face porque fueron entrenados para hacer trampa
    No es algo que se pueda resolver de la noche a la mañana, dice Kai Chen, que dirige el equipo de investigación de alineación de OpenAI. "Hay desafíos que hemos estado siguiendo durante mucho tiempo, y ahora los estamos viendo con mucha mayor precisión.
    Leer entera:laautopsia.com/noticia/agentes
    #LaAutopsia #seguridadia #Ciberseguridad #LLM #PromptInjection #vulnerabilidades #IA #apisdom

  20. Los agentes de OpenAI hackearon Hugging Face porque fueron entrenados para hacer trampa
    No es algo que se pueda resolver de la noche a la mañana, dice Kai Chen, que dirige el equipo de investigación de alineación de OpenAI. "Hay desafíos que hemos estado siguiendo durante mucho tiempo, y ahora los estamos viendo con mucha mayor precisión.
    Leer entera:laautopsia.com/noticia/agentes
    #LaAutopsia #seguridadia #Ciberseguridad #LLM #PromptInjection #vulnerabilidades #IA #apisdom

  21. A Theory of Prompt Injection (and why you should study roles).

    This is a blog-style writeup of a paper.

    We show prompt injections are driven by a flaw in how LLMs perceive roles.

    This lets us create new attacks, explain mech interp results, and predict when attacks succeed.

    We then discuss what roles are and why they matter, and share research ideas for a science of roles.

    role-confusion.github.io/

    #AI #LLM #PromptInjection #Roles

  22. A Theory of Prompt Injection (and why you should study roles).

    This is a blog-style writeup of a paper.

    We show prompt injections are driven by a flaw in how LLMs perceive roles.

    This lets us create new attacks, explain mech interp results, and predict when attacks succeed.

    We then discuss what roles are and why they matter, and share research ideas for a science of roles.

    role-confusion.github.io/

    #AI #LLM #PromptInjection #Roles

  23. A Theory of Prompt Injection (and why you should study roles).

    This is a blog-style writeup of a paper.

    We show prompt injections are driven by a flaw in how LLMs perceive roles.

    This lets us create new attacks, explain mech interp results, and predict when attacks succeed.

    We then discuss what roles are and why they matter, and share research ideas for a science of roles.

    role-confusion.github.io/

    #AI #LLM #PromptInjection #Roles

  24. A Theory of Prompt Injection (and why you should study roles).

    This is a blog-style writeup of a paper.

    We show prompt injections are driven by a flaw in how LLMs perceive roles.

    This lets us create new attacks, explain mech interp results, and predict when attacks succeed.

    We then discuss what roles are and why they matter, and share research ideas for a science of roles.

    role-confusion.github.io/

    #AI #LLM #PromptInjection #Roles

  25. A Theory of Prompt Injection (and why you should study roles).

    This is a blog-style writeup of a paper.

    We show prompt injections are driven by a flaw in how LLMs perceive roles.

    This lets us create new attacks, explain mech interp results, and predict when attacks succeed.

    We then discuss what roles are and why they matter, and share research ideas for a science of roles.

    role-confusion.github.io/

    #AI #LLM #PromptInjection #Roles

  26. Adversa AI disclosed Cryptographic Context Injection, a zero-click attack on Grok that embeds malicious instructions in AES-256-GCM payloads. The model decrypts them in its own runtime, bypassing safety filters and exfiltrating the full chat history.

    #CryptographicContextInjection #LLMSecurity #PromptInjection #AdversaAI

    cyberworldops.eu/en/cryptograp

  27. Adversa AI disclosed Cryptographic Context Injection, a zero-click attack on Grok that embeds malicious instructions in AES-256-GCM payloads. The model decrypts them in its own runtime, bypassing safety filters and exfiltrating the full chat history.

    #CryptographicContextInjection #LLMSecurity #PromptInjection #AdversaAI

    cyberworldops.eu/en/cryptograp

  28. Zero-Click Attack Steals Grok Chat History via Encrypted Payloads

    Researchers demonstrate a new technique that bypasses AI safety guardrails using AES-encrypted instructions, allowing full chat history theft from xAI's Grok with no user interaction.

    pulseofnations.lol/zero-click-

    #Adversa #AiSecurity #Cryptography #DataTheft #Grok #PromptInjection #XAI

  29. How does a prompt injection become a worm? En Klype Salt reports that instructions hidden in white text in a Word file survive a Copilot for Word session. Copilot follows them without saying so and copies them into the document it drafts, which then infects the next session. It spreads only when people reuse each other's documents, slower than a self-moving worm and harder to spot, since every step looks like ordinary work.

    benjaminhan.net/posts/20260822

    #PromptInjection #Security #Microsoft #AI

  30. How does a prompt injection become a worm? En Klype Salt reports that instructions hidden in white text in a Word file survive a Copilot for Word session. Copilot follows them without saying so and copies them into the document it drafts, which then infects the next session. It spreads only when people reuse each other's documents, slower than a self-moving worm and harder to spot, since every step looks like ordinary work.

    benjaminhan.net/posts/20260822

    #PromptInjection #Security #Microsoft #AI

  31. How does a prompt injection become a worm? En Klype Salt reports that instructions hidden in white text in a Word file survive a Copilot for Word session. Copilot follows them without saying so and copies them into the document it drafts, which then infects the next session. It spreads only when people reuse each other's documents, slower than a self-moving worm and harder to spot, since every step looks like ordinary work.

    benjaminhan.net/posts/20260822

    #PromptInjection #Security #Microsoft #AI

  32. How does a prompt injection become a worm? En Klype Salt reports that instructions hidden in white text in a Word file survive a Copilot for Word session. Copilot follows them without saying so and copies them into the document it drafts, which then infects the next session. It spreads only when people reuse each other's documents, slower than a self-moving worm and harder to spot, since every step looks like ordinary work.

    benjaminhan.net/posts/20260822

    #PromptInjection #Security #Microsoft #AI

  33. How does a prompt injection become a worm? En Klype Salt reports that instructions hidden in white text in a Word file survive a Copilot for Word session. Copilot follows them without saying so and copies them into the document it drafts, which then infects the next session. It spreads only when people reuse each other's documents, slower than a self-moving worm and harder to spot, since every step looks like ordinary work.

    benjaminhan.net/posts/20260822

    #PromptInjection #Security #Microsoft #AI

  34. Grok AI Chatbot Tricked Into Leaking Private Chats Through Encrypted Prompt Injection

    Security researchers at Adversa AI found a zero-click flaw in xAI's Grok that hides malicious instructions inside encrypted text to steal names, locations, and chat history. The attack needs no clicks from the victim and exposes a broader weakness in how AI agents handle untrusted content.

    securebulletin.com/grok-ai-cha

  35. Grok AI Chatbot Tricked Into Leaking Private Chats Through Encrypted Prompt Injection

    Security researchers at Adversa AI found a zero-click flaw in xAI's Grok that hides malicious instructions inside encrypted text to steal names, locations, and chat history. The attack needs no clicks from the victim and exposes a broader weakness in how AI agents handle untrusted content.

    securebulletin.com/grok-ai-cha

  36. Grok AI Chatbot Tricked Into Leaking Private Chats Through Encrypted Prompt Injection

    Security researchers at Adversa AI found a zero-click flaw in xAI's Grok that hides malicious instructions inside encrypted text to steal names, locations, and chat history. The attack needs no clicks from the victim and exposes a broader weakness in how AI agents handle untrusted content.

    securebulletin.com/grok-ai-cha