#llm_security — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #llm_security, aggregated by home.social.
-
Летальное трио агентской разработки: как ИИ-агенты открывают доступ к инфраструктуре и что с этим делать
Привет, Хабр! Я Денис Макрушин, работаю в Яндексе, и вместе с командой SourceCraft Security строю платформу для безопасной агентской разработки, а в свободное время ищу уязвимости в ИИ-агентах и иногда рассказываю об исследованиях в своем блоге . Чем дольше этим занимаюсь, тем лучше вижу тенденцию: индустрия обсуждает, что агенты умеют делать, но реже говорит о том, какие решения и как проще внедрять, чтобы сделать агентскую инфраструктуру безопаснее. Вместе с моими коллегами Ратмиром Самархановым и Андреем Погирейчиком мы решили проверить гипотезу: “наши ИИ-агенты в разработке могут быть скомпрометированы и существуют простые средства для контроля их безопасности”. Расскажем о первых результатах.
https://habr.com/ru/companies/oleg-bunin/articles/1071908/
#ai_security #agentic_engineering #ai_red_teaming #ИИагенты #агентская_разработка #безопасность_ИИ #безопасная_разработка #LLM_security #prompt_injection #RAG_poisoning
-
Летальное трио агентской разработки: как ИИ-агенты открывают доступ к инфраструктуре и что с этим делать
Привет, Хабр! Я Денис Макрушин, работаю в Яндексе, и вместе с командой SourceCraft Security строю платформу для безопасной агентской разработки, а в свободное время ищу уязвимости в ИИ-агентах и иногда рассказываю об исследованиях в своем блоге . Чем дольше этим занимаюсь, тем лучше вижу тенденцию: индустрия обсуждает, что агенты умеют делать, но реже говорит о том, какие решения и как проще внедрять, чтобы сделать агентскую инфраструктуру безопаснее. Вместе с моими коллегами Ратмиром Самархановым и Андреем Погирейчиком мы решили проверить гипотезу: “наши ИИ-агенты в разработке могут быть скомпрометированы и существуют простые средства для контроля их безопасности”. Расскажем о первых результатах.
https://habr.com/ru/companies/oleg-bunin/articles/1071908/
#ai_security #agentic_engineering #ai_red_teaming #ИИагенты #агентская_разработка #безопасность_ИИ #безопасная_разработка #LLM_security #prompt_injection #RAG_poisoning
-
Летальное трио агентской разработки: как ИИ-агенты открывают доступ к инфраструктуре и что с этим делать
Привет, Хабр! Я Денис Макрушин, работаю в Яндексе, и вместе с командой SourceCraft Security строю платформу для безопасной агентской разработки, а в свободное время ищу уязвимости в ИИ-агентах и иногда рассказываю об исследованиях в своем блоге . Чем дольше этим занимаюсь, тем лучше вижу тенденцию: индустрия обсуждает, что агенты умеют делать, но реже говорит о том, какие решения и как проще внедрять, чтобы сделать агентскую инфраструктуру безопаснее. Вместе с моими коллегами Ратмиром Самархановым и Андреем Погирейчиком мы решили проверить гипотезу: “наши ИИ-агенты в разработке могут быть скомпрометированы и существуют простые средства для контроля их безопасности”. Расскажем о первых результатах.
https://habr.com/ru/companies/oleg-bunin/articles/1071908/
#ai_security #agentic_engineering #ai_red_teaming #ИИагенты #агентская_разработка #безопасность_ИИ #безопасная_разработка #LLM_security #prompt_injection #RAG_poisoning
-
Самая опасная уязвимость — интеллект атакующего агента?
Сейчас уже не удивляет, что агенты работают с терминалом, БД, браузером и инфраструктурой - без прямого контроля человека. Но чем шире полномочия агента, тем вероятнее, что для достижения цели выбранный им путь обернётся инцидентом. В кибербезопасности сейчас не хватает терминологии и устоявшихся подходов к защите таких систем - особенно в тех случаях, когда агент начинает действовать непредсказуемо. Именно эта неразбериха и подтолкнула нас с автором телеграм-канала OK ML на систематизацию. Традиционные SIEM, DLP, WAF не рассчитаны на системы, умеющие адаптироваться и рассуждать. А может, это им и не нужно? В своей прошлой статье на Хабре я уже рассказывал, как с помощью ловушек сбить с толку пентест-агентов. Но ловушки - точечный инструмент. Кажется, нужно чётко понимать, как защищаться от атак – на всех этапах киллчейна. А для этого нужно охватить весь спектр угроз. Ответом на этот запрос и стала наша таксономия Autonomous Agent Defense Matrix .
https://habr.com/ru/articles/1068248/
#ai_security #llm_security #автономные_агенты #ai_agents #информационная_безопасность #threat_modeling #taxonomies #prompt_injection #goal_hijacking #mitre_attack
-
Самая опасная уязвимость — интеллект атакующего агента?
Сейчас уже не удивляет, что агенты работают с терминалом, БД, браузером и инфраструктурой - без прямого контроля человека. Но чем шире полномочия агента, тем вероятнее, что для достижения цели выбранный им путь обернётся инцидентом. В кибербезопасности сейчас не хватает терминологии и устоявшихся подходов к защите таких систем - особенно в тех случаях, когда агент начинает действовать непредсказуемо. Именно эта неразбериха и подтолкнула нас с автором телеграм-канала OK ML на систематизацию. Традиционные SIEM, DLP, WAF не рассчитаны на системы, умеющие адаптироваться и рассуждать. А может, это им и не нужно? В своей прошлой статье на Хабре я уже рассказывал, как с помощью ловушек сбить с толку пентест-агентов. Но ловушки - точечный инструмент. Кажется, нужно чётко понимать, как защищаться от атак – на всех этапах киллчейна. А для этого нужно охватить весь спектр угроз. Ответом на этот запрос и стала наша таксономия Autonomous Agent Defense Matrix .
https://habr.com/ru/articles/1068248/
#ai_security #llm_security #автономные_агенты #ai_agents #информационная_безопасность #threat_modeling #taxonomies #prompt_injection #goal_hijacking #mitre_attack
-
Самая опасная уязвимость — интеллект атакующего агента?
Сейчас уже не удивляет, что агенты работают с терминалом, БД, браузером и инфраструктурой - без прямого контроля человека. Но чем шире полномочия агента, тем вероятнее, что для достижения цели выбранный им путь обернётся инцидентом. В кибербезопасности сейчас не хватает терминологии и устоявшихся подходов к защите таких систем - особенно в тех случаях, когда агент начинает действовать непредсказуемо. Именно эта неразбериха и подтолкнула нас с автором телеграм-канала OK ML на систематизацию. Традиционные SIEM, DLP, WAF не рассчитаны на системы, умеющие адаптироваться и рассуждать. А может, это им и не нужно? В своей прошлой статье на Хабре я уже рассказывал, как с помощью ловушек сбить с толку пентест-агентов. Но ловушки - точечный инструмент. Кажется, нужно чётко понимать, как защищаться от атак – на всех этапах киллчейна. А для этого нужно охватить весь спектр угроз. Ответом на этот запрос и стала наша таксономия Autonomous Agent Defense Matrix .
https://habr.com/ru/articles/1068248/
#ai_security #llm_security #автономные_агенты #ai_agents #информационная_безопасность #threat_modeling #taxonomies #prompt_injection #goal_hijacking #mitre_attack
-
Prompt Injection Remains the Biggest LLM Security Risk Despite Few Reported Incidents - https://www.redpacketsecurity.com/prompt-injection-remains-biggest-llm-risk-despite-limited-incidents/
-
Prompt Injection Remains the Biggest LLM Security Risk Despite Few Reported Incidents - https://www.redpacketsecurity.com/prompt-injection-remains-biggest-llm-risk-despite-limited-incidents/
-
Prompt Injection Remains the Biggest LLM Security Risk Despite Few Reported Incidents - https://www.redpacketsecurity.com/prompt-injection-remains-biggest-llm-risk-despite-limited-incidents/
-
Prompt Injection Remains the Biggest LLM Security Risk Despite Few Reported Incidents - https://www.redpacketsecurity.com/prompt-injection-remains-biggest-llm-risk-despite-limited-incidents/
-
Prompt Injection Remains the Biggest LLM Security Risk Despite Few Reported Incidents - https://www.redpacketsecurity.com/prompt-injection-remains-biggest-llm-risk-despite-limited-incidents/
-
Атака на LLM, которую нельзя исправить патчем
Что вершит судьбу LLM в этом мире. Некая незримая инструкция или закон, подобно промту Господнему, парящим над миром? По крайне мере истинно то, что LLM не властен даже над своей волей. В 2025 году главной угрозой кибербезопасности по версии OWASP стал не вирус и не баг, а обычный человеческий язык. Многие думают, что проблему можно решить, просто запретить модели нарушать правила. Но это так не работает. Promt Injection нельзя просто выключить, потому что так вы все сломаете. Как работает главная уязвимость LLM? Почему с ней так тяжело бороться? И что все-таки делать, чтобы защититься от нее?
https://habr.com/ru/companies/bothub/articles/1060310/
#llm #promt_injection #ai #ai_security #promt #owasp #llm_security #guardrails #ии_угроза #bothub
-
Атака на LLM, которую нельзя исправить патчем
Что вершит судьбу LLM в этом мире. Некая незримая инструкция или закон, подобно промту Господнему, парящим над миром? По крайне мере истинно то, что LLM не властен даже над своей волей. В 2025 году главной угрозой кибербезопасности по версии OWASP стал не вирус и не баг, а обычный человеческий язык. Многие думают, что проблему можно решить, просто запретить модели нарушать правила. Но это так не работает. Promt Injection нельзя просто выключить, потому что так вы все сломаете. Как работает главная уязвимость LLM? Почему с ней так тяжело бороться? И что все-таки делать, чтобы защититься от нее?
https://habr.com/ru/companies/bothub/articles/1060310/
#llm #promt_injection #ai #ai_security #promt #owasp #llm_security #guardrails #ии_угроза #bothub
-
Атака на LLM, которую нельзя исправить патчем
Что вершит судьбу LLM в этом мире. Некая незримая инструкция или закон, подобно промту Господнему, парящим над миром? По крайне мере истинно то, что LLM не властен даже над своей волей. В 2025 году главной угрозой кибербезопасности по версии OWASP стал не вирус и не баг, а обычный человеческий язык. Многие думают, что проблему можно решить, просто запретить модели нарушать правила. Но это так не работает. Promt Injection нельзя просто выключить, потому что так вы все сломаете. Как работает главная уязвимость LLM? Почему с ней так тяжело бороться? И что все-таки делать, чтобы защититься от нее?
https://habr.com/ru/companies/bothub/articles/1060310/
#llm #promt_injection #ai #ai_security #promt #owasp #llm_security #guardrails #ии_угроза #bothub
-
----------------
🎯 AI
===================Arcanum AI Security Resource Hub is a curated directory of challenge platforms for practicing AI security. The collection spans beginner to advanced levels and covers the core attack surfaces in modern LLM deployments.
Core Features
The directory organizes platforms by difficulty and deployment model. Hosted options like Lakera Gandalf, Wiz AI CTF, and Forces Unseen's prompt injection games require zero setup. Self-hosted labs, including the OWASP LLM Top 10 CTF and the "Juice Shop for Agentic AI," run locally with Python and Ollama using open models such as Mistral and Llama3.
Technical Coverage
• Prompt injection: Direct and indirect techniques, including cross-user data leakage and authentication bypass through LLM manipulation
• Jailbreaking: Progressive challenges from basic password extraction to advanced guardrail circumvention
• Agentic AI attacks: Goal manipulation against tool-using AI agents, multi-step agentic workflow exploitation, and attacks on chained LLM systems performing data transformation in banking contexts
• RAG and document processing: Vulnerabilities in retrieval-augmented generation systems and document-focused AI security
• OWASP LLM Top 10: CTF-style challenges mapped to recognized risk categories
• Adversarial ML: Model inversion, data poisoning, and adversarial attacks via Garak's 80+ challenge setNotable Platforms
• Lakera Gandalf: Classic progressive prompt injection challenge
• PortSwigger Labs: Four labs covering indirect injection, data exfiltration, cross-user leakage, and auth bypass
• OWASP LLM Goat: Deliberately vulnerable chatbot lab for the OWASP LLM Top 10
• Garak: Professional platform with 80+ challenges including DEFCON and Black Hat content
• Wiz AI CTF: Five challenges manipulating a customer-service chatbotStrengths
The directory provides breadth across difficulty levels and attack categories. The mix of hosted and self-hosted options accommodates different environments, including air-gapped setups.
Limitations
Some platforms are marked buggy or offline. The "Juice Shop for Agentic AI" public Render demo is currently down. The directory provides minimal context beyond difficulty level and brief descriptions, so practitioners need to evaluate relevance independently.
🔹 bookmark #prompt_injection #LLM_security #AI_CTF #OWASP_LLM
-
Automated AI Vulnerability Scanning Trust Collapses to 9%: Study Finds Rising False Negatives - https://www.redpacketsecurity.com/trust-in-automated-ai-vulnerability-scanning-collapses-to-9-new-study-finds/
-
Automated AI Vulnerability Scanning Trust Collapses to 9%: Study Finds Rising False Negatives - https://www.redpacketsecurity.com/trust-in-automated-ai-vulnerability-scanning-collapses-to-9-new-study-finds/
-
Automated AI Vulnerability Scanning Trust Collapses to 9%: Study Finds Rising False Negatives - https://www.redpacketsecurity.com/trust-in-automated-ai-vulnerability-scanning-collapses-to-9-new-study-finds/
-
Automated AI Vulnerability Scanning Trust Collapses to 9%: Study Finds Rising False Negatives - https://www.redpacketsecurity.com/trust-in-automated-ai-vulnerability-scanning-collapses-to-9-new-study-finds/
-
Automated AI Vulnerability Scanning Trust Collapses to 9%: Study Finds Rising False Negatives - https://www.redpacketsecurity.com/trust-in-automated-ai-vulnerability-scanning-collapses-to-9-new-study-finds/
-
Иллюзия контроля: почему промпты не защищают ИИ‑агентов
Почему указание вида «не отправляй конфиденциальные данные наружу» не работает? Разбираем уязвимость Permission Boundary Bypass, а также техники scope creep и capability chaining, позволяющие злоумышленникам обходить ограничения через цепочки легитимных действий. В статье приводятся аргументы, почему prompt‑level enforcement проигрывает, зачем математическая строгость (язык Дика) нужна в конфигах политик, и как выстроить безопасную архитектуру, где проверки живут в runtime. В конец статье вы найдете 7 принципов защиты агентов и таблицу‑чеклист для аудита вашей системы.
https://habr.com/ru/articles/1050772/
#redteam #llm #prompt_injection #jailbreak #aiагенты #llm_security #indirect_prompt_injection #безопасность_данных
-
Иллюзия контроля: почему промпты не защищают ИИ‑агентов
Почему указание вида «не отправляй конфиденциальные данные наружу» не работает? Разбираем уязвимость Permission Boundary Bypass, а также техники scope creep и capability chaining, позволяющие злоумышленникам обходить ограничения через цепочки легитимных действий. В статье приводятся аргументы, почему prompt‑level enforcement проигрывает, зачем математическая строгость (язык Дика) нужна в конфигах политик, и как выстроить безопасную архитектуру, где проверки живут в runtime. В конец статье вы найдете 7 принципов защиты агентов и таблицу‑чеклист для аудита вашей системы.
https://habr.com/ru/articles/1050772/
#redteam #llm #prompt_injection #jailbreak #aiагенты #llm_security #indirect_prompt_injection #безопасность_данных
-
Иллюзия контроля: почему промпты не защищают ИИ‑агентов
Почему указание вида «не отправляй конфиденциальные данные наружу» не работает? Разбираем уязвимость Permission Boundary Bypass, а также техники scope creep и capability chaining, позволяющие злоумышленникам обходить ограничения через цепочки легитимных действий. В статье приводятся аргументы, почему prompt‑level enforcement проигрывает, зачем математическая строгость (язык Дика) нужна в конфигах политик, и как выстроить безопасную архитектуру, где проверки живут в runtime. В конец статье вы найдете 7 принципов защиты агентов и таблицу‑чеклист для аудита вашей системы.
https://habr.com/ru/articles/1050772/
#redteam #llm #prompt_injection #jailbreak #aiагенты #llm_security #indirect_prompt_injection #безопасность_данных
-
Почему ИИ-боты более уязвимы, чем их базовые LLM-модели?
В прошлой статье я показал, как защищен Open Source проект телеграм-бота. В комментариях меня спросили о иных инструментах и методах проверки в связи с чем, мы вышли к ключевому вопросу: почему, если основная LLM защищена, кастомные боты на ее основе остаются уязвимыми? Базовые LLM проходят отдельное safety-training и RLHF-выравнивание. Но production-бот, построенный поверх модели, добавляет новый attack surface: system prompts, память диалога, RAG, tools, webhook-логику и внешние API. Именно этот orchestration layer часто становится слабым местом. Вот данные: Из анализа 14 904 кастомных GPT :
https://habr.com/ru/articles/1036854/
#llm_security #prompt_injection #jailbreak #red_teaming #telegram_bot #webhook #rag #ai_safety #gpt
-
Почему ИИ-боты более уязвимы, чем их базовые LLM-модели?
В прошлой статье я показал, как защищен Open Source проект телеграм-бота. В комментариях меня спросили о иных инструментах и методах проверки в связи с чем, мы вышли к ключевому вопросу: почему, если основная LLM защищена, кастомные боты на ее основе остаются уязвимыми? Базовые LLM проходят отдельное safety-training и RLHF-выравнивание. Но production-бот, построенный поверх модели, добавляет новый attack surface: system prompts, память диалога, RAG, tools, webhook-логику и внешние API. Именно этот orchestration layer часто становится слабым местом. Вот данные: Из анализа 14 904 кастомных GPT :
https://habr.com/ru/articles/1036854/
#llm_security #prompt_injection #jailbreak #red_teaming #telegram_bot #webhook #rag #ai_safety #gpt
-
Почему ИИ-боты более уязвимы, чем их базовые LLM-модели?
В прошлой статье я показал, как защищен Open Source проект телеграм-бота. В комментариях меня спросили о иных инструментах и методах проверки в связи с чем, мы вышли к ключевому вопросу: почему, если основная LLM защищена, кастомные боты на ее основе остаются уязвимыми? Базовые LLM проходят отдельное safety-training и RLHF-выравнивание. Но production-бот, построенный поверх модели, добавляет новый attack surface: system prompts, память диалога, RAG, tools, webhook-логику и внешние API. Именно этот orchestration layer часто становится слабым местом. Вот данные: Из анализа 14 904 кастомных GPT :
https://habr.com/ru/articles/1036854/
#llm_security #prompt_injection #jailbreak #red_teaming #telegram_bot #webhook #rag #ai_safety #gpt
-
📢 Réduction du rayon d'impact des agents IA : 7 patterns tactiques contre l'injection de prompt indirecte
📝 ## 🧭 ContextePublié le 12 mai 2026 par Ross McKercha...
📖 cyberveille : https://cyberveille.ch/posts/2026-05-13-reduction-du-rayon-d-impact-des-agents-ia-7-patterns-tactiques-contre-l-injection-de-prompt-indirecte/
🌐 source : https://www.sophos.com/en-gb/blog/inside-the-lethal-trifecta-blast-radius-reduction-in-ai-agent-deployments
#Gitleaks #LLM_security #Cyberveille -
Пентест 2026: как войти в профессию
В пентест часто пытаются войти через список инструментов: выучить Burp, погонять Nmap, пройти пару лабораторий и ждать первой боевой задачи. В 2026 году такой вход всё хуже работает: часть рутины уже забирают AI‑ассистенты и автоматические сканеры, а от специалиста ждут понимания атакующей логики, бизнес‑рисков и умения проверять гипотезы руками. Разбираемся, кому сегодня действительно стоит идти в пентест, какие направления растут быстрее всего и как учиться так, чтобы не конкурировать с автоматизацией за самые простые задачи.
https://habr.com/ru/companies/otus/articles/1029746/
#пентест #кибербезопасность #информационная_безопасность #этичный_хакинг #webпентест #mobile_security #cloud_security #Active_Directory #AI_security #LLM_security
-
Пентест 2026: как войти в профессию
В пентест часто пытаются войти через список инструментов: выучить Burp, погонять Nmap, пройти пару лабораторий и ждать первой боевой задачи. В 2026 году такой вход всё хуже работает: часть рутины уже забирают AI‑ассистенты и автоматические сканеры, а от специалиста ждут понимания атакующей логики, бизнес‑рисков и умения проверять гипотезы руками. Разбираемся, кому сегодня действительно стоит идти в пентест, какие направления растут быстрее всего и как учиться так, чтобы не конкурировать с автоматизацией за самые простые задачи.
https://habr.com/ru/companies/otus/articles/1029746/
#пентест #кибербезопасность #информационная_безопасность #этичный_хакинг #webпентест #mobile_security #cloud_security #Active_Directory #AI_security #LLM_security
-
Пентест 2026: как войти в профессию
В пентест часто пытаются войти через список инструментов: выучить Burp, погонять Nmap, пройти пару лабораторий и ждать первой боевой задачи. В 2026 году такой вход всё хуже работает: часть рутины уже забирают AI‑ассистенты и автоматические сканеры, а от специалиста ждут понимания атакующей логики, бизнес‑рисков и умения проверять гипотезы руками. Разбираемся, кому сегодня действительно стоит идти в пентест, какие направления растут быстрее всего и как учиться так, чтобы не конкурировать с автоматизацией за самые простые задачи.
https://habr.com/ru/companies/otus/articles/1029746/
#пентест #кибербезопасность #информационная_безопасность #этичный_хакинг #webпентест #mobile_security #cloud_security #Active_Directory #AI_security #LLM_security
-
----------------
🛠️ Tool — PwnzzAI Shop: Intentional Vulnerable AI Training Application
===================Opening:
PwnzzAI Shop is an intentionally insecure, Flask-based web application that presents AI security concepts through a pizza-shop scenario. The project maps its content to the OWASP AI Exchange taxonomy and incorporates the OWASP Top 10 for LLMs, making the repo a focused educational tool for practitioners exploring LLM-related weaknesses and mitigations.Key Features:
• Hands-on learning scenarios that demonstrate how prompt flaws, model context misuse, and architecture mistakes lead to data exposure and unauthorised access.
• Pre-built user personas (example credentials like alice/alice, bob/bob) to exercise role-based flows and privilege boundaries without needing separate account creation.
• Integration points and examples showing model interactions and common misconfigurations with model hosts such as Ollama (documented as supported models), enabling exercises across local and hosted model setups.Technical Implementation (conceptual):
• The application is implemented as a Flask web app and intentionally exposes insecure endpoints and flows that illustrate LLM attack vectors such as prompt injection, context leakage, and improper access control around model inputs/outputs.
• The codebase is structured to align exercises with the AI Exchange risk classifications, providing mapping between each lab and specific taxonomy items like data leakage, model poisoning scenarios, and RAG-related issues.Use Cases:
• Red-team/blue-team training focusing on LLM-specific TTPs and exploitation of model interfaces.
• Curriculum development for AI security courses that require reproducible, scenario-driven labs mapped to an industry taxonomy.
• Practitioner onboarding for teams responsible for securing LLM integrations and RAG pipelines.Limitations:
• The project is intentionally insecure by design and intended for controlled, isolated learning environments only; it is not suitable for production deployment.
• The repository documents multiple setup options and model integrations but does not prescribe a single operational model; environment specifics may affect reproducibility.Closing:
PwnzzAI Shop serves as a practical repository for translating OWASP AI Exchange concepts into hands-on labs, helping security practitioners explore LLM-specific failure modes and defensive patterns.🔹 tool #owasp #llm_security
🔗 Source: https://github.com/OWASP/PwnzzAI
-
----------------
🛠️ Tool
===================Opening: LLM Anonymization is a transparent proxy designed to mediate between Claude Code and the Anthropic API in penetration testing engagements. It guarantees that sensitive artifacts — command outputs, file contents, hostnames, IPs and credentials — never leave the tester’s machine in raw form. The proxy substitutes persistent surrogates for any detected PII and restores original values only after Claude’s response returns.
Key Features:
• Persistent surrogate mappings stored in a PII Vault (SQLite) to maintain session-consistent identifiers per engagement.
• Dual detection model combining an LLM-based detector (noted as Ollama qwen3:4b) with a deterministic regex safety net for IPs, CIDRs, hashes, MACs, emails, domains, tokens and JWTs.
• End-to-end behavioral transparency: Claude Code is used unchanged, while the proxy intercepts and transforms inputs and outputs invisibly to the user.Technical Implementation:
• The proxy exposes a local endpoint that Claude Code is pointed at via an overridden ANTHROPIC_BASE_URL setting (conceptual integration only). All traffic passing through the proxy is analyzed by Layer 1 (LLM detector) and Layer 2 (regex safety net). Detected items are replaced with surrogate tokens and stored in the PII Vault keyed by client/engagement ID. Only surrogate data leaves the host; responses from Anthropic containing surrogates are mapped back to original values before presentation to the tester.Use Cases:
• Red team and pentest operations where sharing raw tooling outputs (nmap, crackmapexec, mimikatz, log snippets) with a third‑party LLM is required but must avoid data leakage.
• Safe LLM-assisted analysis of internal files, grep results and credentials during client engagements.Limitations and Considerations:
• Detection coverage relies on regex patterns and the LLM detector; false negatives are possible for novel or obfuscated identifiers.
• Files or screenshots shared outside the proxied channel are out of scope and require separate handling.
• Surrogate consistency depends on the integrity of the PII Vault; protecting that store is critical.Conclusion: LLM Anonymization provides a practical, architecture-oriented approach to reduce sensitive data exposure when leveraging Claude Code for offensive security tasks while preserving the user experience of an unmodified LLM client. #tool #anonymization #llm_security
-
----------------
🛠️ Tool
===================Opening: The GitHub Security Lab released the seclab-taskflows framework and a set of auditing taskflows designed to find web security vulnerabilities across open source repositories. The framework orchestrates LLM-driven tasks described in YAML and preserves intermediate context in a database for follow-on analysis.
Key Features:
• Component decomposition: Repositories are split into functional components and each component is profiled for entry points, intended privileges, and purpose.
• Prompt-driven auditing: Taskflows define prompts and task dependencies; the agent runs tasks sequentially and feeds results forward.
• Context persistence: Results are saved to an SQLite table (audit_results) enabling manual review and follow-up tasks.
• Model flexibility: The auditing process uses LLMs via premium Copilot access and can produce nondeterministic outputs across runs and models (examples cited include GPT-5.2 and Claude Opus 4.6).Technical Implementation:
• Taskflows are authored as YAML files that enumerate tasks and dependencies; the Taskflow Agent executes these tasks and manages state transitions.
• Data captured per component includes entry points for untrusted input, intended privileges, and suggested issues; follow-up tasks perform deeper audits of flagged items.
• Execution artifacts are organized into a local SQLite schema (noted table audit_results) for human verification and reporting.Use Cases:
• Automated reconnaissance of large codebases to surface authorization bypasses and information disclosure vectors.
• Prioritization of manual triage by surfacing high-severity candidate issues for security teams and maintainers.
• Continual improvement of taskflows by adding specialized checks for specific issue classes.Limitations:
• Results are nondeterministic: repeated runs (and runs against different LLMs) can return materially different findings.
• The framework relies on access to premium LLM capabilities (Copilot license) for the prompts described in the blog.
• False positives and unexploitable candidates still require manual verification; the team reports that manual triage remains part of the workflow.Summary: seclab-taskflows demonstrates how orchestrated, prompt-driven taskflows can scale discovery of high-impact web vulnerabilities in open source projects while preserving context for reviewer validation and coordinated disclosure. #tool #seclab_taskflows #GitHub_Security_Lab #LLM_security #security
-
AI Red Teaming: спор с Grok — Часть 4. От атаки к защите: как результаты red team улучшили мой продукт
61 уязвимость бесполезна, если не превращается в защиту. Каждую находку в Grok я превратил в вопрос: «а мы от этого защищаем?» Ответ был неутешительный — 5 из 5 нет. Как результаты red team стали 138 паттернами, правилами и payloads в нашем продукте. Плюс — чем закончился спор с Grok.
https://habr.com/ru/articles/1005306/
#информационная_безопасность #AI #red_team #LLM_security #Sentinel #xAI #Grok #defensive_security
-
AI Red Teaming: спор с Grok — Часть 4. От атаки к защите: как результаты red team улучшили мой продукт
61 уязвимость бесполезна, если не превращается в защиту. Каждую находку в Grok я превратил в вопрос: «а мы от этого защищаем?» Ответ был неутешительный — 5 из 5 нет. Как результаты red team стали 138 паттернами, правилами и payloads в нашем продукте. Плюс — чем закончился спор с Grok.
https://habr.com/ru/articles/1005306/
#информационная_безопасность #AI #red_team #LLM_security #Sentinel #xAI #Grok #defensive_security
-
AI Red Teaming: спор с Grok — Часть 4. От атаки к защите: как результаты red team улучшили мой продукт
61 уязвимость бесполезна, если не превращается в защиту. Каждую находку в Grok я превратил в вопрос: «а мы от этого защищаем?» Ответ был неутешительный — 5 из 5 нет. Как результаты red team стали 138 паттернами, правилами и payloads в нашем продукте. Плюс — чем закончился спор с Grok.
https://habr.com/ru/articles/1005306/
#информационная_безопасность #AI #red_team #LLM_security #Sentinel #xAI #Grok #defensive_security
-
SecureShell - Lớp bảo mật terminal plug-and-play cho agent LLM. Ngăn lệnh nguy hiểm/hỏng, áp dụng chính sách bảo vệ cấu hình, yêu cầu giải thích hợp lý trước khi thực thi. Hỗ trợ đa nền tảng (Linux/macOS/Windows), tích hợp Ollama, llama.cpp, LangChain/MCP. Cài đặt đơn giản qua pip/npm. Bảo vệ hệ thống trước thao tác tự động của AI. #Bảo_mật_AI #LLM_Security #An_toan_he_thống
-
🚑 Incident Response
====================🛠️ Tool
Executive summary: AIDR Bastion is a layered GenAI protection
framework that intercepts user and agent prompts before they reach LLM
applications. The system combines deterministic regex matching,
rule-based detection via Sigma and Roota, vector similarity lookups in
OpenSearch, ML classifiers, and optional LLM review to reduce
successful prompt injection and harmful content flows.Technical implementation: The project exposes a FastAPI endpoint (POST
/api/v1/run_pipeline) that forwards inputs to a Pipeline Manager.
Active pipelines include a Regex Pipeline (Roota), a Similarity
Pipeline (OpenSearch vectors), a Code Analysis Pipeline
(Semgrep-compatible rules via Sigma/Uncoder AI translations), an ML
Pipeline, and an LLM Pipeline for contextual analysis. Rule sources
are extensible, supporting community Sigma rules (~1,200 at release)
and SOC Prime integrations. Processing is asynchronous with
configurable thresholds and JSON-driven pipeline configuration.Detection and response capabilities: Detection logic supports allow,
block, and notify actions based on match severity and policy.
OpenSearch vector similarity is used for prompt classification and
reuse of labeled examples. Semgrep translation of Sigma rules via
Uncoder AI enables standardized code-pattern detection for cases where
local LLMs generate executable code. Logging can operate as a local
sensor for incident discovery and forensic reconstruction.Use cases: The tool is suited for organizations deploying local or
hosted LLMs that require inline prompt sanitization, SOCs seeking
telemetry for LLM misuse, and teams needing a testbed for layered
defensive controls against adversarial prompt engineering.Limitations and considerations: The system relies on rule coverage and
labeled similarity indices; false positives are possible with broad
regexes or low similarity thresholds. LLM-based review introduces cost
and potential latency. Integration with enterprise identity and policy
orchestration may be required for enforcement at scale. Maintenance of
Sigma/Roota rule sets and vector databases is operational overhead.Strategic takeaway: AIDR Bastion presents a modular, extensible
architecture that operationalizes multiple detection modalities for
LLM protection, aligning detection logic with MITRE ATLAS and OWASP
Top 10 for LLMs to improve defensive posture against prompt injection
and AI‑focused attack vectors.🔹 tool #Sigma #OpenSearch #MITRE_ATLAS #LLM_security
🔗 Source: https://github.com/0xAIDR/AIDR-Bastion