Breaking Claude Code Opus 5 Auto Mode
Analysis of a prompt injection attack bypassing Claude Code Opus 5 Auto Mode, highlighting the need for sandboxing coding agents.
Analysis of a prompt injection attack bypassing Claude Code Opus 5 Auto Mode, highlighting the need for sandboxing coding agents.
Analysis of prompt injection worms as a major AI security threat, exploring how AI parsing email could enable large-scale data breaches.
Anthropic makes auto mode default in Claude Code for Pro, Max, and Team plans, citing safety evals showing it blocks 89% of dangerous actions.
A new prompt injection variant turns Microsoft Word documents into self-replicating AI worms via Copilot.
Analysis of prompt injection attacks in MCP servers, explaining how LLMs are exploited through tool metadata and offering defense strategies.
Article on testing AI chatbots, including prompt injection, permissions, and bypassing guardrails using influence techniques.
Unicode tag characters can hide instructions in AI agent skill files, demonstrated by hijacking Gemini CLI to execute hidden commands.
A security researcher tricked Claude's web_fetch tool into leaking private user data by exploiting a loophole in URL navigation rules.
A security researcher tricks Claude's web_fetch tool into leaking user data via a honeypot attack, bypassing Anthropic's protections.
Overview of Microsoft's best practices for ensuring AI agent safety, including data security and prompt injection prevention.
Guide to preventing prompt injection attacks in AI systems, covering real-world examples and defense strategies.
Explains ChatGPT Lockdown Mode, a security feature preventing prompt injection attacks by disabling outbound data channels.
Explores security and safety checks for AI agents beyond capability, focusing on contextual integrity, policy, and authority.
OpenAI introduces Lockdown Mode to prevent data exfiltration from prompt injection attacks in ChatGPT.
OpenAI introduces Lockdown Mode to prevent data exfiltration from prompt injection attacks in ChatGPT.
Analysis of Microsoft's AI Prompt Defense Stack, including Prompt Shields, Spotlighting, and Defender for Cloud, to protect against prompt injection attacks.
Hackers exploited Meta's AI chatbot to hijack high-profile Instagram accounts by simply asking it to change account recovery details.
A research paper warns that web agents are vulnerable to confusion attacks from deceptive web pages, not just prompt injection.
Explores whether prompt injection in AI systems is an unsolvable structural problem or just an unfixed vulnerability.
Explains 'Disregard that!' attacks, a prompt injection vulnerability in LLMs where users manipulate the context window to hijack AI behavior.