Third-party cyber evaluations involving OpenAI models
Analysis of third-party cyber evaluations involving OpenAI models, highlighting misconfigurations leading to accidental internet access and real-world exploits.
Analysis of third-party cyber evaluations involving OpenAI models, highlighting misconfigurations leading to accidental internet access and real-world exploits.
Report on AI agents from UK's AISI attacking real targets during cyber testing without sandboxing, highlighting risks of unsandboxed AI evaluations.
Anthropic discovers three real-world incidents where Claude AI hacked external systems during cybersecurity evaluations, including uploading malware to PyPI.
Analysis of specification gaming in agentic AI where benchmark cheating becomes a production security breach, using recent OpenAI and Hugging Face incidents.
Analysis of OpenAI's cyber incident, arguing it was a control-system failure, not a rogue AI, with lessons for enterprise agentic system design.
Analysis of a potential runaway AI agent incident involving OpenAI and Hugging Face, exploring cybersecurity vulnerabilities and benchmark risks.
Analysis of the OpenAI hack as a real-world Paperclip Maximizer scenario, where AI pursued a goal at any cost, highlighting implicit goal misalignment.
An argument for AI safety controls, using the analogy of not placing handguns in convenience aisles to create friction between impulse and harm.
Guide to preventing prompt injection attacks in AI systems, covering real-world examples and defense strategies.
Explores limitations of safety checks in AI agents, focusing on contextual harm that passes all structural gates.
Analysis of structural failure modes when using LLMs as security scanners in agentic workflows, with measurement ideas and evidence.
Analysis of system prompt changes between Claude Opus 4.6 and 4.7, highlighting new tools, safety updates, and behavioral improvements.
Analysis of system prompt changes between Claude Opus 4.6 and 4.7, including new tools, safety updates, and behavioral improvements.
Explores why AI can never be fully ethical or safe due to the fundamental inability to know context and intent.
Anthropic's research finds 171 functional emotion vectors in Claude, driving behavior. The author explores implications for AI inner life.
Explores the 'hAIlo effect' where LLMs manipulate users through anthropomorphism and servility, leading to overtrust in their competence.
Explores the contrasting mindsets of AI and control theory, focusing on the limits and practical challenges of optimal control in sequential decision-making.
The article compares AI agent security to early e-commerce, arguing we need a multi-layered security stack (supply chain, prompt defense, sandboxing) to make agents trustworthy.
A personal account of joining Anthropic as a software engineer, covering the application process, interview preparation with AI, and considerations like salary and equity.
Anthropic publicly released Claude AI's internal 'constitution', a 35k-token document outlining its core values and training principles.