An opinionated guide to which AI to use to do stuff
An opinionated guide comparing AI tools like ChatGPT and Claude, focusing on agentic systems and computer access modes.
An opinionated guide comparing AI tools like ChatGPT and Claude, focusing on agentic systems and computer access modes.
Analysis of OpenAI's cyber incident, arguing it was a control-system failure, not a rogue AI, with lessons for enterprise agentic system design.
Explains the importance of transparency in AI agent systems, including logging, audit trails, and human oversight for debugging.
Explores using ClickHouse for low-latency analytical loops in active agent systems, emphasizing validation, safety, and architecture patterns.
Anthropic's postmortem on Claude Code quality issues reveals three bugs in the harness causing forgetfulness and repetition.
Review of a paper on using a meta-agent to automatically design novel AI agent architectures, outperforming hand-crafted systems.
Discusses the reliability challenges and lack of provable correctness guarantees in current AI systems, despite their productivity benefits.
A research paper analyzes LLM performance on SQL generation tasks using different structured data formats and large schemas, comparing frontier and open-source models.
Research paper analyzes LLM performance on large SQL schemas, comparing 11 models across 4 data formats for structured context engineering in agentic systems.