A knowledge cutoff is a ceiling, not a snapshot
Explains why a model's knowledge cutoff is a ceiling, not a snapshot, and how this affects AI agent design.
Explains why a model's knowledge cutoff is a ceiling, not a snapshot, and how this affects AI agent design.
AI agent learning: governing durable changes via provenance, scope, evaluation, approval, and revocation.
AI feedback isn't learning: separating memory, runbooks, and model updates for safe agent governance.
Testing AI agent reliability across the full coordination loop, including failure recovery, evidence preservation, and authority enforcement.
Implementing session state for a multi-agent blog system to support search refinement with short-term memory.
How AI agents should verify before acting: targeted diagnostics, external limits, and separating evidence from execution authority.
Build an offline Python grader for AI agents that separates control, behavior, and outcomes—not just explanations.
Designing stable AI agents: separate retries, reconciliation, and recovery to control interventions, not just model retries.
AI agent verification: prove runtime outcomes, not just tool call success. Separate acceptance, config, convergence, and service results.
Armin Ronacher's critical take on AI 'doom' narratives and the push to pace frontier AI development.
Designing a shared workspace for AI agents: bounded, versioned task state, controlled interfaces, and atomic updates to coordinate evidence and proposals.
Kody is a sandboxed cloud runtime where AI agents can author and execute code, integrating with APIs and MCP servers, with secure secret storage.
Simon Willison announces release of llm 0.35, a command-line tool for accessing large language models, with support for new OpenAI model gpt-6-astra.
Interview with Andrew Ambrosino, project lead for OpenAI's Codex, discussing its history, integration with ChatGPT, Computer Use, and user trust.
Agent Package Manager (APM) manages AI agent dependencies like skills, prompts, and MCP servers, ensuring reproducibility and governance for Azure engineering teams.
Explores how AI agents are changing app usage, making traditional UIs less essential as users prefer direct outcomes over navigating interfaces.
Security exploits are now found within minutes of bug rumors due to AI agents, challenging open-source embargo practices.
A reflection on AI use cases, arguing that many popular AI applications solve problems created by technology itself, and that productivity is personal.
Explore integrating GitHub Copilot with Microsoft Agent Framework to build intelligent agents for automation and productivity.
Hands-on review of ChatGPT Work's new cloud browser feature for logging into websites securely, with mixed results on Mac and iPhone.