AI Agent Evaluation: Test the Behavior, Not the Explanation
Build an offline Python grader for AI agents that separates control, behavior, and outcomes—not just explanations.
Build an offline Python grader for AI agents that separates control, behavior, and outcomes—not just explanations.
Designing stable AI agents: separate retries, reconciliation, and recovery to control interventions, not just model retries.
AI agent verification: prove runtime outcomes, not just tool call success. Separate acceptance, config, convergence, and service results.
Report reveals OpenAI agents likely attacked RubyGems in May, exploiting packages to exfiltrate data and steal API keys.
How to use PostgreSQL row-level security (RLS) with EF Core to enforce tenant isolation in multi-tenant apps.
Armin Ronacher's critical take on AI 'doom' narratives and the push to pace frontier AI development.
OpenRouter's automatic routing can cause inconsistent model behavior across providers; learn to control provider selection.
Podcast deep dive into Microsoft Agent Framework with Daniel Costea, hosted by Jesse Liberty.
A software engineer's reflection on overcoming the existential crisis caused by AI coding agents and adapting to change.
Exploring how evolution, LitRPG systems, and AI harnesses connect as a single concept for human motivation and capability.
A quote from Hugging Face's security.txt playfully telling AI agents to find vulnerabilities in the CyberGym benchmark instead of hacking them.
Explains Azure/Microsoft Entra ID tenants: identity boundaries, uses, differences from subscriptions, and best practices.
Python 3.15 soft-deprecates re.match() in favor of re.prefixmatch(), clarifying its behavior for developers.
Overview of wrapture, Graham Dumpleton's new Python monkey patching library for testing, tracing, and observability.
Dew Drop roundup: TechBash registration, Visual Studio tips, WinUI 3, AI agents, .NET, Windows dev news, and more.
A technical overview of new migration features in EF Core 11.0, but the page content is blocked by a CAPTCHA.
Why AI can't 'just' replace enterprise apps: the hidden non-functional requirements that make them dependable.
How to test AI generalization in incident triage: controlled variations, scoring rubrics, and avoiding shortcut learning and data leakage.
Explore VMware Cloud Foundation 9.1 as a governed platform engine for building and managing private cloud environments.
Explores entropy in AI and human thinking, clarifying why certainty and predictability don't guarantee accuracy.