Evaluating Agents Beyond the First Prompt
EvoCode-Bench is a multi-turn coding benchmark that tests AI agents on evolving tasks across 227 rounds, evaluating persistence and adaptability.
EvoCode-Bench is a multi-turn coding benchmark that tests AI agents on evolving tasks across 227 rounds, evaluating persistence and adaptability.
Analysis of rising code review loads due to AI-generated code and the surge of AI tools to manage it.
Explores GitHub Copilot's new Agent Mode, which goes beyond autocomplete to autonomously edit code, run tests, and fix errors.
Explores the concept of software factories using AI loops, comparing human-in-the-loop (light) vs. autonomous (dark) approaches, and the engineering challenges involved.
A philosophical exploration of being wrong in technology, challenging binary thinking and advocating for nuanced judgment.
A software developer's monthly retrospective on writing a book about effective writing for developers, with sales metrics and productivity insights.
Explores the viability of running local AI models for coding, covering hardware, performance, and quality factors.
Explores how AI and supply-chain governance are challenging the 'don't reinvent the wheel' mantra in software development, making custom code potentially safer and cheaper.
Explores the future of product management, predicting that AI and smaller teams will shift PM responsibilities to developers and domain experts.
A set of project management principles emphasizing problem-solving, scope flexibility, and the trade-off between fixed dates and scope.
Personal reflections on the 3rd anniversary of JasperFx Software, a company built around OSS development tools.
A software developer's monthly retrospective on completing a book about effective writing for developers, with metrics and bug bounty updates.
A reflective analysis on whether AI tools truly boost productivity or just create a sense of performative busywork in software development.
Explores how Claude Code Teams Agents enable specialized AI agents to collaborate like human engineering teams, moving beyond single-assistant coding to scalable AI-driven software delivery.
A developer reflects on using Claude Code, noting less coding but more testing and understanding of AI-generated code.
Explores 'vibe coding'—building software by prompting LLMs without reviewing generated code, its benefits, risks, and distinction from agentic programming.
A software developer reflects on balancing writing a book on effective writing for developers with AI-assisted bug bounty hunting.
Explores ZeroVer, a satirical yet realistic alternative to Semantic Versioning for software that stays in 0.x releases.
Explores the difference between simply using AI and achieving AI maturity, focusing on real outcomes and disciplined integration.
How to ask better questions at work by providing context to reduce anxiety and ambiguity.