Exercises in benchmarking and evals, part 7: DeepSWE, Senior SWE-Bench, napkin math, and winter tires
Analysis of benchmarking exercises including napkin math performance estimates, DeepSWE, Senior SWE-Bench, and winter tire analogies.
Analysis of benchmarking exercises including napkin math performance estimates, DeepSWE, Senior SWE-Bench, and winter tire analogies.
A fireside chat with Claude Code team members discussing coding agents, tool design, and Anthropic's internal development practices.
A guide to building cybersecurity evaluations for AI models, covering sandboxed targets, inputs, tools, and grading methods.
Article discusses the importance of using evals to measure AI improvements when adding AI skills to your resume.
A developer details the process of building evaluation systems for two AI-powered developer tools to measure their real-world effectiveness.
Key takeaways from the AI Engineer Summit 2023, focusing on challenges in LLM deployment like evaluation methods and serving costs.