I was wrong: there are bad prompts
Explores why AI prompts can be bad, emphasizing reliability across runs over occasional success, and how to judge prompt quality.
Explores why AI prompts can be bad, emphasizing reliability across runs over occasional success, and how to judge prompt quality.
Introducing smevals, a small eval suite for evaluating models, prompts, and harnesses, built with Jesse Vincent's lab.
Analysis of whether AI labs train models to draw pelicans on bicycles, with no evidence found.
Analysis testing if AI labs trained models to draw pelicans on bicycles, finding no evidence of bias.
A guide to building eval loops for LLM-based apps, focusing on measuring prompt performance across model updates.