Simon Willison 7/31/2026

smevals - a small eval suite for evaluating models, prompts, and harnesses

Read Original

This article introduces smevals, a lightweight evaluation framework for testing AI models, prompts, and harnesses. Developed in collaboration with Jesse Vincent's Prime Radiant lab, smevals allows users to define eval suites as YAML directories, run them against multiple models (e.g., gpt-5.5, claude-opus-4.6), grade results against custom checks, and view reports via localhost or static HTML. The author describes it as their third iteration on evals, highlighting its simplicity and potential for future expansion. The article includes a practical example of evaluating haiku-writing capabilities.

smevals - a small eval suite for evaluating models, prompts, and harnesses

Comments

No comments yet

Be the first to share your thoughts!

Browser Extension

Get instant access to AllDevBlogs from your browser