Evals

Sep 18, 2026

8 min

How we evaluate agents before they ship

Benchmarks you write yourself beat leaderboards you read about.

Leaderboards tell you how a model does on someone else’s problem. Your own benchmarks tell you how it does on yours.

Write the test first

Collect fifty real tasks, define a pass, and run every model and prompt change against them before it ships.

Next step

Build your empire on AI.

Create a free website with Framer, the website builder loved by startups, designers and agencies.