Sketch to Eval
8 chapters.
A short, free course on evaluating AI outputs. From the smallest eval set that's still useful to designing for adversarial cases, reading regression vs noise, and running evals once real users are involved. Read-only, with TypeScript code samples grounded in a single recurring example: extracting structured data from customer support tickets.
8 chapters · 34 units
- Ch 0
Start here
Psychological · light buildWhat this is. Who it's for. The honest test for whether to keep reading.
2 units
- Ch 1
What an eval is, actually
PEAK · most weightStrip the framework noise. An eval at its smallest is a table you can read in fifteen seconds.
5 units
- Ch 2
The vocabulary
Medium intensityAccuracy, recall, precision, drift. The words you need to talk about quality without sounding vague.
5 units
- Ch 3
Designing your eval set
PEAK · most weightThe single most practical chapter. Where your eval set actually comes from, and what goes in.
5 units
- Ch 4
The types of eval
Medium intensityExact-match, semantic, LLM-as-judge, human. Pick the cheapest one that's still honest.
5 units
- Ch 5
Reading the numbers
Medium intensityOnce you have eval results, the new question is which ones matter and which are noise.
4 units
- Ch 6
Evals in production
Emotional high pointThe hardest mode. Continuous evals, shadow evals, drift detection, disagreement with your users.
4 units
- Ch 7
Where to go next
Medium intensityThe eval mindset stays with you. Honest pointers to the next courses in the catalog.
4 units