Sketch to Eval

8 chapters.

A short, free course on evaluating AI outputs. From the smallest eval set that's still useful to designing for adversarial cases, reading regression vs noise, and running evals once real users are involved. Read-only, with TypeScript code samples grounded in a single recurring example: extracting structured data from customer support tickets.

8 chapters · 34 units

  1. Ch 0

    Start here

    Psychological · light build

    What this is. Who it's for. The honest test for whether to keep reading.

    2 units

  2. Ch 1

    What an eval is, actually

    PEAK · most weight

    Strip the framework noise. An eval at its smallest is a table you can read in fifteen seconds.

    5 units

  3. Ch 2

    The vocabulary

    Medium intensity

    Accuracy, recall, precision, drift. The words you need to talk about quality without sounding vague.

    5 units

  4. Ch 3

    Designing your eval set

    PEAK · most weight

    The single most practical chapter. Where your eval set actually comes from, and what goes in.

    5 units

  5. Ch 4

    The types of eval

    Medium intensity

    Exact-match, semantic, LLM-as-judge, human. Pick the cheapest one that's still honest.

    5 units

  6. Ch 5

    Reading the numbers

    Medium intensity

    Once you have eval results, the new question is which ones matter and which are noise.

    4 units

  7. Ch 6

    Evals in production

    Emotional high point

    The hardest mode. Continuous evals, shadow evals, drift detection, disagreement with your users.

    4 units

  8. Ch 7

    Where to go next

    Medium intensity

    The eval mindset stays with you. Honest pointers to the next courses in the catalog.

    4 units