Back

Lesson 5/5

1.4

The eval mindset

1. The shift

Two engineers are looking at a new prompt.

The first one says "I think it's better. The outputs read more naturally and it's catching more of the urgent ones."

The second one says "It's two points higher on category, half a point lower on priority, and the urgent-flag recall went from 78 to 84. I'm going to add three more urgent cases and rerun before we ship."

Same team. Same prompt change. Different mindset.

This chapter has been about installing the second one. Not because the first engineer is wrong (they might be right about the prompt) but because the first engineer can't show their work, and the second one can.

2. What changes when you have evals

Three things shift in your day-to-day.

Arguments become tables. You stop debating whether the new prompt is better. You run both prompts against the same eval set and read the rows. The conversation moves from opinions to numbers, and the numbers move the decision.

Regressions get caught before customers find them. You change something. You rerun the set. The number went down. You investigate before you ship instead of after. The bug never makes it to a user.

Improvements become legible. You change something. You rerun the set. The number went up. You can point at the delta. Your manager doesn't have to take your word for it. Marketing doesn't have to take your word for it. The number is the receipt.

3. What doesn't change

Evals don't make the model better. They tell you whether your changes did.

Evals don't replace product judgment. You still have to decide what counts as a good output. The eval set encodes your judgment, but it doesn't manufacture it. If you can't write down what "right" looks like, no framework on earth will compute it for you.

Evals don't replace looking at outputs. They focus the looking. You stop scrolling through logs hoping to spot something. You read the rows where the score dropped and look at those outputs.

4. The cultural shift

This is the part that's harder than the technical part.

Teams without evals have culture wars about AI quality. Half the team thinks the feature is fine. Half thinks it's getting worse. Nobody can settle it because nobody has the receipts. The arguments are heated, infrequent, and unresolved.

Teams with evals have a different conversation. "The set says category is at 0.91, priority is at 0.84, the summary rubric is at 3.2 out of 5. Where do we want each of these by end of quarter?" That's a planning conversation, not a fight. It happens every week. It resolves.

You'll notice the second team gets more done. Not because they're smarter. Because they're spending their meeting time on decisions instead of vibes.

5. The cost of the mindset

Evals are not free. The cost is real and it's worth naming.

You have to write the cases. The first version takes a few hours. Maintaining them takes ongoing attention. Cases go stale. New failure modes show up. You'll be editing the set as long as the feature exists.

You have to agree on what "right" means. This is hard. Two engineers on the same team will disagree about whether a borderline ticket is "billing" or "general." That disagreement is good — it surfaces something you needed to decide anyway — but it's work.

You have to actually look at the numbers. Evals you never read are worse than no evals, because they give you false confidence. Once a week is enough. Zero times a week is the failure mode.

If you can pay those costs, the payoff is everything described above. Tables instead of arguments. Receipts for your work. Bugs caught before users find them.

6. What's next

Chapter 2 hands you the vocabulary. Accuracy, recall, precision, coverage, regression, ground truth, eval set versus test set versus production. The words the rest of the field uses, so you can read other people's work and write your own without bluffing. Five units. Plain definitions. Same ticket extractor as the running example throughout.