Evals and Observability
How you know an agent works: an eval set that grows from failures, graders you can trust, traces you can search, a taxonomy of failures, and the regression check that runs before every change.
- Lessons
- 5
- Exercises
- 29
- Minutes
- 31
- 1
Building an Eval Set
After this lesson you can start an eval set from the failures you already have, write cases with a checkable expected outcome, and grow it so that every fixed bug stays fixed.
DebugFill blankMultiple choice - 2
Graders You Can Trust
After this lesson you can pick the right grader for a case: exact checks, property checks, or a model as judge, and know when the judge is grading its own homework.
DebugFill blankMultiple choice - 3
Traces, Taxonomy and Regression
After this lesson you can log an agent so its failures can be found, name them with a taxonomy that points at a fix, and run the regression check that keeps them from returning.
DebugMultiple choiceTrace - 4
Checkpoint: ProofCheckpoint
Cases, graders, traces and regression in fresh situations.
Fill blankMultiple choice - 5
Boss: The RegressionBoss
A model upgrade is proposed for your support agent, and you own the decision. One upgrade, decided with evidence: the eval run, the failures that matter, the judge that lied, the trace that explained it, the fix, and the case that keeps it fixed.
Code orderMultiple choiceTrace