CoursesAgents and Reliability

Evals and Observability

How you know an agent works: an eval set that grows from failures, graders you can trust, traces you can search, a taxonomy of failures, and the regression check that runs before every change.

Lessons
5
Exercises
29
Minutes
31

Start Building an Eval Set

  1. 1

    Building an Eval Set

    After this lesson you can start an eval set from the failures you already have, write cases with a checkable expected outcome, and grow it so that every fixed bug stays fixed.

    DebugFill blankMultiple choice
    5 exercises
    5 min
  2. 2

    Graders You Can Trust

    After this lesson you can pick the right grader for a case: exact checks, property checks, or a model as judge, and know when the judge is grading its own homework.

    DebugFill blankMultiple choice
    5 exercises
    5 min
  3. 3

    Traces, Taxonomy and Regression

    After this lesson you can log an agent so its failures can be found, name them with a taxonomy that points at a fix, and run the regression check that keeps them from returning.

    DebugMultiple choiceTrace
    5 exercises
    6 min
  4. 4

    Checkpoint: ProofCheckpoint

    Cases, graders, traces and regression in fresh situations.

    Fill blankMultiple choice
    6 exercises
    6 min
  5. 5

    Boss: The RegressionBoss

    A model upgrade is proposed for your support agent, and you own the decision. One upgrade, decided with evidence: the eval run, the failures that matter, the judge that lied, the trace that explained it, the fix, and the case that keeps it fixed.

    Code orderMultiple choiceTrace
    8 exercises
    9 min