Courses

Agents and Reliability

Build agents that stop, stay inside their permissions, survive prompt injection, and prove it with evals and traces.

Islands
4
Lessons
23
Exercises
133
DebugCode orderFill blankMultiple choiceTrace

Start the first lesson

  1. Island 1

    The Agent Loop

    What an agent is when you strip the word away: a model, a set of tools, a loop, and a stop condition. How state survives across steps, why loops run away, and how to read a trace that did.

    1. What an Agent Is5 exercises, 5 min
    2. State Across Steps5 exercises, 6 min
    3. Stopping5 exercises, 5 min
    4. Reading an Agent Trace5 exercises, 6 min
    5. When One Agent Becomes Several5 exercises, 6 min
    6. Checkpoint: The LoopCheckpoint7 exercises, 6 min
    7. Boss: The Runaway LoopBoss8 exercises, 9 min
  2. Island 2

    Prompt Injection

    Instructions hidden in data, the three ingredients that turn them into a breach, how to see an attack in a trace, and the designs that make it survivable rather than the filters that promise to stop it.

    1. Instructions in the Data5 exercises, 5 min
    2. The Three Ingredients5 exercises, 5 min
    3. Designing Around It5 exercises, 6 min
    4. Checkpoint: InjectionCheckpoint6 exercises, 6 min
    5. Boss: The Inbox AgentBoss8 exercises, 9 min
  3. Island 3

    Guardrails and the Human in the Loop

    Least privilege for tools, the approval boundary that a model cannot talk its way past, sandboxes that limit the blast radius, and the human step placed where it does the most good.

    1. Least Privilege for Tools5 exercises, 5 min
    2. The Approval Boundary5 exercises, 6 min
    3. Sandboxes and Blast Radius5 exercises, 5 min
    4. MCP Servers Are Tool Lists5 exercises, 6 min
    5. Checkpoint: GuardrailsCheckpoint7 exercises, 6 min
    6. Boss: The Deploy BotBoss8 exercises, 9 min
  4. Island 4

    Evals and Observability

    How you know an agent works: an eval set that grows from failures, graders you can trust, traces you can search, a taxonomy of failures, and the regression check that runs before every change.

    1. Building an Eval Set5 exercises, 5 min
    2. Graders You Can Trust5 exercises, 5 min
    3. Traces, Taxonomy and Regression5 exercises, 6 min
    4. Checkpoint: ProofCheckpoint6 exercises, 6 min
    5. Boss: The RegressionBoss8 exercises, 9 min