Agents and Reliability
Build agents that stop, stay inside their permissions, survive prompt injection, and prove it with evals and traces.
- Islands
- 4
- Lessons
- 23
- Exercises
- 133
- Island 1
The Agent Loop
What an agent is when you strip the word away: a model, a set of tools, a loop, and a stop condition. How state survives across steps, why loops run away, and how to read a trace that did.
- Island 2
Prompt Injection
Instructions hidden in data, the three ingredients that turn them into a breach, how to see an attack in a trace, and the designs that make it survivable rather than the filters that promise to stop it.
- Island 3
Guardrails and the Human in the Loop
Least privilege for tools, the approval boundary that a model cannot talk its way past, sandboxes that limit the blast radius, and the human step placed where it does the most good.
- Island 4
Evals and Observability
How you know an agent works: an eval set that grows from failures, graders you can trust, traces you can search, a taxonomy of failures, and the regression check that runs before every change.