Guardrails and the Human in the Loop
Least privilege for tools, the approval boundary that a model cannot talk its way past, sandboxes that limit the blast radius, and the human step placed where it does the most good.
- Lessons
- 6
- Exercises
- 35
- Minutes
- 37
- 1
Least Privilege for Tools
After this lesson you can size an agent's tools to its task: which tools, which scope, which arguments it may choose, and which it may not.
DebugFill blankMultiple choice - 2
The Approval Boundary
After this lesson you can place the human step at the irreversible line and design it so the person sees what they are approving, rather than clicking yes to a summary.
DebugMultiple choiceTrace - 3
Sandboxes and Blast Radius
After this lesson you can bound what an agent can break: where it runs, what it can reach, how much it can spend, and how you undo what it did.
DebugCode orderMultiple choice - 4
MCP Servers Are Tool Lists
After this lesson you can connect a tool server to an agent and size it the way you size any tool list: allow-list the tools, scope the credential, treat the descriptions as untrusted text, and count the ingredients the server adds.
DebugFill blankMultiple choice - 5
Checkpoint: GuardrailsCheckpoint
Privilege, approvals and blast radius in fresh situations.
DebugFill blankMultiple choice - 6
Boss: The Deploy BotBoss
You are designing a deploy bot. Engineers ask it in chat to ship a service version; it runs the checks, opens the change, and deploys. One agent that touches production, made safe: sized tools, bound arguments, an approval that shows the diff, credentials that expire, a rollback, and the message that tries to skip all of it.
DebugFill blankMultiple choiceTrace