Prompt Injection
Instructions hidden in data, the three ingredients that turn them into a breach, how to see an attack in a trace, and the designs that make it survivable rather than the filters that promise to stop it.
- Lessons
- 5
- Exercises
- 29
- Minutes
- 31
- 1
Instructions in the Data
After this lesson you can explain why a model cannot reliably tell your instructions from instructions hidden in a web page or an email, and why that is a property, not a bug to be patched.
DebugFill blankMultiple choice - 2
The Three Ingredients
After this lesson you can look at an agent's tools and say whether an injection can become a breach, using the three ingredients an attacker needs.
Code orderMultiple choiceTrace - 3
Designing Around It
After this lesson you can choose the designs that keep an injected agent harmless: gate the ways out, separate the reader from the actor, limit what each reads, and check outputs against what was allowed.
DebugMultiple choiceTrace - 4
Checkpoint: InjectionCheckpoint
Sources, ingredients and designs in fresh situations.
Fill blankMultiple choiceTrace - 5
Boss: The Inbox AgentBoss
You are reviewing an inbox agent before launch. It triages mail, drafts replies, books meetings and files attachments. One agent with every ingredient, made safe one decision at a time: the tool inventory, the hidden ways out, the split, the scope, the checks, and the attack that arrives on day one.
DebugFill blankMultiple choiceTrace