CoursesBuilding AI Applications
Cost, Latency and Shipping
The decisions between a demo and a product: which model for which request, caching, streaming and latency, what may enter a prompt and a log, and the checks that run before and after launch.
- Lessons
- 7
- Exercises
- 40
- Minutes
- 41
- 1
Routing and Caching
After this lesson you can send each request to the cheapest model that does the job, cache what repeats, and prove the saving with numbers rather than hope.
DebugFill blankMultiple choice - 2
Streaming and Latency
After this lesson you can separate time-to-first-token from total time, decide when to stream, and shave the latency the user actually feels.
Code orderMultiple choiceTrace - 3
The Stream on the Wire
After this lesson you can read a raw server-sent event stream, parse it by event rather than by network chunk, assemble partial tool calls safely, and forward a stream through your own server without holding the API key in the browser.
DebugCode orderMultiple choiceTrace - 4
What May Enter a Prompt
After this lesson you can decide what user data may be sent to a model, what may be logged, how long it may be kept, and how to say so honestly.
DebugMultiple choiceTap - 5
Before and After Launch
After this lesson you can put the small eval, the fallback and the monitors in place that turn a working demo into something you can leave running.
DebugMultiple choiceTrace - 6
Checkpoint: ProductionCheckpoint
Routing, latency, privacy and shipping in fresh situations.
Fill blankMultiple choice - 7
Boss: Launch WeekBoss
One launch from routing plan to the first quality dip: models, caches, streaming, privacy, an outage, a runaway user, and the number that moved.
DebugCode orderMultiple choiceTrace