CoursesDeep AI EngineeringInference and Quantization

Prefill, Decode and the Cache

After this lesson you can explain why the first token is slow and the rest are fast, what the KV cache stores and why it fills memory, and why prefix caching works.

DebugFill blankMultiple choice5 exercises, 6 min

Loading the lesson…