CoursesDeep AI EngineeringInference and Quantization
Prefill, Decode and the Cache
After this lesson you can explain why the first token is slow and the rest are fast, what the KV cache stores and why it fills memory, and why prefix caching works.
Loading the lesson…
Loading the lesson…