Deep AI Engineering
What is under the assistant: tokens, attention and training on one side, decoding, caches, throughput and quantization on the other, for the engineer who has to explain a failure or size a deployment.
- Islands
- 2
- Lessons
- 13
- Exercises
- 74
DebugCode orderFill blankMultiple choiceTapTrace
- Island 1
Tokens, Attention and Training
Enough of the machinery to explain a failure: how text becomes tokens, what attention does with them, what the training stages put in and leave out, and the mechanical reasons behind the mistakes the earlier courses taught you to catch.
- Island 2
Inference and Quantization
What happens when a model runs: how the next token is chosen, why the first token is slow and the rest are fast, what limits throughput, and what quantization buys and costs, for the engineer sizing a deployment.