Courses

Deep AI Engineering

What is under the assistant: tokens, attention and training on one side, decoding, caches, throughput and quantization on the other, for the engineer who has to explain a failure or size a deployment.

Islands
2
Lessons
13
Exercises
74
DebugCode orderFill blankMultiple choiceTapTrace

Start the first lesson

  1. Island 1

    Tokens, Attention and Training

    Enough of the machinery to explain a failure: how text becomes tokens, what attention does with them, what the training stages put in and leave out, and the mechanical reasons behind the mistakes the earlier courses taught you to catch.

    1. Tokens Are Not Words5 exercises, 6 min
    2. Embeddings and Attention5 exercises, 6 min
    3. What Training Puts In5 exercises, 6 min
    4. Every Failure Has a Mechanism5 exercises, 6 min
    5. Checkpoint: The MachineryCheckpoint6 exercises, 6 min
    6. Boss: Explain the FailureBoss8 exercises, 10 min
  2. Island 2

    Inference and Quantization

    What happens when a model runs: how the next token is chosen, why the first token is slow and the rest are fast, what limits throughput, and what quantization buys and costs, for the engineer sizing a deployment.

    1. Choosing the Next Token5 exercises, 5 min
    2. Prefill, Decode and the Cache5 exercises, 6 min
    3. Throughput Versus Latency5 exercises, 6 min
    4. Quantization and Smaller Models5 exercises, 6 min
    5. Serving in Production5 exercises, 6 min
    6. Checkpoint: Running the ModelCheckpoint7 exercises, 6 min
    7. Boss: Size the DeploymentBoss8 exercises, 10 min