CoursesDeep AI EngineeringInference and Quantization
Throughput Versus Latency
After this lesson you can explain why batching makes serving cheaper and each request slower, what continuous batching changes, and which number to optimise for which product.
Loading the lesson…
Loading the lesson…