GenAI System Design
2. Inference, GPUs and Serving

Module 2 quiz

10 questions on inference, gpus and serving. Some ask for an estimate: answers within the stated tolerance count, just as in an interview.

1 / 10
Why is LLM decode at small batch sizes limited by memory bandwidth rather than compute?

Up next: Embeddings for System Designers