Intermediate Lesson 7 of 10

batching

Processing several requests together in one go, to use the hardware more efficiently.

A GPU can handle many inputs at the same time almost as fast as one. Batching groups incoming requests and sends them through the model together. Each request may wait a moment longer, but the cost per request drops, often sharply. Systems tune the batch size to balance each user's wait against total cost.

Think of it as

Like a minibus that waits a minute to fill up before leaving. Each passenger waits slightly longer, but the trip costs far less per person than a taxi each.

Example

A company tags a large pile of product reviews overnight. Sending them to the model in groups rather than one at a time finishes the job much sooner on the same GPU.