Intermediate Lesson 10 of 10

benchmark

A standard test used to measure and compare how well models perform.

A benchmark is a fixed set of tasks with known correct answers. Running different models on it gives each a score, so they can be compared fairly. Scores are useful, but a model that tops a public benchmark may still do worse on your own data. The best check is a small test on your own real cases.

Think of it as

Like a standard driving test. It shows a driver meets a common bar, but not how well they will handle your own daily route in rush hour.

Example

Two tools both claim top scores on a public benchmark. A team runs each on a hundred of its own past support tickets instead, and finds one gets clearly more right. That is the result that matters to them.