benchmark
A standard test used to measure and compare how well models perform.
A benchmark is a fixed set of tasks with known correct answers. Running different models on it gives each a score, so they can be compared fairly. Scores are useful, but a model that tops a public benchmark may still do worse on your own data. The best check is a small test on your own real cases.
Think of it as
Like a standard driving test. It shows a driver meets a common bar, but not how well they will handle your own daily route in rush hour.
Example
Two tools both claim top scores on a public benchmark. A team runs each on a hundred of its own past support tickets instead, and finds one gets clearly more right. That is the result that matters to them.