Benchmark
A benchmark is a standard test used to compare AI models — a set of questions, coding problems or tasks with known answers. Companies often announce new models with benchmark scores.
In one line, for a 12-year-old
A benchmark is a report card test that every AI takes so people can compare them.
An example
A lab says its model scores higher than rivals on a graduate-level science quiz and a software-repair test.
Why it matters to people
Scores can be inflated if test questions leaked into training data, and few benchmarks measure fairness or how well a model serves French speakers. Independent testing matters.