Skip to content
Live
Loading the latest AI news…

Glossary · How models work

Benchmark

Also called Eval, evaluation

A benchmark is a standard test used to compare AI models — a set of questions, coding problems or tasks with known answers. Companies often announce new models with benchmark scores.

In one line, for a 12-year-old

A benchmark is a report card test that every AI takes so people can compare them.

An example

A lab says its model scores higher than rivals on a graduate-level science quiz and a software-repair test.

Why it matters to people

Scores can be inflated if test questions leaked into training data, and few benchmarks measure fairness or how well a model serves French speakers. Independent testing matters.