# Benchmark score passport

Use this before quoting a model score.

- What precise capability is being claimed?
- Does the task resemble the decision or work you care about?
- Who wrote and reviewed the tasks?
- Which benchmark version, split, and date were tested?
- Which items are public, private, searchable, or likely to be in training data?
- What exact model version, prompt, system instruction, reasoning effort, tools, scaffold, and budget were used?
- Is the grader exact, executable, expert, model-based, preference-based, or outcome-based?
- How are refusals, failures, retries, fallbacks, and partial credit treated?
- What is the numerator, denominator, sample population, and confidence interval?
- Are leading results meaningfully separated?
- What did the evaluation cost, and what latency or compute did it consume?
- Has the benchmark saturated, changed, or been retired?
- Who funds and operates the evaluation?
- What does this score explicitly not establish?

MTS — How AI Gets Measured — editorial snapshot 26 August 2026.
