BenchAtlasUsage guide
BenchAtlas field manual

Read benchmark scores with their evaluation context.

BenchAtlas turns model cards and technical reports into an explorable benchmark landscape. Use it to find reported scores, identify which results are genuinely comparable, and inspect the setup behind every number.

Replay the 3-step product tour →
Reported results
Base models
Benchmark families
Source reports

Quick start

A useful comparison usually takes four steps. Begin with the landscape, then narrow the evidence before reading the score.

01

Find a benchmark

Search by benchmark or model name, browse one of the seven fields, or open a landmark on the map.

Search atlas → select landmark
02

Choose a protocol group

Use the evaluation protocol selector in the details panel. Only rows sharing that setup are ranked together.

Comparable ranking → protocol group
03

Inspect the configuration

Check reasoning mode, context length, tools, harness, judge, run count, aggregation, and dataset variant.

Method notes → reported source
04

Open the evidence

Follow the primary-source link and evidence location when a number will influence a decision or publication.

Source evidence → primary source

How the data is organized

The four public layers answer different questions. Map eligibility is only a presentation rule and does not create another data entity.

01

Benchmark family

A canonical benchmark identity that combines naming aliases, metrics, and meaningful variants.

families
02

Benchmark result group

A normalized benchmark, metric, and meaningful variant. Each group maps to one result page.

result groups
03

Comparable setup group

Reported rows that share a sufficiently consistent evaluation setup and can be ranked directly.

comparable groups
04

Reported result

One model or configuration score with its source report, evidence location, and method notes.

results

Views and controls

Landscape

A spatial overview of high-coverage benchmarks. Closer to a field center means broader model coverage and stronger source evidence; direction indicates subfield.

Catalog

A scan-friendly table for benchmark result groups, including model coverage, vendors, best reported score, and protocol signals.

Matrix

A compact model-by-benchmark view based on one documented protocol group per benchmark. Blank cells mean no compatible reported result.

Model filter

Select a base model to highlight the landmarks it covers. The displayed score uses that model's best public configuration within the selected protocol.

Map scope

The landscape shows a curated set of landmarks, not every benchmark in the catalog. The upper-left map label states how many landmarks are shown out of the full registry.

Comparable ranking

BenchAtlas does not place every score with the same benchmark name into one table. A ranking is formed only inside a protocol group that shares a documented evaluation setup.

Shared protocol

Rows have enough matching setup information to support direct comparison.

Source scoped

Rows come from one report or table but the full setup is not sufficiently normalized for strict cross-report comparison.

Protocol variant

A visible warning that benchmark name alone is not enough to establish comparability.

Reported Average Percentile

The overall ranking includes every reported base model without minimum family, field, model-count, or vendor-count thresholds. Comparison groups contribute normalized rank percentiles; a single-model group contributes a neutral 50. Repeated appearances within one Benchmark family are averaged, then all observed family scores are averaged directly. Limited family coverage is shrunk toward 50.

Confidence

Family count, capability-field coverage, and independent report count determine whether a result is high, medium, or provisional confidence. A higher score means stronger average placement in the observed leaderboards, not universal model superiority.

Excluded

Agent systems, CLI products, checkpoints, computational baselines, and composite indexes remain visible as reference evidence but do not enter the overall ranking.

Method notes

Select a reported source in the details panel to see the configuration attached to that exact score. Notes are organized into six evidence types.

Evaluation setup

Harness, task environment, timeout, compute, and implementation details.

Reasoning configuration

Reasoning mode, context window, sampling parameters, and token limits.

Agent / tool scaffold

Agent framework, tools, browser, terminal, repository access, and restrictions.

Dataset variant

Public or internal set, corrected tasks, subset, language coverage, or benchmark version.

Runs and aggregation

Number of attempts, averaging, pass@k, voting, Elo, or judge aggregation.

Source caveat

Missing details, unavailable scores, internal adaptations, or other comparison limits.

Evidence and sources

Every result retains its source report, source URL, evidence location, and a short quote where available. Different reports may publish different scores for the same model and benchmark; BenchAtlas preserves those rows instead of silently overwriting them.

Reported source

Switch between report rows to inspect score differences and their attached configurations.

Shareable state

The URL hash records the selected benchmark, protocol, model filter, and view. Use Share to prepare a link to the current state.

Key terms

Base modelThe canonical model identity used for model counts, filters, coverage, and overall ranking.
ConfigurationA public reasoning or inference setup attached to a base model, not a separate model.
Reference entityAn agent system, checkpoint, or baseline retained for context but excluded from model ranking.
Benchmark familyThe canonical benchmark identity across aliases, metrics, and meaningful variants.
Benchmark result groupA normalized benchmark, metric, and meaningful variant used as one result page.
Comparable setup groupRows that share one sufficiently consistent evaluation setup and may be directly ranked.
Reported resultOne score connected to its model entity, source, evidence, and method notes.