Find a benchmark
Search by benchmark or model name, browse one of the seven fields, or open a landmark on the map.
BenchAtlas turns model cards and technical reports into an explorable benchmark landscape. Use it to find reported scores, identify which results are genuinely comparable, and inspect the setup behind every number.
Replay the 3-step product tour →A useful comparison usually takes four steps. Begin with the landscape, then narrow the evidence before reading the score.
Search by benchmark or model name, browse one of the seven fields, or open a landmark on the map.
Use the evaluation protocol selector in the details panel. Only rows sharing that setup are ranked together.
Check reasoning mode, context length, tools, harness, judge, run count, aggregation, and dataset variant.
Follow the primary-source link and evidence location when a number will influence a decision or publication.
The four public layers answer different questions. Map eligibility is only a presentation rule and does not create another data entity.
A canonical benchmark identity that combines naming aliases, metrics, and meaningful variants.
A normalized benchmark, metric, and meaningful variant. Each group maps to one result page.
Reported rows that share a sufficiently consistent evaluation setup and can be ranked directly.
One model or configuration score with its source report, evidence location, and method notes.
A spatial overview of high-coverage benchmarks. Closer to a field center means broader model coverage and stronger source evidence; direction indicates subfield.
A scan-friendly table for benchmark result groups, including model coverage, vendors, best reported score, and protocol signals.
A compact model-by-benchmark view based on one documented protocol group per benchmark. Blank cells mean no compatible reported result.
Select a base model to highlight the landmarks it covers. The displayed score uses that model's best public configuration within the selected protocol.
The landscape shows a curated set of landmarks, not every benchmark in the catalog. The upper-left map label states how many landmarks are shown out of the full registry.
BenchAtlas does not place every score with the same benchmark name into one table. A ranking is formed only inside a protocol group that shares a documented evaluation setup.
Rows have enough matching setup information to support direct comparison.
Rows come from one report or table but the full setup is not sufficiently normalized for strict cross-report comparison.
A visible warning that benchmark name alone is not enough to establish comparability.
The overall ranking includes every reported base model without minimum family, field, model-count, or vendor-count thresholds. Comparison groups contribute normalized rank percentiles; a single-model group contributes a neutral 50. Repeated appearances within one Benchmark family are averaged, then all observed family scores are averaged directly. Limited family coverage is shrunk toward 50.
Family count, capability-field coverage, and independent report count determine whether a result is high, medium, or provisional confidence. A higher score means stronger average placement in the observed leaderboards, not universal model superiority.
Agent systems, CLI products, checkpoints, computational baselines, and composite indexes remain visible as reference evidence but do not enter the overall ranking.
Select a reported source in the details panel to see the configuration attached to that exact score. Notes are organized into six evidence types.
Harness, task environment, timeout, compute, and implementation details.
Reasoning mode, context window, sampling parameters, and token limits.
Agent framework, tools, browser, terminal, repository access, and restrictions.
Public or internal set, corrected tasks, subset, language coverage, or benchmark version.
Number of attempts, averaging, pass@k, voting, Elo, or judge aggregation.
Missing details, unavailable scores, internal adaptations, or other comparison limits.
Every result retains its source report, source URL, evidence location, and a short quote where available. Different reports may publish different scores for the same model and benchmark; BenchAtlas preserves those rows instead of silently overwriting them.
Switch between report rows to inspect score differences and their attached configurations.
The URL hash records the selected benchmark, protocol, model filter, and view. Use Share to prepare a link to the current state.