Agentic Wikibenchmarks

Reported evaluation results

Benchmark explorer

Compare model scores without treating unlike harnesses as the same test. Pin a model to follow its reported result across every available benchmark.

Current benchmark

SWE-bench Verified

% resolved (SWE-bench Verified, Vals AI harness)

Select model rows to pin them across benchmarks. Hover or focus a row for details.

Source: Vals AI — SWE-bench Verified leaderboard · as of 2026-07-12

Comparison mode

Selected across benchmarks

Select one or more model rows above to compare their reported results. Missing results stay visible as N/A.