Reported evaluation results
Benchmark explorer
Compare model scores without treating unlike harnesses as the same test. Pin a model to follow its reported result across every available benchmark.
Current benchmark
SWE-bench Verified
% resolved (SWE-bench Verified, Vals AI harness)
Select model rows to pin them across benchmarks. Hover or focus a row for details.
Score 96.2 · CI not reportedSource: Vals AI — SWE-bench Verified leaderboard
Score 95 · CI not reportedSource: Vals AI — SWE-bench Verified leaderboard
Score 88.6 · CI not reportedSource: Vals AI — SWE-bench Verified leaderboard
Score 86.6 · CI not reportedSource: Vals AI — SWE-bench Verified leaderboard
Score 82.8 · CI not reportedSource: Vals AI — SWE-bench Verified leaderboard
Score 82.6 · CI not reportedSource: Vals AI — SWE-bench Verified leaderboard
Score 79.6 · CI not reportedSource: Vals AI — SWE-bench Verified leaderboard
Score 78.8 · CI not reportedSource: Vals AI — SWE-bench Verified leaderboard
Score 78.2 · CI not reportedSource: Vals AI — SWE-bench Verified leaderboard
Score 78 · CI not reportedParent model of GPT-5.3-Codex-Spark; Spark itself was absent from the leaderboard (checked 2026-07-12).Source: Vals AI — SWE-bench Verified leaderboard
Score 76.2 · CI not reportedSource: Vals AI — SWE-bench Verified leaderboard
Score 75.2 · CI not reportedSource: Vals AI — SWE-bench Verified leaderboard
Source: Vals AI — SWE-bench Verified leaderboard · as of 2026-07-12
Comparison mode
Selected across benchmarks
Select one or more model rows above to compare their reported results. Missing results stay visible as N/A.