Model VS evidence index

Compare AI models with separate, traceable evidence

Model VS keeps user preference, objective benchmark results, and specialist task data separate. It does not create a synthetic score or declare an overall winner.

Sources and current snapshots

LM Arena dataset lmarena-ai/leaderboard-dataset was synced September 7, 2026; leaderboard published September 2, 2026. LiveBench release 2026-06-25 was synced September 7, 2026. Specialist data comes from Hugging Face official leaderboard API, synced September 7, 2026.

Method: Source records are shown in their native measurement systems. Rank, user-preference rating, benchmark scores, and task-specific percentages are not converted into one score.

Limits: Benchmark measurements are snapshots, not guarantees of real-world behavior. They can change with source updates and should be considered with task requirements, pricing, safety, and licensing.

Current LM Arena records

claude-fable-5

LM Arena overall rank: 1; rating: 1507.16; votes: 27,189.

Organization: anthropic; published September 2, 2026.

claude-opus-4-6-high

LM Arena overall rank: 2; rating: 1504.84; votes: 72,099.

Organization: anthropic; published September 2, 2026.

claude-fable-5.1-max

LM Arena overall rank: 3; rating: 1504.21; votes: 2,906.

Organization: anthropic; published September 2, 2026.

claude-opus-4-7-high

LM Arena overall rank: 4; rating: 1502.11; votes: 60,103.

Organization: anthropic; published September 2, 2026.

Current LiveBench records

claude-opus-4-5-20251101-thinking-64k-high-effort

LiveBench raw metrics: AMPS Hard 99.0; code completion 80.435; code generation 78.873.

Release: 2026-06-25; synced September 7, 2026.

claude-opus-4-6-thinking-auto-high-effort

LiveBench raw metrics: AMPS Hard 97.0; code completion 76.087; code generation 80.282.

Release: 2026-06-25; synced September 7, 2026.

claude-opus-4-7-xhigh-effort

LiveBench raw metrics: AMPS Hard 98.0; code completion 78.261; code generation 85.915.

Release: 2026-06-25; synced September 7, 2026.

claude-sonnet-4-6-thinking-auto-medium-effort

LiveBench raw metrics: AMPS Hard 76.0; code completion 78.261; code generation 80.282.

Release: 2026-06-25; synced September 7, 2026.

Current specialist benchmark records

ornith-ai/Ornith-1.5-397B

SWE-bench Verified: rank 1; % resolved 86.

Dataset: SWE-bench/SWE-bench_Verified; source: Ornith-1.5-397B model card; benchmark data updated August 16, 2026.

mindlab-research/Macaron-V1-Venti

SWE-bench Verified: rank 2; % resolved 85.6.

Dataset: SWE-bench/SWE-bench_Verified; source: Model Card; benchmark data updated August 16, 2026.

mindlab-research/Macaron-V1-Coding-Venti

SWE-bench Verified: rank 3; % resolved 85.6.

Dataset: SWE-bench/SWE-bench_Verified; source: Model Card; benchmark data updated August 16, 2026.

ornith-ai/Ornith-1.0-397B

SWE-bench Verified: rank 4; % resolved 82.4.

Dataset: SWE-bench/SWE-bench_Verified; source: Ornith-1.0-397B model card; benchmark data updated August 16, 2026.

Open the interactive comparison to filter sources and inspect the complete evidence record.