A calmer way to read a noisy market.
The goal is not to crown one model as the universal winner. Model VS keeps distinct questions distinct: human preference, objective task performance, specialist capability, and ecosystem activity each need their own evidence.
Use rankings to narrow the field, compare evidence on its native scale, then test the shortlist on your own workload.
Every surface has a defined job.
LM Arena
Preference ratings, vote counts, and confidence information provide one view of overall user experience.
LiveBench
Release-specific task and category results support a separate view of measurable performance.
Specialist benchmarks
Named records such as SWE-bench Verified and GPQA stay attached to their benchmark context and submitted evidence.
AI Radar & Skills
Public project, release, research, dataset, MCP, and GitHub Skills metadata supports discovery—not a quality score.
Legibility is part of accuracy.
Keep native scales
Scores from incompatible evaluations are not blended into a proprietary overall number.
Name the source
Benchmark releases, public records, and repository metadata remain attributable to their origin.
Separate signal from endorsement
Recency, stars, forks, and release activity describe attention or freshness; they do not prove quality.
Show the boundary
Benchmarks help with shortlisting, but they cannot replace evaluation on a real workload.
Corrections make the evidence stronger.
Questions, source corrections, or partnership enquiries can be sent to [email protected].