Three model views, kept deliberately separate.
LM Arena
Human preference ratings describe comparative user preference. Vote counts and confidence information provide essential context.
LiveBench
Task and category results come from a named public release. Overall is the unweighted mean of its published category averages.
Named benchmarks
SWE-bench Verified and GPQA records preserve their narrower benchmark context and submitted evidence.
Interpret before you compare.
Do not add incompatible scores
Human preference, general task performance, and specialist evaluations are not averaged together.
Read the release context
A result belongs to a named evaluation version and can change with prompts, tools, scaffolds, sampling, or test-time strategy.
Distinguish model from system
A specialist submission may include an agent scaffold, harness, inference configuration, or other system-level support.
Validate the shortlist
A leaderboard position is evidence for shortlisting, not a guarantee on your data, latency budget, or deployment constraints.
Radar and Skills are indexes, not verdicts.
AI Radar groups records from named public sources across projects, releases, models, datasets, MCP, and research. The Skills index organizes valid public Agent Skills metadata from GitHub repositories.
Stars, forks, recency, repository activity, and inclusion in an index are discovery signals. They are not model or skill quality scores.