How we rateEvidence protocol

One score cannot answer every question.

Model VS presents independent evidence lanes and explains what each can—and cannot—support.

Interpretation pathBounded
01 / SRCIdentify the evaluation source
02 / UNITKeep its native score and release
03 / USEMatch evidence to the decision
Evidence lanes

Three model views, kept deliberately separate.

Preference

LM Arena

Human preference ratings describe comparative user preference. Vote counts and confidence information provide essential context.

Objective

LiveBench

Task and category results come from a named public release. Overall is the unweighted mean of its published category averages.

Specialist

Named benchmarks

SWE-bench Verified and GPQA records preserve their narrower benchmark context and submitted evidence.

Reading rules

Interpret before you compare.

01

Do not add incompatible scores

Human preference, general task performance, and specialist evaluations are not averaged together.

02

Read the release context

A result belongs to a named evaluation version and can change with prompts, tools, scaffolds, sampling, or test-time strategy.

03

Distinguish model from system

A specialist submission may include an agent scaffold, harness, inference configuration, or other system-level support.

04

Validate the shortlist

A leaderboard position is evidence for shortlisting, not a guarantee on your data, latency budget, or deployment constraints.

Discovery signals

Radar and Skills are indexes, not verdicts.

AI Radar groups records from named public sources across projects, releases, models, datasets, MCP, and research. The Skills index organizes valid public Agent Skills metadata from GitHub repositories.

Popularity is not quality

Stars, forks, recency, repository activity, and inclusion in an index are discovery signals. They are not model or skill quality scores.