Each benchmark answers a different question.
Model VS does not blend incompatible scores into a proprietary overall number.
LM Arena
Style-controlled ratings derive from human preference votes. Vote counts and confidence information remain part of the reading context.
LiveBench
Overall is the unweighted mean of category averages from the named release; category and task results remain attributable to that release.
Specialist matrix
SWE-bench Verified and GPQA records are synchronized from the Hugging Face official leaderboard API with submitted evidence and verification state.
Build-time checks guard the snapshot.
Retrieve structured sources
Synchronization jobs collect the benchmark and intelligence records used by the site.
Validate before build
Schema, source identity, freshness, uniqueness, and material record-loss checks prevent malformed snapshots from silently replacing verified data.
Preserve verified benchmark data on sync failure
If benchmark synchronization fails, the build process keeps the last verified LM Arena, LiveBench, and specialist snapshots instead of inventing values.
Publish source-attributed output
The interface reads static JSON snapshots and keeps source, release, and verification context close to the result.
Discovery signals from named public sources.
AI Radar organizes public records across GitHub projects and releases, Hugging Face models and datasets, the Official MCP Registry, and arXiv research. Each collection exposes its own source and signal type.
Recency, popularity, and repository activity help people discover what is moving. They do not establish model quality, safety, or production readiness.
Public Agent Skills metadata, indexed for review.
The Skills build searches public GitHub repositories, validates each SKILL.md against name and description requirements, removes forks and archived repositories, and groups valid records with deterministic keyword rules.
Popularity stays labeled
Stars and forks describe the containing repository. They are not presented as a skill-quality score.
Scripts are not run
Third-party scripts are indexed as a review warning and are never executed by Model VS.
Evidence has edges.
No benchmark is a universal measure of quality. Public submissions may be unverified, benchmark contamination can occur, and results can change with prompts, tools, scaffolds, sampling, and evaluation versions.
A specialist result may include an agent scaffold, harness, inference configuration, or test-time strategy; it must not be interpreted automatically as the base model's standalone ability.
Shortlist with the available evidence, then test the candidate systems on your own data, constraints, and failure cases.