MethodologySource architecture

A visible chain from source to screen.

Model VS keeps incompatible evidence separate, validates structured snapshots before publication, and states the boundary of every signal.

Publication pathAuditable
01 / SRCNamed public source
02 / QASchema and integrity checks
03 / PUBTimestamped static snapshot
Model evidence

Each benchmark answers a different question.

Model VS does not blend incompatible scores into a proprietary overall number.

Human preference

LM Arena

Style-controlled ratings derive from human preference votes. Vote counts and confidence information remain part of the reading context.

Objective tasks

LiveBench

Overall is the unweighted mean of category averages from the named release; category and task results remain attributable to that release.

Narrow capability

Specialist matrix

SWE-bench Verified and GPQA records are synchronized from the Hugging Face official leaderboard API with submitted evidence and verification state.

Publication pipeline

Build-time checks guard the snapshot.

GET

Retrieve structured sources

Synchronization jobs collect the benchmark and intelligence records used by the site.

QA

Validate before build

Schema, source identity, freshness, uniqueness, and material record-loss checks prevent malformed snapshots from silently replacing verified data.

HOLD

Preserve verified benchmark data on sync failure

If benchmark synchronization fails, the build process keeps the last verified LM Arena, LiveBench, and specialist snapshots instead of inventing values.

SHOW

Publish source-attributed output

The interface reads static JSON snapshots and keeps source, release, and verification context close to the result.

AI Radar

Discovery signals from named public sources.

AI Radar organizes public records across GitHub projects and releases, Hugging Face models and datasets, the Official MCP Registry, and arXiv research. Each collection exposes its own source and signal type.

Inclusion is not endorsement

Recency, popularity, and repository activity help people discover what is moving. They do not establish model quality, safety, or production readiness.

Skills index

Public Agent Skills metadata, indexed for review.

The Skills build searches public GitHub repositories, validates each SKILL.md against name and description requirements, removes forks and archived repositories, and groups valid records with deterministic keyword rules.

Repository metadata

Popularity stays labeled

Stars and forks describe the containing repository. They are not presented as a skill-quality score.

Execution boundary

Scripts are not run

Third-party scripts are indexed as a review warning and are never executed by Model VS.

Limitations

Evidence has edges.

No benchmark is a universal measure of quality. Public submissions may be unverified, benchmark contamination can occur, and results can change with prompts, tools, scaffolds, sampling, and evaluation versions.

A specialist result may include an agent scaffold, harness, inference configuration, or test-time strategy; it must not be interpreted automatically as the base model's standalone ability.

Use the site as a decision instrument

Shortlist with the available evidence, then test the candidate systems on your own data, constraints, and failure cases.