AboutIndependent observatory

Evidence before verdicts.

Model VS helps people compare AI systems with source-attributed evidence instead of an opaque, all-purpose score.

Evidence pathTraceable
01 / INNamed public sources
02 / QAValidated static snapshots
03 / OUTClear, bounded interpretation
Purpose

A calmer way to read a noisy market.

The goal is not to crown one model as the universal winner. Model VS keeps distinct questions distinct: human preference, objective task performance, specialist capability, and ecosystem activity each need their own evidence.

Our decision rule

Use rankings to narrow the field, compare evidence on its native scale, then test the shortlist on your own workload.

Evidence map

Every surface has a defined job.

Human preference

LM Arena

Preference ratings, vote counts, and confidence information provide one view of overall user experience.

Objective tasks

LiveBench

Release-specific task and category results support a separate view of measurable performance.

Narrow capability

Specialist benchmarks

Named records such as SWE-bench Verified and GPQA stay attached to their benchmark context and submitted evidence.

Ecosystem signals

AI Radar & Skills

Public project, release, research, dataset, MCP, and GitHub Skills metadata supports discovery—not a quality score.

Operating principles

Legibility is part of accuracy.

01

Keep native scales

Scores from incompatible evaluations are not blended into a proprietary overall number.

02

Name the source

Benchmark releases, public records, and repository metadata remain attributable to their origin.

03

Separate signal from endorsement

Recency, stars, forks, and release activity describe attention or freshness; they do not prove quality.

04

Show the boundary

Benchmarks help with shortlisting, but they cannot replace evaluation on a real workload.

Contact

Corrections make the evidence stronger.

Questions, source corrections, or partnership enquiries can be sent to [email protected].