Model VS evidence index

Compare AI models with separate, traceable evidence

Model VS keeps user preference, objective benchmark results, and specialist task data separate. It does not create a synthetic score or declare an overall winner.

Sources and current snapshots

LM Arena dataset lmarena-ai/leaderboard-dataset was synced September 19, 2026; leaderboard published September 13, 2026. LiveBench release 2026-06-25 was synced September 19, 2026. Specialist data comes from Hugging Face official leaderboard API, synced September 19, 2026.

Method: Source records are shown in their native measurement systems. Rank, user-preference rating, benchmark scores, and task-specific percentages are not converted into one score.

Limits: Benchmark measurements are snapshots, not guarantees of real-world behavior. They can change with source updates and should be considered with task requirements, pricing, safety, and licensing.

Current LM Arena records

claude-fable-5

LM Arena overall rank: 1; rating: 1505.68; votes: 30,057.

Organization: anthropic; published September 13, 2026.

claude-opus-4-6-high

LM Arena overall rank: 2; rating: 1504.56; votes: 71,993.

Organization: anthropic; published September 13, 2026.

claude-opus-4-7-high

LM Arena overall rank: 3; rating: 1501.75; votes: 60,002.

Organization: anthropic; published September 13, 2026.

muse-spark-1.2 (xHigh)

LM Arena overall rank: 4; rating: 1499.59; votes: 3,227.

Organization: meta; published September 13, 2026.

Current LiveBench records

claude-opus-4-5-20251101-thinking-64k-high-effort

LiveBench raw metrics: AMPS Hard 99.0; code completion 80.435; code generation 78.873.

Release: 2026-06-25; synced September 19, 2026.

claude-opus-4-6-thinking-auto-high-effort

LiveBench raw metrics: AMPS Hard 97.0; code completion 76.087; code generation 80.282.

Release: 2026-06-25; synced September 19, 2026.

claude-opus-4-7-xhigh-effort

LiveBench raw metrics: AMPS Hard 98.0; code completion 78.261; code generation 85.915.

Release: 2026-06-25; synced September 19, 2026.

claude-sonnet-4-6-thinking-auto-medium-effort

LiveBench raw metrics: AMPS Hard 76.0; code completion 78.261; code generation 80.282.

Release: 2026-06-25; synced September 19, 2026.

Current specialist benchmark records

ornith-ai/Ornith-1.5-397B

SWE-bench Verified: rank 1; % resolved 86.

Community submission — not marked verified. Dataset: SWE-bench/SWE-bench_Verified; source: Ornith-1.5-397B model card; benchmark data updated August 16, 2026.

mindlab-research/Macaron-V1-Venti

SWE-bench Verified: rank 2; % resolved 85.6.

Community submission — not marked verified. Dataset: SWE-bench/SWE-bench_Verified; source: Model Card; benchmark data updated August 16, 2026.

mindlab-research/Macaron-V1-Coding-Venti

SWE-bench Verified: rank 3; % resolved 85.6.

Community submission — not marked verified. Dataset: SWE-bench/SWE-bench_Verified; source: Model Card; benchmark data updated August 16, 2026.

ornith-ai/Ornith-1.0-397B

SWE-bench Verified: rank 4; % resolved 82.4.

Community submission — not marked verified. Dataset: SWE-bench/SWE-bench_Verified; source: Ornith-1.0-397B model card; benchmark data updated August 16, 2026.

Open the interactive comparison to filter sources and inspect the complete evidence record.