By Model VS · Guide text updated . Worked examples are calculated from the site's saved source snapshots; each shows its own retrieval date. These are interpretations of published results, not tests run by Model VS. Report a correction with the source record and the result you expected.
Write the acceptance rule before looking at rank.
“Best AI model” leaves out the work, the constraints, and the cost of being wrong. Write one sentence describing the output you need and one describing what would make that output unacceptable. For a support assistant, that might mean answering only from the supplied policy and declining when the answer is missing. For a coding assistant, it might mean producing a patch that passes the existing tests without changing the public interface. These are evaluation examples, not claims about any listed model.
Next list hard constraints: permitted data handling, deployment region, tool access, context requirements, and budget. Verify them in the provider's current documentation. Our benchmark snapshots do not establish those operational facts. Exclude candidates that fail a hard constraint before using a leaderboard to narrow the remaining set.
Choose the evidence that resembles the task. Conversational preference can help shortlist a general assistant. A coding metric can help shortlist a code tool. A repository-repair submission is relevant only when its agent setup resembles the system you intend to run. A candidate can belong on your shortlist without leading every table.
Three model views, kept deliberately separate.
LM Arena
Human preference ratings describe comparative user preference. Vote counts and confidence information provide essential context.
LiveBench
Task and category results come from a named public release. Overall is the unweighted mean of its published category averages.
Named benchmarks
SWE-bench Verified and GPQA records preserve their narrower benchmark context and submitted evidence.
A small rating gap needs context.
These are the first two ranked records in the saved Arena snapshot, published 2026-09-13 and retrieved 2026-09-19T07:21:27.516Z. Model VS rounds the display to two decimals; the linked JSON retains source precision.
claude-fable-5
Source rank: 1. Rating: 1505.68.
Source 95% interval: 1500.92–1510.45.
Votes: 30057.
claude-opus-4-6-high
Source rank: 2. Rating: 1504.56.
Source 95% interval: 1501.04–1508.08.
Votes: 71993.
Our calculation: the first rating minus the second is 1.12 rating points. Their reported intervals overlap. Interval overlap alone is not a pairwise significance test: the covariance and evaluation design matter. This snapshot cannot establish a universal winner or a probability that one model will solve your task.
Decision: keep both candidates if they meet your constraints. Use preference evidence for shortlisting conversational systems, then compare errors on your own prompts. A greater vote count does not guarantee that the voting population or prompts match your users.
Inspect the task metrics behind the category.
Below are the first two complete coding records in source order from our saved LiveBench snapshot, release 2026-06-25, retrieved 2026-09-19T07:21:06.677Z. This is an arithmetic demonstration, not an editorial top-two selection. Full model identifiers retain the source's effort and thinking configuration.
claude-opus-4-5-20251101-thinking-64k-high-effort
Code completion: 80.44.
Code generation: 78.87.
claude-opus-4-6-thinking-auto-high-effort
Code completion: 76.09.
Code generation: 80.28.
Our calculation: first record minus second record = 4.35 points for code completion and -1.41 points for code generation. A positive difference favors the first record on that metric; a negative difference favors the second. These differences use the unrounded source values. They are score-point differences, not percentages of productivity improvement.
Decision: use the metric closest to your workflow. Filling an existing function and producing a solution from a specification are different tasks. Neither raw metric measures your repository's end-to-end bug-fixing success, and neither is the site's aggregated Coding category score. Do not drop the effort suffix and assume the same result applies to a cheaper or faster configuration.
A benchmark name is not a submission certificate.
The first record in the saved SWE-bench Verified submission snapshot is ornith-ai/Ornith-1.5-397B, reporting 86.00 % resolved. Community submission — not marked verified. Its evidence is the linked Ornith-1.5-397B model card. Snapshot retrieved: 2026-09-19T07:21:06.403Z.
Interpretation: “Verified” in the benchmark name identifies the benchmark subset; it does not certify every submitted score. The separate submission flag above is what our source supplies. Even a source-verified submission is not a result independently reproduced by Model VS.
Decision: open the submitted evidence and check the exact model, agent scaffold, tool access, test split, attempt count, and inference budget before using it in a shortlist. If those details are absent, record them as unknown. Do not treat a system's resolved-issue percentage as the probability of fixing your next production bug.
Compare complete workflows on the same tasks.
Prepare representative cases
Use work you have permission to test: ordinary cases, ambiguous inputs, and known failure cases. Keep a separate holdout set that you do not use to tune prompts. Record the expected output or an explicit rubric before generating answers. A small pilot can reveal obvious problems; it does not establish a population-wide success rate.
Freeze the comparison setup
Save the exact model identifier and date, prompts, supplied context, tool versions, sampling settings, reasoning budget, and retry policy. Give each candidate the same permitted inputs. If one system needs a different scaffold, label this as a comparison of configured systems and account for the extra work and cost.
Score outcomes and retain failures
For code, run the tests and inspect the patch. For grounded answers, check each factual claim against the supplied material and record unsupported claims. For writing, use a predefined rubric and, where feasible, hide model identities from the reviewer. Retain failed, refused, and timed-out attempts in the denominator. Repeated runs on the same prompt are not independent new tasks.
Measure the cost of an accepted result
Record actual usage charges, retries, elapsed time, and human correction time separately. Divide observed usage spend by the number of accepted outputs to estimate cost per accepted result for this workload. If no outputs pass, report no accepted result rather than a zero cost. Use measured billing units and current provider prices; a context-window limit is not an estimate of useful output or value per dollar.
Choose, or keep the result inconclusive
Reject candidates that fail a required privacy or correctness rule even if their average score is higher. If the remaining candidates perform similarly on a small sample, report the uncertainty and collect more representative cases. Recheck after a model version, prompt, tool, or workload change. Keep a fallback for the failure types you observed.
Save: task and acceptance rule; candidate and configuration; source evidence and date; case identifier; pass/fail and reason; elapsed time; usage spend; retries; reviewer corrections; unresolved risks. This turns a shortlist into an auditable decision without inventing a universal score.
Interpret before you compare.
Do not add incompatible scores
Human preference, general task performance, and specialist evaluations are not averaged together.
Read the release context
A result belongs to a named evaluation version and can change with prompts, tools, scaffolds, sampling, or test-time strategy.
Distinguish model from system
A specialist submission may include an agent scaffold, harness, inference configuration, or other system-level support.
Validate the shortlist
A leaderboard position is evidence for shortlisting, not a guarantee on your data, latency budget, or deployment constraints.
Radar and Skills are indexes, not verdicts.
AI Radar groups records from named public sources across projects, releases, models, datasets, MCP, and research. The Skills index organizes valid public Agent Skills metadata from GitHub repositories.
For a project, follow the record to its license, release history, and issue tracker. Check whether recent activity concerns the feature you need, rather than treating a large star count as proof of maintenance. For a Skill or MCP server, inspect its instructions, dependencies, network destinations, and requested permissions before trying it with disposable data. A valid metadata file only establishes that the package meets our indexing format.
For a paper, distinguish an abstract's claim from a reproduced result and look for evaluation data and code. For a dataset, read provenance, consent, license, and split documentation. If essential information is missing, retain that as a reason to postpone adoption. The index helps you find the source; the adoption decision still requires this review.
Stars, forks, recency, repository activity, and inclusion in an index are discovery signals. They are not model or skill quality scores.
Source documentation: Arena leaderboard dataset (CC BY 4.0), LiveBench, SWE-bench, and Hugging Face leaderboard data guide. Model VS contributes the calculations, interpretation, and evaluation workflow above; upstream publishers supply the measurements.