Ranking Methodology
How we decide which AI models are the most useful — and why "useful" is the metric that matters.
Our definition of usefulness
A model is useful to the extent that a competent practitioner can get valuable work done with it, today, at a price and in a form they can actually deploy. That definition deliberately favours shipping models over paper results, and deliberately penalises models that are brilliant but inaccessible.
Scoring
Each model receives a 0–100 score on five dimensions. The dimensions are weighted (Capability 35%, Reliability 20%, Efficiency 15%, Accessibility 15%, Ecosystem 15%) and combined into a single Usefulness Index. Ties are broken by Capability, then Reliability.
Capability (35%)
We run a fixed internal task suite covering reasoning, writing, coding, multimodal understanding and generation quality as appropriate to the category, and combine it with reputable independent leaderboards and blind pairwise community preference data. Vendor-reported benchmarks are not used as inputs.
Reliability (20%)
Consistency across five repeated runs of each task, measured hallucination and refusal rates, API uptime over the edition period, and whether versioned endpoints are available so behaviour doesn't change under users.
Efficiency (15%)
Normalised cost per unit of output (tokens, images, seconds of video or audio), median and p95 latency, and for open models the minimum hardware needed to self-host at acceptable quality.
Accessibility (15%)
Availability (open weights, public API, or both), licence permissiveness, regions and languages served, sign-up friction, and documentation quality.
Ecosystem (15%)
Support in major inference engines and frameworks, number and quality of community fine-tunes, SDKs, plugins, and the velocity of the surrounding community.
Categories
Models are grouped into eight categories so that, for example, a speech model isn't compared on coding tasks. Category-specific task suites feed the Capability score; the other four dimensions are measured the same way across categories. The overall Top 100 is a single list because practitioners choose between categories too.
Cadence and archive
Scores are recomputed monthly. Each edition is archived with its full score table so movements can be audited. Models can enter, exit or move freely between editions.
Independence
Editors do not have access to sponsorship data while scoring. Sponsorship status is appended after scores are locked, purely for display. See the Editorial Policy.
Limitations
Our task suite reflects our editors' judgement of common workloads and will not match every use case. Scores for very new models carry higher uncertainty. We publish the methodology so you can disagree with it precisely.
Changelog
October 2026 — Launched the eight-category structure and the Million Pixel Canvas. Weights unchanged from the pilot edition.