Each role uses a weighted average of outside evaluations on an Arena-like scale. It is a proxy for quality, not a prediction of pairwise wins. The boards use this exact calculation.
Only these rating sources contribute. Numeric scores must be finite and cannot be booleans; percentage scores must be between 0 and 100 inclusive. Arena text and coding are human preference ratings; Epoch AI provides ECI reasoning scores; Aider provides polyglot editing results; LiveBench provides global, coding, instruction-following and data-analysis scores; Terminal-Bench provides agent workflow accuracy. Multiple categories from one publisher count as one independent evaluator.
| Role | Rating family | Weight | Sources averaged within that family |
|---|---|---|---|
| Orchestrator | Objective reasoning | 50% | Epoch AI ECI, LiveBench global |
| Orchestrator | Human preference | 30% | Arena text |
| Orchestrator | Terminal work | 20% | Terminal-Bench |
| Mid-tier worker | Executed edits | 35% | Aider |
| Mid-tier worker | Terminal work | 35% | Terminal-Bench |
| Mid-tier worker | Objective coding | 20% | LiveBench coding |
| Mid-tier worker | Human preference | 10% | Arena coding |
| Rote worker | Routine tasks | 60% | LiveBench instruction following, data analysis |
| Rote worker | Objective coding | 25% | LiveBench coding |
| Rote worker | Human preference | 15% | Arena text |
The orchestrator weights also apply to the open-weight and home-computer boards. Published build prices use the mid-tier weights. Each stack entry uses its own role's weights.
Use each source's highest published raw score for that model instead of averaging reasoning efforts. If raw scores tie, choose the latest publication date, then the lexicographically greatest model spelling and URL. Explicit effort suffixes fold into the base model. A full date suffix folds into the base only when that source has exactly one dated release for that base. Versions, tiers, parameter sizes and checkpoint IDs remain separate. The chosen spelling can differ between sources.
Arena uses the latest publication in each of its overall and coding categories. Epoch uses rows with ECI scores. Aider uses pass_rate_2 from documented single-model whole, diff or diff-fenced runs with 224 or 225 tests, a commit and a date; architect/editor pairs and combined-model runs are excluded. Aider first breaks same-spelling score ties by date, edit format and commit spelling. LiveBench uses its latest linked dated CSV. If category summary columns are absent, it averages the mapped tasks within each category, then averages categories equally for the global score; an incomplete category contributes nothing. Terminal-Bench uses only the latest benchmark version from one pinned repository commit. Same-spelling runs with incompatible agent or agent-version configurations are excluded; otherwise the latest date and URL win before configuration selection. These are workflow proxies, with the named agent retained in the evidence.
Arena text and coding scores stay unchanged. For each other source and release, match the selected models to their selected Arena text ratings. Remove excluded observations and duplicate model keys from the overlap. The fit needs at least 10 overlapping model keys.
mu = mean(raw scores in the overlap)
sigma = population standard deviation(raw scores in the overlap)
z_i = (raw_i - mu) / sigma
alpha = mean(Arena text ratings in the overlap)
beta = sum(z_i * (Arena_i - alpha)) / sum(z_i * z_i)
calibrated(raw) = alpha + beta * (raw - mu) / sigma
A source contributes nothing if the overlap is smaller than 10, sigma is zero, beta is zero or negative, or its observation is excluded. Unselected configurations also contribute nothing. The parameters are fitted when the snapshot is built and saved with it; rendering does not refit them. This calculation has no clipping, minimum quality floor or age multiplier.
For a model and role, use only allowed sources with a positive family weight, a selected configuration, a calibrated score and no exclusion. Average those calibrated scores equally within each family. Then apply the role weights only to the families present for that model:
family_mean[f] = mean(usable calibrated source scores in family f)
R = sum(weight[f] * family_mean[f] for present f) / sum(weight[f] for present f)
displayed score = round(R)
Missing families are omitted from both numerator and denominator; no score is imputed. Python's round rounds to the nearest integer, with exact halfway values going to the even integer. Ranking uses the unrounded R, so two identical displayed integers need not be a tie. Exact equal unrounded scores share a competition rank: rank = 1 + number of ranked rows with a strictly greater R (1, 2, 2, 4). Tied rows are displayed in ascending model spelling order. Price never breaks a quality tie. We never divide quality by dollars.
A model with exactly one contributing independent evaluator goes into New, not enough ratings yet below its board, without a rank. A model with no usable family has no score or rank. Otherwise, fewer than two rating families, fewer than two independent evaluators, or less than 50% of the available family weight adds a partial label; it does not remove an otherwise scored model from ranking. Available families are those with at least one usable observation anywhere in the snapshot, before model or price filters. Coverage weight is the sum of a model's present preset weights divided by all preset weights; available share divides that sum by the snapshot's available weights. A board warns if less than 50% of its preset weight is available and names the missing families.
Movement compares the same model identity and role or machine to the previous snapshot's ranked field; tied ranks use the same rule. The first snapshot or a change of scoring method has no movement label. Rendering a saved snapshot from before the single-source rule suppresses its old movement labels when filtering changes that board's ranking.
The orchestrator board has no price or minimum quality filter. Open-weight models need an Arena weight licence that is present and not marked proprietary; individual licence terms still apply. PirateBench's latest scored epoch supplies our separate build-check column and list build prices, not rating points. A build price is the median of matching job prices when every matching job has a positive numeric list price. It is absent otherwise.
Both worker boards exclude names matching the flagship rule: the tokens opus, sol, astra or fable; names starting muse-spark; Gemini names with a pro token; and Grok names without mini or fast. If a build price exists it takes precedence over token price. Rote workers must cost at most $0.25 per build; mid-tier workers must cost strictly more than $0.25 and strictly less than the cheapest measured flagship build (with no upper bound if no priced flagship exists). Without a build price, rote workers need a positive blended token price at most $0.30 per million tokens; mid-tier workers need more than $0.30 and at most $3.
Token prices use LiteLLM chat entries without a colon in the name. Input and output prices must be finite, numeric, positive and not booleans. The blended price is 1,000,000 * (0.75 * input_cost_per_token + 0.25 * output_cost_per_token). For each folded identity, use the cheapest blended price, then the alphabetically first provider key. The Published build prices board lists every model with a positive measured build price, ordered by quality.
Best stack picks the first non-partial model in each component board, falling back to its first row. The value stack considers only OpenRouter price entries below $5, $1 and $0.30 per million tokens for the three roles, respectively (strict inequalities). It prefers non-partial coverage, then higher R, then model spelling. Its day-of-work estimate is 2 * orchestrator price + 10 * mid-tier price + 50 * rote price, shown only with all three roles. Single-evaluator picks appear in the strip.
Home-computer picks must be open-weight and have a known total checkpoint parameter count. Estimated 4-bit memory is parameters in billions * 0.5 * 1.25 GB, including all experts, and must fit within 80% of the machine's memory. The machines are RTX 3090 (24 GB), Strix Halo (128 GB) and Mac Studio (512 GB). Picks prefer non-partial coverage, then score, then spelling. Quality is measured by API evaluations, not on that hardware; speed is unmeasured.
This is a fixed example from the newest board data on disk when this page was prepared: the 10 October 2026 snapshot, pulled at 2026-10-10T10:45:06.436500+00:00, scoring version aggregate-v2-best-published-zfit-2026-10-03. It uses the Best orchestrator weights. These are actual source values, not invented sample scores.
| Source / variant | Raw | Arena-scaled |
|---|---|---|
| Arena / high | 1507.0758687407206 | 1507.0758687407206 |
| Epoch AI / default | 167.33 | 1544.843625001453 |
| LiveBench / max | 83.21657142857143 | 1493.0178001358406 |
| Terminal-Bench / max | 64.85 | 1491.844345901992 |
| Source | mu | sigma | alpha | beta | Overlap |
|---|---|---|---|---|---|
| epoch_eci | 145.4943925233645 | 14.841039851819973 | 1415.5755747968187 | 87.85980819204782 | 107 |
| livebench_global | 75.3434769047619 | 4.999187697589841 | 1469.3687376103364 | 15.016471868780174 | 50 |
| terminal | 33.606 | 18.318524613079514 | 1477.793604584871 | 8.238024921575137 | 20 |
Epoch's calibrated score is 1415.5755747968187 + 87.85980819204782 * (167.33 - 145.4943925233645) / 14.841039851819973 = 1544.843625001453. Epoch and LiveBench global share the reasoning family, so they are averaged before weighting.
Reasoning = (1544.843625001453 + 1493.0178001358406) / 2 = 1518.9307125686469
Human preference = 1507.0758687407206
Terminal work = 1491.844345901992
R = (0.50 * 1518.9307125686469 + 0.30 * 1507.0758687407206 + 0.20 * 1491.844345901992) / (0.50 + 0.30 + 0.20)
R = 1509.956986086938; displayed score = 1510
The score uses four independent evaluators and all three families. Arena's publication is 8 October, Epoch's model result and Terminal-Bench's run are dated 22 September, and LiveBench's table is dated 25 June; all were fetched on 10 October. Their ages do not reduce their weights.
The page rebuilds every 15 minutes. Board refreshes are scheduled twice daily at 06:45 and 12:45 in the timer's local timezone. Snapshot dates use UTC. The first saved snapshot for a UTC day stays in use; a same-day refresh replaces it only if it has strictly more top-level sources with status ok. Changed scores alone, or more successful rating feeds within the same top-level source count, do not replace that day's snapshot. Thus a twice-daily refresh need not produce a new ranking.
Source publication or run dates are separate from fetch dates; there is no maximum source age cutoff. A board's “as of” is the newest available date among its credited sources, including prices or build checks when present, not the oldest observation. The Sources line credits stored evidence, including sources outside a role or excluded from scoring; listing a source does not mean it contributed. Failed feeds supply no new observations; usable remaining families are renormalised as above. An Arena failure marks the aggregate board stale.
A missing snapshot, invalid or future UTC date, snapshot more than two UTC calendar days old, or snapshot with no rows on any board prevents the live leaderboard swap; the existing news layout stays on the front page and a local preview can show a warning. The rating table is our normalisation and aggregation, shared under CC BY-SA 4.0, with each source's licence and attribution in the board's Sources disclosure.