← The Daily Markdown

How we score model boards

Each role uses a weighted average of outside evaluations on an Arena-like scale. It is a proxy for quality, not a prediction of pairwise wins. The boards use this exact calculation.

Sources and role weights

Only these rating sources contribute. Numeric scores must be finite and cannot be booleans; percentage scores must be between 0 and 100 inclusive. Arena text and coding are human preference ratings; Epoch AI provides ECI reasoning scores; Aider provides polyglot editing results; LiveBench provides global, coding, instruction-following and data-analysis scores; Terminal-Bench provides agent workflow accuracy. Multiple categories from one publisher count as one independent evaluator.

Weights before removing missing families
RoleRating familyWeightSources averaged within that family
OrchestratorObjective reasoning50%Epoch AI ECI, LiveBench global
OrchestratorHuman preference30%Arena text
OrchestratorTerminal work20%Terminal-Bench
Mid-tier workerExecuted edits35%Aider
Mid-tier workerTerminal work35%Terminal-Bench
Mid-tier workerObjective coding20%LiveBench coding
Mid-tier workerHuman preference10%Arena coding
Rote workerRoutine tasks60%LiveBench instruction following, data analysis
Rote workerObjective coding25%LiveBench coding
Rote workerHuman preference15%Arena text

The orchestrator weights also apply to the open-weight and home-computer boards. Published build prices use the mid-tier weights. Each stack entry uses its own role's weights.

Choose one configuration per source and model

Use each source's highest published raw score for that model instead of averaging reasoning efforts. If raw scores tie, choose the latest publication date, then the lexicographically greatest model spelling and URL. Explicit effort suffixes fold into the base model. A full date suffix folds into the base only when that source has exactly one dated release for that base. Versions, tiers, parameter sizes and checkpoint IDs remain separate. The chosen spelling can differ between sources.

Arena uses the latest publication in each of its overall and coding categories. Epoch uses rows with ECI scores. Aider uses pass_rate_2 from documented single-model whole, diff or diff-fenced runs with 224 or 225 tests, a commit and a date; architect/editor pairs and combined-model runs are excluded. Aider first breaks same-spelling score ties by date, edit format and commit spelling. LiveBench uses its latest linked dated CSV. If category summary columns are absent, it averages the mapped tasks within each category, then averages categories equally for the global score; an incomplete category contributes nothing. Terminal-Bench uses only the latest benchmark version from one pinned repository commit. Same-spelling runs with incompatible agent or agent-version configurations are excluded; otherwise the latest date and URL win before configuration selection. These are workflow proxies, with the named agent retained in the evidence.

Normalise onto the Arena scale

Arena text and coding scores stay unchanged. For each other source and release, match the selected models to their selected Arena text ratings. Remove excluded observations and duplicate model keys from the overlap. The fit needs at least 10 overlapping model keys.

mu = mean(raw scores in the overlap)
sigma = population standard deviation(raw scores in the overlap)
z_i = (raw_i - mu) / sigma
alpha = mean(Arena text ratings in the overlap)
beta = sum(z_i * (Arena_i - alpha)) / sum(z_i * z_i)
calibrated(raw) = alpha + beta * (raw - mu) / sigma

A source contributes nothing if the overlap is smaller than 10, sigma is zero, beta is zero or negative, or its observation is excluded. Unselected configurations also contribute nothing. The parameters are fitted when the snapshot is built and saved with it; rendering does not refit them. This calculation has no clipping, minimum quality floor or age multiplier.

Combine families, then rank

For a model and role, use only allowed sources with a positive family weight, a selected configuration, a calibrated score and no exclusion. Average those calibrated scores equally within each family. Then apply the role weights only to the families present for that model:

family_mean[f] = mean(usable calibrated source scores in family f)
R = sum(weight[f] * family_mean[f] for present f) / sum(weight[f] for present f)
displayed score = round(R)

Missing families are omitted from both numerator and denominator; no score is imputed. Python's round rounds to the nearest integer, with exact halfway values going to the even integer. Ranking uses the unrounded R, so two identical displayed integers need not be a tie. Exact equal unrounded scores share a competition rank: rank = 1 + number of ranked rows with a strictly greater R (1, 2, 2, 4). Tied rows are displayed in ascending model spelling order. Price never breaks a quality tie. We never divide quality by dollars.

A model with exactly one contributing independent evaluator goes into New, not enough ratings yet below its board, without a rank. A model with no usable family has no score or rank. Otherwise, fewer than two rating families, fewer than two independent evaluators, or less than 50% of the available family weight adds a partial label; it does not remove an otherwise scored model from ranking. Available families are those with at least one usable observation anywhere in the snapshot, before model or price filters. Coverage weight is the sum of a model's present preset weights divided by all preset weights; available share divides that sum by the snapshot's available weights. A board warns if less than 50% of its preset weight is available and names the missing families.

Movement compares the same model identity and role or machine to the previous snapshot's ranked field; tied ranks use the same rule. The first snapshot or a change of scoring method has no movement label. Rendering a saved snapshot from before the single-source rule suppresses its old movement labels when filtering changes that board's ranking.

What qualifies for each board

The orchestrator board has no price or minimum quality filter. Open-weight models need an Arena weight licence that is present and not marked proprietary; individual licence terms still apply. PirateBench's latest scored epoch supplies our separate build-check column and list build prices, not rating points. A build price is the median of matching job prices when every matching job has a positive numeric list price. It is absent otherwise.

Both worker boards exclude names matching the flagship rule: the tokens opus, sol, astra or fable; names starting muse-spark; Gemini names with a pro token; and Grok names without mini or fast. If a build price exists it takes precedence over token price. Rote workers must cost at most $0.25 per build; mid-tier workers must cost strictly more than $0.25 and strictly less than the cheapest measured flagship build (with no upper bound if no priced flagship exists). Without a build price, rote workers need a positive blended token price at most $0.30 per million tokens; mid-tier workers need more than $0.30 and at most $3.

Token prices use LiteLLM chat entries without a colon in the name. Input and output prices must be finite, numeric, positive and not booleans. The blended price is 1,000,000 * (0.75 * input_cost_per_token + 0.25 * output_cost_per_token). For each folded identity, use the cheapest blended price, then the alphabetically first provider key. The Published build prices board lists every model with a positive measured build price, ordered by quality.

Best stack picks the first non-partial model in each component board, falling back to its first row. The value stack considers only OpenRouter price entries below $5, $1 and $0.30 per million tokens for the three roles, respectively (strict inequalities). It prefers non-partial coverage, then higher R, then model spelling. Its day-of-work estimate is 2 * orchestrator price + 10 * mid-tier price + 50 * rote price, shown only with all three roles. Single-evaluator picks appear in the strip.

Home-computer picks must be open-weight and have a known total checkpoint parameter count. Estimated 4-bit memory is parameters in billions * 0.5 * 1.25 GB, including all experts, and must fit within 80% of the machine's memory. The machines are RTX 3090 (24 GB), Strix Halo (128 GB) and Mac Studio (512 GB). Picks prefer non-partial coverage, then score, then spelling. Quality is measured by API evaluations, not on that hardware; speed is unmeasured.

Example: Claude Opus 5.5

This is a fixed example from the newest board data on disk when this page was prepared: the 10 October 2026 snapshot, pulled at 2026-10-10T10:45:06.436500+00:00, scoring version aggregate-v2-best-published-zfit-2026-10-03. It uses the Best orchestrator weights. These are actual source values, not invented sample scores.

Selected source scores
Source / variantRawArena-scaled
Arena / high1507.07586874072061507.0758687407206
Epoch AI / default167.331544.843625001453
LiveBench / max83.216571428571431493.0178001358406
Terminal-Bench / max64.851491.844345901992
Saved calibration parameters (Arena is unchanged)
SourcemusigmaalphabetaOverlap
epoch_eci145.494392523364514.8410398518199731415.575574796818787.85980819204782107
livebench_global75.34347690476194.9991876975898411469.368737610336415.01647186878017450
terminal33.60618.3185246130795141477.7936045848718.23802492157513720

Epoch's calibrated score is 1415.5755747968187 + 87.85980819204782 * (167.33 - 145.4943925233645) / 14.841039851819973 = 1544.843625001453. Epoch and LiveBench global share the reasoning family, so they are averaged before weighting.

Reasoning = (1544.843625001453 + 1493.0178001358406) / 2 = 1518.9307125686469
Human preference = 1507.0758687407206
Terminal work = 1491.844345901992
R = (0.50 * 1518.9307125686469 + 0.30 * 1507.0758687407206 + 0.20 * 1491.844345901992) / (0.50 + 0.30 + 0.20)
R = 1509.956986086938; displayed score = 1510

The score uses four independent evaluators and all three families. Arena's publication is 8 October, Epoch's model result and Terminal-Bench's run are dated 22 September, and LiveBench's table is dated 25 June; all were fetched on 10 October. Their ages do not reduce their weights.

Freshness and update schedule

The page rebuilds every 15 minutes. Board refreshes are scheduled twice daily at 06:45 and 12:45 in the timer's local timezone. Snapshot dates use UTC. The first saved snapshot for a UTC day stays in use; a same-day refresh replaces it only if it has strictly more top-level sources with status ok. Changed scores alone, or more successful rating feeds within the same top-level source count, do not replace that day's snapshot. Thus a twice-daily refresh need not produce a new ranking.

Source publication or run dates are separate from fetch dates; there is no maximum source age cutoff. A board's “as of” is the newest available date among its credited sources, including prices or build checks when present, not the oldest observation. The Sources line credits stored evidence, including sources outside a role or excluded from scoring; listing a source does not mean it contributed. Failed feeds supply no new observations; usable remaining families are renormalised as above. An Arena failure marks the aggregate board stale.

A missing snapshot, invalid or future UTC date, snapshot more than two UTC calendar days old, or snapshot with no rows on any board prevents the live leaderboard swap; the existing news layout stays on the front page and a local preview can show a warning. The rating table is our normalisation and aggregation, shared under CC BY-SA 4.0, with each source's licence and attribution in the board's Sources disclosure.

View the boards · About · MCP server