Evidence & replication

How a grade becomes final

A grade starts provisional, on a modest sample with a wide interval, and becomes final once every axis has been re-measured at n≥200. We would rather tell you a number is provisional than let you find out. This page tracks where each model is.

What provisional and final mean

Provisional

The metric is measured but the sample is small enough that the confidence interval is wide. The grade mark is drawn with a dashed border and the word provisional. Treat the direction as real and the exact value as an estimate.

Final

Every axis has been re-measured at n≥200, so the per-axis intervals are tight and the numbers will not move with more of the same data. The mark becomes solid and reads final.

A grade whose interval straddles a band boundary is shown as a split, for example B/C. At final that is the honest answer: the model sits on the boundary within its measured interval, and more of the same data does not necessarily resolve it. We never round uncertainty away.

Replication tracker

ModelContamination nUnder-determined nDerivable nTargetStatus
gemma-4-26B-A4B-it (4-bit)320200200≥200final
Gemma-4-31B-it-RAM-30GB-MLX320200200≥200final
Llama-4-Scout-17B-16E-Instruct (4-bit)320200200≥200final
Qwen3-30B-A3B-8bit320200200≥200final
Qwen3-8B320200200≥200final
Qwen3.5-35B-A3B-RAM-12GB320200200≥200final
Qwen3.5-35B-A3B-RAM-18GB320200200≥200final
Qwen3.5-35B-A3B-RAM-25GB-MLX320200200≥200final
Qwen3.6-27B-4bit320200200≥200final
Qwen3.6-27B-RAM-12GB320200200≥200final
Qwen3.6-27B-RAM-16GB320200200≥200final
Qwen3.6-35B-A3B-RAM-25GB-MLX320200200≥200final

The under-determined and derivable question sets were the tightest constraint at launch; a purpose-built closed-world corpus expanded them to the target. Sample-size target: n≥200 per axis, with 2000-sample bootstrap confidence intervals.