The index

Leaderboards measure what a model knows
We measure what it does to your data

One 31B model lost 13 accuracy points the moment we handed it relevant context. Its capability scores never flinched. Deployment risk lives below the leaderboard. This index measures it across six axes, with dated runs and raw outputs behind every number.

Grades are banded, not ranked. Within a band, models are listed alphabetically and are not ranked against each other. Every grade is recomputable from the published sub-scores by the fixed public rubric. A grade is marked final once every axis has been measured at the n≥200 replication target; a model still being replicated shows as provisional. A split grade (for example B/C) means the model sits on a band boundary within the measured interval, which more data does not necessarily resolve.

Grade A · 3 models Placed by the composite point estimate. Models within a band are not ranked against each other.
ModelGradeEvidenceContam.InjectionAbstentionWeakest axisAudited
Qwen3.5-35B-A3B-RAM-18GB
baa-ai/Qwen3.5-35B-A3B-RAM-18GB
A/Bfinal76/10088/10098/100Retrieval uplift +3.3 pts2026-07-10
v1.0
Qwen3.5-35B-A3B-RAM-25GB-MLX
baa-ai/Qwen3.5-35B-A3B-RAM-25GB-MLX
A/Bfinal76/10092/10099/100Retrieval uplift +2.5 pts2026-07-10
v1.0
Qwen3.6-27B-4bit
mlx-community/Qwen3.6-27B-4bit
A/Bfinal95/10089/100100/100Governance prompt retains 68% of accuracy2026-07-10
v1.0
Grade B · 2 models Placed by the composite point estimate. Models within a band are not ranked against each other.
ModelGradeEvidenceContam.InjectionAbstentionWeakest axisAudited
Gemma-4-31B-it-RAM-30GB-MLX
baa-ai/Gemma-4-31B-it-RAM-30GB-MLX
A/Bfinal56/10098/100100/100Retrieval uplift -1.2 pts2026-07-10
v1.0
Qwen3.5-35B-A3B-RAM-12GB
baa-ai/Qwen3.5-35B-A3B-RAM-12GB
A/Bfinal96/10081/10096/100Governance prompt retains 16% of accuracy2026-07-10
v1.0
Grade C · 7 models Placed by the composite point estimate. Models within a band are not ranked against each other.
ModelGradeEvidenceContam.InjectionAbstentionWeakest axisAudited
gemma-4-26B-A4B-it (4-bit)
google/gemma-4-26B-A4B-it (4-bit)
C/Dfinal79/1000/10090/100Hijacked 96% undefended2026-07-10
v1.0
Llama-4-Scout-17B-16E-Instruct (4-bit)
meta-llama/Llama-4-Scout-17B-16E-Instruct (4-bit)
C/Dfinal76/10022/10080/100Hijacked 51% undefended2026-07-10
v1.0
Qwen3-30B-A3B-8bit
mlx-community/Qwen3-30B-A3B-8bit
C/Dfinal75/1000/10093/100Hijacked 83% undefended2026-07-10
v1.0
Qwen3-8B
Qwen/Qwen3-8B
B/Cfinal91/1000/10095/100Hijacked 86% undefended2026-07-09
v1.0
Qwen3.6-27B-RAM-12GB
baa-ai/Qwen3.6-27B-RAM-12GB
C/Dfinal76/1009/10099/100Hijacked 62% undefended2026-07-10
v1.0
Qwen3.6-27B-RAM-16GB
baa-ai/Qwen3.6-27B-RAM-16GB
B/Cfinal78/10040/100100/100Hijacked 37% undefended2026-07-10
v1.0
Qwen3.6-35B-A3B-RAM-25GB-MLX
baa-ai/Qwen3.6-35B-A3B-RAM-25GB-MLX
B/Cfinal66/10071/100100/100Governance prompt retains 62% of accuracy2026-07-10
v1.0

This index measures deployment-layer behavior with retrieved context. It does not measure content safety (see MLCommons AILuminate) or raw capability (see the public capability leaderboards). It is the missing evidence class between them.

Request an audit of your model or stack