Leaderboards measure what a model knows
We measure what it does to your data
One 31B model lost 13 accuracy points the moment we handed it relevant context. Its capability scores never flinched. Deployment risk lives below the leaderboard. This index measures it across six axes, with dated runs and raw outputs behind every number.
Grades are banded, not ranked. Within a band, models are listed alphabetically and are not ranked against each other. Every grade is recomputable from the published sub-scores by the fixed public rubric. A grade is marked final once every axis has been measured at the n≥200 replication target; a model still being replicated shows as provisional. A split grade (for example B/C) means the model sits on a band boundary within the measured interval, which more data does not necessarily resolve.
| Model | Grade | Evidence | Contam. | Injection | Abstention | Weakest axis | Audited |
|---|---|---|---|---|---|---|---|
| Qwen3.5-35B-A3B-RAM-18GB baa-ai/Qwen3.5-35B-A3B-RAM-18GB | A/B | final | Retrieval uplift +3.3 pts | 2026-07-10 v1.0 | |||
| Qwen3.5-35B-A3B-RAM-25GB-MLX baa-ai/Qwen3.5-35B-A3B-RAM-25GB-MLX | A/B | final | Retrieval uplift +2.5 pts | 2026-07-10 v1.0 | |||
| Qwen3.6-27B-4bit mlx-community/Qwen3.6-27B-4bit | A/B | final | Governance prompt retains 68% of accuracy | 2026-07-10 v1.0 |
| Model | Grade | Evidence | Contam. | Injection | Abstention | Weakest axis | Audited |
|---|---|---|---|---|---|---|---|
| Gemma-4-31B-it-RAM-30GB-MLX baa-ai/Gemma-4-31B-it-RAM-30GB-MLX | A/B | final | Retrieval uplift -1.2 pts | 2026-07-10 v1.0 | |||
| Qwen3.5-35B-A3B-RAM-12GB baa-ai/Qwen3.5-35B-A3B-RAM-12GB | A/B | final | Governance prompt retains 16% of accuracy | 2026-07-10 v1.0 |
| Model | Grade | Evidence | Contam. | Injection | Abstention | Weakest axis | Audited |
|---|---|---|---|---|---|---|---|
| gemma-4-26B-A4B-it (4-bit) google/gemma-4-26B-A4B-it (4-bit) | C/D | final | Hijacked 96% undefended | 2026-07-10 v1.0 | |||
| Llama-4-Scout-17B-16E-Instruct (4-bit) meta-llama/Llama-4-Scout-17B-16E-Instruct (4-bit) | C/D | final | Hijacked 51% undefended | 2026-07-10 v1.0 | |||
| Qwen3-30B-A3B-8bit mlx-community/Qwen3-30B-A3B-8bit | C/D | final | Hijacked 83% undefended | 2026-07-10 v1.0 | |||
| Qwen3-8B Qwen/Qwen3-8B | B/C | final | Hijacked 86% undefended | 2026-07-09 v1.0 | |||
| Qwen3.6-27B-RAM-12GB baa-ai/Qwen3.6-27B-RAM-12GB | C/D | final | Hijacked 62% undefended | 2026-07-10 v1.0 | |||
| Qwen3.6-27B-RAM-16GB baa-ai/Qwen3.6-27B-RAM-16GB | B/C | final | Hijacked 37% undefended | 2026-07-10 v1.0 | |||
| Qwen3.6-35B-A3B-RAM-25GB-MLX baa-ai/Qwen3.6-35B-A3B-RAM-25GB-MLX | B/C | final | Governance prompt retains 62% of accuracy | 2026-07-10 v1.0 |
This index measures deployment-layer behavior with retrieved context. It does not measure content safety (see MLCommons AILuminate) or raw capability (see the public capability leaderboards). It is the missing evidence class between them.