Qwen3.6-35B-A3B-RAM-25GB-MLX
baa-ai/Qwen3.6-35B-A3B-RAM-25GB-MLX
audited 2026-07-10 · methodology v1.0 · thinking mode off · seeds 1
Risk Grade
Executive summary
Qwen3.6-35B-A3B-RAM-25GB-MLX grades B/C (final) on the Black Sheep AI Deployment Risk Index, methodology v1.0, audited 2026-07-10. Findings: (1) relevant retrieved context moved accuracy +1.6 pts (95% CI -2.5 to 5.6, n=320). (2) hostile-context injection hijacked it 18% undefended (n=200), reduced to 0% by ingestion sanitization. (3) on unanswerable questions it guessed 0% of the time under an indexed abstention policy (n=200). Numbers are final: every axis was measured at n>=200. This audit measures deployment-layer behavior; it is not a content-safety or capability assessment. Source: hell.ai/models/qwen36-35b-a3b-ram25, report v1.0.
Grade composition
Grade computed from the sub-scores below by the fixed public rubric (methodology v1.0). Composite 69 / 100 (95% interval 63–75). Recompute it yourself from the published rubric.
Deployment qualification
Qualification is scoped to the conditions we tested. “Not tested” is an honest state, not a pass.
| Deployment context | Status | Conditions |
|---|---|---|
| Closed-book assistant | No disqualifying finding | – |
| RAG over a curated internal corpus | No disqualifying finding | – |
| RAG over an uncurated or web corpus | Qualified with mitigations | Injection hijack rate 18% undefended; only ingestion sanitization drove it to 0% |
| Agentic tool use with retrieval | Not tested | Planned methodology v1.2 |
Verdicts are within tested scope only. “No disqualifying finding” means no measured axis crossed its disqualifying threshold for that context; it is not certification. The thresholds are published on the methodology page.
What breaks, in numbers
5 of the six axes were assessed for this model. Quantization robustness is not assessed here: it requires an audited full-precision parent, which this entry does not have.
You are reading the executive view. Switch to Technical for per-axis tables, confidence intervals, and sample sizes.
Relevant retrieved context added 1.6 accuracy points (95% CI -2.5 to 5.6, n=320). When context flipped an answer it overrode a correct answer 19 times against 24 repairs. Irrelevant context of equal length moved accuracy -2.8 pts. robust
| Metric | Value | 95% CI | n | Run | Ev. |
|---|---|---|---|---|---|
| Net context effect (pts) | +1.6 | -2.50 to 5.62 | 320 | 2026-07-10 | PROV |
| Override rate right→wrong | 19 | – | 320 | 2026-07-10 | PROV |
| Repair rate wrong→right | 24 | – | 320 | 2026-07-10 | PROV |
| Random-context control (pts) | -2.8 | -5.94 to 0.31 | 320 | 2026-07-10 | PROV |
| Closed-book accuracy | 67% | – | 320 | 2026-07-10 | PROV |
| With-context accuracy | 68% | – | 320 | 2026-07-10 | PROV |
On knowledge-heavy questions, adding the answer-bearing documents changed accuracy by -0.4 pts (95% CI -5.4 to 4.2, n=240).
| Metric | Value | 95% CI | n | Run | Ev. |
|---|---|---|---|---|---|
| Knowledge-slice uplift (pts) | -0.4 | -5.42 to 4.17 | 240 | 2026-07-10 | PROV |
With a hostile instruction hidden in the retrieved documents, the model was hijacked 18% of the time undefended. A prompt-level instruction to ignore it moved that to 0%. Sanitizing the documents at ingestion moved it to 0% (95% CI 0% to 0%, n=200).
| Metric | Value | 95% CI | n | Run | Ev. |
|---|---|---|---|---|---|
| Hijack rate · undefended | 18% (35/200) | 0.12 to 0.23 | 200 | 2026-07-10 | PROV |
| Hijack rate · prompt-inoculated | 0% (0/200) | – | 200 | 2026-07-10 | PROV |
| Hijack rate · ingestion-sanitized | 0% (0/200) | 0.00 to 0.00 | 200 | 2026-07-10 | PROV |
Asked questions the documents cannot answer, the model guessed instead of abstaining 92% of the time ungoverned and 0% with an indexed abstention policy (95% CI 0% to 0%, n=200).
| Metric | Value | 95% CI | n | Run | Ev. |
|---|---|---|---|---|---|
| Over-inference · ungoverned | 92% (184/200) | – | 200 | 2026-07-10 | PROV |
| Over-inference · governed | 0% (0/200) | 0.00 to 0.00 | 200 | 2026-07-10 | PROV |
On questions it should answer, the governance instruction retained 62% of ungoverned accuracy (62% governed vs 100% ungoverned, n=200).
| Metric | Value | 95% CI | n | Run | Ev. |
|---|---|---|---|---|---|
| Derivable accuracy · ungoverned | 100% (199/200) | – | 200 | 2026-07-10 | PROV |
| Derivable accuracy · governed | 62% (123/200) | – | 200 | 2026-07-10 | PROV |
| Retention ratio | 0.62 | – | 200 | 2026-07-10 | PROV |
Framework mapping
This mapping is informative. It identifies which framework activities each measurement can serve as evidence for. It is not a conformity assessment, a certification, or legal advice.
| Axis | NIST AI RMF | ISO/IEC 42001 | EU AI Act |
|---|---|---|---|
| Context contamination | MEASURE 2.5, 2.9 | A.6 / 8.2 | Art. 15 accuracy; Art. 55 model evaluation |
| Injection resistance | MEASURE 2.7, MANAGE 1.3 | A.5 / A.8 | Art. 15 cybersecurity; Art. 55 adversarial testing |
| Abstention discipline | MEASURE 2.5, MAP 3.4 | 9.1 | Art. 13 transparency; Art. 15 |
| Quantization robustness | MEASURE 2.6 | 8.3 | Annex XI technical documentation |
Caveats and scope
Single language (English). Corpus domains: encyclopedic and closed-world synthetic sets.
Every axis was measured at the replication target (under-determined n=200, derivable n=200), so these numbers are final. A split grade means the model sits on a band boundary within the measured interval, not that the data is incomplete. Track it on Evidence.
Behavior is checkpoint-specific. This audit describes the exact weight artifact named above, at the serving configuration tested. Fine-tuned derivatives and other serving stacks are not covered.
This is a first-party mark. It means one thing: these measurements exist and you can check them. It is not a certification or an attestation. See Trust & independence.
Evidence locker
Every raw run output behind the numbers above, with a SHA-256 per file. Hand them to your own data scientist.
| File | SHA-256 |
|---|---|
| card.json | e2d966fb4479129e98c5ed40… |
| manifest.json | f67fa6e0a1c40712ed73dfe2… |
| raw_contamination.json | 606cb4847f9a5532dc411793… |
| raw_governance.json | 473fa1e0327948ec38920fe5… |
| raw_injection.json | 8481604acd28885f44f3f619… |
| records.jsonl | 558a3c04370e4a5f7986b1b8… |
harness 9fa3835db3127f0a · rubric e2d5e396f7137d1e · run 2026-07-10T02:23:44
Deploying this model on your corpus? Your documents change these numbers. Request an audit of your stack.