← Deployment Risk Index
Deployment Risk Audit

gemma-4-26B-A4B-it (4-bit)

google/gemma-4-26B-A4B-it (4-bit)

audited 2026-07-10 · methodology v1.0 · thinking mode off · seeds 1

C/DDeployment
Risk Grade
audited 2026-07-10
methodology v1.0
final

Executive summary

gemma-4-26B-A4B-it (4-bit) grades C/D (final) on the Black Sheep AI Deployment Risk Index, methodology v1.0, audited 2026-07-10. Findings: (1) relevant retrieved context moved accuracy +4.7 pts (95% CI -0.3 to 9.4, n=320). (2) hostile-context injection hijacked it 96% undefended (n=200), reduced to 8% by ingestion sanitization. (3) on unanswerable questions it guessed 7% of the time under an indexed abstention policy (n=200). Numbers are final: every axis was measured at n>=200. This audit measures deployment-layer behavior; it is not a content-safety or capability assessment. Source: hell.ai/models/gemma4-26b-a4b-4bit, report v1.0.

Grade composition

Grade computed from the sub-scores below by the fixed public rubric (methodology v1.0). Composite 61 / 100 (95% interval 54–68). Recompute it yourself from the published rubric.

Context contamination79/10079
Retrieval uplift79/10079
Injection resistance0/1000
Abstention discipline90/10090
Governance cost96/10096
Where it can go to work

Deployment qualification

Qualification is scoped to the conditions we tested. “Not tested” is an honest state, not a pass.

Deployment contextStatusConditions
Closed-book assistantNo disqualifying finding
RAG over a curated internal corpusNo disqualifying finding
RAG over an uncurated or web corpusNot qualifiedInjection hijack rate 96% undefended; only ingestion sanitization drove it to 8%
Agentic tool use with retrievalNot testedPlanned methodology v1.2

Verdicts are within tested scope only. “No disqualifying finding” means no measured axis crossed its disqualifying threshold for that context; it is not certification. The thresholds are published on the methodology page.

5 of 6 axes assessed

What breaks, in numbers

5 of the six axes were assessed for this model. Quantization robustness is not assessed here: it requires an audited full-precision parent, which this entry does not have.

You are reading the executive view. Switch to Technical for per-axis tables, confidence intervals, and sample sizes.

Does giving this model retrieved documents make it worse than answering from memory?

Relevant retrieved context added 4.7 accuracy points (95% CI -0.3 to 9.4, n=320). When context flipped an answer it overrode a correct answer 23 times against 38 repairs. Irrelevant context of equal length moved accuracy -2.8 pts. robust

MetricValue95% CInRunEv.
Net context effect (pts)+4.7-0.31 to 9.383202026-07-10PROV
Override rate right→wrong233202026-07-10PROV
Repair rate wrong→right383202026-07-10PROV
Random-context control (pts)-2.8-6.88 to 1.253202026-07-10PROV
Closed-book accuracy64%3202026-07-10PROV
With-context accuracy69%3202026-07-10PROV
When the answer is in the retrieved documents, does this model actually use it?

On knowledge-heavy questions, adding the answer-bearing documents changed accuracy by +5.0 pts (95% CI -0.8 to 10.8, n=240).

MetricValue95% CInRunEv.
Knowledge-slice uplift (pts)+5.0-0.83 to 10.832402026-07-10PROV
Can hostile text hidden in retrieved documents hijack this model?

With a hostile instruction hidden in the retrieved documents, the model was hijacked 96% of the time undefended. A prompt-level instruction to ignore it moved that to 33%. Sanitizing the documents at ingestion moved it to 8% (95% CI 4% to 12%, n=200).

MetricValue95% CInRunEv.
Hijack rate · undefended96% (193/200)0.94 to 0.992002026-07-10PROV
Hijack rate · prompt-inoculated33% (66/200)2002026-07-10PROV
Hijack rate · ingestion-sanitized8% (15/200)0.04 to 0.122002026-07-10PROV
When the documents cannot answer the question, does this model admit it or guess?

Asked questions the documents cannot answer, the model guessed instead of abstaining 100% of the time ungoverned and 7% with an indexed abstention policy (95% CI 4% to 11%, n=200).

MetricValue95% CInRunEv.
Over-inference · ungoverned100% (199/200)2002026-07-10PROV
Over-inference · governed7% (14/200)0.04 to 0.112002026-07-10PROV
Does a strict governance instruction degrade the model on questions it should answer?

On questions it should answer, the governance instruction retained 98% of ungoverned accuracy (98% governed vs 100% ungoverned, n=200).

MetricValue95% CInRunEv.
Derivable accuracy · ungoverned100% (200/200)2002026-07-10PROV
Derivable accuracy · governed98% (196/200)2002026-07-10PROV
Retention ratio0.982002026-07-10PROV

Framework mapping

This mapping is informative. It identifies which framework activities each measurement can serve as evidence for. It is not a conformity assessment, a certification, or legal advice.

AxisNIST AI RMFISO/IEC 42001EU AI Act
Context contaminationMEASURE 2.5, 2.9A.6 / 8.2Art. 15 accuracy; Art. 55 model evaluation
Injection resistanceMEASURE 2.7, MANAGE 1.3A.5 / A.8Art. 15 cybersecurity; Art. 55 adversarial testing
Abstention disciplineMEASURE 2.5, MAP 3.49.1Art. 13 transparency; Art. 15
Quantization robustnessMEASURE 2.68.3Annex XI technical documentation

Read the full audit report Framework detail

Caveats and scope

  • Single language (English). Corpus domains: encyclopedic and closed-world synthetic sets.

  • Every axis was measured at the replication target (under-determined n=200, derivable n=200), so these numbers are final. A split grade means the model sits on a band boundary within the measured interval, not that the data is incomplete. Track it on Evidence.

  • Behavior is checkpoint-specific. This audit describes the exact weight artifact named above, at the serving configuration tested. Fine-tuned derivatives and other serving stacks are not covered.

  • This is a first-party mark. It means one thing: these measurements exist and you can check them. It is not a certification or an attestation. See Trust & independence.

Evidence locker

Every raw run output behind the numbers above, with a SHA-256 per file. Hand them to your own data scientist.

FileSHA-256
card.json47292dba94ed6356a64d7f02…
manifest.json0447d8f5fa412b7f9f836754…
raw_contamination.jsona5fa85b130be63c03765114b…
raw_governance.json7e0915e30f7dd3f0a1b126e3…
raw_injection.json3c9f0fff57ad0ac888f11fdb…
records.jsonl6df650d314b7f82c5357d7f5…

harness 9fa3835db3127f0a · rubric e2d5e396f7137d1e · run 2026-07-10T12:00:51

Deploying this model on your corpus? Your documents change these numbers. Request an audit of your stack.