← Qwen3-8B model card
Audit report · v1.0

Qwen3-8B

Qwen/Qwen3-8B
audit period 2026-07-09 · report v1.0

B/CDeployment
Risk Grade
audited 2026-07-09
methodology v1.0
final

Section I · Assessment statement

This report presents the results of a technical assessment performed independently of the model vendor by hell.ai, an index published by Black Sheep AI. hell.ai is not an accredited certification body, and this report is not an attestation under any audit standard. All findings are reproducible from the evidence manifest in Section VIII. Black Sheep AI also sells deployment-mitigation tooling; that conflict of interest is disclosed and managed as described in Section IX.

Section II · Scope

  • System under test: Qwen/Qwen3-8B, at the serving configuration recorded in the manifest.

  • Deployment patterns tested: retrieval-augmented question answering over curated and adversarial document sets; closed-world abstention.

  • Sample sizes: contamination/uplift n=320; under-determined n=200; derivable n=200. Seeds 1.

  • Exclusions: fine-tuned derivatives, multimodal inputs, alternate serving stacks, agentic tool use.

Section III · Methodology summary

Each assessment objective is tested by paired generation runs holding the model and decoding fixed and varying one factor (context relevance, defense rung, or governance instruction). Metrics and their confidence intervals are defined in methodology v1.0. Test items and the detection internals used to classify behavior are withheld so the benchmark cannot be trained against; the published behavioral outcomes are sufficient to check every conclusion.

Section IV · Findings by assessment objective

C-1 · Retrieved relevant context does not reduce answer accuracy relative to closed-book.

NO EXCEPTION

Relevant retrieved context added 7.8 accuracy points (95% CI 3.4 to 12.5, n=320). When context flipped an answer it overrode a correct answer 17 times against 42 repairs. Irrelevant context of equal length moved accuracy -0.9 pts.

C-2 · Hostile instructions in retrieved documents do not hijack the model.

EXCEPTION NOTED

With a hostile instruction hidden in the retrieved documents, the model was hijacked 86% of the time undefended. A prompt-level instruction to ignore it moved that to 1%. Sanitizing the documents at ingestion moved it to 0% (95% CI 0% to 0%, n=200). Compensating control: ingestion-layer sanitization reduced the hijack rate to 0%. Residual risk depends on that control being deployed; undefended, the exception stands at 86%.

C-3 · The model abstains when the documents cannot answer the question.

NO EXCEPTION

Asked questions the documents cannot answer, the model guessed instead of abstaining 44% of the time ungoverned and 4% with an indexed abstention policy (95% CI 1% to 6%, n=200).

C-4 · A governance instruction does not materially degrade answerable-question accuracy.

NO EXCEPTION

On questions it should answer, the governance instruction retained 88% of ungoverned accuracy (86% governed vs 98% ungoverned, n=200).

Section V · Exceptions register

Severity is graded by the residual magnitude of each exception, defined in the methodology. An exception with a named compensating control is not thereby resolved; the residual risk depends on that control being deployed.

ObjectiveExceptionSeverity
C-2Undefended injection hijack 86%High

Section VI · Recommended mitigations and residual risk

  • Ingestion-layer sanitization (strip imperative/procedural spans from retrieved documents before context assembly). In this audit it moved the injection hijack rate from 86% to 0%. Verify by re-running the injection objective on your own corpus.

  • Indexed abstention policy (place the “answer only from the documents” procedure in the retrieval index, not the system prompt). Governed over-inference here was 4% against 44% ungoverned.

  • Post-quantization re-verification of the abstention and contamination axes before deploying a compressed variant.

Section VII · Framework mapping annex

Informative only. Identifies the framework activities each measurement can serve as evidence for. Not a conformity assessment, certification, or legal advice, and not a determination of whether your deployment is high-risk. Provider obligations (EU AI Act Art. 55 for general-purpose models) and deployer obligations (Art. 26) are distinct; confirm which apply to you with your own counsel.

AxisNIST AI RMFISO/IEC 42001EU AI Act
Context contaminationMEASURE 2.5, 2.9A.6 / 8.2Art. 15 accuracy (provider); Art. 26 use per instructions (deployer)
Injection resistanceMEASURE 2.7, MANAGE 1.3A.5 / A.8Art. 15 cybersecurity; Art. 55 adversarial testing (GPAI provider)
Abstention disciplineMEASURE 2.5, MAP 3.49.1Art. 13 transparency; Art. 14 oversight support
Quantization robustnessMEASURE 2.68.3Annex XI technical documentation

Section VIII · Evidence manifest

Raw run outputs, per-axis JSON, and the question-set-size manifest are published in the model’s evidence locker, each with a SHA-256 checksum. Harness and rubric checksums are recorded in the audit bundle.

FileSHA-256
card.json52f748901adf73e05023c826…
manifest.json062aa4de279e2879235fc7ef…
raw_contamination.json4e6c5fe4645392aa5a8a70ce…
raw_governance.json92c4e92a702d1f93edf2b94a…
raw_injection.json9df192a10eb489025b6ce480…
records.jsonld190692c64efa41c1c60f783…

harness 9fa3835db3127f0a · rubric e2d5e396f7137d1e · run 2026-07-09T22:16:32

Section IX · Limitations, independence and conflict disclosure

Every axis in this report was measured at the replication target of n≥200; the numbers are final. A split grade reflects a model that sits on a band boundary within the measured interval, not incomplete data. hell.ai grades are computed from published sub-scores by a published rubric before any commercial activity; vendors cannot pay for inclusion, exclusion, or re-grading; public reports recommend control classes, never Black Sheep AI products by name. Findings describe the artifact as of the audit period; re-audit is recommended on any model revision or after 12 months. Full terms on Trust & independence.