The missing evidence class
Three frameworks tell you to measure deployment risk. None of them tell you how, and the existing public evaluations answer a different question. This page maps each hell.ai axis to the framework activity it can serve as evidence for.
This mapping is informative. It identifies which framework activities each measurement can support. It is not a conformity assessment, a certification, or legal advice, and it is not a determination that your system is high-risk. Provider obligations (EU AI Act Art. 55 for general-purpose models) and deployer obligations (Art. 26) are distinct; which apply to you is a question for your own counsel. Cite this as supporting evidence, not as proof of compliance.
Why capability and content-safety scores do not answer the question
NIST AI RMF asks you to MEASURE validity, reliability, and security in the deployed context. ISO/IEC 42001 clause 9.1 asks for performance evaluation of the AI system, not just the model. The EU AI Act requires accuracy and cybersecurity for high-risk deployers under Article 15, and model evaluation and adversarial testing for general-purpose models with systemic risk under Article 55, with enforcement from August 2 2026.
A capability leaderboard tells you what the model knows. A content-safety benchmark tells you whether it produces harmful text. Neither tells you what the model does when it is wired to a retrieval system and handed a document that is relevant but imperfect, or hostile. That behavior is the deployment-layer evidence these frameworks are asking for, and it is what this index produces.
The mapping
| hell.ai axis | NIST AI RMF | ISO/IEC 42001 | EU AI Act |
|---|---|---|---|
| Context contamination | MEASURE 2.5 validity & reliability; MEASURE 2.9 | A.6 impact assessment inputs; 8.2 | Art. 15 accuracy; Art. 55 model evaluation |
| Injection resistance | MEASURE 2.7 security & resilience; MANAGE 1.3 | A.5 / A.8 control evidence | Art. 15 cybersecurity (provider); Art. 55 adversarial testing (GPAI provider) |
| Abstention discipline | MEASURE 2.5; MAP 3.4 | 9.1 performance evaluation | Art. 13 transparency support; Art. 14 oversight |
| Retrieval uplift | MEASURE 2.6 deployment validity | 9.1 | Art. 15 accuracy (provider) |
| Governance cost | MANAGE 1.3 | 8.2 operational controls | Art. 14 human oversight support |
| Quantization robustness | MEASURE 2.6 | 8.3 change management | Annex XI technical documentation |
Subcategory citations are for orientation. Confirm the exact evidence requirements with your own counsel and assessor. A machine-readable version of this mapping is available on request at audit@hell.ai.