← Notes Compliance

The August 2 problem

On August 2 2026 the EU AI Act’s general-purpose model obligations become enforceable. The question your assessor will ask is one that capability scores and content-safety benchmarks do not answer.

From that date, providers of general-purpose models with systemic risk owe model evaluation, adversarial testing, and technical documentation to the AI Office under Article 55. Those are the model provider’s obligations, not yours as a deployer. Your obligations run through Article 26, which requires you to use a high-risk system per its instructions and to exercise human oversight, and the threshold question is whether your internal assistant is high-risk under Annex III at all. Penalties for the provider tier run to the higher of fifteen million euro or three percent of global turnover. Either way the obligations are documentation obligations, which means someone has to produce evidence, and a deployer who cannot describe how the chosen model behaves has a gap.

Now look at what the common evidence sources actually measure. A capability leaderboard shows what a model knows. A content-safety benchmark shows whether it produces harmful text. Both are real and neither is the deployment question. When your model is connected to a retrieval system and handed a document that is relevant but imperfect, or one that carries a hostile instruction, what does it do? That behavior is what the Article 15 accuracy and cybersecurity requirements on providers, and the Article 55 adversarial-testing requirements on general-purpose model providers, are ultimately about, and it is not on any leaderboard. As a deployer you inherit the consequences of it even where the primary obligation sits upstream.

The evidence you will be asked for

Practically, an assessor wants dated measurements, sample sizes, and confidence intervals, tied to a named methodology, describing how the deployed system behaves under adversarial and imperfect input. That is the shape of a deployment-risk audit: a paired closed-book versus with-context accuracy measurement, an injection defense ladder, and an abstention measurement, each mapped to the framework activity it supports.

We publish that mapping on the framework page. It is informative, not a conformity assessment, and it is written so a compliance reviewer can see which Article 15 and Article 55 activities each measurement can serve as evidence for. The measurements themselves are on the index, per model, with the caveats stated in full.

August 2 is not far. The models are already deployed. The evidence is the part that is missing.