If you are planning AI applications in the Quality or Regulatory area for products on the European market, two questions already constrain your choice of solutions: which uses of AI will hold up under the final Annex 22, and what is required to pass GxP validation?
Annex 22 is still a draft. But the architecture and vendor decisions you take now will have to hold once it lands, and waiting for the final text is not an option in the fast-moving pharma world. Here is how we read the draft, and how we would navigate it.
The short version
In this article
“Annex 22: Artificial Intelligence” is a proposed new annex to EudraLex Volume 4, drafted by the EMA GMP/GDP Inspectors Working Group together with PIC/S. It runs to six pages and ten sections, from Scope to Operation.
How much time that leaves: Annex 11 was published in January 2011 and applied from 30 June 2011; Annex 15 in March 2015, applying from 1 October 2015; Annex 1 in August 2022, applying from August 2023. Six to twelve months is enough to update SOPs and validation plans. It is not enough to replace a system.
Two implications already:
The annex’s influence will not stop at the EU border. PIC/S co-drafted the annex, and PIC/S has 57 participating authorities, including the FDA, the MHRA, Swissmedic, Health Canada, the TGA and Japan’s PMDA. FDA and MHRA were cited as observers. Expect the reasoning to travel well beyond the EU.
The GenAI exclusion is the most debated clause, and it might change. EMA’s own summary of the consultation points of July 2026 mentioned the potential enablement of GenAI and LLMs under conditions rather than excluding them outright. If it softens, it will probably move towards the FDA position: GenAI allowed where output quality can be demonstrated and controlled. In practice that means constraining a probabilistic model until it behaves deterministically or reviewing most outputs. Bottom line: if you can achieve the desired outcomes with deterministic models, it’s an easier path.
Annex 22 is not self-sufficient: it’s a set of refined requirements for AI within the broader Computer System Validation and AI regulation regulation. Here is an overview of the key pieces of regulation.
| Annex 11 (EU) | Annex 22 draft (EU) | 21 CFR Part 11 (US) | FDA AI draft guidance (Jan 2025) | |
|---|---|---|---|---|
| Governs (or will govern when finalised) | The computerised system | The AI model inside the system | Electronic records and electronic signatures, by extension computerised systems | The AI models |
| Applies when | Any computerised system in GMP | Critical applications with direct impact on patient safety, product quality or data integrity | Records in electronic form created, modified, maintained, archived, retrieved or transmitted under any FDA records requirement (§11.1(b)) | AI that produces information or data to support regulatory decision-making regarding safety, effectiveness, or quality for drugs |
| Postition on GenAI / LLMs | Not addressed | “Should not be used” in critical applications; may be used in non-critical ones, with qualified personnel responsible for the outputs (HITL) | Not addressed | Not excluded; demonstrated model performance and adequacy (“credibility”) must be commensurate to the risk of use-case |
| What you must show | Validation, audit trail, data integrity, access control | Intended use, acceptance criteria, independent test data, explainability, confidence, drift monitoring | System validation, secure time-stamped audit trails, limited system access, operational and authority checks, signature controls (see details in section 5 below) | A 7-steps credibility assessment for the defined context of use (see below) |
The FDA position is less prescriptive, not less demanding. Its draft guidance asks for evidence commensurate with model risk rather than ruling out technologies. The first cGMP warning letter citing AI, to Purolea Cosmetics Lab on 2 April 2026, made the principle plain: “If you use AI as an aid in document creation, you must review the AI generated documents to ensure they were accurate and actually compliant with CGMP.” The failure was the missing review, not the AI.
For the Western market, Annex 22 is shaping up as the strictest standard. That makes it the safer baseline to design for.
The wider regulatory frame, in detailEuropean Key Regulations
EudraLex Volume 4, Chapters 4 and 7. Chapter 4 governs GMP records: what is recorded, by whom, how it is controlled and retained. Chapter 7 covers outsourced activities, which includes an AI supplier. A draft revision of Chapter 4 was published in July 2025 alongside the Annex 11 and Annex 22 drafts.
Annex 15 (Qualification and Validation) sets the general validation principles: validation master plan, URS, DQ/IQ/OQ/PQ, traceability, change control and periodic review. It refers to Annex 11 for computerised systems.
Annex 11 (Computerised Systems) sets the functional bar for GMP software: input accuracy checks, audit trails, access control, backup, data export and periodic evaluation. It runs to about three and a half pages; the Annex 22 draft adds considerably more testing and documentation for the model.
US key regulations
21 CFR Part 11 governs electronic records and signatures, and is the core of US CSV practice: authentication, data integrity, audit trails.
FDA Computer Software Assurance guidance (February 2026) is written for medical device production and quality system software, so it does not bind drug GMP directly. It is a useful signal of FDA thinking: classify software by intended use, rate process risk high or low, down to individual features, and set “assurance activities commensurate with the risk”.
The FDA draft guidance on AI to support regulatory decision-making (January 2025) covers all AI, not only ML. It proposes a seven-step credibility assessment:
The documentation expected under step 4 is extensive: training data, model parameters, training choices, quality controls on the software involved and performance evaluation. It has to be built into the project, not assembled afterwards.
The first section of the Annex clarifies the major elements: where it applies, and what kind of AI is allowed.
Is the application critical? The annex applies “where Artificial Intelligence models are used in critical applications with direct impact on patient safety, product quality or data integrity, e.g. to predict or classify data”. Criticality attaches to the application, not the model: the same extraction model is critical when reading a Batch Record whose values feed release, and non-critical when summarising an internal report.
For such applications, two principles apply:
The Annex clarifies that LLMs remain available for non-critical applications, provided "personnel with adequate qualification and training should always be responsible for ensuring that the outputs from such models are suitable for the intended use, i.e. a human-in-the-loop (HITL)".
One clause in the scope section is easy to miss and worth paying attention to: "Models may consist of several individual models, each automating specific process steps in GMP." Ensembles of specialised models are explicitly foreseen. The unit of validation is each model against the step it automates, not the product as a single black box.
Some cases are easy to categorise, and quite a few fall in the grey zone. This is where implementation choices and overall processes will come into the picture. Because the ultimate test is: if the model fails, is patient safety at risk? Is data integrity compromised?
| Document task | Zone | Type of AI / Implications |
|---|---|---|
| Automating production steps such as filling based on computer vision | Critical | Static, deterministic ML only |
| Extracting results from QC analysis and Batch Records to feed release decision | Critical | Static, deterministic ML only |
| Batch Record Review by Exception, where only anomalies are checked by QA | Critical | Static, deterministic ML only |
| Supported Batch Record review: completeness and GDP checks | Grey zone | The software handling the data is GxP relevant, the model may or may not be, depending on the checks and HITL steps |
| Pre-filling a deviation form that an analyst verifies field by field before signing | Grey zone | Treat as critical unless your risk assessment justifies otherwise; reduced testing needs documented operator responsibility |
| Deviation clustering and CAPA recommendations | Non-critical | GenAI acceptable; importance of verified and traceable data; qualified person responsible for the output data |
| Drafting SOP, URS or validation-plan text before review and approval | Non-critical | GenAI acceptable; qualified person responsible for the output |
| Summarising reports for reading, not filing; search for SOP with answers linked to the source | Non-critical | GenAI acceptable, traceability to the source matters for reliability |
For grey-zone and non-critical cases, the quality risk assessment will be an important fall-back when scrutiny happens: keep a detailed record of the rationale and responsible person.
The Purolea warning letter shows what the non-critical column looks like without the human step. AI agents generated specifications, procedures and master production and control records, and nobody verified them. The task was not the violation. The missing review and approval was.
| Clause | What it requires |
|---|---|
| §1 Scope |
|
| §2 Principles |
|
| §3 Intended use |
|
| §4 Acceptance criteria |
|
| §5 Test data |
|
| §6 Test data independency |
|
| §7 Test execution |
|
| §8 Explainability |
|
| §9 Confidence |
|
| §10 Operation |
|
One concept worth emphasising here: human-in-the-loop is a trade-off with depth of testing and model performance. The less a model is tested, the more of its output a person must check, up to every output.
§10.5: “Depending on the criticality of the process and the level of testing of the model, this may imply a consistent review and/or test of every output from the model, according to a procedure.”Annex 22 documentation checklist
Besides regulatory steps and administrative documentation, there are various ways to make AI safer to use and easier to check.
For sensitive but non-critical use cases, here are some solid measures:
Whilst the FDA takes a risk- and responsibility-based approach, the EMA is much more restrictive (no LLMs) and much more demanding on documentation and testing procedure. In the final version of Annex 22, the EMA may soften the stance, but it’s unlikely to significantly alleviate the validation burden and the requirements on deterministic or highly reliable outcomes.
The requirements of Annex 22 also go beyond the validation process itself and impact the whole AI value chain: baseline performance of the pre-AI process, details on training sets, independence of training and test data, independence of staff, which all need to be thought through at the beginning of each AI project. For those building AI applications internally, pay attention to the infrastructure required to provide traceable model training, testing and monitoring.
The question is no longer “what can I do with AI” or even “can I trust AI”, but rather “how to build strong AI foundations across use-cases that will withstand scrutiny and evolve over time”. For us, that means:
Discuss how AI can support your efficiency initiatives in Quality and Regulatory
Acodis was built on deterministic ML models, per-version model freeze, confidence scores and a tamper-proof audit trail before Annex 22 was drafted, so projects on our platform are designed for the strict reading of the draft.
Start with one use case →Editors’ note: this article was written by a human, with support from AI for drafting and visualisations
Key resourcesNo. It is a consultation draft from July 2025 and had not been adopted as of October 2026, with an expected submission to the commission in Q4 2026 and adoption in 2027. Based on recent EU GMP texts, expect six to twelve months between the final text and its application.
In critical GMP applications, the current draft says GenAI and LLMs “should not be used”. In non-critical applications they can be used, with qualified personnel responsible for the output. The consultation surfaced a strong push for enabling GenAI, so this clause may change.
Annex 11 governs the computerised system. Annex 22 adds requirements for the AI model inside it: intended use, acceptance criteria, independent test data, explainability, confidence and monitoring. Where both apply, you need both.
Formally, it is EU GMP guidance. But it was co-drafted with PIC/S, whose 57 participating authorities include the FDA, the MHRA and Swissmedic, so expect its principles to shape inspections beyond the EU.
Yes, if its parameters are frozen after training and it does not learn from production data. Fine-tuning before the freeze does not change that. Retraining creates a new version, which goes through change control.
In our reading, yes, provided a qualified person reviews and approves the draft against the sources and that review is recorded. The LLM should not do the extraction itself. Treat it as a grey zone and document the rationale in your risk assessment.