If you are planning AI applications in the Quality or Regulatory area for products on the European market, two questions already constrain your choice of solutions: which uses of AI will hold up under the final Annex 22, and what is required to pass GxP validation?
Annex 22 is still a draft. But the architecture and vendor decisions you take now will have to hold once it lands, and waiting for the final text is not an option in the fast-moving pharma world. Here is how we read the draft, and how we would navigate it.
The short version
- •Still a draft, already shaping decisions. No adoption date is published, and recent EU GMP texts applied six to twelve months after publication. That is enough time to update procedures, not to rebuild a platform.
- •Critical applications need frozen, deterministic ML. Where AI output affects patient safety, product quality or data integrity, the draft rules out models that keep learning in production and models whose output varies from run to run. Today, that includes LLMs.
- •Generative AI still has a place: for supporting functions that do not impact patient safety or data integrity (drafting, summarising, search), with a qualified person responsible for the output.
- •The final stance on GenAI is highly contested: there are ongoing debates and the wording may change; however, the intent of the regulator regarding deterministic outcomes and the approach to validation are unlikely to change
- •Testing is what earns the time savings. The less a model is tested, the more of its output a person must review, up to every single output.
- •Our recommendation: design for the intent. Deterministic ML for applications in the critical path, GenAI around it under human review. That holds whether or not the final text relaxes the GenAI clause.
In this article
- 1. Where Annex 22 stands, and how much time you have
- 2. Where Annex 22 sits among the rules you already work under
- 3. Annex 22 key principles
- 4. Implications: which type of AI can be used, and where GenAI still fits
- 5. The Annex explained clause by clause
- Practical recommendations for safe AI deployment
- Implications and actions you can take now to build great AI foundations
- FAQ
1. Where Annex 22 stands, and how much time you have
“Annex 22: Artificial Intelligence” is a proposed new annex to EudraLex Volume 4, drafted by the EMA GMP/GDP Inspectors Working Group together with PIC/S. It runs to six pages and ten sections, from Scope to Operation.
- 7 July 2025: draft published for consultation, alongside draft revisions of Annex 11 and Chapter 4.
- 7 October 2025: consultation closed.
- 30 June – 1 July 2026: EMA multistakeholder workshop. EMA noted that the consultation “suggested support for potentially enabling the use of technologies such as generative AI (GenAI) or large language models (LLMs) in medicines manufacturing”, and that it is “still considering the implications”.
- Latest status as of October 2026: the EMA is planning to submit the final text to the commission in Q4 2026, with potential adoption in H1 2027.
How much time that leaves: Annex 11 was published in January 2011 and applied from 30 June 2011; Annex 15 in March 2015, applying from 1 October 2015; Annex 1 in August 2022, applying from August 2023. Six to twelve months is enough to update SOPs and validation plans. It is not enough to replace a system.
Two implications already:
The annex’s influence will not stop at the EU border. PIC/S co-drafted the annex, and PIC/S has 57 participating authorities, including the FDA, the MHRA, Swissmedic, Health Canada, the TGA and Japan’s PMDA. FDA and MHRA were cited as observers. Expect the reasoning to travel well beyond the EU.
The GenAI exclusion is the most debated clause, and it might change. EMA’s own summary of the consultation points of July 2026 mentioned the potential enablement of GenAI and LLMs under conditions rather than excluding them outright. If it softens, it will probably move towards the FDA position: GenAI allowed where output quality can be demonstrated and controlled. In practice that means constraining a probabilistic model until it behaves deterministically or reviewing most outputs. Bottom line: if you can achieve the desired outcomes with deterministic models, it’s an easier path.
2. Where Annex 22 sits among the rules you already work under
Annex 22 is not self-sufficient: it’s a set of refined requirements for AI within the broader Computer System Validation and AI regulation regulation. Here is an overview of the key pieces of regulation.

- Annex 11 deals with the software system, Annex 22 deals with the model. Annex 11 focuses on how to validate and control software systems. It is also being reworked, with a public draft and an expected finalisation in 2027. Annex 22 asks whether the model performs well enough for this use, on this site, for this process, with thoroughly justified choices on training data and acceptance criteria.
- The EU AI Act is a separate track. It is horizontal product law with obligations by risk class. An AI Act conformity assessment is not evidence for a GMP inspector, and an Annex 22 package is no defence under the AI Act.
- Accountability stays with the regulated entity, not the vendor, not the models. An AI output is a GMP record under Eudralex Chapter 4 like any other. Documentation “should be available and reviewed by the regulated user irrespective of whether a model is trained, validated and tested in-house or whether it is provided by a supplier or service provider” (Annex 22 §2.2).
How the EU draft compares with the US
| Annex 11 (EU) | Annex 22 draft (EU) | 21 CFR Part 11 (US) | FDA AI draft guidance (Jan 2025) | |
|---|---|---|---|---|
| Governs (or will govern when finalised) | The computerised system | The AI model inside the system | Electronic records and electronic signatures, by extension computerised systems | The AI models |
| Applies when | Any computerised system in GMP | Critical applications with direct impact on patient safety, product quality or data integrity | Records in electronic form created, modified, maintained, archived, retrieved or transmitted under any FDA records requirement (§11.1(b)) | AI that produces information or data to support regulatory decision-making regarding safety, effectiveness, or quality for drugs |
| Postition on GenAI / LLMs | Not addressed | “Should not be used” in critical applications; may be used in non-critical ones, with qualified personnel responsible for the outputs (HITL) | Not addressed | Not excluded; demonstrated model performance and adequacy (“credibility”) must be commensurate to the risk of use-case |
| What you must show | Validation, audit trail, data integrity, access control | Intended use, acceptance criteria, independent test data, explainability, confidence, drift monitoring | System validation, secure time-stamped audit trails, limited system access, operational and authority checks, signature controls (see details in section 5 below) | A 7-steps credibility assessment for the defined context of use (see below) |
The FDA position is less prescriptive, not less demanding. Its draft guidance asks for evidence commensurate with model risk rather than ruling out technologies. The first cGMP warning letter citing AI, to Purolea Cosmetics Lab on 2 April 2026, made the principle plain: “If you use AI as an aid in document creation, you must review the AI generated documents to ensure they were accurate and actually compliant with CGMP.” The failure was the missing review, not the AI.
For the Western market, Annex 22 is shaping up as the strictest standard. That makes it the safer baseline to design for.
The wider regulatory frame, in detail
European Key Regulations
EudraLex Volume 4, Chapters 4 and 7. Chapter 4 governs GMP records: what is recorded, by whom, how it is controlled and retained. Chapter 7 covers outsourced activities, which includes an AI supplier. A draft revision of Chapter 4 was published in July 2025 alongside the Annex 11 and Annex 22 drafts.
Annex 15 (Qualification and Validation) sets the general validation principles: validation master plan, URS, DQ/IQ/OQ/PQ, traceability, change control and periodic review. It refers to Annex 11 for computerised systems.
Annex 11 (Computerised Systems) sets the functional bar for GMP software: input accuracy checks, audit trails, access control, backup, data export and periodic evaluation. It runs to about three and a half pages; the Annex 22 draft adds considerably more testing and documentation for the model.
US key regulations
21 CFR Part 11 governs electronic records and signatures, and is the core of US CSV practice: authentication, data integrity, audit trails.
FDA Computer Software Assurance guidance (February 2026) is written for medical device production and quality system software, so it does not bind drug GMP directly. It is a useful signal of FDA thinking: classify software by intended use, rate process risk high or low, down to individual features, and set “assurance activities commensurate with the risk”.
FDA draft guidance on AI to support regulatory decision-making
The FDA draft guidance on AI to support regulatory decision-making (January 2025) covers all AI, not only ML. It proposes a seven-step credibility assessment:
- Define the question of interest the model addresses.
- Define the context of use.
- Assess model risk from “model influence” and “decision consequence”.
- Develop a plan to establish the credibility of the model output.
- Execute the plan.
- Document the results and any deviations from the plan.
- Determine whether the model is adequate for the context of use.
The documentation expected under step 4 is extensive: training data, model parameters, training choices, quality controls on the software involved and performance evaluation. It has to be built into the project, not assembled afterwards.
3. Annex 22 key principles
The first section of the Annex clarifies the major elements: where it applies, and what kind of AI is allowed.
Is the application critical? The annex applies “where Artificial Intelligence models are used in critical applications with direct impact on patient safety, product quality or data integrity, e.g. to predict or classify data”. Criticality attaches to the application, not the model: the same extraction model is critical when reading a Batch Record whose values feed release, and non-critical when summarising an internal report.
For such applications, two principles apply:
- Static models: The kind of models that are allowed: "models that do not adapt their performance during use by incorporating new data". Dynamic models that “continuously and automatically learn and adapt performance during use” should not be used in critical GMP applications. Retraining is not forbidden; learning in production is. A retrained model is a new version and a controlled change under §10.1.
- Deterministic outputs and no LLMs: Models that, “when given identical inputs, might not provide identical outputs” should not be used in critical GMP applications. The draft names GenAI and LLMs explicitly.
The Annex clarifies that LLMs remain available for non-critical applications, provided "personnel with adequate qualification and training should always be responsible for ensuring that the outputs from such models are suitable for the intended use, i.e. a human-in-the-loop (HITL)".
One clause in the scope section is easy to miss and worth paying attention to: "Models may consist of several individual models, each automating specific process steps in GMP." Ensembles of specialised models are explicitly foreseen. The unit of validation is each model against the step it automates, not the product as a single black box.
4. Implications: which type of AI can be used, and where GenAI still fits
Some cases are easy to categorise, and quite a few fall in the grey zone. This is where implementation choices and overall processes will come into the picture. Because the ultimate test is: if the model fails, is patient safety at risk? Is data integrity compromised?
| Document task | Zone | Type of AI / Implications |
|---|---|---|
| Automating production steps such as filling based on computer vision | Critical | Static, deterministic ML only |
| Extracting results from QC analysis and Batch Records to feed release decision | Critical | Static, deterministic ML only |
| Batch Record Review by Exception, where only anomalies are checked by QA | Critical | Static, deterministic ML only |
| Supported Batch Record review: completeness and GDP checks | Grey zone | The software handling the data is GxP relevant, the model may or may not be, depending on the checks and HITL steps |
| Pre-filling a deviation form that an analyst verifies field by field before signing | Grey zone | Treat as critical unless your risk assessment justifies otherwise; reduced testing needs documented operator responsibility |
| Deviation clustering and CAPA recommendations | Non-critical | GenAI acceptable; importance of verified and traceable data; qualified person responsible for the output data |
| Drafting SOP, URS or validation-plan text before review and approval | Non-critical | GenAI acceptable; qualified person responsible for the output |
| Summarising reports for reading, not filing; search for SOP with answers linked to the source | Non-critical | GenAI acceptable, traceability to the source matters for reliability |
For grey-zone and non-critical cases, the quality risk assessment will be an important fall-back when scrutiny happens: keep a detailed record of the rationale and responsible person.
The Purolea warning letter shows what the non-critical column looks like without the human step. AI agents generated specifications, procedures and master production and control records, and nobody verified them. The task was not the violation. The missing review and approval was.
5. The Annex explained clause by clause
| Clause | What it requires |
|---|---|
| §1 Scope |
|
| §2 Principles |
|
| §3 Intended use |
|
| §4 Acceptance criteria |
|
| §5 Test data |
|
| §6 Test data independency |
|
| §7 Test execution |
|
| §8 Explainability |
|
| §9 Confidence |
|
| §10 Operation |
|
One concept worth emphasising here: human-in-the-loop is a trade-off with depth of testing and model performance. The less a model is tested, the more of its output a person must check, up to every output.
§10.5: “Depending on the criticality of the process and the level of testing of the model, this may imply a consistent review and/or test of every output from the model, according to a procedure.”
Annex 22 documentation checklist
- Intended use, input sample space and subgroups (§3)
- Quality risk assessment (§2.3)
- Personnel involved and their qualifications (§2.1)
- Supplier documentation and your review of it (§2.2)
- Test data selection, labelling verification, pre-processing and exclusions (§5)
- Controls for test data and staff independence (§6.1–6.5)
- Test data usage records (§6.3)
- Test metrics and SME-approved acceptance criteria (§4)
- Measured performance of the process being replaced (§4.3)
- Approved test plan and test scripts (§7.2)
- Test results, deviations and their justification (§7.3–7.4)
- Feature attribution and feature review (§8)
- Confidence scores and thresholds (§9)
- Operator responsibility and human review records (§3.3, §10.5)
- Change and configuration control records (§10.1–10.2)
- Performance and input-drift monitoring metrics (§10.3–10.4)
Practical recommendations for safe AI deployment
Besides regulatory steps and administrative documentation, there are various ways to make AI safer to use and easier to check.
For sensitive but non-critical use cases, here are some solid measures:
- Ground: feed LLMs with validated, structured data to foster better answers
- Constrain: for document drafting, use fixed templates with built-in rules
- Reference: every statement needs to keep a trace to the source, ideally showing immediately on a split screen view to facilitate effective checks (not fake HITL)
- Attribute: log model, prompt and parameters per output
- Responsibility of the setup: who approved what, when and why
- Break it down: lengthy processes should involve multiple HITL along the chain, to avoid the black box effect and untraceable errors, or said otherwise “don’t overdo it”.
- Guardrail the confidence: uncertainty or lack of source leads to an “n/a” answer, the case goes straight to an SME
Implications and actions you can take now to build great AI foundations
Whilst the FDA takes a risk- and responsibility-based approach, the EMA is much more restrictive (no LLMs) and much more demanding on documentation and testing procedure. In the final version of Annex 22, the EMA may soften the stance, but it’s unlikely to significantly alleviate the validation burden and the requirements on deterministic or highly reliable outcomes.
The requirements of Annex 22 also go beyond the validation process itself and impact the whole AI value chain: baseline performance of the pre-AI process, details on training sets, independence of training and test data, independence of staff, which all need to be thought through at the beginning of each AI project. For those building AI applications internally, pay attention to the infrastructure required to provide traceable model training, testing and monitoring.
The question is no longer “what can I do with AI” or even “can I trust AI”, but rather “how to build strong AI foundations across use-cases that will withstand scrutiny and evolve over time”. For us, that means:
- Starting now and not waiting for regulators to draw up the next paper
- Building for the GMP intent: risk-based, with AI choices adequate to the risk, likely regulation and consequences of model failures
- Finding the right compromise between doing it right and not drowning teams in lengthy administrative documentation
- Building AI validation as an organisational skill that enables safe and reasonably fast deployments
- Consolidating providers of AI systems and building with them an AI validation playbook that creates economies of scale across model versions, sites and use-cases
Discuss how AI can support your efficiency initiatives in Quality and Regulatory
Acodis was built on deterministic ML models, per-version model freeze, confidence scores and a tamper-proof audit trail before Annex 22 was drafted, so projects on our platform are designed for the strict reading of the draft.
Start with one use case →Editors’ note: this article was written by a human, with support from AI for drafting and visualisations
Key resources
- Draft Annex 22 “Artificial Intelligence” (consultation draft, July 2025)
- EudraLex Volume 4, Chapter 4 “Documentation”; draft revision
- Annex 15 “Qualification and Validation”
- Annex 11 “Computerised Systems”; revised draft
- FDA 21 CFR Part 11
- FDA Computer Software Assurance guidance
- FDA draft guidance on AI to support regulatory decision-making
FAQ
Is Annex 22 in force yet?
No. It is a consultation draft from July 2025 and had not been adopted as of October 2026, with an expected submission to the commission in Q4 2026 and adoption in 2027. Based on recent EU GMP texts, expect six to twelve months between the final text and its application.
Does Annex 22 ban generative AI?
In critical GMP applications, the current draft says GenAI and LLMs “should not be used”. In non-critical applications they can be used, with qualified personnel responsible for the output. The consultation surfaced a strong push for enabling GenAI, so this clause may change.
What is the difference between Annex 11 and Annex 22?
Annex 11 governs the computerised system. Annex 22 adds requirements for the AI model inside it: intended use, acceptance criteria, independent test data, explainability, confidence and monitoring. Where both apply, you need both.
Does Annex 22 apply outside the EU?
Formally, it is EU GMP guidance. But it was co-drafted with PIC/S, whose 57 participating authorities include the FDA, the MHRA and Swissmedic, so expect its principles to shape inspections beyond the EU.
Is a document extraction model a static model?
Yes, if its parameters are frozen after training and it does not learn from production data. Fine-tuning before the freeze does not change that. Retraining creates a new version, which goes through change control.
Can an LLM draft GMP documents if the data comes from a validated system?
In our reading, yes, provided a qualified person reviews and approves the draft against the sources and that review is recorded. The LLM should not do the extraction itself. Treat it as a grey zone and document the rationale in your risk assessment.
Sources
- Draft Annex 22 “Artificial Intelligence”, EudraLex Volume 4, consultation draft, 7 July 2025 (all clause and glossary quotes).
- EudraLex Volume 4 Annex 11 “Computerised Systems”, revision 1, operative from 30 June 2011.
- EudraLex Volume 4 Annex 15 “Qualification and Validation”, operative from 1 October 2015.
- EudraLex Volume 4 Annex 1 “Manufacture of Sterile Medicinal Products”, published August 2022, operative from 25 August 2023.
- EudraLex Volume 4 Chapter 4 “Documentation”, and the draft revision published 7 July 2025.
- EMA, Multistakeholder workshop on expert contributions to AI guidance development (Annex 22), 30 June – 1 July 2026
- PIC/S, About PIC/S (57 participating authorities)
- Regulation (EU) 2024/1689 (AI Act), as amended by Regulation (EU) 2026/1744 (Digital Omnibus on AI).
- FDA, Computer Software Assurance for Production and Quality Management System Software, 3 February 2026.
- FDA, Considerations for the Use of AI to Support Regulatory Decision-Making for Drug and Biological Products, draft guidance, January 2025.
- FDA Warning Letter 722591, Purolea Cosmetics Lab, 2 April 2026.