Skip to main content Skip to docs navigation
Benchmark IT Solutions
Back to all blog posts

BlogCan You Trust the AI? Validating AI Tools Under FDA CSA and EU Annex 22

When the FDA issued its first warning letter citing the inappropriate use of AI in drug manufacturing in April 2026, the core problem was not the technology. It was trust without proof. The firm had put AI generated specifications, procedures and production records into use without validation or proper review, and reportedly told the FDA it had not known process validation was required, because its AI tool never raised it.

AI ValidationFDA CSAEU Annex 22
Category
Generative AI
Published
October 5, 2026
Topics
AI ValidationFDA CSAEU Annex 22
Can You Trust the AI? Validating AI Tools Under FDA CSA and EU Annex 22
5 min
Read time
The Article

Can You Trust the AI? Validating AI Tools Under FDA CSA and EU Annex 22

When the FDA issued its first warning letter citing the inappropriate use of AI in drug manufacturing in April 2026, the core problem was not the technology. It was trust without proof. The firm had put AI generated specifications, procedures and production records into use without validation or proper review, and reportedly told the FDA it had not known process validation was required, because its AI tool never raised it.

AI can check labels, review batch records, connect deviations and scan audit trails. But in a GMP setting, none of that counts unless the AI itself can be shown to work as intended. That is the job of validation. This article explains why AI needs a different approach to validation than traditional software, what US and European regulators expect, and how quality teams can put those expectations into practice.

Why AI needs different validation thinking

Traditional software follows fixed rules written by people. If the rules are right and the code follows them, the software behaves the same way every time, and validation confirms exactly that.

AI is different. It learns its behaviour from data rather than from rules written line by line. That raises new questions. Was the training data accurate and representative? How does the model perform on data it has never seen? Does it give the same answer to the same input? What happens if it is retrained? A sound validation approach must answer these questions, not just confirm that the screens work.

What the FDA expects

The FDA finalized its Computer Software Assurance guidance in September 2025. It was written for software used in medical device production and quality systems, but its thinking now shapes practice across life sciences and aligns closely with GAMP 5 Second Edition, which many pharma teams already follow.

The central idea is to focus assurance effort where it matters. Teams define the software's intended use and the decisions that depend on it, then concentrate testing on the features where a failure could affect product quality or patient safety. Lower risk features can be tested with less scripted, more exploratory methods, sound supplier testing can be relied on rather than repeated, and documentation should capture what builds confidence rather than everything that could be written down.

For AI specifically, the FDA's January 2025 draft guidance on using artificial intelligence to support regulatory decision making for drugs and biologics introduced a risk based framework for establishing the credibility of an AI model for a defined context of use. Together with the April 2026 warning letter, the direction is clear: AI is welcome in GMP, provided its output is controlled, verified and approved by the quality unit.

What Europe's Annex 22 will require

In July 2025, EU regulators released a draft Annex 22, the first GMP text written specifically for artificial intelligence, together with a revised Annex 11 on computerised systems and an updated Chapter 4 on documentation. Annex 22 covers AI used in the manufacture of medicines and active substances where it affects patient safety, product quality or data integrity.

The first draft took a cautious line. It limited critical GMP uses to static models with deterministic outputs and excluded generative AI and large language models from critical applications. That position is now evolving. At an EMA workshop in June and July 2026, discussion signalled that the final Annex 22 is likely to cover large language models and to focus on overarching, risk based control principles rather than rules for each technique. A final text is expected by the end of 2026.

Whatever the final wording, the core expectations of the draft are likely to remain. The intended use should be defined with process experts. Test data should be kept separate from training data, so results reflect real world performance. Acceptance criteria should be set before testing and be at least as good as the process the AI replaces. Reviewers should be able to see why a model reached its result and how confident it is. And once live, the model should be monitored, with any change managed through change control.

Putting validation into practice

Validating an AI tool follows a clear path. It begins with a written intended use: what the AI checks, what it flags and who acts on the result. A risk assessment follows, weighing the impact on product quality and patient safety against the level of human review around the AI's output. The more closely a person checks each result, the lower the risk the AI carries on its own.

Next comes test data that reflects real work, kept separate from training data and deliberately including known errors and difficult cases. Acceptance criteria are agreed before any test runs, such as the share of real errors the AI must catch and the highest acceptable rate of false flags. Tests are then run and recorded, with any failures explained.

Consider an AI tool that checks packaging artwork. Its intended use might be to flag differences between final artwork and the approved label copy, along with breaches of defined regulatory and brand rules, for a reviewer to confirm. A sensible test set would include real artwork with deliberately seeded errors: a wrong strength, an allergen not shown in bold, a swapped look alike character, an outdated date. Acceptance criteria would require every critical error to be caught while keeping false flags low enough that reviewers stay focused. Because each finding carries a severity level and a confidence score, and every decision is recorded, performance can be monitored in daily use, not just at validation.

Once the model passes, a single version is approved and locked under change control. The people who review its results are named, and the way their decisions are recorded is defined. Because validation does not end at go live, performance is monitored over time and formally reviewed on a set schedule, so any drift is caught early.

Choosing a vendor you can validate

Most manufacturers will buy AI tools rather than build them, which makes the supplier part of the validation story. A strong vendor can support validation for your intended use, explain how model versions are fixed and how updates are controlled, and show why every result was flagged. The tool should keep a full audit trail with user sign off, and the vendor should share its own testing evidence. Recognised security standards such as ISO 27001 and SOC 2 are a good sign that the supplier takes control seriously.

Validation is what turns a promising AI tool into one that a quality team, and an inspector, can rely on. Done well, it is not a barrier to using AI in GMP. It is the foundation that makes it possible.

Benchmark IT Solutions helps life sciences teams validate computerised systems, including AI, with a risk based approach in line with CSA. Talk to us about validating your next system.

Ready to take your business on the
path of success?