The Article
Data Integrity with AI: Keeping GMP Records Complete and Trustworthy
A batch is released because its test results passed. A process is approved because its validation data met the criteria. A deviation is closed because the investigation found the cause. Every one of these decisions rests on data. If the data cannot be trusted, neither can the decision.
That is why data integrity remains one of the sharpest focus areas in GMP inspections. An independent review of 85 FDA warning letters issued to drug manufacturers in 2025 found that 15% cited data integrity concerns, with some regions far more affected than others. This article explains what data integrity means in practice, why audit trail review is so hard to do well, and how AI helps quality teams find problems before an inspector does.
What data integrity really means
Regulators describe trustworthy data using the ALCOA+ principles. Data should be attributable, so you can tell who created or changed it. It should be legible, now and in the future. It should be contemporaneous, meaning recorded at the time the work was done, and original, or a certified true copy. And it should be accurate. The "plus" adds four further expectations: data should be complete, consistent, enduring and available whenever it is needed.
These principles apply to paper and electronic records alike. For electronic systems, the most important evidence that they are being met lives in the audit trail.
The regulatory picture
Regulators have set out their expectations in detail, including the FDA's 2018 guidance on data integrity and compliance with drug CGMP, the MHRA's 2018 GXP data integrity guidance and PIC/S PI 041. In the US, 21 CFR Part 11 already requires secure, computer generated audit trails for electronic records.
The bar is rising. The draft revision of EU GMP Annex 11, released in July 2025, expands from 5 pages to 19 and places far more weight on audit trails, access control and system security. The EMA has indicated that final texts of the revised Annex 11 and the new Annex 22 on artificial intelligence are expected by the end of 2026.
Why audit trail review is so hard
An audit trail records who did what, when and why within a computerised system. That makes it the best evidence of data integrity. It also makes it enormous. A single laboratory or manufacturing system can log thousands of entries a week, and a site may run dozens of such systems.
Faced with that volume, many teams review summary reports rather than the audit trail itself, or check only a small sample. Hybrid environments, where paper and electronic records are mixed, make the full picture harder still to see. Problems in the entries nobody reviews can stay hidden until an inspector finds them.
How AI changes audit trail review
AI makes it possible to read every audit trail entry, across systems, and flag the patterns that deserve a closer look. Some of these patterns are well known to data integrity specialists: a test result edited or reprocessed after it was first recorded, a chromatography run deleted, aborted or repeated without a documented reason, or a change made with no reason recorded at all.
Other patterns are much harder to spot by hand. Activity at unusual hours or outside a user's normal role. Several people working under one shared login. Long gaps between when work was done and when it was recorded, which raises questions about whether records are truly contemporaneous.
AI ranks each flag by risk and links it directly to the original entries. Reviewers can start with the issues that matter most and see the evidence for themselves, rather than searching for it.
Building AI into a data integrity programme
AI works best inside a clear, risk based review plan. The starting point is deciding which systems and data are most critical to product quality, such as laboratory systems that generate release results. Audit trails from systems such as LIMS, chromatography data systems, MES and the QMS can then be brought together, so patterns across systems become visible for the first time.
The flags themselves should combine firm rules, such as any deletion of a result, with pattern detection for behaviour that looks unusual. A trained reviewer checks each flag and records the outcome. Confirmed issues move into the deviation and CAPA process, so they are investigated and fixed at the root rather than simply noted.
The same principles apply to AI tools themselves, including those used for label and artwork review. Every finding a label check raises, and every reviewer's decision to accept or reject it, is GMP data. It should be attributable to a named user, recorded at the time and kept in a secure audit trail, in line with 21 CFR Part 11 and EU Annex 11. When choosing any AI tool for GMP work, ask how it meets ALCOA+, not only how accurate it is.
Guardrails that matter
AI does not judge intent. A flag shows that something unusual happened, not why. Many flags will have simple, legitimate explanations, and people must investigate before drawing any conclusions.
The AI tool also needs controls of its own. It must be validated for its intended use, protected by access controls and maintain its own audit trail. Because audit trails reveal individual user activity, clear rules on who can view the results and how personal data is handled are essential.
Data integrity has always depended on people doing the right thing and on systems that make the right thing easy. AI adds a third layer: the ability to see every record rather than a sample, and to catch problems early, while they are still easy to fix.
Benchmark builds data and AI solutions for regulated industries, from connecting audit trail sources to flagging high risk activity for review. Our teams work to ISO 27001 and SOC 2 standards. Talk to us about making audit trail review complete rather than sampled.