Sofpact field note

AI in Regulated Life Sciences: Every Claim Needs a Traceable Record

Editorial documentation chain from source evidence to approved record beside a continuous audit trail

In regulated life sciences the deliverable is never really the document. It is the ability to defend the document — years later, to an inspector, against a version of events reconstructed from records. That standard is what makes generative AI both attractive and awkward here. The drafting burden is enormous and genuinely wasteful; the evidentiary burden is unforgiving and does not relax because a machine helped.

The organisations getting this right are not the ones with the best model. They are the ones that decided early what a machine-assisted record has to prove.

Where AI actually lands in the documentation chain

Across the medicinal product lifecycle, the realistic near-term uses are narrower than the marketing suggests and more useful than the sceptics assume: summarising source material into a structured draft, checking a document against a template or style convention, identifying inconsistencies between related documents, retrieving relevant precedent from prior submissions, and translating between registered languages.

What these share is that a human remains the author. The machine compresses effort; it does not assume responsibility. The moment a workflow is designed so that no identifiable person has reviewed the substance of an output before it enters a regulated record, the organisation has created a problem that no amount of model quality resolves.

The regulator has published its thinking

The European Medicines Agency: use of artificial intelligence in the medicinal product lifecycle reflection paper sets out current thinking on AI and machine learning at any step from drug discovery through to the post-authorisation setting. Its framing is risk-based: the obligations scale with the consequence of the model being wrong, and a tool supporting early discovery is not held to the standard of one influencing a regulatory decision.

Two of its observations are directly relevant to documentation. The first is that models with very large numbers of trainable parameters, arranged in architectures that are not transparent, introduce risks that must be actively mitigated to protect patient safety and the integrity of clinical study results. The second is that active measures are required to minimise the integration of bias. Neither is satisfied by a general statement that outputs are reviewed. Both require something written down about how the specific risk was addressed for the specific use.

The accompanying European Medicines Agency: reflection paper on the use of artificial intelligence in the lifecycle of medicines announcement is a shorter entry point for colleagues who will not read the full paper.

Attribution is the requirement most workflows break

Regulated documentation carries expectations about data integrity that predate AI entirely: a record should be attributable to the person responsible for it, legible, contemporaneous, original and accurate, and it should remain complete, consistent and available over its retention period. Generative drafting stresses attribution in particular.

If a section of a report was machine-drafted from a set of source documents and then edited by a named author, the record needs to show what the author actually reviewed and approved, not merely that a review step was ticked. That means retaining the input set, the prompt or configuration, the model version, the raw output, the edits and the approval — and being able to reassemble them. A workflow that discards the intermediate output keeps only an assertion that review happened.

Version control of the model itself becomes part of this. If the tool is updated between the drafting of a document and a question about it eighteen months later, the organisation needs to know which version produced the text. Providers that silently update models are a poor fit for this environment unless the contract addresses notice, versioning and reproducibility explicitly.

Qualification, not just evaluation

The general AI practice of building an evaluation set applies here, but the regulated framing is stricter. A tool used within a quality system is expected to be qualified for its intended use: documented requirements, testing against those requirements, records of the testing, and controlled change thereafter.

Define intended use narrowly enough to test. “Assists with regulatory writing” cannot be qualified. “Produces a first draft of a defined section type from a specified set of approved source documents, for review by a qualified author” can be. Build the test set from representative real cases, including documents with missing sections, conflicting source values, unusual formats and edge terminology. Measure factual consistency with the source, not fluency.

The NIST AI Resource Center: AI RMF Core is useful on measurement discipline — documented methods, conditions resembling deployment, and explicit acknowledgement of the limits of generalising beyond what was tested. For organisations wanting a management-system structure around this, ISO: ISO/IEC 42001 artificial intelligence management systems provides one without requiring certification to be useful.

Automation bias is a documented risk, not a training issue

A fluent, well-structured draft that is confidently wrong is harder to review than a rough one. Reviewers reading machine-generated text consistently find fewer errors than reviewers reading human drafts of equivalent quality, because fluency reads as competence and the review becomes a proofread rather than a verification.

Design against this rather than train against it. Present the source passage alongside each generated claim so the reviewer verifies rather than recalls. Require explicit confirmation of numerical values against source rather than accepting them in flowing text. Sample completed documents independently, by someone who did not perform the original review, and track the error rate over time. If the sampled error rate is indistinguishable from zero across a large sample, that is more likely to indicate a weak sampling process than a perfect one.

Start where the source is already controlled

The best first candidates are documents assembled from material that is already under version control and already approved — a summary drawn from a locked study report, a section built from an approved investigator brochure, a translation of a registered text into another registered language. The provenance question is settled before the tool is involved, which removes the hardest part of the problem.

The worst first candidates are documents that synthesise across uncontrolled sources: working spreadsheets, email threads, drafts at different maturities. Those are exactly the places where the drafting burden feels heaviest, which is why teams reach for them first. They are also where a generated document is least defensible, because the record cannot show what the author actually relied upon.

Where the AI Act intersects

Life sciences organisations sit in a layered position. Existing pharmaceutical regulation applies. Where AI is embedded in a medical device or in vitro diagnostic already covered by Union harmonisation legislation, the European Commission: AI Act adds obligations that apply from 2 August 2028 following the Digital Omnibus agreement, subject to publication in the Official Journal. Internal documentation tooling generally does not fall in that category, but the determination should be made explicitly rather than assumed.

Where personal data is involved — clinical data, safety reports, investigator records — data protection law applies immediately and independently of any AI Act timetable.

What this cannot do

AI cannot resolve a documentation problem that is actually a data problem. If source values conflict between systems, if the study record is incomplete, or if the underlying analysis is contested, a generated document will present the confusion fluently and make it harder to spot. Nor can it carry professional responsibility: the qualified person signing a record is accountable for its content regardless of how the first draft was produced.

Stop where the intended use cannot be stated precisely enough to qualify, where the intermediate record cannot be retained, where the model version cannot be pinned, or where the time saved in drafting is consumed by verification. In that last case the tool has not improved the process — it has moved the work to a different person and added a risk.

Further reading: European Medicines Agency: AI in the medicinal product lifecycle; European Medicines Agency: reflection paper announcement; NIST AI Resource Center: AI RMF Core; ISO: ISO/IEC 42001; European Commission: AI Act.

See how a controlled AI Sprint qualifies one documentation workflow end to end.