Cornerstone

The Regulatory Validation Framework

A verification-architecture framework for validating AI that drafts, verifies, and coordinates clinical trial decisions across FDA, EU, ANVISA, and CDSCO.

Regulatory18 min readUpdated February 1, 2026

Regulatory posture

This analysis reflects a risk-based quality-by-design posture. It describes reliance-oriented practice and does not imply endorsement by any regulatory authority.

A validation framework for AI in clinical trials is the set of rules, artifacts, and verification steps that make an AI-assisted decision reconstructible and defensible under inspection. The framework below proposes one built on a single primitive: a proof certificate produced at the moment of decision, gated by deterministic verification, and attested by a named human. The model proposes. The proof disposes.

Artificial intelligence is entering clinical trial operations at an accelerating pace. AI-assisted regulatory drafting, predictive patient-eligibility analysis, and multi-jurisdictional compliance coordination are no longer speculative. They are in deployment. The validation paradigms governing these systems, however, remain anchored in frameworks designed for diagnostic devices and software as a medical device, not for orchestration systems that draft, verify, and coordinate the documents and decisions of a trial across jurisdictions, from protocol and activation paperwork through consent events and closeout.

This document proposes a validation framework for that gap. It is an architecture proposal intended for critique, refinement, and co-development with the regulatory affairs community, not a product specification and not a claim of regulatory endorsement.

A note on posture. Nothing here asserts or implies endorsement of NexTrial by any regulator, agency, standards body, or individual, nor conformance to any named standard. Regulatory instruments are cited to show design alignment, not certification. Benefits are stated as design intent. The methodology described is GxP-aligned, not GxP-validated. Instruments are cited as in force at the date of publication; any operative system tracks the current text rather than the citation printed here.

The gap: device-era validation meets orchestration-era AI

The United States baseline for data integrity in regulated records is 21 CFR Part 11, and it predates the systems now in question. The FDA's most direct statement on AI in this domain, the draft guidance Considerations for the Use of Artificial Intelligence To Support Regulatory Decision-Making for Drug and Biological Products, was issued in January 2025 and remains a draft. It is built on a seven-step, risk-based credibility assessment framework anchored on a model's Context of Use: the description of where, and for what, a model's output is trusted.

That Context-of-Use spine is important, and this framework adopts it directly. But the draft reserves from its scope AI used for internal operational efficiency where the use does not bear on participant safety, drug quality, or the reliability of results, and it does not specify how a deterministic verification layer that gates regulated outputs is itself validated. A system that produces regulatory submissions, predicts patient trajectories for eligibility, or enforces compliance across jurisdictions bears directly on the reliability of results. It sits inside the guidance's concern. It is not yet addressed by the guidance's mechanics.

The EU has moved further in one direction. The EU AI Act classifies relevant healthcare AI as high-risk and imposes conformity assessment, technical documentation, risk management, and post-market monitoring obligations under Articles 9 through 14. But it provides limited guidance on how formal verification, mathematical proof rather than statistical estimation, maps onto conformity assessment.

The result is a field converging on one question from every jurisdiction at once:

When an inspector asks how an AI-assisted decision was made, what artifact answers?

The thesis: the verification layer is the admissibility layer

Every regulatory framework that governs clinical AI rests on a common requirement. A regulated decision must be reconstructible, from source data through the decision logic to the rule that produced it, at the moment an inspector asks. The requirement appears in 21 CFR Part 11, in ALCOA+, in ICH E6(R3), in EU AI Act Articles 11, 13, and 14, in Brazil's CFM Resolution 2.454, and in India's New Drugs and Clinical Trials Rules, 2019. The phrasing varies. The architectural requirement does not.

Two patterns currently dominate AI deployment in clinical research.

๐Ÿ”น The first treats the model's output as the regulated artifact and applies a compliance overlay to it after the fact.
๐Ÿ”น The second treats the model's output as a proposal that must be gated by deterministic verification before it becomes the regulated artifact.

Only the second pattern is admissible under the frameworks named above, regardless of the model's accuracy, confidence, or training methodology. A compliance overlay applied after the fact documents a decision that has already been made probabilistically. A deterministic gate produces the decision as a verifiable operation in the first place.

The rule that governs the rest of this framework follows from that distinction: the model proposes; the proof disposes. The verification layer, not the model, produces the regulated artifact. A probabilistic model is admissible as a participant in the decision pipeline only because it is gated by a deterministic check it cannot bypass. Anything an inspector or an ethics board sees is produced by the deterministic verification layer, not by the probabilistic substrate that fed it.

The primitive: the eight-property proof certificate

The framework's central architectural primitive is the proof certificate: a structured artifact produced at the moment an AI-assisted decision is made. It is specified to contain eight properties.

๐Ÿ”น 1. Rule invoked. The specific rule, by source, citation, and version: a regulatory provision, a protocol criterion, an SOP step, or a jurisdictional requirement.
๐Ÿ”น 2. Values verified. The exact patient, protocol, site, or operational values checked, listed rather than summarized, attributable to source.
๐Ÿ”น 3. Verification operation. The deterministic procedure that returned pass or fail on this specific decision, inspectable and reproducible.
๐Ÿ”น 4. Boundary statement. What the operation did not check, and the human-judgment factors reserved for the responsible human.
๐Ÿ”น 5. Risk classification. The decision's risk class, frozen at decision time, which indexes the rigor applied and the re-verification cadence.
๐Ÿ”น 6. Human reviewer identity. The identity and role of the human who attested, bound to the attestation.
๐Ÿ”น 7. Override and escalation record. Whether the human accepted, rejected, or asked for revision, with any override rationale and any escalation.
๐Ÿ”น 8. Evidence, not substitution. An explicit declaration that the operation is evidence presented to the reviewer, not a substitution for the reviewer's independent judgment.

The artifact is reproducible, inspectable, and designed to be admissible under inspection. A confidence score is none of these things. It cannot satisfy 21 CFR Part 11 traceability, ALCOA+ data integrity, ICH E6(R3) source verification, EU AI Act Articles 11, 13, and 14, or the analogous requirements of Brazil's CFM Resolution 2.454 and India's NDCTR 2019.

A precise word on what a proof certificate claims, because the precision is the point. It proves structural properties: that required fields are present, that references resolve, that no structural contradictions exist, that defined safety boundaries hold for a specific output. It does not prove that an output is clinically or regulatorily correct in the semantic sense. That judgment remains with the regulatory verification and, finally, the human. The value of the certificate is not that it certifies correctness. It is that it produces a deterministic, independently checkable, tamper-evident record of exactly what was verified and what was not.

The recommendations specify an artifact and a schema, not a vendor. Any system that produces a conforming certificate satisfies the standard, whether or not it resembles NexTrial's.

The three gates

The certificate is produced by a sequence of three verification gates. Each gate is a different substrate, and that difference is deliberate.

๐Ÿ”น Gate 1 โ€” Jurisdiction-specific regulatory compliance. A deterministic regulatory engine checks the proposed decision against the encoded requirements of the governing jurisdiction. Verification is triggered by the jurisdiction of the site an output governs, not applied as a universal checklist. This is where a regulation becomes a computable check.
๐Ÿ”น Gate 2 โ€” Formal structural proof. A formal verification step produces a proof certificate that either succeeds or fails, with no probabilistic middle ground, over the structural properties of the output. This is mathematical proof of structure, not statistical estimation of quality.
๐Ÿ”น Gate 3 โ€” Mandatory human attestation. A named human reviews the evidence and attests. The attestation is bound into the record. The human is the point of accountability, and the certificate is designed to make the human's judgment better informed, not to reduce or replace it.

Across development, two principles surfaced repeatedly and now organize the whole framework.

Risk classification is the architectural primitive. The rigor applied at each gate, the cadence of re-verification, and the level of human attestation required are all indexed to a risk class assigned and frozen at decision time. A high-risk decision and a low-risk one do not receive the same treatment, and the record shows which was applied.

AI verification is evidence, not substitution. The proof certificate is upstream evidence that an inspector or auditor consumes. It does not reduce the human reviewer's burden. This is the single most load-bearing correction to the "human in the loop" platitude: oversight is exercised over a decision, not over a probability. A reviewer cannot meaningfully validate a 0.918. A reviewer can validate a clause-by-clause verification.

Why the evidence has to be uncorrelated

Two properties make the verification layer trustworthy rather than merely present.

Uncorrelated evidence. The three gates are independent substrates that fail in different ways. A deterministic rule check can be wrong in ways a formal structural proof would catch. A structural proof can pass on a determination a human would reject. A human can catch what neither machine operation was scoped to see. Evidence drawn from substrates that fail differently is defensible in a way a single self-reported score is not.

A confidence score offers the opposite. It is generated by the same model whose output it scores, so it inherits that output's blind spots. It is correlated evidence wearing the label of a check. Much of what is presented in clinical AI as quality control, one agent checking another agent trained on the same data, has the same defect.

Structural boundaries, not bolt-on filters. The model in the proposing layer is constrained to non-creative, verifiable operations, so manipulation sits outside the operating envelope rather than being filtered out after the fact. A general-purpose model bolts a refusal filter onto a creative core; the capability remains and the filter is a probabilistic guess about what to suppress. Constraining the operating envelope is a stronger guarantee than filtering its output. The architecture is designed to store no documents and capture no patient data beyond what a verification operation requires, which keeps the attack surface deliberately small.

Proof versus legitimacy: validation under continuous learning

The hardest question in validating adaptive AI is what "validated" even means when the system keeps learning. The framework answers it by binding proof to a system state, not to a model.

A decision and its proof are fixed at a point in time. The proof attests to the exact regulation, protocol, and system state in force when the decision was made. The system continues to learn, but the proof holds for what it certified. To keep that true, each proof must remain reconstructible and defensible after any of three things change underneath it: the regulation, the protocol, or the system itself.

The mechanism has three parts:

๐Ÿ”น A signed state fingerprint that fixes the exact regulation, protocol, and system state at decision time.
๐Ÿ”น A bi-temporal lineage that distinguishes when a fact was true in the world from when the system recorded it, so a decision can always be replayed against the state that produced it.
๐Ÿ”น A predetermined change envelope that bounds how the adaptive system may evolve while remaining auditable, consistent with the direction of FDA's Predetermined Change Control Plan framework.

From this comes a distinction the framework treats as fundamental: proof is what was true of a state at decision time; legitimacy is what remains true as rules, protocols, and context move on. A point-in-time proof does not decay as the world changes. It holds for the state it certified. A "two-clocks" model names when re-certification is required for a decision to remain defensible today, separating the clock of the decision from the clock of the world.

Every proof is bounded within a declared Context of Use, the description of where and for what a given capability is trusted. That bounding is not a limitation to apologize for. It is the honest scope of the claim.

Jurisdiction as a first-class dimension

The architecture is global by construction. A new jurisdiction is not a new platform. It is a new adapter encoding that regime's instruments. The requirement surface each adapter draws on is deliberately specific.

๐Ÿ”น United States (FDA). 21 CFR Parts 11, 50, 54, 56, 312, 314, and 601; ICH E6(R3) and E8(R1); and the FDA AI/ML guidance trajectory, including the Predetermined Change Control Plan and the draft guidance on AI to support regulatory decision-making. The finalized Computer Software Assurance guidance anchors the assurance methodology. HIPAA governs protected health information where in scope.
๐Ÿ”น European Union (EMA / national competent authorities). Regulation (EU) No 536/2014, the Clinical Trials Regulation, submitted through CTIS; GDPR, with particular weight on Articles 6 and 9; ICH E6(R3); and the EU AI Act, which imposes conformity assessment and technical-documentation obligations on high-risk healthcare AI under Articles 9 through 14.
๐Ÿ”น Brazil (ANVISA / CFM). Lei nยบ 14.874/2024 governing trial conduct; the ANVISA regulatory dossier pathway; CFM Resolution 2.454, which requires that AI used in a regulated medical decision be explicable, and defines explicability as traceability of the decision to the rule it applied; and the LGPD. Brazil's significance here is not regulatory complexity. It is that Brazil legislated legal certainty and made decision-level traceability an explicit architectural requirement, which is precisely what a proof certificate produces.
๐Ÿ”น India (CDSCO). The New Drugs and Clinical Trials Rules, 2019, promulgated under the Drugs and Cosmetics Act, 1940; the Indian GCP guidelines; and the Digital Personal Data Protection Act, 2023.
๐Ÿ”น United Kingdom (MHRA). The updated UK clinical-trials framework effective April 2026, with submissions through IRAS; UK GDPR and the Data Protection Act 2018; and ICH E6(R3) as the GCP standard.

A single conforming proof certificate is designed to render against each of these evidentiary standards from one canonical representation, rather than producing a separate artifact per jurisdiction. And because a rule is a rule regardless of its source, the same deterministic operation applies whether the rule is a regulation, a protocol provision, a procedural step, or a jurisdiction-specific requirement. One trustworthiness artifact spans all four without a separate evidentiary regime for each.

Why this is urgent now: the real-time trajectory

The direction of travel sharpens the architectural question rather than softening it.

In April 2026 the FDA announced an initiative to advance real-time clinical trials, pairing oncology proof-of-concept trials that report endpoints and safety signals continuously with a request for information on a proposed pilot program for AI-enabled optimization of early-phase clinical trials. That request, Docket No. FDA-2026-N-4390 (91 FR 23100, April 29, 2026), is the document this framework was written to accompany.

Its significance for verification architecture is structural. When an inspector can see a decision as it is made, the distance between a decision and its inspection compresses toward zero. Procedural verification, which documents conformance after the fact, cannot operate at that cadence. Deterministic verification, which produces an inspectable artifact at the moment of the decision, can. Real-time trials do not make architectural verification optional. They make it the only kind that keeps up.

The framework is an invitation, not a conclusion

This framework is deliberately open on the questions its own architecture cannot resolve alone. The central one leads all the others:

Who certifies that a regulation's encoding into a computable check is correct?

That is the framework's principal unsolved problem. A proof certificate can prove that an encoded rule was applied faithfully to specific values. It cannot, by itself, prove that the encoding of the regulation into a machine-checkable form was a faithful reading of the regulation in the first place. That certification requires a forum that does not yet exist, one that includes regulators, ethics boards, principal investigators, methodologists, and multi-jurisdictional regulatory leads together.

Other questions remain genuinely open: how to validate jurisdiction-specific adapters when each region's rules update on independent cycles; how a system stays verified across model updates, regulatory changes, and data drift; how AI eligibility tools avoid amplifying existing enrollment inequities; how proof artifacts interoperate across borders absent a mutual-recognition agreement for AI-assisted clinical decisions; and how liability is apportioned when an AI-assisted decision is wrong. On the last point, the architecture's answer is the boundary statement itself: if the certificate records what the machine did and did not verify, and reserves the rest to a named human, then liability follows the boundary rather than disappearing into a five-party ambiguity.

The objective is to establish a validation standard through collaboration, before the regulatory landscape mandates one retroactively and the industry scrambles to comply. The disagreement is the work.

Provably right, not probably right. A defined set of structural properties proven with no probabilistic middle ground, and every remaining judgment managed by risk-based controls and accountable human attestation. That is the standard the framework proposes. The certificate is how it is kept.


Frequently asked questions

What is a regulatory validation framework for AI in clinical trials?
It is the set of rules, artifacts, and verification steps that make an AI-assisted clinical decision reconstructible and defensible under regulatory inspection. This framework proposes one built on a proof certificate produced at the moment of decision, gated by deterministic verification, and attested by a named human, so that any regulated decision can be reconstructed on demand.

What are the three verification gates?
Gate 1 is jurisdiction-specific regulatory compliance verification by a deterministic regulatory engine. Gate 2 is formal structural proof that produces a pass-or-fail proof certificate with no probabilistic middle ground. Gate 3 is mandatory attestation by a named human, who remains the point of accountability.

How does the framework validate AI that keeps learning?
Proof is bound to a system state rather than to a model, using a signed state fingerprint, a bi-temporal lineage, and a predetermined change envelope. A point-in-time proof holds for the exact regulation, protocol, and system state it certified, and a two-clocks model determines when a decision requires re-certification to remain defensible.

Does this framework claim FDA or EMA endorsement?
No. It is an architecture proposal offered under an engagement, not endorsement, posture. Regulatory instruments are cited to show design alignment. The methodology is described as GxP-aligned, not GxP-validated, and no regulatory certification is claimed.

How does the framework handle multiple jurisdictions?
Jurisdiction is treated as a first-class dimension. Each jurisdiction is a distinct adapter encoding that regime's instruments, from FDA and the EU AI Act to ANVISA, CFM 2.454, CDSCO, and MHRA. A single conforming proof certificate is designed to render against each jurisdiction's evidentiary standard from one canonical representation.

How does a proof certificate differ from a confidence score?
A confidence score reports the model's internal certainty and is generated by the same model whose output it scores, so it inherits that output's blind spots. A proof certificate records the rule, values, deterministic operation, boundary, risk class, and human attestation for a specific decision, and is reproducible under inspection. See Provably Right, Not Probably Right.

This framework reflects the work of NexTrial's technical and regulatory team and the regulatory affairs, ethics, and clinical operations practitioners who pressure-tested it through public co-development in 2026. Contributions are individual and do not imply institutional endorsement. Capabilities are described as design intent. Correspondence: Steven Thompson, Founder & CEO, NexTrial.ai.