Murali Krishnam

CLINICAL TRIAL DATA ARCHIVING — PART 2/5

Part 1 of this series covered dimension one — cost center to strategic assets. This post covers dimension two: how rear-view mirror data becomes training data for AI/ML models. In January 2026, the FDA and EMA jointly published “Guiding Principles of Good AI Practice in Drug Development,” a shared framework for building, documenting, and validating AI models in drug development. It follows the FDA’s January 2025 draft AI guidance and builds on the 2024 “AI Council” from CDER. CDER has already reviewed over 500 AI-related submissions since 2016.

Read between the lines, though, and the real subject isn’t the model — it’s the data it was trained on. A model is only as credible as its training data: how representative, how consistently labeled, how traceable. Regulators are no longer just asking “does it work.” They’re asking “where did the training data come from, and can you prove it?”

That question already has a real answer — the same one this blog series opened with. AstraZeneca and Merck’s 2020 FDA approval of Koselugo (selumetinib) for pediatric neurofibromatosis type 1, first covered in Part 1, used an external control built from a natural history study and an earlier Phase II trial’s archived placebo arm, instead of enrolling a new placebo cohort in an ultra-rare pediatric population. It’s a published, quantified regulatory outcome — and it depended on the same prerequisite any AI training effort needs: archived data clean and structured enough to train on, and traceable enough to defend with regulators.

None of that happens by just pointing the model at a folder with old clinical trial data. Turning an archive into usable training data follows a fairly consistent pipeline and Koselugo already walked through most of it back then, without anyone calling it that at the time.

  • Harmonize to a common data model. Koselugo’s external control pooled roughly 50 patients from a natural history study with roughly 50 from the earlier SPRINT Phase II placebo arm. Neither source shared a data model, so the tumor volume, visit windows, and endpoint definitions all had to be remapped into one structure before any patient could be compared across sources.
  • Reconcile terminology and re-derive labels. “Response” and disease progression were recalculated against one consistent definition across both studies — not whatever each original study’s statisticians had used at the time.
  • De-identify and link at the patient level. Every record from both sources is tokenized the same way, so patients can be pooled without exposing identity or being double-counted if they appear in more than one source.
  • Engineer a standard baseline feature set. Age, baseline tumor volume, genotype, and disease severity are all extracted into one consistent feature vector — the same fields, measured the same way, for every patient regardless of the source.
  • Filter to an eligible cohort, then pair input and outcome. Apply the new trial’s eligibility criteria to both archives, narrowing the pool to the roughly 50 patients per source who actually qualify, then pair each patient’s baseline features (input) with their recorded outcome (target).
  • Match or model the outcome trajectory. Koselugo used the simpler approach: similarity-based matching, where comparable historical patients’ actual outcomes stand in as the comparator. The more complex alternative, a generative model sampling synthetic trajectories, is the newer, less-tested end of the same spectrum.
  • Hold out entire studies, in miniature. Koselugo pooled just two source studies, but the same principle applied: reviewers needed assurance the comparator patients weren’t influenced by treatment-arm data — the same isolation that holding out a full study protects at larger scale.
  • Calibrate against real outcomes. FDA reviewers had to see the matched patients’ baseline characteristics and trajectories actually line up with the treated cohort before accepting the comparison — the same check any version of this pipeline needs before its output is trusted.
  • Feed results back into the archive. Koselugo’s comparator was built once, for one submission. A reusable pipeline turns that into a continuously updated archive — the “life cycle oversight” the FDA/EMA principles call for, and what Part 1 blog’s Real-Time Clinical Trials pilot is reaching toward.
clinical-trial-data-archiving-part2

The FDA/EMA principles call for “robust management of information quality and provenance” — but stop short of saying how, since the framework has to flex across trial design, pharmacovigilance, and manufacturing alike. That’s where ISO/IEC 42001:2023 fills the gap. Published in December 2023 as the first international AI management standard, it requires certified organizations to run documented risk and bias assessments, keep data traceable across the AI life cycle, and maintain an audit trail regulators can review, which is why it’s already being positioned in life sciences as the practical way to build FDA, EMA, and MHRA trust ahead of any of this becoming mandatory.

That “how” breaks down into concrete requirements that map’s directly onto ISO 42001, and the FDA/EMA principles make them close to mandatory:

  • Representative, bias-audited cohorts (ISO 42001’s bias-assessment requirement). Sponsors must check who the archived trials enrolled and who they didn’t before training, or a “digital twin” just relearns old blind spots as clinical truth.
  • Structured labels, not free text (ISO 42001’s data-integrity requirement). Older trials often recorded outcomes in unstructured case report forms. Normalize them first into consistent, structured labels and not just whatever a coder happened to type years ago.
  • Study-level provenance (ISO 42001’s traceability requirement). Sponsors must know exactly which patients and studies sit in the training set versus the validation set, or a model can look validated while it’s just memorized data it already saw.
  • Versioned “data cards” (ISO 42001’s audit-trail requirement). The FDA/EMA principles push sponsors to document the training data itself — provenance, version, known limitations with the same rigor as the model.
  • Privacy-preserving computation for pooling across sponsors. No single archive has enough patients alone. Federated learning and synthetic data let multiple organizations train a shared model without raw patient data ever leaving its source system.
  • Drift detection (FDA/EMA’s life-cycle oversight principle). A model trained on 2015–2020 data needs a built-in check for when a new trial’s population no longer resembles what it learned from or “validated” simply means “validated on a population that no longer exists.”
  • A documented human review loop (FDA/EMA’s human-oversight expectation). Clinical experts still need a traceable way to challenge or override model outputs before they inform a regulatory decision, so the model doesn’t get the final word.
clinical-trial-data-archiving-regulation-mapping

Skip any one of these and “training data” is just old data with a new label stuck on it , which is exactly the kind of AI use regulators are now positioned to reject.

Getting the archive to “AI-ready” is only half the payoff, though. My colleague’s latest piece picks up right where this leaves off — what happens once that governed data gets asked questions directly, in plain language, with citations attached: From AI-Ready to AI-Activated: Data Sense and Data Ask Are Here.

If you work in biostatistics, data science, or regulatory affairs: when your organization talks about “using AI on our historical trial data,” has anyone documented the training data itself as rigorously as the model — provenance, representativeness, and all?

Next in this series: dimension three — how archival burden from disparate systems and M&A is turning into a unified advantage for cross-therapeutic research.

References

Murali Krishnam

Murali Krishnam

VP - Product Strategy, Enterprise Pharma AI

As the leader of Enterprise Pharma AI at Solix, Murali works across the life sciences ecosystem to drive technological innovations that enable scientific breakthroughs. He is dedicated to blending human knowledge with advanced technology to foster an outcomes-driven approach throughout the entire drug discovery and development process.

At Solix, Murali pioneered Semantic Content Library, an innovative solution that infuses meaning and contextual intelligence into multi-modal patient and clinical datasets, enabling the creation of AI-ready data products spanning drug discovery to post marketing.

DISCLAIMER: THE CONTENT, VIEWS, AND OPINIONS EXPRESSED IN THIS BLOG ARE SOLELY THOSE OF THE AUTHOR(S) AND DO NOT REFLECT THE OFFICIAL POLICY OR POSITION OF SOLIX TECHNOLOGIES, INC., ITS AFFILIATES, OR PARTNERS. THIS BLOG IS OPERATED INDEPENDENTLY AND IS NOT REVIEWED OR ENDORSED BY SOLIX TECHNOLOGIES, INC. IN AN OFFICIAL CAPACITY. ALL THIRD-PARTY TRADEMARKS, LOGOS, AND COPYRIGHTED MATERIALS REFERENCED HEREIN ARE THE PROPERTY OF THEIR RESPECTIVE OWNERS. ANY USE IS STRICTLY FOR IDENTIFICATION, COMMENTARY, OR EDUCATIONAL PURPOSES UNDER THE DOCTRINE OF FAIR USE (U.S. COPYRIGHT ACT § 107 AND INTERNATIONAL EQUIVALENTS). NO SPONSORSHIP, ENDORSEMENT, OR AFFILIATION WITH SOLIX TECHNOLOGIES, INC. IS IMPLIED. CONTENT IS PROVIDED "AS-IS" WITHOUT WARRANTIES OF ACCURACY, COMPLETENESS, OR FITNESS FOR ANY PURPOSE. SOLIX TECHNOLOGIES, INC. DISCLAIMS ALL LIABILITY FOR ACTIONS TAKEN BASED ON THIS MATERIAL. READERS ASSUME FULL RESPONSIBILITY FOR THEIR USE OF THIS INFORMATION. SOLIX RESPECTS INTELLECTUAL PROPERTY RIGHTS. TO SUBMIT A DMCA TAKEDOWN REQUEST, EMAIL INFO@SOLIX.COM WITH: (1) IDENTIFICATION OF THE WORK, (2) THE INFRINGING MATERIAL’S URL, (3) YOUR CONTACT DETAILS, AND (4) A STATEMENT OF GOOD FAITH. VALID CLAIMS WILL RECEIVE PROMPT ATTENTION. BY ACCESSING THIS BLOG, YOU AGREE TO THIS DISCLAIMER AND OUR TERMS OF USE. THIS AGREEMENT IS GOVERNED BY THE LAWS OF CALIFORNIA.