Last week, I posted my blog that mapped five shifts remaking clinical trial data archiving, from a compliance obligation into R&D infrastructure. This is the first deep dive: dimension one, cost center to strategic asset transformation, using archived data to win back time in study start-up and trial closeout.
FDA as part of Operation TrialBlazer, announced its Real-Time Clinical Trials (RTCT) pilot on June 22, 2026, a program where sponsors feed endpoints and safety signals to the agency continuously instead of only at prespecified analysis points, with the goal of speeding up early-phase decisions. While FDA expects to name its first pilot cohort soon, two proof-of-concept trials are already running as working models: AstraZeneca’s TRAVERSE and Amgen’s STREAM-SCLC.
Real-time reporting sounds like the opposite of archiving. It isn’t, a “real-time” signal only means something if the agency and the sponsor can compare it against something — prior historical safety patterns, dose-response data, past trial outcomes in the same or adjacent indications. Take away is that a clean, structured, retrievable record of what came before, and real-time data is just noise arriving faster. The RTCT pilot is a bet that sponsors’ historical trial data is finally organized well enough and easily searchable to be a live reference point, not just a filed report.
The RTCT bet already has a track record, just with a less futuristic name: external (or synthetic) control arms, where historical trial data substitutes for a newly recruited placebo group. Here are few examples from the past:
Koselugo, 2020. The FDA approved Koselugo for pediatric neurofibromatosis type 1 using an external control built by combining roughly 50 patients from a natural history study with roughly 50 patients from the placebo arm of the prior SPRINT Phase II trial. In an ultra-rare pediatric population recruiting a fresh placebo cohort would have been slow and not practical. The approval depended entirely on the older trial’s data being retrievable, patient-level, and clean enough to match from collected and archived data from years ago.
Alecensa, 2015. Roche built a synthetic control arm for its lung cancer drug Alecensa using existing trial data instead of running a new comparator study, and reaching an EU reimbursement decision roughly 18 months sooner than the traditional path.
Blincyto, 2014. Amgen used the same approach. Historical trial data standing in for a new control group to help accelerate approval of Blincyto in a rare leukemia indication, where enrolling a full control arm would have been a significant bottleneck.
None of these are hypothetical AI use cases and are FDA-cleared regulatory strategies, and every one of them only works because someone from years earlier, treated an already closed trial’s data as worth keeping in reusable shape: not just retrievable shape for an inspector, but structured enough for a statistician on a completely different program to match, harmonize, and defend to a regulator.
That’s the real shift behind “cost center to strategic asset.” It isn’t a promise about some future AI capability, it’s that the archive itself: patient-level, structured, harmonized across studies which is already making faster trial start-up and earlier approval possible. RTCT is the next step: instead of only reusing old trial data for a new comparison, it’s structuring data so a trial in progress can be measured against that archive in near real time. That only works if the underlying archive was built to be reusable and searchable in the first place, not just bolted together to survive the next regulatory inspection.
“Built to be reusable and searchable” is not a catchphrase, it’s a specific list of technical capabilities, and most of it is foundational data engineering rather than AI:
- THE FOUNDATION: Queryable & indexed storage instead of cold archives. Data retained in a repository with a multi-week retrieval SLA cannot feed a “real-time” comparison. It has to sit in warm, indexed storage reachable by API or query, built for statisticians and pipelines, not just inspectors.
- THE STRUCTURE: Standardized data models. Historical trial data has to be mapped to common data formats (CDISC’s SDTM and ADaM are the industry defaults), so a 2026 oncology trial and a 2018 oncology trial describe adverse events, doses, and endpoints the same way instead of a manual remapping project.
- THE CONTEXT: Rich, structured metadata, not just the data. Provenance — which protocol version, which site, which analysis population , which assay has to be captured at the record level. That’s what lets a statistician trust that a 2026 patient is comparable to a 2018 comparator and not being superficially similar.
- THE MEANING: A semantic layer across terminologies. Different trials tag the same clinical concept differently: legacy CRF field names, local site codes & MedDRA versions. Without a shared ontology or knowledge graph reconciling those differences, matching historical patients to a live trial’s cohort is a project, not a query.
- THE IDENTITY: Consistent de-identification and patient-level linkage. Comparing baseline characteristics across studies requires patient-level and not summary-level data which is de-identified the same way across trials and systems so a matching algorithm has something uniform to work against.
- THE TRUST: Full audit trail and lineage. Regulators accepting a historical comparator for RTCT or an external control arm, need to trace exactly which records, which version and which transformations fed the comparison. That’s a metadata and versioning requirement, not something bolted-on at submission time.
- THE CONSENT: Governance tied to original consent scope. Reusing a patient’s trial data for a new regulatory purpose years later has to be verified against the patient’s consent, which means consent scope has to exist as structured metadata, not text buried in a scanned PDF.
- POINT OF USE: Direct integration with statistical tooling. None of these matter, if a biostatistician still has to export, clean and reformat archived data by hand before comparing it to anything. The archived data needs to plug into the analytical pipelines such as R & SAS, that process the live trial’s incoming data.
Take any one of these capabilities away and “real-time” collapses back into an ad-hoc and manual reconciliation process: the rear-view-mirror problem this blog series opened with, just running on a faster clock.
If you work in clinical data management or biostatistics: how reusable is your organization’s historical trial data that is harmonizable and matchable by someone outside the original study team, or usable only by the people who built it?
Next in this series: dimension two — how the same archived data is becoming training data for AI/ML models, not just historical comparators.
References
- CASRAI, “FDA Real-Time Clinical Trials (RTCT) Pilot 2026: Where Real-Time Trials Stand” — casrai.org/news/fda-real-time-clinical-trials-rtct-pilot-2026
- Applied Clinical Trials, “HHS Launches Operation TrialBlazer to Restore US Leadership in Clinical Research” — appliedclinicaltrialsonline.com/view/hhs-operation-trialblazer-us-leadership-clinical-research
- Arnold & Porter, “FDA Proposes Expedited Investigational New Drug Pilot Program to Drive Early Phase Clinical Research in the United States” — arnoldporter.com/en/perspectives/advisories/2026/06/fda-proposes-expedited-investigational-new-drug-pilot-program
- Clinical Trials Arena, “External control arms: when can historical data substitute for placebos?” — clinicaltrialsarena.com/features/external-control-arms
- STAT News, “Synthetic control arms are a good option for some clinical trials” — statnews.com/2019/02/05/synthetic-control-arms-clinical-trials
- U.S. FDA, “FDA Actions to Accelerate and Modernize Early and Late-Stage Clinical Development” — fda.gov/industry/fda-actions-accelerate-and-modernize-early-and-late-stage-clinical-development
#ClinicalTrials #ClinicalDataManagement #DataArchiving #LifeSciences #RealWorldEvidence #ClinicalOperations #DrugDevelopment #DataStrategy
DISCLAIMER: THE CONTENT, VIEWS, AND OPINIONS EXPRESSED IN THIS BLOG ARE SOLELY THOSE OF THE AUTHOR(S) AND DO NOT REFLECT THE OFFICIAL POLICY OR POSITION OF SOLIX TECHNOLOGIES, INC., ITS AFFILIATES, OR PARTNERS. THIS BLOG IS OPERATED INDEPENDENTLY AND IS NOT REVIEWED OR ENDORSED BY SOLIX TECHNOLOGIES, INC. IN AN OFFICIAL CAPACITY. ALL THIRD-PARTY TRADEMARKS, LOGOS, AND COPYRIGHTED MATERIALS REFERENCED HEREIN ARE THE PROPERTY OF THEIR RESPECTIVE OWNERS. ANY USE IS STRICTLY FOR IDENTIFICATION, COMMENTARY, OR EDUCATIONAL PURPOSES UNDER THE DOCTRINE OF FAIR USE (U.S. COPYRIGHT ACT § 107 AND INTERNATIONAL EQUIVALENTS). NO SPONSORSHIP, ENDORSEMENT, OR AFFILIATION WITH SOLIX TECHNOLOGIES, INC. IS IMPLIED. CONTENT IS PROVIDED "AS-IS" WITHOUT WARRANTIES OF ACCURACY, COMPLETENESS, OR FITNESS FOR ANY PURPOSE. SOLIX TECHNOLOGIES, INC. DISCLAIMS ALL LIABILITY FOR ACTIONS TAKEN BASED ON THIS MATERIAL. READERS ASSUME FULL RESPONSIBILITY FOR THEIR USE OF THIS INFORMATION. SOLIX RESPECTS INTELLECTUAL PROPERTY RIGHTS. TO SUBMIT A DMCA TAKEDOWN REQUEST, EMAIL INFO@SOLIX.COM WITH: (1) IDENTIFICATION OF THE WORK, (2) THE INFRINGING MATERIAL’S URL, (3) YOUR CONTACT DETAILS, AND (4) A STATEMENT OF GOOD FAITH. VALID CLAIMS WILL RECEIVE PROMPT ATTENTION. BY ACCESSING THIS BLOG, YOU AGREE TO THIS DISCLAIMER AND OUR TERMS OF USE. THIS AGREEMENT IS GOVERNED BY THE LAWS OF CALIFORNIA.
-
White PaperEnterprise Information Architecture for Gen AI and Machine Learning
Download White Paper -
-
-