Murali Krishnam

On June 22, 2026, HHS and the FDA launched Operation TrialBlazer, a sweeping push to speed up U.S. drug development after China overtook the U.S. in registered clinical trials (39% of the global total in 2024, per HHS). Alongside faster Phase 1 pathways and a shift toward single-pivotal-trial approvals, the initiative leans hard into adaptive master protocols — basket, umbrella, and platform trials that share infrastructure and control-arm data across multiple indications at once.

That last piece is the one worth sitting with if you work anywhere near clinical data. Master protocols only function if the data behind them is interoperable across studies and therapeutic areas, which is close to the opposite of how most clinical trial archives are built today.

Evolution of clinical data archiving

For decades, clinical trial data archiving has sat at the bottom of pharma’s priority list, a compliance obligation, not a competitive one. TrialBlazer, layered on top of a wave of other regulatory change, is about to test whether that’s still good enough.

Start with the regulatory floor moving under our feet. The UK’s MHRA just extended the mandatory retention period for Trial Master Files from 5 years to 25 years, effective April 2026, one of the biggest updates to clinical trials regulation in two decades, and one that also pushes alignment with ICH E6(R3). The FDA’s own Data Modernization Action Plan has been pressing the industry toward standardized, cloud-native, AI-ready data infrastructure well before TrialBlazer arrived. The message from regulators everywhere is converging: trial data isn’t something you generate and forget. It’s something you’re accountable for, sometimes for a quarter of a century — and now, something regulators expect to move faster and connect further than it currently does.

At the same time, McKinsey’s research on biopharma R&D IT found that while 40–50% of top-20 pharmaceutical companies have invested heavily in modernizing clinical IT applications, many still haven’t achieved a clear ROI. Legacy, application-centric architecture with fragmented, siloed, hard to connect are the reason why, and every merger or acquisition tends to bolt on another disconnected system rather than remove one. The outcome is familiar to anyone who’s tried to pull an old study’s data for a new analysis: technically retained, practically unusable. Compliant, and completely inert.

Here’s the shift worth watching: leading organizations have stopped treating their archives as a rear-view mirror and started treating them as a forward-looking asset. It’s playing out across five concrete dimensions:

  • Cost center → strategic asset. Data once maintained purely to survive an audit is becoming a lever for faster study start-up and quicker trial closeout, because history informs the design of the next protocol instead of just defending the last one.
  • Rear-view mirror data → training data. Records once opened only for inspections are now being curated, sampled, and subsetted specifically to train and validate the AI/ML models being built across R&D.
  • Fragmentation from disparate systems and M&A → unified archival advantage. Instead of every acquisition bolting on another disconnected repository, a unified archive turns decades of inherited systems into one corpus that supports cross-therapeutic research, not just single-study lookups.
  • Application-centric silos → real interoperability. Architectures built around whichever tool happened to capture the data are giving way to structures organized around the site and the patient, so data moves across systems instead of dying inside them.
  • Application-specific search → study-level retrieval and reproducibility. Instead of hunting through the original capture tool, teams can search, retrieve, and reproduce results at the study level — exactly what both regulators and AI pipelines need.

Novartis’s data42 program is the clearest public proof point of what this looks like at scale. Instead of leaving legacy trial data scattered across disconnected repositories, Novartis consolidated more than 2,000 clinical studies, roughly 2 million patient-years and 20 petabytes of R&D data onto a single platform where machine learning models could be trained across studies instead of trapped within one. Teams have used it to identify disease subtypes in rheumatoid arthritis and map disease progression in oncology — analysis that was effectively impossible when the same data sat siloed by study and by system.

That’s the real transformation underway in clinical data archiving: from a cost center justified by regulatory necessity, to a strategic asset that fuels AI/ML training data, cross-therapeutic research, and faster study start-up and closeout. The organizations pulling ahead aren’t just storing data for longer. They’re making it findable at the study level, interoperable across systems, and structured with the metadata needed for traceability and reuse years later.

The 25-year retention clock is now running in the UK, and similar pressure is building globally. That’s either 25 years of accumulating dead weight, or 25 years of compounding strategic value and the difference comes down to whether archiving is designed as an afterthought or built as infrastructure from day one.

A lot of the interoperability and metadata work above only holds up if it doesn’t quietly become another job your team has to maintain by hand forever. Solix Technologies has a good related read on that exact trap by my colleague Mark Lee — “Is Your AI Working for You — or Are You Working for Your AI?”:

If you work in clinical operations, data management, or R&D IT: is your organization still archiving for compliance, or has it started architecting for reuse?

References:

#ClinicalTrials #ClinicalDataManagement #LifeSciences #DigitalHealth #PharmaAI #ClinicalOperations #DataStrategy

Murali Krishnam

Murali Krishnam

VP - Product Strategy, Enterprise Pharma AI

As the leader of Enterprise Pharma AI at Solix, Murali works across the life sciences ecosystem to drive technological innovations that enable scientific breakthroughs. He is dedicated to blending human knowledge with advanced technology to foster an outcomes-driven approach throughout the entire drug discovery and development process.

At Solix, Murali pioneered Semantic Content Library, an innovative solution that infuses meaning and contextual intelligence into multi-modal patient and clinical datasets, enabling the creation of AI-ready data products spanning drug discovery to post marketing.

DISCLAIMER: THE CONTENT, VIEWS, AND OPINIONS EXPRESSED IN THIS BLOG ARE SOLELY THOSE OF THE AUTHOR(S) AND DO NOT REFLECT THE OFFICIAL POLICY OR POSITION OF SOLIX TECHNOLOGIES, INC., ITS AFFILIATES, OR PARTNERS. THIS BLOG IS OPERATED INDEPENDENTLY AND IS NOT REVIEWED OR ENDORSED BY SOLIX TECHNOLOGIES, INC. IN AN OFFICIAL CAPACITY. ALL THIRD-PARTY TRADEMARKS, LOGOS, AND COPYRIGHTED MATERIALS REFERENCED HEREIN ARE THE PROPERTY OF THEIR RESPECTIVE OWNERS. ANY USE IS STRICTLY FOR IDENTIFICATION, COMMENTARY, OR EDUCATIONAL PURPOSES UNDER THE DOCTRINE OF FAIR USE (U.S. COPYRIGHT ACT § 107 AND INTERNATIONAL EQUIVALENTS). NO SPONSORSHIP, ENDORSEMENT, OR AFFILIATION WITH SOLIX TECHNOLOGIES, INC. IS IMPLIED. CONTENT IS PROVIDED "AS-IS" WITHOUT WARRANTIES OF ACCURACY, COMPLETENESS, OR FITNESS FOR ANY PURPOSE. SOLIX TECHNOLOGIES, INC. DISCLAIMS ALL LIABILITY FOR ACTIONS TAKEN BASED ON THIS MATERIAL. READERS ASSUME FULL RESPONSIBILITY FOR THEIR USE OF THIS INFORMATION. SOLIX RESPECTS INTELLECTUAL PROPERTY RIGHTS. TO SUBMIT A DMCA TAKEDOWN REQUEST, EMAIL INFO@SOLIX.COM WITH: (1) IDENTIFICATION OF THE WORK, (2) THE INFRINGING MATERIAL’S URL, (3) YOUR CONTACT DETAILS, AND (4) A STATEMENT OF GOOD FAITH. VALID CLAIMS WILL RECEIVE PROMPT ATTENTION. BY ACCESSING THIS BLOG, YOU AGREE TO THIS DISCLAIMER AND OUR TERMS OF USE. THIS AGREEMENT IS GOVERNED BY THE LAWS OF CALIFORNIA.