CLINICAL TRIAL DATA ARCHIVING — PART 3/5
Part 1 of this series covered cost center to strategic asset. Part 2 covered rear-view mirror data becoming AI training data. This blog is dimension three: archival burden from disparate systems becoming a unified advantage. What happens to the clinical archives when one acquires the other?
Part 1 of this series covered cost center to strategic asset. Part 2 covered rear-view mirror data becoming AI training data. This blog is dimension three: archival burden from disparate systems becoming a unified advantage. What happens to the clinical archives when one acquires the other?
The acquisition announcement is never the hard part. The hard part comes after close, when someone has to actually merge two archives that were never built to talk to each other. When done properly, the combined archive can support research neither company could have done alone. When done poorly, the acquirer just bought a second set of filing cabinets.
That question already has a real, well-documented answer. In 2018, Roche completed two separate acquisitions in oncology: a $1.9 billion purchase of Flatiron Health, which held EHR derived real-world data from a large U.S. oncology clinic network, and full ownership of Foundation Medicine, a genomic profiling company Roche had already held a stake in since 2015.
The combined Flatiron Health–Foundation Medicine Clinico-Genomic Database (CGDB) now underpins published research spanning far more than either archive could reach alone. A 2024 study in the peer-reviewed journal Nature Communications used it to analyze 78,287 patients across 20 different cancer types, identifying 776 genomic alterations linked to survival outcomes — work that depended on clinical data and genomic data, from two separately acquired companies, sitting in one queryable, patient-level model.
None of that happens automatically the moment two companies sign a merger agreement, though. Turning two acquired archives into one usable resource follows the below 8-step framework.
- Map both archives to one shared data model. Flatiron’s clinical fields and Foundation Medicine’s genomic fields never described a patient the same way. Both had to be remapped into one structure so a single patient record could carry both kinds of data.
- Link patients across systems without duplicating them. The same patient can appear in both: clinical records and genomic records. Consistent, privacy-preserving identifiers are required to let those two records merge into one patient.
- Reconcile terminology and coding standards. Different companies code diagnoses, treatments, and outcomes differently. Everything has to be re-expressed in one consistent vocabulary before it can be queried as a single data-set.
- Preserve provenance through the migration. Every record still needs to show which original company, which system, and which point in time it came from. Without this, a researcher years later has no way to judge how reliable or current it is.
- Carry consent scope forward, not just the data. Patients or trial participants consented to a specific use, under a specific company. That scope has to migrate as structured metadata alongside the record.
- Preserve the trial master file and essential documents. Any active clinical trials the acquired company sponsored come with regulatory obligations that don’t pause for the acquisition — the paperwork proving compliance has to move intact.
- Validate the merged archive before anyone relies on it. Before the combined data-set is used for new research, it has to be confirmed that records from both original sources still produce sensible, consistent results when queried together.
- Keep the unified archive open to the next acquisition. A biotech acquisition is not a one-time event for a large pharma. The data model built to absorb one acquired archive will be the same one that has to absorb the next one.

Neither Flatiron’s clinical data nor Foundation Medicine’s genomic data, on its own, could have supported research spanning 20 cancer types. Unified, together they did.
Two specific frameworks make the unification defensible rather than just convenient. The FAIR principles — findable, accessible, interoperable, reusable, first published by Wilkinson and colleagues in 2016 describe what “usable data” actually means at a technical level, regardless of which company originally created it. Alongside FAIR, the acquisition itself triggers specific regulatory continuity obligations: FDA’s rule on transferring sponsor obligations, and the international standard governing how clinical trial documentation has to survive a change in ownership.
Put together, that gives a concrete list of what a large pharma has to get right after an acquisition closes:
FAIR Principles:
- Persistent identifiers for every migrated record (FAIR — Findable). A record that loses its unique identifier in the migration is effectively lost, even if the underlying data survives the move.
- A documented access layer that survives the systems cutover (FAIR — Accessible). Researchers need a consistent, governed way to reach the combined archive — not two separate logins to two separate legacy systems.
- One shared data model across both legacy systems (FAIR — Interoperable). This is the step that made the Clinico-Genomic Database possible: clinical and genomic data structured so they can be queried together, not just stored together.
- Provenance kept at the record level (FAIR — Reusable). A future researcher has to be able to tell which original company and system a record came from, and trust it enough to reuse it in new work.
Regulatory Continuity:
- A written transfer-of-obligations record (FDA’s 21 CFR 312.52). When a sponsor’s obligations move to the acquiring company, U.S. regulations require a written description of exactly what transferred — anything left undocumented is treated as not transferred at all.
- Trial master file and essential documents preserved intact (ICH E6(R2)). The international Good Clinical Practice standard treats trial documentation as something that has to remain complete and traceable, ownership change or not.
- Data due diligence completed before the deal closes, not after. The UK Information Commissioner’s Office is explicit on this point for any merger or acquisition: figure out what data exists, its original purpose, and the lawful basis for using it, before integration ever begins.

Four requirements trace to a FAIR principle, two to a named regulation. Pre-close due diligence is the one practice that makes every other requirement achievable once the deal closes.
The Biotech’s Side of the Table
That last point — due diligence before close — is really a question for the biotech being acquired, not just the acquirer. Most biotechs eventually face one of two moments: a large pharma wants to partner or collaborate, or a large pharma wants to buy the company outright. Both moments reward the same preparation, done well before either conversation starts.
For partnership or collaboration, a biotech’s data needs to be genuinely FAIR before it’s ever shared outward. That means every data-set carries a persistent identifier, sits behind a documented access and permissions layer, follows a data model a partner’s systems can actually ingest, and comes with enough provenance and consent metadata that a partner’s regulatory team can trust it without a lengthy back-and-forth. A biotech that can hand a prospective partner a clean and well-documented data-set rather than a folder of spreadsheets and an explanation — moves faster into the collaboration and negotiates from a stronger position.
For an eventual acquisition, the same preparation pays off twice. The UK Information Commissioner’s due-diligence guidance applies just as much to the company being acquired as the one doing the acquiring: know what data you hold, what it can lawfully be used for, and whether your systems can integrate without “loss, corruption, or degradation of the data,” in the regulator’s own words. A biotech that keeps its trial master file complete, its consent scope documented as structured metadata, and its data model reasonably close to industry standards isn’t just easier to acquire — it’s worth more at the negotiating table, because the acquirer isn’t pricing in months of costly, uncertain data cleanup.
If you work in business development, data governance, or regulatory affairs at a biotech: if a large pharma wanted to see your data room tomorrow, would your archive hold up as FAIR and audit-ready, or would it need months of cleanup first?
Next in this series: dimension four — what changes when the same governed, unified archive has to support an entire combined portfolio, not just one product at a time.
FAQs
What is a unified clinical archive in biotech acquisitions?
A unified clinical archive is a governed data environment that combines clinical, genomic, trial, and regulatory archives from acquired companies into one interoperable system while preserving provenance, consent, and compliance metadata.
Why is data integration important after a biotech acquisition?
Data integration allows an acquiring organization to connect patient records, research datasets, and clinical trial information across legacy systems, enabling broader research while reducing duplicate data and improving governance.
How do FAIR principles support clinical data archiving?
FAIR principles help ensure that archived clinical data remains findable, accessible, interoperable, and reusable through persistent identifiers, governed access, standardized data models, and preserved provenance.
What data should be preserved during a clinical archive migration?
Organizations should preserve patient identifiers, provenance metadata, consent scope, trial master file documentation, essential regulatory records, coding standards, and historical audit trails during archive migration.
When should data due diligence happen during a biotech acquisition?
Data due diligence should begin before the acquisition closes so organizations understand what data exists, how it can legally be used, how it is structured, and what integration risks need to be addressed.
References
- CNBC, “Biotech M&A hits $106 billion, on track for best year since pre-Covid” (June 2026) — cnbc.com/2026/06/04/biotech-ma-dealmaking-pharma-106-billion.html
- Forbes, “Roche To Buy Big Data Upstart Flatiron Health For $1.9 Billion”; BioSpace, “Roche Acquires Remaining Shares of Foundation Medicine for $2.4 Billion” — forbes.com/sites/matthewherper/2018/02/15/roche-to-buy-big-data-upstart-flatiron-health-for-1-9-billion | biospace.com/roche-acquires-remaining-shares-of-foundation-medicine-for-2-4-billion
- Nature Communications, “Characterizing mutation-treatment effects using clinico-genomics data of 78,287 patients with 20 types of cancers” (2024) — nature.com/articles/s41467-024-55251-5
- Wilkinson et al., “The FAIR Guiding Principles for scientific data management and stewardship,” Scientific Data (2016) — nature.com/articles/sdata201618
- U.S. FDA / eCFR, 21 CFR 312.52, “Transfer of obligations to a contract research organization” — ecfr.gov/current/title-21/chapter-I/subchapter-D/part-312/subpart-D/section-312.52
- U.S. FDA, ICH E6(R2) Good Clinical Practice guidance — fda.gov/regulatory-information/search-fda-guidance-documents/e6r2-good-clinical-practice-integrated-addendum-ich-e6r1
- UK Information Commissioner’s Office, “Due diligence when sharing data following mergers and acquisitions” — ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/data-sharing-a-code-of-practice/due-diligence
#ClinicalTrials #ClinicalDataManagement #LifeSciences #MergersAndAcquisitions #FAIRData #Biotech #DrugDevelopment #DataStrategy
DISCLAIMER: THE CONTENT, VIEWS, AND OPINIONS EXPRESSED IN THIS BLOG ARE SOLELY THOSE OF THE AUTHOR(S) AND DO NOT REFLECT THE OFFICIAL POLICY OR POSITION OF SOLIX TECHNOLOGIES, INC., ITS AFFILIATES, OR PARTNERS. THIS BLOG IS OPERATED INDEPENDENTLY AND IS NOT REVIEWED OR ENDORSED BY SOLIX TECHNOLOGIES, INC. IN AN OFFICIAL CAPACITY. ALL THIRD-PARTY TRADEMARKS, LOGOS, AND COPYRIGHTED MATERIALS REFERENCED HEREIN ARE THE PROPERTY OF THEIR RESPECTIVE OWNERS. ANY USE IS STRICTLY FOR IDENTIFICATION, COMMENTARY, OR EDUCATIONAL PURPOSES UNDER THE DOCTRINE OF FAIR USE (U.S. COPYRIGHT ACT § 107 AND INTERNATIONAL EQUIVALENTS). NO SPONSORSHIP, ENDORSEMENT, OR AFFILIATION WITH SOLIX TECHNOLOGIES, INC. IS IMPLIED. CONTENT IS PROVIDED "AS-IS" WITHOUT WARRANTIES OF ACCURACY, COMPLETENESS, OR FITNESS FOR ANY PURPOSE. SOLIX TECHNOLOGIES, INC. DISCLAIMS ALL LIABILITY FOR ACTIONS TAKEN BASED ON THIS MATERIAL. READERS ASSUME FULL RESPONSIBILITY FOR THEIR USE OF THIS INFORMATION. SOLIX RESPECTS INTELLECTUAL PROPERTY RIGHTS. TO SUBMIT A DMCA TAKEDOWN REQUEST, EMAIL INFO@SOLIX.COM WITH: (1) IDENTIFICATION OF THE WORK, (2) THE INFRINGING MATERIAL’S URL, (3) YOUR CONTACT DETAILS, AND (4) A STATEMENT OF GOOD FAITH. VALID CLAIMS WILL RECEIVE PROMPT ATTENTION. BY ACCESSING THIS BLOG, YOU AGREE TO THIS DISCLAIMER AND OUR TERMS OF USE. THIS AGREEMENT IS GOVERNED BY THE LAWS OF CALIFORNIA.
-
White PaperEnterprise Information Architecture for Gen AI and Machine Learning
Download White Paper -
-
-