Murali Krishnam

 

How accepting fragmented, unstructured biomedical data, rather than fighting it is becoming the foundation of AI-native R&D

This past summer, the FDA began one of the more consequential experiments in modern drug development. Under an initiative housed in a new department-wide program launched in June 2026, two active clinical trials: a Phase 2 study in mantle cell lymphoma and a Phase 1b study in small cell lung cancer started reporting safety and efficacy signals to regulators continuously, rather than at a handful of prespecified checkpoints that have defined trial oversight for decades. The idea is simple to state and hard to execute: if regulators can see a trial’s data as it happens, unsafe or ineffective therapies can be caught sooner, and promising ones can move faster [1, 2].

There’s a catch, though, and it’s an instructive one. As of this writing, the agency still hasn’t published the technical specifications for how that real-time data should actually be formatted, standardized, or transmitted. Two live trials are streaming data into a pipeline whose plumbing isn’t fully built yet. That gap between the ambition to move faster and the unglamorous work of making data consistently interpretable is not a footnote. It is, in miniature, the central challenge facing AI-driven R&D today.

The unstructured-data problem, deep rooted in pharma

There’s a strategic insight on enterprise AI adoption that generalizes well beyond any one sector: most of what an organization actually knows lives in documents, images, transcripts, and notes rather than tidy database rows. The instinct, faced with this, is to try to centralize everything: pull every document, every record, every file into one master repository before doing anything intelligent with it. That instinct is usually wrong. Consolidation projects of that scale can take from months to years, and by the time they finish, the underlying systems have already changed. The more durable approach is to accept that fragmentation is here to stay and rather build a layer that can derive shared meaning, or context, from data wherever it already lives, and to connect that layer to as many sources as possible rather than forcing them into one.

Drug discovery and development is, if anything, a more extreme version of this problem. A single therapeutic program generates multi-omics data, imaging, chemical & pharmacological records, preclinical study notes, and clinical information spanning structured databases and a mountain of unstructured material: physician narratives, pathology & radiology reports, adverse-event descriptions, investigator comments, lab notebooks, and regulatory submission text. These live across electronic lab notebooks, clinical trial management systems, hospital records, and regulatory portals that were never designed to talk to one another. It’s telling that one frequently repeated estimate in the clinical informatics literature puts the unstructured share of electronic health record content as high as 80%, a figure researchers caution is hard to verify precisely, but that matches what anyone who has tried to mine a stack of clinician notes for a research signal already knows [3]. Layer on top of that a McKinsey Global Institute finding that knowledge workers spend close to a fifth of their workweek simply searching for information scattered across systems [4], and the scale of the drag becomes clear: time that, in a research setting, might otherwise go toward evaluating a target hypothesis, flagging a safety signal, or resolving a discrepancy between two studies of the same compound.

Semantics as the connective tissue

This is precisely the gap that semantic technologies: ontologies, knowledge graphs, and natural language understanding are built to close, and it is why they are increasingly described as critical infrastructure for AI-native therapeutic research rather than a back-office curation exercise. Rather than replacing the validated systems that pharmaceutical R&D already depends on, the more durable strategy is semantic layering: introducing a machine-interpretable layer of meaning over existing databases, laboratory systems, and operational processes, so that a foundation model or an autonomous research workflow can reason across all of them without anyone having to rebuild the underlying infrastructure.

Layered across the R&D pipeline, that connective tissue looks different at each stage but solves the same underlying problem. Here are a few examples:

  • Target identification. Knowledge graphs link genetic perturbations to phenotypic outcomes across genomic, transcriptomic, and proteomic datasets that were never collected with each other in mind.
  • Preclinical work. Semantic harmonization aligns study metadata: species, dose, route, model type across in vitro, in vivo, and in silico results so that inconsistencies surface instead of hiding.
  • Clinical data interpretation. Natural language understanding extracts structured signals from unstructured investigator notes and adverse-event narratives, the same kind of free-text material that a purely database-oriented AI system would simply never see.
  • Regulatory and safety work. Ontology-based traceability lets a benefit-risk assessment be reconstructed and defended, mapping a conclusion back through the preclinical and clinical evidence that produced it.
Semantic-harmonization-framework-drug-discovery

Figure: a semantic layer of ontologies, knowledge graphs, and NLU sits between fragmented, unstructured data sources and AI-native capabilities, harmonizing meaning without requiring any of the underlying systems to be replaced.

Where fragmentation is the point, not the problem

Nowhere is the case for accepting fragmentation, rather than fighting it, more evident than in rare and orphan disease research. These conditions are, almost by definition, ones where no single institution has access to patients to build a meaningful dataset on its own. The relevant evidence exists as scattered case reports, disparate patient registries, and natural history studies spread across research centers and countries. Centralizing that data into one warehouse isn’t just slow for many rare diseases, it isn’t realistic at all. A federated semantic layer, one that can connect these fragmented, unstructured sources and reason across them in place, isn’t a workaround here; it may be the only workable path to the kind of biomedical intelligence that can meaningfully accelerate treatments for the patients who have historically waited longest for them.

Back to the trial that’s already running

Which brings us back to those two clinical trials now streaming data to the FDA in real time. Continuous reporting is only as valuable as the interpretability of what’s being reported. A stream of imaging results, lab values, clinician notes, and device data arriving faster doesn’t help a regulator or an AI system meant to support one, unless it arrives harmonized against common definitions of what an endpoint, an adverse event, or a biomarker actually means. Regulators clearly sense this: earlier in 2026, the FDA and its European counterpart jointly published a shared set of guiding principles for AI use in drug development, built around transparent, risk-based evidence and clear traceability from data to decision [5]. Real-time infrastructure without a semantic backbone risks becoming exactly what critics of pure centralization warn about: the same fragmented pile of information, just arriving on a faster clock, with the hard work of reconstructing its meaning still left for later, only now under a regulator’s watch.

The organizations that get ahead of this won’t be the ones with the most data. They’ll be the ones that can make sense of the data they already have, wherever it lives, and connect it into a shared semantic fabric that predictive models, autonomous laboratories, and human researchers can all reason over together. That is the throughline connecting a regulatory pilot testing real-time trial reporting, a broader reckoning with how enterprises treat unstructured content, and a research community trying to build the next generation of AI-native R&D. Unstructured data isn’t the exception in biomedical research, but it’s most of the picture. The question is whether it becomes usable intelligence or stays dark matter.

For a deeper look at the framework behind semantic layering in drug discovery, including how ontology maturity, foundation models, and autonomous research systems fit together, download the full white paper, Semantic Biomedical Intelligence for AI

Frequently Asked Questions

What is semantic harmonization in drug discovery and development?

Semantic harmonization is the process of creating a shared layer of meaning across fragmented biomedical data sources. It connects different data formats, terminology, concepts, and relationships so researchers and AI systems can interpret and reason across data consistently.

Why is semantic harmonization important for AI-native R&D?

Semantic harmonization provides AI-native R&D systems with the context needed to understand fragmented biomedical data. By connecting information across research, preclinical, clinical, and regulatory sources, it helps AI systems and researchers work with data that would otherwise remain difficult to interpret across disconnected systems.

What is a semantic layer in biomedical research?

A semantic layer is a machine-interpretable layer of meaning placed over existing biomedical databases, laboratory systems, and operational processes. It can use ontologies, knowledge graphs, and natural language understanding to connect fragmented data without requiring organizations to replace their existing infrastructure.

How do knowledge graphs support drug discovery?

Knowledge graphs connect relationships between biomedical concepts such as genes, targets, compounds, phenotypes, diseases, and experimental results. This allows researchers and AI systems to discover and reason over relationships across genomic, transcriptomic, proteomic, preclinical, and clinical datasets.

What role do ontologies play in AI-native drug development?

Ontologies provide standardized concepts and relationships that allow different biomedical datasets and systems to share a common vocabulary. They can help establish consistent meanings for concepts such as diseases, biomarkers, endpoints, adverse events, study types, and treatments.

How can semantic technologies make unstructured biomedical data useful for AI?

Semantic technologies such as natural language understanding, ontologies, and knowledge graphs can extract meaningful information from clinical notes, pathology reports, investigator comments, lab notebooks, adverse-event narratives, and regulatory documents. This can make previously difficult-to-query information more accessible to AI systems and researchers.

How can semantic harmonization support clinical trial data interpretation?

Semantic harmonization can align clinical data around shared definitions for endpoints, biomarkers, adverse events, laboratory results, imaging findings, and other research concepts. This can help researchers and AI systems interpret information consistently across different clinical data sources.

Why is a federated semantic layer useful for rare and orphan disease research?

Rare and orphan disease research often relies on evidence distributed across institutions, patient registries, case reports, and natural history studies. A federated semantic layer can connect these fragmented sources and enable researchers and AI systems to reason across them without requiring all data to be centralized in one repository.

How does semantic layering support AI-native R&D?

Semantic layering connects fragmented data sources through shared meaning and context, allowing foundation models, predictive models, autonomous laboratories, and researchers to reason across existing R&D systems. It provides a way to work with existing infrastructure without requiring every underlying system to be replaced.

How can semantic technologies support regulatory and safety decision-making?

Semantic technologies can connect regulatory, preclinical, and clinical evidence through shared concepts and traceable relationships. This can help researchers and AI systems reconstruct how evidence supports a benefit-risk assessment and provide clearer traceability from data to decisions.

Sources

  • [1] U.S. Food and Drug Administration, “FDA Announces Major Steps to Implement Real-Time Clinical Trials,” press announcement, April 28, 2026. https://www.fda.gov/news-events/press-announcements/fda-announces-major-steps-implement-real-time-clinical-trials
  • [2] STAT News, “FDA testing speedier drug development with real-time clinical trials,” April 28, 2026. https://www.statnews.com/2026/04/28/fda-real-time-clinical-trials-pilot-project-astrazeneca-amgen-cancer-drugs/
  • [3] Assale, M., Dui, L.G., Cina, A., Seveso, A., & Cabitza, F. “The Revival of the Notes Field: Leveraging the Unstructured Content in Electronic Health Records.” Frontiers in Medicine, 6:66 (2019). https://www.frontiersin.org/articles/10.3389/fmed.2019.00066/full
  • [4] McKinsey Global Institute, “The Social Economy: Unlocking Value and Productivity Through Social Technologies,” 2012. https://www.mckinsey.com/industries/technology-media-and-telecommunications/our-insights/the-social-economy
  • [5] European Medicines Agency, “EMA and FDA Set Common Principles for AI in Medicine Development,” news announcement, January 14, 2026. https://www.ema.europa.eu/en/news/ema-fda-set-common-principles-ai-medicine-development-0
Murali Krishnam

Murali Krishnam

VP - Product Strategy, Enterprise Pharma AI

As the leader of Enterprise Pharma AI at Solix, Murali works across the life sciences ecosystem to drive technological innovations that enable scientific breakthroughs. He is dedicated to blending human knowledge with advanced technology to foster an outcomes-driven approach throughout the entire drug discovery and development process.

At Solix, Murali pioneered Semantic Content Library, an innovative solution that infuses meaning and contextual intelligence into multi-modal patient and clinical datasets, enabling the creation of AI-ready data products spanning drug discovery to post marketing.

DISCLAIMER: THE CONTENT, VIEWS, AND OPINIONS EXPRESSED IN THIS BLOG ARE SOLELY THOSE OF THE AUTHOR(S) AND DO NOT REFLECT THE OFFICIAL POLICY OR POSITION OF SOLIX TECHNOLOGIES, INC., ITS AFFILIATES, OR PARTNERS. THIS BLOG IS OPERATED INDEPENDENTLY AND IS NOT REVIEWED OR ENDORSED BY SOLIX TECHNOLOGIES, INC. IN AN OFFICIAL CAPACITY. ALL THIRD-PARTY TRADEMARKS, LOGOS, AND COPYRIGHTED MATERIALS REFERENCED HEREIN ARE THE PROPERTY OF THEIR RESPECTIVE OWNERS. ANY USE IS STRICTLY FOR IDENTIFICATION, COMMENTARY, OR EDUCATIONAL PURPOSES UNDER THE DOCTRINE OF FAIR USE (U.S. COPYRIGHT ACT § 107 AND INTERNATIONAL EQUIVALENTS). NO SPONSORSHIP, ENDORSEMENT, OR AFFILIATION WITH SOLIX TECHNOLOGIES, INC. IS IMPLIED. CONTENT IS PROVIDED "AS-IS" WITHOUT WARRANTIES OF ACCURACY, COMPLETENESS, OR FITNESS FOR ANY PURPOSE. SOLIX TECHNOLOGIES, INC. DISCLAIMS ALL LIABILITY FOR ACTIONS TAKEN BASED ON THIS MATERIAL. READERS ASSUME FULL RESPONSIBILITY FOR THEIR USE OF THIS INFORMATION. SOLIX RESPECTS INTELLECTUAL PROPERTY RIGHTS. TO SUBMIT A DMCA TAKEDOWN REQUEST, EMAIL INFO@SOLIX.COM WITH: (1) IDENTIFICATION OF THE WORK, (2) THE INFRINGING MATERIAL’S URL, (3) YOUR CONTACT DETAILS, AND (4) A STATEMENT OF GOOD FAITH. VALID CLAIMS WILL RECEIVE PROMPT ATTENTION. BY ACCESSING THIS BLOG, YOU AGREE TO THIS DISCLAIMER AND OUR TERMS OF USE. THIS AGREEMENT IS GOVERNED BY THE LAWS OF CALIFORNIA.