Everyone in drug discovery agrees that ontologies and knowledge graphs help AI reason across fragmented biomedical data. Almost no one has tested it. One systematic literature mapping found virtually no experimental studies validating the claims made for foundational ontology reuse in biomedical research — which means the industry is investing in semantic infrastructure on the strength of an assumption. This working paper from the SPARK AI Consortium at the San Diego Supercomputer Center sets out to close that gap: a 3×3 analytical framework pairing disease categories against ontology maturity levels, and a phased Semantics Data Co-Design Laboratory built to run ontology-informed and baseline AI pipelines side by side and measure the difference.
What is semantic biomedical intelligence?
Semantic biomedical intelligence is the use of ontologies, knowledge graphs, and reasoning systems to give AI machine-interpretable context across biological, clinical, and regulatory data. Rather than detecting statistical associations, models reason over defined relationships between entities — with provenance and traceability preserved.
Highlights
- ~10% of drugs reach approval, against a capitalized cost estimated near $2.8B per new therapeutic — the economics semantic infrastructure is meant to improve
- A 3×3 research framework pairing three disease categories against three ontology maturity levels, with cancer as the reference condition
- 10 joint FDA–EMA principles, published January 14, 2026 — the first transatlantic regulatory alignment on AI in drug development
- 4 precedent laboratories analyzed — NCATS Biomedical Data Translator / RTX-KG2, KaBOB, TMO/TMKB, and the Biolink Model
What’s inside the paper
- The five knowledge domains × four semantics matrix — where formal, linguistic, computational, and data semantics intersect target identification, compound design, preclinical and clinical interpretation, and regulatory intelligence, with the near-term high-leverage intersections identified
- Regulation moving faster than expected — the joint FDA–EMA guiding principles, the FDA’s seven-step credibility assessment framework and its Context of Use requirement, the real-time clinical trial initiative, the National Priority Voucher, and the 2025 start of the animal-testing phase-out
- Why ontologies have underdelivered — hundreds of biomedical ontologies exist, and prior work has concentrated on cataloguing, alignment, and annotation rather than testing whether they measurably improve AI-driven hypothesis generation, repurposing, or biomarker discovery
- A maturity rubric you can apply — structural richness, community adoption, integration readiness, interoperability, and governance stability, scored to a composite classification
- Three concrete experimental use cases — drug repurposing in Alzheimer’s disease, target discovery in type 2 diabetes, and variant prioritization in Duchenne muscular dystrophy, each run with and without ontology enrichment
- The Semantics Data Co-Design Laboratory — a Phase I “beachhead” design with five goals, and the argument for why a small instrumented environment beats an enterprise rollout for this class of problem
- Four precedents worth learning from — what RTX-KG2 and the NCATS Translator ecosystem, KaBOB, TMO/TMKB, and Biolink each demonstrate about semantic engineering environments maturing into operational platforms
Why it matters
Semantic layering is now standard advice: add a machine-interpretable layer above existing validated systems, preserve legacy infrastructure, gain interoperability and regulatory traceability. The advice is probably right. But “probably right” is a weak basis for infrastructure spend, and the paper is unusually direct about it — semantic reasoning in drug discovery is neither new nor novel, and the evidence on when and how ontology integration contributes to AI effectiveness remains sparse. What this framework contributes is not another semantic model but a way to compare them: parallel pipelines, defined maturity scores, and measurement across predictive performance, interpretability, and operational complexity. For anyone deciding how much to invest in ontologies, knowledge graphs, and harmonization, the useful question is not whether semantics helps but where, by how much, and at what cost in complexity. That is what this is designed to answer.
Download the Whitepaper
A framework, a maturity rubric, and a laboratory design — for testing what the industry has so far assumed.
About the Author:
-
Dr. James Short
Director of the SPARK AI Consortium, San Diego Supercomputer Center
UC San Diego
-
Suresh Mani
Chief AI Architect
Solix Technologies, Inc.
-
Murali Krishnam
VP – Product Strategy, Enterprise Pharma AI
Solix Technologies, Inc.
-
Raju Puspati
VP Life Sciences
Solix Technologies, Inc.