The Conversation Most AI Programmes Are Not Having
The organisations that move from AI pilot to AI programme share one pattern: they resolved the architecture question before scale revealed it. They defined what the data foundation needed to look like – in terms of storage properties, semantic layers, and pipeline controls – before the second or third production application was deployed.
Most do not follow that sequence. The model is chosen. Use cases are defined. Proofs of concept are built. And the data foundation is treated as something that will be addressed as the programme grows. In enterprise environments, that assumption is consistently the source of the delay – and the cost – when scale is eventually attempted.
Accessible Is Not the Same as AI-Ready
The distinction that tends to be missed in early AI programme planning is between data that is accessible and data that is AI-ready.
Accessible data can be queried. It exists in a system, it can be extracted, and a model can be pointed at it. Most enterprise data meets that threshold. AI-ready data is governed from the point of ingestion, carries verified semantic meaning, has been quality-validated before it reaches a model, and can be traced from source to answer without gaps in the lineage.
The gap between those two definitions is an architectural one – not a data quality project. You cannot govern data retroactively at the scale enterprise AI demands. The three layers that determine whether enterprise AI scales – storage, metadata, and pipelines – each have a distinct job, and each creates the conditions the next layer depends on.
Why the Foundation Is Always the Last Thing Built
Most enterprise data environments were not designed with AI in mind. Data lands in storage systems built for transaction processing. Metadata – if it exists at all – is documented in spreadsheets, buried in data dictionaries, or carried in the heads of the engineers who built the original systems. Pipelines are one-off constructions, never designed to feed a production AI application at scale.
The result: every AI initiative effectively starts from scratch on the data layer. Quality checks are applied inconsistently. Governance is bolted on after the fact. Generic ML platforms provide tooling for model development but not the governed ingestion, semantic layer, or production-grade quality controls that make those tools reliable on real enterprise data. Customers build those layers themselves – which is why AI initiatives consistently take longer and cost more than the initial business case assumes.
Three Layers, Three Distinct Jobs
The storage layer is where governance either starts or is permanently deferred. ACID transactions, schema evolution, and time-travel queries are not performance features – they are governance requirements. Schema evolution matters because enterprise source systems change continuously; a storage layer that cannot absorb those changes becomes a bottleneck for every AI initiative above it. Time-travel queries allow regulated businesses to answer historical audit questions without manual data reconstruction. Retention enforced at ingestion removes an entire class of compliance risk that most organisations currently manage through process and documentation rather than architecture. Solix delivers this through an Apache Hudi-based governed lakehouse and a dedicated Preservation Zone for enterprise application data, with retention and legal-hold capability active from the moment data arrives.
The metadata layer is where data acquires meaning – and where most enterprise AI foundations are weakest. A database schema tells you what fields exist, not what they represent or how they relate across applications. An AI system reasoning over raw schema is inferring business meaning it does not have, producing outputs that cannot be explained or traced to a source. An Application Knowledge Graph closes that gap: a semantic layer encoding the business objects, relationships, vocabulary, and tested query patterns specific to an enterprise application. Solix ships pre-built Application Knowledge Graphs for Oracle EBS, SAP ECC/S4HANA, PeopleSoft, and JD Edwards – removing the most time-intensive layer of the semantic build from the customer’s plate. For unstructured content, a semantic index built at ingestion applies the same principle: meaning is encoded before a model reasons over it. The result is AI answers that can be traced to a certified business definition and presented to a regulator or a board with confidence.
The pipeline layer is where the reliability of every downstream AI output is determined. Automated profiling, quality rules, lineage tracking, and data preparation workflows ensure problems are caught at ingestion – not after a flawed AI recommendation has already influenced a business decision. Solix structures this as four sequential governed stages – ingestion, governed lakehouse, data quality, and ML flow – each auditable, each reducing the time between data arriving and being production-ready for AI.
What a Production-Ready Foundation Looks Like
Consider a financial services organisation consolidating Oracle EBS and SAP data alongside real-time operational feeds and archived records from retired applications. The goal: an AI-powered risk and operational intelligence platform with natural-language access for business users and a governed foundation for data science teams.
At the storage layer, data lands in a governed lakehouse with ACID transactions, schema evolution, and time-travel queries for regulatory lookback. Legacy application data moves to a governed Preservation Zone – retention-managed and legal-hold ready from arrival.
At the metadata layer, pre-built Application Knowledge Graphs for Oracle EBS and SAP are configured to the organisation’s specific setup, encoding the Procure-to-Pay, Order-to-Cash, and Record-to-Report value streams. A Unified Asset Catalog maintains certified business definitions and lineage across the estate. When a business analyst queries in natural language or a risk model interrogates the data, the answer traces to a certified definition – not an inference from a raw table.
At the pipeline layer, quality rules fire before data reaches any model. Lineage tracks every element from source to AI-ready product. The data science team builds on governed data – not months of bespoke data engineering undertaken before the first model can be trained.
Three Questions Before the Next AI Initiative Scales
Before committing to the next phase of AI investment, three questions about the data architecture are worth putting to the team directly.
Does data land in a governed environment from the first byte – with quality validation, retention policies, and audit trails enforced at ingestion – or is governance applied after the fact when a problem surfaces? Organisations that cannot answer the first carry compliance risk that scales with every AI application deployed on top of it.
Do the AI systems operating on that data have a semantic layer that encodes business meaning, or are they reasoning directly over raw schemas? Without one, AI outputs cannot be reliably traced, explained, or defended.
Are the pipelines feeding AI models producing governed, lineage-tracked data products, or raw inputs that each new initiative must independently re-engineer? The latter means paying the data engineering cost repeatedly – once per initiative rather than once for the foundation.
If any of those questions is harder to answer than it should be, that is where the architecture review starts – and where the blueprint for scalable enterprise AI begins.
DISCLAIMER: THE CONTENT, VIEWS, AND OPINIONS EXPRESSED IN THIS BLOG ARE SOLELY THOSE OF THE AUTHOR(S) AND DO NOT REFLECT THE OFFICIAL POLICY OR POSITION OF SOLIX TECHNOLOGIES, INC., ITS AFFILIATES, OR PARTNERS. THIS BLOG IS OPERATED INDEPENDENTLY AND IS NOT REVIEWED OR ENDORSED BY SOLIX TECHNOLOGIES, INC. IN AN OFFICIAL CAPACITY. ALL THIRD-PARTY TRADEMARKS, LOGOS, AND COPYRIGHTED MATERIALS REFERENCED HEREIN ARE THE PROPERTY OF THEIR RESPECTIVE OWNERS. ANY USE IS STRICTLY FOR IDENTIFICATION, COMMENTARY, OR EDUCATIONAL PURPOSES UNDER THE DOCTRINE OF FAIR USE (U.S. COPYRIGHT ACT § 107 AND INTERNATIONAL EQUIVALENTS). NO SPONSORSHIP, ENDORSEMENT, OR AFFILIATION WITH SOLIX TECHNOLOGIES, INC. IS IMPLIED. CONTENT IS PROVIDED "AS-IS" WITHOUT WARRANTIES OF ACCURACY, COMPLETENESS, OR FITNESS FOR ANY PURPOSE. SOLIX TECHNOLOGIES, INC. DISCLAIMS ALL LIABILITY FOR ACTIONS TAKEN BASED ON THIS MATERIAL. READERS ASSUME FULL RESPONSIBILITY FOR THEIR USE OF THIS INFORMATION. SOLIX RESPECTS INTELLECTUAL PROPERTY RIGHTS. TO SUBMIT A DMCA TAKEDOWN REQUEST, EMAIL INFO@SOLIX.COM WITH: (1) IDENTIFICATION OF THE WORK, (2) THE INFRINGING MATERIAL’S URL, (3) YOUR CONTACT DETAILS, AND (4) A STATEMENT OF GOOD FAITH. VALID CLAIMS WILL RECEIVE PROMPT ATTENTION. BY ACCESSING THIS BLOG, YOU AGREE TO THIS DISCLAIMER AND OUR TERMS OF USE. THIS AGREEMENT IS GOVERNED BY THE LAWS OF CALIFORNIA.
-
White PaperEnterprise Information Architecture for Gen AI and Machine Learning
Download White Paper -
-
-