Stephen Tallant

Most Enterprise Data Is Not Complex. It Is Ungoverned.

Organisations describe their data environment as complex when they mean something more precise: data has accumulated across systems that were never designed to share it, retired applications hold decades of business history in formats the current team has never seen, metadata exists only in spreadsheets or in the memory of engineers who may no longer be available, and every AI initiative begins the same way — months of data engineering before the first model can be trained.

That is not complexity. It is the absence of governance. And the distinction matters because complexity implies the problem is architectural — requiring new platforms, new investment, new infrastructure. Ungoverned data is a structural problem — one that a defined sequence of decisions can resolve, starting from the moment data arrives.

Three Stages, One Sequence

The path from data swamp to AI-ready asset moves through three stages: ad hoc, managed, and optimised. Most enterprises believe they are further along than they are. The diagnostic is straightforward. If an AI initiative requires months of data engineering before the first model can be trained, the data is ad hoc. If AI outputs cannot be traced to a source, the semantic layer is absent. If the answer to “who can query this data?” is “the data team, after a ticket has been raised,” the optimised stage has not been reached.

The three stages are not a spectrum. They are a sequence. Managed is a prerequisite for optimised. The organisations that reach the optimised state fastest are the ones that treated each stage as a design decision — not an outcome that accumulates by accident.

Stage One: Ad Hoc — Where Most Enterprise Data Lives

At the ad hoc stage, data exists but cannot be reliably used. It lands in storage systems built for transaction processing. Retired applications hold financial records, procurement data, and customer transactions — years or decades of business history — but the application that provided context for that data has been decommissioned. What remains is often cryptic: table names that were meaningful to the original developer, column conventions inherited from software written in a different decade, and relationships encoded in application logic rather than in the database schema.

The cost is paid at the start of every AI initiative. Data engineering teams build bespoke ingestion pipelines for each new project. Quality checks are inconsistent or absent. Governance — retention policies, access controls, classification — is documented somewhere but not enforced in the systems where data actually lives. Each project reinvents the same foundation. The business case for AI investment is undermined before the first model is trained, because the data that model needs is not ready.

Stage Two: Managed — Governance Before Intelligence

The managed stage begins at ingestion. Data arrives in a governed environment where retention policies are enforced from the first byte, quality checks run before data reaches any downstream system, classification tags records against the organisation’s own taxonomy, and legal-hold capability is active from the moment data lands. This is not a storage upgrade. It is a structural change in how data is received — a Preservation Zone architecture that provides a governed, centralised home for data from retired applications, acquired systems, and active production environments under a single policy model.

The enforcement layer sits in the governance platform: retention applied in real time, masking enforced at access, every governance action recorded in an immutable audit trail. The managed stage does not produce AI outputs. It produces data that is trustworthy enough to build on — classified, validated, governed, and legally defensible. Organisations that reach this stage have solved the compliance problem. What they have not yet solved is the meaning problem.

Stage Three: Optimised — When Data Becomes Fuel

Managed data becomes AI-ready when a semantic layer is added. This is the transition most generic platforms omit — and the one that determines whether AI operates on verified business meaning or inferred schema. A database schema identifies what fields exist. It does not encode what they represent, how they relate across the value streams the business runs on, or what constitutes a valid query pattern in that application’s context. An AI system reasoning over raw schema is guessing. The answers it produces cannot be explained, traced, or trusted in a business decision.

An Application Knowledge Graph closes that gap: encoding the business objects, relationships, vocabulary, and tested query patterns specific to an enterprise application. Solix ships pre-built Application Knowledge Graphs for Oracle EBS, SAP ECC/S4HANA, PeopleSoft, and JD Edwards — the ERP systems that hold the majority of enterprise structured data — and builds custom Knowledge Graphs through AI-powered discovery for legacy and acquired applications with no pre-built equivalent. For unstructured content, Content Intelligence builds the semantic layer over documents, contracts, and records that makes them queryable alongside structured data in a single governed interface.

With the semantic layer in place, business users ask questions in natural language and receive governed, auditable answers grounded in certified business definitions. Data science teams build AI applications on a foundation that has already been governed, classified, and semantically enriched — rather than spending months preparing data that should have been ready from arrival. Preserved data stops being an infrastructure cost and becomes, in the language of the Solix Enterprise Data Preservation whitepaper, “fuel for intelligent analysis, discovery, and decision-making.”

A Diagnostic, Not an Aspiration

The maturity model is useful only if applied honestly. Three questions identify where an organisation actually sits.

Is data arriving in a governed environment from the first byte — classified, quality-validated, and retention-enforced at ingestion — or is governance applied after the fact when a compliance question surfaces? Organisations that cannot answer the first are carrying risk that scales with every AI application deployed on top of it.

Is there a semantic layer encoding business meaning for the applications the organisation runs, or are AI systems reasoning over raw schemas and inferring context they do not have? Without one, AI outputs cannot be reliably traced, explained, or presented with confidence.

Can business users query the data directly in natural language — with answers grounded in certified definitions and traceable to a source — or does every question still require a data team and a ticket queue?

The distance between no and yes on each of those questions is the distance between ad hoc and AI-ready. It is a traversable distance — but only in sequence, and only by treating each stage as a deliberate decision rather than an outcome of history.

Stephen Tallant

Stephen Tallant

Vice President of Product Marketing

As the Vice President of Product Marketing at Solix Technologies, I lead the development and communication of the product and solution story to the market. I have over 25 years of experience in product marketing and product management, creating engaging messaging, launch plans, collateral, and content for various software solutions. I live in metro Philadelphia, and am a big sports fan - so much so, I sit on the Board of the Philadelphia Sports Hall of Fame. I attended Villanova University for both my undergraduate and graduate degrees.

DISCLAIMER: THE CONTENT, VIEWS, AND OPINIONS EXPRESSED IN THIS BLOG ARE SOLELY THOSE OF THE AUTHOR(S) AND DO NOT REFLECT THE OFFICIAL POLICY OR POSITION OF SOLIX TECHNOLOGIES, INC., ITS AFFILIATES, OR PARTNERS. THIS BLOG IS OPERATED INDEPENDENTLY AND IS NOT REVIEWED OR ENDORSED BY SOLIX TECHNOLOGIES, INC. IN AN OFFICIAL CAPACITY. ALL THIRD-PARTY TRADEMARKS, LOGOS, AND COPYRIGHTED MATERIALS REFERENCED HEREIN ARE THE PROPERTY OF THEIR RESPECTIVE OWNERS. ANY USE IS STRICTLY FOR IDENTIFICATION, COMMENTARY, OR EDUCATIONAL PURPOSES UNDER THE DOCTRINE OF FAIR USE (U.S. COPYRIGHT ACT § 107 AND INTERNATIONAL EQUIVALENTS). NO SPONSORSHIP, ENDORSEMENT, OR AFFILIATION WITH SOLIX TECHNOLOGIES, INC. IS IMPLIED. CONTENT IS PROVIDED "AS-IS" WITHOUT WARRANTIES OF ACCURACY, COMPLETENESS, OR FITNESS FOR ANY PURPOSE. SOLIX TECHNOLOGIES, INC. DISCLAIMS ALL LIABILITY FOR ACTIONS TAKEN BASED ON THIS MATERIAL. READERS ASSUME FULL RESPONSIBILITY FOR THEIR USE OF THIS INFORMATION. SOLIX RESPECTS INTELLECTUAL PROPERTY RIGHTS. TO SUBMIT A DMCA TAKEDOWN REQUEST, EMAIL INFO@SOLIX.COM WITH: (1) IDENTIFICATION OF THE WORK, (2) THE INFRINGING MATERIAL’S URL, (3) YOUR CONTACT DETAILS, AND (4) A STATEMENT OF GOOD FAITH. VALID CLAIMS WILL RECEIVE PROMPT ATTENTION. BY ACCESSING THIS BLOG, YOU AGREE TO THIS DISCLAIMER AND OUR TERMS OF USE. THIS AGREEMENT IS GOVERNED BY THE LAWS OF CALIFORNIA.