The Second Data Lake
Everyone is racing to build a context layer for their AI agents. Most enterprises are about to repeat 2015, one level up the stack.
In 2015 the enterprise decided the answer was a data lake. Pour everything into one place, sort out the meaning later. It was a reasonable bet at the time. It failed for a reason very few people said out loud: meaning was maintained by hand, and hands do not scale. The lake filled faster than anyone could label it, and inside of three years most of them were swamps.
Eleven years later, the enterprise has decided the answer is a context layer.
Same bet. Same failure mode. Much faster clock.
Ninety days of evidence
Something shifted this summer, and it did not get the coverage it deserved. Five data points, in order.
February 9. Gartner published its second Market Guide for Agentic Analytics with a strategic planning assumption that should have stopped every data team mid sprint: by 2028, 60 percent of agentic analytics projects relying solely on the Model Context Protocol will fail due to the lack of a consistent semantic layer. [1] MCP is the plumbing the entire market is standardizing on right now. Gartner is not saying the plumbing is wrong. It is saying that plumbing carries nothing if there is no shared meaning at the other end of it.
March. At the Gartner Data and Analytics Summit, the opening keynote reported that 14 percent of organizations are confident their data and content assets are properly secured and governed, while 44 percent implemented a semantic layer during 2025 and another 48 percent plan to by 2027. [2][3] Read those two numbers next to each other. Roughly half the market is building a semantic layer on top of a governance posture that one organization in seven actually trusts.
July 7. Forbes revisited the Gartner warning that more than 40 percent of agentic AI projects will be canceled by the end of 2027, and named the mechanism plainly. Not model capability. Governance, undefined business value, and what the analysts call agent washing. [4]
July 29. This is the one that changes the argument. A cross industry analysis of more than 10,000 observed enterprise AI failure events found that hallucination accounted for under 10 percent of them. The largest single failure family, at 31.1 percent, was resolution and escalation breakdown. Execution and action failures were up 62 percent against the 2024 baseline. [5]
August 2. The EU AI Act transparency obligations under Article 50, the enforcement powers over general purpose AI, and the full penalty regime all took effect on schedule. The heavier high risk obligations were deferred to December 2027, and most of the coverage treated the deferral as the story. It was not. The disclosure and accountability duties that reach ordinary enterprises are live right now. [6][7]
Put those five together and a picture forms that is different from the one the industry has been arguing about for two years.
The failure moved, and almost nobody moved with it
Earlier this year I wrote about the confident wrong answer, the output that looks right and is not. I still believe that is the most dangerous single thing an enterprise AI system can produce. But the failure data says I was describing yesterday’s version of the problem.
When AI answered questions, the wrong answer was the entire risk, and a human sat between the answer and the consequence. When AI performs work, the wrong answer stops being the endpoint. It becomes the input to an action. The system does not pause at the bad number. It updates the record, files the ticket, adjusts the price, routes the escalation, and moves on. That is why execution failures climbed 62 percent while hallucination fell to a minority of incidents. The mistake did not become less common. It became less visible, because it stopped appearing as text on a screen and started appearing as a completed transaction.
A hallucination on a dashboard is a reporting problem. The same error inside an agent is a decision problem, and by the time anyone finds it, three downstream systems have already accepted it as true.
That reframe is exactly why the market lunged at the context layer this year. If the agent has to be right before it acts, the agent has to understand the business. The instinct is correct.
The execution is where 2015 comes back.
Why the second lake fails the same way as the first
Walk into most enterprises building a context layer today and here is what you find. A semantic model authored by hand. Metric definitions maintained in YAML by a team of four. An ontology living in a wiki. Business glossary entries written in 2023 by someone who has since left the company. And underneath all of it, tribal knowledge sitting in the heads of the two analysts everyone escalates to.
None of that is wrong. All of it is manual. And manual artifacts have a decay rate.
The data lake did not fail because the storage was bad. It failed because the meaning laid over the top of it was hand maintained while the schema underneath kept moving. Every new source, every ERP upgrade, every column rename opened a gap between what the catalog claimed and what the data actually was. The gap stayed invisible until somebody made a decision on the wrong side of it.
A hand built context layer has the identical structure and a far worse clock speed. A stale dashboard definition gets read by an analyst a few times a week, and the analyst usually notices something is off. A stale context definition gets consumed by an agent a thousand times a day, at machine speed, with nobody reading it at all.
One of the sharpest lines to come out of the Gartner summit came from a data leader writing up his own takeaways, who called the ungoverned semantic layer “the new ungoverned data lake.” [3] We are preparing to make the same mistake one layer up the stack, and we are preparing to make it faster.
Three questions that separate a context layer from a second lake
If someone pitches you a context layer, a semantic layer, an ontology, or an agent ready data initiative in the next two quarters, these three questions will tell you most of what you need to know.
1. Who authors it, and what happens the week they leave?
If the answer involves a named team hand writing definitions, you are buying a maintenance liability, and you are buying it at the exact moment your schema is changing fastest. The layer has to be derived from the environment, not transcribed from it by people.
2. How long between a schema change and the layer reflecting it?
Ask for the number in hours or days. If the honest answer is next quarter’s refresh, then the gap between what the layer says and what your data is will be permanent, because your enterprise never stops changing.
3. When a question is genuinely ambiguous, what does the system do?
This is the one almost nobody asks, and it is the one that matters most now that agents act on their own output. There are exactly two possible behaviors. It guesses, or it asks. A system that guesses on an ambiguous question is a system that will eventually execute on a guess. And under the transparency and accountability duties that went live on August 2, “the model inferred it” is not a sentence you want to say to a regulator, a customer, or your own audit committee.
What we built, and the design choice underneath it
DataSense is our answer to the first two questions, and it was built on a single conviction: the layer should be discovered, not authored.
It profiles the environment as it actually exists, maps the relationships including the implicit ones no documentation ever captured, folds in human business vocabulary where human judgment genuinely adds something, and publishes a working knowledge graph in hours rather than the multi month modeling project most of this market quietly requires before anything works at all. Then it keeps going. Every question your business users ask feeds back into it. It gets more accurate with use instead of less accurate with age, which is the exact opposite of how every hand maintained artifact in your estate behaves. And it does this against real production schemas, the thousand table SAP, EBS, and Workday environments where general purpose approaches come apart, on your data as is, with no schema redesign charged as the entry fee.
DataAsk is our answer to the third question, and I will be direct about it: this is the design decision I am proudest of. When a question is genuinely ambiguous, DataAsk asks a clarifying question. It would rather take one more turn than hand an executive a confident wrong answer. In a world where the output feeds an action instead of a human reader, that one behavior is the difference between AI you can put inside a workflow and AI you have to audit after the fact.
I am not going to explain how it does it. The mechanism is ours. What I will give you is the principle, because the principle is what you should demand from every vendor you evaluate, including us. The system should refuse to guess.
The line I would put in the RFP
Enterprise buyers have gotten very good at asking for accuracy numbers. Accuracy numbers are easy to produce on a curated test set and close to meaningless against a real estate.
Here is the sentence I would put in the next RFP instead:
Demonstrate, on our production schema, what the system does when a question is ambiguous, and show us how the layer stays current when we change the schema underneath it.
Two behaviors, not two benchmarks. Everything else is negotiable.
The first data lake cost the enterprise something like five years and a lot of credibility, and we recovered because the failure was slow and it was visible. The second one will be neither, because the thing consuming it does not pause to wonder whether the definition it just read is still true.
The organizations that get this right over the next four quarters will not be the ones with the best model. That argument is over. They will be the ones whose context layer was never a document in the first place.
References
- 1. Gartner, Market Guide for Agentic Analytics, 9 February 2026, as cited in vendor coverage of the report. [https://www.dremio.com/blog/dremio-named-a-representative-vendor-in-the-2026-gartner-market-guide/]
- 2. Tamr, 5 Key Takeaways from the 2026 Gartner Data & Analytics Summit, March 2026. [https://www.tamr.com/blog/5-key-takeaways-from-the-2026-gartner-data-analytics-summit]
- 3. Juan Sequeda, Gartner Data & Analytics March 2026: My Honest No-BS Takeaways. [https://juansequeda.substack.com/p/gartner-data-and-analytics-march]
- 4. Forbes, Why 40% Of Agentic AI Projects May Be Canceled By 2027, 7 July 2026. [https://www.forbes.com/sites/robertszczerba/2026/07/07/why-40-of-agentic-ai-projects-may-be-canceled-by-2027/]
- 5. ChatSee.ai, State of Enterprise AI Failures 2026, announced 29 July 2026. [https://www.prnewswire.com/news-releases/new-research-finds-enterprise-ai-failures-are-shifting-beyond-hallucinations-as-companies-move-from-chatbots-to-agents-302837907.html]
- 6. Software Improvement Group, A comprehensive EU AI Act Summary, August 2026 update. [https://www.softwareimprovementgroup.com/blog/eu-ai-act-summary/]
- 7. EU AI Act: What Actually Applies on August 2, 2026. [https://accuroai.co/blog/eu-ai-act-what-actually-applies-august-2-2026]
