Everyone’s telling you to build a knowledge graph for AI. Do you actually need one?
Before you launch another data movement project, ask a harder question: what does AI really need from your data?
For decades, our answer to every new workload was the same — move the data. Into warehouses, lakes, lakehouses. Every copy created more pipelines, more governance overhead, more sync failures, more cost. Moving data means managing data, twice.
Now AI context is the new workload, and the old reflex is kicking in: extract everything into a knowledge graph, embed everything into a vector store.
My take: the best context strategy is the one that moves the least data.
Three Rules
Knowledge graphs are powerful — but only where relationships are the value. Entity resolution, lineage, complex domains. Don’t build one because it’s trendy. A well-described semantic layer over data in place often answers the same questions.
Virtualize before you replicate. Let AI reach data where it lives through governed access, metadata, ontologies and semantic descriptions — not another copy.
Materialize only what you must: hot paths where latency or cost demands it.
Unstructured Data Is Finally a First-Class Citizen
And here’s what’s changing the game: with AI, unstructured data — documents, emails, calls, images — is finally a first-class citizen. Vector stores and embeddings make 80% of enterprise data usable for the first time.
That’s why yesterday’s data strategy doesn’t fit. It was built for structured data and centralized copies.
The future of context is federated: structured data accessed in place, unstructured data indexed where it sits, and a semantic layer tying both together — so the AI brings the question to the data, not the data to the question.
Less movement. Less management. More context.