Technology · · ⏱ 7 min read

The entropy of corporate knowledge

A RAG does not eliminate the entropy of organisational knowledge — it displaces it. On embedding drift, corpus staleness and silent vector memory.

Every organisation believes it has solved its memory when it connects an LLM to its internal documentation. A well-set-up RAG, relevant answers in seconds, the apparent end of corporate amnesia. In practice, a different amnesia begins: quieter, harder to diagnose, and with consequences that only show up after months of accumulation. Vectorised knowledge is not preserved. It degrades along paths that engineering is only starting to name.

A corporate library whose books dissolve into vector clouds as they recede on the shelf
Memory was not solved by vectorisation — it changed shape.

The illusion of solved memory

For decades, organisational memory was thought of in two forms of knowledge: explicit (documentable) and tacit (embodied in people and practices). The literature on corporate amnesia described the problem accurately: turnover erodes collective memory, and the tacit part — the largest one — does not document well.

The rise of RAG promised to close the circle. A human no longer needed to read all the documentation: a system could index it, vectorise it, and answer complex questions in real time. Memory went from passive archive to active assistant.

And here comes the displacement. Memory has not been solved — it has been replaced by another one. A vector memory whose degradation laws are different from the ones we knew, and for which most organisations have no instrumentation. Entropy is still there. It changed shape.

How vectorised knowledge degrades

A RAG splits documents into chunks, converts them into vectors via an embedding model, and stores them in an index where queries search by semantic proximity. That mechanism has four known routes of silent degradation.

Embedding drift. Embedding models evolve. Each new version produces vectors in different spaces. When an organisation updates the model without reindexing the corpus, the old vectors are still there but new queries no longer reach them well. The most dangerous form is the gradual one: the system keeps answering, just worse and worse, without producing a visible error. Without specific metrics, nobody notices until a user reports that “the AI has gotten dumber”.

Corpus staleness. The indexed corpus ages. A 2019 manual persists when a 2024 one already exists. Older documents rank higher by pure semantic similarity, even when newer versions exist. The solution is not just to update more often: it is that retrieval logic needs to understand time, not just semantics. Without freshness as a ranking criterion, the index rewards what is old by default.

Semantic drift in reasoning chains. In multi-hop architectures, each sub-query builds on the results of the previous one. Small errors or drifts accumulate, and the final answer ends up responding to a question that is similar but distinct from the original. It is accumulation, not failure — the system ends up far from the starting point without any single step having made a serious error.

Index drift. The index stops reflecting the authoritative source. Deleted documents that remain indexed. Chunks pointing to old versions. Duplicated embeddings. Most production RAG systems do not reconcile their index with the source on a regular basis. When ingestion fails, the evidence stays hidden until somebody looks for it.

What this changes in data engineering

Four concrete decisions emerge from the picture.

Treat the vector index as a dataset with a lifecycle, not as a static archive. The traditional knowledge base mindset — load it and query it — does not work here. A vector index is a living system that requires scheduled reindexing, embedding-model versioning, and traceability of which chunk came from which document at which point.

Add time as an explicit dimension of retrieval. No serious system should retrieve by semantic proximity without considering freshness. Timestamp metadata per chunk, decay policies that penalise old documents when newer versions exist, hybrid retrieval strategies that combine semantics with recency.

Instrument retrieval quality continuously. A RAG without metrics is opaque by default. It needs weekly reviews that compare current retrieval against a set of reference queries — the golden set — and automatic alerts when relevance drops below a threshold. Degradation does not announce its arrival. It is only detected if it is measured.

Reconcile index and source on a schedule. At least once per ingestion cycle, a job should verify that every indexed document still exists in the authoritative source, with the correct version and no duplicates. And that job should produce an actionable report, not a green “OK” that says nothing.

What this approach does not solve

These measures attack the degradation of vectorisable knowledge. Two layers stay out.

The first is tacit knowledge: everything an experienced person knows how to do without being able to explain it, everything that lives in informal conversations. None of that enters an embeddings pipeline. A well-maintained RAG over explicit documentation does not cover that loss — it can even hide it, giving the sense that memory is solved when only the visible part has been.

The second is interpretive context. A chunk retrieved by proximity may be semantically relevant and contextually misleading. If a user asks about Q3 2025 revenue, a vector search may return the 2024 data because the semantic distance between “2024” and “2025” is negligible for the model. One digit of difference between a correct insight and a hallucination. Embeddings are good for concepts, not for specifications.

The other side: what kind of memory we are delegating

So far, engineering. But vectorisation touches something more: what kind of knowledge we are privileging when we decide what deserves to be indexed.

On the bias towards the documentable

The classic literature distinguished between explicit and tacit knowledge precisely because it recognised that not everything an organisation knows can be written. Culture, informal practices, criteria that an experienced professional applies without being able to articulate them — all that constitutes most of real knowledge.

When memory is vectorised, that bias deepens. The system can only retrieve what somebody decided to document, in the format in which it was documented. Tacit knowledge is automatically left out. And worse: the system does not flag its absence. It answers with confidence about what it knows, without signalling that a layer exists that never reached the index and would probably contradict its answer.

The illusion of completeness is the problem. It is not that the system lies — it is that its silence about what it does not know reads as if there was nothing else to know.

On memory as a living ecosystem

A healthy organisational memory is not an archive. It is a living ecosystem where knowledge is generated in conversations, tested in practice, refined by collective correction, and preserved by transmission between people. That ecosystem has its own metabolism: it discards what is obsolete, amplifies what is tested, integrates the new.

A well-built RAG is a useful cognitive infrastructure, but only one layer inside that ecosystem. Replacing the whole ecosystem with it is reducing living memory to its vectorisable fossil. And fossils, however precise, do not evolve on their own.

On the responsibility of the builder

Whoever builds a corporate RAG takes, often without naming it, a structural decision: they decide which part of the organisational knowledge will be accessible and which part will stay outside the main system’s reach. Multiplied by thousands of queries a day, that decision redistributes interpretive power in the organisation. What is indexed gets queried. What is not indexed gets forgotten.

That responsibility has three dimensions in a triad: conscious coverage (knowing what is indexed and what is left out, and why), measured degradation (instrumenting quality loss as an operational metric, not intuition), and honest complementarity (recognising that RAG complements but does not replace living memory).

Open questions

  • If a RAG answers confidently about what it knows but stays silent about what is not indexed, at what point does its apparent completeness turn into structural disinformation?
  • Can an organisation claim to have “solved its memory” when most of what it knows — tacit knowledge — still lives only in the people who are leaving?
  • When the index rewards older documents by default because of their higher historical density, which version of the organisation are we consulting: the current one, or the one that was?

The questions have no closed answer. But one final idea is worth keeping: the entropy of corporate knowledge was not eliminated by RAG — it was displaced. It used to live in staff turnover, in documents nobody could find, in the natural forgetting of practices. Now it lives in embedding drift, in indexes without reconciliation, in the illusion that what is vectorisable is all that is known. The shape changed. The entropy remains.

And that entropy, today, is held by whoever builds the pipeline. With instrumentation, with awareness of what is left out, with humility about what the machine can and cannot remember.

References

  • Openlayer — What are embedding models? A complete guide (March 2026). Drift, retrieval accuracy regression. openlayer.com
  • Leela Desai (Medium) — Knowledge Drift: The Silent AI Killer in RAG models (February 2025). medium.com
  • AI with Aish — All you need to know about RAG (in 2026) (March 2026). Q3 2025 vs 2024 example. substack.com
  • Oracle Developers — How to Detect RAG Index Drift (July 2026). Index/source reconciliation. blogs.oracle.com
  • arXiv — Retrieval-Augmented Generation: A Survey (2024). Semantic drift in multi-hop retrieval. arxiv.org/pdf/2407.13193
  • Wikipedia — Corporate amnesia. Tacit vs explicit knowledge. wikipedia.org
  • Previous articles in this series: Infrastructures that forget, Forgetting in machines.