Retrieval can find a document. It cannot decide whether the document deserves authority.
That distinction disappears easily in a RAG project. Teams discuss chunk sizes, embeddings, hybrid search and reranking. They connect SharePoint sites, policies, manuals and project folders. The first answers look convincing because they contain internal language and citations.
Then the harder cases arrive. Two policies contradict each other. A regional document is newer than the global one. A presentation contains a decision that was never transferred into the official process. The most frequently retrieved source is not the source the organisation would defend.
This is not primarily a retrieval problem. It is an operating-model problem wearing a data label.
Grounding is not the same as truth
RAG reduces one kind of uncertainty by giving a model relevant internal context. It does not remove the uncertainty inside that context.
A grounded answer can still be wrong because the source is outdated, incomplete, unauthorised for the user or simply less authoritative than another source. Citations make the path visible. They do not settle the conflict.
Before tuning retrieval, define what authority means for the domain:
- Which source is normative?
- Who owns its meaning, not only its storage location?
- How quickly must it be updated?
- Which source wins when two documents conflict?
- Which users and agents may retrieve it?
- What should the system do when no authoritative answer exists?
Without these decisions, the index becomes a faster way to surface organisational ambiguity.
Create a truth contract for each knowledge domain
A truth contract does not claim that one database contains every fact. It defines how a domain handles authority.
For one domain—pricing, HR policy, product information or service operations—capture six fields:
- Authoritative source: the system or document class the organisation will stand behind.
- Owner: the person accountable for meaning, conflicts and retirement.
- Freshness rule: when content expires or requires review.
- Access rule: who may retrieve which content in which context.
- Conflict rule: how the system ranks, flags or abstains when sources disagree.
- Evidence rule: which citation, version and timestamp must travel with the answer.
This contract belongs upstream of the vector index. It should shape ingestion, metadata, security trimming, evaluation and the user experience.
Permissions must survive retrieval
An AI application must not turn fragmented access into accidental transparency.
If a user could not open a document in the source system, retrieval should not quietly expose its content in an answer. Document-level access, user context and audit logs are therefore part of answer quality, not a separate security appendix.
The same applies to agents. An agent retrieving on behalf of a person and an autonomous agent operating with its own identity may require different boundaries. The architecture should make that distinction explicit before production.
Test the knowledge system, not only the model
Build an evaluation set from the awkward cases:
- outdated and current versions of the same policy;
- two sources with different authority;
- a user who may see only one of them;
- a question with no approved answer;
- content containing misleading instructions;
- a source that changed after the last index run.
Measure whether the right evidence was retrieved, whether access was respected and whether the application abstained when authority was unclear. A fluent answer is not the acceptance criterion.
The strongest objection: “Our knowledge is too messy”
That may be exactly why the project is valuable.
Do not clean the whole organisation before building anything. Choose one decision domain and make its ambiguity visible. The first useful output may be a map of duplicate sources, missing owners and unresolved conflicts. That is not a failed RAG pilot. It is organisational knowledge the business did not have before.
The next practical move
Select twenty questions the organisation expects an AI assistant or agent to answer. For each question, name the source you would defend in an audit or customer conversation.
Where the team cannot agree, do not tune the prompt. Assign an owner and resolve the truth contract.
RAG can make knowledge easier to reach. Only the organisation can decide which knowledge has the right to direct work.
Deutsche Ausgabe: Bevor RAG skaliert: Wer besitzt die Wahrheit der Organisation?
Sources and framing
- Microsoft Foundry: RAG and indexes
- Azure Architecture Center: secure multitenant RAG
- Azure AI Search: document-level access control
- NIST Generative AI Profile
Editorial note: This is an independent architecture and operating perspective, not a security design for a specific environment.
