Knowledge search with cited sources: cite or decline
Published
The short answer
What should an AI knowledge search do when its documents hold no answer?
Decline. A reliable knowledge search backs every answer with its source, that is file, page and line, and says openly when the corpus holds no passage that supports the answer. A fluent answer without evidence is more dangerous than none, because nobody notices that it is wrong.
The most dangerous answer
The most dangerous answer of a knowledge search is not the wrong one. It is the wrong one that sounds convincing.
An example: someone asks for the notice period in the framework contract. The search finds no matching passage and answers anyway, fluently and with a specific figure. Language models phrase things plausibly even when they lack the basis. Nobody notices until it gets expensive.
A simple rule
- Every answer names its source: file, page and line range.
- If the search finds no passage that supports the answer, it says so, without exception.
- An answer without evidence is never shown.
That does not make the search worse. It makes it checkable: whoever reads an answer can look up what it rests on.
What it looks like technically
The basis is a method known as retrieval-augmented generation: first, matching passages are searched in your documents, then a language model words the answer on their basis. Turning that into an answer with evidence takes three things.
- Provenance on ingestion: every passage is stored with file, page and line range. A passage without provenance is never stored.
- Permissions before the search: only passages the person asking may see are searched.
- A check before output: if none of the passages found supports the answer, the system declines instead of guessing.
How to check that it works
Such a search can be measured. You ask it questions whose answer is in the corpus and questions whose answer is not. Both are measured: how often an answer is correctly backed, and how often the search rightly declines.
The safeguards themselves need tests too. In the open-source project cite-or-decline, each of the 36 guards is removed on purpose, one at a time, to show that its test goes red.
Why declining builds trust
A search that sometimes says it finds nothing in the documents seems less helpful at first. In fact it is the one where you know where you stand: every answer brings its evidence, and every gap is named. A search people trust is a search people use.
Sources and evidence
- cite-or-decline on GitHub
- Lewis et al. (2020): Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- OWASP GenAI Security Project: Top 10 risks for applications built on language models (2025)
- Work index: cite-or-decline and green-but-blind(own evidence)
Why permissions belong in front of the model, not in the promptFocus area: knowledge search and assistants
Planning something along these lines? Briefly describe your project and you will get an honest assessment.
More articles
All articles- Why permissions belong in front of the model, not in the prompt
Is it enough to tell an AI assistant in its prompt not to reveal confidential content?
- AI in your own data centre or an EU cloud: how to keep control of your data
Can a company use AI without its data leaving the building or the EU?
- Systems that maintain themselves: measure, correct, update, approve
How does software stay secure and current after the handover without unreviewed changes going live?