Small follow-up after having this running for a while. One thing I’ve changed in my own approach is that I no longer treat the knowledge base as a static FAQ that the agent simply searches.
In practice, the harder problem is keeping the information useful when the underlying business information changes.
For example, pricing questions can look simple from the user’s side, but a useful answer may depend on several pieces of context: the type of service, amount of material, access conditions, location, and sometimes the way the work is handled. If those details are scattered across different documents, retrieval can return an answer that is technically relevant but incomplete.
I’ve had better results by treating each important business topic as its own maintained knowledge unit, with:
- the actual answer
- supporting details or conditions
- the last updated date
- the service/category it belongs to
- location or audience context where relevant
- a source or owner responsible for keeping it current
This also changed how I think about RAG evaluation. I don’t just ask whether the correct document was retrieved anymore. I check whether the retrieved information contains enough context for the agent to give a complete answer.
Pricing information is a good example of this. A knowledge base can retrieve a pricing page correctly but still produce a poor answer if the actual factors affecting the price are stored somewhere else.
I’ve also found that smaller, well-defined knowledge units are easier to maintain than one huge FAQ document. When one business rule changes, you can update that specific unit and re-index it instead of potentially refreshing the whole knowledge base.
My current workflow is therefore closer to:
business information → structured knowledge units → metadata → retrieval → answer → freshness check
rather than simply:
FAQ document → embedding → answer
I’m still experimenting with how aggressively to re-index after updates. For frequently changing information, event-based updates seem more sensible, while relatively stable content can probably tolerate scheduled refreshes.
Curious if anyone else has tested this in production — particularly whether you found knowledge organization and freshness to be a bigger bottleneck than the actual retrieval method.