Summary
Document-RAG currently relies entirely on semantic (embedding) similarity to find relevant chunks. This works well for natural-language queries but struggles with exact-term matches — error codes, acronyms, product IDs, legal clause references, and other tokens that embeddings tend to fumble.
Adding a sparse/keyword retrieval path (BM25 or TF-IDF) and fusing the results with the existing vector search would significantly improve recall for these cases.
Motivation
A user asking "What does clause 7.3.2 say about indemnification?" may get poor results today because the embedding for "7.3.2" doesn't land near the chunk containing that clause number. A keyword match would find it immediately.
Proposed design
Please PR against the latest release/vX.Y branch.
Retrieval paths
- Vector path (exists today): query →
extract-concepts → embed → doc_embeddings_client.query() → ranked chunk IDs
- Keyword path (new): query → tokenise/stem → BM25 search over chunk text → ranked chunk IDs
Fusion
Use Reciprocal Rank Fusion (RRF) to combine the two ranked lists:
score(doc) = Σ 1 / (k + rank_in_list)
where k is a constant (typically 60). This avoids the need to normalise scores across different retrieval systems.
Where it fits in the codebase
The retrieval logic lives in trustgraph-flow/trustgraph/retrieval/document_rag/document_rag.py, specifically the Query.get_docs() method. The fusion step would sit between chunk retrieval and the synthesis prompt.
Key integration points:
Query.get_docs() — add the BM25 query alongside the existing vector query, fuse results
- A new BM25 index that is populated at chunk ingestion time (alongside the existing embedding index)
- The
doc_limit parameter already controls how many chunks to retrieve — this would apply to the fused result
BM25 index options
Several approaches, roughly ordered by complexity:
- In-process (rank_bm25 library): Simple, no new infrastructure. Rebuild index on startup from chunk store. Good for moderate-scale deployments.
- Elasticsearch/OpenSearch sidecar: Full-featured, scales well, but adds an infrastructure dependency.
- SQLite FTS5: Lightweight, file-based, good middle ground.
The choice of backend could be pluggable behind an interface, similar to how graph stores and vector stores are abstracted today.
Configuration
--retrieval-mode: vector (default, current behaviour), hybrid, or keyword
--bm25-weight: relative weight in fusion (default 1.0)
--vector-weight: relative weight in fusion (default 1.0)
What you'll learn
- How TrustGraph's document-RAG retrieval pipeline works end-to-end
- The chunk ingestion and embedding pipeline
- How to add a new service dependency following existing patterns (embeddings client, triples client, etc.)
References
- Current retrieval:
trustgraph-flow/trustgraph/retrieval/document_rag/document_rag.py
- Chunk ingestion:
trustgraph-flow/trustgraph/chunking/
- Existing embeddings client interface:
trustgraph-base/trustgraph/base/document_embeddings_client.py
- RRF paper: Cormack et al., "Reciprocal Rank Fusion outperforms Condorcet and individual Rank Learning Methods" (SIGIR 2009)
Summary
Document-RAG currently relies entirely on semantic (embedding) similarity to find relevant chunks. This works well for natural-language queries but struggles with exact-term matches — error codes, acronyms, product IDs, legal clause references, and other tokens that embeddings tend to fumble.
Adding a sparse/keyword retrieval path (BM25 or TF-IDF) and fusing the results with the existing vector search would significantly improve recall for these cases.
Motivation
A user asking "What does clause 7.3.2 say about indemnification?" may get poor results today because the embedding for "7.3.2" doesn't land near the chunk containing that clause number. A keyword match would find it immediately.
Proposed design
Please PR against the latest
release/vX.Ybranch.Retrieval paths
extract-concepts→ embed →doc_embeddings_client.query()→ ranked chunk IDsFusion
Use Reciprocal Rank Fusion (RRF) to combine the two ranked lists:
where
kis a constant (typically 60). This avoids the need to normalise scores across different retrieval systems.Where it fits in the codebase
The retrieval logic lives in
trustgraph-flow/trustgraph/retrieval/document_rag/document_rag.py, specifically theQuery.get_docs()method. The fusion step would sit between chunk retrieval and the synthesis prompt.Key integration points:
Query.get_docs()— add the BM25 query alongside the existing vector query, fuse resultsdoc_limitparameter already controls how many chunks to retrieve — this would apply to the fused resultBM25 index options
Several approaches, roughly ordered by complexity:
The choice of backend could be pluggable behind an interface, similar to how graph stores and vector stores are abstracted today.
Configuration
--retrieval-mode:vector(default, current behaviour),hybrid, orkeyword--bm25-weight: relative weight in fusion (default 1.0)--vector-weight: relative weight in fusion (default 1.0)What you'll learn
References
trustgraph-flow/trustgraph/retrieval/document_rag/document_rag.pytrustgraph-flow/trustgraph/chunking/trustgraph-base/trustgraph/base/document_embeddings_client.py