Skip to content

Feature: Hybrid retrieval (BM25 + vector fusion) for document-RAG #875

Description

@cybermaggedon

Summary

Document-RAG currently relies entirely on semantic (embedding) similarity to find relevant chunks. This works well for natural-language queries but struggles with exact-term matches — error codes, acronyms, product IDs, legal clause references, and other tokens that embeddings tend to fumble.

Adding a sparse/keyword retrieval path (BM25 or TF-IDF) and fusing the results with the existing vector search would significantly improve recall for these cases.

Motivation

A user asking "What does clause 7.3.2 say about indemnification?" may get poor results today because the embedding for "7.3.2" doesn't land near the chunk containing that clause number. A keyword match would find it immediately.

Proposed design

Please PR against the latest release/vX.Y branch.

Retrieval paths

  1. Vector path (exists today): query → extract-concepts → embed → doc_embeddings_client.query() → ranked chunk IDs
  2. Keyword path (new): query → tokenise/stem → BM25 search over chunk text → ranked chunk IDs

Fusion

Use Reciprocal Rank Fusion (RRF) to combine the two ranked lists:

score(doc) = Σ  1 / (k + rank_in_list)

where k is a constant (typically 60). This avoids the need to normalise scores across different retrieval systems.

Where it fits in the codebase

The retrieval logic lives in trustgraph-flow/trustgraph/retrieval/document_rag/document_rag.py, specifically the Query.get_docs() method. The fusion step would sit between chunk retrieval and the synthesis prompt.

Key integration points:

  • Query.get_docs() — add the BM25 query alongside the existing vector query, fuse results
  • A new BM25 index that is populated at chunk ingestion time (alongside the existing embedding index)
  • The doc_limit parameter already controls how many chunks to retrieve — this would apply to the fused result

BM25 index options

Several approaches, roughly ordered by complexity:

  1. In-process (rank_bm25 library): Simple, no new infrastructure. Rebuild index on startup from chunk store. Good for moderate-scale deployments.
  2. Elasticsearch/OpenSearch sidecar: Full-featured, scales well, but adds an infrastructure dependency.
  3. SQLite FTS5: Lightweight, file-based, good middle ground.

The choice of backend could be pluggable behind an interface, similar to how graph stores and vector stores are abstracted today.

Configuration

  • --retrieval-mode: vector (default, current behaviour), hybrid, or keyword
  • --bm25-weight: relative weight in fusion (default 1.0)
  • --vector-weight: relative weight in fusion (default 1.0)

What you'll learn

  • How TrustGraph's document-RAG retrieval pipeline works end-to-end
  • The chunk ingestion and embedding pipeline
  • How to add a new service dependency following existing patterns (embeddings client, triples client, etc.)

References

  • Current retrieval: trustgraph-flow/trustgraph/retrieval/document_rag/document_rag.py
  • Chunk ingestion: trustgraph-flow/trustgraph/chunking/
  • Existing embeddings client interface: trustgraph-base/trustgraph/base/document_embeddings_client.py
  • RRF paper: Cormack et al., "Reciprocal Rank Fusion outperforms Condorcet and individual Rank Learning Methods" (SIGIR 2009)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions