Knowledge & Data · Software component
Dense-Sparse Hybrid Retriever
Software componentKnowledge & DataKnowledge & Dataarc:DenseSparseHybridRetriever
A retriever that runs dense vector search and sparse keyword search in parallel and returns the union of their results.
Responsibility. Combines parallel dense and sparse retrieval into one result set.
Also known as: Hybrid dense+sparse retrieval, Hybrid search (semantic + keyword), Hybrid search, Dense+sparse hybrid retriever, Weaviate hybrid query
Variant of Retriever abstract
When to choose. Choose in latency-sensitive RAG agents where retrieval dominates time-to-first-token.
Relationships
invokes dependency
- Fusion Weight Selector Ch6.2B
- Keyword Retriever Ch2.9 Ch6.1 +1
- Vector Retriever abstract Ch2.9 Ch6.1 +1
- Vector Store Query API abstract Ch6.2B
is invoked by dependency
reads dependency
receives data from dynamic
sends data to dynamic
is evaluated by assurance
Design guidance
- SHOULD combine semantic similarity and keyword matching to improve retrieval relevance over pure embedding search.
- SHOULD start dense-only, measure on representative queries, and add hybrid search selectively where lexical precision matters.
- SHOULD start with equal fusion weighting (alpha 0.5) and tune the weight from evaluation metrics for the content and query distribution.
- SHOULD over-fetch candidates from each method (e.g., 2x k) so strong candidates survive fusion.
- SHOULD bias alpha toward semantics (0.7-0.8) for conceptual queries and toward keywords (0.3-0.4) for technical terms and product codes.
- SHOULD log fused and component scores to diagnose unexpected rankings.
Quantitative guidance
As stated by the sources; verify before use.
- Parallel dense+sparse retrieval can cut retrieval time in half where sparse search finds results faster (Ch2.9).
- Hybrid search improves accuracy 15-25% over pure vector search in typical evaluations (Ch6.1).
- RRF constant 60 prevents top ranks from dominating (Ch6.1).
- Hybrid search adds 20-40% latency vs pure vector search for a 15-25% accuracy gain (Ch6.2B).
- Example fusion weight alpha = 0.7 emphasizing semantic search while retaining keyword recall (Ch6.5).
- Example weights 0.7 semantic / 0.3 BM25 with k=10 semantic candidates (Ref7.07).
Classification
- Patterns
- Parallel retrievalResult unionBM25 / SPLADE / BGE-M3 sparse vectors with dense vectors in one collectionReciprocal Rank Fusion (score = weight/(rank+60))Alpha-weighted fusion (alpha=1 dense, 0 sparse, 0.5 equal)Over-fetch k*2 candidates per method before fusionParallel dense and sparse retrievalKeyword matching restricted to content fieldScore explanation (explainScore) logging for diagnosis
- Technologies
- WeaviateMilvusLangChain EnsembleRetriever
- Quality attributes
- Performance efficiency (ISO/IEC 25010)Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)
- Risks mitigated
- Dense retrieval failures on specialized terminology, exact phrases and rare entities
Sources
- Ch2.9: T. Nguyen, "Streaming and Real-Time Responses," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.9. ISBN: 9798244538229.
- Ch3.10: T. Nguyen, "Efficiency Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.10. ISBN: 9798244538229.
- Ch4.1: T. Nguyen, "Introduction to AI Agent Deployment and Scaling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.1. ISBN: 9798244538229.
- Ch6.1: T. Nguyen, "Embeddings and RAG Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.1. ISBN: 9798244538229.
- Ch6.2B: T. Nguyen, "Production Vector Database Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.2B. ISBN: 9798244538229.
- Ch6.5: T. Nguyen, "Production RAG Systems," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.5. ISBN: 9798244538229.
- Ref7.07: E. Li, V. Bellotti, R. Kraus, and R. Kao, "Build a retrieval-augmented generation (RAG) agent with NVIDIA Nemotron," NVIDIA Technical Blog, Sep. 23, 2025. [Online]. Available: https://developer.nvidia.com/blog/build-a-rag-agent-with-nvidia-nemotron/