Knowledge & Data · Software component
Vector Retriever
Software componentKnowledge & DataKnowledge & DataVariation point (abstract)arc:VectorRetriever
A retriever that ranks document chunks by embedding similarity to the query.
Responsibility. Retrieves semantically similar chunks by vector similarity search.
Also known as: Dense Retriever, Vector RAG, Semantic Search, Dense vector search, Dense retriever, Semantic search, Query-driven context retrieval
Variant of Retriever abstract
When to choose. Choose when queries need conceptual understanding or semantic similarity across varied terminology (simple Q&A, documentation search, customer support).
Variants
| Variant | When to choose |
|---|---|
| Approximate Vector Search Retriever | Choose when sub-linear, millisecond-latency search over millions of vectors matters more than exact, reproducible top-K results. |
| Coarse-to-Fine Vector Retriever | Choose when embeddings are MRL-trained and corpus scale makes full-dimensional search latency or memory prohibitive. |
| Exact Vector Search Retriever | Choose when deterministic, reproducible retrieval is required (scientific experiments, compliance audits, A/B testing), accepting slower queries. |
| Metadata-Filtered Retriever | — |
Relationships
invokes dependency
- Embedding Service abstract Ch1.7A Ch2.1 +3
- Lexical Reranker Ch6.2B
- Text Embedding Service Ch5.7 Ch5.8 +1
- Vector Store Query API abstract Ch6.2B
is cached by dependency
is invoked by dependency
reads dependency
- Embedded Vector Index Store Ref7.07
- Vector Index Store abstract Ch1.7A Ch2.1 +10
emits telemetry to dynamic
receives data from dynamic
sends data to dynamic
fails over to control
is orchestrated by control
is evaluated by assurance
alternative to variability
- Graph Retriever Ch1.7A Ch5.7 +1
- Hybrid Retriever abstract Ch1.7A Ch5.7 +1
Design guidance
- SHOULD NOT be relied on alone for multi-hop relational questions, where relationships dissolve into embeddings.
- SHOULD retrieve fewer, higher-quality results rather than maximizing quantity; excess chunks cause truncation and noise.
- SHOULD replace large static context blocks with retrieval of only the segments relevant to each query.
- SHOULD form the retrieval query from both the user question and relevant user state (e.g., portfolio holdings).
- MUST validate that reduced context does not degrade quality (accuracy, CSAT, completion, escalation) before adoption.
Quantitative guidance
As stated by the sources; verify before use.
- Vector RAG only: high semantic coverage, low relational accuracy, p95 ~120ms, low complexity (Ch1.7B).
- LIMIT: fixed-granularity chunking (typically 512 tokens) separates metadata from properties and breaks reasoning links; dense retrievers often return semantic noise (Ch3.3).
- Worked example: top K=5 by cosine similarity (0.92, 0.87, 0.84, 0.79, 0.76); top 3 with similarity >= 0.84 passed to generation (Ch5.8).
- Multi-stage filtering places ~8,000 instead of ~50,000 tokens into context, often performing better with 85% fewer tokens (Ch5.9).
- Top-k retrieval of 5-10 snippets (~1,200 tokens) replaced a 3,000-token market snapshot (60% context cut), reducing input 9,000 -> 7,200 tokens and monthly cost $5,325 -> $4,365 ($960) (Ch8.3).
- Quality after RAG: accuracy 91% -> 92%, CSAT 4.3/5 unchanged, completion 95% -> 94%, escalation 8% unchanged (Ch8.3).
- RAG can replace an 8,000-token static context with ~500-token dynamic retrieval (Ch8.3).
Classification
- Patterns
- Top-k similarity searchMulti-stage selective retrieval, stage 1: broad semantic search for ~top-50 candidatesSemantic code search over source and documentation
- Technologies
- PineconeWeaviateQdrant
- Quality attributes
- Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)Maintainability (ISO/IEC 25010)
- Risks mitigated
- Keyword mismatch between query and document vocabularyContext bloat from static full-dataset inclusion
Sources
- Ch1.7A: T. Nguyen, "Relational Reasoning with Knowledge Graphs - The Fundamentals, Integration, and Extraction," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.7A. ISBN: 9798244538229.
- Ch1.7B: T. Nguyen, "Relational Reasoning with Knowledge Graphs - Hybrid RAG+KG Integration," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.7B. ISBN: 9798244538229.
- Ch2.1: T. Nguyen, "Framework Landscape and Selection," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.1. ISBN: 9798244538229.
- Ch2.2: T. Nguyen, "LangGraph," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.2. ISBN: 9798244538229.
- Ch2.9: T. Nguyen, "Streaming and Real-Time Responses," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.9. ISBN: 9798244538229.
- Ch3.3: T. Nguyen, "Web Navigation and Interaction Benchmarks," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.3. ISBN: 9798244538229.
- Ch5.7: T. Nguyen, "Episodic Memory," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.7. ISBN: 9798244538229.
- Ch5.8: T. Nguyen, "Semantic Memory," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.8. ISBN: 9798244538229.
- Ch5.9: T. Nguyen, "Working Memory," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.9. ISBN: 9798244538229.
- Ch5.13: T. Nguyen, "Hybrid Decision Systems Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.13. ISBN: 9798244538229.
- Ch6.1: T. Nguyen, "Embeddings and RAG Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.1. ISBN: 9798244538229.
- Ch6.5: T. Nguyen, "Production RAG Systems," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.5. ISBN: 9798244538229.
- Ch8.1: T. Nguyen, "Latency Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 8.1. ISBN: 9798244538229.
- Ch8.3: T. Nguyen, "Token Economics and Architecture," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 8.3. ISBN: 9798244538229.
- Ref7.07: E. Li, V. Bellotti, R. Kraus, and R. Kao, "Build a retrieval-augmented generation (RAG) agent with NVIDIA Nemotron," NVIDIA Technical Blog, Sep. 23, 2025. [Online]. Available: https://developer.nvidia.com/blog/build-a-rag-agent-with-nvidia-nemotron/