Knowledge & Data · Data store
Embedding Cache
Data storeKnowledge & DataKnowledge & Dataarc:EmbeddingCache
A cache of previously computed vector embeddings for frequently queried documents or queries, so semantic search avoids recomputing them.
Responsibility. Stores and replays embedding vectors to avoid repeated embedding computation.
Also known as: Embedding caching, Query embedding cache, L2 cache
Relationships
is configured by structural
caches dependency
- Embedding Service abstract Ch3.10 Ch4.7 +2
- Text Embedding Service Ch6.5
is read by dependency
- Embedding Service abstract Ch3.10
is written by dependency
Design guidance
- SHOULD maintain vector representations of frequently queried documents to avoid recalculating embeddings.
- SHOULD match cached query embeddings by cosine similarity with a threshold balancing hit rate against retrieval fidelity.
- SHOULD be invalidated when the embedding model is updated.
- SHOULD be used aggressively with CPU embedding inference; MAY be minimised when GPU throughput makes on-the-fly embedding viable, avoiding cache invalidation complexity.
- SHOULD key on the exact query string and apply TTL expiry to prevent unbounded growth.
Quantitative guidance
As stated by the sources; verify before use.
- Embedding computation is often 20-40% of RAG request time (Ch4.7).
- Thresholds: 0.98+ high fidelity/low hit rate; 0.85-0.90 higher hit rate but risk; production typically 0.92-0.95 (Ch4.7).
- RAG example: ~60% query-embedding hit rate saving 50 ms; 10K queries x 4 KB = 40 MB (Ch4.7).
- Hit rates of 20-40% common, saving 50-100 ms and ~$0.0001 per hit (Ch6.5).
Classification
- Patterns
- Multi-layer cachingSemantic-similarity lookup of cached query embeddings
- Technologies
- RedisMemcached
- Quality attributes
- Cost efficiencyPerformance efficiency (ISO/IEC 25010)
- Risks mitigated
- Redundant embedding computation
Sources
- Ch3.10: T. Nguyen, "Efficiency Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.10. ISBN: 9798244538229.
- Ch4.7: T. Nguyen, "Scaling Strategies," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.7. ISBN: 9798244538229.
- Ch6.1: T. Nguyen, "Embeddings and RAG Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.1. ISBN: 9798244538229.
- Ch6.5: T. Nguyen, "Production RAG Systems," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.5. ISBN: 9798244538229.
- Ref7.15: "Advanced Agentic AI Optimization Techniques," unpublished reference note (15-Advanced-Agentic-Optimization.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note