Model Serving · Data store
Response Cache
Data storeModel ServingModelsVariation point (abstract)arc:ResponseCache
A cache of complete agent/LLM responses keyed by (normalized) query so repeated requests are answered without inference.
Responsibility. Serves previously computed responses for repeated queries without inference.
Also known as: LLM response cache, Query result cache, Cached-result fallback, Application-Level Response Cache, Result caching, Generated response cache, Request caching, Multi-stage caching
Variants
| Variant | When to choose |
|---|---|
| Database Response Cache | Choose when audit requires persisting all outputs, capacity exceeds practical RAM, cache updates must be atomic with other data, or SQL querying of cached data is needed. |
| Distributed Response Cache | Choose for horizontally scaled systems with >20% query repetition or where cache must survive instance restarts. |
| In-Process Response Cache | Choose for single-instance deployments, session-scoped data, or sub-millisecond latency needs where even Redis overhead matters. |
| Semantic Cache | Choose when users ask the same questions in different wording and incremental savings exceed embedding costs. |
Relationships
is configured by structural
caches dependency
is read by dependency
- Agent Controller abstract Ch1.8 Ch2.9
- Model Router Ch2.8
is written by dependency
Design guidance
- SHOULD be introduced only when query repetition exceeds ~20%; below that caching overhead exceeds benefit.
- SHOULD be justified by break-even hit rate H = I / (R x C) (infrastructure cost over uncached inference cost).
- SHOULD combine layers (in-process, distributed, prefix) checked fastest-first when cost optimization justifies complexity.
- SHOULD monitor hit rate, eviction rate and latency percentiles.
- MUST NOT cache responses that depend on user-specific context.
- SHOULD cache results for cost-sensitive applications.
- SHOULD monitor cache hit rate to verify caching effectiveness (Ref8.05).
Quantitative guidance
As stated by the sources; verify before use.
- Multi-layer caching achieves 80-90% hit rates in production FAQ workloads (Ch1.8).
- Break-even hit rate 3.75% for 2M queries/month at $0.002/query with $150/month cache (Ch1.8).
- 80% hit rate cut monthly cost from $4,000 to $950 (76%) in the example (Ch1.8).
- Aggressive response caching reduces API costs 40-60% for FAQ agents; top 100 FAQ queries covered 30% of volume (Ch3.4).
- Cache hit reduces latency from ~2 s to ~10 ms; query caching reduces costs by 60-80% for typical workloads (Ch6.5).
- Same query asked 100 times -> 1 inference + 99 cache hits (99% saved); typical savings 30-60% with a good caching strategy (Ref8.05).
Classification
- Patterns
- Multi-layer cachingWrite-throughWrite-behindCache-asideQuery normalizationLRU evictionTTL expirationAsymmetric TTL per layer
- Quality attributes
- Cost efficiencyPerformance efficiency (ISO/IEC 25010)
- Risks mitigated
- Redundant inference costServing stale answers (with invalidation)Stale cached answers after underlying information changesRedundant LLM callsRunaway generation cost
Sources
- Ch1.8: T. Nguyen, "Scalability and Production Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.8. ISBN: 9798244538229.
- Ch2.8: T. Nguyen, "Error Handling and Resilience," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.8. ISBN: 9798244538229.
- Ch2.9: T. Nguyen, "Streaming and Real-Time Responses," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.9. ISBN: 9798244538229.
- Ch3.4: T. Nguyen, "Tuning Model Parameters for Production Performance," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.4. ISBN: 9798244538229.
- Ch3.10: T. Nguyen, "Efficiency Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.10. ISBN: 9798244538229.
- Ch6.5: T. Nguyen, "Production RAG Systems," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.5. ISBN: 9798244538229.
- Ch10.1: T. Nguyen, "Conversational UI," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.1. ISBN: 9798244538229.
- Ref7.15: "Advanced Agentic AI Optimization Techniques," unpublished reference note (15-Advanced-Agentic-Optimization.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref7.18: "Chapter 7 Summary: NVIDIA Platform Implementation," unpublished reference note (18-Chapter-7-Summary.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref8.05: "Cost Optimization and Resource Monitoring for Agent Systems," unpublished reference note (05-Cost-Optimization-Resource-Monitoring.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note