Model Serving · Data store
In-Process Response Cache
Data storeModel ServingModelsarc:InProcessResponseCache
A cache held in a replica's application memory (e.g., dictionary or LRU), private to that instance and lost on restart.
Responsibility. Serves hot cached entries from the replica's own memory.
Also known as: In-memory cache, Layer 1 cache, Local LRU cache, Exact-match cache (first level)
Variant of Response Cache abstract
When to choose. Choose for single-instance deployments, session-scoped data, or sub-millisecond latency needs where even Redis overhead matters.
Relationships
caches dependency
alternative to variability
Design guidance
- SHOULD NOT be the only cache layer in horizontally scaled deployments because fragmentation lowers effective hit rate and restarts cause cold starts.
Quantitative guidance
As stated by the sources; verify before use.
- Sub-millisecond access; ~2 GB of a 16 GB instance typically available for cache (Ch1.8).
- Five-instance deployment reached 40% hit rate with fragmented in-memory caches vs. 70% with shared cache (Ch1.8).
- Example Layer 1 TTL 15 minutes (Ch1.8).
- Exact-match memory lookup in <10 ms (Ch2.9).
- LRU cache with maxsize = 10,000 entries (Ref8.05).
Classification
- Patterns
- LRU cacheSession-scoped caching
- Technologies
- LangChain InMemoryCachefunctools.lru_cache
- Quality attributes
- Performance efficiency (ISO/IEC 25010)Maintainability (ISO/IEC 25010)
Sources
- Ch1.8: T. Nguyen, "Scalability and Production Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.8. ISBN: 9798244538229.
- Ch2.9: T. Nguyen, "Streaming and Real-Time Responses," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.9. ISBN: 9798244538229.
- Ref7.07: E. Li, V. Bellotti, R. Kraus, and R. Kao, "Build a retrieval-augmented generation (RAG) agent with NVIDIA Nemotron," NVIDIA Technical Blog, Sep. 23, 2025. [Online]. Available: https://developer.nvidia.com/blog/build-a-rag-agent-with-nvidia-nemotron/
- Ref8.05: "Cost Optimization and Resource Monitoring for Agent Systems," unpublished reference note (05-Cost-Optimization-Resource-Monitoring.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note