Model Serving · Data store

Distributed Response Cache

Data storeModel ServingModelsarc:DistributedResponseCache

A network-accessible key-value cache shared by all replicas so a response computed once benefits every instance.

Responsibility. Shares cached responses across all replicas over the network.

Also known as: Redis cache, Layer 2 cache, Exact-match cache, Shared cache, Redis query cache, Exact-match query cache

Variant of Response Cache abstract

When to choose. Choose for horizontally scaled systems with >20% query repetition or where cache must survive instance restarts.

is read by; is written byspecializesalternative toalternative tois target of alternativeToRAG Query Orchestrator: is read by; is written byRAG Query OrchestratorResponse Cache: specializesResponse CacheIn-Process Response Cache: alternative toIn-Process Response CacheSemantic Cache: alternative toSemantic CacheDatabase Response Cache: is target of alternativeToDatabase Response Cache
Direct neighbourhood (hover for relationship types)

Relationships

is read by dependency

is written by dependency

alternative to variability

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
Consistent hashing of keysConnection poolingPipeliningCluster shardingAutomatic failoverMD5-hashed cache keys
Technologies
RedisRedis ClusterRedis SentinelAWS ElastiCacheAzure Cache for RedisGoogle Cloud Memorystore
Quality attributes
Performance efficiency (ISO/IEC 25010)Cost efficiencyReliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)
Risks mitigated
Cache fragmentation across replicasCold start after restart

Sources

  1. Ch1.8: T. Nguyen, "Scalability and Production Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.8. ISBN: 9798244538229.
  2. Ch6.5: T. Nguyen, "Production RAG Systems," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.5. ISBN: 9798244538229.