Model Serving · Data artifact

Cache Policy

Data artifactModel ServingModelsarc:CachePolicy

A configuration specifying cache TTLs per layer, eviction policy, key normalization and coordination strategy.

Responsibility. Defines expiry, eviction and population rules for cache layers.

Also known as: TTL configuration, Eviction policy, Multi-tier cache configuration

configuresconfiguresconfiguresconfiguresconfiguresconfiguresTool Result Cache: configuresTool Result CacheResponse Cache: configuresResponse CacheCache Invalidator: configuresCache InvalidatorReasoning Chain Cache: configuresReasoning Chain CacheEmbedding Cache: configuresEmbedding CacheRetrieval Result Cache: configuresRetrieval Result Cache
Direct neighbourhood (hover for relationship types)

Relationships

configures structural

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
TTL expirationLRU evictionValue-aware evictionAsymmetric TTLMulti-tier cache hierarchy (L1 reasoning templates, L2 embeddings, L3 tool results)Layer-specific invalidationTime-based (TTL) invalidationEvent-based invalidationCapacity-based LRU eviction
Quality attributes
Maintainability (ISO/IEC 25010)
Risks mitigated
Stale responsesMemory exhaustion

Sources

  1. Ch1.8: T. Nguyen, "Scalability and Production Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.8. ISBN: 9798244538229.
  2. Ch3.4: T. Nguyen, "Tuning Model Parameters for Production Performance," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.4. ISBN: 9798244538229.
  3. Ch4.7: T. Nguyen, "Scaling Strategies," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.7. ISBN: 9798244538229.
  4. Ch6.5: T. Nguyen, "Production RAG Systems," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.5. ISBN: 9798244538229.