Tools & Integration · Data store
Tool Result Cache
Data storeTools & IntegrationOrchestration & Toolsarc:ToolResultCache
A store of previously retrieved tool or data-source results that serves as a degraded fallback when live sources fail or are unavailable.
Responsibility. Serves cached results when primary sources are unavailable.
Also known as: Cached results fallback, Cached data source, Cached/historical data fallback, Tool Result Cache (TTL), API result cache, Result caching, Repeat-applicant cache, Customer profile cache, Tool call cache, L3 cache, Tool call memoization cache, LRU tool result cache, Inventory result cache, External API result cache
Relationships
is configured by structural
caches dependency
is read by dependency
is written by dependency
is constrained by control
- Cache Invalidation Policy abstract Ch7.3
Design guidance
- SHOULD NOT serve hours-old cached results as a fallback for a failed live query without clearly indicating that the data is outdated.
- SHOULD cache results of repeated external queries (e.g., customer profiles, recent orders) to eliminate redundant API calls.
- SHOULD serve cached data when APIs time out or rate limits are hit rather than failing requests.
- MUST cache only idempotent READ operations; WRITE operations MUST always execute.
- Cache keys MUST capture every parameter that affects the tool's output, normalized so equivalent representations share an entry.
- SHOULD bound cache size (e.g., LRU) to trade memory for latency.
- SHOULD cache results of slow or unreliable external tool/API calls for frequently queried keys.
- SHOULD estimate the hit rate from actual query logs before deploying; uniform query distributions yield low hit rates and may need other strategies (async execution, timeouts, rerouting).
Quantitative guidance
As stated by the sources; verify before use.
- Cache hits served in milliseconds vs 200-500ms database queries; 5-minute TTL for account lookups/product info (Ch3.4).
- About 80% of e-commerce support queries accessed the same cached customer data repeatedly (Ch3.10).
- Tool call caching reported up to 1.69x latency reduction without accuracy loss for idempotent operations (Ch4.7, unnamed research).
- RAG example: ~40% hit rate saving 100 ms of tool execution; 50K results x 2 KB = 100 MB (Ch4.7).
- 73% hit rate over 100 queries: latency -51% (2,341 to 1,134ms), cost -62% ($0.014 to $0.0053), ~$3,150/month saved (Ch7.3).
- Warm cache (100% hits): 10ms latency, $0.000056 per query, 99.6% cost reduction; hits return in ~0.5ms (Ch7.3).
- maxsize=128 entries uses ~50KB memory (Ch7.3).
- 1,000-entry LRU cache with 5-minute validity; 100 ms cached vs 400 ms external API; 75% of lookups hit the top 200 products -> tool latency 400->175 ms (56%) (Ch8.1).
Classification
- Patterns
- Graceful degradationTTL-based caching of tool resultsREAD/WRITE operation classificationHash-based normalized cache keysTTL + event-driven invalidationLRU caching of external API resultsLocality of reference
- Technologies
- Python functools.lru_cache
- Quality attributes
- Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)
- Risks mitigated
- Total failure on dependency outageRedundant tool callsStale data served to usersSilently serving stale dataTail latency from load spikes in slow legacy external systems
Sources
- Ch1.1A: T. Nguyen, "Designing User Interfaces for Intuitive Human-Agent Interaction," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.1A. ISBN: 9798244538229.
- Ch1.2: T. Nguyen, "Core Agent Patterns," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.2. ISBN: 9798244538229.
- Ch2.8: T. Nguyen, "Error Handling and Resilience," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.8. ISBN: 9798244538229.
- Ch3.4: T. Nguyen, "Tuning Model Parameters for Production Performance," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.4. ISBN: 9798244538229.
- Ch3.7: T. Nguyen, "Tool Usage Auditing," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.7. ISBN: 9798244538229.
- Ch3.10: T. Nguyen, "Efficiency Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.10. ISBN: 9798244538229.
- Ch4.7: T. Nguyen, "Scaling Strategies," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.7. ISBN: 9798244538229.
- Ch7.3: T. Nguyen, "NeMo Agent Toolkit Profiling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.3. ISBN: 9798244538229.
- Ch8.1: T. Nguyen, "Latency Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 8.1. ISBN: 9798244538229.
- Ch8.2A: T. Nguyen, "Error Rates and Reliability," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 8.2A. ISBN: 9798244538229.