Model Serving · Data artifact
Cacheable Prefix Prompt Layout
Data artifactModel ServingModelsarc:CacheablePrefixPromptLayout
A prompt-structure convention that places static, frequently reused content (system instructions, tool definitions, stable user profile) first and variable request-specific content last so provider prefix caches can match it.
Responsibility. Orders prompt segments by cacheability to maximize prefix-cache hits.
Also known as: Cache-friendly prompt structure, Static-first prompt ordering
Relationships
configures structural
Design guidance
- MUST place static, frequently reused content at the beginning of the prompt, followed by variable, request-specific content.
- SHOULD NOT include content that changes more often than it is reused (e.g., daily market data) in the cacheable prefix.
Quantitative guidance
As stated by the sources; verify before use.
- Financial advisor: cacheable prefix of 5,800 tokens (800 system prompt + 5,000 client profile) out of 9,000 input tokens; market data (3,000) and query (200) remain variable (Ch8.3).
- Regulatory context of 12,000 tokens (SEC 3,500, FINRA 4,200, DOL 2,800, state 1,500) repeated in 20,000 of 50,000 monthly requests = 240M input tokens ($600/month) of redundant transmission (Ch8.3).
Classification
- Patterns
- Prompt cachingPrefix matching
- Quality attributes
- Performance efficiency (ISO/IEC 25010)
- Risks mitigated
- Cache misses caused by variable content inside the prompt prefixRedundant retransmission of static context
Sources
- Ch8.3: T. Nguyen, "Token Economics and Architecture," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 8.3. ISBN: 9798244538229.