Memory · Software component
Context Budget Allocator
Software componentMemoryCognition & Memoryarc:ContextBudgetAllocator
A memory-management component that divides an agent's context-window token budget among competing contents (system prompt, history, retrieval, reasoning traces, output) according to task type and observed utilization.
Responsibility. Allocates and re-allocates context-window token budget across context components during an interaction.
Also known as: Dynamic budget allocation, Token budget manager, Working memory budget monitor
Relationships
is configured by structural
invokes dependency
triggers dynamic
constrains control
monitors assurance
- Working Memory Buffer abstract Ch5.9
Design guidance
- SHOULD classify task type (explicitly from metadata or implicitly from query structure) and apply task-specific allocations instead of fixed percentages.
- SHOULD hold 15-25% of capacity as uncommitted reserve for follow-up retrieval, longer reasoning paths or longer output.
- SHOULD escalate compression aggressiveness with observed utilization rather than applying a uniform compression policy.
- MAY signal generous available capacity to the model while enforcing a lower internal usage cap, to avoid context-anxiety behaviour.
- SHOULD account for domain-specific tokenization efficiency (code, medical terminology, non-English text) when budgeting.
Quantitative guidance
As stated by the sources; verify before use.
- Simple QA: ~60% of available tokens to retrieval, ~10% to reasoning; complex analysis: ~50% to reasoning, ~20% to retrieval (Ch5.9).
- Utilization bands: <60% light compression; 60-80% paragraph summaries; >80% aggressive bullet-point compression and restricted retrieval (Ch5.9).
- Reserve example: 100,000-token window with 20,000 fixed overhead -> 65,000 planned + 15,000 reserve (Ch5.9).
- Context anxiety (actual 200k window): signalled 20k vs 100k available gave 127 vs 184 lines, 8% vs 24% comment density, 2.3 vs 5.6 edge cases/function, 73% vs 88% correctness (Ch5.9).
- Production strategy: configure a 1M-token window but cap actual usage at 200,000 tokens (Ch5.9).
Classification
- Patterns
- Dynamic task-type budget allocationToken budgeting with reservesContext-aware (utilization-driven) compressionReal-time budget monitoring during generationPerceived-capacity framing (context-anxiety mitigation)
- Quality attributes
- Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)Performance efficiency (ISO/IEC 25010)Cost efficiency
- Risks mitigated
- Context overflow and truncated responsesUnexpected token-consumption spikesPremature information lossConservative, minimal outputs under perceived context scarcity
Sources
- Ch5.9: T. Nguyen, "Working Memory," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.9. ISBN: 9798244538229.