Memory · Software component
Context Window Manager
Software componentMemoryCognition & MemoryVariation point (abstract)arc:ContextWindowManager
A memory component that tracks and manages how much conversation, reasoning, and observation history fits within the model's context window, signalling when older context will be dropped.
Responsibility. Selects and budgets content to fit the context window.
Also known as: Context limit tracking, Context filtering, Summarization layer, Hierarchical context builder, State pruner, State pruning strategy, Context slimming, Context management, Working memory manager, Conversation history management strategy, Context pruning
Variants
| Variant | When to choose |
|---|---|
| Full Conversation Buffer | Choose for moderate conversation lengths (roughly 10-20 exchanges) where perfect recall of prior turns matters; switch to sliding-window or summarizing variants when history approaches token limits. |
| Hierarchical History Compressor | Choose for long-running conversations where historical context determines current interpretation and users reference earlier turns, accepting retrieval latency and the need to decide which history to fetch. |
| Importance-Weighted History Retainer | Choose for goal-oriented conversations in which certain turns establish critical constraints or preferences that later turns build on, provided an accurate relevance model is available. |
| Sliding-Window History Truncator | Choose as the simplest bound when losing early conversation detail is acceptable. |
| Summarizing History Compressor | Choose when early context (user goals, key facts, binding decisions) remains relevant later and naive truncation would lose it. |
| Trajectory Pruner | Choose for multi-step agents whose trajectories accumulate dead ends, restatements and stale turns. |
Relationships
invokes dependency
is invoked by dependency
reads dependency
- Working Memory Buffer abstract Ch1.2 Ch1.6 +1
writes dependency
- Working Memory Buffer abstract Ch1.4 Ch1.6 +1
emits telemetry to dynamic
is triggered by dynamic
sends data to dynamic
constrains control
Design guidance
- SHOULD surface to users when context is approaching limits and older context might be dropped.
- MUST NOT dump all retrieved memories into prompts; detailed information SHOULD appear only when directly relevant.
- SHOULD supply semantic knowledge as on-demand retrieval rather than permanent context.
- MUST NOT treat state as an append-only log; state must be actively bounded regardless of context window size.
- SHOULD still prune for agents running for hours or days even with long-context models, since window size only delays the problem.
- MUST log token accounting before each LLM call, comparing available tokens with required content and flagging truncation.
- SHOULD count tokens before each agent handoff with fallback pathways for budget-exceeding inputs.
- SHOULD separate short-term working memory from long-term episodic memory to prevent prompt bloat.
- SHOULD expire stale information with time-to-live rules.
- SHOULD establish context sufficiency thresholds below which the agent acknowledges inadequate information instead of speculating.
- SHOULD NOT default to truncation; use it only when old information genuinely becomes irrelevant for the task type.
- SHOULD apply compression before saturation (at a planned threshold) rather than at the capacity-crisis point.
- SHOULD prune irrelevant conversation history as dialogs grow longer to limit repeated input tokens.
Quantitative guidance
As stated by the sources; verify before use.
- Systematic performance drops begin around 3,000 tokens despite much larger context windows (Ch1.4).
- Production data show 40-80% failure rates in multi-agent environments due to memory engineering failures (Ch1.4).
- 3-5 turn conversations hold ~2,000 tokens of state; after 20 tool-heavy turns state can reach 50,000-100,000 tokens (Ch1.6).
- Unbounded state raised latency from <1 s to ~2 s at turn 10, ~5 s at turn 20, ~10 s at turn 30; ~30 s responses at ~200,000 tokens (Ch1.6).
- Fintech case: Unicode-driven truncation affected 3% of transactions; fixes improved fraud detection accuracy 78% within two weeks (Ch3.6).
- Conversation history can exceed 50,000 tokens after ~100 turns in a 200,000-token window (~25% of capacity) (Ch5.9).
- Multi-turn growth example: 500 -> 900 -> 1,400 tokens over the first three turns (Ch5.9).
- Support example (100k window): available response capacity falls from ~99,300 (turns 1-5) to ~86,200 (turn 15), ~74,000 (turn 30) and ~57,000 tokens (turn 50) without intervention (Ch5.9).
Classification
- Patterns
- Hierarchical contextSummarizationOn-demand retrievalBounded stateDynamic memory-retrieval budgetingTruncationHierarchical compressionImportance-weighted retention
- Quality attributes
- Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)Cost efficiencyFunctional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)Performance efficiency (ISO/IEC 25010)
- Risks mitigated
- Context window exhaustionLoss of original task specificationRepetitive actions from forgotten historyLost in the middleContext pollution between agentsContext bloatUnbounded state growthContext window overflowOut-of-memory errorsLatency degradation over long conversationsSilent context truncationLost-in-the-middle effectContext accumulation across multi-turn conversationsLoss of early-conversation constraints
Sources
- Ch1.1A: T. Nguyen, "Designing User Interfaces for Intuitive Human-Agent Interaction," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.1A. ISBN: 9798244538229.
- Ch1.2: T. Nguyen, "Core Agent Patterns," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.2. ISBN: 9798244538229.
- Ch1.4: T. Nguyen, "Memory and Perception Systems," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.4. ISBN: 9798244538229.
- Ch1.6: T. Nguyen, "Stateful Orchestration - Pitfalls, Integration, and Synthesis," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.6. ISBN: 9798244538229.
- Ch3.6: T. Nguyen, "Trace Analysis and Execution Debugging," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.6. ISBN: 9798244538229.
- Ch3.10: T. Nguyen, "Efficiency Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.10. ISBN: 9798244538229.
- Ch5.7: T. Nguyen, "Episodic Memory," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.7. ISBN: 9798244538229.
- Ch5.9: T. Nguyen, "Working Memory," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.9. ISBN: 9798244538229.
- Ch8.3: T. Nguyen, "Token Economics and Architecture," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 8.3. ISBN: 9798244538229.
- Ref5.04: C. Stryker, "What is AI agent memory?," IBM Think. Accessed: Sep. 27, 2026. [Online]. Available: https://www.ibm.com/think/topics/ai-agent-memory