Memory · Software component
Summarizing History Compressor
Software componentMemoryCognition & Memoryarc:SummarizingHistoryCompressor
A context window manager that summarises older turns into compact goal, fact, and decision summaries kept in state before pruning the full messages.
Responsibility. Replaces pruned history with compact summaries.
Also known as: Truncation with summarisation, ConversationSummaryMemory, Prompt compression, Conversation summarization, Progressive summarization, Incremental summarization, Periodic history summarization
Variant of Context Window Manager abstract
When to choose. Choose when early context (user goals, key facts, binding decisions) remains relevant later and naive truncation would lose it.
Relationships
invokes dependency
reads dependency
- Conversation State Store abstract Ch3.5
writes dependency
- Conversation State Store abstract Ch3.10
is triggered by dynamic
is evaluated by assurance
alternative to variability
Design guidance
- SHOULD summarize progressively every N turns rather than retrospectively at capacity crisis, yielding higher-quality summaries and evenly distributed compute.
- SHOULD calibrate summarization frequency: too frequent wastes tokens on compression overhead, too infrequent fails to prevent overflow.
Quantitative guidance
As stated by the sources; verify before use.
- A 10,000-token conversation compressed to a ~500-token summary, freeing ~9,500 tokens (Ch1.6).
- 20-turn history (~10K tokens) compressed to last 3 exchanges (2K) plus summary (1K), a 70% reduction; compression maintains 85-95% of task accuracy (Ch3.4).
- Compressing a ~10K-token 20-turn history to the last 3 exchanges verbatim (~2K) plus a ~1K summary cut context ~70% (Ch3.5).
- Careful compression retains 85-95% of task accuracy (Ch3.5).
- Replacing ~2,000 tokens of verbatim history with ~400-token summaries (Ch3.10).
- E-commerce: summarising messages beyond 5 turns cut history tokens 65% (Ch3.10).
- Summarization reduces token consumption by ~40-60% but introduces hallucination risk (Ch5.9).
Classification
- Patterns
- Summarise-then-pruneRecent-turns-verbatim plus summary of earlier turnsIncremental summarization every N (typically 5-10) turns
- Technologies
- LangChain
- Quality attributes
- Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)Performance efficiency (ISO/IEC 25010)
- Risks mitigated
- Loss of critical early context from truncationExplosive context growth in long conversationsExpensive crisis-point compression
Sources
- Ch1.6: T. Nguyen, "Stateful Orchestration - Pitfalls, Integration, and Synthesis," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.6. ISBN: 9798244538229.
- Ch2.3: T. Nguyen, "LangChain Sequential Frameworks," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.3. ISBN: 9798244538229.
- Ch3.4: T. Nguyen, "Tuning Model Parameters for Production Performance," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.4. ISBN: 9798244538229.
- Ch3.5: T. Nguyen, "Prompt Optimization, Few-Shot Learning, Fine-Tuning," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.5. ISBN: 9798244538229.
- Ch3.10: T. Nguyen, "Efficiency Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.10. ISBN: 9798244538229.
- Ch5.9: T. Nguyen, "Working Memory," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.9. ISBN: 9798244538229.