Knowledge & Data · Software component
Context Compressor
Software componentKnowledge & DataKnowledge & Dataarc:ContextCompressor
A prompt-preparation component that shrinks retrieved context by summarizing documents, extracting bullet points, and truncating to the most relevant passages before generation.
Responsibility. Reduces prompt length to lower model prompt-encoding latency.
Also known as: Prompt compressor, Prompt Compressor, Relevance ranking and progressive summarization, Retrieval compression, Key-fact extraction, Extractive reading, Information density maximization
Relationships
is invoked by dependency
- Context Assembler abstract Ch2.9
- Context Window Manager abstract Ch3.6
- Memory Retriever Ch5.9
is triggered by dynamic
receives data from dynamic
sends data to dynamic
is evaluated by assurance
Design guidance
- SHOULD compress prompts in RAG scenarios where retrieved context pushes prompts beyond ~4,000 tokens.
- SHOULD compress less-critical details while preserving key facts rather than blindly truncating.
- SHOULD favour aggressive fact extraction or structured encoding for lookup-intensive tasks and summarization or preserved prose for nuanced analysis.
Quantitative guidance
As stated by the sources; verify before use.
- Prompt encoding latency becomes significant above 4,000 tokens (Ch2.9).
- Compressing prompts from 6,000 to 2,000 tokens cuts encoding time by 60-70% (Ch2.9).
- Doubling context length often increases latency 2.5-3x; 50K-token contexts take 8-12s vs 2-3s at 5K tokens (Ch3.4).
- Keeping only the most relevant documents saves ~30-50%, summarization ~40-60%, key-fact extraction ~50-70% of tokens (Ch5.9).
- Research example: 25 papers (~250,000 tokens) reduced to 60,000 tokens by extracting abstract, introduction and conclusion from the top 8 (Ch5.9).
- Structured JSON encoding of product specifications ~12 vs ~18 tokens (~33% compression) (Ch5.9).
Classification
- Patterns
- Prompt compressionPassage truncationRelevance-score filtering of retrieved documentsTool description condensationKey-fact extractionDocument summarizationSection-targeted extraction (abstract, introduction, conclusion)Structured encoding of prose into JSON or fact listsCompression of retrieved episodes before injection
- Quality attributes
- Performance efficiency (ISO/IEC 25010)
- Risks mitigated
- Context saturation from verbose retrieved content
Sources
- Ch2.9: T. Nguyen, "Streaming and Real-Time Responses," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.9. ISBN: 9798244538229.
- Ch3.4: T. Nguyen, "Tuning Model Parameters for Production Performance," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.4. ISBN: 9798244538229.
- Ch3.6: T. Nguyen, "Trace Analysis and Execution Debugging," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.6. ISBN: 9798244538229.
- Ch3.10: T. Nguyen, "Efficiency Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.10. ISBN: 9798244538229.
- Ch5.9: T. Nguyen, "Working Memory," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.9. ISBN: 9798244538229.