Knowledge & Data · Software component
Document Chunker
Software componentKnowledge & DataKnowledge & DataVariation point (abstract)arc:DocumentChunker
A software component that splits raw documents into chunks for embedding and entity extraction.
Responsibility. Splits raw documents into processable chunks.
Variants
| Variant | When to choose |
|---|---|
| Fixed-Length Chunker | Choose only when content has no usable structure or timing data; the chapter presents it as the naive baseline that fractures semantic units. |
| Hierarchical Chunker | Choose when document structure (sections, subsections) must be preserved so retrieved chunks keep their structural context, accepting more complex retrieval logic. |
| Overlapping Window Chunker | Choose when important context spans chunk or section boundaries and continuity must be preserved, accepting multiplied storage. |
| Semantic Boundary Chunker | Choose for text documents with structure, splitting at section headers and paragraph breaks rather than arbitrary token counts. |
| Time-Indexed Transcript Chunker | Choose for audio transcripts lacking paragraph breaks or headers, when retrieved segments must link back to exact moments in the recording. |
| Topic-Shift Chunker | Choose when chunk coherence matters more than uniform chunk size. |
Relationships
is configured by structural
invokes dependency
is invoked by dependency
receives data from dynamic
sends data to dynamic
is orchestrated by control
Design guidance
- SHOULD produce 300-500 token segments rather than retrieving entire documents.
- SHOULD NOT split on fixed sizes alone; chunks beginning mid-context lose referents and degrade retrieval.
- SHOULD choose chunk size, overlap and boundary logic from document characteristics and retrieval requirements.
- SHOULD count tokens with the embedding model's tokenizer rather than character approximations, especially for code and non-English text.
- SHOULD discard trailing fragments below a minimum chunk size.
Classification
- Patterns
- Boundary seeking (paragraph, then sentence, then word)Minimum chunk size enforcementToken-accurate counting
- Risks mitigated
- Context loss from undersized chunksGeneric, unspecific embeddings from oversized chunks
Sources
- Ch1.7A: T. Nguyen, "Relational Reasoning with Knowledge Graphs - The Fundamentals, Integration, and Extraction," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.7A. ISBN: 9798244538229.
- Ch1.7B: T. Nguyen, "Relational Reasoning with Knowledge Graphs - Hybrid RAG+KG Integration," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.7B. ISBN: 9798244538229.
- Ch2.7: T. Nguyen, "Multimodal RAG Approaches," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.7. ISBN: 9798244538229.
- Ch3.10: T. Nguyen, "Efficiency Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.10. ISBN: 9798244538229.
- Ch5.8: T. Nguyen, "Semantic Memory," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.8. ISBN: 9798244538229.
- Ch6.2B: T. Nguyen, "Production Vector Database Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.2B. ISBN: 9798244538229.
- Ch6.3A: T. Nguyen, "ETL Pipeline Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.3A. ISBN: 9798244538229.
- Ch6.4: T. Nguyen, "Data Quality Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.4. ISBN: 9798244538229.
- Ch6.5: T. Nguyen, "Production RAG Systems," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.5. ISBN: 9798244538229.
- Ref2.07: NVIDIA Developer, "Building multimodal AI RAG with LlamaIndex, NVIDIA NIM, and Milvus | LLM app development," YouTube. Accessed: Sep. 26, 2026. [Online Video]. Available: https://www.youtube.com/watch?v=NaT5Eo97_I0