Knowledge & Data · Software component
Time-Indexed Transcript Chunker
Software componentKnowledge & DataKnowledge & Dataarc:TimeIndexedTranscriptChunker
A document chunker that groups timestamped transcript segments into duration-bounded windows ending at sentence boundaries, carrying start/end timestamps and source metadata.
Responsibility. Creates temporally anchored transcript chunks for retrieval.
Also known as: Time-indexed chunking, Duration-based chunking
Variant of Document Chunker abstract
When to choose. Choose for audio transcripts lacking paragraph breaks or headers, when retrieved segments must link back to exact moments in the recording.
Relationships
is configured by structural
receives data from dynamic
sends data to dynamic
alternative to variability
Design guidance
- SHOULD finalise each chunk at the last complete sentence once the target duration is exceeded, never truncating mid-sentence.
- SHOULD retain start/end timestamps, audio source path, language, speaker diarization and transcription confidence as chunk metadata.
- SHOULD keep the final shorter chunk to avoid losing the end of the recording.
Quantitative guidance
As stated by the sources; verify before use.
- Target window 3-5 minutes (180-300 s); 5-minute windows balanced context and precision for technical meetings (Ch2.7).
- A 60-minute earnings call produces 8,000-12,000 words of continuous text (Ch2.7).
Classification
- Patterns
- Time-anchored retrievalSentence-boundary alignment
- Quality attributes
- Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)Transparency and accountability (NIST AI RMF: accountable and transparent)
- Risks mitigated
- Mid-sentence truncationLoss of temporal context
Sources
- Ch2.7: T. Nguyen, "Multimodal RAG Approaches," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.7. ISBN: 9798244538229.