Knowledge & Data · Software component
Speech Transcriber
Software componentKnowledge & DataKnowledge & Dataarc:SpeechTranscriber
An automatic speech recognition component that converts recorded audio into transcripts segmented at sentence level with start and end timestamps.
Responsibility. Transcribes audio into timestamped text segments for indexing.
Also known as: Automatic speech recognition (ASR), Audio preprocessing, Speech-to-text, Offline ASR, Batch speech recognition
Variant of Automatic Speech Recognizer abstract
When to choose. Choose for post-hoc analysis (call-center QA, meeting summaries, voicemail), maximum accuracy (legal, medical), speaker diarization or overlapping multi-speaker audio, where higher latency is acceptable.
Relationships
hosts structural
invokes dependency
reads dependency
is routed to by dynamic
sends data to dynamic
is orchestrated by control
produces lifecycle
alternative to variability
Design guidance
- SHOULD NOT process raw audio when professional transcripts with speaker labels and timestamps already exist.
- SHOULD be reserved for cases where speech content adds value beyond written summaries.
Quantitative guidance
As stated by the sources; verify before use.
- Transcription runs at 0.3-0.5x real time on GPUs (Ch2.7).
- A 60-minute meeting yields 400-600 segments averaging 6-9 seconds (Ch2.7).
- Offline recognition is typically 1-2% better WER than streaming; a 30-minute meeting takes 2-3 minutes to process (Ch7.5).
Classification
- Patterns
- Ground-to-text multimodal RAGMulti-pass transcription (raw, punctuation, diarization, terminology correction)
- Technologies
- OpenAI WhisperNVIDIA Riva ASR (offline mode)
- Quality attributes
- Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)Maintainability (ISO/IEC 25010)
- Risks mitigated
- Knowledge in recorded conversations inaccessible to text-only RAG
Sources
- Ch2.7: T. Nguyen, "Multimodal RAG Approaches," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.7. ISBN: 9798244538229.
- Ch7.5: T. Nguyen, "NeMo Curator, Riva Speech AI & Multimodal Integration," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.5. ISBN: 9798244538229.