Knowledge & Data · Software component

Speech Transcriber

Software componentKnowledge & DataKnowledge & Dataarc:SpeechTranscriber

An automatic speech recognition component that converts recorded audio into transcripts segmented at sentence level with start and end timestamps.

Responsibility. Transcribes audio into timestamped text segments for indexing.

Also known as: Automatic speech recognition (ASR), Audio preprocessing, Speech-to-text, Offline ASR, Batch speech recognition

Variant of Automatic Speech Recognizer abstract

When to choose. Choose for post-hoc analysis (call-center QA, meeting summaries, voicemail), maximum accuracy (legal, medical), speaker diarization or overlapping multi-speaker audio, where higher latency is acceptable.

is orchestrated byspecializesis routed to bysends data toreadsalternative toproduceshostsinvokesinvokesIngestion Pipeline Orchestrator: is orchestrated byIngestion Pipeline Orche…Automatic Speech Recognizer: specializesAutomatic Speech Recogni…Multimodal Content Router: is routed to byMultimodal Content RouterTime-Indexed Transcript Chunker: sends data toTime-Indexed Transcript …Source Media Store: readsSource Media StoreStreaming Speech Recognizer: alternative toStreaming Speech Recogni…Raw Document Corpus: producesRaw Document CorpusSpeech Recognition Model: hostsSpeech Recognition ModelPunctuation and Capitalization Restorer: invokesPunctuation and Capitali…Speaker Diarizer: invokesSpeaker Diarizer
Direct neighbourhood (hover for relationship types)

Relationships

hosts structural

invokes dependency

reads dependency

is routed to by dynamic

sends data to dynamic

is orchestrated by control

produces lifecycle

alternative to variability

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
Ground-to-text multimodal RAGMulti-pass transcription (raw, punctuation, diarization, terminology correction)
Technologies
OpenAI WhisperNVIDIA Riva ASR (offline mode)
Quality attributes
Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)Maintainability (ISO/IEC 25010)
Risks mitigated
Knowledge in recorded conversations inaccessible to text-only RAG

Sources

  1. Ch2.7: T. Nguyen, "Multimodal RAG Approaches," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.7. ISBN: 9798244538229.
  2. Ch7.5: T. Nguyen, "NeMo Curator, Riva Speech AI & Multimodal Integration," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.5. ISBN: 9798244538229.