Experience · Software component
Streaming Speech Recognizer
Software componentExperienceExperience & Human Oversightarc:StreamingSpeechRecognizer
A speech recognizer that processes live audio in short chunks and returns provisional partial transcripts immediately, marking segments final once enough context has arrived.
Responsibility. Delivers low-latency incremental transcripts of live speech.
Also known as: Streaming ASR, Real-time recognition
Variant of Automatic Speech Recognizer abstract
When to choose. Choose when real-time interaction is required (live customer-service calls, voice assistants, live accessibility captioning) and latency matters more than perfect accuracy.
Relationships
sends data to dynamic
alternative to variability
Design guidance
- MUST treat intermediate transcripts as provisional; consumers SHOULD either wait for finalized segments or process incrementally with correction.
Quantitative guidance
As stated by the sources; verify before use.
- Processes audio in 80-160ms chunks (Ch7.5).
Classification
- Patterns
- Incremental partial results with is_final flagEarly-context processing
- Technologies
- NVIDIA Riva ASR (streaming mode)
- Quality attributes
- Performance efficiency (ISO/IEC 25010)
- Risks mitigated
- Waiting for silence before agent processing begins
Sources
- Ch7.5: T. Nguyen, "NeMo Curator, Riva Speech AI & Multimodal Integration," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.5. ISBN: 9798244538229.