Experience · Software component

Streaming Speech Recognizer

Software componentExperienceExperience & Human Oversightarc:StreamingSpeechRecognizer

A speech recognizer that processes live audio in short chunks and returns provisional partial transcripts immediately, marking segments final once enough context has arrived.

Responsibility. Delivers low-latency incremental transcripts of live speech.

Also known as: Streaming ASR, Real-time recognition

Variant of Automatic Speech Recognizer abstract

When to choose. Choose when real-time interaction is required (live customer-service calls, voice assistants, live accessibility captioning) and latency matters more than perfect accuracy.

specializesis target of alternativeTosends data toAutomatic Speech Recognizer: specializesAutomatic Speech Recogni…Speech Transcriber: is target of alternativeToSpeech TranscriberVoice Turn Coordinator: sends data toVoice Turn Coordinator
Direct neighbourhood (hover for relationship types)

Relationships

sends data to dynamic

alternative to variability

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
Incremental partial results with is_final flagEarly-context processing
Technologies
NVIDIA Riva ASR (streaming mode)
Quality attributes
Performance efficiency (ISO/IEC 25010)
Risks mitigated
Waiting for silence before agent processing begins

Sources

  1. Ch7.5: T. Nguyen, "NeMo Curator, Riva Speech AI & Multimodal Integration," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.5. ISBN: 9798244538229.