Knowledge & Data · Model asset

Speech Recognition Model

Model assetKnowledge & DataKnowledge & Dataarc:SpeechRecognitionModel

A transformer ASR model that encodes audio spectrograms and decodes text with word- or sentence-level timestamp alignment, offered in size tiers trading accuracy for latency.

Responsibility. Converts speech audio to timestamped text.

Also known as: Acoustic model

deployed ondeployed onAutomatic Speech Recognizer: deployed onAutomatic Speech Recogni…Speech Transcriber: deployed onSpeech Transcriber
Direct neighbourhood (hover for relationship types)

Relationships

deployed on structural

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Technologies
OpenAI WhisperNVIDIA Riva ASR
Quality attributes
Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)

Sources

  1. Ch2.7: T. Nguyen, "Multimodal RAG Approaches," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.7. ISBN: 9798244538229.
  2. Ch7.5: T. Nguyen, "NeMo Curator, Riva Speech AI & Multimodal Integration," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.5. ISBN: 9798244538229.