Knowledge & Data · Model asset
Speech Recognition Model
Model assetKnowledge & DataKnowledge & Dataarc:SpeechRecognitionModel
A transformer ASR model that encodes audio spectrograms and decodes text with word- or sentence-level timestamp alignment, offered in size tiers trading accuracy for latency.
Responsibility. Converts speech audio to timestamped text.
Also known as: Acoustic model
Relationships
deployed on structural
Design guidance
- SHOULD select base or small size tiers for production multimodal RAG to balance accuracy and latency.
Quantitative guidance
As stated by the sources; verify before use.
- Trained on 680,000 hours of multilingual audio; supports 99 languages; five sizes (Ch2.7).
- Tiny 39M parameters near real time on CPU; large 1.5B near-human quality at 0.5x real time on GPU (Ch2.7).
- Base/small: 90%+ accuracy on conversational audio; base (74M) transcribes a 60-minute meeting in under 2 minutes on one GPU (Ch2.7).
Classification
- Technologies
- OpenAI WhisperNVIDIA Riva ASR
- Quality attributes
- Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)
Sources
- Ch2.7: T. Nguyen, "Multimodal RAG Approaches," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.7. ISBN: 9798244538229.
- Ch7.5: T. Nguyen, "NeMo Curator, Riva Speech AI & Multimodal Integration," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.5. ISBN: 9798244538229.