Experience · Software component
Speech Synthesizer
Software componentExperienceExperience & Human Oversightarc:SpeechSynthesizer
A text-to-speech service that converts an agent's text response into natural-sounding audio with controllable voice, pitch, rate and prosody.
Responsibility. Renders agent text output as spoken audio.
Also known as: TTS service, Text-to-speech, Voice output modality handler, Synthetic speech responses
Relationships
deployed on structural
hosts structural
is configured by structural
is invoked by dependency
sends data to dynamic
Design guidance
- SHOULD stream synthesized audio in chunks so playback begins after the first few words.
- SHOULD split long responses into sentences and synthesize them in parallel.
Quantitative guidance
As stated by the sources; verify before use.
- Typical utterance synthesis 100-200ms; users perceive <200ms as instant, 200-500ms acceptable, >500ms sluggish (Ch7.5).
Classification
- Patterns
- Streaming synthesis (chunked audio output)Sentence-level parallel batchingMid-utterance multi-speaker voice switching
- Technologies
- NVIDIA Riva TTS
- Quality attributes
- Performance efficiency (ISO/IEC 25010)Interaction capability (ISO/IEC 25010)
- Risks mitigated
- Sluggish voice responses that feel unnatural
Sources
- Ch7.5: T. Nguyen, "NeMo Curator, Riva Speech AI & Multimodal Integration," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.5. ISBN: 9798244538229.
- Ch10.1: T. Nguyen, "Conversational UI," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.1. ISBN: 9798244538229.