Experience · Software component

Speech Synthesizer

Software componentExperienceExperience & Human Oversightarc:SpeechSynthesizer

A text-to-speech service that converts an agent's text response into natural-sounding audio with controllable voice, pitch, rate and prosody.

Responsibility. Renders agent text output as spoken audio.

Also known as: TTS service, Text-to-speech, Voice output modality handler, Synthetic speech responses

deployed onis invoked bysends data tois invoked byhostsis configured byInference Server: deployed onInference ServerConversational (Chat) Interface: is invoked byConversational (Chat) In…Response Streamer: sends data toResponse StreamerVoice Turn Coordinator: is invoked byVoice Turn CoordinatorSpeech Synthesis Model: hostsSpeech Synthesis ModelVoice Persona Profile: is configured byVoice Persona Profile
Direct neighbourhood (hover for relationship types)

Relationships

deployed on structural

hosts structural

is configured by structural

is invoked by dependency

sends data to dynamic

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
Streaming synthesis (chunked audio output)Sentence-level parallel batchingMid-utterance multi-speaker voice switching
Technologies
NVIDIA Riva TTS
Quality attributes
Performance efficiency (ISO/IEC 25010)Interaction capability (ISO/IEC 25010)
Risks mitigated
Sluggish voice responses that feel unnatural

Sources

  1. Ch7.5: T. Nguyen, "NeMo Curator, Riva Speech AI & Multimodal Integration," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.5. ISBN: 9798244538229.
  2. Ch10.1: T. Nguyen, "Conversational UI," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.1. ISBN: 9798244538229.