Orchestration · Software component
Voice Turn Coordinator
Software componentOrchestrationOrchestration & Toolsarc:VoiceTurnCoordinator
A runtime component that sequences one spoken conversation turn: streams user audio to speech recognition, hands transcripts to the agent, and passes the agent's reply to speech synthesis.
Responsibility. Coordinates the ASR-agent-TTS turn pipeline within the conversational latency budget.
Also known as: Voice agent pipeline, Voice interaction orchestrator
Relationships
invokes dependency
- Agent Controller abstract Ch7.5
- Automatic Speech Recognizer abstract Ch7.5
- Speech Synthesizer Ch7.5
receives data from dynamic
Design guidance
- SHOULD target end-to-end voice turn latency near human turn-taking (~200-400ms) and avoid >1 second.
- SHOULD start agent context retrieval on partial transcripts while the user is still speaking.
Quantitative guidance
As stated by the sources; verify before use.
- Users expect <300ms end-to-end; a simple query totals ~300-900ms across ASR, LLM (100-500ms) and TTS (100-200ms) (Ch7.5).
Classification
- Patterns
- User speaks -> ASR -> LLM agent -> TTS -> user hearsProcess-before-utterance-ends with provisional transcripts
- Technologies
- NVIDIA Riva
- Quality attributes
- Performance efficiency (ISO/IEC 25010)
- Risks mitigated
- End-to-end latency >1s making voice agents feel sluggish
Sources
- Ch7.5: T. Nguyen, "NeMo Curator, Riva Speech AI & Multimodal Integration," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.5. ISBN: 9798244538229.