Model Serving · Data artifact
Throughput-Optimized Deployment Profile
Data artifactModel ServingModelsarc:ThroughputOptimizedDeploymentProfile
A model deployment profile that maximises sustained queries per second, e.g., via FP8 precision and multi-GPU tensor parallelism, at the cost of higher per-request latency.
Responsibility. Configures an inference microservice for maximum throughput on the available GPUs.
Also known as: vllm-fp8-throughput-optimized, High-throughput profile, Multi-GPU tensor-parallel profile
Variant of Model Deployment Profile abstract
When to choose. Choose when maximising queries per second matters more than minimum time-to-first-token and multiple GPUs are available.
Relationships
alternative to variability
Quantitative guidance
As stated by the sources; verify before use.
- FP8 throughput profile gives ~2x throughput with minimal accuracy loss (Ch4.5).
Sources
- Ch4.5: T. Nguyen, "NVIDIA NIM and Triton Inference Server," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.5. ISBN: 9798244538229.