Model Serving · Data artifact

Throughput-Optimized Deployment Profile

Data artifactModel ServingModelsarc:ThroughputOptimizedDeploymentProfile

A model deployment profile that maximises sustained queries per second, e.g., via FP8 precision and multi-GPU tensor parallelism, at the cost of higher per-request latency.

Responsibility. Configures an inference microservice for maximum throughput on the available GPUs.

Also known as: vllm-fp8-throughput-optimized, High-throughput profile, Multi-GPU tensor-parallel profile

Variant of Model Deployment Profile abstract

When to choose. Choose when maximising queries per second matters more than minimum time-to-first-token and multiple GPUs are available.

specializesis target of alternativeTois target of alternativeToModel Deployment Profile: specializesModel Deployment ProfileLatency-Optimized Deployment Profile: is target of alternativeToLatency-Optimized Deploy…Memory-Optimized Deployment Profile: is target of alternativeToMemory-Optimized Deploym…
Direct neighbourhood (hover for relationship types)

Relationships

alternative to variability

Quantitative guidance

As stated by the sources; verify before use.

Sources

  1. Ch4.5: T. Nguyen, "NVIDIA NIM and Triton Inference Server," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.5. ISBN: 9798244538229.