Model Serving · Data artifact

Latency-Optimized Deployment Profile

Data artifactModel ServingModelsarc:LatencyOptimizedDeploymentProfile

A model deployment profile that minimises time-to-first-token, e.g., FP16 precision on a single GPU, sacrificing throughput.

Responsibility. Configures an inference microservice for fastest initial response.

Also known as: vllm-fp16-latency-optimized, Low-latency profile, Single-GPU profile

Variant of Model Deployment Profile abstract

When to choose. Choose for latency-critical interactive applications where time-to-first-token SLOs dominate.

specializesalternative toalternative toModel Deployment Profile: specializesModel Deployment ProfileMemory-Optimized Deployment Profile: alternative toMemory-Optimized Deploym…Throughput-Optimized Deployment Profile: alternative toThroughput-Optimized Dep…
Direct neighbourhood (hover for relationship types)

Relationships

alternative to variability

Sources

  1. Ch4.5: T. Nguyen, "NVIDIA NIM and Triton Inference Server," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.5. ISBN: 9798244538229.