Infrastructure · Infrastructure resource

On-Demand GPU Node

Infrastructure resourceInfrastructureInfrastructurearc:OnDemandGPUNode

A compute instance with guaranteed availability billed at on-demand rates.

Responsibility. Provides guaranteed-availability compute for critical workloads.

Also known as: On-demand instance

Variant of GPU Node abstract

When to choose. Choose for critical real-time production workloads requiring guaranteed availability.

hostsspecializesalternative toalternative tois scaled byLLM Inference Service: hostsLLM Inference ServiceGPU Node: specializesGPU NodeSpot GPU Node: alternative toSpot GPU NodeReserved GPU Node: alternative toReserved GPU NodeSpot Interruption Handler: is scaled bySpot Interruption Handler
Direct neighbourhood (hover for relationship types)

Relationships

hosts structural

is scaled by control

alternative to variability

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Quality attributes
Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)
Risks mitigated
Capacity interruption

Sources

  1. Ch1.8: T. Nguyen, "Scalability and Production Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.8. ISBN: 9798244538229.
  2. Ch4.7: T. Nguyen, "Scaling Strategies," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.7. ISBN: 9798244538229.
  3. Ref8.05: "Cost Optimization and Resource Monitoring for Agent Systems," unpublished reference note (05-Cost-Optimization-Resource-Monitoring.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note