Infrastructure · Infrastructure resource
On-Demand GPU Node
Infrastructure resourceInfrastructureInfrastructurearc:OnDemandGPUNode
A compute instance with guaranteed availability billed at on-demand rates.
Responsibility. Provides guaranteed-availability compute for critical workloads.
Also known as: On-demand instance
Variant of GPU Node abstract
When to choose. Choose for critical real-time production workloads requiring guaranteed availability.
Relationships
hosts structural
is scaled by control
alternative to variability
Design guidance
- SHOULD host real-time, user-facing production workloads with strict uptime requirements.
Quantitative guidance
As stated by the sources; verify before use.
- Tiered example: 100 req/s capacity on on-demand as spot-replacement buffer (no discount) (Ch4.7).
Classification
- Quality attributes
- Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)
- Risks mitigated
- Capacity interruption
Sources
- Ch1.8: T. Nguyen, "Scalability and Production Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.8. ISBN: 9798244538229.
- Ch4.7: T. Nguyen, "Scaling Strategies," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.7. ISBN: 9798244538229.
- Ref8.05: "Cost Optimization and Resource Monitoring for Agent Systems," unpublished reference note (05-Cost-Optimization-Resource-Monitoring.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note