Infrastructure · Infrastructure resource
Spot GPU Node
Infrastructure resourceInfrastructureInfrastructurearc:SpotGPUNode
A discounted, interruptible compute instance reclaimable by the provider on short notice.
Responsibility. Provides low-cost interruptible compute for fault-tolerant work.
Also known as: Spot instance, Preemptible instance, Spot instances for burst capacity
Variant of GPU Node abstract
When to choose. Choose for batch, development, and embarrassingly parallel background tasks that tolerate interruption.
Relationships
hosts structural
is scaled by control
is monitored by assurance
alternative to variability
Design guidance
- SHOULD be used only for workloads that can checkpoint progress and resume elsewhere.
- SHOULD be used only for non-latency-critical workloads to reduce cost.
- SHOULD be used for batch processing, non-critical services, development/testing and burst capacity (Ref8.05).
- SHOULD NOT host user-facing critical requests, long-running inference or time-sensitive operations (Ref8.05).
Quantitative guidance
As stated by the sources; verify before use.
- 60-80% cheaper than on-demand with 2-minute interruption notice (Ch1.8).
- Spot instances offer 70-90% discounts; tiered example puts 300 req/s of peak capacity on spot at 80% discount (Ch4.7).
- Spot $0.60/h vs on-demand $2/h per GPU (70% savings) (Ref8.05).
Classification
- Patterns
- Checkpoint-and-resume
- Quality attributes
- Cost efficiency
Sources
- Ch1.8: T. Nguyen, "Scalability and Production Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.8. ISBN: 9798244538229.
- Ch4.7: T. Nguyen, "Scaling Strategies," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.7. ISBN: 9798244538229.
- Ref4.03: M. Zhang, J. Wyman, I. M. Bhosale, and W. Tan, "Scaling LLMs with NVIDIA Triton and NVIDIA TensorRT-LLM using Kubernetes," NVIDIA Technical Blog, Oct. 22, 2024. [Online]. Available: https://developer.nvidia.com/blog/scaling-llms-with-nvidia-triton-and-nvidia-tensorrt-llm-using-kubernetes/
- Ref7.15: "Advanced Agentic AI Optimization Techniques," unpublished reference note (15-Advanced-Agentic-Optimization.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref7.17: "Scaling Agentic AI Systems: Patterns and Strategies," unpublished reference note (17-Scalability-Patterns.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref8.05: "Cost Optimization and Resource Monitoring for Agent Systems," unpublished reference note (05-Cost-Optimization-Resource-Monitoring.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note