Infrastructure · Infrastructure resource

Spot GPU Node

Infrastructure resourceInfrastructureInfrastructurearc:SpotGPUNode

A discounted, interruptible compute instance reclaimable by the provider on short notice.

Responsibility. Provides low-cost interruptible compute for fault-tolerant work.

Also known as: Spot instance, Preemptible instance, Spot instances for burst capacity

Variant of GPU Node abstract

When to choose. Choose for batch, development, and embarrassingly parallel background tasks that tolerate interruption.

hostshostsspecializesis scaled byhostsis target of alternativeTohostsis target of alternativeTois monitored byLLM Inference Service: hostsLLM Inference ServiceEvaluation Harness: hostsEvaluation HarnessGPU Node: specializesGPU NodeMetric-Driven Autoscaler: is scaled byMetric-Driven AutoscalerDocument Ingestor: hostsDocument IngestorOn-Demand GPU Node: is target of alternativeToOn-Demand GPU NodeDeferred Batch Scheduler: hostsDeferred Batch SchedulerReserved GPU Node: is target of alternativeToReserved GPU NodeSpot Interruption Handler: is monitored bySpot Interruption Handler
Direct neighbourhood (hover for relationship types)

Relationships

hosts structural

is scaled by control

is monitored by assurance

alternative to variability

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
Checkpoint-and-resume
Quality attributes
Cost efficiency

Sources

  1. Ch1.8: T. Nguyen, "Scalability and Production Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.8. ISBN: 9798244538229.
  2. Ch4.7: T. Nguyen, "Scaling Strategies," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.7. ISBN: 9798244538229.
  3. Ref4.03: M. Zhang, J. Wyman, I. M. Bhosale, and W. Tan, "Scaling LLMs with NVIDIA Triton and NVIDIA TensorRT-LLM using Kubernetes," NVIDIA Technical Blog, Oct. 22, 2024. [Online]. Available: https://developer.nvidia.com/blog/scaling-llms-with-nvidia-triton-and-nvidia-tensorrt-llm-using-kubernetes/
  4. Ref7.15: "Advanced Agentic AI Optimization Techniques," unpublished reference note (15-Advanced-Agentic-Optimization.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
  5. Ref7.17: "Scaling Agentic AI Systems: Patterns and Strategies," unpublished reference note (17-Scalability-Patterns.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
  6. Ref8.05: "Cost Optimization and Resource Monitoring for Agent Systems," unpublished reference note (05-Cost-Optimization-Resource-Monitoring.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note