Infrastructure · Infrastructure resource
GPU Node
Infrastructure resourceInfrastructureInfrastructureVariation point (abstract)arc:GPUNode
An accelerated compute host (single or multi-GPU, linked by high-bandwidth interconnect) on which inference and agent workloads run.
Responsibility. Provides accelerated compute capacity for agent and inference workloads.
Also known as: GPU instance, Compute instance, A100 80GB node, GPU-labelled Kubernetes node
Variants
| Variant | When to choose |
|---|---|
| Edge GPU Device | Choose for robotics and autonomous systems needing programmable acceleration across many model types, accepting higher power draw. |
| On-Demand GPU Node | Choose for critical real-time production workloads requiring guaranteed availability. |
| Reserved GPU Node | Choose for the predictable, continuous baseline load present even in low-traffic periods. |
| Spot GPU Node | Choose for batch, development, and embarrassingly parallel background tasks that tolerate interruption. |
Relationships
hosts structural
- Accelerator Container Runtime Ref4.07
- GPU Device Plugin Ch4.5 Ref4.07
- Agent Controller abstract Ch4.7
- Alignment-Score Fact Checker Ch7.1A
- Answer Synthesizer abstract Ch4.2
- Attention Kernel abstract Ch4.4
- Data Curator Ch7.1A
- Distributed Model Executor abstract Ch7.1A
- Distributed Vector Index Store Ch4.1 Ch6.2A
- Engine Builder Ch4.5 Ch4.6
- Full-Parameter Fine-Tuner Ch3.5 Ref7.06
- GPU-Accelerated Dataframe Engine Ch6.3B Ch7.5
- GPU-Accelerated Vector Index Store Ch2.7 Ref2.07
- GPU Partition Manager Ch7.6
- GPU Telemetry Exporter Ch4.4 Ch4.5 +3
- Inference Server Ch2.7 Ch4.2 +10
- KV Cache Store Ch4.4
- LLM Generation Backend Ref7.15 Ref7.17
- LLM Inference Service Ch2.6 Ch4.5 +3
- Large Language Model Tier Ch1.8 Ch4.7
- LoRA Fine-Tuner abstract Ch3.5 Ref7.06
- Node Capability Labeler Ref4.07
- Optimized Inference Engine abstract Ch7.4
- Perception Interpreter Ch5.13
- Policy/Value Network Trainer Ch5.5
- Small Language Model Tier Ch4.7
- Text Embedding Service Ch6.3B Ch6.5
- Worker Agent abstract Ch4.3
is scaled by control
is evaluated by assurance
is monitored by assurance
Design guidance
- SHOULD be sized by profiling CPU, memory, GPU memory and request rates to find the actual bottleneck.
- SHOULD request GPUs explicitly for inference workloads and rely on CPU for lightweight routing and orchestration agents.
- SHOULD label nodes by GPU product and count to enable GPU-aware scheduling.
- SHOULD be right-sized to task requirements (GPU, vCPU, memory, network bandwidth, storage) rather than running lightweight tasks on top-tier hardware.
- SHOULD dedicate GPUs (no sharing) with fixed resource allocation where consistent performance and SLA guarantees are required.
- SHOULD scale vertically (more or faster GPUs) for compute-bound and memory-bound bottlenecks.
Quantitative guidance
As stated by the sources; verify before use.
- L4 $0.60/h vs. A100 $2.50/h (Ch1.8).
- Largest instances offer 8 x A100 with 640 GB combined memory (Ch1.8).
- Replacing a failed single instance takes 3-5 min of model loading (Ch1.8).
- A single A100 with 40GB VRAM is described as serving a 70B LLM, vision inference and vector search concurrently through careful memory partitioning (Ch2.7).
- A100 80GB for 70B-parameter models vs T4 16GB for 7B models; 60-70% of agent tasks are simple enough for lower-tier hardware (Ch4.7).
- Single-request inference uses only ~15-20% of GPU compute capacity (Ch4.7).
- Vertical scaling with tensor parallelism is cited as near-linear (8x with 8 GPUs) (Ref7.17).
- Indicative per-GPU cost: H100 $2-3/hour, A100 $1-2/hour, L40S $1.50-2/hour (Ref7.17).
- A100 GPUs cost $2-4/hour in cloud; teams often run at 40-50% utilization for headroom; sustained ~45% with low memory suggests consolidation (potentially halving cost), sustained 95%+ warrants proactive scaling (Ch8.1).
Classification
- Patterns
- Right-sizingVertical scaling via multi-GPU tensor parallelismHardware right-sizingGPU time-slicingMulti-GPU per container
- Technologies
- NVIDIA L4NVIDIA A100NVLinkNVSwitchCUDANVIDIA A100 40GBNVIDIA DGXRTX workstationsNVIDIA DRIVENVIDIA JetsonA100-80GBH100NVIDIA A100 80GBNVIDIA T4 16GB
- Quality attributes
- Performance efficiency (ISO/IEC 25010)Cost efficiency
Sources
- Ch1.8: T. Nguyen, "Scalability and Production Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.8. ISBN: 9798244538229.
- Ch2.6: T. Nguyen, "Tool Integration and Function Calling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.6. ISBN: 9798244538229.
- Ch2.7: T. Nguyen, "Multimodal RAG Approaches," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.7. ISBN: 9798244538229.
- Ch4.2: T. Nguyen, "Deployment and Scaling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.2. ISBN: 9798244538229.
- Ch4.3: T. Nguyen, "Container Orchestration and Edge Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.3. ISBN: 9798244538229.
- Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
- Ch4.5: T. Nguyen, "NVIDIA NIM and Triton Inference Server," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.5. ISBN: 9798244538229.
- Ch4.7: T. Nguyen, "Scaling Strategies," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.7. ISBN: 9798244538229.
- Ch5.5: T. Nguyen, "Monte Carlo Tree Search Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.5. ISBN: 9798244538229.
- Ch5.13: T. Nguyen, "Hybrid Decision Systems Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.13. ISBN: 9798244538229.
- Ch6.1: T. Nguyen, "Embeddings and RAG Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.1. ISBN: 9798244538229.
- Ch6.3B: T. Nguyen, "ETL Worked Example - Load Phase & Pipeline Integration," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.3B. ISBN: 9798244538229.
- Ch6.5: T. Nguyen, "Production RAG Systems," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.5. ISBN: 9798244538229.
- Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.
- Ch7.1B: T. Nguyen, "Nvidia NIM and Colang," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1B. ISBN: 9798244538229.
- Ch7.2: T. Nguyen, "Performance Optimization and Production Monitoring," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.2. ISBN: 9798244538229.
- Ch7.5: T. Nguyen, "NeMo Curator, Riva Speech AI & Multimodal Integration," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.5. ISBN: 9798244538229.
- Ch8.1: T. Nguyen, "Latency Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 8.1. ISBN: 9798244538229.
- Ref4.07: NVIDIA, "About the NVIDIA GPU Operator," NVIDIA GPU Operator Documentation. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/latest/index.html
- Ref7.04: NVIDIA, "NVIDIA NIM," NVIDIA Docs. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/nim/
- Ref7.12: "Advanced Nemotron Deployment Patterns," unpublished reference note (12-Nemotron-Advanced-Deployment.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref7.15: "Advanced Agentic AI Optimization Techniques," unpublished reference note (15-Advanced-Agentic-Optimization.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref7.16: "Production Monitoring and Operations for Agentic AI," unpublished reference note (16-Production-Monitoring-Operations.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref7.17: "Scaling Agentic AI Systems: Patterns and Strategies," unpublished reference note (17-Scalability-Patterns.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref8.07: "Agent Health Checks and Diagnostics," unpublished reference note (07-Agent-Health-Checks-Diagnostics.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note