Model Serving · Software component
Standard Attention Kernel
Software componentModel ServingModelsarc:StandardAttentionKernel
An attention kernel that materializes the full attention-score matrix in accelerator memory before applying it to values.
Responsibility. Computes attention by materializing the full score matrix.
Also known as: Dense attention, Full all-to-all attention
Variant of Attention Kernel abstract
When to choose. Choose only when benchmarks show fused kernels are slower, e.g., very short sequences (under 256 tokens) with small batches.
Relationships
alternative to variability
Quantitative guidance
As stated by the sources; verify before use.
- 4,096-token sequence needs a 16.8M-element score matrix (33.6 MB per head in FP16); 64 heads at batch 8 total ~17.2 GB (Ch4.4).
- Legal-contract baseline: 2.4 req/s at batch 6, 68.2 GB (Ch4.4).
- Dense attention scales O(N^2): 1k tokens -> 1M comparisons, 10k -> 100M, 100k -> 10B; a 100k-token FP32 attention matrix is ~40 GB vs 16-80 GB on most production GPUs (Ch5.9).
Sources
- Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
- Ch5.9: T. Nguyen, "Working Memory," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.9. ISBN: 9798244538229.