Model Serving · Software component
Attention Kernel
Software componentModel ServingModelsVariation point (abstract)arc:AttentionKernel
An abstract accelerator kernel implementation that computes transformer self-attention over the current sequence and cached key-value projections.
Responsibility. Computes self-attention during model inference.
Variants
| Variant | When to choose |
|---|---|
| Fused Block-wise Attention Kernel | Choose by default for any transformer deployment, especially long contexts on memory-bandwidth-limited GPUs; benchmark against standard attention for very short sequences. |
| Sparse Attention Kernel | Choose when much larger effective contexts are needed and occasionally missing long-range dependencies that full attention would capture is acceptable. |
| Standard Attention Kernel | Choose only when benchmarks show fused kernels are slower, e.g., very short sequences (under 256 tokens) with small batches. |
Relationships
deployed on structural
is configured by structural
is invoked by dependency
reads dependency
is monitored by assurance
Quantitative guidance
As stated by the sources; verify before use.
- Attention consumes 60-80% of inference time; specialized kernels give 2-4x throughput over naive implementations (Ch4.4).
Classification
- Quality attributes
- Performance efficiency (ISO/IEC 25010)
Sources
- Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.