Model Serving · Software component
Fused Block-wise Attention Kernel
Software componentModel ServingModelsarc:FusedBlockwiseAttentionKernel
An exact attention kernel that computes attention in fused blocks without materializing the score matrix, cutting memory traffic while producing identical results.
Responsibility. Computes exact attention block-wise with reduced memory traffic.
Also known as: Flash Attention, Memory-efficient attention, Fused multi-head attention (context FMHA), FlashAttention
Variant of Attention Kernel abstract
When to choose. Choose by default for any transformer deployment, especially long contexts on memory-bandwidth-limited GPUs; benchmark against standard attention for very short sequences.
Relationships
deployed on structural
is configured by structural
reads dependency
alternative to variability
Design guidance
- SHOULD be enabled first for long-context workloads since it has no accuracy trade-off and needs no code changes.
- MUST support block-indirected KV-cache layout when combined with paged allocation.
Quantitative guidance
As stated by the sources; verify before use.
- Reduces attention memory by 40-60% (Ch4.4).
- Speedup ~1.8x at 2,048 tokens, 2.4x at 8,192, 3.2x at 32,768 (Ch4.4).
- Legal-contract example: 2.4 -> 5.2 req/s (2.17x), zero accuracy impact (Ch4.4).
- 10-20% additional speedup on top of paged KV cache (Ch7.4); 2-4x memory-bandwidth reduction (Ref7.05).
- Attention optimization saves 10-20% memory (Ref8.05).
Classification
- Technologies
- TensorRT-LLM GPT Attention PluginFlashAttentionTensorRT-LLMvLLM
- Quality attributes
- Performance efficiency (ISO/IEC 25010)Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)
Sources
- Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
- Ch7.4: T. Nguyen, "TensorRT-LLM Fundamentals and Quantization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.4. ISBN: 9798244538229.
- Ref7.05: S. Verma and N. Vaidya, "Mastering LLM Techniques: Inference Optimization," NVIDIA Technical Blog, Nov. 17, 2023. [Online]. Available: https://developer.nvidia.com/blog/mastering-llm-techniques-inference-optimization/
- Ref8.05: "Cost Optimization and Resource Monitoring for Agent Systems," unpublished reference note (05-Cost-Optimization-Resource-Monitoring.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note