Model Serving · Software component

Fused Block-wise Attention Kernel

Software componentModel ServingModelsarc:FusedBlockwiseAttentionKernel

An exact attention kernel that computes attention in fused blocks without materializing the score matrix, cutting memory traffic while producing identical results.

Responsibility. Computes exact attention block-wise with reduced memory traffic.

Also known as: Flash Attention, Memory-efficient attention, Fused multi-head attention (context FMHA), FlashAttention

Variant of Attention Kernel abstract

When to choose. Choose by default for any transformer deployment, especially long contexts on memory-bandwidth-limited GPUs; benchmark against standard attention for very short sequences.

deployed onspecializesreadsis configured byalternative toLLM Inference Service: deployed onLLM Inference ServiceAttention Kernel: specializesAttention KernelKV Cache Store: readsKV Cache StoreEngine Build Configuration: is configured byEngine Build ConfigurationStandard Attention Kernel: alternative toStandard Attention Kernel
Direct neighbourhood (hover for relationship types)

Relationships

deployed on structural

is configured by structural

reads dependency

alternative to variability

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Technologies
TensorRT-LLM GPT Attention PluginFlashAttentionTensorRT-LLMvLLM
Quality attributes
Performance efficiency (ISO/IEC 25010)Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)

Sources

  1. Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
  2. Ch7.4: T. Nguyen, "TensorRT-LLM Fundamentals and Quantization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.4. ISBN: 9798244538229.
  3. Ref7.05: S. Verma and N. Vaidya, "Mastering LLM Techniques: Inference Optimization," NVIDIA Technical Blog, Nov. 17, 2023. [Online]. Available: https://developer.nvidia.com/blog/mastering-llm-techniques-inference-optimization/
  4. Ref8.05: "Cost Optimization and Resource Monitoring for Agent Systems," unpublished reference note (05-Cost-Optimization-Resource-Monitoring.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note