Observability & Evaluation · Software component
Inference Engine Profiler
Software componentObservability & EvaluationObservability & Evaluationarc:InferenceEngineProfiler
A profiler that breaks down LLM inference-engine execution by operation class, such as attention kernels, KV-cache access and quantized layers, to locate model-level bottlenecks.
Responsibility. Attributes LLM inference time to engine operation classes.
Also known as: Inference-specific profiler, LLM engine profiler, Per-layer engine profile (dumpProfile)
Variant of Performance Profiler abstract
When to choose. Choose when system profiling shows inference dominates and the time split across attention kernels, KV-cache access, quantized operations and batching must be understood.
Relationships
monitors assurance
alternative to variability
Quantitative guidance
As stated by the sources; verify before use.
- Legal-contract workload: attention computation accounted for 77.2% of inference latency (Ch4.4).
- Attention typically consumes 60-80% of inference time for typical sequence lengths (Ch4.4).
Classification
- Technologies
- TensorRT-LLM profilingtrtexec --dumpProfileNsight Deep Learning Designer
- Quality attributes
- Maintainability (ISO/IEC 25010)Performance efficiency (ISO/IEC 25010)
Sources
- Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
- Ch7.4: T. Nguyen, "TensorRT-LLM Fundamentals and Quantization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.4. ISBN: 9798244538229.
- Ref7.01: NVIDIA, "Best practices," NVIDIA TensorRT Documentation. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/deeplearning/tensorrt/latest/performance/best-practices.html