Observability & Evaluation · Software component
Execution Profiler
Software componentObservability & EvaluationObservability & Evaluationarc:ExecutionProfiler
An analysis component that aggregates execution timings hierarchically from workflow down to individual agents and tools to identify performance bottlenecks.
Responsibility. Locates latency bottlenecks across the workflow hierarchy.
Also known as: Workflow profiler, Latency Bottleneck Profiler, Tool-level profiling, Offline efficiency profiler, Workflow instrumentation, Trajectory analysis, Latency component decomposition, Workflow-level profiler, Agent workflow profiler, Per-step latency breakdown, Per-step execution latency analysis
Variant of Performance Profiler abstract
When to choose. Choose to decide whether multi-step agent latency is dominated by LLM inference or external tool latency, and whether multi-agent parallelism is effective or serialized.
Relationships
is invoked by dependency
reads dependency
writes dependency
emits telemetry to dynamic
receives data from dynamic
monitors assurance
- Agent Controller abstract Ch3.4 Ch7.3
- Multi-Agent Coordinator abstract Ch4.4
- ReAct Agent Controller Ch4.4
- Tool Executor Ch3.10 Ch4.4
- Tool Integration Adapter abstract Ref3.05 Ref3.07
- Workflow Orchestrator abstract Ch3.10 Ref1.01
produces lifecycle
alternative to variability
Design guidance
- MUST profile end-to-end latency before optimising and target the dominant bottleneck (model optimisation for inference-bound agents, orchestration for tool-bound agents).
- SHOULD break end-to-end latency into per-operation durations to target the dominant bottleneck.
- MAY wrap individual tools for profiling without modifying agents or workflows, and send profiling metrics to the existing observability platform.
- SHOULD instrument agents and tools to capture model name, input/output tokens per LLM call, tool timings, step count, reasoning trajectory and API call frequency with payload sizes.
- SHOULD decompose end-to-end latency into model inference, API overhead, tool execution and reasoning-step time and target the dominant components.
- SHOULD analyse collected trajectories for repeated reasoning, unused steps and context reconstruction.
- MUST establish per-task-category baselines of tokens, steps and API calls before optimising.
- SHOULD break end-to-end latency into LLM-call and tool-call shares before choosing inference or tool optimizations.
- MUST quantify both latency (ms) and cost (USD per transaction) per operation to enable prioritized optimization.
- SHOULD establish a baseline before any optimization and re-profile after each change to validate its measured impact.
- SHOULD report each step's share of end-to-end latency and its variance so optimization targets the dominant contributor rather than an assumed one.
Quantitative guidance
As stated by the sources; verify before use.
- Pure reasoning agents spend 80-90% of latency in inference; tool-heavy agents 40-60% waiting on tools (Ch3.4).
- Example: 8 s response vs 3 s SLO split into LLM 2.1 s, retrieval 3.8 s, processing 1.2 s, generation 0.9 s (Ch3.6).
- Pareto analysis typically shows ~20% of components consume ~80% of total time (Ch3.10).
- Teams optimising prompts contributing 15% of cost while missing external API calls contributing 60% illustrates the need for profiling (Ch3.10).
- Execution efficiency = contributing steps / total steps; e.g., 6 of 10 steps useful = 60% (Ch3.10).
- If ~80% of latency is LLM calls, target inference (quantization, batching, smaller models); if ~70% is tool calls, target API time, tool-result caching or parallel tool execution (Ch4.4).
- arxiv_search took 1,234ms of 2,341ms total (53%), identified as the bottleneck (Ch7.3).
- Re-profiling after 45% vector-search latency growth may reveal index rebuild, embedding-cache or query-optimization needs (Ch7.3).
- Example: 1.5 s end-to-end with tool execution 900 ms (60%) vs LLM inference 400 ms — caching tool results beats buying a faster GPU (Ch8.1).
- Example breakdown: tool execution 35%, LLM generation 50% of latency (Ref8.07).
Classification
- Patterns
- Hierarchical profilingLatency decomposition (inference, tool execution, orchestration overhead, network)Tool-, agent- and workflow-level integration (opt-in)Pareto (80/20) bottleneck analysisProfile-before-optimiseLLM-vs-tool latency breakdownResource usage forecasting under scale
- Technologies
- NVIDIA Agent Intelligence (AIQ) toolkitNVIDIA Agent Intelligence ToolkitNVIDIA NeMo Agent ToolkitNVIDIA NeMo Agent Toolkit profilerLangfuseNVIDIA Agent Intelligence Toolkit (AIQ)
- Quality attributes
- Performance efficiency (ISO/IEC 25010)Maintainability (ISO/IEC 25010)
- Risks mitigated
- Undetected performance bottlenecksOptimising low-impact components by guessworkTrajectory bloat
Sources
- Ch3.4: T. Nguyen, "Tuning Model Parameters for Production Performance," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.4. ISBN: 9798244538229.
- Ch3.6: T. Nguyen, "Trace Analysis and Execution Debugging," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.6. ISBN: 9798244538229.
- Ch3.10: T. Nguyen, "Efficiency Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.10. ISBN: 9798244538229.
- Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
- Ch7.3: T. Nguyen, "NeMo Agent Toolkit Profiling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.3. ISBN: 9798244538229.
- Ch8.1: T. Nguyen, "Latency Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 8.1. ISBN: 9798244538229.
- Ref1.01: NVIDIA, "NVIDIA NeMo Agent Toolkit overview," NVIDIA NeMo Agent Toolkit Documentation, v1.8. Accessed: Sep. 26, 2026. [Online]. Available: https://docs.nvidia.com/nemo/agent-toolkit/latest/index.html
- Ref3.01: NVIDIA, "Agent Evaluation in NVIDIA NeMo Agent Toolkit," NVIDIA NeMo Agent Toolkit Documentation, v1.8. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/nemo/agent-toolkit/latest/improve-workflows/evaluate.html
- Ref3.05: NVIDIA, "NVIDIA NeMo Agent Toolkit FAQs," NVIDIA NeMo Agent Toolkit Documentation, v1.8. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/nemo/agent-toolkit/latest/resources/faq.html
- Ref3.07: NVIDIA, "NeMo-Agent-Toolkit," GitHub repository. Accessed: Sep. 27, 2026. [Online]. Available: https://github.com/NVIDIA/NeMo-Agent-Toolkit
- Ref8.07: "Agent Health Checks and Diagnostics," unpublished reference note (07-Agent-Health-Checks-Diagnostics.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note