Observability & Evaluation · Software component
Trajectory Matching Evaluator
Software componentObservability & EvaluationObservability & Evaluationarc:TrajectoryMatchingEvaluator
A task success evaluator that compares an agent's executed action sequence against a golden reference trajectory.
Responsibility. Scores runs by conformance to a reference action sequence.
Also known as: Golden path comparison, Trajectory Evaluator, Action Trajectory Evaluator, Sequence Correctness Evaluator
Variant of Task Success Evaluator abstract
When to choose. Choose only when a task has a single valid execution path; the text warns it penalizes valid alternative routes of stochastic agents.
Relationships
is configured by structural
is invoked by dependency
reads dependency
emits telemetry to dynamic
evaluates assurance
- Agent Controller abstract Ch3.8
alternative to variability
Design guidance
- SHOULD NOT be the primary success criterion for non-deterministic agents, as it misclassifies legitimate alternative paths as failures.
- SHOULD use exact match only for compliance-critical procedural tasks, in-order match where sequencing matters but exploration adds value, and any-order match where ordering is irrelevant.
- SHOULD supplement trajectory match with outcome quality and update references when agents consistently find better valid paths.
- MUST augment generic trajectory metrics with domain-specific prerequisite and ordering constraints in safety-critical domains.
- SHOULD measure error detection, recovery attempt, recovery success and recovery efficiency separately from first-attempt accuracy.
Quantitative guidance
As stated by the sources; verify before use.
- Ocado case: precision 68%, recall 96%, step utility 0.74, exact match 31%, in-order 78%; after optimization trajectory length 8.7->5.8 actions, precision 68%->89%, step utility 0.74->0.91 (Ch3.8).
- Agent with 76% first-attempt accuracy reached 94% task completion via recovery vs 88%-accuracy agent at 89% (Ch3.8).
Classification
- Patterns
- Trajectory matchingExact matchIn-order matchAny-order matchTrajectory precision and recallStep utility scoring (task state before/after each action)Context-aware scoring with conditioned reference actionsPrerequisite/ordering safety constraint checksMulti-attempt error recovery analysis
- Technologies
- Google Agent Development Kit
- Quality attributes
- Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)Performance efficiency (ISO/IEC 25010)Safety (ISO/IEC 25010 | NIST AI RMF: safe)
- Risks mitigated
- Task success masking inefficient or policy-violating action pathsExact match trap penalizing valid alternative pathsSkipped prerequisite steps in regulated workflows
Sources
- Ch3.3: T. Nguyen, "Web Navigation and Interaction Benchmarks," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.3. ISBN: 9798244538229.
- Ch3.8: T. Nguyen, "Action Accuracy Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.8. ISBN: 9798244538229.