Human Oversight · Data artifact

Pairwise Comparison Format

Data artifactHuman OversightExperience & Human Oversightarc:PairwiseComparisonFormat

An annotation task format presenting two candidate responses to one prompt and asking which better satisfies the stated criteria.

Responsibility. Elicits reliable comparative judgments that map directly to Bradley-Terry reward modeling.

Variant of Annotation Task Format abstract

When to choose. Choose as the default (gold standard) elicitation format: comparative judgments are more reliable than absolute ratings and scale well.

specializesis target of alternativeTois target of alternativeToAnnotation Task Format: specializesAnnotation Task FormatAbsolute Rating Format: is target of alternativeToAbsolute Rating FormatMulti-Way Comparison Format: is target of alternativeToMulti-Way Comparison For…
Direct neighbourhood (hover for relationship types)

Relationships

alternative to variability

Classification

Quality attributes
Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)Performance efficiency (ISO/IEC 25010)

Sources

  1. Ch10.3: T. Nguyen, "RLHF Methodology," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.3. ISBN: 9798244538229.