Human Oversight · Data artifact
Pairwise Comparison Format
Data artifactHuman OversightExperience & Human Oversightarc:PairwiseComparisonFormat
An annotation task format presenting two candidate responses to one prompt and asking which better satisfies the stated criteria.
Responsibility. Elicits reliable comparative judgments that map directly to Bradley-Terry reward modeling.
Variant of Annotation Task Format abstract
When to choose. Choose as the default (gold standard) elicitation format: comparative judgments are more reliable than absolute ratings and scale well.
Relationships
alternative to variability
Classification
- Quality attributes
- Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)Performance efficiency (ISO/IEC 25010)
Sources
- Ch10.3: T. Nguyen, "RLHF Methodology," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.3. ISBN: 9798244538229.