Human Oversight · Data artifact
Annotation Task Format
Data artifactHuman OversightExperience & Human OversightVariation point (abstract)arc:AnnotationTaskFormat
An abstract configuration specifying how human judgments are elicited for each prompt, such as absolute ratings, pairwise comparisons or multi-way rankings of candidate responses.
Responsibility. Determines the shape and reliability of the preference signal collected from annotators.
Also known as: Preference elicitation format
Variants
| Variant | When to choose |
|---|---|
| Absolute Rating Format | Used in early RLHF work; generally avoid for preference learning because annotators interpret scales inconsistently across people and over time. |
| Multi-Way Comparison Format | Choose when more information per annotation is worth higher annotator cognitive load and potentially lower judgment consistency. |
| Pairwise Comparison Format | Choose as the default (gold standard) elicitation format: comparative judgments are more reliable than absolute ratings and scale well. |
Relationships
configures structural
Sources
- Ch10.3: T. Nguyen, "RLHF Methodology," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.3. ISBN: 9798244538229.