Human Oversight · Data artifact

Annotation Task Format

Data artifactHuman OversightExperience & Human OversightVariation point (abstract)arc:AnnotationTaskFormat

An abstract configuration specifying how human judgments are elicited for each prompt, such as absolute ratings, pairwise comparisons or multi-way rankings of candidate responses.

Responsibility. Determines the shape and reliability of the preference signal collected from annotators.

Also known as: Preference elicitation format

configuresis specialized byis specialized byis specialized byPreference Annotation Console: configuresPreference Annotation Co…Absolute Rating Format: is specialized byAbsolute Rating FormatMulti-Way Comparison Format: is specialized byMulti-Way Comparison For…Pairwise Comparison Format: is specialized byPairwise Comparison Format
Direct neighbourhood (hover for relationship types)

Variants

VariantWhen to choose
Absolute Rating FormatUsed in early RLHF work; generally avoid for preference learning because annotators interpret scales inconsistently across people and over time.
Multi-Way Comparison FormatChoose when more information per annotation is worth higher annotator cognitive load and potentially lower judgment consistency.
Pairwise Comparison FormatChoose as the default (gold standard) elicitation format: comparative judgments are more reliable than absolute ratings and scale well.

Relationships

configures structural

Sources

  1. Ch10.3: T. Nguyen, "RLHF Methodology," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.3. ISBN: 9798244538229.