Model Adaptation · Software component
Preference Agreement Filter
Software componentModel AdaptationModelsarc:PreferenceAgreementFilter
A data-curation filter that removes preference comparisons whose annotator votes are near-random while retaining high- and moderate-agreement examples.
Responsibility. Removes noise from ambiguous or confusing comparisons without discarding legitimate disagreement.
Also known as: Disagreement analysis filter
Relationships
writes dependency
receives data from dynamic
Design guidance
- SHOULD NOT filter for unanimity; unanimous-only data teaches the dominant, often superficial criterion.
- SHOULD remove only near-random (roughly 50-50) comparisons, which typically indicate ambiguous items or problematic prompts or responses.
Quantitative guidance
As stated by the sources; verify before use.
- Reward models trained with moderate-agreement data (roughly 60-70% annotator agreement) often outperformed those trained on unanimous preferences in a summarization study (Ch10.3).
Classification
- Risks mitigated
- Reward model learning only dominant superficial criteriaNoisy preference labels
Sources
- Ch10.3: T. Nguyen, "RLHF Methodology," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.3. ISBN: 9798244538229.