Safety & Security · Software component
LLM-Judge Bias Detector
Software componentSafety & SecuritySafety, Security & Governancearc:LLMJudgeBiasDetector
A bias detector that prompts a large language model, given an output and its demographic context, to judge whether the output contains stereotypes, demographic assumptions, or differential treatment.
Responsibility. Judges fairness of an output semantically using an LLM as evaluator.
Also known as: LLM-as-judge bias detection, Semantic bias judge
Variant of Output Bias Detector abstract
When to choose. Choose for high-stakes decisions needing maximum detection accuracy, accepting a full LLM inference of latency and per-evaluation cost.
Relationships
invokes dependency
alternative to variability
Quantitative guidance
As stated by the sources; verify before use.
- Highest accuracy of the three approaches, 94-97% in research studies (Ch9.4).
- Cost about $0.002-0.005 per evaluation plus one full LLM inference of latency (Ch9.4).
Classification
- Patterns
- LLM-as-a-judge
- Quality attributes
- Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)
Sources
- Ch9.4: T. Nguyen, "Fairness and Bias Mitigation," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 9.4. ISBN: 9798244538229.