Knowledge & Data · Software component
Data Quality Validator
Software componentKnowledge & DataKnowledge & Dataarc:DataQualityValidator
A validation engine that applies a composable set of schema, type, range, format and cross-field checks to each incoming document and returns severity-graded, structured validation results.
Responsibility. Detects structural and value-level quality defects in documents before they enter the knowledge pipeline.
Also known as: Quality Validator, Validation engine, Validation pipeline, Ingestion validator, Structural input validation, Value range validation
Relationships
is configured by structural
invokes dependency
is invoked by dependency
- Production Quality Monitor abstract Ref8.02
writes dependency
receives data from dynamic
sends data to dynamic
guards control
is orchestrated by control
Design guidance
- SHOULD reject structurally invalid data at ingestion, before embedding and indexing, rather than discovering problems after loading.
- SHOULD accumulate all validation results per document instead of failing on the first error, so remediation can be batched.
- SHOULD grade failures by severity so that minor issues flag documents for review instead of excluding valuable content.
- SHOULD be extensible through custom validators without modifying core validation logic.
- MUST validate that source documents are accurate and well organized before RAG use (Ref6.01: quality knowledge base best practice).
- SHOULD act as a pre-prediction quality gate (missing values, schema, type, range, outliers) so bad predictions are rejected rather than silently served (Ref8.02).
Quantitative guidance
As stated by the sources; verify before use.
- A robust ingestion process might reject 15-20% of documents on first run, triggering source-system remediation (Ch6.4).
- With 5% corrupted data, top-10 retrieval over 100,000 documents can surface at least one corrupted chunk ~40% of the time for frequently asked questions (Ch6.4).
- Example rule: alert if missing-value percentage >5% (Ref8.02).
- Alert when > 5% of records miss required fields, > 2% have invalid types, or > 1% violate constraints; query_text length 1-5,000 characters (Ref8.04).
Classification
- Patterns
- Fail-fast ingestion validationCollect-all-errors validationSeverity-graded results (ERROR/WARNING/INFO)Pluggable custom validatorsSeparation of validation logic from enforcement policyRegex format validationRange checksCross-field validation
- Technologies
- JSON SchemaXML Schema
- Quality attributes
- Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)Maintainability (ISO/IEC 25010)Transparency and accountability (NIST AI RMF: accountable and transparent)
- Risks mitigated
- Schema violations entering the knowledge baseMalformed input causing embedding failuresEncoding errors corrupting textMissing required metadata breaking citations and date filteringLogically impossible values (negative ages, percentages over 100%)
Sources
- Ch6.4: T. Nguyen, "Data Quality Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.4. ISBN: 9798244538229.
- Ref6.01: S. Schürch, "How to Make Your LLM More Accurate with RAG & Fine-Tuning," Towards Data Science, Mar. 11, 2025. [Online]. Available: https://towardsdatascience.com/how-to-make-your-llm-more-accurate-with-rag-fine-tuning/
- Ref8.02: "Machine Learning Monitoring in Production," unpublished reference note (02-ML-Monitoring-Production.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref8.04: "Data Quality and Drift Detection for Agent Systems," unpublished reference note (04-Data-Quality-Drift-Detection.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note