Knowledge & Data · Model asset
Text Embedding Model
Model assetKnowledge & DataKnowledge & DataVariation point (abstract)arc:TextEmbeddingModel
A trained encoder model that maps text (queries, knowledge chunks, episode summaries) to fixed-dimension dense vectors whose cosine similarity approximates semantic relatedness.
Responsibility. Encodes text into dense semantic vectors.
Also known as: Embedding model, Dense embedding model
Variants
| Variant | When to choose |
|---|---|
| Cost-Optimized Embedding Model | Choose for cost-sensitive applications with moderate accuracy needs and as the default starting point to establish baseline retrieval performance; open-source variants when self-hosting without API dependencies is required. |
| Domain-Specific Embedding Model | Choose for specialized applications with technical or industry-specific terminology, where domain models outperform general-purpose ones. |
| General-Purpose Embedding Model | Choose for broad coverage across diverse content types. |
| High-Accuracy General Embedding Model | Choose when retrieval accuracy directly drives user experience and passages are short to medium length, justifying higher per-token cost. |
| Long-Context Embedding Model | Choose for enterprise deployments processing long technical documents, contracts or papers that routinely exceed 8,000 tokens, typically self-hosted for data sovereignty. |
| Retrieval-Optimized Embedding Model | Choose for search-specific applications over short passages where ranking precision and near-duplicate discrimination matter more than long context. |
Relationships
deployed on structural
is evaluated by assurance
is optimized by lifecycle
Design guidance
- MUST use the same embedding model for indexing and querying.
- SHOULD be validated empirically on sampled domain pairs before trusting similarity as relevance, and fine-tuned on domain data when systematic errors appear.
- SHOULD NOT use higher embedding dimensionality by default; very high dimensions make distances near-uniform and similarity less discriminative.
- SHOULD select dimensionality as an explicit trade-off between retrieval quality and storage/compute cost.
- MAY use MRL-trained models to truncate vectors (e.g., to 256 dimensions) for storage-constrained or billion-scale deployments.
- SHOULD NOT be relied on alone for specialized terminology, acronyms, exact phrases and rare entity names; complement with sparse retrieval.
- SHOULD select the embedding model deliberately, since model choice profoundly affects retrieval quality.
Quantitative guidance
As stated by the sources; verify before use.
- text-embedding-3-small produces 1536-dimensional vectors (Ch5.8).
- Embedding dimension sweet spot typically 256-1536; 4096 dimensions do not improve precision over 1536 (Ch5.8).
- Embedding dimensions range from 384 (MiniLM) to 3,072 (text-embedding-3-large) (Ch6.1).
- MRL: 256-dim truncation (8x smaller) retains ~99.9% of retrieval performance; 5-10x storage reduction at billion scale (Ch6.1).
- 100M documents: 1.2 TB full-dimensional vs 120 GB at 256 dimensions (Ch6.1).
- Embedding dimensionality typically ranges from 384 to 1,536; the worked example uses 1,024 dimensions (Ch6.3A).
Classification
- Patterns
- Multilingual embeddings for language-agnostic retrievalDense fixed-dimension embeddingsMatryoshka Representation Learning (truncatable dimensions)Distance metrics: cosine similarity, L2, dot product
- Technologies
- OpenAI text-embedding-3-smalltext-embedding-3-largetext-embedding-3-smallNV-Embed-v2nv-embedqa-e5-v5Cohere embed-english-v3E5-Large-V2MiniLM
- Quality attributes
- Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)Performance efficiency (ISO/IEC 25010)
- Risks mitigated
- Keyword-mismatch retrieval failuresKeyword-only search missing paraphrase and synonymy
Sources
- Ch5.7: T. Nguyen, "Episodic Memory," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.7. ISBN: 9798244538229.
- Ch5.8: T. Nguyen, "Semantic Memory," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.8. ISBN: 9798244538229.
- Ch6.1: T. Nguyen, "Embeddings and RAG Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.1. ISBN: 9798244538229.
- Ch6.3A: T. Nguyen, "ETL Pipeline Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.3A. ISBN: 9798244538229.
- Ch6.5: T. Nguyen, "Production RAG Systems," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.5. ISBN: 9798244538229.