Knowledge & Data · Data store

Distributed Vector Index Store

Data storeKnowledge & DataKnowledge & Dataarc:DistributedVectorIndexStore

A self-operated, horizontally distributed vector store built for maximum throughput over billions of vectors, supporting sparse and dense vectors per collection, multi-level tenant isolation and hot/cold storage tiering.

Responsibility. Provides high-throughput multi-tenant vector search at very large scale.

Also known as: High-performance vector database

Variant of Vector Index Store abstract

When to choose. Choose when query throughput, storage efficiency, multi-tenancy or on-premises constraints demand capabilities simpler vector databases lack and an infrastructure team can absorb the operational complexity.

specializesdeployed ondeployed onalternative tosends data tosends data toalternative toalternative toalternative toalternative tosends data toVector Index Store: specializesVector Index StoreGPU Node: deployed onGPU NodeContainer Orchestrator: deployed onContainer OrchestratorSelf-Managed Vector Index Store: alternative toSelf-Managed Vector Inde…Event Stream Log: sends data toEvent Stream LogObject Store: sends data toObject StoreEmbedded Vector Index Store: alternative toEmbedded Vector Index St…Managed Vector Index Store: alternative toManaged Vector Index StoreFilter-Optimized Vector Index Store: alternative toFilter-Optimized Vector …Relational Vector Extension Store: alternative toRelational Vector Extens…Cluster Consensus Coordinator: sends data toCluster Consensus Coordi…
Direct neighbourhood (hover for relationship types)

Relationships

deployed on structural

sends data to dynamic

alternative to variability

Design guidance

Classification

Patterns
Sparse-dense hybrid collectionsHot/cold storage tieringMulti-tenancy by database/collection/partition/partition keyGPU-accelerated indexingMultiple index types (HNSW, IVF, Annoy, DiskANN)GPU-accelerated index build and searchTime-travel queriesPartition keysChange data capture for downstream synchronization
Technologies
MilvusetcdMinIOAmazon S3Apache KafkaApache PulsarWeaviate
Quality attributes
Performance efficiency (ISO/IEC 25010)Security (ISO/IEC 25010 | NIST AI RMF: secure and resilient)
Risks mitigated
Memory cost growing proportionally with collection sizeCross-tenant data exposure

Sources

  1. Ch4.1: T. Nguyen, "Introduction to AI Agent Deployment and Scaling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.1. ISBN: 9798244538229.
  2. Ch6.2A: T. Nguyen, "Vector Database Selection," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.2A. ISBN: 9798244538229.
  3. Ref7.07: E. Li, V. Bellotti, R. Kraus, and R. Kao, "Build a retrieval-augmented generation (RAG) agent with NVIDIA Nemotron," NVIDIA Technical Blog, Sep. 23, 2025. [Online]. Available: https://developer.nvidia.com/blog/build-a-rag-agent-with-nvidia-nemotron/