Prev Next

Database / Qdrant Vector DB Interview questions

How do you implement multi-vector (late interaction / ColBERT-style) search in Qdrant?

Multi-vector search stores several vectors per point — for instance, one embedding per token in a ColBERT-style late-interaction model, rather than a single pooled embedding for the whole document — and scores a query against all of them together using a specialized comparison, typically a MaxSim-style aggregation.

client.create_collection(
    collection_name="documents",
    vectors_config=models.VectorParams(
        size=128,
        distance=models.Distance.COSINE,
        multivector_config=models.MultiVectorConfig(
            comparator=models.MultiVectorComparator.MAX_SIM
        ),
    ),
)

client.upsert(
    collection_name="documents",
    points=[models.PointStruct(id=1, vector=[[0.1, 0.2, ...], [0.3, 0.4, ...], ...])],
)

Instead of a single flat vector, each point's vector field holds a list of vectors (one per token, in the ColBERT case), and Qdrant's MAX_SIM comparator scores a query (also a list of vectors, one per query token) against a stored point by finding, for each query token vector, its best match among the point's token vectors, then summing those best matches into a final relevance score.

This late-interaction approach generally captures finer-grained relevance than a single pooled embedding, since it can recognize that specific parts of a query strongly match specific parts of a document even when the documents' overall pooled representations wouldn't rank as closely; the trade-off is higher storage and compute cost, since a document with many tokens now stores and compares many vectors instead of just one, which is why multi-vector search is often used specifically as a reranking stage over a smaller candidate set from an initial, cheaper dense retrieval pass, rather than as the primary search method over an entire large collection.

What does each point's vector field contain in a multi-vector/ColBERT-style setup?
How does the MAX_SIM comparator score a query against a point?

More Related questions...

What is Qdrant? What is the purpose of Qdrant? What are the key features of Qdrant? What is a collection in Qdrant? What is a point in Qdrant? What is a payload in Qdrant? What is the HNSW algorithm? What distance metrics does Qdrant support? How do you create a collection in Qdrant? How do you insert/upsert points into a collection? What is the difference between REST and gRPC APIs in Qdrant? What client libraries are available for Qdrant? What is payload filtering in Qdrant? Define scalar quantization in Qdrant? What is a segment in Qdrant? How do you perform a similarity search in Qdrant? What is the purpose of a payload index? List the supported field types for payload indexing? What is Qdrant Cloud? What is memmap storage in Qdrant? What is the difference between Qdrant and Pinecone? What is the difference between Qdrant and Weaviate? Why is Qdrant implemented in Rust? How does the HNSW graph work internally in Qdrant? What is the difference between scalar, binary, and product quantization? Explain the internal working of binary quantization and why it's fast? What is oversampling and rescoring in quantized search? How does Qdrant implement filtering during HNSW traversal? Explain the internal working of Qdrant's sharding and replication? What consensus protocol does Qdrant use for distributed clusters, and how does it work? What are named vectors, and when should you use them? What are sparse vectors in Qdrant? Explain hybrid search using the Query API and Prefetch? What is Reciprocal Rank Fusion (RRF) versus Distribution-Based Score Fusion (DBSF)? How do you implement multitenancy in Qdrant? Explain the lifecycle of a write operation in Qdrant (WAL, segments, optimizers)? What is the role of the Write-Ahead Log (WAL) in Qdrant? How do you take and restore snapshots in Qdrant? What is the difference between keeping vectors on-disk versus the HNSW index in RAM? Explain the execution flow of a filtered vector search query in Qdrant? When should you choose binary quantization versus scalar quantization? How do you optimize Qdrant for high-throughput production workloads? What is the ACORN-1 method, and why does it matter for filtered search? Explain the internal working of Qdrant's segment optimizer/merging? How do you implement multi-vector (late interaction / ColBERT-style) search in Qdrant? What is the role of the payload index in query planning? How does Qdrant handle consistency during a node failure? Explain the execution flow of a RAG pipeline built with Qdrant as the retrieval layer? What are the trade-offs of self-hosting Qdrant versus using Qdrant Cloud? What is Qdrant's Discovery/Recommendation API used for?
Show more question and Answers...


Comments & Discussions