Prev Next

Cloud / Amazon S3 Vectors Interview questions

Last updated

1. What is Amazon S3 Vectors? 2. What is a vector embedding? 3. What is a vector bucket in Amazon S3 Vectors? 4. What is a vector index in Amazon S3 Vectors? 5. What are the distance metrics supported by S3 Vectors? 6. What is the purpose of the dimension setting on an index? 7. What does a vector record contain in S3 Vectors? 8. What is filterable metadata in S3 Vectors? 9. What is non-filterable metadata in S3 Vectors? 10. How do you create a vector bucket and index using boto3? 11. How do you insert vectors with PutVectors? 12. How do you query similar vectors with QueryVectors? 13. What is top-K in a QueryVectors request? 14. What are the main use cases for Amazon S3 Vectors? 15. What are the main API operations in Amazon S3 Vectors? 16. What are the encryption options for S3 Vectors? 17. What are the key service limits in Amazon S3 Vectors? 18. How do you delete vectors, an index, or a vector bucket? 19. What are the pricing dimensions of Amazon S3 Vectors? 20. How do you list all vectors in an index? 21. What is the difference between S3 Vectors and a general purpose S3 bucket? 22. How does S3 Vectors integrate with Amazon Bedrock Knowledge Bases? 23. How does the S3 Vectors integration with Amazon OpenSearch work? 24. What is the difference between S3 Vectors and Amazon OpenSearch for vector search? 25. Why is S3 Vectors cheaper than memory-based vector databases? 26. When should you choose S3 Vectors over other vector stores? 27. When should you not use Amazon S3 Vectors? 28. How does metadata filtering work in QueryVectors? 29. Why can a filtered query return fewer than top-K results? 30. How does S3 Vectors handle consistency for newly written vectors? 31. How can you optimize PutVectors ingestion throughput? 32. How do you design metadata to stay within S3 Vectors limits? 33. Why can't you change the dimension or distance metric of an index? 34. Which is better for S3 Vectors: cosine or Euclidean distance? 35. How do you build a RAG pipeline with S3 Vectors? 36. Explain the execution flow of a QueryVectors request? 37. How do you control access to a vector index with IAM? 38. How do you secure data in Amazon S3 Vectors? 39. What changed between S3 Vectors preview and general availability? 40. How do you design multi-tenancy in S3 Vectors? 41. What happens when you PutVectors with an existing key? 42. How do you paginate QueryVectors results when top-K is above 100? 43. How can you reduce the cost of running S3 Vectors? 44. How do you migrate vectors from another vector database to S3 Vectors? 45. How do you evaluate recall in S3 Vectors? 46. How do you implement hybrid search with S3 Vectors? 47. How do you troubleshoot PutVectors 400 errors? 48. How do you troubleshoot throttling or slow queries in S3 Vectors? 49. What are the common pitfalls of using S3 Vectors with Bedrock Knowledge Bases? 50. How do you structure indexes for a very large dataset in S3 Vectors?

1. What is Amazon S3 Vectors?

Amazon S3 Vectors is a capability of Amazon S3 that stores and queries vector embeddings natively. It was the first cloud object store to offer this, announced in preview in July 2025 and made generally available later that year.

You work with it through its own API: create a vector bucket, add vector indexes to it, write vectors with PutVectors, and run similarity searches with QueryVectors. There are no clusters or instances to size, and storage grows as you add data.

It keeps S3's durability and elasticity, which makes it a good fit for large embedding collections used in RAG and semantic search where cost matters more than the very lowest latency.

Take quiz
What does Amazon S3 Vectors natively store and query?
Relational rows with SQL joins
Streaming events with ordered shards
Vector embeddings with similarity search
Parquet files for Athena only
How is capacity handled in S3 Vectors?
You choose an instance family per index
You reserve memory for each index
You pre-allocate shards before ingesting
There is nothing to provision; storage grows as vectors are added

2. What is a vector embedding?

A vector embedding is a list of floating-point numbers that a machine learning model produces to represent the meaning of a piece of data such as text, an image, or audio. Items with similar meaning end up with vectors that are close to each other.

For example, an embedding model turns "reset my password" and "I forgot my login" into vectors that sit near each other, even though the words differ. That closeness is what a vector search measures.

The length of the list is the dimension, and it is fixed by the model. Amazon Titan Text Embeddings V2, for instance, can output 256, 512, or 1024 dimensions. Whatever model you pick, the S3 Vectors index must be created with the same dimension.

Take quiz
What do embeddings of similar content look like?
They are sorted alphabetically by source text
They always share the same dimension and values
They are identical byte for byte
Their vectors are close to each other
Who decides the number of dimensions in an embedding?
The embedding model you use
The distance metric
The size of the vector bucket
The S3 console at query time

3. What is a vector bucket in Amazon S3 Vectors?

A vector bucket is a new S3 bucket type built specifically for vector data. It does not hold objects like a general purpose bucket does. It is a container for vector indexes.

A single account can create up to 10,000 vector buckets per Region, and each bucket can hold up to 10,000 indexes. The bucket is also where you set the default encryption (SSE-S3 or SSE-KMS) and attach a bucket policy.

Its ARN uses the s3vectors service prefix, for example arn:aws:s3vectors:us-east-1:123456789012:bucket/my-vectors.

Take quiz
What does a vector bucket directly contain?
Vector indexes
Lambda deployment packages
Plain objects addressed by key prefix
Parquet table snapshots
Which ARN service prefix identifies S3 Vectors resources?
s3:vectors
s3vectors
vectordb
s3-vector-store
How many vector indexes can one vector bucket hold by default?
2 billion
100
10,000
50

4. What is a vector index in Amazon S3 Vectors?

A vector index is the structure inside a vector bucket where vectors are stored and searched. Every similarity query runs against exactly one index.

When you create it you fix three things: the dimension, the distance metric (cosine or Euclidean), and an optional list of non-filterable metadata keys. These cannot be changed later.

One index can hold up to 2 billion vectors. Each index has its own ARN, so you can control access, tag it, and even override encryption at the index level.

Take quiz
Which setting is NOT chosen when you create a vector index?
The vector dimension
The number of replicas
Non-filterable metadata keys
The distance metric
How many vectors can a single index hold?
Up to 100,000
Up to 50 million
Up to 2 billion
Up to 10,000

5. What are the distance metrics supported by S3 Vectors?

S3 Vectors supports two distance metrics, chosen when the index is created.

Cosine Euclidean
Compares the angle between two vectors and ignores their length. Measures the straight-line distance between two points.
Common for text embeddings. Useful when magnitude carries meaning.

The query returns a distance, so a smaller value means a closer match. Pick the metric your embedding model was trained for; the model's documentation usually says which one.

Both metrics are fixed per index, so a single search cannot mix them. If two embedding models need different metrics, give each its own index.

Take quiz
Which two distance metrics does S3 Vectors support?
Jaccard and Levenshtein
Dot product and Chebyshev
Cosine and Euclidean
Manhattan and Hamming
In a QueryVectors result, a smaller distance value means:
A vector stored more recently
A less relevant match
A vector with fewer dimensions
A closer match

6. What is the purpose of the dimension setting on an index?

The dimension tells the index how many numbers every vector must contain. S3 Vectors accepts values from 1 to 4,096, and each vector you write must match the index dimension exactly.

It has to match the embedding model. If a model outputs 1024 values, the index is created with dimension=1024. A vector with 768 values sent to that index is rejected with a validation error.

Because the dimension is fixed at creation, switching to a model with a different output size means creating a new index and re-embedding your data.

Take quiz
What happens if you PutVectors a 768-value vector into a 1024-dimension index?
It goes to a separate overflow index
It is padded with zeros automatically
It is truncated to 1024 values
The request is rejected
What is the allowed dimension range per vector?
1 to 4,096
128 to 2,048
1 to 512
1 to 65,536

7. What does a vector record contain in S3 Vectors?

Each vector record has three parts: a key, the data, and optional metadata.

  • key: a string you choose, unique within the index, used to get, overwrite, or delete the vector.
  • data: the embedding as float32 values, matching the index dimension.
  • metadata: JSON key-value pairs such as source, tenant, or category.
{
  "key": "doc-42#chunk-3",
  "data": {"float32": [0.12, -0.08, 0.33, ...]},
  "metadata": {"source": "faq.pdf", "lang": "en", "page": 3}
}

Choosing a meaningful key, like a document ID plus chunk number, makes later updates and deletes much simpler.

Take quiz
Which numeric type does S3 Vectors use for vector data?
float32
decimal128
int8
float64 only
What is the vector key used for?
Encrypting that vector with its own KMS key
Identifying a vector to get, overwrite, or delete it
Choosing which Region stores it
Setting the distance metric for that vector

8. What is filterable metadata in S3 Vectors?

Filterable metadata is metadata you can use in the filter of a QueryVectors request. Metadata is filterable by default unless you list the key as non-filterable when creating the index.

It supports string, number, boolean, and list values, and each vector can carry up to 2 KB of it. A filter like {"lang": "en"} narrows the search to vectors where that field matches; with no operator given, equality ($eq) is assumed.

Keep it small and useful for narrowing results: tenant ID, category, language, year.

Take quiz
What is the default behavior for a metadata key?
It is non-filterable
It is filterable
It is encrypted separately
It is dropped at query time
How much filterable metadata can a single vector carry?
There is no limit
Up to 100 bytes
Up to 2 KB
Up to 40 KB

9. What is non-filterable metadata in S3 Vectors?

Non-filterable metadata is metadata that is stored and can be returned with query results but cannot be used in a filter. You declare these keys in metadataConfiguration when you create the index.

An index can have up to 10 non-filterable keys, each name up to 63 characters. A vector can carry up to 40 KB of total metadata across both kinds and 50 keys in all.

The usual use is storing the original text chunk, so RAG code can read it straight from the search response. It is also left out of the index size used for query charges.

Take quiz
When must non-filterable metadata keys be declared?
At QueryVectors time
After 24 hours of index warm-up
At index creation
At every PutVectors call
What is a typical use of non-filterable metadata?
Filtering by tenant ID
Choosing the encryption key
Setting the distance metric
Storing the original text chunk to return with results
How many non-filterable metadata keys can one index have?
Up to 10
Up to 2
Unlimited
Up to 50

10. How do you create a vector bucket and index using boto3?

Use the s3vectors client. Create the bucket first, then the index inside it.

import boto3

s3v = boto3.client("s3vectors", region_name="us-east-1")

s3v.create_vector_bucket(vectorBucketName="docs-vectors")

s3v.create_index(
    vectorBucketName="docs-vectors",
    indexName="faq-index",
    dataType="float32",
    dimension=1024,
    distanceMetric="cosine",
    metadataConfiguration={"nonFilterableMetadataKeys": ["source_text"]},
)

Dimension and distance metric must match your embedding model. The IAM role needs s3vectors:CreateVectorBucket and s3vectors:CreateIndex.

Run this once from a setup script or infrastructure code rather than on every application start. Creating a resource that already exists raises a conflict error, so catch it or check with get_index first. Use the same Region for the bucket, the index, and every later call, because vector buckets are regional.

Take quiz
Which boto3 client name is used for S3 Vectors?
s3-vectors-runtime
s3
bedrock-vectors
s3vectors
In which order are the resources created?
Vector bucket first, then the index
Vectors first, then the bucket
Both in one combined call
Index first, then the bucket

11. How do you insert vectors with PutVectors?

Call put_vectors with the bucket name, index name, and a list of vectors. Each vector needs a key and float32 data, and may include metadata.

s3v.put_vectors(
    vectorBucketName="docs-vectors",
    indexName="faq-index",
    vectors=[
        {"key": "faq-1", "data": {"float32": emb1},
         "metadata": {"lang": "en", "source_text": "How do I reset..."}},
        {"key": "faq-2", "data": {"float32": emb2},
         "metadata": {"lang": "en", "source_text": "Where is my invoice..."}},
    ],
)

A single call accepts up to 500 vectors, so batch your ingestion. New vectors become searchable right after the write succeeds.

Use stable keys, so a retried batch simply overwrites the same vectors instead of creating duplicates.

Take quiz
How many vectors can one PutVectors call carry?
Up to 500
Up to 100,000
Up to 10,000
Up to 5
Which field is required on every vector in PutVectors?
A KMS key ARN
A key and the float32 data
A tenant ID in metadata
A timestamp

12. How do you query similar vectors with QueryVectors?

Embed the user's query with the same model used for the stored vectors, then call query_vectors with that vector and a topK.

resp = s3v.query_vectors(
    vectorBucketName="docs-vectors",
    indexName="faq-index",
    queryVector={"float32": query_embedding},
    topK=5,
    filter={"lang": "en"},
    returnMetadata=True,
    returnDistance=True,
)
for v in resp["vectors"]:
    print(v["key"], v["distance"], v["metadata"]["source_text"])

Results come back ordered from nearest to farthest. Set returnMetadata and returnDistance only when you need them.

Embedding the query with a different model than the one used at ingestion is the most common cause of poor results, because the two vector spaces do not line up. Also confirm the query vector has exactly the index dimension, or the call fails validation.

Take quiz
Which embedding model must be used to embed the query?
A model with fewer dimensions
The same model used for the stored vectors
A model chosen by S3 automatically
Any model, dimensions do not matter
In what order does QueryVectors return results?
Alphabetically by key
Newest to oldest
Nearest to farthest
Random order

13. What is top-K in a QueryVectors request?

topK is the number of nearest neighbors you ask the index to return. A request with topK=5 returns the five closest vectors, fewer if filters leave fewer matches.

The limit used to be 100. AWS has since raised it to 10,000, with results returned in pages of up to 100. The higher ceiling helps when you want a wide candidate set, for example to re-rank with a stronger model.

Keep topK as small as your application really needs; larger values return more data to process.

Take quiz
What is the current maximum topK for QueryVectors?
100 only
10
10,000
1,000,000
What can happen to the number of results when a filter is applied?
The query fails with a 400 error
More than topK are returned
The filter is ignored
Fewer than topK may be returned

14. What are the main use cases for Amazon S3 Vectors?

S3 Vectors suits workloads that need similarity search over a large and growing embedding set without paying for always-on compute.

  • RAG: the vector store behind Amazon Bedrock Knowledge Bases or your own retrieval code.
  • Semantic search over documents, tickets, or product catalogs.
  • Image, audio, and video similarity using multimodal embeddings.
  • Agent memory: long-term recall for AI agents.
  • Archives and cold tiers: keeping the full corpus cheaply while a faster engine serves the hot subset.

It is less suited to very high sustained query rates, where a provisioned engine is the better choice.

Take quiz
Which workload is a natural fit for S3 Vectors?
Transactional order processing
Sub-millisecond high-QPS ad ranking
Full-text keyword search with scoring
RAG over a large document corpus
Which pairing is a common tiered pattern?
S3 Vectors for the full corpus, OpenSearch for hot data
S3 Vectors for hot data, S3 Glacier for queries
DynamoDB as the primary vector index
RDS as a replica of the vector bucket

15. What are the main API operations in Amazon S3 Vectors?

The API is split into bucket, index, vector, and tagging operations, all under the s3vectors namespace.

Group Operations
Vector buckets CreateVectorBucket, GetVectorBucket, ListVectorBuckets, DeleteVectorBucket, PutVectorBucketPolicy, GetVectorBucketPolicy, DeleteVectorBucketPolicy
Indexes CreateIndex, GetIndex, ListIndexes, DeleteIndex
Vectors PutVectors, GetVectors, ListVectors, QueryVectors, DeleteVectors
Tags TagResource, UntagResource, ListTagsForResource

Each maps to an IAM action such as s3vectors:QueryVectors, so you can grant read-only or write-only access cleanly.

A handy way to remember them: PutVectors writes, QueryVectors searches, and GetVectors or ListVectors read back by key. Everything else is control-plane setup.

Take quiz
Which operation runs a similarity search?
QueryVectors
GetVectors
ListVectors
PutVectorBucketPolicy
What IAM prefix do S3 Vectors actions use?
s3-vector:
s3vectors:
s3:
vectordb:

16. What are the encryption options for S3 Vectors?

Data at rest is always encrypted. The default is SSE-S3, using AES-256 keys managed by S3. You can instead choose SSE-KMS with your own AWS KMS key for tighter control and audit.

The choice is made on the vector bucket when it is created, and the bucket's encryption type cannot be changed afterward. An individual index can override the bucket setting with its own configuration, which is useful for per-tenant keys.

Traffic to the service is protected in transit with TLS.

Take quiz
What is the default encryption for a new vector bucket?
No encryption
SSE-S3
SSE-C
Client-side encryption only
Can the bucket-level encryption type be changed after the bucket is created?
Yes, but only once a year
Yes, by recreating indexes only
No
Yes, at any time

17. What are the key service limits in Amazon S3 Vectors?

These are the limits to remember for interviews and design reviews.

Item Limit
Vector buckets per Region per account 10,000
Indexes per vector bucket 10,000
Vectors per index Up to 2 billion
Dimensions per vector 1 to 4,096
Vectors per PutVectors / DeleteVectors call 500
Vectors per GetVectors call 100
Total metadata per vector 40 KB (50 keys)
Filterable metadata per vector 2 KB
Non-filterable keys per index 10
Write rate per index 1,000 requests/s, 2,500 vectors/s
Request payload 20 MiB

Limits change over time, so confirm against the current AWS documentation.

Take quiz
Which request-size cap applies to a single API payload?
20 KB
2 GiB
20 MiB
200 bytes
What is the combined limit on vectors inserted and deleted per second per index?
1,000,000
25
250,000
2,500
What is the maximum total metadata per vector?
40 KB
2 KB
400 bytes
4 MB

18. How do you delete vectors, an index, or a vector bucket?

Deletion works from the inside out.

  1. delete_vectors removes vectors by key, up to 500 per call.
  2. delete_index removes an index and every vector in it.
  3. delete_vector_bucket removes the bucket once its indexes are gone.
s3v.delete_vectors(vectorBucketName="docs-vectors",
                   indexName="faq-index",
                   keys=["faq-1", "faq-2"])
s3v.delete_index(vectorBucketName="docs-vectors", indexName="faq-index")
s3v.delete_vector_bucket(vectorBucketName="docs-vectors")

Deleting an index is permanent, so many teams restrict s3vectors:DeleteIndex to a small admin role.

To clear one document's chunks, delete by key rather than dropping the index. To start over completely, deleting the index is faster than removing vectors in batches of 500. Neither can be undone, so export first if you may need the data again.

Take quiz
How many keys can one DeleteVectors call remove?
Up to 50
Exactly 1
Unlimited
Up to 500
What must be removed before deleting a vector bucket?
Its vector indexes
Its IAM role
Its CloudTrail trail
Its KMS key

19. What are the pricing dimensions of Amazon S3 Vectors?

Pricing has three parts, and you pay for what you use rather than for provisioned capacity.

  • Storage: the logical size of stored vectors, keys, and metadata.
  • PUT requests: charged by the amount of data you write.
  • Queries: a per-API-call charge plus a per-TB charge based on index size. Non-filterable metadata is excluded from that size, and the per-TB rate drops once an index passes about 100,000 vectors.

Check the AWS pricing page for current rates, which vary by Region.

Take quiz
Which of these is a billing dimension for S3 Vectors?
Query charges
Reserved memory units
Provisioned shard hours
Per-index monthly license
What is excluded from the index size used in query charges?
Vector keys
Non-filterable metadata
Filterable metadata
The vector data

20. How do you list all vectors in an index?

Use list_vectors. It returns up to 1,000 vectors per page, with a pagination token to fetch the next page, and can include data and metadata if you ask for them.

For big indexes, split the work with segments. You set a segment count (up to 16) and have each worker read a different segment in parallel.

token = None
while True:
    kw = {"vectorBucketName": "docs-vectors", "indexName": "faq-index"}
    if token: kw["nextToken"] = token
    resp = s3v.list_vectors(**kw)
    for v in resp["vectors"]:
        print(v["key"])
    token = resp.get("nextToken")
    if not token: break

Listing is the practical way to audit, export, or migrate an index.

Take quiz
How many vectors can a ListVectors page return?
Exactly 100
Up to 1,000
Up to 100,000
Up to 10
How can ListVectors be parallelized?
By opening 16 indexes
It cannot be parallelized
By reading multiple segments, up to 16
By using 16 buckets

21. What is the difference between S3 Vectors and a general purpose S3 bucket?

Both live in S3 and share durability, but they store and retrieve different things.

Aspect General purpose bucket Vector bucket
Stores Objects of any type Vectors with keys and metadata
Retrieval By object key or prefix By similarity search, or by vector key
Structure Bucket then prefixes Bucket then vector indexes
API and IAM prefix s3: s3vectors:
Fixed at creation Region, name Index dimension and distance metric

You cannot put regular objects in a vector bucket, and vector indexes do not appear as objects in a general purpose bucket.

In practice this means familiar S3 tooling, such as lifecycle rules, Inventory, and Athena, does not apply to vector data. Plan backups and exports through the s3vectors API instead.

Take quiz
How do you find data in a vector bucket?
By S3 Select over objects
By listing object prefixes only
By similarity search against an index
By Athena SQL on the bucket
Which IAM prefix applies to vector buckets?
s3tables:
s3:
s3express:
s3vectors:

22. How does S3 Vectors integrate with Amazon Bedrock Knowledge Bases?

When you create a knowledge base you can pick S3 Vectors as the vector store. Bedrock can create the vector bucket and index for you or use ones you provide.

During ingestion Bedrock chunks your documents, embeds each chunk with the chosen model, and writes the vectors to the index. The chunk text is kept in non-filterable metadata. At question time it embeds the query, runs QueryVectors, and passes the top chunks to the foundation model.

flowchart LR
  A["Source documents in S3"] --> B["Chunk and embed"]
  B --> C[(S3 Vectors index)]
  D["User question"] --> E["Embed question"]
  E --> C
  C --> F["Top matching chunks"]
  F --> G["Foundation model answer"]

You can also pick SSE-KMS for the vector store in the knowledge base setup. Hybrid keyword-plus-vector search is not available with this store.

Take quiz
What does Bedrock keep in non-filterable metadata for each chunk?
The index dimension
The KMS key policy
The IAM role ARN
The chunk text
Which search style is not supported when S3 Vectors backs a knowledge base?
Hybrid keyword plus vector search
Metadata filtering by key
Top-K retrieval
Semantic vector search

23. How does the S3 Vectors integration with Amazon OpenSearch work?

There are two paths, both aimed at a hot/cold design where S3 Vectors holds the large, cheaper copy.

  1. Export: in the S3 console choose a vector index, then Advanced search export. It builds an OpenSearch Ingestion pipeline that copies the index into an OpenSearch Serverless vector collection, along with an IAM role and a dead-letter bucket for failures.
  2. S3 Vectors as the storage layer for OpenSearch, so OpenSearch provides search and analytics while the vectors rest in S3.

After an export, the S3 index stays intact and OpenSearch serves the high-QPS or hybrid-search traffic. It is a copy, not a move, so plan how updates reach both sides.

Take quiz
What does the S3 console 'Advanced search export' create?
An OpenSearch Ingestion pipeline into OpenSearch Serverless
A Glue crawler over the bucket
A Kinesis stream into DynamoDB
A Redshift cluster
After exporting to OpenSearch, what happens to the S3 vector index?
It is renamed with a suffix
It remains intact
It is deleted
It becomes read-only forever

24. What is the difference between S3 Vectors and Amazon OpenSearch for vector search?

They trade cost against speed and features.

Aspect S3 Vectors OpenSearch
Capacity No planning, storage-elastic Sized for data and peak load
Cost model Pay per storage, write, and query Pay for provisioned or serverless compute
Latency About 100 ms for frequent queries, sub-second otherwise Tuned for low latency at high QPS
Search features Vector search with metadata filters Hybrid, full-text, aggregations, analytics
Write rate Up to 1,000 requests/s per index Scales with the cluster

Many teams use both: S3 Vectors for the full corpus and OpenSearch for the slice that needs top speed.

A practical rule: if your latency budget is a few hundred milliseconds and queries are occasional, start with S3 Vectors. Move only the indexes that prove they need more speed to OpenSearch.

Take quiz
Which feature is stronger in OpenSearch than in S3 Vectors?
Pay-per-query pricing
Hybrid and full-text search
Zero capacity planning
Bucket-level SSE-S3
Which typically needs the least capacity planning?
An RDS instance with pgvector
A self-managed vector database on EC2
S3 Vectors
A provisioned OpenSearch domain

25. Why is S3 Vectors cheaper than memory-based vector databases?

Memory-based engines keep indexes in RAM on nodes that run all the time, so you pay for them whether or not anyone is searching. S3 Vectors keeps the data in durable object storage and charges for what you store, write, and query.

AWS has said it can cut vector storage and query costs by up to 90% compared with specialized databases. The saving is largest for big collections with infrequent or bursty queries, such as RAG long tails and archives.

The trade-off is query speed and throughput. If you need sustained high QPS at very low latency, the always-on engine may still be the better value.

Take quiz
Why does S3 Vectors cost less for large, rarely queried data?
Vectors are compressed to one dimension
Data is deleted after 30 days
You pay for storage and use, not always-on provisioned nodes
Queries are free
Which workload benefits least from the S3 Vectors cost model?
Infrequent RAG lookups
Archived embeddings
Bursty batch evaluation
Sustained very high QPS at lowest latency

26. When should you choose S3 Vectors over other vector stores?

Choose S3 Vectors when the dataset is large or growing, and queries are occasional or bursty rather than constant.

  • You want no capacity planning from a few vectors up to billions per index.
  • Cost is a bigger concern than single-digit-millisecond latency.
  • You are using Bedrock Knowledge Bases and want the cheapest store behind it.
  • You need S3-style durability, IAM, and KMS control at the index level.
  • You want a cold tier now and the option to export hot data to OpenSearch later.

Queries around 100 ms are fine for most RAG and semantic search apps, which is why it works well there.

Take quiz
Which situation favors S3 Vectors?
A workload needing full-text scoring
A feed ranker needing 2 ms responses
A tiny dataset with 50,000 queries per second
A large embedding archive queried occasionally
What option does S3 Vectors keep open for hot data?
Exporting it to OpenSearch
Moving it to Glacier
Switching it to Redshift
Converting it into Parquet

27. When should you not use Amazon S3 Vectors?

Avoid it when the workload hits one of its boundaries.

  • Sustained writes above the per-index limits: 1,000 combined put/delete requests per second and 2,500 vectors per second.
  • Very low-latency, high-QPS serving, where a provisioned engine fits better.
  • Hybrid or keyword search, which S3 Vectors does not provide on its own.
  • Queries across several indexes, since each query targets one index.
  • Vectors with more than 4,096 dimensions.

In many of these cases a tiered design, or sharding across indexes, can still keep S3 Vectors in the picture.

If one of these is only a partial match, test before ruling it out. A filtered RAG lookup at moderate traffic usually fits comfortably, while a recommendation feed at thousands of queries per second does not.

Take quiz
Which requirement is a poor fit for S3 Vectors alone?
Keyword plus vector hybrid search
Semantic search over documents
Metadata filtering by tenant
Storing 1 billion vectors in an index
What is the maximum dimension S3 Vectors supports?
1,024
4,096
65,536
8,192

28. How does metadata filtering work in QueryVectors?

You pass a filter with the query. S3 Vectors evaluates the filter while it searches, checking candidate vectors against your conditions as it looks for the top K, rather than filtering after the fact.

Equality is the default operator, and comparison and logical operators such as $gt, $in, $and, and $or can be combined for richer conditions.

filter={"$and": [
    {"tenant": "acme"},
    {"year": {"$gte": 2024}}
]}

Only filterable metadata can be used. A key marked non-filterable at index creation cannot appear in a filter.

Keep filter values consistent at write time. If one writer stores year as a string and another as a number, comparisons like $gte will not behave as you expect. Decide the types up front and validate them in the ingestion code.

Take quiz
When does S3 Vectors apply the metadata filter?
It never applies filters
While searching for the top K candidates
Only when the index is created
Only after returning all vectors
Which metadata can be used in a filter?
Only the vector key
Any key including non-filterable ones
Filterable metadata only
Only numeric metadata

29. Why can a filtered query return fewer than top-K results?

A filter can only return vectors that match it. If the index holds fewer matching vectors than topK, you simply get all of them.

Say you ask for topK=20 with {"tenant": "acme"} and that tenant has only 12 vectors. The response contains 12. This is expected and not an error.

Code that builds prompts or pages from results should handle a short list. If you consistently see sparse results, check that the filter keys and values match what you wrote, since a typo in a value silently matches nothing.

For prompts, do not assume you always have K chunks. Build the context from whatever comes back, and fall back gracefully when a new tenant has nothing indexed yet.

Take quiz
A query asks for topK=20 and only 12 vectors match the filter. What is returned?
20 results padded with nearest non-matches
An error
12 results
0 results
What is a common cause of unexpectedly empty filtered results?
Using SSE-S3
Having too many dimensions
Using cosine distance
A mismatch between filter values and stored metadata

30. How does S3 Vectors handle consistency for newly written vectors?

S3 Vectors offers strong consistency. Once a PutVectors call returns successfully, the vectors can be found by queries and reads that follow.

This matters for apps that write and then immediately search, such as an agent saving a memory and recalling it on the next turn, or a user uploading a document and asking about it right away. You do not need wait loops or polling.

Writes are still bounded by the per-index rate limits, so heavy concurrent writers should batch and retry.

Consistency applies to the index you wrote to. A copy exported to OpenSearch is updated by the pipeline, so it can lag behind the S3 index and should not be treated as strongly consistent.

Take quiz
What consistency can you rely on after a successful PutVectors?
They appear after about 15 minutes
They appear only after a manual refresh
They appear after the next index rebuild
New vectors are searchable immediately
Which app benefits directly from this behavior?
An agent that saves a memory and recalls it next turn
A nightly batch report
A static image gallery
A log archive

31. How can you optimize PutVectors ingestion throughput?

Work within the per-index ceilings: 1,000 combined put/delete requests per second, 2,500 vectors inserted and deleted per second, 500 vectors and 20 MiB per request.

  1. Batch up to 500 vectors per call instead of sending one at a time.
  2. Keep payloads under 20 MiB; large metadata shrinks how many vectors fit.
  3. Use parallel workers but cap total vectors per second near 2,500 per index.
  4. Retry throttled calls with exponential backoff and jitter.
  5. For more throughput, spread data across several indexes.

Remember that embedding generation is usually the real bottleneck, so batch that call too.

Load test before launch, because the effective ceiling depends on batch size and metadata weight. Smaller metadata means more vectors fit under the payload cap.

Take quiz
What is the per-index limit on vectors inserted and deleted per second?
2,500
25
100,000
500
What is the best fix when one index cannot absorb your write rate?
Switch to Euclidean distance
Spread data across multiple indexes
Increase the vector dimension
Use smaller float values

32. How do you design metadata to stay within S3 Vectors limits?

Split metadata by purpose. Put only what you will filter on in filterable keys, and keep that under 2 KB per vector. Put everything else, especially long text, in non-filterable keys.

The totals to respect are 40 KB of metadata and 50 keys per vector, plus a maximum of 10 non-filterable keys per index. Exceeding the size limits makes PutVectors return a 400 error.

Metadata Store as
tenant_id, category, language, year Filterable
Original text chunk, summary Non-filterable (e.g. source_text)
Full document body Keep in S3 and store only a pointer

Remember that non-filterable keys must be declared when the index is created, so decide this early.

Take quiz
Where should a long text chunk be stored in metadata?
In a filterable key
In a non-filterable key
In the bucket name
In the vector key
What happens when metadata exceeds the allowed size?
It is silently truncated
It is compressed automatically
PutVectors returns a 400 error
The index is rebuilt

33. Why can't you change the dimension or distance metric of an index?

Both are part of how the index is built and how every stored vector is interpreted, so S3 Vectors fixes them at creation.

A vector produced for 1024 dimensions has no meaning in a 768-dimension space, and a distance metric changes what counts as nearest. Switching either would invalidate the stored vectors.

If you move to a new embedding model, the normal procedure is to create a new index with the new settings, re-embed the source data, load it, and cut traffic over once results are verified. Keeping the old index until then makes rollback easy.

A blue-green approach works well: keep serving from the old index while the new one loads, compare recall on test queries, then switch the index name in configuration.

Take quiz
What is the right way to adopt a new embedding model with a different dimension?
Update the dimension on the existing index
Truncate old vectors to fit
Create a new index and re-embed the data
Change the distance metric only
Why is dimension fixed at index creation?
To support object versioning
Because IAM requires it
To save KMS costs
Stored vectors only make sense in that space

34. Which is better for S3 Vectors: cosine or Euclidean distance?

Neither is better in general. Use the one your embedding model was trained and evaluated with.

Most text embedding models, including common sentence and Titan-style models, are designed for cosine similarity, which compares direction and ignores vector length. Choose Euclidean when the length of the vector carries meaning, or when the model documentation says to use it.

If vectors are already normalized to length one, the two give the same ranking, so the choice matters little. Since the metric cannot be changed later, check the model docs before creating the index.

If you are unsure, run a small recall test with both metrics on the same sample queries. That is cheap compared with rebuilding a large index later.

Take quiz
If vectors are normalized to unit length, how do cosine and Euclidean rankings compare?
Cosine always ranks higher
Euclidean fails on unit vectors
They give reversed rankings
They give the same ranking
What should drive your choice of metric?
What the embedding model was trained for
The bucket name
The number of indexes
The AWS Region

35. How do you build a RAG pipeline with S3 Vectors?

A RAG pipeline has an ingestion side and a query side, both using the same embedding model.

flowchart TD
  A[Documents] --> B["Split into chunks"]
  B --> C["Embed each chunk"]
  C --> D["PutVectors with chunk text as metadata"]
  D --> E[(S3 Vectors index)]
  Q["User question"] --> R["Embed question"]
  R --> S["QueryVectors topK with filter"]
  E --> S
  S --> T["Build prompt with retrieved chunks"]
  T --> U["LLM generates answer"]
  1. Chunk documents and embed each chunk.
  2. Write vectors with a stable key like docId#chunkNo, and chunk text in non-filterable metadata.
  3. At query time, embed the question and call QueryVectors, adding filters such as tenant or language.
  4. Pass the returned chunks to the LLM as context.

Bedrock Knowledge Bases can do all of this for you if you prefer a managed route.

Take quiz
What should be stored with each vector so the LLM can use it as context?
The chunk text, as metadata
The bucket policy
The index ARN
The KMS key ID
Which model must embed the user's question?
No model, it is keyword matched
The same model used to embed the chunks
A larger LLM only
Any model with a different dimension

36. Explain the execution flow of a QueryVectors request?

A query carries an already-embedded vector, so the embedding step happens before S3 Vectors is involved.

sequenceDiagram
  participant App
  participant Model as Embedding model
  participant S3V as S3 Vectors
  App->>Model: Embed user question
  Model-->>App: Query vector
  App->>S3V: QueryVectors(index, vector, topK, filter)
  S3V->>S3V: Authorize, search index, apply filter
  S3V-->>App: Keys, distances, metadata (paged)
  1. IAM is checked for s3vectors:QueryVectors on the index (and KMS access when SSE-KMS is used).
  2. The service searches the index for the nearest neighbors, checking filter conditions on candidates as it goes.
  3. It returns up to topK results in pages of up to 100, nearest first, with the keys, distances, and metadata you requested.

Search is approximate, so the very nearest vector is almost always but not guaranteed to be returned.

Take quiz
What does QueryVectors receive as input?
An image file
An already embedded query vector
Raw text to embed
A SQL statement
What kind of search does S3 Vectors perform?
Exact brute-force scan only
Keyword inverted index lookup
Approximate nearest neighbor
Regular-expression match

37. How do you control access to a vector index with IAM?

Use IAM actions in the s3vectors: namespace and scope them to bucket or index ARNs. Reads and writes can be separated cleanly.

{
  "Effect": "Allow",
  "Action": ["s3vectors:QueryVectors", "s3vectors:GetVectors"],
  "Resource": "arn:aws:s3vectors:us-east-1:123456789012:bucket/docs-vectors/index/faq-index"
}

A search service can get read-only access, while an ingestion job gets PutVectors and DeleteVectors. You can add a vector bucket policy for cross-account access, and use tag-based (ABAC) conditions on indexes to manage access at scale. With SSE-KMS, callers also need permission on the key.

Separate roles for ingestion, search, and administration is a good baseline. Deny DeleteIndex and DeleteVectorBucket for everyone except an admin role, and scope resource ARNs to the index name so one team cannot read another team's index. Cross-account callers need an identity policy on their side and a bucket policy on yours.

Take quiz
Which policy lets a search service read vectors but never change them?
Allow s3vectors:* on all resources
Allow DeleteIndex only
Allow QueryVectors and GetVectors only
Allow PutVectors only
What can scale index permissions based on index tags?
S3 Inventory
Object Lock
Lifecycle rules
Attribute-based access control (ABAC)

38. How do you secure data in Amazon S3 Vectors?

Layer the usual AWS controls, since S3 Vectors works with them.

  • Identity: least-privilege IAM, bucket policies, and tag-based ABAC.
  • Encryption at rest: SSE-S3 by default, SSE-KMS for key control, and per-index overrides.
  • Network: an interface VPC endpoint (service com.amazonaws.<region>.s3vectors) with an endpoint policy keeps traffic off the public internet.
  • Audit: AWS CloudTrail. Data events must be turned on explicitly, and they record that a call happened but not the vector contents.

Also treat stored text as data: embeddings and chunks can leak sensitive content, so restrict who can query.

Review the endpoint policy, bucket policy, and IAM policy together, since a request must pass every layer. Also decide early which principals can use the KMS key, because losing key access makes the vectors unreadable.

Take quiz
Which feature keeps S3 Vectors traffic on the AWS network?
A NAT gateway only
Object Lock
A public bucket ACL
An interface VPC endpoint
Are CloudTrail data events for S3 Vectors on by default?
No, you must enable them
Only for SSE-KMS buckets
Only in the us-east-1 Region
Yes, always

39. What changed between S3 Vectors preview and general availability?

The preview launched in July 2025, and general availability followed later that year with much bigger limits.

Aspect Preview General availability
Vectors per index 50 million Up to 2 billion
Vectors per bucket Lower ceiling Up to 20 trillion
Regions 5 14 at launch, with more added since
Query latency Sub-second About 100 ms for frequent queries
Write rate Lower Up to 1,000 PUTs per second per index

Since GA, AWS has also raised the maximum topK from 100 to 10,000. Always check current docs for limits and Region availability.

For interviews, remember the headline numbers: 50 million to 2 billion vectors per index, 5 to 14 Regions at launch, and the later jump in topK from 100 to 10,000. They show how quickly the service matured.

Take quiz
How many vectors per index did the preview allow?
50 million
10 trillion
2 billion
1 million
What is the maximum number of vectors per index at GA?
50 million
2 billion
100 million
20 trillion

40. How do you design multi-tenancy in S3 Vectors?

There are two common patterns, and the right one depends on how strict the isolation must be.

Pattern Index per tenant Shared index with tenant filter
Isolation Strong: separate IAM, tags, even KMS key Logical only, depends on always sending the filter
Scale Up to 10,000 indexes per bucket Up to 2 billion vectors in one index
Cost and ops More indexes to manage Simple, one index
Risk Low A missed filter exposes other tenants

Choose index-per-tenant for regulated or high-value tenants, and the shared index with a tenant_id filter for many small ones. Enforce the filter in a service layer, never in client code.

Take quiz
Which pattern gives the strongest tenant isolation?
One vector key prefix per tenant
A separate index per tenant
A shared index with a filter
A shared index with larger topK
What is the main risk of a shared index with a tenant filter?
Slower PutVectors
Higher storage cost per vector
A missed filter can expose other tenants' data
Fewer dimensions allowed

41. What happens when you PutVectors with an existing key?

The new vector replaces the old one stored under that key. PutVectors acts as an upsert, so there is no separate update call.

This is how you re-embed a changed chunk. Write it again with the same key and the new data and metadata, and later queries use the new version.

Deleted or shortened documents need care: if a document shrinks from 10 chunks to 7, overwriting keys 1 to 7 leaves chunks 8 to 10 behind. Use a predictable key pattern like docId#n and call DeleteVectors for the leftovers.

Because writes replace by key, retries are safe: sending the same batch twice leaves one copy of each vector. This also makes pipelines easy to rerun after a failure.

Take quiz
What does PutVectors do when the key already exists?
Stores a second copy with the same key
Appends metadata to the old vector
Replaces the stored vector
Rejects the call with a conflict
A document shrinks from 10 chunks to 7. What should you also do?
Change the distance metric
Nothing, they expire automatically
Rebuild the index
Delete the stale chunks 8 to 10

42. How do you paginate QueryVectors results when top-K is above 100?

With a topK above 100, results come back in pages of up to 100, and you follow a pagination token to fetch more until you have collected what you need. The maximum topK is 10,000.

Use it when you want a wide candidate pool, for example to retrieve 500 chunks and re-rank them with a cross-encoder before choosing the best 10 for the prompt.

Be careful with latency and payload size: more pages means more round trips and more data. Ask for only the metadata you need, and keep topK near what you will really use.

If your code only needs the best 10 matches, skip paging entirely and set topK to 10. Paging is for re-ranking or analytics-style retrieval, not the typical RAG lookup.

Take quiz
How many results does one QueryVectors page contain at most?
10
1,000
10,000
100
Why would you request a topK of 500?
To re-rank a wide candidate set with a stronger model
To reduce query cost
To change the distance metric
To bypass metadata limits

43. How can you reduce the cost of running S3 Vectors?

Costs follow storage, writes, and queries, so each has its own lever.

  • Storage: use a smaller embedding dimension if quality holds, and keep large text out of filterable metadata or in S3 itself.
  • Writes: send only changed chunks rather than re-embedding everything.
  • Queries: cache frequent results, avoid oversized topK, and use non-filterable keys for bulky data since they are excluded from the index size used for query pricing.
  • Hygiene: delete unused indexes and stale chunks.

Test the dimension tradeoff with a recall check before committing, since fewer dimensions can lower accuracy.

Tag indexes for cost allocation. That shows which team or tenant drives the bill and where pruning will pay off.

Take quiz
Why store bulky text in non-filterable metadata?
It is excluded from the index size used for query charges
It increases the topK limit
It makes queries faster by definition
It removes the need for KMS
Which change reduces write cost?
Rewriting the whole corpus nightly
Re-embedding and writing only changed chunks
Using larger batches of duplicate keys
Increasing metadata per vector

44. How do you migrate vectors from another vector database to S3 Vectors?

Treat it as a bulk export, transform, and verified load.

  1. Create the index with the same dimension and a compatible distance metric. If you change the model, re-embed instead.
  2. Export vectors, keys, and metadata from the source.
  3. Map metadata: filterable fields (tenant, category) and non-filterable fields (text), and trim anything over the limits.
  4. Load in batches of up to 500 with parallel workers, staying under the write limits.
  5. Compare counts using ListVectors and spot-check keys.

Before cutover, run the same set of test queries against both stores and compare recall and latency. Keep the old store running until the results look right.

Take quiz
What must match when you copy existing vectors into a new index?
The Region's availability zones
The vector dimension
The number of indexes
The source database vendor name
What should you do before switching traffic over?
Delete the source immediately
Lower topK to 1
Compare test query results on both stores
Disable IAM on the index

45. How do you evaluate recall in S3 Vectors?

S3 Vectors does approximate nearest neighbor search, so the top results can differ slightly from an exact search. Measure it instead of assuming.

  1. Take a sample of a few hundred real queries.
  2. Compute the exact nearest neighbors offline with a brute-force scan on the same vectors and metric.
  3. Run the same queries through QueryVectors with the same K.
  4. Calculate recall@K: the share of exact neighbors that also appear in the S3 Vectors results.

Repeat with your real filters, since filtering can change what comes back. If recall is low for your use case, try a larger topK with re-ranking.

Take quiz
What does recall@K measure here?
The storage cost per vector
The number of indexes used
The share of exact nearest neighbors found in the approximate results
The query latency in milliseconds
How do you get the ground truth for comparison?
A random sample of vectors
The index's distance metric name
The previous day's query logs only
A brute-force exact search offline

46. How do you implement hybrid search with S3 Vectors?

S3 Vectors does semantic vector search with metadata filters. It does not do keyword scoring, so hybrid search needs another component.

  • OpenSearch: export the index, or use S3 Vectors as its storage layer, and run keyword plus vector queries there.
  • Two-stage in your app: run a keyword search elsewhere, run QueryVectors, and merge the lists with a method such as reciprocal rank fusion.
  • Metadata prefilters for exact attributes like category or year, which covers many needs.

Note that Bedrock Knowledge Bases with S3 Vectors does not offer hybrid search, so use the direct API route if you need it.

Take quiz
What do you need for keyword plus vector search over S3 Vectors data?
Setting hybrid=true on QueryVectors
A larger dimension
Cosine distance with SSE-KMS
Another engine such as OpenSearch, or an app-level merge
Which method merges two ranked lists in an app-level hybrid design?
Reciprocal rank fusion
Raft consensus
Bloom filter union
Consistent hashing

47. How do you troubleshoot PutVectors 400 errors?

A 400 means the request failed validation. Check these in order.

  1. Dimension mismatch: the number of values must equal the index dimension.
  2. Metadata too large: over 2 KB filterable, over 40 KB total, or more than 50 keys.
  3. Batch too big: more than 500 vectors or a payload over 20 MiB.
  4. Wrong data type: values must be float32 numbers, not strings or nulls such as NaN.
  5. Key problems: a missing or empty key.

Log the first failing vector and its sizes, and validate against the index settings with get_index before sending. Handle throttling (429-style) errors separately with backoff.

Take quiz
A 768-value vector sent to a 1024-dimension index causes:
A validation (400) error
A silent drop
Automatic padding
A new index to be created
Which is NOT a likely cause of a PutVectors 400?
More than 500 vectors in the call
Using SSE-S3 instead of SSE-KMS
A dimension mismatch
Metadata over the size limit

48. How do you troubleshoot throttling or slow queries in S3 Vectors?

Separate write throttling from query latency.

Throttled writes usually mean the index hit 1,000 requests per second or 2,500 vectors per second. Batch up to 500 vectors per call, add backoff with jitter, and spread load over more indexes.

Slow queries have a few common causes:

  • A very large topK or returning metadata you do not need.
  • Infrequently queried indexes, which respond in the sub-second range rather than the roughly 100 ms of frequent ones.
  • Network path: calling from another Region or without a nearby VPC endpoint.
  • Overly restrictive filters, which can make the search work harder.

Measure end-to-end time too. Embedding the question is often slower than the search itself.

Take quiz
What is the first fix for write throttling on one index?
Increase the dimension
Batch writes, back off, and spread across indexes
Switch distance metric
Disable encryption
Which makes a query slower?
topK of 5
A small payload
A very large topK with metadata returned
Calling from the same Region

49. What are the common pitfalls of using S3 Vectors with Bedrock Knowledge Bases?

A few behaviors surprise teams the first time.

  • Metadata limits: Bedrock stores chunk text and its own fields as metadata, so big chunks or lots of custom filterable metadata can hit the size limits.
  • No hybrid search: only semantic retrieval is available with this store.
  • Deletion behavior: when a knowledge base is deleted, the data deletion policy (delete or retain) decides whether the vectors go too.
  • Fixed index settings: dimension must match the embedding model, and changing models means a new index.
  • Encryption choice: choose SSE-KMS at setup, since bucket encryption cannot be changed later.

Test with representative documents early to catch metadata and chunk size issues.

Take quiz
Which feature is NOT available with Bedrock Knowledge Bases on S3 Vectors?
Semantic retrieval
Metadata filtering
Hybrid keyword plus vector search
SSE-KMS encryption
What controls whether vectors are removed when a knowledge base is deleted?
The distance metric
The vector dimension
The bucket's Region
The data deletion policy

50. How do you structure indexes for a very large dataset in S3 Vectors?

One index can hold up to 2 billion vectors, and a query only searches one index. That shapes how you split data.

Strategy Use when Trade-off
Single index Under 2 billion vectors, one query scope Simplest; write rate capped per index
Index per tenant or domain Isolation or separate query scopes More indexes; up to 10,000 per bucket
Index per time window Data ages out or is searched by period Querying many windows needs several calls
Sharded indexes (hash of key) Writes exceed one index's limits Fan out queries and merge top K in your code

If you fan out across indexes, query them in parallel and merge by distance. Choose a split that matches your most common query so most requests hit only one index.

Take quiz
Why might you shard data across several indexes?
To change the distance metric per query
To avoid using IAM
To increase vector dimension
To exceed single-index write limits
What must your code do when it queries several indexes?
Merge the results by distance
Nothing, S3 merges them
Switch to a single index automatically
Wait for an index sync job
«
»

Comments & Discussions