Cloud / Amazon S3 Vectors Interview questions
Last updated
1. What is Amazon S3 Vectors?
Amazon S3 Vectors is a capability of Amazon S3 that stores and queries vector embeddings natively. It was the first cloud object store to offer this, announced in preview in July 2025 and made generally available later that year.
You work with it through its own API: create a vector bucket, add vector indexes to it, write vectors with PutVectors, and run similarity searches with QueryVectors. There are no clusters or instances to size, and storage grows as you add data.
It keeps S3's durability and elasticity, which makes it a good fit for large embedding collections used in RAG and semantic search where cost matters more than the very lowest latency.
Take quiz
Relational rows with SQL joins
Streaming events with ordered shards
Vector embeddings with similarity search
Parquet files for Athena only
You choose an instance family per index
You reserve memory for each index
You pre-allocate shards before ingesting
There is nothing to provision; storage grows as vectors are added
2. What is a vector embedding?
A vector embedding is a list of floating-point numbers that a machine learning model produces to represent the meaning of a piece of data such as text, an image, or audio. Items with similar meaning end up with vectors that are close to each other.
For example, an embedding model turns "reset my password" and "I forgot my login" into vectors that sit near each other, even though the words differ. That closeness is what a vector search measures.
The length of the list is the dimension, and it is fixed by the model. Amazon Titan Text Embeddings V2, for instance, can output 256, 512, or 1024 dimensions. Whatever model you pick, the S3 Vectors index must be created with the same dimension.
Take quiz
They are sorted alphabetically by source text
They always share the same dimension and values
They are identical byte for byte
Their vectors are close to each other
The embedding model you use
The distance metric
The size of the vector bucket
The S3 console at query time
3. What is a vector bucket in Amazon S3 Vectors?
A vector bucket is a new S3 bucket type built specifically for vector data. It does not hold objects like a general purpose bucket does. It is a container for vector indexes.
A single account can create up to 10,000 vector buckets per Region, and each bucket can hold up to 10,000 indexes. The bucket is also where you set the default encryption (SSE-S3 or SSE-KMS) and attach a bucket policy.
Its ARN uses the s3vectors service prefix, for example arn:aws:s3vectors:us-east-1:123456789012:bucket/my-vectors.
Take quiz
Vector indexes
Lambda deployment packages
Plain objects addressed by key prefix
Parquet table snapshots
s3:vectors
s3vectors
vectordb
s3-vector-store
2 billion
100
10,000
50
4. What is a vector index in Amazon S3 Vectors?
A vector index is the structure inside a vector bucket where vectors are stored and searched. Every similarity query runs against exactly one index.
When you create it you fix three things: the dimension, the distance metric (cosine or Euclidean), and an optional list of non-filterable metadata keys. These cannot be changed later.
One index can hold up to 2 billion vectors. Each index has its own ARN, so you can control access, tag it, and even override encryption at the index level.
Take quiz
The vector dimension
The number of replicas
Non-filterable metadata keys
The distance metric
Up to 100,000
Up to 50 million
Up to 2 billion
Up to 10,000
5. What are the distance metrics supported by S3 Vectors?
S3 Vectors supports two distance metrics, chosen when the index is created.
| Cosine | Euclidean |
| Compares the angle between two vectors and ignores their length. | Measures the straight-line distance between two points. |
| Common for text embeddings. | Useful when magnitude carries meaning. |
The query returns a distance, so a smaller value means a closer match. Pick the metric your embedding model was trained for; the model's documentation usually says which one.
Both metrics are fixed per index, so a single search cannot mix them. If two embedding models need different metrics, give each its own index.
Take quiz
Jaccard and Levenshtein
Dot product and Chebyshev
Cosine and Euclidean
Manhattan and Hamming
A vector stored more recently
A less relevant match
A vector with fewer dimensions
A closer match
6. What is the purpose of the dimension setting on an index?
The dimension tells the index how many numbers every vector must contain. S3 Vectors accepts values from 1 to 4,096, and each vector you write must match the index dimension exactly.
It has to match the embedding model. If a model outputs 1024 values, the index is created with dimension=1024. A vector with 768 values sent to that index is rejected with a validation error.
Because the dimension is fixed at creation, switching to a model with a different output size means creating a new index and re-embedding your data.
Take quiz
It goes to a separate overflow index
It is padded with zeros automatically
It is truncated to 1024 values
The request is rejected
1 to 4,096
128 to 2,048
1 to 512
1 to 65,536
7. What does a vector record contain in S3 Vectors?
Each vector record has three parts: a key, the data, and optional metadata.
- key: a string you choose, unique within the index, used to get, overwrite, or delete the vector.
- data: the embedding as
float32values, matching the index dimension. - metadata: JSON key-value pairs such as source, tenant, or category.
{ "key": "doc-42#chunk-3", "data": {"float32": [0.12, -0.08, 0.33, ...]}, "metadata": {"source": "faq.pdf", "lang": "en", "page": 3} }
Choosing a meaningful key, like a document ID plus chunk number, makes later updates and deletes much simpler.
Take quiz
float32
decimal128
int8
float64 only
Encrypting that vector with its own KMS key
Identifying a vector to get, overwrite, or delete it
Choosing which Region stores it
Setting the distance metric for that vector
8. What is filterable metadata in S3 Vectors?
Filterable metadata is metadata you can use in the filter of a QueryVectors request. Metadata is filterable by default unless you list the key as non-filterable when creating the index.
It supports string, number, boolean, and list values, and each vector can carry up to 2 KB of it. A filter like {"lang": "en"} narrows the search to vectors where that field matches; with no operator given, equality ($eq) is assumed.
Keep it small and useful for narrowing results: tenant ID, category, language, year.
Take quiz
It is non-filterable
It is filterable
It is encrypted separately
It is dropped at query time
There is no limit
Up to 100 bytes
Up to 2 KB
Up to 40 KB
9. What is non-filterable metadata in S3 Vectors?
Non-filterable metadata is metadata that is stored and can be returned with query results but cannot be used in a filter. You declare these keys in metadataConfiguration when you create the index.
An index can have up to 10 non-filterable keys, each name up to 63 characters. A vector can carry up to 40 KB of total metadata across both kinds and 50 keys in all.
The usual use is storing the original text chunk, so RAG code can read it straight from the search response. It is also left out of the index size used for query charges.
Take quiz
At QueryVectors time
After 24 hours of index warm-up
At index creation
At every PutVectors call
Filtering by tenant ID
Choosing the encryption key
Setting the distance metric
Storing the original text chunk to return with results
Up to 10
Up to 2
Unlimited
Up to 50
10. How do you create a vector bucket and index using boto3?
Use the s3vectors client. Create the bucket first, then the index inside it.
import boto3 s3v = boto3.client("s3vectors", region_name="us-east-1") s3v.create_vector_bucket(vectorBucketName="docs-vectors") s3v.create_index( vectorBucketName="docs-vectors", indexName="faq-index", dataType="float32", dimension=1024, distanceMetric="cosine", metadataConfiguration={"nonFilterableMetadataKeys": ["source_text"]}, )
Dimension and distance metric must match your embedding model. The IAM role needs s3vectors:CreateVectorBucket and s3vectors:CreateIndex.
Run this once from a setup script or infrastructure code rather than on every application start. Creating a resource that already exists raises a conflict error, so catch it or check with get_index first. Use the same Region for the bucket, the index, and every later call, because vector buckets are regional.
Take quiz
s3-vectors-runtime
s3
bedrock-vectors
s3vectors
Vector bucket first, then the index
Vectors first, then the bucket
Both in one combined call
Index first, then the bucket
11. How do you insert vectors with PutVectors?
Call put_vectors with the bucket name, index name, and a list of vectors. Each vector needs a key and float32 data, and may include metadata.
s3v.put_vectors( vectorBucketName="docs-vectors", indexName="faq-index", vectors=[ {"key": "faq-1", "data": {"float32": emb1}, "metadata": {"lang": "en", "source_text": "How do I reset..."}}, {"key": "faq-2", "data": {"float32": emb2}, "metadata": {"lang": "en", "source_text": "Where is my invoice..."}}, ], )
A single call accepts up to 500 vectors, so batch your ingestion. New vectors become searchable right after the write succeeds.
Use stable keys, so a retried batch simply overwrites the same vectors instead of creating duplicates.
Take quiz
Up to 500
Up to 100,000
Up to 10,000
Up to 5
A KMS key ARN
A key and the float32 data
A tenant ID in metadata
A timestamp
12. How do you query similar vectors with QueryVectors?
Embed the user's query with the same model used for the stored vectors, then call query_vectors with that vector and a topK.
resp = s3v.query_vectors( vectorBucketName="docs-vectors", indexName="faq-index", queryVector={"float32": query_embedding}, topK=5, filter={"lang": "en"}, returnMetadata=True, returnDistance=True, ) for v in resp["vectors"]: print(v["key"], v["distance"], v["metadata"]["source_text"])
Results come back ordered from nearest to farthest. Set returnMetadata and returnDistance only when you need them.
Embedding the query with a different model than the one used at ingestion is the most common cause of poor results, because the two vector spaces do not line up. Also confirm the query vector has exactly the index dimension, or the call fails validation.
Take quiz
A model with fewer dimensions
The same model used for the stored vectors
A model chosen by S3 automatically
Any model, dimensions do not matter
Alphabetically by key
Newest to oldest
Nearest to farthest
Random order
13. What is top-K in a QueryVectors request?
topK is the number of nearest neighbors you ask the index to return. A request with topK=5 returns the five closest vectors, fewer if filters leave fewer matches.
The limit used to be 100. AWS has since raised it to 10,000, with results returned in pages of up to 100. The higher ceiling helps when you want a wide candidate set, for example to re-rank with a stronger model.
Keep topK as small as your application really needs; larger values return more data to process.
Take quiz
100 only
10
10,000
1,000,000
The query fails with a 400 error
More than topK are returned
The filter is ignored
Fewer than topK may be returned
14. What are the main use cases for Amazon S3 Vectors?
S3 Vectors suits workloads that need similarity search over a large and growing embedding set without paying for always-on compute.
- RAG: the vector store behind Amazon Bedrock Knowledge Bases or your own retrieval code.
- Semantic search over documents, tickets, or product catalogs.
- Image, audio, and video similarity using multimodal embeddings.
- Agent memory: long-term recall for AI agents.
- Archives and cold tiers: keeping the full corpus cheaply while a faster engine serves the hot subset.
It is less suited to very high sustained query rates, where a provisioned engine is the better choice.
Take quiz
Transactional order processing
Sub-millisecond high-QPS ad ranking
Full-text keyword search with scoring
RAG over a large document corpus
S3 Vectors for the full corpus, OpenSearch for hot data
S3 Vectors for hot data, S3 Glacier for queries
DynamoDB as the primary vector index
RDS as a replica of the vector bucket
15. What are the main API operations in Amazon S3 Vectors?
The API is split into bucket, index, vector, and tagging operations, all under the s3vectors namespace.
| Group | Operations |
| Vector buckets | CreateVectorBucket, GetVectorBucket, ListVectorBuckets, DeleteVectorBucket, PutVectorBucketPolicy, GetVectorBucketPolicy, DeleteVectorBucketPolicy |
| Indexes | CreateIndex, GetIndex, ListIndexes, DeleteIndex |
| Vectors | PutVectors, GetVectors, ListVectors, QueryVectors, DeleteVectors |
| Tags | TagResource, UntagResource, ListTagsForResource |
Each maps to an IAM action such as s3vectors:QueryVectors, so you can grant read-only or write-only access cleanly.
A handy way to remember them: PutVectors writes, QueryVectors searches, and GetVectors or ListVectors read back by key. Everything else is control-plane setup.
Take quiz
QueryVectors
GetVectors
ListVectors
PutVectorBucketPolicy
s3-vector:
s3vectors:
s3:
vectordb:
16. What are the encryption options for S3 Vectors?
Data at rest is always encrypted. The default is SSE-S3, using AES-256 keys managed by S3. You can instead choose SSE-KMS with your own AWS KMS key for tighter control and audit.
The choice is made on the vector bucket when it is created, and the bucket's encryption type cannot be changed afterward. An individual index can override the bucket setting with its own configuration, which is useful for per-tenant keys.
Traffic to the service is protected in transit with TLS.
Take quiz
No encryption
SSE-S3
SSE-C
Client-side encryption only
Yes, but only once a year
Yes, by recreating indexes only
No
Yes, at any time
17. What are the key service limits in Amazon S3 Vectors?
These are the limits to remember for interviews and design reviews.
| Item | Limit |
| Vector buckets per Region per account | 10,000 |
| Indexes per vector bucket | 10,000 |
| Vectors per index | Up to 2 billion |
| Dimensions per vector | 1 to 4,096 |
| Vectors per PutVectors / DeleteVectors call | 500 |
| Vectors per GetVectors call | 100 |
| Total metadata per vector | 40 KB (50 keys) |
| Filterable metadata per vector | 2 KB |
| Non-filterable keys per index | 10 |
| Write rate per index | 1,000 requests/s, 2,500 vectors/s |
| Request payload | 20 MiB |
Limits change over time, so confirm against the current AWS documentation.
Take quiz
20 KB
2 GiB
20 MiB
200 bytes
1,000,000
25
250,000
2,500
40 KB
2 KB
400 bytes
4 MB
18. How do you delete vectors, an index, or a vector bucket?
Deletion works from the inside out.
delete_vectorsremoves vectors by key, up to 500 per call.delete_indexremoves an index and every vector in it.delete_vector_bucketremoves the bucket once its indexes are gone.
s3v.delete_vectors(vectorBucketName="docs-vectors", indexName="faq-index", keys=["faq-1", "faq-2"]) s3v.delete_index(vectorBucketName="docs-vectors", indexName="faq-index") s3v.delete_vector_bucket(vectorBucketName="docs-vectors")
Deleting an index is permanent, so many teams restrict s3vectors:DeleteIndex to a small admin role.
To clear one document's chunks, delete by key rather than dropping the index. To start over completely, deleting the index is faster than removing vectors in batches of 500. Neither can be undone, so export first if you may need the data again.
Take quiz
Up to 50
Exactly 1
Unlimited
Up to 500
Its vector indexes
Its IAM role
Its CloudTrail trail
Its KMS key
19. What are the pricing dimensions of Amazon S3 Vectors?
Pricing has three parts, and you pay for what you use rather than for provisioned capacity.
- Storage: the logical size of stored vectors, keys, and metadata.
- PUT requests: charged by the amount of data you write.
- Queries: a per-API-call charge plus a per-TB charge based on index size. Non-filterable metadata is excluded from that size, and the per-TB rate drops once an index passes about 100,000 vectors.
Check the AWS pricing page for current rates, which vary by Region.
Take quiz
Query charges
Reserved memory units
Provisioned shard hours
Per-index monthly license
Vector keys
Non-filterable metadata
Filterable metadata
The vector data
20. How do you list all vectors in an index?
Use list_vectors. It returns up to 1,000 vectors per page, with a pagination token to fetch the next page, and can include data and metadata if you ask for them.
For big indexes, split the work with segments. You set a segment count (up to 16) and have each worker read a different segment in parallel.
token = None while True: kw = {"vectorBucketName": "docs-vectors", "indexName": "faq-index"} if token: kw["nextToken"] = token resp = s3v.list_vectors(**kw) for v in resp["vectors"]: print(v["key"]) token = resp.get("nextToken") if not token: break
Listing is the practical way to audit, export, or migrate an index.
Take quiz
Exactly 100
Up to 1,000
Up to 100,000
Up to 10
By opening 16 indexes
It cannot be parallelized
By reading multiple segments, up to 16
By using 16 buckets
21. What is the difference between S3 Vectors and a general purpose S3 bucket?
Both live in S3 and share durability, but they store and retrieve different things.
| Aspect | General purpose bucket | Vector bucket |
| Stores | Objects of any type | Vectors with keys and metadata |
| Retrieval | By object key or prefix | By similarity search, or by vector key |
| Structure | Bucket then prefixes | Bucket then vector indexes |
| API and IAM prefix | s3: | s3vectors: |
| Fixed at creation | Region, name | Index dimension and distance metric |
You cannot put regular objects in a vector bucket, and vector indexes do not appear as objects in a general purpose bucket.
In practice this means familiar S3 tooling, such as lifecycle rules, Inventory, and Athena, does not apply to vector data. Plan backups and exports through the s3vectors API instead.
Take quiz
By S3 Select over objects
By listing object prefixes only
By similarity search against an index
By Athena SQL on the bucket
s3tables:
s3:
s3express:
s3vectors:
22. How does S3 Vectors integrate with Amazon Bedrock Knowledge Bases?
When you create a knowledge base you can pick S3 Vectors as the vector store. Bedrock can create the vector bucket and index for you or use ones you provide.
During ingestion Bedrock chunks your documents, embeds each chunk with the chosen model, and writes the vectors to the index. The chunk text is kept in non-filterable metadata. At question time it embeds the query, runs QueryVectors, and passes the top chunks to the foundation model.
flowchart LR A["Source documents in S3"] --> B["Chunk and embed"] B --> C[(S3 Vectors index)] D["User question"] --> E["Embed question"] E --> C C --> F["Top matching chunks"] F --> G["Foundation model answer"]
You can also pick SSE-KMS for the vector store in the knowledge base setup. Hybrid keyword-plus-vector search is not available with this store.
Take quiz
The index dimension
The KMS key policy
The IAM role ARN
The chunk text
Hybrid keyword plus vector search
Metadata filtering by key
Top-K retrieval
Semantic vector search
23. How does the S3 Vectors integration with Amazon OpenSearch work?
There are two paths, both aimed at a hot/cold design where S3 Vectors holds the large, cheaper copy.
- Export: in the S3 console choose a vector index, then Advanced search export. It builds an OpenSearch Ingestion pipeline that copies the index into an OpenSearch Serverless vector collection, along with an IAM role and a dead-letter bucket for failures.
- S3 Vectors as the storage layer for OpenSearch, so OpenSearch provides search and analytics while the vectors rest in S3.
After an export, the S3 index stays intact and OpenSearch serves the high-QPS or hybrid-search traffic. It is a copy, not a move, so plan how updates reach both sides.
Take quiz
An OpenSearch Ingestion pipeline into OpenSearch Serverless
A Glue crawler over the bucket
A Kinesis stream into DynamoDB
A Redshift cluster
It is renamed with a suffix
It remains intact
It is deleted
It becomes read-only forever
24. What is the difference between S3 Vectors and Amazon OpenSearch for vector search?
They trade cost against speed and features.
| Aspect | S3 Vectors | OpenSearch |
| Capacity | No planning, storage-elastic | Sized for data and peak load |
| Cost model | Pay per storage, write, and query | Pay for provisioned or serverless compute |
| Latency | About 100 ms for frequent queries, sub-second otherwise | Tuned for low latency at high QPS |
| Search features | Vector search with metadata filters | Hybrid, full-text, aggregations, analytics |
| Write rate | Up to 1,000 requests/s per index | Scales with the cluster |
Many teams use both: S3 Vectors for the full corpus and OpenSearch for the slice that needs top speed.
A practical rule: if your latency budget is a few hundred milliseconds and queries are occasional, start with S3 Vectors. Move only the indexes that prove they need more speed to OpenSearch.
Take quiz
Pay-per-query pricing
Hybrid and full-text search
Zero capacity planning
Bucket-level SSE-S3
An RDS instance with pgvector
A self-managed vector database on EC2
S3 Vectors
A provisioned OpenSearch domain
25. Why is S3 Vectors cheaper than memory-based vector databases?
Memory-based engines keep indexes in RAM on nodes that run all the time, so you pay for them whether or not anyone is searching. S3 Vectors keeps the data in durable object storage and charges for what you store, write, and query.
AWS has said it can cut vector storage and query costs by up to 90% compared with specialized databases. The saving is largest for big collections with infrequent or bursty queries, such as RAG long tails and archives.
The trade-off is query speed and throughput. If you need sustained high QPS at very low latency, the always-on engine may still be the better value.
Take quiz
Vectors are compressed to one dimension
Data is deleted after 30 days
You pay for storage and use, not always-on provisioned nodes
Queries are free
Infrequent RAG lookups
Archived embeddings
Bursty batch evaluation
Sustained very high QPS at lowest latency
26. When should you choose S3 Vectors over other vector stores?
Choose S3 Vectors when the dataset is large or growing, and queries are occasional or bursty rather than constant.
- You want no capacity planning from a few vectors up to billions per index.
- Cost is a bigger concern than single-digit-millisecond latency.
- You are using Bedrock Knowledge Bases and want the cheapest store behind it.
- You need S3-style durability, IAM, and KMS control at the index level.
- You want a cold tier now and the option to export hot data to OpenSearch later.
Queries around 100 ms are fine for most RAG and semantic search apps, which is why it works well there.
Take quiz
A workload needing full-text scoring
A feed ranker needing 2 ms responses
A tiny dataset with 50,000 queries per second
A large embedding archive queried occasionally
Exporting it to OpenSearch
Moving it to Glacier
Switching it to Redshift
Converting it into Parquet
27. When should you not use Amazon S3 Vectors?
Avoid it when the workload hits one of its boundaries.
- Sustained writes above the per-index limits: 1,000 combined put/delete requests per second and 2,500 vectors per second.
- Very low-latency, high-QPS serving, where a provisioned engine fits better.
- Hybrid or keyword search, which S3 Vectors does not provide on its own.
- Queries across several indexes, since each query targets one index.
- Vectors with more than 4,096 dimensions.
In many of these cases a tiered design, or sharding across indexes, can still keep S3 Vectors in the picture.
If one of these is only a partial match, test before ruling it out. A filtered RAG lookup at moderate traffic usually fits comfortably, while a recommendation feed at thousands of queries per second does not.
Take quiz
Keyword plus vector hybrid search
Semantic search over documents
Metadata filtering by tenant
Storing 1 billion vectors in an index
1,024
4,096
65,536
8,192
28. How does metadata filtering work in QueryVectors?
You pass a filter with the query. S3 Vectors evaluates the filter while it searches, checking candidate vectors against your conditions as it looks for the top K, rather than filtering after the fact.
Equality is the default operator, and comparison and logical operators such as $gt, $in, $and, and $or can be combined for richer conditions.
filter={"$and": [ {"tenant": "acme"}, {"year": {"$gte": 2024}} ]}
Only filterable metadata can be used. A key marked non-filterable at index creation cannot appear in a filter.
Keep filter values consistent at write time. If one writer stores year as a string and another as a number, comparisons like $gte will not behave as you expect. Decide the types up front and validate them in the ingestion code.
Take quiz
It never applies filters
While searching for the top K candidates
Only when the index is created
Only after returning all vectors
Only the vector key
Any key including non-filterable ones
Filterable metadata only
Only numeric metadata
29. Why can a filtered query return fewer than top-K results?
A filter can only return vectors that match it. If the index holds fewer matching vectors than topK, you simply get all of them.
Say you ask for topK=20 with {"tenant": "acme"} and that tenant has only 12 vectors. The response contains 12. This is expected and not an error.
Code that builds prompts or pages from results should handle a short list. If you consistently see sparse results, check that the filter keys and values match what you wrote, since a typo in a value silently matches nothing.
For prompts, do not assume you always have K chunks. Build the context from whatever comes back, and fall back gracefully when a new tenant has nothing indexed yet.
Take quiz
20 results padded with nearest non-matches
An error
12 results
0 results
Using SSE-S3
Having too many dimensions
Using cosine distance
A mismatch between filter values and stored metadata
30. How does S3 Vectors handle consistency for newly written vectors?
S3 Vectors offers strong consistency. Once a PutVectors call returns successfully, the vectors can be found by queries and reads that follow.
This matters for apps that write and then immediately search, such as an agent saving a memory and recalling it on the next turn, or a user uploading a document and asking about it right away. You do not need wait loops or polling.
Writes are still bounded by the per-index rate limits, so heavy concurrent writers should batch and retry.
Consistency applies to the index you wrote to. A copy exported to OpenSearch is updated by the pipeline, so it can lag behind the S3 index and should not be treated as strongly consistent.
Take quiz
They appear after about 15 minutes
They appear only after a manual refresh
They appear after the next index rebuild
New vectors are searchable immediately
An agent that saves a memory and recalls it next turn
A nightly batch report
A static image gallery
A log archive
31. How can you optimize PutVectors ingestion throughput?
Work within the per-index ceilings: 1,000 combined put/delete requests per second, 2,500 vectors inserted and deleted per second, 500 vectors and 20 MiB per request.
- Batch up to 500 vectors per call instead of sending one at a time.
- Keep payloads under 20 MiB; large metadata shrinks how many vectors fit.
- Use parallel workers but cap total vectors per second near 2,500 per index.
- Retry throttled calls with exponential backoff and jitter.
- For more throughput, spread data across several indexes.
Remember that embedding generation is usually the real bottleneck, so batch that call too.
Load test before launch, because the effective ceiling depends on batch size and metadata weight. Smaller metadata means more vectors fit under the payload cap.
Take quiz
2,500
25
100,000
500
Switch to Euclidean distance
Spread data across multiple indexes
Increase the vector dimension
Use smaller float values
32. How do you design metadata to stay within S3 Vectors limits?
Split metadata by purpose. Put only what you will filter on in filterable keys, and keep that under 2 KB per vector. Put everything else, especially long text, in non-filterable keys.
The totals to respect are 40 KB of metadata and 50 keys per vector, plus a maximum of 10 non-filterable keys per index. Exceeding the size limits makes PutVectors return a 400 error.
| Metadata | Store as |
| tenant_id, category, language, year | Filterable |
| Original text chunk, summary | Non-filterable (e.g. source_text) |
| Full document body | Keep in S3 and store only a pointer |
Remember that non-filterable keys must be declared when the index is created, so decide this early.
Take quiz
In a filterable key
In a non-filterable key
In the bucket name
In the vector key
It is silently truncated
It is compressed automatically
PutVectors returns a 400 error
The index is rebuilt
33. Why can't you change the dimension or distance metric of an index?
Both are part of how the index is built and how every stored vector is interpreted, so S3 Vectors fixes them at creation.
A vector produced for 1024 dimensions has no meaning in a 768-dimension space, and a distance metric changes what counts as nearest. Switching either would invalidate the stored vectors.
If you move to a new embedding model, the normal procedure is to create a new index with the new settings, re-embed the source data, load it, and cut traffic over once results are verified. Keeping the old index until then makes rollback easy.
A blue-green approach works well: keep serving from the old index while the new one loads, compare recall on test queries, then switch the index name in configuration.
Take quiz
Update the dimension on the existing index
Truncate old vectors to fit
Create a new index and re-embed the data
Change the distance metric only
To support object versioning
Because IAM requires it
To save KMS costs
Stored vectors only make sense in that space
34. Which is better for S3 Vectors: cosine or Euclidean distance?
Neither is better in general. Use the one your embedding model was trained and evaluated with.
Most text embedding models, including common sentence and Titan-style models, are designed for cosine similarity, which compares direction and ignores vector length. Choose Euclidean when the length of the vector carries meaning, or when the model documentation says to use it.
If vectors are already normalized to length one, the two give the same ranking, so the choice matters little. Since the metric cannot be changed later, check the model docs before creating the index.
If you are unsure, run a small recall test with both metrics on the same sample queries. That is cheap compared with rebuilding a large index later.
Take quiz
Cosine always ranks higher
Euclidean fails on unit vectors
They give reversed rankings
They give the same ranking
What the embedding model was trained for
The bucket name
The number of indexes
The AWS Region
35. How do you build a RAG pipeline with S3 Vectors?
A RAG pipeline has an ingestion side and a query side, both using the same embedding model.
flowchart TD A[Documents] --> B["Split into chunks"] B --> C["Embed each chunk"] C --> D["PutVectors with chunk text as metadata"] D --> E[(S3 Vectors index)] Q["User question"] --> R["Embed question"] R --> S["QueryVectors topK with filter"] E --> S S --> T["Build prompt with retrieved chunks"] T --> U["LLM generates answer"]
- Chunk documents and embed each chunk.
- Write vectors with a stable key like
docId#chunkNo, and chunk text in non-filterable metadata. - At query time, embed the question and call
QueryVectors, adding filters such as tenant or language. - Pass the returned chunks to the LLM as context.
Bedrock Knowledge Bases can do all of this for you if you prefer a managed route.
Take quiz
The chunk text, as metadata
The bucket policy
The index ARN
The KMS key ID
No model, it is keyword matched
The same model used to embed the chunks
A larger LLM only
Any model with a different dimension
36. Explain the execution flow of a QueryVectors request?
A query carries an already-embedded vector, so the embedding step happens before S3 Vectors is involved.
sequenceDiagram participant App participant Model as Embedding model participant S3V as S3 Vectors App->>Model: Embed user question Model-->>App: Query vector App->>S3V: QueryVectors(index, vector, topK, filter) S3V->>S3V: Authorize, search index, apply filter S3V-->>App: Keys, distances, metadata (paged)
- IAM is checked for
s3vectors:QueryVectorson the index (and KMS access when SSE-KMS is used). - The service searches the index for the nearest neighbors, checking filter conditions on candidates as it goes.
- It returns up to
topKresults in pages of up to 100, nearest first, with the keys, distances, and metadata you requested.
Search is approximate, so the very nearest vector is almost always but not guaranteed to be returned.
Take quiz
An image file
An already embedded query vector
Raw text to embed
A SQL statement
Exact brute-force scan only
Keyword inverted index lookup
Approximate nearest neighbor
Regular-expression match
37. How do you control access to a vector index with IAM?
Use IAM actions in the s3vectors: namespace and scope them to bucket or index ARNs. Reads and writes can be separated cleanly.
{ "Effect": "Allow", "Action": ["s3vectors:QueryVectors", "s3vectors:GetVectors"], "Resource": "arn:aws:s3vectors:us-east-1:123456789012:bucket/docs-vectors/index/faq-index" }
A search service can get read-only access, while an ingestion job gets PutVectors and DeleteVectors. You can add a vector bucket policy for cross-account access, and use tag-based (ABAC) conditions on indexes to manage access at scale. With SSE-KMS, callers also need permission on the key.
Separate roles for ingestion, search, and administration is a good baseline. Deny DeleteIndex and DeleteVectorBucket for everyone except an admin role, and scope resource ARNs to the index name so one team cannot read another team's index. Cross-account callers need an identity policy on their side and a bucket policy on yours.
Take quiz
Allow s3vectors:* on all resources
Allow DeleteIndex only
Allow QueryVectors and GetVectors only
Allow PutVectors only
S3 Inventory
Object Lock
Lifecycle rules
Attribute-based access control (ABAC)
38. How do you secure data in Amazon S3 Vectors?
Layer the usual AWS controls, since S3 Vectors works with them.
- Identity: least-privilege IAM, bucket policies, and tag-based ABAC.
- Encryption at rest: SSE-S3 by default, SSE-KMS for key control, and per-index overrides.
- Network: an interface VPC endpoint (service
com.amazonaws.<region>.s3vectors) with an endpoint policy keeps traffic off the public internet. - Audit: AWS CloudTrail. Data events must be turned on explicitly, and they record that a call happened but not the vector contents.
Also treat stored text as data: embeddings and chunks can leak sensitive content, so restrict who can query.
Review the endpoint policy, bucket policy, and IAM policy together, since a request must pass every layer. Also decide early which principals can use the KMS key, because losing key access makes the vectors unreadable.
Take quiz
A NAT gateway only
Object Lock
A public bucket ACL
An interface VPC endpoint
No, you must enable them
Only for SSE-KMS buckets
Only in the us-east-1 Region
Yes, always
39. What changed between S3 Vectors preview and general availability?
The preview launched in July 2025, and general availability followed later that year with much bigger limits.
| Aspect | Preview | General availability |
| Vectors per index | 50 million | Up to 2 billion |
| Vectors per bucket | Lower ceiling | Up to 20 trillion |
| Regions | 5 | 14 at launch, with more added since |
| Query latency | Sub-second | About 100 ms for frequent queries |
| Write rate | Lower | Up to 1,000 PUTs per second per index |
Since GA, AWS has also raised the maximum topK from 100 to 10,000. Always check current docs for limits and Region availability.
For interviews, remember the headline numbers: 50 million to 2 billion vectors per index, 5 to 14 Regions at launch, and the later jump in topK from 100 to 10,000. They show how quickly the service matured.
Take quiz
50 million
10 trillion
2 billion
1 million
50 million
2 billion
100 million
20 trillion
40. How do you design multi-tenancy in S3 Vectors?
There are two common patterns, and the right one depends on how strict the isolation must be.
| Pattern | Index per tenant | Shared index with tenant filter |
| Isolation | Strong: separate IAM, tags, even KMS key | Logical only, depends on always sending the filter |
| Scale | Up to 10,000 indexes per bucket | Up to 2 billion vectors in one index |
| Cost and ops | More indexes to manage | Simple, one index |
| Risk | Low | A missed filter exposes other tenants |
Choose index-per-tenant for regulated or high-value tenants, and the shared index with a tenant_id filter for many small ones. Enforce the filter in a service layer, never in client code.
Take quiz
One vector key prefix per tenant
A separate index per tenant
A shared index with a filter
A shared index with larger topK
Slower PutVectors
Higher storage cost per vector
A missed filter can expose other tenants' data
Fewer dimensions allowed
41. What happens when you PutVectors with an existing key?
The new vector replaces the old one stored under that key. PutVectors acts as an upsert, so there is no separate update call.
This is how you re-embed a changed chunk. Write it again with the same key and the new data and metadata, and later queries use the new version.
Deleted or shortened documents need care: if a document shrinks from 10 chunks to 7, overwriting keys 1 to 7 leaves chunks 8 to 10 behind. Use a predictable key pattern like docId#n and call DeleteVectors for the leftovers.
Because writes replace by key, retries are safe: sending the same batch twice leaves one copy of each vector. This also makes pipelines easy to rerun after a failure.
Take quiz
Stores a second copy with the same key
Appends metadata to the old vector
Replaces the stored vector
Rejects the call with a conflict
Change the distance metric
Nothing, they expire automatically
Rebuild the index
Delete the stale chunks 8 to 10
42. How do you paginate QueryVectors results when top-K is above 100?
With a topK above 100, results come back in pages of up to 100, and you follow a pagination token to fetch more until you have collected what you need. The maximum topK is 10,000.
Use it when you want a wide candidate pool, for example to retrieve 500 chunks and re-rank them with a cross-encoder before choosing the best 10 for the prompt.
Be careful with latency and payload size: more pages means more round trips and more data. Ask for only the metadata you need, and keep topK near what you will really use.
If your code only needs the best 10 matches, skip paging entirely and set topK to 10. Paging is for re-ranking or analytics-style retrieval, not the typical RAG lookup.
Take quiz
10
1,000
10,000
100
To re-rank a wide candidate set with a stronger model
To reduce query cost
To change the distance metric
To bypass metadata limits
43. How can you reduce the cost of running S3 Vectors?
Costs follow storage, writes, and queries, so each has its own lever.
- Storage: use a smaller embedding dimension if quality holds, and keep large text out of filterable metadata or in S3 itself.
- Writes: send only changed chunks rather than re-embedding everything.
- Queries: cache frequent results, avoid oversized topK, and use non-filterable keys for bulky data since they are excluded from the index size used for query pricing.
- Hygiene: delete unused indexes and stale chunks.
Test the dimension tradeoff with a recall check before committing, since fewer dimensions can lower accuracy.
Tag indexes for cost allocation. That shows which team or tenant drives the bill and where pruning will pay off.
Take quiz
It is excluded from the index size used for query charges
It increases the topK limit
It makes queries faster by definition
It removes the need for KMS
Rewriting the whole corpus nightly
Re-embedding and writing only changed chunks
Using larger batches of duplicate keys
Increasing metadata per vector
44. How do you migrate vectors from another vector database to S3 Vectors?
Treat it as a bulk export, transform, and verified load.
- Create the index with the same dimension and a compatible distance metric. If you change the model, re-embed instead.
- Export vectors, keys, and metadata from the source.
- Map metadata: filterable fields (tenant, category) and non-filterable fields (text), and trim anything over the limits.
- Load in batches of up to 500 with parallel workers, staying under the write limits.
- Compare counts using
ListVectorsand spot-check keys.
Before cutover, run the same set of test queries against both stores and compare recall and latency. Keep the old store running until the results look right.
Take quiz
The Region's availability zones
The vector dimension
The number of indexes
The source database vendor name
Delete the source immediately
Lower topK to 1
Compare test query results on both stores
Disable IAM on the index
45. How do you evaluate recall in S3 Vectors?
S3 Vectors does approximate nearest neighbor search, so the top results can differ slightly from an exact search. Measure it instead of assuming.
- Take a sample of a few hundred real queries.
- Compute the exact nearest neighbors offline with a brute-force scan on the same vectors and metric.
- Run the same queries through
QueryVectorswith the same K. - Calculate recall@K: the share of exact neighbors that also appear in the S3 Vectors results.
Repeat with your real filters, since filtering can change what comes back. If recall is low for your use case, try a larger topK with re-ranking.
Take quiz
The storage cost per vector
The number of indexes used
The share of exact nearest neighbors found in the approximate results
The query latency in milliseconds
A random sample of vectors
The index's distance metric name
The previous day's query logs only
A brute-force exact search offline
46. How do you implement hybrid search with S3 Vectors?
S3 Vectors does semantic vector search with metadata filters. It does not do keyword scoring, so hybrid search needs another component.
- OpenSearch: export the index, or use S3 Vectors as its storage layer, and run keyword plus vector queries there.
- Two-stage in your app: run a keyword search elsewhere, run
QueryVectors, and merge the lists with a method such as reciprocal rank fusion. - Metadata prefilters for exact attributes like category or year, which covers many needs.
Note that Bedrock Knowledge Bases with S3 Vectors does not offer hybrid search, so use the direct API route if you need it.
Take quiz
Setting hybrid=true on QueryVectors
A larger dimension
Cosine distance with SSE-KMS
Another engine such as OpenSearch, or an app-level merge
Reciprocal rank fusion
Raft consensus
Bloom filter union
Consistent hashing
47. How do you troubleshoot PutVectors 400 errors?
A 400 means the request failed validation. Check these in order.
- Dimension mismatch: the number of values must equal the index dimension.
- Metadata too large: over 2 KB filterable, over 40 KB total, or more than 50 keys.
- Batch too big: more than 500 vectors or a payload over 20 MiB.
- Wrong data type: values must be
float32numbers, not strings or nulls such as NaN. - Key problems: a missing or empty key.
Log the first failing vector and its sizes, and validate against the index settings with get_index before sending. Handle throttling (429-style) errors separately with backoff.
Take quiz
A validation (400) error
A silent drop
Automatic padding
A new index to be created
More than 500 vectors in the call
Using SSE-S3 instead of SSE-KMS
A dimension mismatch
Metadata over the size limit
48. How do you troubleshoot throttling or slow queries in S3 Vectors?
Separate write throttling from query latency.
Throttled writes usually mean the index hit 1,000 requests per second or 2,500 vectors per second. Batch up to 500 vectors per call, add backoff with jitter, and spread load over more indexes.
Slow queries have a few common causes:
- A very large
topKor returning metadata you do not need. - Infrequently queried indexes, which respond in the sub-second range rather than the roughly 100 ms of frequent ones.
- Network path: calling from another Region or without a nearby VPC endpoint.
- Overly restrictive filters, which can make the search work harder.
Measure end-to-end time too. Embedding the question is often slower than the search itself.
Take quiz
Increase the dimension
Batch writes, back off, and spread across indexes
Switch distance metric
Disable encryption
topK of 5
A small payload
A very large topK with metadata returned
Calling from the same Region
49. What are the common pitfalls of using S3 Vectors with Bedrock Knowledge Bases?
A few behaviors surprise teams the first time.
- Metadata limits: Bedrock stores chunk text and its own fields as metadata, so big chunks or lots of custom filterable metadata can hit the size limits.
- No hybrid search: only semantic retrieval is available with this store.
- Deletion behavior: when a knowledge base is deleted, the data deletion policy (delete or retain) decides whether the vectors go too.
- Fixed index settings: dimension must match the embedding model, and changing models means a new index.
- Encryption choice: choose SSE-KMS at setup, since bucket encryption cannot be changed later.
Test with representative documents early to catch metadata and chunk size issues.
Take quiz
Semantic retrieval
Metadata filtering
Hybrid keyword plus vector search
SSE-KMS encryption
The distance metric
The vector dimension
The bucket's Region
The data deletion policy
50. How do you structure indexes for a very large dataset in S3 Vectors?
One index can hold up to 2 billion vectors, and a query only searches one index. That shapes how you split data.
| Strategy | Use when | Trade-off |
| Single index | Under 2 billion vectors, one query scope | Simplest; write rate capped per index |
| Index per tenant or domain | Isolation or separate query scopes | More indexes; up to 10,000 per bucket |
| Index per time window | Data ages out or is searched by period | Querying many windows needs several calls |
| Sharded indexes (hash of key) | Writes exceed one index's limits | Fan out queries and merge top K in your code |
If you fan out across indexes, query them in parallel and merge by distance. Choose a split that matches your most common query so most requests hit only one index.