Cloud / Amazon Bedrock Interview questions
Last updated
1. What is Amazon Bedrock?
Amazon Bedrock is AWS's fully managed, serverless service for building generative AI applications. You get one API layer over foundation models from several providers, so you can try Claude, Nova, Llama, Mistral or OpenAI models without hosting anything yourself.
Beyond model access, Bedrock bundles the pieces most teams end up needing: Guardrails for safety controls, Knowledge Bases for RAG, model customization (fine-tuning and distillation), evaluation tools, and AgentCore for building and running agents.
Your prompts and outputs aren't used to train the base models and aren't handed to the model providers. Access is governed by IAM, and you can add VPC endpoints (PrivateLink) and KMS encryption. Bedrock became generally available in September 2023.
Take quiz
Serverless, with no model infrastructure to provision or patch
It only runs models that Amazon trained itself
You launch and patch GPU instances for every model you call
They are used for training only when Guardrails is turned off
They are used to retrain the base models every night
They are not used to train the base foundation models
They are forwarded to each model provider for product improvement
2. List the model providers available on Amazon Bedrock?
The catalog keeps growing and it differs by Region, but the main providers are:
- Amazon - Nova (text, image, video, speech) and Titan (mostly embeddings now)
- Anthropic - Claude
- Meta - Llama
- Mistral AI
- Cohere - Command, Embed and Rerank
- AI21 Labs - Jamba
- Stability AI - image generation and editing
- OpenAI - GPT models, which arrived on Bedrock in 2026
- DeepSeek and Qwen - open-weight models
Don't hard-code a model list in an application. Check the model catalog in the console, or call ListFoundationModels, for the Region you're deploying to.
Take quiz
Llama and Mistral
Nova and Titan
Jamba and Command
Assume every provider's full lineup is identical in every Region
Read the list off the Bedrock pricing page only
Ask AWS Support to enable the list for your account
Check the model catalog or call ListFoundationModels for that Region
3. How do you invoke a model using the Converse API?
The Converse API gives you one request and response shape for chat-style models, so switching models mostly means changing the modelId. You send messages, an optional system prompt, inference settings and, if needed, tool definitions or a guardrail.
import boto3 client = boto3.client("bedrock-runtime", region_name="us-east-1") response = client.converse( modelId="us.amazon.nova-lite-v1:0", system=[{"text": "You answer in two sentences."}], messages=[{"role": "user", "content": [{"text": "What is Amazon Bedrock?"}]}], inferenceConfig={"maxTokens": 300, "temperature": 0.3}, ) print(response["output"]["message"]["content"][0]["text"]) print(response["usage"])
The usage block reports input, output and total tokens, which is handy for cost tracking. For streaming, call converse_stream instead. The us. prefix above is an inference profile ID, which you use when a model needs cross-Region inference.
Because the request shape is identical across providers, you can keep the modelId in configuration and swap models per environment without touching the code that builds the request.
Take quiz
The system block
modelId
The role value inside messages
inferenceConfig.maxTokens
The stopReason field
The metrics.latencyMs field
The usage block
4. How do you get access to foundation models in Amazon Bedrock?
Bedrock retired the old Model access page in October 2025. Serverless foundation models are now enabled automatically the first time you invoke them, and you control who can use what through IAM policies and service control policies.
A few exceptions still apply:
- Anthropic models need a one-time use-case form, submitted in the console or with the
PutUseCaseForModelAccessAPI. Do it from the organization's management account and member accounts inherit it. - Marketplace-sold models are subscribed automatically on first call, but that first call needs AWS Marketplace permissions and a valid payment method.
- The model has to be offered in your Region, either directly or through an inference profile.
If you hit an AccessDeniedException, walk through these three points before opening a support case.
Take quiz
Automatic enablement on first invocation, governed by IAM and SCPs
A Provisioned Throughput purchase required for each model
A separate Bedrock account type reserved for production use
A per-model approval ticket handled by AWS Support staff
Request a dedicated VPC and private endpoint from AWS Support
Submit a one-time use-case form, inherited organization-wide via the management account
Purchase a Reserved tier commitment for the Anthropic model family
5. Describe the pricing models in Amazon Bedrock?
Bedrock bills by usage, but the mode you pick changes the price. Text models charge per input and output token, embedding models per input token, and image models per image.
| Mode | How it works | Good for |
| On-demand (service tiers) | Pay per token. Standard is the default, with Flex, Priority and Reserved tiers on supported models | Most applications |
| Batch | Asynchronous jobs through S3, about 50% cheaper on select models | Bulk, offline work |
| Provisioned Throughput | Reserved model units billed hourly for a chosen term | Steady high volume |
| Customization | Charged per token trained, plus monthly storage for the custom model | Fine-tuning and distillation |
Guardrails policies, Knowledge Base retrieval and AgentCore services are billed separately from model tokens, so include them in any estimate.
Take quiz
Provisioned Throughput
Reserved tier
The Priority tier
Batch inference
Per image generated
Per GPU hour of a dedicated endpoint
Per input token only
6. What is Provisioned Throughput in Amazon Bedrock?
Provisioned Throughput reserves dedicated capacity for a model in units called model units. Each unit delivers a set amount of input and output tokens per minute, and you pay an hourly rate per unit whether you use it or not.
You choose a term: no commitment (billed hourly, handy for testing), or a longer commitment of one or six months at a lower hourly price. You then call the model with the ARN of the provisioned resource instead of the normal model ID.
It makes sense when traffic is steady enough that reserved capacity beats per-token pricing, or when you need guaranteed throughput. Some customized models also rely on it. Before buying, compare it with the Reserved service tier and with a quota increase on standard on-demand, since fine-tuned Nova models can now run on-demand.
Take quiz
Cache checkpoints
Endpoint replicas
Model units
vCPU hours
The ARN of the provisioned model resource
The base model ID, and capacity attaches automatically
The ARN of the calling IAM role
7. What is Amazon Bedrock Guardrails?
Amazon Bedrock Guardrails is a configurable safety layer that checks what goes into a model and what comes back out. You define policies once, version them, and attach the guardrail to model calls, agents or Knowledge Base queries.
When a check trips, the guardrail can block the message and return a canned reply, or mask sensitive data and let the rest through. Each intervention appears in the trace, so you can see which policy fired.
The point is that it works independently of the model's own training. Because it's enforced outside the prompt, a clever user message can't simply tell it to stand down. It also works with models outside Bedrock through the ApplyGuardrail API.
Take quiz
Only the final model output
Only the system prompt at deployment time
Both user inputs and model outputs
A guardrail is just a longer system prompt with stricter wording
The policies are enforced outside the prompt, not as model instructions
The guardrail retrains the model after each request it blocks
The guardrail quietly rewrites the system prompt on every turn
8. What are the types of safeguards in Amazon Bedrock Guardrails?
A guardrail is built from independent policies, and you can turn on any mix:
- Content filters - hate, insults, sexual, violence, misconduct and prompt attacks, each with adjustable strength. They cover text and, for supported categories, images.
- Denied topics - subjects you describe in plain language, such as investment advice.
- Word filters - custom words and phrases, plus a managed profanity list.
- Sensitive information filters - built-in PII types and custom regex, with block or mask actions.
- Contextual grounding checks - grounding and relevance scores to catch hallucinations in RAG answers.
- Automated Reasoning checks - logic-based validation of answers against a formal policy.
Prompt attack detection is an input-side control and contextual grounding is output-side. The others can be applied to both directions.
Take quiz
Denied topics for off-limits subjects
Prompt attack filter for jailbreak attempts
Word filter for blocked terms and phrases
Contextual grounding check
Content filters
Sensitive information filters
Automated Reasoning checks
9. What is the purpose of the ApplyGuardrail API?
The ApplyGuardrail API evaluates content against a guardrail without calling a model. You send the text (or image) and mark it as an input or an output, and you get back the action taken and the detailed assessments.
That decoupling is the whole point. You can guard models hosted elsewhere, such as SageMaker endpoints, EC2, on-premises servers or third-party APIs. You can also check content at different stages of a pipeline, for example screening a question before retrieval and validating the answer afterwards.
It's also a quick way to test policy changes on their own. When you call a Bedrock model directly, APIs like Converse accept a guardrailConfig that performs the same check inline.
Take quiz
It fine-tunes the guardrail on whatever content you send it
It only works with Amazon Nova models hosted on Bedrock
It checks content without invoking a model, so it works with any model
Whether the content is an INPUT or an OUTPUT
The chunking strategy
The model's tokens-per-minute quota
The Provisioned Throughput ARN
10. What are Knowledge Bases for Amazon Bedrock?
Knowledge Bases for Amazon Bedrock is the managed RAG feature. It connects a foundation model to your own documents so answers are based on your data and come with citations back to the source.
The service handles the plumbing: pulling content from sources like S3, SharePoint or Confluence, parsing and chunking it, creating embeddings, and storing them in a vector index. At query time it embeds the question, finds the closest chunks, and either returns them or passes them to a model that writes the answer.
AWS now offers two types. In a managed knowledge base, Bedrock runs storage and retrieval for you. In a customer-managed one, you pick and operate the vector store yourself.
Take quiz
A certificate guaranteeing the answer is correct
Citations pointing to the source chunks
A fine-tuned copy of the model
Retraining the base model on the new documents
Rebuilding every chunk from the data source
Compressing the vector index into model weights
Embedding the user's question and searching the vector index
11. What are the chunking strategies in Bedrock Knowledge Bases?
Chunking controls how documents are split before embedding, and it has a big effect on answer quality. Bedrock offers these options:
- Fixed-size - you set a maximum token count and an overlap. Simple and predictable.
- Hierarchical - small child chunks are matched, but the larger parent chunk is returned for context. You set both sizes and the overlap.
- Semantic - splits on meaning instead of length, so related sentences stay together. It adds ingestion cost.
- No chunking - each file is treated as one chunk, useful when you've already split the content.
- Custom - a Lambda function runs your own chunking logic.
Hierarchical chunking isn't recommended with an S3 vector bucket because of metadata size limits.
Take quiz
Hierarchical chunking
No chunking
Custom Lambda chunking
Fixed-size chunking
Your documents are very long unstructured PDFs
You want the service to detect topic boundaries for you
Your files are already split into retrieval-sized pieces
12. What is Amazon Bedrock AgentCore?
Amazon Bedrock AgentCore is AWS's platform for building, deploying and operating AI agents in production. It's modular, so you can use one piece or all of them, and it's framework and model agnostic: LangGraph, CrewAI, Strands, LlamaIndex, the OpenAI Agents SDK or plain Python all work, with models inside or outside Bedrock.
Think of it as the infrastructure around the agent rather than the agent logic itself: a secure place to run it, memory, a tool gateway, identity, observability and governance controls.
It targets a gap teams keep hitting. A prototype runs fine on a laptop, but production needs session isolation, credential handling, tracing and access controls. AgentCore provides those as managed services billed on consumption.
Take quiz
It works with any agent framework and any model
It only supports models hosted on Bedrock
It only runs agents written with one AWS-specific framework
Replacing vector databases for enterprise search
Training foundation models from scratch on customer data
Moving agents from prototype to secure, observable production
Translating documents in bulk across many languages
13. List the core services of Amazon Bedrock AgentCore?
AgentCore is a set of services you can mix and match:
- Runtime - serverless, session-isolated hosting for agents and MCP servers
- Memory - short-term session events and long-term memory across sessions
- Gateway - turns APIs, Lambda functions and existing MCP servers into agent tools
- Identity - workload identities, OAuth flows and a token vault for acting on a user's behalf
- Code Interpreter - a sandbox where the agent can run code
- Browser - a managed browser the agent can drive
- Observability - traces, metrics and dashboards, OpenTelemetry compatible
- Policy - rules on what an agent is allowed to do, enforced outside the model
- Evaluations - scoring agent quality and catching regressions
You pay only for the services you use.
Take quiz
Evaluations
Gateway
Browser
Memory
Code Interpreter
Identity
Evaluations
14. What is Amazon Bedrock Agents Classic?
Amazon Bedrock Agents, launched in November 2023, is now called Agents Classic. It's the configuration-first approach: you supply instructions, a foundation model, action groups (Lambda functions or OpenAPI-described APIs) and optionally a knowledge base, and Bedrock runs the reasoning loop for you.
It also offered session and memory settings, guardrail attachment, and multi-agent collaboration with supervisor and collaborator agents.
As of July 30, 2026, Agents Classic is in maintenance mode and closed to new customers. Existing agents keep running and receive fixes and security patches, and AWS hasn't announced an end-of-life date. New agent projects should start on AgentCore.
Take quiz
It is in maintenance mode and closed to new customers since July 30, 2026
It was retired and existing agents stopped in July 2026
It is open only to customers with Provisioned Throughput
Chunking strategies
Cache points
Model units
Action groups
15. What are Amazon Nova models?
Amazon Nova is Amazon's own family of foundation models on Bedrock. It covers several sizes and modalities so you can match cost and quality to the task.
- Nova Micro, Lite, Pro and Premier - understanding models, from cheapest and fastest to most capable
- Nova Canvas and Reel - image and video generation
- Nova Sonic - real-time speech-to-speech
- Nova 2 - the next generation, for example Nova 2 Lite
Nova models can be customized in Bedrock with supervised fine-tuning, reinforcement fine-tuning and distillation, and text prompts get automatic caching on several of them. That makes Nova a common pick for high-volume, cost-sensitive work. Titan is the earlier Amazon family, still used mainly for embeddings.
Take quiz
Nova Canvas
Nova Sonic
Nova Micro
Nova Reel
Only prompt templates, with no weight changes at all
Only continued pre-training on raw text corpora
Supervised fine-tuning, reinforcement fine-tuning and distillation
16. What is Custom Model Import in Amazon Bedrock?
Custom Model Import lets you bring a model you trained or fine-tuned elsewhere into Bedrock and call it through the same serverless APIs. You upload the weights to S3 in a supported open architecture (for example the Llama, Mistral or Qwen families), start an import job, and Bedrock hosts the result.
After the import you invoke it with InvokeModel, or Converse where the architecture supports it, using the imported model's ARN. There are no instances to manage.
Billing is based on active model copies, measured in 5-minute windows, rather than per token. If nobody calls the model for a while it scales down, and the next request can hit a cold start.
Use it when you've already tuned an open-weight model in SageMaker or on your own GPUs and want Bedrock's operating model without retraining.
Take quiz
Per image generated only
By active model copies, measured in 5-minute windows
A flat monthly license fee per imported model
An Amazon S3 location you own
A Bedrock Marketplace endpoint
A Knowledge Base data source
A Guardrails policy version
17. What is Amazon Bedrock Data Automation?
Amazon Bedrock Data Automation (BDA) turns unstructured content into structured output. You feed it documents, images, audio or video, and it returns things like extracted fields, tables, transcripts, summaries and scene descriptions.
You can take the standard output for a general result, or define blueprints that spell out exactly which fields to extract (invoice number, total, due date, for example) and in what format. Blueprints are reusable across a project.
BDA also plugs into Knowledge Bases as a parser, which helps when your PDFs contain tables, charts or diagrams that plain text extraction mangles. It suits intelligent document processing, call analytics and media indexing pipelines.
Take quiz
A denied-topic definition
An inference profile
A blueprint
A cache checkpoint
Extracting structured fields from scanned documents, audio or video
Serving chat completions at the lowest possible latency
Storing embeddings for vector search
18. What is Amazon Bedrock Marketplace?
Amazon Bedrock Marketplace is a catalog of specialized and emerging models (over 100 at launch) that aren't part of the standard serverless lineup. You pick a model, deploy it to an endpoint, and call it with the usual Bedrock APIs.
The difference from serverless models is the capacity and billing model. Marketplace models run on endpoints you configure, choosing the instance type and count, so you pay for the infrastructure while it runs, plus any provider software fee, rather than per token.
Once deployed, the endpoint works with Converse, Guardrails, Knowledge Bases and Agents. Reach for it when you need a niche model, such as a domain-specific or regional-language one, that Bedrock doesn't offer serverless. Don't confuse it with AWS Marketplace subscriptions that some serverless models (like Claude) use for billing.
Take quiz
Per token at the same rate as Nova
For the endpoint infrastructure you run, not per token
Free until one million requests per month
Only the batch inference API
Only the SageMaker console playground
Only the Knowledge Base Retrieve API
The same Bedrock runtime APIs, such as Converse
19. Define model evaluation in Amazon Bedrock?
Model evaluation in Bedrock means measuring how well a model, or a whole RAG pipeline, does on your task before you commit to it. There are four common ways to run it:
- Automatic evaluation - built-in or custom prompt datasets scored with metrics such as accuracy, robustness and toxicity.
- LLM-as-a-judge - a judge model rates responses on qualities like correctness, completeness, helpfulness and harmfulness.
- Human evaluation - your own team or an AWS-managed workforce rates responses.
- RAG evaluation - scores retrieval alone, or retrieval plus generation, against a knowledge base.
Use it to compare candidate models, to catch regressions when you swap a model or edit a prompt, and to justify a cheaper model when quality holds up. In the redesigned console, evaluations are grouped under projects.
Take quiz
Human evaluation
Batch evaluation
LLM-as-a-judge
RAG evaluation
Retrieval quality, and optionally the generated answer too
Only the GPU utilization of the vector store
Only the model's token throughput
20. What is Amazon Bedrock Prompt Management?
Prompt Management lets you create, test, version and run prompts as managed resources instead of strings buried in application code.
You define a prompt with variables, choose the model and inference settings, and compare up to three variants side by side. When you're happy, you create an immutable version and reference it from your app by ARN, so a prompt update doesn't need a code deploy.
Prompt Optimization can rewrite a prompt for a chosen model to improve accuracy or brevity, and you compare the result to the original before saving. Saved prompts can also be reused in Bedrock Flows and agents.
The main benefit is discipline: a change history, side-by-side testing, and the ability to roll back to a known-good version.
Take quiz
By caching it with a one-hour TTL
By referencing a created prompt version
By copying the prompt into every Lambda environment variable
As many as the model has parameters
Exactly one
Up to ten
Up to three
21. What is the difference between InvokeModel and Converse in Amazon Bedrock?
Both call a model synchronously. The difference is how much of the request format you own.
| Aspect | InvokeModel | Converse |
| Request body | Model-specific JSON defined by each provider | One unified schema for messages, system prompt and tools |
| Switching models | Rewrite the payload | Usually change only the modelId |
| Multi-turn chat | You build the history format yourself | A messages list with roles is built in |
| Tools and guardrails | Provider-specific fields | toolConfig and guardrailConfig, same shape for every model |
| Best for | Embedding and image models, model-specific features | Chat and agent-style text and multimodal models |
Converse doesn't expose every model-specific parameter. Use additionalModelRequestFields for those, or fall back to InvokeModel. InvokeModel also remains the route for embeddings and image generation. Both have streaming versions and both accept inference profile IDs.
Take quiz
Converse
ApplyGuardrail
RetrieveAndGenerate
InvokeModel with hand-built provider payloads
Inside the Retrieve API's filter block
Through additionalModelRequestFields
By editing the guardrail version
As a cachePoint block
22. What is the difference between the bedrock-runtime and bedrock-mantle endpoints?
Bedrock exposes several endpoints, and the right one depends on the API style your code already uses.
| Endpoint | What it serves | Typical use |
| bedrock | Control plane: guardrails, inference profiles, customization jobs, Provisioned Throughput | Creating and managing resources |
| bedrock-runtime | InvokeModel, Converse and their streaming variants; for some OpenAI models also Responses and Chat Completions | Native Bedrock inference |
| bedrock-mantle | OpenAI Responses and Chat Completions APIs, plus the Anthropic Messages API | Reusing OpenAI or Anthropic SDK code by changing the base URL and key |
| bedrock-agent-runtime | Knowledge Base Retrieve and RetrieveAndGenerate, agent invocation | RAG and agent calls |
The mantle endpoint adds asynchronous inference for long-running work and stateful conversations, so you don't resend history on every turn. Per the current docs, Anthropic models called through it also skip the first-time-use form.
Feature support differs by endpoint and model, so check the model card, and AWS's current guidance on which endpoint to prefer for new applications, before you commit.
Take quiz
bedrock-agent-runtime
The bedrock control-plane endpoint
bedrock-runtime with InvokeModel only
bedrock-mantle
The bedrock-runtime endpoint used for InvokeModel and Converse
The bedrock-mantle endpoint, which serves OpenAI-style APIs
The bedrock control-plane endpoint
23. How does cross-Region inference work in Amazon Bedrock?
Cross-Region inference (CRIS) lets Bedrock spread requests across several Regions, so one Region's capacity limits matter less. You opt in by passing an inference profile ID as the modelId, such as us.anthropic.claude-... or global.anthropic.claude-....
Your request enters at the source Region, where you call the API. Bedrock then picks a destination Region from those in the profile based on available capacity, runs inference there, and sends the result back. Quotas, logs and CloudTrail events stay tied to the source Region.
flowchart LR
A["App calls Converse with a profile ID in the source Region"] --> B["Bedrock routing layer"]
B --> C{"Choose a destination Region with capacity"}
C --> D["Destination Region 1"]
C --> E["Destination Region 2"]
C --> F["Destination Region 3"]
D --> G["Response returned through the source Region"]
E --> G
F --> G
IAM matters here: the caller needs permission on the inference profile and on the underlying model in every possible destination Region. There's no extra routing charge. Pricing depends on the profile type and model, and for some models global profiles cost less per token than geographic ones.
Take quiz
Pass an inference profile ID as the modelId
Buy a Provisioned Throughput commitment
Create a second bedrock-runtime client in every Region
Add a cachePoint to the request
Only bedrock:ListFoundationModels, because AWS handles the routing
Access to the inference profile and the model in each destination Region
Access to the source Region's model only, since routing is internal
24. When should you choose a global inference profile over a geographic one?
Decide on data residency first, then on throughput.
A geographic profile (prefixes like us., eu. or apac.) routes only inside a defined geography. Choose it when regulation or contract says prompts must be processed within the US, the EU or another boundary.
A global profile (global.) can route to any supported commercial Region. Choose it when you want the highest throughput and best resilience during demand spikes and residency isn't a constraint. For some models it's also cheaper per token.
| Geographic | Global | |
| Routing scope | Regions inside one geography | Any supported commercial Region |
| Best for | Residency and compliance needs | Peak throughput and resilience |
| Price | Model and profile specific | Lower for some models |
One gotcha: destination Regions can include opt-in Regions you never enabled, so Region-restricting SCPs can block calls. Add scoped exceptions, ideally conditioned on bedrock:InferenceProfileArn.
Take quiz
Any profile, because routing never leaves the source Region
A global profile, for maximum throughput
A geographic EU profile
SCPs cannot be applied to Bedrock actions
A destination Region may be one the SCP denies
Global profiles require Provisioned Throughput
Global profiles ignore IAM entirely
25. What is the difference between Bedrock's Standard, Flex, Priority and Reserved tiers?
Bedrock has four service tiers for model inference. You choose one per request with the service_tier parameter, using default, flex, priority or reserved. Leave it out and the request goes to Standard.
| Tier | Trade-off | Use it for |
| Standard | Regular pay-as-you-go price and performance | Everyday workloads |
| Flex | Discounted price, but requests can be slower | Evaluations, summarization, background agent steps |
| Priority | Premium price and preferential compute, up to about 25% better output tokens per second on most supported models | Customer-facing, latency-sensitive traffic |
| Reserved | Reserved tokens-per-minute capacity for 1 or 3 months, fixed monthly price per 1K TPM | Mission-critical steady traffic |
Support isn't universal. Standard works with every model, while Flex, Priority and Reserved depend on the model, so check its card. Standard, Priority and Flex share the model's on-demand quota, whereas Reserved has its own capacity pool. Premiums and discounts vary by model, so read the pricing page instead of assuming a fixed percentage.
Take quiz
It is served on the Standard tier
It is rejected with a ValidationException
It is served on Flex to save money
Reserved
Priority
Flex
Standard with Provisioned Throughput
26. When should you use batch inference instead of on-demand in Amazon Bedrock?
Choose batch inference when nobody is waiting on the answer and the volume is high. Bedrock prices batch jobs on select models at roughly 50% below on-demand.
How it works: you put records in a JSONL file in S3 (each with a recordId and a modelInput), create a job with CreateModelInvocationJob, and Bedrock processes it asynchronously and writes results to an S3 output location. You poll the job status or react to events.
Good fits are classifying millions of tickets, generating embeddings for a corpus, summarizing an archive, running evaluation sets, and producing synthetic training data.
Avoid it for anything interactive. Jobs can take hours and finish times aren't guaranteed to the minute. If you want something in between, the Flex tier gives a discount with the normal on-demand API shape, so no S3 plumbing, but its discount and model support differ from batch.
Take quiz
In the Converse request body
As JSONL files in Amazon S3
In Provisioned Throughput model units
A voice assistant that needs sub-second replies
An agent that must answer inside a 2-second SLA
A live customer-facing chatbot
Classifying two million archived support tickets overnight
27. How does prompt caching work in Amazon Bedrock?
Prompt caching lets Bedrock reuse the already-processed prefix of a prompt so repeated context isn't recomputed. For supported models it can cut input token cost by up to 90% and noticeably reduce latency.
You mark a boundary with a cachePoint block in the Converse API. Everything before it (tools, system prompt, messages) becomes the cacheable prefix. Keep these rules in mind:
- The prefix must be identical between calls. Change one token before the checkpoint and you get a miss.
- Each model has a minimum token count per checkpoint, for example 1,024 for some Claude Sonnet models and more for recent Opus-class ones. Below it the call still works but nothing is cached.
- The TTL is 5 minutes by default and resets on every hit. Some Claude models also offer 1 hour.
- Most models allow up to 4 checkpoints per request.
Watch cacheReadInputTokens and cacheWriteInputTokens in the usage block to confirm it works. Cache writes cost more than normal input tokens, and caches are regional, so cross-Region routing can cause misses. Put static content first and user-specific content last.
Take quiz
Changing any token before the cache checkpoint
Streaming the response instead of waiting for it
Calling from a different IAM role in the same account
Adding new tokens after the checkpoint
The guardrail assessment counts shown in the trace
latencyMs and stopReason from the response metadata
cacheReadInputTokens and cacheWriteInputTokens
28. How can you optimize cost in Amazon Bedrock?
Start with the biggest lever, model choice, then stack the rest.
- Right-size the model. Send simple tasks to a small model such as Nova Micro or a Haiku-class model and keep large models for hard prompts. Intelligent Prompt Routing can automate this within a model family.
- Cache prompts. Long system prompts, tool definitions and shared documents are prime candidates.
- Use batch or Flex for anything non-interactive.
- Trim tokens. Shorter prompts, a sensible
maxTokenscap, and fewer but better chunks in RAG. - Distill or fine-tune. A tuned small model often replaces a large one on a narrow task.
- Compare profiles. Global inference profiles can be cheaper than geographic ones for some models.
- Attribute and monitor. Use application inference profiles or request-level attribution, and watch usage in CloudWatch.
Measure before and after. Cost per successful task matters more than cost per token.
Take quiz
Moving it to a smaller model
Switching every call to the Priority tier
Raising maxTokens to allow longer outputs
Splits one large request across several Regions at once
Sends each request to a cheaper or stronger model in the same family
Compresses long prompts into embeddings before sending them
Caches full responses keyed by a hash of the prompt
29. How does an application inference profile help with cost tracking in Bedrock?
By default it's hard to split Bedrock cost by team or product, because everyone calls the same model ID. Application inference profiles fix that. You create a profile that wraps a model (or a cross-Region profile) and tag it with cost allocation tags such as team or environment.
Your app then calls the profile's ARN as the modelId. Usage in CloudWatch metrics and billing can be filtered by profile and tag, so finance can see who spent what. Because the profile is its own resource, you can also grant access to it through IAM.
In May 2026 Bedrock added request-level usage attribution for InvokeModel and InvokeModelWithResponseStream. It attributes usage to teams, applications, environments or experiments on each request, without provisioning extra resources.
A practical split: use profiles when you want stable per-application resources and IAM boundaries, and request-level attribution when the dimensions change too often to justify a profile for each.
Take quiz
Create one guardrail per application and read its trace
Buy Provisioned Throughput for each application separately
Create a profile per application and tag it with cost allocation tags
A separate billing account per model
Free tokens for any tagged request
Automatic fine-tuning for each team
Per-request usage attribution on the InvokeModel APIs
30. Why should you use Guardrails instead of relying only on system prompts?
A system prompt is an instruction the model may follow, not a rule it must follow. A determined user, or a poisoned document in a RAG pipeline, can push the model off script. A guardrail is an independent enforcement layer that checks content no matter what the model decided to do.
Reasons to add one:
- Defense in depth. It screens input and output, so a jailbreak that fools the model can still be caught on the way out.
- Consistent controls. Word filters, denied topics and PII masking apply the same way on every request.
- One policy, many models. The same guardrail applies when you swap models, and through ApplyGuardrail it can cover non-Bedrock models too.
- Auditability. Traces record which policy triggered and why.
It isn't a silver bullet. Filters have false positives and negatives, they add some latency, and they need tuning against real traffic. Keep good prompts, least-privilege tool access and evaluation in place alongside it.
Take quiz
System prompts are ignored by every hosted model on Bedrock
Guardrails make the underlying model larger and more accurate
The model can still be steered off script, while a guardrail checks content independently
System prompts cannot contain any safety instructions at all
Filters can misfire and need tuning against real traffic
They cannot be attached to agents or Knowledge Bases
They only work in a single Region per AWS account
31. How does a contextual grounding check work in Bedrock Guardrails?
The contextual grounding check targets RAG hallucinations. It scores a model's answer on two axes:
- Grounding - is the answer supported by the source text you provided?
- Relevance - does the answer actually address the user's query?
Each score runs from 0 to 1. You set a threshold for each, and if a score falls below it the guardrail blocks the response. A higher threshold is stricter: at 0.85 a loosely paraphrased answer may be blocked, while at 0.5 it probably passes.
The guardrail needs three labeled inputs: the grounding source (the retrieved passages), the query, and the model content to check. In Converse you tag these with guardContent qualifiers such as grounding_source and query. With Knowledge Bases, the service supplies them.
Start with moderate thresholds, review what gets blocked, then tighten. The check catches unsupported claims, but it can't catch an answer that is faithfully grounded in wrong or outdated source text.
Take quiz
The model retrieves more chunks
More borderline answers get blocked
More answers pass because the check is looser
The relevance score stops being evaluated
The grounding source, the query, and the content to evaluate
The agent alias, the action group and the session ID
The embedding model, the vector index and the reranker
32. What is the difference between Automated Reasoning checks and content filters?
Content filters ask "is this harmful or off-topic?" using statistical classifiers. Automated Reasoning checks ask "is this answer logically consistent with our rules?" using formal logic, so the outcome is a proof-backed finding rather than a confidence score.
The setup: you upload a source document such as an HR or refund policy, Bedrock extracts variables and rules into an Automated Reasoning policy, you test and refine it, and then you attach it to a guardrail. At runtime the checker translates the question and answer into logic and tests them against the rules.
Typical findings are valid (provably consistent), invalid (contradicts a rule, with the broken rule and a suggested fix), satisfiable (could be true or false depending on unstated conditions) and translation ambiguous. An invalid result is what your app uses to trigger a rewrite loop.
It suits rule-heavy domains like insurance, benefits and compliance. It isn't for subjective judgments, and it depends on a precise source policy. Callers need the bedrock:InvokeAutomatedReasoningPolicy permission on the policy version. A June 2026 update added automated refinement workflows that improve a policy against your test cases.
Take quiz
The grounding score fell below the threshold
The response contained profanity
The model timed out before answering
The answer contradicts a rule in the policy
Casual brand-voice chit-chat
Insurance claim eligibility rules
Creative poetry generation
33. Explain the execution flow of a Bedrock Knowledge Base from ingestion to retrieval?
A knowledge base works in two phases: ingestion, which runs when you sync a data source, and retrieval, which runs for every user question.
Ingestion. Bedrock reads the documents, parses them (optionally with Data Automation or a foundation model for tables and images), splits them into chunks, creates embeddings with your chosen model, and writes vectors and metadata to the vector store. Syncs are incremental, so only changed files are reprocessed, but content and permission changes appear only after the next sync.
Retrieval. The question is embedded, the vector store returns the nearest chunks (semantic or hybrid search, with optional metadata filters), and an optional reranker reorders them. Retrieve hands those chunks back as they are, while RetrieveAndGenerate adds them to a prompt and returns an answer with citations.
flowchart LR
subgraph Ingestion
A["Data source"] --> B["Parse"] --> C["Chunk"] --> D["Embed"] --> E[("Vector store")]
end
subgraph Retrieval
Q["User question"] --> R["Embed query"] --> S["Vector or hybrid search"] --> T["Optional rerank"] --> U["Chunks or cited answer"]
end
E --> S
Take quiz
Only after retraining the embedding model
Instantly, the moment you save the file
After the next data source sync
Chunking the documents
Embedding the user's question
Returning citations
Reranking the retrieved passages
34. What is the difference between Retrieve and RetrieveAndGenerate?
Retrieve returns the relevant chunks with scores, source locations and metadata, and stops there. What happens next is up to you.
RetrieveAndGenerate goes further. It retrieves, builds the prompt, calls the foundation model you specify, and returns the answer with citations. It can also keep a session so follow-up questions work.
| Retrieve | RetrieveAndGenerate | |
| Output | Ranked chunks with metadata | Generated answer with citations |
| Prompt assembly | Entirely in your code | Done by Bedrock, with a customizable template |
| Model choice | Any model or API you like | A model supported by the API |
| Code needed | More | Minimal |
Pick Retrieve when you need full control: your own prompt assembly, extra reranking, merging several sources, or a model the other API doesn't support. Pick RetrieveAndGenerate for quick, cited Q&A with little code. For complicated multi-part questions on managed knowledge bases, agentic retrieval is a third option.
Take quiz
ApplyGuardrail
RetrieveAndGenerate
Retrieve
CreateModelInvocationJob
The raw embedding of every chunk
A generated answer with citations
The shard map of the vector index
35. What is the difference between managed and customer-managed knowledge bases?
The difference is who runs the retrieval stack.
| Managed | Customer-managed | |
| Who operates it | Bedrock runs ingestion, storage, indexing and retrieval | You run the vector store and pipeline choices |
| Vector store | Handled and scaled for you | Your pick: S3 Vectors, OpenSearch, Aurora pgvector, Neptune Analytics, Pinecone, Redis or MongoDB Atlas |
| Models | Bedrock selects and maintains embedding model, reranker and generation model by default | You choose the embedding model, parser and chunking |
| Standout features | Native connectors, hybrid search, agentic retrieval | GraphRAG, full pipeline control |
Managed knowledge bases reached general availability on June 17, 2026, and AWS recommends them for most workloads. Choose customer-managed when you need a specific vector store, GraphRAG, cross-cloud portability or full control of the pipeline.
The choice also matters if you're replacing Amazon Kendra, which entered maintenance mode on June 30, 2026 and stopped taking new customers on July 30. Kendra had more connectors than Managed Knowledge Base, so check coverage first and use S3 ingestion for gaps.
Take quiz
Wanting Bedrock to choose the embedding model and reranker
Wanting the agentic retrieval mode
Wanting zero infrastructure decisions
Using a specific vector store such as Aurora pgvector or Pinecone
Use a managed knowledge base
Use customer-managed for every new project
Use Kendra for all new workloads
36. How does agentic retrieval improve answers to multi-part questions?
Standard retrieval runs one similarity search. That works for "what's our refund window?" but struggles with "compare our 2020 and 2023 strategy", because a single query embedding can't represent several intents and the results come back as a blurry average.
Agentic retrieval in Managed Knowledge Bases (the AgenticRetrieveStream API) runs a planning loop driven by a foundation model:
- The model breaks the question into sub-queries.
- It retrieves against each one.
- It judges whether the evidence is enough.
- If not, it refines the queries and loops again.
- It de-duplicates and re-ranks, then writes a grounded answer.
The API streams trace events, each with a step and a status, so you can see what the planner did and why. A speculative retrieval runs before the first planning step to hide some latency. Set generateResponse to false if you only want the retrieved results.
The trade-off is extra model calls, so latency and cost are higher than a plain Retrieve. Use it for comparative, multi-part or multi-hop questions, and stay with Retrieve for simple lookups.
Take quiz
The vector store can't hold data from two different years
Chunks are never allowed to contain dates
One query embedding can't represent several intents well
It turns off the planning loop's streamed trace events
It returns retrieval results without generating an answer
It forces the search to use keyword matching only
It moves the request into asynchronous batch mode
37. When should you use GraphRAG in Bedrock Knowledge Bases?
Use GraphRAG when answers depend on relationships that span documents, and plain vector search keeps returning chunks that are each only partly useful. Typical cases: "which suppliers are affected if this plant shuts down?", root-cause analysis across incident reports, or compliance questions that chain regulations to products.
In Bedrock it runs on Neptune Analytics in a customer-managed knowledge base. You pick an embedding model and a graph-construction model. During ingestion the second model extracts entities and relationships from the chunks and Bedrock builds the graph for you, with no graph modeling on your side. At query time it does the vector search first, then traverses relationships to pull in connected content, which also makes answers easier to explain.
Know the limits. S3 is the only supported data source, there's a default cap of 1,000 files per data source, and ingestion costs more because of the extra model calls. If most of your questions are single-document lookups, plain vector RAG is cheaper and simpler.
Take quiz
Which products are affected if supplier X fails, across many documents?
What time does the on-site cafeteria open on Fridays?
What is the office WiFi password for the visitor network?
Summarize the single memo I just uploaded
Any web crawler or SharePoint connector
Amazon S3 only
Only Amazon RDS tables
38. How can you improve retrieval accuracy in a Bedrock Knowledge Base?
Diagnose first: is the right chunk missing from the results (a retrieval problem), or present but ignored (a generation problem)? Bedrock's RAG evaluation can score the two separately. Then pull the matching lever:
- Fix chunking. Hierarchical chunking helps when small chunks match well but need surrounding context. Semantic chunking helps with narrative text.
- Improve parsing. Use Data Automation or a foundation model parser for tables, charts and scans.
- Add metadata filters. Tag documents (department, year, product) and filter on them, or let a model generate implicit filters from the question.
- Use hybrid search so exact terms like part numbers or error codes still match.
- Rerank. Retrieve more candidates, then let a reranker model pick the best few.
- Reformulate queries. Break complex questions into sub-queries, or move to agentic retrieval.
- Tune embeddings and top-k. Try a stronger embedding model or a different
numberOfResults.
Change one variable at a time against a fixed test set. Otherwise you can't tell what actually helped.
Take quiz
Switching the knowledge base to no chunking at all
Raising the generation temperature for more variety
Lowering the reranker's candidate count to just one
Hybrid search that combines keyword and vector matching
Raise the guardrail's grounding threshold
Immediately fine-tune the generator model
Check whether the right chunk was retrieved at all
39. How can you migrate from Bedrock Agents Classic to AgentCore?
Start with the status. Agents Classic has been closed to new customers since July 30, 2026, but existing agents keep working and AWS hasn't announced an end-of-life date. So this is a planned migration, not an emergency.
| Agents Classic feature | AgentCore equivalent |
| Action groups (Lambda or OpenAPI) | Tools exposed through AgentCore Gateway |
| Knowledge base | Unchanged; reconnect it through Gateway |
| Session and long-term memory | AgentCore Memory |
| Built-in code interpreter action group | AgentCore Code Interpreter |
| Guardrails attached to the agent | Guardrails config plus Policy enforcement on Gateway |
| Configuration-first orchestration | AgentCore managed harness (defined by model, prompt and tools) |
| Custom orchestrator | Your own code on AgentCore Runtime |
| Multi-agent collaboration | Limited: supervisor pattern via agents exposed as tools, or your own framework code |
AWS provides an agent skill that walks you through moving a Classic agent to the harness. It was in preview when announced in April 2026, so confirm its current status. A simple agent (model, action groups, knowledge base) is typically hours of work. Agents with custom orchestration or multi-agent collaboration need real code changes, so tackle those early.
Take quiz
As Guardrails denied topics
As Knowledge Base data sources
As tools behind AgentCore Gateway
As Provisioned Throughput model units
Multi-agent collaboration with supervisor routing
A knowledge base attached to the agent
The agent's instruction prompt
40. How does AgentCore Runtime isolate and scale agent sessions?
AgentCore Runtime hosts your agent code, or an MCP server, in a serverless environment where each session gets its own microVM. Compute, memory and filesystem are isolated per session by the hypervisor, so one user's session can't reach another's. When the session ends, the microVM is terminated.
Key behaviors:
- Long-running sessions. Up to 8 hours per session, which suits multi-step agents that would outgrow short function timeouts.
- Scale to zero. No pre-provisioning, and you pay for active use.
- Protocols. HTTP, MCP and A2A (agent-to-agent).
- Framework freedom. Deploy a container or code built with any framework.
The next-generation Runtime released on September 18, 2026 adds elastic memory that gives back what a session no longer uses, so you pay for actual usage instead of peak, plus consistent cold starts regardless of image size or concurrency.
For work that outlives eight hours, runtime instances run agents on managed EC2 in your account for up to fourteen days, with shared file systems and GPU instance types, while keeping the same AgentCore APIs, identity and observability.
Take quiz
A separate AWS account for every user
A dedicated microVM for each session
A shared container with per-user environment variables
15 minutes
14 days
24 hours
8 hours
41. How does AgentCore Memory work for short-term and long-term context?
Agents forget everything between calls unless you give them memory. AgentCore Memory provides two layers.
Short-term memory stores the raw events of a session, such as user and assistant turns and tool results, so the agent can rebuild context within a conversation.
Long-term memory holds distilled knowledge that persists across sessions. A background pipeline runs memory strategies over the stored events. Built-in strategies, such as semantic facts, summaries and user preferences, use an LLM to extract and consolidate records. If you want full control, a self-managed strategy lets you run your own extraction and consolidation.
Records are organized by namespaces (for example per user or per session) and retrieved with semantic search. Since April 2026 you can also attach structured metadata, such as priority or department, and filter retrieval on it. Indexed metadata keys must be declared when the memory is created and can't be removed later, so plan the schema up front.
A typical use is remembering a customer's preferences and open issues, so the next session starts informed instead of blank.
Take quiz
Distilled facts and preferences that persist across sessions
The IAM role the runtime uses to call tools
The agent's container image and its dependencies
The raw token stream of the current turn, discarded afterwards
Which Region hosts the summarization model used for extraction
The maximum session length in hours
Which metadata keys to index, since they can't be removed later
42. How does AgentCore Gateway expose tools to agents?
Gateway is a managed tool server. Agents connect to one MCP endpoint and discover every tool behind it, so you don't wire each API into each agent.
You add targets and Gateway translates them into MCP tools. Supported targets include:
- REST APIs described with OpenAPI
- Lambda functions
- Existing MCP servers
Security is two-sided. For inbound calls, the agent authenticates to Gateway with OAuth or IAM. For outbound calls, each target has exactly one auth configuration, such as an API key or OAuth client credentials (2LO), with secrets held by AgentCore Identity's credential providers.
With many tools, listing them all bloats the model's context, so Gateway offers semantic tool selection: the agent searches for relevant tools by intent. Gateway can also apply Policy rules to tool calls and emit traces to CloudWatch. With VPC egress support, agents can reach private resources such as an MCP server running on EKS.
Take quiz
A cache checkpoint
A Lambda target
A Knowledge Base chunker
A Flex tier request
Too many tool definitions bloating the model's context
Slow vector index rebuilds after every tool change
Missing IAM permissions on the agent's execution role
43. How does AgentCore Identity let an agent act on a user's behalf?
An agent often needs to call GitHub, Salesforce or Slack as the user, not as itself. AgentCore Identity handles that delegation without tokens living in your agent code.
The flow:
- The agent runs under a workload identity, so it's a first-class principal you can audit.
- A user signs in through your identity provider, and the inbound JWT tells the agent who's calling.
- For a third-party service, Identity runs the OAuth 2.0 three-legged (3LO) flow. The user consents once, and Identity keeps the access and refresh tokens in a secure token vault.
- On later calls the agent asks Identity for a token for that user and service, and Identity returns or refreshes it.
Since September 2026, each Gateway can provide a managed consent portal with its own hosted web client, so you no longer build and run custom OAuth callback infrastructure. For machine-to-machine access, the 2LO client-credentials flow needs no user consent.
Take quiz
In a public S3 bucket
In the agent's system prompt
In a secure token vault managed by the service
In the short-term events of AgentCore Memory
The need for users to sign in at all
The need to host custom OAuth callback infrastructure
The need for IAM roles on Lambda targets
44. How does AgentCore govern production agents using Policy and Observability?
Once an agent is live, two questions matter: "is it allowed to do that?" and "what did it actually do?" AgentCore answers them separately.
Policy defines which actions an agent may take and is enforced at the Gateway, outside the model's own reasoning. That matters because a prompt injection can change what the model wants to do, but it can't rewrite a rule the model never sees. For example, allow a refund tool only below a set amount, or only for a specific user role.
Observability gives end-to-end traces, metrics and dashboards through CloudWatch. It's OpenTelemetry compatible, so you can also send data to Datadog, Dynatrace, Langfuse, LangSmith or Arize Phoenix. You can inspect each model call, tool call and latency inside a session.
Evaluations closes the loop by scoring agent behavior on quality dimensions, so you catch regressions before users do.
Put simply: Policy prevents, Observability explains, and Evaluations measures.
Take quiz
Gateways cache every tool result so rules are checked only once
Gateways run a larger model than the agent, so they reason better
System prompts are invisible to the model during inference
The rule applies outside the model, so a prompt injection can't rewrite it
The JSONL batch format
OpenTelemetry
OAuth 2.0
45. When should you fine-tune a model instead of using RAG or prompt engineering?
They solve different problems, and the classic mistake is using fine-tuning to inject facts.
| Approach | What it changes | Best for | Weak spot |
| Prompt engineering | Instructions and examples | Format, tone, quick fixes | Token cost, limited depth |
| RAG | What the model sees at query time | Fresh, private, citable knowledge | Retrieval quality, added latency |
| Supervised fine-tuning | Weights, from labeled examples | Consistent style, format, narrow tasks | Needs quality data; facts go stale |
| Reinforcement fine-tuning | Weights, from reward signals | Tasks with checkable answers | Needs a reliable reward function |
| Distillation | A small student learns from a large teacher | Cutting cost and latency on a stable task | Capped by the teacher's outputs |
A rule of thumb: if the problem is "the model doesn't know our data", use RAG. If it's "the model knows but doesn't behave the way we need", fine-tune. If it's "this works but costs too much", distill. Start with prompting and RAG because they're cheapest to iterate on, and combine them when it helps, for example a fine-tuned small model plus retrieval.
In Bedrock, customization is available for selected models, and the supported methods differ by model and Region, so check the model's page before planning.
Take quiz
Add the pricing documents through RAG
Distill a smaller student model
Raise the temperature
Fine-tune the model weekly on the pricing text
Semantic chunking
Reinforcement fine-tuning
Distillation
46. How does reinforcement fine-tuning work in Amazon Bedrock?
Reinforcement fine-tuning (RFT) trains a model with scores instead of answers. You give Bedrock prompts and a reward function, and it nudges the model toward whatever the function rewards.
The loop:
- Bedrock samples several responses per prompt from the current model.
- Your reward function scores each one. It can be custom code (usually a Lambda function) or an LLM acting as judge.
- Bedrock updates the model with a policy-based method, Group Relative Policy Optimization (GRPO), favoring higher-scored responses.
- It repeats until the data runs out or you stop at a checkpoint.
Guidance from the docs: start with 100-200 examples and prove the reward function correct before scaling (Nova RFT accepts up to 20K prompts). Test the base model first. If rewards sit near 0%, do supervised fine-tuning first. If they're already above 95%, RFT may be unnecessary. Keep reward functions fast and watch for overfitting, where training reward rises while validation reward falls. The finished model can run on-demand or on Provisioned Throughput.
RFT fits tasks with a checkable outcome: code that passes tests, math, structured extraction, tool-call accuracy. It's a poor fit for subjective quality unless you have a dependable judge. Model support is limited and grows over time, so check which models are eligible.
Take quiz
Human-written reference answers for every single prompt
Vector similarity between responses and a knowledge base
Token counts taken from the usage block
Scores from a reward function, applied through GRPO
Run supervised fine-tuning first to build basic capability
Skip evaluation and deploy immediately
Raise the learning rate and keep going with RFT alone
47. How can you secure data and access in Amazon Bedrock for enterprise workloads?
Think in layers.
- Identity and access. Use least-privilege IAM: scope
bedrock:InvokeModelto specific model and inference profile ARNs, add conditions such asbedrock:InferenceProfileArn, and use SCPs to restrict models or Regions across the organization. With the Model access page gone, IAM and SCPs are your gate. - Network. Use PrivateLink interface VPC endpoints so traffic to Bedrock stays off the public internet, and tighten them with endpoint policies.
- Encryption. Data is encrypted in transit and at rest. Use customer-managed KMS keys for custom models, knowledge bases, agents and evaluation jobs.
- Data handling. Prompts and outputs aren't used to train base models or shared with model providers, and customization data and custom models stay private to your account.
- Content safety. Put Guardrails on inputs and outputs, with PII masking where needed.
- Auditing. CloudTrail records API activity. Model invocation logging can send prompts and responses to CloudWatch Logs or S3. Enable it deliberately, because those logs may hold sensitive text and need their own protection.
Also check data residency: prefer geographic over global cross-Region profiles when the processing location matters.
Take quiz
Enable the Flex tier for all requests
Use PrivateLink interface VPC endpoints
Turn on prompt caching with a one-hour TTL
Attach a denied-topics policy to the guardrail
It disables Guardrails on the account until logging is turned off again
It converts every request into a Provisioned Throughput reservation
The logs can hold sensitive prompts, so they need their own access controls and encryption
48. How do you troubleshoot a ThrottlingException in Amazon Bedrock?
A ThrottlingException (HTTP 429) means you exceeded a quota. Work through it in this order.
- Identify the quota. Requests per minute or tokens per minute, for that model, Region and tier. Service Quotas and the CloudWatch
InvocationThrottlesmetric show which one. - Check maxTokens. Bedrock reserves capacity based on
maxTokensat the start of a request, and for some models output tokens count several times against the TPM quota (a burndown multiplier). An oversized maxTokens can throttle you even when answers are short. Set it close to what you need. - Add retries. Use exponential backoff with jitter. The AWS SDKs have retry modes, and adaptive mode reacts to throttling.
- Spread the load. A cross-Region inference profile gives more headroom, and a queue smooths bursts.
- Move non-urgent work. Send background jobs to batch or Flex so they don't compete with interactive traffic.
- Raise the ceiling. Request a quota increase, or buy Reserved capacity or Provisioned Throughput if traffic is steady.
Don't confuse it with ServiceUnavailableException (capacity or an outage; retry and fail over) or ModelTimeoutException (the request ran too long).
Take quiz
Whether the guardrail has too many denied topics
Whether maxTokens is set far higher than needed
Whether the knowledge base has enough chunks
Never retry; fail the request
Immediate retries in a tight loop
Exponential backoff with jitter
Retrying only after switching to Provisioned Throughput
49. How do you troubleshoot an AccessDeniedException when invoking a Bedrock model?
An AccessDeniedException on InvokeModel or Converse nearly always falls into one of these buckets. Check them in order.
- IAM permissions. The caller needs
bedrock:InvokeModel(and the streaming variant if used) on the right resource. With an inference profile, you need permission on the profile ARN and on the foundation model ARN in every destination Region. Granting only the source Region's model is a classic miss. - Anthropic first-time-use form. It hasn't been submitted for the account or organization.
- Marketplace enablement. For Marketplace-sold models, the first invocation needs AWS Marketplace permissions and a valid payment method.
- SCPs and permission boundaries. A Region-deny SCP can block cross-Region destinations, and a boundary can cap the role.
- Wrong model or Region. The model ID isn't offered in that Region, or you used a base model ID where an inference profile is required.
- Account eligibility. Some models have account-level requirements that the console doesn't show. If everything else checks out, open a support case.
Read the exception text and the CloudTrail event. The denied action and resource ARN usually point straight at the missing statement.
Take quiz
Missing permission on the model in a destination Region
The guardrail's grounding threshold is too high
The prompt exceeded the context window
On the model's pricing page for that Region
In the Knowledge Base sync log for the data source
In the Provisioned Throughput console under model units
In the exception message and the CloudTrail event
50. Explain the execution flow of a guarded RAG request on Amazon Bedrock?
Here's the request path for a production RAG chatbot built on Bedrock services.
sequenceDiagram participant U as User participant A as App backend participant G as ApplyGuardrail participant K as Knowledge Base participant M as Model via Converse U->>A: Question A->>G: Check the input G-->>A: Allowed or blocked A->>K: Retrieve with metadata filters K-->>A: Ranked chunks with sources A->>M: Cached system prompt, chunks, question, guardrailConfig M-->>A: Answer checked by the guardrail A-->>U: Answer with citations
- Screen the input. ApplyGuardrail checks for prompt attacks, denied topics and PII before any retrieval happens.
- Retrieve. Query the knowledge base with metadata filters that reflect the user's permissions, and take the top chunks.
- Build the prompt. Put the static system instructions first with a
cachePoint, then the chunks and question. Tag the chunks as the grounding source and the question as the query. - Generate with guardrails. Call Converse with a
guardrailConfig, so the output is checked for policy, PII and contextual grounding. - Respond. Return the answer with citations, or a safe fallback message if a check blocked it.
- Record. Keep traces, token usage, cache hits and per-application cost tags for later tuning.
Around this core, add retries with backoff for throttling, a fallback model, and RAG evaluation against a fixed test set. For multi-part questions, swap Retrieve for agentic retrieval.