Prev Next

Cloud / Amazon Bedrock Interview questions

Last updated

1. What is Amazon Bedrock? 2. List the model providers available on Amazon Bedrock? 3. How do you invoke a model using the Converse API? 4. How do you get access to foundation models in Amazon Bedrock? 5. Describe the pricing models in Amazon Bedrock? 6. What is Provisioned Throughput in Amazon Bedrock? 7. What is Amazon Bedrock Guardrails? 8. What are the types of safeguards in Amazon Bedrock Guardrails? 9. What is the purpose of the ApplyGuardrail API? 10. What are Knowledge Bases for Amazon Bedrock? 11. What are the chunking strategies in Bedrock Knowledge Bases? 12. What is Amazon Bedrock AgentCore? 13. List the core services of Amazon Bedrock AgentCore? 14. What is Amazon Bedrock Agents Classic? 15. What are Amazon Nova models? 16. What is Custom Model Import in Amazon Bedrock? 17. What is Amazon Bedrock Data Automation? 18. What is Amazon Bedrock Marketplace? 19. Define model evaluation in Amazon Bedrock? 20. What is Amazon Bedrock Prompt Management? 21. What is the difference between InvokeModel and Converse in Amazon Bedrock? 22. What is the difference between the bedrock-runtime and bedrock-mantle endpoints? 23. How does cross-Region inference work in Amazon Bedrock? 24. When should you choose a global inference profile over a geographic one? 25. What is the difference between Bedrock's Standard, Flex, Priority and Reserved tiers? 26. When should you use batch inference instead of on-demand in Amazon Bedrock? 27. How does prompt caching work in Amazon Bedrock? 28. How can you optimize cost in Amazon Bedrock? 29. How does an application inference profile help with cost tracking in Bedrock? 30. Why should you use Guardrails instead of relying only on system prompts? 31. How does a contextual grounding check work in Bedrock Guardrails? 32. What is the difference between Automated Reasoning checks and content filters? 33. Explain the execution flow of a Bedrock Knowledge Base from ingestion to retrieval? 34. What is the difference between Retrieve and RetrieveAndGenerate? 35. What is the difference between managed and customer-managed knowledge bases? 36. How does agentic retrieval improve answers to multi-part questions? 37. When should you use GraphRAG in Bedrock Knowledge Bases? 38. How can you improve retrieval accuracy in a Bedrock Knowledge Base? 39. How can you migrate from Bedrock Agents Classic to AgentCore? 40. How does AgentCore Runtime isolate and scale agent sessions? 41. How does AgentCore Memory work for short-term and long-term context? 42. How does AgentCore Gateway expose tools to agents? 43. How does AgentCore Identity let an agent act on a user's behalf? 44. How does AgentCore govern production agents using Policy and Observability? 45. When should you fine-tune a model instead of using RAG or prompt engineering? 46. How does reinforcement fine-tuning work in Amazon Bedrock? 47. How can you secure data and access in Amazon Bedrock for enterprise workloads? 48. How do you troubleshoot a ThrottlingException in Amazon Bedrock? 49. How do you troubleshoot an AccessDeniedException when invoking a Bedrock model? 50. Explain the execution flow of a guarded RAG request on Amazon Bedrock?

1. What is Amazon Bedrock?

Amazon Bedrock is AWS's fully managed, serverless service for building generative AI applications. You get one API layer over foundation models from several providers, so you can try Claude, Nova, Llama, Mistral or OpenAI models without hosting anything yourself.

Beyond model access, Bedrock bundles the pieces most teams end up needing: Guardrails for safety controls, Knowledge Bases for RAG, model customization (fine-tuning and distillation), evaluation tools, and AgentCore for building and running agents.

Your prompts and outputs aren't used to train the base models and aren't handed to the model providers. Access is governed by IAM, and you can add VPC endpoints (PrivateLink) and KMS encryption. Bedrock became generally available in September 2023.

Take quiz
What best describes Amazon Bedrock's hosting model?
Serverless, with no model infrastructure to provision or patch
It only runs models that Amazon trained itself
You launch and patch GPU instances for every model you call
How are customer prompts and completions treated with respect to the base models?
They are used for training only when Guardrails is turned off
They are used to retrain the base models every night
They are not used to train the base foundation models
They are forwarded to each model provider for product improvement

2. List the model providers available on Amazon Bedrock?

The catalog keeps growing and it differs by Region, but the main providers are:

  • Amazon - Nova (text, image, video, speech) and Titan (mostly embeddings now)
  • Anthropic - Claude
  • Meta - Llama
  • Mistral AI
  • Cohere - Command, Embed and Rerank
  • AI21 Labs - Jamba
  • Stability AI - image generation and editing
  • OpenAI - GPT models, which arrived on Bedrock in 2026
  • DeepSeek and Qwen - open-weight models

Don't hard-code a model list in an application. Check the model catalog in the console, or call ListFoundationModels, for the Region you're deploying to.

Take quiz
Which pair are both Amazon's own model families on Bedrock?
Llama and Mistral
Nova and Titan
Jamba and Command
What is the reliable way to find which models you can call in a Region?
Assume every provider's full lineup is identical in every Region
Read the list off the Bedrock pricing page only
Ask AWS Support to enable the list for your account
Check the model catalog or call ListFoundationModels for that Region

3. How do you invoke a model using the Converse API?

The Converse API gives you one request and response shape for chat-style models, so switching models mostly means changing the modelId. You send messages, an optional system prompt, inference settings and, if needed, tool definitions or a guardrail.

import boto3

client = boto3.client("bedrock-runtime", region_name="us-east-1")

response = client.converse(
    modelId="us.amazon.nova-lite-v1:0",
    system=[{"text": "You answer in two sentences."}],
    messages=[{"role": "user", "content": [{"text": "What is Amazon Bedrock?"}]}],
    inferenceConfig={"maxTokens": 300, "temperature": 0.3},
)

print(response["output"]["message"]["content"][0]["text"])
print(response["usage"])

The usage block reports input, output and total tokens, which is handy for cost tracking. For streaming, call converse_stream instead. The us. prefix above is an inference profile ID, which you use when a model needs cross-Region inference.

Because the request shape is identical across providers, you can keep the modelId in configuration and swap models per environment without touching the code that builds the request.

Take quiz
In a Converse call, which field do you change to point the same code at a different model?
The system block
modelId
The role value inside messages
inferenceConfig.maxTokens
Where do you read the token counts from a Converse response?
The stopReason field
The metrics.latencyMs field
The usage block

4. How do you get access to foundation models in Amazon Bedrock?

Bedrock retired the old Model access page in October 2025. Serverless foundation models are now enabled automatically the first time you invoke them, and you control who can use what through IAM policies and service control policies.

A few exceptions still apply:

  1. Anthropic models need a one-time use-case form, submitted in the console or with the PutUseCaseForModelAccess API. Do it from the organization's management account and member accounts inherit it.
  2. Marketplace-sold models are subscribed automatically on first call, but that first call needs AWS Marketplace permissions and a valid payment method.
  3. The model has to be offered in your Region, either directly or through an inference profile.

If you hit an AccessDeniedException, walk through these three points before opening a support case.

Take quiz
What replaced the manual per-model enablement page?
Automatic enablement on first invocation, governed by IAM and SCPs
A Provisioned Throughput purchase required for each model
A separate Bedrock account type reserved for production use
A per-model approval ticket handled by AWS Support staff
What extra step do first-time Anthropic model users complete?
Request a dedicated VPC and private endpoint from AWS Support
Submit a one-time use-case form, inherited organization-wide via the management account
Purchase a Reserved tier commitment for the Anthropic model family

5. Describe the pricing models in Amazon Bedrock?

Bedrock bills by usage, but the mode you pick changes the price. Text models charge per input and output token, embedding models per input token, and image models per image.

Mode How it works Good for
On-demand (service tiers) Pay per token. Standard is the default, with Flex, Priority and Reserved tiers on supported models Most applications
Batch Asynchronous jobs through S3, about 50% cheaper on select models Bulk, offline work
Provisioned Throughput Reserved model units billed hourly for a chosen term Steady high volume
Customization Charged per token trained, plus monthly storage for the custom model Fine-tuning and distillation

Guardrails policies, Knowledge Base retrieval and AgentCore services are billed separately from model tokens, so include them in any estimate.

Take quiz
Which mode is about 50% cheaper on select models but runs asynchronously?
Provisioned Throughput
Reserved tier
The Priority tier
Batch inference
How are image-generation models billed on demand?
Per image generated
Per GPU hour of a dedicated endpoint
Per input token only

6. What is Provisioned Throughput in Amazon Bedrock?

Provisioned Throughput reserves dedicated capacity for a model in units called model units. Each unit delivers a set amount of input and output tokens per minute, and you pay an hourly rate per unit whether you use it or not.

You choose a term: no commitment (billed hourly, handy for testing), or a longer commitment of one or six months at a lower hourly price. You then call the model with the ARN of the provisioned resource instead of the normal model ID.

It makes sense when traffic is steady enough that reserved capacity beats per-token pricing, or when you need guaranteed throughput. Some customized models also rely on it. Before buying, compare it with the Reserved service tier and with a quota increase on standard on-demand, since fine-tuned Nova models can now run on-demand.

Take quiz
Provisioned Throughput capacity is measured in:
Cache checkpoints
Endpoint replicas
Model units
vCPU hours
After buying Provisioned Throughput, what do you pass as the model identifier?
The ARN of the provisioned model resource
The base model ID, and capacity attaches automatically
The ARN of the calling IAM role

7. What is Amazon Bedrock Guardrails?

Amazon Bedrock Guardrails is a configurable safety layer that checks what goes into a model and what comes back out. You define policies once, version them, and attach the guardrail to model calls, agents or Knowledge Base queries.

When a check trips, the guardrail can block the message and return a canned reply, or mask sensitive data and let the rest through. Each intervention appears in the trace, so you can see which policy fired.

The point is that it works independently of the model's own training. Because it's enforced outside the prompt, a clever user message can't simply tell it to stand down. It also works with models outside Bedrock through the ApplyGuardrail API.

Take quiz
Which sides of a model interaction can a guardrail evaluate?
Only the final model output
Only the system prompt at deployment time
Both user inputs and model outputs
Why can't a user simply talk a guardrail out of its rules?
A guardrail is just a longer system prompt with stricter wording
The policies are enforced outside the prompt, not as model instructions
The guardrail retrains the model after each request it blocks
The guardrail quietly rewrites the system prompt on every turn

8. What are the types of safeguards in Amazon Bedrock Guardrails?

A guardrail is built from independent policies, and you can turn on any mix:

  • Content filters - hate, insults, sexual, violence, misconduct and prompt attacks, each with adjustable strength. They cover text and, for supported categories, images.
  • Denied topics - subjects you describe in plain language, such as investment advice.
  • Word filters - custom words and phrases, plus a managed profanity list.
  • Sensitive information filters - built-in PII types and custom regex, with block or mask actions.
  • Contextual grounding checks - grounding and relevance scores to catch hallucinations in RAG answers.
  • Automated Reasoning checks - logic-based validation of answers against a formal policy.

Prompt attack detection is an input-side control and contextual grounding is output-side. The others can be applied to both directions.

Take quiz
Which safeguard is designed to catch hallucinated answers in a RAG app?
Denied topics for off-limits subjects
Prompt attack filter for jailbreak attempts
Word filter for blocked terms and phrases
Contextual grounding check
Which policy lets you block or mask PII and custom regex patterns?
Content filters
Sensitive information filters
Automated Reasoning checks

9. What is the purpose of the ApplyGuardrail API?

The ApplyGuardrail API evaluates content against a guardrail without calling a model. You send the text (or image) and mark it as an input or an output, and you get back the action taken and the detailed assessments.

That decoupling is the whole point. You can guard models hosted elsewhere, such as SageMaker endpoints, EC2, on-premises servers or third-party APIs. You can also check content at different stages of a pipeline, for example screening a question before retrieval and validating the answer afterwards.

It's also a quick way to test policy changes on their own. When you call a Bedrock model directly, APIs like Converse accept a guardrailConfig that performs the same check inline.

Take quiz
What makes ApplyGuardrail different from attaching a guardrail to a Converse call?
It fine-tunes the guardrail on whatever content you send it
It only works with Amazon Nova models hosted on Bedrock
It checks content without invoking a model, so it works with any model
When calling ApplyGuardrail, what do you specify along with the content?
Whether the content is an INPUT or an OUTPUT
The chunking strategy
The model's tokens-per-minute quota
The Provisioned Throughput ARN

10. What are Knowledge Bases for Amazon Bedrock?

Knowledge Bases for Amazon Bedrock is the managed RAG feature. It connects a foundation model to your own documents so answers are based on your data and come with citations back to the source.

The service handles the plumbing: pulling content from sources like S3, SharePoint or Confluence, parsing and chunking it, creating embeddings, and storing them in a vector index. At query time it embeds the question, finds the closest chunks, and either returns them or passes them to a model that writes the answer.

AWS now offers two types. In a managed knowledge base, Bedrock runs storage and retrieval for you. In a customer-managed one, you pick and operate the vector store yourself.

Take quiz
What does a knowledge base return alongside generated answers?
A certificate guaranteeing the answer is correct
Citations pointing to the source chunks
A fine-tuned copy of the model
Which step happens at query time in a knowledge base?
Retraining the base model on the new documents
Rebuilding every chunk from the data source
Compressing the vector index into model weights
Embedding the user's question and searching the vector index

11. What are the chunking strategies in Bedrock Knowledge Bases?

Chunking controls how documents are split before embedding, and it has a big effect on answer quality. Bedrock offers these options:

  1. Fixed-size - you set a maximum token count and an overlap. Simple and predictable.
  2. Hierarchical - small child chunks are matched, but the larger parent chunk is returned for context. You set both sizes and the overlap.
  3. Semantic - splits on meaning instead of length, so related sentences stay together. It adds ingestion cost.
  4. No chunking - each file is treated as one chunk, useful when you've already split the content.
  5. Custom - a Lambda function runs your own chunking logic.

Hierarchical chunking isn't recommended with an S3 vector bucket because of metadata size limits.

Take quiz
Which strategy matches a small child chunk but returns a larger parent chunk?
Hierarchical chunking
No chunking
Custom Lambda chunking
Fixed-size chunking
When does the 'no chunking' option make sense?
Your documents are very long unstructured PDFs
You want the service to detect topic boundaries for you
Your files are already split into retrieval-sized pieces

12. What is Amazon Bedrock AgentCore?

Amazon Bedrock AgentCore is AWS's platform for building, deploying and operating AI agents in production. It's modular, so you can use one piece or all of them, and it's framework and model agnostic: LangGraph, CrewAI, Strands, LlamaIndex, the OpenAI Agents SDK or plain Python all work, with models inside or outside Bedrock.

Think of it as the infrastructure around the agent rather than the agent logic itself: a secure place to run it, memory, a tool gateway, identity, observability and governance controls.

It targets a gap teams keep hitting. A prototype runs fine on a laptop, but production needs session isolation, credential handling, tracing and access controls. AgentCore provides those as managed services billed on consumption.

Take quiz
Which statement about AgentCore is accurate?
It works with any agent framework and any model
It only supports models hosted on Bedrock
It only runs agents written with one AWS-specific framework
What problem is AgentCore mainly aimed at?
Replacing vector databases for enterprise search
Training foundation models from scratch on customer data
Moving agents from prototype to secure, observable production
Translating documents in bulk across many languages

13. List the core services of Amazon Bedrock AgentCore?

AgentCore is a set of services you can mix and match:

  • Runtime - serverless, session-isolated hosting for agents and MCP servers
  • Memory - short-term session events and long-term memory across sessions
  • Gateway - turns APIs, Lambda functions and existing MCP servers into agent tools
  • Identity - workload identities, OAuth flows and a token vault for acting on a user's behalf
  • Code Interpreter - a sandbox where the agent can run code
  • Browser - a managed browser the agent can drive
  • Observability - traces, metrics and dashboards, OpenTelemetry compatible
  • Policy - rules on what an agent is allowed to do, enforced outside the model
  • Evaluations - scoring agent quality and catching regressions

You pay only for the services you use.

Take quiz
Which AgentCore service turns Lambda functions and REST APIs into MCP tools?
Evaluations
Gateway
Browser
Memory
Which AgentCore service scores how well an agent is performing?
Code Interpreter
Identity
Evaluations

14. What is Amazon Bedrock Agents Classic?

Amazon Bedrock Agents, launched in November 2023, is now called Agents Classic. It's the configuration-first approach: you supply instructions, a foundation model, action groups (Lambda functions or OpenAPI-described APIs) and optionally a knowledge base, and Bedrock runs the reasoning loop for you.

It also offered session and memory settings, guardrail attachment, and multi-agent collaboration with supervisor and collaborator agents.

As of July 30, 2026, Agents Classic is in maintenance mode and closed to new customers. Existing agents keep running and receive fixes and security patches, and AWS hasn't announced an end-of-life date. New agent projects should start on AgentCore.

Take quiz
What is the status of Agents Classic for brand-new customers?
It is in maintenance mode and closed to new customers since July 30, 2026
It was retired and existing agents stopped in July 2026
It is open only to customers with Provisioned Throughput
What do you attach to a Classic agent so it can call your business APIs?
Chunking strategies
Cache points
Model units
Action groups

15. What are Amazon Nova models?

Amazon Nova is Amazon's own family of foundation models on Bedrock. It covers several sizes and modalities so you can match cost and quality to the task.

  • Nova Micro, Lite, Pro and Premier - understanding models, from cheapest and fastest to most capable
  • Nova Canvas and Reel - image and video generation
  • Nova Sonic - real-time speech-to-speech
  • Nova 2 - the next generation, for example Nova 2 Lite

Nova models can be customized in Bedrock with supervised fine-tuning, reinforcement fine-tuning and distillation, and text prompts get automatic caching on several of them. That makes Nova a common pick for high-volume, cost-sensitive work. Titan is the earlier Amazon family, still used mainly for embeddings.

Take quiz
Which Nova model targets real-time speech-to-speech conversations?
Nova Canvas
Nova Sonic
Nova Micro
Nova Reel
Which customization methods can Bedrock run on Nova models?
Only prompt templates, with no weight changes at all
Only continued pre-training on raw text corpora
Supervised fine-tuning, reinforcement fine-tuning and distillation

16. What is Custom Model Import in Amazon Bedrock?

Custom Model Import lets you bring a model you trained or fine-tuned elsewhere into Bedrock and call it through the same serverless APIs. You upload the weights to S3 in a supported open architecture (for example the Llama, Mistral or Qwen families), start an import job, and Bedrock hosts the result.

After the import you invoke it with InvokeModel, or Converse where the architecture supports it, using the imported model's ARN. There are no instances to manage.

Billing is based on active model copies, measured in 5-minute windows, rather than per token. If nobody calls the model for a while it scales down, and the next request can hit a cold start.

Use it when you've already tuned an open-weight model in SageMaker or on your own GPUs and want Bedrock's operating model without retraining.

Take quiz
How is Custom Model Import billed?
Per image generated only
By active model copies, measured in 5-minute windows
A flat monthly license fee per imported model
Where do the model weights come from for an import job?
An Amazon S3 location you own
A Bedrock Marketplace endpoint
A Knowledge Base data source
A Guardrails policy version

17. What is Amazon Bedrock Data Automation?

Amazon Bedrock Data Automation (BDA) turns unstructured content into structured output. You feed it documents, images, audio or video, and it returns things like extracted fields, tables, transcripts, summaries and scene descriptions.

You can take the standard output for a general result, or define blueprints that spell out exactly which fields to extract (invoice number, total, due date, for example) and in what format. Blueprints are reusable across a project.

BDA also plugs into Knowledge Bases as a parser, which helps when your PDFs contain tables, charts or diagrams that plain text extraction mangles. It suits intelligent document processing, call analytics and media indexing pipelines.

Take quiz
In BDA, what tells the service which fields to pull out of an invoice?
A denied-topic definition
An inference profile
A blueprint
A cache checkpoint
Which job fits BDA best?
Extracting structured fields from scanned documents, audio or video
Serving chat completions at the lowest possible latency
Storing embeddings for vector search

18. What is Amazon Bedrock Marketplace?

Amazon Bedrock Marketplace is a catalog of specialized and emerging models (over 100 at launch) that aren't part of the standard serverless lineup. You pick a model, deploy it to an endpoint, and call it with the usual Bedrock APIs.

The difference from serverless models is the capacity and billing model. Marketplace models run on endpoints you configure, choosing the instance type and count, so you pay for the infrastructure while it runs, plus any provider software fee, rather than per token.

Once deployed, the endpoint works with Converse, Guardrails, Knowledge Bases and Agents. Reach for it when you need a niche model, such as a domain-specific or regional-language one, that Bedrock doesn't offer serverless. Don't confuse it with AWS Marketplace subscriptions that some serverless models (like Claude) use for billing.

Take quiz
How are Bedrock Marketplace models typically billed?
Per token at the same rate as Nova
For the endpoint infrastructure you run, not per token
Free until one million requests per month
Which APIs can call a deployed Bedrock Marketplace model?
Only the batch inference API
Only the SageMaker console playground
Only the Knowledge Base Retrieve API
The same Bedrock runtime APIs, such as Converse

19. Define model evaluation in Amazon Bedrock?

Model evaluation in Bedrock means measuring how well a model, or a whole RAG pipeline, does on your task before you commit to it. There are four common ways to run it:

  • Automatic evaluation - built-in or custom prompt datasets scored with metrics such as accuracy, robustness and toxicity.
  • LLM-as-a-judge - a judge model rates responses on qualities like correctness, completeness, helpfulness and harmfulness.
  • Human evaluation - your own team or an AWS-managed workforce rates responses.
  • RAG evaluation - scores retrieval alone, or retrieval plus generation, against a knowledge base.

Use it to compare candidate models, to catch regressions when you swap a model or edit a prompt, and to justify a cheaper model when quality holds up. In the redesigned console, evaluations are grouped under projects.

Take quiz
Which method has a judge model rate responses for qualities like helpfulness?
Human evaluation
Batch evaluation
LLM-as-a-judge
RAG evaluation
What does RAG evaluation measure?
Retrieval quality, and optionally the generated answer too
Only the GPU utilization of the vector store
Only the model's token throughput

20. What is Amazon Bedrock Prompt Management?

Prompt Management lets you create, test, version and run prompts as managed resources instead of strings buried in application code.

You define a prompt with variables, choose the model and inference settings, and compare up to three variants side by side. When you're happy, you create an immutable version and reference it from your app by ARN, so a prompt update doesn't need a code deploy.

Prompt Optimization can rewrite a prompt for a chosen model to improve accuracy or brevity, and you compare the result to the original before saving. Saved prompts can also be reused in Bedrock Flows and agents.

The main benefit is discipline: a change history, side-by-side testing, and the ability to roll back to a known-good version.

Take quiz
How does an application pin a specific revision of a prompt?
By caching it with a one-hour TTL
By referencing a created prompt version
By copying the prompt into every Lambda environment variable
How many prompt variants can you compare side by side in Prompt Management?
As many as the model has parameters
Exactly one
Up to ten
Up to three

21. What is the difference between InvokeModel and Converse in Amazon Bedrock?

Both call a model synchronously. The difference is how much of the request format you own.

Aspect InvokeModel Converse
Request body Model-specific JSON defined by each provider One unified schema for messages, system prompt and tools
Switching models Rewrite the payload Usually change only the modelId
Multi-turn chat You build the history format yourself A messages list with roles is built in
Tools and guardrails Provider-specific fields toolConfig and guardrailConfig, same shape for every model
Best for Embedding and image models, model-specific features Chat and agent-style text and multimodal models

Converse doesn't expose every model-specific parameter. Use additionalModelRequestFields for those, or fall back to InvokeModel. InvokeModel also remains the route for embeddings and image generation. Both have streaming versions and both accept inference profile IDs.

Take quiz
You want to swap Claude for Nova with minimal code changes. Which API helps most?
Converse
ApplyGuardrail
RetrieveAndGenerate
InvokeModel with hand-built provider payloads
How do you pass a model-specific parameter that Converse doesn't model natively?
Inside the Retrieve API's filter block
Through additionalModelRequestFields
By editing the guardrail version
As a cachePoint block

22. What is the difference between the bedrock-runtime and bedrock-mantle endpoints?

Bedrock exposes several endpoints, and the right one depends on the API style your code already uses.

Endpoint What it serves Typical use
bedrock Control plane: guardrails, inference profiles, customization jobs, Provisioned Throughput Creating and managing resources
bedrock-runtime InvokeModel, Converse and their streaming variants; for some OpenAI models also Responses and Chat Completions Native Bedrock inference
bedrock-mantle OpenAI Responses and Chat Completions APIs, plus the Anthropic Messages API Reusing OpenAI or Anthropic SDK code by changing the base URL and key
bedrock-agent-runtime Knowledge Base Retrieve and RetrieveAndGenerate, agent invocation RAG and agent calls

The mantle endpoint adds asynchronous inference for long-running work and stateful conversations, so you don't resend history on every turn. Per the current docs, Anthropic models called through it also skip the first-time-use form.

Feature support differs by endpoint and model, so check the model card, and AWS's current guidance on which endpoint to prefer for new applications, before you commit.

Take quiz
An app already uses the OpenAI SDK. Which Bedrock endpoint lets you keep it by changing only the base URL and key?
bedrock-agent-runtime
The bedrock control-plane endpoint
bedrock-runtime with InvokeModel only
bedrock-mantle
Which endpoint do you call to create a guardrail or a customization job?
The bedrock-runtime endpoint used for InvokeModel and Converse
The bedrock-mantle endpoint, which serves OpenAI-style APIs
The bedrock control-plane endpoint

23. How does cross-Region inference work in Amazon Bedrock?

Cross-Region inference (CRIS) lets Bedrock spread requests across several Regions, so one Region's capacity limits matter less. You opt in by passing an inference profile ID as the modelId, such as us.anthropic.claude-... or global.anthropic.claude-....

Your request enters at the source Region, where you call the API. Bedrock then picks a destination Region from those in the profile based on available capacity, runs inference there, and sends the result back. Quotas, logs and CloudTrail events stay tied to the source Region.

flowchart LR
  A["App calls Converse with a profile ID in the source Region"] --> B["Bedrock routing layer"]
  B --> C{"Choose a destination Region with capacity"}
  C --> D["Destination Region 1"]
  C --> E["Destination Region 2"]
  C --> F["Destination Region 3"]
  D --> G["Response returned through the source Region"]
  E --> G
  F --> G

IAM matters here: the caller needs permission on the inference profile and on the underlying model in every possible destination Region. There's no extra routing charge. Pricing depends on the profile type and model, and for some models global profiles cost less per token than geographic ones.

Take quiz
What do you change in your code to use cross-Region inference?
Pass an inference profile ID as the modelId
Buy a Provisioned Throughput commitment
Create a second bedrock-runtime client in every Region
Add a cachePoint to the request
Which IAM permissions must the caller hold for a cross-Region call?
Only bedrock:ListFoundationModels, because AWS handles the routing
Access to the inference profile and the model in each destination Region
Access to the source Region's model only, since routing is internal

24. When should you choose a global inference profile over a geographic one?

Decide on data residency first, then on throughput.

A geographic profile (prefixes like us., eu. or apac.) routes only inside a defined geography. Choose it when regulation or contract says prompts must be processed within the US, the EU or another boundary.

A global profile (global.) can route to any supported commercial Region. Choose it when you want the highest throughput and best resilience during demand spikes and residency isn't a constraint. For some models it's also cheaper per token.

Geographic Global
Routing scope Regions inside one geography Any supported commercial Region
Best for Residency and compliance needs Peak throughput and resilience
Price Model and profile specific Lower for some models

One gotcha: destination Regions can include opt-in Regions you never enabled, so Region-restricting SCPs can block calls. Add scoped exceptions, ideally conditioned on bedrock:InferenceProfileArn.

Take quiz
Compliance says prompts must stay processed within the EU. Which profile fits?
Any profile, because routing never leaves the source Region
A global profile, for maximum throughput
A geographic EU profile
Why might an SCP break a call made with a global profile?
SCPs cannot be applied to Bedrock actions
A destination Region may be one the SCP denies
Global profiles require Provisioned Throughput
Global profiles ignore IAM entirely

25. What is the difference between Bedrock's Standard, Flex, Priority and Reserved tiers?

Bedrock has four service tiers for model inference. You choose one per request with the service_tier parameter, using default, flex, priority or reserved. Leave it out and the request goes to Standard.

Tier Trade-off Use it for
Standard Regular pay-as-you-go price and performance Everyday workloads
Flex Discounted price, but requests can be slower Evaluations, summarization, background agent steps
Priority Premium price and preferential compute, up to about 25% better output tokens per second on most supported models Customer-facing, latency-sensitive traffic
Reserved Reserved tokens-per-minute capacity for 1 or 3 months, fixed monthly price per 1K TPM Mission-critical steady traffic

Support isn't universal. Standard works with every model, while Flex, Priority and Reserved depend on the model, so check its card. Standard, Priority and Flex share the model's on-demand quota, whereas Reserved has its own capacity pool. Premiums and discounts vary by model, so read the pricing page instead of assuming a fixed percentage.

Take quiz
What happens to a request that has no service_tier parameter?
It is served on the Standard tier
It is rejected with a ValidationException
It is served on Flex to save money
Which tier suits a nightly evaluation run that can tolerate slower responses?
Reserved
Priority
Flex
Standard with Provisioned Throughput

26. When should you use batch inference instead of on-demand in Amazon Bedrock?

Choose batch inference when nobody is waiting on the answer and the volume is high. Bedrock prices batch jobs on select models at roughly 50% below on-demand.

How it works: you put records in a JSONL file in S3 (each with a recordId and a modelInput), create a job with CreateModelInvocationJob, and Bedrock processes it asynchronously and writes results to an S3 output location. You poll the job status or react to events.

Good fits are classifying millions of tickets, generating embeddings for a corpus, summarizing an archive, running evaluation sets, and producing synthetic training data.

Avoid it for anything interactive. Jobs can take hours and finish times aren't guaranteed to the minute. If you want something in between, the Flex tier gives a discount with the normal on-demand API shape, so no S3 plumbing, but its discount and model support differ from batch.

Take quiz
Where do batch inference inputs and outputs live?
In the Converse request body
As JSONL files in Amazon S3
In Provisioned Throughput model units
Which workload is best suited to batch inference?
A voice assistant that needs sub-second replies
An agent that must answer inside a 2-second SLA
A live customer-facing chatbot
Classifying two million archived support tickets overnight

27. How does prompt caching work in Amazon Bedrock?

Prompt caching lets Bedrock reuse the already-processed prefix of a prompt so repeated context isn't recomputed. For supported models it can cut input token cost by up to 90% and noticeably reduce latency.

You mark a boundary with a cachePoint block in the Converse API. Everything before it (tools, system prompt, messages) becomes the cacheable prefix. Keep these rules in mind:

  1. The prefix must be identical between calls. Change one token before the checkpoint and you get a miss.
  2. Each model has a minimum token count per checkpoint, for example 1,024 for some Claude Sonnet models and more for recent Opus-class ones. Below it the call still works but nothing is cached.
  3. The TTL is 5 minutes by default and resets on every hit. Some Claude models also offer 1 hour.
  4. Most models allow up to 4 checkpoints per request.

Watch cacheReadInputTokens and cacheWriteInputTokens in the usage block to confirm it works. Cache writes cost more than normal input tokens, and caches are regional, so cross-Region routing can cause misses. Put static content first and user-specific content last.

Take quiz
What invalidates a cached prompt prefix?
Changing any token before the cache checkpoint
Streaming the response instead of waiting for it
Calling from a different IAM role in the same account
Adding new tokens after the checkpoint
Which usage fields tell you whether caching actually worked?
The guardrail assessment counts shown in the trace
latencyMs and stopReason from the response metadata
cacheReadInputTokens and cacheWriteInputTokens

28. How can you optimize cost in Amazon Bedrock?

Start with the biggest lever, model choice, then stack the rest.

  1. Right-size the model. Send simple tasks to a small model such as Nova Micro or a Haiku-class model and keep large models for hard prompts. Intelligent Prompt Routing can automate this within a model family.
  2. Cache prompts. Long system prompts, tool definitions and shared documents are prime candidates.
  3. Use batch or Flex for anything non-interactive.
  4. Trim tokens. Shorter prompts, a sensible maxTokens cap, and fewer but better chunks in RAG.
  5. Distill or fine-tune. A tuned small model often replaces a large one on a narrow task.
  6. Compare profiles. Global inference profiles can be cheaper than geographic ones for some models.
  7. Attribute and monitor. Use application inference profiles or request-level attribution, and watch usage in CloudWatch.

Measure before and after. Cost per successful task matters more than cost per token.

Take quiz
Which change usually saves the most on a high-volume, simple classification task?
Moving it to a smaller model
Switching every call to the Priority tier
Raising maxTokens to allow longer outputs
What does Intelligent Prompt Routing do?
Splits one large request across several Regions at once
Sends each request to a cheaper or stronger model in the same family
Compresses long prompts into embeddings before sending them
Caches full responses keyed by a hash of the prompt

29. How does an application inference profile help with cost tracking in Bedrock?

By default it's hard to split Bedrock cost by team or product, because everyone calls the same model ID. Application inference profiles fix that. You create a profile that wraps a model (or a cross-Region profile) and tag it with cost allocation tags such as team or environment.

Your app then calls the profile's ARN as the modelId. Usage in CloudWatch metrics and billing can be filtered by profile and tag, so finance can see who spent what. Because the profile is its own resource, you can also grant access to it through IAM.

In May 2026 Bedrock added request-level usage attribution for InvokeModel and InvokeModelWithResponseStream. It attributes usage to teams, applications, environments or experiments on each request, without provisioning extra resources.

A practical split: use profiles when you want stable per-application resources and IAM boundaries, and request-level attribution when the dimensions change too often to justify a profile for each.

Take quiz
How do you track spend per application with application inference profiles?
Create one guardrail per application and read its trace
Buy Provisioned Throughput for each application separately
Create a profile per application and tag it with cost allocation tags
What did the May 2026 update add for cost visibility?
A separate billing account per model
Free tokens for any tagged request
Automatic fine-tuning for each team
Per-request usage attribution on the InvokeModel APIs

30. Why should you use Guardrails instead of relying only on system prompts?

A system prompt is an instruction the model may follow, not a rule it must follow. A determined user, or a poisoned document in a RAG pipeline, can push the model off script. A guardrail is an independent enforcement layer that checks content no matter what the model decided to do.

Reasons to add one:

  1. Defense in depth. It screens input and output, so a jailbreak that fools the model can still be caught on the way out.
  2. Consistent controls. Word filters, denied topics and PII masking apply the same way on every request.
  3. One policy, many models. The same guardrail applies when you swap models, and through ApplyGuardrail it can cover non-Bedrock models too.
  4. Auditability. Traces record which policy triggered and why.

It isn't a silver bullet. Filters have false positives and negatives, they add some latency, and they need tuning against real traffic. Keep good prompts, least-privilege tool access and evaluation in place alongside it.

Take quiz
Why isn't a strong system prompt a full replacement for a guardrail?
System prompts are ignored by every hosted model on Bedrock
Guardrails make the underlying model larger and more accurate
The model can still be steered off script, while a guardrail checks content independently
System prompts cannot contain any safety instructions at all
Which is a genuine limitation of guardrails?
Filters can misfire and need tuning against real traffic
They cannot be attached to agents or Knowledge Bases
They only work in a single Region per AWS account

31. How does a contextual grounding check work in Bedrock Guardrails?

The contextual grounding check targets RAG hallucinations. It scores a model's answer on two axes:

  • Grounding - is the answer supported by the source text you provided?
  • Relevance - does the answer actually address the user's query?

Each score runs from 0 to 1. You set a threshold for each, and if a score falls below it the guardrail blocks the response. A higher threshold is stricter: at 0.85 a loosely paraphrased answer may be blocked, while at 0.5 it probably passes.

The guardrail needs three labeled inputs: the grounding source (the retrieved passages), the query, and the model content to check. In Converse you tag these with guardContent qualifiers such as grounding_source and query. With Knowledge Bases, the service supplies them.

Start with moderate thresholds, review what gets blocked, then tighten. The check catches unsupported claims, but it can't catch an answer that is faithfully grounded in wrong or outdated source text.

Take quiz
A grounding threshold is raised from 0.5 to 0.85. What is the effect?
The model retrieves more chunks
More borderline answers get blocked
More answers pass because the check is looser
The relevance score stops being evaluated
Which three inputs does the contextual grounding check need?
The grounding source, the query, and the content to evaluate
The agent alias, the action group and the session ID
The embedding model, the vector index and the reranker

32. What is the difference between Automated Reasoning checks and content filters?

Content filters ask "is this harmful or off-topic?" using statistical classifiers. Automated Reasoning checks ask "is this answer logically consistent with our rules?" using formal logic, so the outcome is a proof-backed finding rather than a confidence score.

The setup: you upload a source document such as an HR or refund policy, Bedrock extracts variables and rules into an Automated Reasoning policy, you test and refine it, and then you attach it to a guardrail. At runtime the checker translates the question and answer into logic and tests them against the rules.

Typical findings are valid (provably consistent), invalid (contradicts a rule, with the broken rule and a suggested fix), satisfiable (could be true or false depending on unstated conditions) and translation ambiguous. An invalid result is what your app uses to trigger a rewrite loop.

It suits rule-heavy domains like insurance, benefits and compliance. It isn't for subjective judgments, and it depends on a precise source policy. Callers need the bedrock:InvokeAutomatedReasoningPolicy permission on the policy version. A June 2026 update added automated refinement workflows that improve a policy against your test cases.

Take quiz
What does an INVALID finding from Automated Reasoning checks mean?
The grounding score fell below the threshold
The response contained profanity
The model timed out before answering
The answer contradicts a rule in the policy
Which domain is the best fit for Automated Reasoning checks?
Casual brand-voice chit-chat
Insurance claim eligibility rules
Creative poetry generation

33. Explain the execution flow of a Bedrock Knowledge Base from ingestion to retrieval?

A knowledge base works in two phases: ingestion, which runs when you sync a data source, and retrieval, which runs for every user question.

Ingestion. Bedrock reads the documents, parses them (optionally with Data Automation or a foundation model for tables and images), splits them into chunks, creates embeddings with your chosen model, and writes vectors and metadata to the vector store. Syncs are incremental, so only changed files are reprocessed, but content and permission changes appear only after the next sync.

Retrieval. The question is embedded, the vector store returns the nearest chunks (semantic or hybrid search, with optional metadata filters), and an optional reranker reorders them. Retrieve hands those chunks back as they are, while RetrieveAndGenerate adds them to a prompt and returns an answer with citations.

flowchart LR
  subgraph Ingestion
    A["Data source"] --> B["Parse"] --> C["Chunk"] --> D["Embed"] --> E[("Vector store")]
  end
  subgraph Retrieval
    Q["User question"] --> R["Embed query"] --> S["Vector or hybrid search"] --> T["Optional rerank"] --> U["Chunks or cited answer"]
  end
  E --> S
Take quiz
When do edited source documents become searchable in a knowledge base?
Only after retraining the embedding model
Instantly, the moment you save the file
After the next data source sync
Which step happens during ingestion but not at query time?
Chunking the documents
Embedding the user's question
Returning citations
Reranking the retrieved passages

34. What is the difference between Retrieve and RetrieveAndGenerate?

Retrieve returns the relevant chunks with scores, source locations and metadata, and stops there. What happens next is up to you.

RetrieveAndGenerate goes further. It retrieves, builds the prompt, calls the foundation model you specify, and returns the answer with citations. It can also keep a session so follow-up questions work.

Retrieve RetrieveAndGenerate
Output Ranked chunks with metadata Generated answer with citations
Prompt assembly Entirely in your code Done by Bedrock, with a customizable template
Model choice Any model or API you like A model supported by the API
Code needed More Minimal

Pick Retrieve when you need full control: your own prompt assembly, extra reranking, merging several sources, or a model the other API doesn't support. Pick RetrieveAndGenerate for quick, cited Q&A with little code. For complicated multi-part questions on managed knowledge bases, agentic retrieval is a third option.

Take quiz
Which API returns chunks only and leaves generation to your own code?
ApplyGuardrail
RetrieveAndGenerate
Retrieve
CreateModelInvocationJob
What does RetrieveAndGenerate return that Retrieve doesn't?
The raw embedding of every chunk
A generated answer with citations
The shard map of the vector index

35. What is the difference between managed and customer-managed knowledge bases?

The difference is who runs the retrieval stack.

Managed Customer-managed
Who operates it Bedrock runs ingestion, storage, indexing and retrieval You run the vector store and pipeline choices
Vector store Handled and scaled for you Your pick: S3 Vectors, OpenSearch, Aurora pgvector, Neptune Analytics, Pinecone, Redis or MongoDB Atlas
Models Bedrock selects and maintains embedding model, reranker and generation model by default You choose the embedding model, parser and chunking
Standout features Native connectors, hybrid search, agentic retrieval GraphRAG, full pipeline control

Managed knowledge bases reached general availability on June 17, 2026, and AWS recommends them for most workloads. Choose customer-managed when you need a specific vector store, GraphRAG, cross-cloud portability or full control of the pipeline.

The choice also matters if you're replacing Amazon Kendra, which entered maintenance mode on June 30, 2026 and stopped taking new customers on July 30. Kendra had more connectors than Managed Knowledge Base, so check coverage first and use S3 ingestion for gaps.

Take quiz
Which need points you toward a customer-managed knowledge base?
Wanting Bedrock to choose the embedding model and reranker
Wanting the agentic retrieval mode
Wanting zero infrastructure decisions
Using a specific vector store such as Aurora pgvector or Pinecone
What is AWS's stated recommendation for most workloads?
Use a managed knowledge base
Use customer-managed for every new project
Use Kendra for all new workloads

36. How does agentic retrieval improve answers to multi-part questions?

Standard retrieval runs one similarity search. That works for "what's our refund window?" but struggles with "compare our 2020 and 2023 strategy", because a single query embedding can't represent several intents and the results come back as a blurry average.

Agentic retrieval in Managed Knowledge Bases (the AgenticRetrieveStream API) runs a planning loop driven by a foundation model:

  1. The model breaks the question into sub-queries.
  2. It retrieves against each one.
  3. It judges whether the evidence is enough.
  4. If not, it refines the queries and loops again.
  5. It de-duplicates and re-ranks, then writes a grounded answer.

The API streams trace events, each with a step and a status, so you can see what the planner did and why. A speculative retrieval runs before the first planning step to hide some latency. Set generateResponse to false if you only want the retrieved results.

The trade-off is extra model calls, so latency and cost are higher than a plain Retrieve. Use it for comparative, multi-part or multi-hop questions, and stay with Retrieve for simple lookups.

Take quiz
Why does single-shot retrieval struggle with 'compare our 2020 and 2023 strategy'?
The vector store can't hold data from two different years
Chunks are never allowed to contain dates
One query embedding can't represent several intents well
What does setting generateResponse to false do in agentic retrieval?
It turns off the planning loop's streamed trace events
It returns retrieval results without generating an answer
It forces the search to use keyword matching only
It moves the request into asynchronous batch mode

37. When should you use GraphRAG in Bedrock Knowledge Bases?

Use GraphRAG when answers depend on relationships that span documents, and plain vector search keeps returning chunks that are each only partly useful. Typical cases: "which suppliers are affected if this plant shuts down?", root-cause analysis across incident reports, or compliance questions that chain regulations to products.

In Bedrock it runs on Neptune Analytics in a customer-managed knowledge base. You pick an embedding model and a graph-construction model. During ingestion the second model extracts entities and relationships from the chunks and Bedrock builds the graph for you, with no graph modeling on your side. At query time it does the vector search first, then traverses relationships to pull in connected content, which also makes answers easier to explain.

Know the limits. S3 is the only supported data source, there's a default cap of 1,000 files per data source, and ingestion costs more because of the extra model calls. If most of your questions are single-document lookups, plain vector RAG is cheaper and simpler.

Take quiz
Which question benefits most from GraphRAG?
Which products are affected if supplier X fails, across many documents?
What time does the on-site cafeteria open on Fridays?
What is the office WiFi password for the visitor network?
Summarize the single memo I just uploaded
Which data source does Bedrock GraphRAG support?
Any web crawler or SharePoint connector
Amazon S3 only
Only Amazon RDS tables

38. How can you improve retrieval accuracy in a Bedrock Knowledge Base?

Diagnose first: is the right chunk missing from the results (a retrieval problem), or present but ignored (a generation problem)? Bedrock's RAG evaluation can score the two separately. Then pull the matching lever:

  1. Fix chunking. Hierarchical chunking helps when small chunks match well but need surrounding context. Semantic chunking helps with narrative text.
  2. Improve parsing. Use Data Automation or a foundation model parser for tables, charts and scans.
  3. Add metadata filters. Tag documents (department, year, product) and filter on them, or let a model generate implicit filters from the question.
  4. Use hybrid search so exact terms like part numbers or error codes still match.
  5. Rerank. Retrieve more candidates, then let a reranker model pick the best few.
  6. Reformulate queries. Break complex questions into sub-queries, or move to agentic retrieval.
  7. Tune embeddings and top-k. Try a stronger embedding model or a different numberOfResults.

Change one variable at a time against a fixed test set. Otherwise you can't tell what actually helped.

Take quiz
Part numbers and error codes keep getting missed by semantic search. What helps most?
Switching the knowledge base to no chunking at all
Raising the generation temperature for more variety
Lowering the reranker's candidate count to just one
Hybrid search that combines keyword and vector matching
What is the first diagnostic step for a bad RAG answer?
Raise the guardrail's grounding threshold
Immediately fine-tune the generator model
Check whether the right chunk was retrieved at all

39. How can you migrate from Bedrock Agents Classic to AgentCore?

Start with the status. Agents Classic has been closed to new customers since July 30, 2026, but existing agents keep working and AWS hasn't announced an end-of-life date. So this is a planned migration, not an emergency.

Agents Classic feature AgentCore equivalent
Action groups (Lambda or OpenAPI) Tools exposed through AgentCore Gateway
Knowledge base Unchanged; reconnect it through Gateway
Session and long-term memory AgentCore Memory
Built-in code interpreter action group AgentCore Code Interpreter
Guardrails attached to the agent Guardrails config plus Policy enforcement on Gateway
Configuration-first orchestration AgentCore managed harness (defined by model, prompt and tools)
Custom orchestrator Your own code on AgentCore Runtime
Multi-agent collaboration Limited: supervisor pattern via agents exposed as tools, or your own framework code

AWS provides an agent skill that walks you through moving a Classic agent to the harness. It was in preview when announced in April 2026, so confirm its current status. A simple agent (model, action groups, knowledge base) is typically hours of work. Agents with custom orchestration or multi-agent collaboration need real code changes, so tackle those early.

Take quiz
Where do Classic action groups usually end up after migration?
As Guardrails denied topics
As Knowledge Base data sources
As tools behind AgentCore Gateway
As Provisioned Throughput model units
Which Classic capability is the least straightforward to migrate?
Multi-agent collaboration with supervisor routing
A knowledge base attached to the agent
The agent's instruction prompt

40. How does AgentCore Runtime isolate and scale agent sessions?

AgentCore Runtime hosts your agent code, or an MCP server, in a serverless environment where each session gets its own microVM. Compute, memory and filesystem are isolated per session by the hypervisor, so one user's session can't reach another's. When the session ends, the microVM is terminated.

Key behaviors:

  • Long-running sessions. Up to 8 hours per session, which suits multi-step agents that would outgrow short function timeouts.
  • Scale to zero. No pre-provisioning, and you pay for active use.
  • Protocols. HTTP, MCP and A2A (agent-to-agent).
  • Framework freedom. Deploy a container or code built with any framework.

The next-generation Runtime released on September 18, 2026 adds elastic memory that gives back what a session no longer uses, so you pay for actual usage instead of peak, plus consistent cold starts regardless of image size or concurrency.

For work that outlives eight hours, runtime instances run agents on managed EC2 in your account for up to fourteen days, with shared file systems and GPU instance types, while keeping the same AgentCore APIs, identity and observability.

Take quiz
What isolates one user's agent session from another in AgentCore Runtime?
A separate AWS account for every user
A dedicated microVM for each session
A shared container with per-user environment variables
What is the standard maximum duration of a microVM session?
15 minutes
14 days
24 hours
8 hours

41. How does AgentCore Memory work for short-term and long-term context?

Agents forget everything between calls unless you give them memory. AgentCore Memory provides two layers.

Short-term memory stores the raw events of a session, such as user and assistant turns and tool results, so the agent can rebuild context within a conversation.

Long-term memory holds distilled knowledge that persists across sessions. A background pipeline runs memory strategies over the stored events. Built-in strategies, such as semantic facts, summaries and user preferences, use an LLM to extract and consolidate records. If you want full control, a self-managed strategy lets you run your own extraction and consolidation.

Records are organized by namespaces (for example per user or per session) and retrieved with semantic search. Since April 2026 you can also attach structured metadata, such as priority or department, and filter retrieval on it. Indexed metadata keys must be declared when the memory is created and can't be removed later, so plan the schema up front.

A typical use is remembering a customer's preferences and open issues, so the next session starts informed instead of blank.

Take quiz
What does long-term memory hold that short-term memory does not?
Distilled facts and preferences that persist across sessions
The IAM role the runtime uses to call tools
The agent's container image and its dependencies
The raw token stream of the current turn, discarded afterwards
What must you decide up front to use metadata filtering on long-term memory?
Which Region hosts the summarization model used for extraction
The maximum session length in hours
Which metadata keys to index, since they can't be removed later

42. How does AgentCore Gateway expose tools to agents?

Gateway is a managed tool server. Agents connect to one MCP endpoint and discover every tool behind it, so you don't wire each API into each agent.

You add targets and Gateway translates them into MCP tools. Supported targets include:

  • REST APIs described with OpenAPI
  • Lambda functions
  • Existing MCP servers

Security is two-sided. For inbound calls, the agent authenticates to Gateway with OAuth or IAM. For outbound calls, each target has exactly one auth configuration, such as an API key or OAuth client credentials (2LO), with secrets held by AgentCore Identity's credential providers.

With many tools, listing them all bloats the model's context, so Gateway offers semantic tool selection: the agent searches for relevant tools by intent. Gateway can also apply Policy rules to tool calls and emit traces to CloudWatch. With VPC egress support, agents can reach private resources such as an MCP server running on EKS.

Take quiz
What do you add to Gateway to expose a Lambda function as an agent tool?
A cache checkpoint
A Lambda target
A Knowledge Base chunker
A Flex tier request
What problem does semantic tool selection solve?
Too many tool definitions bloating the model's context
Slow vector index rebuilds after every tool change
Missing IAM permissions on the agent's execution role

43. How does AgentCore Identity let an agent act on a user's behalf?

An agent often needs to call GitHub, Salesforce or Slack as the user, not as itself. AgentCore Identity handles that delegation without tokens living in your agent code.

The flow:

  1. The agent runs under a workload identity, so it's a first-class principal you can audit.
  2. A user signs in through your identity provider, and the inbound JWT tells the agent who's calling.
  3. For a third-party service, Identity runs the OAuth 2.0 three-legged (3LO) flow. The user consents once, and Identity keeps the access and refresh tokens in a secure token vault.
  4. On later calls the agent asks Identity for a token for that user and service, and Identity returns or refreshes it.

Since September 2026, each Gateway can provide a managed consent portal with its own hosted web client, so you no longer build and run custom OAuth callback infrastructure. For machine-to-machine access, the 2LO client-credentials flow needs no user consent.

Take quiz
Where does AgentCore Identity keep refresh tokens for delegated access?
In a public S3 bucket
In the agent's system prompt
In a secure token vault managed by the service
In the short-term events of AgentCore Memory
What does the managed consent portal remove?
The need for users to sign in at all
The need to host custom OAuth callback infrastructure
The need for IAM roles on Lambda targets

44. How does AgentCore govern production agents using Policy and Observability?

Once an agent is live, two questions matter: "is it allowed to do that?" and "what did it actually do?" AgentCore answers them separately.

Policy defines which actions an agent may take and is enforced at the Gateway, outside the model's own reasoning. That matters because a prompt injection can change what the model wants to do, but it can't rewrite a rule the model never sees. For example, allow a refund tool only below a set amount, or only for a specific user role.

Observability gives end-to-end traces, metrics and dashboards through CloudWatch. It's OpenTelemetry compatible, so you can also send data to Datadog, Dynatrace, Langfuse, LangSmith or Arize Phoenix. You can inspect each model call, tool call and latency inside a session.

Evaluations closes the loop by scoring agent behavior on quality dimensions, so you catch regressions before users do.

Put simply: Policy prevents, Observability explains, and Evaluations measures.

Take quiz
Why is enforcing policy at the Gateway safer than relying on the system prompt?
Gateways cache every tool result so rules are checked only once
Gateways run a larger model than the agent, so they reason better
System prompts are invisible to the model during inference
The rule applies outside the model, so a prompt injection can't rewrite it
Which standard makes AgentCore telemetry portable to tools like Datadog or Langfuse?
The JSONL batch format
OpenTelemetry
OAuth 2.0

45. When should you fine-tune a model instead of using RAG or prompt engineering?

They solve different problems, and the classic mistake is using fine-tuning to inject facts.

Approach What it changes Best for Weak spot
Prompt engineering Instructions and examples Format, tone, quick fixes Token cost, limited depth
RAG What the model sees at query time Fresh, private, citable knowledge Retrieval quality, added latency
Supervised fine-tuning Weights, from labeled examples Consistent style, format, narrow tasks Needs quality data; facts go stale
Reinforcement fine-tuning Weights, from reward signals Tasks with checkable answers Needs a reliable reward function
Distillation A small student learns from a large teacher Cutting cost and latency on a stable task Capped by the teacher's outputs

A rule of thumb: if the problem is "the model doesn't know our data", use RAG. If it's "the model knows but doesn't behave the way we need", fine-tune. If it's "this works but costs too much", distill. Start with prompting and RAG because they're cheapest to iterate on, and combine them when it helps, for example a fine-tuned small model plus retrieval.

In Bedrock, customization is available for selected models, and the supported methods differ by model and Region, so check the model's page before planning.

Take quiz
Your assistant doesn't know this month's pricing changes. What's the best first step?
Add the pricing documents through RAG
Distill a smaller student model
Raise the temperature
Fine-tune the model weekly on the pricing text
Which technique transfers a large teacher's behavior into a cheaper student model?
Semantic chunking
Reinforcement fine-tuning
Distillation

46. How does reinforcement fine-tuning work in Amazon Bedrock?

Reinforcement fine-tuning (RFT) trains a model with scores instead of answers. You give Bedrock prompts and a reward function, and it nudges the model toward whatever the function rewards.

The loop:

  1. Bedrock samples several responses per prompt from the current model.
  2. Your reward function scores each one. It can be custom code (usually a Lambda function) or an LLM acting as judge.
  3. Bedrock updates the model with a policy-based method, Group Relative Policy Optimization (GRPO), favoring higher-scored responses.
  4. It repeats until the data runs out or you stop at a checkpoint.

Guidance from the docs: start with 100-200 examples and prove the reward function correct before scaling (Nova RFT accepts up to 20K prompts). Test the base model first. If rewards sit near 0%, do supervised fine-tuning first. If they're already above 95%, RFT may be unnecessary. Keep reward functions fast and watch for overfitting, where training reward rises while validation reward falls. The finished model can run on-demand or on Provisioned Throughput.

RFT fits tasks with a checkable outcome: code that passes tests, math, structured extraction, tool-call accuracy. It's a poor fit for subjective quality unless you have a dependable judge. Model support is limited and grows over time, so check which models are eligible.

Take quiz
What does Bedrock use to update the model during reinforcement fine-tuning?
Human-written reference answers for every single prompt
Vector similarity between responses and a knowledge base
Token counts taken from the usage block
Scores from a reward function, applied through GRPO
Baseline rewards are almost always 0%. What do the docs suggest?
Run supervised fine-tuning first to build basic capability
Skip evaluation and deploy immediately
Raise the learning rate and keep going with RFT alone

47. How can you secure data and access in Amazon Bedrock for enterprise workloads?

Think in layers.

  1. Identity and access. Use least-privilege IAM: scope bedrock:InvokeModel to specific model and inference profile ARNs, add conditions such as bedrock:InferenceProfileArn, and use SCPs to restrict models or Regions across the organization. With the Model access page gone, IAM and SCPs are your gate.
  2. Network. Use PrivateLink interface VPC endpoints so traffic to Bedrock stays off the public internet, and tighten them with endpoint policies.
  3. Encryption. Data is encrypted in transit and at rest. Use customer-managed KMS keys for custom models, knowledge bases, agents and evaluation jobs.
  4. Data handling. Prompts and outputs aren't used to train base models or shared with model providers, and customization data and custom models stay private to your account.
  5. Content safety. Put Guardrails on inputs and outputs, with PII masking where needed.
  6. Auditing. CloudTrail records API activity. Model invocation logging can send prompts and responses to CloudWatch Logs or S3. Enable it deliberately, because those logs may hold sensitive text and need their own protection.

Also check data residency: prefer geographic over global cross-Region profiles when the processing location matters.

Take quiz
How do you keep Bedrock API traffic off the public internet?
Enable the Flex tier for all requests
Use PrivateLink interface VPC endpoints
Turn on prompt caching with a one-hour TTL
Attach a denied-topics policy to the guardrail
What should you consider before enabling model invocation logging?
It disables Guardrails on the account until logging is turned off again
It converts every request into a Provisioned Throughput reservation
The logs can hold sensitive prompts, so they need their own access controls and encryption

48. How do you troubleshoot a ThrottlingException in Amazon Bedrock?

A ThrottlingException (HTTP 429) means you exceeded a quota. Work through it in this order.

  1. Identify the quota. Requests per minute or tokens per minute, for that model, Region and tier. Service Quotas and the CloudWatch InvocationThrottles metric show which one.
  2. Check maxTokens. Bedrock reserves capacity based on maxTokens at the start of a request, and for some models output tokens count several times against the TPM quota (a burndown multiplier). An oversized maxTokens can throttle you even when answers are short. Set it close to what you need.
  3. Add retries. Use exponential backoff with jitter. The AWS SDKs have retry modes, and adaptive mode reacts to throttling.
  4. Spread the load. A cross-Region inference profile gives more headroom, and a queue smooths bursts.
  5. Move non-urgent work. Send background jobs to batch or Flex so they don't compete with interactive traffic.
  6. Raise the ceiling. Request a quota increase, or buy Reserved capacity or Provisioned Throughput if traffic is steady.

Don't confuse it with ServiceUnavailableException (capacity or an outage; retry and fail over) or ModelTimeoutException (the request ran too long).

Take quiz
Answers are short but throttling persists at a low request rate. What should you check first?
Whether the guardrail has too many denied topics
Whether maxTokens is set far higher than needed
Whether the knowledge base has enough chunks
Which retry approach is recommended for throttled Bedrock calls?
Never retry; fail the request
Immediate retries in a tight loop
Exponential backoff with jitter
Retrying only after switching to Provisioned Throughput

49. How do you troubleshoot an AccessDeniedException when invoking a Bedrock model?

An AccessDeniedException on InvokeModel or Converse nearly always falls into one of these buckets. Check them in order.

  1. IAM permissions. The caller needs bedrock:InvokeModel (and the streaming variant if used) on the right resource. With an inference profile, you need permission on the profile ARN and on the foundation model ARN in every destination Region. Granting only the source Region's model is a classic miss.
  2. Anthropic first-time-use form. It hasn't been submitted for the account or organization.
  3. Marketplace enablement. For Marketplace-sold models, the first invocation needs AWS Marketplace permissions and a valid payment method.
  4. SCPs and permission boundaries. A Region-deny SCP can block cross-Region destinations, and a boundary can cap the role.
  5. Wrong model or Region. The model ID isn't offered in that Region, or you used a base model ID where an inference profile is required.
  6. Account eligibility. Some models have account-level requirements that the console doesn't show. If everything else checks out, open a support case.

Read the exception text and the CloudTrail event. The denied action and resource ARN usually point straight at the missing statement.

Take quiz
A cross-Region call fails even though the source Region's model is allowed. What's the likely cause?
Missing permission on the model in a destination Region
The guardrail's grounding threshold is too high
The prompt exceeded the context window
Where can you usually find the exact denied action and resource ARN?
On the model's pricing page for that Region
In the Knowledge Base sync log for the data source
In the Provisioned Throughput console under model units
In the exception message and the CloudTrail event

50. Explain the execution flow of a guarded RAG request on Amazon Bedrock?

Here's the request path for a production RAG chatbot built on Bedrock services.

sequenceDiagram
  participant U as User
  participant A as App backend
  participant G as ApplyGuardrail
  participant K as Knowledge Base
  participant M as Model via Converse
  U->>A: Question
  A->>G: Check the input
  G-->>A: Allowed or blocked
  A->>K: Retrieve with metadata filters
  K-->>A: Ranked chunks with sources
  A->>M: Cached system prompt, chunks, question, guardrailConfig
  M-->>A: Answer checked by the guardrail
  A-->>U: Answer with citations
  1. Screen the input. ApplyGuardrail checks for prompt attacks, denied topics and PII before any retrieval happens.
  2. Retrieve. Query the knowledge base with metadata filters that reflect the user's permissions, and take the top chunks.
  3. Build the prompt. Put the static system instructions first with a cachePoint, then the chunks and question. Tag the chunks as the grounding source and the question as the query.
  4. Generate with guardrails. Call Converse with a guardrailConfig, so the output is checked for policy, PII and contextual grounding.
  5. Respond. Return the answer with citations, or a safe fallback message if a check blocked it.
  6. Record. Keep traces, token usage, cache hits and per-application cost tags for later tuning.

Around this core, add retries with backoff for throttling, a fallback model, and RAG evaluation against a fixed test set. For multi-part questions, swap Retrieve for agentic retrieval.

Take quiz
Why is the static system prompt placed before the retrieved chunks?
Because Converse rejects prompts that start with retrieved text
Because chunks must always come last for billing reasons
So the identical prefix can be cached with a cachePoint
So the guardrail can skip reading it
Which input plays the grounding source role in the contextual grounding check?
The retrieved chunks
The user's original question
The model's previous answer
«
»

Comments & Discussions