Prev Next

Cloud / Amazon SageMaker Interview questions

Last updated

1. What is Amazon SageMaker AI? 2. What are the main stages of an ML workflow in SageMaker? 3. What is SageMaker Studio? 4. What is a SageMaker training job? 5. What is a SageMaker endpoint? 6. List the built-in algorithms available in SageMaker? 7. What is SageMaker JumpStart? 8. What is a SageMaker Processing job? 9. What is SageMaker Canvas? 10. What is SageMaker Ground Truth? 11. What is SageMaker Feature Store? 12. What is the SageMaker Model Registry? 13. What is SageMaker Pipelines? 14. What is SageMaker Automatic Model Tuning? 15. What is SageMaker Autopilot? 16. What is SageMaker Clarify? 17. What is SageMaker Model Monitor? 18. List the inference options available in SageMaker? 19. What is the SageMaker execution role? 20. How do you launch a training job with the SageMaker Python SDK? 21. What is SageMaker HyperPod? 22. What is the difference between File, Pipe, and FastFile input modes? 23. What is the difference between SageMaker AI and SageMaker Unified Studio? 24. How does SageMaker Managed Spot Training work? 25. Explain the lifecycle of a SageMaker training job? 26. What is the difference between real-time, serverless, and asynchronous inference? 27. When should you use Batch Transform instead of a real-time endpoint? 28. Why use multi-model endpoints in SageMaker? 29. How does auto scaling work for SageMaker real-time endpoints? 30. How do inference components differ from multi-model endpoints? 31. What happens when you deploy a model to a SageMaker endpoint? 32. How do production variants enable A/B testing and shadow testing? 33. How do SageMaker deployment guardrails work? 34. What is the difference between online and offline stores in SageMaker Feature Store? 35. How does SageMaker Model Monitor detect data drift? 36. How does SageMaker Clarify detect bias and explain predictions? 37. How do you build a CI/CD workflow with SageMaker Pipelines and Model Registry? 38. How do Bayesian and Hyperband tuning strategies differ in SageMaker? 39. What is the difference between data parallelism and model parallelism in SageMaker? 40. How does SageMaker HyperPod recover from hardware failures? 41. When should you choose HyperPod over SageMaker training jobs? 42. How do you secure SageMaker workloads inside a VPC? 43. How do you bring your own container to SageMaker? 44. How can you optimize SageMaker costs? 45. How do you troubleshoot a SageMaker endpoint that fails to deploy? 46. How do you troubleshoot a failed SageMaker training job? 47. How do you deploy and optimize large language models on SageMaker? 48. Which is better for generative AI: SageMaker or Amazon Bedrock, and why? 49. How does serverless model customization work in SageMaker AI? 50. How do you track experiments in SageMaker with managed MLflow?

1. What is Amazon SageMaker AI?

Amazon SageMaker AI is AWS's fully managed service for building, training, and deploying machine learning models, including foundation models. You bring the data and the code; AWS handles provisioning, scaling, and shutting down the compute underneath.

It launched in November 2017 as "Amazon SageMaker". In December 2024 AWS reused the SageMaker name for a broader data-and-AI platform built around Unified Studio, and the original ML service became SageMaker AI.

The service covers the whole ML lifecycle: labeling (Ground Truth), data prep (Processing jobs, Canvas), training and tuning, a model registry, several hosting options (real-time, serverless, asynchronous, batch), and monitoring. You pay for the compute and storage you use, and most jobs are billed per second.

Take quiz
What happened to the original Amazon SageMaker service name in December 2024?
It was retired and replaced by Amazon Bedrock
It was renamed SageMaker Studio Classic
It became SageMaker AI, while SageMaker grew into a broader data and AI platform
How is most SageMaker training compute billed?
Per second for the instances a job runs on
A flat monthly license per data scientist
Per token generated by the trained model

2. What are the main stages of an ML workflow in SageMaker?

A typical SageMaker workflow has six stages, and each one has a dedicated capability:

  1. Label and prepare data - Ground Truth for labeling, Processing jobs and Canvas for cleaning and feature engineering, Feature Store for reusable features.
  2. Explore and build - Studio notebooks, JumpStart pretrained models, or Autopilot for AutoML.
  3. Train and tune - training jobs with built-in algorithms or your own containers, plus Automatic Model Tuning.
  4. Evaluate and register - Clarify for bias and explainability checks, then Model Registry for versioned approval.
  5. Deploy - real-time, serverless, or asynchronous endpoints, or Batch Transform.
  6. Monitor - Model Monitor watches data quality and drift and can trigger retraining.

SageMaker Pipelines can stitch all of these stages into one automated, repeatable workflow.

Take quiz
Which capability catalogs model versions and tracks their approval before deployment?
Model Registry
Ground Truth
Model Monitor
Where do Ground Truth and Feature Store fit in the ML workflow?
Endpoint deployment
Data labeling and preparation
Hyperparameter tuning

3. What is SageMaker Studio?

SageMaker Studio is the web-based IDE for SageMaker AI. From one browser tab you can write code, launch training jobs, track experiments, deploy endpoints, and manage pipelines.

The current experience is built on domains. An admin creates a domain and adds users through IAM or IAM Identity Center, and each user works in spaces that have their own compute and storage. Inside a space you pick an app such as JupyterLab, Code Editor (based on Code-OSS), or RStudio. The no-code Canvas app lives in the same domain.

The older Studio Classic interface is the legacy option, and AWS steers new projects to the newer Studio. Turn on idle shutdown for your apps so a forgotten notebook does not keep billing you.

Take quiz
Which IDE options are available inside the current SageMaker Studio?
Eclipse and NetBeans
JupyterLab, Code Editor, and RStudio
Only the legacy Jupyter Classic notebook
What top-level admin construct do Studio users belong to?
A model package group
A feature group
A domain

4. What is a SageMaker training job?

A training job is a managed, one-off run of your training code on SageMaker-provisioned instances. You call CreateTrainingJob (directly or through the SDK) and describe four things: the container image or algorithm, the input data channels in S3, the instance type and count, and an S3 output path.

SageMaker then starts the instances, pulls the image, copies the data over, runs training, uploads the resulting model.tar.gz to S3, and terminates the compute. You are billed per second for the time the instances ran, so nothing keeps costing money after the job ends.

Logs stream to CloudWatch, and any metric your code prints can be captured with a regex and graphed.

Take quiz
What does SageMaker write to the S3 output path when a training job succeeds?
A running HTTPS endpoint
The Docker image used for training
The model artifacts, packaged as model.tar.gz
By default, what happens to the training instances after a job finishes?
SageMaker terminates them, so billing stops
They keep running idle until the next job is queued
They become a permanent inference endpoint

5. What is a SageMaker endpoint?

A SageMaker endpoint is a managed HTTPS service that hosts a trained model for real-time predictions. Clients call it through InvokeEndpoint on the SageMaker Runtime API, and SageMaker load-balances requests across the instances behind it.

Creating one takes three resources, in this order:

  1. Model - the inference container image, the S3 location of the model artifacts, and an execution role.
  2. Endpoint configuration - one or more production variants, each with an instance type, instance count, and traffic weight.
  3. Endpoint - the live resource built from that configuration.

The instances run, and bill, for as long as the endpoint exists, so delete endpoints you no longer need.

Take quiz
Which resource defines the instance type and count behind an endpoint?
The endpoint configuration (production variants)
The model artifact tarball
The execution role
Which API do clients call to get a real-time prediction?
CreateTrainingJob
InvokeEndpoint on SageMaker Runtime
DescribeEndpointConfig

6. List the built-in algorithms available in SageMaker?

SageMaker ships with built-in algorithms you can train without writing model code. They are grouped by task:

Category Algorithms
Supervised (tabular) XGBoost, Linear Learner, k-NN, Factorization Machines
Unsupervised and anomaly detection K-Means, PCA, Random Cut Forest, IP Insights, LDA, NTM
Time series DeepAR
Text and sequences BlazingText, Seq2Seq, Object2Vec
Computer vision Image Classification, Object Detection, Semantic Segmentation

In practice XGBoost is the workhorse for tabular data. When nothing on the list fits, use a framework container (PyTorch, TensorFlow, Hugging Face) or bring your own image.

Take quiz
Which built-in algorithm is designed for probabilistic time series forecasting?
Object2Vec
DeepAR
PCA
Which built-in algorithm is aimed at detecting anomalies in data streams?
Image Classification
Seq2Seq
Random Cut Forest

7. What is SageMaker JumpStart?

SageMaker JumpStart is a model hub inside SageMaker AI with hundreds of pretrained models that you can deploy or fine-tune with a few clicks or a few lines of code. Models come from open-source hubs like Hugging Face and from model providers, and cover both foundation models and classic vision and NLP models.

Beyond models, it offers prebuilt solution templates (fraud detection and demand forecasting, for example) and example notebooks. Deploying a model gives you an endpoint running an optimized container, and fine-tuning runs as a managed training job on your own data.

Organizations can also create a private hub to curate exactly which models their teams are allowed to use.

Take quiz
What is a private hub in JumpStart used for?
Storing encrypted training data only
Hosting the SageMaker control plane
Curating the set of models a team is allowed to use
What can you do with a JumpStart pretrained model?
Deploy it to an endpoint or fine-tune it on your own data
Only read its model card without running it
Only download it to a laptop

8. What is a SageMaker Processing job?

A Processing job runs a script or container on managed instances for work that sits around training: data cleaning, feature engineering, and model evaluation. It is the SageMaker way to run pandas, scikit-learn, or Spark code at scale without keeping a server around.

SageMaker copies data from S3 into the container at /opt/ml/processing/input, runs your code, and uploads whatever it writes under /opt/ml/processing/output back to S3. Then the instances shut down and billing stops.

You can use the built-in scikit-learn or Spark processors or supply your own image. Processing jobs are also the engine underneath Clarify and Model Monitor.

Take quiz
Where does a processing container read its input data from by default?
/opt/ml/processing/input
/opt/ml/model
/opt/ml/checkpoints
Which SageMaker features run on top of Processing jobs?
Ground Truth and JumpStart
Clarify and Model Monitor
Studio Classic and Notebook Instances

9. What is SageMaker Canvas?

SageMaker Canvas is a visual, no-code workspace that lets business analysts build ML models and generate predictions without writing code. You import data from sources such as S3, Redshift, or Snowflake and work through a point-and-click interface.

The usual path is to prepare data with visual transformations, choose a target column, and let Canvas run AutoML to train and rank candidate models. It supports tabular prediction, time-series forecasting, and image and text tasks, and it can also chat with foundation models.

When an analyst needs help, a model can be shared with data scientists in Studio, who can inspect or extend it. That handoff between analysts and data scientists is a big part of its appeal.

Take quiz
Who is SageMaker Canvas primarily designed for?
Kernel developers writing custom CUDA code
Business analysts who want ML without writing code
Network engineers managing VPC routing
How does Canvas train candidate models for a chosen target column?
By asking a Ground Truth workforce to write the model
By requiring manual gradient tuning from an administrator
Through AutoML that tries and ranks candidates

10. What is SageMaker Ground Truth?

SageMaker Ground Truth is a managed data-labeling service. You point it at unlabeled data in S3, pick a task type, and it coordinates human annotators to produce a labeled dataset.

Built-in task types include image classification, bounding boxes, semantic segmentation, text classification, named entity recognition, video, and 3D point clouds. You choose the workforce: Amazon Mechanical Turk, a vetted vendor, or your own private team.

To cut labeling cost, Ground Truth supports automated data labeling (active learning). A model trains on the human labels and takes over the items it is confident about, sending only the hard ones to people. Ground Truth Plus is the turnkey option where AWS manages the workforce and quality for you.

Take quiz
In Ground Truth automated data labeling, which items still go to human annotators?
Only the easiest items
Every item, always
Items the model is not confident about
Which workforce options can Ground Truth use?
Mechanical Turk, a vetted vendor, or a private team
Only AWS employees
Only an on-premises team with no console access

11. What is SageMaker Feature Store?

SageMaker Feature Store is a managed repository for ML features. Teams write a feature once and reuse it for both training and real-time inference, which prevents training-serving skew.

Features live in feature groups. Each group needs a record identifier (such as customer_id) and an event time feature that records when each value was valid. A group can have two stores:

  • an online store for low-latency lookups of the latest record at inference time
  • an offline store in S3 that keeps full history for training and batch jobs

Because the offline store keeps timestamped versions, you can build training sets as of a specific point in time and avoid data leakage.

Take quiz
What two things must every feature group define?
A record identifier and an event time feature
An endpoint name and a variant weight
A container image and an execution role
Why does the offline store keep timestamped history?
To serve millisecond lookups to applications
To build point-in-time correct training datasets
To store trained model artifacts

12. What is the SageMaker Model Registry?

The SageMaker Model Registry is a catalog of model versions and the approval workflow around them. It stores what you need to deploy a model later: the container image, the S3 artifact location, training metrics, and metadata.

Models are grouped into model package groups, one per use case, and each registration becomes a numbered model package version. Every version has an approval status: PendingManualApproval, Approved, or Rejected.

A status change fires an event on Amazon EventBridge, which is how teams trigger deployment automatically once someone approves a version. Registries can also be shared across AWS accounts, which suits separate dev, staging, and prod setups.

Take quiz
Which of these is a valid Model Registry approval status?
InService
PendingManualApproval
Creating
How do teams usually trigger automatic deployment after approving a model version?
Restarting the Studio domain
Editing the Dockerfile of the training image
An EventBridge rule reacting to the approval status change

13. What is SageMaker Pipelines?

SageMaker Pipelines is the workflow orchestration service for ML, built for CI/CD-style automation. You describe a directed acyclic graph (DAG) of steps in Python, and SageMaker runs them in the right order while tracking lineage for every execution.

Common step types are Processing, Training, Tuning, Transform, CreateModel, RegisterModel, Condition, Lambda, and Fail. Dependencies are inferred when one step consumes another step's output, so you rarely wire the order by hand.

Pipelines support parameters (instance type or input S3 path, for example), step caching to skip unchanged steps, and retry policies. Runs can start on a schedule, from EventBridge, or from a CI system.

Take quiz
What does a Condition step do in SageMaker Pipelines?
Encrypts data flowing between steps
Scales the endpoint up or down
Branches the workflow based on a value such as an accuracy metric
How does Pipelines usually decide the order of steps?
From data dependencies between one step's outputs and another's inputs
Alphabetically by step name
Randomly, then merges the results

14. What is SageMaker Automatic Model Tuning?

Automatic Model Tuning (hyperparameter optimization) searches for the hyperparameter values that give the best model. It launches many training jobs with different settings and tracks which one scores best on an objective metric you choose, such as validation AUC.

You define the search ranges (continuous, integer, or categorical), the objective metric and whether to maximize or minimize it, and limits for total jobs and parallel jobs. The metric can be parsed from logs with a regex or emitted by a built-in algorithm.

Search strategies include Bayesian optimization, random, grid, and Hyperband. You can also turn on early stopping or warm start from earlier tuning jobs to save cost.

Take quiz
What does a tuning job launch under the hood?
Multiple training jobs with different hyperparameter values
One endpoint per hyperparameter
A single job that rewrites your source code
How does SageMaker decide which trial is best?
By picking the shortest training time only
By comparing the objective metric you specified
By picking the smallest model file

15. What is SageMaker Autopilot?

SageMaker Autopilot is SageMaker's AutoML capability. Give it a tabular dataset in S3 and a target column, and it infers the problem type (regression, binary, or multiclass classification), preprocesses the data, tries several algorithms, tunes them, and ranks the candidates on a leaderboard.

What sets it apart from a black-box AutoML tool is transparency. Autopilot generates notebooks that show the data exploration it did and the candidate pipelines it built, so you can read, edit, and rerun them. It offers ensembling and hyperparameter-optimization training modes, and the V2 API also covers time-series forecasting, image and text classification, and text generation fine-tuning.

Take quiz
What does Autopilot need from you for a tabular problem?
A hand-written model architecture
A dataset in S3 and the name of the target column
A pretrained checkpoint of the final model
What makes Autopilot a 'white box' AutoML tool?
It hides every intermediate step from you
It requires you to write the algorithm code
It generates notebooks showing the data exploration and candidate pipelines

16. What is SageMaker Clarify?

SageMaker Clarify helps you detect bias in data and models and explain individual predictions. It runs as a managed Processing job, so there is nothing to install.

It covers three jobs: measuring pre-training bias in the dataset (class imbalance across a sensitive group, for example), measuring post-training bias in the model's predictions, and computing feature attributions with SHAP to show which inputs drove a prediction. Clarify also plugs into Model Monitor to watch bias and attribution drift in production, and it offers foundation model evaluation.

Take quiz
What technique does Clarify use to compute feature attributions?
K-means clustering
Principal component analysis
SHAP (Shapley values)
How does Clarify run?
As a managed Processing job
As a permanent endpoint you must provision yourself
As a browser extension inside Studio

17. What is SageMaker Model Monitor?

SageMaker Model Monitor continuously checks the data and predictions of a deployed model and alerts you when quality drifts. It compares production traffic against a baseline computed from your training data.

It offers four monitor types:

  • Data quality - schema violations, missing values, distribution shifts
  • Model quality - accuracy, precision, RMSE and similar metrics, once ground-truth labels arrive
  • Bias drift - changes in bias metrics over time
  • Feature attribution drift - changes in which features drive predictions

Monitors run on a schedule, write violation reports to S3, and publish metrics to CloudWatch so you can set alarms.

Take quiz
Which monitor type needs ground-truth labels to work?
Model quality
Data quality
Feature attribution drift
What is each Model Monitor run compared against?
The previous hour's CPU utilization
A baseline built from training data statistics and constraints
The instance type of the endpoint

18. List the inference options available in SageMaker?

SageMaker offers four main ways to serve a model:

  • Real-time inference - a persistent endpoint for low-latency, steady traffic.
  • Serverless inference - no instances to manage; scales to zero, which suits spiky or infrequent traffic.
  • Asynchronous inference - queues requests, handles large payloads and long processing times, and can scale to zero.
  • Batch Transform - offline scoring of a whole dataset with no persistent endpoint.

On top of these you can use multi-model endpoints, inference components, and inference pipelines to host several models efficiently or to chain containers together.

Take quiz
Which option scores a whole S3 dataset without a persistent endpoint?
Real-time inference
Batch Transform
Serverless inference
Which options can scale down to zero when idle?
Only real-time inference
Only multi-model endpoints
Serverless and asynchronous inference

19. What is the SageMaker execution role?

The SageMaker execution role is an IAM role that SageMaker assumes to act on your behalf during jobs, notebooks, and endpoints. Without it, a training job could not read your S3 data, pull an image from ECR, or write logs to CloudWatch.

The role's trust policy must allow sagemaker.amazonaws.com to assume it. Its permissions should grant only what the workload needs: specific S3 buckets, ECR pulls, CloudWatch Logs, and KMS keys if you use encryption. The managed AmazonSageMakerFullAccess policy is fine for learning but too broad for production.

Whoever creates a job must also hold iam:PassRole for that role. Its absence is behind many "not authorized to perform iam:PassRole" errors.

Take quiz
Which service principal must the execution role's trust policy allow?
ec2.amazonaws.com only
lambda.amazonaws.com
sagemaker.amazonaws.com
Which permission must a user hold to hand an execution role to a SageMaker job?
iam:PassRole
sagemaker:DeleteDomain
s3:PutBucketPolicy

20. How do you launch a training job with the SageMaker Python SDK?

With SageMaker Python SDK v3 you use the ModelTrainer class. You give it a training image, your source code, compute settings, and a role, then call train() with the input data channels.

from sagemaker.train import ModelTrainer
from sagemaker.train.configs import Compute, SourceCode, InputData

trainer = ModelTrainer(
    training_image="ACCOUNT_ID.dkr.ecr.us-east-1.amazonaws.com/my-train:latest",
    source_code=SourceCode(source_dir="src", entry_script="train.py"),
    compute=Compute(instance_type="ml.m5.xlarge", instance_count=1),
    role="arn:aws:iam::ACCOUNT_ID:role/SageMakerExecutionRole",
)

trainer.train(
    input_data_config=[
        InputData(channel_name="train", data_source="s3://my-bucket/train/")
    ]
)

SageMaker starts the instances, runs train.py, and stores the model artifacts in S3. Each channel appears inside the container under /opt/ml/input/data/<channel_name>, so the channel named train shows up at /opt/ml/input/data/train.

Older v2 code used an Estimator and fit(). v3 replaced those framework-specific classes with ModelTrainer.

Take quiz
In SDK v3, which class replaces the framework-specific Estimator classes for training?
ModelTrainer
ModelBuilder
Predictor
Where does an input channel named train appear inside the training container?
/opt/ml/model/train
/opt/ml/input/data/train
/opt/ml/output/train

21. What is SageMaker HyperPod?

SageMaker HyperPod is purpose-built infrastructure for training and serving large foundation models on clusters of hundreds or thousands of accelerators. Unlike a training job, a HyperPod cluster is persistent: the instances stay up between runs, and you connect to them like any other cluster.

You pick an orchestrator, either Slurm or Amazon EKS. HyperPod adds health monitoring, automatic replacement of faulty nodes, and job auto-resume, so a single GPU failure does not sink a multi-week run. Newer features include checkpointless training, elastic training, task governance for sharing capacity across teams, and flexible training plans for reserving capacity ahead of time.

It also ships with tuned distributed training libraries and ready-made recipes for popular open models.

Take quiz
Which orchestrators can you choose for a HyperPod cluster?
Apache Airflow or Oozie
Slurm or Amazon EKS
AWS Batch or Elastic Beanstalk
What distinguishes a HyperPod cluster from a standard training job?
It is serverless and billed per prediction
It can only run on a single instance
It is a persistent cluster that keeps running between jobs

22. What is the difference between File, Pipe, and FastFile input modes?

The three modes differ in when and how data moves from S3 to the training instance. File copies everything up front, Pipe streams it through a named pipe, and FastFile exposes S3 objects as if they were local files but fetches them on demand.

Mode How it works Trade-off
File Downloads the full dataset to instance storage before training starts Simple, but startup is slow and disk must fit the data
Pipe Streams data from S3 through Unix named pipes No local copy and quick start, but sequential access and pipe-aware code
FastFile Presents S3 objects as files and streams them lazily Quick start, works with normal file APIs, no code changes

FastFile is a good default for large datasets because your code just opens files. Choose File for small data you read repeatedly, and consider FSx for Lustre when many jobs share a huge dataset and need sustained throughput.

Take quiz
Which input mode needs no code changes yet starts without downloading the whole dataset first?
File
Pipe
FastFile
What is the main drawback of File mode on very large datasets?
A long startup while everything downloads, plus enough disk for all of it
It only supports CSV files
It cannot read from S3

23. What is the difference between SageMaker AI and SageMaker Unified Studio?

SageMaker AI is the ML-focused service for building, training, and deploying models. SageMaker Unified Studio is a wider workspace that adds data engineering, analytics, and generative AI app building around it, with SageMaker AI capabilities as one part of the picture.

Aspect SageMaker AI (with its Studio) SageMaker Unified Studio
Focus Build, train, and deploy ML models Data, analytics, ML, and gen AI in one place
Tools Notebooks, training jobs, pipelines, endpoints, HyperPod EMR, Glue, Athena, Redshift, Bedrock, and SageMaker AI
Organizing unit Domains, user profiles, spaces Domains with shared projects
Data access S3 and sources you connect Lakehouse and catalog with built-in governance

Use the SageMaker AI Studio when the work is purely ML. Reach for Unified Studio when teams need to share governed data across analytics, ML, and gen AI projects. Most SageMaker AI capabilities are available in both.

Take quiz
When is SageMaker Unified Studio the better fit?
When teams share governed data across analytics, ML, and generative AI
When you only need to run one GPU training job
When you want to avoid S3 completely
Which services does Unified Studio bring together with SageMaker AI?
CodeCommit, Cloud9, and Elastic Beanstalk
EMR, Glue, Athena, Redshift, and Bedrock
WorkSpaces, AppStream, and Lightsail

24. How does SageMaker Managed Spot Training work?

Managed Spot Training runs your job on spare EC2 capacity at a discount of up to 90%, and SageMaker handles the bookkeeping when AWS reclaims an instance. Your part is to write checkpoints, so a restarted job resumes instead of starting over.

When an instance is reclaimed, the job goes into an interrupted state, SageMaker waits for capacity, restarts the container, and restores the checkpoint files from S3 into the local checkpoint path. Your script must look for existing checkpoints on startup and continue from them. You pay only for the seconds the job actually trained, and the job details show the savings.

The relevant CreateTrainingJob settings:

EnableManagedSpotTraining=True,
StoppingCondition={"MaxRuntimeInSeconds": 7200, "MaxWaitTimeInSeconds": 14400},
CheckpointConfig={"S3Uri": "s3://my-bucket/checkpoints/", "LocalPath": "/opt/ml/checkpoints"}

MaxWaitTimeInSeconds must be at least MaxRuntimeInSeconds, because it covers both training time and time spent waiting for capacity.

Take quiz
What must your training script do to survive a Spot interruption?
Disable the StoppingCondition
Save checkpoints to the checkpoint path and resume from them on restart
Rebuild and re-upload the Docker image
How must MaxWaitTimeInSeconds relate to MaxRuntimeInSeconds?
It must be smaller than MaxRuntimeInSeconds
Both must be set to zero
It must be equal to or greater than MaxRuntimeInSeconds

25. Explain the lifecycle of a SageMaker training job?

A training job moves through a series of secondary statuses while its primary status stays InProgress. It ends as Completed, Failed, or Stopped.

  1. Starting - SageMaker provisions the ML instances and pulls the training image from ECR.
  2. Downloading - input channels are copied (File mode) or mounted and streamed (Pipe and FastFile).
  3. Training - your container runs, and logs and metrics stream to CloudWatch.
  4. Uploading - the contents of /opt/ml/model are compressed into model.tar.gz and written to S3.
  5. Completed - the instances are terminated.
flowchart LR
  A["Starting: provision instances, pull image"] --> B["Downloading: fetch input data"]
  B --> C["Training: run container"]
  C --> D["Uploading: model.tar.gz to S3"]
  D --> E{Container exit code 0?}
  E -- Yes --> F[Completed]
  E -- No --> G["Failed with FailureReason"]

If the container exits with a non-zero code, the job is marked Failed, and the FailureReason field plus the CloudWatch logs are the first places to look.

Take quiz
Which secondary status comes right after Downloading?
Uploading
Starting
Training
What happens if the training container exits with a non-zero code?
The job ends as Failed and a FailureReason is recorded
SageMaker retries forever until the code succeeds
The job is marked Completed if any model file exists

26. What is the difference between real-time, serverless, and asynchronous inference?

The choice mostly follows traffic shape, latency needs, and payload or processing size.

Aspect Real-time Serverless Asynchronous
Best for Steady, latency-sensitive traffic Spiky or infrequent traffic Large payloads, long processing
Compute Provisioned instances, always on Managed for you, scales with requests Provisioned, can scale to zero
Latency Lowest and most consistent Cold starts after idle periods Seconds to minutes, queue-based
Limits Modest payloads, short timeout 1 to 6 GB memory, no GPUs Up to 1 GB payload, up to 1 hour processing
Billing Instance-hours Compute time and data processed Instance-hours, none when scaled to zero

Serverless endpoints offer optional provisioned concurrency to soften cold starts. Asynchronous endpoints write results to S3 and can notify you through SNS when a request finishes.

Take quiz
A model receives a handful of requests per day and can tolerate a cold start. Which option is the most economical?
Serverless inference
A real-time endpoint on two always-on GPU instances
A multi-container endpoint with a fixed instance count
Which option suits a 500 MB video file that takes 20 minutes to process?
Serverless inference
Asynchronous inference
Real-time inference

27. When should you use Batch Transform instead of a real-time endpoint?

Use Batch Transform when you need predictions for a large, known dataset and nobody is waiting on an individual response. A nightly scoring run over a million records is the classic case. A live checkout page is not.

A transform job spins up instances, loads your model, reads input objects from S3, writes predictions to an S3 prefix, and shuts everything down. There is no endpoint to keep running, so you pay only while the job runs.

Useful settings:

  • BatchStrategy - SingleRecord or MultiRecord to pack several records into one request
  • MaxPayloadInMB and MaxConcurrentTransforms - tune throughput per instance
  • DataProcessing - filter input columns and join the original input with the prediction in the output

If results must come back in milliseconds, or requests arrive one at a time, stay with a real-time endpoint.

Take quiz
What is the main cost advantage of Batch Transform for nightly scoring?
Instances are free when scoring runs at night
No persistent endpoint, so instances run only for the duration of the job
It skips loading the model entirely
Which Batch Transform setting lets you keep the original input columns next to each prediction?
CheckpointConfig
The production variant weight
DataProcessing with the input joined to the output

28. Why use multi-model endpoints in SageMaker?

Multi-model endpoints (MMEs) let one endpoint host hundreds or thousands of models on shared instances, so you do not pay for a separate always-on endpoint per model. They fit best when you have many similar models that are each used occasionally, such as one model per customer, region, or store.

All model artifacts sit under one S3 prefix. On the first request for a model, SageMaker downloads it to the instance and loads it into memory, and later requests hit the cached copy. When memory or disk fills up, the least recently used models are unloaded. The caller names the model in each request:

runtime.invoke_endpoint(
    EndpointName="per-customer-models",
    TargetModel="customer-42.tar.gz",
    ContentType="text/csv",
    Body=payload,
)

The trade-off is cold-start latency for any model that is not loaded yet, and every model must work with the same container (same framework and serving stack). GPU-backed MMEs are supported too, using Triton.

Take quiz
Which request parameter tells a multi-model endpoint which model to run?
TargetVariant
InferenceId
TargetModel
What happens to a rarely used model when an MME runs low on memory?
It is unloaded, with least recently used models going first
It is permanently deleted from S3
The endpoint stops accepting new requests

29. How does auto scaling work for SageMaker real-time endpoints?

Endpoint auto scaling uses Application Auto Scaling to change the instance count of a production variant based on a CloudWatch metric. The most common setup is target tracking on SageMakerVariantInvocationsPerInstance, which is the average number of invocations each instance handles per minute.

import boto3
aas = boto3.client("application-autoscaling")
rid = "endpoint/my-endpoint/variant/AllTraffic"

aas.register_scalable_target(
    ServiceNamespace="sagemaker", ResourceId=rid,
    ScalableDimension="sagemaker:variant:DesiredInstanceCount",
    MinCapacity=1, MaxCapacity=6)

aas.put_scaling_policy(
    PolicyName="invocations-target", ServiceNamespace="sagemaker",
    ResourceId=rid, ScalableDimension="sagemaker:variant:DesiredInstanceCount",
    PolicyType="TargetTrackingScaling",
    TargetTrackingScalingPolicyConfiguration={
        "TargetValue": 70.0,
        "PredefinedMetricSpecification": {
            "PredefinedMetricType": "SageMakerVariantInvocationsPerInstance"},
        "ScaleOutCooldown": 60, "ScaleInCooldown": 300})

To choose TargetValue, load-test a single instance to find its maximum sustainable invocations per minute, then set the target below that with a safety margin. Use a short scale-out cooldown and a longer scale-in cooldown so the endpoint reacts fast to spikes but does not flap.

Take quiz
Which predefined metric is commonly used for target tracking on an endpoint variant?
SageMakerVariantInvocationsPerInstance
TrainingJobDuration
FeatureGroupRecordCount
What does TargetValue represent in this policy?
The maximum number of instances allowed
The desired average number of invocations per instance per minute
The accuracy the model must maintain

30. How do inference components differ from multi-model endpoints?

A multi-model endpoint packs many small models into one shared container and loads them on demand. Inference components instead let you deploy each model as its own component on a shared endpoint, with resources reserved for it and its own scaling.

Aspect Multi-model endpoint Inference components
Unit of hosting Many small models in one container, loaded on demand Each model is a component with its own container and settings
Resources Shared pool; models compete for memory Reserved per copy: accelerators, CPU, memory
Scaling At endpoint level Per component, including scaling copies to zero
Best for Thousands of small, sparsely used models A few large models such as LLMs sharing a GPU fleet

Lazy loading works badly for multi-gigabyte LLMs because a cold start means minutes of loading. Inference components keep each model's copies resident and let you scale a busy model without touching a quiet one, which uses the shared accelerators more efficiently.

Take quiz
Which is a strength of inference components compared with multi-model endpoints?
Hosting thousands of models on a single CPU with LRU eviction
Per-model resource reservation and independent scaling
No need for an execution role
You must serve three large LLMs on a shared GPU fleet and scale each one independently. Best fit?
Batch Transform
A multi-model endpoint relying on lazy loading
Inference components

31. What happens when you deploy a model to a SageMaker endpoint?

Deployment is a three-call sequence that ends with SageMaker building a fleet of containers behind a managed load balancer.

  1. CreateModel records the image URI, the model artifact location, and the execution role. Nothing runs yet.
  2. CreateEndpointConfig defines the production variants: instance type, count, and traffic weight.
  3. CreateEndpoint provisions the instances, pulls the image, extracts the artifact into /opt/ml/model, and starts the container.

SageMaker then polls GET /ping on port 8080. Once the container answers 200, the endpoint moves from Creating to InService and POST /invocations starts receiving traffic. If the health check never succeeds within the startup timeout, the endpoint ends up Failed.

sequenceDiagram
  participant You
  participant SM as SageMaker
  participant C as Container
  You->>SM: CreateModel
  You->>SM: CreateEndpointConfig
  You->>SM: CreateEndpoint
  SM->>C: Start container, mount /opt/ml/model
  SM->>C: GET /ping
  C-->>SM: 200 OK
  SM-->>You: Endpoint InService
  You->>SM: InvokeEndpoint
  SM->>C: POST /invocations
  C-->>SM: Prediction

Later changes through UpdateEndpoint use blue/green by default: a new fleet comes up first, and traffic switches only after it is healthy.

Take quiz
Which call actually provisions the instances that serve the model?
CreateModel
CreateEndpointConfig
CreateEndpoint
Which path must the container answer with 200 before the endpoint goes InService?
/ping
/health
/ready

32. How do production variants enable A/B testing and shadow testing?

Both techniques put more than one variant behind the same endpoint, but they treat live traffic differently.

Aspect A/B testing (production variants) Shadow testing (shadow variants)
Traffic Requests are split by variant weight, and callers get responses from each variant A share of requests is copied to the shadow variant, and only the production response goes back
User impact Some users see the new model None
Compare with Per-variant CloudWatch metrics and business KPIs Latency, errors, and predictions side by side
Control Variant weights, or the TargetVariant header to force a variant Sampling percentage

For A/B tests you can change weights without downtime using UpdateEndpointWeightsAndCapacities, moving from 90/10 to 50/50 and finally 0/100 when the new variant wins. A sensible order is shadow first, to catch latency and error problems safely, then A/B to measure real business impact.

Take quiz
In shadow testing, what happens to the shadow variant's response?
It is captured for comparison but not returned to the caller
It replaces the production response
It is averaged into the production response
Which API shifts traffic weights between variants without downtime?
CreateTrainingJob
UpdateEndpointWeightsAndCapacities
UpdateDomain

33. How do SageMaker deployment guardrails work?

Deployment guardrails let you update a live endpoint gradually and roll back automatically if CloudWatch alarms fire. You configure them in the DeploymentConfig of UpdateEndpoint, and they build on blue/green deployment, where the new fleet is created next to the old one.

The traffic-shifting modes are:

  • All at once - move all traffic to the new fleet, then keep the old one through a baking period so a rollback is instant.
  • Canary - send a small slice first (say 10%); if the alarms stay quiet, move the rest.
  • Linear - shift traffic in equal steps, such as 25% every few minutes.

There is also a rolling deployment option that replaces instances in batches instead of doubling the fleet, which helps when GPU capacity is tight. In every mode, if a configured alarm goes off during the process, SageMaker reverts to the old fleet on its own.

Take quiz
Which traffic mode sends a small slice to the new fleet first and then moves the remainder?
Linear
Canary
All at once
What triggers an automatic rollback during a guarded deployment?
The endpoint name changing
A training job finishing in the same account
A configured CloudWatch alarm going into the alarm state

34. What is the difference between online and offline stores in SageMaker Feature Store?

The online store answers "what is this customer's latest value right now?" in milliseconds. The offline store answers "what were all the values over time?" for training and analysis.

Aspect Online store Offline store
Purpose Real-time lookup at inference Training sets and batch scoring
Contents Latest record per record identifier Full history of every record
Storage Low-latency key-value store, Standard or InMemory tier S3 in Parquet or Iceberg format, cataloged in Glue
Access GetRecord, BatchGetRecord Athena, Spark, dataset builder
Latency Single-digit milliseconds typical Seconds or more

A single PutRecord call writes to both stores when both are enabled, though the offline copy shows up after a short delay. Enable the online store only for features you actually need at inference time, since it costs more to keep.

Take quiz
Which store holds only the latest record for each record identifier?
The offline store
The Model Registry
The online store
Where does the offline store keep its data?
In Amazon S3, cataloged in the Glue Data Catalog
In the endpoint instance's memory
Inside the Model Registry

35. How does SageMaker Model Monitor detect data drift?

Model Monitor works as a loop: capture live traffic, compare it with a baseline on a schedule, and raise alarms when constraints are violated.

  1. Enable data capture on the endpoint so a sample of requests and responses is written to S3.
  2. Create a baseline from training data with a baselining Processing job. It outputs statistics.json and constraints.json.
  3. Create a monitoring schedule, for example hourly. Each run analyzes the newly captured data.
  4. Review results. Violations land in constraint_violations.json in S3, and metrics go to CloudWatch.
  5. React. An alarm can notify a team or start a retraining pipeline.
flowchart LR
  A["Endpoint with data capture"] --> B["Captured requests in S3"]
  B --> C["Scheduled monitoring job"]
  D["Baseline statistics and constraints"] --> C
  C --> E["Violations report in S3"]
  C --> F["CloudWatch metrics"]
  F --> G[Alarm]
  G --> H["Notify or start retraining pipeline"]

Model quality monitoring needs one extra piece: you must supply ground-truth labels and merge them with the captured predictions, because accuracy cannot be computed without them.

Take quiz
Which file holds the baseline rules that production data is compared against?
constraints.json
model.tar.gz
hyperparameters.json
What must be enabled on an endpoint before Model Monitor can analyze its traffic?
Managed spot training
Data capture
Multi-model loading

36. How does SageMaker Clarify detect bias and explain predictions?

Clarify works on two fronts: it measures bias against a facet (a sensitive attribute such as age or gender) and it explains predictions with SHAP values.

Bias. Pre-training metrics look at the data alone, such as class imbalance (CI) and difference in proportions of labels (DPL). Post-training metrics compare the model's predictions across groups, such as difference in positive proportions in predicted labels (DPPL) and disparate impact (DI). Running both tells you whether the bias comes from the data or is amplified by the model.

Explainability. Clarify computes SHAP values by perturbing inputs against a baseline and watching how the prediction changes. You get global importance (which features matter overall) and local attributions (why this one record scored high).

In code, you create a SageMakerClarifyProcessor and pass a DataConfig (dataset and label), a BiasConfig (facet and favorable outcome), a ModelConfig (a temporary endpoint Clarify calls), and a SHAPConfig. Reports land in S3 and in Studio.

Take quiz
Which Clarify metric is measured on the dataset before any model is trained?
Difference in positive proportions in predicted labels (DPPL)
Class imbalance (CI)
Disparate impact (DI)
What is the purpose of BiasConfig?
Setting the instance type for training
Choosing the hyperparameter tuning strategy
Naming the sensitive facet and the favorable label value to evaluate

37. How do you build a CI/CD workflow with SageMaker Pipelines and Model Registry?

The usual pattern is a pipeline that trains and evaluates a model, gates on a metric, registers a new version, and leaves deployment to an approval-driven step.

  1. A commit or schedule starts the pipeline.
  2. Processing and Training steps produce a candidate model.
  3. An evaluation step writes metrics to a JSON report.
  4. A Condition step checks a metric against a threshold. Candidates that fail stop there or hit a Fail step.
  5. Passing candidates are registered in the Model Registry as PendingManualApproval.
  6. A reviewer approves the version. The state change triggers an EventBridge rule that deploys to staging and, after tests pass, to production through CodePipeline or Lambda.
flowchart LR
  A["Commit or schedule"] --> B["Pipeline: process and train"]
  B --> C[Evaluate]
  C --> D{AUC above threshold?}
  D -- No --> E["Fail step"]
  D -- Yes --> F["Register: PendingManualApproval"]
  F --> G["Reviewer approves"]
  G --> H["EventBridge rule"]
  H --> I["Deploy to staging, then prod"]

SageMaker Projects provide MLOps templates that wire this up with your Git provider and AWS CodePipeline, if you would rather not assemble it by hand.

Take quiz
Which step decides whether a candidate model should be registered?
The Transform step
The Tuning step
The Condition step
What event usually kicks off deployment after human review?
The model package's approval status changing to Approved
The endpoint reaching 100% CPU
A new Studio user being created

38. How do Bayesian and Hyperband tuning strategies differ in SageMaker?

Bayesian search decides each new trial by learning from the trials already finished. Hyperband instead starts many trials cheaply and cuts the weak ones early.

Strategy How it picks trials Best when
Bayesian Fits a surrogate model to past results and picks promising values Each trial is expensive and you can afford only a few
Random Samples independently You want high parallelism or a quick baseline
Grid Tries every combination of categorical values The space is small and purely categorical
Hyperband Stops underperforming trials using intermediate metrics and gives resources to the leaders Training is iterative and reports a metric every epoch

Bayesian search learns from history, so running too many jobs in parallel hurts: each new choice has less feedback to work with. Hyperband needs your script to emit the objective metric as training progresses, and it handles early stopping itself, so you do not enable the separate early-stopping option. Warm start can seed any strategy with results from earlier tuning jobs.

Take quiz
Which strategy stops underperforming trials using intermediate metrics?
Hyperband
Grid search
Random search
Why can very high parallelism hurt Bayesian tuning?
Bayesian tuning cannot run more than one job at a time
New trials get less feedback from completed ones
Parallel jobs are billed at double the price

39. What is the difference between data parallelism and model parallelism in SageMaker?

Data parallelism copies the whole model onto every GPU and splits the batch between them. Model parallelism splits the model itself across GPUs because it does not fit on one.

Aspect Data parallelism Model parallelism
What is split Training batches Layers, tensors, or parameter and optimizer state
Model per GPU Full copy Only a shard
Communication Gradient sync (AllReduce) every step Activations or sharded parameters exchanged between GPUs
Use when The model fits in one GPU and you want faster epochs The model or optimizer state exceeds one GPU's memory
SageMaker library Distributed data parallelism (SMDDP) Model parallelism library (SMP)

SMDDP swaps in collective communication operations tuned for the AWS network. SMP v2 works with PyTorch FSDP and adds tensor parallelism, so you can combine sharding with scale-out. Start with data parallelism and move to model parallelism only when memory forces you to. Many large runs end up combining both.

Take quiz
When is model parallelism needed?
When the dataset is too small for a single epoch
When the model or its optimizer state does not fit on one GPU
When training runs on a single CPU core
Which library optimizes gradient communication for data-parallel training on AWS?
SageMaker Clarify
SageMaker Neo
SageMaker distributed data parallelism (SMDDP)

40. How does SageMaker HyperPod recover from hardware failures?

HyperPod treats hardware faults as routine and automates the detect, replace, and resume loop.

Detection. Deep health checks (GPU, network, and stress tests) run when nodes are created, and a health-monitoring agent watches running nodes for GPU, interconnect, and other hardware faults. When it flags a bad node, HyperPod replaces or reboots it.

Resume. On Slurm you enable job auto-resume with srun --auto-resume=1, and on EKS the HyperPod training operator handles the restart. The job then continues from its latest checkpoint.

Checkpointless training goes a step further. It keeps state across the cluster, swaps the faulty node, and copies model and optimizer state peer-to-peer from healthy accelerators. AWS reports recovery in minutes instead of hours and goodput above 95% on clusters with thousands of accelerators.

flowchart TD
  A["Training job running"] --> B["Health agent detects faulty node"]
  B --> C["Node replaced or rebooted"]
  C --> D{Checkpointless training enabled?}
  D -- No --> E["Restart job from last checkpoint in S3"]
  D -- Yes --> F["Peer-to-peer state copy from healthy nodes"]
  E --> G["Training resumes"]
  F --> G
Take quiz
How does checkpointless training restore state after a faulty node is replaced?
It re-reads the whole dataset from scratch
It restores a nightly EBS snapshot of the cluster
It copies model and optimizer state peer-to-peer from healthy accelerators
Which Slurm option turns on job auto-resume on HyperPod?
srun --auto-resume=1
sbatch --retry-forever
scontrol --resume-all

41. When should you choose HyperPod over SageMaker training jobs?

Choose HyperPod when you run long, large-scale training on many accelerators, or when several teams share one GPU pool. Choose plain training jobs for shorter, bursty, or experimental runs.

Aspect Training job HyperPod
Lifetime Ephemeral, one per job Persistent cluster
Access API or SDK only, no node access SSH or kubectl on Slurm or EKS
Fault handling Retry the job, optionally from checkpoints Health checks, node replacement, auto-resume, checkpointless training
Sharing capacity Each job requests its own instances Task governance, queues, preemption, elastic training
Billing Per second while the job runs Instances bill while the cluster exists

HyperPod pays off for multi-day runs across many nodes, custom schedulers or libraries, and shared research clusters. The catch is that an idle cluster still costs money, so it makes less sense for occasional small experiments.

Take quiz
Which scenario points toward HyperPod?
A multi-week foundation model run on hundreds of GPUs shared by several teams
A ten-minute XGBoost experiment
A one-off nightly batch scoring job
How does billing differ between a training job and a HyperPod cluster?
Both bill per prediction served
Training jobs bill only while running; HyperPod instances bill while the cluster exists
Training jobs bill monthly; HyperPod bills only when GPUs are busy

42. How do you secure SageMaker workloads inside a VPC?

Securing SageMaker means controlling the network path, the identity, and the data at rest and in transit. The main controls:

  • VPC configuration - launch training, processing, endpoints, and Studio domains (VPC-only mode) into private subnets with security groups.
  • Network isolation - EnableNetworkIsolation blocks all outbound network calls from the container, so untrusted code cannot exfiltrate data.
  • VPC endpoints - interface endpoints (PrivateLink) for the SageMaker API and Runtime, and a gateway endpoint for S3, keep traffic off the public internet.
  • Encryption - KMS keys for training volumes, S3 outputs, and endpoint storage; TLS in transit; optional inter-container traffic encryption for distributed training.
  • IAM - least-privilege execution roles, plus condition keys such as sagemaker:VpcSubnets, sagemaker:NetworkIsolation, and sagemaker:VolumeKmsKey to force these settings.
  • Auditing - CloudTrail for API calls and CloudWatch for logs.

One gotcha: VPC-only mode with no internet path breaks package installs, so add a NAT gateway or a private package mirror.

Take quiz
What does enabling network isolation on a training job do?
It automatically encrypts the S3 bucket
It blocks outbound network access from the container
It moves the job into a public subnet
Which IAM condition key can force jobs to use a KMS key for their storage volumes?
sagemaker:ModelCardName
aws:RequestedRegion
sagemaker:VolumeKmsKey

43. How do you bring your own container to SageMaker?

Bring your own container (BYOC) means packaging your code in a Docker image that follows SageMaker's runtime contract, pushing it to ECR, and passing the image URI to a job or model.

Purpose Path or interface
Hyperparameters (training) /opt/ml/input/config/hyperparameters.json
Input data channels /opt/ml/input/data/<channel>
Model output (training) /opt/ml/model
Failure message /opt/ml/output/failure
Health check (hosting) GET /ping on port 8080, return 200
Predictions (hosting) POST /invocations on port 8080

SageMaker runs the image with the argument train for training and serve for hosting, so your entrypoint must handle both, or you can set the entrypoint explicitly. Most teams extend a prebuilt AWS image and add their packages instead of starting from scratch, since it already honors the contract. The SageMaker Training Toolkit and Inference Toolkit help if you build your own base.

Take quiz
On which port and path does a hosting container answer health checks?
Port 80, GET /health
Port 443, POST /status
Port 8080, GET /ping
Where must a training container write its final model files?
/opt/ml/model
/opt/ml/input/data/train
/tmp/model

44. How can you optimize SageMaker costs?

SageMaker cost work falls into a few buckets: idle compute, oversized compute, on-demand pricing, and duplicate hosting.

  • Kill idle resources - delete unused endpoints (they bill hourly while InService), enable idle shutdown on Studio apps, and stop notebook instances.
  • Right-size - use Inference Recommender or the newer inference recommendations to benchmark instance types, and check CloudWatch utilization.
  • Change the purchase model - Managed Spot Training for fault-tolerant jobs, and SageMaker Savings Plans (up to 64% off for a one- or three-year usage commitment) for steady workloads.
  • Match hosting to traffic - serverless or asynchronous inference with scale-to-zero for sparse traffic, Batch Transform for offline scoring, auto scaling for variable load.
  • Consolidate - multi-model endpoints or inference components instead of one endpoint per model.
  • Use efficient silicon - Inferentia, Trainium, or Graviton instances when your model supports them.
  • Track spend - cost allocation tags and AWS Budgets per team.

Begin with idle resources. Forgotten endpoints are a common source of unpleasant bills.

Take quiz
Which purchase option gives a discount in return for a committed hourly spend on SageMaker usage?
SageMaker Savings Plans
Managed Spot Training
Serverless provisioned concurrency
An endpoint receives requests only a few times a day. Which change most reduces its cost?
Adding more instances to the real-time endpoint
Moving to serverless or asynchronous inference that scales to zero
Switching to a larger GPU instance type

45. How do you troubleshoot a SageMaker endpoint that fails to deploy?

Start with facts, not guesses. DescribeEndpoint returns a FailureReason, and the CloudWatch log group /aws/sagemaker/Endpoints/<endpoint-name> shows what the container printed.

Frequent causes and fixes:

  1. Container never becomes healthy - it must listen on port 8080 and return 200 on /ping. A slow model load may need a larger ContainerStartupHealthCheckTimeoutInSeconds.
  2. Bad model artifact - a wrong model.tar.gz layout (for framework containers, inference code belongs in a code/ folder) or a slow download; adjust ModelDataDownloadTimeoutInSeconds if needed.
  3. Permissions - the execution role cannot read the S3 artifact or pull from ECR, or a KMS key policy blocks access.
  4. Memory - model plus runtime exceeds instance RAM or GPU memory and the container is killed. Use a bigger instance or quantize the model.
  5. Capacity or quota - the instance type limit is reached or capacity is short. Request a quota increase or try another instance type or Region.

After the endpoint is InService, 4xx errors from InvokeEndpoint usually mean bad input or content type, while 5xx or ModelError points you to the container logs for the stack trace.

Take quiz
Where do you find the container's output when an endpoint fails to start?
The Model Registry approval history
The CloudWatch log group /aws/sagemaker/Endpoints/endpoint-name
The Feature Store offline table
Which setting gives a slow-loading model more time to pass health checks?
MaxWaitTimeInSeconds
InitialVariantWeight
ContainerStartupHealthCheckTimeoutInSeconds

46. How do you troubleshoot a failed SageMaker training job?

Read FailureReason from DescribeTrainingJob first. Its prefix sorts most failures into three families:

  • ClientError - configuration problems such as S3 access denied, a missing channel, an image that cannot be found, or ResourceLimitExceeded when a service quota is hit. Fix the role, the path, or the quota.
  • AlgorithmError - your training code exited with an error. Open the log stream under /aws/sagemaker/TrainingJobs in CloudWatch for the traceback and exit code.
  • CapacityError - AWS could not supply the requested instances. Retry, use an instance preference list so the job can accept several instance types, or reserve capacity with a flexible training plan.

Common code-level culprits are CUDA out-of-memory (lower the batch size, use mixed precision or gradient checkpointing, or move to a larger GPU), wrong channel names, and scripts that look for files in the wrong path. To iterate quickly, reproduce the failure on a small instance with a data sample, or use local mode, before paying for GPUs. Add a RetryStrategy and checkpointing to absorb transient infrastructure failures.

Take quiz
A job fails with CUDA out of memory. Which change is most likely to help?
Adding more input channels
Switching the input mode to Pipe
Reducing the batch size or enabling mixed precision
What does a CapacityError indicate?
SageMaker could not obtain the requested instance type at that time
The training script has a syntax error
The IAM role lacks S3 permissions

47. How do you deploy and optimize large language models on SageMaker?

Deploying an LLM well comes down to choosing the right serving stack, sizing GPUs correctly, and using the optimizations SageMaker provides.

  1. Serving stack - use a Large Model Inference (LMI) container, which runs DJL Serving with engines such as vLLM or TensorRT-LLM, or a Hugging Face TGI container. JumpStart and ModelBuilder can pick one for you.
  2. Sizing - weights need roughly parameters times bytes per parameter (a 7B model in FP16 is about 14 GB) plus room for the KV cache. Use tensor parallelism to shard across the GPUs of an instance such as ml.g5 or ml.p4/p5.
  3. Optimize - quantization (FP8 or AWQ, for example), speculative decoding, and compilation through the inference optimization toolkit cut latency and cost.
  4. Find a configuration - inference recommendations benchmark instance types against a cost, latency, or throughput goal and return validated metrics such as time to first token.
  5. Serve well - use response streaming (InvokeEndpointWithResponseStream), prefix-aware routing to reuse the KV cache for repeated prompts, and auto scaling with inference components.

Measure with realistic prompt lengths. Throughput and latency numbers from short prompts rarely carry over to production traffic.

Take quiz
Roughly how much GPU memory do the weights of a 7B-parameter model need in FP16?
About 14 GB
About 1.4 GB
About 140 GB
What does prefix-aware routing improve?
Training data shuffling
KV cache reuse, by sending repeated prompt prefixes to the same instance
IAM policy evaluation speed

48. Which is better for generative AI: SageMaker or Amazon Bedrock, and why?

Neither is better in general. Bedrock is the faster route when you want to consume a foundation model through an API. SageMaker AI wins when you need control over the model and the runtime.

Aspect Amazon Bedrock SageMaker AI
Access model Serverless API to hosted foundation models You deploy models on instances you choose
Billing Per token (or provisioned throughput) Per instance-hour
Model choice Curated provider models, plus custom model import Any model: open-weight, custom architecture, your own
Control Limited runtime control Full control of container, GPUs, scaling, networking
Extras Guardrails, agents, knowledge bases Training, tuning, Pipelines, MLOps tooling

Pick Bedrock for quick prototypes, sporadic workloads, and standard provider models. Pick SageMaker when you train or heavily fine-tune your own models, need a specific open model or inference engine, or have steady high volume where instance pricing beats per-token pricing. The two also combine well: you can fine-tune on SageMaker and deploy to Bedrock or to a SageMaker endpoint.

Take quiz
Which offering is billed per instance-hour rather than per token?
On-demand Amazon Bedrock model invocation
A SageMaker AI endpoint
Amazon Bedrock Guardrails
You must serve a custom open-weight model on a specific inference engine. Best fit?
A standard provider model on Amazon Bedrock
SageMaker Ground Truth
SageMaker AI

49. How does serverless model customization work in SageMaker AI?

Serverless model customization lets you fine-tune a foundation model without provisioning or managing a training cluster. You choose a base model and a technique, and SageMaker AI handles the infrastructure and orchestration.

  1. Prepare data in S3 in the format the technique requires.
  2. Pick a base model from supported families such as Amazon Nova, Llama, Qwen, DeepSeek, gpt-oss, Gemma, and NVIDIA Nemotron.
  3. Pick a technique - supervised fine-tuning (SFT), direct preference optimization (DPO), reinforcement learning from verifiable rewards (RLVR), or from AI feedback (RLAIF). Multi-turn reinforcement learning covers agentic, multi-step tasks.
  4. Train from Studio or with the SDK v3 trainers: SFTTrainer, DPOTrainer, RLVRTrainer, and RLAIFTrainer.
  5. Evaluate with metrics including LLM-as-a-judge, then deploy to a SageMaker endpoint or Amazon Bedrock.
flowchart LR
  A["Data in S3"] --> B["Choose base model and technique"]
  B --> C["Serverless training"]
  C --> D[Evaluate]
  D --> E{Meets success criteria?}
  E -- No --> B
  E -- Yes --> F["Deploy to endpoint or Bedrock"]

The payoff is usually a smaller, cheaper model that matches a larger general model on your specific task. Supported models and techniques change often and vary by Region, so check the current list before planning.

Take quiz
Which technique trains against a reward that can be checked programmatically?
Supervised fine-tuning on labeled pairs
Grid search over hyperparameters
Reinforcement learning from verifiable rewards (RLVR)
What does SageMaker AI handle for you in serverless customization?
Provisioning and orchestrating the training infrastructure
Writing your dataset's labels
Defining your business success criteria

50. How do you track experiments in SageMaker with managed MLflow?

SageMaker AI offers a managed MLflow tracking server, so you get the familiar MLflow API without running your own server, database, or artifact store. It records experiments, parameters, metrics, and artifacts, and it integrates with the SageMaker Model Registry.

  1. Create a tracking server from Studio or the API, with an S3 artifact location and an IAM role.
  2. Install mlflow and the sagemaker-mlflow plugin, and set the server's ARN as the tracking URI.
  3. Log runs with regular MLflow calls.
  4. Compare runs in the MLflow UI, then register the best model so it also appears in the Model Registry.
import mlflow

mlflow.set_tracking_uri("arn:aws:sagemaker:us-east-1:ACCOUNT_ID:mlflow-tracking-server/my-server")
mlflow.set_experiment("churn-model")

with mlflow.start_run():
    mlflow.log_param("max_depth", 6)
    mlflow.log_metric("val_auc", 0.91)

Because the tracking server sits outside any single training job, the same server can collect runs from notebooks, training jobs, and pipelines.

Take quiz
What do you set as the tracking URI for SageMaker managed MLflow?
The tracking server's ARN
The Studio domain ID
An S3 bucket policy
Where does the managed tracking server keep run artifacts?
In the Feature Store online store
In an Amazon S3 location you specify
Inside Lambda function code
«
»

Comments & Discussions