Cloud / Amazon Trainium Interview questions
Last updated
1. What is AWS Trainium?
AWS Trainium is a family of custom machine-learning accelerators designed by Amazon's Annapurna Labs. It was built mainly to train deep learning and generative AI models at a lower cost than general-purpose GPU instances.
You get to it through Amazon EC2 Trn instances (Trn1, Trn2, Trn3) and program it with the AWS Neuron SDK, which plugs into PyTorch and JAX. Despite the name, Trainium also runs inference, and AWS uses it to serve models on Amazon Bedrock.
Each chip carries several NeuronCores, on-package HBM and NeuronLink connections to neighbouring chips. The selling point is price-performance and energy efficiency rather than raw peak numbers.
Take quiz
Lab126, the Kindle and Echo hardware group
The EC2 Graviton compiler team
Annapurna Labs
CUDA Toolkit
AWS Neuron SDK
ROCm
TensorRT
2. What are the generations of Trainium chips?
There are three generally available generations, with a fourth announced. Each one raised on-chip memory, compute and the size of the connected domain.
| Generation | EC2 family | Key facts |
| Trainium1 | Trn1, Trn1n | 2 NeuronCore-v2 and 32 GB HBM per chip; 16 chips per instance; launched 2022 |
| Trainium2 | Trn2, Trn2 UltraServer | 8 NeuronCore-v3 and 96 GiB HBM per chip; 16 chips per instance, 64 per UltraServer; launched Dec 2024 |
| Trainium3 | Trn3 UltraServer | 3nm chip, 144 GB HBM3e; up to 144 chips per UltraServer; GA Dec 2025 |
| Trainium4 | Announced | Previewed with NVLink Fusion support |
Software stays largely common across them because all use the same Neuron SDK.
Take quiz
Trainium1
Trainium2
Trainium3
Trainium2
Trainium1
Trainium4
Trainium3
3. What is Amazon EC2 Trn1?
Trn1 is the first EC2 instance family powered by Trainium chips, aimed at training deep learning models. It launched in October 2022.
There are two main sizes: trn1.2xlarge with a single chip, and trn1.32xlarge with 16 chips, 512 GB of HBM and up to roughly 3.4 petaflops of BF16/FP16 compute. The 32 in the name describes the instance size, not the chip count.
trn1n.32xlarge is the network-optimised variant: it doubles EFA bandwidth to 1600 Gbps, which helps large multi-node jobs. Chips inside an instance talk over NeuronLink-v2. AWS positioned Trn1 as offering up to 50% lower cost-to-train than comparable GPU instances.
Take quiz
8
32
16
A newer NeuronCore generation
Double the EFA network bandwidth
More HBM per chip
It has no NeuronLink
4. What is Amazon EC2 Trn2?
Trn2 is the instance family built on Trainium2, generally available since December 2024. The standard size is trn2.48xlarge.
One instance holds 16 Trainium2 chips joined by NeuronLink-v3. That gives about 1.5 TB of HBM, 46 TB/s of memory bandwidth and roughly 20.8 PFLOPS of dense FP8 compute, with 3.2 Tbps of EFAv3 networking.
Each chip has 8 NeuronCore-v3 and 96 GiB of HBM. AWS said Trn2 gives 30-40% better price-performance than the GPU-based P5e and P5en instances, and that is the main reason teams evaluate it.
Take quiz
96 GiB
144 GiB
32 GiB
64
8
4
16
5. What is a Trn2 UltraServer?
A Trn2 UltraServer (trn2u.48xlarge) links four Trn2 instances into one scale-up system of 64 Trainium2 chips, using NeuronLink between the instances.
Together they offer about 83.2 PFLOPS of dense FP8 compute, 6 TB of HBM3 and 185 TB/s of memory bandwidth. Software sees one large, tightly connected domain.
That matters for models too big for 16 chips, or for tensor and expert parallelism that would otherwise cross slower network links. Beyond one UltraServer you scale out over EFA inside EC2 UltraClusters.
Take quiz
64
32
144
16
Only standard EFA over the data-center network
NeuronLink
A shared PCIe switch on the host CPU
6. What is Trainium3?
Trainium3 is AWS's third-generation Trainium chip and its first on a 3nm process. It powers Trn3 UltraServers, which reached general availability at re:Invent in December 2025.
A fully configured Trn3 UltraServer holds up to 144 chips, delivering 362 FP8 PFLOPS, 20.7 TB of HBM3e and 706 TB/s of memory bandwidth. Divided per chip that works out to roughly 2.5 PFLOPS FP8 and 144 GB of HBM3e.
Against Trn2 UltraServers, AWS claims up to 4.4x performance, 3.9x memory bandwidth and 4x better performance per watt. Inside the server, the NeuronSwitch-v1 fabric doubles interchip bandwidth.
Take quiz
5nm
3nm
7nm
256
64
144
72
7. What is the AWS Neuron SDK?
The AWS Neuron SDK is the software stack that lets you train and deploy models on Trainium and Inferentia. It sits between your framework and the hardware.
- Framework plugins: torch-neuronx for PyTorch, plus JAX support.
- Compiler:
neuronx-ccturns graphs into executables for NeuronCores. - Runtime and driver: load and run compiled models on the devices.
- Libraries: NeuronX Distributed (training) and NxD Inference (serving).
- NKI: a Python-based kernel language for custom kernels.
- Tools: neuron-ls, neuron-top, neuron-monitor and profilers.
It ships in Deep Learning AMIs and containers, so you rarely install it by hand.
Take quiz
nxd-train
neuron-top
nrt-loader
neuronx-cc
NeuronX Distributed
Neuron Fabric Trainer
NeuronSwitch Core
NeuronLink Lightning
8. What is a NeuronCore?
A NeuronCore is the independent compute unit inside a Trainium or Inferentia chip. Think of it roughly like a very large processing core with its own engines and on-chip memory, not like a GPU thread.
Each NeuronCore contains four engines: Tensor (matrix multiplication), Vector, Scalar (activation functions) and GpSimd (programmable general-purpose). It also has on-chip SRAM and DMA engines for moving data in from HBM.
A Trainium1 chip has 2 NeuronCore-v2, while a Trainium2 chip has 8 NeuronCore-v3. Frameworks target NeuronCores, so your parallelism degree is usually counted in cores.
Take quiz
GpSimd engine
Tensor engine
Vector engine
Scalar engine
4
2
8
16
9. What is NeuronLink?
NeuronLink is the chip-to-chip interconnect that lets Trainium chips exchange data directly, without going through the host CPU or the instance network.
Trn1 uses NeuronLink-v2 between the 16 chips in an instance. Trn2 uses NeuronLink-v3, which also stretches across the four instances of an UltraServer. On Trn3 the NeuronSwitch-v1 fabric provides all-to-all connectivity.
It is not the same as EFA. NeuronLink serves the scale-up domain (inside an instance or UltraServer), while EFA carries traffic between separate instances and UltraServers.
Take quiz
EC2 instances to S3
Trainium chips to each other
NeuronCores to the host CPU only
Containers to the Neuron runtime
NVLink
NeuronLink-v1
EFA
PCIe Gen5 only
10. What is the Neuron compiler?
The Neuron compiler, neuronx-cc, converts a framework graph (XLA HLO from PyTorch/XLA or JAX) into a NEFF, the Neuron Executable File Format that NeuronCores run.
It decides tiling, memory placement in on-chip SRAM, and the instruction schedule across the engines. Because that happens ahead of execution, tensor shapes must be known at compile time.
You tune it with the NEURON_CC_FLAGS environment variable:
export NEURON_CC_FLAGS="--model-type=transformer --auto-cast=matmul --auto-cast-type=bf16"
Compiled results are cached, so repeated runs skip the compile step. Compile time grows with graph size, so large models can take many minutes the first time.
Take quiz
PTX
SavedModel
NEFF
ONNX
NEURON_RT_FLAGS
XLA_COMPILER_OPTS
NEFF_BUILD_ARGS
NEURON_CC_FLAGS
11. What is the Neuron Kernel Interface (NKI)?
NKI is a Python-embedded kernel language for writing custom kernels that run directly on NeuronCores. It plays the role that CUDA or Triton plays on GPUs.
You write at the tile level: load tiles from HBM into on-chip memory, compute with the tensor, vector and scalar engines, and store results back. Kernels are decorated with @nki.jit and can be called from PyTorch or JAX.
Reach for NKI when the compiler's output for an operation, such as a fused attention variant, leaves performance on the table. For everyday models you rarely need it.
Take quiz
Provisioning Trn instances
Converting ONNX models
Monitoring chip temperature
Writing custom kernels for NeuronCores
Python
Java
Verilog
C++ templates
12. What data types does Trainium support?
Trainium handles the usual training formats and adds some low-precision options aimed at LLM workloads.
- FP32 and TF32 for higher-precision accumulation and sensitive operations.
- BF16 and FP16 for most training and inference compute.
- FP8 variants (configurable FP8) for higher throughput on newer chips.
- Microscaling formats such as MXFP8 and MXFP4 arrive with Trainium3.
BF16 is the common default because it keeps FP32's exponent range, so it rarely needs loss scaling. Matrix multiplications accumulate in higher precision on chip, and stochastic rounding helps when weights are stored in BF16.
Take quiz
It needs no hardware support
It is only used for inference
It stores values with more mantissa bits than FP32
It keeps the same exponent range as FP32
INT2 and INT3
Only INT8
FP64 and FP128
MXFP8 and MXFP4
13. What is NeuronX Distributed (NxD)?
NeuronX Distributed is a PyTorch library on top of torch-neuronx that supplies the building blocks for large-scale training and inference on Trainium.
It provides tensor parallel layers such as ColumnParallelLinear and RowParallelLinear, pipeline parallelism, sequence parallelism, expert parallelism for MoE, and ZeRO-1 optimizer sharding. It also handles distributed checkpointing.
NxD Training wraps these pieces into YAML-configured recipes for popular models, while NxD Inference is the serving counterpart.
Use it when a model is too large for one NeuronCore: you choose the parallel degrees, and NxD maps the ranks onto the chips and issues the needed collectives.
Take quiz
ColumnParallelLinear
ShardedBatchNorm
RowSplitEmbedding
StripeLinearReduce
Only the learning rate
Compiled NEFF files
Optimizer states
Input tokens
14. What is NxD Inference?
NxD Inference (NxDI) is the Neuron library for serving large language models on Trainium and Inferentia2. You point it at a Hugging Face style checkpoint, and it compiles and loads a sharded, optimised copy.
It includes continuous batching, KV-cache management, bucketing, quantization, speculative decoding and other LLM serving features. It integrates with vLLM, so you get an OpenAI-compatible server without writing your own scheduler.
The first run compiles NEFFs, which can take a while. Save them and reuse them in later deployments.
Take quiz
TensorFlow Serving
Triton Inference Server only
vLLM
TorchServe only
Compiling NEFFs
Training a draft model
Warming up EFA
Downloading CUDA kernels
15. What is Project Rainier?
Project Rainier is AWS's large-scale AI compute cluster built for Anthropic. It links Trainium2 UltraServers into an EC2 UltraCluster so Anthropic can train and serve Claude models.
AWS reported roughly 500,000 Trainium2 chips when it came online in October 2025, with plans to grow well beyond that. A data center campus in Indiana was the first site.
It is the standard example for answering questions on whether Trainium can scale: the same chips and Neuron stack that customers rent are running frontier-model workloads.
Take quiz
Anthropic
OpenAI
Stability AI
Meta
Inferentia2
Trainium2
Trainium1
Trainium3
16. What are the Neuron command-line tools?
Neuron ships a small set of CLI tools for checking devices and watching workloads:
neuron-lslists the Neuron devices and their NeuronCores on the instance.neuron-topshows live NeuronCore utilization and memory usage, similar in spirit totop.neuron-monitorstreams JSON telemetry, which exporters can forward to Prometheus or CloudWatch.neuron-profilecaptures and visualizes execution traces; newer releases fold this into Neuron Explorer.
First thing to run on a fresh instance is neuron-ls. If it shows no devices, the driver or runtime is not set up correctly.
Take quiz
neuron-profile
neuron-monitor
neuron-ls
neuron-top
neuron-ls
neuron-nki
neuron-cc
neuron-monitor
17. What is stochastic rounding in Trainium?
Stochastic rounding is a hardware rounding mode used when a value is cast down to a lower precision such as BF16. Rather than always rounding to the nearest value, it rounds up or down at random, with probability proportional to the discarded fraction.
Why bother? BF16 has only about 8 bits of mantissa. A tiny weight update, say 0.0001 added to 1.0, would be rounded away every time, and training would stall. With stochastic rounding, the update shows up on average, so expected value is preserved.
It is controlled by the runtime variable NEURON_RT_STOCHASTIC_ROUNDING_EN. It is most useful when you keep weights in BF16 without FP32 master copies.
Take quiz
Compiler cache misses
Small updates being rounded away
Missing exponent bits
Slow disk reads
Always rounds up
Randomly, weighted by the discarded fraction
It alternates up and down in turn
Always rounds toward zero
18. What is Trainium4?
Trainium4 is the next Trainium generation that AWS previewed at re:Invent 2025. At announcement it was not generally available.
Compared with Trainium3, AWS cited roughly 6x FP4 performance, 3x FP8 performance, 4x memory bandwidth and 2x memory capacity. The most striking item is support for NVLink Fusion, which would let Trainium4 sit alongside NVIDIA GPUs in the same rack-scale design.
Treat the figures as vendor targets, not measured results. Check AWS's announcements for current availability and instance names before planning around it, since details tend to change between preview and launch.
Take quiz
CXL only
Intel UPI
NVLink Fusion
Infinity Fabric
FP64
INT16
FP4
TF32
19. How do you launch a Trainium instance?
The quickest route is an EC2 instance with the Neuron Deep Learning AMI:
- Pick a Region that offers the Trn family you need, then secure capacity: On-Demand, a Savings Plan, or an EC2 Capacity Block for ML for a time-boxed reservation.
- Choose an instance type, for example
trn1.32xlargeortrn2.48xlarge. - Select a Deep Learning AMI (Neuron) so drivers, runtime and framework environments are preinstalled.
- Add EFA networking and fast storage if you plan multi-node jobs.
- SSH in and run
neuron-lsto confirm the devices are visible, then activate the PyTorch Neuron virtual environment.
For containers, use the Neuron DLC images with ECS or EKS instead.
Take quiz
lspci -v neuron
neuron-ls
aws neuron list
nvidia-smi
Bottlerocket ECS AMI
Amazon Linux minimal AMI
Deep Learning AMI (Neuron)
Windows Server base AMI
20. What are Neuron Deep Learning AMIs and containers?
They are prebuilt images that bundle everything you need on Trainium so you avoid dependency mismatches between drivers, runtime and frameworks.
The Deep Learning AMIs include the Neuron driver, runtime, tools and Python virtual environments for PyTorch and JAX. You boot an instance and start working.
The Deep Learning Containers package the same pieces as Docker images in Amazon ECR, with separate flavours for training and inference. They suit ECS, EKS and SageMaker, where the host only needs the Neuron driver.
Version pairing matters. Always match the container's Neuron SDK release to the driver installed on the host.
Take quiz
AWS CodeArtifact
Amazon S3 Glacier
Amazon ECR
Docker Hub only
The EBS volume type
The VPC CIDR block
The Linux desktop theme
The Neuron SDK release and driver version
21. How do you compile a PyTorch model for Trainium?
There are two routes, depending on whether you are training or serving.
For training, run the model on the XLA device through torch-neuronx. Graphs compile lazily the first time they run, or ahead of time with neuron_parallel_compile.
For inference, trace the model with an example input. The result is a TorchScript module that wraps the compiled NEFF:
import torch, torch_neuronx model.eval() example = torch.rand(1, 3, 224, 224) neuron_model = torch_neuronx.trace(model, example) neuron_model.save("model_neuron.pt")
The example input fixes the tensor shapes, so later inputs must match them.
Take quiz
torch_neuronx.trace
torch.neuron.export_cuda
torch_neuronx.fit
torch.compile_neff
Which chip generation is used
The learning rate
The fixed tensor shapes of the compiled model
The Region of the instance
22. What is the difference between Trainium and Inferentia2?
Both use NeuronCore-v2 and the same Neuron SDK, but they target different jobs. Trainium handles training and inference; Inferentia2 is tuned for inference only.
| Aspect | Trainium (Trn1) | Inferentia2 (Inf2) |
| Primary use | Training, plus inference | Inference only |
| Chips in largest instance | 16 (trn1.32xlarge) | 12 (inf2.48xlarge) |
| Accelerator memory | 512 GB HBM | 384 GB HBM |
| Can train models? | Yes | No |
If you only serve models and want the lowest inference cost, Inf2 is usually the first thing to try. Later Trainium generations are the newer line AWS is investing in for both phases. Models compiled for one Neuron target can usually be moved to the other with a recompile, because the SDK and NeuronCore generation are shared.
Take quiz
Trainium2
Trainium1
Inferentia2
Trainium3
16
24
12
8
23. What is the difference between Trn1 and Trn2?
Trn2 is a large step up in per-chip memory, NeuronCore count and, most importantly, the size of the connected domain.
| Feature | Trn1 | Trn2 |
| NeuronCore generation | v2 | v3 |
| NeuronCores per chip | 2 | 8 |
| HBM per chip | 32 GB | 96 GiB |
| Chips per instance | 16 | 16 |
| Instance accelerator memory | 512 GB | 1.5 TB |
| Largest scale-up domain | One instance (16 chips) | UltraServer (64 chips) |
AWS quoted up to 4x the performance of Trn1 for Trn2. In practice, Trn2's bigger memory per chip also means you need less model parallelism for the same model. An 8B model that needed tensor parallelism on Trn1 may fit on a few Trn2 cores, which cuts collective traffic and simplifies the setup.
Take quiz
8 versus 2
4 versus 16
2 versus 4
2 versus 8
A 4-chip node
A 64-chip UltraServer
A 144-chip UltraServer
A single 16-chip instance only
24. What is the difference between Trn2 and Trn3?
Trn3 raises chip memory, compute and domain size, and replaces the point-to-point style NeuronLink fabric with a switched design.
| Feature | Trn2 UltraServer | Trn3 UltraServer |
| Chips | 64 | 144 |
| Memory per chip | 96 GiB HBM | 144 GB HBM3e |
| Total HBM | About 6 TB | About 20.7 TB |
| FP8 compute | 83.2 PFLOPS (dense) | 362 PFLOPS |
| Intra-server fabric | NeuronLink-v3 | NeuronSwitch-v1 (all-to-all) |
AWS claims up to 4.4x higher performance and 4x better performance per watt. The bigger domain lets you keep larger models, or more experts, inside one NeuronLink-class system. For teams already on Trn2, the Neuron SDK carries most code across, though bucket sizes and parallel degrees are worth retuning to the new memory and bandwidth.
Take quiz
PCIe daisy chain
NeuronSwitch-v1
InfiniBand NDR
NeuronLink-v1 ring
6 TB
512 GB
20.7 TB
1.5 TB
25. How does Trainium differ from NVIDIA GPUs?
The biggest difference is the programming model. GPUs run many small kernels launched eagerly by the host, whereas Trainium compiles whole graphs ahead of time into one schedule for each NeuronCore.
| Area | Trainium | NVIDIA GPU |
| Compute unit | NeuronCore with Tensor, Vector, Scalar, GpSimd engines | Streaming multiprocessors with CUDA/Tensor cores |
| On-chip memory | Software-managed SRAM (SBUF/PSUM) | Hardware-managed caches plus shared memory |
| Execution | Ahead-of-time compiled graph (NEFF) | Eager kernel launches, optional graphs |
| Custom kernels | NKI | CUDA, Triton |
| Availability | AWS only | All major clouds and on-premises |
So Trainium rewards stable shapes and standard architectures, while GPUs are more forgiving of dynamic or unusual code.
Take quiz
It compiles whole graphs ahead of time into one schedule
It launches every kernel eagerly from the host
It only runs precompiled ONNX files
It interprets Python bytecode on the chip
CUDA C
HIP
NKI
Triton
26. Why does Trainium need static shapes?
Because the Neuron compiler fixes everything at compile time: tensor dimensions, memory layout in on-chip SRAM, DMA transfers and the instruction schedule on each engine. A NEFF is built for exact shapes, with no runtime kernel selection to adapt.
The consequences differ by workflow:
- Training: a new shape produces a new graph, which triggers another compile. That can mean minutes of stall in the middle of a run.
- Inference: a traced model rejects inputs that do not match its traced shape.
The usual workarounds are padding, drop_last=True on data loaders, avoiding data-dependent control flow, and bucketing for variable-length inputs. Treat shape stability as a design rule from day one, not a late optimization.
Take quiz
The NeuronCores reshape their SRAM automatically
The runtime picks a different kernel instantly
Nothing, shapes are dynamic by default
The graph is recompiled
drop_last=True
pin_memory=True
shuffle=False
num_workers=0
27. How do you handle variable-length inputs with bucketing on Trainium?
Bucketing means compiling a handful of fixed-length versions of the model and routing each request to the smallest one that fits. You pad the input up to that bucket length and mask the padding.
BUCKETS = [128, 256, 512, 1024, 2048] def pick_bucket(n_tokens): return next(b for b in BUCKETS if b >= n_tokens)
The trade-off is padding waste versus resource cost. More buckets mean less wasted compute, but each bucket is another NEFF to compile and hold in HBM. Choose boundaries from your real length distribution rather than powers of two by habit.
NxD Inference lets you configure separate buckets for the prompt-processing and token-generation phases.
Take quiz
Longer prompts get truncated
More NEFFs to compile and keep in memory
The model loses accuracy
NeuronLink bandwidth drops
To the smallest compiled length that fits it
To whichever chip is coolest
To the largest bucket always
To a random bucket
28. How does the Neuron compilation cache work?
Neuron caches compiled NEFFs keyed on the input graph, the compiler flags and the compiler version. If all three match, the next run skips neuronx-cc and loads the stored artifact.
By default the cache lives in a local directory (/var/tmp/neuron-compile-cache). Set NEURON_COMPILE_CACHE_URL to an S3 location and every node in a cluster can share it, so only the first run pays for compilation.
A change to the model graph, flags or SDK version means a cache miss. That is why a harmless-looking code edit sometimes triggers a long recompile. neuron_parallel_compile can warm the cache before the real run starts.
Take quiz
Setting NEURON_RT_LOG_LEVEL=debug
Pointing NEURON_COMPILE_CACHE_URL at S3
Mounting the cache on tmpfs
Copying NEFFs through EFA at runtime
Changing the log file path
Running neuron-ls
Upgrading the Neuron SDK version
Restarting the instance with the same SDK
29. What is the Logical NeuronCore configuration on Trn2?
A Trn2 chip has 8 physical NeuronCores. The Logical NeuronCore (LNC) setting decides how they are presented to software.
- LNC=2 (default): two physical cores are fused into one logical core, so each chip shows 4 larger logical cores with access to more HBM and combined compute.
- LNC=1: all 8 cores are exposed separately, each with a smaller memory share.
You set it with NEURON_LOGICAL_NC_CONFIG. A trn2.48xlarge therefore shows 64 logical cores at LNC=2, or 128 at LNC=1.
Use LNC=2 for big models that need memory per rank. Use LNC=1 when the model fits comfortably and you want more data-parallel replicas.
Take quiz
64
8
128
16
NEURON_LNC_MODE_FLAG
XLA_NEURON_LNC
NEURON_LOGICAL_NC_CONFIG
NEURON_CORE_SPLIT
30. How does tensor parallelism work in NxD?
NxD splits individual weight matrices across NeuronCores so each core holds only a slice. The key pair is ColumnParallelLinear, which splits the output dimension, and RowParallelLinear, which splits the input dimension.
In a transformer MLP, the first projection is column-parallel, so each rank computes its own partial activations with no communication. The second projection is row-parallel, so each rank produces a partial sum, and one all-reduce combines them. That is one collective per block.
Attention heads are divided across ranks the same way. With sequence parallelism, the all-reduce is replaced by a reduce-scatter and an all-gather, which also shards activations.
Keep the tensor-parallel group inside one NeuronLink domain, because these collectives run on every layer.
Take quiz
A pipeline stage boundary
RowParallelLinear
A replicated embedding
Another ColumnParallelLinear
Broadcast
Barrier
Scatter
All-reduce
31. When would you choose Trainium over GPU instances?
Choose Trainium when cost per token or per training run is what matters most, and your model and team fit its constraints.
- The model is a mainstream architecture (Llama-style, Mixtral-style MoE and similar) with Neuron support.
- The workload is large and steady, so porting effort pays back through lower cost and reserved capacity.
- You already run on AWS and want integration with EKS, SageMaker HyperPod and Bedrock.
- Power efficiency and supply availability are part of the decision.
Stay on GPUs when you depend on CUDA-only libraries, run research code with highly dynamic shapes, or need portability across clouds. A short proof of concept with real shapes and real cost per million tokens settles the question better than spec-sheet comparison.
Take quiz
Your code depends on CUDA-only libraries
You serve a mainstream LLM at steady high volume
You already use SageMaker HyperPod
You want the lowest cost per token on AWS
Use whichever instance the team used last
Compare only peak TFLOPS on the datasheet
Benchmark cost per million tokens with real shapes
Pick whichever has a shorter name
32. How do you run Trainium workloads on Amazon EKS?
You add Trn nodes to the cluster and install the Neuron device plugin, which advertises Neuron resources to Kubernetes. Pods then request them like any other resource.
- Create a node group with a Trn instance type and an AMI that has the Neuron driver.
- Deploy the Neuron device plugin DaemonSet. Add the Neuron scheduler extension so multi-core pods get contiguous cores.
- Install the EFA device plugin for multi-node training.
- Build pods from a Neuron DLC image and request the resource:
resources: limits: aws.amazon.com/neuroncore: 8
You can request whole chips with aws.amazon.com/neuron instead. Deploy neuron-monitor as a DaemonSet to export metrics.
Take quiz
aws.amazon.com/efa-core
nvidia.com/gpu
aws.amazon.com/neuroncore
neuron.io/chip-slice
Encrypts pod traffic
Compiles NEFF files for pods
Advertises Neuron devices to the Kubernetes scheduler
Replaces the kubelet
33. How does SageMaker HyperPod support Trainium?
SageMaker HyperPod provides persistent, resilient clusters for long training runs, and it supports Trn instances alongside GPU instances. You can orchestrate with Slurm or Amazon EKS.
The value is reliability. HyperPod runs health checks on nodes, replaces faulty ones, and can automatically resume a job from the latest checkpoint. On a run lasting weeks across hundreds of chips, hardware failures are expected, so this saves a lot of manual work.
You still bring your own training code, typically NxD Training or a PyTorch script, and your own checkpoints on shared storage such as FSx for Lustre or S3. Test a restore before the long run, so auto-resume actually works on the day a node fails.
Take quiz
Restarts training from step zero always
Replaces the node and can resume from a checkpoint
Ignores the failure and continues with fewer chips silently
Moves the whole job to GPUs
Only Docker Swarm
Slurm or Amazon EKS
Only Apache Airflow
Only AWS Lambda
34. How does mixed precision with auto-cast work on Trainium?
Auto-cast is a compiler feature. With --auto-cast you tell neuronx-cc to lower FP32 operations to a faster type without editing your model.
--auto-cast=matmulcasts only matrix-multiplication inputs.--auto-cast=allcasts every FP32 operation.--auto-cast=noneleaves precision untouched.--auto-cast-typepicks the target: bf16, fp16, tf32 or an FP8 format.
The matmul option is a safe default because accumulation still happens in higher precision on chip, and sensitive operations such as softmax and normalization stay in FP32. Many teams skip auto-cast and cast the model to torch.bfloat16 directly, which is more explicit. Whichever route you pick, compare loss curves against an FP32 baseline for a few hundred steps before committing to a long run.
Take quiz
--auto-cast=all
--model-type=matmul
--auto-cast=none
--auto-cast=matmul
It removes the need for a compiler
It disables compilation caching
It forces FP8 everywhere
Sensitive ops like softmax stay in FP32
35. How do you serve LLMs on Trainium with vLLM?
vLLM has a Neuron backend that uses NxD Inference underneath, so you get continuous batching and an OpenAI-compatible endpoint on Trn instances.
- Launch a Trn instance and pull the Neuron vLLM container (or use the Neuron DLAMI).
- Choose a tensor-parallel degree that fits the model across the NeuronCores.
- Set
max-model-lenandmax-num-seqs. These feed bucket and batch sizes, which are compiled in. - Start the server and wait for the first-run compile.
vllm serve meta-llama/Llama-3.1-8B-Instruct \ --tensor-parallel-size 8 \ --max-model-len 4096 \ --max-num-seqs 16
Exact flags vary by Neuron release, so check the release notes. Cache the compiled artifacts so restarts are fast.
Take quiz
They set the AWS Region
They determine compiled shapes and bucket sizes
They disable continuous batching
They control EFA firmware
Warming up a CDN
Compiling the model into NEFFs
Downloading CUDA drivers
Provisioning a GPU
36. How does ZeRO-1 reduce memory on Trainium?
ZeRO-1 shards the optimizer states across data-parallel ranks instead of copying them to every rank. Weights and gradients stay replicated, so it is the cheapest ZeRO stage to apply.
Adam keeps an FP32 master copy plus two moments, about 12 bytes per parameter. For an 8-billion-parameter model that is around 96 GB. With 16 data-parallel ranks, each holds roughly 6 GB instead.
The flow: each rank updates only its shard, then an all-gather redistributes the updated weights. The added communication is modest, and the freed HBM often allows a larger micro-batch or less tensor parallelism.
NxD provides a ZeRO-1 optimizer wrapper that you apply around your regular optimizer.
Take quiz
A global barrier only
A checkpoint reload
A broadcast from rank 0
An all-gather
About 600 MB
About 96 GB
About 6 GB
About 24 GB
37. Explain the execution flow of PyTorch/XLA lazy tensors on Trainium?
On Trainium, PyTorch runs through the XLA backend, so operations on XLA tensors do not execute immediately. They are recorded into a lazy graph, and nothing runs on the NeuronCores until the graph is cut and submitted.
flowchart LR
A["Python ops on XLA tensors"] --> B["Lazy graph recorded"]
B --> C["mark_step or sync point"]
C --> D{Graph hash in cache?}
D -- Yes --> F["Load cached NEFF"]
D -- No --> E["HLO to neuronx-cc to NEFF"]
E --> F
F --> G["Runtime executes on NeuronCores"]
G --> H["Results available to host"]
The graph is cut at xm.mark_step() (called for you by Neuron's parallel data loader and optimizer helpers) or when the host needs a value, as with .item() or print(tensor). The hash of the graph decides whether a compiled NEFF already exists.
Two practical effects follow. First, a stray .item() inside the step splits it into several smaller graphs and adds sync stalls. Second, anything that changes the graph between steps, such as shapes or Python-side branching on tensor values, creates a new hash and a recompile.
for batch in loader: # MpDeviceLoader marks steps for you loss = model(batch).loss loss.backward() xm.optimizer_step(optimizer) optimizer.zero_grad()
Take quiz
When the graph is cut at a mark_step or sync point
Only at program exit
As soon as each Python operation is written
Whenever a new Python module is imported
It deletes the compile cache
It switches the job to CPU permanently
It forces a sync and splits the step into extra graphs
It disables BF16
38. Explain the internal working of the NeuronCore engines?
A NeuronCore is built around four engines that run in parallel and share one on-chip SRAM, so good performance comes from keeping all of them busy.
- Tensor engine: a 128x128 systolic array for matrix multiplication and convolutions. Results land in PSUM, a small accumulation buffer.
- Vector engine: element-wise work, reductions and normalizations across the partition dimension.
- Scalar engine: activation functions such as exp, GELU and sigmoid, applied in a pipeline.
- GpSimd engine: small programmable cores for operations that do not map to the other three, including custom ops.
Data comes from HBM through DMA engines into SBUF, the main software-managed SRAM, organised as 128 partitions. The compiler schedules DMA and compute so they overlap, with semaphores enforcing dependencies.
flowchart LR H[HBM] <--> D["DMA engines"] D <--> S["SBUF, 128 partitions"] S --> T["Tensor engine"] T --> R["PSUM accumulators"] R --> V["Vector engine"] R --> C["Scalar engine"] V --> S C --> S G["GpSimd engine"] <--> S
A softmax after attention is a typical case: matmul on Tensor, exp on Scalar, sums on Vector, all pipelined over tiles.
Take quiz
Holding the host OS page tables
Storing the compiled NEFF
Accumulating tensor engine matmul results
Buffering EFA packets
DMA engine
GpSimd engine
Tensor engine
Scalar engine
39. Explain the lifecycle of a NEFF from compilation to execution?
A NEFF (Neuron Executable File Format) is the compiled package a NeuronCore runs. Its life has two halves: building it, and loading and running it.
flowchart TD A["Framework graph: HLO"] --> B["neuronx-cc passes: tiling, layout, scheduling"] B --> C["Code generation per engine"] C --> D["NEFF: instruction streams, DMA descriptors, weight layout, metadata"] D --> E["Runtime load: allocate HBM, copy weights"] E --> F["Runtime execute: bind inputs, run, return outputs"]
- Compile: the compiler tiles operations to fit SBUF, chooses memory layouts and schedules instructions across the four engines.
- Package: the NEFF holds per-engine instruction streams, DMA descriptors, weight data and metadata about inputs and outputs.
- Load: the Neuron runtime allocates HBM, copies weights in and prepares queues. This is why each loaded NEFF costs device memory.
- Execute: the runtime binds input and output buffers and starts execution. Collective operations go through the Neuron collectives library.
Loading is the slow, one-time step. Execution just replays a fixed schedule, which is why steady-state latency is predictable.
Take quiz
Allocates HBM and copies weights onto the device
Opens a connection to the compile cache on S3 per token
Recompiles the graph for the current batch
Rewrites the Python source
Per-engine instruction streams
The Python training script
Weight layout data
DMA descriptors
40. How do you troubleshoot slow training on Trainium?
Work from the cheapest check to the deepest one, and confirm each theory with a number.
- Look at utilization. Run
neuron-top. Low NeuronCore utilization with a busy CPU points to a host bottleneck. - Check for recompiles. Repeated compiler messages or cache misses mean shapes or graph structure are unstable. Fix padding or
drop_last. - Inspect the data pipeline. Slow tokenization or storage reads starve the chips. Use the parallel device loader and prefetch.
- Look at sync points. Remove
.item(), logging of tensors and gradient-norm checks from the inner loop, or do them every N steps. - Profile. Use neuron-profile (or Neuron Explorer) to see idle engines, DMA-bound regions and collective wait time.
- Revisit parallelism and precision. Excess tensor-parallel degree makes collectives dominate. Training in FP32 wastes the tensor engine, so move to BF16.
Change one thing at a time and compare tokens per second, since several fixes interact.
Take quiz
High NeuronCore utilization with idle CPU
A short compile cache path
Large HBM free space
Low NeuronCore utilization while the CPU is busy
Unstable input shapes between steps
Running neuron-ls
Having too many CPU cores
Using BF16 weights
41. How do you troubleshoot long compile times or compiler failures on Trainium?
Start by finding out whether the compile is slow because the graph is large, or because the cache is being missed.
- Read the compiler log.
neuronx-ccwrites a log and workdir; the failing pass and operator are usually named there. - Shrink the graph. Unrolled loops and very deep models produce huge graphs. Reduce Python-side unrolling and use model-level wrappers where Neuron supports them.
- Try faster flags. Use a lower optimization level while iterating, then raise it for the final run.
- Compile in parallel.
neuron_parallel_compileextracts graphs from a trial run and compiles them concurrently. - Reuse artifacts. Point
NEURON_COMPILE_CACHE_URLat S3 so one node's compile benefits the others. - Give it memory. The compiler is RAM-hungry, so a larger instance (or fewer parallel jobs) avoids out-of-memory kills.
For an operator the compiler rejects, check the release notes for support, then either rewrite it with supported ops or implement it in NKI.
Take quiz
nki.jit
neuron-ls
neuron-top
neuron_parallel_compile
Delete the NEURON_CC_FLAGS variable and retry blindly
Read the compiler log to find the failing pass or operator
Switch to a GPU immediately
Reinstall the operating system
42. How do you troubleshoot HBM out-of-memory errors on Trainium?
First, work out what is consuming memory, because the fix differs. neuron-top shows device memory split into categories such as model code, tensors and scratch space.
| Likely cause | Fix |
| Weights, gradients and optimizer states too big | Raise tensor or pipeline parallel degree; enable ZeRO-1; use BF16 |
| Activations too big | Enable activation checkpointing; lower micro-batch; use sequence parallelism |
| KV cache too big (inference) | Reduce max sequence length or batch size; quantize the KV cache |
| Too many loaded NEFFs | Cut the number of buckets; unload unused models |
| Cores too small on Trn2 | Use LNC=2 so each logical core has more HBM |
Remember that every loaded NEFF reserves device memory for instructions and scratch buffers. A server with ten buckets can run out of memory even though weights alone fit.
Take quiz
Buckets duplicate the S3 cache in HBM
Buckets disable ZeRO-1
Each loaded NEFF reserves device memory
Padding tokens are stored on the CPU
LNC=2
NEURON_RT_STOCHASTIC_ROUNDING_EN=0
LNC=1
--model-type=generic
43. How can you optimize LLM inference throughput on Trainium?
Decode is memory-bandwidth bound, so most gains come from keeping the chips busy with useful batches and moving fewer bytes per token.
| Technique | What it helps |
| Continuous batching | Fills batch slots as sequences finish, raising utilization |
| Chunked prefill | Stops long prompts from stalling ongoing decodes (better TPOT) |
| Prefix caching | Skips recomputing shared system prompts |
| Right-sized buckets | Cuts padding waste, tuned to real length histograms |
| FP8 or INT8 weights | Moves fewer bytes per token and fits more KV cache |
| Speculative decoding | Produces several tokens per target-model pass |
| Tuned tensor-parallel degree | Balances per-token latency against chips used per replica |
Measure tokens per second per chip, time to first token and time per output token under realistic traffic. A configuration that wins on a single request often loses once batching is on.
Take quiz
Each generated token must read the full weights and KV cache
NeuronLink is too slow for single tokens
The compiler disables DMA
The Tensor engine cannot do matmul at small batch
ZeRO-1
Stochastic rounding
Prefix caching
Bucket padding
44. How does speculative decoding work on Trainium?
A small draft model guesses the next k tokens cheaply. The large target model then checks all k in a single forward pass and keeps the longest prefix it agrees with. Output quality is unchanged, but several tokens can be produced per target pass.
sequenceDiagram participant D as Draft model participant T as Target model D->>D: Propose k tokens D->>T: Send k candidate tokens T->>T: One forward pass scores all k T-->>D: Accept longest matching prefix plus 1 new token D->>D: Continue from the accepted position
On Trainium, NxD Inference supports draft-model and EAGLE-style speculation. Both models are compiled and loaded, and the speculation length k is baked into the NEFF because shapes are static. Changing k means recompiling.
It pays off when decode is memory-bound and the draft agrees often. A low acceptance rate wastes the draft's work, so measure it on your own prompts before enabling.
Take quiz
The driver restricts k to 4 at boot
NeuronLink must be retrained
The draft model is stored in S3
k is part of the compiled static shape
Generates the draft tokens itself one by one
Compresses the KV cache
Verifies all proposed tokens in one forward pass
Selects the instance type
45. How do you write an efficient NKI kernel?
Think in tiles. SBUF has 128 partitions, so the first dimension of a tile is the partition dimension and is capped at 128. Everything else is the free dimension.
import neuronxcc.nki as nki import neuronxcc.nki.language as nl @nki.jit def add_kernel(a, b): out = nl.ndarray(a.shape, dtype=a.dtype, buffer=nl.shared_hbm) a_t = nl.load(a[0:128, 0:512]) # HBM -> SBUF b_t = nl.load(b[0:128, 0:512]) nl.store(out[0:128, 0:512], value=a_t + b_t) # SBUF -> HBM return out
Import paths shift between Neuron releases, so check your version's docs. The ideas stay the same:
- Size tiles to fit SBUF while keeping the partition dimension full, since a half-empty tile wastes lanes.
- Loop with
nl.affine_rangeso the compiler can pipeline iterations. - Accumulate matmul partial results in PSUM instead of writing back after each step.
- Overlap DMA with compute by double-buffering tiles.
- Profile and check which engine is idle.
Take quiz
128
512
32
1024
The EFA send buffer
PSUM
HBM after every step
The host CPU memory
46. How does collective communication work across Trainium chips?
Collectives such as all-reduce, all-gather, reduce-scatter and all-to-all are issued by the framework as XLA collective operations. The Neuron collectives library maps them onto the physical links, and the compiler overlaps them with compute where it can.
| Parallelism | Typical collective | Preferred link |
| Tensor parallel | All-reduce or reduce-scatter/all-gather | NeuronLink or NeuronSwitch |
| Expert parallel (MoE) | All-to-all | NeuronLink or NeuronSwitch |
| Data parallel | Gradient all-reduce or reduce-scatter | EFA across instances |
| Pipeline parallel | Point-to-point sends | EFA across stages |
Replica groups come from NxD's parallel state, which maps ranks to NeuronCores. Place the chatty groups (tensor and expert) inside one NeuronLink domain, and let the less frequent data-parallel traffic cross EFA.
Within a Trn2 instance, chips form a 2D torus over NeuronLink-v3. On Trn3, NeuronSwitch-v1 offers all-to-all connectivity, which helps patterns like MoE dispatch.
Take quiz
Log shipping
Tensor parallelism
Checkpoint replication
Data parallelism across Regions
All-to-all
Barrier only
Gather to rank 0
Broadcast of weights
47. How do you scale training across multiple Trn2 UltraServers?
Use the fast scale-up domain for the chatty parallelism and the scale-out network for the quiet kind.
- Inside each UltraServer (64 chips over NeuronLink): tensor parallelism and, for MoE, expert parallelism.
- Across UltraServers (EFA in an EC2 UltraCluster): data parallelism and, for very deep models, pipeline parallelism.
- Use ZeRO-1 to cut optimizer memory across the data-parallel dimension.
- Launch inside a cluster placement setup so nodes get low-latency EFA paths.
- Compile once and share the cache through S3, so hundreds of nodes do not each recompile.
- Write sharded checkpoints to shared storage such as FSx for Lustre, and test restore early.
- Validate with a collective benchmark before the real run, to catch slow links.
NxD Training expresses this as a YAML config: degrees for tensor, pipeline, expert and data parallelism, plus batch sizes. Scale in stages, checking tokens per second at each size, because efficiency often dips when you first cross the EFA boundary.
Take quiz
Tensor parallelism
Data parallelism
Sequence parallelism within attention
Expert all-to-all at every layer
So nodes do not each repeat the same compilation
To increase NeuronLink bandwidth
To replace checkpoints
To enable FP8
48. What happens when a Trainium node fails during a long training run?
Usually the failing rank's process dies or the runtime reports a device error, and the remaining ranks hang inside a collective waiting for it. After a timeout, the whole job fails; collective training does not continue with a missing rank.
flowchart TD A["Node or device error"] --> B["Other ranks block in collective"] B --> C["Timeout, job fails"] C --> D["Orchestrator replaces node"] D --> E["Job restarts"] E --> F["Load latest checkpoint"] F --> G["Training resumes"]
Recovery depends on tooling. On SageMaker HyperPod, health checks replace the node and auto-resume restarts the job. On EKS or Slurm, you rely on elastic launchers or your own restart scripts.
What you lose is the work since the last checkpoint, so checkpoint frequency is a cost trade-off: frequent saves cost throughput, rare ones cost redone steps. Sharded or asynchronous checkpoints to shared storage keep that overhead small.
Take quiz
Switch to CPU training
Silently redistribute the missing shard
Ignore the collective and proceed
Block waiting, then time out
How long since the last checkpoint
The number of buckets
The EFA firmware date
The Neuron SDK version
49. How can you reduce the cost of training on Trainium?
Cost per run is price per hour multiplied by hours, so attack both.
- Lower precision: train in BF16 and evaluate FP8 where the recipe supports it, so each chip does more work per hour.
- Avoid recompiles: stable shapes and a shared compile cache stop expensive instances sitting idle while compiling.
- Right-size parallelism: over-sharding makes collectives dominate. Pick the smallest tensor-parallel degree that fits memory.
- Buy capacity wisely: Savings Plans for steady use, EC2 Capacity Blocks for scheduled runs.
- Use Spot for tolerant jobs: fine-tuning with frequent checkpoints can handle interruptions.
- Pack small jobs: use LNC=1 to run more replicas or experiments per chip.
- Track the right metric: cost per trained token, not just hourly price.
Compile time is a quiet cost: a 30-minute compile across 16 idle instances is real money.
Take quiz
Disabling checkpoints
Using FP32 everywhere
Raising tensor parallel degree to the maximum
Sharing a warmed compile cache
Size of the AMI
Cost per trained token
Number of NeuronCores
Instance hourly price alone
50. Why do Mixture-of-Experts models benefit from Trn2 and Trn3 UltraServers?
MoE layers route each token to a few experts, and those experts usually live on different chips. Every MoE layer therefore needs an all-to-all to dispatch tokens and another to combine results, in both the forward and backward passes.
That traffic is heavy and latency-sensitive. Over a general network it can dominate step time. Inside an UltraServer, all chips share the NeuronLink domain (64 chips on Trn2, up to 144 on Trn3 with NeuronSwitch-v1), so the all-to-all stays on the fast fabric.
Example: a model with 64 experts can place one expert per chip in a Trn2 UltraServer, and each token's top-2 selection becomes two hops inside the domain.
The large pooled HBM matters as well. A Trn2 UltraServer holds about 6 TB and a Trn3 UltraServer about 20.7 TB, enough for big expert sets plus their KV caches during serving.