Prev Next

Cloud / Amazon Trainium Interview questions

Last updated

1. What is AWS Trainium? 2. What are the generations of Trainium chips? 3. What is Amazon EC2 Trn1? 4. What is Amazon EC2 Trn2? 5. What is a Trn2 UltraServer? 6. What is Trainium3? 7. What is the AWS Neuron SDK? 8. What is a NeuronCore? 9. What is NeuronLink? 10. What is the Neuron compiler? 11. What is the Neuron Kernel Interface (NKI)? 12. What data types does Trainium support? 13. What is NeuronX Distributed (NxD)? 14. What is NxD Inference? 15. What is Project Rainier? 16. What are the Neuron command-line tools? 17. What is stochastic rounding in Trainium? 18. What is Trainium4? 19. How do you launch a Trainium instance? 20. What are Neuron Deep Learning AMIs and containers? 21. How do you compile a PyTorch model for Trainium? 22. What is the difference between Trainium and Inferentia2? 23. What is the difference between Trn1 and Trn2? 24. What is the difference between Trn2 and Trn3? 25. How does Trainium differ from NVIDIA GPUs? 26. Why does Trainium need static shapes? 27. How do you handle variable-length inputs with bucketing on Trainium? 28. How does the Neuron compilation cache work? 29. What is the Logical NeuronCore configuration on Trn2? 30. How does tensor parallelism work in NxD? 31. When would you choose Trainium over GPU instances? 32. How do you run Trainium workloads on Amazon EKS? 33. How does SageMaker HyperPod support Trainium? 34. How does mixed precision with auto-cast work on Trainium? 35. How do you serve LLMs on Trainium with vLLM? 36. How does ZeRO-1 reduce memory on Trainium? 37. Explain the execution flow of PyTorch/XLA lazy tensors on Trainium? 38. Explain the internal working of the NeuronCore engines? 39. Explain the lifecycle of a NEFF from compilation to execution? 40. How do you troubleshoot slow training on Trainium? 41. How do you troubleshoot long compile times or compiler failures on Trainium? 42. How do you troubleshoot HBM out-of-memory errors on Trainium? 43. How can you optimize LLM inference throughput on Trainium? 44. How does speculative decoding work on Trainium? 45. How do you write an efficient NKI kernel? 46. How does collective communication work across Trainium chips? 47. How do you scale training across multiple Trn2 UltraServers? 48. What happens when a Trainium node fails during a long training run? 49. How can you reduce the cost of training on Trainium? 50. Why do Mixture-of-Experts models benefit from Trn2 and Trn3 UltraServers?

1. What is AWS Trainium?

AWS Trainium is a family of custom machine-learning accelerators designed by Amazon's Annapurna Labs. It was built mainly to train deep learning and generative AI models at a lower cost than general-purpose GPU instances.

You get to it through Amazon EC2 Trn instances (Trn1, Trn2, Trn3) and program it with the AWS Neuron SDK, which plugs into PyTorch and JAX. Despite the name, Trainium also runs inference, and AWS uses it to serve models on Amazon Bedrock.

Each chip carries several NeuronCores, on-package HBM and NeuronLink connections to neighbouring chips. The selling point is price-performance and energy efficiency rather than raw peak numbers.

Take quiz
Which AWS team designs the Trainium chips?
Lab126, the Kindle and Echo hardware group
The EC2 Graviton compiler team
Annapurna Labs
Which software stack do you use to program Trainium?
CUDA Toolkit
AWS Neuron SDK
ROCm
TensorRT

2. What are the generations of Trainium chips?

There are three generally available generations, with a fourth announced. Each one raised on-chip memory, compute and the size of the connected domain.

Generation EC2 family Key facts
Trainium1 Trn1, Trn1n 2 NeuronCore-v2 and 32 GB HBM per chip; 16 chips per instance; launched 2022
Trainium2 Trn2, Trn2 UltraServer 8 NeuronCore-v3 and 96 GiB HBM per chip; 16 chips per instance, 64 per UltraServer; launched Dec 2024
Trainium3 Trn3 UltraServer 3nm chip, 144 GB HBM3e; up to 144 chips per UltraServer; GA Dec 2025
Trainium4 Announced Previewed with NVLink Fusion support

Software stays largely common across them because all use the same Neuron SDK.

Take quiz
Which Trainium generation is the first AWS AI chip built on 3nm?
Trainium1
Trainium2
Trainium3
Which generation was previewed with NVLink Fusion support?
Trainium2
Trainium1
Trainium4
Trainium3

3. What is Amazon EC2 Trn1?

Trn1 is the first EC2 instance family powered by Trainium chips, aimed at training deep learning models. It launched in October 2022.

There are two main sizes: trn1.2xlarge with a single chip, and trn1.32xlarge with 16 chips, 512 GB of HBM and up to roughly 3.4 petaflops of BF16/FP16 compute. The 32 in the name describes the instance size, not the chip count.

trn1n.32xlarge is the network-optimised variant: it doubles EFA bandwidth to 1600 Gbps, which helps large multi-node jobs. Chips inside an instance talk over NeuronLink-v2. AWS positioned Trn1 as offering up to 50% lower cost-to-train than comparable GPU instances.

Take quiz
How many Trainium chips are in a trn1.32xlarge instance?
8
32
16
What distinguishes trn1n.32xlarge from trn1.32xlarge?
A newer NeuronCore generation
Double the EFA network bandwidth
More HBM per chip
It has no NeuronLink

4. What is Amazon EC2 Trn2?

Trn2 is the instance family built on Trainium2, generally available since December 2024. The standard size is trn2.48xlarge.

One instance holds 16 Trainium2 chips joined by NeuronLink-v3. That gives about 1.5 TB of HBM, 46 TB/s of memory bandwidth and roughly 20.8 PFLOPS of dense FP8 compute, with 3.2 Tbps of EFAv3 networking.

Each chip has 8 NeuronCore-v3 and 96 GiB of HBM. AWS said Trn2 gives 30-40% better price-performance than the GPU-based P5e and P5en instances, and that is the main reason teams evaluate it.

Take quiz
How much HBM does each Trainium2 chip have?
96 GiB
144 GiB
32 GiB
How many Trainium2 chips are in a trn2.48xlarge instance?
64
8
4
16

5. What is a Trn2 UltraServer?

A Trn2 UltraServer (trn2u.48xlarge) links four Trn2 instances into one scale-up system of 64 Trainium2 chips, using NeuronLink between the instances.

Together they offer about 83.2 PFLOPS of dense FP8 compute, 6 TB of HBM3 and 185 TB/s of memory bandwidth. Software sees one large, tightly connected domain.

That matters for models too big for 16 chips, or for tensor and expert parallelism that would otherwise cross slower network links. Beyond one UltraServer you scale out over EFA inside EC2 UltraClusters.

Take quiz
How many Trainium2 chips make up one Trn2 UltraServer?
64
32
144
16
What links the four instances inside a Trn2 UltraServer?
Only standard EFA over the data-center network
NeuronLink
A shared PCIe switch on the host CPU

6. What is Trainium3?

Trainium3 is AWS's third-generation Trainium chip and its first on a 3nm process. It powers Trn3 UltraServers, which reached general availability at re:Invent in December 2025.

A fully configured Trn3 UltraServer holds up to 144 chips, delivering 362 FP8 PFLOPS, 20.7 TB of HBM3e and 706 TB/s of memory bandwidth. Divided per chip that works out to roughly 2.5 PFLOPS FP8 and 144 GB of HBM3e.

Against Trn2 UltraServers, AWS claims up to 4.4x performance, 3.9x memory bandwidth and 4x better performance per watt. Inside the server, the NeuronSwitch-v1 fabric doubles interchip bandwidth.

Take quiz
What process node is Trainium3 built on?
5nm
3nm
7nm
What is the maximum number of chips in a Trn3 UltraServer?
256
64
144
72

7. What is the AWS Neuron SDK?

The AWS Neuron SDK is the software stack that lets you train and deploy models on Trainium and Inferentia. It sits between your framework and the hardware.

  • Framework plugins: torch-neuronx for PyTorch, plus JAX support.
  • Compiler: neuronx-cc turns graphs into executables for NeuronCores.
  • Runtime and driver: load and run compiled models on the devices.
  • Libraries: NeuronX Distributed (training) and NxD Inference (serving).
  • NKI: a Python-based kernel language for custom kernels.
  • Tools: neuron-ls, neuron-top, neuron-monitor and profilers.

It ships in Deep Learning AMIs and containers, so you rarely install it by hand.

Take quiz
Which Neuron SDK component compiles graphs for NeuronCores?
nxd-train
neuron-top
nrt-loader
neuronx-cc
What is the Neuron SDK's library for distributed training called?
NeuronX Distributed
Neuron Fabric Trainer
NeuronSwitch Core
NeuronLink Lightning

8. What is a NeuronCore?

A NeuronCore is the independent compute unit inside a Trainium or Inferentia chip. Think of it roughly like a very large processing core with its own engines and on-chip memory, not like a GPU thread.

Each NeuronCore contains four engines: Tensor (matrix multiplication), Vector, Scalar (activation functions) and GpSimd (programmable general-purpose). It also has on-chip SRAM and DMA engines for moving data in from HBM.

A Trainium1 chip has 2 NeuronCore-v2, while a Trainium2 chip has 8 NeuronCore-v3. Frameworks target NeuronCores, so your parallelism degree is usually counted in cores.

Take quiz
Which NeuronCore engine performs matrix multiplication?
GpSimd engine
Tensor engine
Vector engine
Scalar engine
How many NeuronCores does one Trainium2 chip have?
4
2
8
16

NeuronLink is the chip-to-chip interconnect that lets Trainium chips exchange data directly, without going through the host CPU or the instance network.

Trn1 uses NeuronLink-v2 between the 16 chips in an instance. Trn2 uses NeuronLink-v3, which also stretches across the four instances of an UltraServer. On Trn3 the NeuronSwitch-v1 fabric provides all-to-all connectivity.

It is not the same as EFA. NeuronLink serves the scale-up domain (inside an instance or UltraServer), while EFA carries traffic between separate instances and UltraServers.

10. What is the Neuron compiler?

The Neuron compiler, neuronx-cc, converts a framework graph (XLA HLO from PyTorch/XLA or JAX) into a NEFF, the Neuron Executable File Format that NeuronCores run.

It decides tiling, memory placement in on-chip SRAM, and the instruction schedule across the engines. Because that happens ahead of execution, tensor shapes must be known at compile time.

You tune it with the NEURON_CC_FLAGS environment variable:

export NEURON_CC_FLAGS="--model-type=transformer --auto-cast=matmul --auto-cast-type=bf16"

Compiled results are cached, so repeated runs skip the compile step. Compile time grows with graph size, so large models can take many minutes the first time.

Take quiz
What file format does neuronx-cc produce?
PTX
SavedModel
NEFF
ONNX
Which environment variable passes extra flags to the Neuron compiler?
NEURON_RT_FLAGS
XLA_COMPILER_OPTS
NEFF_BUILD_ARGS
NEURON_CC_FLAGS

11. What is the Neuron Kernel Interface (NKI)?

NKI is a Python-embedded kernel language for writing custom kernels that run directly on NeuronCores. It plays the role that CUDA or Triton plays on GPUs.

You write at the tile level: load tiles from HBM into on-chip memory, compute with the tensor, vector and scalar engines, and store results back. Kernels are decorated with @nki.jit and can be called from PyTorch or JAX.

Reach for NKI when the compiler's output for an operation, such as a fused attention variant, leaves performance on the table. For everyday models you rarely need it.

Take quiz
What is NKI mainly used for?
Provisioning Trn instances
Converting ONNX models
Monitoring chip temperature
Writing custom kernels for NeuronCores
Which language is NKI embedded in?
Python
Java
Verilog
C++ templates

12. What data types does Trainium support?

Trainium handles the usual training formats and adds some low-precision options aimed at LLM workloads.

  • FP32 and TF32 for higher-precision accumulation and sensitive operations.
  • BF16 and FP16 for most training and inference compute.
  • FP8 variants (configurable FP8) for higher throughput on newer chips.
  • Microscaling formats such as MXFP8 and MXFP4 arrive with Trainium3.

BF16 is the common default because it keeps FP32's exponent range, so it rarely needs loss scaling. Matrix multiplications accumulate in higher precision on chip, and stochastic rounding helps when weights are stored in BF16.

Take quiz
Why is BF16 popular for training on Trainium?
It needs no hardware support
It is only used for inference
It stores values with more mantissa bits than FP32
It keeps the same exponent range as FP32
Which low-precision microscaling formats come with Trainium3?
INT2 and INT3
Only INT8
FP64 and FP128
MXFP8 and MXFP4

13. What is NeuronX Distributed (NxD)?

NeuronX Distributed is a PyTorch library on top of torch-neuronx that supplies the building blocks for large-scale training and inference on Trainium.

It provides tensor parallel layers such as ColumnParallelLinear and RowParallelLinear, pipeline parallelism, sequence parallelism, expert parallelism for MoE, and ZeRO-1 optimizer sharding. It also handles distributed checkpointing.

NxD Training wraps these pieces into YAML-configured recipes for popular models, while NxD Inference is the serving counterpart.

Use it when a model is too large for one NeuronCore: you choose the parallel degrees, and NxD maps the ranks onto the chips and issues the needed collectives.

Take quiz
Which layer type does NxD provide for splitting a weight matrix by output dimension?
ColumnParallelLinear
ShardedBatchNorm
RowSplitEmbedding
StripeLinearReduce
What does ZeRO-1 shard across data-parallel ranks?
Only the learning rate
Compiled NEFF files
Optimizer states
Input tokens

14. What is NxD Inference?

NxD Inference (NxDI) is the Neuron library for serving large language models on Trainium and Inferentia2. You point it at a Hugging Face style checkpoint, and it compiles and loads a sharded, optimised copy.

It includes continuous batching, KV-cache management, bucketing, quantization, speculative decoding and other LLM serving features. It integrates with vLLM, so you get an OpenAI-compatible server without writing your own scheduler.

The first run compiles NEFFs, which can take a while. Save them and reuse them in later deployments.

Take quiz
Which serving framework does NxD Inference integrate with?
TensorFlow Serving
Triton Inference Server only
vLLM
TorchServe only
What takes the extra time on the first run of an NxDI model?
Compiling NEFFs
Training a draft model
Warming up EFA
Downloading CUDA kernels

15. What is Project Rainier?

Project Rainier is AWS's large-scale AI compute cluster built for Anthropic. It links Trainium2 UltraServers into an EC2 UltraCluster so Anthropic can train and serve Claude models.

AWS reported roughly 500,000 Trainium2 chips when it came online in October 2025, with plans to grow well beyond that. A data center campus in Indiana was the first site.

It is the standard example for answering questions on whether Trainium can scale: the same chips and Neuron stack that customers rent are running frontier-model workloads.

Take quiz
Which company is the anchor customer of Project Rainier?
Anthropic
OpenAI
Stability AI
Meta
Which chip generation powered Project Rainier at launch?
Inferentia2
Trainium2
Trainium1
Trainium3

16. What are the Neuron command-line tools?

Neuron ships a small set of CLI tools for checking devices and watching workloads:

  • neuron-ls lists the Neuron devices and their NeuronCores on the instance.
  • neuron-top shows live NeuronCore utilization and memory usage, similar in spirit to top.
  • neuron-monitor streams JSON telemetry, which exporters can forward to Prometheus or CloudWatch.
  • neuron-profile captures and visualizes execution traces; newer releases fold this into Neuron Explorer.

First thing to run on a fresh instance is neuron-ls. If it shows no devices, the driver or runtime is not set up correctly.

Take quiz
Which tool lists Neuron devices on an instance?
neuron-profile
neuron-monitor
neuron-ls
neuron-top
Which tool streams JSON telemetry for Prometheus or CloudWatch exporters?
neuron-ls
neuron-nki
neuron-cc
neuron-monitor

17. What is stochastic rounding in Trainium?

Stochastic rounding is a hardware rounding mode used when a value is cast down to a lower precision such as BF16. Rather than always rounding to the nearest value, it rounds up or down at random, with probability proportional to the discarded fraction.

Why bother? BF16 has only about 8 bits of mantissa. A tiny weight update, say 0.0001 added to 1.0, would be rounded away every time, and training would stall. With stochastic rounding, the update shows up on average, so expected value is preserved.

It is controlled by the runtime variable NEURON_RT_STOCHASTIC_ROUNDING_EN. It is most useful when you keep weights in BF16 without FP32 master copies.

Take quiz
What problem does stochastic rounding address in BF16 training?
Compiler cache misses
Small updates being rounded away
Missing exponent bits
Slow disk reads
How does stochastic rounding choose between two neighbouring values?
Always rounds up
Randomly, weighted by the discarded fraction
It alternates up and down in turn
Always rounds toward zero

18. What is Trainium4?

Trainium4 is the next Trainium generation that AWS previewed at re:Invent 2025. At announcement it was not generally available.

Compared with Trainium3, AWS cited roughly 6x FP4 performance, 3x FP8 performance, 4x memory bandwidth and 2x memory capacity. The most striking item is support for NVLink Fusion, which would let Trainium4 sit alongside NVIDIA GPUs in the same rack-scale design.

Treat the figures as vendor targets, not measured results. Check AWS's announcements for current availability and instance names before planning around it, since details tend to change between preview and launch.

Take quiz
Which interconnect technology is Trainium4 planned to support?
CXL only
Intel UPI
NVLink Fusion
Infinity Fabric
Which precision did AWS highlight with the biggest claimed gain for Trainium4?
FP64
INT16
FP4
TF32

19. How do you launch a Trainium instance?

The quickest route is an EC2 instance with the Neuron Deep Learning AMI:

  1. Pick a Region that offers the Trn family you need, then secure capacity: On-Demand, a Savings Plan, or an EC2 Capacity Block for ML for a time-boxed reservation.
  2. Choose an instance type, for example trn1.32xlarge or trn2.48xlarge.
  3. Select a Deep Learning AMI (Neuron) so drivers, runtime and framework environments are preinstalled.
  4. Add EFA networking and fast storage if you plan multi-node jobs.
  5. SSH in and run neuron-ls to confirm the devices are visible, then activate the PyTorch Neuron virtual environment.

For containers, use the Neuron DLC images with ECS or EKS instead.

Take quiz
Which command confirms the Neuron devices are visible after you log in?
lspci -v neuron
neuron-ls
aws neuron list
nvidia-smi
Which AMI type comes with Neuron drivers and frameworks preinstalled?
Bottlerocket ECS AMI
Amazon Linux minimal AMI
Deep Learning AMI (Neuron)
Windows Server base AMI

20. What are Neuron Deep Learning AMIs and containers?

They are prebuilt images that bundle everything you need on Trainium so you avoid dependency mismatches between drivers, runtime and frameworks.

The Deep Learning AMIs include the Neuron driver, runtime, tools and Python virtual environments for PyTorch and JAX. You boot an instance and start working.

The Deep Learning Containers package the same pieces as Docker images in Amazon ECR, with separate flavours for training and inference. They suit ECS, EKS and SageMaker, where the host only needs the Neuron driver.

Version pairing matters. Always match the container's Neuron SDK release to the driver installed on the host.

Take quiz
Where are Neuron Deep Learning Containers published?
AWS CodeArtifact
Amazon S3 Glacier
Amazon ECR
Docker Hub only
What must match between a Neuron container and its host?
The EBS volume type
The VPC CIDR block
The Linux desktop theme
The Neuron SDK release and driver version

21. How do you compile a PyTorch model for Trainium?

There are two routes, depending on whether you are training or serving.

For training, run the model on the XLA device through torch-neuronx. Graphs compile lazily the first time they run, or ahead of time with neuron_parallel_compile.

For inference, trace the model with an example input. The result is a TorchScript module that wraps the compiled NEFF:

import torch, torch_neuronx

model.eval()
example = torch.rand(1, 3, 224, 224)
neuron_model = torch_neuronx.trace(model, example)
neuron_model.save("model_neuron.pt")

The example input fixes the tensor shapes, so later inputs must match them.

Take quiz
Which torch-neuronx call compiles a model for inference?
torch_neuronx.trace
torch.neuron.export_cuda
torch_neuronx.fit
torch.compile_neff
What does the example input passed to trace() determine?
Which chip generation is used
The learning rate
The fixed tensor shapes of the compiled model
The Region of the instance

22. What is the difference between Trainium and Inferentia2?

Both use NeuronCore-v2 and the same Neuron SDK, but they target different jobs. Trainium handles training and inference; Inferentia2 is tuned for inference only.

Aspect Trainium (Trn1) Inferentia2 (Inf2)
Primary use Training, plus inference Inference only
Chips in largest instance 16 (trn1.32xlarge) 12 (inf2.48xlarge)
Accelerator memory 512 GB HBM 384 GB HBM
Can train models? Yes No

If you only serve models and want the lowest inference cost, Inf2 is usually the first thing to try. Later Trainium generations are the newer line AWS is investing in for both phases. Models compiled for one Neuron target can usually be moved to the other with a recompile, because the SDK and NeuronCore generation are shared.

Take quiz
Which chip family is designed for inference only?
Trainium2
Trainium1
Inferentia2
Trainium3
How many chips does the largest Inf2 instance, inf2.48xlarge, have?
16
24
12
8

23. What is the difference between Trn1 and Trn2?

Trn2 is a large step up in per-chip memory, NeuronCore count and, most importantly, the size of the connected domain.

Feature Trn1 Trn2
NeuronCore generation v2 v3
NeuronCores per chip 2 8
HBM per chip 32 GB 96 GiB
Chips per instance 16 16
Instance accelerator memory 512 GB 1.5 TB
Largest scale-up domain One instance (16 chips) UltraServer (64 chips)

AWS quoted up to 4x the performance of Trn1 for Trn2. In practice, Trn2's bigger memory per chip also means you need less model parallelism for the same model. An 8B model that needed tensor parallelism on Trn1 may fit on a few Trn2 cores, which cuts collective traffic and simplifies the setup.

Take quiz
How many NeuronCores does a Trn1 chip have compared with a Trn2 chip?
8 versus 2
4 versus 16
2 versus 4
2 versus 8
What is the largest scale-up domain offered with Trn2?
A 4-chip node
A 64-chip UltraServer
A 144-chip UltraServer
A single 16-chip instance only

24. What is the difference between Trn2 and Trn3?

Trn3 raises chip memory, compute and domain size, and replaces the point-to-point style NeuronLink fabric with a switched design.

Feature Trn2 UltraServer Trn3 UltraServer
Chips 64 144
Memory per chip 96 GiB HBM 144 GB HBM3e
Total HBM About 6 TB About 20.7 TB
FP8 compute 83.2 PFLOPS (dense) 362 PFLOPS
Intra-server fabric NeuronLink-v3 NeuronSwitch-v1 (all-to-all)

AWS claims up to 4.4x higher performance and 4x better performance per watt. The bigger domain lets you keep larger models, or more experts, inside one NeuronLink-class system. For teams already on Trn2, the Neuron SDK carries most code across, though bucket sizes and parallel degrees are worth retuning to the new memory and bandwidth.

Take quiz
Which fabric connects chips inside a Trn3 UltraServer?
PCIe daisy chain
NeuronSwitch-v1
InfiniBand NDR
NeuronLink-v1 ring
Approximately how much total HBM does a Trn3 UltraServer have?
6 TB
512 GB
20.7 TB
1.5 TB

25. How does Trainium differ from NVIDIA GPUs?

The biggest difference is the programming model. GPUs run many small kernels launched eagerly by the host, whereas Trainium compiles whole graphs ahead of time into one schedule for each NeuronCore.

Area Trainium NVIDIA GPU
Compute unit NeuronCore with Tensor, Vector, Scalar, GpSimd engines Streaming multiprocessors with CUDA/Tensor cores
On-chip memory Software-managed SRAM (SBUF/PSUM) Hardware-managed caches plus shared memory
Execution Ahead-of-time compiled graph (NEFF) Eager kernel launches, optional graphs
Custom kernels NKI CUDA, Triton
Availability AWS only All major clouds and on-premises

So Trainium rewards stable shapes and standard architectures, while GPUs are more forgiving of dynamic or unusual code.

Take quiz
How does Trainium execute models compared with a typical GPU workflow?
It compiles whole graphs ahead of time into one schedule
It launches every kernel eagerly from the host
It only runs precompiled ONNX files
It interprets Python bytecode on the chip
Which kernel language targets Trainium rather than NVIDIA GPUs?
CUDA C
HIP
NKI
Triton

26. Why does Trainium need static shapes?

Because the Neuron compiler fixes everything at compile time: tensor dimensions, memory layout in on-chip SRAM, DMA transfers and the instruction schedule on each engine. A NEFF is built for exact shapes, with no runtime kernel selection to adapt.

The consequences differ by workflow:

  • Training: a new shape produces a new graph, which triggers another compile. That can mean minutes of stall in the middle of a run.
  • Inference: a traced model rejects inputs that do not match its traced shape.

The usual workarounds are padding, drop_last=True on data loaders, avoiding data-dependent control flow, and bucketing for variable-length inputs. Treat shape stability as a design rule from day one, not a late optimization.

Take quiz
What happens in training when a new tensor shape appears mid-run?
The NeuronCores reshape their SRAM automatically
The runtime picks a different kernel instantly
Nothing, shapes are dynamic by default
The graph is recompiled
Which dataloader setting helps avoid a smaller last batch triggering a recompile?
drop_last=True
pin_memory=True
shuffle=False
num_workers=0

27. How do you handle variable-length inputs with bucketing on Trainium?

Bucketing means compiling a handful of fixed-length versions of the model and routing each request to the smallest one that fits. You pad the input up to that bucket length and mask the padding.

BUCKETS = [128, 256, 512, 1024, 2048]

def pick_bucket(n_tokens):
    return next(b for b in BUCKETS if b >= n_tokens)

The trade-off is padding waste versus resource cost. More buckets mean less wasted compute, but each bucket is another NEFF to compile and hold in HBM. Choose boundaries from your real length distribution rather than powers of two by habit.

NxD Inference lets you configure separate buckets for the prompt-processing and token-generation phases.

Take quiz
What is the main downside of adding many buckets?
Longer prompts get truncated
More NEFFs to compile and keep in memory
The model loses accuracy
NeuronLink bandwidth drops
How is a request routed in bucketing?
To the smallest compiled length that fits it
To whichever chip is coolest
To the largest bucket always
To a random bucket

28. How does the Neuron compilation cache work?

Neuron caches compiled NEFFs keyed on the input graph, the compiler flags and the compiler version. If all three match, the next run skips neuronx-cc and loads the stored artifact.

By default the cache lives in a local directory (/var/tmp/neuron-compile-cache). Set NEURON_COMPILE_CACHE_URL to an S3 location and every node in a cluster can share it, so only the first run pays for compilation.

A change to the model graph, flags or SDK version means a cache miss. That is why a harmless-looking code edit sometimes triggers a long recompile. neuron_parallel_compile can warm the cache before the real run starts.

Take quiz
Which setting lets multiple nodes share a Neuron compile cache?
Setting NEURON_RT_LOG_LEVEL=debug
Pointing NEURON_COMPILE_CACHE_URL at S3
Mounting the cache on tmpfs
Copying NEFFs through EFA at runtime
Which change causes a compile cache miss?
Changing the log file path
Running neuron-ls
Upgrading the Neuron SDK version
Restarting the instance with the same SDK

29. What is the Logical NeuronCore configuration on Trn2?

A Trn2 chip has 8 physical NeuronCores. The Logical NeuronCore (LNC) setting decides how they are presented to software.

  • LNC=2 (default): two physical cores are fused into one logical core, so each chip shows 4 larger logical cores with access to more HBM and combined compute.
  • LNC=1: all 8 cores are exposed separately, each with a smaller memory share.

You set it with NEURON_LOGICAL_NC_CONFIG. A trn2.48xlarge therefore shows 64 logical cores at LNC=2, or 128 at LNC=1.

Use LNC=2 for big models that need memory per rank. Use LNC=1 when the model fits comfortably and you want more data-parallel replicas.

Take quiz
How many logical NeuronCores does a trn2.48xlarge expose at the default LNC=2?
64
8
128
16
Which environment variable selects the logical NeuronCore configuration?
NEURON_LNC_MODE_FLAG
XLA_NEURON_LNC
NEURON_LOGICAL_NC_CONFIG
NEURON_CORE_SPLIT

30. How does tensor parallelism work in NxD?

NxD splits individual weight matrices across NeuronCores so each core holds only a slice. The key pair is ColumnParallelLinear, which splits the output dimension, and RowParallelLinear, which splits the input dimension.

In a transformer MLP, the first projection is column-parallel, so each rank computes its own partial activations with no communication. The second projection is row-parallel, so each rank produces a partial sum, and one all-reduce combines them. That is one collective per block.

Attention heads are divided across ranks the same way. With sequence parallelism, the all-reduce is replaced by a reduce-scatter and an all-gather, which also shards activations.

Keep the tensor-parallel group inside one NeuronLink domain, because these collectives run on every layer.

Take quiz
In a tensor-parallel MLP, which layer type usually follows ColumnParallelLinear?
A pipeline stage boundary
RowParallelLinear
A replicated embedding
Another ColumnParallelLinear
Which collective combines the partial sums after a row-parallel layer?
Broadcast
Barrier
Scatter
All-reduce

31. When would you choose Trainium over GPU instances?

Choose Trainium when cost per token or per training run is what matters most, and your model and team fit its constraints.

  • The model is a mainstream architecture (Llama-style, Mixtral-style MoE and similar) with Neuron support.
  • The workload is large and steady, so porting effort pays back through lower cost and reserved capacity.
  • You already run on AWS and want integration with EKS, SageMaker HyperPod and Bedrock.
  • Power efficiency and supply availability are part of the decision.

Stay on GPUs when you depend on CUDA-only libraries, run research code with highly dynamic shapes, or need portability across clouds. A short proof of concept with real shapes and real cost per million tokens settles the question better than spec-sheet comparison.

Take quiz
Which situation favours choosing GPUs over Trainium?
Your code depends on CUDA-only libraries
You serve a mainstream LLM at steady high volume
You already use SageMaker HyperPod
You want the lowest cost per token on AWS
What is the best way to decide between the two for a real workload?
Use whichever instance the team used last
Compare only peak TFLOPS on the datasheet
Benchmark cost per million tokens with real shapes
Pick whichever has a shorter name

32. How do you run Trainium workloads on Amazon EKS?

You add Trn nodes to the cluster and install the Neuron device plugin, which advertises Neuron resources to Kubernetes. Pods then request them like any other resource.

  1. Create a node group with a Trn instance type and an AMI that has the Neuron driver.
  2. Deploy the Neuron device plugin DaemonSet. Add the Neuron scheduler extension so multi-core pods get contiguous cores.
  3. Install the EFA device plugin for multi-node training.
  4. Build pods from a Neuron DLC image and request the resource:
resources:
  limits:
    aws.amazon.com/neuroncore: 8

You can request whole chips with aws.amazon.com/neuron instead. Deploy neuron-monitor as a DaemonSet to export metrics.

Take quiz
Which Kubernetes resource name requests individual NeuronCores?
aws.amazon.com/efa-core
nvidia.com/gpu
aws.amazon.com/neuroncore
neuron.io/chip-slice
What does the Neuron device plugin do on EKS?
Encrypts pod traffic
Compiles NEFF files for pods
Advertises Neuron devices to the Kubernetes scheduler
Replaces the kubelet

33. How does SageMaker HyperPod support Trainium?

SageMaker HyperPod provides persistent, resilient clusters for long training runs, and it supports Trn instances alongside GPU instances. You can orchestrate with Slurm or Amazon EKS.

The value is reliability. HyperPod runs health checks on nodes, replaces faulty ones, and can automatically resume a job from the latest checkpoint. On a run lasting weeks across hundreds of chips, hardware failures are expected, so this saves a lot of manual work.

You still bring your own training code, typically NxD Training or a PyTorch script, and your own checkpoints on shared storage such as FSx for Lustre or S3. Test a restore before the long run, so auto-resume actually works on the day a node fails.

Take quiz
What does HyperPod do when a node in a long run fails?
Restarts training from step zero always
Replaces the node and can resume from a checkpoint
Ignores the failure and continues with fewer chips silently
Moves the whole job to GPUs
Which orchestrators can HyperPod use with Trn instances?
Only Docker Swarm
Slurm or Amazon EKS
Only Apache Airflow
Only AWS Lambda

34. How does mixed precision with auto-cast work on Trainium?

Auto-cast is a compiler feature. With --auto-cast you tell neuronx-cc to lower FP32 operations to a faster type without editing your model.

  • --auto-cast=matmul casts only matrix-multiplication inputs.
  • --auto-cast=all casts every FP32 operation.
  • --auto-cast=none leaves precision untouched.
  • --auto-cast-type picks the target: bf16, fp16, tf32 or an FP8 format.

The matmul option is a safe default because accumulation still happens in higher precision on chip, and sensitive operations such as softmax and normalization stay in FP32. Many teams skip auto-cast and cast the model to torch.bfloat16 directly, which is more explicit. Whichever route you pick, compare loss curves against an FP32 baseline for a few hundred steps before committing to a long run.

Take quiz
Which flag setting casts only matrix-multiplication inputs?
--auto-cast=all
--model-type=matmul
--auto-cast=none
--auto-cast=matmul
Why is --auto-cast=matmul considered the safer choice?
It removes the need for a compiler
It disables compilation caching
It forces FP8 everywhere
Sensitive ops like softmax stay in FP32

35. How do you serve LLMs on Trainium with vLLM?

vLLM has a Neuron backend that uses NxD Inference underneath, so you get continuous batching and an OpenAI-compatible endpoint on Trn instances.

  1. Launch a Trn instance and pull the Neuron vLLM container (or use the Neuron DLAMI).
  2. Choose a tensor-parallel degree that fits the model across the NeuronCores.
  3. Set max-model-len and max-num-seqs. These feed bucket and batch sizes, which are compiled in.
  4. Start the server and wait for the first-run compile.
vllm serve meta-llama/Llama-3.1-8B-Instruct \
  --tensor-parallel-size 8 \
  --max-model-len 4096 \
  --max-num-seqs 16

Exact flags vary by Neuron release, so check the release notes. Cache the compiled artifacts so restarts are fast.

Take quiz
Why do max-model-len and max-num-seqs matter for vLLM on Neuron?
They set the AWS Region
They determine compiled shapes and bucket sizes
They disable continuous batching
They control EFA firmware
What slows down the very first start of a vLLM Neuron server?
Warming up a CDN
Compiling the model into NEFFs
Downloading CUDA drivers
Provisioning a GPU

36. How does ZeRO-1 reduce memory on Trainium?

ZeRO-1 shards the optimizer states across data-parallel ranks instead of copying them to every rank. Weights and gradients stay replicated, so it is the cheapest ZeRO stage to apply.

Adam keeps an FP32 master copy plus two moments, about 12 bytes per parameter. For an 8-billion-parameter model that is around 96 GB. With 16 data-parallel ranks, each holds roughly 6 GB instead.

The flow: each rank updates only its shard, then an all-gather redistributes the updated weights. The added communication is modest, and the freed HBM often allows a larger micro-batch or less tensor parallelism.

NxD provides a ZeRO-1 optimizer wrapper that you apply around your regular optimizer.

Take quiz
Which step redistributes the updated weights after each rank updates its own optimizer shard?
A global barrier only
A checkpoint reload
A broadcast from rank 0
An all-gather
Roughly how much optimizer memory does each rank hold for an 8B model with Adam sharded over 16 ranks?
About 600 MB
About 96 GB
About 6 GB
About 24 GB

37. Explain the execution flow of PyTorch/XLA lazy tensors on Trainium?

On Trainium, PyTorch runs through the XLA backend, so operations on XLA tensors do not execute immediately. They are recorded into a lazy graph, and nothing runs on the NeuronCores until the graph is cut and submitted.

flowchart LR
  A["Python ops on XLA tensors"] --> B["Lazy graph recorded"]
  B --> C["mark_step or sync point"]
  C --> D{Graph hash in cache?}
  D -- Yes --> F["Load cached NEFF"]
  D -- No --> E["HLO to neuronx-cc to NEFF"]
  E --> F
  F --> G["Runtime executes on NeuronCores"]
  G --> H["Results available to host"]

The graph is cut at xm.mark_step() (called for you by Neuron's parallel data loader and optimizer helpers) or when the host needs a value, as with .item() or print(tensor). The hash of the graph decides whether a compiled NEFF already exists.

Two practical effects follow. First, a stray .item() inside the step splits it into several smaller graphs and adds sync stalls. Second, anything that changes the graph between steps, such as shapes or Python-side branching on tensor values, creates a new hash and a recompile.

for batch in loader:            # MpDeviceLoader marks steps for you
    loss = model(batch).loss
    loss.backward()
    xm.optimizer_step(optimizer)
    optimizer.zero_grad()

Take quiz
When does lazy-tensor work actually execute on the NeuronCores?
When the graph is cut at a mark_step or sync point
Only at program exit
As soon as each Python operation is written
Whenever a new Python module is imported
Why is a stray .item() call inside a training step harmful?
It deletes the compile cache
It switches the job to CPU permanently
It forces a sync and splits the step into extra graphs
It disables BF16

38. Explain the internal working of the NeuronCore engines?

A NeuronCore is built around four engines that run in parallel and share one on-chip SRAM, so good performance comes from keeping all of them busy.

  • Tensor engine: a 128x128 systolic array for matrix multiplication and convolutions. Results land in PSUM, a small accumulation buffer.
  • Vector engine: element-wise work, reductions and normalizations across the partition dimension.
  • Scalar engine: activation functions such as exp, GELU and sigmoid, applied in a pipeline.
  • GpSimd engine: small programmable cores for operations that do not map to the other three, including custom ops.

Data comes from HBM through DMA engines into SBUF, the main software-managed SRAM, organised as 128 partitions. The compiler schedules DMA and compute so they overlap, with semaphores enforcing dependencies.

flowchart LR
  H[HBM] <--> D["DMA engines"]
  D <--> S["SBUF, 128 partitions"]
  S --> T["Tensor engine"]
  T --> R["PSUM accumulators"]
  R --> V["Vector engine"]
  R --> C["Scalar engine"]
  V --> S
  C --> S
  G["GpSimd engine"] <--> S

A softmax after attention is a typical case: matmul on Tensor, exp on Scalar, sums on Vector, all pipelined over tiles.

Take quiz
What is PSUM used for in a NeuronCore?
Holding the host OS page tables
Storing the compiled NEFF
Accumulating tensor engine matmul results
Buffering EFA packets
Which engine typically applies activation functions like GELU or exp?
DMA engine
GpSimd engine
Tensor engine
Scalar engine

39. Explain the lifecycle of a NEFF from compilation to execution?

A NEFF (Neuron Executable File Format) is the compiled package a NeuronCore runs. Its life has two halves: building it, and loading and running it.

flowchart TD
  A["Framework graph: HLO"] --> B["neuronx-cc passes: tiling, layout, scheduling"]
  B --> C["Code generation per engine"]
  C --> D["NEFF: instruction streams, DMA descriptors, weight layout, metadata"]
  D --> E["Runtime load: allocate HBM, copy weights"]
  E --> F["Runtime execute: bind inputs, run, return outputs"]
  1. Compile: the compiler tiles operations to fit SBUF, chooses memory layouts and schedules instructions across the four engines.
  2. Package: the NEFF holds per-engine instruction streams, DMA descriptors, weight data and metadata about inputs and outputs.
  3. Load: the Neuron runtime allocates HBM, copies weights in and prepares queues. This is why each loaded NEFF costs device memory.
  4. Execute: the runtime binds input and output buffers and starts execution. Collective operations go through the Neuron collectives library.

Loading is the slow, one-time step. Execution just replays a fixed schedule, which is why steady-state latency is predictable.

Take quiz
What does the Neuron runtime do during the load phase of a NEFF?
Allocates HBM and copies weights onto the device
Opens a connection to the compile cache on S3 per token
Recompiles the graph for the current batch
Rewrites the Python source
Which item is NOT stored inside a NEFF?
Per-engine instruction streams
The Python training script
Weight layout data
DMA descriptors

40. How do you troubleshoot slow training on Trainium?

Work from the cheapest check to the deepest one, and confirm each theory with a number.

  1. Look at utilization. Run neuron-top. Low NeuronCore utilization with a busy CPU points to a host bottleneck.
  2. Check for recompiles. Repeated compiler messages or cache misses mean shapes or graph structure are unstable. Fix padding or drop_last.
  3. Inspect the data pipeline. Slow tokenization or storage reads starve the chips. Use the parallel device loader and prefetch.
  4. Look at sync points. Remove .item(), logging of tensors and gradient-norm checks from the inner loop, or do them every N steps.
  5. Profile. Use neuron-profile (or Neuron Explorer) to see idle engines, DMA-bound regions and collective wait time.
  6. Revisit parallelism and precision. Excess tensor-parallel degree makes collectives dominate. Training in FP32 wastes the tensor engine, so move to BF16.

Change one thing at a time and compare tokens per second, since several fixes interact.

Take quiz
Which observation suggests the host CPU, not the NeuronCores, is the bottleneck?
High NeuronCore utilization with idle CPU
A short compile cache path
Large HBM free space
Low NeuronCore utilization while the CPU is busy
What is a common cause of repeated recompilation during training?
Unstable input shapes between steps
Running neuron-ls
Having too many CPU cores
Using BF16 weights

41. How do you troubleshoot long compile times or compiler failures on Trainium?

Start by finding out whether the compile is slow because the graph is large, or because the cache is being missed.

  • Read the compiler log. neuronx-cc writes a log and workdir; the failing pass and operator are usually named there.
  • Shrink the graph. Unrolled loops and very deep models produce huge graphs. Reduce Python-side unrolling and use model-level wrappers where Neuron supports them.
  • Try faster flags. Use a lower optimization level while iterating, then raise it for the final run.
  • Compile in parallel. neuron_parallel_compile extracts graphs from a trial run and compiles them concurrently.
  • Reuse artifacts. Point NEURON_COMPILE_CACHE_URL at S3 so one node's compile benefits the others.
  • Give it memory. The compiler is RAM-hungry, so a larger instance (or fewer parallel jobs) avoids out-of-memory kills.

For an operator the compiler rejects, check the release notes for support, then either rewrite it with supported ops or implement it in NKI.

Take quiz
Which tool compiles graphs concurrently by extracting them from a trial run?
nki.jit
neuron-ls
neuron-top
neuron_parallel_compile
What is a good first step when neuronx-cc fails on a model?
Delete the NEURON_CC_FLAGS variable and retry blindly
Read the compiler log to find the failing pass or operator
Switch to a GPU immediately
Reinstall the operating system

42. How do you troubleshoot HBM out-of-memory errors on Trainium?

First, work out what is consuming memory, because the fix differs. neuron-top shows device memory split into categories such as model code, tensors and scratch space.

Likely cause Fix
Weights, gradients and optimizer states too big Raise tensor or pipeline parallel degree; enable ZeRO-1; use BF16
Activations too big Enable activation checkpointing; lower micro-batch; use sequence parallelism
KV cache too big (inference) Reduce max sequence length or batch size; quantize the KV cache
Too many loaded NEFFs Cut the number of buckets; unload unused models
Cores too small on Trn2 Use LNC=2 so each logical core has more HBM

Remember that every loaded NEFF reserves device memory for instructions and scratch buffers. A server with ten buckets can run out of memory even though weights alone fit.

Take quiz
Why can a server with many buckets run out of HBM even when the weights fit?
Buckets duplicate the S3 cache in HBM
Buckets disable ZeRO-1
Each loaded NEFF reserves device memory
Padding tokens are stored on the CPU
Which setting gives each logical core more HBM on Trn2?
LNC=2
NEURON_RT_STOCHASTIC_ROUNDING_EN=0
LNC=1
--model-type=generic

43. How can you optimize LLM inference throughput on Trainium?

Decode is memory-bandwidth bound, so most gains come from keeping the chips busy with useful batches and moving fewer bytes per token.

Technique What it helps
Continuous batching Fills batch slots as sequences finish, raising utilization
Chunked prefill Stops long prompts from stalling ongoing decodes (better TPOT)
Prefix caching Skips recomputing shared system prompts
Right-sized buckets Cuts padding waste, tuned to real length histograms
FP8 or INT8 weights Moves fewer bytes per token and fits more KV cache
Speculative decoding Produces several tokens per target-model pass
Tuned tensor-parallel degree Balances per-token latency against chips used per replica

Measure tokens per second per chip, time to first token and time per output token under realistic traffic. A configuration that wins on a single request often loses once batching is on.

Take quiz
Why is decode usually limited by memory bandwidth rather than compute?
Each generated token must read the full weights and KV cache
NeuronLink is too slow for single tokens
The compiler disables DMA
The Tensor engine cannot do matmul at small batch
Which technique avoids recomputing a shared system prompt across requests?
ZeRO-1
Stochastic rounding
Prefix caching
Bucket padding

44. How does speculative decoding work on Trainium?

A small draft model guesses the next k tokens cheaply. The large target model then checks all k in a single forward pass and keeps the longest prefix it agrees with. Output quality is unchanged, but several tokens can be produced per target pass.

sequenceDiagram
  participant D as Draft model
  participant T as Target model
  D->>D: Propose k tokens
  D->>T: Send k candidate tokens
  T->>T: One forward pass scores all k
  T-->>D: Accept longest matching prefix plus 1 new token
  D->>D: Continue from the accepted position

On Trainium, NxD Inference supports draft-model and EAGLE-style speculation. Both models are compiled and loaded, and the speculation length k is baked into the NEFF because shapes are static. Changing k means recompiling.

It pays off when decode is memory-bound and the draft agrees often. A low acceptance rate wastes the draft's work, so measure it on your own prompts before enabling.

Take quiz
Why does changing the speculation length require recompilation on Trainium?
The driver restricts k to 4 at boot
NeuronLink must be retrained
The draft model is stored in S3
k is part of the compiled static shape
What does the target model do in speculative decoding?
Generates the draft tokens itself one by one
Compresses the KV cache
Verifies all proposed tokens in one forward pass
Selects the instance type

45. How do you write an efficient NKI kernel?

Think in tiles. SBUF has 128 partitions, so the first dimension of a tile is the partition dimension and is capped at 128. Everything else is the free dimension.

import neuronxcc.nki as nki
import neuronxcc.nki.language as nl

@nki.jit
def add_kernel(a, b):
    out = nl.ndarray(a.shape, dtype=a.dtype, buffer=nl.shared_hbm)
    a_t = nl.load(a[0:128, 0:512])   # HBM -> SBUF
    b_t = nl.load(b[0:128, 0:512])
    nl.store(out[0:128, 0:512], value=a_t + b_t)  # SBUF -> HBM
    return out

Import paths shift between Neuron releases, so check your version's docs. The ideas stay the same:

  • Size tiles to fit SBUF while keeping the partition dimension full, since a half-empty tile wastes lanes.
  • Loop with nl.affine_range so the compiler can pipeline iterations.
  • Accumulate matmul partial results in PSUM instead of writing back after each step.
  • Overlap DMA with compute by double-buffering tiles.
  • Profile and check which engine is idle.
Take quiz
What is the maximum size of the partition dimension of a tile?
128
512
32
1024
Where should partial matmul results be accumulated in an NKI kernel?
The EFA send buffer
PSUM
HBM after every step
The host CPU memory

46. How does collective communication work across Trainium chips?

Collectives such as all-reduce, all-gather, reduce-scatter and all-to-all are issued by the framework as XLA collective operations. The Neuron collectives library maps them onto the physical links, and the compiler overlaps them with compute where it can.

Parallelism Typical collective Preferred link
Tensor parallel All-reduce or reduce-scatter/all-gather NeuronLink or NeuronSwitch
Expert parallel (MoE) All-to-all NeuronLink or NeuronSwitch
Data parallel Gradient all-reduce or reduce-scatter EFA across instances
Pipeline parallel Point-to-point sends EFA across stages

Replica groups come from NxD's parallel state, which maps ranks to NeuronCores. Place the chatty groups (tensor and expert) inside one NeuronLink domain, and let the less frequent data-parallel traffic cross EFA.

Within a Trn2 instance, chips form a 2D torus over NeuronLink-v3. On Trn3, NeuronSwitch-v1 offers all-to-all connectivity, which helps patterns like MoE dispatch.

Take quiz
Which parallelism strategy should usually stay inside a single NeuronLink domain?
Log shipping
Tensor parallelism
Checkpoint replication
Data parallelism across Regions
Which collective is typical for Mixture-of-Experts token dispatch?
All-to-all
Barrier only
Gather to rank 0
Broadcast of weights

47. How do you scale training across multiple Trn2 UltraServers?

Use the fast scale-up domain for the chatty parallelism and the scale-out network for the quiet kind.

  1. Inside each UltraServer (64 chips over NeuronLink): tensor parallelism and, for MoE, expert parallelism.
  2. Across UltraServers (EFA in an EC2 UltraCluster): data parallelism and, for very deep models, pipeline parallelism.
  3. Use ZeRO-1 to cut optimizer memory across the data-parallel dimension.
  4. Launch inside a cluster placement setup so nodes get low-latency EFA paths.
  5. Compile once and share the cache through S3, so hundreds of nodes do not each recompile.
  6. Write sharded checkpoints to shared storage such as FSx for Lustre, and test restore early.
  7. Validate with a collective benchmark before the real run, to catch slow links.

NxD Training expresses this as a YAML config: degrees for tensor, pipeline, expert and data parallelism, plus batch sizes. Scale in stages, checking tokens per second at each size, because efficiency often dips when you first cross the EFA boundary.

Take quiz
Which parallelism type is best kept across UltraServers over EFA?
Tensor parallelism
Data parallelism
Sequence parallelism within attention
Expert all-to-all at every layer
Why share the compile cache through S3 in a large cluster?
So nodes do not each repeat the same compilation
To increase NeuronLink bandwidth
To replace checkpoints
To enable FP8

48. What happens when a Trainium node fails during a long training run?

Usually the failing rank's process dies or the runtime reports a device error, and the remaining ranks hang inside a collective waiting for it. After a timeout, the whole job fails; collective training does not continue with a missing rank.

flowchart TD
  A["Node or device error"] --> B["Other ranks block in collective"]
  B --> C["Timeout, job fails"]
  C --> D["Orchestrator replaces node"]
  D --> E["Job restarts"]
  E --> F["Load latest checkpoint"]
  F --> G["Training resumes"]

Recovery depends on tooling. On SageMaker HyperPod, health checks replace the node and auto-resume restarts the job. On EKS or Slurm, you rely on elastic launchers or your own restart scripts.

What you lose is the work since the last checkpoint, so checkpoint frequency is a cost trade-off: frequent saves cost throughput, rare ones cost redone steps. Sharded or asynchronous checkpoints to shared storage keep that overhead small.

Take quiz
What do surviving ranks typically do when one rank dies mid-collective?
Switch to CPU training
Silently redistribute the missing shard
Ignore the collective and proceed
Block waiting, then time out
What determines how much work is lost after a node failure?
How long since the last checkpoint
The number of buckets
The EFA firmware date
The Neuron SDK version

49. How can you reduce the cost of training on Trainium?

Cost per run is price per hour multiplied by hours, so attack both.

  • Lower precision: train in BF16 and evaluate FP8 where the recipe supports it, so each chip does more work per hour.
  • Avoid recompiles: stable shapes and a shared compile cache stop expensive instances sitting idle while compiling.
  • Right-size parallelism: over-sharding makes collectives dominate. Pick the smallest tensor-parallel degree that fits memory.
  • Buy capacity wisely: Savings Plans for steady use, EC2 Capacity Blocks for scheduled runs.
  • Use Spot for tolerant jobs: fine-tuning with frequent checkpoints can handle interruptions.
  • Pack small jobs: use LNC=1 to run more replicas or experiments per chip.
  • Track the right metric: cost per trained token, not just hourly price.

Compile time is a quiet cost: a 30-minute compile across 16 idle instances is real money.

Take quiz
Which practice keeps expensive instances from idling during startup?
Disabling checkpoints
Using FP32 everywhere
Raising tensor parallel degree to the maximum
Sharing a warmed compile cache
Which metric best reflects the true cost efficiency of a training run?
Size of the AMI
Cost per trained token
Number of NeuronCores
Instance hourly price alone

50. Why do Mixture-of-Experts models benefit from Trn2 and Trn3 UltraServers?

MoE layers route each token to a few experts, and those experts usually live on different chips. Every MoE layer therefore needs an all-to-all to dispatch tokens and another to combine results, in both the forward and backward passes.

That traffic is heavy and latency-sensitive. Over a general network it can dominate step time. Inside an UltraServer, all chips share the NeuronLink domain (64 chips on Trn2, up to 144 on Trn3 with NeuronSwitch-v1), so the all-to-all stays on the fast fabric.

Example: a model with 64 experts can place one expert per chip in a Trn2 UltraServer, and each token's top-2 selection becomes two hops inside the domain.

The large pooled HBM matters as well. A Trn2 UltraServer holds about 6 TB and a Trn3 UltraServer about 20.7 TB, enough for big expert sets plus their KV caches during serving.

Take quiz
Which collective does an MoE layer rely on to dispatch tokens to experts?
Barrier
Broadcast from rank 0
All-to-all
Point-to-point checkpoint copy
Why is the UltraServer domain valuable for expert parallelism?
Expert traffic stays on the fast NeuronLink fabric
It removes the need for a router
It stores experts in S3
It disables all-to-all entirely
«
»

Comments & Discussions