Prev Next

AI / LLM Basics Interview Questions

What is the purpose of the Feed-Forward Network inside a Transformer block?

Each Transformer block contains a feed-forward network that processes every token's representation individually after the attention step has mixed in context from the rest of the sequence.

  • Applies the same two-layer transformation to each token's vector independently
  • Adds additional representational capacity and non-linearity that attention alone doesn't provide
  • Works alongside attention in every block, attention gathers context across tokens, the feed-forward network then further processes each token's resulting representation

Together, the attention and feed-forward steps in each block are what let a Transformer refine its understanding of a sequence layer by layer as information flows through the network.

When does the feed-forward network process a token's representation?
Does the feed-forward network process tokens individually or all mixed together?

More Related questions...

What is a Large Language Model (LLM)? What are Tokens in an LLM? What is Tokenization? What is Byte-Pair Encoding (BPE)? What is the Vocabulary of an LLM? What is an Embedding in the context of LLMs? Define the Transformer architecture? What is the Attention Mechanism? What is Self-Attention? What is Multi-Head Attention? What is Positional Encoding? What is the purpose of the Feed-Forward Network inside a Transformer block? What is Layer Normalization? What is Causal Masking? Describe the role of Softmax in an LLM's output layer? What is Cross-Entropy Loss? What is the purpose of a Loss Function during LLM training? What is Pretraining in the context of LLMs? What is Fine-Tuning? What is Instruction Tuning? Describe Reinforcement Learning from Human Feedback (RLHF)? What is Alignment in the context of LLMs? Define In-Context Learning? Define Zero-Shot Learning? Define Few-Shot Learning? What is Chain-of-Thought Prompting? What is Prompt Engineering? What is the purpose of a System Prompt? What is a Context Window? What is the Max Tokens parameter? What are Stop Sequences? What is Temperature in LLM sampling? What are Top-k and Top-p Sampling? What is Hallucination in LLMs? What is Perplexity in language modeling? What is Overfitting in the context of training an LLM? What are Parameters in an LLM? What is Quantization? What is Model Distillation? What is a Foundation Model? What is a Decoder-Only model? What is an Encoder-Only model? What is an Encoder-Decoder model? What is a Mixture of Experts (MoE) architecture? What is Retrieval-Augmented Generation (RAG)? How are Embeddings used beyond text generation? What are Guardrails in LLM applications? What is Prompt Injection?
Show more question and Answers...


Comments & Discussions