Prev Next

AI / LLM Basics Interview Questions

What is Causal Masking?

Causal masking is a technique used in decoder-style LLMs that prevents a token from attending to any tokens that come after it in the sequence.

  • Ensures the model can only use information from earlier in the text when predicting the next token
  • Necessary because generation happens one token at a time, left to right, so a token genuinely can't have access to future tokens it hasn't generated yet
  • Implemented by masking out, effectively zeroing out, the attention scores between a token and any later position

This is what keeps an LLM's next-token prediction honest during training, it can't "cheat" by peeking ahead at the answer it's supposed to be predicting.

What does causal masking prevent a token from doing?
Why is causal masking necessary for generation?

More Related questions...

What is a Large Language Model (LLM)? What are Tokens in an LLM? What is Tokenization? What is Byte-Pair Encoding (BPE)? What is the Vocabulary of an LLM? What is an Embedding in the context of LLMs? Define the Transformer architecture? What is the Attention Mechanism? What is Self-Attention? What is Multi-Head Attention? What is Positional Encoding? What is the purpose of the Feed-Forward Network inside a Transformer block? What is Layer Normalization? What is Causal Masking? Describe the role of Softmax in an LLM's output layer? What is Cross-Entropy Loss? What is the purpose of a Loss Function during LLM training? What is Pretraining in the context of LLMs? What is Fine-Tuning? What is Instruction Tuning? Describe Reinforcement Learning from Human Feedback (RLHF)? What is Alignment in the context of LLMs? Define In-Context Learning? Define Zero-Shot Learning? Define Few-Shot Learning? What is Chain-of-Thought Prompting? What is Prompt Engineering? What is the purpose of a System Prompt? What is a Context Window? What is the Max Tokens parameter? What are Stop Sequences? What is Temperature in LLM sampling? What are Top-k and Top-p Sampling? What is Hallucination in LLMs? What is Perplexity in language modeling? What is Overfitting in the context of training an LLM? What are Parameters in an LLM? What is Quantization? What is Model Distillation? What is a Foundation Model? What is a Decoder-Only model? What is an Encoder-Only model? What is an Encoder-Decoder model? What is a Mixture of Experts (MoE) architecture? What is Retrieval-Augmented Generation (RAG)? How are Embeddings used beyond text generation? What are Guardrails in LLM applications? What is Prompt Injection?
Show more question and Answers...


Comments & Discussions