AI / LLM Basics Interview Questions
What is Causal Masking?
Causal masking is a technique used in decoder-style LLMs that prevents a token from attending to any tokens that come after it in the sequence.
- Ensures the model can only use information from earlier in the text when predicting the next token
- Necessary because generation happens one token at a time, left to right, so a token genuinely can't have access to future tokens it hasn't generated yet
- Implemented by masking out, effectively zeroing out, the attention scores between a token and any later position
This is what keeps an LLM's next-token prediction honest during training, it can't "cheat" by peeking ahead at the answer it's supposed to be predicting.
More Related questions...