Prev Next

AI / LLM Basics Interview Questions

What is Multi-Head Attention?

Multi-Head Attention runs several self-attention operations in parallel, each with its own learned parameters, then combines their results together.

  • Each "head" can learn to focus on a different kind of relationship, one might track grammatical structure, another might track topical relevance
  • Running these in parallel, rather than one attention operation alone, lets the model capture multiple types of relationships between tokens at once
  • The outputs from all heads are concatenated and passed through a final linear layer to produce one combined representation

This is a standard part of every Transformer block, and it's a major reason a single attention layer can capture such a rich, multi-faceted understanding of a sequence.

What does each attention head in Multi-Head Attention potentially focus on?
What happens to the outputs of all the heads?

Invest now in Acorns!!! 🚀 Join Acorns and get your $5 bonus!
Acorns Logo

Invest now in Acorns!!! 🚀
Join Acorns and get your $5 bonus!

Earn passively and while sleeping

Acorns is a micro-investing app that automatically invests your "spare change" from daily purchases into diversified, expert-built portfolios of ETFs. It is designed for beginners, allowing you to start investing with as little as $5. The service automates saving and investing. Disclosure: I may receive a referral bonus.

Robinhood Logo

Invest now!!! Get Free equity stock (US, UK only)!

Use Robinhood app to invest in stocks. It is safe and secure. Use the Referral link to claim your free stock when you sign up!.

The Robinhood app makes it easy to trade stocks, crypto and more.


Webull Logo

Webull! Receive free stock by signing up using the link: Webull signup.

More Related questions...

What is a Large Language Model (LLM)? What are Tokens in an LLM? What is Tokenization? What is Byte-Pair Encoding (BPE)? What is the Vocabulary of an LLM? What is an Embedding in the context of LLMs? Define the Transformer architecture? What is the Attention Mechanism? What is Self-Attention? What is Multi-Head Attention? What is Positional Encoding? What is the purpose of the Feed-Forward Network inside a Transformer block? What is Layer Normalization? What is Causal Masking? Describe the role of Softmax in an LLM's output layer? What is Cross-Entropy Loss? What is the purpose of a Loss Function during LLM training? What is Pretraining in the context of LLMs? What is Fine-Tuning? What is Instruction Tuning? Describe Reinforcement Learning from Human Feedback (RLHF)? What is Alignment in the context of LLMs? Define In-Context Learning? Define Zero-Shot Learning? Define Few-Shot Learning? What is Chain-of-Thought Prompting? What is Prompt Engineering? What is the purpose of a System Prompt? What is a Context Window? What is the Max Tokens parameter? What are Stop Sequences? What is Temperature in LLM sampling? What are Top-k and Top-p Sampling? What is Hallucination in LLMs? What is Perplexity in language modeling? What is Overfitting in the context of training an LLM? What are Parameters in an LLM? What is Quantization? What is Model Distillation? What is a Foundation Model? What is a Decoder-Only model? What is an Encoder-Only model? What is an Encoder-Decoder model? What is a Mixture of Experts (MoE) architecture? What is Retrieval-Augmented Generation (RAG)? How are Embeddings used beyond text generation? What are Guardrails in LLM applications? What is Prompt Injection?
Show more question and Answers...

Database

Comments & Discussions