AI / LLM Basics Interview Questions
What is a Large Language Model (LLM)?
A Large Language Model is a neural network trained on massive amounts of text to predict the next word, or more precisely the next token, in a sequence. That simple prediction task, repeated across billions of examples, is enough to teach the model grammar, facts, reasoning patterns, and style.
- Built on the Transformer architecture, which processes text using an attention mechanism rather than reading word by word
- "Large" refers to both the size of the training data and the number of parameters, often in the billions
- Can be adapted after initial training through fine-tuning for specific tasks or behaviors
Examples include models like GPT, Claude, and Llama, all of which share this same core next-token-prediction foundation despite differing in size, training data, and fine-tuning approach.
More Related questions...