AI / LLM Basics Interview Questions
What is the Max Tokens parameter?
Max Tokens is a setting that limits how many tokens a model is allowed to generate in a single response, capping the output length.
- Doesn't affect how much input the model can read, only how much it's allowed to write back
- Setting it too low can cut off a response mid-sentence before the model finishes its thought
- Commonly used to control cost and response time, since longer outputs take more compute and more time to generate
This parameter is typically set by whoever is calling the model through an API, giving developers direct control over how long a generated response is allowed to run.
More Related questions...