AI / Claude OPUS5 Interview questions
Why is thinking on by default a breaking change when migrating to Claude Opus 5?
Because max_tokens caps thinking and the visible reply together as one shared budget, a request that previously produced no thinking at all on Opus 4.8 now consumes part of that same budget on thinking before Opus 5 even starts on the visible answer.
If an application's max_tokens value was tuned assuming the full budget went to the visible reply, as it did by default on Opus 4.8, that same value on Opus 5 can leave less room than expected for the answer - the request still returns successfully with HTTP 200, but with stop_reason: max_tokens and the answer cut off mid-sentence, rather than throwing an obvious error.
This makes it a quiet breaking change: nothing in the API contract technically fails, so it's easy to miss during migration testing unless truncated responses are specifically checked for on requests that previously fit comfortably within budget.
More Related questions...