Why an answer gets shorter or slower
An output budget caps how long a single reply may be. A context window caps how much of the earlier conversation the model still sees. Throttling slows delivery when hardware is busy. The first shortens the answer, the second makes it forgetful, and the third only makes it late.
People describe all 3 as the assistant getting worse, and the 3 have nothing in common except the feeling. Separating them turns a vague complaint into something you can act on.
On this page
The output budget cuts the reply, not the memory
Every service sets a maximum length for one reply, measured in tokens rather than in words. When the reply reaches it, generation stops, sometimes mid-sentence. The model still knows the rest; it was not allowed to write it. Asking it to continue usually works, because the budget applies per reply rather than per conversation.
The context window makes the model forget the start
The window is how much earlier text the model can still read. When a long conversation passes it, the oldest messages fall out of view. The model then answers as if it never saw them, which reads as carelessness rather than as a ceiling. Starting a new conversation with a short summary fixes it; asking the model to try harder does not.
Throttling changes speed and nothing else
When hardware is busy, tokens arrive more slowly. The answer is the same answer, delivered late. If the text is complete and correct but crawls, the ceiling you met is capacity, and it lifts on its own when load drops.
Which one you met, in one question
Ask whether the answer is short, forgetful, or slow. Short points at the output budget, and continuing usually helps. Forgetful points at the window, and a fresh conversation helps. Slow points at throttling, and waiting helps. The 3 fixes are different, which is why naming the ceiling is worth the moment it takes.
Fair questions
Does a longer question shorten the answer?
It can. Question and answer share the same window, so a very long question leaves less room for the reply on services that budget them together.
Is a cut-off answer a bug?
Usually it is the output budget doing its job. A bug looks like an answer that stops and cannot be continued at all.
Do paid plans get bigger windows?
Often, because a larger window costs more to run. This is one of the few places where paying changes the capability rather than the quantity.
Ask something long and see where this chat stops, if it stops.
Open unrestricted chat