Context window
Context window - how much text (in tokens) the model holds at a time: both your prompt and its response. Overfilled - the model “forgets” the beginning.
The context window is the amount of text that the model “sees” at the same time: the system prompt, the dialogue history, attached documents and the response itself. Measured in tokens. For modern 2026 models, these are hundreds of thousands of tokens—entire books.
Why is this important in practice? If you dump a huge brief, three articles into the chat and ask to also take into account the past dialogue, you can hit the ceiling of the window, and the model will begin to lose what it had at the beginning. Symptom - responses float, early instructions are ignored.
My approach: don’t confuse “big window” with “throw everything in there.” The clearer and shorter the context, the more accurate the answer. It is better to give the model 2 necessary documents than 20 just in case.
Related terms
Where is it understood in practice?
Need to set this up on your project?
I analyze metrics, calculate unit economics and collect funnels on real budgets. 30 minutes on call - free.