Here is something that surprises most business owners the first time they hit it: you can paste a 40-page supplier contract into an AI chatbot, ask a sharp question about clause 12, and get a perfect answer. Ask a follow-up an hour later in the same chat, and the AI acts like it has never seen the contract at all. Nothing broke. You simply ran into the edge of the context window, the single most important limit that decides what your AI can and cannot handle.
What is a context window in AI?
A context window is the maximum amount of text an AI model can hold in its working memory at one time, measured in units called tokens. Everything counts toward it: your question, any documents you paste, and the AI's own replies. It is the model's short-term memory for a single conversation, and once it fills up, the oldest information drops out of view.
Think of it like a whiteboard in a meeting room. The board holds a fixed amount of writing. As long as your notes fit, everyone can see the full picture. When the board fills up and you keep writing, you have to erase the top to make room at the bottom. The AI does the same thing, and the erased part is simply gone from that conversation.
How does a context window actually work?
A context window works by converting all text into tokens, then processing only as many tokens as the window allows. According to IBM, a token is roughly 3 to 4 characters, or about 0.75 of an English word. A 128,000-token window can hold several hundred pages of text at once.
Every time you send a message, the AI re-reads the entire conversation inside the window from scratch to decide its reply. It has no memory outside that window.
Here is what the window has to hold at the same time:
- Your instructions and questions
- Any files, emails, or data you pasted in
- The AI's previous answers in the same chat
- Hidden system rules that shape how it behaves
When the total crosses the limit, the model quietly forgets the earliest material to keep the newest in view.
Why does AI seem to forget what you just told it?
AI forgets because the conversation has grown longer than the context window, so the earliest details have dropped out to make space. It is not a glitch and it is not the AI being lazy. The information physically no longer fits inside the memory the model can see.
There is a second, subtler reason called the lost in the middle effect. Research consistently shows that models recall information best from the beginning and the end of the window, and worst from the middle. So even when a detail technically still fits, an AI can overlook it if it is buried in the centre of a very long conversation.
For a business owner, the practical lesson is simple: put your most important instruction at the start or the very end of your message, not sandwiched in the middle of a wall of text.
How big are context windows in 2026?
In 2026, context windows range from around 128,000 tokens on entry-level models to over 1 million on the top tier. To picture 1 million tokens, imagine roughly 1,500 pages of text, the length of several thick novels, held in view at once.
The jump has been dramatic and recent. In July 2026, Meta's Muse Spark 1.1 shipped with a 1-million-token window, and DeepSeek released V4 models built around the same 1-million-token capacity for long, multi-step tasks. Google's Gemini has pushed to 2 million tokens, while flagship models from OpenAI and Anthropic sit in the 400,000 to 1-million-token range.
For your business, a bigger window means the AI can read a whole year of invoices, an entire staff handbook, or a long customer chat history in one pass, instead of you feeding it a few pages at a time.
It helps to know that this size is usually a choice, not a fixed fact about "AI." The same brand often sells a small-window model for quick, cheap tasks and a large-window model for heavy documents. Part of using AI well is picking the right size for the job, rather than assuming every model can swallow anything you throw at it.
What does the context window mean for a small business?
For a small business, the context window decides how much of your real-world material the AI can handle in a single task, from long contracts to full customer histories. Matching the window to the job is the difference between a useful answer and a confident but incomplete one.
Here are concrete examples a Hong Kong SME owner will recognise:
- A restaurant owner pastes a 30-page tenancy agreement and asks the AI to flag every clause about rent increases. A large window reads all 30 pages; a small one may only see the first few.
- A retail shop owner feeds three months of WhatsApp customer chats to spot the most common complaint. This only works if the window is big enough to hold the whole log.
- A property agent uploads a full building management report to draft a summary for clients. A long report can exceed a small window and get cut off mid-way.
The takeaway is not that bigger is always necessary. For a quick email reply, a small window is plenty. The skill is knowing when the job needs a bigger one.
There is a cost angle too. Because most tools charge by the token, pushing a giant document into a huge window on every small question quietly runs up your bill. A shop owner who understands the window will save the heavyweight model for the heavyweight jobs, and reach for a cheap, small-window model for the dozens of quick questions in between.
How can you work around a small context window?
You can work around a small context window by feeding the AI less at once, summarising as you go, and starting fresh chats for new topics. These simple habits often matter more than which model you pay for.
Four practical habits that make any AI tool behave better:
- Break big documents into parts. Instead of pasting a 60-page report at once, feed one section, get the answer, then move to the next. Smaller chunks stay well inside the window.
- Ask for a running summary. Before a long chat fills up, ask the AI to summarise the key points so far. Paste that summary into a new chat and you carry the essentials forward without the bulk.
- Start a new chat for a new task. Do not stretch one endless conversation across unrelated jobs. A fresh window keeps the AI focused and avoids the lost in the middle problem.
- Put the ask first. Lead with what you want, then supply the background. The model reads the opening most reliably.
A restaurant owner reviewing a year of supplier emails does not need the biggest model on the market. Feeding one month at a time into a modest window, with a summary carried between rounds, gets the job done at a fraction of the cost.
What are the most common misconceptions about context windows?
The most common misconception is that a context window is permanent memory. It is not. A context window is temporary and single-conversation only. When you start a fresh chat, the window is wiped clean and the AI remembers nothing from before.
Three more myths worth clearing up:
- "Bigger is always better." A larger window costs more to run and can dilute focus. A 1-million-token window is wasted on a one-line question.
- "A 1-million-token window means the AI reads everything perfectly." The lost in the middle effect means detail buried in a huge input can still be missed.
- "Context window and training are the same thing." Training is the general knowledge the model learned before you met it. The context window is only what you show it right now.
Frequently asked questions about context windows
Short, direct answers to the questions business owners ask most about context windows, covering memory, cost, and everyday use.
Does the AI remember my last conversation?
No, not through the context window alone. Once a chat ends, that window is cleared. Some tools add a separate "memory" feature on top, but the basic context window resets every new conversation.
Does a bigger context window cost more?
Usually yes. Most AI services charge by the token, so feeding a huge document into a large window uses more tokens and costs more than a short prompt.
How many pages is 1 million tokens?
Roughly 1,500 pages of plain text, though images, tables, and formatting use tokens too, so the real figure varies.
What happens if I go over the limit?
The AI drops the oldest content to make room, or the tool warns you and asks you to start a new chat. Either way, the material that falls out is no longer considered.
The bottom line for your business
The context window is the invisible ruler behind almost every AI task you will ever run. Understand it, and the mysterious moments when AI "forgets" stop being frustrating and start being predictable. You will know when to start a fresh chat, where to place your key instruction, and which jobs need a heavyweight model versus a quick one.
None of this requires a technical background. It just requires someone to explain it in plain language, which is exactly how we think good technology should be shared. We understand AI. UD stands with you.
Put this into practice with UD
Knowing what a context window is helps you use everyday AI tools better. The next step is turning that understanding into real workflows for your business, from feeding AI the right documents to picking the right tool for each job. UD has walked Hong Kong SMEs through exactly this for 28 years, and we will walk you through it step by step, from your first question to a working setup.