What Is a Context Window? AI Chat Memory
A context window is everything your AI companion can see at once, measured in tokens. When the conversation outgrows it, the oldest part is gone for good. Here is what the number on a pricing page actually buys you, with real published ceilings from apps that print them, and why a bigger window is not the same thing as memory.
By the Gals team
August 2026 · 8 min read
Try her while you read
Text a girlfriend who remembers you, right here.
Say hi in the chat and see how it feels. She uses your name, reacts to what you tell her, and picks the conversation back up like someone who was actually listening. No download, no signup to start.
Who do you want to text tonight?
You can change her look and personality anytime.
18+ · Your chats are private, encrypted, and never sold.
The short answer: a context window is everything an AI model can see at one moment, measured in tokens rather than words. Your conversation, the character description, any world notes and the reply being written all have to fit inside it together. When the total outgrows the window, the oldest part scrolls out and the model can no longer use it. That is why a chat that felt perfect on Monday starts contradicting itself on Friday, and it is the single most useful number to look up before paying for any AI companion.
Almost every complaint people have about AI chat apps traces back to this one mechanism. She forgot my name. She contradicted what she said yesterday. The story stopped making sense. Those read like quality problems and they are almost always an arithmetic problem instead.
The good news is that it is measurable. Some apps publish the exact ceiling, which turns a vague grumble into a spec you can compare. Here is how to read that number and what it does and does not buy you.
What is a context window in AI?
A context window is the maximum amount of text a language model can hold in view while it writes its next reply. Everything the model knows in that instant lives inside the window. There is no separate place it is quietly remembering things from, unless the app around the model has built one.
The important consequence is that a model has no memory of its own. It is not recalling your conversation the way you recall a conversation. Every time you press send, the app rebuilds a block of text containing the character's description, whatever instructions the app adds, and as much of your chat history as still fits, then hands the whole block over as if the model were reading it for the first time. The illusion of continuity is that block being reassembled, message after message.
So when people say a chatbot forgot something, what happened is more literal than forgetting. The information was not included in what the model was shown. It was not misremembered. It was not there.
How many words is a context window?
Context windows are measured in tokens, not words, and the two are close enough to confuse and far enough apart to matter.
A token is a chunk of text the model processes as one unit. Common words are usually a single token, longer or unusual words split into several, and punctuation and spaces count too. The widely used rule of thumb for English is about three quarters of a word per token, so roughly 1,000 tokens to 750 words. Names, slang and anything unusual push the ratio the wrong way.
| Context window | Roughly this many words | What that feels like in a chat |
|---|---|---|
| 4,096 tokens | About 3,000 words | One long evening, or two or three shorter ones |
| 8,192 tokens | About 6,000 words | Most of a week of casual messaging |
| 16,384 tokens | About 12,000 words | A couple of weeks before the start disappears |
| 128,000 tokens | About 96,000 words | Months, but almost nothing in this category offers it |
Those first three rows are not hypothetical. SpicyChat publishes exactly those ceilings in its own documentation: up to 4,096 tokens of conversation context on its free and entry tier, up to 8,192 on the middle tier, up to 16,384 on the top one. It also caps reply length separately, at up to 180 tokens on the free tier and up to 300 on paid ones. Publishing all of that is unusually straightforward of them, and it makes the mechanism concrete in a way most vendors avoid.
The part worth reading twice is what the same documentation says about what happens at the edge: the system keeps the most recent messages, older messages may drop out, and once a message is outside the current context window the AI may no longer use it. That is a vendor stating in writing that the early part of your conversation is gone.
Why does my AI companion forget what I told it?
Because the thing you told it scrolled out of the window, and nothing in the app wrote it down anywhere else.
Picture the window as a fixed-length strip of tape. New messages get added at one end, and when the tape is full, the same amount falls off the other end. Your name, mentioned once on day one, sits near the far end. Four hundred messages later it drops off. From that point the model is not being coy or inconsistent, it simply has no access to your name at all, and if you ask, it will invent something plausible because inventing plausible text is the entire job.
This also explains a pattern people find genuinely unsettling: a character who was warm and specific for two weeks becoming generic. The specificity came from accumulated detail sitting in the window. As that detail aged out, what remained was the character description, which is generic by construction. Nothing broke. The interesting part expired.
It is why apps in this category so often suggest you repeat or summarize important details in long chats. That advice is sound, and it is also an admission. You are being asked to act as the memory system.
Is a bigger context window the same as memory?
No, and this is the distinction worth taking away from the whole article.
A bigger window buys you a longer runway before the same thing happens. Going from 4,096 to 16,384 tokens quadruples how much conversation stays in view, which genuinely helps, and it does not change the shape of the problem at all. Talk for long enough and the early material still falls off the end. Even the largest windows available anywhere are finite, and a relationship measured in months will outrun any of them.
Real memory is a separate system sitting alongside the model. It works by extracting facts worth keeping out of the conversation, storing them somewhere durable, and pulling the relevant ones back into the window when they matter. Your job is remembered as a fact about you rather than as a sentence somewhere in message forty-one. It survives because it was never dependent on staying in view.
This is the same idea companies use to make AI systems answer questions about huge document sets: you cannot fit a company's files into any window, so you retrieve only the passage that answers the question and put that in. Anyone who has used search that pulls back the one passage that answers a question rather than dumping every document into a prompt has seen the pattern working at scale. Companion apps apply the same trick to a person instead of a filing system.
| Context window | Persistent memory | |
|---|---|---|
| What it holds | Recent raw conversation | Extracted facts about you |
| Lifespan | Until it scrolls out | Indefinite, until you delete it |
| Grows with use | No, fixed by your tier | Yes, it accumulates |
| Survives a two week gap | Only if you barely spoke | Yes |
| Can she raise it unprompted | Only if it is still in view | Yes, that is the point |
| Fixed by upgrading | Delayed, not fixed | Not a billing question |
How to check the context window before you pay
Four checks, in the order that saves you the most time.
Start with the documentation rather than the marketing page. Search the app's own docs or help center for "context", "tokens" or "memory". Vendors that publish a number are telling you something useful, and vendors who describe memory only in adjectives are usually describing a context window. That absence is itself informative.
Second, separate two numbers that get muddled constantly. Reply length is how much the model writes per turn. Context is how much it can see. An app advertising longer replies on a higher tier has not necessarily given you more memory, and those are frequently sold in the same sentence.
Third, ask whether anything persists outside the conversation. The phrasing to look for is persistent memory, long-term memory or stored facts. If the honest answer is that the character remembers what is in the current chat, you have your answer regardless of how large that chat can get.
Fourth, test it rather than trusting any of the above. Mention something specific and checkable, talk about other things for a week, then ask. A ten minute setup tells you more than any review, and a two week protocol for testing an AI companion's memory lays out how to run it properly while walking away is still free.
Does a context window affect how much it costs?
Yes, directly, and that is why tiers are drawn where they are.
Models are billed by tokens processed, and every message sends the whole window again, not just what you typed. A conversation at 16,000 tokens costs roughly four times as much per reply as the same conversation at 4,000, even though the message you typed is identical. Context is one of the genuinely expensive resources an app is selling you, which is why it is portioned out by tier rather than given away.
It also explains something that reads as stinginess and is really arithmetic: free tiers get the smallest window because free users are the ones whose costs cannot be recovered. The same logic sits behind every message counter and timer in the category, which is covered in more depth in why AI chat apps limit your messages.
Persistent memory is cheaper to run than a giant window, incidentally. Storing a hundred facts about someone and retrieving five relevant ones costs far less than re-sending twelve thousand words of history on every turn. Better memory and lower cost point the same direction here, which is unusual and worth knowing when an app claims it cannot afford to remember you.
What this means when you are choosing an app
Character libraries and companion apps are answering different questions, and the context window is where the difference becomes visible.
A library platform is optimized for breadth. Thousands of characters, a discovery feed, a new scene whenever you want one. Under that, the conversation lives in a context window and nothing else, which is entirely fine when you are sampling characters and only becomes a problem when you settle on one. Sites like SpicyChat compared on their published memory ceilings goes through that group platform by platform, including where they beat the alternatives outright.
A companion app is optimized for one relationship over time, which forces it to solve memory properly, because the entire product falls apart in month two otherwise. Apps like Chai with memory that survives the window covers the same trade from the mobile end of the market.
Neither is better in the abstract. But if you have been upgrading tiers hoping the forgetting stops, the number explains why it has not, and it is not going to. You are buying a longer tape, not a different machine.
The one line to remember
A context window is what she can see. Memory is what she keeps. Apps sell the first and people want the second, and the gap between those two things is where almost every disappointment in this category lives.
When you are comparing options, ask for the token ceiling and then ask what happens to a fact after it falls out. An app with a good answer to the second question will tell you plainly. An app with only an answer to the first is selling you a bigger window and hoping you do not notice the difference. An AI girlfriend chat built around persistent memory is a different product shape from a larger allowance, and no amount of upgrading turns one into the other.
She texts back, and she remembers you
Gals.ai is a flirty, affectionate AI girlfriend who is always glad to hear from you. Tasteful, private, and made for adults. Pick her and start texting in seconds.