How AI companion memory actually works — and what we could prove
Every platform in this category advertises memory, and almost every one of them means something narrower than the word suggests. What is usually being sold is a context window — the model can see the current conversation — dressed in language that implies a companion who knows you.
The distinction is testable, so we test it. We plant three specific facts, change the subject, ask for them back, then correct one of them and check which version returns. This guide explains the mechanism underneath, sets out what that protocol established platform by platform, and is honest about the part nobody has proven: what happens tomorrow.
See how these platforms rank → Independent editorial scores from the Atlas deskThree different things called “memory”
The context window. The current conversation, held in front of the model while it writes. Recall inside it is cheap and nearly universal — it is not really memory, it is the model reading back up the page. This is what most “3/3 recall” demonstrations actually show, including ours.
Persistent facts. A stored profile the companion can consult regardless of the conversation: a backstory, key details, directives about how to behave. This is closer to memory in the everyday sense, and it is usually written or curated rather than learned.
Retrieval across time. Reaching back into old conversations, days or weeks later, and surfacing the relevant part. This is the hard one, the one plan cards imply, and the one none of our testing has established.
Most disappointment in this category comes from paying for the third and receiving the first.
Documented is not the same as demonstrated
Kindroid publishes the most detailed memory vocabulary in the category. Its subscription guide describes longer short-term context for subscribers, Cascaded Memory, enhanced long-term recall and larger backstory and memory fields, with the Ultra and MAX add-ons advertising progressively larger context still. Its August update log adds Learned Context and memory-consolidation work.
That is more structure than anyone else names publicly, and it is genuinely useful for understanding how these systems are put together: a short-term window, a medium-term layer that has to decide what is worth keeping, and a long-term store that has to be searched.
It is also, in our hands, entirely unmeasured. No companion was created in our connected Kindroid account, so initial recall, a corrected fact, recall after another turn and recall after a new session are all NT — not tested. The free tier is documented as offering basic long-term memory without Cascaded Memory, which means the layer doing the interesting work is the one you cannot reach without paying.
We are stating this plainly because it is the trap the rest of this guide is about. A published architecture is evidence that a company has thought about memory. It is not evidence that the memory works, and the two are easy to confuse when only one of them is written down.
The test almost everyone passes
Seed three distinctive facts — a name, a pet, a plan — change the subject, then ask for them back in the same conversation. Nine of the platforms we tested returned all three.
Secrets.ai, Candy AI, OurDream AI and Dreamz.ai each returned 3/3, and several did it while staying in character rather than reciting a list — OurDream AI‘s companion folded the details back in while maintaining a demanding-boss persona.
Swipey AI passed it too, in French: the companion returned the seeded name, the dog and a Saturday 20:00 appointment.
This result is worth less than it looks. It establishes that the context window is working, which is close to table stakes. Passing it does not distinguish a good product from a mediocre one.
The test that actually separates them
Change one of the seeded facts mid-conversation — a meeting moves from 2:30 to 3:15 — then ask again later. Now the model has to prefer the newer value over an older one sitting in the same context, rather than replaying the first thing it saw. This is a much better proxy for whether conversational state is being updated rather than merely stored.
Four platforms passed it: Joi AI, GPTGirlfriend, FLIRTcam.AI and Xotic AI all returned 3:15 rather than 2:30. On FLIRTcam.AI the corrected value survived an adult-boundary test in between, and on Joi AI it survived a direct stop instruction — the three facts came back intact immediately after the roleplay was ended, so complying with the stop did not cost the conversation its state.
Swipey AI sits between the two outcomes and is worth recording precisely. We changed the pet to Nova and moved the appointment to Sunday 19:30; the companion acknowledged the updated set, and the revised day survived a later turn. The revised time was not independently rechecked after that turn, so we score it as a partial rather than borrowing the benefit of the doubt.
On the four platforms where we recorded 3/3 without a correction, the correction step simply was not run. That is a gap in our coverage, not a finding about the product.
One platform failed, and how it failed matters
Promptchan returned two of three seeded details, and the planted 2:30 appointment came back as 5:00 in English and 8:00 in French. Not a refusal, not an “I don’t recall” — two different confident wrong answers in two languages.
That failure mode is the one worth understanding, because it is the one you will not notice. A companion that forgets is obvious. A companion that confidently substitutes a plausible value looks exactly like a companion that remembers, right up until the detail matters. Anyone relying on one of these products to hold personal context accurately should test for substitution, not for silence.
What the twelve did on the same protocol
| Platform | Same-session recall | Correction replaced the old value | Beyond the session |
|---|---|---|---|
| Joi AI | 3/3 | Pass — 3:15 returned, and survived a stop instruction | NT · “8,000-message smart memory” advertised on longer plans |
| GPTGirlfriend | 3/3 | Pass — 3:15 replaced 2:30 | NT · Elite advanced-memory claim not compared |
| FLIRTcam.AI | 3/3 | Pass — survived an adult-boundary test in between | NT · Premium’s better-memory model not tested |
| Xotic AI | 3/3 | Pass — 3:15 returned | NT · no return visit, no model switch |
| Secrets.ai | 3/3 | NT — step not run | NT · basic memory documented on the free tier |
| OurDream AI | 3/3 | NT — step not run | NT |
| Dreamz.ai | 3/3 | NT — step not run | NT · “unlimited chat & memory” is plan wording |
| Candy AI | 3/3 | NT — step not run | NT |
| Swipey AI | 3/3 — tested in French | Partial — the revised day survived; the revised time was not rechecked | NT · long-delay and cross-character memory not tested |
| Promptchan | 2/3 | Fail — 2:30 became 5:00 in English, 8:00 in French | NT |
| Kindroid | NT — no companion created in our account | NT | NT · Cascaded Memory, enhanced long-term recall and Learned Context documented, none measured |
| DarLink AI | NT — free allowance exhausted first | NT | NT · Basic, Enhanced and Living memory sold as plan tiers |
NT means not tested, not failed — an untested platform earns nothing here, in either direction. Seven platforms we rank are absent from this table altogether: Xtease AI, Lovescape, Dondi AI, Dream Companion, Kupid AI, GoLove AI and Get Harder were reviewed after this protocol ran and have not been through it. Every result is one conversation on one day with one character, and does not certify other characters, models or future versions. Tested between 14 and 21 August 2026.
Nobody has proven memory across days
Look down the third column: it is entirely NT. That is not an oversight, it is a structural consequence of how these products are sold.
Testing memory across days requires returning to the same conversation tomorrow, which requires an allowance that is still there tomorrow. Most free tiers are exhausted in a single session — five lifetime messages on Candy AI and DarLink AI, four on Dreamz.ai, ten on Swipey AI before the paywall. Kindroid documents unlimited messages on its Lite model, which would be the obvious place to run the test — but no companion was created in our account, and Kindroid’s own update log schedules Lite for retirement on 1 October 2026. Our free tier guide sets out where each one stops.
So the honest state of the evidence is this: same-session recall is common and cheap to demonstrate, corrected recall is meaningfully harder and passed by four platforms, and durable memory remains a claim on a plan card everywhere. We will report it when we have paid for it and returned the next day.
Memory as a plan feature
Because it is hard to verify, memory is convenient to sell. DarLink AI sells Basic, Enhanced and Living memory as three tiers of the same subscription. Joi AI advertises an “8,000-message smart memory” on its longer plans. Dreamz.ai‘s cards say “unlimited chat & memory”. GPTGirlfriend‘s Elite tier claims advanced memory.
None of these is necessarily untrue. All of them are unverified in our testing, and all of them are the kind of claim that is straightforward to make and awkward for a buyer to check. Treat a memory tier as a reason to subscribe monthly first rather than as a reason to commit to a year.
Test it yourself in five minutes
Plant three specific things in one message — a name, a detail with a number, a plan with a date. Specific beats plausible: a dog called Pixel is a better test than a dog, because a model can guess the second.
Change the subject properly for several exchanges. Recall immediately after planting proves nothing.
Ask for all three back, then correct one and continue talking.
Ask again. Getting the corrected value is the pass. Getting the original is a partial. Getting a third value that was never said — as we saw on Promptchan — is the result that should decide your answer.
Then come back tomorrow and ask once more. That is the test that matters most, the one we cannot run for you on a free tier, and the one worth spending the first paid month on.
Memory is the feature most likely to be oversold in this category, because the cheap version of it — recall inside the current conversation — is indistinguishable from the expensive version for about ten minutes. Nine platforms passed that ten-minute test. Four passed the harder one, and one came close. One returned confident wrong answers in two languages. The platform with the most detailed published architecture is one we could not measure at all.
What none of them has shown us is the thing the marketing implies. Until a platform demonstrates that it remembers you across days, treat every memory claim on a plan card as a claim, run the correction test yourself in the first session, and buy a month before you buy a year.
AI companion memory — FAQ
⚠ Every result described here is a single conversation, with one character, on one day, and does not certify other characters, models, languages or future versions. Platform behaviour and plan claims change regularly — everything was tested between 14 and 21 August 2026. These services are intended for adults, and AI companions are synthetic systems, not people or professional support.