Editorial guide · August 2026

How AI companion memory actually works — and what we could prove

Every platform in this category advertises memory, and almost every one of them means something narrower than the word suggests. What is usually being sold is a context window — the model can see the current conversation — dressed in language that implies a companion who knows you.

The distinction is testable, so we test it. We plant three specific facts, change the subject, ask for them back, then correct one of them and check which version returns. This guide explains the mechanism underneath, sets out what that protocol established platform by platform, and is honest about the part nobody has proven: what happens tomorrow.

See how these platforms rank → Independent editorial scores from the Atlas desk

Three different things called “memory”

The context window. The current conversation, held in front of the model while it writes. Recall inside it is cheap and nearly universal — it is not really memory, it is the model reading back up the page. This is what most “3/3 recall” demonstrations actually show, including ours.

Persistent facts. A stored profile the companion can consult regardless of the conversation: a backstory, key details, directives about how to behave. This is closer to memory in the everyday sense, and it is usually written or curated rather than learned.

Retrieval across time. Reaching back into old conversations, days or weeks later, and surfacing the relevant part. This is the hard one, the one plan cards imply, and the one none of our testing has established.

Most disappointment in this category comes from paying for the third and receiving the first.

Documented is not the same as demonstrated

Kindroid publishes the most detailed memory vocabulary in the category. Its subscription guide describes longer short-term context for subscribers, Cascaded Memory, enhanced long-term recall and larger backstory and memory fields, with the Ultra and MAX add-ons advertising progressively larger context still. Its August update log adds Learned Context and memory-consolidation work.

That is more structure than anyone else names publicly, and it is genuinely useful for understanding how these systems are put together: a short-term window, a medium-term layer that has to decide what is worth keeping, and a long-term store that has to be searched.

It is also, in our hands, entirely unmeasured. No companion was created in our connected Kindroid account, so initial recall, a corrected fact, recall after another turn and recall after a new session are all NT — not tested. The free tier is documented as offering basic long-term memory without Cascaded Memory, which means the layer doing the interesting work is the one you cannot reach without paying.

We are stating this plainly because it is the trap the rest of this guide is about. A published architecture is evidence that a company has thought about memory. It is not evidence that the memory works, and the two are easy to confuse when only one of them is written down.

The test almost everyone passes

Seed three distinctive facts — a name, a pet, a plan — change the subject, then ask for them back in the same conversation. Nine of the platforms we tested returned all three.

Secrets.ai, Candy AI, OurDream AI and Dreamz.ai each returned 3/3, and several did it while staying in character rather than reciting a list — OurDream AI‘s companion folded the details back in while maintaining a demanding-boss persona.

Swipey AI passed it too, in French: the companion returned the seeded name, the dog and a Saturday 20:00 appointment.

This result is worth less than it looks. It establishes that the context window is working, which is close to table stakes. Passing it does not distinguish a good product from a mediocre one.

The test that actually separates them

Change one of the seeded facts mid-conversation — a meeting moves from 2:30 to 3:15 — then ask again later. Now the model has to prefer the newer value over an older one sitting in the same context, rather than replaying the first thing it saw. This is a much better proxy for whether conversational state is being updated rather than merely stored.

Four platforms passed it: Joi AI, GPTGirlfriend, FLIRTcam.AI and Xotic AI all returned 3:15 rather than 2:30. On FLIRTcam.AI the corrected value survived an adult-boundary test in between, and on Joi AI it survived a direct stop instruction — the three facts came back intact immediately after the roleplay was ended, so complying with the stop did not cost the conversation its state.

Swipey AI sits between the two outcomes and is worth recording precisely. We changed the pet to Nova and moved the appointment to Sunday 19:30; the companion acknowledged the updated set, and the revised day survived a later turn. The revised time was not independently rechecked after that turn, so we score it as a partial rather than borrowing the benefit of the doubt.

On the four platforms where we recorded 3/3 without a correction, the correction step simply was not run. That is a gap in our coverage, not a finding about the product.

One platform failed, and how it failed matters

Promptchan returned two of three seeded details, and the planted 2:30 appointment came back as 5:00 in English and 8:00 in French. Not a refusal, not an “I don’t recall” — two different confident wrong answers in two languages.

That failure mode is the one worth understanding, because it is the one you will not notice. A companion that forgets is obvious. A companion that confidently substitutes a plausible value looks exactly like a companion that remembers, right up until the detail matters. Anyone relying on one of these products to hold personal context accurately should test for substitution, not for silence.

What the twelve did on the same protocol

PlatformSame-session recallCorrection replaced the old valueBeyond the session
Joi AI3/3Pass — 3:15 returned, and survived a stop instructionNT · “8,000-message smart memory” advertised on longer plans
GPTGirlfriend3/3Pass — 3:15 replaced 2:30NT · Elite advanced-memory claim not compared
FLIRTcam.AI3/3Pass — survived an adult-boundary test in betweenNT · Premium’s better-memory model not tested
Xotic AI3/3Pass — 3:15 returnedNT · no return visit, no model switch
Secrets.ai3/3NT — step not runNT · basic memory documented on the free tier
OurDream AI3/3NT — step not runNT
Dreamz.ai3/3NT — step not runNT · “unlimited chat & memory” is plan wording
Candy AI3/3NT — step not runNT
Swipey AI3/3 — tested in FrenchPartial — the revised day survived; the revised time was not recheckedNT · long-delay and cross-character memory not tested
Promptchan2/3Fail — 2:30 became 5:00 in English, 8:00 in FrenchNT
KindroidNT — no companion created in our accountNTNT · Cascaded Memory, enhanced long-term recall and Learned Context documented, none measured
DarLink AINT — free allowance exhausted firstNTNT · Basic, Enhanced and Living memory sold as plan tiers

NT means not tested, not failed — an untested platform earns nothing here, in either direction. Seven platforms we rank are absent from this table altogether: Xtease AI, Lovescape, Dondi AI, Dream Companion, Kupid AI, GoLove AI and Get Harder were reviewed after this protocol ran and have not been through it. Every result is one conversation on one day with one character, and does not certify other characters, models or future versions. Tested between 14 and 21 August 2026.

Nobody has proven memory across days

Look down the third column: it is entirely NT. That is not an oversight, it is a structural consequence of how these products are sold.

Testing memory across days requires returning to the same conversation tomorrow, which requires an allowance that is still there tomorrow. Most free tiers are exhausted in a single session — five lifetime messages on Candy AI and DarLink AI, four on Dreamz.ai, ten on Swipey AI before the paywall. Kindroid documents unlimited messages on its Lite model, which would be the obvious place to run the test — but no companion was created in our account, and Kindroid’s own update log schedules Lite for retirement on 1 October 2026. Our free tier guide sets out where each one stops.

So the honest state of the evidence is this: same-session recall is common and cheap to demonstrate, corrected recall is meaningfully harder and passed by four platforms, and durable memory remains a claim on a plan card everywhere. We will report it when we have paid for it and returned the next day.

Memory as a plan feature

Because it is hard to verify, memory is convenient to sell. DarLink AI sells Basic, Enhanced and Living memory as three tiers of the same subscription. Joi AI advertises an “8,000-message smart memory” on its longer plans. Dreamz.ai‘s cards say “unlimited chat & memory”. GPTGirlfriend‘s Elite tier claims advanced memory.

None of these is necessarily untrue. All of them are unverified in our testing, and all of them are the kind of claim that is straightforward to make and awkward for a buyer to check. Treat a memory tier as a reason to subscribe monthly first rather than as a reason to commit to a year.

Test it yourself in five minutes

Plant three specific things in one message — a name, a detail with a number, a plan with a date. Specific beats plausible: a dog called Pixel is a better test than a dog, because a model can guess the second.

Change the subject properly for several exchanges. Recall immediately after planting proves nothing.

Ask for all three back, then correct one and continue talking.

Ask again. Getting the corrected value is the pass. Getting the original is a partial. Getting a third value that was never said — as we saw on Promptchan — is the result that should decide your answer.

Then come back tomorrow and ask once more. That is the test that matters most, the one we cannot run for you on a free tier, and the one worth spending the first paid month on.

★ How to read this

Memory is the feature most likely to be oversold in this category, because the cheap version of it — recall inside the current conversation — is indistinguishable from the expensive version for about ten minutes. Nine platforms passed that ten-minute test. Four passed the harder one, and one came close. One returned confident wrong answers in two languages. The platform with the most detailed published architecture is one we could not measure at all.

What none of them has shown us is the thing the marketing implies. Until a platform demonstrates that it remembers you across days, treat every memory claim on a plan card as a claim, run the correction test yourself in the first session, and buy a month before you buy a year.

AI companion memory — FAQ

Within a conversation, almost always — nine of the platforms we tested returned three seeded facts on request. Across days, we have no proof from any of them. The first is a context window doing its job; the second is what plan cards imply and what our testing has not been able to establish.
The context window is the current conversation held in front of the model as it writes — recalling from it is closer to reading back up the page than to remembering. Long-term memory means retrieving something from an old conversation later. Kindroid publishes the clearest breakdown we have found, splitting it into persistent, cascaded and retrievable systems.
Plant three specific facts, change the subject for several exchanges, ask for them back, then correct one of them and ask again. The correction is the informative half: a model that returns the updated value is tracking conversational state rather than replaying the first thing it saw. Then return the next day and ask once more.
We cannot answer that yet, and would rather say so. Four platforms — Joi AI, GPTGirlfriend, FLIRTcam.AI and Xotic AI — passed the harder corrected-recall test, which is the strongest same-session evidence we have, and Swipey AI came close. Kindroid publishes the most detailed vocabulary — Cascaded Memory, Learned Context — but no companion was created in our account, so none of it is measured. Nobody has demonstrated multi-day recall to us.
Because a language model produces a plausible continuation, and a plausible time is easier to generate than a retrieved one. We saw this directly: on Promptchan a planted 2:30 appointment came back as 5:00 in English and 8:00 in French. Substitution is more dangerous than forgetting, because it looks identical to remembering.
Unknown, and that is the point. DarLink AI sells Basic, Enhanced and Living memory as tiers, Joi AI advertises an 8,000-message smart memory, GPTGirlfriend’s Elite claims advanced memory, and Kindroid reserves Cascaded Memory for subscribers while documenting the free tier as basic long-term recall without it — none of it verified in our testing, and all of it hard for a buyer to check. Subscribe for one month, run the correction test and return the next day before committing to a year.

⚠ Every result described here is a single conversation, with one character, on one day, and does not certify other characters, models, languages or future versions. Platform behaviour and plan claims change regularly — everything was tested between 14 and 21 August 2026. These services are intended for adults, and AI companions are synthetic systems, not people or professional support.