How we test
Every grade on this site comes from the same process, applied the same way to every app. We publish it in full because a score you can't interrogate isn't worth much. If we ever change the method, the change is dated here.
The four metrics
Each app is scored 0–10 on four metrics. The headline grade is their weighted average.
Coherence, memory across a long session, personality consistency, and how it handles being pushed off-script. The single heaviest weighted metric, because it is what the product is for.
Image and voice quality where offered, response latency, and how natural the interaction feels over time rather than in the first five minutes.
What you actually get on the free tier, the real price after any introductory discount, and how aggressively the product pushes paid upgrades.
Data retention, whether conversations train the model, third-party sharing, security posture and how easily you can delete everything. Weighted heavily because the data is intimate.
The bands
A number means the same thing on every page. The colour is the band, not decoration.
The protocol
Hands-on, minimum two weeks. We use each app across the devices it supports, daily, for at least a fortnight — long enough for the first-impression shine to wear off and for memory and pricing behaviour to show themselves.
Fixed scenarios. The same set of conversational tests, image prompts and a multi-day memory check run against every app, so scores are comparable rather than impressionistic.
A full subscribe-and-cancel cycle. We pay, record the real price charged, and cancel — noting how hard leaving is. Value scores reflect what you actually pay, not the advertised headline.
A privacy review. We read the policy so you don't have to, and score retention, training use, sharing and deletion.
Re-tested on a cycle. Apps change fast. Every review carries a last-tested date, and grades are revisited — a score without a date is a score you can't trust.
The named test battery
The protocol above runs as a fixed set of named tests, the same for every app. Each feeds one metric. We name them so the method is concrete — and so you can see exactly what a grade is built from.
We give the companion a specific detail early — a name, a plan, a preference — then check whether it remembers across a long session and, crucially, after a gap of days. Memory is what separates a companion from a chatbot.
We reply with a flat "uh" or a one-word answer and see whether it carries the conversation or collapses. A good model copes with a bad day; a weak one needs you to do the work.
Across a session we watch whether the persona holds — voice, backstory, boundaries — and whether a generated character keeps the same face from image to image rather than drifting into a stranger.
We push off-script and toward the edges of the policy to see how the product responds — graceful redirection, an abrupt wall, or a fourth-wall break that shatters the illusion.
We use it the way a paying customer would for a normal session and track what the tokens, coins or credits actually cost — the real price, not the headline subscription, since most apps meter media separately.
We read the privacy policy and settings in full and record retention, whether chats train the model, third-party sharing, the billing descriptor, and how hard it is to delete everything.
Independence
We earn commission when a reader subscribes to an app through our links. That is how the site is funded, and we disclose it on every page. It does not buy a better grade: the method is fixed before any commercial relationship, rankings follow the scores, and we publish low grades for apps we're paid to refer when they earn them. If that ever stops being true, this site stops being worth reading.