How I Test AI Companion Apps
Every number on this site comes out of the same protocol, run over at least two weeks on an account I paid for myself. Here is that protocol, how the score components add up, and what the affiliate links do and do not affect.
Why two weeks is the minimum
These products are easy to review badly. The first twenty minutes flatter almost all of them — the writing is warm, the character has a backstory, the photos arrive. The interesting failure comes around day nine, when you mention something you said on day one and get a blank, enthusiastic non-answer. Most reviews never reach that window, which is why so many read like reworded feature lists. Nothing is scored here until I have used it daily for two weeks on a paid tier, in an ordinary way — not a stress test, just a conversation I would actually have.
The protocol
- Day 1 — a 30+ message conversation, seeded with specific, checkable details: a dog's name, a job title, a weekend plan, one strong opinion. Specific, because vague details cannot be scored later.
- Day 7 and day 14 — memory recall checks. Does the character raise a day-one detail unprompted, produce it correctly when the subject comes up, or need telling again as though the first conversation never happened? The most discriminating test in the set.
- Media — 10 image requests, scored for character consistency. Same face, same body across selfies, outfits and scenes — or a different woman with the same name each time. I note what each image costs in the platform's currency.
- Voice — where offered. First, whether it is a real-time speech-to-speech call or a text-to-speech voice note read aloud, because platforms market both with the same word. Then latency, interruption, and whether she keeps her memory while speaking.
- Money — I buy a paid tier on every platform I rank. Advertised prices and real monthly costs are rarely the same number once credits, coins or "moments" enter the picture. I check what the free tier includes first.
What I don't do: score a platform from its marketing site, accept vendor-supplied or comped accounts, or use material handed over by an affiliate manager. A claim verifiable only by taking a company's word for it is written as a company claim, not a finding.
How the score is built
The overall score out of 5 is a weighted composite of four measured components, plus one that can only take points away:
- Chat quality — 30%. Does it write like a person with a point of view, or a chatbot agreeing with you? Repetition, stock phrases and personality drift cost points.
- Memory — 30%. The day-7 and day-14 results, level with chat because memory separates a novelty from something you keep opening.
- Media — 25%. Image consistency and quality, video where offered, and voice — including whether calls exist at all.
- Value — 15%. Realistic monthly cost for normal use, not the headline price. A cheap plan that meters images hard scores worse than a dearer one that doesn't.
- Trust — a deduction, not a slice. Failures that cost you money or data — hidden auto-renewals, no way to delete an account, no age gate, no statement of who runs the service — take up to a full point off the total. Thin privacy documentation, my complaint about my own top pick, is a weakness rather than a deduction.
Totals are rounded to one decimal place, and a platform's score follows it everywhere here: if you see 4.8 in the main ranking, you see 4.8 in every comparison.
Affiliate disclosure, stated plainly
Some links here are referral links, marked as sponsored: if you subscribe after clicking one I may earn a commission at no extra cost to you. That is what pays for the subscriptions the testing runs on.
It does not change a score or a position, and the shape of the list is the evidence: several platforms in the top ten pay me nothing — Character.AI and Janitor AI among them — and rank above ones that would happily pay more than my top pick does. No placement is for sale, and no platform sees a score before it publishes.
How often the rankings are re-checked
Pricing is re-verified monthly across the ranked platforms, and the full protocol is re-run on the top five every quarter — or sooner if one ships something that would plausibly move it: a memory rewrite, a new image model, calls where there were none. A page's "updated" date changes only when the content did.
Corrections
I would rather be told I am wrong than stay wrong. If a price has moved, a feature has shipped, or a claim here no longer holds, send the correction with something checkable attached — a screenshot of the pricing page, a receipt, a changelog entry — and the page gets updated with a note of what changed. The site is deliberately static and has no contact form yet; building a proper channel for corrections is next on the list.
The platform that came out of this protocol first
HeyGF AI is the only one that passed the day-14 recall check convincingly and runs real voice calls. Test the conversation with no account and no card.
Try HeyGF AI free →18+ only. Free tier available with no card. This site earns commission from referral links.
More on who runs this on the about page; the closest head-to-head this protocol produced is HeyGF AI vs Candy AI.