AI Companion App Review Template: How We Score Apps
Our reviews score every AI companion app on seven weighted criteria. Here's exactly what we test for safety, privacy, pricing, realism, memory, support and cancellation.

Table of contents
Every review we publish answers one question: would we trust this app with our time, money and personal data? To keep that answer consistent across very different products, we score each AI companion app against the same seven criteria, each with a fixed weight. This page explains exactly what we test, why each factor matters, and how the weights add up — so you can see our reasoning and, if your priorities differ from ours, re-weight the scores yourself.
How scoring works
We assess every app on a 0–10 scale per criterion, then combine those into a weighted overall score out of 100. Weights reflect what we think protects users most: safety, privacy and pricing transparency together account for more than half the total, because those are the areas where a bad app does real harm. Feature quality — realism, memory, support — matters, but it earns a smaller share, since a polished experience built on shady data practices is not a good product. We re-test after major updates, because companion apps change fast and a score is only a snapshot of a moment.
The criteria and weights
| Criterion | Weight | What we check |
|---|---|---|
| Safety | 20% | Age gating and verification, clear non-explicit defaults, crisis/self-harm handling, refusal to encourage harmful behaviour, guardrails against emotional manipulation and dependency-driven design |
| Privacy | 20% | Existence and clarity of a privacy policy, what data is collected, whether chats train models, third-party sharing, encryption claims, and a real data-deletion path |
| Pricing transparency | 15% | Whether prices and renewal terms are shown before sign-up, free-tier honesty, no surprise paywalls, and no manufactured urgency |
| Realism | 15% | Conversation quality, consistency of persona, tone control, and how natural voice or image features feel in normal use |
| Memory | 10% | Whether the app remembers facts across sessions, lets you view/edit/delete what it stored, and avoids confidently inventing past events |
| Support | 10% | Reachable human or responsive help, working contact channels, documentation, and how quickly issues are acknowledged |
| Cancellation | 10% | Whether cancelling is as easy as subscribing, in the same medium, with no dark-pattern hoops |
Why safety and privacy lead
These two carry the most weight because they affect users who can least afford a mistake. On safety, we look for sensible defaults rather than marketing claims: does the app verify age, keep mainstream conversations non-explicit unless a verified adult opts in, and respond responsibly when a user signals distress? On privacy, we read the policy in full and note what's missing. Vague language, no stated deletion process, or chats silently used for training all cost points. Regulators have made clear that deletion promises must be true, so we treat an unverifiable claim as a red flag, not a feature.
Pricing, cancellation and dark patterns
We score pricing transparency and cancellation separately because an app can be honest about price yet trap you on the way out. Under the FTC's click-to-cancel approach, if you can subscribe in the app you should be able to cancel in the app, with no forced phone calls, confusing screens, or hidden menus. We test the real flow where we can, and we never quote a number we haven't confirmed — if pricing isn't clearly published, our review says it varies and explains where to verify it. Free trials that auto-convert without a clear reminder lose points here.
Realism, memory and support
The experience criteria reward craft. Realism covers whether conversations stay coherent, whether the persona holds its tone, and whether voice or image features add something rather than feeling like a gimmick. Memory is scored on user control as much as capability: an app that remembers your details but won't let you see or erase them scores worse than one with simpler memory and full transparency. Support rewards apps you can actually reach when a payment or data issue goes wrong — a working contact channel and timely replies beat a glossy FAQ that answers nothing.
How to read our scores
A high overall number means an app cleared our safety and privacy bars first, then earned the rest on merit. Because the weights are published above, you can recompute any review for your own priorities — if cancellation matters more to you than realism, raise its weight and re-add. We also flag any single criterion that scores very low even when the total looks healthy, because one serious failure (no privacy policy, no age gate, a cancellation trap) can outweigh an otherwise pleasant experience. The goal isn't a ranking to obey; it's a transparent framework you can argue with.
Bottom line
Our template exists so reviews stay comparable, honest and re-checkable. Safety, privacy and pricing lead the weighting because they're where users get hurt; realism, memory, support and cancellation fill in the experience. Whenever the data isn't public, we say so plainly rather than guess. Use the weights as a starting point, adjust them to your own life, and treat any app that fails the safety or privacy bar as a no regardless of how good it feels to talk to.
Sources and further reading
Sources


