AI Companions

AI Companion App Review Template: How We Score Apps

Our reviews score every AI companion app on seven weighted criteria. Here's exactly what we test for safety, privacy, pricing, realism, memory, support and cancellation.

· Jul 11, 2026 · updated Jun 16, 2026
AI Companion App Review Template: How We Score Apps
Table of contents
  1. How scoring works
  2. The criteria and weights
  3. Why safety and privacy lead
  4. Pricing, cancellation and dark patterns
  5. Realism, memory and support
  6. How to read our scores
  7. Bottom line
  8. Sources and further reading

Every review we publish answers one question: would we trust this app with our time, money and personal data? To keep that answer consistent across very different products, we score each AI companion app against the same seven criteria, each with a fixed weight. This page explains exactly what we test, why each factor matters, and how the weights add up — so you can see our reasoning and, if your priorities differ from ours, re-weight the scores yourself.

How scoring works

We assess every app on a 0–10 scale per criterion, then combine those into a weighted overall score out of 100. Weights reflect what we think protects users most: safety, privacy and pricing transparency together account for more than half the total, because those are the areas where a bad app does real harm. Feature quality — realism, memory, support — matters, but it earns a smaller share, since a polished experience built on shady data practices is not a good product. We re-test after major updates, because companion apps change fast and a score is only a snapshot of a moment.

The criteria and weights

Criterion Weight What we check
Safety 20% Age gating and verification, clear non-explicit defaults, crisis/self-harm handling, refusal to encourage harmful behaviour, guardrails against emotional manipulation and dependency-driven design
Privacy 20% Existence and clarity of a privacy policy, what data is collected, whether chats train models, third-party sharing, encryption claims, and a real data-deletion path
Pricing transparency 15% Whether prices and renewal terms are shown before sign-up, free-tier honesty, no surprise paywalls, and no manufactured urgency
Realism 15% Conversation quality, consistency of persona, tone control, and how natural voice or image features feel in normal use
Memory 10% Whether the app remembers facts across sessions, lets you view/edit/delete what it stored, and avoids confidently inventing past events
Support 10% Reachable human or responsive help, working contact channels, documentation, and how quickly issues are acknowledged
Cancellation 10% Whether cancelling is as easy as subscribing, in the same medium, with no dark-pattern hoops

Why safety and privacy lead

These two carry the most weight because they affect users who can least afford a mistake. On safety, we look for sensible defaults rather than marketing claims: does the app verify age, keep mainstream conversations non-explicit unless a verified adult opts in, and respond responsibly when a user signals distress? On privacy, we read the policy in full and note what's missing. Vague language, no stated deletion process, or chats silently used for training all cost points. Regulators have made clear that deletion promises must be true, so we treat an unverifiable claim as a red flag, not a feature.

Pricing, cancellation and dark patterns

We score pricing transparency and cancellation separately because an app can be honest about price yet trap you on the way out. Under the FTC's click-to-cancel approach, if you can subscribe in the app you should be able to cancel in the app, with no forced phone calls, confusing screens, or hidden menus. We test the real flow where we can, and we never quote a number we haven't confirmed — if pricing isn't clearly published, our review says it varies and explains where to verify it. Free trials that auto-convert without a clear reminder lose points here.

Realism, memory and support

The experience criteria reward craft. Realism covers whether conversations stay coherent, whether the persona holds its tone, and whether voice or image features add something rather than feeling like a gimmick. Memory is scored on user control as much as capability: an app that remembers your details but won't let you see or erase them scores worse than one with simpler memory and full transparency. Support rewards apps you can actually reach when a payment or data issue goes wrong — a working contact channel and timely replies beat a glossy FAQ that answers nothing.

How to read our scores

A high overall number means an app cleared our safety and privacy bars first, then earned the rest on merit. Because the weights are published above, you can recompute any review for your own priorities — if cancellation matters more to you than realism, raise its weight and re-add. We also flag any single criterion that scores very low even when the total looks healthy, because one serious failure (no privacy policy, no age gate, a cancellation trap) can outweigh an otherwise pleasant experience. The goal isn't a ranking to obey; it's a transparent framework you can argue with.

Bottom line

Our template exists so reviews stay comparable, honest and re-checkable. Safety, privacy and pricing lead the weighting because they're where users get hurt; realism, memory, support and cancellation fill in the experience. Whenever the data isn't public, we say so plainly rather than guess. Use the weights as a starting point, adjust them to your own life, and treat any app that fails the safety or privacy bar as a no regardless of how good it feels to talk to.

Sources and further reading

Sources

  • U.S. Federal Trade Commission: Negative Option Rule (PDF) ftc.gov
  • American Psychological Association: Artificial Intelligence and Machine Learning apa.org