Skip to content

How reviews become findings

The governing principle is traceability. A pattern can be followed down to apps, quotes and source ratings; weak or ambiguous text never gets an artificially precise theme.

Data version dated August 13, 2026

Corpus

The current snapshot holds 1,451,072 reviews from public sources about 4,623 apps in 71 niches. The complete text, star rating and app of every review are kept.

Per-review labelling is complete: every review carries one or more labels. Topic labels in total: 1,757,106.

Source access

Dating is fully open as a verifiable sample: category → app → reviews by topic. The rest of the archive is part of Plus. The lock applies to both the pages and the API; category names and corpus size stay visible before purchase.

Three labelling layers

  • Every topic in each review

    We identify every explicit story: charges, ads, crashes, login, delivery, quality and other narrow signals. One text may get several labels. If it has no specifics, only an overall assessment is kept, without an invented reason.

  • Niche patterns

    A story must appear in at least 8 signals and in at least 3 different apps. The apps, the direction and verifiable quotes are kept for every pattern.

  • App themes

    Recurring stories are formed separately within each app. The same issue can therefore be named differently for two competitors and keep its product context.

How to read the numbers

Signal
A review that supports a specific pattern. One substantive review can touch several stories, so signal totals do not equal unique review counts and are not market share.
Theme direction
“Mostly praised”, “mostly criticised” or “opinions differ” describes the theme as a whole, not every text. A five-star review may mention a drawback, and a one-star review a useful feature.
Specific theme
A substantive story that can be named without speculation. 707,333 unique reviews — 48.7% of the complete corpus — have one or more; the rest stays in explicit “without a specific reason” buckets.
Coverage
Texts and per-review labels exist for all 1,451,072 reviews. Market patterns are ready for each of the 71 niches. The additional deep product layer is ready for 937 of 4,623 apps, or 19.9% of the complete corpus. These figures are deliberately reported separately.

Limitations

  • Reviewers are self-selected: more often people with an especially good or bad experience. This is the audience’s voice, not a representative survey.
  • The corpus is a snapshot in time. Product version, country, language and store policy can affect what is visible.
  • We do not verify purchases or who wrote a review. Suspicious activity should not be treated as proven real-world experience.
  • Mention frequency shows signal strength within the corpus, but does not measure how common a problem is among all users of an app.

The best way to verify a finding

Open a category and an app, pick a topic and read every source text under it. Any label can be checked against the stars and the exact words — the labelling never hides the corpus behind a summary.

Go to the reviews