inApp

Methodology

How reviews become findings

The governing principle is traceability. A pattern can be followed down to apps, quotes, and source ratings; weak or ambiguous text is not assigned an artificially precise theme.

Data version dated August 13, 2026

1. Corpus

The current snapshot contains 1,451,072 public reviews across 4,623 mobile apps in 71 niches. Complete text, star rating, and app attribution are available for every review directly in the catalogue.

Per-review labelling is complete: 1,451,072 of 1,451,072 texts carry exactly one label.

2. Three labelling layers

Text layer

One label per review

We first look for an explicit product story or cross-product mechanism: charges, ads, crashes, login, sync, and other narrow signals. If the text is insufficient, only sentiment is retained, without inventing a reason.

Market layer

Niche patterns

A story must appear in at least 8 signals and across at least 3 different apps. Apps, direction, and verifiable quotes are retained for every pattern.

Product layer

App themes

Recurring stories are formed separately within each app. The same issue can therefore be named differently for two competitors while preserving product context.

3. How to read the metrics

Signal
A review supporting a specific pattern. One substantive review may touch several stories, so signal totals do not equal unique review counts and are not market share.
Theme direction
“Mostly praised,” “mostly criticised,” or “opinions differ” describes the theme in aggregate, not every individual text. A five-star review may mention a drawback, while a one-star review may mention a useful feature.
Specific theme
A substantive story that can be named without speculation. 550,502 reviews, or 37.9% of the complete corpus, have one; the remainder stays in explicit “without a specific reason” sentiment buckets.
Coverage
Texts and per-review labels are available for all 1,451,072 reviews. Market patterns are ready for all 71 niches. The additional deep product layer is ready for 937 of 4,623 apps, or 19.9% of the complete corpus. These figures are deliberately reported separately.

4. Limitations

  • Reviewers are self-selected and often had an especially good or bad experience. This is the audience voice, not a representative survey.
  • The corpus is a snapshot. Product version, country, language, and store policy can affect what is visible.
  • We do not verify purchase status or reviewer identity. Suspicious activity should not be treated as proven real-world experience.
  • Mention frequency shows signal strength within the corpus, but does not measure prevalence among all users of an app.