This methodology defines the records, fields and denominators used in Testers Community's 2026 Google Play research. The unit is a unique app ID, discovery is kept separate from launch, and collection failures are kept separate from collected records. This is methodology version 1.1 for the database maintained since 2024.
The counting rule
One unique Google Play app ID is one listing record. The current corpus contains 3,037,501 unique app IDs observed between 27 July and 25 August 2026. It is an observed corpus accumulated across that window, not a simultaneous census of the store and not an official Google Play total.
Unit of analysis
Google Play identifies an Android app with an app ID, also called a package ID. We use the stored
appId field as the primary identifier. Two app IDs from one publisher are two listing
records. The same app ID is not counted again because its listing is available in another
country.
Validation of the 2026 corpus found 3,037,501 records, 3,037,501 distinct app IDs, no missing app IDs and no duplicate app IDs. This is the operational rule behind the count, replacing the vague phrase “counted one by one.”
Collection context
The stored request URLs use English-language listing pages. They use the US region for 3,037,485 records; 16 records have no stored region parameter. The dataset therefore describes the listing fields returned in that language and regional context. Availability, price and some presentation fields can differ elsewhere.
The saved research artifacts establish the language and region used for listing fields. They do not establish device- or account-specific presentation, and the discovery records do not provide a complete coverage denominator for every public Google Play listing.
Collection flow
The store-wide corpus and weekly discovery batches are different datasets. A corpus record is a listing successfully observed during the study window. A newly discovered ID is an identifier present in a dated discovery source but absent from the historical seen-ID set. Discovery is not a launch date.
Discovery and collection outcomes
| Term | Definition | What it does not establish |
|---|---|---|
| Observed corpus | Unique app-ID records successfully collected during the stated study window. | A simultaneous total for Google Play. |
| Previously unseen ID | An ID in the dated discovery source that is absent from the historical seen-ID set. | Submission, publication or launch date. |
| Successful collection | The collector saved a parsed listing record for the discovered ID. | Availability in every language, region, device or account context. |
| Collection failed | The collector did not save a parsed listing record for the discovered ID. | Removal, delisting, suspension, policy action or any other specific cause. |
The current weekly aggregate supplies successful and failed counters but no verified breakdown of failure causes, redirects or retry outcomes. We therefore report the counter as “collection failed.” We do not translate it into “gone” or “removed.”
Field definitions and denominators
Most descriptive percentages use all 3,037,501 corpus records. When a field can be missing, the denominator is stated explicitly. Derived measures are labeled as derived rather than presented as fields shown directly on a store page.
| Measure | Stored source and type | Null handling | Default denominator | Derived rule |
|---|---|---|---|---|
| Listing record | appId, string | Missing IDs fail identity validation | 3,037,501 unique IDs | One row per distinct app ID |
| Lifetime installs | realInstalls, number; installs, display-band string | Excluded from a threshold if the numeric field is not numeric | Records with a numeric value | Threshold studies use realInstalls; the public band is retained separately |
| Shows a rating | score, number or null | Null means no score was displayed in the collected response | All corpus records | Presence test, not sentiment or quality |
| Release date | released, date string or null | Missing or unparseable values are excluded from release-year shares | 2,311,689 records with a parseable date | Year parsed from the stored date |
| Last update | updated, timestamp | Missing values are excluded from recency shares | 3,036,563 records with an update date | Age equals observation timestamp minus update timestamp |
| Game | genreId, string | Missing values are not classified as games | All corpus records | Game when the genre ID begins with GAME |
| Category | genre and genreId, strings | Missing values are reported outside category totals | Records in the stated category set | Some public tables use a declared subset and publish its coverage |
| Promo video | video, URL string or null | Null or empty means no populated video field | All corpus records | Presence only; playability, views and effect are not tested |
| Screenshots | screenshots, array | Empty array counts as zero | All corpus records | Array length |
| Description length | description, string or null | Missing values are excluded from the median | 3,037,486 records with a description | Unicode string length in characters; median 774 |
| Permissions | permissions, array | Empty array counts as zero | All corpus records | Array length |
| Contains ads and IAP | containsAds and offersIAP, booleans | Reported from the stored listing declaration | All corpus records | No revenue or ad-load inference |
| Price | free, boolean; price, number; currency, string | Reported as observed in the request context | All corpus records or paid subset | Paid when free is false |
| Publisher | developerId, string | Missing IDs are excluded from unique-publisher counts | Records with a publisher ID | Distinct publisher IDs |
Observed window, not simultaneous total
The 3,037,501 records were observed from 27 July to 25 August 2026 according to their saved observation timestamps. Because collection spans a month, the number is not proof that all 3,037,501 pages were simultaneously available on one day. We call it the study corpus, not the total number of apps on Google Play.
Google does not publish a complete official listing dataset. The research process used a sitemap source for ID discovery, but a discovery source is not evidence of complete coverage. Without a seeding manifest and an official denominator, recall cannot be calculated.
Validation and reproducibility
The corpus is stored across 134 compressed JSONL parts accompanied by a 134-entry SHA-256 manifest. Validation for this revision parsed all 3,037,501 records, found no malformed JSON, no missing app IDs and no duplicate app IDs. These checks establish file readability and ID integrity. They do not validate every field against the live store or replace manual content review.
Weekly tracker generation now stops if it encounters malformed JSON or if the number of parsed records does not match the successful-collection counter. This prevents percentages from being computed over a silently smaller base.
Public supporting files contain aggregates only: promo-video summary CSV and category listing-pressure CSV. Raw records are not published because they include personal developer fields.
Observed fields and constructed measures
Most research outputs are direct counts, shares, medians or thresholds calculated from stored listing fields. That does not mean every published metric is observed directly. The listing-pressure index, for example, is a constructed average of two observed percentages and is labeled experimental.
No result in this research should be read as a causal model unless a separate experimental design is documented. Cross-sectional associations, such as promo-video presence rising with install band, do not establish what caused the installs.
Methodology version history
| Version | Date | Change |
|---|---|---|
| 1.1 | 15 Sep 2026 | Defined app-ID deduplication, request locale and region, corpus versus discovery batches, collection failures, field types, null rules, install fields, validation checks and constructed metrics. |
| 1.0 | 6 Sep 2026 | Initial public definitions and denominator notes. |
Limits
- Coverage is not measurable. There is no official complete denominator and the seeding manifest is unavailable.
- Regional context matters. The snapshot reflects English-language, primarily US-region requests.
- The corpus is cross-sectional. It does not by itself measure growth, survival or causal effects.
- Missing fields change bases. Release-year and update-recency calculations use their stated non-missing subsets.
- Collection failure is not removal. A missing saved record does not identify the cause.
- Aggregates do not validate interpretation. Every article must still match its claim to the field and denominator used.
Studies using this method
Every study below reports aggregate findings from the corpus defined on this page. Each one states its own denominator in its own text, because several work from a reduced base where a field is missing. This list covers the studies that are published; more are in preparation.
- Google Play App Discovery: One Weekly Batch in 2026
- Inside Google Play in 2026: What 3M+ Listings Reveal
- We Analyzed 3 Million Google Play Listings in 2026
- Why Most Google Play Apps Remain Below 1,000 Installs
Frequently asked questions
What is the unit of analysis in this Google Play research?
One unique Google Play app ID, also called a package ID, is one listing record. The 2026 corpus contains 3,037,501 records, with no missing or duplicate app IDs in the validated snapshot.
Does one app available in multiple countries count more than once?
No. The 2026 corpus contains one record per app ID and uses English-language, US-region listing requests. Country availability is not counted as a separate app.
Is 3,037,501 a simultaneous total for Google Play?
No. It is the number of unique app IDs observed in the study corpus between 27 July and 25 August 2026, not a simultaneous census or an official Google Play total.
What is a newly discovered app ID?
It is an app ID found in a dated discovery source that was not in the historical seen-ID set. Discovery does not establish the app's launch date.
What does collection failed mean?
It means the collector did not save a parsed listing record for that discovered app ID. The counter does not establish why collection failed and is not evidence that the app was removed.
What denominators do the research percentages use?
Most use all 3,037,501 corpus records. Release-year figures use 2,311,689 records with a parseable release date, and update-recency figures use 3,036,563 records with an update date.
Citing this methodology
Testers Community, How We Count Google Play Apps: Definitions and Limits, methodology version 1.1, 16 September 2026. Apply this page to research based on the 3,037,501-record corpus observed from 27 July to 25 August 2026, and cite the individual study page for its result.
Copy a reference
- Plain
- Testers Community (2026). How We Count Google Play Apps: Definitions and Limits, version 1.1. https://www.testerscommunity.com/research/how-we-count-google-play-listings
- APA 7
- Testers Community. (2026, September 16). How we count Google Play apps: Definitions and limits (Version 1.1). https://www.testerscommunity.com/research/how-we-count-google-play-listings
- MLA 9
- “How We Count Google Play Apps: Definitions and Limits.” Testers Community, version 1.1, 16 Sept. 2026, www.testerscommunity.com/research/how-we-count-google-play-listings.
Sources
- Metadata, Google Play Console Help
- Choose a category and tags for your app or game, Google Play Console Help
- Publish your app, Google Play Console Help