
App Store Listings
Part of App store listing experiments
Interpreting store test results with limited traffic
Read App Store and Google Play experiment results with small audiences, separate lift from uncertainty and make an honest inconclusive decision.
With limited traffic, a store test may remain unable to distinguish its variants. Start with the store’s reported status, available counts and uncertainty. A large-looking percentage based on sparse observations is not, by itself, a winning result.
Check who entered the comparison
Record the store, listing, variants, dates, traffic setting, selected localisations or languages and target metric. Apple Product Page Optimization uses a treatment traffic proportion and selected localisations. Check whether the configured audience or available report covers the people you want to understand; an English-language test can include people elsewhere.
Apple test results appear in Analytics after at least five first-time downloads are attributed to the test. Apple may later label a treatment “Likely to be Inconclusive” when current results suggest it is unlikely to reach statistical significance. This status does not mean that the variants perform equally.
Read magnitude with uncertainty
Apple reports estimated conversion, relative lift and confidence. When a treatment reaches 90% confidence, it may be labelled “Performing Better” or “Performing Worse” compared with the baseline. Keep each store’s terms with its figure. A range that still allows outcomes materially better and worse than the current listing leaves the decision uncertain.
Show available counts and the selected metric’s definition. Do not treat percentages from different stores as though they came from one experiment.
Key Metrics in Store Test Interpretation
- Confidence Level
- 90% or higher
- Uncertainty Status
- Likely to be Inconclusive (if not reaching significance)
Decide whether to continue or close
Apple’s setup estimate is a guide; its tests run for 90 days or until manually stopped. If the planned window can still add useful evidence, let it continue. Avoid repeatedly checking an early favourable number and stopping only when it looks good.
A later test with a clearer, still truthful creative difference may be more informative when the current comparison is unlikely to resolve. If a decision is needed sooner, the team may choose on product clarity or accessibility grounds. Record that as a judgement under uncertainty, not an experiment win.
Before closing the review, note any release, campaign, promotion, eligibility or listing change during the test. Apple warns that a version containing assets or metadata under test may affect its results. These events should inform interpretation without automatically invalidating the comparison.
State what the store reported and what the team decided: favour a treatment, keep the original, accept a reported draw under the store’s rules or leave the result unresolved. Preserve the counts and uncertainty so an inconclusive test is not later retold as proof of no effect.



