Close-up of App Store icon on iPhone screen with notification badge, highlighting app updates.
Photo by Brett Jordan on Pexels

App Store Listings

App store listing experiments

Plan an App Store or Google Play listing experiment, define its audience and outcome, read uncertainty and record a defensible decision.

An app store listing experiment compares a current listing with a defined alternative shown to eligible visitors. Before launch, choose a decision the result can support, keep changed assets identifiable and define the store outcome. Treat Apple and Google Play as separate implementations, and compare only where populations and outcome definitions support it. A stronger store result does not establish better use after installation.

Define the question

Name the listing, eligible audience and change. State the decision the result will inform before choosing treatments. For example, an Australian booking app might test whether a screenshot showing a confirmed booking outperforms its current opening screenshot. Check that both versions describe features available to the people who will see them.

Apple Product Page Optimization compares the default iOS or iPadOS product page with up to three treatments using icons, screenshots or app previews. It does not test Apple custom product pages.

For Google Play, open the store listing experiment workflow in Play Console and follow the options shown for the selected experiment. Record the listing, control and alternatives as labelled in Console, along with each alternative’s change and any audience, traffic or result settings presented. Do not assume Apple’s options or treatment limit apply to Google Play.

Check Apple test prerequisites

Apple Product Page Optimization requires the app to be live on the App Store and in Pre-Order Ready for Distribution or Ready for Distribution status. Treatments appear to users on iOS 15 or iPadOS 15 and later, so the test does not represent visitors on earlier operating-system versions.

Access is limited to users with the Account Holder, Admin, App Manager or Marketing role in App Store Connect. Check the app’s status and your access before treating a platform limitation as a test-design problem.

Before planning a Google Play test, check the selected experiment workflow in Play Console for its eligibility requirements. Record any restriction it presents rather than assuming Apple’s app-status, role or operating-system requirements apply.

Keep the change identifiable

Save the original and each variant before launch. Apple treatments start as copies of the original page, so an item left unedited continues to display the original asset. For a clearer comparison, change one asset type at a time where practical, keeping other listing assets steady where the setup permits.

A screenshot message may require several images to change together. That compares the whole sequence, so a better result cannot identify which caption or frame made the difference.

For Apple, treatment metadata must be approved before testing. You can submit it without a new app version, but any test icon must already be in the current app binary.

Choose the audience and outcome

Apple lets you select localisations and the proportion of eligible traffic shown treatments. Record both settings with the test so the audience and allocation are part of the result.

A language selection does not, by itself, isolate Australian visitors or people with a particular task in mind. Do not treat localisation as a substitute for an audience definition.

For Google Play, note any audience or traffic settings offered for the selected experiment and record the settings you choose. If an option is not available, do not assume it can be configured.

Apple reports an estimated product page conversion rate—the percentage of viewers who download or pre-order—relative lift against the baseline and confidence. Keep these definitions together so the result can support the decision set before launch.

For Google Play, record the outcome measure and definition shown for the selected experiment, along with any reported status or uncertainty. Use the wording and definition shown in Play Console rather than assuming they match Apple’s measures.

Compare each alternative with the control as labelled in Play Console. Keep the control and alternatives clearly identified when interpreting the result.

Avoid calling a result decisive if the reported evidence does not support that conclusion. A preferred alternative is not automatically a supported decision.

Do not assume audience settings or outcome definitions transfer between stores. Record each platform’s configured audience, treatment allocation and metric definition, and compare results across stores only where the populations and outcome definitions align.

For the booking-app example, record Apple’s audience and traffic settings with its estimated conversion rate and relative lift. Record Google Play’s selected audience and traffic settings with the outcome name and definition shown in Console. Compare the results only if their populations and measured events align; otherwise keep separate findings.

Keep the selected metric’s name and definition in the result; a store outcome alone does not show whether someone completed a useful task in the app. Do not relabel a result as an equivalent measure on the other store unless the definitions match.

If later use matters to the decision, assess it separately with a defined cohort. Do not combine store and app analytics counts into a single rate unless their populations can be connected reliably.

Allocate traffic deliberately

Apple’s traffic proportion is the share of users randomly shown a treatment instead of the original page. For example, with three treatments and a 30% traffic proportion, each treatment receives 10% of total traffic.

Apple selects all supported localisations by default, but you can exclude localisations. Users whose displayed localisation is excluded are not included in the test, so check that the selected markets match the question before launch.

For Google Play, record the traffic allocation offered and selected in the workflow. Do not assume Apple’s default, allocation method or traffic split applies.

Allow for uncertainty

For Apple, review the estimated duration and impressions needed to reach the desired improvement; the estimate uses existing performance data such as daily impressions and new downloads. It is a guide and does not affect the test.

An Apple test runs for up to 90 days or until you manually stop it within that time. The desired improvement may not be reached within 90 days.

Apple results appear in Analytics after at least five first-time downloads are attributed to the test. More variants divide experimental traffic and can delay a useful result.

Read the status and uncertainty alongside the apparent lift. “Likely to be Inconclusive” on Apple is not evidence that the variants perform equally.

Note any release, campaign or listing change during the test. Apple warns that a new version containing assets or metadata under test may affect results.

In Apple Analytics, the original product page is the baseline, with conversion-rate cards for the baseline and each treatment. The Improvement Trend graph shows how treatment performance changes over the test period against a dotted baseline. This can help reveal whether the apparent difference is steady or varies over time.

Apple labels a treatment Performing Better or Performing Worse when it reaches 90% confidence. Treat those labels as evidence about the measured conversion outcome, not proof of downstream app use. The result still needs to answer the decision defined before launch.

For Google Play, use any duration or data guidance displayed for the selected experiment, and record its status or uncertainty wording. Interpret the result using the evidence Play Console provides rather than borrowing Apple’s guidance.

Do not apply Apple’s 90-day limit, five-download reporting point or 90% confidence label to Google Play. If Console does not provide a particular duration or data-sufficiency estimate, do not substitute an Apple threshold.

Make and record the decision

For each store, apply a supported treatment, retain the current listing, plan a narrower follow-up or leave the question unresolved. Make the decision from that store’s reported result, and check that the preferred variant still represents the current product. A result on one store is not automatically a winner on the other.

Keep assets, audience settings, dates, the metric, reported result, concurrent changes and the decision together. Where populations or outcome definitions do not align, record separate findings rather than a shared cross-store winner.

In this guide

  1. Comparing screenshot messages by audience needFrame a screenshot message test around a specific user need, keep the images accurate and distinguish randomised tests from audience targeting.
  2. Interpreting store test results with limited trafficRead App Store and Google Play experiment results with small audiences, separate lift from uncertainty and make an honest inconclusive decision.
  3. Recording what a listing experiment changedKeep a useful record of listing variants, eligible traffic, store metrics, overlapping changes, uncertainty and the final decision.

More from App Store Listings

App Store Listings

App store optimisation

Review app store search terms, descriptions, screenshots and listing reports without confusing interest with completed installs.