Most Shopify stores have never run a controlled test. The change goes live on a Tuesday, the conversion rate moves, and whoever made the change claims the move. Rollouts, Shopify's built-in tool for scheduling and testing storefront changes, was meant to end that, and on 22 September 2026 Shopify rebuilt its setup so that the question it asks first is what kind of change this is.
This is what Rollouts does now, in Shopify's words, how to choose between a launch, an event and an experiment, what is worth testing in a store and what is not, and how to read the result honestly, which is harder than it was a month ago because Shopify changed what a session is on 21 September.
Shopify's definition: "A rollout is a scheduled set of changes to your online store's main theme, checkout and accounts pages, discounts, product catalogs, or a combination of these changes." It lives under Markets in the admin, and since 1 October 2026 discounts can be included in a rollout alongside the theme and checkout changes.
The three types, as the help page states them. A Launch "merges your changes into your store permanently." An Event "publishes your changes for a set period, and then rolls them back." An Experiment "compares a treatment against a control so that you can measure the impact of changes before a full Launch."
The 22 September redesign is about making that choice explicit. Shopify says you now "choose the timing and traffic percentage for your rollout" and the settings adjust to reflect the intent, so an experiment and a launch are set up differently from the first screen. A single rollout can "combine multiple changes" to the theme, checkout and customer account pages. The setup shows what percentage of traffic the rollout reaches and how it is split between versions, and surfaces conflicting settings in the flow rather than after scheduling.
A launch is for a change you have already decided on: a new theme version, a fixed bug, the checkout change you have to make anyway. Scheduling it is the value; nothing is being learned.
An event is for the sale weekend, the collection takeover, the free shipping promotion with a date on it. The roll-back is the value: the store returns to its normal state without somebody remembering to do it at midnight.
The September checklist for a store that holds up through BFCM is a list of events, and Rollouts is where they belong.
An experiment is for the change you believe in and cannot prove: a different product page layout, a different order of information above the fold, a shipping threshold, a checkout configuration. The split is the value, and so is the discipline of writing down, before it starts, what number would make you keep it.
Test the things that decide a purchase and that the store controls: the product page's first screen, the answers to the three questions every buyer of the product asks, the shipping promise and where it appears, the checkout's configuration, the free shipping threshold, whose arithmetic
the free shipping threshold article sets out. These move conversion by amounts a test can see.
Do not test things that are not in doubt. Page speed is not a hypothesis; a slow page loses shoppers and
what actually slows a Shopify store is a repair, not an experiment. Signed-in checkout for returning customers is a configuration with a known effect, covered in
Shopify's 2026 checkout changes for signed-in customers; turn it on. And do not test two things at once in one experiment, however convenient the new combine-changes setting makes it. A rollout that changes the theme and the checkout together tells you the pair moved the number, and nothing about which half did it.
An experiment compares the treatment's conversion with the control's, and both are session-based numbers. On 21 September 2026 Shopify changed how sessions are measured, with the stated aim of "a cleaner view of real shopper activity": bot and system traffic that used to inflate session counts is now filtered more consistently. The practical effect is that a conversion rate from before 21 September and one from after are not the same measure, so an experiment should start after that date and compare only within itself, and a store's history before it is not the baseline.
Shopify traffic but no sales: bot sessions or a store problem is the longer explanation of why the old count misled.
Then the three rules that keep a test honest. Decide the stopping rule before the start: a number of orders per version, not a number of days, because a quiet fortnight produces a confident-looking result from nothing. Keep the split fixed; changing the percentage during the test mixes two populations. And read the whole funnel, not only the conversion rate: an experiment that raises conversion and lowers average order value may have moved revenue the wrong way. Shopify's help pages point to analytics for reviewing a rollout; whether the admin reports statistical significance is not something those pages state, so the stopping rule is the brand's to set.
One experiment at a time on the page that matters most, which for most stores is the best-selling product's page. Write the hypothesis, the single change, the metric and the stopping rule in a sentence each before scheduling. Run it until the stopping rule is met. Launch the winner as a launch, so it is permanent and dated. Then the next one. Four honest experiments in a quarter beat forty opinions, and the record of what was tried and what it did is the store's most valuable document by spring.
We run store experiments in Rollouts to a written plan: the hypothesis, the one change, the split, the stopping rule, the full-funnel read, the launch of what won and the record of what did not. That is the
conversion rate optimisation we do for stores, and the tool's September redesign makes the mechanics cheaper without changing what makes a test worth running. If you have a change you believe in and no proof, that is exactly the right place to start.