Skip to content
Store slow, dated, or hard to edit? Request a Store Diagnostic
Running B2B on spreadsheets or email? Plan Your B2B Project
CRO

Shopify A/B Testing: What You Can (and Can’t) Test on Shopify vs Shopify Plus

A/B testing gets talked about as if it’s a single, universally available capability, install a tool, split your traffic, wait for a winner. On Shopify, that’s only partly true. What you can actually test, and how reliably, depends heavily on whether you’re on standard Shopify or Shopify Plus, because Shopify’s checkout has historically been locked down for good reason (payment security, PCI compliance, platform stability), and that lock changes what testing tools can touch.

This post breaks down what’s realistically testable on each plan, what tools you need, and, just as importantly, when A/B testing isn’t actually the right tool for the question you’re trying to answer.

The Core Constraint: Shopify’s Checkout

On standard Shopify plans, the checkout is a fixed, Shopify-hosted flow. You can customise branding (colours, logo, some layout elements) via Settings > Checkout, but you cannot inject custom code, run scripts, or serve different checkout experiences to different visitor segments. This is deliberate, Shopify controls checkout tightly to maintain PCI compliance and conversion performance across its entire merchant base.

Shopify Plus changes this materially through checkout extensibility, a framework that lets developers build custom checkout UI extensions and functions using Shopify’s supported APIs, without touching the underlying checkout.liquid (which Shopify has deprecated in favour of extensions anyway). This opens up genuine checkout-level testing that simply isn’t possible on standard plans.

What You Can Test on Standard Shopify

Plenty is still testable, most of the highest-impact conversion factors actually live outside checkout:

  • Product page layout and content: image order, description length and format, review placement, size guide presentation, badge/trust signal placement
  • Collection page structure: grid density, filtering options, sort order defaults, featured product placement
  • Homepage messaging and hero content: value proposition wording, hero image/video, featured collections
  • Pricing display: showing/hiding compare-at pricing, bundle framing, currency formatting
  • Cart page and cart drawer: upsell placement, shipping threshold messaging, cart page versus slide-out cart behaviour
  • Pop-ups and on-site messaging: exit intent offers, free shipping bars, urgency messaging
  • Pre-checkout steps: anything up to the “Check out” button, including guest vs account prompts on the cart

These are typically run through third-party testing apps available on the Shopify App Store, which work by showing different theme content or app-injected content to different visitor segments and measuring downstream conversion. Some run as native theme app extensions; others rely on client-side script injection, which is worth checking for speed impact since a slow test can quietly cost you more in performance than it teaches you in insight.

What You Cannot Test on Standard Shopify

  • Checkout page fields, layout, or flow, this is fixed and cannot be split-tested
  • Different payment method presentations or ordering at checkout
  • Custom checkout upsells or post-purchase offers, not available without Plus
  • Shipping method presentation logic beyond what Shopify’s native settings allow

If your hypothesis is “I think changing something about checkout itself would improve conversion,” standard Shopify simply isn’t the environment to test that in. You can test everything that leads up to checkout, and infer some things indirectly (e.g. testing shipping threshold messaging on the cart page to see if it changes checkout completion), but you can’t touch checkout directly.

What Shopify Plus Adds

Checkout extensibility on Shopify Plus allows testing of things that were previously locked entirely:

  • Custom checkout fields and layout changes via UI extensions
  • Post-purchase upsell offers (a genuine incremental AOV lever many Plus merchants underuse)
  • Custom shipping and payment method presentation logic
  • Checkout branding and content blocks tailored by customer segment, product type, or cart contents, within what Shopify’s extension APIs support
  • Functions for custom discount, shipping, and payment logic that can be tested against default behaviour

This is meaningfully more powerful, but it also comes with more implementation overhead, checkout extensions are built with Shopify’s specific UI extension framework, not arbitrary custom code, so testing here usually needs proper development resourcing rather than a plug-and-play app. It’s one of the clearer cases where the jump to Shopify Plus development unlocks capability that standard Shopify structurally can’t offer, though it’s worth being honest that not every store needs it, if your biggest conversion problems live in your product pages or traffic quality, Plus checkout testing won’t move those numbers.

Tools Commonly Used for Each Type of Test

For pre-checkout testing on standard Shopify, most merchants use a dedicated A/B testing or personalisation app from the Shopify App Store. These generally work in one of two ways: theme app extensions that integrate more cleanly with Shopify’s native theme architecture and tend to have less speed impact, or script-injection tools that overlay changes client-side, which are often faster to set up but can introduce a brief “flash” of the original content before the variant loads, and add a script weight cost worth testing for.

For Shopify Plus checkout testing, there isn’t yet a mature ecosystem of plug-and-play testing apps working directly within checkout extensions the way there is for pre-checkout pages, most checkout-level testing on Plus is built as custom development work: two versions of a UI extension or Function, with traffic split at the code level and results tracked against Shopify’s own order and analytics data rather than a dedicated third-party testing dashboard. This is a genuine gap in the tooling ecosystem worth knowing about before assuming checkout testing will be as turnkey as testing a product page headline.

Common Mistakes That Invalidate Shopify Test Results

  • Running a test during a sale period or major campaign, where external demand spikes swamp the actual effect of the variable being tested.
  • Testing on too small a page or too low-traffic a segment to ever reach a meaningful sample size within a reasonable timeframe.
  • Changing the test mid-flight, editing variant B after a test has started invalidates the comparison, since early and late traffic saw different things.
  • Ignoring device split, a change that helps desktop conversion can hurt mobile conversion (or vice versa) and a blended result can mask both effects cancelling out.
  • Declaring a winner too early, before reaching statistical significance, based on an early lead that later reverses as more data comes in.

Statistical Reality Check Before You Start

A/B testing only produces reliable answers with enough traffic and enough conversions per variant. This is the part most guides skip. If your store gets a few hundred sessions a day and a modest conversion rate, running a proper split test to statistical significance can take weeks, and testing something with a small expected effect size may never reach a clean result at all.

Before setting up any test, be honest about:

  • Your baseline traffic volume to the page or step being tested
  • Your baseline conversion rate, to estimate how long a test needs to run
  • The minimum effect size worth detecting, a 2% lift might not be practically meaningful even if it were detectable

For lower-traffic stores, sequential testing (change something, measure before/after over a comparable period, control for seasonality) is often more practical than a true simultaneous split test, even though it’s statistically weaker. It’s a trade-off worth making deliberately rather than running an underpowered A/B test and treating an inconclusive result as a “loss.”

Qualitative Data Should Inform What You Test

A/B testing tells you which of two specific options performs better, but it doesn’t tell you what to test in the first place. That’s where qualitative input earns its place, session recordings, on-site surveys, and support ticket themes often surface the actual hypothesis worth testing, rather than guessing based on best-practice lists from other stores. If several customers mention in support tickets that they weren’t sure about sizing, that’s a stronger starting hypothesis for a product page test than an arbitrary layout change picked because a competitor does it that way.

A Simple Framework for Choosing What to Test

  1. Identify the funnel stage with the clearest drop-off (see your Shopify Analytics conversion funnel), test there first, not wherever’s easiest.
  2. Check whether the test requires checkout access. If yes and you’re on standard Shopify, either reframe the hypothesis to something pre-checkout, or evaluate whether Plus is warranted.
  3. Estimate required sample size roughly, using your current conversion rate and traffic, don’t launch a test you can’t realistically finish in a reasonable window.
  4. Test one meaningful variable at a time. Bundling multiple changes into one “variant B” makes the result impossible to act on specifically.
  5. Let it run a full business cycle (at least one to two weeks, ideally including a weekend and a weekday) before reading results, to avoid day-of-week bias.

FAQ

Can I A/B test Shopify checkout without Shopify Plus?
No, standard Shopify’s checkout is fixed and not customisable at the code level, so it can’t be split-tested directly. You can still test everything up to the “Check out” button, including cart page content and pre-checkout messaging.

What’s the difference between A/B testing and just changing something and watching the numbers?
A/B testing splits traffic simultaneously between two versions, controlling for time-based factors like day of week or a concurrent ad campaign. Before/after testing changes one thing and compares periods sequentially, simpler to set up but more vulnerable to external factors muddying the result. Both have a place depending on your traffic volume.

Do I need a dedicated app for A/B testing on Shopify?
For most standard Shopify testing (product pages, collection pages, on-site messaging), yes, a testing app from the Shopify App Store handles the traffic split and reporting. For Shopify Plus checkout testing, this typically requires custom development using Shopify’s checkout extensibility framework rather than an off-the-shelf app.

How long should a Shopify A/B test run?
Long enough to reach statistical significance for your traffic volume and expected effect size, which varies enormously by store, commonly a minimum of one to two weeks to account for weekly buying patterns, but lower-traffic stores may need considerably longer or should consider sequential testing instead.

Is A/B testing worth it for a small Shopify store?
Often not in the strict simultaneous-split-test sense, simply due to traffic constraints. Smaller stores are usually better served by structured before/after testing of higher-confidence changes (informed by funnel data and user behaviour, not guesswork) rather than running underpowered formal tests.

Next Step

If you’re not sure whether your store has the traffic to test properly, or whether a checkout-level test actually needs Shopify Plus, it’s worth talking it through before building anything. Book a call and we’ll help you work out a realistic testing plan for where your store is right now.

Niraj Raut
Written by Niraj Raut SEO Manager

Niraj Raut is the SEO Manager and co-founder at Nexly. He helps Australian Shopify and Shopify Plus brands earn durable organic growth through technical SEO, search-led store architecture and content that ranks. He writes about what actually moves rankings for ecommerce.

Connect on LinkedIn
Have a Shopify project? Chat with us, takes 30 seconds.