Zum Inhalt springen

For new clients: we design your mobile menu free of charge.See mobile designs

A/B testing for online stores

A/B testing: every test needs a reason.

We are an A/B testing agency. We test changes to your online store before they are built in for good. Every test starts with a hypothesis and ends with a decision: build it in or discard it.

  • Hypothesis written down before the start
  • Required order volume calculated in advance
  • Result: build it in or discard it

djoonA/B tests and user flow optimization: only what won in the test was built in for good.Watch the video

How we test

A/B testing with the 545 method: five steps.

The A/B test is the last of five steps. The four before it decide whether the test says anything at all. Behind this is our 545 method: 5 decisions lead to a purchase, 4 forces act on each one, 5 steps lead from measurement to decision.

  1. Schritt 1: Measure

    Are the numbers the test is built on correct?

    If the tracking counts wrong, the test measures wrong. So we first check whether your tracking records visits and orders completely. Then we collect your order volume and the metric for each of the five decisions your buyers make, split by device and traffic source.

    • Check tracking of visits and orders
    • One metric per decision, from click-through rate to repeat purchase rate
    • Split by mobile, desktop and traffic source
    • Orders per month as the basis for every test plan

    ResultNumbers you can trust, and the answer to whether your order volume is enough for A/B tests.

  2. Schritt 2: Narrow down

    Where does your store lose the most buyers?

    We test at the decision where you lose the most buyers in absolute terms. A weak rate on a rarely visited page can cost less than a mediocre one in the cart. At the same time we check whether enough visitors see this spot for a test to deliver a result there in a reasonable time.

    • The decision with the largest absolute loss
    • Visitors and orders at exactly this spot
    • Order of test locations, with reasons

    ResultThe spot where we test, with reasons. You see it before we begin.

  3. Schritt 3: Explain

    Why do you lose buyers exactly there?

    We check the spot against the four forces ease, framing, loss and reassurance and show which one is at work there: with customer voices, data and a UX audit. In our experience this is the step most often skipped. Then a test may examine a force that is not even at work at this spot.

    • Check against ease, framing, loss and reassurance
    • Customer voices from surveys, reviews and forums
    • UX audit, meaning a review of the page for usability
    • Hypothesis: if we change X, we expect Y, because Z

    ResultA written hypothesis. Or the honest finding that the cause is price, range or traffic.

  4. Schritt 4: Design

    One hypothesis, one change.

    We design the variant with copy, images and layout and code it for the test. If you bring your own ideas and designs, we build them as the variant. Before the start we check the variant on mobile and desktop and make sure both groups are measured cleanly.

    • Design of the variant: copy, images, layout
    • Or your own design as the variant
    • Coding of the variant for the test
    • Quality check before the start

    ResultA tested variant that differs only in the one point the hypothesis names.

  5. Schritt 5: Test

    Did the change work?

    Your visitors are split in half between the current page and the variant. How many orders the test needs is fixed before the start. We evaluate once that number is reached: by statistical significance, and where needed split by device and by new versus returning customers.

    • Visitors split 50/50, the split stays the same
    • No stopping at the first good interim result
    • Evaluation by statistical significance
    • Winners built in for good, losers documented

    ResultA decision backed by numbers: build it in or discard it. Then step 1 begins again with the next measurement.

What's included

See what a change really delivers.

Hypothesis, design, technology and evaluation work together. Both versions run at the same time. If the difference is significant, it comes from your change, not from season or chance.

  1. Hypothesis first

    If we change X, we expect Y, because Z. We write this sentence down before the start. That way the result cannot be reinterpreted afterwards.

  2. Sample size calculated in advance

    We calculate how many orders the test needs before it starts. Anyone who stops at the first good interim result easily declares chance the winner.

  3. One goal per test

    Before the start we define what the test is measured on, for example clicks on the add-to-cart button or orders. The goal does not change during the test.

  4. Quality check

    Before visitors see the variant, we check its display and function on mobile and desktop. A bug in the variant would distort the result.

  5. Clean split

    Visitors are split 50/50, and the split stays the same until the end. Where needed we test mobile and desktop separately.

  6. Evaluation by significance

    Significant means: the difference is very unlikely to be chance. Where needed we evaluate new and returning customers separately. The same change can work in opposite directions for the two groups.

  7. Clear decision

    Winners are built in for good, by us or by your development team depending on the store platform. Losers we document and tell you about openly.

  8. Documentation

    Hypothesis, runtime, numbers and decision are recorded for every test. So the next test builds on the last one. The list of hypotheses belongs to you.

Honest assessment

When A/B testing pays off and when it does not.

For an A/B test, the number of orders matters more than the number of visitors. Whether a test pays off at your order volume, we tell you before you hire us.

  1. Enough orders

    Rule of thumb: to detect a difference of 20 percent, a test needs around 400 orders per variant, for 10 percent around 1,600. Effects half as large need four times the volume.

  2. Reliable measurement

    If the tracking counts orders incompletely or the store is technically unstable, the test measures errors instead of behavior. Then we fix the measurement first.

  3. Enough time

    A test runs for at least one full business cycle. At 200 to 300 orders per month, even proving a 20 percent difference takes around three to four months.

  4. A real question

    A test pays off when its result decides something: a change that many visitors see and that you would otherwise build in on a hunch.

  5. Too few orders

    Then we work with customer surveys, UX audits and before-and-after comparisons and name the uncertainty openly. The A/B test is one method among several.

  6. Obvious blockers

    Usability errors, missing trust elements, unclear navigation: what clearly blocks, we fix without a test. A test would let half of your visitors keep running into the error.

Behind every test: a team that wants to know for sure.

Strategy and hypothesis, design, development, quality check, evaluation: several disciplines work on every test, and none of them settles for a good interim result. Your contact stays one of the two founders. In the end you get a decision backed by numbers: build it in or discard it.

  1. Strategy and hypothesis
  2. Design
  3. Development
  4. Quality check
  5. Evaluation

As featured in:

Selected trade magazines and media have covered our work and results.

  • This is Marketing
  • Kameleoon
  • Founders League
  • CRO.CAFE
  • Merchant Inspiration

Questions about A/B testing

Common questions about A/B testing with an agency.

Cost, runtime, order volume: here are the short answers. Whatever stays open, we clear up in the intro call.

Ask your open questions

What is an A/B test?

In an A/B test, part of your visitors see the current page (A), the other part sees a changed version (B). Both run at the same time, under the same conditions.

The numbers show which version sells better and how certain the difference is. So you only build in a change once it has proven itself.

What does A/B testing with an agency cost?

It depends on how many tests should run and whether we only code and evaluate or also develop hypotheses and designs. A flat price that fits every store would be a guess.

You get a precise estimate in the intro call, once we know your store and your order volume.

How long does an A/B test take?

We set this before the start: from the number of orders the test needs and the number your store reaches in that time. A test runs for at least one full business cycle.

At around 1,000 orders per month, a difference of 20 percent can be proven in about four weeks, at 300 orders only after around three months. We do not stop a test at the first good interim result, because early values fluctuate strongly.

How many orders do you need for A/B testing to pay off?

As a rule of thumb: to prove a difference of 20 percent, a test needs around 400 orders per variant, for 10 percent around 1,600.

If the order volume is not enough, we work with customer surveys, UX audits and before-and-after comparisons and name the uncertainty openly. Whether an A/B test pays off for you, we tell you before you place the order.

What is the difference between A/B testing and conversion optimization?

Conversion optimization (CRO) is the whole program: we measure where your store loses buyers, narrow down the spot, explain the cause, redesign the page and test the result.

A/B testing is the measurement method within it. You can also book it on its own if your team already has its own ideas and designs and wants to test them cleanly.

Can you test our own ideas and designs?

Yes, that is what A/B testing as a standalone service is for. We record your idea as a hypothesis, code your design as the variant, check it before the start and evaluate the test.

Beforehand we tell you whether the test can deliver a reliable result at your order volume.

Do you also test on store platforms other than Shopify?

Yes, we run analyses and A/B tests for stores on other platforms too, because how people decide does not depend on the store platform.

In Shopify we build winners in ourselves. On other platforms, your own development team implements the change.

Is an A/B test worth it for your store?

In the intro call you talk for 30 minutes with David Pehl or Anas Elimani. You describe your situation and your order volume, and we tell you openly whether A/B tests will deliver a result for you or whether another approach fits better.

Book an intro call

  • 30 minutes, no obligation
  • With one of the two founders
  • Experience from 180+ stores we have worked on

We are proud to be a Shopify Premier Partner.

Clicking loads the chat of our provider. No data is sent to them before that.