Skip to content
Confidence
  • Documentation
  • Blog
  • Bootcamp
  • Status
  • Confidence Bootcamp
    • My learning
    • Intro to experimentation
      • Introduction
      • Lesson 1: Why you should experiment
      • Lesson 2: Experiment hypothesis
      • Lesson 3: Success and guardrail metrics
      • Lesson 4: Success metrics
      • Lesson 5: Set up your experiment
      • Lesson 6: Calculation frequency
      • Lesson 7: Target audience
      • Lesson 8: Sample size
      • Lesson 9: Quality assurance
      • Lesson 10: Run your experiment
      • Lesson 11: Evaluate your experiment and make a decision
      • Lesson 12: A/B tests and rollouts
      • Course wrap up
    • Intro to metrics
      • Introduction
      • Lesson 1: What is a metric?
      • Lesson 2: Metric roles
      • Lesson 3: Time considerations
      • Lesson 4: Capturing behavior
      • Lesson 5: Strategic metrics
      • Lesson 6: Interpretability
      • Lesson 7: Feasibility and sensitivity
      • Lesson 8: Variance reduction and metric selection
      • Lesson 9: Select metrics
      • Lesson 10: Segment-level analysis
      • Course wrap up
    • Scientific product development
      • Introduction
      • Lesson 1: Why you should experiment
      • Lesson 2: The scientific method
      • Lesson 3: Randomized controlled trials
      • Lesson 4: Experiment hypothesis
      • Lesson 5: Case study
        • Case study
        • Answers to case study
      • Lesson 6: Why do we need statistics?
      • Lesson 7: Success metrics
      • Lesson 8: Detectable effects and sample size
      • Lesson 9: Make a decision
      • Course wrap up
    • A primer on hypothesis testing
      • Introduction
      • Lesson 1: Introduction to hypothesis testing
      • Lesson 2: True vs estimated effects
      • Lesson 3: Sampling distribution of the difference-in-means estimator
      • Lesson 4: Z-tests and how to reject the null hypothesis
      • Lesson 5: False postive rate and alpha
      • Lesson 6: True positive rate, MDE, and power
      • Course wrap up
    • Intro to Feature Flags
      • Introduction
      • Lesson 1: What is a feature flag?
      • Lesson 2: Lifecycle of a feature flag
      • Lesson 3: Clients
      • Lesson 4: Evaluation context and targeting
    • Sample size calculation - I
      • Introduction
      • Lesson 1: What is the required sample size?
      • Lesson 2: Alpha and power
      • Lesson 3: Baseline mean and variance
      • Lesson 4: Sample size playground - I
    • Sample size calculation - II
      • Introduction
      • Lesson 1: Multi-metric decision making
      • Lesson 2: Number of success metrics
      • Lesson 3: Number of guardrail metrics
      • Lesson 4: Number of comparisons
      • Lesson 5: Sample size playground - II
    • Sample size calculation - III
      • Introduction
      • Lesson 1: Binary metrics
      • Lesson 2: Treatment group proportions
      • Lesson 3: Variance reduction
      • Lesson 4: Sequential testing and sample size
      • Lesson 5: Sample size playground - III
    • Advance your experimentation
      • Introduction
      • Lesson 1: Guardrail metrics with non-inferiority margins
      • Lesson 2: Choose evaluation frequency
      • Lesson 3: Metrics' roles in experiments
      • Lesson 4: Cumulative holdback evaluations
    • Experimentation culture
      • Introduction
      • Lesson 1: Onboarding into experimentation
      • Lesson 2: Empowering experimentation champions
      • Lesson 3: Sustaining the experimentation culture
    • Videos

Lesson 12: A/B tests and rollouts

Summary

Use A/B tests to identify the winning variant, use rollouts to ship the winner. For technical changes, like major refactors and migrations, use rollouts to avoid risk and to be able to roll back instantly.

A/B tests and rollouts are two tools for product evaluation. Although similar in some ways, they are usually used for different stages of evaluation.

A/B tests

The main characteristics of A/B tests are that they

  • Can have more than two variants
  • Can have both success metrics and guardrail metrics
  • Have a fixed allocation of the total population
  • Allow for different evaluation frequencies for calculating results

Use A/B tests to

  • Decide a winner among two or more variants of a product
  • Explore and learn about how different settings affect use behavior

A/B tests are flexible and rich product evaluation tools. They help you ensure that the winning variant is better than the losing variants for the business, by allowing you to consider a complete set of success and guardrail metrics.

Rollouts

The main characteristics of a rollout is that it

  • Can only have two variants of which one is the current default and one is the variant that you want to roll out
  • Can only have guardrail metrics
  • Has an allocation that you can gradually increase
  • Always displays results continuously

Use rollouts to

  • Gradually ship a variant while monitoring important guardrail metrics
  • Gradually ship technical changes to the system, for example major refactors and migrations

A great benefit of using rollouts to ship changes is that the change that you roll out is behind a feature flag. This makes it easy to roll back a change. In other words, if you start rolling something out and get an alert that the rollout harms the end-user experience, it is only a button click away to revert back to the earlier experience. This saves engineers a lot of time and agony at Spotify, and has made rollouts the default way for engineers to release changes.

Combine A/B tests and rollouts for important changes

Significant results from experiments with small sample sizes tend to over estimate the treatment effect. Some online experimenters propose that you should replicate the results from such experiments by rerunning the experiment on other users to confirm the result. In practice, it is hard to know which experiments are underpowered. One way to think about it is that you should scrutinize unexpected results harder, and replicate them, to believe in them. A practical way to get confirmation of results from an A/B test is to ship the winner with a rollout. The metric results in the rollout works as a replication of the A/B test results and you can be even more certain about making the right decision.

At Spotify, most A/B tests that identify a winning variant use a rollout to ship that variant, which means that most results are replicated while released to everyone.

In Confidence

In Confidence, A/B tests and rollouts are the two main experiment types. Rollouts are behind feature flags, making it easy to revert with a single click if something goes wrong. Read more about A/B tests and rollouts in the documentation.

Reader exercise

You should choose a rollout instead of an A/B test when you want to:

Was this page helpful?

PreviousLesson 11: Evaluate your experiment and make a decision
NextCourse wrap up

© Copyright 2026. All rights reserved.

Follow us on TwitterFollow us on GitHub

On this page

  1. A/B tests

  2. Rollouts

  3. Combine A/B tests and rollouts for important changes