Skip to content
Confidence
  • Documentation
  • Blog
  • Bootcamp
  • Status
  • Confidence Bootcamp
    • My learning
    • Intro to experimentation
      • Introduction
      • Lesson 1: Why you should experiment
      • Lesson 2: Experiment hypothesis
      • Lesson 3: Success and guardrail metrics
      • Lesson 4: Success metrics
      • Lesson 5: Set up your experiment
      • Lesson 6: Calculation frequency
      • Lesson 7: Target audience
      • Lesson 8: Sample size
      • Lesson 9: Quality assurance
      • Lesson 10: Run your experiment
      • Lesson 11: Evaluate your experiment and make a decision
      • Lesson 12: A/B tests and rollouts
      • Course wrap up
    • Intro to metrics
      • Introduction
      • Lesson 1: What is a metric?
      • Lesson 2: Metric roles
      • Lesson 3: Time considerations
      • Lesson 4: Capturing behavior
      • Lesson 5: Strategic metrics
      • Lesson 6: Interpretability
      • Lesson 7: Feasibility and sensitivity
      • Lesson 8: Variance reduction and metric selection
      • Lesson 9: Select metrics
      • Lesson 10: Segment-level analysis
      • Course wrap up
    • Scientific product development
      • Introduction
      • Lesson 1: Why you should experiment
      • Lesson 2: The scientific method
      • Lesson 3: Randomized controlled trials
      • Lesson 4: Experiment hypothesis
      • Lesson 5: Case study
        • Case study
        • Answers to case study
      • Lesson 6: Why do we need statistics?
      • Lesson 7: Success metrics
      • Lesson 8: Detectable effects and sample size
      • Lesson 9: Make a decision
      • Course wrap up
    • A primer on hypothesis testing
      • Introduction
      • Lesson 1: Introduction to hypothesis testing
      • Lesson 2: True vs estimated effects
      • Lesson 3: Sampling distribution of the difference-in-means estimator
      • Lesson 4: Z-tests and how to reject the null hypothesis
      • Lesson 5: False postive rate and alpha
      • Lesson 6: True positive rate, MDE, and power
      • Course wrap up
    • Intro to Feature Flags
      • Introduction
      • Lesson 1: What is a feature flag?
      • Lesson 2: Lifecycle of a feature flag
      • Lesson 3: Clients
      • Lesson 4: Evaluation context and targeting
    • Sample size calculation - I
      • Introduction
      • Lesson 1: What is the required sample size?
      • Lesson 2: Alpha and power
      • Lesson 3: Baseline mean and variance
      • Lesson 4: Sample size playground - I
    • Sample size calculation - II
      • Introduction
      • Lesson 1: Multi-metric decision making
      • Lesson 2: Number of success metrics
      • Lesson 3: Number of guardrail metrics
      • Lesson 4: Number of comparisons
      • Lesson 5: Sample size playground - II
    • Sample size calculation - III
      • Introduction
      • Lesson 1: Binary metrics
      • Lesson 2: Treatment group proportions
      • Lesson 3: Variance reduction
      • Lesson 4: Sequential testing and sample size
      • Lesson 5: Sample size playground - III
    • Advance your experimentation
      • Introduction
      • Lesson 1: Guardrail metrics with non-inferiority margins
      • Lesson 2: Choose evaluation frequency
      • Lesson 3: Metrics' roles in experiments
      • Lesson 4: Cumulative holdback evaluations
    • Experimentation culture
      • Introduction
      • Lesson 1: Onboarding into experimentation
      • Lesson 2: Empowering experimentation champions
      • Lesson 3: Sustaining the experimentation culture
    • Videos

Lesson 3: How to measure impact with success and guardrail metrics

Summary

Use success metrics to capture what you want to improve. Use guardrail metrics to capture what you don't want to affect negatively.

When you run an experiment, such as an A/B test or a rollout, the ultimate goal is to learn about the impact of the change you made. To know what the impact is, you need to measure the outcome on a relevant set of metrics. The metrics you select can serve different purposes, and even be subject to different statistical tests. This page describes the two main types of metrics you can use to measure impact, and how to select them.

The two types of metrics are:

  • Success metrics. Metrics that you aim to improve with your change.
  • Guardrail metrics. Metrics that you don't expect to improve, but that you want to make sure you don't have a negative impact on.

Success and guardrail metrics

Success metrics are the metrics that you aim to improve with your change. They're what you use to prove that your change had a positive impact.

In companion to success metrics, you should also select guardrail metrics. Guardrail metrics are metrics that help you make sure that your change doesn't have a negative impact on other aspects of your product. This means a hypothesis for an experiment includes two criteria: one for the success metric and one for the guardrail metric. Both criteria need evidence to support the decision to launch the change.

Let's look at some examples of success and guardrail metrics.

Example: Checkout flow

You run an A/B test with an improvement to the checkout flow of your e-commerce website. Your goal is to make the checkout flow more efficient so that your visitors spend less time in the checkout flow. You want to make sure that the improvement in the checkout flow doesn't come at the expense of the number of purchases.

  • Success metric: Average time to completed checkout per visitor.
  • Guardrail metric: Number of purchases per visitor.

Example: Search algorithm

With your new Spotify search algorithm, your hypothesis is that users get better podcast recommendations. You want to measure the impact of the new algorithm on the consumption of podcasts. You want to make sure that the new algorithm change doesn't reduce the number of users that listen to music.

  • Success metric: The average number of podcast minutes played per user.
  • Guardrail metric : The average number of music minutes played per user.

Example: Dating app

You have a dating app that requires new users to complete their profile before they can interact with others. You run a test where your hypothesis is that, showing a dialog with advice for how to onboard, increases the number of users that complete their setup. The dialog shown in your dating app experiment uses new technology, and you want to make sure you don't introduce any bugs.

  • Success metric: Share of users that complete their profile setup.
  • Guardrail metric: Number of crashes per user.
Recommendation

When you start out with experimentation, it is a good idea just to select some guardrail metrics and start experimenting. After you got the hang of it, you can make guardrail metrics even more valuable to your decision making by specifying so-called non-inferior margins (NIM). Learn more about the tests used for guardrail metrics in this Advance your experimentation course lesson.

Example

You are trying to increase the engagement in your product with a new variant that is aiming to increase the engagement in a certain view, say A. To ensure that an increase in engagement in view A doesn't come at the expense of engagement in a competing view B. You should use the engagement in view A as the success metric and the engagement in view B as the guardrail metric. If the new variant increases the engagement in view A and doesn't decrease the engagement in view B, the variant is successful and should be shipped.

How success and guardrail metrics together define a successful variant

For a change to be worth shipping, at least one success metric must have improved significantly, while all guardrail metrics must not have regressed beyond acceptable limits. Both conditions must hold: a win on the success metric doesn't override a failure on a guardrail.

In Confidence

Confidence uses both success and guardrail metrics to identify a successful variant in A/B tests. Read more about how success and guardrail metrics feed into the overall recommendation for a decision.

Reader exercise

What is a guardrail metric used for?

Reader exercise

What is a correct statement?

Was this page helpful?

PreviousLesson 2: Experiment hypothesis
NextLesson 4: Success metrics

© Copyright 2026. All rights reserved.

Follow us on TwitterFollow us on GitHub

On this page

  1. Success and guardrail metrics

  2. How success and guardrail metrics together define a successful variant