Skip to content
Confidence
  • Documentation
  • Blog
  • Bootcamp
  • Status
  • Confidence Bootcamp
    • My learning
    • Intro to experimentation
      • Introduction
      • Lesson 1: Why you should experiment
      • Lesson 2: Experiment hypothesis
      • Lesson 3: Success and guardrail metrics
      • Lesson 4: Success metrics
      • Lesson 5: Set up your experiment
      • Lesson 6: Calculation frequency
      • Lesson 7: Target audience
      • Lesson 8: Sample size
      • Lesson 9: Quality assurance
      • Lesson 10: Run your experiment
      • Lesson 11: Evaluate your experiment and make a decision
      • Lesson 12: A/B tests and rollouts
      • Course wrap up
    • Intro to metrics
      • Introduction
      • Lesson 1: What is a metric?
      • Lesson 2: Metric roles
      • Lesson 3: Time considerations
      • Lesson 4: Capturing behavior
      • Lesson 5: Strategic metrics
      • Lesson 6: Interpretability
      • Lesson 7: Feasibility and sensitivity
      • Lesson 8: Variance reduction and metric selection
      • Lesson 9: Select metrics
      • Lesson 10: Segment-level analysis
      • Course wrap up
    • Scientific product development
      • Introduction
      • Lesson 1: Why you should experiment
      • Lesson 2: The scientific method
      • Lesson 3: Randomized controlled trials
      • Lesson 4: Experiment hypothesis
      • Lesson 5: Case study
        • Case study
        • Answers to case study
      • Lesson 6: Why do we need statistics?
      • Lesson 7: Success metrics
      • Lesson 8: Detectable effects and sample size
      • Lesson 9: Make a decision
      • Course wrap up
    • A primer on hypothesis testing
      • Introduction
      • Lesson 1: Introduction to hypothesis testing
      • Lesson 2: True vs estimated effects
      • Lesson 3: Sampling distribution of the difference-in-means estimator
      • Lesson 4: Z-tests and how to reject the null hypothesis
      • Lesson 5: False postive rate and alpha
      • Lesson 6: True positive rate, MDE, and power
      • Course wrap up
    • Intro to Feature Flags
      • Introduction
      • Lesson 1: What is a feature flag?
      • Lesson 2: Lifecycle of a feature flag
      • Lesson 3: Clients
      • Lesson 4: Evaluation context and targeting
    • Sample size calculation - I
      • Introduction
      • Lesson 1: What is the required sample size?
      • Lesson 2: Alpha and power
      • Lesson 3: Baseline mean and variance
      • Lesson 4: Sample size playground - I
    • Sample size calculation - II
      • Introduction
      • Lesson 1: Multi-metric decision making
      • Lesson 2: Number of success metrics
      • Lesson 3: Number of guardrail metrics
      • Lesson 4: Number of comparisons
      • Lesson 5: Sample size playground - II
    • Sample size calculation - III
      • Introduction
      • Lesson 1: Binary metrics
      • Lesson 2: Treatment group proportions
      • Lesson 3: Variance reduction
      • Lesson 4: Sequential testing and sample size
      • Lesson 5: Sample size playground - III
    • Advance your experimentation
      • Introduction
      • Lesson 1: Guardrail metrics with non-inferiority margins
      • Lesson 2: Choose evaluation frequency
      • Lesson 3: Metrics' roles in experiments
      • Lesson 4: Cumulative holdback evaluations
    • Experimentation culture
      • Introduction
      • Lesson 1: Onboarding into experimentation
      • Lesson 2: Empowering experimentation champions
      • Lesson 3: Sustaining the experimentation culture
    • Videos

Lesson 9: Make a Decision

Summary

In this lesson, you learn about decision making in the context of experimentation.

To benefit from experimentation in your decision making you should:

  • Have a pre-determined decision rule
  • Ship successful variants, iterate on non-successful variants using explorations and experiments

Define a decision rule before you run the experiment

It is important to have pre-defined decision rule that maps any possible outcome of the experiment to a product decision. For example, what will you do if one guardrail metric moved in the wrong direction, but everything else looks good? Predetermining the rule makes it easier to not change the goal after you see the results.

In Confidence

In Confidence, there is a default decision rule that gives decision recommendations throughout the experiments. Read more about how the recommendations are constructed in the documentation.

Iterate on a product with experimentation

The following chart shows how to apply the scientific method when you make a decision based on an experiment.

Decision making

If the experiment confirms the hypothesis and there is evidence that your change works well, then you can proceed and roll out the change to all users.

If the experiment is not successful, that is, if there is a negative effect detected on a guardrail metric or no success metric has improved, you should not roll out the change. Instead, you should try and understand why this iteration didn't have the intended effect, fix the problem, and then re-run the experiment to see if your new fix actually fixed the problem.

There are several ways of trying to understand why an iteration didn't have the intended effect. For example, you can do exploratory analysis by diving into segments and additional metrics to better understand your results. It is also common to do user research to get more in-depth understanding of how users experienced the change. All this information can then be used to formulate a new hypothesis, and iterate on the product. You then test the new iteration in a new experiment.

In Confidence

In Confidence, you can do exploratory analysis directly in the explore tab, diving into segments and additional metrics to better understand your results.

Although exploratory research and analysis is an important and natural step to inform new iterations, don't use it to make a decision on the finished experiment. Changing the prediction to match the results invalidates the conclusion.

Of course, it is not a good idea to iterate forever. If repeated experiments fail to show improvement, that is a signal to abandon the hypothesis rather than keep refining it.

Reader exercise

What should you do if you detect a negative effect on a guardrail metric?

Was this page helpful?

PreviousLesson 8: Detectable effects and sample size
NextCourse wrap up

© Copyright 2026. All rights reserved.

Follow us on TwitterFollow us on GitHub

On this page

  1. Define a decision rule before you run the experiment

  2. Iterate on a product with experimentation