Skip to content
Confidence
  • Documentation
  • Blog
  • Bootcamp
  • Status
  • Confidence Bootcamp
    • My learning
    • Intro to experimentation
      • Introduction
      • Lesson 1: Why you should experiment
      • Lesson 2: Experiment hypothesis
      • Lesson 3: Success and guardrail metrics
      • Lesson 4: Success metrics
      • Lesson 5: Set up your experiment
      • Lesson 6: Calculation frequency
      • Lesson 7: Target audience
      • Lesson 8: Sample size
      • Lesson 9: Quality assurance
      • Lesson 10: Run your experiment
      • Lesson 11: Evaluate your experiment and make a decision
      • Lesson 12: A/B tests and rollouts
      • Course wrap up
    • Intro to metrics
      • Introduction
      • Lesson 1: What is a metric?
      • Lesson 2: Metric roles
      • Lesson 3: Time considerations
      • Lesson 4: Capturing behavior
      • Lesson 5: Strategic metrics
      • Lesson 6: Interpretability
      • Lesson 7: Feasibility and sensitivity
      • Lesson 8: Variance reduction and metric selection
      • Lesson 9: Select metrics
      • Lesson 10: Segment-level analysis
      • Course wrap up
    • Scientific product development
      • Introduction
      • Lesson 1: Why you should experiment
      • Lesson 2: The scientific method
      • Lesson 3: Randomized controlled trials
      • Lesson 4: Experiment hypothesis
      • Lesson 5: Case study
        • Case study
        • Answers to case study
      • Lesson 6: Why do we need statistics?
      • Lesson 7: Success metrics
      • Lesson 8: Detectable effects and sample size
      • Lesson 9: Make a decision
      • Course wrap up
    • A primer on hypothesis testing
      • Introduction
      • Lesson 1: Introduction to hypothesis testing
      • Lesson 2: True vs estimated effects
      • Lesson 3: Sampling distribution of the difference-in-means estimator
      • Lesson 4: Z-tests and how to reject the null hypothesis
      • Lesson 5: False postive rate and alpha
      • Lesson 6: True positive rate, MDE, and power
      • Course wrap up
    • Intro to Feature Flags
      • Introduction
      • Lesson 1: What is a feature flag?
      • Lesson 2: Lifecycle of a feature flag
      • Lesson 3: Clients
      • Lesson 4: Evaluation context and targeting
    • Sample size calculation - I
      • Introduction
      • Lesson 1: What is the required sample size?
      • Lesson 2: Alpha and power
      • Lesson 3: Baseline mean and variance
      • Lesson 4: Sample size playground - I
    • Sample size calculation - II
      • Introduction
      • Lesson 1: Multi-metric decision making
      • Lesson 2: Number of success metrics
      • Lesson 3: Number of guardrail metrics
      • Lesson 4: Number of comparisons
      • Lesson 5: Sample size playground - II
    • Sample size calculation - III
      • Introduction
      • Lesson 1: Binary metrics
      • Lesson 2: Treatment group proportions
      • Lesson 3: Variance reduction
      • Lesson 4: Sequential testing and sample size
      • Lesson 5: Sample size playground - III
    • Advance your experimentation
      • Introduction
      • Lesson 1: Guardrail metrics with non-inferiority margins
      • Lesson 2: Choose evaluation frequency
      • Lesson 3: Metrics' roles in experiments
      • Lesson 4: Cumulative holdback evaluations
    • Experimentation culture
      • Introduction
      • Lesson 1: Onboarding into experimentation
      • Lesson 2: Empowering experimentation champions
      • Lesson 3: Sustaining the experimentation culture
    • Videos

Lesson 11: Evaluate your experiment and make a decision

Summary

In this lesson, you learn how to interpret the results of your experiment and make a decision based on the results and how Confidence helps you do that. Use exploration to learn more about the results you got and to get inspiration for new hypotheses.

A good experimentation platform calculates results for you and displays the performance of each variant, taking care of the statistical details so you can focus on learning from the experiment. With those insights, you make a decision on how to proceed with the change you tested.

At this stage, your experiment has successfully run for a period of time and has no visible errors. Congratulations! Now it's time for the fun part. You have at least one result to interpret, but often there are more than just one. More precisely, your experiment has T x M results to interpret, where T = Number of treatment groups (excluding control) and M = Number of metrics.

For an experiment with 3 treatment groups and 4 metrics, you have 12 results to interpret.

Overall decision recommendations

A good experimentation platform provides overall decision recommendations that use the outcomes of all metrics to suggest whether a specific treatment is worth rolling out.

The shipping recommendation recommends you to ship a change if at least one success metric has moved in the desired direction with significance. Simultaneously, all guardrail metrics must be significantly non-inferior, meaning that they're all within the acceptable margin you set using the non-inferiority margin. The test must also be in a healthy state, with no significant negative changes in any of the metrics, and no sign that there is a problem with the quality of the test.

In Confidence

Confidence provides overall decision recommendations on each treatment card on the results page.

Metric results

For each metric, you see a comparison between the control group and each treatment group. You can dig deeper into the results to see metric values, confidence intervals, variances, and more. If you ran your experiment with results delivered continuously, you can also view the results over time.

Exploration

If at the end of the experiment you find things that you would like to dig deeper into, you can do exploratory analysis. Here you can add any metric and see how it performed for each of the treatment groups, and split the results by dimensions.

Note

This type of explorations in which you look at many metrics, perhaps until you find an "interesting" result, severely increases the risk for finding false positives. This means you risk that results are significant only by chance.

For that reason, you shouldn't use exploratory analysis to make decisions about whether an experiment was successful or not. Use it to get inspiration for new hypotheses.

In Confidence

Use the Explore tab to add any metric and split results by dimensions.

Reader exercise

After ending your experiment, the results page in Confidence tells you that 'Shipping might be recommended' for the treatment. What does this mean?

Reader exercise

You ran an A/B-test and it turned out to show no significant difference between the treatment and control. However, you then go into the Exploratory tab, and after looking at about 10 different metrics, you find that there is a significant difference for one of the metrics for the treatment group. What do you do?

Was this page helpful?

PreviousLesson 10: Run your experiment
NextLesson 12: A/B tests and rollouts

© Copyright 2026. All rights reserved.

Follow us on TwitterFollow us on GitHub

On this page

  1. Overall decision recommendations

  2. Metric results

  3. Exploration