Skip to content
Confidence
  • Documentation
  • Blog
  • Bootcamp
  • Status
  • Confidence Bootcamp
    • My learning
    • Intro to experimentation
      • Introduction
      • Lesson 1: Why you should experiment
      • Lesson 2: Experiment hypothesis
      • Lesson 3: Success and guardrail metrics
      • Lesson 4: Success metrics
      • Lesson 5: Set up your experiment
      • Lesson 6: Calculation frequency
      • Lesson 7: Target audience
      • Lesson 8: Sample size
      • Lesson 9: Quality assurance
      • Lesson 10: Run your experiment
      • Lesson 11: Evaluate your experiment and make a decision
      • Lesson 12: A/B tests and rollouts
      • Course wrap up
    • Intro to metrics
      • Introduction
      • Lesson 1: What is a metric?
      • Lesson 2: Metric roles
      • Lesson 3: Time considerations
      • Lesson 4: Capturing behavior
      • Lesson 5: Strategic metrics
      • Lesson 6: Interpretability
      • Lesson 7: Feasibility and sensitivity
      • Lesson 8: Variance reduction and metric selection
      • Lesson 9: Select metrics
      • Lesson 10: Segment-level analysis
      • Course wrap up
    • Scientific product development
      • Introduction
      • Lesson 1: Why you should experiment
      • Lesson 2: The scientific method
      • Lesson 3: Randomized controlled trials
      • Lesson 4: Experiment hypothesis
      • Lesson 5: Case study
        • Case study
        • Answers to case study
      • Lesson 6: Why do we need statistics?
      • Lesson 7: Success metrics
      • Lesson 8: Detectable effects and sample size
      • Lesson 9: Make a decision
      • Course wrap up
    • A primer on hypothesis testing
      • Introduction
      • Lesson 1: Introduction to hypothesis testing
      • Lesson 2: True vs estimated effects
      • Lesson 3: Sampling distribution of the difference-in-means estimator
      • Lesson 4: Z-tests and how to reject the null hypothesis
      • Lesson 5: False postive rate and alpha
      • Lesson 6: True positive rate, MDE, and power
      • Course wrap up
    • Intro to Feature Flags
      • Introduction
      • Lesson 1: What is a feature flag?
      • Lesson 2: Lifecycle of a feature flag
      • Lesson 3: Clients
      • Lesson 4: Evaluation context and targeting
    • Sample size calculation - I
      • Introduction
      • Lesson 1: What is the required sample size?
      • Lesson 2: Alpha and power
      • Lesson 3: Baseline mean and variance
      • Lesson 4: Sample size playground - I
    • Sample size calculation - II
      • Introduction
      • Lesson 1: Multi-metric decision making
      • Lesson 2: Number of success metrics
      • Lesson 3: Number of guardrail metrics
      • Lesson 4: Number of comparisons
      • Lesson 5: Sample size playground - II
    • Sample size calculation - III
      • Introduction
      • Lesson 1: Binary metrics
      • Lesson 2: Treatment group proportions
      • Lesson 3: Variance reduction
      • Lesson 4: Sequential testing and sample size
      • Lesson 5: Sample size playground - III
    • Advance your experimentation
      • Introduction
      • Lesson 1: Guardrail metrics with non-inferiority margins
      • Lesson 2: Choose evaluation frequency
      • Lesson 3: Metrics' roles in experiments
      • Lesson 4: Cumulative holdback evaluations
    • Experimentation culture
      • Introduction
      • Lesson 1: Onboarding into experimentation
      • Lesson 2: Empowering experimentation champions
      • Lesson 3: Sustaining the experimentation culture
    • Videos

Lesson 6: Calculation frequency

Summary
  • Calculating results only once ("Upon Conclusion") is the most efficient way to set up your experiments. You only get to see the results at the end of your experiment, which gives you higher sensitivity.
  • Calculating results continuously sacrifices sensitivity for faster results. Select this if getting a rough estimate early on is important to you.

When you set up an experiment, you need to decide when you want results to be displayed.

Deliver results continuously or upon conclusion

There are two options for when to calculate results:

  • Continuously display results to see new results every hour or day when the experiment is live.
  • Upon Conclusion to see results when you end the experiment.

For both settings, all experiments run until you decide to stop them. The largest difference between the two options is that in experiments with results delivered Upon Conclusion, you can't view the results for your metrics during the experiment. The results show up when you end your experiment. For experiments using results calculated Continuously, you can follow the experiment results during the course of the experiment and decide to end it whenever you want.

If you're wondering why not always choose to view results continuously—stay tuned! We'll explain that shortly.

Two strategies to avoid being fooled by randomness

Imagine that you run an experiment that truly has no impact whatsoever on any metric. You just split users in two groups, measure some outcome, and calculate the difference between the groups every day. Even if the experiment didn't actually change anything, you can expect to see results fluctuate over time, just by random chance. If you check the results every day, you might on some days see a difference that's so large that you wrongly assume that this experiment impacts the metric. Such results are called false positive results. Checking the results every day means there are multiple chances to find a false significant result. To avoid getting fooled by randomness and draw the wrong conclusions based on a false positive result, you can use either of two strategies to keep this risk under control.

  • Calculate result upon conclusion and use standard statistical tests

    Before starting your experiment, you define a point in time when you plan to calculate the results. To make sure that you have a good chance of finding an impact, you calculate what sample size you need to reliably detect an effect of a certain size. By only calculating the results this one time, you have only one chance to be fooled by a false positive result. This method is the most efficient way to minimize the risks of false positives.

  • Calculate the results continuously and use sequential tests

    You calculate results every day or every hour, and correct for the increased risk of being fooled by randomness. The statistical methods used, known as sequential tests, correct for the multiple peeks at the data by using a stricter threshold to conclude significance. With this approach, you can check the results daily without hesitation. Because the tests use a stricter threshold for calling significance, the impact needs to be larger to be reliably detected. This means that for an experiment with results updated continuously, you either need to increase the sample size, or accept that you lose some sensitivity to detect changes. Sequential tests trade off sensitivity for faster results.

Learn about difference between different evaluation frequencies in 2 minutes and 14 seconds.

Automatic sequential monitoring

Modern experimentation platforms run automatic sequential monitoring checks regardless of the evaluation frequency you choose. This means that you do not have to choose to calculate results continuously to sleep well at night—the platform monitors your experiment for deterioration in all your metrics.

In Confidence

Confidence monitors your experiment using sequential tests for all checks, regardless of the evaluation frequency you choose.

What you should choose

The trade-off is different in all experiments. Sequential experiments offer you the opportunity to stop experiments early at the expense of some certainty. If speed is more important than precise estimates of the treatment effect, continuously updating the results is the better choice.

For a given experiment length, like say two weeks, only calculating the results once at the end gives more precise estimates. If you are going to run the experiment for a fixed amount of time regardless, select Upon Conclusion to maximize your chances of finding effects.

Example

At Spotify, most experiments use Upon Conclusion and only display results when the experiment ends. For early abortion of experiments due to errors or negative user experiences, experimenters rely on Confidence's monitoring of their experiments. For the product decision, most teams want at least two weeks of data to make decisions to ship.

Reader exercise

What setup should you select if you need to be able to see results every day?

Was this page helpful?

PreviousLesson 5: Set up your experiment
NextLesson 7: Target audience

© Copyright 2026. All rights reserved.

Follow us on TwitterFollow us on GitHub

On this page

  1. Deliver results continuously or upon conclusion

  2. Two strategies to avoid being fooled by randomness

  3. Automatic sequential monitoring

  4. What you should choose