Skip to content
Confidence
  • Documentation
  • Blog
  • Bootcamp
  • Status
  • Confidence Bootcamp
    • My learning
    • Intro to experimentation
      • Introduction
      • Lesson 1: Why you should experiment
      • Lesson 2: Experiment hypothesis
      • Lesson 3: Success and guardrail metrics
      • Lesson 4: Success metrics
      • Lesson 5: Set up your experiment
      • Lesson 6: Calculation frequency
      • Lesson 7: Target audience
      • Lesson 8: Sample size
      • Lesson 9: Quality assurance
      • Lesson 10: Run your experiment
      • Lesson 11: Evaluate your experiment and make a decision
      • Lesson 12: A/B tests and rollouts
      • Course wrap up
    • Intro to metrics
      • Introduction
      • Lesson 1: What is a metric?
      • Lesson 2: Metric roles
      • Lesson 3: Time considerations
      • Lesson 4: Capturing behavior
      • Lesson 5: Strategic metrics
      • Lesson 6: Interpretability
      • Lesson 7: Feasibility and sensitivity
      • Lesson 8: Variance reduction and metric selection
      • Lesson 9: Select metrics
      • Lesson 10: Segment-level analysis
      • Course wrap up
    • Scientific product development
      • Introduction
      • Lesson 1: Why you should experiment
      • Lesson 2: The scientific method
      • Lesson 3: Randomized controlled trials
      • Lesson 4: Experiment hypothesis
      • Lesson 5: Case study
        • Case study
        • Answers to case study
      • Lesson 6: Why do we need statistics?
      • Lesson 7: Success metrics
      • Lesson 8: Detectable effects and sample size
      • Lesson 9: Make a decision
      • Course wrap up
    • A primer on hypothesis testing
      • Introduction
      • Lesson 1: Introduction to hypothesis testing
      • Lesson 2: True vs estimated effects
      • Lesson 3: Sampling distribution of the difference-in-means estimator
      • Lesson 4: Z-tests and how to reject the null hypothesis
      • Lesson 5: False postive rate and alpha
      • Lesson 6: True positive rate, MDE, and power
      • Course wrap up
    • Intro to Feature Flags
      • Introduction
      • Lesson 1: What is a feature flag?
      • Lesson 2: Lifecycle of a feature flag
      • Lesson 3: Clients
      • Lesson 4: Evaluation context and targeting
    • Sample size calculation - I
      • Introduction
      • Lesson 1: What is the required sample size?
      • Lesson 2: Alpha and power
      • Lesson 3: Baseline mean and variance
      • Lesson 4: Sample size playground - I
    • Sample size calculation - II
      • Introduction
      • Lesson 1: Multi-metric decision making
      • Lesson 2: Number of success metrics
      • Lesson 3: Number of guardrail metrics
      • Lesson 4: Number of comparisons
      • Lesson 5: Sample size playground - II
    • Sample size calculation - III
      • Introduction
      • Lesson 1: Binary metrics
      • Lesson 2: Treatment group proportions
      • Lesson 3: Variance reduction
      • Lesson 4: Sequential testing and sample size
      • Lesson 5: Sample size playground - III
    • Advance your experimentation
      • Introduction
      • Lesson 1: Guardrail metrics with non-inferiority margins
      • Lesson 2: Choose evaluation frequency
      • Lesson 3: Metrics' roles in experiments
      • Lesson 4: Cumulative holdback evaluations
    • Experimentation culture
      • Introduction
      • Lesson 1: Onboarding into experimentation
      • Lesson 2: Empowering experimentation champions
      • Lesson 3: Sustaining the experimentation culture
    • Videos

Lesson 2: The Spotlight

Summary

In this lesson, you learn what each Spotlight recommendation means, what drives each one, and how the Spotlight synthesizes many metric results into a single actionable signal.

The Spotlight is the first thing you see on the results page. It gives a recommendation for each treatment variant: Ship, Continue, End, or Abort. Each recommendation has a precise meaning, and understanding what drives each one lets you immediately understand the state of your experiment when you open it.

The four recommendations

Ship

Ship means Confidence has found sufficient evidence that the treatment variant is worth rolling out. Three conditions must all be true for a Ship recommendation:

  1. At least one success metric has improved significantly in the intended direction.
  2. All guardrail metrics meet their tolerance levels: either significantly non-inferior (if you use non-inferiority margins), or showing no evidence of deterioration (if you do not).
  3. No health check has failed. There is no SRM and no evidence of metric deterioration.

Ship is a positive signal, but it is a recommendation, not a mandate. You should still exercise judgment. Consider whether the effect size is practically meaningful, whether the experiment ran long enough to rule out novelty effects, and whether the results make sense given your understanding of the product.

Continue

Continue means there is not yet enough evidence to ship, and no reason to stop. This is the most common recommendation for an experiment that is on track. It means:

  • No success metric has significantly improved yet.
  • No health checks have failed.
  • No metrics have deteriorated.

The right response to Continue is to let the experiment run until it reaches the required sample size. Stopping an experiment early because results are not yet significant is a common mistake: it produces biased, unreliable results.

End

End appears when an experiment has collected enough data to be considered powered for all success metrics, but none of those metrics has shown a significant improvement. This is different from Continue.

Continue means "we do not know yet." End means "we have collected enough data, and there is no signal."

An End recommendation is a genuine null result. The treatment variant does not appear to improve the metrics you care about, and you have enough data to be fairly confident in that conclusion. The right response is to stop the experiment and treat this as a real finding: the treatment variant did not work as hypothesized.

Note

A null result is a valuable result. Knowing that a change did not improve metrics saves engineering and design resources that would otherwise go into shipping and maintaining a change that does not help users. Do not dismiss null results.

Abort

Abort means something has gone wrong that makes continuing the experiment harmful or pointless. This recommendation appears when:

  • A health check has failed (most commonly a sample ratio mismatch), which means the results cannot be trusted.
  • One or more metrics have deteriorated significantly, meaning there is statistical evidence the treatment variant is harming something you care about.

When you see Abort, stop the experiment. If the cause is an SRM, investigate the exposure logic before relaunching. If the cause is metric deterioration, the treatment variant may be harmful and should not be shipped.

After the experiment ends

After you end an experiment, the Continue and End recommendations merge into a single Don't ship label. The Abort recommendation may also appear if there was a health check failure. The Ship recommendation remains if the evidence for shipping was already established before the experiment ended.

The full picture

The Spotlight is a synthesis. It takes the outcomes of all success metrics, all guardrail metrics, and all health checks, and compresses them into one recommendation per treatment variant. Understanding the individual components (significance, CI width, health checks, evaluation strategy) gives you the ability to look at a Spotlight recommendation and trace it back to its causes.

Example

An experiment shows a Continue recommendation in the Spotlight, with one success metric showing "Not significant" (+1.4%) and all guardrail metrics showing "Has not deteriorated." The experiment is halfway through its planned duration.

The correct interpretation: the experiment is healthy and on track. The success metric has not yet moved enough to be statistically significant, but "not significant" at the halfway point is expected. Continue running until the required sample size is reached before drawing conclusions.

Reader exercise

The Spotlight shows 'End' for a running experiment. What is the most accurate interpretation?

Reader exercise

Which of the following must be true for Confidence to recommend 'Ship'?

Notes for nerds

The Spotlight's synthesis of many metric results into a single recommendation is a non-trivial statistical problem. When you have many metrics, each tested at some significance level, the probability that at least one shows a spurious significant result grows quickly. Spotify has published research on how to make principled, risk-aware product decisions under exactly this kind of multi-metric setting. The underlying ideas are described in a paper by Schultzberg, Ankargren, and Frånberg (2024). You can read the engineering post at engineering.atspotify.com.

Was this page helpful?

PreviousLesson 1: The anatomy of the results page
NextLesson 3: Means and relative effects

© Copyright 2026. All rights reserved.

Follow us on TwitterFollow us on GitHub

On this page

  1. The four recommendations

  2. After the experiment ends

  3. The full picture

  4. Notes for nerds