Skip to content
  • Pricing
  • Blog
  • Bootcamp
  • Contact us
  • Login
Start free trial

Blog

42% of Spotify experiments get rolled back. That's the point.

At Spotify, 42% of experiments are rolled back after guardrail metrics detect regressions. The discipline to discard is more valuable than the ability to ship.

Read article
August 3, 2026/Johan Rydberg

Accurate Sample Size for Always-Valid Inference

Most A/B testing tools either overestimate the sample size needed for always-valid inference or do not adjust for the sequential test in the sample size calculation at all. We derived a closed-form correction that requires no simulation.

July 14, 2026/Mårten Schultzberg

When A/B tests tell you what you want to hear

A/B testing maturity has a dangerous middle phase: teams test, dislike the answer, and explain it away. How to spot it and learn to trust your data.

July 8, 2026/Johan Rydberg

When AI writes the code, who decides what ships?

AI-accelerated code production increases the need for experimentation. The validation bottleneck grows with build speed. The fastest learners will win.

June 8, 2026/Johan Rydberg

What Makes a Good Sample Size Calculator?

A good sample size calculator must match your analysis: sequential testing, multiple metrics, and variance reduction all change the number it returns.

May 27, 2026/Mårten Schultzberg

The Judgment Gap

AI made building cheap. It also made bad decisions cheaper to ship. The distance between execution speed and validation speed is the judgment gap.

May 20, 2026/Johan Rydberg

The Real ROI of Experimentation

The ROI of experimentation goes beyond counting winners: shipped wins, prevented regressions, and faster organizational learning all add measurable value.

May 19, 2026/Johan Rydberg

Spotify's Experimentation Bootcamp is now free: Introducing Confidence Bootcamp

The A/B testing curriculum Spotify built over ten years to train thousands of experimenters is now free and open to everyone.

April 27, 2026/Mårten Schultzberg

Powered ≠ Trustworthy

Statistical power does not guarantee trustworthy results. Why powered experiments still inflate effects, and what to ask before trusting a significant win.

April 21, 2026/Mårten Schultzberg

Are Optimal Multiple Testing Corrections Optimal for You?

Bonferroni's conservatism reputation is mostly a denominator mistake. Here is why it holds up once you correct only the metrics that need it.

April 14, 2026/Mårten Schultzberg

All posts

When Proxy Metrics Break: How Optimizing for Proxies Can Backfire

View post

Why We Use Separate Tech Stacks for Personalization and Experimentation

View post

Two Questions Every Experiment Should Answer

View post

The Feature Flag Toolbox: Cloud, Edge, and Local

View post

How Experimental Evidence Travels Through Your Organization: Why Better May Be Worse

View post

Beyond Winning: Spotify's Experiments with Learning Framework

View post

A/B Test Bandwidth: The Currency of Innovation

View post

Experiments with Smaller Samples

View post

Reduce Dilution and Improve Sensitivity with Trigger Analysis

View post

Fixed-Power Designs: It's Not IF You Peek, It's WHAT You Peek at

View post
Spotify

Learn more

  • Read our blog
  • Take the bootcamp
  • See comparisons
  • Glossary
  • RFP guides
  • Listen to us
  • Read our docs
  • Status page

Need help

  • Contact us

Legal

  • Terms of Service
  • Data Protection Agreement
  • Privacy Policy

© 2026 Spotify

The Confidence name and logo are registered trademarks of Spotify.