
42% of Spotify experiments get rolled back. That's the point.
At Spotify, 42% of experiments are rolled back after guardrail metrics detect regressions. The discipline to discard is more valuable than the ability to ship.
Read article
At Spotify, 42% of experiments are rolled back after guardrail metrics detect regressions. The discipline to discard is more valuable than the ability to ship.
Read article
Most A/B testing tools either overestimate the sample size needed for always-valid inference or do not adjust for the sequential test in the sample size calculation at all. We derived a closed-form correction that requires no simulation.

A/B testing maturity has a dangerous middle phase: teams test, dislike the answer, and explain it away. How to spot it and learn to trust your data.

AI-accelerated code production increases the need for experimentation. The validation bottleneck grows with build speed. The fastest learners will win.

A good sample size calculator must match your analysis: sequential testing, multiple metrics, and variance reduction all change the number it returns.

AI made building cheap. It also made bad decisions cheaper to ship. The distance between execution speed and validation speed is the judgment gap.

The ROI of experimentation goes beyond counting winners: shipped wins, prevented regressions, and faster organizational learning all add measurable value.

The A/B testing curriculum Spotify built over ten years to train thousands of experimenters is now free and open to everyone.

Statistical power does not guarantee trustworthy results. Why powered experiments still inflate effects, and what to ask before trusting a significant win.

Bonferroni's conservatism reputation is mostly a denominator mistake. Here is why it holds up once you correct only the metrics that need it.









