# Confidence blog

Canonical source: https://confidence.spotify.com/blog
Owner: Spotify AB

> If this file and the page at the canonical URL disagree, the page is authoritative.

- 2026-08-19 — [The product loop](https://confidence.spotify.com/blog/the-product-loop.md): Discovery and delivery are no longer sequential phases. AI collapsed the build phase, and now the teams winning are the ones with the tightest loops.
- 2026-08-03 — [42% of Spotify experiments get rolled back. That's the point.](https://confidence.spotify.com/blog/42-percent-rolled-back.md): At Spotify, 42% of experiments are rolled back after guardrail metrics detect regressions. The discipline to discard is more valuable than the ability to ship.
- 2026-07-14 — [Accurate Sample Size for Always-Valid Inference](https://confidence.spotify.com/blog/accurate-sample-size-always-valid.md): Most A/B testing tools either overestimate the sample size needed for always-valid inference or do not adjust for the sequential test in the sample size calculation at all. We derived a closed-form correction that requires no simulation.
- 2026-07-08 — [When A/B tests tell you what you want to hear](https://confidence.spotify.com/blog/when-ab-tests-tell-you-what-you-want-to-hear.md): A/B testing maturity has a dangerous middle phase: teams test, dislike the answer, and explain it away. How to spot it and learn to trust your data.
- 2026-06-10 — [What experiments actually teach you](https://confidence.spotify.com/blog/what-experiments-teach-you.md): The primary output of a mature experimentation program is better judgment. At Spotify, the learning rate is 64%. The win rate is 12%.
- 2026-06-08 — [When AI writes the code, who decides what ships?](https://confidence.spotify.com/blog/when-ai-writes-the-code.md): AI-accelerated code production increases the need for experimentation. The validation bottleneck grows with build speed. The fastest learners will win.
- 2026-05-27 — [What Makes a Good Sample Size Calculator?](https://confidence.spotify.com/blog/good-sample-size-calculator.md): A good sample size calculator must match your analysis: sequential testing, multiple metrics, and variance reduction all change the number it returns.
- 2026-05-20 — [The Judgment Gap](https://confidence.spotify.com/blog/the-judgment-gap.md): AI made building cheap. It also made bad decisions cheaper to ship. The distance between execution speed and validation speed is the judgment gap.
- 2026-05-19 — [The Real ROI of Experimentation](https://confidence.spotify.com/blog/experimentation-roi.md): The ROI of experimentation goes beyond counting winners: shipped wins, prevented regressions, and faster organizational learning all add measurable value.
- 2026-04-27 — [Spotify's Experimentation Bootcamp is now free: Introducing Confidence Bootcamp](https://confidence.spotify.com/blog/confidence-bootcamp.md): The A/B testing curriculum Spotify built over ten years to train thousands of experimenters is now free and open to everyone.
- 2026-04-21 — [Powered ≠ Trustworthy](https://confidence.spotify.com/blog/powered-trustworthy.md): Statistical power does not guarantee trustworthy results. Why powered experiments still inflate effects, and what to ask before trusting a significant win.
- 2026-04-14 — [Are Optimal Multiple Testing Corrections Optimal for You?](https://confidence.spotify.com/blog/multiple-testing-corrections.md): Bonferroni's conservatism reputation is mostly a denominator mistake. Here is why it holds up once you correct only the metrics that need it.
- 2026-01-26 — [When Proxy Metrics Break: How Optimizing for Proxies Can Backfire](https://confidence.spotify.com/blog/proxy-metrics.md): Learn how wrong things can go when proxy metrics start to influence product development, and how to use them safely.
- 2026-01-07 — [Why We Use Separate Tech Stacks for Personalization and Experimentation](https://confidence.spotify.com/blog/personalization-experimentation-stacks.md): Why Spotify runs personalization and experimentation on separate tech stacks: ML systems need their own infrastructure, and A/B tests evaluate them.
- 2025-12-15 — [Two Questions Every Experiment Should Answer](https://confidence.spotify.com/blog/two-questions.md): Neutral A/B test results often mean the experiment was set up to fail. Learn the two questions every experiment should answer so you always learn something.
- 2025-12-05 — [The Feature Flag Toolbox: Cloud, Edge, and Local](https://confidence.spotify.com/blog/feature-flag-tool-box.md): Feature flagging with Confidence: compare cloud, local, and edge flag resolution, the latency and resilience tradeoffs, and how to pick the right setup.
- 2025-11-24 — [How Experimental Evidence Travels Through Your Organization: Why Better May Be Worse](https://confidence.spotify.com/blog/experimental-evidence.md): Learn about how experimental evidence travels through your organization, and why better may be worse.
- 2025-09-23 — [Beyond Winning: Spotify's Experiments with Learning Framework](https://confidence.spotify.com/blog/experiments-with-learning.md): How the Experiments with Learning (EwL) framework measures experimentation success beyond win rates.
- 2025-09-04 — [A/B Test Bandwidth: The Currency of Innovation](https://confidence.spotify.com/blog/ab-testing-bandwidth.md): Learn about how experiment bandwidth is the currency of innovation, and how you can improve yours.
- 2024-11-26 — [Experiments with Smaller Samples](https://confidence.spotify.com/blog/smaller-sample-experiments.md): Learn about how you can benefit from experimentation even when your samples are smaller than you wish.
- 2024-06-07 — [Reduce Dilution and Improve Sensitivity with Trigger Analysis](https://confidence.spotify.com/blog/trigger-analysis.md): Read about how trigger analysis can help improve sensitivity by narrowing the exposure definition.
- 2024-05-15 — [Fixed-Power Designs: It's Not IF You Peek, It's WHAT You Peek at](https://confidence.spotify.com/blog/fixed-power-designs.md): A new experimental design that lets you estimate sample size during the experiment without compromising inference.
- 2024-04-11 — [Better Product Decisions with Guardrail Metrics](https://confidence.spotify.com/blog/better-decisions-with-guardrails.md): What guardrail metrics are and how to use them in A/B tests: the metrics that catch regressions — at Spotify they trigger rollbacks in 42% of experiments.
- 2024-03-20 — [Collaboration Fuels Efficient Experimentation](https://confidence.spotify.com/blog/collaboration.md): Collaboration fuels efficient experimentation. See how Confidence supports every role, from hypothesis to analysis, so cross-functional teams learn faster.
- 2024-03-05 — [Risk-Aware Product Decisions in A/B Tests with Multiple Metrics](https://confidence.spotify.com/blog/risk-aware-decisions.md): Risk-aware product decisions in A/B tests: how Spotify combines success and guardrail metrics into one ship or no-ship call while controlling error rates.
- 2024-02-15 — [Experiment like Spotify: Analysis of Experiments](https://confidence.spotify.com/blog/experiment-analysis.md): Experiment analysis is a common bottleneck. See how Confidence automates analysis with trusted statistical methods so your experimentation program scales.
- 2024-01-25 — [Experiment like Spotify: Feature Flags](https://confidence.spotify.com/blog/feature-flags.md): Feature flags in Confidence let you ship behind a flag, validate changes in A/B tests, and roll out winners safely. See how Spotify uses flags every day.
- 2024-01-11 — [Experiment like Spotify: A/B Tests and Rollouts](https://confidence.spotify.com/blog/ab-tests-and-rollouts.md): Read more about the difference between an A/B test and a rollout, and how they're used at Spotify.
- 2023-12-18 — [Experiment like Spotify: With Confidence](https://confidence.spotify.com/blog/experiment-like-spotify.md): Confidence brings Spotify experimentation to your team: a warehouse-native platform built from 10+ years of testing at scale, designed to grow with you.
- 2023-07-25 — [The Peeking Problem 2.0 (Part 2): Sequential Testing](https://confidence.spotify.com/blog/peeking-problem-part-2.md): Sequential testing with longitudinal data can inflate false positive rates. Part 2 covers modeling approaches that handle repeated observations correctly.
- 2023-07-18 — [The Peeking Problem 2.0 (Part 1)](https://confidence.spotify.com/blog/peeking-problem-part-1.md): A new challenge in sequential testing with longitudinal data that can inflate false positive rates.
- 2023-03-21 — [Choosing a Sequential Testing Framework — Comparisons and Discussions](https://confidence.spotify.com/blog/sequential-testing-comparison.md): Sequential testing frameworks compared through simulation: group sequential tests, always-valid inference, and how your data infrastructure drives the choice.
- 2022-02-28 — [Search Journey Towards Better Experimentation Practices](https://confidence.spotify.com/blog/experimentation-journey.md): Building experimentation practices takes sustained effort. How the Spotify Search team went from ad hoc testing to data-informed product development.
