Skip to content
  • Pricing
  • Blog
  • Bootcamp
  • Contact us
  • Login
Start free trial

Blog

Introducing Confidence Agent

Confidence Agent is an AI collaborator built into Confidence that works with your flags, experiments, metrics, and project documents.

Read article
September 9, 2026/Sebastian Ankargren

The rise of the product builder

As AI increases individual leverage across the product development stack, the boundaries between PM, design, analysis, and engineering are blurring.

September 1, 2026/Johan Rydberg

The product loop

Discovery and delivery are no longer sequential phases. AI collapsed the build phase, and now the teams winning are the ones with the tightest loops.

August 19, 2026/Johan Rydberg

42% of Spotify experiments get rolled back. That's the point.

At Spotify, 42% of experiments are rolled back after guardrail metrics detect regressions. The discipline to discard is more valuable than the ability to ship.

August 3, 2026/Johan Rydberg

Accurate Sample Size for Always-Valid Inference

Most A/B testing tools either overestimate the sample size needed for always-valid inference or do not adjust for the sequential test in the sample size calculation at all. We derived a closed-form correction that requires no simulation.

July 14, 2026/Mårten Schultzberg

When A/B tests tell you what you want to hear

A/B testing maturity has a dangerous middle phase: teams test, dislike the answer, and explain it away. How to spot it and learn to trust your data.

July 8, 2026/Johan Rydberg

What experiments actually teach you

The primary output of a mature experimentation program is better judgment. At Spotify, the learning rate is 64%. The win rate is 12%.

June 10, 2026/Johan Rydberg

When AI writes the code, who decides what ships?

AI-accelerated code production increases the need for experimentation. The validation bottleneck grows with build speed. The fastest learners will win.

June 8, 2026/Johan Rydberg

What Makes a Good Sample Size Calculator?

A good sample size calculator must match your analysis: sequential testing, multiple metrics, and variance reduction all change the number it returns.

May 27, 2026/Mårten Schultzberg

The Judgment Gap

AI made building cheap. It also made bad decisions cheaper to ship. The distance between execution speed and validation speed is the judgment gap.

May 20, 2026/Johan Rydberg

All posts

The Real ROI of Experimentation

View post

Spotify's Experimentation Bootcamp is now free: Introducing Confidence Bootcamp

View post

Powered ≠ Trustworthy

View post

Are Optimal Multiple Testing Corrections Optimal for You?

View post

When Proxy Metrics Break: How Optimizing for Proxies Can Backfire

View post

Why We Use Separate Tech Stacks for Personalization and Experimentation

View post

Two Questions Every Experiment Should Answer

View post

The Feature Flag Toolbox: Cloud, Edge, and Local

View post

How Experimental Evidence Travels Through Your Organization: Why Better May Be Worse

View post

Beyond Winning: Spotify's Experiments with Learning Framework

View post
Spotify

Learn more

  • About Confidence
  • Read our blog
  • Take the bootcamp
  • See comparisons
  • Glossary
  • Why not Bayes?
  • RFP guides
  • Listen to us
  • Read our docs
  • Status page

Need help

  • Contact us

Legal

  • Terms of Service
  • Data Protection Agreement
  • Privacy Policy
  • Security

© 2026 Spotify