Skip to content
  • Pricing
  • Blog
  • Bootcamp
  • Contact us
  • Login
Start free trial

The Experimentation RFP Series

Feature lists tend to oversimplify experimentation offerings. Not all implementations of the same feature are alike, and a vague RFP leads to frustration and friction when the platform does not deliver what the checkmark promised. Here we describe how we would specify an RFP for various capabilities, and what to look for to ensure the implementation is worth committing to.

Features Confidence has built

Topics where we have a connected implementation and can show what the connected version looks like.

We built this
Experimentation RFP: Sample Size

Most platforms have a sample size calculator. Almost none connect it to the analysis the experiment will actually use. Here is what to ask vendors instead.

Read more
We built this
Experimentation RFP: Sequential Testing

Every platform claims to support sequential testing. The claim is almost always incomplete. Here is what to ask vendors beyond the feature checklist.

Read more
We built this
Experimentation RFP: Multi-Metric Decisions

Every platform lets you add multiple metrics. Most will display results for each one. What they rarely tell you is what to do next. Here is what to ask.

Read more
We built this
Experimentation RFP: Multiple Testing

Every platform says it corrects for multiple comparisons. Most do, partially. Here is what to ask instead of "do you correct for multiple testing?"

Read more
We built this
Experimentation RFP: Variance Reduction

Every platform offering variance reduction claims 20-50% runtime cuts. What is missing is how far the reduction reaches. Here is what to ask vendors.

Read more
We built this
Experimentation RFP: Ratio Metrics

Revenue per user. Streams per session. Most platforms support ratio metrics, and most get the variance wrong. Here is what to ask vendors instead.

Read more
We built this
Experimentation RFP: Fixed-Power Designs

Every platform estimates sample size before an experiment starts. Almost none revisit it after. Here is what to ask about in-flight power monitoring.

Read more
We built this
Experimentation RFP: Time-in Metrics

Every platform measures user behavior after exposure. The question most buyers never ask is: over what time period? Here is what to ask about observation windows.

Read more
We built this
Experimentation RFP: Monitoring & Alerting

Every platform lets you look at results. The question is whether the platform looks for you. Here is what to ask about monitoring and alerting.

Read more
We built this
Experimentation RFP: Clustered Randomization

Most platforms let you randomize by account or store. The question is what happens in the analysis after randomization. Here is what to ask.

Read more
We built this
Experimentation RFP: Metric Zero-Handling

When a user generates zero events, the platform makes a choice that changes what the experiment measures. Most vendors do not document which choice they make.

Read more
We built this
Experimentation RFP: Exploratory Analysis

Every platform lets you slice by dimension. The question is whether the platform controls the false positive rate when you do.

Read more
We built this
Experimentation RFP: Experiment Review

Every platform lets you configure an experiment. Few require structured design review before launch. Here is what to ask vendors about review workflows.

Read more
We built this
Experimentation RFP: Experiment Coordination

Most experiments can run overlapping. For the ones that cannot, precise control matters. Here is what to ask beyond "do you support mutual exclusion?"

Read more

Features we chose not to build

Topics where we deliberately chose not to ship the feature, and what to look for if you need it.

We chose not to
Experimentation RFP: Percentile Metrics

Most vendors support percentile metrics. Most implementations break down exactly when you need them. Here is what to ask instead of "do you support it?"

Read more
We chose not to
Experimentation RFP: Geo-Lift & Synthetic Control

Most experimentation programs do not need geo-lift. If yours does, ask whether the platform confronts the assumptions that make or break the analysis.

Read more
We chose not to
Experimentation RFP: Switchback Experiments

Switchback experiments exist because standard A/B tests break down when users interact with each other. Only two vendors offer support. Here is what to ask.

Read more
We chose not to
Experimentation RFP: Bayesian Inference

The label Bayesian does not tell you what stopping rule is used. Here is what to ask instead.

Read more
Spotify

Learn more

  • Read our blog
  • Take the bootcamp
  • See comparisons
  • Glossary
  • RFP guides
  • Listen to us
  • Read our docs
  • Status page

Need help

  • Contact us

Legal

  • Terms of Service
  • Data Protection Agreement
  • Privacy Policy

© 2026 Spotify

The Confidence name and logo are registered trademarks of Spotify.