Skip to content
  • Pricing
  • Success stories
  • Blog
  • Bootcamp
  • Contact us
  • Login
Start free trial
January 7, 2026· Updated May 12, 2026/Yu Zhao, Staff Machine Learning Engineer, Mårten Schultzberg, Staff Data Scientist

Why We Use Separate Tech Stacks for Personalization and Experimentation

Why Spotify runs personalization and experimentation on separate tech stacks: ML systems need their own infrastructure, and A/B tests evaluate them.

Diagram showing separate tech stacks for personalization and experimentation

Want to experiment like Spotify? Sign up for a 30 day free trial.

Start your free trial

TL;DR

At Spotify, we build personalization systems using our ML stack and evaluate them through our experimentation stack. Each tech stack does what it's good at.

Personalization systems have strict infrastructure requirements: access to diverse model types (neural networks, boosting, bandits), rich feature sets, low-latency inference, and real-time data collection. These don't fit naturally inside an experimentation tool. And even if you use a contextual bandit, you still need to evaluate that bandit as a system through A/B tests on different bandit versions. When A/B tests and multi-armed bandits live in the same tool, you get confusing dependencies between instances of the same tool.

Keeping a clean separation of concerns helps us scale with less friction for product teams.

Read the full post on Spotify Engineering: Why We Use Separate Tech Stacks for Personalization and Experimentation

PreviousWhen Proxy Metrics Break: How Optimizing for Proxies Can BackfireNextTwo Questions Every Experiment Should Answer
Spotify

Learn more

  • Read our blog
  • Take the bootcamp
  • See comparisons
  • Glossary
  • RFP guides
  • Listen to us
  • Read our docs
  • Status page

Need help

  • Contact us

Legal

  • Terms of Service
  • Data Protection Agreement
  • Privacy Policy

© 2026 Spotify

The Confidence name and logo are registered trademarks of Spotify.