Better Experiments with LLM Evals — A funnel, not a fork | Spotify Engineering (opens in new tab)
LLM evaluations and A/B experiments serve different purposes and should work as a funnel rather than alternatives. Evals verify whether a change meets qualitative standards and filter out weak ideas, while experiments validate whether it improves real user and business outcomes. Their greatest value comes from a feedback loop that continuously calibrates eval scores against online results.
Evals as Verification, Experiments as Validation
- LLM judges can assess relevance, coherence, tone, intent alignment, and other qualities at scale.
- Evals identify promising candidates before they consume experiment capacity.
- Experiments determine whether users actually respond better to the change.
- An eval can uncover unexpected failure patterns, guide product fixes, and later verify that violations have decreased.
- It cannot establish whether those fixes improve long-term outcomes such as engagement, trust, or retention.
The Importance of Guardrails
- Evals typically measure selected quality dimensions, not every possible business risk.
- Spotify reports that roughly 42% of launched experiments are rolled back because of regressions in secondary metrics such as crashes, session length, or retention.
- Online experiments expose production effects and monitor guardrail metrics that offline evals may miss.
- Although only about 12% of Spotify’s A/B tests produce a shipped positive result, approximately 64% generate useful learning.
Calibrating Evals Against Outcomes
- Evals are proxies, so their scores must be validated against the outcomes they are intended to predict.
- LLM judges introduce an additional calibration layer on top of conventional metrics such as precision, recall, and ranking scores.
- Judges can reward superficial patterns or miss improvements that matter in real-world use. For example, Qodo’s coding evals did not detect improvements in Anthropic’s Opus 4.5 on longer tasks.
- Both judges and quantitative metrics can drift, making offline-to-online comparison essential.
- Continuous calibration clarifies when evals are reliable verification tools and whether they may eventually become strong enough to support validation.
Closing the Evaluation Loop
- Run evals early and frequently to select the strongest treatments.
- Use experiments to confirm user impact, assess business outcomes, and detect unintended regressions.
- Apply the same LLM evals to A/B test data to compare qualitative scores with actual user results.
- A mismatch between eval scores and experiment outcomes is valuable diagnostic information:
- Improved scores and improved outcomes validate the judge.
- Improved scores without better outcomes show that the judge is measuring something that may not drive value.
- Use lightweight experiments for iteration and learning, and more rigorous tests for launch decisions.
The practical recommendation is to treat LLM evals as an upstream filter and a downstream calibration tool—not a replacement for experimentation. This funnel improves test efficiency while making both the judges and product decisions more trustworthy over time.