Before: Data Science as the Bottleneck
Every A/B test ran through our data science team. Engineers who wanted to test a change had no way to configure an experiment or read out results on their own, so every test meant filing a request and waiting on us to set it up.
That dependency surfaced three recurring pain points: experiment setup and result readout required our team end-to-end; the results view was dense enough that engineers, product, and business stakeholders struggled to read it and make a decision together; and teams often didn't learn until they were deep into a request whether they had the right dataset, the right metrics, or were even eligible to run a test in BaseLine.
The goal: let engineers without A/B expertise create and run experiments without our assistance. Experiments still need viable parameters and accurate KPI tracking, and results still need to be clear enough to actually drive a decision.
Guardrails caught invalid parameters, but some tests still needed a human check — ones that touched shared traffic or crossed into another team's goals. We added a lightweight approval step so a test's manager could sign off before launch, keeping that oversight in place without pulling our team back into the loop.