We have several AI features launching soon in the Statsig console, and more in development, so I wanted to share a bit about our approach to AI for experimentation.
I think it’s important to set the stage because there are a lot of differing views on AI out there. Many of our experiment platform peers are adding AI everywhere, supposedly to make it so you can test faster. Some data scientists we talk to are skeptical of AI anywhere.
High-level, we’re focusing not only on breadth and “speed,” but also impact. We want to add AI in the most impactful places, making sure you can trust its output, so that you can learn better and learn accurately.
If you’re incorporating AI into a product for a highly-skilled, technical audience who are AI-skeptical, we hope this could be interesting for your own strategy too.
From a user-experience perspective, we think of AI as a natural extension of the self-serve ethos that makes Statsig so powerful.
Say you need to analyze experiment results. Statsig makes that easier: the results are accessible by anyone on your team, so you can dig into the data yourself and share what you find with cross-functional partners. Plus, your data scientist gets time back to focus on the most impactful analysis.
But we know processes like that still have friction. Collecting results for a report can turn into an Easter egg hunt across experiments and dashboards. Your data scientist may still have to help pick out the signal from the noise. And for your cross-functional partners, it may take a little extra manual effort to make the impact clear.
From implementing AI in Amplitude, we’ve seen firsthand that AI can drastically reduce this friction. It can bring the data together for you. Draw connections between disparate signals. Explain answers in a way that gives your data scientist back the time you borrowed.
AI helps fulfill the promise of self-serve experimentation. But it can only have that impact if it’s implemented correctly.
We’ve learned from talking with our data scientist customers, the very heavy users of Statsig, that you don’t always trust AI. You’re very wary of hallucinations. You want to be able to trace and support every step of causal reasoning. You have solid ways of working, and you don’t want AI to mess things up.
For AI to be impactful, it has to be trustworthy, just like the rest of the Statsig platform. We’re approaching that from three directions: rigorous process to mitigate hallucinations, control over autonomy, and not putting AI where it doesn’t belong.
Incorrect results lead to wrong decisions, and then you’re screwed. To avoid hallucinations, we build in extensive guardrails so AI features only look at real data. We also evaluate and train AI features’ analysis abilities, so their insights are grounded in that data instead of an LLM’s natural flair for the dramatic.
But yes, there’s no way to avoid hallucinations entirely. Our AI features will always be transparent, so you can read through their results and verify yourself if things went right or wrong. “Prove it” are two of the strongest words for data scientists; we’ll make sure you can.
Currently, AI in the Statsig console has a tight scope of action that’s very close to the user’s control. Like with the AI hypothesis advisor: you write a hypothesis, you click a button, and it outputs what, if anything, you should change.
As we develop agentic features with larger scopes of action, we’re making sure control stays strong by building customizable autonomy levels. Similar to role controls, you’ll be able to say, “Don’t do anything,” or, “Do anything for these certain circumstances,” or maybe even, “Automatically do everything, knock yourself out.” But the default will always be conservative.
In one of my testing sessions for an AI feature, the AI recommended a next step to take based on the experiment results. My tester responded, “What is this Next Step section? I would never use this.” My answer: “Cool, you don’t have to! These are just suggestions.”
Part of trust is us extending trust to you, that you can use Statsig the way you want to. You don’t have to use AI if you don’t want to. We’re rolling out AI so it can only augment, and never break, your workflows. And for newer users, we hope that AI will help you get started with the best workflows more easily.
The question of impact is one we actively evaluate for any feature we develop. For AI, we focus on where its strengths lie: synthesizing information, speeding up slow processes, and improving access for groups and less-technical users.
Something like creating an experiment shell with AI—that’s pretty cool, but honestly, creating an experiment is one of the easiest parts of using Statsig. You click a few buttons. It’s done. It takes like five minutes.
Where we see a bigger impact is with things like AI building the code, analyzing the results, or prioritizing the next ideas to test. Tasks like these often still require engineering and data science teams’ time and effort, regardless of the payoff. AI can reduce that burden when it isn’t as necessary.
And as we implement AI into deliberate parts of the experiment cycle, we’ll also be building it out to help with your whole experimentation program. Back in the day, an old team of mine had to track our whole program in Confluence. To answer top-line questions from execs, we’d have to dig through hundreds of experiments manually. Those answers live in Statsig now; AI can help find them.
By pursuing AI with a focus on impact and trustworthy results, our overall aim is the same as we’ve always had: to help improve the quality of your experiments so you can learn better.
If you have an experimentation program, you’re already on that journey to better learning. But it’s very easy to get caught up in the cycle of test and ship, test and ship; gradually, you lose momentum toward improving how you test.
We’re building out AI in the Statsig console, so it’s easier not just to run experiments, but to run better experiments. We’re excited for you to try it out soon.