I’ve touched on different topics since coming on Substack, but I’ve yet to dive into what makes experiments so exciting.
As I embark on my journey of becoming a freelancer, I thought the 12 years of ups and downs can be summarized here.
What does Product Analytics do to differentiate itself?
I’m glad you (didn’t) ask!
Most skills used are typical and used across many different roles: Marketing Analyst, Systems Analyst, Business Analyst. You can already tell what they differ which is the type of data that they work with. Product Analyst is what has been my favourite part. Over the years, it was called Data Scientist, Product Data Scientist, eventually landing in Product Analytics.
We’re particularly obsessed with experiments and AB test design. And the reason is because it’s the most scientific approach to understand whether features we deliver have any impact.
I remember being in a job where they called alternatives such as pre-post analysis, which is literally drawing a line in the time-series chart, and just saying “well, we released this at this point, as such, that uptick is due to our release”.
But it’s not enough. If you’re running a global company, it’s even worse. Events across all over the world can impact drastically how we read anything.
You might wonder: but AB test designs are easy. Feature on/feature off. Why is this … needed?
I’m glad you (didn’t) ask, again!
In my years of doing this, the experiments became as complex as the product and context became. You might get interesting data from one experiment, but it’s all the more interesting, to get data from 2 experiments combined.
It might be interesting to design one on/off experiment, but equally, I might want to see the impact of many moving pieces.
But this experiment has 200 different additions
A typical conversation with Product Managers goes like:
Analyst: You can’t have 200 additions.
Product Manager: Okay, let’s test each addition separately
Analyst: Well, we don’t have enough data to measure them individually
Product Manager: Okay, then let’s add them together
Analyst: Okay, but keep in mind we won’t know which one will impact the results
Product Manager: Okay, what can we do?
Analyst: We can run individual variants separately, for a long time?
Product Manager: How long?
Analyst: … 3 months?
Product Manager: So, I have nothing to show to management this quarter
Analyst: yeah….
I know that pain. Because I’ve lived it for years. This isn’t to say we hit a dead end. On the contrary. Each side has to give and take.
It’s a fun process determining what makes sense to test and what not. It’s also fun to break a product into milestones, in order to understand if it’s bringing the gain we hoped it did.
Experiments are not just for the sake of confirmation.
We don’t do experiments to confirm what we know. We do them to understand what works for our users, and how we can use these insights to move forward, or modify the roadmap.
This job comes with a lot of disappointed faces.
Rightfully so.
The case for an experimentation culture is an emotionally grueling piece of convincing. Not only do you have to slow down releases, by adding an experiment at the end of the development work, but then only ~34% of experiments produce a statistically significant & positive result.
But it doesn’t take away the many successful experiments. The famous 41 shades of blue from Google estimates that it added around $200M annually.
Most companies don’t have enough users to split their population into 41 (or 42 if you consider the control group), but it points to the idea that the possibilities of how to do experiments is endless.
So Miruna, how do we solve this?
As with many things in life, we can solve this with a pizza.
Meaning: you look at your user base, and understand whom of them is impacted. As such, how many pizza slices will be eaten.
There are a few golden rules to understand the design of an experiment:
What is the change done? (or many changes…)
What is the randomization unit? (Let’s say in this case we’re speaking of users)
What is your success metric (aligned with the randomization unit)
How many datapoints are needed in order to determine an increase of x% ? (so called, Minimum detectable effect)
If your user base impacted turns out to be just a slice of pizza, you might have to wait a little bit longer to make sense of your results.
In broad strokes, that’s the template we use for any experiment, after which negotiations start.
Will there be notifications for the user to be aware of it?
Will there be marketing budget to justify the awareness?
Will there be cannibalization from other products that are important to the business?
And so on and so forth.
I still remember to this day the experiment that had to be killed after 2 years of back and forth. It was a feature where the group of users who were engaging with it, were very keen and eager ones, but not a large amount.
Their engagement metrics were off the charts, but for the overall numbers, they came flat.
A large part of the issues we dealt with weren’t even statistical, but rather, marketing budget. On one side, marketing didn’t want to invest, because the numbers were laying flat, but on our side, the numbers were lying flat, because marketing wouldn’t invest.
It’s the beautiful catch-22: if users don’t adopt the feature, what can we even say about the success of it?
I thought you’ll tell me how to run experiments
As you can tell, I love doing experiments. I do them in my day to day. Heck, people have changed their entire way of life, after reading Tiny experiments. It made my boring Statistics courses from uni to be something I go on to obsess over for 12 years.
But more practical, I love helping people put order into their 20 ideas that they don’t know how to start.
Experiments are a great tool to organise those 20 ideas. And also to bring numbers and impact on what you’re trying to build.
Over the past few months, I keep hitting the same hopeful dream from many entrepreneurs: “I have so many ideas, I don’t know which one to start”.
When people spend their time waiting to see which idea does the gut scream loudest to, I spend time in systems, in such a way that …everything can be done.
Sometimes, the result is in between: While designing the systems/experiments, it becomes clear that in reality, some ideas don’t scream loud enough to be neatly organised into a roadmap.
But that clarity is closer than you think.



