The two-stage check for metrics quality: Design Quality and Practical Quality
Stay up to date with the latest insights
Recently ChatGPT made me aware of something in my Success Metrics Workbook that wasn't obvious to me: How I split metrics quality!
As all of you, I use ChatGPT and Claude, too. They help me refine my workshops, sometimes even build up new ones from scratch. Create my talks. Create learning material. Improve my writing. And lots more.
I am in control of the input, so I share everything that's necessary, tell them what I expect and what good looks like, evaluate the output, and manually build or refine the final version.
So you can imagine that they know a lot about me and my work.
One day, I was working on my custom GPT the "Success Metrics Lighthouse" (you can find it here) when ChatGPT just told me:
"Something interesting I noticed. The workbook and the metric-quality PDF together suggest a two-stage evaluation model: Stage 1: Design Quality. Stage 2: Practical Quality."
Chatty is right, I see two types of quality levels that you have to be aware of and a specific order of how to assess them.
Let's have a look at what they are and how to assess them.

Stage 1: Design Quality
When you design or define a success metrics, check whether each of them is:
Correlated
Actionable
Sensitive
Comparative
Related

Have you ticket all boxes? Yes? Good.
Then check another point that is important for a well designed metric:
Is it clearly formulated?
This has two aspects:
Is it idiot proof?
Meaning: Is it extremely easy to understand, what it stands for, and protected against human error? I will contradict myself now by saying that you might want to add a short explanation because in my experience even the metrics that we categorise as idiot proof are not. We should have an explanation ready what the metric means, what it measures, how it measures it and how & from whom the data is collected.Is it short but powerful?
When the metric is too long, nobody will remember it. When it's too short, it typically leaves too much room for interpretation. Even "impressions" (yes, my love-hate relationship with LinkedIn continues) can mean different things. How many times was the post seen - by unique users or does repeat view count? Just as an example. My favourite example though is Weekly Active Users. For a scheduling SaaS "active" means something different to an online magazine. It's short and powerful when you add a definition to it (see the point "idiot proof"). You shouldn't call it "Number of users who create or edit or delete a poll or participate in a poll, or manage their contacts and settings, or create a scheduling link per week" to make it idiot proof. You should keep Weekly Active Users and define what "active" means. This way, people remember the metric.
Stage 2: Practical Quality
Alright, your metrics passed the standard check that is important for any metric definition. In the second step, you need to check if those metrics really focus on the situation that you have designed them for and are helpful in the way you intended.
Ask yourself:
Is it gameable?
A practical test: Can your metric easily improve without you achieving the goal / intention behind the metric? Yes? Well then, it's super gameable.Is it measuring the relevant part?
Does the metric capture the relevant behaviour or activity that creates the biggest impact? Not only but particularly when your metric measures an absolute number (which it shouldn't if you followed stage 1, or it should at least be a comparison of absolute numbers), you must ask for relevance or quality of the measured amount. Do you care if you have a lot of free users in your new product, or enough of the right free users who will eventually convert to the paid offer?Does it help make decisions?
In stage 1, you have checked already if the metric is actionable. Good. Here you check if the generally actionable metric can help you make decisions in this specific context. Is that metric applicable in this use case, or are you just carrying it with you only because you have defined it once?Does it work in the user's context?
Is the metric meaningful and realistically measurable for this specific user / customer group, business model, product, and use case? Take traffic for example. For a SaaS, more traffic on the sign up page means nothing if people don't sign up. But for an information website like a blog, traffic is very important. Therefore, it works in that context, and is measurable.Does it require segmentation?
I often see teams discussing about a what I call "schizophrenic" metric. In case A improving the metric means increasing it, in case B it means decreasing it. Or in case A it's relevant, in case B it's not. Or in case A it additionally needs another metric, in case B it doesn't, in case C it needs a different additional metric. That's a sign that you might need to talk about different segments of your target group. It might be demographic, it might be based on the job to be done, it might simply be new vs. repeat visitors or buyers, or even based on the hardware they use, or whatever is meaningful. Instead of dismissing the metric entirely, think if segmentation could help. If not: dismiss it.Does it need proxy metrics?
Is the intention really measurable through one specific metric or would do rather need proxy metrics that would provide a more accurate picture of the intention you want to measure? If it's not directly measurable, use proxies. When you use proxies, be explicit about what you are trying to measure with those proxies and where they could be misleading.
Conclusion
I like this split a lot. I wasn't aware that I was doing this and I'm thankful that Chatty surfaced it. For one because it's true that I do that. And for two because this approach mirrors how an experienced Product Manager thinks, and should be the standard of what a Product Leader excepts from any Product Manager in their team: Is this a well-designed metric, and is it actually useful in the real world?
Thank you Chatty, I'll add this to my teaching material 🫶
And thank you all for reading so far.
When you try it out and have questions, reach out and ask me anything. I read every email.
Product management insights, delivered to your inbox
Sign up for weekly product insights. No spam.