Writing

Proving It Holds Up: Turning Product Usage Into Statistically Defensible Impact Evidence

August 1, 2026  · #measurement#experimentation#data-leadership

Every product team has a dashboard that proves the product works. Usage is up. Engaged users do better than disengaged ones. The line goes in the right direction. It’s persuasive, it’s on the wall in the business review, and it is not evidence. It’s a correlation you’re rooting for — and a correlation you’re rooting for is the most seductive kind of wrong.

I’ve spent a lot of my career on the gap between “it looks like it works” and “it provably works,” because that gap is where renewals, budgets, and credibility are actually won or lost. Closing it is less about fancier models than about a discipline most teams skip: designing a measurement that could prove you wrong, and then running it anyway.

Why “usage correlates with outcomes” is a trap

Take the most common impact claim there is: people who use our product more get better results. It’s almost always true in the data. It’s also almost always unconvincing, because the people who use a product more are different — more motivated, better resourced, further along to begin with. The usage didn’t necessarily cause the outcome; the same thing that caused the usage may have caused the outcome too.

This is the confounding problem, and it’s fatal to the dashboard version of impact. Until you’ve dealt with it, “more usage, better outcomes” tells you these two things move together — not that one moves the other. Anyone who’s going to write you a check, or stake a budget on your product, eventually asks the question that breaks that chart: how do you know it wasn’t just the motivated users?

Rigorous evidence is designed up front

You don’t find rigorous evidence in your existing data by slicing it harder. You design for it before you collect anything, around a few non-negotiables:

  • A counterfactual. What would have happened anyway, without the product? Every real impact claim is a comparison to that baseline — a control group, a matched comparison, a pre/post with a credible untreated benchmark. No counterfactual, no causal claim.
  • Effect size, not just significance. “Statistically significant” only says the effect probably isn’t zero. The question that matters is how big, in units a decision-maker cares about — not a p-value.
  • Confounders named and handled, not hoped away. Motivation, prior performance, access — model them, match on them, or your effect is just them wearing a costume.
  • A willingness to find nothing. This is the real tell. If a test is built so that a positive result is the only thing it can produce, it can’t tell you whether the product worked — it can only confirm what you already assumed. Real rigor means designing the test so that “no effect” is a genuine possible answer, and reporting that answer if it’s what you get.

The strongest version of all this is handing it to someone with no stake in the result. Independent, third-party validation is the highest bar there is, precisely because they have no reason to be kind to your chart.

Built to be attacked

The most rigorous impact work I’ve led was an independent efficacy study of roughly 1,600 students, designed from the start to survive the scrutiny above and submitted to an outside evaluator against a formal evidence standard.

The central threat was the one from earlier: selection. Students who engage more with a learning platform tend to be more motivated and higher-achieving to begin with, so a raw usage-to-outcome correlation would mostly measure who those students already were. Handling that was most of the work:

  • Control for the starting point. Each student’s outcome was adjusted for their prior achievement, so the estimate reflected change from where they started rather than where they started.
  • Control for the classroom. Teacher fixed effects absorbed differences between instructors, so the effect wasn’t just good teachers who happened to have engaged classes.
  • Standardize against the right population. Scores were expressed relative to everyone who took the same assessment, not relative to a handful of classmates.
  • Write the plan down first. Hypotheses, the model, and the bar for significance were specified up front, so the analysis couldn’t quietly turn into a hunt for a positive result.
  • Try to break it. The finding had to hold across alternative model specifications and different subsamples before I’d trust it.

It held. The result was a statistically significant positive link between usage intensity and measurable learning gains — on the order of 13 percentile points tied to an additional daily question. But the number isn’t the point. The point is that it was built to be attacked and it didn’t fall over. That’s what lets you put a figure in front of a skeptical buyer, an outside evaluator, or a board and not flinch.

Markets that require formal evidence before anyone will spend — where buyers have to show what they purchased actually works — make this concrete. In those rooms a dashboard is table stakes, and a defensible study is what moves money.

The error that looks like nothing

Here’s how quietly a measurement can betray you. Deep in that study, outcome scores were being standardized within each teacher’s individual class — groups sometimes as small as a few students — instead of across everyone who took the same assessment. Everything still ran. The numbers still looked reasonable. But standardizing against three classmates instead of a whole population makes the units unstable and the effect sizes non-comparable: a strong-looking result in one class and a weak one in another could be identical underneath, or reversed.

Nothing in the output announced the problem. You only catch it if you go looking — if you treat “a number came out” as the start of the checking, not the end. Correcting it was what separated an effect size that would have collapsed under review from one that held. That’s the whole discipline in miniature: the code ran fine, the measurement was quietly wrong, and only grounding it correctly made the result mean anything.

Why this wins renewals

Here’s the business version, plainly: buyers don’t renew on dashboards. They renew on proof.

Predictive analytics makes this sharper, not softer. I’ve built model-driven features that score and explain outcomes — the kind of predictive work that looks impressive in a demo. But a prediction is a claim about reality, and a claim about reality is only worth what your measurement can back. Predictive analytics without measurement rigor is just a confident guess with a nicer chart, and sophisticated buyers can smell the difference. The features that contributed to a multi-million-dollar renewal earned it because the impact underneath them was defensible — not because the visualization was slick.

The fundamentals don’t change

This connects to a conviction I keep coming back to: the tools change, the fundamentals don’t. Better models, LLMs, more data — none of it excuses you from the basic question of whether the thing actually works, and none of it answers that question for you. If anything, the easier it gets to generate a confident-looking result, the more the discipline of measuring it against something real is what separates work you can defend from work that just looks good.

There’s a real difference between a result that looks good in a review and one that holds up when someone with every incentive to doubt you is in the room. The entire job is producing the second kind — and being honest enough to run the test that could have said otherwise.


← All writing