Why marketers should focus on business impact, not just statistical significance
- Tim Hefner

- Jul 15
- 6 min read

I know I’m going to make a lot of people mad, but hey, I might also make some new friends: I hate statistical significance.
It’s not the math. The math is good; the math is great. The smart people who came up with it did so for good reason. What I hate about statistical significance is how I see it used in marketing. At some point, statistical significance began to feel like the gold standard, a substitute for judgment, and sometimes it even becomes the point where people stop thinking and move on to the next box to check.
If something is statistically significant, some marketers act like the case is closed, the winner has been crowned, roll it out, and pop that bottle of wine we’ve been saving. But that’s not really what happened.
At its core, statistical significance tells you something actually pretty simple: The difference you observed in a test probably wasn’t random chance. This is useful, but it doesn’t tell the whole story. Will it happen again? Will it hold at scale? Does it even matter to the business? Is this result a universal, undeniable truth? But that’s how it’s often used, and that’s where I become the old man shaking his fist at the clouds.
It happens all the time. You run a test, and version B beats version A. The read … well, it’s statistically significant, hooray! Everyone circles the winner, updates the deck, and moves on.
But before we celebrate a 0.03% lift, how about asking this: “Did it actually matter?” Because frankly, something can be statistically significant while also being completely unimportant.
If response goes from 0.30% to 0.33%, that might be real, but if it doesn’t change volume, cost, revenue, or ROI in a meaningful way … did it actually matter? Too often, statistical significance is shorthand for good, and those two things just aren’t the same.
I have a friend who’s far smarter than I am (and a much better mountain biker for what it’s worth) who once said to me: “Do you want me to tell you it’s statistically significant, or do you want me to tell you that it made us money?” That has always stuck with me because that’s the decision we’re actually trying to make.
Let me give you a credit card example because that’s the world that I live in. Let’s say you’re running a direct mail campaign looking to acquire new cardholders. You test:
Offer A – $200 bonus with the new account opening
Offer B – $250 bonus with the new account opening
Your results show that offer B outperformed offer A 0.36% to 0.30%. That’s a 0.06 percentage point improvement and a 20% lift! And guess what? Statistically significant! Before you shout, “Offer B wins; let’s roll it out!” let’s pump the brakes for a moment. Let’s say that lift equates to 60 more accounts per 100,000 pieces mailed, nice job! But … now you’re paying $50 more for each of those accounts, so the question becomes “Did those 60 accounts generate more than $3,000 in incremental value?” Sometimes the answer is yes, but often it’s no (it’s even more often no when we’re talking about checking account acquisition instead of credit card acquisition). In this case, statistical significance did what it’s supposed to do, it told you that the lift was real, and in that specific iteration of the test, likely not random. What it didn’t tell you is whether adding $50 to the incentive was worth it to the business.
You see similar things in other test scenarios. Same approach as the credit card example, but this time you’re testing imagery instead of offer.
Creative A – clean, product forward
Creative B – a smiling couple, holding an adorable golden retriever puppy, sitting on the front porch of their brand-
new house
The results: B wins again, this time with a 0.04 percentage point improvement and a 10% lift over A … and still statistically significant. The conclusion is that lifestyle imagery is the future, more cute puppies. But what was actually learned here? It’s not that puppies sell mortgages; it’s that this creative worked slightly better this time. This is merely a clue, an insight; it’s not a rollout strategy.
To be fair, those in the pro-statistical significance camp who are currently lighting their torches and picking up their pitch forks have a valid point. Without some rigor, marketing teams will absolutely talk themselves into some nonsense decisions. If you remove this rigor completely, random bumps become an insight and every opinion in the room becomes a new strategy. Statistical significance helps separate signals from noise, and that matters. It’s significantly (see what I did there?) important in direct response marketing where response rates are often small and random variations can make mediocre ideas look brilliant.
I’ll even make some arguments for the pro-stat sig mob:
“Confidence is already included in the framework and provides us with a range of likely outcomes!”
That’s right, they are, and the range of likely outcomes is actually far more useful than the thumbs-up/thumbs-down vote from statistical significance. But let’s be honest, they’re not often used; tests typically end with a binary conclusion of significant or not significant.
I’ll go one step further. Let’s say your test shows a 0.08 percentage point lift and it’s statistically significant. What’s missing is that range, and that range is 0.01 to 0.15 percentage points. The takeaway from this test shouldn’t be “the test was statistically significant”; it should be “this looks positive, but the exact impact isn’t known.”
“We’re experienced marketers. Of course we’re not going to roll out after one test; we’re obviously going to test it again!”
Great, see we found our common ground, something very rare these days. A repeated test can strengthen confidence that a result isn’t random, and it can help gain traction that something actually is there.
But it doesn’t uncover the durable truth about what works. Maybe you’ve validated a specific execution, a specific image, or a specific layout, and those are all useful learnings, but they’re not “we now know this approach works.”
In other words, retesting doesn’t give stat sig magical powers; it just means you’re using it as part of the learning process or how it was intended to be used.
So no, I’m not arguing that we should throw statistical significance out the window. I’m arguing that we should stop asking it to do more than it was ever intended to do.
OK, smart guy, so what do you propose we do differently? Not just in theory, but the next time we’re sitting in a test readout.
Well, instead of treating statistical significance as the final answer, treat it as one of many inputs into a business decision that should also include:
Practical significance – Not just lift, but is that lift meaningful? Does it actually matter?
Uncertainty – Not just whether the results were able to cross the stat sig bridge guarded by the troll, but how stable those results actually look.
Repeatability – One test is a data point; a pattern across tests is a signal. If the test works repeatedly across campaigns, audiences, and time periods, well, I think we’re getting somewhere now.
Economics – At the end of the day, this is what matters most for me. Are acquisition costs lower? Can we see an impact on LTV? Have our retention rates improved? These are all revenue-driving questions. No executive is going to pop into a room and ask, “Was it statistically significant?” They’re going to ask, “Did we make any money?”
If you’re a CMO or a marketing lead sitting in a test readout, here are a few simple rules of thumb:
Don’t stop at “statistical significance.”
Always ask: “So what?”
Look for patterns in tests, not one-off wins.
Make sure everything ladders up to business goals and objectives.
These aren’t just nice ideas to keep in mind, they’re also what should shape your testing approach. Build a testing agenda where each test is grounded in a clear hypothesis that’s not “We think the improvement of test B will be statistically significant compared to test A.” Develop a back test or retest approach so you can validate insights and identify what really matters to the business. You want to know whether a test practically matters to your business, has stability, can be repeated, and ultimately increases revenue or value.
OK, I’ve talked myself off the ledge. I don’t actually hate statistical significance, and I’m sorry for the friends I may have lost who read the first line and dropped off. I hate how statistical significance is used. I hate how it has become a shortcut. I hate that it gets confused with business value. And I hate when one test is statistically significant, a broad strategy is then built from it.
Stat sig is useful, but it’s not magic. Sure, it will tell you if results are real, but it can’t tell you if it actually matters, whether it will happen again, or if you should bet your budget on it. That part requires judgment, real, human judgment. And, unfortunately, there’s no p-value for real human judgment.
By: Tim Hefner, Senior Director, Strategy
At Pragmatic, we help our clients stop chasing statistically significant wins and start building testing programs that actually drive growth. Our approach is always grounded in real hypotheses, repeatability, and the impact on your business. If you’re ready to turn test results into better business decisions, let’s connect, let’s connect.

Comments