Experimentation & Growth Testing
Better tests, not more tests.
65–85% of A/B tests in under-structured programs are statistically underpowered. They can't produce a decision-quality result no matter what they show. We fix the program, not just the tests.
Most teams have a testing culture without a learning culture.
Growth teams run experiments. That part is true. Few teams examine whether those experiments generate evidence strong enough to act on. Tests end early when a positive trend appears. Sample sizes get set by gut feel. Success metrics get chosen because they're available, not because they connect to what the business is trying to learn.
The result is a program that reports wins without reliably producing them. Win rates look acceptable until you audit the methodology. Then it becomes clear: most of the "wins" would not survive a rigorous statistical review, and many of the decisions made from them were not supported by the data.
This is a program design problem. Running better tests comes down to what happens before the test runs and what criteria determine how results get interpreted.
Designed for evidence, not for velocity.
- Program audit We review your existing testing program: test selection, statistical methodology, win rate analysis, and whether results are informing decisions.
- Hypothesis development Each test starts with a specific, testable hypothesis linked to a mechanism. A weak hypothesis asks 'will this layout convert better?' A strong one states 'reducing cognitive load at this step will improve completion rate because...'
- Statistical design We set sample sizes using power analysis, choose primary and guardrail metrics before the test runs, and select the appropriate statistical method for the question being asked.
- Decision criteria We pre-specify what constitutes a win, a loss, and an inconclusive result, and what the team does in each case. Results without a decision rule aren't evidence. They're data.
- Results interpretation We review results in context: statistical significance alone is not sufficient. We assess effect size, confidence interval, and whether the result is robust enough to act on.
Who this is for
- Growth leaders at $5M–$150M companies Who run experiments but can't confidently say which decisions were supported by the data and which were made despite it.
- Heads of experimentation Who need to improve program quality, not velocity, and want analytical support to build that case internally.
- CMOs and VPs of marketing Who are investing in a testing program and want to ensure the methodology produces results worth acting on.
FAQ
We already have an experimentation platform. Is that still relevant?
The platform is rarely the problem. Statistical rigor, hypothesis quality, and decision criteria are program design questions. They apply regardless of which tool you use.
How do we know if our current program is working?
We look at win rates in context (most well-run programs show 20–30%), test replication rates, and most importantly whether decisions are actually being made from test results.
Do you work with internal data science teams?
Often. We complement analytical capability rather than replace it. Our role is usually the program design and decision framework layer, not the data infrastructure.
What's a minimum detectable effect, and why does it matter?
It's the smallest result your test is designed to detect. Setting it too low means your test will run for months. Setting it too high means you'll miss real but modest improvements. Getting it right starts with understanding the commercial value of the improvement.
Can you help us report experimentation results to leadership?
Yes. We often help teams frame results in commercial language, turning statistical outputs into business decisions. That's a separate skill from running the tests.