One practical way to begin scaling an experimentation program is to start with a small set of decision-relevant tests, run them cleanly, and show leadership both the evidence and its limitations. This is a sequence I have used; it is not a guarantee that every organization will prove financial value in 90 days.

The teams that fail to scale do the opposite. They try to run everything at once, produce a flurry of activity with unclear business impact, and then wonder why leadership is reluctant to give them more budget, more headcount, or better tools.

Proving value before you scale is not caution. It is leverage.

"Building an experimentation program is really about initially getting some quick winnings and running clean tests — so we can prove that experimentation has a dollar value attached to the stuff we're doing."
— Atticus Li

Why Scaling Without Proof Fails

When a program scales before defining how value will be measured, several risks rise: uneven idea quality, more work in progress, higher operating cost, and a portfolio that is hard to explain. The outcome is not predetermined, but the team has less evidence for deciding where to invest.

Eventually, someone in finance looks at the spend and says "what are we getting for this?" And you do not have a clean answer because the program's value was never framed in dollars from day one. At that point, budget arguments become defensive instead of offensive, and the program stalls.

The inverse path is easier to audit. Start with a small, qualified set of tests. For each one, separate the observed test-window result, the modeled annual impact, and any revenue or contribution actually recognized after launch. Build a scorecard that preserves those evidence classes instead of combining them into one “generated revenue” total.

The Pre-Test Revenue Calculation (in Detail)

This is the single most important habit I teach experimentation leads. Before a test is approved to run, calculate its expected revenue impact.

The math is simple:

Inputs you need:

  • Baseline conversion rate on the primary metric
  • Weekly sessions on the page being tested
  • Revenue per converted user (or LTV if appropriate)
  • Minimum detectable effect you are powering for (MDE)

The calculation:

Weekly baseline conversions = Weekly sessions × Baseline conversion rate

Weekly baseline revenue from this page = Weekly baseline conversions × Revenue per conversion

If the variant wins at the MDE you are powering for:

Weekly lift in revenue = Weekly baseline revenue × MDE

Annualized projected impact = Weekly lift × 52

Now you have a single number: "this experiment is projected to generate $X in annualized revenue if it wins at the minimum detectable effect." That number is your executive translation.

The Post-Test Impact Update

The post-test version is equally important. After the test concludes, report the observed test-window effect and update the modeled impact range. That is not yet the same as realized revenue.

  • If the result is negative or inconclusive under the precommitted rule, report that decision outcome without booking a prevented loss as revenue.
  • If the result is positive, replace the planning MDE with the observed effect and its uncertainty.
  • Apply eligible traffic, contribution per conversion, adoption, and persistence assumptions to produce a range rather than a single guaranteed total.

The key discipline is to not pretend. If a model projected $440K and post-launch finance reporting later recognized $180K, preserve both numbers with their dates and definitions. Credibility depends on reconciliation, not on choosing the larger figure.

For material changes, define a persistence check before launch. A holdout may be appropriate when the value of better evidence justifies the traffic and operational cost; otherwise use a dated post-launch review. Novelty, seasonality, and later product changes can all alter the test-window effect.

"Using revenue per customer, we can calculate — before the test runs — the projected value of this test based on MDE. After the test runs, we look at the actual stats, the actual lift, how much it generates during the test, and how much it's going to over a year."
— Atticus Li

Where Experimentation Strengthens the ROI Argument

Here is a conversation I have had with CFOs more than once. The finance team is comparing experimentation to brand marketing and asking why they should fund more of the former.

The answer is not that brand marketing does not work. It does. But the ROI of brand marketing comes with a lot of asterisks — brand lift studies, estimated impression counts, multi-touch attribution assumptions, share-of-voice models. Every one of those numbers has a range of uncertainty around it.

Experimentation still has uncertainty, but randomization can reduce some attribution problems. A valid test estimates the effect for the exposed population, period, and metric. Extending that estimate to a full year or the whole customer base adds modeling assumptions that should remain visible.

This is not an argument against brand marketing. Brand lift studies, geo experiments, marketing-mix models, and randomized digital tests answer different questions. The useful comparison is which design fits the decision—not which function always “wins.”

"A/B testing has scientific rigor. We know exactly what was lifted, what changes were made. There's way less noise compared to brand marketing — where companies spend a lot of time and money, but can't really get a very clear answer on the impact."
— Atticus Li

The First 90 Days

If you are scaling an experimentation program from scratch, here is how I would structure the first quarter:

Weeks 1-2: Choose your proving-ground tests.

Pick a small set of high-traffic, decision-relevant opportunities that can meet a precomputed sample requirement in the available window. Avoid selecting only ideas you expect to win; the goal is clean evidence, not a curated success rate.

Weeks 3-4: Instrumentation audit.

Before running anything, validate your tracking. Run A/A tests. Check for sample ratio mismatch. Audit the data pipeline from event fire to dashboard. You cannot report dollar impact if your instrumentation is broken.

Weeks 5-10: Run the tests.

With pre-test revenue calculations attached to each one. Document every hypothesis, every projected impact, every assumption. Treat this phase as the foundation for the narrative you will tell later.

Weeks 11-12: Build the scorecard.

One page. Top section: observed test-window outcomes. Next: modeled impact ranges with assumptions. Next: recognized post-launch outcomes, if available. Then: what the team learned and what additional capacity would enable.

Week 13: Present to leadership.

Not to report. To ask for the next level of investment, with a specific plan for what you would do with it.

The Standardization Advantage

One reason proving value works is that it forces the team to standardize early. When you are calculating projected and realized impact for every test, you have to agree on:

  • How baseline conversion is measured
  • How revenue per conversion is attributed
  • Which MDE is acceptable for which types of tests
  • How long the post-test holdout runs
  • Who signs off on the final realized-value number

These decisions become your program's standard operating procedure. And because they are tied to reportable dollar value from day one, they do not get skipped. Standards imposed after a program is already scaling are almost impossible to enforce. Standards built into the reporting from the start survive.

FAQ

What if the first few tests lose?

That can still be valuable if the decision was worth the implementation and traffic cost. Report the cost and the decision made. Do not book the modeled downside of a change you did not ship as realized revenue or “savings.”

How conservative should pre-test projections be?

Use a range built from explicit assumptions and track forecast error over time. There is no universal acceptable multiplier. Calibration means comparing prior ranges with later evidence and updating the model.

What if I do not have a clean revenue-per-conversion number?

Work with finance to get one. If it is unavailable, begin with an explicitly provisional range and make better unit economics a shared measurement task. Do not invent precision to fill the gap.

How do you handle tests whose value is indirect?

For tests that affect engagement, retention, or activation, use a two-step translation. Step one: calculate the lift in the proximal metric. Step two: apply a learned or assumed conversion rate from that metric to revenue. Document the assumption. Leadership will push back on it, and that conversation is healthy.

Turn Your Program Into a Revenue Function

If your experimentation program feels like a cost center, the problem is almost always framing. You have the data. You have the results. What is missing is the discipline to translate every test into the language leadership uses to allocate capital.

I built GrowthLayer with pre-test and post-test impact fields as a first-class feature. The point is to preserve the evidence label—planning scenario, test-window estimate, or recognized outcome—so executive reporting does not collapse unlike numbers into one total.

If you are hiring for experimentation leaders who know how to frame programs in revenue terms, or looking to build those skills, explore open roles on Jobsolv.

Or book a consultation and I will help you design a 90-day proving-ground plan for your program.

Share this article
LinkedIn (opens in new tab)X / Twitter (opens in new tab)
Atticus Li

Experimentation and growth leader. CXL-certified CRO practitioner, Mindworx-certified in behavioral economics. Led 100+ in-house experiments at NRG in 2025, with project evidence and limits documented in the case studies.