Atticus Li designed an internal NRG Energy model that translates test-window evidence into assumption-labeled financial estimates. In 2025, NRG's internal program reporting recorded $30M+ in impact across 100+ experiments. That is a historical company readout, not booked EBITDA, an externally audited result, or a forecast for another company.
The Language Problem That Kills Experimentation Programs
Every experimentation team I've seen struggle has the same root cause. It's not bad ideas. It's not insufficient traffic. It's not the wrong tools.
It's that they can't speak finance's language.
Here's what a typical experimentation report looks like: "We ran a test on the homepage hero banner. The variant achieved a 4.2% conversion rate versus the control's 3.8%. This is a 10.5% relative lift, statistically significant at 95% confidence."
And here's what the CFO thinks when they read that: "So what?"
The CFO doesn't care about relative lift. They don't care about confidence levels. They care about one thing: how does this affect the P&L? And if you can't answer that question in dollars, your experimentation program will always be fighting for budget, fighting for headcount, and fighting for survival.
I learned this lesson at NRG. When I arrived and started building the experimentation program, the first tests I ran were technically solid. Good hypotheses, clean execution, valid statistical methods. But when I brought results to leadership, the reaction was polite indifference. Nice work. Keep it up. But no additional budget, no additional tools, no additional team.
That changed when I built the EBITDA impact model.
The Formula
The model translates test results into projected financial impact using inputs that finance already understands:
Incremental conversions = Eligible annual traffic × Baseline conversion rate × Observed relative effect
Projected contribution = Incremental conversions × Contribution per conversion × Adoption × Persistence adjustment
Let me break each component down.
Contribution per conversion is the finance-approved economic value of an incremental conversion. Depending on the business, that may be contribution margin, expected customer value, or another unit-economics input. It should not be substituted with total brand EBITDA.
Eligible annual traffic is the population expected to see the shipped change, not automatically all company traffic. Annualizing a test-window result is a model, so the traffic forecast and rollout percentage must be visible.
Baseline Conversion Rate is the control group's conversion rate during the test period. This is the starting point that the lift is measured against.
Observed relative effect is the test-window estimate, shown with its uncertainty rather than treated as a perfectly known lift.
Adoption and persistence adjustments account for rollout coverage and the possibility that an effect changes after launch.
The output is a projected contribution range in dollars. It is useful for prioritization, but it is not booked revenue.
Pre-Test Projections: Knowing the Value Before You Run the Test
The model isn't just retrospective. I use it prospectively to prioritize which tests are worth running.
Before any test gets greenlit, I calculate the projected impact using the MDE (Minimum Detectable Effect) as a stand-in for the relative lift. The MDE is determined by the available traffic and the desired statistical power — it tells you the smallest lift the test can reliably detect.
So the pre-test calculation becomes:
Projected contribution at MDE = Eligible annual traffic × Baseline conversion rate × MDE × Contribution per conversion × Planned adoption × Persistence adjustment
For an illustrative planning scenario, suppose the smallest detectable effect translates to $150K in modeled annual contribution under the stated traffic, unit-economics, adoption, and persistence assumptions. A different funnel might model only $8K. Those hypothetical amounts are prioritization inputs, not project readouts or guaranteed outcomes.
This is how I compare opportunities across brands with different traffic volumes and unit economics. A smaller effect on a high-volume revenue path can matter more than a larger effect on a low-traffic informational page. The model makes those assumptions explicit and reviewable.
I add revenue per customer data to the projection to make it even more granular. If we know that the average Reliant residential customer generates a specific revenue figure over their lifetime, we can project the customer acquisition impact of a conversion lift and translate it directly into customer lifetime value.
Post-Test Validation: What Actually Happened
After a test concludes, the model updates the estimate:
- Replace the MDE scenario with the observed test-window effect and interval
- Recalculate the annual impact range using explicit traffic, unit-economics, adoption, and persistence inputs
- Use post-launch monitoring or holdouts where feasible to test whether the effect persists
The holdout step is important when it is feasible. A test-window effect may not persist after implementation. Novelty, seasonality, traffic-mix changes, and later product releases can all change the result. A holdout can estimate persistence under its own design, while post-launch monitoring provides a weaker but still useful check.
Not every test gets a holdout. It depends on the magnitude of the result, the strategic importance of the page, and whether the lift was surprising given our prior expectations. But for any test that's going to be cited in a board presentation or used to justify additional investment, holdout validation is non-negotiable.
How This Changed the Conversation
Before the EBITDA model, the experimentation program's narrative was about activity: "We ran 100 tests this year."
After the model, the narrative became about impact: "NRG's internal program reporting recorded $30M+ in impact, with the assumptions and evidence class shown alongside the total."
That's the difference between a cost center and a revenue driver. And it changed everything downstream.
Budget conversations shifted. Instead of defending the experimentation line item with activity counts, I could show observed test evidence, modeled impact ranges, and fully loaded program costs on the same page. That made the assumptions inspectable rather than turning the result into a guaranteed ROI multiple.
Stakeholder engagement increased. Brand marketing teams that previously saw testing as an inconvenience started requesting test slots. Showing the estimated business value and the uncertainty made prioritization more concrete without claiming that every tested idea would create revenue.
Tool and team investment cases followed. The model supported internal scenarios for experimentation, behavioral-analytics, and customer-data capabilities. Those scenarios made assumptions about capacity and decision quality visible; they do not prove that any tool caused a higher outcome rate.
This is the same principle behind Atticus Li's PRISM Method — measurement drives investment, not the other way around. You build the measurement framework first, establish an evidence trail at small scale, and then use the results and their limitations to make a scaling decision.
Honest About the Limitations
I want to be transparent about what this model is and isn't.
These are estimates. The EBITDA impact model produces projections, not audited financial results. There are assumptions baked in — that traffic levels will hold, that the lift is durable, that the competitive environment stays roughly stable, that the revenue per customer figure is accurate for the projection period.
More accuracy is possible. You can build more sophisticated models with time-decay adjustments, segment-level projections, competitive response modeling, and econometric controls. But more accuracy costs flexibility and speed. The EBITDA model is designed to be fast enough to run on every test and simple enough that a brand marketer or a finance analyst can understand and challenge the inputs.
The model is practical enough to apply consistently across a large portfolio, but its credibility depends on visible inputs, sensitivity ranges, and reconciliation with post-launch evidence.
The Organizational Impact
Beyond the numbers, the EBITDA model changed how the organization thinks about experimentation.
Before: Experimentation was perceived as a UX optimization activity — something the digital team did to make pages look better. It was evaluated on activity metrics (tests run, pages tested) and subjective assessments ("that new homepage looks great").
After: In my internal operating observation, experimentation was discussed alongside paid media, SEO, and product launches with its modeled assumptions and evidence limits visible. Test results entered forums where P&L decisions were reviewed.
That shift supported budget and staffing conversations alongside the portfolio evidence, delivery backlog, and measurement system. It was one input, not proof of a single cause.
I've written about why A/B tests fail, and one recurring failure mode in my experience is organizational misalignment. The model does not fix that directly, but it gives a common response to a common objection: “we do not know what assumptions make this worth funding.”
An assumption-labeled range makes that discussion more concrete without pretending the value is known exactly.
Building the Model for Your Organization
If you want to implement something similar, here's the practical sequence:
Step 1: Get the financial inputs. You need a finance-approved contribution per conversion or other unit-economics value, eligible traffic, adoption, and persistence assumptions. Do not multiply total brand EBITDA by page traffic.
Step 2: Build the pre-test projection template. A spreadsheet can expose eligible traffic, baseline conversion, MDE scenario, contribution per conversion, adoption, persistence, uncertainty, and implementation cost. Use the resulting range as one prioritization input.
Step 3: Add post-test updating. After each test, replace the planning scenario with the observed test-window estimate and recalculate the modeled range. Track recognized post-launch outcomes separately.
Step 4: Define persistence checks for material changes. Use a holdout where the value of better evidence justifies the traffic and operational cost. Otherwise, set a post-launch review date and report the weaker evidence class plainly.
Step 5: Present the evidence in finance's format. Show the observed effect, modeled annual range, recognized outcome when available, program cost, and alternatives without collapsing them into one certainty score.
The model doesn't have to be perfect on day one. My first version at NRG was a Google Sheet with manual inputs. It evolved over time as we refined the financial inputs, added segment-level projections, and automated the post-test calculations. The important thing is starting with a financial framing from the beginning, not bolting it on after you've already been running tests for a year.
What Comes Next
The next work explores customer-value scenarios, cross-test interactions, and whether historical portfolio patterns can improve planning. Those are model-development questions, not validated forecasts.
But the principle stays the same: every experiment should connect to a decision-relevant business outcome. When the primary metric is upstream of revenue, show the remaining assumptions rather than inventing a dollar value.
_Questions about building financial models for experimentation programs? Reach me at atticus@atticusli.com._
FAQ
Is modeled EBITDA impact the same as realized EBITDA?
No. A model translates observed evidence through stated assumptions. Realized impact requires post-launch reconciliation with the relevant finance or operating record.
Which assumptions should be visible?
Show the eligible population, effect estimate, rollout rate, unit economics, time horizon, costs, and any decay or substitution assumptions.
How should uncertainty be reported?
Carry an effect interval or scenario range through the model and label the output as an estimate. Microsoft Research explains why experiment metrics are easy to misinterpret, and Google Research shows why long-term effects may differ from test-window effects.