Atticus Li is the experimentation lead at NRG Energy. By 2025, the company's testing program was running 100+ experiments per year after scaling from roughly 20, with 150+ historical experiments across five retail energy brands. The exact figures below come from internal project readouts. Projected annual lift is modeled from test-period evidence; the broader $30M+ program figure is impact validated through internal NRG reporting, not an external audit or a client forecast.
The State of Things When I Arrived
When I joined NRG, the experimentation program existed but it was small. Around 20 tests a year, mostly driven by agency partners. The tests weren't bad, but they suffered from a problem I've seen at every enterprise I've worked with: the HiPPO problem.
HiPPO stands for Highest-Paid Person's Opinion. And at a company managing brands like Reliant Energy, Direct Energy, Green Mountain Energy, Cirro Energy, and Discount Power, there were a lot of HiPPOs with a lot of opinions.
The typical test request looked like this: a VP sees a competitor's homepage and says "let's test that layout." No data backing the hypothesis. No understanding of whether the current design was actually underperforming. No pre-test analysis to determine if the test could even reach statistical significance with available traffic.
I knew from my time at SVB that experimentation only scales when you remove opinion from the equation and replace it with process. The question was how to do that across five-plus brands with different traffic volumes, different customer segments, and different stakeholder groups.
Step One: Bringing Everything In-House
The first major shift was bringing experimentation fully in-house. The agency model had served its purpose, but it created a dependency that slowed everything down. Turnaround times were long. Context was lost between handoffs. And the agency team didn't live inside our data the way an internal team needs to.
I built out the internal capability piece by piece. That meant standardizing our tooling — Adobe Analytics for measurement, Optimizely for test execution, Contentsquare for behavioral analytics, and Tealium as our CDP layer. Each tool had a specific role, and I documented exactly how they connected so that anyone joining the team could get productive within their first week.
This wasn't glamorous work. It was writing SOPs, building test request templates, creating pre-test analysis frameworks, and establishing review cadences. But it was the foundation everything else was built on.
Step Two: Tying Every Test to Revenue
Here's the thing most CRO content on the internet won't tell you: at real companies, you don't have millions of monthly users. You don't get to run 50 concurrent tests with clean traffic allocation. You face sample size constraints that make many "best practice" recommendations from CRO influencers completely irrelevant.
When you're working with the traffic volumes of a regional energy brand — not Google, not Amazon — every test slot is precious. You can't waste one on a test that has no chance of reaching significance, or worse, one that reaches significance but can't be tied back to dollars.
So I built a pre-test framework centered on Minimum Detectable Effect (MDE) projections tied to revenue per customer. Before any test gets greenlit, we calculate:
- Current conversion rate for the page or flow being tested
- Traffic volume over the planned test duration
- Finance-approved contribution per conversion for that brand and product
- MDE and business threshold — what the traffic can detect and the smallest effect worth implementing
- Modeled annual contribution range under explicit adoption and persistence assumptions
This changed the conversation from seniority to inspectable inputs: eligible traffic, business threshold, detectable effect, contribution economics, implementation cost, and uncertainty.
Finance could challenge that language. The evidence supported a more concrete budget discussion; it did not make investment automatic.
Step Three: Building a Reviewable Value Case
Growing from roughly 20 tests to 100+ per year required an internally reviewable value case: test-period outcomes, explicit modeling assumptions, and program reporting the finance team could challenge.
Here are four selected internal readouts from 2025 that contributed to that case. Their dollar figures are modeled annual impact, not audited realized revenue:
Enrollment Flow Optimization (Reliant): We redesigned the enrollment flow based on Contentsquare session replay analysis that showed users dropping off at the plan comparison step. The internal readout recorded a 12% lift in enrollment confirmations, translating to approximately $299K in modeled annual impact.
Homepage Phone Number Placement (Green Mountain Energy): An internal readout recorded roughly 3× call sales after phone prominence, a utility selector, and mobile sticky navigation changed. The model projected approximately $523K in annual impact. Because call tracking was not instrumented from the start, the movement is not perfectly isolated incremental revenue, and the bundled variant did not prove which element caused it. The enrollment retrospective preserves the full test detail and limitation.
Home Bundle Enrollment Experience (Direct Energy): A bundled copy and visual-hierarchy variant on Direct Energy's home battery experience page recorded a 60% lift in mobile enrollment during the test window. Modeled annual impact: $177K. This test was born from a dead-click analysis in Contentsquare—users were tapping on non-interactive elements, signaling confusion about what was clickable. Because multiple elements changed together, the readout did not isolate which element drove the movement.
Hero Layout Test (Green Mountain Energy): A hero section layout change on GME's homepage recorded a 7% secondary lift in enrollment starts during the internal test window, with approximately $212K in modeled annual impact. The hypothesis came from heat map data showing that users weren't scrolling past the hero on mobile—the CTA wasn't visible without scrolling on most phone screens.
Together, these four readouts modeled more than $1.2M in annual impact under their stated assumptions. That made the staffing and tooling case more concrete; it did not guarantee approval or realized value.
The System That Makes It Repeatable
The hardest part of scaling experimentation isn't running more tests. It's making the program run without you being the bottleneck.
I built standardized processes so that anyone joining the team can follow the PRISM framework from hypothesis to post-test analysis without needing to reinvent the wheel. That includes:
- Test request intake forms that force requestors to state the problem, not the solution
- Pre-test analysis templates with MDE calculations built in
- QA checklists for every test deployment across Optimizely
- Post-test reporting templates that include statistical results, projected annual lift, and recommendations for holdout testing
- A centralized test repository so stakeholders across all five brands can see what's been tested, what won, and what lost
The repository is especially important in a multi-brand environment. A test that wins on Reliant might inform a hypothesis for Direct Energy. A pattern that fails on Stream might save GME from wasting a test slot.
What I've Learned About Experimentation at Scale
Most CRO advice is written for companies with millions of users. If you're at a real enterprise with five brands and varying traffic levels, you need a completely different playbook. Sample size constraints are your biggest strategic challenge, not your testing tool's feature set.
Experimentation isn't about running tests. It's about making better decisions. Every test is a decision-making tool. The output isn't a green or red result — it's a dollar value attached to a specific change, with a confidence interval that tells you how much to trust it.
The HiPPO problem never fully goes away. But when you have a framework that translates every hypothesis into projected revenue impact, you give stakeholders a common language. They can still advocate for their ideas, but now those ideas compete on merit, not org-chart position.
Your win rate needs a denominator and evidence context. The NRG program internally recorded a roughly 24% positive-primary-outcome rate in 2025 under its own counting rules. We invested in heat maps, session replays, click-rate analysis, rage-click detection, and qualitative research where available. That process made hypotheses more explicit; it does not prove that research caused the portfolio rate or establish an industry benchmark.
These are best estimates with available data. I want to be honest: projected revenue numbers are estimates. They're calculated using the best data we have — conversion lifts, traffic volumes, revenue per customer — but they're not audited financial statements. More precision is possible, but it comes at the cost of flexibility and speed. In a fast-moving experimentation program, I'd rather be directionally right and fast than precisely right and slow.
What's Next
The program continues to grow. We're expanding into personalization testing, exploring AI-assisted hypothesis generation, and building more sophisticated holdout methodologies to validate long-term impact beyond the initial test window.
If you're building an experimentation program at a real company — not a Silicon Valley unicorn with unlimited traffic — I'd love to hear how you're handling sample size constraints. Reach out at atticus@atticusli.com.
You can see more about my work at NRG on my NRG case study page, or read about the PRISM framework I developed for running revenue-driven experiments.
FAQ
Is the $30M figure externally audited revenue?
No. It is the historical impact recorded in NRG's internal 2025 experiment readouts and program reporting. It is not an external audit or a forecast for another company.
What enabled the increase from roughly 20 to 100+ annual tests?
The program combined intake, prioritization, reusable standards, cross-functional delivery, and explicit measurement. More volume was useful only while the evidence bar remained visible.
Can another company expect the same timeline or outcome?
No. Traffic, funnel economics, team capacity, instrumentation, and decision rights vary. Microsoft Research documents common experiment pitfalls, and Google Research explains why test-window and long-term effects can differ.