Atticus Li shares five historical enrollment-flow A/B tests from NRG Energy. The test-period metrics come from internal project readouts; the $1M+ portfolio total is modeled annual impact based on company inputs and persistence assumptions, not externally audited or booked revenue.
Why I'm Publishing Real Numbers
Most experimentation content on the internet falls into two categories: theoretical frameworks with no data, or case studies with vague results like "significant improvement in conversions." Neither is useful to a practitioner who needs to understand what a real test looks like from hypothesis to revenue projection.
These five records come from internal enrollment-flow readouts across NRG Energy's retail portfolio during my time building the experimentation program. They are historical company records, not an external audit or a benchmark another program should expect to reproduce.
For each test, I'll share: the hypothesis, the behavioral mechanism that informed it, the results, the revenue projection methodology, and what I'd do differently in retrospect. All five followed Atticus Li's PRISM Method — Probe, Revenue Rank, Implement, Score, Multiply.
The Pre-Work That Makes Tests Succeed
Before diving into the five experiments, I need to explain what happens before a test gets greenlit. This pre-work is what separates tests that generate insight from tests that waste traffic.
Minimum Detectable Effect (MDE) calculations: For every test, I calculate the smallest lift we'd need to detect given the available traffic and test duration. If the MDE is 15% but we're hypothesizing a 3% lift, the test can't detect the expected effect. It shouldn't run.
Revenue per customer projections: Each brand has a known average revenue per enrolled customer. This lets me translate a conversion rate lift into annual dollar impact before the test even starts. If the projected impact doesn't justify the test slot, we test something else.
Behavioral hypothesis: Every test starts with an observed behavior, not an opinion. "I think this page should look different" is not a hypothesis. "Session replay data shows 35% of users abandon the enrollment flow at step 3, with dead-click analysis suggesting confusion about form fields" is a hypothesis.
In this workflow, the pre-work usually took a few hours per test. Its value was decision quality: it exposed tests whose detectable effect, traffic, or projected business value did not justify the slot. That is a process observation, not a universal ROI claim.
Experiment 1: Streamlined Enrollment Flow (Reliant)
Brand: Reliant Energy
Hypothesis: The existing enrollment flow contains unnecessary steps, redundant form fields, and visual distractions that create friction for users who have already decided to enroll. Removing these elements will reduce abandonment and increase enrollment completions.
Behavioral mechanism: In this project, Contentsquare session replay showed users who had already selected a plan dropping off during enrollment. The form asked them to re-confirm information, navigate unnecessary steps, and work through visual noise. That observation motivated the experiments; it did not prove which mechanism caused the drop-off.
The behavioral principle here is cognitive load reduction. Once a user has committed to a decision, every additional step is friction that gives them a reason to reconsider or postpone. The hypothesis was grounded in specific observations: users were pausing at form fields that asked for information already captured in earlier steps, and back-button usage spiked at the review screen, suggesting users were unsure whether they'd entered everything correctly.
What we changed: We removed redundant form fields, eliminated an unnecessary confirmation step, reduced visual distractions on the enrollment screens, and streamlined the progress indicator to make the remaining steps feel shorter.
Results:
- 12% lift in enrollment confirms (p < 0.001, 99% Bayesian probability)
- ~100 additional enrollment confirmations recorded in the 46-day readout
- Projected annual revenue: ~$299K
What the data suggested: The observed movement was concentrated in mid-flow completion rather than traffic entering the flow. That pattern was consistent with the friction hypothesis, but the bundled variant did not isolate which individual change caused the effect.
What I'd do differently: I would have run a more granular test — isolating the form field reduction from the step elimination to understand which change drove the majority of the lift. The bundled approach gave us a clear winner but limited our ability to prioritize future optimizations within the flow.
Experiment 2: Homepage Acquisition Phone Number + Mobile Sticky Nav (Green Mountain Energy)
Brand: Green Mountain Energy
Hypothesis: GME's customer base includes a significant segment that prefers phone enrollment over digital. Making the phone number more prominent on the homepage and adding a sticky mobile navigation bar will increase call-driven enrollments without cannibalizing web enrollments.
Behavioral mechanism: This hypothesis came from an unexpected data intersection. Call center data showed that a meaningful percentage of enrollments happened via phone. Web analytics showed that users who eventually called often visited the website first — they used the site for research but preferred the phone for the actual transaction.
Cross-referencing these datasets revealed that GME's customer demographic skewed older than other NRG brands. This segment was comfortable browsing online but trusted a phone conversation for a purchase decision that involved a long-term energy contract. The phone number on the existing homepage was small, in the header, and disappeared when you scrolled on mobile.
What we changed: We placed a prominent, high-contrast phone number in the hero section. We added a utility selector widget to the homepage. And we implemented a sticky navigation bar on mobile that kept the phone number visible regardless of scroll position.
Results:
- Internal readout recorded roughly 3× call sales after the bundled change
- 948 additional call sales recorded in the internal 57-day readout
- Modeled annual impact: ~$523K
What the data suggested: The phone path mattered for this audience, and web enrollment did not show a corresponding decline in the internal readout. Because call tracking was not instrumented from the start, the recorded call-sales movement should not be treated as perfectly isolated incremental revenue.
What I'd do differently: I would have instrumented call tracking from day one of the test instead of relying on the call center's existing tracking. Better attribution data would have given us more confidence in the incremental nature of the call lift.
Experiment 3: Home Bundle Enrollment Free Copy and Visual Improvements (Direct Energy)
Brand: Direct Energy
Hypothesis: The home bundle enrollment page's copy and visual design were confusing users about the value proposition, particularly on mobile. Redesigning the copy to emphasize the "free" benefit and improving visual hierarchy would increase mobile enrollment completions.
Behavioral mechanism: Dead-click analysis in Contentsquare told the story here. Users were tapping on elements that weren't interactive — specifically on copy blocks that described pricing benefits. This is a strong signal that users were trying to learn more about something that interested them but couldn't. The design was creating interest it wasn't fulfilling.
Additionally, heatmap data showed that mobile users weren't reaching the CTA below the fold. They were reading the copy, getting confused or losing interest, and leaving. The copy itself was technically accurate but used industry language that required energy market knowledge to parse.
What we changed: We restructured the copy to lead with the clearest benefit statement — "free" — in plain language. We improved visual hierarchy so the eye naturally moved from headline to benefit to CTA. We reduced the amount of copy above the fold to ensure the CTA was visible without scrolling on most mobile devices.
Results:
- 60% mobile enrollment lift
- 70 additional enrollments recorded in the 36-day readout
- Projected annual revenue: ~$177K
What the data suggested: The larger mobile movement was consistent with a layout problem in which critical content sat below the fold. Because copy and visual hierarchy changed together, the experiment did not isolate which element produced the result.
What I'd do differently: I'd run a follow-up test that isolated copy changes from visual changes. The 60% lift was compelling, but we couldn't attribute it cleanly to the copy rewrite versus the visual restructuring. Understanding which lever drove more of the lift would inform how we approach other brand pages.
Experiment 4: Hero Layout Optimization (Green Mountain Energy)
Brand: Green Mountain Energy
Hypothesis: Restructuring the hero section on mobile to prioritize product information and CTA visibility will increase enrollment starts and improve the path from hero engagement to the product comparison chart.
Behavioral mechanism: Heatmap data showed a clear problem on mobile: users were not scrolling past the hero section. The hero occupied the entire viewport on most phone screens, and the CTA to explore plans was below the fold. Users who did scroll showed high engagement with the product chart, suggesting the content below the hero was effective — it just wasn't being seen.
This is a classic "first screen" problem. On desktop, the hero and the first section of content are both visible without scrolling. On mobile, the hero takes up everything, and users have to actively scroll to discover that there's more to the page. If the hero doesn't compel a scroll or provide a visible CTA, you lose users who never see your best content.
What we changed: We restructured the hero layout on mobile to reduce its viewport footprint, ensuring that the top of the next section (including the plan comparison CTA) was visible without scrolling. We also adjusted the visual hierarchy within the hero to direct attention toward the primary CTA.
Results:
- 7% secondary lift in enrollment starts
- 35% improvement in product chart views-to-enrollment rate
- Projected annual revenue: ~$212K
What the data suggested: Enrollment starts and the chart-to-enrollment ratio both moved in the intended direction. The secondary ratio was diagnostic, not proof that the layout created “higher-quality” users; selection into chart viewing can also change the denominator.
This is why Atticus Li's PRISM Method separates the primary decision metric from secondary diagnostics. The latter can suggest a mechanism and shape a follow-up test without being promoted into causal proof.
What I'd do differently: I'd test progressive disclosure on the hero—showing a teaser of plan pricing directly in the hero—to examine the scroll-gap hypothesis more directly.
Experiment 5: Product Chart Value Prop CTAs (Green Mountain Energy)
Brand: Green Mountain Energy
Hypothesis: Adding value proposition CTAs directly in the product comparison chart will increase enrollments by reducing the cognitive gap between plan evaluation and enrollment action.
Behavioral mechanism: The product chart was where users made their plan selection decision. Analytics showed strong engagement with the chart — users compared plans, toggled between options, and spent significant time on the page. But the enrollment CTA was separated from the chart. Users had to finish evaluating plans, scroll past the chart, and find the enrollment button.
This separation created a cognitive gap. Users were in evaluation mode while looking at the chart. By the time they reached the enrollment CTA, they had left evaluation mode and needed to re-motivate themselves to take action. The hypothesis was that embedding CTAs with value prop messaging directly in the chart would capture users at their moment of peak interest.
What we changed: We added value proposition CTAs — brief messaging highlighting key benefits — directly within the product comparison chart, adjacent to each plan option. The CTA appeared in context, while the user was actively evaluating that specific plan.
Results:
- 3% enrollment lift
- 35% improvement in product chart-to-enrollment rate
- Projected annual revenue: ~$86K
What the data suggested: The overall enrollment metric and the chart-to-enrollment ratio moved in the intended direction. The pattern was consistent with reducing the gap between comparison and action, but the test did not independently identify which copy or placement element drove it.
This test illustrates a principle I come back to repeatedly: small percentage lifts on high-volume pages translate to meaningful revenue. A 3% lift sounds unimpressive in a case study. $86K in projected annual revenue sounds like a business decision worth making.
What I'd do differently: I'd test different value prop messages by plan type. The winning variant used the same messaging framework across all plans. Tailoring the value prop to the specific differentiator of each plan (price for the budget plan, renewable sourcing for the green plan) could compound the effect.
What Ties All Five Experiments Together
Looking across these five tests, several patterns emerge that inform how I prioritize and design experiments:
Pre-test analysis made the trade-offs visible: Each test had an MDE calculation and a revenue-per-customer scenario. Those inputs did not guarantee a useful result; they showed whether the test could answer a decision-relevant question with the available traffic.
Behavioral mechanisms proposed before testing: Session replays, heatmaps, dead-click analysis, and call-center data informed each hypothesis. The experiments tested interventions; they did not prove every proposed mechanism.
Financial translation with an evidence label: Each test can be paired with a modeled annual-impact range so finance can inspect the assumptions. The observed test result, modeled annualization, and recognized post-launch revenue should remain separate.
The PRISM Method as connective tissue: Atticus Li's PRISM Method provided the framework for every test. Probe (behavioral analysis), Revenue Rank (MDE and revenue projections), Implement (test design and deployment), Score (statistical analysis and segmentation), Multiply (scaling winners and informing the backlog). That consistency supported the 100+ experiment annual cadence alongside staffing, traffic, tooling, and stakeholder capacity; it did not prove the framework alone caused the scale-up.
Portfolio reporting changes the conversation: Five internal models summed to $1.2M+ in projected annual impact. That historical scenario helped translate the portfolio for executive stakeholders, but it should not be read as booked revenue or a forecast for another team.
The Invitation
If you're running experiments and not tying every test to revenue projection, you're leaving program growth on the table. The tests themselves might be excellent. But without financial framing, the people who control budget won't understand why your work matters.
I'm happy to walk through how to build revenue-per-customer models for your experimentation program, or how to structure your executive reporting to drive buy-in. Reach out at atticus@atticusli.com.
The data is only as valuable as the decisions it enables. Make sure the decision-makers can see what you see.
FAQ
Does the projected annual impact equal collected revenue?
No. The figures in this article are historical internal projections tied to NRG test readouts. They are not an external audit, recognized revenue, or a forecast for another company.
What makes an enrollment test decision-ready?
Define eligibility, the primary enrollment outcome, guardrails, sample plan, stopping rule, and the unit-economics assumptions used after the test.
What should happen after rollout?
Reconcile the test-window estimate with observed enrollment and financial records, while checking for decay or delayed effects. Microsoft Research documents metric-interpretation pitfalls, and Google Research explains why long-term measurement matters.