More Experiments.
The Same Evidence Bar.
An in-house experimentation program across 5 regulated energy retail brands—Reliant Energy, Direct Energy, Green Mountain Energy, Cirro Energy, and Discount Power—with shared standards for prioritization, QA, analysis, and financial modeling.
Five Brands. Zero
Shared Methodology.
NRG Energy operates multiple regulated energy retail brands, each with different customer bases, regulatory environments, and acquisition funnels. When I joined, experimentation was fragmented—roughly 20 tests per year, with inconsistent standards for connecting test evidence to revenue models.
The central business question was: what evidence supports the value of the experimentation program? Making test-period outcomes, assumptions, and uncertainty visible gave finance and leadership a more reviewable basis for capacity and investment decisions.
From Foundation
to Compound Growth
Consolidating a Fragmented Program
Inherited testing across 5 brands without one shared operating method. Reviewed the existing portfolio, documented measurement gaps, and built a hypothesis-to-revenue evidence chain. The operating standard included minimum detectable effect, power calculations, documented stopping rules, and guardrail checks.
From Roughly 20 to 100+ Tests by 2025
Rolled out shared intake, prioritization, QA, and readout standards across Reliant Energy, Direct Energy, Green Mountain Energy, Cirro Energy, and Discount Power. Modeled impact made assumptions reviewable alongside the experimental evidence; it did not turn every metric movement into realized revenue.
AI-Augmented Experimentation
Added approved AI-assisted steps for candidate-pattern review, analysis support, and drafting. A first-person, uncontrolled workflow estimate moved from roughly eight to five hours per readout; statistical checks, source-data validation, and human review remained mandatory. The estimate does not isolate AI as the cause.
Inside the NRG Experiment Portfolio
Test-level details stay internal—but the themes show how the program operated. The documented standard required a named behavioral mechanism, pre-specified primary outcome, power and minimum-detectable-effect checks, guardrails, and an approved stopping and readout procedure. Those controls make the evidence reviewable; they do not guarantee a positive result.
Reliant Enrollment-Flow Simplification
A selected internal readout recorded a 12% lift in enrollment confirmations, about 100 additional confirmations over 46 days, and approximately $299K in modeled annual impact. The bundled variant did not isolate one causal element.
Direct Energy Home-Bundle Experience
A bundled copy and visual-hierarchy variant recorded a 60% mobile enrollment lift, 70 additional enrollments over 36 days, and approximately $177K in modeled annual impact. The readout did not isolate copy from layout.
Green Mountain Energy Call-Channel Capture
A selected internal readout recorded roughly 3× call sales and 948 additional call sales over 57 days after a bundled change, with approximately $523K in modeled annual impact. Incomplete call tracking means the movement is not perfectly isolated incremental revenue.
Green Mountain Energy Hero Layout
A selected internal readout recorded a 7% secondary lift in enrollment starts, a 35% change in the chart-to-enrollment ratio, and approximately $212K in modeled annual impact. The secondary ratio was diagnostic, not proof of higher-quality users.
Credit and Deposit-Step Friction
Internal tests evaluated identity-verification framing and the salience of lower-commitment options. Their readouts were interpreted as funnel-specific evidence, not a reusable conversion or cost-savings benchmark.
Personalization via CDP Segmentation
A bundled Tealium and Optimizely treatment recorded a 23% lift in lead acquisition versus control in an internal readout. The result does not isolate AI, segmentation, content, or delivery tooling as the sole cause.
Every test moves through the same 12-step operating pipeline — from idea intake to executive readout and handoff.See the pipeline →
The Principles
Behind the Numbers
The NRG program used named behavioral mechanisms where customer evidence supported them. These four examples generated testable hypotheses; none was treated as a guaranteed outcome.
Loss Aversion
Used truthful loss-versus-gain framing as a hypothesis for rate-plan messaging, then evaluated the result within each brand instead of assuming the mechanism would transfer.
Choice Architecture
Tested whether fewer, better-explained plan choices and clearly disclosed defaults improved qualified enrollment decisions.
Social Proof
Tested substantiated local choice cues on rate-comparison pages. One internal readout recorded a 22% shorter time-to-decision during its measurement window.
Anchoring
Tested plan order as an anchoring hypothesis while retaining average revenue per customer as a downstream guardrail.
Built on the
PRISM Method
The NRG experimentation program uses PRISM to frame a pre-test value range, document its assumptions, and compare the post-test evidence with that range. The result retains its uncertainty; a metric movement is not automatically booked as realized revenue.
See the PRISM Method →See all case studies
Explore the geo-experimentation work at Silicon Valley Bank, see individual test-level evidence in the Experiments hub, or read weekly breakdowns on Lean Experiments.
Make Growth Decisions
You Can Defend
Know what to test, when to trust the result, and what to do next. Practical decision guides for analysts, growth teams, and founders.
Opens Substack to confirm · Free · Unsubscribe anytime
Read Better Decisions
Practical guides on experimentation, analytics, commercial judgment, and using AI without outsourcing the decision.
Open Substack (opens in new tab)