Atticus Li built the governance framework used when NRG’s internal 2025 program readouts recorded 100+ tests across five retail brands: Reliant Energy, Direct Energy, Green Mountain Energy, Cirro Energy, and Discount Power. The framework standardized briefs, prioritization, execution, and evidence labels across different traffic volumes and stakeholder groups; it does not prove that governance alone caused the program’s scale or outcomes.
The Challenge Nobody Warns You About
Much CRO guidance assumes one website, one team, and one stakeholder group. My NRG context involved five brands with separate priorities, traffic profiles, and decision owners.
The operating challenge was not that one brand was “better” than another; it was that the same test design, effect size, and calendar could not be copied across materially different funnels.
When I started building the experimentation program, the biggest challenge wasn't statistical methodology. It was organizational. How do you create a testing culture across five brands where each brand's marketing team has its own roadmap, its own VP, and its own definition of success?
The answer is governance. Not bureaucracy — governance. There's a difference, and getting it right determines whether your program scales or stalls.
Standardizing the Fundamentals
The first thing I established was a set of non-negotiable standards that apply across every brand. These aren't suggestions. They're requirements.
Decision threshold and power: pre-specified per procedure. Each brief states the error thresholds, power target, meaningful effect, and analysis method before launch. A conventional fixed-horizon design may use 5% alpha and 80% power, but those are documented inputs—not universal definitions of truth.
Minimum Detectable Effect (MDE): calculated per test, justified in the brief. There's no universal MDE threshold because traffic volumes differ dramatically across brands. What's detectable on Reliant's enrollment page in two weeks might require three months on a smaller brand's site. The MDE is calculated for every test, documented in the test brief, and reviewed before launch.
Test duration: cover the relevant operating cycle. A fixed-horizon retail-energy test often included complete weekday/weekend cycles, but the appropriate duration depended on traffic, seasonality, billing context, novelty, and the chosen procedure.
When to stop: follow the pre-specified rule. Fixed-horizon tests ran to their planned end. Sequential tests could stop only at pre-specified boundaries. Material guardrail harm or invalid instrumentation could stop either. An attractive dashboard point estimate was not, by itself, a stopping rule.
These standards exist so that any person joining the team — whether they're a new analyst, a brand marketer submitting a test request, or an agency partner executing a test — can follow the process without needing to derive the methodology from first principles.
Test Collision Avoidance
When multiple teams are running tests simultaneously across overlapping user populations, you get test collisions. Two tests competing for the same traffic. One test's variant affecting another test's control group. Interaction effects that make both tests' results uninterpretable.
At the scale of 100+ tests per year across five brands, this isn't a theoretical concern — it's a weekly operational challenge.
I built a collision avoidance system with three layers:
Layer 1: The test calendar. Every test across all brands goes on a shared calendar with traffic allocation percentages. Before any test launches, I review the calendar for overlaps. If two tests target the same page or the same audience segment, they get sequenced — not run simultaneously.
Layer 2: Traffic allocation rules. In this program, no single test received more than 50% of page traffic unless a critical business page had a documented risk/reward justification. The remaining traffic stayed on the default experience. That was an internal operating convention, not a universal statistical rule; allocation still had to support the approved design and power requirement.
Layer 3: Mutual exclusion groups. When tests had to run concurrently, mutual exclusion could prevent the same user from receiving both treatments. It reduced effective sample and co-exposure risk, but it did not remove spillovers, shared-market effects, or instrumentation interference.
The calendar is the most important layer. It sounds low-tech, and it is. But having visibility into what's running where, across all five brands, prevents more collisions than any automated system I've seen.
Stakeholder Management Across Brands
Here's where governance gets political.
Each brand has a marketing team that believes their priorities should come first. Each brand has a VP who saw a competitor's landing page and wants to test a copycat version. The HiPPO problem — Highest-Paid Person's Opinion — doesn't just multiply across five brands. It compounds.
I manage this through three mechanisms:
Mechanism 1: The test brief. Every test request, from every brand, goes through the same test brief template. The brief requires: a data-backed hypothesis (not an opinion), a projected EBITDA impact (not a gut feeling), and a pre-test power analysis (not a hope).
The brief is the common record. When someone proposes a new hero image, it requires the observed problem, decision metric, detectable effect, guardrails, and financial assumptions. In my operating experience, this often surfaced a weak foundation early enough to strengthen or withdraw the request; I do not publish a defensible rejection percentage.
Mechanism 2: The prioritization framework. All approved test briefs get scored on three dimensions: projected financial impact (from the EBITDA model), test feasibility (can we actually run this test cleanly?), and strategic alignment (does this test address a known conversion bottleneck?). Tests are ranked and scheduled in priority order, not in "who-asked-first" order.
The framework does not remove judgment, but it makes that judgment inspectable. Anyone can challenge the traffic, financial, risk, or strategic inputs rather than relying only on seniority.
Mechanism 3: Cross-functional workshops. I ran 10+ cross-functional workshops with Marketing, Product, and UX teams across the brands. These workshops served two purposes: educating stakeholders on how the testing process works, and surfacing test ideas from people who are closest to the customer.
The workshops created a shared vocabulary for decision rules, uncertainty, and the difference between diagnostic research and causal evidence. They also helped turn session-replay observations into falsifiable hypotheses rather than automatic conclusions.
I observed broader participation in briefs and readouts after the workshops, but this retrospective does not publish a defined stakeholder-buy-in metric, denominator, or comparison window.
When New People Join
Process documentation matters most when the person who built it isn't in the room. I designed the governance framework to be transferable, not dependent on my institutional knowledge.
The standards, templates, and decision framework gave new analysts a defined onboarding path. Documentation reduced dependence on oral history; it did not guarantee productivity within a fixed number of sprints.
But here's the nuance: the process isn't a religion. When new people join with experience from other programs — maybe they ran testing at a company with different traffic profiles, or they've used different statistical methodologies — I want to hear their recommendations.
The standards I set are defaults, not commandments. If someone can make a compelling case that we should use sequential testing instead of fixed-horizon testing for certain test types, or that we should adopt a Bayesian framework for low-traffic brands, I'm open to that. The test brief template has been revised four times based on team feedback. The prioritization scoring weights have been adjusted twice.
Good governance absorbs better methods without losing its structural integrity. Rigidity isn't a feature of governance — it's a failure of governance.
Guardrail Metrics: What to Watch When Scaling
Optimizing a primary metric without guardrails can move harm downstream. In an explicitly illustrative example, an enrollment-confirmation increase paired with a material retention decline would not qualify as a win.
At NRG, the guardrail metrics I monitor across every test are:
Enrollment confirmations. The primary metric for most enrollment funnel tests. But it's also a guardrail for tests on upstream pages — if we test a new homepage layout and enrollment confirmations drop, the homepage test failed regardless of what the homepage metrics show.
Deposit completion rates. For energy retail, the enrollment isn't complete until the customer completes their deposit. A test that increases enrollment starts but decreases deposit completions is moving friction, not removing it.
NPS impact. We monitor post-enrollment NPS scores segmented by test exposure. This is a lagging indicator, but it catches tests that optimize conversion at the expense of customer experience. A confusing but high-converting enrollment flow will show up in NPS before it shows up in churn — and by the time it shows up in churn, the damage is done.
Revenue per customer post-enrollment. This catches tests that attract lower-value customers. If a test increases conversion rate but the customers it converts generate less revenue, the EBITDA impact may be negative even though the conversion rate went up.
Guardrails aren't optional. They're part of the test definition. Every test brief specifies which guardrail metrics will be monitored and what threshold constitutes a stop condition.
What Governance Actually Looks Like Day to Day
Abstract frameworks are easy to describe and hard to sustain. Here's what governance looks like in practice:
Monday: Review the test calendar for the week. Check for any collisions or allocation issues. Review any tests that reached their planned end date over the weekend.
Tuesday: Test brief reviews for new submissions. This usually involves conversations with brand teams about hypothesis refinement, sample size calculations, and scheduling.
Wednesday: Mid-week check on running tests. Review guardrail metrics. Flag any anomalies for investigation.
Thursday: Results review for completed tests. Record the observed effect and uncertainty, update any modeled impact range, and document the decision and limitations.
Friday: Stakeholder communications. Weekly summary of results, upcoming tests, and program metrics. This is where the brand teams get visibility into what's happening across the portfolio.
This cadence provided operational support for a program that ran 100+ tests in 2025. The retrospective does not claim the calendar alone caused that throughput or publish a pre/post quality score.
The Payoff
The governance framework I built at NRG serves Atticus Li's PRISM Method in practice — structured measurement and iterative improvement applied at enterprise scale. It's not about controlling people. It's about creating an environment where good experimentation happens consistently, not accidentally.
When a program runs many tests but cannot show decision or business impact, one recurring root cause I have observed is weak governance—not necessarily bad people or ideas.
Good governance means every test has a purpose, every result has context, and every stakeholder understands both.
Microsoft Research's catalog of common experiment pitfalls is a useful review checklist for readout errors, while Google Research's work on long-term experiment effects explains why a test-window result and a durable business result can differ.
_Building a multi-brand or multi-team experimentation program? Reach me at atticus@atticusli.com._
FAQ
Does governance guarantee that an experimentation program will scale?
No. Governance can make briefs, decisions, risks, and readouts more consistent. Traffic, staffing, implementation capacity, instrumentation, and the quality of the opportunity set still constrain throughput.
Should every test use the same statistical method?
No. The approved method should fit the decision, traffic, stopping behavior, and error tolerance. The transferable rule is to specify the procedure before reading the result and then follow it.
Does mutual exclusion eliminate interaction risk?
No. It can prevent the same identified user from receiving two treatments, but it does not remove spillovers, shared-market effects, anonymous cross-device exposure, or instrumentation interference.