Atticus Li's PRISM Method is a five-step experimentation framework — Probe, Revenue Rank, Implement, Score, Multiply — designed for enterprise teams that need every test to expose its business assumptions. It was developed through 150+ historical NRG Energy experiments across five brands. In 2025, the NRG program internally recorded a roughly 24% positive-primary-outcome rate; that project result is not an industry benchmark, a causal claim about the framework, or a forecast for another team.
Why I Built the PRISM Method
Most experimentation frameworks I've encountered fall into one of two traps. Either they're so rigid that they slow teams down to the point where you're running 15 tests a year instead of 100. Or they're so loose that every test is basically a coin flip — someone has a hunch, they test it, and win or lose, nobody learns anything systematic.
When I was scaling NRG's experimentation program from 20 tests per year to 100+, I needed something in between. The framework had to preserve a consistent evidence bar while handling multiple brands, varying traffic volumes, stakeholders with strong opinions, and leadership that needed the business assumptions—not only the p-values.
That's how Atticus Li's PRISM Method was born. Not in a conference talk or a blog post, but in the messy reality of trying to convince a CFO that experimentation deserves more budget.
The Five Steps
P — Probe
The Probe phase is where weak evidence often enters the process. Skipping it can leave a team with an unexamined problem statement, but it does not by itself explain a portfolio's outcome rate.
Before you ever write a test hypothesis, you need to understand the problem. Not the solution you want to test. The problem.
This means gathering both qualitative and quantitative data:
Quantitative signals:
- Heat maps showing where users actually click versus where you expect them to click
- Click-rate analysis across the entire page
- Dead-click detection — users clicking on non-interactive elements, which signals confusion about the interface
- Rage clicks — repeated rapid clicking that indicates frustration
- Scroll depth analysis to understand what content users actually see
- Funnel drop-off analysis in Adobe Analytics or your analytics tool of choice
Qualitative signals:
- Session replays watched in bulk, not cherry-picked
- Customer support ticket themes related to the page or flow
- Call center feedback (especially valuable for brands with phone-heavy customer segments)
- User testing sessions when available
- AI-powered session analysis tools like Contentsquare's AI summaries
The cardinal rule of the Probe phase: don't solutionize. You're not here to figure out what to change. You're here to figure out what's broken and why.
I've seen teams skip the Probe phase hundreds of times. A stakeholder says "the button should be green" and the team tests green vs. blue. That's not experimentation. That's decoration.
R — Revenue Rank
Before a test is approved, we build an annual-impact scenario using the minimum detectable effect (MDE) and visible financial assumptions. This is a planning model, not a forecast that the treatment will win.
Here's the actual process:
- Identify the conversion metric — enrollment starts, form submissions, call initiations, whatever maps to revenue for this specific page and brand
- Pull current performance — baseline conversion over a relevant pre-period, with seasonality and traffic-mix risks documented
- Calculate available sample size — how much traffic this page gets over the planned test duration
- Determine MDE — under the pre-agreed analysis procedure, power, and error thresholds, what effect is the design equipped to detect?
- Map to unit economics — using a finance-approved contribution per conversion, translate the MDE scenario into dollars
- Calculate a modeled annual range — if an effect of that size held after rollout, what would the contribution range be under explicit adoption and persistence assumptions?
Tests are ranked using modeled impact alongside evidence strength, feasibility, implementation cost, customer risk, and learning value.
In the historical NRG workflow, we prioritized tests whose pre-test models showed roughly $200K–$500K in annual-impact scenarios if the specified effect held. Those were decision estimates, not known outcomes.
A critical nuance: these models are only as useful as their inputs. Traffic, contribution per conversion, adoption, and persistence can all change. The purpose is to expose a reviewable scenario before consuming a test slot, then reconcile it with the test-window estimate and later recognized outcomes.
I — Implement
Implementation is where you actually design and build the test. But there's a crucial sub-step most teams miss: diverse opinions.
Before finalizing any test design, I bring in perspectives from across disciplines:
- UX researchers who understand user mental models
- Designers who can identify visual hierarchy issues
- Developers who know what's technically feasible and what might introduce bugs
- Analytics engineers who can flag measurement challenges
- Marketing stakeholders who understand brand voice and positioning
Why? Because a hypothesis built from one person's perspective is a sample size of one. And we all know what happens with a sample size of one.
The implementation itself follows strict QA protocols. Every test deployed in Optimizely goes through:
- Visual QA across devices and browsers
- Analytics validation — are events firing correctly?
- Traffic allocation verification
- Edge case testing — what happens when users navigate away and return?
I've seen tests where a measurement bug meant the "winning" variant was actually capturing duplicate conversions. QA isn't optional. It's insurance against making expensive wrong decisions.
S — Score
After the test reaches its pre-agreed stopping or readout rule—or stops because of invalidity or material guardrail harm—we score the evidence. This goes beyond “winner” or “loser.”
The scoring process includes:
Statistical rigor:
- The decision rule used, whether the stopping rule was followed, and the effect estimate with uncertainty
- How the observed test-window effect compares with the pre-agreed meaningful effect and design MDE
- Pre-specified segment results, with post-hoc cuts labeled exploratory
Revenue projection:
- The observed test-window effect and uncertainty, kept separate from any financial model
- A modeled annual-impact range using eligible traffic, contribution per conversion, adoption, and persistence
- Later recognized post-launch outcomes, compared with the earlier model to calibrate planning assumptions
Learning documentation:
- What did we learn about user behavior?
- Does this validate or contradict previous test results?
- Are there implications for other brands in the portfolio?
A valid negative or null result can still improve a decision when the design was capable of answering the question. An invalid test cannot support the decision even when the point estimate looks favorable.
M — Multiply
The Multiply phase is what separates a testing program from an experimentation program.
When a readout supports rollout, we do not assume the same result will transfer unchanged. We ask:
- Can this insight be applied to other brands? A hero layout win on Green Mountain Energy might translate to Reliant or Stream.
- Can we isolate the mechanism? If a readout moves after repositioning the CTA, a follow-up can test whether placement, copy, or their combination mattered.
- Should we run a holdout test? For high-impact wins, we hold back a percentage of traffic on the original experience to validate that the lift persists over months, not just weeks.
- What does this suggest about user behavior more broadly? An internal GME readout recorded roughly 3× call sales after phone prominence and mobile navigation changed. Because call tracking was not instrumented from the start, that movement is not perfectly isolated incremental revenue; it supports a follow-up hypothesis, not a universal rule. The enrollment retrospective preserves the full limitation.
The Multiply phase tests whether an insight transfers. Each rollout or brand needs its own evidence; portfolio impact should be accumulated from labeled readouts rather than assumed from the first result.
What Atticus Li's PRISM Method Is Designed to Control
The NRG program's internally tracked 2025 rate was roughly 24% under its own counting rules. Research was deliberately front-loaded, but the observed rate does not prove that PRISM caused the result or that another portfolio should target the same percentage.
By the time a test launches, we've already:
- Diagnosed the actual problem through behavioral data
- Sized the test for a pre-agreed meaningful effect under the chosen procedure
- Built an assumption-labeled impact scenario to compare the opportunity with cost and risk
- Gathered diverse perspectives to strengthen the hypothesis
- QA'd the implementation to reduce measurement risk
Pre-launch work reduces avoidable risk; the readout and post-launch reconciliation still require the same evidence discipline.
Common Mistakes I See
Testing solutions instead of problems. "Let's test a new button color" is a solution. "Users aren't seeing the CTA on mobile because it's below the fold" is a problem. Test the problem, not a specific solution.
Ignoring sample size constraints. As an illustrative example, a page with 50K monthly visitors needs a different test design and opportunity set than a page with millions. The framework makes that constraint visible rather than promising the same cadence everywhere.
Hiding the path to a business outcome. When the primary metric is upstream of revenue, show the remaining assumptions instead of inventing a precise dollar value. Revenue Rank makes that chain inspectable.
Running tests in isolation. A new brief should inspect relevant prior evidence. The Multiply phase creates a record and reuse loop that can support learning across brands and test cycles; it cannot ensure that an insight transfers or compounds.
Treating experimentation as a feature of a tool. The tool is not the program. The program is the people, process, measurement, and decision framework. PRISM is tool-agnostic by design, while each implementation still needs local validation.
Applying PRISM to Your Team
The framework can be adapted to different program sizes. If you run 10 tests a year, PRISM helps compare candidate decisions under stated traffic, value, effort, and evidence inputs; it cannot prove those are the highest-impact tests possible.
Start with the Probe phase. Use approved quantitative and qualitative tools to understand the problem before testing a solution.
Then add the Revenue Rank step. An illustrative comparison—one page has 10× the eligible traffic and a finance-approved $500 contribution per conversion—can change prioritization, provided uncertainty and implementation cost remain visible.
The remaining steps do not emerge automatically. Implementation, scoring, and transfer each require their own owners, controls, and local validation.
If you want to see the PRISM Method in action, read about how I applied it to NRG's experimentation program or explore the framework page for a visual overview.
Questions about applying the PRISM Method to your team? Reach out at atticus@atticusli.com.
FAQ
Does PRISM guarantee a revenue lift?
No. PRISM makes assumptions, decision rules, and follow-up evidence visible; it cannot guarantee that a hypothesis will win. Microsoft Research's catalog of experiment pitfalls explains why a disciplined process still needs careful interpretation.
How should a team choose its primary metric?
Choose the closest measurable outcome to the business decision, then document the remaining steps to recognized revenue. Google's research on long-term experiment effects is a useful reminder that the test-window metric may not capture the full effect.
What is the smallest useful way to adopt PRISM?
Start with one priority funnel. Write the user problem, revenue assumption, primary metric, sample plan, and ship-or-stop rule before development begins.