The most-cited finding in behavioral pricing research predicted a clean win here, and a well-powered experiment against real customers returned a shrug.

TL;DR

  • We tested the textbook "choice overload" prediction directly: would showing fewer pricing plans (9) convert better than showing more (14) on a live product-chart page? (Exp-057)
  • Result: a small directional lift toward the shorter list, in the 0% to 5% range — but not statistically significant.
  • This ran on one of the largest traffic samples in this entire portfolio, which matters: an underpowered sample can't explain away this null the way it could elsewhere.
  • The finding worth publishing isn't a near-miss — it's that a genuine, well-powered null on a landmark psychological effect rules something out more convincingly than a smaller inconclusive test, or simply trusting the literature, ever could.
  • Lesson for leaders: choice overload is real research, not a universal law you can design around unverified.
Longer list (control)Shorter list (variant)
Plans shown149
HypothesisBaseline experienceFewer options reduce decision fatigue, lift conversion
Duration~8 weeks~8 weeks
SampleOne of the largest in this portfolioSame
OutcomeDirectional lift, 0–5% range, not statistically significant

The decision behind the plan count

Somewhere in nearly every pricing or catalog review, someone cites the same idea: people convert better when they have fewer choices. It shows up in board decks, in redesign briefs, in "best practice" checklists handed to design teams with no experiment attached. It's treated as settled enough that teams will restructure a pricing page, cut a product line, or collapse a menu on the strength of the citation alone.

That's the decision this experiment was built to test, not assume. A regional utility brand's product-chart page listed a relatively wide set of plans — a normal outcome of years of product proliferation, not a deliberate choice architecture. The hypothesis on the table was straightforward: cut the list roughly in half, from 14 options down to 9, and conversion should rise because visitors spend less energy comparing and more energy deciding.

The business stakes here aren't abstract. Trimming a plan list has real downstream cost — it means suppressing plan types some segment of visitors specifically wants, and it changes the mix of what gets sold, not just how much. A team shouldn't absorb that cost on the strength of a well-known finding it has never tested in its own funnel, with its own customers, on its own page. That's the exact posture this experiment was designed to avoid: don't inherit someone else's effect size and call it a decision.

Choice overload, the paradox of choice, and why we expected a win

The prediction here comes from real, well-regarded research, not a marketing myth. Sheena Iyengar and Mark Lepper's jam-tasting studies — the foundational work behind what's now called choice overload — found that shoppers offered a smaller assortment at a tasting table were dramatically more likely to actually make a purchase than shoppers offered a much larger one. Barry Schwartz took the same idea further in _The Paradox of Choice_, arguing that beyond a certain point, more options don't just fail to help decision quality — they actively degrade satisfaction and follow-through.

It's a genuinely compelling mechanism, and it's the reason the hypothesis for Exp-057 was reasonable to fund in the first place: reduce the plan list, reduce the friction, expect the lift. If you only read the popular summary of this literature, running this experiment might look unnecessary — the outcome seems obvious.

It isn't obvious, and this is where intellectual honesty about the underlying science matters. Choice overload is also one of the more prominent casualties of the replication crisis in behavioral science. A well-known meta-analysis of choice-overload studies (Scheibehenne, Greifeneder, and Todd, 2010) found the effect is far less consistent than the popular narrative suggests — it appears strongly in some contexts and vanishes or reverses in others, depending on category familiarity, how options are presented, and who's choosing. A famous finding from a jam-tasting table doesn't automatically transfer to a utility customer comparing rate plans. That gap between "well-established in the literature" and "confirmed in your context" is precisely what this experiment was built to close.

Why we ran this at genuine scale, not a quick read

The methodology choice that matters most in this experiment isn't the 9-versus-14 split — it's the decision to run it on one of the largest traffic samples in this portfolio, for a full run of roughly eight weeks, instead of treating it as a quick confirmatory check.

That choice was deliberate. A senior practitioner's job isn't just to design the right test — it's to size it so the result actually means something regardless of which way it lands. Elsewhere in this same portfolio, an inconclusive outcome could plausibly be explained away by an underpowered sample: not enough traffic, not enough time, no real signal to detect either way. That excuse isn't available here. This experiment had the traffic and duration to detect a lift if the underlying mechanism was operating the way the literature predicts, and it still came back inconclusive.

That distinction is the entire value of the exercise. A small, underpowered test that comes back null tells you almost nothing — it's equally consistent with "no effect" and "we couldn't see the effect." A large, well-powered test that comes back null is a much stronger claim: it means the effect, if it exists in this context, is probably too small to matter commercially, not merely hidden by noise. That's a different, more useful kind of answer, and it's only available if you spend the sample size to earn it.

The result: a real, well-powered null

Exp-057 came back inconclusive. There was a small directional lift toward the shorter, 9-plan list — the direction the choice-overload hypothesis predicted — but it landed in a 0% to 5% range and did not reach statistical significance, despite running on one of the largest, best-powered samples in this experiment set.

To be direct about what that means: the data was inconclusive, and I'm comfortable saying so. This wasn't a test that ran too short, too small, or on too thin a slice of traffic to know what happened. It ran long enough and large enough that the honest read is: at this scale, plan count alone wasn't the lever the hypothesis expected it to be. That's a genuine result, not an artifact of noise or a rounding error in the experiment design.

The paradigm most teams don't question

Here's the uncomfortable part for anyone who has built a strategy deck around choice overload: the field tends to treat this effect as close to universal — a safe assumption you can design a pricing page or catalog around without re-testing it, because "the research already proved it." Exp-057 argues against that posture, at least as a blanket rule.

A well-powered null on a landmark finding like this doesn't mean choice overload is false — it means plan count in isolation isn't guaranteed to be the dominant lever in every context, even when the literature says it should be. Utility customers comparing rate plans are not lab participants tasting jam at a table, and the mechanism that drives one may not drive the other in the same way, at the same magnitude, or at all. The right conclusion isn't "ignore the research." It's "treat published effects as a hypothesis worth testing in your context, not a conclusion you're entitled to skip straight to."

That's a paradigm shift in posture more than in fact: from behavioral economics as a menu of assumptions you apply, to behavioral economics as a set of testable hypotheses you verify — especially the famous ones, precisely because their fame makes them the ones people are most tempted to skip testing.

FAQ

Does this mean choice overload isn't real?

No. Iyengar and Lepper's original research and Schwartz's synthesis are well-regarded work, and the mechanism is plausible and sometimes strong. What this experiment shows is that the effect didn't clear a significance bar in this specific context, at this scale, on this page. That's a statement about this application of the theory, not a refutation of the underlying research.

If a large sample still came back inconclusive, why not just run it again?

Because the honest read of a well-powered null isn't "try harder" — it's "the lever probably isn't plan count, at least not on its own, in this funnel." Re-running the identical test on the same page is more likely to burn traffic confirming the same result than to reveal a different one. The better next step is testing a different lever entirely, informed by what this result ruled out.

When should I trust published behavioral-economics research versus test it myself?

Treat published findings as strong hypotheses, not settled facts about your funnel. The stronger and more famous the finding, the more valuable it is to verify locally — not because the researchers were wrong, but because context, audience, and category change how (and whether) an effect shows up. Save your testing capacity for decisions with real cost if you get them wrong, and this one — restructuring a pricing page — qualifies.

Isn't publishing a null result a wasted experiment?

No — it's the opposite. A poorly-run test that comes back inconclusive wastes traffic because you learn nothing. A well-powered test that comes back inconclusive rules out a plausible, popular explanation with real confidence, which changes what you should test next. That's a return on the traffic spent, not a loss.

What would you actually recommend a founder or CMO do with this?

Don't let "well-known research says so" be the last step before a page redesign, especially for changes with a real cost to reverse. Build a testing program sized to answer the question definitively either way, and be willing to publish the null when that's what the data says.

Bottom line

Choice overload is real research worth taking seriously, but this experiment is a reminder that a landmark finding from a lab isn't a substitute for a locally-run test — at this scale, plan count on its own wasn't the lever the literature predicted it would be, and a well-powered null taught us more than a smaller test or blind trust in the research ever could.

If you're deciding whether to redesign a pricing page, a catalog, or a checkout flow around a behavioral-economics finding you've never tested in your own funnel, that's exactly the kind of decision I help teams get evidence on before they commit. If you want a second opinion on what's actually worth testing versus what's safe to assume, let's talk.

Evidence sources and free next step

The published choice-overload meta-analysis shows why the mechanism is plausible while emphasizing substantial variation across contexts. Compare this null with the A/B testing examples evidence table and sample-size guide. Then try GrowthLayer free to document a minimum detectable effect and preserve an honest null.

Share this article
LinkedIn (opens in new tab) X / Twitter (opens in new tab)
Atticus Li

Experimentation and growth leader. CXL-certified CRO practitioner, Mindworx-certified behavioral economist (1 of ~1,000 worldwide). 200+ A/B tests across energy, SaaS, fintech, e-commerce, and marketplace verticals.