Behavioral Science for Marketing
Diagnose why a behavior is not happening, choose an intervention with a defensible mechanism, recognize when it may fail, and test commercial and customer outcomes together.
- Modules
- 10
- Audit steps
- 9
- Research sources
- 49

From bias tactics to decision strategy.
This course follows a slide-by-slide audit of the original 52-page class. It keeps the practical Behavioral Audit, corrects blended mechanisms, removes unsupported stories and metrics, and adds current research on generalizability, durability, personalization, and AI-mediated choice.
- Every empirical claim links to a source and names its boundary.
- Claims are labeled established mechanism, context-dependent effect, emerging evidence, or testable hypothesis.
- Every intervention includes customer, commercial, and ethical guardrails.
- The advanced field chapter audits GrowthLayer archive records behind prior case-study claims, including negative, inconclusive, and integrity-limited results.
Evidence before effects
Treat each principle as a mechanism-level prior, then earn the right to generalize through diagnosis and measurement.
What you will be able to do
- Separate a theory, a finding, and a marketing hypothesis.
- Match claim strength to evidence strength.
- Name the population, setting, outcome, and time horizon before transfer.
Evidence and boundaries
Behavioral economics integrates psychological evidence into economic analysis of judgment and decision-making.
Evidence: established mechanismBoundary: This does not mean people are uniformly irrational or that every deviation is a mistake.
A comprehensive portfolio of 126 government RCTs produced smaller average take-up effects than the selected published literature used for comparison.
Evidence: context-dependent effectBoundary: This is evidence about two nudge units, not a universal correction factor for marketing.
Choice-architecture effects can be ineffective or counterproductive in some conditions, and current knowledge does not reliably predict every transfer.
Evidence: context-dependent effectBoundary: A useful prior still needs a local counterfactual and outcome measure.
Source-backed example: From anecdote to testable claim
Setting:A launch team sees one countdown campaign outperform its predecessor and credits urgency.
What it shows:The comparison confounds audience, offer, season, creative, and timing. Scarcity is one candidate mechanism, not the finding.
Testable application:Evidence: testable hypothesisRewrite it as: for this audience and offer, a truthful deadline may increase completed purchases versus the same page without a deadline.
Failure mode and guardrails before lift
- Failure or backfire
- A famous finding is converted into a lift forecast, sending the team toward an untested tactic and away from the actual constraint.
- Customer guardrail
- Track comprehension, access, welfare, and subgroup outcomes so a business result cannot conceal customer harm or exclusion.
- Commercial guardrail
- Predefine the primary outcome, uncertainty, unit economics, and scale decision before reading the experiment result.
- Never convert an anecdote into a causal estimate.
- Report absolute outcomes and uncertainty, not only relative lift.
- Keep observed evidence separate from the next test hypothesis.
Map the decision gap before naming a bias.
What would the stated objective imply?
State the customer objective, constraints, information set, and the rule a normative model uses as the comparison point.
What happens in the actual journey?
Measure choices, completion, downstream outcomes, timing, and variation by journey context or relevant segment.
A difference identifies a question—not a bias. Specify competing explanations and the evidence that would distinguish them.
- 01
First-party durable outcome
A credible local counterfactual on the target population and business or customer outcome, measured after the intervention ends.
Watch: Inspect uncertainty, implementation, guardrails, attrition, spillovers, and whether the outcome—not only behavior—persisted. - 02
First-party causal outcome
A credible randomized or quasi-experimental estimate on the target population, behavior, and decision-relevant outcome.
Watch: A short-run effect can decay, displace another outcome, or fail when implementation changes. - 03
Scaled trials or synthesis
Multi-site field evidence, a high-quality meta-analysis, or a large coordinated trial with relevant outcomes.
Watch: Average effects can conceal contexts, segments, nulls, harms, and implementation differences. - 04
Relevant field evidence
A well-designed study in a comparable population, decision, channel, and incentive environment.
Watch: Transfer remains a hypothesis when any material dimension changes. - 05
Direct replication
A close repeat tests whether a reported pattern survives a new sample, team, time, or geography.
Watch: Replication strengthens a pattern but does not make a different marketing application equivalent. - 06
Classic demonstration
A controlled laboratory or classroom contrast makes a candidate mechanism visible under simplified conditions.
Watch: Useful for teaching and treatment generation—not for projecting local lift, revenue, or welfare. - 07
Mechanism hypothesis
A reasoned prediction grounded in journey evidence, practitioner experience, or theory.
Watch: Label it before testing; an anecdote, pattern name, or confident story is not validation.
Full experiment decision scorecard
Pre-register the decision inputs, then record what the experiment can and cannot support.
- 01Outcome
- Pre-register
Name one primary outcome, its unit and window, and the smallest change that would alter the business decision.
Read after the testReport the estimate on that outcome and whether the result is practically useful—not only whether a threshold was crossed.
- 02Uncertainty
- Pre-register
Specify the analysis, sample plan, data-quality checks, and which heterogeneous effects are decision-relevant.
Read after the testReport the uncertainty interval, data-integrity findings, plausible alternatives, and limits on transfer.
- 03Guardrails
- Pre-register
Define safety, autonomy, accessibility, quality, and downstream measures with stop or rollback thresholds.
Read after the testCheck every guardrail, including distribution across users; a primary-outcome gain does not erase a guardrail failure.
- 04Decision rule
- Pre-register
Precommit what evidence leads to scale, revise, rerun, stop, or rollback—and who owns that call.
Read after the testApply the rule to the complete evidence packet, record exceptions, and state the next decision the result enables.
Move down the ladder to generate ideas; move up it before forecasting, scaling, or claiming an effect.
Decision lab
Claim ladder
Take one strong claim from your marketing materials and label it observation, mechanism, causal estimate, or transfer hypothesis.
Deliverable:A four-line rewrite with setting, audience, outcome, and uncertainty.
Knowledge check
What is the strongest defensible conclusion after one campaign beats the previous campaign?
Reveal the answer and reasoning
Answer: The new campaign was associated with a better result and deserves a controlled test
Without a credible counterfactual, the mechanism and causal effect remain uncertain.
Diagnose before you design
Start with a precise behavior and its context, diagnose constraints, then choose an intervention and a counterfactual.
What you will be able to do
- Define one observable target behavior.
- Map decision moments and competing explanations.
- Write a falsifiable intervention hypothesis with guardrails.
Evidence and boundaries
COM-B characterizes behavior as arising through interactions among capability, opportunity, and motivation.
Evidence: established mechanismBoundary: The model organizes diagnosis; it does not identify the causal bottleneck without evidence.
Hands-on application assistance increased aid applications and college outcomes in one randomized financial-aid study, while information alone did not produce the same pattern.
Evidence: context-dependent effectBoundary: The study combines a specific population, process, and assistance model; it is not proof that all service help converts.
The nine-step Behavioral Audit is a practitioner synthesis for translating diagnosis into an ethical experiment.
Evidence: testable hypothesisBoundary: Its value should be judged by decision quality and downstream experiments, not treated as a validated scientific scale.
Source-backed example: Financial-aid help versus more information
Setting:Families received either information or information plus direct help completing an application in a randomized study.
What it shows:The stronger intervention changed practical capability and process friction, not merely message wording.
Testable application:Evidence: testable hypothesisBefore rewriting a form, test whether users lack understanding, documents, time, confidence, or a workable next action.
Failure mode and guardrails before lift
- Failure or backfire
- The team labels a bias before diagnosing the journey, so an interface treatment misses a structural, capability, or access constraint.
- Customer guardrail
- Include accessibility, effort, error recovery, and the downstream customer outcome in the audit—not completion alone.
- Commercial guardrail
- Require a credible comparison, one decision-relevant primary metric, and explicit stop or revise criteria before launch.
- Diagnose structural constraints before adding persuasion.
- Do not use a behavioral label as a substitute for customer evidence.
- Include a no-change or simpler alternative in the test design.
Private working area. These fields have no save or submit action and are not sent by this course. Copy your notes before you leave or reload the page.
- 01
What observable action should change, by whom, where, and by when?
- 02
What do behavioral data, interviews, support logs, and accessibility findings show?
- 03
Where do attention, understanding, action, and recovery break down?
- 04
What outcome is the customer hiring this experience for, and how does the business benefit?
- 05
Which physical or psychological capability, physical or social opportunity, and reflective or automatic motivation may constrain behavior?
- 06
Which frictions obstruct the customer, and which protect consent, security, comprehension, or fit?
- 07
What mechanism could change behavior, and in which conditions should it fail or reverse?
- 08
What changes, what stays equivalent, and what counterfactual identifies impact?
- 09
What result is practically valuable, durable, safe, and worth operational cost?
Complete the diagnosis before selecting an intervention.
Decision lab
Run the nine boxes
Choose one stalled SaaS activation behavior and complete every Behavioral Audit step without naming a bias until the mechanism step.
Deliverable:A one-page audit ending in intervention, counterfactual, primary outcome, guardrails, and stop rule.
Knowledge check
Which diagnosis is actionable and testable?
Reveal the answer and reasoning
Answer: Users abandon after identity verification because recovery instructions and document requirements are unclear
It identifies a moment, behavior, and plausible capability or friction constraint that can be investigated.
COM-B and the friction ledger
Map capability, opportunity, and motivation, then classify each friction by whose outcome it protects or obstructs.
What you will be able to do
- Use the full COM-B taxonomy.
- Distinguish helpful friction from sludge.
- Map friction to customer and business outcomes.
Evidence and boundaries
COM-B divides capability into physical and psychological, opportunity into physical and social, and motivation into reflective and automatic processes.
Evidence: established mechanismBoundary: The categories overlap in practice and should organize inquiry rather than force a single label.
Making sales tax salient on grocery shelf labels changed demand in a field experiment.
Evidence: context-dependent effectBoundary: A tax display intervention does not establish that every price disclosure has the same magnitude or direction.
Friction is not inherently harmful: confirmation, comparison, consent, and security steps can protect comprehension and autonomy.
Evidence: context-dependent effectBoundary: Whether a step is protective or obstructive depends on its purpose, burden, placement, and distribution across users.
Source-backed example: A price the buyer can actually compare
Setting:A grocery field experiment displayed tax-inclusive prices before checkout rather than leaving tax less salient until purchase.
What it shows:The interface changed the physical opportunity to evaluate total cost; it did not change product quality or tax rate.
Testable application:Evidence: testable hypothesisInventory every fee, requirement, delay, and recovery path at the decision moment, then ask which ones users can anticipate.
Failure mode and guardrails before lift
- Failure or backfire
- Removing protective steps increases mistakes or fraud, while adding seller-serving friction suppresses refusal, correction, or exit.
- Customer guardrail
- Measure time, errors, accessibility, comprehension, reversals, and the effort required to refuse, correct, or leave.
- Commercial guardrail
- Balance completion against fraud, support cost, rework, refunds, and durable activation rather than optimizing clicks in isolation.
- Measure errors, reversals, support contacts, and regret alongside completion.
- Do not optimize friction away from consent or security.
- Audit who bears the burden, especially in recovery and cancellation.
Diagnose capability, opportunity, and motivation as conditions for the behavior before choosing a tactic.
- Capability
- Can the person understand and perform the action?
- Opportunity
- Do access, time, tools, and social conditions allow it?
- Motivation
- Does it feel valuable, safe, salient, and worth doing now?
Use COM-B to diagnose what must change; it does not prescribe one tactic for every context.
Decision lab
Friction ledger
List every action, wait, decision, information request, and surprise between intent and completion.
Deliverable:A table with COM-B diagnosis, customer cost, protective value, owner, evidence, and testable change.
Knowledge check
Which change is most likely helpful friction?
Reveal the answer and reasoning
Answer: Showing a clear irreversible-action confirmation with an undo window
A proportionate confirmation can prevent costly errors while keeping the choice understandable and reversible.
Risk, reference points, and framing
Use prospect theory to ask how outcomes are encoded relative to a reference point under risk, without turning parameters into universal copy rules.
What you will be able to do
- Explain reference dependence, diminishing sensitivity, and decision weighting.
- Separate risky-choice framing from generic negative wording.
- Treat loss asymmetry as contextual, not fixed.
Evidence and boundaries
Prospect theory models risky choice using reference-dependent value, diminishing sensitivity, and nonlinear weighting of probabilities.
Evidence: established mechanismBoundary: The original model concerns decisions under risk; not every gain-versus-loss sentence is a prospect-theory test.
A 19-country replication reproduced most original choice patterns but also found attenuation and heterogeneity.
Evidence: context-dependent effectBoundary: The replication supports core patterns, not one universal loss-to-gain ratio.
A certain-gain versus probabilistic-gain contrast tests the reflection or certainty pattern, not loss aversion by itself.
Evidence: established mechanismBoundary: Loss aversion requires comparison around a reference point; probability and certainty also change the choice.
Source-backed example: What the replication actually supports
Setting:Participants across 19 countries completed structured monetary choices adapted from the original prospect-theory items.
What it shows:Most theoretical contrasts replicated, while country and person-level variation remained meaningful.
Testable application:Evidence: testable hypothesisFor a renewal decision, first establish the customer reference point and uncertainty; then test gain and loss frames with identical facts.
Failure mode and guardrails before lift
- Failure or backfire
- A loss or certainty frame changes comprehension, perceived risk, or trust instead of improving a fact-equivalent decision.
- Customer guardrail
- Keep outcomes and probabilities identical, disclose material uncertainty, and measure comprehension, perceived pressure, and fit.
- Commercial guardrail
- Read choice together with retention, refunds, support demand, and trust; a short-term selection shift is not sufficient.
- Never exaggerate risk or invent a loss.
- Predefine the reference point rather than inferring it after results.
- Measure comprehension and trust as well as choice.
01 Value around a reference point
02 Probability weighting
03 Fourfold risk pattern
Combine the outcome frame with the probability range before forming a hypothesis.
Higher probability
Gain
Common pattern: risk aversionA sure gain may be preferred to a gamble with a larger possible gain.Higher probability
Loss
Common pattern: risk seekingA gamble may be preferred to accepting a sure loss.Lower probability
Gain
Common pattern: risk seekingA small chance of a large gain can support lottery-like choices.Lower probability
Loss
Common pattern: risk aversionProtection can appeal when it removes a small chance of a large loss.Choose first. Then reveal the arithmetic.
There is no scored answer. This compares two equal-expected-value choices; it does not predict what you or a customer will choose elsewhere.
Reveal expected values
Option A: 100% × $500 = $500
Option B: (50% × $1,000) + (50% × $0) = $500
Option A: 100% × −$500 = −$500
Option B: (50% × −$1,000) + (50% × $0) = −$500
Equal expected value does not make the options psychologically identical. Treat any contrast you notice as a prompt for diagnosis—not a measured effect or a marketing forecast.
- Diagnose
- What expectation, status quo, comparison, or goal sets the reference point?
- Test
- Hold the offer and probabilities constant; vary the frame without hiding information.
These are common prospect-theory patterns, not forecasts. Framing, stakes, experience, ambiguity, and the way probabilities are learned can change a decision.
Decision lab
Reference-point map
For one choice, list the current state, expected state, promised state, and competitor state. Identify which may serve as the reference point.
Deliverable:Two fact-equivalent frames plus a prediction, manipulation check, and trust guardrail.
Knowledge check
What must stay constant in a clean framing test?
Reveal the answer and reasoning
Answer: The underlying outcomes and probabilities
A framing comparison changes presentation while preserving the substantive outcomes and probabilities.
Defaults, choice, and autonomy
Use defaults and assortment design as hypotheses about effort, endorsement, and uncertainty while preserving meaningful choice.
What you will be able to do
- Explain effort, endorsement, and endowment or reference-point default mechanisms.
- Diagnose when assortment complexity may create overload.
- Separate choice uptake from user outcome.
Evidence and boundaries
A meta-analysis found that default effects varied substantially and were partly associated with effort or ease, endorsement, and endowment or reference-point mechanisms.
Evidence: context-dependent effectBoundary: Some studies found null or negative effects, and changing a choice does not guarantee a better outcome.
Automatic enrollment strongly changed retirement-plan participation and contribution patterns in one employer setting.
Evidence: context-dependent effectBoundary: The welfare implications depend on suitability of the default and the path from enrollment to long-term outcomes.
Choice overload is conditional rather than a rule that fewer options always perform better.
Evidence: context-dependent effectBoundary: Complexity, task difficulty, preference uncertainty, and decision goal moderate the effect.
Source-backed example: Default changes choice, then what?
Setting:Automatic enrollment increased participation in an employer retirement plan and clustered contributions at the preset rate and fund.
What it shows:The default reduced action requirements and may have signaled a recommendation, while also shaping the selected outcome.
- Brigitte C. Madrian and Dennis F. Shea (2001)
- Jon M. Jachimowicz, Shannon Duncan, Elke U. Weber, and Eric J. Johnson (2019)
Testable application:Evidence: testable hypothesisFor onboarding, evaluate whether the preset is safe, representative, visible, editable, and good for the user if left unchanged.
Failure mode and guardrails before lift
- Failure or backfire
- A default or reduced assortment increases uptake while steering people into poor-fit, hard-to-reverse, or insufficiently considered choices.
- Customer guardrail
- Keep the default visible, suitable, editable, and reversible, with enough comparison support for preference formation.
- Commercial guardrail
- Measure activation quality, retention, support, downgrade, and reversal instead of treating initial option share as success.
- Make defaults visible and easy to change.
- Evaluate user outcomes after the choice, not uptake alone.
- Do not hide a dominated or unsuitable option inside complexity.
Consent
- Default state
- No preselection
- Decision friction
- One informed choice
- Guardrail
- Comprehension and reversibility
Deletion
- Default state
- No destructive action
- Decision friction
- Confirm and offer undo
- Guardrail
- Error prevention
Onboarding
- Default state
- Safe, editable setup
- Decision friction
- Review consequential choices
- Guardrail
- Fit after activation
- Can people notice the choice?
- Can they understand the consequence?
- Can they reverse it without disproportionate effort?
Name who the extra effort protects.
- Comprehension before commitment
- Security and identity checks
- Protection from irreversible error
- Enough information to compare
- Hidden or hard-to-find exit
- Repeated requests after refusal
- Surprise fees late in the flow
- Unequal effort serving only the seller
Compare the paired paths, not just the entry path.
- Join
- Leave
- Accept
- Refuse
- Enable
- Disable or correct
Symmetry means proportionate visibility, comprehension, and effort—not necessarily an identical number of clicks.
A default changes what happens without action. Evaluate ease, comprehension, autonomy, and error costs together.
Decision lab
Choice architecture inventory
Document the default, option set, order, labels, comparison aid, exit, and recovery path; contrast a repeat shopper with high preference certainty and a first-time shopper with low preference certainty.
Deliverable:A redesign that reduces task difficulty where needed while keeping alternatives and consequences visible for both shopper states.
Knowledge check
When should a team reduce an option set?
Reveal the answer and reasoning
Answer: When diagnosis shows complexity or uncertainty is obstructing the target decision
Assortment size interacts with the task and decision-maker; there is no universal optimal count.
Scarcity, urgency, and credibility
Use scarcity only when a real limit changes the decision, and make the reason and consequence legible.
What you will be able to do
- Separate quantity scarcity from time scarcity.
- Identify persuasion-knowledge and trust risks.
- Design an equivalent-offer scarcity test.
Evidence and boundaries
A marketing meta-analysis found that scarcity tactics can increase purchase intentions, with effects moderated by scarcity and product conditions.
Evidence: context-dependent effectBoundary: Purchase intention is not purchase, retention, margin, or customer welfare.
Online time-scarcity promotions did not consistently outperform equivalent controls in a program of meta-analytic and experimental research.
Evidence: context-dependent effectBoundary: Credible externally grounded reasons sometimes improved responses; results still stopped short of a universal online advantage.
A truthful deadline is best treated as a testable information intervention, not as a guaranteed persuasion device.
Evidence: testable hypothesisBoundary: The hypothesis fails if buyers ignore it, distrust it, rush into poor-fit purchases, or defer until the next promotion.
Source-backed example: A reason for the clock
Setting:Online promotions with time limits were compared with controls and with limits justified by externally grounded events.
What it shows:The online setting can activate persuasion knowledge; a credible reason may matter more than visual intensity.
Testable application:Evidence: testable hypothesisState what ends, when, why, and what remains available. Test against the same offer without time pressure.
Failure mode and guardrails before lift
- Failure or backfire
- Pressure activates suspicion, delays the decision, or accelerates a poor-fit purchase that later becomes regret, cancellation, or distrust.
- Customer guardrail
- Show the real constraint, source, expiry behavior, consequence, and remaining alternatives without inventing total loss.
- Commercial guardrail
- Measure net completed value after cancellations, refunds, support, and trust—not clicks or checkout starts alone.
- No resettable clocks or invented inventory.
- Keep the underlying offer equivalent across test cells.
- Track cancellations, regret, refunds, and trust after purchase.
[Verified capacity, inventory, or deadline]
Quantity, time, capacity, eligibility, or another specific limit.
[Named system or operational owner]
A current record and a person accountable for accuracy.
[Exact time and what changes afterward]
The interface, offer, and saved work behave as described.
[What remains available to the customer]
The message explains what ends without inventing total loss.
Reject: a timer, stock count, or deadline that resets, cannot be sourced, or leaves the real consequence ambiguous.
Separate the constraint from the customer’s interpretation.
A verifiable limit in inventory, capacity, access, or production.
Current demand changes remaining availability; report the denominator and update rule.
A real deadline changes price, eligibility, fulfillment, or another stated condition.
People may treat limited availability as information about popularity, quality, or fit.
Pressure, vague provenance, or a resetting limit may instead reduce credibility.
Measure the interpretation. Do not assume that a real constraint creates urgency or trust.
Illustrative audit—not a live scarcity claim. Replace every bracketed field with verifiable operational facts.
Decision lab
Constraint proof
Choose one capacity, inventory, or deadline claim and document its operational source and customer consequence.
Deliverable:A fact-equivalent control, truth owner, expiry behavior, trust measure, and post-purchase quality guardrail.
Knowledge check
Which scarcity message is most defensible?
Reveal the answer and reasoning
Answer: Applications close Friday because reviewers start Saturday; saved drafts remain available
The constraint, timing, reason, and consequence are specific and verifiable.
Pricing context, decoys, and free offers
Treat anchors, option context, and zero price as conditional mechanisms while preserving transparent comparison.
What you will be able to do
- Distinguish anchor, compromise, and attraction effects.
- Interpret field magnitudes without folklore.
- Include monetary and nonmonetary costs in offer design.
Evidence and boundaries
Anchoring estimates can hide opposing responses across people and experimental conditions.
Evidence: context-dependent effectBoundary: An arbitrary high number is not a dependable or ethically sufficient pricing strategy.
An archival analysis of 3.6 million grocery purchases detected a modest, history-dependent attraction effect, while a field experiment across 146,851 flight searches found no general effect from a recommended decoy.
Evidence: context-dependent effectBoundary: Different products, choice sets, outcomes, and study designs prevent either result from becoming a universal pricing rule.
Zero-price experiments found a discontinuous response to free options, but free samples show heterogeneous acceleration, expansion, and cannibalization effects.
Evidence: context-dependent effectBoundary: A zero sticker price can coexist with shipping, time, data, switching, or opportunity costs.
Partitioned and drip prices can change evaluation and search by changing total-price visibility and when fees appear.
Evidence: context-dependent effectBoundary: Effects vary by presentation, surcharge, category, and regulation; transparent all-in comparison is the safety baseline.
Source-backed example: Observed shift, experimental null
Setting:Researchers found a modest association across millions of wine purchases, while a randomized decoy test across more than 140,000 flight searches found no general attraction effect.
What it shows:The contrast makes implementation, setting, study design, and customer history part of the mechanism—not footnotes after the result.
- Sean Devine, James Goulding, John Harvey, Anya Skatova, and A. Ross Otto (2025)
- Ismael Rafai, Zakaria Babutsidze, Thierry Delahaye, Nobuyuki Hanaki, and Rodrigo Acuna-Agost (2022)
Testable application:Evidence: testable hypothesisValidate the dominance structure, then test whether a third option improves comprehension and fit rather than assuming it will move target-plan share.
Failure mode and guardrails before lift
- Failure or backfire
- The option set or free offer creates confusion, hidden cost, cannibalization, or low-fit selection rather than better comparison.
- Customer guardrail
- Show mandatory total and nonmonetary costs before commitment, with truthful attributes and meaningful alternatives.
- Commercial guardrail
- Evaluate contribution margin, cost to serve, expansion, retention, refunds, and support—not target-plan share alone.
- Show mandatory total cost before commitment.
- Do not use an option designed only to confuse or misdirect.
- Measure plan fit, downgrade, refund, and support outcomes.
- AlternativealternativeLower cost · Core outcome
- TargettargetHigher cost · Expanded outcome
- Illustrative decoydecoyHighest cost · Lower than target
Mechanism map
Start with the diagnostic condition; do not infer a mechanism from layout alone.
- Starting valueAnchor
- An initial number becomes an input to a later estimate. It need not be a credible prior price.
- Comparison baselineReference price
- An expected, remembered, or displayed price becomes the baseline for judging the current price.
- PositionCompromise
- An option becomes intermediate on meaningful attributes, not merely centered on screen.
- AttentionSalience
- A feature becomes easier to notice or compare. Attention is not dominance.
- RelationAsymmetrically dominated decoy
- The target dominates the decoy; the competitor does not. Impact still requires a test.
Decoy configuration check
For this diagram only: lower cost and higher benefit are preferred.
- PassTarget dominates decoy. No worse on either dimension; better on at least one.
- PassCompetitor does not dominate decoy. The alternative and decoy retain a trade-off.
- PassPrimary options retain a trade-off. Neither primary option beats the other on both dimensions.
- Trade-off option: it beats the target on one dimension.
- Universally dominated: the competitor also dominates it.
- Trivial set: the target already dominates the competitor.
Schematic choice set. The plotted logic uses lower cost and higher benefit; real decisions may include omitted attributes. Any change in choice remains an empirical question.
- Advertised priceShown at selection
- $72
- Mandatory service feeMust also be visible now
- $8
- Time, data, or switching costNonmonetary but material
- Name it
- Comparable total
- $80 + nonmonetary costs
Which costs remain?
Include shipping, time, data, switching, attention, and opportunity cost.
What changed later?
Separate trial, accelerated purchase, category expansion, and cannibalization.
When was it knowable?
Test comprehension, comparison, and trust when mandatory fees appear before the choice.
Choose the model by the learning and risk it creates.
- Customer job
- Experience a bounded unit before buying more.
- Business risk
- Cannibalization or attracting use without learning.
- Measure next
- Incremental category purchase and repeat behavior.
- Customer job
- Test the full workflow long enough to judge fit.
- Business risk
- Deadline-driven activation without durable value.
- Measure next
- Activation quality, informed conversion, and retention.
- Customer job
- Use a durable core before needs expand.
- Business risk
- Service cost without a credible expansion path.
- Measure next
- Value realization, upgrade trigger, and cost to serve.
- Customer job
- Reduce implementation risk with a scoped commitment.
- Business risk
- Custom work that does not generalize or expand.
- Measure next
- Decision quality, implementation cost, and expansion.
This is a strategy-selection map, not evidence that one model wins across contexts.
Illustrative receipt. The numbers are not observed performance; the lesson is total-cost visibility.
Decision lab
Total-cost option map
Map each option on price, mandatory fees, time, data, reversibility, and customer outcome. Mark every anchor and dominated relationship.
Deliverable:A transparent comparison plus hypotheses for choice, comprehension, margin, and regret.
Knowledge check
What does the recent large purchase analysis justify?
Reveal the answer and reasoning
Answer: Context effects can appear in real choices but may be modest and person-dependent
The evidence supports detectable conditional context effects, not a universal merchandising recipe.
Motivation over time and durable effects
Design progress and commitment around a real goal, then measure whether behavior and outcomes survive the intervention.
What you will be able to do
- Connect present bias to a specific timing conflict.
- Design honest progress and reversible commitment.
- Measure persistence and mechanism after treatment ends.
Evidence and boundaries
Loyalty-program studies found effort accelerated as people approached a reward, and an endowed head start increased completion when requirements were equivalent.
Evidence: context-dependent effectBoundary: Clear progress toward one reward is not evidence that every progress bar improves durable outcomes.
A 2026 meta-analysis found present-bias estimates differed between money, effort, and other nonmonetary rewards.
Evidence: context-dependent effectBoundary: A population average cannot select a message, incentive, or reminder schedule for a specific journey.
A review found demand for commitment devices and their effectiveness vary, while commitment can also impose costs when people misforecast or need flexibility.
Evidence: context-dependent effectBoundary: The device, population, time horizon, exit terms, and alternative supports determine whether commitment helps.
Pooled evidence from 38 home-energy field experiments found that part of the reduction persisted after treatment and was consistent with technology adoption.
Evidence: context-dependent effectBoundary: This mechanism is context-specific and does not imply that all behavioral interventions form habits.
Large multi-arm field studies can compare many plausible messages under one shared outcome and population.
Evidence: context-dependent effectBoundary: A winner ranking in vaccination reminders is not a reusable ranking for commercial messages.
Source-backed example: Persistence through changed infrastructure
Setting:Home-energy studies followed outcomes beyond active norm feedback and after resident moves.
What it shows:Part of the effect remained with the home, consistent with durable technology adoption rather than a message-only habit story.
Testable application:Evidence: testable hypothesisFor onboarding, distinguish temporary prompts from lasting changes in setup, skill, or environment and measure both after prompts stop.
Failure mode and guardrails before lift
- Failure or backfire
- A progress cue or commitment increases exposure-period activity while creating pressure, lock-in, low-quality completion, or no durable outcome.
- Customer guardrail
- Show only earned progress, make commitment informed and voluntary, and preserve a proportionate exit or revision path.
- Commercial guardrail
- Measure durable activation, retention, incremental cost, and post-treatment value rather than prompt-period clicks.
- Do not imply progress that has not occurred.
- Make commitment optional, comprehensible, and easy to exit.
- Measure post-treatment outcomes and unintended lock-in.
- T0During treatment
Immediate target behavior
Did exposure change action? - T1After treatment
Behavior without the prompt
Did the effect persist? - T2Later outcome
Customer and business result
Was persistence useful and safe?
Structural adoption: Track setup, skill, environment, or technology changes that could carry the behavior forward.
Temporary response: behavior fades when reminders, rewards, or social feedback stop.
A persistent outcome does not prove habit. Measure competing mechanisms and keep a credible comparison over time.
Decision lab
Three-horizon scorecard
Define immediate, intermediate, and post-treatment outcomes for one lifecycle intervention.
Deliverable:A measurement plan with mechanism indicators, holdout or withdrawal design, and a date for the keep-revise-stop decision.
Knowledge check
What is the cleanest evidence of durability?
Reveal the answer and reasoning
Answer: The target outcome remains after the intervention ends, with a design that addresses alternative causes
Durability concerns downstream behavior after treatment, not exposure-period response alone.
Advanced lab: AI, personalization, and agency
Audit what the system knows, chooses, explains, and lets the user change before optimizing recommendation acceptance.
What you will be able to do
- Map data provenance and inference risk.
- Design choice and control autonomy while anticipating reactance.
- Test acceptance alongside fit, correction, and trust.
Evidence and boundaries
Personalized advertising can increase relevance while covert data collection increases felt vulnerability and can reduce response.
Evidence: context-dependent effectBoundary: The studies predate generative AI and do not quantify every modern personalization context.
Three online vacation-choice experiments found greater recommendation acceptance when users received multiple recommendations and some control over recommendation format.
Evidence: emerging evidenceBoundary: Multiple options can add overload in harder tasks, and acceptance is not the same as decision quality.
Psychological reactance research describes how a perceived threat to freedom can motivate counterarguing, resistance, or efforts to restore the threatened choice.
Evidence: established mechanismBoundary: Not every refusal, correction, or low response is reactance; perceived freedom, threat, context, and individual differences matter.
Three controlled song experiments found that randomized or perturbed recommendation signals shifted willingness-to-pay judgments beyond measured preferences.
Evidence: context-dependent effectBoundary: Song valuations in experiments are not completed purchases, durable satisfaction, or evidence that recommendation influence improves welfare.
Large-scale interface audits and experiments document manipulative design patterns that can impair autonomous choice.
Evidence: context-dependent effectBoundary: Pattern labels do not replace design-specific causal evidence or legal advice.
Digital choice interfaces can face explicit restrictions on manipulative design in applicable jurisdictions.
Evidence: established mechanismBoundary: Scope depends on service, conduct, and jurisdiction; obtain qualified legal review.
Source-backed example: One answer versus a choice with control
Setting:Online experiments compared a single algorithmic recommendation with multiple recommendations and user control over the format.
What it shows:Choice and control autonomy increased acceptance in the simulated vacation decisions.
Testable application:Evidence: testable hypothesisFor an AI plan selector, expose assumptions, offer alternatives, let users correct inputs, and compare acceptance with fit and reversal.
Failure mode and guardrails before lift
- Failure or backfire
- A persuasive recommendation is accepted despite poor fit, opaque inference, reduced autonomy, or unequal errors across groups.
- Customer guardrail
- Expose material inputs and rationale, offer meaningful alternatives, allow correction, and preserve a no-personalization path.
- Commercial guardrail
- Read acceptance with fit, willingness to pay, reversals, support, retention, trust, and subgroup outcomes.
- Disclose material data use and inferred attributes.
- Provide meaningful alternatives and correction controls.
- Never optimize acceptance without fit, reversal, trust, and subgroup checks.
Low comparison effort, but the ranking carries more of the decision. Expose why and an exit.
Preserves comparison. Audit set size, meaningful differences, ordering, and omitted options.
Let the person set goals, filters, or constraints and revise the resulting choice set.
Known or inferred signals
- Declared goals
- Observed behavior
- Uncertain inference
Recommendations shown
- Primary recommendation
- Meaningful alternative
- No personalization
User controls
- View the rationale
- Correct inputs
- Choose alternatives
- Turn personalization off
Expose the rationale → let the user correct inputs → preserve alternatives → learn from the outcome.
Acceptance is incomplete evidence. Also measure fit, correction, reversal, trust, and subgroup outcomes.
Decision lab
Recommendation-system audit
Trace one AI recommendation from collected data through inference, ranking, explanation, user correction, choice, and downstream outcome.
Deliverable:A test matrix for single versus multiple recommendations, visible rationale, edit controls, and a no-personalization baseline.
Knowledge check
Which metric set best protects agency?
Reveal the answer and reasoning
Answer: Acceptance, fit, corrections, reversals, trust, and subgroup outcomes
Acceptance can rise even when the recommendation is poor or manipulative; downstream quality and control complete the picture.
What the archive can support—and what it cannot.
This is a public-safe audit of historical experiments stored in the GrowthLayer evidence system and interpreted alongside current behavioral-science research. GrowthLayer stores and audits the evidence; it did not independently run every historical test.
The practical standard is intentionally strict: treatment and outcome can be historical facts, while the proposed psychological mechanism remains a hypothesis unless the design measured or isolated it.
The source-row labels and canonical evidence classifications use different units. Keep the denominators visible so a record label cannot quietly become a portfolio win rate.
- 01 · Ingest142source records
- 3duplicate rows
- 02 · Adjudicate139canonical experiments
Source-row outcome labels
- Positive
- 41
- Negative
- 32
- Inconclusive
- 69
These are inherited record labels—not independently recomputed outcomes and not canonical experiment-level winners or losers.
Canonical evidence classifications
- Quantitative-ready
- 80
- Narrative-ready
- 37
- Review-required
- 22
Readiness describes what each canonical record can responsibly support. It does not encode whether a treatment helped, harmed, or produced an uncertain result.
No pooling. Do not average lifts, combine p-values, or infer a cross-experiment effect without compatible outcomes, units, assignment checks, and a prespecified synthesis method.
Wins, losses, non-detections, and records that must stop the story.
Public ranges protect prior organizations and clients. The interpretation beside each result is narrower than the headline because accuracy is more valuable than a cleaner anecdote.
01Source-reported resultA pricing architecture package moved completed transactionsReported +15% to +20%; p < .01
- What changed
- The variant combined product-first ordering, visibility of all price points, and a plan callout.
- Primary readout
- Completed transactions
- Public-safe scope
- 30,000–35,000 sessions · Four to five weeks
- Archive read
- The historical record reports a positive result for the package on the transaction outcome.
This bundled pricing-page package outperformed its control in this recorded experiment.
The archive cannot attribute the result to price order, a decoy, the callout, or any single behavioral mechanism.
Evidence boundary: Treat the package as one intervention. A follow-up component test would be required to isolate which element mattered.
02Source-reported resultMaking the cheapest price prominent moved the primary outcome downReported -10% to -5%; no public p-value
- What changed
- The variant showed all price points and made the cheapest option the strongest visual price anchor.
- Primary readout
- Funnel starts
- Public-safe scope
- 25,000–30,000 sessions · Four to five weeks
- Archive read
- The historical record labels the variant negative; the public projection does not include the arm-level table needed to check the estimate.
This implementation did not support shipping stronger cheapest-price prominence on this page.
It does not show that anchoring caused the result or that prominent low prices generally reduce conversion.
Evidence boundary: Visual prominence, copy, page hierarchy, and the specific choice set travel together in this treatment.
03Directional evidenceA checkout countdown produced directional negative evidenceReported -5% to 0%; p ≈ .08
- What changed
- The variant added a rate-lock countdown during a considered checkout.
- Primary readout
- Checkout completion
- Public-safe scope
- 5,000–10,000 sessions · About four weeks
- Archive read
- The estimate leaned negative but did not clear the conventional two-sided .05 threshold in the archived summary.
The result was sufficient evidence not to roll out this countdown unchanged in this context.
It does not establish reactance as the cause or show that accurate urgency cues fail in other journeys.
Evidence boundary: Decision policy and mechanism attribution are separate questions: a cautious product decision can coexist with uncertainty about why.
04Directional evidenceA progress indicator produced an exploratory positive readReported +5% to +10%; p ≈ .09
- What changed
- A three-arm test compared progress treatments during product selection.
- Primary readout
- Completed transactions
- Public-safe scope
- 30,000–35,000 sessions · Seven to nine weeks
- Archive read
- The archive labels a positive variant, while its rounded p-value supports an exploratory rather than confirmatory interpretation.
The treatment generated a follow-up-worthy signal in this funnel stage.
It is not a verified replication and does not show that progress indicators reliably increase completion.
Evidence boundary: A three-arm analysis also requires a prespecified comparison and multiplicity policy, neither of which is available in the public projection.
05Non-detectionA linear-versus-radial stepper test showed why magnitude without precision misleadsReported -20% to -10%; p ≈ .38
- What changed
- The variant changed the progress treatment from a radial display to a linear one.
- Primary readout
- Completed transactions
- Public-safe scope
- 2,000–3,000 sessions · About six weeks
- Archive read
- A large negative point estimate coexisted with weak evidence against the null in a small sample.
The experiment did not resolve whether the linear-versus-radial change affected the primary outcome.
The point estimate alone cannot justify a stable negative effect or a claim that the treatment had no effect.
Evidence boundary: The missing per-arm counts and interval prevent independent precision, power, or sample-ratio checks.
06Non-detectionReducing fourteen plans to nine did not produce a clear detectionReported 0% to +5%; p ≈ .14
- What changed
- The variant reduced a product set from fourteen plans to nine.
- Primary readout
- Completed transactions
- Public-safe scope
- 250,000–300,000 sessions · About eight weeks
- Archive read
- The large traffic count did not turn the observed estimate into a clear detection under the archived read.
This experiment did not provide clear evidence that this reduction improved the primary outcome.
It neither establishes no effect nor overturns the conditional choice-overload literature.
Evidence boundary: Traffic volume is not power by itself; baseline rate, allocation, variance, minimum detectable effect, and stopping policy are also required.
07Integrity-limitedConflicting navigation metadata blocks the attractive storyMetadata conflict: result withheld
- What changed
- The variant consolidated competing mobile navigation systems.
- Primary readout
- Orders
- Public-safe scope
- 200,000–250,000 sessions · About five weeks
- Archive read
- The recorded winner field conflicts with the variation-level conversion ordering and the archived p-value does not support the headline.
The record requires reconciliation against the experiment platform before any outcome claim is published.
No lift, winner, Hick-law effect, or causal navigation claim is supportable from the conflicting projection.
Evidence boundary: A large sample cannot rescue internally inconsistent metadata. Integrity checks precede interpretation.
08Integrity-limitedA four-change landing page failed the allocation checkSRM flag: causal read blocked
- What changed
- One variant bundled four landing-page changes.
- Primary readout
- Completed transactions
- Public-safe scope
- 65,000–70,000 sessions · Four to five weeks
- Archive read
- The archive flags a sample-ratio mismatch, so the allocation does not match the expected randomization split.
The correct action is to diagnose instrumentation or allocation before interpreting the outcome.
No causal effect, component attribution, or reusable lift is supportable while the mismatch remains unresolved.
Evidence boundary: Even without the mismatch, four simultaneous changes would identify only the package, not which component mattered.
What changed in the evidence—not just what got published.
Recent work is less useful as a stream of new tactics than as a correction to how marketers transfer evidence: replication criteria matter, hypothetical magnitude is noisy, interfaces change welfare as well as clicks, and AI autonomy depends on the job and context.
- 01274 replication attempts across 164 papers
Published effects often shrink under direct replication
Replication success depends on the criterion, and the project reported a lower median effect size in replications than in original studies.
Boundary: The project sampled selected positive claims from earlier publications; its aggregate rates do not forecast a specific marketing intervention.
- 0220 preregistered hypothetical experiments; n = 16,114
Hypothetical tests may help with direction, not dependable magnitude
In this study set, reconstructed scenarios often recovered the direction of field effects while magnitude estimates varied substantially.
Boundary: The comparison covered five field interventions; stated or imagined choices remain weaker than observed behavior in the target setting.
- 03Three vaccination field RCTs; n = 314,824
Transfer can occur for one component and fail for another
Selected reminder and ownership-language elements transferred, while several additions chosen from intentions or predictions had no detectable benefit.
Boundary: The results concern booster uptake in one health-system setting and should inform a transfer test, not an automatic rollout.
- 04Retail-platform field experiment involving 1.6 million consumers
Choice overload can follow a conditional, non-linear pattern
Purchase probability followed an inverted-U pattern as recommendation-set size increased, with search initiation explaining much of the later decline.
Boundary: One recommender system does not determine the best assortment size for another product, audience, or decision journey.
- 05Emerging working paper; 23,685 cookie choices
Small consent frictions can materially alter recorded choices
Hiding rejection behind an extra step reduced rejection by 17.8 percentage points in survey visits and 9.4 points during organic browsing; subtler ordering and visual treatments moved behavior less.
Boundary: The study is a working paper in one browsing environment; behavioral movement does not establish informed consent or consumer welfare.
- 0667 completed subscriptions across 34 news sites in four countries
Cancellation asymmetry can be audited as product behavior
The audit documented cancellation journeys requiring many steps and, in some cases, a phone call.
Boundary: An observational interface audit does not estimate causal retention, customer welfare, or the legal status of every flow.
- 07Marketplace field experiment plus an emerging randomized study, n = 1,608
Price presentation can change quantity, quality, and the value of decision support
Upfront mandatory fees changed purchase quantity and tier selection in a ticket marketplace; in a newer randomized study, a digital shopping assistant reduced the reported incidence of overpayment by 63%.
Boundary: The newer result is a working paper using endowment-funded gift-card purchases, and neither marketplace estimate transfers automatically to subscriptions or ordinary retail.
- 08Randomized recommendation studies plus one ad field study and four experiments
AI nudges depend on transparency, preference match, autonomy, and context
Transparency badges and preference match shaped choice confidence in recommendation studies, while scarcity increased perceived benefits and acceptance of higher-autonomy assistance in a separate study set.
Boundary: Student shopping tasks, clicks, and hypothetical acceptance do not establish objective decision quality or lasting adoption, and the findings do not justify fabricated scarcity.
From principle to decision in seven records.
A mature behavioral strategy is a chain of auditable records. If one link is missing, the claim should get weaker—not more eloquent.
- 01Decision
Name who must decide what, by when, and what happens if evidence stays uncertain.
- 02Observed behavior
Define the consequential action, denominator, time window, and current journey.
- 03Competing diagnosis
Write at least two explanations that predict different observable patterns.
- 04Treatment contract
State exactly what changes, what stays fixed, and whether the treatment is a bundle.
- 05Evidence plan
Precommit the primary metric, guardrails, allocation, sample logic, and stopping rule.
- 06Integrity read
Check instrumentation, allocation, contamination, missingness, and decision-rule fit before lift.
- 07Decision and reuse
Store result, evidence strength, business decision, mechanism status, and next test separately.
Inspect the evidence, not just the lesson.
The registry distinguishes study design from claim strength and keeps each limitation next to the citation. Links go to original papers, publisher records, or official sources.
Browse all 49 sources and their limitations
official recognition · Field definition The Prize in Economic Sciences 2002: Foundations for Behavioral and Experimental Economics
The Royal Swedish Academy of Sciences · 2002
Use: Foundations
Limit: A prize summary, not a systematic review or an estimate of intervention effects.
review · Framework development and evaluation The behaviour change wheel: A new method for characterising and designing behaviour change interventions
Susan Michie, Maartje M. van Stralen, and Robert West · 2011
Use: Behavioral Audit, COM-B
Limit: A general intervention-design framework; marketing applications still require local diagnosis and testing.
foundational paper · Foundational experiments and model Prospect Theory: An Analysis of Decision under Risk
Daniel Kahneman and Amos Tversky · 1979
Use: Risk and reference points
Limit: Original stylized risky-choice studies; parameters and effects are not universal constants.
multi-country replication · Preregistered 19-country replication Replicating patterns of prospect theory for decision under risk
Kai Ruggeri et al. · 2020
Use: Risk and reference points, Evidence boundaries
Limit: Replicated structured monetary choices with attenuation and cross-country heterogeneity; not a marketing lift forecast.
program evaluation · 126 RCTs covering 23 million people RCTs to Scale: Comprehensive Evidence From Two Nudge Units
Stefano DellaVigna and Elizabeth Linos · 2022
Use: Foundations, Experimentation
Limit: Government nudge-unit portfolio; channel mix and outcomes differ from commercial marketing.
review · Current generalizability review Generalizability of choice architecture interventions
Barnabas Szaszi, Daniel G. Goldstein, Dilip Soman, et al. · 2025
Use: Foundations, Experimentation
Limit: Synthesizes heterogeneity and transfer problems; it does not identify one best intervention.
meta-analysis · 58 studies, pooled n = 73,675 When and why defaults influence decisions: a meta-analysis of default effects
Jon M. Jachimowicz, Shannon Duncan, Elke U. Weber, and Eric J. Johnson · 2019
Use: Choice architecture
Limit: Substantial heterogeneity; defaults can be ineffective or negative and can change choices without improving outcomes.
meta-analysis · 99 observations across 7,202 participants Choice overload: A conceptual review and meta-analysis
Alexander Chernev, Ulf Bockenholt, and Joseph Goodman · 2015
Use: Choice architecture
Limit: Overload is conditional on complexity, task difficulty, preference uncertainty, and decision goal.
natural experiment · Natural field evidence The Power of Suggestion: Inertia in 401(k) Participation and Savings Behavior
Brigitte C. Madrian and Dennis F. Shea · 2001
Use: Choice architecture
Limit: One employer retirement context; participation and contribution choices are not consumer conversion.
field experiment · Field experiment and observational evidence Salience and Taxation: Theory and Evidence
Raj Chetty, Adam Looney, and Kory Kroft · 2009
Use: COM-B and friction, Pricing
Limit: Grocery tax salience and alcohol-tax evidence; transfer to other prices or interfaces must be tested.
field experiment · Randomized field experiment The Role of Application Assistance and Information in College Decisions
Eric P. Bettinger, Bridget Terry Long, Philip Oreopoulos, and Lisa Sanbonmatsu · 2012
Use: Behavioral Audit, COM-B and friction
Limit: Financial-aid application context with hands-on assistance; not evidence that copy simplification alone is sufficient.
field experiment · Household field experiment The Constructive, Destructive, and Reconstructive Power of Social Norms
P. Wesley Schultz et al. · 2007
Use: Social influence
Limit: Residential energy setting; descriptive norms can move low performers in the unwanted direction without an injunctive cue.
meta-analysis · Cross-study analysis of norm interventions The critical role of second-order normative beliefs in predicting energy conservation
Jon M. Jachimowicz, Oliver P. Hauser, Julia D. OBrien, Erin Sherman, and Adam D. Galinsky · 2018
Use: Social influence
Limit: Energy-conservation evidence; the relevant reference group and beliefs must be measured locally.
meta-analysis · Marketing scarcity meta-analysis Scarcity tactics in marketing: A meta-analysis of product scarcity effects on consumer purchase intentions
Rebecca Hamilton et al. · 2022
Use: Scarcity and credibility
Limit: Focuses largely on purchase intentions; scarcity source, product, and message conditions moderate effects.
meta-analysis · Meta-analytic and experimental evidence Examining the Efficacy of Time Scarcity Marketing Promotions in Online Retail
Jillian Hmurovic, Cait Lamberton, and Kelly Goldsmith · 2023
Use: Scarcity and credibility
Limit: Online time limits did not consistently outperform controls; credible exogenous reasons were an important moderator.
meta-analysis · Theory-guided anchoring meta-analysis Recovering the Anchoring of Economic Valuations
Liang Guo · 2024
Use: Pricing and context
Limit: Shows opposing and moderated treatment effects; an arbitrary number is not a dependable pricing tactic.
observational study · 3.6 million observed purchases How decoy options ferment choice biases in real-world consumer decision-making
Sean Devine, James Goulding, John Harvey, Anya Skatova, and A. Ross Otto · 2025
Use: Pricing and context
Limit: A single grocery category; the average preference shift was modest and varied with personal purchase history.
field experiment · 146,851 flight-search sessions No evidence of attraction effect among recommended options: A large-scale field experiment on an online flight aggregator
Ismael Rafai, Zakaria Babutsidze, Thierry Delahaye, Nobuyuki Hanaki, and Rodrigo Acuna-Agost · 2022
Use: Pricing and context
Limit: One flight aggregator and recommended-options implementation; a null in this setting does not rule out every context effect.
foundational paper · Controlled zero-price experiments Zero as a Special Price: The True Value of Free Products
Kristina Shampanier, Nina Mazar, and Dan Ariely · 2007
Use: Pricing and free offers
Limit: Specific low-cost product choices; free does not erase time, shipping, privacy, or switching costs.
field experiment · Two consumer-brand field experiments The Effects of Free Sample Promotions on Incremental Brand Sales
Kapil Bawa and Robert Shoemaker · 2004
Use: Pricing and free offers
Limit: Effects varied widely between two brands and included acceleration, expansion, and cannibalization.
quasi-experiment · 55,000+ samples in an e-commerce field study Commercializing the Package Flow: Cross-Sampling Physical Products Through E-Commerce Warehouses
Brian Rongqing Han, Leon Yang Chu, Tianshu Sun, and Lixia Wu · 2025
Use: Pricing and free offers
Limit: Six brands on one platform; profitability depends on sample, distribution, and incremental-spend costs.
meta-analysis · 149 observations across 27 papers When Does Partitioned Pricing Lead to More Favorable Consumer Preferences?: Meta-Analytic Evidence
Ajay T. Abraham and Rebecca W. Hamilton · 2018
Use: Pricing and free offers
Limit: Responses vary with total-price visibility, surcharge magnitude, benefit, and category; preference is not welfare.
controlled experiment · Controlled market experiment Drip pricing and its regulation: Experimental evidence
Alexander Rasch, Miriam Thone, and Tobias Wenzel · 2020
Use: Pricing and free offers
Limit: Experimental market; legal obligations and buyer expectations vary by jurisdiction and category.
field experiment · Loyalty-program field and laboratory evidence The Goal-Gradient Hypothesis Resurrected: Purchase Acceleration, Illusionary Goal Progress, and Customer Retention
Ran Kivetz, Oleg Urminsky, and Yuhuang Zheng · 2006
Use: Motivation over time
Limit: Progress can increase effort near a clear reward; it does not prove every progress bar causes durable behavior.
field experiment · Car-wash loyalty field experiment The Endowed Progress Effect: How Artificial Advancement Increases Effort
Joseph C. Nunes and Xavier Dreze · 2006
Use: Motivation over time
Limit: A specific completed-stamps design; perceived head starts must still leave requirements truthful and equivalent.
meta-analysis · 86-study meta-analysis A Meta-Analysis of Quasi-Hyperbolic Discounting
Stephen L. Cheung, Agnieszka Tymula, and Xueting Wang · 2026
Use: Motivation over time
Limit: Discounting estimates vary by reward domain and selective reporting; they do not prescribe a single reminder cadence.
review · Economics review Commitment Devices
Gharad Bryan, Dean Karlan, and Scott Nelson · 2010
Use: Motivation over time
Limit: Commitment can help sophisticated users but can also impose costs; easy exit and informed consent matter.
field experiment · Preregistered multi-arm field megastudy A 680,000-person megastudy of nudges to encourage vaccination in pharmacies
Katherine L. Milkman et al. · 2021
Use: Experimentation, Motivation over time
Limit: Vaccination reminders in a pharmacy context; rankings and magnitudes need not transfer to marketing funnels.
field experiment · 38 natural field experiments Do The Effects of Nudges Persist? Theory and Evidence from 38 Natural Field Experiments
Alec Brandon, Paul J. Ferraro, John A. List, Robert D. Metcalfe, Michael K. Price, and Florian Rundhammer · 2026
Use: Experimentation, Motivation over time
Limit: Home-energy intervention; persistence was consistent with technology adoption and should not be generalized to all nudges.
controlled experiment · Field evidence and three experiments Unraveling the personalization paradox: The effect of information collection and trust-building strategies on online advertisement effectiveness
Elizabeth Aguirre, Dominik Mahr, Dhruv Grewal, Ko de Ruyter, and Martin Wetzels · 2015
Use: AI and personalization
Limit: Advertising studies predate generative AI; findings concern overt versus covert data collection and felt vulnerability.
online experiment · Three online recommendation experiments Let me decide: Increasing user autonomy increases recommendation acceptance
Lior Fink, Leorre Newman, and Uriel Haran · 2024
Use: AI and personalization
Limit: Simulated vacation choices; multiple recommendations can also add complexity in other tasks.
review · Psychological reactance review Understanding Psychological Reactance: New Developments and Findings
Christina Steindl, Eva Jonas, Sandra Sittenthaler, Eva Traut-Mattausch, and Jeff Greenberg · 2015
Use: AI and personalization, Choice architecture
Limit: Synthesizes varied settings and outcomes; a freedom threat does not imply that every refusal, correction, or low conversion is reactance.
controlled experiment · Three controlled recommendation experiments Effects of Online Recommendations on Consumers’ Willingness to Pay
Gediminas Adomavicius, Jesse C. Bockstedt, Shawn P. Curley, and Jingjing Zhang · 2018
Use: AI and personalization
Limit: Digital-song willingness-to-pay judgments in experiments; stated valuations are not completed purchases or a welfare measure.
observational study · Large-scale observational crawl Dark Patterns at Scale: Findings from a Crawl of 11K Shopping Websites
Arunesh Mathur et al. · 2019
Use: COM-B and friction, Ethics
Limit: Observed interface patterns, not causal estimates of every pattern or a substitute for legal review.
online experiment · Two experimental studies Shining a Light on Dark Patterns
Jamie Luguri and Lior Jacob Strahilevitz · 2021
Use: Ethics, AI and personalization
Limit: Online experiments on dark-pattern exposure; effects and legal standards vary by design and population.
official guidance · International practice review Fixing Frictions: Sludge Audits Around the World
OECD · 2024
Use: COM-B and friction, Ethics
Limit: Public-policy cases and audit guidance; commercial implementation needs local customer and legal review.
official guidance · International enforcement sweep Results of Review of Dark Patterns Affecting Subscription Services and Privacy
US Federal Trade Commission, ICPEN, and GPEN · 2024
Use: Ethics, AI and personalization
Limit: A regulatory sweep, not a causal study; applicability depends on jurisdiction and facts.
regulation · Binding EU regulation Regulation (EU) 2022/2065: Digital Services Act
European Parliament and Council of the European Union · 2022
Use: Ethics, AI and personalization
Limit: Legal scope is service-, conduct-, and jurisdiction-specific; this course is not legal advice.
multi-country replication · 274 attempted replications; 55.1% were significant in the original direction Investigating the replicability of the social and behavioural sciences
Andrew H. Tyner et al. · 2026
Use: Evidence boundaries, Experimentation
Limit: The project sampled positive claims published from 2009 to 2018; median Pearson’s r fell from 0.25 to 0.10 and success ranged from 28.6% to 74.8% across criteria, so no single rate forecasts a specific intervention.
online experiment · 20 preregistered hypothetical experiments, n = 16,114 Hypothetical nudges provide directional but noisy estimates of real behavior change
Linnea Gandhi, Anoushka Kiyawat, Colin Camerer, and Duncan J. Watts · 2025
Use: Evidence boundaries, Experimentation
Limit: The scenarios reconstructed five field experiments and usually recovered direction in this set, but magnitude estimates varied widely; hypothetical choices are not dependable field-effect forecasts.
field experiment · Three preregistered field RCTs, n = 314,824; selected reminders added 1.13–1.93 percentage points Field testing the transferability of behavioural science knowledge on promoting vaccinations
Silvia Saccardo et al. · 2024
Use: Evidence boundaries, Experimentation, Motivation over time
Limit: COVID-19 booster messages in one health-system setting; reminders and ownership language transferred, while several intention- or prediction-selected additions had no detectable benefit.
field experiment · Platform field experiment involving 1.6 million consumers The Choice Overload Effect in Online Recommender Systems
Xiaoyang Long et al. · 2024
Use: Choice architecture, Evidence boundaries
Limit: One retail platform and retargeting recommender implementation; up to 64% of the later purchase decline was attributed to fewer consumers starting search, and patterns varied by category, price, and timing.
field experiment · Emerging NBER working paper; 563 users and 23,685 cookie choices Designing Consent: Choice Architecture and Consumer Welfare in Data Sharing
Chiara Farronato, Andrey Fradkin, and Tesary Lin · 2025
Use: Choice architecture, Ethics
Limit: A working paper in one browsing environment; hiding rejection reduced rejection by 17.8 percentage points in survey visits and 9.4 points during organic browsing, but welfare conclusions rely partly on modelling and do not replace consent requirements.
observational study · 67 completed subscriptions across 34 news sites in four countries Staying at the Roach Motel: Cross-Country Analysis of Manipulative Subscription and Cancellation Flows
Ashley Sheil, Gunes Acar, Hanna Schraffenberger, Raphael Gellert, and David Malone · 2024
Use: COM-B and friction, Ethics
Limit: Some cancellation flows required up to ten steps and four required a call; this observational audit does not estimate causal conversion, retention, welfare, or the legal status of every design.
controlled experiment · Emerging working paper; randomized marketplace experiment, n = 1,608 Measuring and Mitigating Drip Pricing Overcharge: Evidence from an Online Marketplace Experiment with a Digital Shopping Assistant
Benjamin Lu, Daniel Markovits, Andrew Miller, and Rory Van Loo · 2025
Use: Pricing and free offers, AI and personalization, Ethics
Limit: A working paper using endowment-funded gift-card purchases; drip pricing raised average prices paid by up to about 10% and an assistant reduced overpayment incidence by 63%, but effects may differ in other markets.
field experiment · Large-scale ticket-marketplace field experiment Price Salience and Product Choice
Tom Blake, Sarah Moshary, Kane Sweeney, and Steve Tadelis · 2021
Use: Pricing and free offers, Ethics
Limit: One ticket marketplace with product-quality and seller responses; quality substitution accounted for at least 28% of the reported revenue decline, which is not a universal estimate for other categories.
online experiment · Randomized recommendation experiments, n = 837 AI nudging and decision quality: Evidence from randomized experiments in online recommendation setting
Yuxiao Luo, Nanda Kumar, and Adel Yazdanmehr · 2025
Use: AI and personalization, Evidence boundaries
Limit: College-student shopping experiments; badge transparency and preference match shaped choice confidence, which is not the same as objective decision quality or completed purchases.
field experiment · One Google Ads field study and four online experiments with about 2,000 people Consumer Acceptance of High-Autonomy AI Assistants Is Driven by Perceived Benefits in Online Shopping Settings Characterized by Scarcity
Darius-Aurel Frank, Michal Folwarczny, and Tobias Otterbring · 2026
Use: Scarcity and credibility, AI and personalization
Limit: Scarcity increased perceived benefits and acceptance of high-autonomy assistance in this study set, but ad clicks and hypothetical choices do not establish adoption or justify manufactured scarcity.
Build an evidence-led Behavioral Audit
Choose one consequential marketing decision and produce an intervention portfolio that improves customer and business outcomes without hiding cost, constraining agency, or overstating evidence.
- Behavior and baselineTarget behavior, population, journey moment, baseline rate, and measurement window.
- Evidence packetBehavioral data, customer evidence, source ladder, and documented unknowns.
- DiagnosisCOM-B map, friction ledger, reference points, and at least two competing explanations.
- Intervention portfolioA structural fix, a communication treatment, and a no-change or simpler comparator.
- ExperimentUnit, allocation, primary outcome, minimum useful effect, guardrails, and analysis plan.
- Ethics and operationsTruth owner, consent and autonomy review, accessibility check, failure recovery, and rollback.
- Decision memoPrecommitted scale, revise, and stop rules plus the next decision each result enables.
Precommit the stop rules
- Stop when the primary outcome misses the predeclared practical threshold.
- Stop when trust, error, refund, complaint, or subgroup guardrails breach their limits.
- Stop when the mechanism check fails even if a vanity metric improves.
- Do not scale beyond the tested population or context without a transfer test.
Turn the audit into a live working session.
Review the training approach, or bring a real decision to a private conversation.
Continue with the behavioral economics experimentation guide or the behavioral economics in marketing article.
Social influence without social-proof theater
Design norm information around a relevant comparison and a desired direction, then watch for boomerang effects.
What you will be able to do
Evidence and boundaries
Descriptive norm feedback reduced high household energy use but increased use among some already-efficient households; an injunctive cue countered that boomerang pattern.
Evidence: context-dependent effectBoundary: This was a household energy field experiment, not a license to add approval symbols to every metric.
Beliefs about what others approve can help explain variation in norm-based energy interventions.
Evidence: context-dependent effectBoundary: The relevant group and second-order beliefs need measurement; broad popularity counts may be irrelevant.
Large-scale energy-program evidence finds average conservation effects alongside important heterogeneity.
Evidence: context-dependent effectBoundary: An average program effect does not predict the response of a marketing segment or individual.
Source-backed example: The boomerang hidden in an average
Setting:Households were shown how their energy use compared with nearby homes.
What it shows:A descriptive comparison could motivate high users while signaling permission for low users to consume more; an approval cue helped preserve the desired direction.
Testable application:Evidence: testable hypothesisFor usage, giving, or completion norms, analyze customers above and below the reference separately and predefine unwanted movement.
Failure mode and guardrails before lift
The comparison may prompt correction.
Confirm the audience sees the group as relevant and credible.
The same comparison may license regression.
Possible correction: Recognize already-desired behavior instead of licensing regression.
Move from a cue to inspectable evidence.
A descriptive norm can move different audiences in opposite directions. Test the reference group and the full message.
Decision lab
Norm-message pre-mortem
Draft one truthful norm message, then identify the reference group, data window, desired direction, and segment that could move the wrong way.
Deliverable:A message, evidence note, subgroup analysis plan, and kill criterion.
Knowledge check
Why can “the average customer uses four features” backfire?
Reveal the answer and reasoning
Answer: Customers already above the average may see it as permission to do less
Descriptive information can pull behavior toward the norm from either side of the comparison.