Evaluate an A/B testing case study by reconstructing the experiment before judging the headline. Find the primary source, identify the sample and allocation, verify the primary metric and duration, locate the stopping rule and SRM check, then separate what worked once from what has actually been replicated.

An A/B testing case study is a narrative report of an experiment; it is not automatically a complete experiment record.

DataForSEO estimates only about 20 US searches per month for “A/B testing case study,” but the query is commercially valuable. People searching it are choosing ideas, vendors, or methods. A trustworthy evaluation framework can prevent them from spending traffic and engineering time on a story that cannot support its headline.

Key takeaways

  • Use the primary source, not a chain of roundups.
  • A winning outcome does not determine evidence quality.
  • Missing sample, duration, stopping rule, or SRM should remain visibly missing.
  • Range-bucketed private data supports directional synthesis, not exact meta-analysis.
  • Every case should end in an original test plan, not a cloned design.

Why most case studies are incomplete

Case studies are written to communicate a result. Experiment records are written to preserve a decision. Those incentives produce different documents.

A marketing case often emphasizes:

  • The dramatic visual change
  • A relative lift
  • Confidence or significance
  • A short explanation of why the winner won

A reusable experiment record also needs:

  • Eligibility and assignment unit
  • Raw arm sizes and outcomes
  • Metric definitions
  • Planned allocation
  • Duration and business cycles
  • Stopping rule
  • SRM and instrumentation checks
  • Guardrails and downstream effects
  • Contradictory segments
  • Decision and limitations

In my portfolio audit, records that looked complete at the headline level often became narrative-only after these fields were checked. That is not a reason to discard them. It is a reason to use them for mechanisms and hypotheses instead of pooled effect estimates.

The 12-point Primary Evidence Reconstruction

Use this framework for every case.

#FieldQuestion
1Primary sourceIs this the organization’s report, platform case, paper, or a summary?
2PopulationWho was eligible, and what traffic sources were included?
3Assignment unitUser, session, account, device, region, or something else?
4SampleHow many eligible units entered each arm?
5AllocationWhat split was planned and observed?
6TreatmentWhat changed, and what else changed with it?
7Primary metricWas success defined before the result?
8DurationDid the test cover relevant cycles and novelty?
9Stopping ruleFixed sample, fixed time, or documented sequential method?
10SRM and qualityDid assignment and instrumentation behave correctly?
11GuardrailsWhat cost could the primary win hide?
12LimitationsWhat cannot be inferred or transferred?

Call this the Primary Evidence Reconstruction. A blank cell is a finding, not an invitation to guess.

Worked example: navigation removal

The VWO case report says removing navigation from a registry landing page moved registrations from 3% to 6%. The source report describes the page, traffic sources, treatment, goal, and rates.

It does not report sample size, allocation ratio, duration, stopping rule, SRM, raw counts, or a confidence interval. That earns a C — partially reported.

Calibrated conclusion: the treatment worked once in the reported setting. It suggests that global navigation can compete with a narrow registration goal. It does not prove a 100% expected lift for another landing page.

This sentence is less exciting than the headline and more useful to someone allocating traffic.

Worked example: pricing-page registration

Buttondown’s first-party pricing experiment reports registrations changing from 6.7% to 9.5%. It identifies the downstream registration metric and explains the implementation. It does not publish the sample, duration, allocation, stopping rule, or SRM.

The treatment also changed two elements: it removed a callout and repositioned the signup action. The report supports the business conclusion that the combined page performed better. It cannot isolate which element caused the change.

This is a classic separation:

  • Shipping evidence: useful
  • Mechanism evidence: limited
  • Replication claim: unsupported

The original thesis most summaries miss is that experiment usefulness has two axes: decision value and learning resolution. A case can be high on the first and low on the second.

How to grade evidence

Use five grades:

GradeDefinitionSafe use
A — ReproducibleRaw arms, metric, duration, uncertainty, stopping, and quality checks availableQuantitative synthesis when outcomes are comparable
B — VerifiedFirst-party or independently reviewed evidence, publicly sanitizedDirectional and mechanism synthesis
C — PartialCredible report with important fields missingSupporting case and hypothesis
D — AnecdotalClaim with little methodological detailInspiration only
X — ExcludedDuplicate, contradictory, unverifiable, or improperly sourcedDo not count as evidence

Do not award a better grade because the lift is large or significant. An inconclusive A-grade test is more reusable than a dramatic D-grade winner.

My own public-safe navigation example receives a B. The repository preserves more than 100,000 observations, three approximately equal arms, completed orders, a four-to-eight-week duration, no detected SRM, and a 10%–20% winning range. The stopping rule is missing, and exact data are intentionally not public. That supports a directional first-party contribution but not independent reanalysis.

When can you call the article a meta-analysis?

A formal meta-analysis requires comparable effect estimates and uncertainty. At minimum, you generally need arm-level sample and outcomes or an effect size with a standard error or confidence interval. You also need a defensible reason the populations, treatments, and metrics belong in one model.

Do not pool:

  • A checkout order with a landing-page click
  • Revenue per session with registration rate without a coherent transformation
  • Multiple records from the same underlying experiment as independent tests
  • Range-bucketed confidential effects as exact values
  • Vendor winners while silently excluding nulls and losses

If those conditions are not met, use systematic evidence review, research synthesis, or case-study comparison. The portfolio evidence audit uses that language because heterogeneous metrics and dependent records make a pooled lift misleading.

How should public and private evidence coexist?

Use two transparent policies.

Public cases

Describe the subject generically when prominence is unnecessary, but keep the publisher citation visible. Quote only what is needed, and evaluate omissions directly.

Private portfolio cases

Anonymize the organization, remove proprietary screenshots, bucket samples and effects, and disclose that public readers cannot reproduce the estimate. Never use range-bucketing to upgrade evidence quality.

This lets the database protect commercial confidentiality without pretending the source is public. It also follows the same separation used in decisions with incomplete data.

Turn every case into an original test plan

An article adds value only if its thesis goes beyond the source. Use this sequence:

  1. State exactly what the source supports.
  2. Identify the missing evidence fields.
  3. Name at least one competing mechanism.
  4. Add a portfolio pattern, counterexample, calculation, or enterprise constraint.
  5. Specify where transfer is plausible and where it may fail.
  6. Define a new control, treatment, primary metric, guardrails, sample plan, and stopping rule.

For navigation, a stronger follow-up compares visible, collapsed, and enclosed states rather than cloning the winning screenshot. See the three-arm navigation test plan and calculate feasibility with the sample-size guide.

Use calibrated claim language

Use these terms consistently:

  • Worked once: one credible experiment favored the treatment.
  • Suggests: incomplete or heterogeneous evidence points in a direction.
  • Supports: multiple credible sources or methods align with the mechanism.
  • Replicated: comparable independent experiments reproduce the claim under specified conditions.

“Has been replicated” should be the rarest label in a public case library. Two vendor blog posts that both declare winners are not automatically replications.

Start your evidence review in GrowthLayer

Start your free GrowthLayer workspace to capture the 12 reconstruction fields, assign an evidence grade, link the primary source, and turn the case into an actionable experiment brief.

FAQ

What should an A/B testing case study include?

It should include the population, assignment, sample, allocation, treatment, primary metric, duration, stopping rule, SRM status, guardrails, result, decision, and limitations.

Can I trust a case study without sample size?

Use it as a hypothesis source, not a precise effect estimate. Without sample and uncertainty, you cannot judge how stable the reported rates are.

Is statistical confidence enough to validate a case study?

No. Confidence does not fix poor assignment, an undocumented stopping rule, multiple comparisons, instrumentation errors, weak metrics, or missing guardrails.

Can anonymized internal experiments be included?

Yes, when the author has permission and discloses the sanitization. Range-bucketed private results should support directional synthesis rather than exact public meta-analysis.

When has an A/B test result been replicated?

Use that label when comparable independent experiments reproduce the effect under defined conditions. Similar-looking winners from incomplete reports are only supporting cases.

Share this article
LinkedIn (opens in new tab) X / Twitter (opens in new tab)
Atticus Li

Experimentation and growth leader. CXL-certified CRO practitioner, Mindworx-certified behavioral economist (1 of ~1,000 worldwide). 200+ A/B tests across energy, SaaS, fintech, e-commerce, and marketplace verticals.