What Intelligence Analysts Know About Evidence That Growth Teams Don't
Intelligence tradecraft solved the problem growth teams face daily: weighing evidence when no single source is conclusive. Here is the playbook.
Know what to test, when to trust the result, and what to do next. Practical guides for analysts, growth teams, and founders.
Four short routes through the strongest work, organized around what you are trying to do.
See what changes when experimentation becomes an operating system rather than a sequence of isolated tests.
Follow this path For analysts, optimizers, and new experimentation managersStrengthen the judgment around test setup, evidence quality, and shipping decisions—not just the mechanics.
Follow this path For marketers, product teams, and behavioral-science practitionersApply behavioral principles as testable hypotheses while protecting trust and staying honest about the evidence.
Follow this path For bootstrapped founders and hands-on operatorsFollow the constraints, experiments, and distribution choices behind building products with a small team.
Follow this path38 practitioner-written articles covering everything the official Optimizely docs skip — Stats Engine mechanics, QA checklists, targeting gotchas, results interpretation, and program scaling. Built for experimentation leads who need answers, not tutorials.
Intelligence tradecraft solved the problem growth teams face daily: weighing evidence when no single source is conclusive. Here is the playbook.
Twenty years of forecasting tournaments measured what actually produces good judgment. The most valuable habit is one almost no business leader practices.
Being specific about what 'finished' looks like matters more than finding magic wording — Claude Code fills in any gap you leave, not always the way you meant.
Most of what makes someone effective with Claude Code is clear thinking, not fluency in a programming language. Here's what actually matters instead.
Describing what you want in plain English and having an AI build it produces working software, not a toy version of programming. Here's why that holds up.
Claude Code writes, edits, and runs real code on your computer. Here's the difference that makes, and what it means if you've never written a line of code.
A per-day API spend cap enforced in code still failed, because every scheduled run started from a clean checkout with no memory of prior spend.
Prompt caching is supposed to be free money. On calls spaced further apart than the cache actually lasts, it's a straight surcharge with nothing recouping it.
We named our cost-control setting something that collided with Claude Code's own environment. Every automated run quietly inherited the most expensive option.
A twice-daily job billed against the API drained an account balance for days. Moving it to a subscription runner fixed the bill and the blind spot it created.
A pricing test made the cheapest of three plans the visual anchor -- and conversion dropped. Why anchoring on price can backfire.
A rate-lock countdown timer worked at ticket checkout. It backfired at checkout for a recurring service. Why urgency is category-conditional.
A -20% topline result looked like a clear loss. It wasn't statistically significant. Why a big number and a real result aren't the same claim.
A heatmap showed most homepage visitors ignored the extra pathways offered to them. Removing those paths, not adding more, won.
A progress bar that won at checkout got re-tested earlier in the funnel, not assumed. What transferred, and why it wasn't automatic.
A well-powered test of 'choice overload' came back null. What a landmark behavioral-economics finding looks like when it doesn't transfer.
Sometimes making a price harder to notice outperforms making it easier to justify. A seasonal pricing experiment explains why.
Reordering three prices on a pricing page outperformed a full redesign -- a decoy-effect lesson in testing cheap before expensive.
Four bundled changes in one experiment came back inconclusive, and couldn't have told us anything either way. A confounded-test-design lesson.
Deleting a few sentences from a mobile modal lifted conversion by double digits -- what cognitive load teaches about 'helpful' copy.
A decade-old mobile UX principle got tested in production instead of assumed on reputation. It held up -- here's the discipline behind why.
See why a customer-selector pop-up can fail, how two first-party chooser tests compare, and how to test useful personalization without adding friction.
See why brochure previews may increase downloads, how to grade the evidence, and how to test lead magnet clarity without mistaking images for proof.
See what a product-color matching test really suggests, what its source omits, and how to test color congruence without relying on folklore.
Know what to test, when to trust the result, and what to do next. Practical decision guides for analysts, growth teams, and founders. Free. Weekly.
Opens Substack to confirm your subscription · Free · Unsubscribe anytime