How to Tell If a Growth Hire Understands Risk
Meta description: A great win story tells you almost nothing about judgment. Two borrowed interview probes — from forecasting research and intelligence tradecraft — do.
TL;DR
- A confident story about a big win is one of the weakest signals available in a growth-leadership interview — it's vivid, it's memorable, and it tells you almost nothing about the judgment that produced it versus the luck that carried it.
- Most interviews test for articulateness under a friendly, low-stakes question. Almost none test for the specific skill that actually separates a good experimentation leader from a confident one: how someone reasons under uncertainty, in real time.
- Two techniques, borrowed from forecasting research and intelligence tradecraft, convert that skill from a vibe into something you can actually probe for in forty-five minutes.
- Neither probe can be prepared for with a rehearsed answer, because neither asks the candidate to tell a story — both ask them to do something, live, that a good story can't substitute for.
- The candidate with the thinner resume who passes both probes is very often the better long-term hire. The skill you're actually buying is how their next ten ambiguous calls go, not how their best past one turned out.
A hiring manager sits across from a growth-lead candidate who tells a great story: a redesign that lifted conversion, a pricing change that held up, a launch that hit its number. It's specific, it's confident, it's told with the easy fluency of someone who's told it many times before. Everyone in the room nods. It's also almost worthless as evidence, and not because the candidate is lying — because a vivid, well-rehearsed story about a past win is exactly the kind of input that's easiest to be fooled by and hardest to verify in the room. It could be the record of genuinely excellent judgment. It could be one lucky call, retold with the confidence of a pattern. From the interview seat, the two are indistinguishable, because the story format itself doesn't carry the information that would let you tell them apart.
That's a real gap, not a minor one. Growth and product leadership is a job that consists almost entirely of making calls under incomplete evidence — and most interview processes have no mechanism for testing that skill directly. They test for confidence, for fluency, for a track record that may or may not have been earned by the judgment being evaluated. What follows are two techniques that test the actual thing.
Why the resume story is the wrong evidence to weigh heavily
The reason a compelling war story is a weak signal isn't specific to hiring — it's a well-documented feature of how people process vivid information generally. A single, detailed, emotionally satisfying account is disproportionately persuasive regardless of whether it's been corroborated, because coherence feels like truth even when it isn't evidence of it. Intelligence analysts are trained specifically to resist this, because the same trap — a single vivid, uncorroborated source outweighing several duller, more reliable ones — is one of the best-documented causes of analytical failure. A hiring interview has the identical structure: one polished anecdote, delivered with total conviction, against no corroborating information at all.
The practical implication isn't to distrust every story a candidate tells. It's to stop treating story quality as a proxy for judgment quality, because the two are only loosely correlated — a genuinely lucky call and a genuinely well-reasoned one get told with the same confident cadence after the fact, especially by someone who's had years to polish the retelling. If the story is the only evidence on the table, you're not evaluating judgment. You're evaluating narrative skill, which is a real skill, just not the one the role actually requires.
Probe one: ask for a number, not a story
Forecasting research — most famously the Good Judgment Project's multi-year tournaments — has spent years measuring what actually separates good judgment from confident judgment, and the finding that matters most here is almost embarrassingly simple: people with genuinely good judgment under uncertainty routinely assign real probabilities to their beliefs and check those probabilities against what actually happened. Almost everyone else just remembers whether they were right, in the vague, self-flattering way memory tends to work.
The interview version: when a candidate describes a call they made with real conviction, stop them and ask two follow-up questions. First — "What percentage would you have put on that, at the time, before you knew the answer?" Not "were you confident," a number. Second — "How did you find out whether you were right, and when?"
Watch for three things. Whether they can produce an actual number instead of reaching for "pretty confident" — most candidates, asked for the first time, visibly struggle with this, and that struggle is itself informative. Whether the number sounds calibrated rather than reflexively extreme — someone who says "95%" about a genuinely uncertain call before they had evidence is telling you something different than someone who says "65%." And whether they have an actual answer to the second question, or whether the honest answer is that they moved on to the next project and never checked. That last gap is the most common finding, and it's the more important one — a candidate who has never once gone back and scored their own past confidence against reality has never built the one habit most reliably associated with good judgment in this specific line of work.
Probe two: give them an ambiguous drop and listen for the alternatives
The second technique comes from intelligence tradecraft rather than forecasting research, and it targets a different failure mode: not overconfidence, but premature certainty about _why_ something happened.
Give the candidate a short, realistic scenario: "A core conversion metric dropped by a meaningful margin the week after a launch. Walk me through how you'd figure out what happened." Then just listen, and count how many genuinely distinct explanations they generate before they commit to one. The specific, structured version of this — Analysis of Competing Hypotheses: listing every plausible explanation before looking closely at any of them, then actively hunting for evidence that would rule each one out rather than evidence that confirms the first one — is a real technique taught throughout the intelligence community precisely because unaided judgment reliably jumps to the first plausible story and then spends its remaining effort defending it rather than testing it.
A weak answer commits early: "sounds like the new checkout flow broke something," followed by a description of how they'd confirm that specific theory. A strong answer holds off: the launch, yes, but also a tracking regression, a traffic-mix shift from a paid campaign change that landed the same week, a seasonal pattern, a competitor promotion — several candidate explanations, named before any of them gets investigated, with a specific idea of what evidence would rule each one out rather than just support the favorite. The gap between those two answers is not intelligence or experience. It's a discipline, and it's one of the more reliable predictors of whether someone's future root-cause calls will be trustworthy or just confidently plausible.
What each source of evidence actually tells you
| Evidence source | What it reveals | What it can't tell you |
|---|---|---|
| A resume win story | That something good happened once, and that the candidate can narrate it persuasively | Whether the outcome reflected their judgment or favorable noise they got credit for |
| The calibration probe | Whether they've ever quantified their own confidence and checked it against reality | Domain expertise, or how they'll perform on a problem type they haven't faced before |
| The competing-hypotheses probe | Whether their default mode under ambiguity is to search for the truth or to defend the first plausible story | Whether they'll actually apply that discipline under real time pressure, not just in an interview |
None of the three is sufficient alone. But the first one is the only one most interview processes actually collect, which is precisely backwards — it's the one that's easiest for a good storyteller to pass regardless of the judgment underneath it.
The diagnostic catch this surfaces
Here's the finding that makes this worth doing, stated plainly: a candidate with an excellent-sounding track record can fail both probes, and a candidate with a noticeably thinner one can pass both. That's not a paradox — it's exactly what you'd expect once you separate outcome from process. A strong track record can be a small number of calls that happened to land well, retold with the confidence of a pattern, by someone who's never once gone back to check their own calibration or practiced distinguishing a favorite explanation from a verified one. A thinner track record can belong to someone earlier in their career who nonetheless already has both habits — and habits, unlike a specific past win, are the thing that transfers to your business, under your constraints, on problems that look nothing like the ones on their resume.
That's the actual hire you're making. Not their best past call. How their next ten ambiguous ones are likely to go, somewhere you can't yet see, on evidence you haven't yet collected. A win story can't tell you that. A number and a list of ruled-out alternatives, produced live under a question they couldn't have rehearsed, comes a lot closer.
FAQ
Isn't this too abstract to actually use in a normal interview?
No — both probes are two follow-up questions, not a separate exercise. Ask a candidate for a story you'd ask for anyway, then push with "what percentage, at the time" and "how did you find out." Ask a candidate a scenario question you'd ask for anyway, then just count the alternatives before they commit. The discipline is in what you listen for, not in redesigning the interview.
What if the candidate has never explicitly tracked their calibration? Is that disqualifying?
Not by itself — almost nobody has, which is exactly the point of asking. What matters is what happens next: do they recognize the gap immediately and reason about why it matters, or do they deflect with a version of "I just have good instincts." The first response is a promising sign in someone who hasn't built the habit yet. The second is the more useful warning, regardless of how strong their resume looks.
Does this work for evaluating an existing team, not just a candidate?
Yes, and arguably it's more useful there, because you have real, current decisions to probe instead of a rehearsed history. Pick a live ambiguous result from the last quarter and ask the same two questions of whoever owned the call. The same gaps show up, and unlike an interview, you can act on what you find.
How is this different from a standard behavioral interview question?
A standard behavioral question ("tell me about a time you...") asks for a story, which is exactly the format that's easiest to prepare and hardest to verify. Both probes here ask for something a rehearsed story doesn't contain — a specific number produced in the moment, or a live count of alternative explanations — which is much harder to fake convincingly on the spot.
What's the actual pass/fail signal here?
There isn't a hard pass/fail — the value is comparative. A candidate who produces a real number and checks it, and who generates several genuine alternatives before committing to one, is demonstrating a specific, transferable discipline. A candidate who can only offer conviction and a single confident theory is demonstrating storytelling. Weigh the difference the same way you'd weigh any other job-relevant skill you tested for directly.
Related reading: Calibration Training, What Intelligence Analysts Know About Evidence, The Confidence Tier Model.
Bottom line
A great story is the easiest thing for a strong candidate to prepare and the hardest thing for an interviewer to verify — which makes it close to the worst evidence available for the judgment a growth-leadership role actually requires. Asking for a number instead of a feeling, and counting alternatives instead of accepting the first plausible one, tests something a rehearsed answer can't fake. What you're really buying isn't the candidate's best past call. It's how the next ambiguous one, on your business, is likely to go — and that's a specific, probeable habit, not a vibe you pick up from a good story.
If you're trying to tell whether a growth hire — or a fractional leader you're evaluating — actually has this judgment, not just a good story about it, get in touch.