Atticus Li designed a three-market evaluation at Silicon Valley Bank to test whether OOH advertising increased digital demand. The result is useful precisely because it did not support the easy success story: Austin and Miami grew during the flight, but the Seattle control grew faster. The aggregate comparison therefore did not establish causal web-traffic lift from OOH.

The Problem Nobody Wanted to Solve

SVB was spending serious money on out-of-home advertising. Billboards in Austin. Taxi cab wraps in Miami. Bus stop displays in Boston. Airport signage in other campaign periods. The creative looked great. The placements were premium. And nobody could tell you whether any of it worked. During the three-market test described below, Seattle was the digital-only control with no planned OOH treatment.

OOH is often evaluated with reach, frequency, and brand measures because individual exposure is difficult to connect to a later digital action. That makes a credible incrementality design especially valuable.

At SVB, I wasn't willing to accept that. Not because I thought OOH was a waste — I actually suspected it was working. But "I think it's working" is not a sentence that survives a budget review when finance is looking for places to cut.

The challenge was real. OOH doesn't have click-through rates. There's no cookie. No UTM parameter on a highway billboard. The standard digital attribution toolchain is completely useless here.

So I designed something different.

Why Geo-Incrementality Is the Right Framework for OOH

When you can't track individual user journeys from a billboard to your website, you have to change the unit of analysis. Instead of tracking people, you track markets.

Geo-incrementality testing compares treated markets with untreated markets. When the assignment, pre-period fit, exposure, and analysis are credible, the difference in outcomes can estimate incremental impact.

The logic is simple: if a treated market improves more than a credible counterfactual during the same period, that difference may support an incremental effect. If the control improves more, the aggregate result does not support a positive lift claim.

You need markets that behave similarly before the test starts. You need to account for seasonality. You need to normalize for baseline differences. And you need enough statistical power to detect a real signal through the noise of normal business variation.

This is the kind of measurement problem I find genuinely interesting — the kind where getting the design right matters more than having sophisticated tools. I've written about the difference between tool proficiency and analytical capability in my piece on enterprise analytics at SVB and NRG, and this project was a perfect example.

The Experiment Design

I structured the test around three markets:

  • Austin — test market (billboards active)
  • Miami — test market (billboards active)
  • Seattle — control market (no OOH spend)

Why These Markets

Market selection is the most important decision in a geo-test and the one that gets the least attention. You can't just pick cities at random. They need to be structurally similar enough that differences in outcomes are attributable to the treatment, not to underlying market dynamics.

I selected Austin and Miami as test markets because they had comparable baseline web traffic patterns, similar seasonal trends in SVB's core banking products, and both were active OOH markets where we could place media efficiently. Seattle served as the control because it shared traffic profile characteristics with the test markets but had no planned OOH activity during the test window.

Before running anything, I used pre-test data to assess whether the markets could support a comparison. Pre-period similarity is necessary, but it does not override what happens during the test.

Handling Seasonality and Noise

Financial services traffic isn't steady. It fluctuates with earnings seasons, market events, Fed announcements, and startup funding cycles. A naive pre/post comparison would be contaminated by these factors.

I built seasonality adjustments into the baseline using the prior year's traffic patterns to create expected values for each market during the test window. The lift calculation compared actual performance against the seasonally adjusted baseline, not against raw pre-period numbers.

I also established matched market baselines—normalizing each test market's performance against the control's performance over the same period. That comparison was intended to adjust for shocks shared across markets, but one control cannot absorb every difference. Credibility still depends on pre-period fit, no major concurrent interventions, limited spillover, and sensitivity to model and control choices.

The Results

The recorded market-level changes were:

MarketOOH treatmentObserved traffic change
AustinYes+94.4%
MiamiYes+97.8%
SeattleNo+114.3%

Those are observed changes relative to each market's baseline—not treatment lift versus Seattle. The control outgrew Austin by 19.9 percentage points and Miami by 16.5 percentage points. On this aggregate outcome, the test did not demonstrate positive OOH incrementality.

That does not prove OOH had no effect on every outcome. It means this comparison cannot honestly be presented as evidence that OOH caused the reported web-traffic growth.

Building the Offline-to-Online Attribution Pipeline

The aggregate geo result did not answer “does OOH work?” with a positive causal estimate. A separate question was whether direct-response traces could connect a specific placement to a lead.

This is a harder problem, but not an impossible one.

I built an offline-to-online attribution pipeline using QR-enabled billboard tracking. Each tracked placement routed through a dedicated landing page instrumented to capture the source placement and connect subsequent form activity to the CRM.

This produced direct-response records for scans and subsequent tracked actions. Those records are evidence about the people who used the QR path; they are not evidence about everyone exposed to OOH.

QR tracking cannot recover unobserved paths such as later search or direct visits. The geo comparison was intended to measure the broader market effect, but this run did not produce a positive aggregate lift estimate. The two methods therefore answered different, limited questions rather than forming a complete causal picture.

The Paid-Digital Halo Was a Hypothesis

One of the most interesting findings wasn't about OOH in isolation — it was about what OOH did to our other channels.

Paid-digital metrics also moved in the treatment markets, which made an OOH halo worth investigating. But a concurrent improvement is not enough to claim that billboards caused the change. The comparison needs the same counterfactual discipline, uncertainty analysis, and protection against channel-mix differences as the primary outcome.

The defensible conclusion was a follow-up hypothesis, not a proven halo: OOH exposure may change response to later digital ads, and a future design should isolate that interaction directly.

From Measurement to Budget Reallocation

Data without decisions is just decoration. The real value of this work wasn't the charts — it was the budget conversation it enabled.

Before the geo-test, OOH budget discussions were based on brand proxy metrics. Estimated impressions. Reach and frequency models from the media agency. CPM comparisons against other "awareness" channels. These metrics aren't useless, but they don't answer the question finance actually cares about: is this spend generating demand?

After the geo-test, I built executive-ready scenarios that combined media cost, observed market movement, conversion assumptions, and estimated pipeline value. Because the control outgrew both treatment markets, those scenarios could not honestly book the treatment-market traffic growth as incremental ROI.

The decision value was the evidence boundary: leadership could distinguish direct-response traces, observed market movement, and modeled scenarios from a demonstrated causal return.

What Most Companies Get Wrong About OOH Measurement

Having gone through this process, I see three mistakes companies make consistently:

Mistake 1: They don't try to measure it at all. OOH gets treated as unmeasurable brand spend, and the budget survives or dies based on executive opinion rather than evidence. Geo testing is one candidate when a company has enough comparable units, stable outcome measurement, a credible no-treatment comparison, and enough signal to distinguish an effect from normal variation. Without those conditions, another design may be more honest.

Mistake 2: They only measure direct response. QR codes and vanity URLs observe only trackable response paths. They do not establish the size or direction of unobserved effects. Use direct-response tracking alongside a credible aggregate design, and keep the conclusions separate.

Mistake 3: They run the test badly. Poor market selection. No pre-test normalization. No seasonality adjustment. No control market at all. I've seen "geo-tests" where the test and control markets were so different that the results were meaningless regardless of what happened. The design rigor matters as much as the execution.

Why This Matters Beyond OOH

The geo-incrementality framework I used at SVB is not limited to billboards. It can be a useful candidate for channels where individual-level tracking is impossible or unreliable—when the comparison and power requirements are credible.

TV advertising. Podcast sponsorships. Conference sponsorships. PR campaigns. Brand partnerships. Diffuse channels may benefit from a geo design, but only where treatment assignment, comparable controls, spillover, timing, and measurement make the counterfactual defensible.

The core questions transfer: define the causal question, assess pre-period fit, specify treatment and outcomes, measure the differential result, and test sensitivity to competing explanations. The actual design must change with the channel, market count, spillover risk, and available power.

This is also why I built the measurement capabilities at SVB the way I did — not as one-off analyses, but as repeatable frameworks. When I moved to NRG and started building the experimentation program there, the same measurement principles applied even though the industry, the channels, and the tools were completely different. Good analytical frameworks transfer across contexts.

The Bigger Lesson

Marketing measurement isn't about having the perfect tool or the perfect data. It's about asking clear causal questions and designing studies that can answer them.

The SVB geo experiment was a straightforward test/control design. Its most important lesson was that large pre/post gains in treatment markets can still fail an incrementality test when the control grows faster.

That's the standard I hold every measurement project to: did it change a decision? If the analysis is methodologically beautiful but doesn't influence how the company allocates resources, it's academic exercise, not applied analytics.

The project changed the quality of the decision: observed growth could no longer be presented as causal lift without the control comparison. That is what a useful measurement framework should do, even when the answer is not the one the campaign team hoped for.


_Have a question about measuring offline advertising impact or designing geo-incrementality tests? Reach me at atticus@atticusli.com._

FAQ

Did the SVB geo test prove that OOH reduced growth?

No. The Seattle control outgrew the treatment markets in the observed period, so the comparison did not establish positive causal lift from the campaign. It does not prove a universal negative OOH effect.

What makes a geo control credible?

Document the selection period, pre-treatment fit, concurrent interventions, market spillover, outcome definition, and sensitivity to alternative controls.

Can one geo test produce a permanent ROI estimate?

No. The estimate is specific to its markets, timing, spend, implementation, and measurement window. Microsoft Research's experiment-pitfall catalog and Google's work on long-term effects explain two important classes of limitation.

Share this article
LinkedIn (opens in new tab)X / Twitter (opens in new tab)
Atticus Li

Experimentation and growth leader. CXL-certified CRO practitioner, Mindworx-certified in behavioral economics. Led 100+ in-house experiments at NRG in 2025, with project evidence and limits documented in the case studies.