The day eBay turned
its ads off
The 2012 eBay experiment is widely regarded as one of the most significant natural experiments in the history of digital advertising. By deliberately switching off its own paid search ads across a large portion of the United States, eBay was able to measure the true incremental value of that spending, and the results fundamentally challenged the industry’s assumptions about the effectiveness of paid search.
Background and Methodology
1. The Experiment
In 2012, a team of economists working with eBay designed a large-scale field experiment to determine whether the company’s brand keyword advertising was actually generating additional sales. Rather than relying on the correlational data provided by standard attribution tools, they used a controlled, randomized approach.
- Randomized Design: eBay paused its brand-keyword search advertising in 68 of the 210 designated market areas in the United States, selected at random, while continuing to advertise as normal in the remaining 142.
- Duration: The test ran for 60 days, allowing researchers to compare sales in the “dark” markets against the control markets over a meaningful period.
- Scale: The experiment covered roughly a third of eBay’s US traffic, making it one of the largest advertising experiments ever conducted.
2. The Key Findings
The outcome was unambiguous, and it contradicted what eBay’s internal reporting had indicated for years.
- 99.5% Substitution: Approximately 99.5% of the clicks eBay had been paying for arrived anyway through organic search results once the paid ads were removed. Users who typed “eBay” into a search engine simply clicked the free link instead of the sponsored one.
- Negative Return: The measured return on investment for brand search advertising was -63%, meaning that for every dollar spent, eBay was losing money rather than generating incremental revenue.
- Misleading Attribution: When the researchers applied the industry-standard observational method to the very same data, it reported a return of over 1,400%. The method was not simply imprecise; it was directionally wrong.
Why Standard Measurement Failed
- Correlation Versus Causation: Standard attribution models credit an ad with every purchase that follows a click. They have no mechanism for identifying which of those purchases would have occurred without the ad.
- Selection Bias: The users most likely to click a brand ad are, by definition, those already searching for the brand. These are the customers least in need of persuasion, which inflates the apparent effectiveness of the advertising.
Broader Validation
The eBay result was not an isolated finding. In 2019, researchers at Facebook conducted a comparative analysis of fifteen large-scale advertising experiments, encompassing approximately 500 million user-experiment observations.
- Consistent Failure: Observational measurement methods failed to recover the true experimental effect in every single case, even after controlling for extensive demographic and behavioral variables.
- Bidirectional Error: The errors did not follow a predictable pattern; the standard tools sometimes overstated and sometimes understated the true effect of the advertising.
Implications for Your Business
- Advertising Still Works, Selectively: The eBay data also demonstrated that ads did generate incremental sales among new and infrequent users. The spending was wasted only on customers who were already coming.
- Audit Brand-Term Spending: If your account is bidding on your own name, the eBay findings suggest that a significant portion of that spending may be purchasing traffic you would have received for free.
- Establish Measurement Rules in Advance: A written measurement standard, defined and dated before a campaign begins, provides a far more reliable basis for evaluation than a retrospective attribution report.
In Conclusion
The eBay experiment remains a landmark study because it exposed a gap between what advertising reports show and what advertising actually accomplishes. For any business evaluating its own paid acquisition, the lesson is clear: the number on the dashboard is not the same as the number that matters, and only a properly designed test can tell the difference.
1 · Blake, Nosko and Tadelis, “Consumer Heterogeneity and Paid Search Effectiveness,” Econometrica, 2015. Field experiment run at eBay in 2012; brand search paused across roughly a third of US traffic; about 99.5 percent of paid clicks substituted to free clicks; positive effects concentrated in new and infrequent users. 2 · Gordon, Zettelmeyer, Bhargava and Chapsky, “A Comparison of Approaches to Advertising Measurement,” Marketing Science, 2019. Fifteen US advertising experiments at Facebook; roughly 500 million user-experiment observations; observational methods failed to recover experimental lift.