Event-study plots have become one of the most recognizable tools in applied causal inference.
They appear to offer several things at once: a visual assessment of pre-treatment trends, an estimate of treatment dynamics, and an intuitive picture of how outcomes evolve before and after an intervention.
But in staggered-adoption settings, the familiar event-study regression can become much harder to interpret.
A convincing graph is not the same thing as credible identification.
The Familiar Event-Study Specification
A conventional event-study regression often takes the form:
Event time is measured relative to treatment adoption. Negative values of k correspond to periods before treatment, while positive values correspond to periods after treatment.
Researchers commonly plot the estimated coefficients and look for two patterns: approximately flat coefficients before treatment and economically meaningful changes afterward.
In the simplest settings, this can be informative. Under staggered treatment timing and heterogeneous effects, however, the coefficients may not have the interpretation we expect.
Where Contamination Comes From
When units adopt treatment at different dates, units at different stages of treatment exposure coexist in the same calendar period.
Some units may be untreated. Some may have just entered treatment. Others may have been treated for several periods.
In a conventional fixed-effects event study, these observations can contribute to the same regression comparisons. If treatment effects vary across cohorts or evolve over time, coefficients for one event period can therefore reflect treatment effects from other event periods.
The coefficient labeled “two years after treatment” need not be a clean average of effects exactly two years after treatment.
Why Pre-Treatment Coefficients Require Care
Event-study leads are frequently interpreted as tests of the parallel-trends assumption.
The intuition is appealing: if treated and comparison units were already moving differently before treatment, the pre-treatment coefficients should reveal those differences.
But under staggered adoption and heterogeneous treatment effects, conventional event-study coefficients can themselves be affected by the structure of treatment timing and the comparisons embedded in the regression.
This means that apparently small pre-treatment coefficients do not, by themselves, establish the identifying assumption.
“No significant pre-trend” is not equivalent to “parallel trends has been proven.”
Parallel Trends Is an Assumption About Counterfactuals
The central identification problem in difference-in-differences concerns an outcome we never observe: what would have happened to treated units had they remained untreated?
Parallel trends provides a way to construct that counterfactual using the evolution of an appropriate comparison group.
Pre-treatment data can help us assess whether the design is plausible. But no graph can directly reveal the unobserved post-treatment counterfactual.
That is why institutional knowledge, treatment timing, comparison group construction, anticipation, and possible confounding shocks remain central to identification.
Event Time Is Economically Useful
None of this means that event studies should be abandoned.
Event time is often exactly what we care about when treatment effects are dynamic.
A policy may take time to affect behavior. A new product may require learning. An advertising intervention may generate an immediate response that later fades—or a small initial effect that grows as adoption matures.
The objective is therefore not to eliminate event-time analysis, but to estimate event-time effects using comparisons that match the intended causal estimand.
Build the Event Study From Group-Time Effects
A useful approach is to begin with group-time average treatment effects:
These effects make treatment cohort and calendar time explicit. They can then be aggregated according to event time to study treatment dynamics.
This changes the workflow. Instead of beginning with a collection of lead and lag coefficients and asking afterward what they mean, we first define the relevant causal comparisons.
Estimand → Comparison Group → Identification → Estimation → Event-Time Aggregation
Statistical Inference Still Matters
Even with a well-defined estimand, uncertainty must be handled carefully.
Event-study graphs display many coefficients simultaneously. Looking at each confidence interval independently can encourage overinterpretation of isolated estimates.
The dependence structure in panel data also matters. Standard errors should reflect the level at which treatment assignment and shocks are correlated, and inference should match the structure of the research design.
Identification and inference solve different problems. A correctly calculated standard error cannot rescue a contaminated causal comparison.
A Better Way to Read an Event-Study Graph
Before interpreting the shape of the graph, ask what comparisons generated each point.
Who is treated at that event time? Who is serving as the comparison group? Are already-treated units being used as controls? Could anticipation affect the periods immediately before treatment? Are treatment effects likely to vary across cohorts?
Only after those questions are addressed should the visual pattern be given a causal interpretation.
Takeaway
Event studies remain extremely useful because they make treatment dynamics visible.
But visualization does not replace identification.
In staggered-adoption designs, a clean-looking set of pre-trends and post-treatment coefficients may conceal complicated comparisons.
The graph should be the visualization of a credible causal design—not the evidence that the design is credible.