For decades, two-way fixed effects regressions have been one of the standard tools for estimating difference-in-differences models.
The specification looks familiar: unit fixed effects absorb time-invariant differences across units, time fixed effects absorb common shocks, and the treatment coefficient appears to summarize the effect of the intervention.
In the canonical two-group, two-period setting, this interpretation is often straightforward. But many real policy and business settings are more complicated. Treatment begins at different times for different units, effects may evolve after adoption, and the magnitude of the effect may differ across cohorts.
In these settings, the main question is no longer simply whether a regression contains unit and time fixed effects. The question is: what causal effect does the coefficient actually identify?
The Identification Problem
Suppose different states, cities, firms, or users adopt a policy at different dates. Some units are never treated, some are not-yet-treated, and others have already been treated.
A conventional TWFE regression can implicitly combine comparisons across all of these groups. The difficulty is that an already-treated unit may become part of the comparison group for a unit treated later.
If treatment effects are constant across cohorts and over time, this may not create a serious interpretation problem. But when treatment effects are heterogeneous or dynamic, already-treated observations are no longer clean untreated counterfactuals.
The resulting coefficient can therefore combine economically different comparisons into a single number.
From a Single β to Cohort-Time Treatment Effects
A more transparent way to formulate the causal question is to define treatment effects by treatment cohort and calendar time.
Here, g denotes the period in which a group first receives treatment, while t denotes the period in which the outcome is measured.
Instead of immediately asking for one treatment coefficient, we first ask a more precise question:
What is the treatment effect for units first treated in period g, measured at time t?
This distinction matters because a policy may have a small effect immediately after adoption and a larger effect several periods later. Early adopters may also respond differently from later adopters.
Why Aggregation Matters
Once cohort-time effects are estimated, we may still want a single summary measure. But aggregation should be an explicit economic choice rather than an accidental consequence of a regression specification.
We might average effects across treated units, across cohorts, or by event time. Each answers a different question.
For example, averaging by event time can tell us how treatment effects evolve one, two, or three periods after adoption. Averaging across cohorts may instead summarize the experience of groups that entered treatment at different times.
The weights are part of the estimand. They should reflect the question we want to answer.
A Marketplace Example
Consider a platform that introduces a new advertising product across markets at different times.
San Francisco receives the product first, Chicago several months later, and additional markets adopt afterward. The platform wants to know whether the advertising product increases seller revenue.
A conventional TWFE regression may appear natural because markets adopt at different dates. But suppose the advertising effect grows with time as sellers learn how to use the product.
When Chicago becomes treated, San Francisco is already treated and may have experienced several months of treatment effects. Using San Francisco as part of Chicago's comparison group can therefore contaminate the counterfactual.
The economic question is not simply: What is β?
It is closer to: What would seller revenue have been in each treated market at each point in time had the advertising product not yet been introduced?
Modern Difference-in-Differences
Modern staggered-adoption methods make these comparisons more explicit. Rather than automatically using all available observations as controls, they construct treatment effects using appropriate comparison groups such as never-treated or not-yet-treated units under the relevant identifying assumptions.
The workflow becomes:
Causal Question → Estimand → Identification → Comparison Group → Estimator → Aggregation
This ordering is important. An estimator cannot tell us what causal question we intended to answer.
The Broader Econometric Lesson
The staggered-adoption problem illustrates a broader principle in applied econometrics: familiar regression specifications do not automatically produce familiar causal estimands.
Fixed effects can control for important sources of variation, but they do not by themselves determine whether the underlying comparisons form credible counterfactuals.
Before estimating a model, we should be able to state clearly: Who is treated? Who provides the counterfactual? When is the effect measured? How are heterogeneous effects aggregated? And what assumptions allow the comparison to be interpreted causally?
The goal is not to abandon TWFE. It is to understand when the coefficient corresponds to the causal object we care about—and when it does not.
Takeaway
In staggered-adoption settings with heterogeneous treatment effects, the central problem is not simply choosing a more sophisticated estimator.
It is defining the causal object first.
Once the estimand is clear, the identification strategy, comparison groups, estimator, and aggregation rule can be chosen to match it.