Generalized Synthetic Control (GSYNTH)#
When to Use This Method#
You have a panel, several units adopt a policy at dates that need not coincide, and a good number of units never adopt at all. The units move together through shared unobserved forces — a business cycle, a national mood, a regional shock — but not in parallel, because each one responds to those forces with its own sensitivity. Difference-in-differences assumes the treated and untreated units share a common time effect, which is exactly what heterogeneous responses break. Plain synthetic control builds a convex combination of donors for one treated unit, and with nine treated units adopting on four different dates it is not obvious what to do.
The generalized synthetic control method of Xu ([Xu2017]) covers that regime. It writes the untreated outcome as a small number of latent common factors \(\mathbf{f}_t\) weighted by unit-specific loadings \(\boldsymbol{\lambda}_i\), plus additive unit and period effects and any covariates. The units that are never treated identify the factors and the covariate coefficients. Each treated unit is then placed in that estimated factor space using its own pre-treatment history, and its untreated potential outcome follows.
The result is a synthetic control that is fit by projection instead of by a weighted average of donors: the treated unit is matched to the estimated factors, not to particular control units. Nothing constrains the treated unit to lie inside the convex hull of the donor pool, and units adopting at different dates need no special handling, because each one’s loadings come from the pre-period block it happens to have.
Reach for GSYNTH when#
several units are treated, possibly at different dates, and a substantial group is never treated;
the units co-move through latent common factors but not in parallel, so parallel trends is hard to defend;
the pre-treatment histories are long enough to estimate a loading vector for each treated unit;
you want the number of factors chosen by the data instead of asserted;
you want the treated units aggregated into one ATT with a matching confidence interval and an effect path by time since adoption.
Do not use GSYNTH when#
every unit is eventually treated. The factor space here comes from the never-treated units alone, and the estimator raises when there are none. Matrix Completion with Nuclear Norm Minimization (MCNNM) handles that design;
treatment turns on and off. Step 2 projects onto one contiguous pre-treatment block per unit, which a reversing treatment does not define. Matrix Completion with Nuclear Norm Minimization (MCNNM) again, or Rolling-Transformation DiD (ROLLDID);
a treated unit’s pre-period is shorter than the number of factors you want. Its loading vector is then unidentified and the fit raises instead of returning a number;
the intervention plausibly moved the common factors themselves. The factors are estimated off units assumed to be unaffected, so a general equilibrium response contaminates them.
Notation#
Let \(\mathcal{T}\) index the \(N_{tr}\) treated units and \(\mathcal{C}\) the \(N_{co}\) never-treated units, observed for \(t = 1, \dots, T\). Unit \(i \in \mathcal{T}\) adopts at \(T_{0i} + 1\), so its pre-treatment block is \(t \le T_{0i}\). The untreated potential outcome follows
with \(\mathbf{f}_t\) an \(r \times 1\) vector of latent common factors, \(\boldsymbol{\lambda}_i\) unit \(i\)’s loadings, \(\mathbf{x}_{it}\) observed time-varying covariates entering with a common coefficient vector \(\boldsymbol{\beta}\), \(\alpha_i\) and \(\xi_t\) additive unit and period effects, and \(\varepsilon_{it}\) idiosyncratic noise. Write \(\mathbf{F} = (\mathbf{f}_1, \dots, \mathbf{f}_T)^\top\) and let a superscript \(0\) restrict a quantity to a unit’s pre-treatment periods.
The estimand is the average effect on the treated,
and its path \(\widehat{ATT}_h\) in time since adoption, with \(h = 0\) the first treated period.
The estimator in three steps#
Step 1 fits the interactive fixed effects model to the never-treated units alone, minimizing
subject to Bai’s normalization \(\mathbf{F}^\top \mathbf{F} / T = \mathbf{I}_r\) with \(\boldsymbol{\Lambda}_{co}^\top \boldsymbol{\Lambda}_{co}\) diagonal. Additive effects are swept out by demeaning, and what remains alternates between a pooled regression for \(\boldsymbol{\beta}\) and principal components for \((\mathbf{F}, \boldsymbol{\Lambda})\) until the coefficient vector stops moving.
Step 2 recovers each treated unit’s loadings from its own pre-treatment periods, holding \(\widehat{\boldsymbol{\beta}}\) and \(\widehat{\mathbf{F}}\) fixed:
Step 3 imputes and differences: \(\widehat{Y}_{it}(0) = \mathbf{x}_{it}^\top \widehat{\boldsymbol{\beta}} + \widehat{\boldsymbol{\lambda}}_i^\top \widehat{\mathbf{f}}_t\).
When unit effects are in the specification a column of ones is appended to \(\widehat{\mathbf{F}}\) before Step 2, so a treated unit’s own level is estimated jointly with its loadings off its pre-periods. The control units’ unit effects say nothing about the level of a unit outside that group, so this is the only place \(\alpha_i\) for \(i \in \mathcal{T}\) can come from.
Which additive effects#
The force option chooses which of \(\alpha_i\) and \(\xi_t\)
accompany the factors, named and coded as in gsynth and fect:
|
unit effects |
period effects |
what it means |
|---|---|---|---|
|
no |
no |
the grand mean and the factors carry everything |
|
yes |
no |
levels differ across units, common shocks left to the factors |
|
no |
yes |
a common time path, with unit levels left to the factors |
|
yes |
yes |
the default, and the specification Xu (2017) Table 2 reports |
An effect that is switched off is not estimated and not removed, so it stays
in the residual and the factors absorb what they can of it. The choice
therefore moves the estimate instead of merely relabelling it, and it moves
the estimate most where the pre-treatment fit is worst. Applied work varies
it deliberately: Lang et al. (2026) run their preregistered specification at
force="none" and sweep all four in a multiverse.
Two settings differ in what they can identify, not only in what they fit. A
treated unit with \(T_0\) pre-periods supports \(T_0\) regressors in
Step 2, and under "unit" or "two-way" one of those is the intercept.
So the largest rank the cross-validation will consider is one lower for those
two than for "none" and "time".
two_way was the first release’s spelling of this option. It still
resolves — True to "two-way" and False to "time" — with a
DeprecationWarning. Passing both is an error.
Assumptions#
Functional form. The untreated potential outcome follows the interactive fixed effects model above, with the same factors and the same \(\boldsymbol{\beta}\) for treated and untreated units.
Remark. This is what replaces parallel trends. The units need not move together; they need only respond to the same latent shocks, with loadings that may differ arbitrarily across units.
Strict exogeneity. The idiosyncratic error is mean-independent of treatment assignment, the covariates, the factors and the loadings, for all units and periods.
Remark. Treatment may be assigned on the loadings — states that adopt a reform may be systematically more exposed to a national trend — which is what a plain two-way fixed effects regression cannot accommodate. What is ruled out is assignment on the idiosyncratic shock itself.
Weak serial dependence. The errors have finite variance and are weakly dependent over time, so the pre-period average of the noise vanishes as the pre-period grows.
Remark. The treated loadings are identified from pre-period time-series variation, so a treated unit with a short history yields a noisy loading and, through it, a noisy counterfactual. The per-unit
prefit_rmseon the result reports how well each unit’s pre-period is tracked.Regularity and a factor structure that is learnable. The factors are non-degenerate, the loadings have a well-behaved second moment, and the number of never-treated units and the pre-period lengths both grow.
Remark. The width of the never-treated pool determines how sharply the factors are estimated; the length of each treated unit’s pre-period determines how sharply its loadings are. A wide pool and short pre-periods gives good factors and bad loadings, and the estimator does not warn about that on its own.
No anticipation and absorbing adoption. A unit’s pre-treatment outcomes are untreated outcomes, and once treated it stays treated.
Remark. Both are enforced at ingestion. A panel where treatment reverses raises, because Step 2 needs one contiguous pre-treatment block per unit; anticipation is not detectable from the data and is the reader’s to defend.
Choosing the number of factors#
The rank is not a nuisance parameter here. The estimate is not monotone in it, and on the paper’s own application it moves the headline number by more than a quarter over the plausible range, so how it is chosen decides the answer.
GSYNTH chooses it by the paper’s Algorithm 1, a leave-one-period-out
cross-validation that scores a rank on data the treated units actually
supply. For each candidate \(r\): fit Step 1 once; then for each
pre-treatment period \(s\) of each treated unit, drop that period, refit
the loadings on the rest,
predict the held-out cell and save the error.
The scores are then walked upward from the smallest rank, and a larger rank is taken only when it beats the running minimum by more than 1 percent of it. The guard is what makes this a selection rule instead of a minimization. Prediction error is not convex in \(r\) and its minimum is often a near-tie, so without it the criterion spends a degree of freedom to buy an improvement indistinguishable from noise. Both R references guard the step the same way.
The procedure draws no random numbers, so the selected rank is a property of the panel and not of a seed. Two consequences follow, and both are checked in the test suite. The set of scored cells is the same at every rank, so the criterion is comparable across \(r\); and repeated calls return the same rank.
Which reference this is. Algorithm 1 above is what gsynth implements
through version 1.3.x. From 1.4.0 the R package is a wrapper that forwards to
fect, whose cross-validation is a different scheme: it masks a random share
of observed cells over \(k\) rounds and weights the error across relative
periods, dropping periods with too few treated observations. On a panel where
the criterion is flat the two can land on different ranks — on the Ronczewski
(2026) cannabis application they choose 4 and 1, and the estimate differs by a
factor of 2.6. Fix r when the comparison is meant to be about the
estimators, which agree to 4e-17 at a common rank.
Pass r to fix the count instead, and r_max to bound the search. The
estimator lowers r_max further when the shortest treated pre-period
history cannot identify that many loadings, since a unit with \(T_0\)
pre-periods supports at most \(T_0\) regressors and one fewer once a
period is held back.
Inference#
The uncertainty comes from the paper’s Algorithm 2, a parametric bootstrap that holds the estimated conditional mean fixed and resamples errors around it. Treated and control cells draw from different pools, and that asymmetry is the mechanism.
The factors and the coefficient vector are estimated from the control units, so the fitted surface tracks a control unit better than it tracks a treated counterfactual; the treated prediction error is the larger of the two. Loop 1 measures it. Pick a control unit, treat it as if it had adopted on the treated units’ dates, predict it from a resample of the remaining controls, and keep the difference. Repeating that builds an empirical distribution of prediction errors for a unit the model did not fit.
Loop 2 then rebuilds panels from the fitted values plus resampled errors — treated columns drawing from Loop 1’s pool, control columns from the in-sample residuals — refits, and collects the ATT. The simulated treated columns carry no treatment effect, so the spread of those refits is the sampling distribution of the estimator under the estimated data-generating process, and the point estimate is added back when the percentile intervals are formed. The same machinery produces a band by horizon.
Set inference=False to skip it; n_bootstrap controls both loops and
seed makes the result reproducible. Draws whose refit fails are discarded
and counted in inference_detail.n_failed, and the routine raises when too
few survive to estimate a variance from.
Example#
import pandas as pd
from mlsynth import GSYNTH
# Election Day Registration and voter turnout, 1920-2012.
# Nine states adopt EDR across four dates; thirty-eight never do.
df = pd.read_parquet("basedata/xu_edr_turnout.parquet")
res = GSYNTH({
"df": df, "outcome": "turnout", "treat": "policy_edr",
"unitid": "abb", "time": "year",
"covariates": ["policy_mail_in", "policy_motor"],
"inference": True, "n_bootstrap": 2000, "seed": 2139,
"display_graphs": False,
}).fit()
res.att # 4.90 percentage points
res.att_ci # bootstrap percentile interval
res.design.r_selected # 2, chosen by Algorithm 1
res.design.cv.mspe # its criterion by rank
res.design.beta # covariate coefficients: 0.15, -1.05
res.event_study.tau # effect by period since adoption
res.per_unit["MN"].prefit_rmse # how well one state's pre-period is tracked
Passing r fixes the count instead, and design.cv is then None
because no selection was run. Adding "force": "none" drops the additive
effects and leaves the factors to carry them, which is how a good deal of
applied work runs this estimator; design.force records the setting used.
time_series carries the treated-average observed and imputed paths in
calendar time, with the gap between them; event_study carries the same
effect in time since adoption, where the staggered adopters line up. Horizon
\(0\) is the first treated period. Reference implementations in the
gsynth and fect family number it \(1\), so their paths are this one
shifted by one.
GSYNTH is not Continuous-Treatment Synthetic Control (CTSC). That is Cao, Lu and Wu’s generalized synthetic control, a differently constructed estimator that shares the name.
Verification#
GSYNTH reproduces Xu (2017) Table 2 columns (3) and (4) on the author’s own
data, and matches a live fect 2.4.5 reference across every rank from
zero to five on both specifications. See Generalized synthetic control — Election Day Registration (Xu 2017) for the
cell-by-cell comparison and for what the replication turned up about rank
selection, and the durable case
benchmarks/cases/gsynth_xu_turnout.py.
A second case takes the estimator to a panel six times longer, at weekly
frequency, with ten staggered adoption dates: the state age-verification laws
of Lang et al. (2026). All four force settings match a commit-pinned
gsynth 1.2.1 to 6.4e-14 over 96 fits, and Algorithm 1 selects the same rank
in all sixteen outcome-by-force combinations. See
Generalized synthetic control — age verification laws (Lang et al. 2026) and
benchmarks/cases/gsynth_av_laws.py.
A third case is a cross-package one on real data: the cannabis-alcohol panel
of Ronczewski (2026). At a fixed factor number GSYNTH matches the paper’s
published gsynth estimate to 6.9e-18. The cross-validation is where the
two part – Algorithm 1 selects 4 where the paper’s gsynth 1.4.0 selects
1, and the estimate differs by a factor of 2.6 – so the case validates the
estimator at a fixed rank and records the selected rank as its own metric.
See benchmarks/cases/ronczewski_cannabis.py.
Core API#
- class mlsynth.GSYNTH(config: GSYNTHConfig | dict)#
Generalized synthetic control (Xu 2017) estimator.
- Parameters:
config (GSYNTHConfig or dict) – Configuration object. See
mlsynth.config_models.GSYNTHConfig.- Returns:
GSYNTHResults – Frozen container with the fitted factor structure, the effect by period since adoption, the per-unit breakdown, and the parametric bootstrap.
- fit() GSYNTHResults#
Run the GSYNTH pipeline end to end.
- class mlsynth.config_models.GSYNTHConfig(*, df: ~pandas.DataFrame, outcome: str, treat: str, unitid: str, time: str, display_graphs: bool = True, save: bool | str = False, counterfactual_color: ~typing.List[str] = <factory>, treated_color: str = 'black', plot: ~mlsynth.config_models.PlotConfig = <factory>, covariates: ~typing.List[str] | None = None, r: ~typing.Annotated[int | None, ~annotated_types.Ge(ge=0)] = None, r_max: ~typing.Annotated[int, ~annotated_types.Ge(ge=0)] = 5, force: ~typing.Literal['none', 'unit', 'time', 'two-way'] = 'two-way', two_way: bool | None = None, inference: bool = True, n_bootstrap: ~typing.Annotated[int, ~annotated_types.Ge(ge=1)] = 200, alpha: ~typing.Annotated[float, ~annotated_types.Gt(gt=0.0), ~annotated_types.Lt(lt=1.0)] = 0.05, seed: ~typing.Annotated[int, ~annotated_types.Ge(ge=0)] = 0, tol: ~typing.Annotated[float, ~annotated_types.Gt(gt=0)] = 1e-05, max_iter: ~typing.Annotated[int, ~annotated_types.Ge(ge=1)] = 500)#
Configuration for the generalized synthetic control (GSYNTH) estimator.
Implements Xu, Y. (2017), Generalized Synthetic Control Method: Causal Inference with Interactive Fixed Effects Models, Political Analysis 25(1):57-76. An interactive fixed effects model is fit to the never-treated units, each treated unit’s factor loadings are recovered from its own pre-treatment periods, and the untreated potential outcome is imputed from the two.
- Parameters:
covariates (list of str or None) – Time-varying covariates entering the outcome equation with a common coefficient vector.
Nonefits the outcome on fixed effects and factors alone.r (int or None) – Number of unobserved factors.
Noneselects it by the paper’s Algorithm 1, a leave-one-period-out cross-validation over the treated units’ pre-treatment periods.r_max (int) – Largest rank the cross-validation considers. The estimator lowers this further when the shortest treated pre-period history cannot identify that many loadings.
force ({“none”, “unit”, “time”, “two-way”}) – Which additive effects accompany the factors, named and coded as in gsynth and fect:
"none"keeps the grand mean alone,"unit"adds unit effects,"time"adds period effects,"two-way"adds both. Two-way is the specification Xu (2017) Table 2 reports and the default here. An effect that is switched off is not removed from the panel, so the factors absorb it; the choice moves the estimate, and applied work varies it deliberately.two_way (bool or None) – Superseded by
force, and kept so code written against the first release keeps running:Trueresolves to"two-way"andFalseto"time", with aDeprecationWarning. Supplying both is an error.inference (bool) – Run Algorithm 2, the parametric bootstrap.
n_bootstrap (int) – Bootstrap draws, used for both the prediction-error pool and the resampling loop. The paper uses 2,000.
alpha (float) – Two-sided significance level for the percentile intervals.
seed (int) – Seed for the bootstrap.
tol (float) – Convergence tolerance on the coefficient vector in the alternating least squares that fits the control-unit model.
max_iter (int) – Iteration cap for the same loop.
- model_config: ClassVar[ConfigDict] = {'arbitrary_types_allowed': True, 'extra': 'forbid'}#
Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].
References#
Xu, Y. (2017). “Generalized Synthetic Control Method: Causal Inference with Interactive Fixed Effects Models.” Political Analysis 25(1):57-76.
Bai, J. (2009). “Panel Data Models with Interactive Fixed Effects.” Econometrica 77(4):1229-1279.
Efron, B. (2012). “Bayesian Inference and the Parametric Bootstrap.” Annals of Applied Statistics 6(4):1971-1997.