SRC — Synthetic Regressing Control (Zhu 2023)#
- Estimator:
Synthetic Regressing Control (SRC) —
mlsynth.SRC- Source:
Zhu, Rong J. B. (2023), “Synthetic Regressing Control,” arXiv:2306.02584 [SRC2023].
- Replication type:
cross-validation against the author’s reference R implementation (
Code_SMC.R), plus Path A — the Basque / ETA study.- Status:
verified — the weight computation matches the reference to machine precision.
Validation strategy#
SRC ships an R reference implementation (Code_SMC.R): a function SMCV
that, given a matching matrix and predictor weights, computes the per-donor
univariate coefficients \(\widehat{\theta}_j\), the plug-in
\(\widehat{\sigma}^2\), and the Mallows / \(C_p\) box weights via
quadprog::solve.QP. That function is the ground truth. We reproduce it two
ways: a cell-by-cell numeric match of the weight computation, and the Basque
counterfactual it produces.
Cross-validation — machine precision#
We build the Abadie-Gardeazabal Basque matching matrix through
Synth::dataprep (the reference’s own path) and run SMCV with the
predictor weights fixed to one, so the oracle is deterministic. The mlsynth
weight computation mlsynth.utils.src_helpers.src_weights() is fed the
identical matrix. Every quantity agrees:
Quantity |
max |mlsynth − R| |
|---|---|
\(\widehat{\theta}_j\) (all donors) |
5.3e-15 |
box weights \(\mathbf{w}\) |
1.9e-14 |
combined \(\widehat{\theta}_j w_j\) |
2.0e-14 |
|
6.3e-14 |
\(\widehat{\sigma}^2\) |
2.4e-14 |
counterfactual (43 years) |
1.7e-13 |
The synthesis QP is the load-bearing step. mlsynth solves it with an exact
active-set box solver (mlsynth.utils.src_helpers.solver.solve_box_qp()),
the box-[0, 1] analogue of the repository’s simplex active-set solver. On
the Basque problem it reproduces solve.QP to 2e-14 with a KKT residual
below 1.5e-14 — tighter than solve.QP’s own residual — and pins the box
bounds exactly, so a dropped donor’s weight is exactly zero. Against a first-
order QP (OSQP) it is an order of magnitude faster at synthetic-control donor
sizes; against an interior-point QP (CLARABEL) it agrees on the objective but,
unlike CLARABEL, leaves no donor microscopically off its bound. The solver is
fuzz-tested against an independent cvxpy oracle over hundreds of random
problems (mlsynth/tests/test_src.py).
Path A — the Basque study#
Run through the public estimator on basedata/basque_data.csv (outcome-only
matching, treatment in 1970), SRC reproduces the Abadie-Gardeazabal result:
import pandas as pd
from mlsynth import SRC
df = pd.read_csv("basedata/basque_data.csv")
df["treat"] = ((df["regionname"] == "Basque Country (Pais Vasco)")
& (df["year"] >= 1970)).astype(int)
res = SRC({"df": df, "outcome": "gdpcap", "treat": "treat",
"unitid": "regionname", "time": "year",
"display_graphs": False}).fit()
gives a pre-period RMSE of about 0.048, a mean post-1969 ATT of about
\(-0.858\), and a 1997 gap of about \(-0.848\) (thousand-1986-USD per
capita), with the combined donor coefficients concentrated on Murcia
(\(0.63\)), Madrid (\(0.37\)) and Castilla y León (\(0.24\)). The
divergence traces the familiar economic cost of ETA terrorism. The durable case
is benchmarks/cases/src_basque.py.
The covariate / predictor-weight variant (Table 5)#
The paper’s Basque tables use the covariate-augmented variant (Algorithm 3) with
an Abadie predictor-weight (\(V\)) search. mlsynth exposes it as a seeded
opt-in — covariates with covariate_windows, the fit_window for the
outcome rows, and v_search="de". Configured this way the estimator rebuilds
the paper’s matching matrix from the raw panel and reproduces its Basque result:
res = SRC({
"df": df, "outcome": "gdpcap", "treat": "treat",
"unitid": "regionname", "time": "year",
"covariates": [...], # the 12 Abadie predictors
"covariate_windows": {...}, # 1964-69 schooling/invest, 1961-69 sectors
"fit_window": (1960, 1969), # time.optimize.ssr
"v_search": "de", "v_seed": 0,
}).fit()
gives Rioja-dominant weights with Madrid second and an ATT reaching \(\approx -2.4\) by 1997 — the Table 5 / Figure 1 donor structure and magnitude, in place of the outcome-only default’s Murcia / Madrid / Castilla y León and \(-0.86\).
Why it is opt-in, and what not to over-read. The \(V\) optimum is not
identified: on the Basque matching matrix a large manifold of \(V\) achieves
essentially the same pre-outcome fit while producing post-period effects that
range over a wide band, so the exact split among the top donors is not a
well-defined function of the data — a global differential-evolution search
reproduces the paper’s average placebo MSPE (Table 4) but not its per-region
cells, and the split shifts with the seed. The search is therefore seeded (a
given call is reproducible) and opt-in, leaving the \(C_p\)-identified
Algorithm 1 as the default. Read the v_search weights as one draw from the
paper’s non-identified \(V\) manifold, not as identified quantities.