Validation dashboard#

Every estimator in mlsynth is checked against the original authors’ code on real data. This page is generated from the pinned reference bundles the test suite asserts against, so the numbers here cannot drift from what CI enforces. Each row links to the reference implementation, the dataset (with checksum), and the mlsynth case that runs the check.

Coverage: 85 cross-validation checks against original implementations across 44 estimators – 32 reproduce the reference to display precision, 29 to within two percent. A further 4 are captured on the next daily run (see Pending capture). Per-estimator paper replications (Path A / Path B) are catalogued in Replications.

Legend: exact (agreement to display precision), tight (worst relative deviation \(\le 2\%\)), close (\(\le 10\%\)), and documented (looser, with a stated reason on the estimator’s replication page – typically an intrinsically extrapolated or weakly-identified quantity).

Summary#

Estimator

Checks

Agreement

Worst max |Δ|

?

2

1 tight · 1 documented

9e+03

BEAST

1

1 tight

0.15

BFSC

1

1 close

1

BVSS

1

1 tight

0.00041

CLUSTERSC

5

4 exact · 1 tight

0.036

COMPSC

2

1 exact · 1 tight

0.047

CSCM

1

1 tight

0.014

DPSC

1

1 exact

0

DROSC

1

1 exact

0

DSC

1

1 tight

0.01

FDID

1

1 exact

0.00032

GEOX

1

1 close

0.71

LINF

2

1 tight · 1 close

0.39

MAREX

1

1 tight

0.016

MASC

1

1 exact

4.6e-05

MCNNM

1

1 tight

0.81

MLSC

1

1 exact

1e-06

MTGP

1

1 close

0.03

MVBBSC

1

1 tight

20

MicroSynth

2

1 tight · 1 documented

1e+02

NSC

1

1 close

1.9

ORTHSC

2

1 exact · 1 close

0.037

PDA

5

4 exact · 1 close

0.056

PPSCM

2

1 tight · 1 documented

3

PROPSC

1

1 exact

0

PROXIMAL

3

1 exact · 2 tight

41

RESCM

2

2 tight

0.0013

ROLLDID

1

1 exact

0

RRSC

1

1 exact

0

SBC

2

1 close · 1 documented

1e+06

SCD

1

1 exact

0

SCMO

1

1 tight

0.011

SCUL

1

1 tight

0.14

SDID

1

1 tight

0.0016

SI

1

1 exact

0

SNN

1

1 exact

0

SPILLSYNTH

4

1 exact · 1 close · 2 documented

3.7e+02

SPSC

2

2 exact

0.0007

SSC

1

1 tight

0.001

SpSyDiD

2

1 exact · 1 close

0.094

TASC

1

1 documented

25

TSSC

1

1 tight

0.0004

VanillaSC

19

5 exact · 7 tight · 6 close · 1 documented

4.1

mlsynth.utils.inferutils.rae

1

1 exact

0

?#

Reference

Dataset

#

max |Δ|

Verdict

Case

14

0.028

tight

brabander_brexit_table1

R package tidysynth 0.2.0 (live run); the authors’ published numbers come from tidysynth <= 0.1.0 and are recorded in docs/replications/lamba_tigers.rst rather than pinned

tiger_reserves.csv (a529e3de9e6f…)

9

9e+03

documented — see notes

lamba_tigers

BEAST#

Reference

Dataset

#

max |Δ|

Verdict

Case

jeremylhour/alternative-synthetic-control-sparsity R (CalibrationLasso/OrthogonalityReg/ImmunizedATT)

augmented_cali_long.csv (974ae6ad6ab7…)

13

0.15

tight

beast_prop99

BFSC#

Reference

Dataset

#

max |Δ|

Verdict

Case

author appendix Stan (via Rscript + rstan)

3

1

close

bfsc_prop99

BVSS#

Reference

Dataset

#

max |Δ|

Verdict

Case

authors’ two-coordinate Gibbs (example2_fspda_2.R primitives), live run, captured

china_watches_long.csv (1ce8146af9a9…)

6

0.00041

tight

bvss_watches

CLUSTERSC#

Reference

Dataset

#

max |Δ|

Verdict

Case

SucreRouge/synth_control learn(method=’bayesian’) (live run, captured), num_sv=3

smoking_data.csv (a13dd4d5d6e4…)

8

0

exact — matches to display precision

bayesian_rsc_ref

Bayani RPCA-SC – the author’s own code, vendored verbatim (vendor/bayani_rpca_synth: FPCA.R + RPCA_2.py)

9

0

exact — matches to display precision

clustersc_rpca_germany

jehangiramjad/tslib RobustSyntheticControl (live run, captured), modelType=’svd’, kSingularValuesToKeep=3

smoking_data.csv (a13dd4d5d6e4…)

8

0.036

tight

pcr_rsc_ref

deshen24/panel-data-regressions var.var_est (homoskedastic + jackknife)

6

0

exact — matches to display precision

rsc_shen_coverage

scpi_pkg scest(w_constr={‘name’:’ridge’}) + df_EST

scpi_germany.csv (10b150fbcc2c…)

3

0

exact — matches to display precision

scpi_ridge_germany

COMPSC#

Reference

Dataset

#

max |Δ|

Verdict

Case

Boussim (2026) published tables (no replication package released)

34

0.047

tight

compsc_pennsylvania

Boussim (2026) csc_replication.R (quadprog::solve.QP on the stacked ALR log-odds), run live on basedata/pa_aeps_generation.csv

pa_aeps_generation.csv (7d7c0aebf62e…)

23

0

exact — matches to display precision

compsc_pennsylvania_r

CSCM#

Reference

Dataset

#

max |Δ|

Verdict

Case

Bonander CSCM_helper_functions.R (OSF osf.io/uvt5p, live run, captured)

4

0.014

tight

cscm_viszero

DPSC#

Reference

Dataset

#

max |Δ|

Verdict

Case

srho1/dpsc PrivateSC (differentially private SC)

4

0

exact — matches to display precision

dpsc_prop99

DROSC#

Reference

Dataset

#

max |Δ|

Verdict

Case

authors’ helpers.R sc() + DRoSC() (limSolve::lsei, live via Rscript)

10

0

exact — matches to display precision

drosc_basque

DSC#

Reference

Dataset

#

max |Δ|

Verdict

Case

Davidvandijcke/DiSCos DiSCo(), mean of 40 seeds at M = 10,000 (the package’s runif quadrature makes a single run a Monte Carlo draw)

dube_minwage.parquet (b93b3cbff573…)

66

0.01

tight

dsc_disco_xval

FDID#

Reference

Dataset

#

max |Δ|

Verdict

Case

Kathleen T. Li’s Fun_FDID.R (MKSC replication, live run, captured)

GDP.csv (487c7007cad0…)

7

0.00032

exact — matches to display precision

fdid_hongkong

GEOX#

Reference

Dataset

#

max |Δ|

Verdict

Case

mlsynth SDID (the engine GEOX wraps)

3

0.71

close

geox_sdid_equivalence

LINF#

Reference

Dataset

#

max |Δ|

Verdict

Case

LinfinitySC our(method=’inf’|’l1-inf’) (Wang, Xing & Ye 2025), BioAlgs/LinfinitySC

40

0.00041

tight

linf_crossval_ref

LinfinitySC our(method=’inf’) (Wang, Xing & Ye 2025), BioAlgs/LinfinitySC, lambda via param_selector(method=’inf’, n_folds=10)

smoking_data.csv (a13dd4d5d6e4…)

43

0.39

close

linf_prop99

MAREX#

Reference

Dataset

#

max |Δ|

Verdict

Case

jinglongzhao2/SCDesign (cardinality-K design, open quadprog, live run)

walmart_weekly_sales_covariates.csv (906fb3cd9e2f…)

7

0.016

tight

marex_walmart

MASC#

Reference

Dataset

#

max |Δ|

Verdict

Case

maxkllgg/masc masc(…, nogurobi=TRUE) (LowRankQP), live run, captured

basque_jasa.csv (b3f957771c8e…)

8

4.6e-05

exact — matches to display precision

masc_crossval

MCNNM#

Reference

Dataset

#

max |Δ|

Verdict

Case

susanathey/MCPanel R (mcnnm_cv, defaults)

smoking_data.csv (a13dd4d5d6e4…)

13

0.81

tight

mcnnm_prop99

MLSC#

Reference

Dataset

#

max |Δ|

Verdict

Case

leabottmer/multi-level-sc-estimator (mlSC_estimator, cvxpy+SCS)

4

1e-06

exact — matches to display precision

mlsc_bottmer

MTGP#

Reference

Dataset

#

max |Δ|

Verdict

Case

replication-package Stan (via Rscript + rstan)

3

0.03

close

mtgp_california

MVBBSC#

Reference

Dataset

#

max |Δ|

Verdict

Case

R package bsynth (bayesianSynth, predictor_match=FALSE, live run)

german_reunification.csv (f431666efbf3…)

4

20

tight

mvbbsc_germany

MicroSynth#

Reference

Dataset

#

max |Δ|

Verdict

Case

R microsynth (config A: match.out=trajectory)

56

1e+02

documented — see notes

microsynth_baltimore

R package microsynth (via Rscript)

6

0.097

tight

microsynth_seattle

NSC#

Reference

Dataset

#

max |Δ|

Verdict

Case

Tian (2023) NSC.R (vendored, live run, captured), a*=0.3, b*=0.7

smoking_data.csv (a13dd4d5d6e4…)

23

1.9

close

nsc_prop99

ORTHSC#

Reference

Dataset

#

max |Δ|

Verdict

Case

Fry GMM-SCE.R GMMSC() (R, live run, captured)

carbontax_fullsample_data.dta.txt (0df075d00fcc…)

15

0.037

close

gmmsce_carbontax

Fry OrthogonalizedSyntheticControl (R, live run, captured)

carbontax_fullsample_data.dta.txt (0df075d00fcc…)

5

0

exact — matches to display precision

orthsc_carbontax

PDA#

Reference

Dataset

#

max |Δ|

Verdict

Case

Authors’ Fun/L2relax.R (ishwang1/L2relax-PDA), per UK firm, reproduced via cvxpy/ECOS, live run captured

brexit_long.parquet (1e5997075c1e…)

4

1e-06

exact — matches to display precision

pda_brexit

R package pampe (pampe(), live run, captured)

HongKong.csv (ad5b35ff563a…)

6

3.6e-05

exact — matches to display precision

pda_hcw_hongkong

Authors’ Fun/L2relax.R (ishwang1/L2relax-PDA) reproduced via cvxpy/ECOS, live run captured

HongKong.csv (ad5b35ff563a…)

24

0

exact — matches to display precision

pda_hongkong

Shi & Huang fsPDA application script (zhentaoshi/fsPDA, live run, captured)

china_watches_long.csv (1ce8146af9a9…)

4

0.056

close

pda_luxurywatch

Authors’ Fun/L2relax.R (ishwang1/L2relax-PDA) reproduced via cvxpy/ECOS, live run captured

china_ppi_long.csv (cc4cda27e17b…)

64

1e-06

exact — matches to display precision

pda_ppi

PPSCM#

Reference

Dataset

#

max |Δ|

Verdict

Case

R augsynth::multisynth (live run, captured)

Teachingaugsynth.scv (59573a2dd46f…)

49

0.0022

tight

ppscm_paglayan

Ronczewski (2026) replication package, Results/csv/

14

3

documented — see notes

ronczewski_cannabis

PROPSC#

Reference

Dataset

#

max |Δ|

Verdict

Case

R package propsdid (via Rscript)

6

0

exact — matches to display precision

propsc_spain

PROXIMAL#

Reference

Dataset

#

max |Δ|

Verdict

Case

R gmm (authors’ analysis.Rmd, commit 3bcb5ec, reltol=1e-13)

4

41

tight

dr_proximal_brazil

KenLi93/proximal_sc_manuscript NC_nocov + NC_nocov_gmm (over-identified, Newey-West q=10), live run, captured

scpi_germany.csv (10b150fbcc2c…)

11

1e-06

exact — matches to display precision

proximal_germany_oid

authors’ proximal code (freshtaste/proximal, cloned)

3

0.014

tight

proximal_panic1907

RESCM#

Reference

Dataset

#

max |Δ|

Verdict

Case

scmrelax L2RelaxationCV (Liao-Shi-Zheng; github.com/metricshilab/scmrelax = github.com/YapengZheng/Relaxed_SC; MOSEK->CLARABEL; live run, captured)

balanced_gdp.csv (26fee37d55d9…)

6

0.0013

tight

rescm_balanced_gdp

scmrelax L2RelaxationCV (Liao-Shi-Zheng; github.com/metricshilab/scmrelax = github.com/YapengZheng/Relaxed_SC; MOSEK->CLARABEL; live run, captured)

10

0.00036

tight

rescm_relax_ref

ROLLDID#

Reference

Dataset

#

max |Δ|

Verdict

Case

lwdid.lwdid (Lee & Wooldridge DiD, live run, captured): prop99 common-timing (d, post, vce=None); castle staggered (gvar, control_group=’never_treated’, aggregate=’overall’, vce=None/hc3)

smoking_data.csv (a13dd4d5d6e4…)

9

0

exact — matches to display precision

rolldid_lw

RRSC#

Reference

Dataset

#

max |Δ|

Verdict

Case

reference R implementation of both RRSC regimes (fa.em + Huber-IRLS; stats::factanal + robustbase::ltsReg + Donoho-Johnstone selection), run LIVE via Rscript

8

0

exact — matches to display precision

rrsc_reference

SBC#

Reference

Dataset

#

max |Δ|

Verdict

Case

authors’ Germany.R (lsq detrend + trend_predict + Synth::synth ipop), live run, captured

german_reunification.csv (f431666efbf3…)

15

3.3e+04

documented — see notes

sbc_germany

authors’ SBC_HK.R (lsq detrend + trend_predict + Synth::synth ipop), live run, captured

hong_kong_handover.csv (4f3fea9b93ba…)

11

1e+06

close

sbc_hongkong

Notes (sbc_germany): The deviation is the reference solver’s, and it is in one place: the cyclical weight solve. The detrending and trend-forecast rows agree to 1.7e-14 of each series’ scale (R’s lm QR against numpy’s lstsq). On the weight solve the program is strictly convex with a unique optimum, mlsynth attains it – certified to 1.4e-6 by the convexity of the objective, and a cyclical sum of squares 2.6% lower than the authors’ Synth::synth ipop reaches at any tolerance – so the ATT and weight rows differ because the reference does not converge to the optimum. See docs/replications/sbc.rst.

Notes (sbc_hongkong): Same shape as the German panel. The detrending rows agree to 2.6e-14 of each series’ scale; the ATT, objective and weight rows differ because the authors’ Synth::synth ipop converges to a point about 6% worse in cyclical SSE on the identical strictly-convex program, where mlsynth attains the optimum (certified to 9.8e-8). See docs/replications/sbc.rst.

SCD#

Reference

Dataset

#

max |Δ|

Verdict

Case

base-R SCD (point estimator + corrected RC variance + in_C projection QP), reproduced on public CPS microdata

cps_lawa_arizona.parquet (87a228da0307…)

5

0

exact — matches to display precision

scd_cps

SCMO#

Reference

Dataset

#

max |Δ|

Verdict

Case

Tian-Lee-Panchenko Germany.R (fn_W solve.QP, live run, captured)

repgermany.csv (61a624e307e6…)

6

0.011

tight

scmo_germany

SCUL#

Reference

Dataset

#

max |Δ|

Verdict

Case

authors’ SCUL() (R, via Rscript + glmnet)

3

0.14

tight

scul_prop99

SDID#

Reference

Dataset

#

max |Δ|

Verdict

Case

synth-inference/synthdid R (synthdid_estimate)

smoking_data.csv (a13dd4d5d6e4…)

1

0.0016

tight

sdid_prop99

SI#

Reference

Dataset

#

max |Δ|

Verdict

Case

authors’ SI code (INFORMS opre.2025.1590.cd), vendored benchmarks/reference/synth_iv_OR25

20

0

exact — matches to display precision

si_prop99

SNN#

Reference

Dataset

#

max |Δ|

Verdict

Case

deshen24/syntheticNN (live run, captured), SyntheticNearestNeighbors(n_neighbors=1)

smoking_data.csv (a13dd4d5d6e4…)

14

0

exact — matches to display precision

snn_prop99

SPILLSYNTH#

Reference

Dataset

#

max |Δ|

Verdict

Case

Melnychuk-Andrii/Spillover-SCM inclusive SCM (scm_weights / runInclusiveSCM). The first four rows solve the same program to the simplex; the rows marked ‘reference as shipped’ are the authors’ ipop output, whose weights sum to 0.9666 and 1.1933

10

3.7e+02

documented — see notes

spillsynth_iscm_xval

jcao0/synthetic-control-spillover MATLAB spillover.csv (CA row)

13

5.7e-05

exact — matches to display precision

spillsynth_prop99

Mendez tutorial Rcpp sc_spillover (cmg777)

california_panel.csv (9d6a73e21f1a…)

4

3.4

documented — see notes

spillsynth_prop99_sar

Sakaguchi-Tagawa RcppArmadillo sc_spillover (method=sar, live run on the nonproprietary panel, captured)

sudan_panel.csv (722471a42b6c…)

5

0.41

close

spillsynth_sudan

SPSC#

Reference

Dataset

#

max |Δ|

Verdict

Case

qkrcks0218/SPSC R (single-proxy synthetic control)

4

0

exact — matches to display precision

spsc_panic

qkrcks0218/SPSC R (single-proxy synthetic control)

31

0.0007

exact — matches to display precision

spsc_prop99

SSC#

Reference

Dataset

#

max |Δ|

Verdict

Case

jcao0/staggered_synthetic_control (committed results_ssc.csv / Table1_eigenvalue.csv)

364

0.001

tight

ssc_guanajuato

SpSyDiD#

Reference

Dataset

#

max |Δ|

Verdict

Case

authors’ functions_ssdid fit_unit_weights / fit_time_weights under the canonical SDID convention (1/T_post post weights, 1/N_sp affected weights)

2

0

exact — matches to display precision

spsydid_lawa_diff

authors’ SDID weight functions (serenini/spatial_SDID functions_ssdid) + the notebook’s spatial WLS, via benchmarks.reference.spsydid_ref

20

0.094

close

spsydid_state_mc

TASC#

Reference

Dataset

#

max |Δ|

Verdict

Case

srho1/tasc TimeAwareSC (live run, captured; em_pre, naive init, set_seed(1))

smoking_data.csv (a13dd4d5d6e4…)

15

25

documented — see notes

tasc_prop99

TSSC#

Reference

Dataset

#

max |Δ|

Verdict

Case

authors’ _aux.R :: synth_control_est_demean (quadprog QP, live via Rscript)

basque_data.csv (54dfc49971c9…)

6

0.0004

tight

ferman_demeaned_basque

Notes (ferman_demeaned_basque): MSCa (TSSC’s simplex+intercept variant) IS Ferman-Pinto’s demeaned SC. Treatment 1975 is the identified regime (20 pre-periods > 16 donors); at 1970 (C>n) the demeaned-SC weights are non-unique and the two implementations legitimately diverge – see docs.

VanillaSC#

Reference

Dataset

#

max |Δ|

Verdict

Case

R package augsynth (live run, Kansas study)

kansas_ascm.csv (b026c651760c…)

8

0.0091

tight

ascm_kansas

R package scinference (conformal, live run, captured), cross-checked against the JASA supplement’s own functions

logfemrate.txt (fcdf30c41522…)

17

0.0002

tight

cwz_conformal

R package scinference (conformal) driven by the authors’ own simulation design

14

0.02

tight

cwz_conformal_mc

R package scinference (sc.cf t-test, live run, captured)

carbontax_data.dta (815787c1e448…)

3

0

exact — matches to display precision

cwz_ttest

R package scinference (ttest) driven by the authors’ own calibrated simulation design

carbontax_data.dta (815787c1e448…)

21

0.073

close

cwz_ttest_mc

Ferman (2021) JASA Table 1 (SC columns 1-4, OLS se col 5-8)

12

0.24

close

ferman_manyperiods

authors’ _aux.R synth_control_est + synth_control_est_demean (quadprog QPs, live via Rscript)

12

0.0003

tight

ferman_pinto_mc

mharoruiz/ibex R replication (01_functions/sc.R: limSolve::lsei simplex SC; scinference)

ibex_day_ahead_price.csv (18c69704e7ee…)

6

0

exact — matches to display precision

ibex_dap

Malo et al. scm.corner (SCM-Debug, live run, captured)

basque_mscmt.csv (3aca35dc9b55…)

3

0.00048

tight

malo_basque

Malo et al. scm.corner (SCM-Debug, live run, captured)

augmented_cali_long.csv (974ae6ad6ab7…)

6

0.0048

tight

malo_prop99

R package MSCMT (live run, captured)

basque_mscmt.csv (3aca35dc9b55…)

4

4e-05

exact — matches to display precision

mscmt_basque

authors’ wsoll1 (R, via Rscript + LowRankQP)

3

0.00098

exact — matches to display precision

pensynth_prop99

scpi_pkg scdata(cointegrated_data=True)+scpi CI_all_gaussian

scpi_germany.csv (10b150fbcc2c…)

13

0.11

close

scpi_germany_pi

Python package scpi_pkg

15

4.3e-05

exact — matches to display precision

scpi_staggered

authors’ replication: SyntheticControlMethods (Synth, pen=’auto’ + covariates)

secession_autonomy.csv (b3509727778e…)

4

4.1

documented — see notes

secession_scm

R Synth (j-hai/Synth, synth + synth_inference)

9

0.26

close

synth_jhai_prop99

R package Synth::synth

california_panel.csv (9d6a73e21f1a…)

8

0.02

tight

synth_prop99

Andersson (2019) AEJ:EP 11(4), Section III reported values

carbontax_data.dta (815787c1e448…)

4

0.028

close

vanillasc_carbontax

Synth (uniform custom.v) + tidysynth (ADH spec)

7

3.6

close

vanillasc_xval_references

mlsynth.utils.inferutils.rae#

Reference

Dataset

#

max |Δ|

Verdict

Case

the authors’ RAE.R (JPE replication package), live run

12

0

exact — matches to display precision

cwz_rae

Pending capture#

These cross-validation cases are wired up but their reference had not been captured when this page was last generated; the daily action records them once its toolchain provisions.

Case

Reference

brazil_vaccine_scm_vs_proximal

lto_refined_placebo

independent reproduction of tsudijon/LeaveTwoOutSCI LTO pair loop (outcome-only SC via LowRankQP), all three empirical applications

marex_scdesign_sim

jinglongzhao2/SCDesign (live run: Section 5 generation block + Synthetic_Experiment_Cardinality_Constraint on the open quadprog backend)

ppscm_cs_real_panels