Atlas
statminds
Causal GLM (Comparative Case Study Model)The underlying model family class (e.g. GLM, linear model, categorical matrix, log-linear).Parametric ReferenceStatistical methods that assume a specific probability distribution family (typically normal).12-stage workflow

Synthetic Control Method

The engine for Comparative Discovery in Single-Unit Trials. This model audits a single treated unit (e.g., a city or clinic) by constructing a 'Synthetic Shadow' from a donor pool of controls, reveal the definitive impact of local policy or intervention.

Model familyCausal GLM (Comparative Case Study Model)
Hypothesistwo-tailed
AliasesSynthetic Controls · SCM · Abadie’s Comparative Model
G1
Shadow-Unit Construction
Mathematically assemble a combination of control units that perfectly mimics the treated unit's pre-intervention pulse.
G2
Dynamic Causal Audit
Reveal the 'Trajectory Gap' between the treated unit and its synthetic counterfactual after the intervention date.
G3
Policy Sensitivity Discovery
Isolate the impact of a large-scale change when randomized control trials are logistically impossible.
Visual Overview Dashboard
1

What is it?

Synthetic Control Method is a causal inference method designed to estimate treatment effects by adjusting for confounding in observational studies.

The engine for Comparative Discovery in Single-Unit Trials. This model audits a single treated unit (e.g., a city or clinic) by constructing a 'Synthetic Shadow' from a donor pool of controls, reveal the definitive impact of local policy or intervention.

2

Goals & Indications

  • Shadow-Unit Construction: Mathematically assemble a combination of control units that perfectly mimics the treated unit's pre-intervention pulse.
  • Dynamic Causal Audit: Reveal the 'Trajectory Gap' between the treated unit and its synthetic counterfactual after the intervention date.
  • Policy Sensitivity Discovery: Isolate the impact of a large-scale change when randomized control trials are logistically impossible.
3

Core Idea Diagram

InterventionSynthetic ControlActual Treated
4

Claims tested

H₀: H₀: No treatment effect (treated unit's post-intervention outcome equals synthetic control's counterfactual)
Hₐ: Hₐ: Treatment effect exists (treated unit diverges from synthetic control post-intervention)
5

How it works

  1. Identify donor pool units resembling treated unit in pre-treatment variables.
  2. Compute optimal weights W mapping donor outcomes to treated pre-treatment path.
  3. Construct the synthetic counterfactual path using weighted donor pool outcomes.
  4. Compare post-treatment outcomes between actual treated unit and synthetic control.
6

Assumptions

Single Treated Unit: Method designed for one or few treated units
No Interference Between Units: Treatment of one unit does not affect outcomes of other units
Good Pre-Intervention Fit: Synthetic control closely matches treated unit before treatment
7

Important Note

SCM estimates the counterfactual outcome for a treated unit by constructing a weighted combination of control units that closely approximates the treated unit's pre-intervention characteristics. The treatment effect is the post-intervention gap between the treated unit and its synthetic control. Inference typically relies on permutation-based placebo tests rather than classical hypothesis testing.

8

Worked Example

Donor UnitWeightPre-MSPE MatchPost-Treatment Gap
State A0.420.021-8.45%
State B0.35
State C0.23
Interactive Sandbox

Synthetic Control Counterfactual Path

Observe how the actual treated unit path diverges from the weighted synthetic counterfactual control after the intervention.

Post-Intervention Effect1.20
Synthetic Control Estimate
Pre-intervention years: 6 years (2018–2023)
Post-intervention years: 5 years (2024–2028)
Final year treatment gap: 1.200
Intervention20182019202020212022202320242025202620272028
Treated Unit
Synthetic Control
The 12-Stage Precision Workflow
01Trajectory Departure
Hypotheses
We test if the treated unit 'Peels Away' from its synthetic shadow after the intervention strike.
02Pre-Trend Alignment
Assumptions
The ultimate prerequisite: the synthetic shadow must perfectly track the treated unit *before* the intervention—the 'Parallel Pulse' mandate.
03MSPE Forensics
Diagnostics
Utilizing Mean Squared Prediction Error (MSPE) to audit the quality of the 'Pre-Fit'—a high MSPE indicates a failing shadow.
04focus
Auditing the impact of a city-wide FlowMotion mandate by constructing a 'Synthetic City' from 10 other similar urban centers.
05Diff-in-Diff Pivot
Alternatives
Knowing when to switch to Difference-in-Differences if you have multiple treated units and a consistent group-level trend.
06Placebo Strikes
Significance
Executing 'Iterative Placebo Tests'—running the shadow math on every control unit to see if the treated effect is truly unique or just a random ripple.
07The Gap Magnitude
Effect Size
Interpreting the vertical distance between the Treated and Synthetic lines as the definitive clinical or economic discovery.
08Donor Density
Sample Size
Calculating the balance: you need a rich 'Donor Pool' of controls to assembly a shadow that reaches statistical authority.
09The Shadow Plot
Reporting
Providing the 'Intervention vs. Synthetic' trajectory plot—the only valid way to tell the story of a single-unit discovery.
10Synth / SC Logic
Software
Executing 'synth()' commands, ensuring the 'V-weights' correctly represent the importance of each pre-intervention marker.
11focus
The fatal error of excluding donors just because they don't help the story—which inadvertently contaminates the causal audit.
12focus
Tracing the model back to Alberto Abadie (2003) and the revolution of comparative forensics in social science.
01Hypothesis test logic

Hypotheses

Pragmatic null and alternative hypotheses defined in mathematical notation.

A hypothesis is a question sharpened to a point. Ambiguity is the enemy of inference.
Logic Core
Null · H₀

H₀: No treatment effect (treated unit's post-intervention outcome equals synthetic control's counterfactual)

Alternative · Hₐ

Hₐ: Treatment effect exists (treated unit diverges from synthetic control post-intervention)

Why it matters two-tailed

SCM estimates the counterfactual outcome for a treated unit by constructing a weighted combination of control units that closely approximates the treated unit's pre-intervention characteristics. The treatment effect is the post-intervention gap between the treated unit and its synthetic control. Inference typically relies on permutation-based placebo tests rather than classical hypothesis testing.

02Model diagnostics

Assumptions

The core mathematical criteria needed to ensure that statistical testing remains unbiased and valid.

Build your analysis on rock, not sand. Verify the mathematical foundation before building the model.
Integrity Shield
7
Assumptions
4
Critical / High Severity
How to check
Quick
Count number of treated units in dataset. SCM works best with 1 treated unit and many (20+) control units. With multiple treated units, aggregate or use separate analyses.
Rigorous
Assess trade-off: more treated units → fewer donor pool controls → worse pre-intervention fit. Consider whether units received treatment at same time or staggered.
If violated
If many treated units: (1) Use difference-in-differences (DiD) instead, which handles multiple treated units naturally; (2) Aggregate treated units into single composite unit if appropriate; (3) Use generalized synthetic control methods for multiple treated units; (4) Conduct separate SCM analyses for each treated unit if sufficient controls. If all units treated: SCM not applicable - use interrupted time series or within-unit comparisons.
How to check
Quick
Assess spillover potential: Are units geographically proximate? Do they trade/compete? Could treatment in one unit affect others? Example: tobacco tax in California might affect neighboring states if residents cross borders to buy cigarettes.
Rigorous
Examine outcomes in geographically/economically close control units for unusual patterns post-treatment; test whether distance from treated unit predicts outcome changes; review literature on spillover mechanisms
If violated
If spillovers detected: (1) Exclude affected controls from donor pool (e.g., exclude neighboring states); (2) Model spillovers explicitly by distance or network connections; (3) Increase geographic distance criteria for donor pool; (4) Use spatial SCM methods accounting for spillovers; (5) Acknowledge limitation and interpret as 'direct effect' not total effect. If severe spillovers, SCM may not be appropriate.
How to check
Quick
Plot treated vs synthetic control in pre-treatment period. Calculate Root Mean Squared Prediction Error (RMSPE) for pre-treatment period. Visual inspection: trajectories should be nearly identical. RMSPE should be small relative to outcome variance.
Rigorous
Calculate pre-treatment RMSPE and compare to placebo distribution; examine individual predictor balance (synthetic weights should produce covariate match); test whether pre-treatment trends are parallel; assess whether synthetic control requires extreme weights (convexity constraint may fail)
If violated
If poor pre-fit: (1) Expand donor pool (include more controls); (2) Adjust predictor variables - add lagged outcomes, different time periods, or omitted predictors; (3) Use different time windows for matching; (4) Try augmented synthetic control (SC + ridge penalty); (5) Use matrix completion methods; (6) If fit remains poor, acknowledge SCM may not be appropriate for this setting. Poor fit indicates donor pool cannot replicate treated unit - inference unreliable.
How to check
Quick
Examine pre-treatment trend for unusual patterns before treatment date. Plot outcome over time and look for breaks/jumps before official treatment. Check if treatment was publicly announced months/years before implementation.
Rigorous
Conduct placebo tests with pre-treatment cutoff dates; test for structural breaks in pre-treatment period; examine news/policy announcements for anticipation signals; run event study with leads (pre-treatment indicators)
If violated
If anticipation detected: (1) Move treatment date backward to when anticipation began (require sufficient pre-anticipation data); (2) Exclude anticipation period from analysis entirely; (3) Model anticipation explicitly with separate effect estimates; (4) Acknowledge bias toward zero (if anticipation causes units to converge before official treatment). Cannot fix if anticipation began too early with insufficient pre-treatment data.
How to check
Quick
Assess whether external shocks (recessions, policy changes, natural disasters) affected treated unit differently than controls after treatment. Check for concurrent interventions or events.
Rigorous
Examine residuals for control units post-treatment (should remain stable); test whether other policies/events occurred simultaneously; conduct robustness checks excluding periods with external shocks; use time-varying covariates if available
If violated
If unstable relationships: (1) Use time-varying synthetic controls that re-weight over time; (2) Adjust for concurrent shocks if observable; (3) Conduct sensitivity analysis excluding shock periods; (4) Use interactive fixed effects models that allow factor loadings to vary; (5) Restrict inference to periods without shocks. Acknowledge limitation if major concurrent confounders unaccounted for.
How to check
Quick
Count number of potential donor units. Check if donor pool is diverse enough to span treated unit's covariate space. SCM works best with 20+ donors, struggles with < 10.
Rigorous
Examine convex hull of donor pool covariates - does treated unit fall within? Check whether synthetic weights are distributed across multiple controls or concentrated on 1-2 units (indicates sparse donor pool); calculate predictor balance after weighting
If violated
If small donor pool: (1) Expand eligibility (relax exclusion criteria); (2) Use matching on key covariates followed by SCM; (3) Use penalized synthetic control allowing negative weights (removes convexity); (4) Try matrix completion methods; (5) Acknowledge extrapolation and increased uncertainty. With < 5 donors, SCM likely unreliable - consider alternative methods.
How to check
Quick
Check missingness patterns: % missing for each unit and time period. SCM requires outcome data for treated unit and donors across all pre-treatment periods used for matching.
Rigorous
Test whether missingness is related to outcome levels (MNAR); examine whether missingness varies systematically across units; assess sensitivity to imputation methods if used
If violated
If missing data: (1) Exclude time periods with high missingness from predictor set; (2) Use imputation (linear interpolation, matrix completion, multiple imputation) - but test sensitivity; (3) Exclude donors with high missingness; (4) Use matrix completion SCM variant designed for missing data; (5) Report sensitivity to missingness assumptions. Avoid if missingness > 20% in pre-period.
03Residual Forensics

Diagnostics

Checking residual plots and indices to examine model deviations and ensure standard error integrity.

Trust, but verify. The outliers often hold more truth than the averages.
System Health
Essential checks
  1. Pre-treatment fit plot (treated vs synthetic control over time)
  2. Pre-treatment RMSPE (root mean squared prediction error)
  3. In-space placebo tests (apply SCM to control units, compare effect sizes)
  4. Predictor balance table (covariates: treated vs synthetic control)
Recommended checks
  1. In-time placebo tests (move treatment date backward, test for false effects)
  2. Placebo distribution plot (treated effect vs all placebo effects)
  3. Leave-one-out robustness (exclude each donor sequentially)
  4. Weights distribution (which donors contribute most to synthetic control)
  5. Post-treatment gap plot with trend extrapolation
  6. Permutation-based inference p-value
04Live Instances

Applied Minds

Review concrete study examples, data layout guidelines, and copy executable syntax scripts.

Theory is the map. Practice is the terrain. Simulation bridges the gap.
Applied Wisdom
Example 01

Effect of Tobacco Control Program on Cigarette Sales (California Proposition 99)

Research question: Did California's 1988 tobacco control program (Proposition 99) causally reduce cigarette consumption? Design: Comparative case study with California as treated unit and 38 other US states as donor pool. Outcome: Per-capita cigarette sales (packs per capita) 1970-2000. Pre-treatment: 1970-1988 (19 years). Post-treatment: 1989-2000 (12 years). Predictors: Lagged cigarette sales, beer consumption, retail price, income. This replicates the seminal Abadie, Diamond & Hainmueller (2010) study.

Total n39
Outcome ScalePer-capita cigarette sales (packs/person)
# Synthetic Control Method: California Prop 99 → cigarette sales
# Based on Abadie, Diamond & Hainmueller (2010)

library(Synth)         # Core SCM package
library(tidyverse)
library(reshape2)

set.seed(2025)

# === STEP 1: Simulate realistic panel data ===
# 39 states (California + 38 donors), 1970-2000 (31 years)

states <- c("California", paste0("State", 1:38))
years <- 1970:2000
n_states <- 39
n_years <- 31
treatment_year <- 1989  # Prop 99 implemented

# Generate baseline cigarette sales (downward trend + state effects)
data_list <- list()

for (i in 1:n_states) {
  state_effect <- rnorm(1, 120, 15)  # State-specific baseline (packs/capita)
  time_trend <- -1.2  # General decline over time
  volatility <- runif(1, 3, 6)
  
  # Pre-treatment
  sales_pre <- state_effect + time_trend * (1:(treatment_year - 1970)) + 
               rnorm(treatment_year - 1970, 0, volatility)
  
  # Post-treatment
  if (states[i] == "California") {
    # California: additional -20 pack decline due to Prop 99
    treatment_effect <- -20
    sales_post <- state_effect + time_trend * ((treatment_year - 1970 + 1):n_years) + 
                  treatment_effect + 
                  rnorm(n_years - (treatment_year - 1970), 0, volatility)
  } else {
    # Controls: continue baseline trend
    sales_post <- state_effect + time_trend * ((treatment_year - 1970 + 1):n_years) + 
                  rnorm(n_years - (treatment_year - 1970), 0, volatility)
  }
  
  sales <- c(sales_pre, sales_post)
  
  data_list[[i]] <- data.frame(
    state = states[i],
    state_num = i,
    year = years,
    sales = pmax(sales, 20),  # Floor at 20 packs
    beer = rnorm(n_years, 25, 5),  # Predictor: beer consumption
    price = seq(1.5, 3.5, length.out = n_years) + rnorm(n_years, 0, 0.2),
    income = seq(20, 40, length.out = n_years) + rnorm(n_years, 0, 2)
  )
}

data <- bind_rows(data_list)

# Add treatment indicator
data$treated <- ifelse(data$state == "California" & data$year >= treatment_year, 1, 0)

cat("=== DATA STRUCTURE ===", "\n")
cat("States:", n_states, "(1 treated, 38 donors)\n")
cat("Years:", min(years), "-", max(years), "\n")
cat("Pre-treatment periods:", treatment_year - min(years), "\n")
cat("Post-treatment periods:", max(years) - treatment_year + 1, "\n\n")

# === STEP 2: Prepare data for Synth package ===
dataprep_out <- dataprep(
  foo = as.data.frame(data),
  predictors = c("beer", "price", "income"),
  predictors.op = "mean",
  time.predictors.prior = 1970:(treatment_year - 1),
  dependent = "sales",
  unit.variable = "state_num",
  unit.names.variable = "state",
  time.variable = "year",
  treatment.identifier = 1,  # California
  controls.identifier = 2:39,
  time.optimize.ssr = 1970:(treatment_year - 1),
  time.plot = 1970:2000
)

cat("=== SYNTHETIC CONTROL OPTIMIZATION ===", "\n")

# === STEP 3: Estimate synthetic control weights ===
synth_out <- synth(dataprep_out)

cat("\nConvergence:", ifelse(synth_out$solution.w.star[1] > 0, "SUCCESS", "FAILED"), "\n")

# Extract weights
weights <- data.frame(
  state = dataprep_out$Y0names,
  weight = synth_out$solution.w
)

cat("\n=== DONOR WEIGHTS(top 5) ===", "\n")
print(head(weights %>% arrange(desc(weight)), 10))

cat("\nNumber of donors with weight > 0.01:", sum(weights$weight > 0.01), "\n")

# === STEP 4: Assess pre-treatment fit ===
cat("\n=== PRE-TREATMENT FIT ===", "\n")

# Calculate RMSPE
pre_treatment_years <- 1970:(treatment_year - 1)
treated_pre <- dataprep_out$Y1plot[as.character(pre_treatment_years)]
synthetic_pre <- dataprep_out$Y0plot %*% synth_out$solution.w

rmspe_pre <- sqrt(mean((treated_pre - synthetic_pre)^2))
cat("Pre-treatment RMSPE:", round(rmspe_pre, 3), "packs\n")

# Predictor balance
cat("\n=== PREDICTOR BALANCE ===", "\n")
synth_tables <- synth.tab(dataprep.res = dataprep_out, synth.res = synth_out)
print(synth_tables$tab.pred)

# === STEP 5: Visualize results ===
cat("\n=== GENERATING PLOTS ===", "\n")

# Main path plot
path.plot(synth.res = synth_out, dataprep.res = dataprep_out,
          Ylab = "Per-Capita Cigarette Sales(packs)",
          Xlab = "Year",
          Legend = c("California", "Synthetic California"),
          Legend.position = "topright",
          Main = "California vs Synthetic Control: Cigarette Sales")
abline(v = treatment_year, lty = 2, col = "red")

# Gap plot (treatment effect over time)
gaps <- dataprep_out$Y1plot - (dataprep_out$Y0plot %*% synth_out$solution.w)
gaps.plot(synth.res = synth_out, dataprep.res = dataprep_out,
          Ylab = "Gap in Cigarette Sales(packs)",
          Xlab = "Year",
          Main = "Gap: California - Synthetic California")
abline(v = treatment_year, lty = 2, col = "red")
abline(h = 0, lty = 2, col = "gray")

# === STEP 6: Calculate treatment effects ===
post_treatment_years <- treatment_year:2000
post_gaps <- gaps[as.character(post_treatment_years)]

cat("\n=== TREATMENT EFFECTS ===", "\n")
cat("Average post-treatment effect:", round(mean(post_gaps), 2), "packs\n")
cat("Effect in final year(2000):", round(post_gaps[length(post_gaps)], 2), "packs\n")

# === STEP 7: In-space placebo tests (permutation inference) ===
cat("\n=== IN-SPACE PLACEBO TESTS ===", "\n")
cat("Running placebo SCM for all 38 donor states(may take 30-60 seconds)...\n")

placebo_effects <- numeric(38)
placebo_rmspe_pre <- numeric(38)
placebo_rmspe_post <- numeric(38)

for (i in 1:38) {
  # Run SCM with donor i as "treated"
  tryCatch({
    dataprep_placebo <- dataprep(
      foo = as.data.frame(data),
      predictors = c("beer", "price", "income"),
      predictors.op = "mean",
      time.predictors.prior = 1970:(treatment_year - 1),
      dependent = "sales",
      unit.variable = "state_num",
      unit.names.variable = "state",
      time.variable = "year",
      treatment.identifier = i + 1,  # Placebo treated
      controls.identifier = setdiff(2:39, i + 1),
      time.optimize.ssr = 1970:(treatment_year - 1),
      time.plot = 1970:2000
    )
    
    synth_placebo <- synth(dataprep_placebo, verbose = FALSE)
    
    # Calculate placebo effect (mean post-treatment gap)
    gaps_placebo <- dataprep_placebo$Y1plot - 
                    (dataprep_placebo$Y0plot %*% synth_placebo$solution.w)
    placebo_effects[i] <- mean(gaps_placebo[as.character(post_treatment_years)])
    
    # Pre and post RMSPE
    placebo_rmspe_pre[i] <- sqrt(mean(gaps_placebo[as.character(pre_treatment_years)]^2))
    placebo_rmspe_post[i] <- sqrt(mean(gaps_placebo[as.character(post_treatment_years)]^2))
  }, error = function(e) {
    placebo_effects[i] <- NA
  })
}

# True California effect
true_effect <- mean(post_gaps)

# p-value: proportion of placebo effects as extreme as true effect
p_value_twosided <- mean(abs(placebo_effects) >= abs(true_effect), na.rm = TRUE)
p_value_onesided <- mean(placebo_effects <= true_effect, na.rm = TRUE)

cat("\nTrue California effect:", round(true_effect, 2), "packs\n")
cat("Placebo effects range:", round(range(placebo_effects, na.rm = TRUE), 2), "\n")
cat("P-value(two-sided):", round(p_value_twosided, 3), "\n")
cat("P-value(one-sided):", round(p_value_onesided, 3), "\n")
cat("Rank:", sum(placebo_effects <= true_effect, na.rm = TRUE), "out of", sum(!is.na(placebo_effects)) + 1, "\n")

# Plot placebo distribution
hist(placebo_effects, breaks = 15, col = "lightblue", 
     main = "Placebo Distribution(In-Space Tests)",
     xlab = "Average Post-Treatment Effect(packs)",
     xlim = c(min(c(placebo_effects, true_effect), na.rm = TRUE) - 5,
              max(c(placebo_effects, true_effect), na.rm = TRUE) + 5))
abline(v = true_effect, col = "red", lwd = 3, lty = 2)
text(true_effect, par("usr")[4] * 0.9, "California", col = "red", pos = 4)

# === STEP 8: Pre/Post RMSPE ratio test ===
# Exclude placebos with poor pre-fit (RMSPE ratio > 2)
rmspe_ratio_california <- sqrt(mean(post_gaps^2)) / rmspe_pre
rmspe_ratios_placebo <- placebo_rmspe_post / placebo_rmspe_pre

cat("\n=== RMSPE RATIO TEST ===", "\n")
cat("California RMSPE ratio(post/pre):", round(rmspe_ratio_california, 2), "\n")
cat("Mean placebo RMSPE ratio:", round(mean(rmspe_ratios_placebo, na.rm = TRUE), 2), "\n")

# Filter placebos with good pre-fit
good_prefit <- which(placebo_rmspe_pre < 2 * rmspe_pre)
cat("Placebos with good pre-fit:", length(good_prefit), "/", length(placebo_effects), "\n")

if (length(good_prefit) > 0) {
  p_value_filtered <- mean(rmspe_ratios_placebo[good_prefit] >= 
                           rmspe_ratio_california, na.rm = TRUE)
  cat("P-value(RMSPE ratio, filtered):", round(p_value_filtered, 3), "\n")
}

cat("\n=== APA-STYLE REPORTING ===", "\n")
cat(paste0(
  "We used the synthetic control method to estimate the causal effect of ",
  "California's Proposition 99 tobacco control program on per-capita cigarette ",
  "sales. Using 38 US states as donors, we constructed a synthetic California ",
  "that closely matched pre-intervention sales(1970-1988, RMSPE = ",
  round(rmspe_pre, 2), " packs). ",
  "The synthetic control was a weighted combination of ",
  sum(weights$weight > 0.01), " states. ",
  "Post-intervention(1989-2000), California's sales diverged substantially from ",
  "the synthetic control, with an average reduction of ",
  abs(round(true_effect, 1)), " packs per capita. ",
  "Permutation-based inference using in-space placebo tests yielded p = ",
  round(p_value_onesided, 3), ", indicating the effect was unlikely due to chance. ",
  "These findings provide strong evidence that Proposition 99 causally reduced ",
  "cigarette consumption in California."
))
Interpretation Blueprint

The synthetic control analysis estimated that California's Proposition 99 tobacco control program causally reduced per-capita cigarette sales by approximately 20 packs per year post-intervention (1989-2000). The synthetic California (constructed from 38 donor states) provided excellent pre-treatment fit (RMSPE = 2-4 packs), indicating the donor pool could credibly approximate California's counterfactual trajectory. Post-intervention, California's sales diverged substantially from the synthetic control, with the gap widening over time. Permutation-based inference using in-space placebo tests showed that California's effect was more extreme than 95% of placebo effects (p < 0.05), providing strong evidence against the null hypothesis of no effect. The RMSPE ratio (post/pre) for California was substantially larger than most placebos, further supporting causal interpretation. These findings replicate the seminal Abadie et al. (2010) study and demonstrate SCM's power for comparative case studies.

05Tactical Pivots

Alternatives

Structured fallback pathways for choosing alternative tests when normality or slopes requirements fail.

When the path is blocked, pivot. Rigor is not rigidity; it is the intelligent adaptation to reality.
Adaptive Strategy
Measurement Precision Ladder Ideal · Continuous Unit Matrix
Ratio
Maintain SCM logic. Optimal for auditing the 'Shadow Gap' in high-fidelity aggregate markers (e.g., mortality rates).
Peak Precision
Interval
Ideal for Policy Evaluation. Ensure the 'Pre-Trend' alignment is clinically meaningful.
Standard Signal
Temporal Trajectory Audit Long-Term Single-Unit Snapshot
Single-Unit Strike
Policy change.
Stay with SCM. Construct a synthetic counterfactual city/clinic from donor pools.
Multi-Unit Strike
Widespread change.
Pivot to Difference-in-Differences (DiD) if you have multiple treated units and stable parallel trends.
Adaptive Technical Safeguards · adaptive safeguards
poor pre trend alignment
  • Augmented Synthetic Control — Incorporate negative weights or ridge penalties to force better pre-fit.
  • Generalized Synthetic Control — Utilize latent factors to model the underlying pulse of the donor pool.
donor pool is sparse
  • Interrupted Time Series (ITS) — Pivot if you have no donors and must rely purely on the treated unit's own history.
  • Synthetic DiD — A hybrid model that increases power when donor counts are low.
non stationary trajectories
  • Differencing Audit — Mathematically 'level' the treated and synthetic lines before the gap strike.
06Adjusted Comparisons

Post-hoc

Group mean comparisons and correction controls (e.g. Tukey HSD, Bonferroni) to protect against Family-Wise Error Rates.

The omnibus test opens the door; post-hoc analysis explores the room.
Forensic Detail
Adjusted Comparisons
  • Placebo tests: apply method to untreated units (permutation inference)
  • Leave-one-out: exclude each donor unit and check robustness
  • In-time placebo: apply treatment at fake pre-treatment dates
  • Compare pre-treatment fit (RMSPE ratio)
  • Sensitivity to donor pool selection
Interpretation Guidelines

Synthetic control compares treated unit to weighted combination of controls. Traditional post-hoc tests are not applicable.

07Standardized scale impact

Effect Size

Understanding effect sizes (e.g., Cohen's d, Partial Eta-Squared) and clinical impact benchmarks.

Significance is noise. Magnitude is the signal. Measure the impact, not just the probability.
Impact Magnitude

Raw difference between treated unit and synthetic control at each time point. Most transparent measure.

Mean gap across post-treatment period. Summarizes overall magnitude but loses temporal dynamics.

Sum of all post-treatment gaps. Useful for assessing total impact (e.g., total lives saved, revenue lost).

Relative effect size. Facilitates comparison across studies with different scales.

Recommended Metric: Report both average gap and trajectory plot. Average gap provides single summary statistic; trajectory shows whether effect grows/shrinks/stabilizes over time.
Small
0.2
Medium
0.5
Large
0.8
0.50
Report both average gap and trajectory plot. Average gap provides single summary statistic; trajectory shows whether effect grows/shrinks/stabilizes over time.
Recommended Measure
4
Available Metrics
ReportUse Report both average gap and trajectory plot. Average gap provides single summary statistic; trajectory shows whether effect grows/shrinks/stabilizes over time. to represent clinical impact magnitude.
08Statistical Power

Sample Size

Guidelines for minimum sample requirements and power analysis parameters.

An underpowered study is an ethical failure. Respect the data by collecting enough of it.
Power Protocol
Floor Requirements

The 'Shadow Stability' Mandate: A donor pool of at least 10-15 control units is essential to construct a high-fidelity synthetic shadow. Single-unit causal discovery fails if the donor diversity is too shallow.

Effect SizeParametersRequired n
Small EffectLow Signal (10% shift)T0 ≈ 20, J ≈ 30
Medium EffectModerate Signal (25% shift)T0 ≈ 10, J ≈ 15
Large EffectStrong Signal (50% shift)T0 ≈ 5, J ≈ 8
Key considerations

The 'Placebo Strike': SCM significance is proven through placebo iterations on all donor units. If your donor pool is small (J < 10), your p-value resolution is capped at 1/11 (p=0.09), making it impossible to reach 'Elite' significance thresholds.

G*Power StrategyBenchmark: Comparative Case Study (SCM). Parameters: Pre-intervention periods (T0), Post-intervention periods (T1), Number of donor units (J), Signal magnitude. Note: Power is dictated by the 'Pre-Fit Accuracy' (RMSPE).
09APA narrative blueprint

Reporting

How to compile statistical results into publication prose matching APA and journal style guides.

Data does not speak for itself. It requires a translator. Be clear, be precise, be honest.
Narrative Arc
Worked APA paragraph example
We used the synthetic control method to estimate the causal effect of California's Proposition 99 tobacco control program (implemented 1989) on per-capita cigarette sales. The donor pool consisted of 38 US states. We constructed a synthetic California by optimizing weights to minimize pre-treatment mean squared prediction error over 19 years (1970-1988), matching on lagged cigarette sales (1975, 1980, 1988), beer consumption, retail price, and income per capita. The synthetic control achieved excellent pre-treatment fit (RMSPE = 2.1 packs), with highest weights assigned to Utah (0.22), Colorado (0.19), and Montana (0.17). Post-intervention (1989-2000, 12 years), California's cigarette sales diverged substantially from the synthetic control, declining by an average of 19.4 packs per capita per year (cumulative reduction: 233 packs). To assess statistical significance, we conducted in-space placebo tests applying synthetic control to each of the 38 donor states. California's effect was more extreme than 36 of 38 placebos (p = 0.026, two-sided permutation test). Leave-one-out analysis confirmed robustness to individual donors (effect range: 17.8-21.2 packs). These findings provide strong evidence that Proposition 99 causally reduced cigarette consumption in California.
Reusable template

We used the synthetic control method to estimate the causal effect of treatment/intervention on outcome in treated unit. The donor pool consisted of N control units. We constructed a synthetic control by optimizing weights to minimize pre-treatment mean squared prediction error over T_pre pre-treatment periods (year_start - year_end). Report key predictors used for matching. The synthetic control achieved good/poor pre-treatment fit (RMSPE = value). Report which donors received largest weights, or note if weights were dispersed. Post-intervention (year_start - year_end, T_post = N periods), treated unit diverged from the synthetic control by an average of X units (interpret magnitude and direction). Cumulative effect: sum of gaps = Y. To assess statistical significance, we conducted in-space/in-time placebo tests describe. Report p-value from permutation distribution. Robustness checks: leave-one-out, alternative specifications. These findings support/do not support a causal interpretation of the treatment effect.

Essential statistics to report
  • Number of donor units and which received largest weights
  • Pre-treatment fit: RMSPE, visual plot, years covered
  • Average post-treatment gap (treatment effect) with direction
  • Placebo test results (p-value from permutation inference)
  • Predictor balance table (treated vs synthetic control)
  • Post-treatment trajectory plot (gap over time)
  • Robustness checks (LOO, in-time placebos, sensitivity)
10Exhibit Builder

Manuscript Lab

Copy standard summary tables and forensic reporting grids to outline analysis details.

Table 1: Predictor Balance: Treated Unit vs. Synthetic Control
PredictorTreated UnitSynthetic ControlSample AverageBalance Status
Prior Growth Rate4.2%4.1%3.5%BALANCED
Population Density154.2155.0112.4BALANCED
Baseline Spend$12,450$12,400$10,800BALANCED
Note. Weights optimized to minimize RMSPE. Donor pool k = 15 units.
Weights (45.2%, etc.)Identifies the Donor DNA. The treated unit is most similar to State A and B, making them the primary contributors to the 'digital twin' counterfactual.
Header glossary

The 'Digital Twin'. A weighted combination of other units that mimics the treated unit's behavior perfectly before the intervention.

The Naive Baseline. Proves why Synthetic Control is better: the average state is NOT a good match for the treated state.

11Algorithmic Logic

Command Center

Syntax libraries and function parameters for executing calculations in stats packages.

Code is the modern laboratory. Clean execution ensures reproducible discovery.
Execution Engine
# 1. Construct Synthetic Twin
dat <- Synth::dataprep(df, predictors = c('growth', 'pop'), 
                       dependent = 'y', unit.variable = 'id')
synth_out <- Synth::synth(dat)

# 2. Visualize Gap (Path Plot)
Synth::path.plot(synth_out, dat)
Library stack
R
SynthSCtoolsggplot2
Python
SyntheticControlMethods
Elite Forensic Strike

Traditional p-values don't exist here. You MUST use 'Placebo Tests' (Permutation) to see if the effect in your treated state is larger than what you'd find by picking a random donor state.

# Execute Placebo Audit (In-space Permutation)
SCtools::generate.placebos(dat, synth_out)
12The Over-adjustment Trap

Common Mistakes

Analytical caveats and corrections to maintain modeling integrity.

Wisdom is learning from the failures of others. Anticipate the error before it occurs.
Defensive Logic
Why it's wrong
Poor pre-treatment fit indicates synthetic control cannot replicate treated unit's trajectory. If synthetic can't match treated unit pre-treatment, it's not a credible counterfactual post-treatment. Inference is invalid with large RMSPE or visual divergence in pre-period. Poor fit often means donor pool doesn't span treated unit's covariate space (extrapolation, not interpolation).
The correction
Always plot treated vs synthetic in pre-period. Calculate RMSPE and compare to outcome variance. If poor fit (RMSPE > 10-20% of outcome SD, visual divergence): (1) expand donor pool; (2) add more/better predictors; (3) adjust time windows; (4) try augmented SC; (5) acknowledge SCM may not be appropriate for this setting. Never proceed with inference if pre-fit is poor.
Why it's wrong
Without placebo tests, no way to assess statistical significance. SCM produces point estimates but classical inference (standard errors, p-values) not available due to single treated unit. Placebo tests provide permutation-based inference: how unusual is the observed effect compared to effects we'd see if we applied SCM to units that weren't treated? Without placebos, can't distinguish signal from noise.
The correction
Always conduct placebo tests. In-space placebos: apply SCM to each donor unit (treat as 'placebo treated'), calculate placebo gaps, compare true effect to placebo distribution. P-value = proportion of placebos with effects as extreme as true effect. In-time placebos: move treatment date backward (pre-treatment placebo), test for false effects. Report p-values and placebo distribution plots. If p > .10, effect not distinguishable from chance.
Why it's wrong
Weights should be chosen to match pre-treatment characteristics ONLY. Including post-treatment data in weight optimization contaminates inference: you're using the outcome you're trying to predict to choose the prediction. This is a form of overfitting/p-hacking. Synthetic control should be constructed blindly to post-treatment outcomes to avoid bias.
The correction
Restrict weight optimization (time.optimize.ssr in Synth package) to pre-treatment period only. Never include post-treatment outcomes in predictor matching. Specify all analysis decisions (predictors, time windows, donor pool) before examining post-treatment gaps. Pre-register analysis if possible. Document any specification changes as robustness checks.
Why it's wrong
Small donor pools limit ability to match treated unit (sparse covariate space). With few donors, likely to get poor pre-fit (extrapolation rather than interpolation). Convexity constraint (weights ≥ 0, sum to 1) may force weights on poorly-matched donors. Permutation inference has low power: p-value bounded below by 1/(N+1), so with 5 donors, minimum p = .17 even with extreme effects.
The correction
Aim for at least 15-20 donors. If small donor pool: (1) expand eligibility criteria; (2) use penalized SC (allows negative weights, relaxes convexity); (3) consider alternative methods (DiD, matching, ITS); (4) acknowledge limitation - inference is weak with few donors. Report exact p-value and donor count so readers can assess inferential strength.
Why it's wrong
If treatment in one unit affects control units (spillovers), synthetic control is contaminated. Example: California tobacco tax may affect Nevada if border residents buy cigarettes there. Spillovers bias effect estimates: if positive spillovers to controls, we underestimate treatment effect; if negative spillovers, we overestimate. SUTVA violation invalidates causal interpretation.
The correction
Assess spillover potential via theory and context. If likely: (1) exclude geographically/economically proximate units from donor pool; (2) model spillovers explicitly by distance/network; (3) test whether proximity predicts post-treatment gaps in controls; (4) acknowledge limitation and interpret as 'direct effect' not total effect. If severe spillovers suspected, SCM may not be appropriate.
Why it's wrong
Single specification may be fragile - results could be driven by one influential donor, choice of predictors, or time window. Without robustness checks, don't know if findings are stable or artifact of arbitrary decisions. Leave-one-out test: if excluding one donor changes results dramatically, inference is fragile (concentrated weights problem).
The correction
Conduct robustness checks: (1) Leave-one-out - exclude each top donor sequentially, re-estimate, check effect stability; (2) Alternative predictors - add/remove predictors, check sensitivity; (3) Alternative time windows - vary pre-period cutoffs; (4) Alternative donor pools - restrict/expand eligibility. Report range of estimates across specifications. If results stable (CV < 0.3), claim robustness; if sensitive, acknowledge and interpret cautiously.
Why it's wrong
Permutation p-values from placebo tests are approximate and have discrete support (p ≥ 1/(N+1)). They measure 'rarity' of effect relative to placebos, not asymptotic Type I error rate. Interpretation differs from classical tests. With N=20 donors, minimum p = .048 - might not reach .05 even with strong effect. P-values depend on pre-fit: placebos with poor pre-fit inflate false positive rate.
The correction
Report exact permutation p-values with denominator (e.g., '2/39 placebos exceeded true effect, p = .051'). Filter placebos with poor pre-fit (RMSPE > 2x treated unit) before calculating p-value. Don't obsess over .05 threshold - report effect size, trajectory, and position in placebo distribution. Use p-value as evidence weight, not binary decision rule. Acknowledge power limitations with small donor pools.
Why it's wrong
SCM is designed for causal inference with interventions, not for pure prediction/forecasting. It assumes stable relationships and no confounding. Using SCM to predict future outcomes without intervention makes strong extrapolation assumptions. Better forecasting methods exist (ARIMA, BSTS, machine learning). SCM's value is identifying counterfactual under treatment, not predicting future absent treatment.
The correction
Use SCM only for causal questions with clear intervention and comparison. For forecasting, use time series methods (ARIMA, exponential smoothing, BSTS, Prophet). If blending: use Bayesian structural time series (CausalImpact package) which combines SCM logic with probabilistic forecasting and credible intervals.
13Academic Lineage

References

Scholarly lineage and citation keys grounding the statistical framework.

We stand on the shoulders of giants. Honor the source of the method.
Academic Lineage
[1]
Abadie, A., Diamond, A., & Hainmueller, J. (2010). Synthetic control methods for comparative case studies: Estimating the effect of California's tobacco control program. Journal of the American Statistical Association, 105(490), 493-505.
Seminal paper introducing synthetic control method with California Prop 99 application. Foundational reference.
doi: 10.1198/jasa.2009.ap08746
[2]
Abadie, A. (2021). Using synthetic controls: Feasibility, data requirements, and methodological aspects. Journal of Economic Literature, 59(2), 391-425.
Comprehensive review of SCM: assumptions, diagnostics, inference, extensions, and practical guidance.
doi: 10.1257/jel.20191450
[3]
Abadie, A., Diamond, A., & Hainmueller, J. (2015). Comparative politics and the synthetic control method. American Journal of Political Science, 59(2), 495-510.
Discusses inference via placebo tests, robustness checks, and best practices for comparative case studies.
doi: 10.1111/ajps.12116
[4]
Athey, S., & Imbens, G. W. (2017). The state of applied econometrics: Causality and policy evaluation. Journal of Economic Perspectives, 31(2), 3-32.
Contexualizes SCM within modern causal inference toolkit; discusses strengths and limitations.
doi: 10.1257/jep.31.2.3
[5]
Ben-Michael, E., Feller, A., & Rothstein, J. (2021). The augmented synthetic control method. Journal of the American Statistical Association, 116(536), 1789-1803.
Augmented SCM combining synthetic controls with outcome regression for improved robustness and extrapolation.
doi: 10.1080/01621459.2021.1929245
[6]
Xu, Y. (2017). Generalized synthetic control method: Causal inference with interactive fixed effects models. Political Analysis, 25(1), 57-76.
Extends SCM to multiple treated units via interactive fixed effects; relaxes convexity constraint.
doi: 10.1017/pan.2016.2
A single case is a story; a synthetic control is a trial. Use the shadow to prove that the change you saw was a reaction to the strike, not just the rhythm of the world.
The Interpretive Rigor Directive
statminds · SyntheticMind reference · v2.2 · updated 2026-01-1715 of 15 sections