Atlas
statminds
Hypothesis Framework (TOST Model)The underlying model family class (e.g. GLM, linear model, categorical matrix, log-linear).Parametric ReferenceStatistical methods that assume a specific probability distribution family (typically normal).12-stage workflow

Equivalence & Non-Inferiority

The engine for Clinical Parity. These models audit if a new treatment is 'Just as Good' as the gold standard, reveal if differences fall within a pre-defined safety margin (Δ) rather than just being 'Not Significant'.

Model familyHypothesis Framework (TOST Model)
HypothesisEquivalence: two one-sided tests (TOST). Non-inferiority: one-sided test
AliasesEquivalence Trial · Non-Inferiority (NI) Test · TOST (Two One-Sided Tests)
G1
Clinical Parity Audit
Prove that a new intervention performs within a strict 'Acceptability Window' compared to the established norm.
G2
Non-Inferiority Discovery
Verify that a novel treatment is 'Not Meaningfully Worse' than the standard, justifying its use for lower cost or side effects.
G3
Margin-Based Precision
Shift the burden of proof from 'No Difference' to 'Verified Similarity' using Two One-Sided Tests (TOST).
Visual Overview Dashboard
1

What is it?

Equivalence & Non-Inferiority is a specialized statistical test used to evaluate proportions, multivariate mean vectors, or clinical equivalence margins.

The engine for Clinical Parity. These models audit if a new treatment is 'Just as Good' as the gold standard, reveal if differences fall within a pre-defined safety margin (Δ) rather than just being 'Not Significant'.

2

Goals & Indications

  • Clinical Parity Audit: Prove that a new intervention performs within a strict 'Acceptability Window' compared to the established norm.
  • Non-Inferiority Discovery: Verify that a novel treatment is 'Not Meaningfully Worse' than the standard, justifying its use for lower cost or side effects.
  • Margin-Based Precision: Shift the burden of proof from 'No Difference' to 'Verified Similarity' using Two One-Sided Tests (TOST).
3

Core Idea Diagram

EstimateEquivalence Zone
4

Claims tested

H₀: H₀ (Equivalence): |μ₁ - μ₂| ≥ Δ (difference exceeds margin). H₀ (Non-inferiority): μ₁ - μ₂ ≤ -Δ (test treatment inferior by margin Δ)
Hₐ: Hₐ (Equivalence): |μ₁ - μ₂| < Δ (difference within margin). Hₐ (Non-inferiority): μ₁ - μ₂ > -Δ (test treatment non-inferior)
5

How it works

  1. Pre-specify equivalence margin Delta based on clinical significance.
  2. Establish hypotheses: H₀ states difference exceeds Delta.
  3. Calculate two one-sided t-tests (TOST) at Delta bounds.
  4. If both TOST reject, conclude treatment equivalence.
6

Assumptions

Equivalence/non-inferiority margin: Margin reflects meaningful clinical difference
Adequate power for demonstrating equivalence/non-inferiority: Sample size powered for equivalence claim
ITT: All randomized participants analyzed as assigned
7

Important Note

REVERSAL OF TYPICAL HYPOTHESIS TESTING: Here we want to REJECT H₀ to claim equivalence/non-inferiority. Equivalence tests BOTH upper and lower bounds (-Δ, +Δ). Non-inferiority tests ONLY lower bound (-Δ). Margin Δ must be pre-specified based on clinical relevance, not data-driven. Use 90% CI (not 95%) to test at α = .05 level (two one-sided tests at α/2 = .025 each).

8

Worked Example

MetricEstimatep-value
Test Statistic3.120.015

TOST Equivalence & Non-Inferiority Laboratory

TOST (Two One-Sided Tests) verifies whether the difference in group means lies completely within pre-specified equivalence bounds ($\pm\Delta$).

The 12-Stage Precision Workflow
01The Margin (Δ)
Hypotheses
We test the null of 'Inequivalence' (Difference > Δ) against the discovery of a non-zero shift that stays within the safety walls.
02Clinical Relevance
Assumptions
The ultimate requirement: the 'Margin of Indifference' (Δ) must be defined by clinical consensus before the data is observed.
03The Alpha Shift
Diagnostics
Ensuring the test uses appropriate significance levels—Equivalence typically uses two 5% strikes, resulting in a 90% Confidence Interval.
04focus
Testing if FlowMotion at home is equivalent to FlowMotion in-clinic for overall pain reduction.
05Superiority Pivot
Alternatives
Knowing when to switch to standard T-tests if the goal is to prove 'Better' rather than just 'Equivalent'.
06The TOST Strike
Significance
Executing two separate one-sided tests—if BOTH are significant, the intervention is mathematically verified as equivalent.
07The Gap Confidence
Effect Size
Interpreting the 90% Confidence Interval: if the entire interval sits between -Δ and +Δ, equivalence is established.
08High-Power Mandate
Sample Size
Accounting for the massive N required: proving 'Same' often demands 2-4x more participants than proving 'Different'.
09The Margin Statement
Reporting
Reporting the CI alongside the clinical margin: 'The treatment was equivalent within the 2-point margin, 90% CI [-0.5, 0.8].'
10TOST / equivalence Logic
Software
Executing 'TOST()' commands, ensuring the 'low' and 'high' margins are explicitly specified in the algorithm.
11focus
The fatal error of claiming 'Equivalence' because a standard t-test was non-significant—which only proves a lack of evidence, not parity.
12focus
Tracing the model back to Schuirmann (1987) and the foundational shift toward regulatory bioequivalence forensics.
01Hypothesis test logic

Hypotheses

Pragmatic null and alternative hypotheses defined in mathematical notation.

A hypothesis is a question sharpened to a point. Ambiguity is the enemy of inference.
Logic Core
Null · H₀

H₀ (Equivalence): |μ₁ - μ₂| ≥ Δ (difference exceeds margin). H₀ (Non-inferiority): μ₁ - μ₂ ≤ -Δ (test treatment inferior by margin Δ)

Alternative · Hₐ

Hₐ (Equivalence): |μ₁ - μ₂| < Δ (difference within margin). Hₐ (Non-inferiority): μ₁ - μ₂ > -Δ (test treatment non-inferior)

Why it matters Equivalence: two one-sided tests (TOST). Non-inferiority: one-sided test

REVERSAL OF TYPICAL HYPOTHESIS TESTING: Here we want to REJECT H₀ to claim equivalence/non-inferiority. Equivalence tests BOTH upper and lower bounds (-Δ, +Δ). Non-inferiority tests ONLY lower bound (-Δ). Margin Δ must be pre-specified based on clinical relevance, not data-driven. Use 90% CI (not 95%) to test at α = .05 level (two one-sided tests at α/2 = .025 each).

02Model diagnostics

Assumptions

The core mathematical criteria needed to ensure that statistical testing remains unbiased and valid.

Build your analysis on rock, not sand. Verify the mathematical foundation before building the model.
Integrity Shield
6
Assumptions
5
Critical / High Severity
How to check
Quick
Verify margin was specified in protocol BEFORE data collection. Check if margin is based on: (1) clinical expert consensus on minimal clinically important difference (MCID); (2) percentage of active control effect preserved (e.g., retain ≥50% of control effect); (3) regulatory guidance (FDA recommends preserving ≥50% of control effect for NI trials)
Rigorous
Review protocol registration (ClinicalTrials.gov) for pre-specification. Compare margin to historical effect sizes of active control vs placebo (meta-analysis). Ensure margin smaller than smallest effect that would be clinically meaningful. For NI, verify margin preserves substantial fraction (typically 50%) of control's efficacy vs placebo. FDA guidance: Δ should be smaller than lower bound of 95% CI for historical control effect
If violated
If margin not pre-specified → cannot conduct valid equivalence/NI test; post-hoc margin selection inflates Type I error. If margin too large (not clinically justified) → equivalence/NI claim is meaningless even if statistically significant. If margin not based on historical data → conduct meta-analysis of control vs placebo to establish appropriate margin. NEVER select margin based on observed data or to achieve desired result. For NI trials, calculate 'true M1' margin using FDA formula: M1 = fraction × (historical control effect)
How to check
Quick
Verify sample size calculation in protocol specified equivalence/NI testing (not superiority). Check assumed effect size: for equivalence, typically assume difference = 0; for NI, assume difference = 0 or small positive. Confirm power ≥ 80% to detect equivalence/NI within margin Δ at α = .05 (using 90% CI approach). Typical formula: n ≈ 2(Z₁₋α + Z₁₋β)²σ²/Δ²
Rigorous
Recalculate power using observed standard deviation and actual sample size. Equivalence trials require LARGER samples than superiority trials (paradoxically, because proving 'no difference' is harder than proving 'a difference'). Check if power calculation accounted for dropout; effective sample size may be lower than planned. For NI, verify power calculation assumed small or zero true difference, not large superiority
If violated
If underpowered → 90% CI will be wide, potentially crossing margin even if true difference is near zero; study is inconclusive, not proof of equivalence/NI. If powered for superiority instead of equivalence → sample size likely too small; recalculate required n for equivalence. Report achieved power based on observed data. Acknowledge underpowering as limitation. For pivotal trials, FDA may reject underpowered equivalence/NI claims. Consider meta-analysis with similar trials to increase precision
How to check
Quick
Verify all randomized participants included in primary analysis regardless of: (1) protocol adherence; (2) treatment discontinuation; (3) missing data (imputed appropriately). Check if any participants excluded post-randomization; exclusions must be pre-specified and minimal. ITT is CONSERVATIVE for equivalence/NI (biases toward null = no difference between treatments, making it HARDER to show equivalence/NI)
Rigorous
Compare randomized n vs analyzed n by arm. Investigate reasons for exclusions (e.g., protocol violations, loss to follow-up). For NI trials, FDA requires ITT as primary; per-protocol (PP) as sensitivity analysis. ITT protects against bias from non-adherence but may obscure true treatment differences. Check imputation method for missing data: last observation carried forward (LOCF), multiple imputation, or mixed models with REML
If violated
If per-protocol (PP) analysis used instead of ITT → bias toward finding equivalence/NI (both treatments look more similar when non-adherers excluded). FDA typically requires BOTH ITT (primary) and PP (sensitivity) to agree for NI claim. If substantial missing data (>20%) → conduct multiple imputation or use mixed models to handle missingness. If ITT and PP results conflict → study is inconclusive; investigate reasons for discordance. Report both analyses with full transparency about excluded participants
How to check
Quick
For NI trials comparing Test vs Active Control, assay sensitivity means: if a placebo arm were included (hypothetically), the active control would demonstrate superiority over placebo in this trial's setting. Evidence: (1) Historical data showing control beats placebo; (2) Similar trial design, population, and endpoint to historical trials; (3) Adequate control group response rate. Without assay sensitivity, failing to show difference between Test and Control could mean BOTH are ineffective, not that Test is non-inferior
Rigorous
Systematic review of historical active control vs placebo trials: extract effect sizes, CIs, heterogeneity (I²). Compare current trial's control group outcomes to historical control group outcomes (similar if assay sensitivity present). Check for threats to assay sensitivity: (1) overly healthy population (floor effect); (2) overly sick population (ceiling effect); (3) concomitant medications reducing effect size; (4) different endpoints than historical trials. FDA guidance emphasizes assay sensitivity as essential for valid NI inference
If violated
If assay sensitivity questionable → NI claim is invalid even if statistical test shows non-inferiority. Cannot distinguish 'Test = Control because both work' from 'Test = Control because neither works'. Fixes: (1) Include active placebo arm (three-arm trial: Test vs Control vs Placebo) to directly demonstrate assay sensitivity; (2) Conduct rigorous systematic review/meta-analysis of historical Control vs Placebo; (3) Verify trial conditions match historical trials (population, endpoints, concomitant care); (4) Monitor control group outcomes closely and compare to historical benchmarks. FDA may reject NI claim without demonstrated assay sensitivity
How to check
Quick
For continuous outcomes with t-test TOST: check normality (Q-Q plots, Shapiro-Wilk) and homogeneity of variance (Levene's test). For binary outcomes with risk difference TOST: check adequate sample size (n ≥ 30 per group). For survival outcomes (hazard ratio NI): check proportional hazards assumption (Schoenfeld residuals, log-log plots). Violations can distort equivalence/NI conclusions
Rigorous
Diagnostic plots and formal tests for each model type: (1) ANOVA: normality per group, Levene's test, outliers; (2) Logistic/binomial: sufficient events (≥10 per group), no complete separation; (3) Cox: proportional hazards (test time × treatment interaction), log-linearity for continuous covariates. For violated assumptions, use robust alternatives or transformations. Ensure primary analysis method was pre-specified in protocol
If violated
If normality violated with continuous outcome → use non-parametric equivalence test (e.g., Mann-Whitney U-based TOST, or bootstrap TOST). If proportional hazards violated → use restricted mean survival time (RMST) difference instead of HR for NI test. If heteroscedasticity → use Welch's t-test variant of TOST. If model assumptions cannot be met → consider different outcome scale or statistical model. Pre-specify sensitivity analyses using different models in protocol
How to check
Quick
For NI trials: review historical Control vs Placebo trials chronologically. Check if control's effect size has declined over time (biocreep = gradual erosion of treatment efficacy in successive NI trials as increasingly inferior treatments are approved and become new controls). Calculate weighted average effect size and temporal trend. Biocreep occurs when: Trial 1 (Control vs Placebo, effect d); Trial 2 (Test1 vs Control, NI by small margin); Trial 3 (Test2 vs Test1, NI by small margin) → cumulative margin erosion
Rigorous
Meta-regression of historical Control vs Placebo effect sizes vs publication year. Test for temporal trend (p < .05 indicates biocreep risk). Compare margin Δ to lower bound of 95% CI for historical control effect (FDA guidance: margin should preserve ≥50% of control's effect). If historical effect is shrinking, current margin may be too large, allowing truly inferior treatment to pass NI test. For pivotal NI trials, FDA reviews historical data comprehensively to assess biocreep risk
If violated
If biocreep suspected (declining historical control effects) → use more conservative (smaller) margin based on recent trials only, not all historical data. Alternatively, use the LOWER bound of 95% CI for historical control effect to set margin (conservative approach). For new drug classes or indications, prioritize recent trials (within 5-10 years) that reflect current standards of care. FDA may require three-arm trial (Test vs Control vs Placebo) to avoid biocreep issues entirely. Acknowledge biocreep as limitation if using historical data spanning decades
03Residual Forensics

Diagnostics

Checking residual plots and indices to examine model deviations and ensure standard error integrity.

Trust, but verify. The outliers often hold more truth than the averages.
System Health
Essential checks
  1. Verify margin Δ pre-specified in protocol (not data-driven)
  2. Check 90% CI for mean difference (or relevant parameter)
  3. Compare 90% CI bounds to equivalence margins (-Δ, +Δ) or NI margin (-Δ)
  4. Verify ITT analysis conducted (all randomized participants)
Recommended checks
  1. Per-protocol sensitivity analysis (concordance with ITT)
  2. Systematic review/meta-analysis of historical control vs placebo
  3. Forest plot showing 90% CI relative to margin(s)
  4. Power calculation verification (was study adequately powered?)
  5. Assay sensitivity assessment (control group outcomes vs historical benchmarks)
  6. Model assumption diagnostics (normality, proportional hazards, etc.)
04Live Instances

Applied Minds

Review concrete study examples, data layout guidelines, and copy executable syntax scripts.

Theory is the map. Practice is the terrain. Simulation bridges the gap.
Applied Wisdom
Example 01

Generic Drug Bioequivalence (Equivalence Test with TOST)

Research question: Is a generic formulation bioequivalent to the brand-name drug? Design: Crossover bioequivalence study with 24 healthy volunteers. Each participant receives both Generic and Brand formulations in random order with washout period. Outcome: AUC (area under curve) for drug concentration. Equivalence margin: ±20% (80-125% ratio, FDA standard for bioequivalence). Use 90% CI for AUC ratio; if entirely within [0.80, 1.25], conclude bioequivalence.

DesignCrossover equivalence trial (within-subjects)
Total n48
Outcome ScaleAUC (area under curve, log-transformed)
# TOST Equivalence Test: Generic vs Brand-name Bioequivalence
# FDA bioequivalence testing using Two One-Sided Tests (TOST)

library(equivalence)  # For TOST
library(TOSTER)       # Alternative TOST package
library(ggplot2)

# Set seed
set.seed(2025)

# === Simulate crossover bioequivalence data ===
# True ratio ≈ 1.05 (5% higher generic, but within margin)
# Log-normal distribution for AUC

n <- 24  # Number of subjects in crossover design

# Each subject gets both treatments
# True geometric mean ratio = 1.05 (Generic 5% higher than Brand)
log_mean_brand <- 4.5    # log(AUC) for brand
log_mean_generic <- log(exp(log_mean_brand) * 1.05)  # 5% higher
sd_within <- 0.25        # Within-subject SD on log scale

# Subject-specific random effects
subject_effect <- rnorm(n, mean=0, sd=0.3)

# Generate paired data (crossover)
data <- data.frame(
  subject = rep(1:n, each=2),
  treatment = rep(c("Brand", "Generic"), times=n),
  log_AUC = c(
    log_mean_brand + subject_effect + rnorm(n, 0, sd_within),      # Brand
    log_mean_generic + subject_effect + rnorm(n, 0, sd_within)     # Generic
  )
)

# Convert to wide format for paired test
library(tidyr)
data_wide <- data %>%
  pivot_wider(id_cols=subject, names_from=treatment, values_from=log_AUC)

print("=== Sample Data(first 6 subjects) ===")
print(head(data_wide))

# === STEP 1: Check Assumptions ===
cat("\n=== Assumption Checks ===\n")

# 1. Normality of differences (log scale)
differences <- data_wide$Generic - data_wide$Brand
shapiro_result <- shapiro.test(differences)
cat("Shapiro-Wilk test for normality: W =", round(shapiro_result$statistic, 3), 
    ", p =", round(shapiro_result$p.value, 3), 
    ifelse(shapiro_result$p.value > 0.05, "✓", "✗"), "\n")

# Q-Q plot
qqnorm(differences, main="Q-Q Plot: Log(AUC) Differences")
qqline(differences)

cat("\nCrossover design: Each subject as own control ✓\n")
cat("Adequate washout period assumed ✓\n")

# === STEP 2: Define Equivalence Margin ===
# FDA bioequivalence: 80-125% ratio on original scale
# On log scale: log(0.80) to log(1.25)
lower_margin <- log(0.80)  # -0.223
upper_margin <- log(1.25)  # +0.223

cat("\n=== Equivalence Margins(FDA Standard) ===\n")
cat("Original scale: 80% to 125% ratio\n")
cat("Log scale: ", round(lower_margin, 3), " to ", round(upper_margin, 3), "\n")

# === STEP 3: Calculate Mean Difference and 90% CI ===
cat("\n=== Mean Difference(Generic - Brand) ===\n")

mean_diff <- mean(differences)
sd_diff <- sd(differences)
se_diff <- sd_diff / sqrt(n)

# 90% CI (for α = .05 TOST)
t_crit <- qt(0.95, df=n-1)  # One-sided 5%, two-sided 10%
ci_90_lower <- mean_diff - t_crit * se_diff
ci_90_upper <- mean_diff + t_crit * se_diff

cat("Mean log(AUC) difference:", round(mean_diff, 4), "\n")
cat("90% CI:", round(ci_90_lower, 4), "to", round(ci_90_upper, 4), "\n")

# Back-transform to ratio scale
ratio <- exp(mean_diff)
ratio_ci_lower <- exp(ci_90_lower)
ratio_ci_upper <- exp(ci_90_upper)

cat("\nGeometric mean ratio(Generic/Brand):", round(ratio, 3), "\n")
cat("90% CI for ratio:", round(ratio_ci_lower * 100, 1), "% to", 
    round(ratio_ci_upper * 100, 1), "%\n")

# === STEP 4: TOST (Two One-Sided Tests) ===
cat("\n=== TOST Equivalence Test ===\n")

# Test 1: H₀: μ_diff ≤ lower_margin vs Hₐ: μ_diff > lower_margin
t1 <- (mean_diff - lower_margin) / se_diff
p1 <- pt(t1, df=n-1, lower.tail=FALSE)

cat("Test 1 (lower bound): t =", round(t1, 3), ", p =", round(p1, 4), "\n")

# Test 2: H₀: μ_diff ≥ upper_margin vs Hₐ: μ_diff < upper_margin
t2 <- (mean_diff - upper_margin) / se_diff
p2 <- pt(t2, df=n-1, lower.tail=TRUE)

cat("Test 2 (upper bound): t =", round(t2, 3), ", p =", round(p2, 4), "\n")

# TOST p-value = max(p1, p2)
p_tost <- max(p1, p2)
cat("\nTOST p-value(max of two tests):", round(p_tost, 4), "\n")

# === STEP 5: Equivalence Decision ===
cat("\n=== Equivalence Decision ===\n")

if (ci_90_lower > lower_margin & ci_90_upper < upper_margin) {
  cat("CONCLUSION: BIOEQUIVALENT ✓\n")
  cat("90% CI [", round(ci_90_lower, 4), ",", round(ci_90_upper, 4), 
      "] entirely within margins [", round(lower_margin, 3), ",", 
      round(upper_margin, 3), "]\n")
  cat("On ratio scale: 90% CI [", round(ratio_ci_lower * 100, 1), "%, ", 
      round(ratio_ci_upper * 100, 1), "%] within [80%, 125%]\n")
  cat("TOST p-value =", round(p_tost, 4), "< .05: Reject H₀ of non-equivalence\n")
} else {
  cat("CONCLUSION: NOT BIOEQUIVALENT ✗\n")
  cat("90% CI crosses equivalence margin\n")
  cat("Cannot conclude bioequivalence at α = .05 level\n")
}

# === STEP 6: Visualizations ===

# 6.1 Forest plot with equivalence margins
forest_data <- data.frame(
  Study = "Generic vs Brand",
  Estimate = mean_diff,
  CI_lower = ci_90_lower,
  CI_upper = ci_90_upper
)

ggplot(forest_data, aes(y=Study, x=Estimate)) +
  geom_rect(aes(xmin=lower_margin, xmax=upper_margin, ymin=-Inf, ymax=Inf),
            fill="lightgreen", alpha=0.3) +
  geom_point(size=5, color="darkblue") +
  geom_errorbarh(aes(xmin=CI_lower, xmax=CI_upper), height=0.2, linewidth=1.5) +
  geom_vline(xintercept=0, linetype="solid", color="black", linewidth=1) +
  geom_vline(xintercept=lower_margin, linetype="dashed", color="red", linewidth=1) +
  geom_vline(xintercept=upper_margin, linetype="dashed", color="red", linewidth=1) +
  annotate("text", x=lower_margin, y=1.3, label="Lower Margin", 
           size=3, hjust=1.1) +
  annotate("text", x=upper_margin, y=1.3, label="Upper Margin", 
           size=3, hjust=-0.1) +
  labs(title="Bioequivalence: Generic vs Brand-name Drug",
       subtitle="90% CI for log(AUC) difference with FDA equivalence margins",
       x="Log(Generic) - Log(Brand)\n[Equivalence Zone in Green]",
       y="") +
  theme_classic(base_size=14) +
  theme(plot.title=element_text(hjust=0.5, face="bold"),
        plot.subtitle=element_text(hjust=0.5))

ggsave("bioequivalence_forest.png", width=12, height=6)

# 6.2 Ratio scale visualization
ggplot(data.frame(x=1), aes(x=x)) +
  geom_rect(aes(xmin=0.80, xmax=1.25, ymin=-Inf, ymax=Inf),
            fill="lightgreen", alpha=0.3) +
  geom_point(aes(x=ratio, y=1), size=5, color="darkblue") +
  geom_errorbarh(aes(y=1, xmin=ratio_ci_lower, xmax=ratio_ci_upper), 
                 height=0.3, linewidth=1.5) +
  geom_vline(xintercept=1, linetype="solid", color="black", linewidth=1) +
  geom_vline(xintercept=0.80, linetype="dashed", color="red", linewidth=1) +
  geom_vline(xintercept=1.25, linetype="dashed", color="red", linewidth=1) +
  scale_x_continuous(limits=c(0.7, 1.4), breaks=seq(0.7, 1.4, 0.1)) +
  labs(title="Bioequivalence: Geometric Mean Ratio",
       subtitle="90% CI must be entirely within 80-125% for bioequivalence",
       x="Ratio: Generic AUC / Brand AUC(%)",
       y="") +
  theme_classic(base_size=14) +
  theme(plot.title=element_text(hjust=0.5, face="bold"),
        plot.subtitle=element_text(hjust=0.5),
        axis.text.y=element_blank(),
        axis.ticks.y=element_blank())

ggsave("bioequivalence_ratio.png", width=12, height=6)

# === APA-Style Reporting ===
cat("\n=== APA-Style Report ===\n")
cat(sprintf(
  "A crossover bioequivalence study(n=%d healthy volunteers) compared a generic\nformulation to the brand-name drug using AUC(area under concentration-time curve).\n\nUsing FDA's Two One-Sided Tests(TOST) procedure with equivalence margins of\n80%%-125%% (±20%% on ratio scale), the geometric mean ratio was %.2f (90%% CI [%.2f%%, %.2f%%]).\n\nThe 90%% confidence interval fell entirely within the pre-specified equivalence\nmargins of 80%%-125%% (TOST p = %.4f < .05), supporting bioequivalence of the\ngeneric to the brand-name formulation.\n\nOn the log scale, the mean difference was %.4f (90%% CI [%.4f, %.4f]), within\nthe margins of [%.3f, %.3f].\n\nConclusion: The generic formulation is bioequivalent to the brand-name drug,\nmeeting FDA criteria for therapeutic equivalence.",
  n, ratio, ratio_ci_lower * 100, ratio_ci_upper * 100, p_tost,
  mean_diff, ci_90_lower, ci_90_upper, lower_margin, upper_margin
))
Interpretation Blueprint

Geometric mean ratio = 1.05, 90% CI [0.94, 1.17]. The 90% confidence interval for the Generic/Brand AUC ratio falls entirely within the FDA-specified equivalence margins of 80-125% (TOST p = .01 < .05). This supports bioequivalence, meaning the generic formulation can be considered therapeutically equivalent to the brand-name drug. The point estimate of 1.05 (5% higher generic) is clinically trivial and the CI demonstrates that the true difference is within acceptable regulatory limits. FDA requires BOTH 90% CI bounds to be within [80%, 125%] for bioequivalence approval. This example demonstrates the TOST procedure: testing simultaneously that Generic is not too low (<80%) AND not too high (>125%) compared to Brand.

05Tactical Pivots

Alternatives

Structured fallback pathways for choosing alternative tests when normality or slopes requirements fail.

When the path is blocked, pivot. Rigor is not rigidity; it is the intelligent adaptation to reality.
Adaptive Strategy
Measurement Precision Ladder Ideal · Ratio / Interval
Ratio
Maintain TOST logic. Proving parity in high-fidelity markers requires maximum numerical precision.
Peak Signal
Interval
Ideal for Primary Metrics. Ensure the 'Margin of Indifference' (Δ) is clinically justified.
Standard Precision
Ordinal / Nominal
Abandon TOST. Use Non-Inferiority tests for proportions or log-odds to model categorical parity.
Model Collapse
Temporal Trajectory Audit Static Parity Snapshot
Static Equivalence
Cross-sectional audit.
Stay with TOST. Prove two interventions are 'Just as Good' as each other.
Longitudinal Parity
Stable recovery.
Pivot to Equivalence tests within a Mixed ANOVA or LMM framework to model trajectory similarity.
Adaptive Technical Safeguards · adaptive safeguards
normality violated
  • Non-Parametric TOST — Execute the Two One-Sided Tests using the Mann-Whitney U basis.
  • Bootstrap Equivalence — Generate a 90% CI for the mean difference using 1,000 resamples.
margin too narrow for N
  • Non-Inferiority Pivot — Relax the mandate to prove 'Better or Same' rather than 'Exactly Same'.
  • Informal Parity Audit — Report the Confidence Interval and its overlap with the margin without a formal p-value strike.
06Adjusted Comparisons

Post-hoc

Group mean comparisons and correction controls (e.g. Tukey HSD, Bonferroni) to protect against Family-Wise Error Rates.

The omnibus test opens the door; post-hoc analysis explores the room.
Forensic Detail
Adjusted Comparisons

Post-hoc pairwise tests defined for this model.

Interpretation Guidelines

Equivalence post-hoc is an audit of 'Closeness'. Use margin-sensitivity to prove that your intervention is not just 'Within the Limit', but 'Centrally Balanced' against the gold standard.

07Standardized scale impact

Effect Size

Understanding effect sizes (e.g., Cohen's d, Partial Eta-Squared) and clinical impact benchmarks.

Significance is noise. Magnitude is the signal. Measure the impact, not just the probability.
Impact Magnitude

For equivalence: 90% CI must be ENTIRELY within [-Δ, +Δ]. For NI: 90% CI lower bound must be > -Δ (upper bound can exceed +Δ, indicating superiority).

Larger distance between CI bound and margin = stronger evidence for equivalence/NI. If CI barely inside margin, evidence is weak.

Even if statistically equivalent/non-inferior, assess clinical importance: is observed difference clinically trivial? For example, mean difference of 0.5 points on 100-point scale may be statistically equivalent but still clinically meaningful if margin was set too large.

Recommended Metric: Report point estimate with 90% CI, distance from margin(s), and clinical interpretation. Include forest plot showing CI relative to equivalence zone.
Small
0.2
Medium
0.5
Large
0.8
0.50
Report point estimate with 90% CI, distance from margin(s), and clinical interpretation. Include forest plot showing CI relative to equivalence zone.
Recommended Measure
4
Available Metrics
ReportUse Report point estimate with 90% CI, distance from margin(s), and clinical interpretation. Include forest plot showing CI relative to equivalence zone. to represent clinical impact magnitude.
08Statistical Power

Sample Size

Guidelines for minimum sample requirements and power analysis parameters.

An underpowered study is an ethical failure. Respect the data by collecting enough of it.
Power Protocol
Floor Requirements

The 'High-Power Mandate': Proving 'Similarity' requires significantly more data than proving 'Difference'. A minimum of 50 participants per group is recommended for even moderate clinical margins.

Effect SizeParametersRequired n
Small EffectNarrow Margin (Δ=0.2)n ≈ 450 total
Medium EffectStandard Margin (Δ=0.5)n ≈ 80 total
Large EffectWide Margin (Δ=0.8)n ≈ 35 total
Key considerations

The 'Two-Strike' Penalty: Equivalence uses Two One-Sided Tests (TOST). To maintain an 80% global power, each individual strike must achieve ~90% power, effectively doubling the required sample size compared to superiority tests.

G*Power StrategyBenchmark: T-tests → Equivalence (TOST). Parameters: Margin (Δ), Expected Difference (δ), α = .05, Power = .80. Note: If your expected difference is zero, N is minimized. If you expect a small difference within the margin, N explodes.
09APA narrative blueprint

Reporting

How to compile statistical results into publication prose matching APA and journal style guides.

Data does not speak for itself. It requires a translator. Be clear, be precise, be honest.
Narrative Arc
Worked APA paragraph example
A crossover bioequivalence study (n=24 healthy volunteers) compared a generic formulation to the brand-name drug using AUC (area under the concentration-time curve). The pre-specified equivalence margins were 80-125% on the ratio scale (±20%), consistent with FDA bioequivalence guidance. Using the Two One-Sided Tests (TOST) procedure, the geometric mean ratio (Generic/Brand) was 1.05 (90% CI [0.94, 1.17], TOST p = .01). The 90% confidence interval fell entirely within the pre-specified margins of 80-125%, supporting bioequivalence at the α = .05 significance level. The observed 5% difference in AUC is clinically trivial and well within the acceptable range for therapeutic equivalence. This study demonstrates that the generic formulation meets FDA criteria for bioequivalence and can be considered therapeutically interchangeable with the brand-name drug.
Reusable template

An equivalence/non-inferiority trial (n = total N) compared Test treatment to Control/Standard for outcome. The pre-specified equivalence margin was Δ (or lower Δ, upper Δ for two-sided), justified by clinical rationale / % of control effect preserved / regulatory guidance. State study design: RCT, crossover, parallel-group. Primary analysis: ITT or per-protocol. The mean difference / risk difference / hazard ratio was point estimate (90% CI [lower, upper], TOST p-value or statement about CI relative to margins). For equivalence: The 90% confidence interval fell entirely within the pre-specified margins of [−Δ, +Δ], supporting equivalence (TOST p < .05). For NI: The 90% CI lower bound exceeded the NI margin of −Δ, supporting non-inferiority (p < .05). Clinical interpretation: The observed difference of X is clinically trivial / meaningful and does / does not represent a meaningful change in patient outcomes. Report sensitivity analyses if applicable: per-protocol concordance, assay sensitivity evidence. For NI with superiority claim: Additionally, the point estimate and CI suggest superiority of Test over Control (point estimate > 0 and 95% CI excludes 0).

Essential statistics to report
  • Point estimate (mean difference, RD, RR, HR, etc.) with 90% CI
  • Pre-specified equivalence/NI margin(s) with justification
  • TOST p-value or statement of CI location relative to margins
  • Study design (RCT, crossover, parallel-group)
  • Sample sizes per group (ITT and per-protocol)
  • Primary outcome definition
  • ITT and per-protocol results (if both conducted)
  • Power calculation verification (was study adequately powered?)
  • Assay sensitivity evidence (for NI trials)
  • Clinical interpretation of observed difference magnitude
10Exhibit Builder

Manuscript Lab

Copy standard summary tables and forensic reporting grids to outline analysis details.

Table 1: Non-Inferiority Audit for New vs. Standard Treatment
MetricM_diff (New-Std)95% CI (Upper Bound)Margin (Limit)p (Non-Inf)
Efficacy Score-1.2-3.8-5.0.004
Note. Non-Inferiority Margin (delta) = -5.0. Analysis via TOST (Two One-Sided Tests).
p = .004Powerful Non-Inferiority Proof. Since the entire 95% CI (-3.8) is above the margin (-5.0), we conclude the new treatment is non-inferior to standard care.
Header glossary

The Performance Gap. A negative value means the new treatment performed slightly worse than the standard.

The 'Worst Case' Scenario. We are 95% confident that the new treatment is, at worst, only 3.8 points lower than the standard.

The Clinical Tolerance. The pre-defined threshold of what we consider 'close enough' to the standard.

11Algorithmic Logic

Command Center

Syntax libraries and function parameters for executing calculations in stats packages.

Code is the modern laboratory. Clean execution ensures reproducible discovery.
Execution Engine
# 1. Execute TOST (Two One-Sided Tests) for Equivalence
TOSTER::tost(m1 = 75, m2 = 76, sd1 = 8, sd2 = 8, n1 = 100, n2 = 100, 
             low_eqbound = -5, high_eqbound = 5)

# 2. Non-Inferiority Audit
# (Check if Lower bound of CI > delta)
Library stack
R
TOSTequivalence
Python
statsmodels.stats.weightstats
Elite Forensic Strike

Statistical Significance does NOT equal Equivalence. You can have a p < .05 (Difference) AND still be Equivalent if the difference is smaller than your clinical margin. Always look at the Confidence Interval, not the p-value.

# Visualize Equivalence Bounds vs CI
TOSTER::plot_tost(m1 = 75, m2 = 76, sd1 = 8, sd2 = 8, n1 = 100, n2 = 100, 
                  low_eqbound = -5, high_eqbound = 5)
12The Over-adjustment Trap

Common Mistakes

Analytical caveats and corrections to maintain modeling integrity.

Wisdom is learning from the failures of others. Anticipate the error before it occurs.
Defensive Logic
Why it's wrong
Margin must be pre-specified before seeing data; data-driven margin selection invalidates statistical inference and inflates Type I error. If margin is set post-hoc to be just wide enough to capture observed CI, 'equivalence' claim is circular and meaningless. This is scientific misconduct (HARKing - Hypothesizing After Results are Known). Regulatory agencies (FDA, EMA) require pre-specified margins in trial protocols.
The correction
Margin MUST be specified in study protocol before enrollment begins. Register protocol on ClinicalTrials.gov with margin specified. Base margin on: (1) clinical expert consensus on MCID; (2) historical effect sizes (e.g., for NI, preserve ≥50% of active control's effect vs placebo); (3) regulatory guidance. Never adjust margin after seeing data. If unsure about margin, conduct pilot study or use more conservative (smaller) margin.
Why it's wrong
TOST at α = .05 level requires 90% CI, not 95% CI. Each of the two one-sided tests is conducted at α/2 = .025, which corresponds to 90% CI (not 95% CI). Using 95% CI effectively tests at α = .10 (more liberal), increasing Type I error. This is analogous to two-tailed test requiring α/2 per tail. For equivalence, we have two boundaries (lower and upper), each tested at .025 level.
The correction
For TOST equivalence test at α = .05: use 90% CI. Check if 90% CI is ENTIRELY within equivalence margins [-Δ, +Δ]. For non-inferiority at α = .05: use 90% CI and check if lower bound > -Δ. If reporting 95% CI for other purposes, clearly state that equivalence decision was based on 90% CI. Software output: verify CI level matches intended α.
Why it's wrong
Statistical equivalence (CI within margin) does NOT guarantee clinical equivalence if margin was set too large. For example, if margin Δ = 20 points on 100-point scale but MCID is only 5 points, observed difference of 15 points would be 'statistically equivalent' but clinically meaningful. Regulators care about BOTH statistical AND clinical equivalence. Statistical significance is necessary but not sufficient.
The correction
Justify margin based on clinical relevance, not just statistical convenience. Margin should be smaller than smallest clinically important difference (SCID or MCID). Consult clinical experts, patient advocacy groups, and prior literature. Report observed effect size and discuss clinical importance independently of statistical equivalence conclusion. Example: 'Although statistically equivalent (CI within margins), the observed 15-point difference may be clinically meaningful for some patients.'
Why it's wrong
Per-protocol (PP) analysis excludes non-adherers and protocol violators, which BIASES TOWARD finding equivalence/NI (both groups look more similar when non-compliers removed). This is opposite to superiority trials where ITT is conservative. For NI trials, ITT is conservative (more likely to show difference, making NI harder to prove) while PP is anti-conservative (easier to show NI). FDA requires ITT as primary for NI trials; PP as sensitivity analysis only.
The correction
For regulatory NI trials: ITT is primary analysis; PP is supportive sensitivity analysis. BOTH ITT and PP should agree (show NI) for robust conclusion. If ITT shows NI but PP does not (or vice versa), study is inconclusive. Investigate reasons for discordance (differential dropout, non-adherence patterns). Pre-specify both analyses in protocol. Report both with full transparency about excluded participants and reasons.
Why it's wrong
Assay sensitivity = ability to distinguish effective from ineffective treatment. If assay sensitivity is lacking, failing to find difference between Test and Control could mean BOTH are ineffective (not that Test is non-inferior to effective Control). For example, if active control doesn't work in this trial's setting (wrong population, wrong dose, concomitant medications masking effect), NI conclusion is invalid. Without placebo arm or historical validation, cannot distinguish 'Test = Control because both work' from 'Test = Control because neither works'.
The correction
For NI trials: (1) BEST: Include placebo arm (three-arm trial: Test vs Control vs Placebo) to directly demonstrate assay sensitivity by showing Control > Placebo; (2) Conduct systematic review/meta-analysis of historical Control vs Placebo trials to establish Control's efficacy; (3) Ensure trial conditions (population, endpoints, concomitant care) match historical trials; (4) Monitor control group outcomes and compare to historical benchmarks. FDA guidance emphasizes assay sensitivity as essential for valid NI claims.
Why it's wrong
Absence of evidence is not evidence of absence. Failing to reject H₀ in superiority test (p > .05) does NOT prove equivalence; it only shows insufficient evidence for difference. This is 'accepting the null' fallacy. Study may be underpowered. Wide CI including both large positive and large negative differences is NOT evidence of equivalence. Example: mean difference = 2, 95% CI [-10, 14], p = .7 → not significant, but CI includes clinically important differences up to 14 points. This is NOT equivalence.
The correction
To claim equivalence, must conduct formal equivalence test (TOST) with pre-specified margin. Check if 90% CI is ENTIRELY within equivalence margins. 'p > .05' in superiority test is IRRELEVANT for equivalence claim. Never say 'no significant difference, therefore equivalent'. Correct phrasing: 'TOST shows equivalence (90% CI entirely within margins, p < .05)' OR 'superiority test shows no significant difference (p > .05), but formal equivalence test was not conducted'.
Why it's wrong
Some incorrectly claim 'equivalence' if 95% CI excludes the margins (e.g., CI = [-3, 5] with margins ±10, claiming 'CI doesn't reach margins, so equivalent'). This is WRONG. Equivalence requires CI to be INSIDE margins, not just excluding them. The correct test is: 90% CI entirely within margins. Using 95% CI to check if it EXCLUDES margins is testing a different hypothesis and uses wrong error rate.
The correction
Correct equivalence test: 90% CI (not 95%) must be ENTIRELY WITHIN margins [-Δ, +Δ]. Both bounds of CI must be inside the margins. If 90% CI = [-3, 5] and margins = [-10, +10], this IS equivalent (✓). If 90% CI = [-3, 12] and margins = [-10, +10], this is NOT equivalent because upper bound (12) exceeds upper margin (10) (✗). Never confuse 'CI excludes margin' with 'CI within margin'.
Why it's wrong
Equivalence trials require LARGER sample sizes than superiority trials (paradoxically). If sample size was calculated for superiority (detecting difference from zero), it will be underpowered for equivalence (proving difference is within margin). Underpowered equivalence study will have wide CI likely to cross margins even if true difference is small. Result: inconclusive study, wasted resources. Example: n=50/group powered for superiority with d=0.5; but for equivalence with Δ=0.3, need n=120/group.
The correction
Calculate sample size specifically for equivalence/NI test with pre-specified margin. Use specialized software (PASS, nQuery, PowerTOST in R). Formula differs from superiority power calculation. Specify: margin Δ, assumed true difference (typically 0 for equivalence), SD, α=.05, power=80-90%. If study already conducted with superiority-based n, report achieved power for equivalence and acknowledge underpowering as major limitation. Avoid claiming equivalence from underpowered study.
13Academic Lineage

References

Scholarly lineage and citation keys grounding the statistical framework.

We stand on the shoulders of giants. Honor the source of the method.
Academic Lineage
[1]
U.S. Food and Drug Administration. (2016). Non-Inferiority Clinical Trials to Establish Effectiveness: Guidance for Industry. FDA Center for Drug Evaluation and Research.
FDA guidance on NI trial design, margin justification, assay sensitivity, ITT vs PP analysis. Essential reference for regulatory NI trials. Recommends preserving ≥50% of control effect.
[2]
European Medicines Agency. (2005). Guideline on the Choice of the Non-Inferiority Margin. CHMP/EWP/2158/99. EMA Committee for Medicinal Products for Human Use.
EMA guidance on selecting and justifying NI margins. Emphasizes clinical relevance and statistical considerations. Requires margin to be smaller than lower bound of 95% CI for historical control effect.
[3]
Schuirmann, D. J. (1987). A comparison of the two one-sided tests procedure and the power approach for assessing the equivalence of average bioavailability. Journal of Pharmacokinetics and Biopharmaceutics, 15(6), 657-680.
Classic paper introducing TOST procedure for bioequivalence testing. Basis for FDA bioequivalence guidance using 90% CI within 80-125% margins.
doi: 10.1007/BF01068419
[4]
Piaggio, G., Elbourne, D. R., Pocock, S. J., Evans, S. J., Altman, D. G., & CONSORT Group. (2012). Reporting of noninferiority and equivalence randomized trials: Extension of the CONSORT 2010 statement. JAMA, 308(24), 2594-2604.
CONSORT extension for reporting NI and equivalence trials. Emphasizes pre-specification of margins, ITT/PP analyses, and assay sensitivity. Essential for trial reporting standards.
doi: 10.1001/jama.2012.87802
[5]
Rothmann, M., Li, N., Chen, G., Chi, G. Y., Temple, R., & Tsou, H. H. (2012). Design and analysis of non-inferiority mortality trials in oncology. Statistics in Medicine, 31(2), 239-250.
Discusses NI trial challenges in oncology where mortality is outcome. Addresses margin selection when control effect varies by subgroup and issue of biocreep in successive NI trials.
doi: 10.1002/sim.4250
[6]
Walker, E., & Nowacki, A. S. (2011). Understanding equivalence and noninferiority testing. Journal of General Internal Medicine, 26(2), 192-196.
Accessible tutorial on equivalence and NI testing for clinicians. Explains TOST, margin selection, and interpretation. Good introduction for applied researchers.
doi: 10.1007/s11606-010-1513-8
To prove two things are the same is the hardest task in science. A non-significant p-value is a surrender; a significant equivalence test is a victory.
The Interpretive Rigor Directive
statminds · EquivalenceMind reference · v2.2 · updated 2026-01-1715 of 15 sections