Atlas
statminds
Frequentist InferenceThe underlying model family class (e.g. GLM, linear model, categorical matrix, log-linear).Parametric ReferenceStatistical methods that assume a specific probability distribution family (typically normal).12-stage workflow

Hypothesis Foundations

The logic of discovery. Master the forensic framework for judging evidence against the Null hypothesis.

Model familyFrequentist Inference
HypothesisMean Equality
AliasesSignificance Testing · Decision Theory · Null Hypothesis Significance Testing (NHST) · Inference Logic
G1
Forensic Benchmarking
Establish the 'Null' state as the skeptical baseline for all discovery.
G2
Probability Auditing
Use P-values to measure the strength of evidence against the status quo.
G3
Error Management
Quantify the risk of 'False Positives' (Alpha) and 'False Negatives' (Beta).
G4
Decision Integrity
Define strict mathematical thresholds for rejecting coincidence in favor of truth.
Visual Overview Dashboard
1

What is it?

Hypothesis Foundations represents a core statistical conceptual framework required to understand research design, data mapping, and analytical models.

The logic of discovery. Master the forensic framework for judging evidence against the Null hypothesis.

2

Goals & Indications

  • Forensic Benchmarking: Establish the 'Null' state as the skeptical baseline for all discovery.
  • Probability Auditing: Use P-values to measure the strength of evidence against the status quo.
  • Error Management: Quantify the risk of 'False Positives' (Alpha) and 'False Negatives' (Beta).
  • Decision Integrity: Define strict mathematical thresholds for rejecting coincidence in favor of truth.
3

Core Idea Diagram

Alpha (0.05)Rejection RegionRetain H₀ AreaHypothesis significance boundary check
4

Key Elements

  • Null (H₀): No difference.
  • Alternative (Hₐ): Research claim.
  • Error control: Type I (α) and Type II (β).
5

How it works

  1. State H₀ null status-quo and research alternative Hₐ.
  2. Set Alpha significance limit (Type I error tolerance).
  3. Evaluate observed test statistic from sample data.
  4. Compute p-value to retain or reject the null claim.
6

Defensive Pitfall

Warning: Claiming the null hypothesis is proven true when p > 0.05, rather than failing to reject it.

7

Expert Directive

Hypothesis testing does not declare absolute truth; it manages error rates under repeated trials.

8

Quick Reference

Error TypeNameRate Control
Type I ErrorFalse AlarmAlpha (α = 0.05)
Type II ErrorMissed SignalBeta (β = 0.20)
PowerHit Rate1 - Beta (0.80)
Interactive Sandbox

Hypothesis Power & Error Simulator

Adjust significance levels, effect size, and sample size to watch statistical Power (1 - Beta) expand.

Alpha level (Type I error)0.050
True Effect Size (d)1.20
Sample Size (N)40

Type I Error (Alpha) 5%
Type II Error (Beta) 0.0%
Statistical Power (1 - Beta) 100.0%
H0 curve (grey) vs Ha curve (blue)
Alpha boundary
Inference Command Briefing
01
Alpha (α)
The 'Threshold of Skepticism'—usually 0.05. Your willingness to be wrong if you claim a discovery.
02
P-Value
The 'Surprise Index'. The probability of seeing your data if the Null were actually true.
03
Power (1-β)
The 'Detection Strength'. Your ability to find a real effect if it actually exists.
Briefing Logic

Start with the assumption of 'No Effect' (Null). Calculate the probability of your evidence. If P < Alpha, the evidence is strong enough to 'overthrow' the Null and claim a discovery.

01The Hypotheses

The Core Duel

The logical framework comparing the Null Hypothesis (H₀), which assumes no effect, against the Alternative Hypothesis (Hₐ), which represents the research discovery.

The Null is the gravity of science. It holds our claims to the ground until the signal is strong enough to let them fly.
The Skeptical Anchor
Why it matters

The Duel is the 'Skeptical Filter' of science. It prevents us from making false claims by establishing a baseline of 'No Change'. Discovery is not simply finding something new; it is successfully overthrowing the existing Null belief with overwhelming evidence.

When to use

The mandatory starting point for every inferential statistical test, including T-tests, ANOVA, and Regression.

Defensive Warning

Proving the Null. You can never 'Prove' the Null is true; you can only fail to find enough evidence to reject it. Absence of evidence is not evidence of absence.

Analogy & Core Concept

"Think of it as a criminal trial. The suspect (the treatment) is 'Innocent until proven guilty' (The Null). You are the prosecutor. You must find enough evidence to prove 'beyond a reasonable doubt' that the treatment actually works."

Worked Cases
Clinical Trial

Null: 'Drug A is no better than sugar water.' Alt: 'Drug A significantly reduces patient fever.'

Safety Audit

Null: 'The bridge is structurally sound.' Alt: 'The bridge has a dangerous crack that needs repair.'

The Null is the gravity of science. It holds our claims to the ground until the signal is strong enough to let them fly.
The Skeptical Anchor
02The Evidence

The P-Value

The probability of obtaining research results at least as extreme as those observed, assuming that the Null Hypothesis is actually true.

A P-value doesn't measure truth; it measures coincidence. It tells you when the silence of the Null has been broken by a scream of evidence.
The Surprise Index
Why it matters

The P-value is the 'Surprise Index'. It quantifies exactly how 'weird' your data is under the status quo. If P is low (e.g., 0.01), it means there's only a 1% chance this was a coincidence, making the treatment look like a genuine discovery.

When to use

Standard for deciding whether to 'Reject' or 'Fail to Reject' the Null hypothesis in all frequentist statistics.

Defensive Warning

The Binary Trap. Thinking that P=0.049 is 'True' and P=0.051 is 'False'. P-values are a continuous measure of evidence; don't let a hard cutoff blind you to clinical context.

Analogy & Core Concept

"Think of it as a 'Coincidence Meter'. If you flip a coin 10 times and it's always heads, the P-value is the tiny probability that a 'Fair Coin' would do that. If that probability is small enough, you stop believing the coin is fair."

Worked Cases
P = 0.04

There is a 4% chance this result is just noise. This is usually strong enough to claim a discovery.

P = 0.65

There is a 65% chance this was just luck. The evidence is far too weak to reject the Null.

A P-value doesn't measure truth; it measures coincidence. It tells you when the silence of the Null has been broken by a scream of evidence.
The Surprise Index
03The Evidence

Alpha Threshold (α)

The pre-specified level of significance (usually 0.05) representing the maximum risk of a False Positive discovery the researcher is willing to accept.

Integrity is defined before the evidence arrives. Set your Alpha to reflect the weight of the consequences, not the habits of the crowd.
The Line in the Sand
Why it matters

Alpha is the 'Line in the Sand'. By deciding the threshold BEFORE the study begins, you prevent 'P-Hacking'—the unethical practice of changing the rules after you see the results to force a significant finding.

When to use

Must be defined in the 'Methods' section of every research protocol before data collection begins.

Defensive Warning

The Traditionalist's Blindness. Just because 0.05 is standard doesn't mean it's right for every study. If the cost of a 'False Alarm' is massive, Alpha must be much smaller.

Analogy & Core Concept

"It is your budget for being wrong. If Alpha is 0.05, you are saying: 'I am okay with claiming a discovery that turns out to be fake in 5 out of every 100 studies'."

Worked Cases
α = 0.05

The standard scientific threshold. Balances the need for discovery with the need for caution.

α = 0.01

A conservative threshold for high-stakes research, like new neurosurgery techniques, where a False Positive could be fatal.

Integrity is defined before the evidence arrives. Set your Alpha to reflect the weight of the consequences, not the habits of the crowd.
The Line in the Sand
04Risk Management

Type I & II Errors

The two primary ways an inference can fail: Type I (claiming an effect that doesn't exist) and Type II (missing an effect that truly does exist).

Inference is an audit of risk. Choose your error carefully, for the one you don't fear is the one that will blind you.
The Double-Edged Sword
Why it matters

Understanding these errors allows for 'Informed Gambling'. You can never eliminate error entirely, so you must choose which risk is more dangerous for your patients: a False Alarm or a Missed Discovery.

When to use

Essential during the 'Discussion' section of a paper to explain why a result might have been non-significant (Low Power).

Defensive Warning

The Sensitivity Paradox. You cannot lower the risk of Type I errors (by lowering Alpha) without automatically increasing the risk of Type II errors, unless you increase your sample size.

Analogy & Core Concept

"Think of it as a smoke detector. Type I: The alarm goes off when there is no fire (annoying, but safe). Type II: There is a fire, but the alarm stays silent (deadly). You adjust the sensitivity based on what you fear most."

Worked Cases
Type I (False Positive)

Approving a drug that is actually useless. Risk: Patients waste money and suffer side effects for no benefit.

Type II (False Negative)

Rejecting a cure for a rare disease because the sample was too small. Risk: A life-saving treatment is lost forever.

Inference is an audit of risk. Choose your error carefully, for the one you don't fear is the one that will blind you.
The Double-Edged Sword
05Risk Management

Statistical Power (1-β)

The probability that a study will correctly detect a true effect if one actually exists in the population.

Success is a function of strength. Do not start the hunt until your flashlight is bright enough to see the prey.
The Discovery Engine
Why it matters

Power is the 'Discovery Engine'. A study with 80% power has an 80% chance of finding the 'Truth'. Conducting a low-powered study is unethical because it wastes time and money on a project likely to fail even if the theory is right.

When to use

Must be calculated and reported in the 'Sample Size Calculation' section of every professional research protocol.

Defensive Warning

The Post-Hoc Power Trap. Calculating power AFTER the study is finished using the observed results. This is mathematically redundant and often provides misleading justification for 'near-significant' findings.

Analogy & Core Concept

"Think of it as the resolution of a microscope. If you are looking for tiny bacteria (small effect), you need a high-power lens (large sample). If you use a magnifying glass, you'll see nothing and wrongly claim the bacteria don't exist."

Worked Cases
Grant Funding

A grant board rejects a study because its Power is only 40%, meaning it has a coin-flip chance of missing the target.

Sample Planning

Determining that you need exactly 120 patients to achieve 90% power to find a meaningful reduction in recovery time.

Success is a function of strength. Do not start the hunt until your flashlight is bright enough to see the prey.
The Discovery Engine
06Defensive Logic

Logical Faults

Common pitfalls, logical fallacies, and structural warnings to watch out for.

Why it's wrong
Treating a P > 0.05 as 'Proving the Null'. A non-significant result only means you failed to find enough evidence to reject the Null. It does NOT mean the Null is true.
The correction
Always state: 'We failed to detect a significant effect.' Never state: 'There was no effect.' Use Confidence Intervals to show how small the effect might actually be.
Why it's wrong
Checking P-values during the study and stopping as soon as P < 0.05, or analyzing 10 different outcomes and only reporting the one that was significant.
The correction
Pre-register your primary outcome and your Sample Size. Follow the protocol regardless of the early results to maintain the integrity of the Alpha threshold.
Why it's wrong
Focusing entirely on whether P is less than 0.05 and ignoring the Effect Size. A result can be 'Statistically Significant' (P < 0.05) but 'Clinically Meaningless' if the actual difference is tiny.
The correction
Always report the Effect Size (e.g., Cohen's d) alongside the P-value. Ask: 'Even if this is real, does it matter to the patient?'
statminds · HypothesisMind reference · v1.9.0 · updated 2026-01-187 of 7 sections