Effect Size Foundations
Beyond the P-value. Master the art of measuring the physical magnitude and clinical impact of your results.
What is it?
Effect Size Foundations represents a core statistical conceptual framework required to understand research design, data mapping, and analytical models.
Beyond the P-value. Master the art of measuring the physical magnitude and clinical impact of your results.
Goals & Indications
- Magnitude Isolation: Measure the physical size of a treatment effect, independent of sample size.
- Clinical Relevance: Determine if a 'statistically significant' result actually matters to a patient's life.
- Standardized Metrics: Use Cohen's d and Pearson's r to compare results across different measurement scales.
- Predictive Power: Quantify how much of the outcome variance is directly controlled by your intervention.
Core Idea Diagram
Key Elements
- Distance: Cohen's d, Hedges' g.
- Association: Pearson r, Odds Ratio.
- Variance: Eta-squared, R² values.
How it works
- Subtract group mean difference to find raw delta.
- Divide delta by standard deviation to scale standard units.
- Compute variance explained proportion (R²).
- Reference index coefficient against benchmark guidelines.
Defensive Pitfall
Warning: Confusing statistical significance (p-value) with practical clinical impact magnitude.
Expert Directive
“Significance only flags occurrence chance; effect sizes quantify raw clinical significance.”
Quick Reference
| Index | Small | Medium | Large |
|---|---|---|---|
| Cohen's d | 0.20 | 0.50 | 0.80 |
| Pearson r | 0.10 | 0.30 | 0.50 |
| Eta-sq (η²) | 0.01 | 0.06 | 0.14 |
Statistical Power vs. Effect Size Lab
Vary Cohen's d and sample size to watch overlap contract and significance (p-value) shift.
A significant P-value is only the invitation to look closer. The Effect Size is the actual discovery. Always prioritize the magnitude of the confidence interval over the binary P < 0.05 cutoff.
Magnitude vs. Significance
The critical distinction between statistical significance (likelihood of chance) and effect size (physical size of the result).
This is the 'Master Key' of research integrity. Significance (P-value) only tells you if a result is lucky. Magnitude (Effect Size) tells you if it actually matters to the patient. A drug can be 'Significant' but clinically useless if the change is too small to feel.
Must be addressed in the 'Discussion' section of every paper to justify the clinical value of the findings.
The P-Value Blindfold. Focusing entirely on 'P < 0.05' and ignoring the actual clinical change. This leads to 'Statistically Significant' findings that fail to replicate or provide patient value.
"Think of it as the volume vs. the clarity of a radio station. P-value is the clarity (is there a signal?). Effect Size is the volume (how loud is it?). You can have a crystal clear signal that is so quiet you can't hear the music."
In a study of 100,000 people, a 0.5 mmHg drop in blood pressure is 'Significant' (P < 0.001) but provides zero real-world health benefit.
In a study of 10 people, a massive 20 mmHg drop might not be 'Significant' (P = 0.08), but the magnitude suggests a breakthrough worth investigating further.
Cohen's d
A standardized measure of the difference between two means, expressed in units of standard deviation.
Cohen's d is the 'Universal Ruler'. It strips away original units (like kg or cm) and tells you how far the treatment 'kicked' the group. This allows you to compare the magnitude of a yoga study against a drug study on the same scale.
Primary choice for reporting standardized differences between two independent groups (e.g., T-tests).
Ignoring Variance. If your data is extremely spread out (high SD), the 'd' will be small even if the raw difference is large. Always interpret d in the context of the group's natural variation.
"It measures the physical separation of groups. A d = 1.0 means the average treated person is now 1 full standard deviation away from the control group—a massive shift in biological terms."
Comparing the 'kick' of an anti-depressant (d=0.3) against regular exercise (d=0.5) to see which has more physical impact.
Averaging 50 different studies into one master effect size to find the global truth of a treatment.
Explained Variance (R²)
The proportion of the variance in the dependent variable that is predictable from the independent variable(s).
R² is the 'Pie of Control'. It tells the researcher exactly how much of the outcome 'belongs' to the treatment. If R² is 0.30, your treatment owns 30% of the result, while 70% is still uncontrolled chaos or noise.
Standard for regression models and correlation analysis to show the 'Strength of Association'.
The Overfitting Illusion. In small samples with many variables, R² can look artificially high. Always check 'Adjusted R²' to get the honest truth.
"It measures predictive power. If you know the treatment dose, how well can you guess the patient outcome? High R² means the treatment is the primary driver of change."
Determining that 'Time Spent Studying' explains 40% (R²=0.40) of the variance in final exam scores.
Finding that a specific gene variant explains only 2% (R²=0.02) of the risk for heart disease—meaning other factors are 98% responsible.
Odds Ratio (OR)
A measure of association between an exposure and an outcome, representing the odds that an outcome will occur given a particular exposure, compared to the odds of the outcome occurring in the absence of that exposure.
The Odds Ratio is the 'Risk Multiplier'. It is the most powerful way to communicate clinical danger or benefit. Telling a doctor that a drug 'doubles the odds of recovery' (OR=2.0) is far more meaningful than reporting a P-value.
Mandatory for Case-Control studies and Logistic Regression where the outcome is binary (Yes/No).
The Probability Confusion. Odds are NOT the same as Probability. An OR of 2.0 does not always mean the 'Risk' has doubled, especially if the outcome is very common.
"It's a ratio of likelihoods. If OR = 1, there is no difference. If OR > 1, the exposure increases the odds. If OR < 1, the exposure is protective (reduces the odds)."
Calculating that smokers have 15x the odds (OR=15.0) of developing lung cancer compared to non-smokers.
Determining that vaccinated individuals have only 0.1x the odds (OR=0.1) of severe hospitalization.
Clinical Thresholds
The use of standardized benchmarks (like Cohen's 0.2, 0.5, 0.8) to categorize the practical significance of an effect size.
Thresholds are the 'Scientific Adjectives'. They bridge the gap between sterile numbers and clinical judgment. Telling a researcher that 'd = 0.85' is helpful, but telling them it is a 'Large' effect provides instant decision-making context.
Essential for setting 'Minimal Clinically Important Differences' (MCID) during the design phase of a study.
The Rigid Rule Trap. These thresholds are just guidelines. In life-saving heart surgery, even a 'Small' effect (0.2) is a massive victory. Always interpret thresholds in the context of the field.
"It provides a common language. By agreeing on what constitutes a 'Small' vs 'Large' effect, researchers across the world can align their expectations and resource allocation."
A real effect, but so small it might not be noticeable to the individual patient without a large population study.
A massive clinical shift. The difference is large enough to be obvious to the naked eye of a practitioner.
The P-Value Trap
Common pitfalls, logical fallacies, and structural warnings to watch out for.