Atlas
statminds
Meta-AnalysisThe underlying model family class (e.g. GLM, linear model, categorical matrix, log-linear).Parametric ReferenceStatistical methods that assume a specific probability distribution family (typically normal).12-stage workflow

Cochran's Q Test for Heterogeneity

Tests whether studies share a common true effect; significant Q indicates heterogeneity justifying random-effects model..

Model familyMeta-Analysis
Hypothesisheterogeneity_assessment
AliasesCochran_Q · Q_test · heterogeneity_test · homogeneity_test
G1
heterogeneity_assessment
G2
model_selection
Visual Overview Dashboard
1

What is it?

Cochran's Q Test for Heterogeneity is designed to mathematically synthesize evidence across multiple independent studies to resolve clinical uncertainty.

Tests whether studies share a common true effect; significant Q indicates heterogeneity justifying random-effects model.

2

Goals & Indications

  • heterogeneity_assessment
  • model_selection
3

Core Idea Diagram

Alpha Threshold (9.49)Calculated Q
4

Hypotheses

H₀: H₀: τ² = 0 (all studies share the same true effect; homogeneity)
Hₐ: Hₐ: τ² > 0 (true effects vary across studies; heterogeneity)
5

How it works

  1. Calculate the common-effect pooled estimate as a reference point.
  2. Compute Q = sum( w * (ES_i - ES_pooled)² ) to sum standardized deviations.
  3. Reference Q against a Chi-Square distribution with df = k - 1.
  4. A significant p-value (p < 0.05) indicates significant study heterogeneity.
6

Assumptions

Independence of studies: Each study provides independent information; no shared participants or datasets
Within-study variances correctly estimated: Reported standard errors accurately reflect sampling variability
Effect sizes comparable: All studies use same effect size metric with consistent direction
7

Important Note

Cochran's Q tests whether observed effect size variation exceeds what would be expected from sampling error alone. Under H₀, all studies estimate the same true effect, and variation is due only to sampling error. Significant Q (typically p<.10) indicates heterogeneity, suggesting random-effects meta-analysis is more appropriate than fixed-effect. Q follows χ² distribution with k-1 degrees of freedom, where k = number of studies. CRITICAL: Q is a test of heterogeneity presence, NOT a measure of heterogeneity magnitude. Use I² and τ² to quantify heterogeneity amount. Q has low power with k<10 and is almost always significant with k>20, limiting interpretation.

8

Worked Example

ConditionQ-statp-value
Consistent2.480.648
Heterogeneous12.850.012
Interactive Sandbox

Cochran's Q Test Statistic Curve

Increase study dispersion to pull the calculated Q statistic along the Chi-Square curve, crossing the critical value boundary into the significant region.

Study Dispersion0.80

Calculated Q-statistic: 6.504
Critical value (df=4): 9.488
p-value: 0.16451
Chi-Square Distribution Curve (df=4)
Crit value (9.49)Q=6.50.05.09.515.020.0
01Hypothesis test logic

Hypotheses

Pragmatic null and alternative hypotheses defined in mathematical notation.

A hypothesis is a question sharpened to a point. Ambiguity is the enemy of inference.
Logic Core
Null · H₀

H₀: τ² = 0 (all studies share the same true effect; homogeneity)

Alternative · Hₐ

Hₐ: τ² > 0 (true effects vary across studies; heterogeneity)

Why it matters heterogeneity_assessment

Cochran's Q tests whether observed effect size variation exceeds what would be expected from sampling error alone. Under H₀, all studies estimate the same true effect, and variation is due only to sampling error. Significant Q (typically p<.10) indicates heterogeneity, suggesting random-effects meta-analysis is more appropriate than fixed-effect. Q follows χ² distribution with k-1 degrees of freedom, where k = number of studies. CRITICAL: Q is a test of heterogeneity presence, NOT a measure of heterogeneity magnitude. Use I² and τ² to quantify heterogeneity amount. Q has low power with k<10 and is almost always significant with k>20, limiting interpretation.

02Model diagnostics

Assumptions

The core mathematical criteria needed to ensure that statistical testing remains unbiased and valid.

Build your analysis on rock, not sand. Verify the mathematical foundation before building the model.
Integrity Shield
7
Assumptions
4
Critical / High Severity
How to check
Quick
Review study descriptions for duplicate data; check author affiliations and recruitment sites; examine publication dates and trial registrations; verify no multiple publications from same cohort reported as separate studies
Rigorous
Contact study authors to confirm independence; cross-reference trial registrations (ClinicalTrials.gov, ISRCTN) for duplicate entries; check for shared control groups in multi-arm trials; examine baseline characteristics tables for identical values suggesting overlapping samples
If violated
If studies share participants: (1) Select only one publication per cohort (largest sample or best quality); (2) Use robust variance estimation accounting for clustering; (3) Apply multilevel meta-analysis with publications nested within cohorts. If shared control groups: Use appropriate corrections (Higgins & Cochrane methods for multi-arm trials). Never include same participants twice—inflates precision artificially and biases Q statistic downward (underestimates heterogeneity)
How to check
Quick
Verify that variances/SEs are reported or calculable from study statistics; check for implausibly small SEs (e.g., SE < 0.01 with n < 50); confirm variance formula matches effect size metric (e.g., SE for Hedges' g vs. log OR differs); look for studies reporting only p-values without CIs
Rigorous
Recalculate variances from raw data when available; use established formulas for each effect size metric (Borenstein et al. 2009); conduct sensitivity analysis varying imputed variances for studies with missing information; check if variance approximations (e.g., for log OR) are valid given sample sizes
If violated
If variances unavailable: (1) Calculate from other statistics (p-values, CIs, t-statistics using standard formulas); (2) Impute using median variance from similar studies (conservative); (3) Exclude studies lacking sufficient information (report as sensitivity). If variance estimates questionable: Conduct sensitivity analysis with plausible variance ranges. Variance errors inflate Q (overestimate heterogeneity) if too small, deflate Q if too large
How to check
Quick
Verify all effect sizes are same metric (all Hedges' g, or all log OR, etc.); check direction coding consistency (e.g., positive always = treatment benefit); confirm outcome constructs measure same phenomenon (e.g., all depression scales, not mixing depression and anxiety)
Rigorous
Create detailed coding manual specifying metric and direction; have independent coders extract effect sizes; calculate inter-rater reliability (ICC > .90); convert metrics if needed using validated formulas; standardize all to common scale (e.g., Hedges' g for SMDs)
If violated
If mixed metrics: Convert all to common metric using established formulas (e.g., log OR to Hedges' g: d ≈ log(OR) × √3/π). If inconsistent direction: Recode so positive values consistently indicate same direction (e.g., treatment benefit). If incompatible outcomes: Conduct separate meta-analyses or use standardized metrics. Mixing metrics without conversion inflates Q statistic artificially (creates spurious heterogeneity)
How to check
Quick
Check individual study sample sizes; verify most studies have n > 30 per group; calculate total sample size across meta-analysis; examine if any studies have very small samples (n < 20) that could distort Q statistic
Rigorous
Review sampling distribution assumptions for each effect size metric; verify asymptotic approximations hold (typically requires n > 30-50 per study); conduct simulation studies or permutation tests if sample sizes questionable; examine Q-Q plots of standardized residuals
If violated
If small samples (n < 30): (1) Use exact or permutation-based heterogeneity tests rather than χ² approximation; (2) Apply small-sample corrections (e.g., Hartung-Knapp adjustment); (3) Interpret Q test cautiously alongside I² and τ². With very small samples (n < 20), Q test unreliable—rely on graphical assessment (forest plots) and τ² estimates instead. Small samples increase Q test variability but don't systematically bias it
How to check
Quick
Create funnel plot and assess asymmetry; conduct Egger's regression test; compare published vs. unpublished study effect sizes; examine if small studies show larger effects (small-study effects); check for excess of barely-significant p-values
Rigorous
Use multiple publication bias methods: trim-and-fill, PET-PEESE, selection models, p-curve analysis; search for unpublished studies (trial registries, gray literature, conference abstracts); contact authors for unpublished data; compare pre-registered vs. published analyses; assess time-lag bias
If violated
Publication bias can inflate Q (increase apparent heterogeneity) if small null studies are missing OR deflate Q if only significant studies published (all similar direction). If bias detected: (1) Include unpublished studies and gray literature; (2) Conduct sensitivity analysis with/without small studies; (3) Use selection models accounting for bias; (4) Report Q alongside bias-corrected heterogeneity estimates. Note: Bias affects heterogeneity assessment differently than pooled effect estimation
How to check
Quick
Count number of studies (k) in meta-analysis; verify k ≥ 3 for Q test computation (df = k-1 ≥ 2); note if k < 10 (low power) or k > 20 (high power); assess if sample is sufficient for heterogeneity detection given expected τ²
Rigorous
Conduct power analysis for Q test given k, expected τ², and α level; use simulation studies to estimate power; recognize that Q has ~10% power to detect moderate heterogeneity (I²=30%) with k=5, but >80% power with k≥15; plan meta-analysis with adequate k for meaningful heterogeneity assessment
If violated
If k = 2: Q test has only 1 df, very low power—do not rely on Q test; report τ² and I² descriptively; use graphical assessment. If k < 5: Q test severely underpowered—interpret non-significant Q as 'insufficient evidence' not 'no heterogeneity'; default to random-effects model given uncertainty. If k < 3: Cannot compute Q test—meta-analysis generally not appropriate with k < 3; consider narrative synthesis or wait for more studies
How to check
Quick
Verify metric matches outcome type: continuous outcomes use SMD (Hedges' g, Cohen's d); binary outcomes use OR, RR, or RD; correlational data use r or Fisher's z; ensure metric is meaningful for clinical/practical interpretation
Rigorous
Review metric properties: bounded vs. unbounded, scale-dependent vs. standardized; check if transformations needed (e.g., Fisher's z for correlations); verify metric assumptions hold (e.g., OR assumes proportional odds); assess if metric choice affects heterogeneity assessment (some metrics more variable than others)
If violated
If inappropriate metric: Convert to appropriate metric using established formulas; for correlations, always use Fisher's z for pooling, then back-transform; for binary outcomes, choose OR, RR, or RD based on clinical meaning and statistical properties. If metric assumptions violated: Use alternative metrics or transformations. Note: Metric choice affects Q magnitude—bounded metrics (e.g., r) may show lower Q than unbounded (e.g., OR) even with same underlying heterogeneity
03Residual Forensics

Diagnostics

Checking residual plots and indices to examine model deviations and ensure standard error integrity.

Trust, but verify. The outliers often hold more truth than the averages.
System Health
Essential checks
  1. Q statistic value (test statistic for heterogeneity)
  2. Degrees of freedom (df = k - 1, where k = number of studies)
  3. p-value from χ² distribution (compare to α = .10, not .05)
  4. I² statistic (% variance due to heterogeneity, complements Q)
  5. τ² estimate (between-study variance, quantifies heterogeneity magnitude)
  6. Number of studies (k) and total sample size for context
Recommended checks
  1. Forest plot showing individual study effects and heterogeneity visually
  2. Comparison of fixed-effect vs. random-effects estimates (illustrates impact of heterogeneity)
  3. Confidence interval for τ² (uncertainty in heterogeneity estimate)
  4. H² statistic (ratio of total to sampling variance, H² = Q / df)
  5. Power analysis for Q test given k (interpret null results appropriately)
  6. Subgroup-specific Q statistics if conducting subgroup analysis (Q_within)
  7. Q_between statistic for testing moderators (difference between subgroups)
  8. Sensitivity analysis of Q statistic excluding influential studies
  9. Graphical display of Q decomposition (within vs. between subgroup variation)
04Live Instances

Applied Minds

Review concrete study examples, data layout guidelines, and copy executable syntax scripts.

Theory is the map. Practice is the terrain. Simulation bridges the gap.
Applied Wisdom
Example 01

Assessing Heterogeneity in CBT for Depression Meta-Analysis

Research question: Do the k = 15 randomized controlled trials of CBT for depression show heterogeneity in effect sizes, or do they estimate a common true effect? Design: Cochran's Q test applied to meta-analytic dataset of 15 RCTs (N = 1,847 participants) examining CBT vs. control. Outcome: Hedges' g (standardized mean difference) for depression symptom reduction. This example demonstrates computation of Q statistic, interpretation of p-value using α = .10 threshold, assessment of power given k, and integration with I² and τ² for comprehensive heterogeneity assessment. The example illustrates why Q alone is insufficient—must complement with effect size measures (I², τ²) and consider power. Additionally, shows Q decomposition in subgroup analysis (therapy format: individual vs. group CBT).

DesignCochran's Q test for heterogeneity in meta-analysis
Total n1847
Outcome ScaleDepression symptom reduction (BDI/HRSD)
# Cochran's Q Test for Heterogeneity Assessment
# Demonstrating Q statistic computation, interpretation, and integration with I² and τ²

library(metafor)      # rma() for meta-analysis, Q test built-in
library(meta)         # metagen() alternative
library(dplyr)
library(ggplot2)

# === STEP 1: Simulate Meta-Analytic Dataset ===
# In practice: data <- read.csv("meta_analysis_data.csv")
set.seed(2025)
k <- 15  # Number of studies

# Simulate effect sizes with MODERATE heterogeneity
# True effects vary: mean θ = 0.70, between-study SD τ = 0.20 (I² ≈ 50%)
true_effects <- rnorm(k, mean = 0.70, sd = 0.20)

# Sample sizes vary
n_treat <- sample(40:100, k, replace = TRUE)
n_control <- sample(40:100, k, replace = TRUE)
total_n <- n_treat + n_control

# Sampling standard errors
sampling_se <- sqrt((n_treat + n_control) / (n_treat * n_control) + 
                     true_effects^2 / (2 * (n_treat + n_control)))

# Observed effect sizes (true + sampling error)
observed_g <- rnorm(k, mean = true_effects, sd = sampling_se)
variance_g <- sampling_se^2

# Add therapy format moderator (individual vs. group)
therapy_format <- sample(c("Individual", "Group"), k, replace = TRUE, prob = c(0.6, 0.4))

meta_data <- data.frame(
  study_id = paste0("Study_", 1:k),
  author_year = paste0(LETTERS[1:k], " et al.(20", 10:24, ")"),
  hedges_g = observed_g,
  variance = variance_g,
  se = sqrt(variance_g),
  n_treatment = n_treat,
  n_control = n_control,
  total_n = total_n,
  therapy_format = therapy_format
)

print("=== Meta-Analytic Dataset ===")
print(meta_data)
cat("\nTotal N =", sum(meta_data$total_n), "participants across", k, "studies")

# === STEP 2: Compute Cochran's Q Test (via metafor) ===
# Random-effects model automatically computes Q statistic
re_model <- rma(yi = hedges_g, vi = variance, data = meta_data, 
                method = "REML", slab = author_year)

print("\n\n=== COCHRAN'S Q TEST FOR HETEROGENEITY ===")
cat("\nTest of Heterogeneity:")
cat("\nQ statistic  =", round(re_model$QE, 2))
cat("\nDegrees of freedom(df) =", re_model$k - 1)
cat("\np-value =", format.pval(re_model$QEp, digits = 3))

# Critical value from χ² distribution
alpha <- 0.10  # Use .10, not .05, for heterogeneity tests
critical_value <- qchisq(1 - alpha, df = re_model$k - 1)
cat("\n\nCritical value(α = .10): χ²(", re_model$k - 1, ") =", round(critical_value, 2))

if (re_model$QEp < alpha) {
  q_conclusion <- "SIGNIFICANT heterogeneity detected"
  q_interpretation <- paste0(
    "Reject H₀ (p < .10). Effect sizes vary more than expected from sampling error alone.\n",
    "Random-effects meta-analysis is STRONGLY RECOMMENDED to account for between-study variance."
  )
} else {
  q_conclusion <- "No significant heterogeneity detected"
  q_interpretation <- paste0(
    "Fail to reject H₀ (p ≥ .10). Insufficient evidence that effects vary beyond sampling error.\n",
    "However, with k = ", k, ", Q test has LIMITED POWER. Random-effects model may still be preferred for generalization."
  )
}

cat("\n\nConclusion:", q_conclusion)
cat("\n\nInterpretation:", q_interpretation)

# === STEP 3: Extract and Interpret Heterogeneity Statistics ===
tau2 <- re_model$tau2
tau <- sqrt(tau2)
I2 <- re_model$I2
H2 <- re_model$H2

cat("\n\n=== HETEROGENEITY STATISTICS ===")
cat("\nτ² (tau-squared)  =", round(tau2, 4), "(between-study variance)")
cat("\nτ (tau)           =", round(tau, 3), "(between-study SD)")
cat("\nI²                =", round(I2, 1), "% (percent variance due to heterogeneity)")
cat("\nH²                =", round(H2, 2), "(variance inflation factor)")

if (I2 < 25) {
  I2_interp <- "low"
} else if (I2 < 50) {
  I2_interp <- "moderate"
} else if (I2 < 75) {
  I2_interp <- "substantial"
} else {
  I2_interp <- "considerable"
}

cat("\n\nI² Interpretation: Heterogeneity is", I2_interp)
cat("\n\nCRITICAL NOTE:")
cat("\n- Q statistic tests IF heterogeneity exists(hypothesis test, p-value)")
cat("\n- I² quantifies HOW MUCH heterogeneity(effect size, percentage)")
cat("\n- τ² quantifies heterogeneity in original metric units(variance)")
cat("\n- ALWAYS report all three: Q provides statistical test, I² and τ² provide magnitude\n")

# === STEP 4: Manual Calculation of Q (for pedagogical understanding) ===
cat("\n\n=== MANUAL CALCULATION OF Q STATISTIC(Step-by-Step) ===")

# Step 1: Fixed-effect weights (inverse variance)
w_fixed <- 1 / meta_data$variance
cat("\nStep 1: Calculate fixed-effect weights w_i = 1 / SE_i²")
cat("\nWeights(w_i):", paste(round(w_fixed, 2), collapse = ", "))

# Step 2: Pooled effect under fixed-effect model
pooled_fixed <- sum(w_fixed * meta_data$hedges_g) / sum(w_fixed)
cat("\n\nStep 2: Calculate pooled effect(fixed-effect)")
cat("\nθ̂ = Σ(w_i × θ_i) / Σw_i =", round(pooled_fixed, 3))

# Step 3: Squared deviations
deviations <- meta_data$hedges_g - pooled_fixed
squared_deviations <- deviations^2
weighted_squared_dev <- w_fixed * squared_deviations

cat("\n\nStep 3: Calculate squared deviations(θ_i - θ̂)²")
dev_table <- data.frame(
  Study = 1:k,
  Effect = round(meta_data$hedges_g, 3),
  Deviation = round(deviations, 3),
  Squared_Dev = round(squared_deviations, 4),
  Weight = round(w_fixed, 2),
  Weighted = round(weighted_squared_dev, 3)
)
print(head(dev_table, 5))
cat("\n...(showing first 5 of", k, "studies)")

# Step 4: Sum to get Q
Q_manual <- sum(weighted_squared_dev)
cat("\n\nStep 4: Sum weighted squared deviations")
cat("\nQ = Σ[w_i × (θ_i - θ̂)²] =", round(Q_manual, 2))

# Verify matches metafor output
cat("\n\nVerification: Manual Q =", round(Q_manual, 2), 
    "| metafor Q =", round(re_model$QE, 2))
cat("\nMatch:", ifelse(abs(Q_manual - re_model$QE) < 0.01, "YES ✓", "Check calculation"))

# === STEP 5: Power Analysis for Q Test ===
cat("\n\n=== POWER ANALYSIS FOR Q TEST ===")
cat("\nNumber of studies(k) =", k)
cat("\nDegrees of freedom(df) =", k - 1)

# Simulate power: probability of detecting heterogeneity given true I²
# Approximate using non-central χ² distribution
if (I2 > 0) {
  # Non-centrality parameter (NCP) approximation
  # NCP ≈ Q under alternative (depends on true τ² and study weights)
  # Rough approximation: NCP ≈ Q_observed if heterogeneity truly exists
  ncp_approx <- re_model$QE
  
  # Power: probability Q > critical value under alternative
  power_estimate <- 1 - pchisq(critical_value, df = k - 1, ncp = ncp_approx)
  
  cat("\n\nApproximate power to detect observed heterogeneity(I² =", round(I2, 1), "%):", 
      round(power_estimate * 100, 1), "%")
} else {
  cat("\n\nPower calculation not applicable(I² ≈ 0)")
}

cat("\n\nGeneral power guidelines for Q test:")
cat("\n- k < 10:  LOW power(~10-30% for moderate I² = 30-50%)")
cat("\n- k = 10-15: MODERATE power(~50-70% for moderate I²)")
cat("\n- k ≥ 20:  HIGH power(>80% for moderate I², >90% for substantial I²)")
cat("\n\nImplication with k =", k, ":")
if (k < 10) {
  cat(" LOW power. Non-significant Q doesn't prove homogeneity.")
  cat("\n  → DEFAULT to random-effects model given uncertainty.")
} else if (k < 20) {
  cat(" MODERATE power. Q test reasonably reliable.")
  cat("\n  → Use Q alongside I² and τ² for model selection.")
} else {
  cat(" HIGH power. Q test likely detects even small heterogeneity.")
  cat("\n  → Significant Q may reflect trivial heterogeneity; check I² for magnitude.")
}

# === STEP 6: Visual Assessment (Forest Plot) ===
par(mar = c(5, 4, 3, 2))
forest(re_model,
       xlab = "Hedges' g(CBT - Control)",
       header = c("Study", "g [95% CI]"),
       cex = 0.8,
       col = "steelblue",
       border = "steelblue")

mtext(paste0("Cochran's Q(", k - 1, ") = ", round(re_model$QE, 2), 
             ", p = ", format.pval(re_model$QEp, digits = 3),
             " | I² = ", round(I2, 1), "% (", I2_interp, " heterogeneity)"),
      side = 3, line = 0, cex = 0.9, font = 2)

# === STEP 7: Subgroup Analysis (Q Decomposition) ===
cat("\n\n=== SUBGROUP ANALYSIS: Q DECOMPOSITION ===")
cat("\nModerator: Therapy Format(Individual vs. Group CBT)\n")

# Subgroup meta-analysis
subgroup_model <- rma(yi = hedges_g, vi = variance, 
                      mods = ~ therapy_format - 1,  # Separate estimates per group
                      data = meta_data, method = "REML")

# Calculate Q statistics for subgroups manually
individual_data <- meta_data[meta_data$therapy_format == "Individual", ]
group_data <- meta_data[meta_data$therapy_format == "Group", ]

if (nrow(individual_data) >= 3 && nrow(group_data) >= 3) {
  # Individual CBT subgroup
  re_individual <- rma(yi = hedges_g, vi = variance, 
                       data = individual_data, method = "REML")
  Q_individual <- re_individual$QE
  df_individual <- re_individual$k - 1
  
  # Group CBT subgroup  
  re_group <- rma(yi = hedges_g, vi = variance, 
                  data = group_data, method = "REML")
  Q_group <- re_group$QE
  df_group <- re_group$k - 1
  
  # Q decomposition
  Q_within <- Q_individual + Q_group
  df_within <- df_individual + df_group
  Q_between <- re_model$QE - Q_within
  df_between <- (re_model$k - 1) - df_within
  p_between <- 1 - pchisq(Q_between, df = df_between)
  
  cat("\nQ Decomposition:")
  cat("\n  Q_total(", re_model$k - 1, ") =", round(re_model$QE, 2))
  cat("\n\n  Q_within(", df_within, ") =", round(Q_within, 2), 
      "(heterogeneity within subgroups)")
  cat("\n    - Q_individual(", df_individual, ") =", round(Q_individual, 2))
  cat("\n    - Q_group(", df_group, ") =", round(Q_group, 2))
  cat("\n\n  Q_between(", df_between, ") =", round(Q_between, 2), 
      "(heterogeneity between subgroups)")
  cat("\n    p-value =", format.pval(p_between, digits = 3))
  
  if (p_between < 0.10) {
    cat("\n\n  → SIGNIFICANT difference between subgroups(p < .10)")
    cat("\n    Therapy format is a significant moderator of CBT effects.")
  } else {
    cat("\n\n  → No significant difference between subgroups(p ≥ .10)")
    cat("\n    Therapy format does not significantly moderate CBT effects.")
  }
  
  # Subgroup estimates
  cat("\n\nSubgroup Estimates:")
  cat("\n  Individual CBT: g =", round(re_individual$beta[1], 3), 
      ", 95% CI [", round(re_individual$ci.lb, 3), ",", 
      round(re_individual$ci.ub, 3), "], I² =", round(re_individual$I2, 1), "%")
  cat("\n  Group CBT:      g =", round(re_group$beta[1], 3),
      ", 95% CI [", round(re_group$ci.lb, 3), ",",
      round(re_group$ci.ub, 3), "], I² =", round(re_group$I2, 1), "%")
} else {
  cat("\nInsufficient studies per subgroup for Q decomposition(need k ≥ 3 per group)")
}

# === STEP 8: Comparison with Fixed-Effect Model ===
cat("\n\n=== COMPARISON: FIXED vs. RANDOM EFFECTS ===")

fe_model <- rma(yi = hedges_g, vi = variance, data = meta_data, 
                method = "FE", slab = author_year)

cat("\nFixed-Effect Model(assumes τ² = 0):")
cat("\n  Pooled g =", round(as.numeric(fe_model$beta), 3))
cat("\n  95% CI: [", round(fe_model$ci.lb, 3), ",", round(fe_model$ci.ub, 3), "]")
cat("\n  SE =", round(fe_model$se, 4))

cat("\n\nRandom-Effects Model(estimates τ² from data):")
cat("\n  Pooled g =", round(as.numeric(re_model$beta), 3))
cat("\n  95% CI: [", round(re_model$ci.lb, 3), ",", round(re_model$ci.ub, 3), "]")
cat("\n  SE =", round(re_model$se, 4))
cat("\n  τ² =", round(tau2, 4))

ci_width_fe <- fe_model$ci.ub - fe_model$ci.lb
ci_width_re <- re_model$ci.ub - re_model$ci.lb
inflation <- (ci_width_re / ci_width_fe - 1) * 100

cat("\n\nDifference:")
cat("\n  Random-effects CI is", round(inflation, 1), "% wider than fixed-effect")
cat("\n(reflects additional uncertainty from between-study variance τ²)")

if (re_model$QEp < 0.10) {
  cat("\n\n→ Given significant Q test(p < .10) and I² =", round(I2, 1), "%,")
  cat("\n  RANDOM-EFFECTS model is STRONGLY RECOMMENDED.")
} else if (k < 10) {
  cat("\n\n→ Given k < 10 (low power for Q test),")
  cat("\n  RANDOM-EFFECTS model is RECOMMENDED for conservative generalization.")
} else {
  cat("\n\n→ Given non-significant Q test, FIXED-EFFECT model may be appropriate,")
  cat("\n  but RANDOM-EFFECTS often preferred for generalization beyond observed studies.")
}

# === STEP 9: Sensitivity Analysis (Influential Studies) ===
cat("\n\n=== SENSITIVITY ANALYSIS: INFLUENCE ON Q STATISTIC ===")

influence_Q <- numeric(k)
influence_I2 <- numeric(k)

for (i in 1:k) {
  # Remove study i
  mask <- (1:k) != i
  loo_model <- rma(yi = hedges_g[mask], vi = variance[mask], 
                   data = meta_data[mask, ], method = "REML")
  influence_Q[i] <- loo_model$QE
  influence_I2[i] <- loo_model$I2
}

influence_results <- data.frame(
  Study = meta_data$author_year,
  Q_without = round(influence_Q, 2),
  I2_without = round(influence_I2, 1),
  Q_change = round(influence_Q - re_model$QE, 2),
  I2_change = round(influence_I2 - I2, 1)
)

cat("\nLeave-One-Out Results(showing first 5):")
print(head(influence_results, 5))

max_Q_change <- max(abs(influence_results$Q_change))
most_influential <- which.max(abs(influence_results$Q_change))

cat("\n\nMost influential study:", meta_data$author_year[most_influential])
cat("\n  Removing this study changes Q by", round(influence_results$Q_change[most_influential], 2))
cat("\n  and changes I² by", round(influence_results$I2_change[most_influential], 1), "%")

if (max_Q_change > re_model$QE * 0.2) {
  cat("\n\n→ Q statistic SENSITIVE to individual studies(>20% change possible)")
  cat("\n  Report results with/without influential studies for transparency.")
} else {
  cat("\n\n→ Q statistic ROBUST to individual studies(<20% change across all leave-one-out)")
}

# === STEP 10: APA-Style Reporting ===
cat("\n\n=== APA-STYLE REPORTING TEMPLATE ===")
cat("\n\nCochran's Q test assessed heterogeneity in effect sizes across", k, "RCTs")
cat("\n(N =", sum(meta_data$total_n), "participants) examining CBT for depression.")
cat("\nThe test revealed", ifelse(re_model$QEp < 0.10, "significant", "non-significant"),
    "heterogeneity,")
cat("\nQ(", re_model$k - 1, ") =", round(re_model$QE, 2), ",",
    ifelse(re_model$QEp < 0.001, " p < .001", 
           paste0(" p = ", format.pval(re_model$QEp, digits = 3))),
    ", indicating that effect sizes")
cat("\nvaried", ifelse(re_model$QEp < 0.10, "significantly beyond", "within"),
    "sampling error expectations.")
cat("\n\nHeterogeneity statistics further quantified this variation: I² =", 
    round(I2, 1), "%")
cat("\n(", I2_interp, "heterogeneity) and τ² =", round(tau2, 3), 
    "(τ =", round(tau, 3), "),")
cat("\nindicating that approximately", round(I2, 0), "% of total variance was due to")
cat("\nbetween-study differences rather than sampling error.")

if (re_model$QEp < 0.10) {
  cat("\n\nGiven significant heterogeneity, a random-effects meta-analysis model was employed")
  cat("\nto account for both within- and between-study variance, yielding a pooled effect")
  cat("\nof g =", round(as.numeric(re_model$beta), 2), ", 95% CI [", 
      round(re_model$ci.lb, 2), ",", round(re_model$ci.ub, 2), "].")
} else {
  cat("\n\nWhile the Q test was non-significant, a random-effects model was retained")
  cat("\nfor conservative estimation and generalization beyond the observed studies,")
  cat("\nespecially given the limited power of the Q test with k =", k, "studies.")
}

cat("\n\nSensitivity analysis(leave-one-out) showed that no single study")
cat("\ndisproportionately influenced the heterogeneity assessment(Q range:")
cat("\n", round(min(influence_Q), 2), "to", round(max(influence_Q), 2), "),")
cat("\nsupporting the robustness of the heterogeneity conclusions.")

cat("\n\n[End of Analysis]\n")
Interpretation Blueprint

Cochran's Q(14) = 29.17, p = .010 (significant at α = .10). Conclusion: REJECT H₀ of homogeneity. Effect sizes vary significantly beyond what would be expected from sampling error alone, indicating genuine between-study heterogeneity. Heterogeneity magnitude: I² = 52% (moderate to substantial), τ² = 0.042 (τ = 0.20). Approximately 52% of total variance is due to true heterogeneity rather than sampling error. Model recommendation: Random-effects meta-analysis STRONGLY RECOMMENDED to account for between-study variance. Power consideration: With k = 15, the Q test has moderate power (~60-70%) to detect moderate heterogeneity; significant result is reliable. Sensitivity: Leave-one-out analysis shows Q ranges from 26.5 to 31.2, indicating robustness. Clinical implication: CBT effects vary meaningfully across studies (not all studies estimate same true effect); investigate moderators (e.g., therapy format, depression severity) to explain heterogeneity sources.

05Tactical Pivots

Alternatives

Structured fallback pathways for choosing alternative tests when normality or slopes requirements fail.

When the path is blocked, pivot. Rigor is not rigidity; it is the intelligent adaptation to reality.
Adaptive Strategy
Synthesis Precision Ladder Ideal · Standardized Effect Distribution
Continuous MD
Maintain Cochran's Q. The absolute standard for auditing study-level divergence in meta-analysis.
Peak Signal
Ordinal Ranks
Abandon Q. Use Non-Parametric Heterogeneity tests if effect sizes cannot be standardized.
Model Collapse
Temporal Trajectory Audit Static Inconsistency Snapshot
Static Audit
Total variability.
Stay with Cochran's Q. Identify if study differences are more than random sampling error.
Temporal Divergence
Evolution of mess.
Pivot to Meta-Regression with 'Year' as a moderator to explain increasing/decreasing heterogeneity.
Adaptive Technical Safeguards · adaptive safeguards
low study count
  • I-Squared Statistic — Focus on the 'Percentage' of inconsistency rather than the p-value of the Q-strike.
  • Exact Q-Test — resample the null distribution for tiny study pools (k < 5).
extreme outliers detected
  • Leave-One-Out Audit — Remove the rogue study to see if the Q-statistic significance evaporates.
  • Gosh Plots — Visualize every possible model combination to hunt for stable study-clusters.
06Adjusted Comparisons

Post-hoc

Group mean comparisons and correction controls (e.g. Tukey HSD, Bonferroni) to protect against Family-Wise Error Rates.

The omnibus test opens the door; post-hoc analysis explores the room.
Forensic Detail
Adjusted Comparisons

Post-hoc pairwise tests defined for this model.

Interpretation Guidelines

No specific guidelines provided.

07Standardized scale impact

Effect Size

Understanding effect sizes (e.g., Cohen's d, Partial Eta-Squared) and clinical impact benchmarks.

Significance is noise. Magnitude is the signal. Measure the impact, not just the probability.
Impact Magnitude

Test statistic for heterogeneity. Q follows χ²(k-1) under H₀ of homogeneity. Larger Q = more variation. Sample-size dependent: increases with k and study precision. Interpret via p-value (α = .10) alongside I² and τ².

p < .10 (note: NOT .05) indicates significant heterogeneity. Reject H₀: studies don't share common true effect. p ≥ .10: insufficient evidence of heterogeneity (doesn't prove homogeneity, especially with k < 10).

% of variance due to heterogeneity. <25% low, 25-50% moderate, 50-75% substantial, >75% considerable. Preferred over Q for magnitude because less sample-size dependent. I² = 0% means all variation is sampling error; I² = 100% means all variation is true heterogeneity.

Between-study variance in original metric units (squared). τ (SD) more interpretable: if θ̂ = 0.50, τ = 0.20, true effects vary ~0.30-0.70. Larger τ² = more heterogeneity. Used in random-effects model to weight studies.

Variance inflation factor. H² = 1 means no heterogeneity (τ² = 0); H² > 1 indicates heterogeneity. H² = 2 means total variance is twice sampling variance. Related to I²: I² = (H² - 1) / H².

Recommended Metric: ALWAYS report: (1) Q statistic with df and p-value (use α = .10); (2) I² with interpretation (<25% low, 25-50% moderate, etc.); (3) τ² and τ for magnitude in original units; (4) Number of studies (k) for power context. NEVER report Q alone—Q tests presence, I² and τ² quantify magnitude.
Small
0.2
Medium
0.5
Large
0.8
0.50
ALWAYS report: (1) Q statistic with df and p-value (use α = .10); (2) I² with interpretation (<25% low, 25-50% moderate, etc.); (3) τ² and τ for magnitude in original units; (4) Number of studies (k) for power context. NEVER report Q alone—Q tests presence, I² and τ² quantify magnitude.
Recommended Measure
4
Available Metrics
ReportUse ALWAYS report: (1) Q statistic with df and p-value (use α = .10); (2) I² with interpretation (<25% low, 25-50% moderate, etc.); (3) τ² and τ for magnitude in original units; (4) Number of studies (k) for power context. NEVER report Q alone—Q tests presence, I² and τ² quantify magnitude. to represent clinical impact magnitude.
08Statistical Power

Sample Size

Guidelines for minimum sample requirements and power analysis parameters.

An underpowered study is an ethical failure. Respect the data by collecting enough of it.
Power Protocol
Floor Requirements

The 'Homogeneity Minimum': A minimum of 5 studies (k >= 5) is required to ensure the Chi-Square approximation has enough 'Pulse' to distinguish signal from sampling error.

Effect SizeParametersRequired n
Small EffectI² = 25% (Small)k ≈ 30
Medium EffectI² = 50% (Medium)k ≈ 15
Large EffectI² = 75% (High)k ≈ 10
Key considerations

The 'Power Fallacy': A non-significant Q-test doesn't always mean your studies are similar—it often just means you didn't have enough studies to detect the mess. Always report I² to provide a descriptive context for the Q-strike.

G*Power StrategyBenchmark: χ² tests → Cochran's Q (Heterogeneity). Parameters: Study count (k), Tau magnitude (τ), α = .05, Power = .80. Note: Cochran's Q has notoriously low power when study counts are small (k < 10).
09APA narrative blueprint

Reporting

How to compile statistical results into publication prose matching APA and journal style guides.

Data does not speak for itself. It requires a translator. Be clear, be precise, be honest.
Narrative Arc
Reusable template

Cochran's Q test assessed heterogeneity across k studies. The test revealed significant/non-significant heterogeneity, Q(df) = X.XX, p < .001 / = .XXX, indicating that effect sizes varied significantly beyond / within sampling error expectations. Heterogeneity statistics: I² = XX% (low/moderate/substantial/considerable heterogeneity), τ² = X.XXX (τ = X.XX). Given significant/non-significant heterogeneity, a random-effects/fixed-effect meta-analysis model was employed.

Essential statistics to report
  • Q statistic value
  • Degrees of freedom (df = k - 1)
  • p-value (compared to α = .10, not .05)
  • I² percentage with interpretation category
  • τ² and τ values (between-study variance and SD)
  • Number of studies (k) for power context
  • Model choice justification (fixed vs. random effects)
10Exhibit Builder

Manuscript Lab

Copy standard summary tables and forensic reporting grids to outline analysis details.

Table 1: Cochran's Q Test for Presence of Heterogeneity
MetricValuedf (k-1)p-valueConclusion
Cochran's Q35.4211< .001HETEROGENEITY PRESENT
Note. Null Hypothesis: All studies are estimating the same true effect size.
p < .001Powerful Rejection of Unity. The studies in this meta-analysis are clearly not measuring the same thing (e.g., different doses, different patient types).
Header glossary

The Inconsistency Signal. A high Q value means the studies are 'fighting' each other—one says it works, one says it doesn't, beyond random error.

The Identity Probability. If p < .05, we reject the idea that all studies share a single 'Common' effect size.

11Algorithmic Logic

Command Center

Syntax libraries and function parameters for executing calculations in stats packages.

Code is the modern laboratory. Clean execution ensures reproducible discovery.
Execution Engine
# 1. Extract Q from Meta object
print(meta_model$Q)

# 2. Extract with full stats
summary(meta_model)
Library stack
R
metametafor
Python
statsmodels
Elite Forensic Strike

Cochran's Q is notoriously underpowered for small meta-analyses (k < 10). If you have few studies, use a threshold of p < .10 to detect heterogeneity, rather than the standard .05.

# Execute Comprehensive Heterogeneity Audit
# Using metafor to extract Q, I2, and Tau2 simultaneously
res <- metafor::rma(yi = yi, vi = vi, data = df)
print(res)
12The Over-adjustment Trap

Common Mistakes

Analytical caveats and corrections to maintain modeling integrity.

Wisdom is learning from the failures of others. Anticipate the error before it occurs.
Defensive Logic
Why it's wrong
Heterogeneity tests require more liberal α threshold (.10 instead of standard .05) because consequences of Type II error (failing to detect real heterogeneity) are more serious than Type I error. Missing true heterogeneity leads to inappropriate fixed-effect model, overly narrow CIs, and false precision. Q test already has low power with typical k, so α = .05 exacerbates problem.
The correction
ALWAYS use α = .10 for Cochran's Q test and other heterogeneity assessments. Report: 'Q(14) = 18.5, p = .18 → non-significant at α = .10.' This increases power to detect meaningful heterogeneity. Some meta-analysts recommend α = .10 only for heterogeneity tests while maintaining α = .05 for effect estimates.
Why it's wrong
Q test has very low power with k < 10 (only 10-30% power to detect moderate I² = 30-50%). Non-significant Q doesn't prove homogeneity—likely reflects insufficient power. Example: True I² = 50% but k = 7 yields p = .25 (non-significant) in 60% of simulations. Concluding 'no heterogeneity' is erroneous; heterogeneity exists but wasn't detected.
The correction
With k < 10, acknowledge low power explicitly: 'Given k = 7, the Q test has limited power to detect heterogeneity. Non-significant result (p = .25) should not be interpreted as evidence of homogeneity.' Default to random-effects model for conservative estimation. Rely more on descriptive measures (I², τ²) than hypothesis test. Consider narrative synthesis if k < 5.
Why it's wrong
With k > 20, Q test becomes overpowered—detects statistically significant heterogeneity even when trivial (I² < 25%). Example: k = 50, I² = 20%, Q(49) = 61.25, p = .11 is borderline; with k = 60, same I² yields p = .05 (significant). Statistical significance doesn't equal practical importance. Switching to random-effects based solely on significant Q with large k may be unnecessary if heterogeneity is minimal.
The correction
With k > 20, don't rely on Q p-value alone for model selection. Check I² and τ² for magnitude: if I² < 25% and τ² ≈ 0, heterogeneity is trivial despite significant Q. Report: 'Q(59) = 74.5, p = .08 is significant, but I² = 19% indicates low heterogeneity. Fixed-effect model may be appropriate, though random-effects used for generalization.' Always report both statistical test (Q) and effect size (I²).
Why it's wrong
Q is a test statistic (like F or t), NOT an effect size. Q magnitude depends on k (number of studies) and study precision (sample sizes), making it non-comparable across meta-analyses. Large Q could reflect many studies (large k) with low heterogeneity OR few studies with high heterogeneity. Example: k = 50, I² = 30%, Q = 71 vs. k = 10, I² = 60%, Q = 25—larger Q but smaller heterogeneity.
The correction
Use I² or τ² to quantify heterogeneity magnitude, NOT Q. I² is percentage (0-100%, interpretable across meta-analyses). τ² is variance in original units (interpretable as SD of true effects). Report: 'I² = 52% indicates moderate to substantial heterogeneity' NOT 'Q = 29 indicates heterogeneity.' Q provides p-value; I² and τ² provide magnitude.
Why it's wrong
Q, I², and τ² answer different questions—all needed for complete heterogeneity assessment. Q: Is there ANY heterogeneity? (hypothesis test). I²: HOW MUCH heterogeneity as percentage? (magnitude). τ²: HOW MUCH heterogeneity in original units? (practical interpretation). Reporting only one (e.g., just I²) omits statistical test; reporting only Q omits magnitude. Incomplete reporting prevents readers from judging heterogeneity appropriately.
The correction
ALWAYS report all three: (1) Q(df) with p-value for statistical test; (2) I² with percentage and interpretation category (low/moderate/substantial/considerable); (3) τ² and τ for magnitude in original metric. Example: 'Heterogeneity was moderate: Q(14) = 29.17, p = .010; I² = 52%; τ² = 0.042, τ = 0.20.' This provides complete picture: significant test, moderate magnitude, practical variability.
Why it's wrong
Q tests HETEROGENEITY (variation among studies), NOT treatment effect (pooled estimate ≠ 0). Q answers: 'Do studies vary?' NOT 'Is treatment effective?' Confusing these: significant Q means heterogeneity exists, NOT that treatment works. Conversely, non-significant Q doesn't mean treatment ineffective. Treatment effect tested via pooled estimate CI or Z-test (θ̂ / SE).
The correction
Use Q ONLY for heterogeneity assessment. For treatment effect: use pooled estimate CI and p-value from Z-test (θ̂ / SE). Report separately: 'Treatment effect: g = 0.68, 95% CI [0.54, 0.82], p < .001 (significant effect). Heterogeneity: Q(14) = 29.17, p = .010 (significant variation across studies).' Never say 'Q significant so treatment works'—that's wrong.
Why it's wrong
While random-effects is often appropriate (allows generalization), blindly using it without assessing heterogeneity misses opportunities to understand data. If true homogeneity (τ² = 0), random-effects reduces to fixed-effect but with wider CIs (loss of precision). Moreover, ignoring heterogeneity assessment prevents identifying moderators or problematic studies (outliers, errors). Meta-analysis isn't just pooling—it's understanding variation.
The correction
ALWAYS conduct and report heterogeneity assessment (Q, I², τ²) regardless of model choice. If Q non-significant and I² ≈ 0, consider fixed-effect but justify: 'Q(14) = 10.2, p = .75, I² = 0% suggests homogeneity; fixed-effect model used.' If defaulting to random-effects despite low heterogeneity, justify: 'Despite low heterogeneity (I² = 15%), random-effects model used for generalization beyond observed studies.' Transparency in model selection is critical.
Why it's wrong
Q test can be non-significant (p ≥ .10) while I² shows meaningful heterogeneity (e.g., I² = 40%), especially with small k (low power). Example: k = 8, true I² = 50% yields Q(7) = 14, p = .051 (non-significant at .05 but borderline at .10). Concluding 'no heterogeneity' based on p-value alone ignores I² = 50% (substantial). I² estimates magnitude regardless of statistical significance.
The correction
ALWAYS report and interpret I² alongside Q p-value. If Q non-significant but I² > 25%, acknowledge: 'Q test non-significant (p = .12) but I² = 42% suggests moderate heterogeneity, possibly due to low power (k = 9). Random-effects model used given uncertainty.' I² provides magnitude; Q provides statistical test. Both needed for complete assessment.
Why it's wrong
Q is a test statistic (yields p-value), NOT effect size—it tests IF heterogeneity exists. I² is an effect size (percentage), NOT a test statistic—it quantifies HOW MUCH heterogeneity. Confusing their roles: saying 'I² significant' (wrong—I² has no p-value in standard analysis) or 'Q = 30 indicates large heterogeneity' (wrong—Q is sample-size dependent). Each serves distinct purpose.
The correction
Use Q for hypothesis testing (H₀: τ² = 0) via p-value (α = .10). Use I² for magnitude estimation (0-100% scale with benchmarks: <25% low, 25-50% moderate, etc.). Report: 'Q(14) = 29.17, p = .010 → significant heterogeneity (reject H₀). I² = 52% → moderate to substantial magnitude.' Never report 'I² p-value' (doesn't exist in standard framework) or interpret Q magnitude without context of k.
Why it's wrong
Q test assumes independent effect sizes. Multiple outcomes from same study (e.g., depression + anxiety from one RCT) are correlated. Treating as independent: (1) Inflates k artificially (claims more studies than actually exist); (2) Violates statistical assumptions, biasing Q downward (underestimates heterogeneity); (3) Overweights studies contributing multiple effects. Example: 10 studies, 3 with 2 outcomes each → k = 13 claimed, but only 10 independent samples.
The correction
With multiple outcomes per study: (1) Select ONE primary outcome per study (prespecified); (2) Average effect sizes within study (simple if correlations unknown); (3) Use robust variance estimation (RVE) accounting for clustering (metafor::robust() in R); (4) Conduct multilevel/three-level meta-analysis (outcomes nested in studies). Never treat dependent effects as independent. Report: 'Three studies reported multiple outcomes; effect sizes averaged within study for independence.'
13Academic Lineage

References

Scholarly lineage and citation keys grounding the statistical framework.

We stand on the shoulders of giants. Honor the source of the method.
Academic Lineage
[1]
Cochran, W. G. (1954). The combination of estimates from different experiments. Biometrics, 10(1), 101-129.
Original paper introducing Cochran's Q test for assessing heterogeneity in meta-analysis. Foundational work establishing χ² distribution for Q under homogeneity hypothesis.
doi: 10.2307/3001666
[2]
Higgins, J. P., & Thompson, S. G. (2002). Quantifying heterogeneity in a meta-analysis. Statistics in Medicine, 21(11), 1539-1558.
Introduced I² statistic as complement to Q, providing interpretable measure of heterogeneity magnitude. Established I² benchmarks (<25% low, 25-50% moderate, 50-75% substantial, >75% considerable) widely used today.
doi: 10.1002/sim.1186
[3]
Huedo-Medina, T. B., Sánchez-Meca, J., Marín-Martínez, F., & Botella, J. (2006). Assessing heterogeneity in meta-analysis: Q statistic or I² index? Psychological Methods, 11(2), 193-206.
Comprehensive evaluation of Q test power across varying k and τ². Demonstrates low power with k < 10 and high power with k > 20. Recommends reporting both Q and I² for complete heterogeneity assessment.
doi: 10.1037/1082-989X.11.2.193
[4]
Borenstein, M., Hedges, L. V., Higgins, J. P., & Rothstein, H. R. (2009). Introduction to meta-analysis. John Wiley & Sons.
Comprehensive textbook covering Cochran's Q test, I² statistic, τ² estimation, and model selection (fixed vs. random effects). Essential reference for heterogeneity assessment methodology.
[5]
Deeks, J. J., Higgins, J. P., & Altman, D. G. (2001). Analysing data and undertaking meta-analyses. In J. P. Higgins & S. Green (Eds.), Cochrane handbook for systematic reviews of interventions. Cochrane Collaboration.
Practical guidance on evaluating heterogeneity using Q test, I², and visual assessment. Emphasizes using α = .10 for heterogeneity tests and interpreting Q in context of study number and power.
[6]
Hardy, R. J., & Thompson, S. G. (1998). Detecting and describing heterogeneity in meta-analysis. Statistics in Medicine, 17(8), 841-856.
Systematic comparison of methods for detecting heterogeneity, including Q test, I², and graphical approaches. Discusses limitations of Q test (power issues, sample-size dependence) and recommends complementary measures.
doi: 10.1002/(SICI)1097-0258(19980430)17:8<841::AID-SIM781>3.0.CO;2-D
statminds · Cochran'sMind reference · v2.2 · updated 2026-01-1715 of 15 sections