How to calculate Cohen's d and Hedges' g
Article overview
Overview
Compute Cohen's d and Hedges' g from two group means and SDs, or one-sample d, and read the conventional bands without treating them as a clinical rule.
A statistically significant difference can still be too small to matter. That is why clinical and human-factors reports for medical devices should show an effect size next to a p value. Cohen's d and Hedges' g are standardized mean differences: they put the gap between two means into standard-deviation units so you can compare studies that used different scales.
Use the free effect size calculator when you already have means, standard deviations, and sample sizes. This article shows the formulas the page uses, works an example you can reproduce with the default inputs, and explains why the small / medium / large bands are conventions rather than a clinical cutoff.
It is written for clinical scientists, biostatistics partners, and regulatory writers who have to describe a bench or clinical comparison without over-claiming.
Why a p value is not the effect
A p value answers how surprising the data are under a stated null. It does not say how large the difference is, whether the difference would matter to a clinician, or whether the study was large enough to detect the difference you care about.
We already walked through that trap in P values in diagnostics: when small is not meaningful. The companion question is the one this page answers: given two means and their scatter, how large is the standardized gap?
Effect size also feeds sample-size planning. If you cannot say what standardized difference is worth detecting, you cannot defend n. See Getting n right: two-arm sample size basics once you have a target effect.
Nothing in FDA device regulations or ISO 14155 requires Cohen's d specifically. What reviewers do expect is a pre-specified analysis and an estimate they can interpret. d and g are one transparent way to provide that estimate for mean differences.
The two-group formulas
For two independent groups, Cohen's d is:
d = (mean1 − mean2) / s_pooled
The pooled standard deviation is:
s_pooled = sqrt( ((n1 − 1) × SD1² + (n2 − 1) × SD2²) / (n1 + n2 − 2) )
That is the common two-sample form that uses the same denominator degrees of freedom as a Student t test, df = n1 + n2 − 2.
Hedges' g applies a small-sample correction:
g = d × (1 − 3 / (4 × df − 1))
Cohen's d is slightly upward-biased in small samples. g reduces that bias. With large n the two numbers are nearly the same. Report both when n is modest, and say which one you used in the statistical analysis plan.
Each n must be a whole number of at least 2, and each SD must be above 0. If either SD is 0, you do not have a defined standardized difference with this formula.
The sign follows the mean order you entered. If group 1 is the new algorithm and group 2 is the predicate, a negative d means the new algorithm’s mean is lower. That can be desirable for error and undesirable for a benefit score. Write the direction in words. Do not leave a reviewer to guess.
Worked example: the calculator defaults
Open the effect size calculator and leave the two-group defaults:
- Group 1: mean 12.4, SD 3.1, n = 30
- Group 2: mean 10.9, SD 3.6, n = 28
Degrees of freedom are 30 + 28 − 2 = 56.
The pooled variance is (29 × 3.1² + 27 × 3.6²) / 56 = (29 × 9.61 + 27 × 12.96) / 56 = 628.61 / 56 = 11.225. The pooled SD is sqrt(11.225) ≈ 3.350.
Cohen's d is (12.4 − 10.9) / 3.350 ≈ 0.448.
Hedges' g is 0.448 × (1 − 3 / (4 × 56 − 1)) = 0.448 × (1 − 3/223) ≈ 0.448 × 0.9865 ≈ 0.442.
The page labels |d| = 0.448 as Small, because its bands are: below 0.2 very small, 0.2 to under 0.5 small, 0.5 to under 0.8 medium, and 0.8 or more large. Those cut points follow Cohen’s conventional 0.2 / 0.5 / 0.8 values. They are not a clinical decision rule.
Suppose this comparison was mean absolute heart-rate error, in beats per minute, for a wearable versus a reference. A standardized difference of about 0.45 says the group means are less than half a pooled SD apart. Whether that matters depends on the pre-specified tolerance, the intended use, and the risk of being wrong — not on the word “small.”
Worked example: one-sample d
Switch the calculator to one sample. The defaults are sample mean 5.4, SD 1.2, n = 25, hypothesized mean 5.
One-sample Cohen's d is (mean − μ0) / SD = (5.4 − 5) / 1.2 = 0.333. The page does not apply a Hedges correction in this mode.
|d| = 0.333 is still in the Small band. A one-sample test against a specification (for example, a mean bias of 0, or a claimed operating time of 5 days) should still be reported with the raw mean, the SD, a confidence interval, and the specification. The standardized value is a companion, not a replacement.
Use one-sample d when you have a single series and a number you committed to in a protocol. Use two-group d when you have a comparator arm.
How to read Cohen's bands
Cohen proposed 0.2, 0.5, and 0.8 as rough conventions for the behavioral sciences when no better context exists. Medical device outcomes usually have better context: a minimal clinically important difference, a recognized performance band, or a risk-based tolerance in the verification plan.
Do not write “the effect was large, therefore the device is effective.” A large d on a surrogate that is not clinically meaningful is still a large d on a surrogate. A small d on a safety endpoint can still be important.
Do not hide a sign. If the new method is worse, say so.
Do not mix unstandardized and standardized claims. If the protocol said you care about a 2 mmHg bias, report mmHg first. Then, if useful, report d so a reader can compare studies.
What to put in a protocol or report
State the formula. Two-group pooled d and Hedges' g as above are enough for most two-arm mean comparisons. If you use a different denominator (for example, the control-group SD only, sometimes called Glass’s Δ), name it.
State the group order and the units of the raw means.
State whether you will interpret d against a pre-specified clinically important difference, against Cohen’s bands, or not at all. If you do not need the bands, do not display them as if they were a regulatory threshold.
Pair the effect size with a confidence interval for the mean difference when you have the data. The confidence interval page can sketch a t interval for a mean. Pair it with a p value only after the analysis is defined; the p value calculator turns a z, t, or chi-square statistic into one-sided and two-sided p values.
If you are still sizing the study, take the standardized effect you actually care about to the clinical sample-size calculator. Sizing on an optimistic d from a pilot is a common way to underpower a pivotal comparison.
Use the effect size calculator
Enter the two group means, SDs, and ns, or the one-sample mean, SD, n, and hypothesized value, in the effect size calculator. Reproduce the default two-group result (d ≈ 0.448, g ≈ 0.442, Small band) once so you trust the page, then replace the numbers with your own.
Report the raw difference, the standardized difference, and the clinical tolerance you pre-specified. Use P values in diagnostics: when small is not meaningful when you need to keep significance and importance from collapsing into one sentence.