P values in diagnostics: when small is not meaningful
Article overview
Overview
A low p value is not the finish line. Pair it with the size of the effect before making a claim.
A p value says how surprising the data are if there is no real difference. It does not say how large or useful the difference is. For a diagnostic device, look at sensitivity, specificity, how precise the estimate is, and whether the difference matters clinically.
Worked example. A test of mean heart-rate error against zero bias gives t = 2.3 with 98 degrees of freedom. The two-sided p is about 0.023, so it is statistically different from zero. The mean error is 0.6 bpm with a standard deviation of 2.5 bpm, a small effect. The intended-use tolerance was set at plus or minus 5 bpm. The result is statistically detectable and clinically trivial. Do not overstate it.
Common mistakes. Switching to a one-sided p value after seeing the direction. Reading p below 0.05 as proof of clinical benefit without the size of the effect. Running many tests without a plan.
Turn a test statistic into a p value: https://nexamedtech.com/tools/p-value