
Last modified: 4/17/25
ANESTHESIOLOGY
Multiple Comparisons: Should we worry about p-value adjustment?
The typical goal of p-value adjustment is to compensate for possible increases in type I errors (false
positive results) when multiple outcome measures are used.
Classicists believe that the chance of finding at least one test statistically significant due to
chance and incorrectly declaring a difference increases as the number of comparisons increases.
Rationalists object to this theory for two reasons:
1. P-value adjustments are calculated based on how many tests are to be considered, and
that number has been defined arbitrarily and variably.
2. P-value adjustments reduce the chance of making type I errors, but they increase the
chance of making type II errors or needing to increase the sample size.
Why were adjustments for multiple tests developed at all?
Such adjustments are correct in the original framework of statistical test theory proposed in
the 1920’s. This theory was intended to aid decisions in repetitive situations.
For which situations do statistical p-value adjustments make sense?
1. The universal hypothesis is occasionally of interest. For example, to verify that a disease in not
associated with an HLA phenotype, we may compare available HLA antigens (perhaps 40) in a case
control study. If no association existed, at least one test would be significant with a probability
0.87, and p-value adjustment would protect against making excessive claims.
2. The same test is repeated in many subsamples (e.g. analyses conducted without an a priori
hypothesis that the primary association should differ between subgroups). Note: this is
reminiscent of repeated sampling of the same lot (Tukey, Bland and Altman’s justification)
3. Searching for significant associations without pre-established hypotheses.
Summary and recommendations:
In summary, p-value adjustments have limited application in medical research and should not be
used when assessing evidence about specific hypotheses. Readers should balance a study’s
statistical significance with the magnitude of the effect and the quality of the study design and
compare study results with findings from other studies. Researchers facing multiple outcome
measures may want to select a primary outcome measure or use global assessment measure,
rather than adjusting the p-value. If adjustment is still needed, methods such as Benjamini-
Hochberg can be used for controlling the false discovery rate.