Last modified: 11/3/25
ANESTHESIOLOGY
p-values and Standardized Differences
What is a p-value?
According to the American Statistical Association (ASA)’s 2016 statement on p-values, “a p-value
is the probability under a specified statistical model that a statistical summary of the data (e.g., the
sample mean difference between two compared groups) would be equal to or more extreme than
its observed value.”
1
In the context of statistical hypothesis testing, “more extreme” refers to
values more compatible with the alternative hypothesis.
1,2
In the statement, the ASA provided six principles related to p-values
1
:
1. P-values can indicate how incompatible the data are with a specified statistical model
2. P-values do not measure the probability that the studied hypothesis is true, or the probability that
the data were produced by random chance alone
3. Scientific conclusions and business or policy decisions should not be based only on whether a p-
value passes a specific threshold
4. Proper inference requires full reporting and transparency
5. A p-value, or statistical significance, does not measure the size of an effect or the importance of a
result
6. By itself, a p-value does not provide a good measure of evidence regarding a model or hypothesis.
What is a standardized difference?
The standardized difference is used to assess balance between groups.
3
Standardized
differences are a measure of effect size that are not a function of sample size.
3
This makes them
appropriate to use even for large samples that result in p-values indicating statistical significance
despite small effect sizes (that may also lack clinical significance).
3,4
Standardized differences can be calculated to compare continuous variables (variables
summarized using means and standard deviations), binary variables (variables with two levels,
often yes/no variables), and categorical variables (variables with more than two levels summarized
by counts and percentages).
3,5
One widely used type of standardized difference is Cohen’s d.
5,6
Absolute standardized differences values of 0.2-0.5 are considered small, values of 0.5-0.8 are
considered medium, and values > 0.8 are considered large.
5
The rule-of-thumb for interpreting standardized differences is that a standardized difference
greater than 0.2 suggests the variable is unbalanced between the groups being compared.
5
Last modified: 11/3/25
References:
1. Wasserstein RL, Lazar NA. The ASA's statement on P-values: Context, process, and purpose.
Am Stat. 2016 70:12933. Available
at: https://amstat.tandfonline.com/doi/full/10.1080/00031305.2016.1154108#
2. Bonovas S, Piovani D. On p-Values and Statistical Significance. Journal of Clinical Medicine.
2023; 12(3):900. https://doi.org/10.3390/jcm12030900
3. Yang D, Dalton JE. A unified approach to measuring the effect size between two groups using
SAS®. Proceedings of the SAS® Global Forum 2012 Conference. Cary, NC: SAS Institute Inc.
Available at: https://support.sas.com/resources/papers/proceedings12/335-2012.pdf
4. Sullivan GM, Feinn R. Using Effect Size-or Why the P Value Is Not Enough. J Grad Med Educ.
2012 Sep;4(3):279-82. doi: 10.4300/JGME-D-12-00156.1. PMID: 23997866; PMCID:
PMC3444174.
5. Huntington-Klein N. The Effect: An Introduction to Research Design and Causality. CRC Press;
2021.
6. Lakens D. Calculating and reporting effect sizes to facilitate cumulative science: a practical
primer for t-tests and ANOVAs. Front Psychol. 2013;4:863. doi:10.3389/fpsyg.2013.00863.
Additional Information:
1. Greenland S, Senn SJ, Rothman KJ, Carlin JB, Poole C, Goodman SN, Altman DG. Statistical
tests, P values, confidence intervals, and power: a guide to misinterpretations. Eur J Epidemiol.
2016 Apr;31(4):337-50. doi: 10.1007/s10654-016-0149-3. Epub 2016 May 21. PMID: 27209009;
PMCID: PMC4877414.