
Last modified: 11/3/25
ANESTHESIOLOGY
p-values and Standardized Differences
What is a p-value?
According to the American Statistical Association (ASA)’s 2016 statement on p-values, “a p-value
is the probability under a specified statistical model that a statistical summary of the data (e.g., the
sample mean difference between two compared groups) would be equal to or more extreme than
its observed value.”
1
In the context of statistical hypothesis testing, “more extreme” refers to
values more compatible with the alternative hypothesis.
1,2
In the statement, the ASA provided six principles related to p-values
1
:
1. P-values can indicate how incompatible the data are with a specified statistical model…
2. P-values do not measure the probability that the studied hypothesis is true, or the probability that
the data were produced by random chance alone…
3. Scientific conclusions and business or policy decisions should not be based only on whether a p-
value passes a specific threshold…
4. Proper inference requires full reporting and transparency…
5. A p-value, or statistical significance, does not measure the size of an effect or the importance of a
result…
6. By itself, a p-value does not provide a good measure of evidence regarding a model or hypothesis.
What is a standardized difference?
The standardized difference is used to assess balance between groups.
3
Standardized
differences are a measure of effect size that are not a function of sample size.
3
This makes them
appropriate to use even for large samples that result in p-values indicating statistical significance
despite small effect sizes (that may also lack clinical significance).
3,4
Standardized differences can be calculated to compare continuous variables (variables
summarized using means and standard deviations), binary variables (variables with two levels,
often yes/no variables), and categorical variables (variables with more than two levels summarized
by counts and percentages).
3,5
One widely used type of standardized difference is Cohen’s d.
5,6
Absolute standardized differences values of 0.2-0.5 are considered small, values of 0.5-0.8 are
considered medium, and values > 0.8 are considered large.
5
The rule-of-thumb for interpreting standardized differences is that a standardized difference
greater than 0.2 suggests the variable is unbalanced between the groups being compared.
5