Statistical Significance and Clinical Significance
Statistical significance and clinical significance answer different questions. A statistical test examines data under a specified model. Clinical interpretation asks what an observed difference means for a person's health, functioning or valued activities. A result can be statistically significant but too small to matter in practice; an important possible benefit can also remain statistically uncertain.
Clear reporting keeps the size of change, its uncertainty, its relevance and its attribution separate. These distinctions apply to both formal research and the outcome claims made in routine therapy.
What a p-value does and does not mean
A p-value describes how incompatible the observed data are with a specified statistical model, commonly one including a null hypothesis of no difference, under the assumptions used to calculate it. It is not the probability that the null hypothesis is true, or the probability that an observed improvement was caused by chance.
The American Statistical Association states:
A p-value, or statistical significance, does not measure the size of an effect or the importance of a result.
— Wasserstein and Lazar (2016), principle 5. A threshold such as p < .05 is a decision convention, not a boundary separating important effects from unimportant ones or true explanations from false ones.1)
In a hypothetical large study, a treatment could produce a precisely estimated one-point difference on a scale where such a difference has little practical meaning. A small study could suggest a much larger benefit while remaining too imprecise to support a confident conclusion. The p-value alone cannot resolve either situation.
Effect size and uncertainty
An effect estimate describes the magnitude and direction of a difference or association. It may be expressed in the original units, as a standardised difference, as a risk ratio or in another suitable form. Original units are often especially useful when readers understand the scale and its relevance.
A confidence interval describes uncertainty under the statistical model. In the usual frequentist interpretation, a 95% confidence procedure would cover the true parameter in 95% of repeated applications under its assumptions. A particular observed interval is not a 95% probability statement about a fixed parameter. Greenland and colleagues explain why intervals and p-values require attention to assumptions, study design and potential bias.2)
For interpretation, ask which effects remain compatible with the estimate. An interval spanning appreciable benefit and appreciable harm is substantively different from a narrow interval around a negligible difference. Neither can be adequately summarised by saying that the study “found no effect”.
Group averages and individual outcomes
A group mean does not describe every participant. The same average improvement could arise from modest improvement in most people, large improvement in a minority, or improvement in some people combined with deterioration in others. Outcome reports should therefore make the distribution of experiences visible when the data permit it.
Likewise, a before-and-after reduction within a treated group is not the same quantity as a difference between treatment and comparison groups. If both groups improve by a similar amount, the within-group changes may be substantial while the estimated comparative benefit is small. The choice of comparison determines the question being answered.
Consider an invented example: Group A improves by eight points and Group B by six. The unadjusted difference in mean change is two points, not eight. Further analysis may be required to account for baseline differences, missing data and the trial design, but the example shows why the treatment group's change should not be presented as its entire added benefit over the comparator.
Reliable change and clinically significant change
Jacobson and Truax distinguish whether an individual's change is large enough to exceed expected measurement error from whether the person's post-treatment score has crossed a meaningful clinical boundary. Their framework combines reliable change with movement towards a functional reference distribution. Classification depends on an appropriate instrument, reliability estimate and reference data.3)
In a conventional form, the reliable change index is:
RCI = (post-treatment score − pre-treatment score) / Sdiff
Sdiff = SD × sqrt(2 × (1 − r))
Here SD is the relevant reference standard deviation and r is an appropriate reliability estimate. Under the method's assumptions, an absolute RCI exceeding approximately 1.96 is conventionally treated as reliable change at the 95% level. The sign indicates direction according to the scale. Instrument-specific guidance should take precedence over applying this formula mechanically.4)
A numerical example
All values in this example are invented and do not belong to a validated questionnaire. Suppose higher scores indicate greater difficulty. A person's score falls from 35 to 20; the relevant SD is 10 and reliability is .90.
The standard error of the difference is 10 × sqrt(2 × .10), approximately 4.47. The RCI is (20 − 35) / 4.47, approximately −3.35. Under the stated assumptions, the reduction exceeds the conventional reliable-change threshold.
Now suppose the chosen clinical boundary requires a score of 18 or below. The person has improved reliably but remains above that boundary. Calling this “recovery” would misstate the classification. Conversely, a person starting just above a cut-off could cross it with a very small change that does not exceed measurement error.
Neither classification tells us whether this person can now undertake an activity they value. That requires additional information. Nor does a reliable change establish that therapy caused it: reliable measurement and causal attribution are separate issues.
Minimally important differences
A minimally important difference concerns a change considered important in the relevant context. It is not automatically equal to a statistically detectable difference, a reliable-change threshold or movement across a diagnostic cut-off.
Revicki and colleagues recommend drawing on meaningful external anchors, such as patient or clinical judgements, supported by distribution-based information. Estimates depend on the population, instrument and purpose. A published threshold should therefore be checked for applicability rather than transferred automatically to a different group or a different use of the scale.5)
This distinction matters for individual care. A small score change might accompany a highly valued practical achievement, while a large score reduction might leave the person's main difficulty unresolved. These observations do not make scales useless; they show why scores and goals should be interpreted together.
A compact interpretation framework
| Question | Information needed | What it does not establish by itself |
|---|---|---|
| Did the measured value change? | Comparable observations before and after. | That the change exceeds measurement error. |
| Was change reliable? | Suitable reliability and reference information. | That the person has recovered or that therapy caused it. |
| Was change clinically meaningful? | Relevant thresholds, functioning and the person's priorities. | A specific mechanism of action. |
| Did treatment add benefit? | An appropriate comparison and credible causal design. | That every participant benefited. |
| Did benefit persist? | Follow-up with clear timing and adequate reporting. | Permanence beyond the period observed. |
“No significant difference” is not equivalence
A non-significant comparison may reflect a small difference, imprecise data or both. Demonstrating equivalence requires defining an acceptable range of differences and using a design and analysis capable of addressing that range. Lakens explains equivalence testing using prespecified bounds linked to the smallest effect considered meaningful.6)
For example, a trial comparing two therapies cannot establish that they are equally effective merely because p exceeds .05. If its confidence interval includes differences large enough to matter, the comparison remains unresolved. Conversely, a suitably precise estimate within justified equivalence bounds supports a more specific conclusion about similarity under the studied conditions.
Reporting therapeutic change responsibly
A useful report identifies the outcome, assessment times, sample size, comparison, effect estimate and uncertainty. Where individual classifications are used, it explains the instrument, reliability information, cut-off and treatment of missing observations. It also reports deterioration and unwanted effects when these were assessed, rather than presenting only improvement.
For IEMT, EMI and EMDR, a lower distress rating after a session may document an immediate change in that rating. It should not be automatically converted into a recovery rate, a comparative treatment effect or proof of a proposed memory mechanism. Each stronger claim requires the additional evidence appropriate to it.