What a P-Value Actually Answers
A p-value answers a narrow, specific question: if there were truly no effect, how likely would a result at least this extreme be, just from random chance? By convention, most fields treat p<0.05 as a threshold worth calling “statistically significant.” The Falutz et al. tesamorelin trial covered earlier reported visceral fat reductions at p<0.001 — meaning that result would be extremely unlikely to occur by chance alone if tesamorelin had no real effect. But a p-value says nothing about how large or clinically meaningful that effect actually is — that’s a separate question entirely.
Three Concepts That Aren’t the Same Thing
- Statistical significance (p-value): how unlikely the result is to be due to chance alone
- Effect size: how large the actual difference or change is — independent of statistical significance
- Confidence interval: a range of plausible values for the true effect, given the data — narrower generally means more precise
- Clinical / practical significance: whether the effect size is large enough to actually matter in practice
Why a Large Study Can Find “Significant” Small Effects
With a large enough sample size, even a very small, practically trivial effect can reach statistical significance — because a bigger sample makes it easier to distinguish a small real effect from pure chance. This is why reading past the p-value to the actual effect size matters: “statistically significant” tells you an effect probably exists, not that it’s large enough to matter. Conversely, a small study can fail to reach statistical significance even when a real, meaningful effect is present, simply because it didn’t have enough participants to detect it reliably.
✅ Quick Recap
- A p-value measures how likely a result is to be due to chance, not how large or important the effect is
- Effect size and confidence intervals give you the magnitude and precision a p-value alone doesn’t
- Large studies can find statistically significant effects that are too small to matter practically
- Small studies can miss real effects simply from having too few participants to detect them