What the P-Value Actually Says

Statistical significance is determined by a probability value — the p-value. If a study finds that participants in the active group scored lower on an anxiety scale than participants in the placebo group, the p-value expresses how likely that difference would be to occur by chance if the intervention had no real effect. The conventional threshold is p < 0.05: if the probability of observing the difference by chance alone is less than 5%, the result is labeled statistically significant.

What the p-value does not tell you is how large the difference was, whether it was large enough to matter in practice, or whether it would appear again in a different sample. Statistical significance is a statement about the reliability of a signal — not a statement about the size or importance of what that signal represents.

Three Concepts That Don't Mean the Same Thing

All three questions are worth asking separately. A finding that answers "yes" to the first but "no" to the third is genuinely less important than a headline will suggest. A finding that answers "yes" to all three and has been replicated across multiple studies is genuinely more important than a single trial's press release will convey.

Why Sample Size Creates Paradoxes

The relationship between sample size and statistical significance produces counterintuitive results that matter specifically for cannabinoid research, where most trials are small.

The Sample Size Problem — Both Directions

Large Sample

Even trivially small differences can reach statistical significance. A study of 10,000 participants might find a statistically significant 0.3-point reduction on a 21-point anxiety scale — real, unlikely to be chance, and probably not experientially meaningful.

Small Sample

Even meaningful effects may fail to reach statistical significance because natural variability across participants overwhelms the signal. A real effect in 30 participants may look like noise. This is why small trials that do show significant effects are notable — they clear a higher bar.

Most cannabinoid anxiety and stress trials have enrolled under 50 participants. The Cuttler et al. (2024) CBG trial enrolled 34. When a small crossover trial achieves statistically significant effects, the finding is meaningful precisely because the design and sample size make significance harder to reach — not in spite of the small sample. The appropriate response is neither to dismiss the finding because of the sample size nor to treat it as definitive because it crossed the significance threshold.

Confidence Intervals: More Information Than P-Values Alone

Many published studies report confidence intervals alongside p-values — and the confidence interval is often more informative. A 95% confidence interval expresses the range within which the true effect probably falls if the study were replicated many times. A narrow confidence interval around a meaningful effect size suggests a precise and reliable finding. A wide confidence interval — even around a statistically significant p-value — signals uncertainty about how large the true effect actually is.

When the confidence interval for an effect includes zero, the true effect may be zero — which means the statistically significant p-value may be reflecting a boundary case rather than a robust signal. Reading confidence intervals alongside p-values gives a richer picture of what the data actually show.

The Replication Crisis Context

The p < 0.05 threshold is a convention established decades before the research community understood how many published significant findings would fail to replicate in subsequent studies. In fields where samples are small, outcome measures are subjective, and researchers have latitude in how they analyze data, false positives are a documented problem. This is not unique to cannabinoid research — it affects psychology, nutrition, and behavioral medicine broadly. It means that a single statistically significant finding, regardless of how well the study was designed, is a signal worth investigating rather than a conclusion worth announcing.

How this applies to the archive's evidence tiers

The archive's Tier 1 designation — applied to human randomized controlled trials — reflects study design quality, not certainty of the finding. A Tier 1 result in a small, single-site, acute trial is meaningfully different from a Tier 1 result that has been replicated across multiple independent studies in diverse populations. The tier system captures the level of evidence, not the degree of confidence. This article is the explanation of why those are different things.