Understanding — How to Read Research
What Statistical Significance Means — and What It Does Not
One of the most misunderstood concepts in science writing — what this phrase actually means, and what it tells you (and doesn't tell you) about a study's findings.

When a cannabinoid study reports "statistically significant reductions in anxiety," the phrase carries a specific technical meaning that most headlines don't explain. Statistically significant does not mean large. It does not mean clinically important. It does not mean the finding will replicate. It means one precise thing, and understanding that one thing changes how every study in this archive should be read.
What the P-Value Actually Says
Statistical significance is determined by a probability value — the p-value. If a study finds that participants in the active group scored lower on an anxiety scale than participants in the placebo group, the p-value expresses how likely that difference would be to occur by chance if the intervention had no real effect. The conventional threshold is p < 0.05: if the probability of observing the difference by chance alone is less than 5%, the result is labeled statistically significant.
What the p-value does not tell you is how large the difference was, whether it was large enough to matter in practice, or whether it would appear again in a different sample. Statistical significance is a statement about the reliability of a signal — not a statement about the size or importance of what that signal represents.
Three Concepts That Don't Mean the Same Thing
Concept 1
Statistical Significance
The observed difference is unlikely to reflect chance variation. Expressed as a p-value below a threshold — conventionally p < 0.05. Tells you the signal is probably real.
Question it answers: Is the effect likely real?
Concept 2
Effect Size
The magnitude of the observed difference — how large the change actually was. Commonly reported as Cohen's d or a similar metric. A large effect size in a small sample may not reach statistical significance. A tiny effect size in a large sample easily does.
Question it answers: How big is the effect?
Concept 3
Clinical Significance
Whether the magnitude of change is large enough to make a meaningful difference in how someone functions or experiences their condition. A 2-point shift on a 100-point anxiety scale may be statistically significant and experientially negligible.
Question it answers: Does the effect matter?
All three questions are worth asking separately. A finding that answers "yes" to the first but "no" to the third is genuinely less important than a headline will suggest. A finding that answers "yes" to all three and has been replicated across multiple studies is genuinely more important than a single trial's press release will convey.
Why Sample Size Creates Paradoxes
The relationship between sample size and statistical significance produces counterintuitive results that matter specifically for cannabinoid research, where most trials are small.
The Sample Size Problem — Both Directions
Large Sample
Even trivially small differences can reach statistical significance. A study of 10,000 participants might find a statistically significant 0.3-point reduction on a 21-point anxiety scale — real, unlikely to be chance, and probably not experientially meaningful.
Small Sample
Even meaningful effects may fail to reach statistical significance because natural variability across participants overwhelms the signal. A real effect in 30 participants may look like noise. This is why small trials that do show significant effects are notable — they clear a higher bar.
Most cannabinoid anxiety and stress trials have enrolled under 50 participants. The Cuttler et al. (2024) CBG trial enrolled 34. When a small crossover trial achieves statistically significant effects, the finding is meaningful precisely because the design and sample size make significance harder to reach — not in spite of the small sample. The appropriate response is neither to dismiss the finding because of the sample size nor to treat it as definitive because it crossed the significance threshold.
Confidence Intervals: More Information Than P-Values Alone
Many published studies report confidence intervals alongside p-values — and the confidence interval is often more informative. A 95% confidence interval expresses the range within which the true effect probably falls if the study were replicated many times. A narrow confidence interval around a meaningful effect size suggests a precise and reliable finding. A wide confidence interval — even around a statistically significant p-value — signals uncertainty about how large the true effect actually is.
When the confidence interval for an effect includes zero, the true effect may be zero — which means the statistically significant p-value may be reflecting a boundary case rather than a robust signal. Reading confidence intervals alongside p-values gives a richer picture of what the data actually show.
The Replication Crisis Context
The p < 0.05 threshold is a convention established decades before the research community understood how many published significant findings would fail to replicate in subsequent studies. In fields where samples are small, outcome measures are subjective, and researchers have latitude in how they analyze data, false positives are a documented problem. This is not unique to cannabinoid research — it affects psychology, nutrition, and behavioral medicine broadly. It means that a single statistically significant finding, regardless of how well the study was designed, is a signal worth investigating rather than a conclusion worth announcing.
How this applies to the archive's evidence tiers
The archive's Tier 1 designation — applied to human randomized controlled trials — reflects study design quality, not certainty of the finding. A Tier 1 result in a small, single-site, acute trial is meaningfully different from a Tier 1 result that has been replicated across multiple independent studies in diverse populations. The tier system captures the level of evidence, not the degree of confidence. This article is the explanation of why those are different things.
Statistical significance is a threshold, not a verdict. It tells you that something probably happened — not that it was large, not that it matters, not that it will happen again in a different sample. Those are three separate questions requiring three separate evaluations, and most popular science reporting collapses all of them into the single phrase "statistically significant results."
Reading the studies in this archive with these distinctions held clearly produces a more accurate picture than any headline summary can offer — and it also produces a more honest appreciation of what the early cannabinoid research has actually found, which is a consistent and promising signal in a field that has not yet delivered the scale of evidence it is building toward.
These statements have not been evaluated by the Food and Drug Administration. J.P. Hemp Company products are not intended to diagnose, treat, cure, or prevent any disease.