Skip to content
Peptideworth
Menu

Evidence review

Why n=30 Is Not the Same as n=3,000

Enrolment size is not a quality badge — it is precision, and it is the ability to see rare harm. Worked on real peptide trials, with the arithmetic shown.

By Grant Delaney, Research Editor

Two things n buys, and one it doesn't

Precision. A bigger trial gives you a narrower range around the estimate, not a different estimate.

Visibility of rare events. A harm that happens in one person per thousand is invisible in thirty people, and nearly invisible in three hundred.

What n does not buy is validity. A large trial with the wrong comparator, or the wrong endpoint, is a precisely measured answer to a question no reader was asking.

The arithmetic, on a real trial

The published phase 3 of a thymosin beta-4 ophthalmic solution in neurotrophic keratopathy enrolled 18 patients. Complete healing at four weeks occurred in 6 of 10 treated patients and 1 of 8 on placebo, which the authors reported at p = 0.0656 and described as a strong efficacy trend rather than a significant difference1.

Take those two fractions and put a 95% interval around each.

6 of 10 is 60%. The 95% interval runs from about 31% to about 83%.

1 of 8 is 12.5%. The 95% interval runs from about 2% to about 47%.

Those two ranges overlap across a span from 31% to 47%. The point estimates look like a large difference. The intervals are consistent with a large difference, a small one, or none. That is not a criticism of the trial — it is what eighteen patients can resolve, and the authors said so.

Why small trials read as more impressive than they are

A meta-epidemiological study of critical care meta-analyses compared effect sizes from small and large trials within the same questions and found that small studies may overestimate effect sizes2.

The mechanism is not fraud. It is selection. A small trial that finds nothing is easy to leave in a drawer; a small trial that finds a large effect gets written up. Over time the small-trial literature is enriched for large effects, and the large-trial literature is not.

Which means: when a compound's case rests entirely on small studies, the expected direction of error is toward overstatement.

"No adverse events reported" at n=42

There is a shortcut for this, and it is worth memorising. If zero events are observed in n participants, the upper bound of the 95% interval is roughly 3 divided by n.

n = 18 — zero events is consistent with a true rate as high as about 17%.

n = 42 — zero events is consistent with a true rate as high as about 7%.

n = 1,961 — zero events is consistent with a true rate up to about 0.15%.

n = 17,604 — zero events is consistent with a true rate up to about 0.02%.

So a phase 1 safety trial in 42 healthy volunteers that reports no adverse events has ruled out a common harm and has said essentially nothing about an uncommon one. A ClinicalTrials.gov search for BPC-157 returned 3 studies in August 2026, the oldest of which is a 42-participant phase 1 safety and pharmacokinetics trial with no results posted3. Our BPC-157 page reads that file line by line.

The other end of the scale

STEP 1 randomised 1,961 adults to once-weekly semaglutide 2.4 mg or placebo over 68 weeks and reported a mean weight change of −14.9% against −2.4%4.

SELECT randomised 17,604 participants and ran for years, because its question was not weight — it was whether cardiovascular events happen less often5.

The jump from roughly two thousand to roughly eighteen thousand is not a jump in ambition. It is the price of measuring an outcome that is rare per person per year. If the event you care about happens to 2% of people annually, you need a great many person-years before the difference between two arms is readable.

Precision, quantified

For a continuous outcome such as percent weight change, the width of the interval shrinks with the square root of n — which is slower than most people expect.

Assume a standard deviation of 8 percentage points, purely as an illustration. At n = 30, the half-width of a 95% interval around the mean is about 2.9 points. At n = 300 it is about 0.9. At n = 3,000 it is about 0.3.

Going from 30 to 3,000 — a hundredfold increase in cost — buys a tenfold narrowing. That is the whole trade, and it is why a sponsor will not fund a 3,000-person trial to answer a question a 30-person trial could settle.

What to do with a small trial

Do not dismiss it. A small, well-conducted, blinded, placebo-controlled trial is worth more than a large open-label case series.

Do read the interval, not the point estimate. If the paper reports one, use it. If it does not, the fractions are usually enough to reconstruct it.

Do ask what the trial was powered for. A trial powered to detect a 20-point difference will report "no significant difference" for a 10-point one, and that is not evidence of no effect.

Do check whether it is the only trial. One small positive result is a hypothesis. Two independent ones start to be a finding. Where a compound sits on that scale is what our evidence grading is trying to capture.

The question that cuts through

Ask what would have to be true for this result to be a fluke — and then ask whether the trial was large enough to rule that out. For eighteen patients, quite a lot could be a fluke. For seventeen thousand, much less.

Frequently asked questions

Is a phase 3 trial always large?

No. Phase is a stage of development, not a size. The published phase 3 of a thymosin beta-4 ophthalmic solution in neurotrophic keratopathy enrolled 18 patients, and its primary healing comparison was reported at p = 0.0656 — described by the authors as a strong efficacy trend rather than a significant difference. STEP 1, also a phase 3, enrolled 1,961.

What does 'no adverse events were reported' mean in a small trial?

Much less than it sounds. With zero events observed in n participants, the upper bound of the 95% interval is roughly 3 divided by n. In 42 participants that upper bound is about 7%; in 18 participants it is about 17%. A small trial reporting no adverse events has ruled out a common harm, not an uncommon one.

Why are cardiovascular outcome trials so much bigger than weight-loss trials?

Because the outcome is rarer. Weight can be measured in every participant at every visit; a heart attack happens to a small percentage of people per year. SELECT randomised 17,604 participants to answer that question, against 1,961 in STEP 1, which measured weight.

Do small trials tend to exaggerate?

A meta-epidemiological study of critical care meta-analyses found that small studies may overestimate effect sizes relative to larger trials asking the same question. The likeliest mechanism is selective publication rather than misconduct — small null results are easier to leave unpublished than small positive ones.

References

  1. Sosne G et al (2022). 0.1% RGN-259 (Thymosin ß4) Ophthalmic Solution Promotes Healing and Improves Comfort in Neurotrophic Keratopathy Patients in a Randomized, Placebo-Controlled, Double-Masked Phase III Clinical Trial. Int J Mol Sci. https://pubmed.ncbi.nlm.nih.gov/36613994/
  2. Zhang Z; Xu X; Ni H (2013). Small studies may overestimate the effect sizes in critical care meta-analyses: a meta-epidemiological study. Crit Care. https://pubmed.ncbi.nlm.nih.gov/23302257/
  3. U.S. National Library of Medicine (2026). NCT02637284 — PCO-02: Safety and Pharmacokinetics Trial. ClinicalTrials.gov. https://clinicaltrials.gov/study/NCT02637284
  4. Wilding JPH et al (2021). Once-Weekly Semaglutide in Adults with Overweight or Obesity. N Engl J Med. https://pubmed.ncbi.nlm.nih.gov/33567185/
  5. Lincoff AM et al (2023). Semaglutide and Cardiovascular Outcomes in Obesity without Diabetes. N Engl J Med. https://pubmed.ncbi.nlm.nih.gov/37952131/

Medical disclaimer: This content is for general educational purposes only and is not medical advice, diagnosis, or treatment. Always consult a licensed healthcare professional before starting, stopping, or changing any treatment.

Continue reading