Evidence review
Why n=30 Is Not the Same as n=3,000
Enrolment size is not a quality badge — it is precision, and it is the ability to see rare harm. Worked on real peptide trials, with the arithmetic shown.
Two things n buys, and one it doesn't
Precision. A bigger trial gives you a narrower range around the estimate, not a different estimate.
Visibility of rare events. A harm that happens in one person per thousand is invisible in thirty people, and nearly invisible in three hundred.
What n does not buy is validity. A large trial with the wrong comparator, or the wrong endpoint, is a precisely measured answer to a question no reader was asking.
The arithmetic, on a real trial
The published phase 3 of a thymosin beta-4 ophthalmic solution in neurotrophic keratopathy enrolled 18 patients. Complete healing at four weeks occurred in 6 of 10 treated patients and 1 of 8 on placebo, which the authors reported at p = 0.0656 and described as a strong efficacy trend rather than a significant difference1.
Take those two fractions and put a 95% interval around each.
6 of 10 is 60%. The 95% interval runs from about 31% to about 83%.
1 of 8 is 12.5%. The 95% interval runs from about 2% to about 47%.
Those two ranges overlap across a span from 31% to 47%. The point estimates look like a large difference. The intervals are consistent with a large difference, a small one, or none. That is not a criticism of the trial — it is what eighteen patients can resolve, and the authors said so.
Why small trials read as more impressive than they are
A meta-epidemiological study of critical care meta-analyses compared effect sizes from small and large trials within the same questions and found that small studies may overestimate effect sizes2.
The mechanism is not fraud. It is selection. A small trial that finds nothing is easy to leave in a drawer; a small trial that finds a large effect gets written up. Over time the small-trial literature is enriched for large effects, and the large-trial literature is not.
Which means: when a compound's case rests entirely on small studies, the expected direction of error is toward overstatement.
"No adverse events reported" at n=42
There is a shortcut for this, and it is worth memorising. If zero events are observed in n participants, the upper bound of the 95% interval is roughly 3 divided by n.
n = 18 — zero events is consistent with a true rate as high as about 17%.
n = 42 — zero events is consistent with a true rate as high as about 7%.
n = 1,961 — zero events is consistent with a true rate up to about 0.15%.
n = 17,604 — zero events is consistent with a true rate up to about 0.02%.
So a phase 1 safety trial in 42 healthy volunteers that reports no adverse events has ruled out a common harm and has said essentially nothing about an uncommon one. A ClinicalTrials.gov search for BPC-157 returned 3 studies in August 2026, the oldest of which is a 42-participant phase 1 safety and pharmacokinetics trial with no results posted3. Our BPC-157 page reads that file line by line.
The other end of the scale
STEP 1 randomised 1,961 adults to once-weekly semaglutide 2.4 mg or placebo over 68 weeks and reported a mean weight change of −14.9% against −2.4%4.
SELECT randomised 17,604 participants and ran for years, because its question was not weight — it was whether cardiovascular events happen less often5.
The jump from roughly two thousand to roughly eighteen thousand is not a jump in ambition. It is the price of measuring an outcome that is rare per person per year. If the event you care about happens to 2% of people annually, you need a great many person-years before the difference between two arms is readable.
Precision, quantified
For a continuous outcome such as percent weight change, the width of the interval shrinks with the square root of n — which is slower than most people expect.
Assume a standard deviation of 8 percentage points, purely as an illustration. At n = 30, the half-width of a 95% interval around the mean is about 2.9 points. At n = 300 it is about 0.9. At n = 3,000 it is about 0.3.
Going from 30 to 3,000 — a hundredfold increase in cost — buys a tenfold narrowing. That is the whole trade, and it is why a sponsor will not fund a 3,000-person trial to answer a question a 30-person trial could settle.
What to do with a small trial
Do not dismiss it. A small, well-conducted, blinded, placebo-controlled trial is worth more than a large open-label case series.
Do read the interval, not the point estimate. If the paper reports one, use it. If it does not, the fractions are usually enough to reconstruct it.
Do ask what the trial was powered for. A trial powered to detect a 20-point difference will report "no significant difference" for a 10-point one, and that is not evidence of no effect.
Do check whether it is the only trial. One small positive result is a hypothesis. Two independent ones start to be a finding. Where a compound sits on that scale is what our evidence grading is trying to capture.
The question that cuts through
Ask what would have to be true for this result to be a fluke — and then ask whether the trial was large enough to rule that out. For eighteen patients, quite a lot could be a fluke. For seventeen thousand, much less.
Frequently asked questions
Is a phase 3 trial always large?
No. Phase is a stage of development, not a size. The published phase 3 of a thymosin beta-4 ophthalmic solution in neurotrophic keratopathy enrolled 18 patients, and its primary healing comparison was reported at p = 0.0656 — described by the authors as a strong efficacy trend rather than a significant difference. STEP 1, also a phase 3, enrolled 1,961.
What does 'no adverse events were reported' mean in a small trial?
Much less than it sounds. With zero events observed in n participants, the upper bound of the 95% interval is roughly 3 divided by n. In 42 participants that upper bound is about 7%; in 18 participants it is about 17%. A small trial reporting no adverse events has ruled out a common harm, not an uncommon one.
Why are cardiovascular outcome trials so much bigger than weight-loss trials?
Because the outcome is rarer. Weight can be measured in every participant at every visit; a heart attack happens to a small percentage of people per year. SELECT randomised 17,604 participants to answer that question, against 1,961 in STEP 1, which measured weight.
Do small trials tend to exaggerate?
A meta-epidemiological study of critical care meta-analyses found that small studies may overestimate effect sizes relative to larger trials asking the same question. The likeliest mechanism is selective publication rather than misconduct — small null results are easier to leave unpublished than small positive ones.
References
- Sosne G et al (2022). 0.1% RGN-259 (Thymosin ß4) Ophthalmic Solution Promotes Healing and Improves Comfort in Neurotrophic Keratopathy Patients in a Randomized, Placebo-Controlled, Double-Masked Phase III Clinical Trial. Int J Mol Sci. https://pubmed.ncbi.nlm.nih.gov/36613994/
- Zhang Z; Xu X; Ni H (2013). Small studies may overestimate the effect sizes in critical care meta-analyses: a meta-epidemiological study. Crit Care. https://pubmed.ncbi.nlm.nih.gov/23302257/
- U.S. National Library of Medicine (2026). NCT02637284 — PCO-02: Safety and Pharmacokinetics Trial. ClinicalTrials.gov. https://clinicaltrials.gov/study/NCT02637284
- Wilding JPH et al (2021). Once-Weekly Semaglutide in Adults with Overweight or Obesity. N Engl J Med. https://pubmed.ncbi.nlm.nih.gov/33567185/
- Lincoff AM et al (2023). Semaglutide and Cardiovascular Outcomes in Obesity without Diabetes. N Engl J Med. https://pubmed.ncbi.nlm.nih.gov/37952131/
Medical disclaimer: This content is for general educational purposes only and is not medical advice, diagnosis, or treatment. Always consult a licensed healthcare professional before starting, stopping, or changing any treatment.
Continue reading
Amylin Analogs: The Next Class After Incretins
An amylin analog has been FDA-approved since 2005. What changed is not the hormone — it is how long the molecule lasts. The trials, in order.
ReadCertificates of Analysis: What They Prove
A COA is a claim about a lot, not a guarantee about a vial. What each test answers, what a purity figure cannot say, and the gap the document never closes.
ReadCompounded Is Not the Trial Drug
A trial tests one defined article: a molecule, a form, a concentration, a container. Compounding changes some of those, and the evidence travels only that far.
Read503A vs 503B: What Is Actually in the Vial
One scheme is inspected against manufacturing standards; the other is not. What each may legally start from, and what the inspection record shows in practice.
ReadGrowth-Hormone Secretagogues: What the Evidence Actually Shows
These compounds reliably raise IGF-1. That is the surrogate. When trials measured function, fracture recovery or disease progression, the results changed.
ReadHealing Peptides and the Gap Between Anecdote and Trial
BPC-157 and TB-500 have thirty years of animal work and a very thin human record. Here is exactly what is registered, what is published, and what neither shows.
ReadHow to Read a ClinicalTrials.gov Record
A registration is a filing, not a review. Fifteen fields, two real peptide records, and the specific places where a record quietly tells you it has stalled.
ReadIncretins Explained: GLP-1, GIP and Glucagon
Three hormones, three receptors, and a class of drugs built on them. What each one does, and why the receptor diagram predicts less than it looks like.
ReadMuscle During Rapid Weight Loss: What Has Been Measured
Roughly a quarter of the weight lost is lean mass — and that was true on placebo too. What the DXA substudies actually show, and how small they are.
ReadOpen-Label vs Blinded: What You Can Conclude
Masking is not a quality score — it decides which explanations a trial can rule out. Five real registry records, and the honest evidence on blinding itself.
ReadOral vs Injectable Peptides: What Changes
One molecule, two routes, one label — and roughly 73 times the milligrams per week. What the gut does to a peptide, and what it costs to get past it.
ReadHalf-Life and Dosing Frequency
From 1.5 minutes to a week, on labels for the same hormone family. Half-life is the number that decides whether a peptide is dosed before meals or once a month.
ReadReconstitution: The Arithmetic, Not the Advice
Concentration is mass divided by volume — except the naive division is wrong, and an FDA label's own numbers prove it. The maths, and what it cannot tell you.
ReadWhat a Phase 2 Result Does and Does Not Tell You
Phase 2 is a dose-finding experiment, not a verdict. Two peptide programmes where the phase 3 exists show exactly which parts of the number survive.
ReadRegistered but Never Reported: The Trials That Vanish
Completed trials that post no results are common and measurable. How to check a compound's registry file, and what the silence does and doesn't tell you.
ReadSurrogate Endpoints vs Outcomes That Matter
Most peptide claims rest on a marker moving, not on anything happening to a person. Two paired trials show what the difference costs to establish.
ReadWhy a Triple Agonist Is Not Simply Better Than a Dual
Adding a receptor adds a mechanism, a side-effect profile, and a reason a trial might read differently. The numbers do not rank the way the labels suggest.
ReadWhat a Peptide Is, and What the Word Hides
Three amino acids or sixty-three thousand daltons — both get called peptides. The definitional confusions that cause the most misreading, settled by labels.
ReadWhat “Research Chemical” Actually Means
The phrase is a sales category, not a regulatory one. What FDA has actually published about seventeen of these peptides, and what the label is doing instead.
ReadHow to Tell a Peptide Claim Is Outrunning Its Evidence
Twelve tells, each with a real case from the registry or the literature. Overstated peptide claims tend to fail on one of them, and usually the same one.
ReadWhat Happens When a Peptide Trial Fails
Four negative results, and what each one killed. A trial can fail an indication, a molecule, a surrogate, or nothing at all — and the difference matters.
ReadWho Funded the Trial, and Why It Matters
Sponsorship shifts conclusions more than it shifts methods. Where to find the funding line, what the Cochrane evidence measured, and what it did not.
ReadWhy Most Peptide Trials Are Small and Short
A registry census, August 2026: one compound has 759 registered studies, another has zero. The reasons are economic, not about whether the drug works.
Read