Evidence review
Open-Label vs Blinded: What You Can Conclude
Masking is not a quality score — it decides which explanations a trial can rule out. Five real registry records, and the honest evidence on blinding itself.
What masking is for
A blinded trial is not trying to be rigorous for its own sake. It is trying to remove a specific set of alternative explanations for its own result.
If a participant knows they got the active drug, their report of pain, energy, sleep or recovery is measuring the drug plus their expectation of the drug. Blinding does not make the expectation go away. It puts an equal one in the other arm.
The five labels, decoded
ClinicalTrials.gov records carry a masking field with five values, plus a list of who was masked.
None — an open-label trial. Everyone knows who got what.
Single — one group is masked, usually the participants.
Double — two groups. On the retatrutide phase 2 record, NCT04881760, that is participants and investigators1.
Triple — three groups.
Quadruple — participants, care providers, investigators and outcome assessors. STEP 1, NCT03548935, is quadruple-masked2.
The label alone is not the whole story. The record for the BPC 157 hamstring trial, NCT07437547, adds a masking description: identical-appearing prefilled syringes prepared by an investigational pharmacy under a randomisation code, and MRI injury volume assessed by blinded central radiology review3. That is a much more specific claim than the word "quadruple", and it is the sort of detail worth reading for.
The uncomfortable evidence about blinding itself
Here is where an honest page has to slow down.
The largest study of whether blinding actually changes measured treatment effects — MetaBLIND, covering 142 meta-analyses and 1,153 trials from the Cochrane Database — found no evidence for an average difference in estimated treatment effect between trials with and without blinded patients, healthcare providers or outcome assessors4.
For trials not reported as double blind versus those that were, the ratio of odds ratios was 1.02, with a 95% credible interval of 0.90 to 1.13, across 74 meta-analyses4.
The authors' own reading is worth quoting in shape if not in full: these results could reflect that blinding is less important than often believed, or they could reflect limitations of the method, such as residual confounding and imprecision. They recommend replication, and they say blinding should remain a methodological safeguard4.
So the correct position is not "open-label trials inflate results by X". It is that the average inflation, if any, has not been demonstrated, while the specific mechanisms by which an open-label design can mislead remain entirely real and are visible case by case.
Where the mechanism bites hardest
Subjective primary endpoints. Pain scores, recovery ratings, sleep quality, "energy". These are reports, and a report is made by a person who knows something.
Decisions made by clinicians. Return-to-sport dates, hospital discharge, dose escalation, rescue medication. Someone chooses, and the chooser has a belief.
Everything that follows from expectation. Dropout, adherence, co-interventions, how hard a physiotherapist pushes.
Against that, blinding matters much less for a laboratory value read by a machine from a coded sample, or for all-cause mortality.
This is the practical test: name the primary endpoint, then ask whether a person who knew the assignment could have nudged it. If the answer is yes, the masking field is load-bearing.
Worked: the record that shows the problem
Consider NCT07752381, registered in 2026 for a study completed in November 2025. It is interventional, 40 participants, allocation not applicable, a single group, masking none. The primary outcome measures include change in high-sensitivity C-reactive protein and interleukin-6 — laboratory values — alongside change in self-reported swelling5.
Every person in that study received the product. There is no comparison arm. The self-reported measures are made by participants who know exactly what they took, over eight weeks, in a design with no way to separate the product from time, from rehabilitation, from regression to the mean, or from expectation.
None of which makes it a bad study. It is a legitimate design for its purpose. It simply cannot support a sentence of the form "this reduced inflammation", and if such a sentence is being built on it, the design is the reason to doubt it.
An extension is not automatically unmasked
A common shortcut is to assume that the long-term data from any drug programme are open-label. Sometimes they are. Often they are not, and it is worth checking rather than assuming.
The tesamorelin extension study, NCT00608023, ran 263 participants for a further period with quadruple masking retained, and re-randomised participants across tesamorelin-to-tesamorelin, tesamorelin-to-placebo and placebo-to-tesamorelin sequences, with fasting glucose and oral glucose tolerance at week 52 as its primary endpoints6.
That design is why its 52-week metabolic data can be read as a comparison rather than a time series. Our tesamorelin page covers what that programme did and did not establish.
What an open-label result is good for
Generating a hypothesis worth testing properly.
Measuring things a person cannot influence — pharmacokinetics, an assay, an imaging measurement read blind.
Describing safety events that are unambiguous and serious.
What it is not good for is settling whether a compound works, when the outcome is something a hopeful person reports about themselves.
The reading rule
Open the registry record. Find the masking field, then find the primary endpoint. If the endpoint is subjective and the masking is none, the trial has measured the drug and the expectation together and has no way to tell you which is which. That is not a small caveat — it is the whole conclusion. The same discipline applies when you read any registry record.
Frequently asked questions
Does an open-label trial always exaggerate the effect?
Not demonstrably, on average. MetaBLIND, covering 142 meta-analyses and 1,153 trials, found no evidence for an average difference in estimated treatment effect between blinded and non-blinded trials; for trials not reported as double blind versus those that were, the ratio of odds ratios was 1.02 (95% credible interval 0.90 to 1.13). The authors noted this could reflect study limitations, recommended replication, and said blinding should remain a methodological safeguard.
What does 'quadruple masking' mean?
That participants, care providers, investigators and outcome assessors were all masked to assignment. STEP 1 (NCT03548935) is recorded that way. Some records go further and describe the mechanism — the BPC 157 hamstring record (NCT07437547) specifies identical-appearing prefilled syringes prepared under a randomisation code and MRI injury volume read by blinded central radiology review.
When does blinding matter least?
When the primary endpoint cannot be influenced by knowing the assignment — a laboratory value read from a coded sample, an imaging measurement read blind, or all-cause mortality. It matters most for pain scores, recovery ratings, sleep quality and any clinician decision such as discharge or return to sport.
Is a single-group study with no comparison arm worthless?
No, but it answers a narrower question. NCT07752381 enrolled 40 participants in a single group with masking recorded as none and self-reported swelling among its primary measures. With no comparison arm, there is no way to separate the product from time, rehabilitation, regression to the mean or expectation.
References
- U.S. National Library of Medicine (2023). NCT04881760 — A Study of LY3437943 in Participants Who Have Obesity or Are Overweight. ClinicalTrials.gov. https://clinicaltrials.gov/study/NCT04881760
- U.S. National Library of Medicine (2021). NCT03548935 — STEP 1: Research Study Investigating How Well Semaglutide Works in People Suffering From Overweight or Obesity. ClinicalTrials.gov. https://clinicaltrials.gov/study/NCT03548935
- U.S. National Library of Medicine (2026). NCT07437547 — BPC 157 for Acute Hamstring Muscle Strain Repair. ClinicalTrials.gov. https://clinicaltrials.gov/study/NCT07437547
- Moustgaard H et al (2020). Impact of blinding on estimated treatment effects in randomised clinical trials: meta-epidemiological study. BMJ. https://pubmed.ncbi.nlm.nih.gov/31964641/
- U.S. National Library of Medicine (2026). NCT07752381 — A Clinical Trial to Evaluate the Effects of Peptide Gummies on Markers of Inflammation, Physical Performance, and Recovery. ClinicalTrials.gov. https://clinicaltrials.gov/study/NCT07752381
- U.S. National Library of Medicine (2022). NCT00608023 — TH9507 Extension Study in Patients With HIV-Associated Lipodystrophy. ClinicalTrials.gov. https://clinicaltrials.gov/study/NCT00608023
Medical disclaimer: This content is for general educational purposes only and is not medical advice, diagnosis, or treatment. Always consult a licensed healthcare professional before starting, stopping, or changing any treatment.
Continue reading
Amylin Analogs: The Next Class After Incretins
An amylin analog has been FDA-approved since 2005. What changed is not the hormone — it is how long the molecule lasts. The trials, in order.
ReadCertificates of Analysis: What They Prove
A COA is a claim about a lot, not a guarantee about a vial. What each test answers, what a purity figure cannot say, and the gap the document never closes.
ReadCompounded Is Not the Trial Drug
A trial tests one defined article: a molecule, a form, a concentration, a container. Compounding changes some of those, and the evidence travels only that far.
Read503A vs 503B: What Is Actually in the Vial
One scheme is inspected against manufacturing standards; the other is not. What each may legally start from, and what the inspection record shows in practice.
ReadGrowth-Hormone Secretagogues: What the Evidence Actually Shows
These compounds reliably raise IGF-1. That is the surrogate. When trials measured function, fracture recovery or disease progression, the results changed.
ReadHealing Peptides and the Gap Between Anecdote and Trial
BPC-157 and TB-500 have thirty years of animal work and a very thin human record. Here is exactly what is registered, what is published, and what neither shows.
ReadHow to Read a ClinicalTrials.gov Record
A registration is a filing, not a review. Fifteen fields, two real peptide records, and the specific places where a record quietly tells you it has stalled.
ReadIncretins Explained: GLP-1, GIP and Glucagon
Three hormones, three receptors, and a class of drugs built on them. What each one does, and why the receptor diagram predicts less than it looks like.
ReadMuscle During Rapid Weight Loss: What Has Been Measured
Roughly a quarter of the weight lost is lean mass — and that was true on placebo too. What the DXA substudies actually show, and how small they are.
ReadOral vs Injectable Peptides: What Changes
One molecule, two routes, one label — and roughly 73 times the milligrams per week. What the gut does to a peptide, and what it costs to get past it.
ReadHalf-Life and Dosing Frequency
From 1.5 minutes to a week, on labels for the same hormone family. Half-life is the number that decides whether a peptide is dosed before meals or once a month.
ReadReconstitution: The Arithmetic, Not the Advice
Concentration is mass divided by volume — except the naive division is wrong, and an FDA label's own numbers prove it. The maths, and what it cannot tell you.
ReadWhat a Phase 2 Result Does and Does Not Tell You
Phase 2 is a dose-finding experiment, not a verdict. Two peptide programmes where the phase 3 exists show exactly which parts of the number survive.
ReadRegistered but Never Reported: The Trials That Vanish
Completed trials that post no results are common and measurable. How to check a compound's registry file, and what the silence does and doesn't tell you.
ReadWhy n=30 Is Not the Same as n=3,000
Enrolment size is not a quality badge — it is precision, and it is the ability to see rare harm. Worked on real peptide trials, with the arithmetic shown.
ReadSurrogate Endpoints vs Outcomes That Matter
Most peptide claims rest on a marker moving, not on anything happening to a person. Two paired trials show what the difference costs to establish.
ReadWhy a Triple Agonist Is Not Simply Better Than a Dual
Adding a receptor adds a mechanism, a side-effect profile, and a reason a trial might read differently. The numbers do not rank the way the labels suggest.
ReadWhat a Peptide Is, and What the Word Hides
Three amino acids or sixty-three thousand daltons — both get called peptides. The definitional confusions that cause the most misreading, settled by labels.
ReadWhat “Research Chemical” Actually Means
The phrase is a sales category, not a regulatory one. What FDA has actually published about seventeen of these peptides, and what the label is doing instead.
ReadHow to Tell a Peptide Claim Is Outrunning Its Evidence
Twelve tells, each with a real case from the registry or the literature. Overstated peptide claims tend to fail on one of them, and usually the same one.
ReadWhat Happens When a Peptide Trial Fails
Four negative results, and what each one killed. A trial can fail an indication, a molecule, a surrogate, or nothing at all — and the difference matters.
ReadWho Funded the Trial, and Why It Matters
Sponsorship shifts conclusions more than it shifts methods. Where to find the funding line, what the Cochrane evidence measured, and what it did not.
ReadWhy Most Peptide Trials Are Small and Short
A registry census, August 2026: one compound has 759 registered studies, another has zero. The reasons are economic, not about whether the drug works.
Read