Nearly every question researchers ask about a peptide — does it work, how well, compared to what — is really a question about evidence. And the sentence "studies show" hides an enormous range. It can mean a receptor assay in a dish of immortalized cells. It can mean twelve mice. It can mean a 290-participant randomized trial that missed its primary endpoint. Those three things are not interchangeable, but a vendor description, a forum post, and a summary abstract will often flatten them into the same three words.
This piece is about the rungs of that ladder — what each type of study is designed to answer, what it structurally cannot answer, and the specific reading habits that keep you from over-crediting a result. It is the companion to our primers on pharmacokinetics and receptor pharmacology: those describe what a compound does, this one describes how confidently anyone knows it.
Why the Evidence Question Comes Before the Mechanism Question
A clean mechanism is persuasive. It gives you a story with a beginning and an end: this peptide binds that receptor, which triggers this cascade, which produces that outcome. The problem is that a plausible mechanism is the cheapest thing in pharmacology to produce, and the least predictive of results.
The base rates are unforgiving. In the BIO/Informa/QLS analysis of 12,728 clinical phase transitions across 9,704 drug development programs between 2011 and 2020, the overall likelihood of approval measured from Phase 1 was 7.9% — and that is the success rate for compounds that already had enough preclinical support to justify dosing humans under an IND. Roughly half of programs did not survive Phase 1 to Phase 2 (52.0% average transition rate). Every one of the failures had a mechanism story.
So mechanism tells you what to test. Evidence tells you what happened when someone did.
Rung One: In Vitro — Necessary, Rarely Sufficient
Cell-based and biochemical assays establish that an interaction is possible. A binding assay shows the compound engages the target. A reporter assay shows engagement produces a signal. This is real information, and for compounds whose target is intracellular and has no receptor at all — FOXO4-DRI disrupting a protein-protein interaction, for instance — it is where the mechanism was first demonstrated at all.
Three structural limits are worth internalizing:
Concentration is not dose. The concentration bathing a cell in a well is chosen by the experimenter. It is frequently far above anything a whole organism would sustain at the target tissue. When the cosmetic literature reports catecholamine-release inhibition by SNAP-8-type sequences in chromaffin cells, the working range is in the tens of micromolar — informative about the mechanism, silent on whether that concentration is ever reached through skin.
System artifacts inflate the answer. Receptor reserve means an overexpressing cell line can make a partial agonist look like a full agonist, and can shift potency by orders of magnitude relative to native tissue. This is why EC50 values are not comparable across papers.
There is no ADME. A dish has no liver, no kidney, no plasma proteases, and no barriers. A peptide with a one-minute plasma half-life, like VIP, performs identically to a stable one in a well.
Rung Two: Animal Models — Where Translation Usually Breaks
Animal work adds the whole organism: absorption, distribution, metabolism, clearance, and an intact physiology that can push back. It is the first place a compound can fail for reasons that have nothing to do with target engagement. When reading it, ask three questions.
What is the model actually modelling? Chemically induced colitis, D-galactose-accelerated "aging," or a progeroid mouse strain are constructs. They reproduce some features of a human condition and not others, and a compound that reverses the construct has demonstrated activity against the construct. Some models have strong track records; many do not.
Did anyone outside the originating lab reproduce it? This is the single highest-yield question, and on this shelf it separates compounds sharply. FOXO4-DRI has independent groups working across Leydig cells, chondrocytes, keloid fibroblasts, and vascular endothelium — different labs, different tissues, converging. By contrast, Epithalon and Thymalin rest largely on a single national research lineage published largely in one language, and the tissue-repair literature around BPC-157 is dominated by one originating group. That is not a claim the findings are wrong. It is a claim that the usual error-correction mechanism has not run.
Has it already failed to translate? Sometimes the answer is on record. AICAR produced a striking endurance phenotype in sedentary mice in 2008 that was never reproduced in humans, and the related clinical program was halted at a prespecified futility analysis. That history is more informative than the mouse paper.
Rung Three: Human Trials — Read the Endpoint, Not the Headline
Human studies are the only rung that answers the question people actually care about, and they are the rung most often summarized inaccurately. Four distinctions do most of the work.
Phase tells you the question, not the quality. Phase 1 asks about safety and pharmacokinetics in a small group. Phase 2 looks for a signal and a workable dose range. Phase 3 is the adequately powered efficacy test. "Phase 2 results were encouraging" is a statement about a hypothesis-generating study.
Prespecified primary endpoints are the result. Everything else is a hypothesis. A trial declares in advance what it will measure and how it will be judged. If that measure is missed, the trial is negative — regardless of what the secondary and post-hoc analyses show. The senolytic UBX0101 had an encouraging Phase 1 and then failed to beat placebo on pain in an adequately powered Phase 2. The SOARS-B oxytocin trial (n=290) missed its prespecified social-withdrawal endpoint; a later secondary analysis using data-driven composite outcomes reported significant effects, which is a legitimate hypothesis for a new trial and not a reversal of the old one. The registered endpoint stands.
Surrogate endpoints are not outcomes. A change in a biomarker, an imaging measure, or a functional test is a proxy for something that matters. Tesamorelin's randomized evidence in HIV-associated lipodystrophy is built on visceral fat area, hepatic fat fraction, and waist circumference — real, measured, and surrogate; there is no cardiovascular outcomes trial. Similarly, the September 2025 accelerated approval of SS-31 (elamipretide) for Barth syndrome rested on an intermediate muscle-strength endpoint, in a narrow population, while the same compound missed co-primary endpoints in larger trials in primary mitochondrial myopathy and dry AMD.
Observational is not randomized. When a study compares people who took something to people who did not, the reason they took it is a candidate explanation for any difference found. Confounding by indication and reverse causality are not footnotes; they are frequently the whole finding.
Cross-Cutting Checks That Apply at Every Rung
Publication status. A peer-reviewed paper, a preprint, and a conference abstract are three different levels of scrutiny. Conference abstracts in particular are often a few hundred words with no methods section and no reviewer, and they circulate as if they were papers.
Who is counted, and how many. A first-in-human study with 2 participants and a Phase 3 with 3,000 both appear in citations as "a study."
Whether the study finished. A terminated trial with 4 of 39 planned participants enrolled and no posted results — the fate of Adipotide's only human study — establishes almost nothing, in either direction. It is especially not evidence of the specific toxicity the internet has assigned to it.
Empty files. For several catalog compounds the honest summary is that there are no human data at all. That is a finding worth stating plainly rather than filling with extrapolation.
A Five-Question Checklist
When you encounter a claim about a research peptide, ask in order:
- What species, and in what system? Dish, rodent, or human.
- Was the measured outcome the one declared in advance? Primary endpoint or post-hoc.
- Is the outcome the thing itself or a proxy for it? Outcome or surrogate.
- Has anyone independent reproduced it? Multiple labs or a single lineage.
- Where was it published, and did the study complete? Journal, preprint, abstract, or nothing.
Five questions, and most confident claims about research compounds do not survive them.
FAQ
Does preclinical-only mean a compound doesn't work? No. It means nobody knows. Absence of human data is absence of information, not evidence of failure — and equally, it is not permission to assume success. The correct posture toward a preclinical-only compound is that it is an open research question.
Is a failed trial worse than no trial? For evidence purposes, a completed negative trial is more informative than silence. AOD-9604's Phase IIb weight-loss miss and Thymosin Alpha-1's negative Phase 3 sepsis result tell you something specific: this compound, in this population, at the doses tested, did not produce the effect. That is knowledge. A compound with no trials has neither the finding nor its absence.
Why do class primers keep sorting compounds by evidence maturity? Because the category label is a research question, not a chemical class, and members of the same shelf can sit six rungs apart — approved drug, Phase 3, failed Phase 3, missed Phase IIb, terminated first-in-human, preclinical-only. Grouping by mechanism tells you what a compound is meant to do. Grouping by evidence tells you how much anyone knows about whether it does. Both are covered across our compound library and quality standards pages.
This article is educational and for the laboratory research community. Trulogic Labs products are sold for laboratory and research use only and are not for human consumption.