PeptideStack
5.2kquestions
20kanswers
220users

Is there an honest apples-to-apples comparison of the STEP and SURMOUNT headline results?

Asked 11 Feb 2025Modified 15 months agoViewed 16k times
28

Every comparison I find is either a marketing graphic putting -20.9% next to -14.9% with no caveats, or a methodology piece telling me cross-trial comparison is invalid and then refusing to say anything further. Both are useless to me.

I accept that indirect comparison is weaker than a head-to-head. What I want is a table that actually lays out the trials with the features that differ - sample size, dose, duration, population, what the placebo arm was doing - so that I can see for myself which comparisons are nearly fair and which are not.

Specific things I do not understand yet: SURMOUNT-1 ran 72 weeks and STEP 1 ran 68, which is a 4-week difference on a curve that has not fully flattened, and I have no idea whether that matters. STEP 2 and SURMOUNT-2 both enrolled people with type 2 diabetes and both show smaller losses than the non-diabetes trials, which I assume is real, but the size of that penalty differs between programmes. And I have read that SURMOUNT-5 is a direct head-to-head, which if true should settle the primary question and make the indirect arithmetic mostly moot.

Can someone lay the programme out properly and say which numbers can be compared and which cannot?

semaglutide
semaglutide

A GLP-1 receptor agonist with a fatty-acid-acylated backbone and a roughly one-week half-life, marketed for type 2 diabetes and for weight…

360 questions
tirzepatide
tirzepatide

A dual GIP and GLP-1 receptor agonist. Questions here cover the SURPASS and SURMOUNT programmes, the practical differences from a pure GLP-1…

162 questions
clinical-trials
clinical-trials

Reading the primary literature properly: estimands, intention-to-treat versus per-protocol, confidence intervals, absolute versus relative…

913 questions
vendor-comparison
vendor-comparison

Side-by-side comparison of suppliers on measurable axes - independently confirmed purity and content, lead time, cold-chain handling,…

351 questions
mazdutide
mazdutide

A GLP-1 and glucagon receptor dual agonist developed primarily in China, with a distinct dose range and a fast-moving publication record.…

14 questions
shareeditfollowflag
ZM
askedzainab_mustafa16k1711 Feb 2025
5The 4-week duration gap is smaller than it looks, but the population differences are larger than most graphics admit. – rota_site 32 days ago
add a comment

3 Answers

Accepted answer first, then by votes
88

Accepted answer

Here is the programme laid out. Every figure below is the primary published estimand for that trial, so the column is internally consistent; where a trial used a withdrawal or lead-in design the weight change is over the randomised period only, which is noted.

TrialnAgent and doseDurationPopulationMean weight changeComparator arm
STEP 11961Semaglutide 2.4 mg weekly68 wkBMI 30+, no diabetes-14.9%-2.4% placebo
STEP 21210Semaglutide 2.4 mg weekly68 wkType 2 diabetes-9.6%-3.4% placebo; -7.0% at 1.0 mg
STEP 3611Semaglutide 2.4 mg weekly68 wkBMI 30+, intensive behavioural therapy in both arms-16.0%-5.7% placebo
STEP 4803Semaglutide 2.4 mg, randomised withdrawal after 20-wk run-inwk 20 to 68BMI 30+, all already on drug-7.9% further+6.9% regain on placebo
STEP 5304Semaglutide 2.4 mg weekly104 wkBMI 30+, no diabetes-15.2%-2.6% placebo
STEP 6401Semaglutide 2.4 mg weekly68 wkEast Asian, BMI 27+-13.2%-2.1% placebo; -9.6% at 1.7 mg
STEP 8338Semaglutide 2.4 mg weekly68 wkBMI 30+, active comparator-15.8%-6.4% liraglutide 3.0 mg daily
SURMOUNT-12539Tirzepatide 5 / 10 / 15 mg weekly72 wkBMI 30+, no diabetes-15.0% / -19.5% / -20.9%-3.1% placebo
SURMOUNT-2938Tirzepatide 10 / 15 mg weekly72 wkType 2 diabetes-12.8% / -14.7%-3.2% placebo
SURMOUNT-3806Tirzepatide max tolerated, after 12-wk lifestyle lead-in72 wk post-lead-inAlready lost 5%+ on lifestyle alone-18.4% further+2.5% regain on placebo
SURMOUNT-4670Tirzepatide, randomised withdrawal after 36-wk open-label lead-inwk 36 to 88BMI 30+, all already on drug-5.5% further+14.0% regain on placebo
SURMOUNT-5751Tirzepatide max tolerated vs semaglutide 2.4 mg72 wkBMI 30+, no diabetes-20.2% tirzepatide-13.7% semaglutide

Sources for the rows above, in order: [1] [2] [3] [4] [5] [6] [7] [8] [9] [10] [11] [12].

Which comparisons are nearly fair

STEP 1 versus SURMOUNT-1 is the closest indirect pair: similar entry criteria, similar mean baseline weight in the region of 105 kg, similar mean baseline BMI around 38, both placebo-controlled with lifestyle counselling rather than intensive therapy, and placebo arms that landed within a percentage point of each other at -2.4% and -3.1%. That last point is the useful diagnostic. When two placebo arms behave the same, the background conditions were probably comparable, and the difference between active arms is more likely to be the drug.

The 4-week duration gap is the smallest of your worries. Both curves are shallow but not flat by then; extrapolating STEP 1 forward 4 weeks would move it by well under a percentage point, and STEP 5 at 104 weeks suggests the semaglutide plateau sits near -15% rather than continuing to descend.

Which comparisons are not fair

STEP 3 versus anything, because both arms got intensive behavioural therapy, which is why its placebo arm lost 5.7% - more than double the placebo arms elsewhere. SURMOUNT-3 versus anything, because participants had already lost at least 5% before randomisation, so the -18.4% is additional loss on top of that and the placebo arm regained rather than lost. The two withdrawal trials versus anything, for the same reason in reverse.

The diabetes penalty is real in both programmes but not equal in size: semaglutide drops from -14.9% to -9.6%, a loss of roughly a third of the effect, whereas tirzepatide 15 mg drops from -20.9% to -14.7%, closer to a 30% relative reduction. Those are similar enough that I would not read a mechanistic story into the difference.

SURMOUNT-5 does mostly settle it

You are right that the head-to-head makes the indirect arithmetic largely redundant for this specific question. Randomised, 72 weeks, maximum tolerated tirzepatide against semaglutide 2.4 mg in people without diabetes, and the separation was about 6.5 percentage points in favour of tirzepatide [12]. Two caveats worth holding: it was open-label, which matters more for tolerability reporting and adherence than for scale weight, and the tirzepatide arm was titrated to maximum tolerated dose while semaglutide was fixed at its licensed 2.4 mg. That is the fair real-world comparison but it is not a milligram-for-milligram pharmacological comparison, and it does not tell you what would happen against a semaglutide dose above 2.4 mg.

edited 6 May 2025 by tobias_maartens — added a caveat about sampling

shareimprove this answerflag
TM
answered · acceptedtobias_maartens94k25816 Apr 2025
6Using the placebo arms as a comparability diagnostic is the single most transferable idea in this thread. – kelvin_lam 9 months ago
7Worth flagging SURMOUNT-5 was open-label - it changes how you read the adverse-event columns more than the weight column. – Dr_Otto_Lindqvist 8 days ago
4The STEP 5 104-week data is the strongest argument that the semaglutide plateau is real rather than an artefact of stopping at 68 weeks. – plate_count_9k 2 months ago
add a comment
Sponsored

Janoshik Analytical - Independent Third-Party Testing

HPLC purity, identity confirmation and quantified content on the vial you actually hold. Reports arrive with the chromatogram attached, not just a number.

Submit a sample
Sponsored — paired listing

GL Biochem (Shanghai) Ltd. - Direct Synthesis

Founded 1998. ISO 9001 and cGMP certified, 1,500+ staff and 200+ patents. The synthesis house behind a great many of the vials that get sent out for testing - batch-specific documentation with every order.

Visit GL Biochem
29

A structural point about the table above that is easy to miss: the two programmes did not use the same titration philosophy, and this leaks into the headline numbers in a way that has nothing to do with receptor pharmacology.

Semaglutide in STEP has one obesity dose. You escalate to 2.4 mg weekly over 16 weeks and that is the dose, with the protocol allowing you to sit at a lower dose if you cannot tolerate it. Tirzepatide in SURMOUNT-1 randomised participants to three fixed maintenance doses and reported them separately, and SURMOUNT-3, -4 and -5 used maximum tolerated dose instead.

Consequences for reading the numbers:

  • Quoting -20.9% for tirzepatide against -14.9% for semaglutide compares the top of a dose-response curve against a single licensed dose. That is a legitimate practical comparison but it is not a comparison at equipotent exposure, and nobody knows what semaglutide at a genuinely higher exposure does in this population.
  • The 5 mg tirzepatide arm at -15.0% and semaglutide 2.4 mg at -14.9% are indistinguishable. Any story that has tirzepatide categorically superior has to explain why its lowest maintenance dose ties.
  • Maximum-tolerated-dose designs bake tolerability into the efficacy endpoint. If a drug is easier to tolerate, more participants reach the top dose, and the arm mean rises for a reason that is real but is not potency.

So the honest summary is a dose-response statement, not a ranking: across the published range, tirzepatide reaches higher mean weight loss than the licensed semaglutide obesity dose, and the head-to-head confirms that at maximum tolerated dosing. Whether that reflects GIP co-agonism, higher achievable exposure, or both, the trials as designed cannot separate.

shareimprove this answerflag
DB
answeredDr_Fatima_Belkacem52k13827 Apr 2025
14

Two housekeeping items for anyone building this table themselves.

First, baseline weight is not constant across the programme and percentages hide that. STEP 6 enrolled an East Asian population with a lower BMI entry threshold of 27 and a substantially lower mean baseline weight, so its -13.2% is a smaller absolute number of kilograms than a -13.2% in STEP 1 would be [1]. The same applies to the later East Asian trial in the programme [2]. If you care about kilograms rather than percentages, record baseline weight as its own column and do the multiplication yourself.

Second, waist circumference, HbA1c and blood-pressure changes are reported as confirmatory secondary endpoints in most of these trials, and they are frequently more comparable across programmes than weight is, because they are less sensitive to the lifestyle background. If two trials disagree on weight but agree on waist and HbA1c, that is a hint that the disagreement is about trial conduct rather than about the drugs.

Third, check whether a trial reported body composition. Very few in this programme did, and the ones that used DEXA reported it in a subset rather than the full population. Since the fraction of lost mass that is lean tissue is one of the live questions about this class, a table of percentage weight loss with no composition column is silently treating 15% loss from one agent as equivalent to 15% from another, and nothing in the published data establishes that.

Fourth, note the run-in and screening attrition. Several of these trials screened substantially more people than they randomised, and the withdrawal designs additionally required completing an open-label lead-in on active drug. That means the randomised population in STEP 4 and SURMOUNT-4 consists of people who already tolerated the drug for months, which is a tolerability-selected group. Their regain figures are therefore estimates of what happens to tolerators who stop, not to a general population.

Neither point changes the ranking. All of them change how much confidence you should attach to small differences, and the last one changes how you should read the withdrawal trials specifically.

shareimprove this answerflag
MI
answeredmicron2236k13825 Mar 2025

Your answer

Ask PeptideStack is a static archive. Posting is closed, but the norms are worth stating: answer the question that was asked, show your working, cite the trial or the certificate, and say plainly where the evidence runs out.

Not medical advice. Research-use-only compounds are not approved for human use.