PeptideStack
5.2kquestions
20kanswers
220users

How much of the placebo-arm result in these trials is the lifestyle intervention, and does it distort the drug effect?

Asked 3 Sept 2025Modified 8 months agoViewed 6.5k times
21

I noticed that the placebo arms in this literature do not behave consistently. STEP 1 placebo loses 2.4%. STEP 3 placebo loses 5.7%. SURMOUNT-3 placebo gains 2.5%. SURMOUNT-4 placebo gains 14%. These are all placebo arms in obesity trials by the same two sponsors within a few years of each other, so the variation has to be design rather than chance.

What I am trying to work out is whether that variation distorts the drug effect I care about. If the placebo arm in one trial is getting a genuinely effective behavioural intervention and the placebo arm in another is getting a leaflet, then the between-arm difference - which is the number everyone quotes as "the drug effect" - is measuring different things in the two trials.

Worse, I can imagine it cutting both ways. A strong lifestyle arm should shrink the difference, making the drug look weaker. But it might also raise the active arm, if drug and behaviour change add. STEP 3 has the highest active-arm number in the whole semaglutide obesity programme at -16.0%, which is consistent with the second story.

Is there a defensible way to read across trials with such different backgrounds, or do I have to treat each trial as its own island?

clinical-trials
clinical-trials

Reading the primary literature properly: estimands, intention-to-treat versus per-protocol, confidence intervals, absolute versus relative…

913 questions
semaglutide
semaglutide

A GLP-1 receptor agonist with a fatty-acid-acylated backbone and a roughly one-week half-life, marketed for type 2 diabetes and for weight…

360 questions
tirzepatide
tirzepatide

A dual GIP and GLP-1 receptor agonist. Questions here cover the SURPASS and SURMOUNT programmes, the practical differences from a pure GLP-1…

162 questions
exercise
exercise

Training during substantial weight loss: cardiorespiratory work, step count, low energy availability, and how exercise interacts with the appetite…

36 questions
weight-regain
weight-regain

Regain after stopping or reducing: the trajectory reported in the withdrawal extensions, how much is fluid, and what the maintenance arms tell us…

14 questions
shareeditfollowflag
BA
askedben_akintola14k283 Sept 2025
7SURMOUNT-4 placebo gaining 14% is not a lifestyle effect at all - that is a withdrawal design and needs separating out. – a_lindgren 8 months ago
add a comment

3 Answers

Sorted by votes
62

Your four placebo arms are actually three different phenomena, and separating them dissolves most of the confusion.

1. Background intervention intensity (STEP 1 versus STEP 3)

STEP 1 gave both arms brief lifestyle counselling: monthly contact, a modest deficit target, activity advice. Its placebo arm lost 2.4% [1]. STEP 3 gave both arms intensive behavioural therapy - roughly 30 counselling contacts across 68 weeks plus an 8-week low-calorie meal-replacement period at the start. Its placebo arm lost 5.7% [2]. That 3.3-percentage-point gap is a clean, within-programme estimate of what intensive behavioural therapy adds on top of brief counselling, and it is a useful number to carry around.

Now the important part. The active arms were -14.9% and -16.0%. So adding an intervention worth 3.3 points to the placebo arm added about 1.1 points to the semaglutide arm. The effects are not additive; they are strongly sub-additive. The between-arm difference therefore shrank from about -12.4 points in STEP 1 to about -10.3 points in STEP 3.

That is the answer to your central question, and it is not the answer most people expect. A stronger control arm makes the between-arm difference smaller without meaning the drug did less. Both readings you proposed are partly right - the active arm did rise - but the rise was small relative to the placebo rise, so the net effect on the headline difference is negative.

2. Sequential design (SURMOUNT-3)

SURMOUNT-3 is not a comparison of drug against lifestyle. Everyone did 12 weeks of intensive lifestyle intervention first, and only people who had already lost at least 5% were randomised [3]. So the placebo arm consists of demonstrated lifestyle responders who had already banked their loss, and what you are watching over the next 72 weeks is what happens next: a 2.5% regain. The 18.4% in the active arm is additional loss layered on top of the lead-in loss.

This design answers a genuinely different question - what does adding drug to someone already succeeding on lifestyle achieve - and it also quietly demonstrates that lifestyle responders regain in the absence of pharmacology. Do not put its numbers in the same column as STEP 1.

3. Randomised withdrawal (STEP 4, SURMOUNT-4)

Nothing to do with lifestyle. Both trials ran an open-label lead-in on active drug, then randomised participants to continue or switch to placebo. STEP 4's placebo arm gained 6.9% over the following 48 weeks; SURMOUNT-4's gained 14.0% over 52 weeks [4] [5]. The gain is regain after drug withdrawal, and its magnitude tracks how much had been lost during the lead-in - SURMOUNT-4's lead-in was longer and deeper, so there was more to regain.

How to read across anyway

Three rules that work:

  • Use the placebo arm as a calibration instrument. If two placebo arms landed within about a percentage point of each other, the backgrounds were comparable and the active-arm difference is interpretable. STEP 1 at -2.4% and SURMOUNT-1 at -3.1% pass this test. STEP 3 at -5.7% fails it against both.
  • Compare differences, not active arms, but only within a design class. Placebo-controlled parallel trials with each other; withdrawal trials with each other; lead-in trials with each other.
  • Expect sub-additivity. If you are mentally adding a diet effect to a drug effect, halve your estimate of the sum. The physiological reason is straightforward: both interventions work partly through the same final pathway of reduced energy intake, and there is a floor on how little a person will eat.

None of this is a reason to treat each trial as an island. It is a reason to record design class alongside every number, and to distrust any figure quoted without it.

shareimprove this answerflag
DA
answeredDr_Yusuf_Adeyemi95k24818 Nov 2025
8The sub-additivity estimate from STEP 1 versus STEP 3 is the most quotable thing here - 3.3 points added to placebo, 1.1 to active. – Dr_Yusuf_Adeyemi 6 months ago
Also worth noting STEP 3 randomised 2:1 with a smaller n, so its point estimates carry wider intervals than STEP 1. – Dr_Sara_Kuusela 7 months ago
add a comment
Sponsored

Sigma-Aldrich - Certified Reference Materials

Analytical standards and reagents with traceable certificates. Every quantitative result you read inherits the accuracy of the standard behind it.

Shop standards
24

Supplementing the accepted answer with the part that gets people into trouble in practice: blinded placebo arms in these trials are not a good model for no treatment.

Participants in a placebo arm know they are in a weight-loss trial, get weighed regularly by someone who writes the number down, receive counselling of some intensity, and have a one-in-two or one-in-three chance of being on the drug, which affects behaviour. Every one of those is an intervention. The measured -2.4% in STEP 1 is therefore an estimate of "trial participation with brief counselling", not of "doing nothing".

Observational cohorts of untreated people with obesity over a comparable interval typically show weight roughly stable to slightly increasing. So if you want an estimate of drug effect against genuine no-treatment, the between-arm difference in a placebo-controlled trial is arguably a slight underestimate rather than an overestimate. That works in the opposite direction to the lifestyle-confounding worry in the question, and the two partially cancel.

The other practical consequence is for anybody trying to reason about outcomes outside a trial. Trial participants get structured contact, free drug, protocol-driven titration, and monitoring. Retention in routine practice is markedly worse than in these trials, and the treatment-policy estimand already reflects trial-grade retention rather than real-world retention. Whatever number you take from the table, real-world mean outcomes will sit below it - and that is before any question of product identity or content, which for research-use-only material is not established at all. If you are making decisions about your own health, that conversation belongs with a clinician who can see your bloodwork.

shareimprove this answerflag
BU
answeredbufferline4249k13829 Nov 2025
11

Small addition on the arithmetic of the withdrawal trials, because the percentages are relative to shifting baselines and this trips people up.

SURMOUNT-4's lead-in produced about -20.9% by week 36. The continuation arm then lost a further 5.5% over 52 weeks and the placebo arm gained 14.0%. Those later percentages are expressed relative to the week-36 weight, not the original baseline. So for a participant starting at 100 kg:

  • Week 36 weight: 100 - 20.9 = 79.1 kg.
  • Continuation arm at week 88: 79.1 x (1 - 0.055) = 74.8 kg, which is -25.2% from original baseline.
  • Withdrawal arm at week 88: 79.1 x (1 + 0.140) = 90.2 kg, which is -9.8% from original baseline.

The gap between arms at the end is about 15.4 kg on a 100 kg starting weight. Chaining the percentages naively - 20.9 minus 5.5 versus 20.9 plus 14.0 - gives -26.4% and -6.9%, which is wrong in both arms, and wrong by more in the arm that gained. Always convert to absolute mass before combining, then convert back once at the end.

The same trap catches STEP 4. Its 20-week run-in produced roughly -10.6% before randomisation, so a 100 kg participant was at 89.4 kg at week 20. The continuation arm's further -7.9% gives 89.4 x 0.921 = 82.3 kg, or -17.7% from original baseline. The withdrawal arm's +6.9% gives 89.4 x 1.069 = 95.6 kg, or -4.4% from baseline. Between-arm gap of about 13.3 kg.

Two things this arithmetic makes visible that the percentages hide. First, the withdrawal arms in both trials end up well above where the continuation arms are but still below their original baseline - regain was substantial and incomplete over the follow-up periods studied, which is a different claim from "the weight all came back". Second, the reason SURMOUNT-4's regain looks so much more dramatic than STEP 4's is largely that its lead-in was longer and deeper: more loss banked means more available to regain, and the percentages are computed against a smaller denominator. Comparing the two regain figures as if they measured the same quantity is not defensible.

shareimprove this answerflag
TA
answeredtri_gly_ala48k3811 Dec 2025

Your answer

Ask PeptideStack is a static archive. Posting is closed, but the norms are worth stating: answer the question that was asked, show your working, cite the trial or the certificate, and say plainly where the evidence runs out.

Not medical advice. Research-use-only compounds are not approved for human use.