Accepted answer
They answer two different questions, and neither is "intention-to-treat versus per-protocol" in the classical sense. Both analyses use every randomised participant. What differs is how each one handles the events that happen after randomisation: stopping the drug, and starting rescue therapy.
The two estimands
- Treatment policy. The question is: what happens to weight if you assign this drug as a policy, and then let real life happen? Data collected after a participant stops the drug or adds another weight-loss intervention still counts, at the value it actually was. Discontinuation is treated as part of the effect of the policy, not as a nuisance to be removed. This is the primary estimand in STEP 1 and it produced -14.9% versus -2.4% for placebo, a difference of about -12.4 percentage points [1].
- Trial product. The question is: what is the pharmacological effect of the molecule if it is taken as the protocol intends, without rescue medication? Observations after discontinuation or rescue are handled as if that deviation had not occurred, using a model that borrows information from participants who remained on drug. In STEP 1 this shifts the semaglutide arm to roughly -17% while leaving the placebo arm essentially where it was, because placebo participants who stopped were not losing much anyway.
Note what that asymmetry tells you. The gap between the two numbers in the active arm is almost entirely the arithmetic of dilution: roughly one in six participants was off drug by week 68, and their weight had partly come back. In the placebo arm there was nothing to come back from, so the estimand choice barely moves it.
Which one you want
If you are asking "what does prescribing this drug to a population achieve", treatment policy is the honest answer, because discontinuation is a real and large part of what happens. If you are asking "what does this molecule do to adipose tissue in someone who keeps taking it", the trial-product number is closer. Regulators generally want the first. People comparing molecules head-to-head usually want the second, because it is less contaminated by trial-conduct differences.
The failure mode to avoid is mixing them across trials. Quoting the trial-product figure for one drug and the treatment-policy figure for another manufactures a difference out of nothing but analysis convention.
Yes, SURMOUNT does the same thing
SURMOUNT-1 reports a treatment-regimen estimand of -15.0%, -19.5% and -20.9% at 5, 10 and 15 mg versus -3.1% for placebo, and an efficacy estimand of about -16.1%, -21.4% and -22.5% [2]. Same structure, different labels. "Treatment regimen" maps to treatment policy; "efficacy" maps to trial product. The vocabulary is not standardised between sponsors, which is a large part of why this confuses everyone.
Practical rule for your spreadsheet: add a column recording which estimand each row came from, and refuse to compare rows that disagree. If a source quotes a number without saying which, treat the number as unusable rather than guessing.
edited 9 Aug 2024 by Dr_Nadia_Farsi — tightened the wording; no substantive change
2The asymmetry point is the one most reviews miss - the estimand choice moves the active arm far more than placebo. – fib4_reader 30 days ago Adding an estimand column to my own table immediately killed three comparisons I had been making. – Dr_Malik_Osei 9 months ago add a comment