Accepted answer
Your suspicion is correct and the effect is large: as a rough rule from meta-regressions across glucose-lowering trials, each additional 1.0 percentage point of baseline A1c buys roughly 0.4 to 0.5 additional percentage points of reduction, whatever the agent. Comparing headline reductions across programmes with different baselines is therefore measuring recruitment as much as pharmacology.
The numbers, with baselines attached
| Trial | Agent and dose | Baseline A1c | A1c change | Comparator change |
| SURPASS-1 | Tirzepatide 5 / 10 / 15 mg, monotherapy | ~7.9% | −1.87 / −1.89 / −2.07 | Placebo −0.04 |
| SURPASS-2 | Tirzepatide 5 / 10 / 15 mg | ~8.3% | −2.01 / −2.24 / −2.30 | Semaglutide 1 mg −1.86 |
| STEP 2 | Semaglutide 2.4 mg, T2DM + obesity | ~8.1% | −1.6 | Placebo −0.4 |
| PIONEER 1 | Oral semaglutide 3 / 7 / 14 mg | ~8.0% | −0.6 / −0.9 / −1.1 | Placebo −0.3 |
| SUSTAIN-6 | Semaglutide 0.5 / 1.0 mg | ~8.7% | −1.1 / −1.4 | Placebo −0.4 / −0.4 |
Sources: [1] [2] [3] [4] [5].
Notice that the ordering is not purely baseline-driven either. PIONEER 1 had a baseline of 8.0% and produced −1.1 at its top dose; STEP 2 had a similar baseline and produced −1.6. So baseline explains part of the spread and exposure explains the rest. Both matter, and neither alone lets you rank the drugs.
Why baseline dependence happens
Three mechanisms, all real:
- A floor. Glucose-lowering agents cannot push A1c much below roughly 5.5% in a person with functioning counter-regulation. Someone starting at 7.2% has at most 1.7 points of headroom before the floor; someone at 9.5% has nearly four. Group mean reductions inherit that ceiling on effect size.
- Curvature in the underlying physiology. At higher A1c, more of the excess comes from fasting hyperglycaemia driven by hepatic glucose output, which responds strongly to these agents. At A1c near 7%, more of the residual excess is postprandial and harder to shift.
- Regression to the mean at the group level. Trials enrol on the basis of an A1c above some entry threshold, which selects people whose measured value was on a high draw. Some of the fall in both arms is that selection unwinding — which is exactly why the placebo arm in these trials never sits at zero, and why the SURPASS-1 placebo change of −0.04 is unusual enough to be worth noticing.
That third point is the reason you must never quote a single-arm change. The interpretable quantity is always the between-arm difference, because the placebo arm absorbs the selection effect, the trial-participation effect and any secular drift in background therapy.
What to compare instead
- The placebo-subtracted difference, not the within-arm change. STEP 2's −1.6 becomes a treatment effect of −1.2 against placebo. PIONEER 1's −1.1 becomes −0.8.
- The proportion reaching a target, such as A1c under 7.0% or at or below 6.5%. This is more clinically legible, but it is even more baseline-sensitive than the mean change, so it only permits comparison when the baseline distributions are similar.
- Head-to-head arms only. SURPASS-2 is the one entry in that table that licenses a between-drug conclusion, because it randomised participants between tirzepatide and semaglutide 1 mg within a single trial. Everything else is an indirect comparison resting on the exchangeability of placebo arms recruited in different years under different background therapy.
- Note the comparator dose. SURPASS-2's semaglutide arm used 1.0 mg, which was the highest approved glycaemic dose at the time and is not the highest dose now available. That is a fair trial design and an unfair citation when the sentence "tirzepatide beat semaglutide" is written without the dose.
On the non-inferiority margin
The conventional 0.3 to 0.4 percentage point margin looks small next to a single person's reference change value of roughly 7% relative — about 0.5 points at an A1c of 7.4%. The apparent contradiction dissolves once you notice they are measurements of different things.
The RCV governs whether one person's two draws differ. The trial margin governs whether two group means differ, and the standard error of a group mean falls as one over the square root of the sample size. With several hundred participants per arm, the standard error on a mean A1c change is on the order of 0.05 points, so a 0.3-point difference is enormous in that currency — roughly six standard errors.
The margin is not chosen for statistical reasons anyway. It is chosen as the largest loss of efficacy that would be clinically tolerable in exchange for whatever the new agent offers, and regulators have historically settled on 0.3 to 0.4 by convention rather than derivation. It is worth being sceptical of that convention: nothing establishes that a 0.35-point A1c difference is clinically unimportant, and a chain of successive non-inferiority trials each conceding 0.3 points can drift a long way from the original comparator.
edited 25 Jul 2026 by mz_4113 — removed a claim I could not source
3The point about successive non-inferiority trials drifting is the classic biocreep argument and it applies here. – ekaterina_volk 26 days ago 2SURPASS-2 using semaglutide 1 mg is the most commonly omitted detail in every comparison I have read. – j_wierzbicki 9 months ago add a comment