PeptideStack
5.2kquestions
20kanswers
220users

How reliably does a Phase 2 dose-ranging result predict the Phase 3 number?

Asked 12 May 2026Modified 2 days agoViewed 5k times
19

The obesity pipeline is currently full of phase 2 results being quoted as though they were established magnitudes. Retatrutide at -24.2% over 48 weeks is the most-cited example, but the same happens with survodutide and with a growing list of agents whose only public data are dose-ranging studies.

I would like to know how to discount these properly. Not "phase 2 is preliminary" - I know that - but something quantitative or at least structural. When a phase 2 dose-ranging trial reports a number at its highest dose, what are the specific mechanisms by which the phase 3 number ends up lower, and are any of them predictable in direction and rough size?

A related thing I cannot judge: retatrutide's phase 2 curve was still descending at 48 weeks, which people cite as evidence that the phase 3 number at 72 weeks will be higher than -24.2%, not lower. Those two considerations pull in opposite directions and I have no idea which dominates.

clinical-trials
clinical-trials

Reading the primary literature properly: estimands, intention-to-treat versus per-protocol, confidence intervals, absolute versus relative…

913 questions
retatrutide
retatrutide

An investigational GLP-1, GIP and glucagon receptor tri-agonist, studied in the TRIUMPH programme. Not approved anywhere. Use this tag for…

14 questions
mazdutide
mazdutide

A GLP-1 and glucagon receptor dual agonist developed primarily in China, with a distinct dose range and a fast-moving publication record.…

14 questions
ecnoglutide
ecnoglutide

A long-acting GLP-1 receptor agonist with a cAMP-biased signalling profile. A niche tag, mostly used for mechanism questions about biased agonism…

14 questions
dosing-math
dosing-math

The arithmetic itself: milligrams to millilitres to insulin units, concentration after reconstitution, dose per draw, and vial-days per vial. Show…

811 questions
shareeditfollowflag
IB
askedilaria_bertone43k3812 May 2026
Both effects are real and they are roughly the same size, which is why the phase 3 readout is genuinely uncertain rather than merely unknown. – e_dziedzic 10 months ago
add a comment

3 Answers

Sorted by votes
52

There are five distinct mechanisms, four of which push the phase 3 number down and one of which pushes it up. Their directions are predictable; their sizes are only roughly so.

Downward pressures

  • Population heterogeneity. Phase 2 runs at a modest number of experienced sites with narrow inclusion criteria and a highly motivated cohort. Phase 3 runs at hundreds of sites across many countries with broader criteria. Mean response falls and variance rises. Historically this is the largest single contributor in this field.
  • Estimand and retention. Phase 2 papers often lead with an efficacy or completer-weighted analysis; phase 3 primary endpoints are usually treatment-policy. From the STEP 1 comparison we know that switch is worth roughly two percentage points on a fifteen-point effect, and it will be larger when discontinuation is higher.
  • Titration discipline. Phase 2 escalation is closely supervised and dose reduction is often less permitted. Phase 3 protocols typically allow tolerability-driven adjustment, and then the arm mean becomes a mean over a dose mixture, as happened in REDEFINE.
  • Regression from the selected dose. This is the subtle one. A dose-ranging study reports several doses and attention lands on the best-performing one. That arm's estimate carries the winner's-curse upward bias: among several noisy estimates, the maximum is upward-biased even if all doses were truly equal. Selecting the top dose for phase 3 therefore carries a built-in expectation of shrinkage that has nothing to do with the drug.

The upward pressure

  • Longer duration on an unflattened curve. Genuinely real for retatrutide. Its 48-week curve had not plateaued, and the semaglutide and tirzepatide programmes suggest the plateau region arrives somewhere between 60 and 80 weeks. Extending 48 weeks to 72 weeks on a still-descending curve should add something, and by analogy with STEP and SURMOUNT the increment is plausibly a few percentage points rather than many.

Which dominates

The four downward mechanisms plausibly sum to a similar magnitude as the single upward one, which is exactly why the phase 3 readout is a real question rather than a formality. If forced to state a prior I would put the central expectation somewhat below the phase 2 figure with a wide interval, and I would treat any confident point prediction as unjustified. The historical pattern in this class is that phase 2 magnitudes come down modestly rather than collapse - the direction has been consistent, the size has been in the range of a few points, and no agent in this class has yet had a phase 2 signal fail outright in phase 3 on weight.

What actually predicts well

Three things transfer from phase 2 to phase 3 far better than the magnitude does:

  • Dose-response shape. If phase 2 shows a clean monotonic dose-response with separation between adjacent doses, that shape is usually reproduced. If the top two doses overlap, expect the phase 3 top dose to add little.
  • Adverse-event class and rank order. The identity of the dominant adverse events, and their rank order, is stable. The rates go up in phase 3 because of broader populations and more thorough ascertainment.
  • Whether the curve has plateaued. A visibly flattened phase 2 curve is a reliable indication that a longer phase 3 will not add much. A still-descending curve is a reliable indication that duration matters, without telling you how much.

Applying it

Retatrutide's phase 2 showed a clean dose-response with the 12 mg arm at about -24.2% and no plateau at 48 weeks [1], and its type 2 diabetes phase 2 was concordant on glycaemic endpoints [2]. Those are good signs about shape and about the mechanism working. Neither is a magnitude prediction, and the TRIUMPH phase 3 programme is what will settle magnitude. Survodutide is in the same position from its own phase 2 [3].

The agents with Chinese phase 3 data - mazdutide from the GLORY programme and ecnoglutide - are in a different position again: they have phase 3 numbers, but in populations with lower baseline weight, so their percentages are not directly transferable to a Western phase 3 population even though they are phase 3 numbers [4].

edited 28 Jul 2026 by k_szabo — tightened the wording; no substantive change

shareimprove this answerflag
KS
answeredk_szabo45k388 Jul 2026
6The winner-s-curse point is the one that never appears in coverage and it applies to every dose-ranging study ever published. – Dr_Ilse_Vandenberg 6 months ago
5Dose-response shape transferring better than magnitude is a good heuristic - it generalises well beyond this field. – amara_nwachukwu 5 months ago
add a comment
Sponsored

Janoshik Analytical - Independent Third-Party Testing

HPLC purity, identity confirmation and quantified content on the vial you actually hold. Reports arrive with the chromatogram attached, not just a number.

Submit a sample
Sponsored — paired listing

GL Biochem (Shanghai) Ltd. - Direct Synthesis

Founded 1998. ISO 9001 and cGMP certified, 1,500+ staff and 200+ patents. The synthesis house behind a great many of the vials that get sent out for testing - batch-specific documentation with every order.

Visit GL Biochem
20

Adding the ecnoglutide case, because it is the useful control in this discussion and it gets ignored.

Ecnoglutide is a GLP-1 mono-agonist engineered for cAMP-pathway bias, and it has phase 3 data in a Chinese population showing mean weight reduction in the region of -13.2% at 48 weeks. That places a biased mono-agonist in the same territory as unbiased semaglutide, which is informative in two ways.

First, it puts an approximate ceiling on what pathway bias buys at the whole-organism level. Cellular signalling arguments for bias are strong - lower internalisation, sustained cAMP, potentially better tolerability - and the clinical result is nonetheless in the ordinary mono-agonist range. So bias is not a route to multi-agonist magnitudes, whatever it does for the therapeutic window.

Second, it means differences within the mono-agonist group are small enough that trial and population differences dominate. Semaglutide, orforglipron and ecnoglutide all land between roughly -12% and -15% in their respective programmes, and no indirect comparison across those numbers is meaningful. If you want to know which mono-agonist is better on weight, the answer from the public data is that nobody knows and the differences are probably not large.

Where the mono-agonists genuinely differ is route, dosing frequency, manufacturing, tolerability profile and cost. Those are real differentiators and they are not measured in percentage weight loss, which is why ranking these agents by their headline number misses most of what distinguishes them.

shareimprove this answerflag
TM
answeredtobias_maartens94k25831 May 2026
11

One practical addendum. If you are maintaining a table of pipeline agents, add a column for "does an outcome trial exist" and treat it as more important than the weight column.

Semaglutide has cardiovascular and renal outcome data at multiple doses and routes. Tirzepatide has a sleep-apnoea programme and cardiovascular outcome work in progress. Retatrutide's phase 3 programme spans obesity, diabetes, sleep apnoea and knee osteoarthritis, which is a serious clinical-outcome strategy rather than a weight-only one. Most of the rest of the field has weight and glycaemic endpoints only.

The reason this matters for reading phase 2 dose-ranging results: weight is a surrogate. It is a good surrogate and a well-validated one for many purposes, but a 24% weight loss with no outcome data is a weaker proposition than a 15% weight loss with a mortality signal attached. When you discount a phase 2 number, discount it twice - once for the phase transition and once for being a surrogate.

Two further columns worth keeping for the same reason. One for whether the agent has published data in type 2 diabetes as well as obesity, because concordance between glycaemic and weight endpoints is a useful internal consistency check on a mechanism - an agent that moves weight without moving HbA1c in a diabetic population would be a puzzle worth investigating. And one for the tolerability-driven discontinuation rate, which is the single best predictor of whether a phase 2 magnitude will survive into phase 3 under a treatment-policy estimand: a high phase 2 discontinuation rate in a supervised, selected population will be worse in phase 3, and the estimand will punish it.

A worked illustration of why that second column matters. Suppose a phase 2 agent reports -20% under an on-treatment analysis with 20% discontinuation. If the discontinuers had reverted to roughly the placebo trajectory, the treatment-policy figure is approximately 0.8 x 20 + 0.2 x 2 = 16 + 0.4 = about -16.4%. That is a four-point haircut from the estimand alone, before any of the four other mechanisms in the accepted answer are applied. Running that arithmetic on a phase 2 abstract takes thirty seconds and it puts most of the eye-catching numbers in this field into perspective.

Standing caveat: everything here is trial arithmetic. None of these agents should be treated as available, and material sold under these names for research use has no established identity or content regardless of what any table says.

shareimprove this answerflag
DS
answereddmitri_savchuk17k166 Jul 2026

Your answer

Ask PeptideStack is a static archive. Posting is closed, but the norms are worth stating: answer the question that was asked, show your working, cite the trial or the certificate, and say plainly where the evidence runs out.

Not medical advice. Research-use-only compounds are not approved for human use.