Accepted answer
Fair depends on the comparator arm, and in STEP 4 that means asking whether the comparator was titrated to the same ambition as the experimental one. A head-to-head that runs its comparator to a dose below the one it is licensed at is not measuring the two agents, it is measuring one agent against a handicapped version of the other. Check three things: the maximum comparator dose reached, the proportion of the comparator arm that reached it, and whether the titration schedules had the same duration. If those match, the comparison is fair on dosing and the argument moves to the endpoint. If they do not, the effect size is partly an artefact of the protocol.
This is answerable from the published record, but only if you take the placebo arm seriously rather than reading the active arm alone.
Duration decides what can be seen. A 68-week trial can measure weight and glycaemia; it cannot measure anything whose event rate is one per cent per year without enrolling tens of thousands.
It helps to be literal here: intention-to-treat and per-protocol analyses answer different questions. ITT asks what happens if you offer the treatment; per-protocol asks what happens if it is taken as directed. The gap between the two is a measure of how tolerable the protocol was.
Where a result is quoted from a conference abstract rather than a peer-reviewed publication, the numbers routinely move between the two. It is worth checking which one you are reading.
The short version: check the endpoint, check the comparator, check who was excluded, then look at the number.