Peptide basics

How to read a sex split in a trial result without over-reading it

On this page
  1. The mistake the checklist prevents
  2. The five questions
  3. Running the checklist on the GLP-1 sex splits
  4. What a pooled estimate adds, and what it does not
  5. Common questions

When a trial reports one result for women and another for men, those are two separate estimates. A difference between them is not automatically a real difference in effect.

The standard test is whether the trial pre-specified the subgroup, tested a small number of them, and reported a formal statistical test of interaction.

A 2012 review of 207 randomised trials reporting subgroup analyses found that only 9% of subgroup claims were backed by a statistically significant interaction test.

Applied to the GLP-1 sex splits, most published figures are subgroup or post-hoc estimates and should be read as observations rather than findings.

This page contains no dose, schedule or route and makes no claim about any compound. It is a method page: how to judge whether a sex-split result you have been shown supports the conclusion attached to it.

The mistake the checklist prevents

A trial reports a result in women and a result in men. One looks bigger. The natural next sentence — that the drug works better in women — is a claim about an interaction between sex and treatment, and it does not follow from two separate within-group estimates.1

The reason is that each estimate carries its own uncertainty, and the smaller group carries more of it. Two intervals that overlap substantially can still produce point estimates that look far apart, especially when one group is a third the size of the other.14

The five questions

The NEJM methods paper on subgroup reporting sets out what an author should provide and what a reader should look for. Reduced to questions you can ask of any sex-split figure, it comes to five.1

  1. Was the subgroup analysis pre-specified in the protocol or statistical analysis plan, before the data were unblinded?
  2. Was a formal test of interaction reported, rather than only two separate estimates?
  3. How many subgroups were examined in total? Many comparisons raise the chance of a spurious one.
  4. Was the subgroup variable measured at baseline, and was it used to stratify randomisation?
  5. Has the same direction of effect been seen in other trials, or is this the only place it appears?12

The BMJ review that formalised these criteria applied them to 207 randomised trials published in core clinical journals that reported subgroup analyses. Sixty-four made a subgroup claim for the primary outcome.2

Of those 64 claims, 26 (41%) had clearly pre-specified their hypotheses, 4 (6%) had correctly pre-specified the direction, 13 (20%) had used the subgroup variable as a stratification factor at randomisation, and 6 (9%) reported a test of interaction that proved statistically significant.2

Running the checklist on the GLP-1 sex splits

The most widely repeated sex split in this literature is the STEP 1 comparison reported in a 2023 review: an estimated treatment difference in weight reduction of −14.0% in women against −8.0% in men. Two within-group estimates, presented side by side.5

Question one: we are not aware of documentation that sex was a pre-specified subgroup with a stated directional hypothesis in that trial. Question two: no formal interaction test accompanies the figures as usually quoted. Question four: STEP 1's registry record does not indicate sex was used as a randomisation stratification factor.45

Question five is where the claim does better. A 2025 meta-analysis pooled 14 randomised trials reporting weight change by sex across five compounds and found a consistent direction: females lost more weight than males, with a pooled mean difference of 1.04 kg and 1.69 percentage points.3

Replication across trials and compounds is the strongest thing this claim has going for it. It is also the criterion the BMJ review found satisfied least often in practice, so it is worth weighting.23

What a pooled estimate adds, and what it does not

Pooling across trials increases the number of participants behind an estimate and shows whether the direction is stable. Both are genuine gains over a single trial's split.3

It does not repair the underlying design. If none of the pooled trials pre-specified sex or tested an interaction, the pooled result is a summary of many weak analyses rather than one strong one. The 2025 authors flagged this directly, noting that incorporating post-hoc analyses raised the risk of bias and that many included studies lacked granular sex-stratified data.3

It also cannot tell you anything about an individual. A one-kilogram average difference between two large groups is compatible with enormous overlap between them.3

Semaglutide's sex-split figures are the ones most often reproduced. Before repeating one, check whether the source reports an interaction test or only two separate estimates — in the widely circulated version, it is the latter.

Common questions

If a subgroup result is not confirmed, should I ignore it?

No. Treat it as a hypothesis rather than a finding. It is a reason to look for replication and for a properly designed test, not a reason to conclude anything about how a compound behaves in one sex.

What is a test of interaction?

A single statistical test of whether the treatment effect differs between the subgroups, rather than two separate tests of whether the treatment had an effect within each. It is the only one of the two that addresses the question being asked.

Does a large trial make its subgroups reliable?

Not by itself. A trial powered for its primary outcome is usually underpowered for interactions, because detecting a difference between effects requires substantially more participants than detecting the effect itself.

Sources

  1. Secondary source
    Statistics in Medicine — Reporting of Subgroup Analyses in Clinical TrialsWang R, Lagakos SW, Ware JH, Hunter DJ, Drazen JM. New England Journal of Medicine, 2007 · doi:10.1056/NEJMsr077003 · PMID 18032770doi.org/10.1056/NEJMsr077003Back to text
  2. Systematic review
    Credibility of Claims of Subgroup Effects in Randomised Controlled Trials: Systematic ReviewSun X, Briel M, Busse JW, You JJ, Akl EA, et al.. BMJ, 2012 · doi:10.1136/bmj.e1553 · PMID 22422832doi.org/10.1136/bmj.e1553Back to text
  3. Meta-analysis
    Sex Differences in the Efficacy of Glucagon-Like Peptide-1 Receptor Agonists for Weight Reduction: A Systematic Review and Meta-AnalysisYang Y, He L, Han S, Yang N, Liu Y. Journal of Diabetes, 2025 · doi:10.1111/1753-0407.70063 · PMID 40040445doi.org/10.1111/1753-0407.70063Back to text
  4. Primary study
    STEP 1 (NCT03548935) — posted results, baseline characteristicsClinicalTrials.gov, U.S. National Library of Medicine, 2021clinicaltrials.gov/study/NCT03548935Back to text
  5. Secondary source
    Semaglutide in Obesity: Unmet Needs in MenJensterle M, Rizzo M, Janež A. Diabetes Therapy, 2023 · doi:10.1007/s13300-022-01360-7 · PMID 36609945doi.org/10.1007/s13300-022-01360-7Back to text