FDA Approvals Based on Prespecified Subgroup Results After a Failed Overall Trial
When a pivotal trial misses its primary endpoint in the overall population, sponsors face one of the hardest judgment calls in drug development: whether a favorable result in a prespecified subgroup can still support approval, or whether the program must return to the clinic. The answer shapes go/no-go decisions, the design of confirmatory trials, and how subgroup hypotheses are specified and statistically protected in the first place.
The analysis below examines how FDA has actually handled this situation in practice, drawing on approval precedents, statistical and clinical reviews, and Complete Response Letters. It looks at when subgroup findings have supported an approval, when they have not, and what distinguishes the two sets of outcomes.
Want to ask Rhizome your own regulatory questions? Try it for free.
When the overall trial fails: can FDA approve a drug on a prespecified subgroup?
The short answer is that FDA can, but rarely does, and almost never on the strength of a subgroup analysis alone. When a pivotal trial misses its primary endpoint in the overall (intent-to-treat) population, a favorable result in a subgroup is treated as hypothesis-generating rather than as substantial evidence of effectiveness. The approvals that do rest on subgroup findings share a specific feature: the subgroup hypothesis was carried forward into a dedicated, prospectively designed confirmatory trial, or a defensible mechanistic and statistical case narrowed the indication to the population where the effect was real. Where sponsors instead ask the agency to accept a post hoc subgroup as the pivotal evidence, the reviews and Complete Response Letters are consistent: it is not enough.
The governing principle FDA applies
Across statistical and clinical reviews, FDA articulates a stable position. Subgroup analyses are useful for assessing consistency of a demonstrated effect and for generating hypotheses, but they are generally not confirmatory on their own, particularly when they are post hoc, underpowered, or not prospectively specified with type I error control 2527314144. When the overall primary endpoint fails, the agency does not treat a favorable subgroup signal as sufficient by itself unless that effect was prospectively defined, statistically robust, replicated across studies, and not merely a nominal signal arising from multiplicity and small samples 25274041.
The recurring statistical objections are worth stating precisely, because they are the same ones a review division will raise in a pre-submission meeting:
- Multiplicity. Testing treatment effect across many subgroups inflates false-positive rates, so nominal p-values must be read in the context of the number of comparisons made 232426272932334144.
- Post hoc selection. Results identified after seeing the data are exploratory, not confirmatory 262731344044.
- Power and instability. Small subgroups produce unstable estimates and make interaction tests unreliable; the absence of a significant interaction does not prove the effect is uniform, and an isolated nominal signal does not establish a true subgroup effect 2324272932334144.
- Replication. FDA looks for consistency across studies, doses, and populations, and ideally independent prospective confirmation. One trial's post hoc subgroup finding should not be used to "cross-validate" another post hoc analysis 25274041.
Concrete illustrations of the policy in review documents include Praluent (alirocumab), where subgroup analyses were treated as exploratory given multiplicity and low power 25; Olumiant (baricitinib), where the sponsor's subgroup interpretation was deemed post hoc and hypothesis-generating 31; Tanzeum (albiglutide), where the reviewer noted that subgroup results not planned to be confirmatory are not confirmatory 26; Fasenra (benralizumab), where an adolescent subgroup signal was discounted for multiplicity and very small sample size 29; Alimta (pemetrexed), where FDA rejected the logic of using one post hoc subgroup analysis to confirm another and stressed the need for independent prospective validation 41; and the Vafseo (vadadustat) file, where post hoc regional (US vs ex-US) subgroup analyses were characterized as exploratory and hypothesis-generating because of insufficient power and uncontrolled multiplicity 2743.
Precedents where subgroup findings supported approval
The affirmative precedents do not contradict the principle; they show the sanctioned way to act on a subgroup signal.
BiDil (isosorbide dinitrate/hydralazine): subgroup signal, then a confirmatory trial in that subgroup
BiDil is the cleanest example of a subgroup pathway to approval. The earlier heart failure trials, V-HeFT I and V-HeFT II, were not overall positive for the combination, but the data showed a signal in self-identified Black patients. In V-HeFT I, Black patients had lower annual mortality with isosorbide dinitrate plus hydralazine than placebo (9.7% vs 17.3%, nominal p=0.04), while White patients did not show a clear benefit 75. That was a post hoc subgroup observation. Rather than seek approval on it, the sponsor designed A-HeFT specifically to test safety and efficacy in African-American patients with moderate to severe heart failure on standard therapy 75. A-HeFT was statistically significant on its primary composite endpoint (adjusted p=0.021 against an adjusted alpha of 0.044), with significant reductions in all-cause mortality (p=0.012) and first heart failure hospitalization (p<0.001) 75. FDA accepted the pathway: non-definitive overall trials, a clinically relevant subgroup signal, and then a dedicated confirmatory trial in that same subgroup 75.
Radicava (edaravone): a failed broad-population trial, then a prospectively enriched confirmatory trial
Radicava followed the same logic in ALS. The pivotal Study MCI186-16 failed its prespecified primary endpoint in the broad population: change in ALSFRS-R at Week 24 was not significant (approximately -6.35 placebo vs -5.70 edaravone, p=0.41) 6974. Post hoc exploratory analyses identified a more narrowly defined responder subgroup, in which the effect was nominally significant (for example, placebo -7.59 vs edaravone -4.58, p=0.027 in the restricted definition) 697174. The sponsor then prospectively designed Study MCI186-19 to enroll that enriched population (definite or probable ALS, defined severity grade, preserved respiratory function, within two years of diagnosis), and Study 19 was positive on the same endpoint in that population 717273. Importantly, FDA did not treat the failed Study 16 as a positive trial. Approval rested mainly on the positive Study 19, with Study 16 regarded as supportive but not confirmatory, and reviewers openly disagreed about how much weight the Study 16 post hoc findings deserved 7273. The lesson is procedural: the subgroup hypothesis earned its place by being re-tested prospectively, not by being credited retrospectively.
Noctiva (desmopressin): narrowing the indication to the population studied
Noctiva shows a related but distinct move, narrowing the labeled population rather than rescuing a failed endpoint. Here the 1.5 mcg dose was statistically superior to placebo on both co-primary endpoints in the overall randomized population, but FDA concluded the trials did not support a general nocturia indication because of restrictive enrollment and limited generalizability 7. The agency narrowed the indication to nocturia due to nocturnal polyuria, noting the subgroup results were essentially identical to the overall population 7. This is subgroup reasoning used to define, not manufacture, the effective population.
What the Complete Response Letters show
The CRLs are where the boundary is enforced, and recent letters are unusually explicit.
- Troriluzole (Biohaven, NDA 210862). Study 206 failed its prespecified primary endpoint (f-SARA at 48 weeks) in spinocerebellar ataxia. The sponsor relied on post hoc analyses in an SCA genotype 3 subgroup. FDA found the subgroup was not prespecified and, after adjustment for multiple comparisons and prespecified covariates, was neither statistically nor clinically persuasive, and required an adequate and well-controlled study demonstrating an effect on a clinically meaningful endpoint 45.
- Tolebrutinib (Genzyme/Sanofi, NDA 219624). In secondary progressive MS, the sponsor argued a larger effect in the roughly 13% of subjects with baseline gadolinium-enhancing lesions ("active SPMS"). FDA noted the subgroup was not pre-defined in the enrolled population, the analysis did not permit definitive conclusions, and, against the drug's serious drug-induced liver injury risk, substantial evidence of effectiveness had not been established in any clinically identifiable population 464748.
- Wakix (pitolisant, Bioprojet, NDA 211150). For cataplexy, the supportive HARMONY I subgroup analysis was post hoc, cataplexy was only a secondary endpoint with no prospective type I error control, and significance depended on the missing-data handling. FDA judged the single positive HARMONY CTP trial too small and regionally limited to stand alone and required an additional adequate and well-controlled trial 66.
- Sulopenem (Iterum, NDA 213972). FDA required a second adequate and well-controlled trial because the single positive Study 301 was insufficient on its own, particularly when other sulopenem trials were negative or showed inferiority 67.
The through-line is that FDA repeatedly asks for the same remedy it accepted in BiDil and Radicava: a prospectively designed, adequately powered, type I error-controlled study in the population of interest. A post hoc subgroup, however biologically plausible, is treated as the hypothesis for that study, not as its result.
Practical implications for sponsors
For teams facing an overall endpoint miss with a promising subgroup, the precedents point to a narrow but real path. The subgroup should be biologically and clinically coherent, ideally prespecified, and its effect should be reproduced in a dedicated confirmatory trial designed around that population, with multiplicity handled prospectively. Retrospective arguments that a subgroup "was really the target population all along" have a poor track record in the CRLs, especially when weighed against a meaningful safety liability. Where the overall effect is genuinely present but the studied population is narrower than the proposed label, the more successful move is to narrow the indication to match the evidence, as in Noctiva, rather than to elevate a subgroup above a failed primary analysis.