Rhizome

FDA Endpoint Selection, Composite Endpoints, and Biomarkers in Ultra-Rare Disease Trials

Chetan Mishra
Chetan Mishra
Sep 22, 2026

Ultra-rare disease programs rarely support a conventional adequate and well-controlled trial: the eligible population is small, phenotypically heterogeneous, and often spread across disease stages that progress at different rates. Because FDA does not lower the substantial-evidence standard for these programs, the weight of the review shifts onto endpoint choice, measurement rigor, and how credibly a sponsor can interpret a result from a small, single-arm, or externally controlled study. Endpoint decisions made early in development therefore determine whether an ultra-rare program has a viable approval pathway at all.

The analysis below draws on FDA's rare-disease and endpoint guidance together with the Drugs@FDA review record to trace how the agency has actually reasoned in three recurring areas: selecting endpoints for heterogeneous populations, constructing and defending composite endpoints, and grounding approval on biomarker-based evidence. Each section cites the underlying guidance and review documents so the precedent can be checked against a specific program.

Want Rhizome's help on your own question? Try it for free.

Endpoint selection, composite endpoints, and biomarkers in ultra-rare disease trials: how FDA has actually decided

Ultra-rare disease programs force a collision between two regulatory imperatives: the statutory requirement for substantial evidence from adequate and well-controlled investigations, and the reality that too few sufficiently similar patients exist to run a conventional trial. FDA's practice, visible across its guidance and its Drugs@FDA review record, is not to relax the evidence standard but to move the burden onto endpoint design, measurement rigor, and analytic honesty. The agency accepts that the trial may be small, single-arm, or externally controlled, and in exchange it insists that the endpoint be objective, disease-relevant, meaningful to patients, and interpretable despite heterogeneity. This article maps how those principles have played out in three areas reviewers repeatedly wrestle with: choosing endpoints for heterogeneous populations, building composite endpoints, and grounding approval in biomarkers.

Endpoint selection when the population is heterogeneous

FDA's framing guidance, Rare Diseases: Considerations for the Development of Drugs and Biological Products, tells sponsors to select endpoints from the disease's own manifestations and course, the specific target population, and the aspects of disease that matter to patients and caregivers across stages and severities 57. The central caution is that one endpoint rarely fits an entire rare disease: endpoint validity, sensitivity, reliability, and interpretability can differ between mild, early, slowly progressive disease and severe, late, rapidly progressive disease, so a clinical outcome assessment (COA) appropriate for some patients may be inappropriate for others 5051. Broad enrollment across stages or phenotypes is often practical, but FDA wants sponsors to evaluate how those differences affect power and interpretation and to discuss the population and enrollment criteria for a given endpoint with the review division 5051. Where only a subgroup is enrolled, FDA asks for a plan to evaluate other subgroups so generalizability to the broader disease can be assessed 61.

Two methodological preferences recur. First, characterize the heterogeneity before designing the endpoint: natural-history and registry data should map genotypic and phenotypic subtypes, involved organ systems, severity, and rate of deterioration, and those data should drive inclusion criteria, disease stage treated, trial duration, assessment frequency, and endpoint choice 58. Second, preserve information in small samples: FDA advises against dichotomizing continuous measures whenever possible, because the loss of information is especially costly when N is small 50.

The review record shows reviewers acting on exactly these instructions. In Mepsevii (vestronidase alfa) for MPS VII, the sponsor used a multi-domain responder index, but the clinical reviewer judged several components (forced vital capacity, visual acuity, shoulder flexion) inappropriate given heterogeneity in presentation and progression, and instead treated the six-minute walk test and selected BOT-2 balance and running-speed/agility measures as interpretable, noting 6MWT precedent in related enzyme-replacement programs 132. In Myalept (metreleptin) for lipodystrophy, FDA found that combining multiple phenotypes created heterogeneity that compromised interpretation, and warned that further subgrouping produced "tiny" samples in which valid between-group comparisons were impossible; absent a control group, only consistent data and a dramatic effect could support inference 119. Firazyr (icatibant) in hereditary angioedema drew the opposite-direction caution: FDA warned against extrapolating low-variability biomarker data from homogeneous healthy volunteers to the heterogeneous patient population, and pressed for population-PK and mixed-effects PK/PD modeling to avoid an underpowered Phase III 116.

Composite and multi-domain endpoints: two distinct practices

The reviews reveal two clearly separable uses of "composite," and FDA treats them very differently.

Pattern 1: confirmatory composite responder endpoints that require improvement plus protection against worsening. These are multiplicity-controlled primary endpoints in which a patient counts as a responder only if they improve on one axis and do not deteriorate on others, usually with rules that classify treatment failure, rescue, or discontinuation as non-response.

  • Ilaris (canakinumab) in the periodic fever syndromes (FMF, HIDS/MKD, TRAPS) used a complete-responder composite: resolution of the index flare at Day 15 and no new flare through Week 16, where resolution required both clinical control (physician global assessment <2) and inflammatory-marker control (normal or ≥70%-reduced CRP), and dose escalation, placebo escape, or early discontinuation counted as non-response 8.
  • Sylvant (siltuximab) in multicentric Castleman disease used durable tumor and symptomatic response, so tumor shrinkage alone was not sufficient: complete response required disappearance of disease plus resolution of MCD-related symptoms sustained at least 18 weeks 9.
  • Benlysta (belimumab) in SLE used the SLE Responder Index-4 (SRI-4), requiring a ≥4-point SELENA-SLEDAI reduction, no new BILAG A or two new BILAG B organ scores, and no PGA worsening, so that global improvement is demonstrated while ruling out organ-specific or overall deterioration 19. Notably, FDA flagged that the SRI was derived from exploratory analyses after a failed Phase 2 and was not validated before the pivotal trials, and that the clinically important between-treatment difference in SRI responder rate was unknown 1920.
  • Saphnelo (anifrolumab) in SLE used SRI-4 in one Phase 3 trial and BICLA in the other, and FDA explicitly contrasted these multiplicity-controlled composites against FACIT-F fatigue, which was only exploratory because it was not multiplicity-adjusted 23.

Of these, Ilaris and Sylvant are the rare and ultra-rare confirmatory-composite precedents; Benlysta and Saphnelo are SLE programs, not ultra-rare, and are cited only for how FDA treats composite structure, multiplicity control, and pre-validation, not as population-size precedents.

Pattern 2: multi-domain COA batteries whose weight depends on measurement validity, meaningful-change thresholds, anchors, and multiplicity control. Here FDA does not accept the battery as confirmatory unless the psychometric and statistical scaffolding is in place.

  • Yorvipath (palopegteriparatide) in hypoparathyroidism carried two multiplicity-controlled HPES secondary endpoints for which FDA accepted content and construct validity but required further work on reliability and responsiveness, and evaluated meaningful within-patient change using anchor-based analysis supplemented by empirical CDF and probability-density curves 5.
  • Ogsiveo (nirogacestat) in desmoid tumor used several separate, multiplicity-adjusted multi-domain secondary endpoints (worst pain, Desmoid Tumor Symptom Scale, Desmoid Tumor Impact Scale physical functioning, EORTC QLQ-C30 domains) with PGI-S/PGI-C anchors, rather than one combined composite 6.
  • Loargys (pegzilarginase) in arginase-1 deficiency assessed mobility and adaptive behavior across the Functional Mobility Assessment and VABS-II domains, but FDA noted no multiplicity control was planned, limiting these to non-confirmatory interpretation 1.
  • Voxzogo (vosoritide) in achondroplasia treated its multi-domain HRQoL instruments (PedsQL, QoLISSY, WeeFIM-II) as exploratory, non-multiplicity-controlled endpoints with no labeling claims 17.

The through-line: a composite earns confirmatory status when it is prespecified, multiplicity-controlled, and structured to prevent a favorable signal in one domain from masking harm in another. A multi-domain battery without that structure is read as supportive at best.

Biomarker-based evidence and surrogate endpoints

FDA's guidance defines a surrogate endpoint as generally a biomarker not itself a direct measure of clinical benefit but intended to predict it, and sorts biomarkers by evidentiary strength: a validated surrogate can support traditional approval, a surrogate reasonably likely to predict benefit can support accelerated approval, and a biomarker that merely changes with disease cannot support approval at all 158. Whether a surrogate is "reasonably likely" is a case-specific judgment resting on biological plausibility plus empirical evidence linking disease, endpoint, and effect; pharmacologic activity alone is insufficient, and the strongest support comes from data showing that the magnitude of surrogate change tracks the magnitude of clinical improvement across interventions 149162.

Crucially for ultra-rare programs, FDA acknowledges the strongest correlation data are often unavailable, and says it may then weigh the convergence of other evidence (preclinical models, epidemiology, and available clinical data) to decide whether the surrogate is reasonably likely to predict benefit 140149. In select genetic-disorder programs where the mechanistic link between target and surrogate is compelling, particularly some gene therapies, FDA has signaled that clinical correlation data linking the surrogate to outcome may be sparse and that the totality of evidence can still support a reasonably-likely finding, while confirmatory clinical evidence is still expected 140. The agency points sponsors toward the Rare Disease Endpoint Advancement (RDEA) Pilot Program to collaborate on novel surrogate or intermediate clinical endpoints intended for accelerated approval 140.

The review record shows the biomarker-as-surrogate pathway in action, and also its contentiousness:

  • Exondys 51 (eteplirsen) in Duchenne muscular dystrophy was granted accelerated approval on skeletal-muscle dystrophin production as reasonably likely to predict benefit. The rationale was biologic (near-absence of dystrophin is the proximal cause of DMD), but the decision split the agency: the review team found roughly 1%-of-normal dystrophin too low and the evidence insufficient, while CDER leadership concluded the small increase met the accelerated-approval standard 100.
  • Viltepso (viltolarsen) in exon-53-amenable DMD followed the established antisense precedent, using epidemiologic/pathophysiologic evidence that dystrophin amounts like those produced are associated with the milder Becker phenotype, with FDA reiterating that pharmacologic activity alone is not enough 102.
  • Qalsody (tofersen) in SOD1-ALS was approved on plasma neurofilament light chain (NfL) reduction as a reasonably likely surrogate, supported by NfL's association with progression and survival and by correlation between NfL reduction and slower clinical decline, with SOD1 protein reduction as confirmatory mechanistic evidence 96.
  • Avlayah (tividenofusp alfa) for the neurologic manifestations of MPS II was approved on cerebrospinal-fluid heparan sulfate reduction (91.4% at Week 24) as reasonably likely to predict neurologic benefit, with downstream CSF GM2/GM3 and NfL as corroboration, and residual uncertainty over which patients benefit left to the confirmatory trial 107.

Two related primary biliary cholangitis programs, Iqirvo (elafibranor) and Livdelzi (seladelpar), were randomized trials in a rare rather than ultra-rare disease, but illustrate the composite biochemical surrogate construct that Exondys 51, Viltepso, Qalsody, and Avlayah applied in small populations: an alkaline phosphatase threshold plus a ≥15% ALP reduction (to guard against spontaneous fluctuation) plus a bilirubin criterion (to ensure no worsening), each accepted as reasonably likely with confirmatory studies required 10193. And Vitrakvi (larotrectinib) shows a different biomarker role entirely: NTRK fusions defined a tissue-agnostic population rather than serving as the surrogate outcome, with approval resting on durable objective response 97.

Handling heterogeneity in the analysis, not just the design

Even with a well-chosen endpoint, small heterogeneous samples constrain what statistics can honestly claim, and reviewers have repeatedly downgraded analyses the data cannot support. In Strensiq (asfotase alfa) for juvenile-onset hypophosphatasia, FDA regarded the natural-history/historical-control approach as weaker than a randomized trial and said effectiveness would depend substantially on clinical judgment and graphical individual-patient profiles rather than usual statistical rigor, adding that historical controls persuade most when untreated disease is predictably uniform, eligibility and assessments match, the endpoint is objective, and the treated outcome is markedly different 118. Soliris (eculizumab) in a 15-patient PNH cohort leaned on repeated-measures longitudinal estimation and exact binomial confidence intervals rather than large-sample subgroup claims 130. Pedmark (sodium thiosulfate) pooled heterogeneous pediatric tumor types for event-free-survival monitoring and expressly labeled log-rank and Cox survival comparisons exploratory, because power sufficed only to detect a very large adverse difference 137.

FDA's guidance on external controls frames the boundary conditions: it generally prefers randomized concurrent controls, reserves external-control designs for well-characterized natural history with highly predictable morbidity or mortality and a large, self-evident treatment effect, and warns that many external-control studies cannot credibly demonstrate effectiveness, in which case a more suitable design should be chosen 828684. Where external controls are used, the effect should be large relative to likely bias and disease variability, objectively measured, collected to minimize bias, temporally associated with treatment, and consistent with expected pharmacology 68. For self-controlled trials of cell and gene therapies in small populations, FDA stresses reliable baseline data and objectively measured, non-effort-dependent endpoints, especially when blinding is absent 72.

What this means for a rare-disease program

The consistent message across guidance and reviews is that FDA compensates for small, heterogeneous populations by tightening endpoint quality rather than loosening the evidence standard. Practically, that means: characterize disease heterogeneity through natural history before locking an endpoint, and expect to justify that the endpoint behaves across the enrolled subgroups 5850; prefer objective, non-effort-dependent, patient-meaningful endpoints and avoid dichotomizing continuous data 5072; if using a composite, make it multiplicity-controlled and structured so improvement cannot mask worsening, and do the COA validation and meaningful-change/anchor work up front 8195; and if building on a biomarker surrogate, assemble the convergent biological, epidemiologic, and clinical evidence that the marker is reasonably likely to predict benefit, and plan the confirmatory trial from the outset 14910096. Early engagement with the review division and, where relevant, the RDEA pilot is FDA's recommended route for the novel-endpoint and novel-surrogate questions these programs inevitably raise 14051.

Want Rhizome's help on your own question? Try it for free.