Topical dermatology sponsors planning a new chemical entity have to settle several pivotal-trial questions early: the choice of vehicle comparator, randomization ratio, responder endpoint definition and maximal-use pharmacokinetic package. Recent FDA approvals of novel topical creams are the closest available precedent for how the agency judges those choices, and review files show where reviewers pushed back even when efficacy against vehicle was clear.
The analysis below compares the Phase 3 programs behind roflumilast cream (Zoryve), tapinarof cream (Vtama) and ruxolitinib cream (Opzelura) across plaque psoriasis, atopic dermatitis and nonsegmental vitiligo. It sets out the shared design template, the primary and key secondary endpoints, and the statistical, patient-reported outcome and systemic-exposure questions FDA reviewers raised during review.
Want Rhizome's help on your own question? Try it for free.
Vehicle-controlled pivotal trials for novel topical dermatology drugs: roflumilast, ruxolitinib and tapinarof creams
Three recent novel topical creams used almost the same pivotal template: roflumilast cream 0.3% (Zoryve) and tapinarof cream 1% (Vtama) for plaque psoriasis, and ruxolitinib cream 1.5% (Opzelura) for atopic dermatitis and later nonsegmental vitiligo. Each program ran two identically designed, randomized, double-blind, vehicle-controlled Phase 3 trials, with a skewed randomization ratio that put more patients on active drug. The psoriasis and atopic dermatitis programs used a static investigator or physician global assessment (IGA/PGA) "success" endpoint; the vitiligo program used F-VASI75. FDA's review files show that reviewers spent less time on whether the drugs beat vehicle, which they clearly did, and more time on four other questions. How robust was the result to missing data? Were site effects hidden? Were secondary and patient-reported endpoints clinically interpretable? And how far does systemic exposure under maximal-use conditions extend beyond the body surface area (BSA) studied in the controlled trials?
The shared design template
| Design element | Roflumilast cream 0.3% (psoriasis) | Tapinarof cream 1% (psoriasis) | Ruxolitinib cream 1.5% (atopic dermatitis) | Ruxolitinib cream 1.5% (vitiligo) |
|---|---|---|---|---|
| Pivotal trials | ARQ-151-301 and ARQ-151-302, identical Phase 3, parallel-group, double-blind, vehicle-controlled 87 | DMVT-505-3001 and DMVT-505-3002, identical multicenter Phase 3, US and Canada 4144 | TRuE-AD1 (Study 303) and TRuE-AD2 (Study 304) 34 | TRuE-V1 and TRuE-V2 (Studies 306 and 307) 154163 |
| Randomization | 2:1 active to vehicle, stratified by site, baseline IGA and intertriginous involvement 87 | 2:1 active to vehicle 4041 | 2:2:1 (0.75% BID, 1.5% BID, vehicle BID) 24 | 2:1 active to vehicle 154 |
| Regimen and controlled duration | Once daily, 8 weeks 87 | Once daily, 12 weeks 4044 | Twice daily, 8 weeks 2 | Twice daily, 24 weeks 154 |
| Disease extent | 2% to 20% BSA, IGA of 2 or more, PASI of 2 or more 87 | 3% to 20% BSA; PGA 2 to 4, with mild and severe disease each capped at about 10% of enrollment 41 | 3% to 20% BSA excluding scalp; IGA 2 or 3 34 | Facial depigmentation of 0.5% BSA or more, non-facial involvement of 3% or more, total up to 10% BSA 154 |
| Primary endpoint | IGA 0/1 plus a 2-grade or greater improvement at Week 8 87 | PGA 0/1 plus a 2-grade or greater improvement at Week 12 4041 | IGA 0/1 plus a 2-grade or greater improvement at Week 8 45 | F-VASI75 at Week 24 154163 |
| Randomized sample | 439 and 442 86 | 510 and 515 (1,025 total) 44 | 631 and 618 (1,249 total) 2 | 674 total 154 |
Two features of the template recur in every program. First, the composite responder definition (clear or almost clear and at least a 2-grade improvement) is the same for all three products' global-assessment endpoints 87404. Second, every program followed the controlled period with open-label or re-randomized long-term treatment. Roflumilast had open-label extensions with exposure of at least 52 weeks in 195 subjects 112. Tapinarof had a 40-week open-label extension with 763 rollover subjects 50. Ruxolitinib in atopic dermatitis had a 44-week double-blind extension in which vehicle patients were re-randomized to active treatment 23. Ruxolitinib in vitiligo had a 28-week all-active extension 154.
Primary efficacy results against vehicle
| Program | Trial | Active | Vehicle | Difference |
|---|---|---|---|---|
| Roflumilast 0.3% | ARQ-151-301 | 41.5% | 5.8% | 39.7 points (95% CI 32.4 to 47.0) 137 |
| Roflumilast 0.3% | ARQ-151-302 | 36.7% | 7.1% | 29.5 points (95% CI 21.5 to 37.6) 137 |
| Tapinarof 1% | DMVT-505-3001 | 36% | 6% | 29 points (95% CI 22 to 36) 164 |
| Tapinarof 1% | DMVT-505-3002 | 40% | 6% | 34 points (95% CI 27 to 41) 164 |
| Ruxolitinib 1.5% (AD) | TRuE-AD1 | 53.8% (136/253) | 15.1% (19/126) | 38.9 points 11 |
| Ruxolitinib 1.5% (AD) | TRuE-AD2 | 51.3% (117/228) | 7.6% (9/118) | 44.1 points 11 |
| Ruxolitinib 1.5% (vitiligo) | TRuE-V1 | 29.9% | 7.5% | 22.5 points (95% CI 14.2 to 30.8) 150 |
| Ruxolitinib 1.5% (vitiligo) | TRuE-V2 | 29.9% | 12.9% | 16.9 points (95% CI 7.8 to 26.0) 150 |
All primary comparisons were statistically significant. The roflumilast reviewer noted that the treatment effect was larger in ARQ-151-301 than in ARQ-151-302, but found the applicant's and reviewer's analyses very similar 137. Vehicle response rates stayed in the single digits to mid-teens across programs.
What FDA reviewers questioned
Missing data and tipping-point robustness
Missing-data handling was the most consistent statistical theme across the three programs.
- Tapinarof. Week-12 primary-endpoint missingness was 16% to 22% across arms and trials. The primary analysis used 100 multiple imputations, and the review also examined nonresponder imputation, LOCF, observed cases and tipping-point analyses 19. In DMVT-505-3001, significance could be lost only under an extreme assumption: very high proportions of missing vehicle outcomes imputed as successes and missing tapinarof outcomes as failures. DMVT-505-3002 stayed significant even in the worst case 19. Reviewers also noted that the reasons for discontinuation differed by arm. Adverse-event discontinuations were higher with tapinarof (6% in each trial vs 0% and 1% on vehicle), while subject-request withdrawals were higher on vehicle (11% and 10% vs 6% and 4%) 27.
- Roflumilast. Week-8 missingness was 11% and 9% on active vs 14% on vehicle in both trials, with COVID-19 accounting for only 1% to 2% 86. The primary method was regression-based multiple imputation with predictive mean matching 132. The statistical reviewer ran a worst-case analysis (missing active outcomes set as nonresponses, missing vehicle outcomes as responses). Roflumilast stayed superior in both trials, so no tipping-point analysis was needed 86. During development, FDA had advised that at least two alternative imputation methods based on different assumptions be used 134. FDA also requested executable SAS programs for the multiple imputation and the tipping-point analysis 82.
- Ruxolitinib (atopic dermatitis). FDA recommended nonresponder imputation plus at least two sensitivity analyses and a tipping-point analysis 143140. Even the extreme tipping-point scenario remained nominally significant, with all sensitivity p-values below 0.002 139.
- Ruxolitinib (vitiligo). This was the one program where the most extreme tipping-point scenario was not nominally significant in either study. Reviewers still judged the F-VASI75 result robust because losing significance required implausible assumptions: vehicle missing-data responders at least 38 percentage points higher than ruxolitinib missing-data responders 141. The reviewers also examined changes to the analysis plan. The SAP switched the primary missing-data method from the protocol's nonresponder imputation to multiple imputation. It also initially narrowed the primary population to patients with at least one post-baseline assessment, but the applicant restored the all-randomized population after FDA comments 146. FDA found the submitted imputation program inconsistent with the SAP (30 vs 10 imputations) and asked for corrected programs 142. Reviewers also found small SAS desktop-versus-server differences in imputed estimates. They concluded these did not affect conclusions for the labeled endpoints 145147.
Site effects, stratification and data integrity
- For roflumilast, FDA advised that site-to-site variability be assessed using the original sites, and cautioned that merging sites could mask site effects 82. The SAP nonetheless pooled sites with fewer than 10 randomized participants 136. The reviewer presented site-level IGA-success plots showing variation across sites 131. The reviewer also criticized discrepancies between IVRS and eCRF baseline stratification data, stating that these "should not occur in well-designed clinical trials." Efficacy was reanalyzed using eCRF values, with very similar results 137.
- For ruxolitinib, FDA had questioned before the trials whether stratifying by region alone was adequate, noting that center-level stratification can account for center-to-center variability 143. During review, one Study 304 site (41 subjects) was removed from efficacy analyses for serious GCP source-documentation noncompliance. Excluding it did not change the primary conclusion 144. At a separate inspected site, FDA found documentation lapses (non-original ECGs with identifiers added after copying, and incomplete investigational product storage-access records). Primary IGA and EASI data were verified without discrepancies and no Form FDA-483 was issued 61.
- For tapinarof, FDA chose sites for inspection because of high site efficacy effects, high enrollment, high discontinuation rates or no prior GCP inspection history 28. At one DMVT-505-3002 site, percent-BSA values in the EDC differed from source records. Monitors had told the investigator to round to whole numbers even though the protocol did not require rounding 30. Week-12 PGA data were verified, and OSI judged overall data quality adequate 2830.
Clinical meaningfulness of endpoints
Reviewers repeatedly asked whether secondary and patient-reported endpoints could be interpreted in the populations actually enrolled.
- PASI in mild psoriasis. FDA reiterated its concern that PASI-based endpoints "may not be clinically meaningful" in mild psoriasis, which mattered because the roflumilast population included mild-to-moderate disease 82. The approved label reports IGA, intertriginous IGA (I-IGA) and itch (WI-NRS) results and no PASI-75 results 127.
- Subset endpoints. FDA questioned whether the roughly 20% of roflumilast subjects per arm with baseline I-IGA of 2 or more would be enough for a meaningful intertriginous assessment 82. It also said qualitative labeling claims would depend on robust data and endpoint content validity 138. The final I-IGA analysis sets were small (63/32 and 53/31 active/vehicle) but showed large, significant differences 135.
- EASI-75 and PROMIS in atopic dermatitis. The ruxolitinib clinical reviewer questioned how to interpret EASI-75 in patients with very low baseline EASI (some as low as 0.6). The reviewer also found a flaw in the applicant's PROMIS sleep threshold: a baseline score of 6 or more on a scale running from 8 to 40 could include patients unable to improve by 6 points. The reviewer reanalyzed PROMIS in patients with a baseline score of 14 or more 144.
- T-VASI50 in vitiligo. Reviewers accepted F-VASI75 as clinically relevant, citing patient-focused drug development input that the face matters most to patients. They raised concerns about the precision of the VASI scale and about whether a 50% reduction in total-body VASI is meaningful 156. Clinical Outcome Assessment reviewers concluded that T-VASI50 was clinically meaningful at Week 24 based on anchor-based analyses and exit interviews. They also noted that the Voice of the Patient report said many patients preferred about 80% to 100% repigmentation 156, and that the studies lacked a patient global impression of severity scale 158.
Multiplicity and SAP changes
Multiplicity control was generally accepted. Examples are the sequential gatekeeping for tapinarof 20 and the hierarchical testing of ruxolitinib key secondary endpoints in both indications 144145. The notable exception was roflumilast. The sponsor corrected an error in the multiple-testing procedure by memo to file after database lock but before unblinding, and did not tell FDA until the NDA was submitted. The reviewer therefore treated IGA success at Week 4 as post hoc rather than prespecified 136135.
COVID-19 disruptions
For tapinarof, FDA rejected an observed-case analysis that excluded virtual Week-12 PGA assessments as the only COVID sensitivity analysis. FDA instead asked for an ITT analysis that treats virtual assessments as missing and imputes them with the prespecified method, plus dataset flags for COVID-related missed and virtual visits 24. In the end the impact was small. Fewer than 1% of subjects had COVID-related missing data, and results were very similar when alternative-visit data were treated as missing 2521. The roflumilast review likewise used an mITT sensitivity population that excluded Week-8 assessments missed because of COVID-19 126.
Long-term extensions and durability
- FDA agreed to the ruxolitinib atopic dermatitis extension but questioned how it would be interpreted. Patients could enter regardless of IGA or BSA, including poor responders who would continue ineffective treatment, so FDA recommended rescue or withdrawal criteria after Week 8 143.
- For vitiligo, FDA said a 30-day post-treatment follow-up was too short to assess durability of repigmentation. It asked for longer off-treatment follow-up and a way to assess maintenance treatment 162. T-VASI75 at Week 24 was not statistically significant. Rates kept rising after Week 24, but all patients were on active drug by then, so these Week-52 analyses were exploratory 156147.
- The tapinarof reviewers noted that open-label extensions without a placebo arm are inherently harder to interpret 51.
Systemic exposure: where the vehicle-controlled data stop
For each product, reviewers used maximal-use PK studies to judge risk beyond the BSA studied in the pivotal trials.
- Ruxolitinib. In the atopic dermatitis maximal-use trial, patients with 40% or more BSA involvement had about 14-fold higher Day-1 Cmax and AUC than those below 40% 78. One subject with 90% BSA reached exposures in the range reported after oral ruxolitinib 79. Because the Phase 3 trials covered only up to 22% BSA, reviewers considered the proposed 20% BSA labeling limit reasonable 64. The label carries a Boxed Warning and extensive Warnings and Precautions language on systemic JAK-inhibitor exposure 63. For vitiligo, reviewers kept the 60 g/week limit, noting that more than 97% of subjects used less than that 69. Vitiligo also brought an application-site acne signal (5.8% vs 0.9% on vehicle through Week 24) that had not appeared in the atopic dermatitis program 70.
- Roflumilast. Adult maximal-use exposure exceeded published oral roflumilast 500 mcg exposure (AUC 2.2-fold higher for the parent drug and 1.7-fold higher for the N-oxide) 99. The reviewer noted that maximal-use dosing (mean 6.5 g in adults) was far above Phase 3 use (2.24 g) and concluded that actual clinical use would most likely not increase risk 116. Reviewers looked specifically at oral PDE-4 class events: weight loss, psychiatric events and gastrointestinal events. No new signal emerged 98110. FDA required drug-interaction and renal and hepatic impairment assessments because of the higher-than-oral exposure 84.
- Tapinarof. Systemic exposure was low and declined by Day 29, with no accumulation 56. More than 95% of Phase 3 PK samples were below the limit of quantification 48. The main safety issue was local: folliculitis in 19% and contact dermatitis in 5% of treated patients during the controlled trials (20% and 7% in the label's grouped-term table), with similar rates in the extension 50. Reviewers concluded that labeling and routine pharmacovigilance were adequate 52.
All three programs carried pediatric postmarketing requirements for long-term safety studies with maximal-use PK components 1015368. For ruxolitinib, reviewers pushed back on a short infant follow-up proposal. They noted that a prompt response to an acute flare does not rule out chronic atopic dermatitis in infants, so it does not answer the long-term safety question 73.
Implications for topical development programs
Several points from these reviews apply directly to sponsors planning vehicle-controlled topical trials:
- Pre-specify at least two alternative missing-data approaches and a tipping-point analysis, and submit executable code. FDA asked for this in all three programs 13414382. Differential discontinuation between active and vehicle arms, such as adverse-event dropouts on active versus withdrawal by subject request on vehicle, will be examined 27.
- Don't let site pooling or stratification errors hide site effects. FDA's roflumilast advice on keeping original sites 82 and its ruxolitinib question about region-only stratification 143 show reviewers expect site-level evaluability. IVRS versus eCRF discrepancies will draw comment 137.
- Check that secondary endpoint thresholds can be reached from the enrolled baseline. The EASI-75 and PROMIS critiques 144, the PASI-in-mild-disease concern 82 and the I-IGA subset question 82 each came from a mismatch between the endpoint and the baseline population.
- Report any change to the multiple-testing procedure to FDA before unblinding. An unreported correction can downgrade an endpoint to post hoc 136135.
- Match the labeled BSA and weekly-quantity limits to the controlled-trial experience. Maximal-use data showing exposure near oral levels at high BSA shaped ruxolitinib's labeling 787964.
- Build durability and extension designs that reviewers can interpret. That means rescue or withdrawal criteria in extensions 143 and a meaningful off-treatment follow-up when durability matters, as in vitiligo 162.