Critical Quality Attributes and Analytical Similarity in FDA Biosimilar 351(k) Reviews
For teams planning a 351(k) biosimilar program, the analytical similarity package carries the largest share of the evidentiary burden, and how FDA has evaluated that package in prior reviews directly shapes attribute risk-ranking, tiering decisions, and the methods a sponsor must bring to bear. Understanding the agency's recurring expectations helps regulatory and clinical teams scope comparability studies, anticipate reviewer scrutiny, and calibrate how much residual uncertainty clinical data will need to resolve.
This analysis reads across approved 351(k) reviews to trace how FDA has approached critical quality attribute identification, statistical tiering, and the orthogonal analytical methods used to compare a proposed biosimilar to its US-licensed reference product. It maps the recurring structure of these assessments and illustrates each element with examples drawn from the FDA review record.
Want to ask Rhizome your own regulatory questions? Try it for free.
Critical quality attributes and analytical similarity in FDA 351(k) biosimilar reviews: what recurs across approved products
Analytical similarity is the foundation of every 351(k) biosimilar demonstration. Under the "totality of the evidence" standard, FDA reviewers expect the structural and functional comparison of the proposed biosimilar to the US-licensed reference product to carry the largest analytical burden, with clinical pharmacology and comparative clinical studies confirming that any residual uncertainty is not clinically meaningful. Reading across approved 351(k) reviews, a consistent methodology emerges: attributes are identified and ranked by criticality, sorted into statistical tiers, compared with a standardized panel of orthogonal analytical methods, and any observed differences are adjudicated against mechanism of action and functional potency. This article summarizes the recurring elements of that approach, with examples drawn from the FDA review record.
The criticality ranking that drives everything downstream
Before any statistics are run, the sponsor risk-ranks each quality attribute for its potential impact on activity, PK/PD, safety, and immunogenicity, and FDA scrutinizes that ranking closely. In the bevacizumab-awwb (Mvasi) review, FDA laid out the general expectation plainly: attributes are assigned to tiers based on criticality/risk, with the highest-risk attributes going to Tier 1, lower-risk to Tier 2, and lowest-risk to Tier 3, and the assignment must consider not only criticality but also the levels of the attribute, assay sensitivity, and the limitations of the statistical analysis 144146.
The criticality ranking is where reviewers most often push back on sponsors. Two reviews show FDA resisting attempts to soften a high-risk classification:
- For trastuzumab-dkst (Ogivri), the applicant's risk-ranking tool assigned criticality based on impact and uncertainty, and could downgrade an attribute when uncertainty was low. The FDA reviewer explicitly disagreed, stating that clearly high or very high risk attributes should remain high/very high criticality rather than being reduced because uncertainty was low 139.
- For adalimumab-atto (Amjevita), FDA rejected a one-size-fits-all statistical approach, reserving Tier 1 equivalence testing for clinically relevant attributes and those tied to mechanism of action, and re-assigning some comparisons to higher tiers than the applicant proposed. For certain orthogonal aggregate measures FDA considered Tier 3 visual criteria more appropriate than the applicant's proposed quality-range approach 183186.
The practical lesson is that criticality ranking is a review-managed exercise, not a sponsor declaration, and FDA will move attributes between tiers and change the statistical treatment applied to them.
The three-tier statistical framework
Approved reviews consistently apply the same three-tier structure to the analytical similarity dataset. The clearest articulations appear in the etanercept-szzs (Erelzi) and adalimumab-adbm (Cyltezo) reviews:
- Tier 1 (equivalence testing) for the very high / high-risk attributes tied to the clinical mode of action. The proposed biosimilar must fall within an equivalence margin derived from reference-product variability, typically evaluated as a 90% two-sided confidence interval for the mean difference lying within the margin 86113135137.
- Tier 2 (quality ranges) for lower-impact attributes, expressed as the reference-product mean plus/minus a multiple of the standard deviation. Erelzi used mean plus 3 SD 86; Cyltezo described the multiplier as typically 2 to 4 113; Mvasi described Tier 2 as "mean +/- X" with a justified range 144. A high percentage of biosimilar lots (FDA cited "more than 90%" for Cyltezo) must fall within the range 109125.
- Tier 3 (graphical / visual comparison) for the lowest-risk attributes or those not amenable to quantitative analysis, or present at very low levels, assessed by side-by-side raw-data or orthogonal graphical comparison 86113144.
The tier multiplier itself is treated as a criticality signal. In the rituximab-abbs (Truxima) review, FDA noted that the Tier 2 quality-range multiplier was generally 3, but the CD20 binding CELISA and the ADCC reporter assay used a tighter multiplier of 2 because of the higher criticality of those attributes; FDA found the justifications acceptable except for one exception involving afucosylated species 97. Equivalence margins in Tier 1 are anchored to reference-product lot variability: Truxima set its margin from Rituxan lot variability in the CDC and ADCC potency assays 97, and Cyltezo used 1.5x the reference-product SD, with FDA applying confirmatory analysis to confirm the 90% CIs fell within the prespecified margins 113117125.
FDA's independent statistical review typically concentrates on the Tier 1 mechanism-of-action attributes. Examples of the Tier 1 attributes selected:
- Bevacizumab-awwb: % relative potency by proliferation-inhibition bioassay and VEGF-A binding by ELISA 142.
- Adalimumab-atto: apoptosis-inhibition bioassay (potency) and soluble TNF-alpha binding 184.
- Adalimumab-adbm: soluble TNF-alpha neutralization and TNF-alpha binding by SPR 113117.
- Etanercept-szzs: TNF-alpha binding by SPR and TNF-alpha neutralization by NF-kB reporter gene assay 8790.
- Trastuzumab-dkst: relative potency and relative ADCC activity, plus HER2 binding 136138.
FDA also corrects the sponsor's statistics when needed. In the Mvasi review, FDA accepted the applicant's general approach but corrected the handling of sample-size imbalance in its own analysis before concluding the Tier 1 comparisons passed equivalence 141142.
The recurring analytical method panel
Across products, the same categories of orthogonal methods appear, matched to attribute class. The filgrastim-sndz (Zarxio) and adalimumab-atto (Amjevita) reviews together illustrate the standard toolkit:
| Attribute class | Representative methods cited in FDA reviews |
|---|---|
| Primary structure | N-terminal sequencing; peptide mapping with UV and MS/MS; intact mass by ESI-MS and MALDI-TOF; DNA sequencing of the construct 147152 |
| Higher order structure | Far- and near-UV circular dichroism; 1H NMR and 1H-15N HSQC NMR; LC-MS of disulfide bonds 147152 |
| Charge variants | Cation-exchange HPLC (CEX-HPLC); IEF; AEX/CEX 147191 |
| Size variants / aggregates | SEC / SE-HPLC with light scattering; reduced and non-reduced SDS-PAGE or CGE; AUC sedimentation velocity; field-flow fractionation; dynamic light scattering; light obscuration (HIAC); microflow imaging (MFI); 90-degree light scattering 147152191 |
| Glycosylation | N-linked oligosaccharide profiling for afucosylation, high-mannose, sialylation, galactosylation 8991109 |
| Binding / potency | Cell-based bioassays (e.g., NFS-60 proliferation for filgrastim 152; apoptosis-inhibition and A549 NF-kB assays for adalimumab 182184); SPR receptor binding 147152 |
| Fc effector function | ADCC (reporter and PBMC-based), CDC, ADCP; FcgammaRIIIa / FcgammaRIIa and C1q binding 9597109125 |
For antibody products, the functional and Fc-effector panel is central because it interrogates mechanism of action directly, which is why these assays are the ones most often placed in Tier 1 or given tighter Tier 2 multipliers 97138.
Similarity findings that recur, and how FDA adjudicates the differences
Approved biosimilars are not identical to their reference products, and the reviews are candid about the differences. The recurring pattern is that FDA identifies minor quantitative differences, then judges them against functional data and mechanism of action rather than rejecting the application.
Glycosylation, afucosylation, and Fc effector function. This is the most common site of observed difference for monoclonal antibodies. In the infliximab-dyyb (Inflectra) review, FDA found subtle shifts in glycan composition and product-related variants (including H2L1) that produced a roughly 20% reduction in Fc gamma receptor binding versus the reference, with correspondingly reduced in vitro activity, yet still concluded CT-P13 was analytically similar across primary sequence, TNF binding, structure, potency, and effector function, with the WEHI potency bioassay and TNF-binding ELISA meeting Tier 1 equivalence 105. In the etanercept-szzs review, small differences in afucosylation and high-mannose forms were noted but considered not clinically relevant, supported in part by observed PK similarity 8991. For adalimumab-adbm, minor Tier 2 differences in N-linked oligosaccharides (fucosylation, glycosylation, sialylation), charge variants, size variants, and aggregates sat just outside the 3 SD range and were mirrored in Tier 3 graphical analysis, but FDA judged the residual uncertainty mitigated by functional potency assays (CDC, ADCC, ADCP) that showed no meaningful differences 109125.
Charge and size variants. These appear repeatedly as Tier 2 attributes with small offsets. Beyond the Cyltezo example above, the Amjevita review documented that some attributes failed acceptance criteria between EU- and US-licensed Humira (main-peak charge variants, non-glycosylated heavy chain, and % heavy chain + % light chain), and that these differences appeared in greater magnitude for the biosimilar relative to US Humira, yet the functional data still supported high similarity 185.
Disulfide and structural variants affecting potency. The etanercept-szzs review is the clearest case of a Tier 1 failure that was resolved analytically. GP2015 initially did not meet statistical equivalence for TNF-alpha neutralization by the NF-kB reporter assay. FDA attributed the shortfall to a post-peak hydrophobic product-related impurity caused by wrongly bridged disulfide bonds (notably the T7 peptide) that reduced potency; the sponsor showed the variants could refold under mild redox conditions, and a computed potency model adjusted for T7 levels brought the products into equivalence. FDA concluded the difference did not preclude high similarity 8992.
Impurities and process-related species. Small quantitative differences in impurities recur. In the filgrastim-sndz review, FDA noted slightly more truncated variants (about 0.5 to 0.7%), a small difference in N-leucine species (about 0.8%), and minor oxidation differences, all judged unlikely to affect activity, alongside equal or lower high-molecular-weight and sub-visible particle levels 159162165167169170. In the etanercept-szzs review, host cell proteins were somewhat higher in the biosimilar, but reviewers noted the HCP assay reagents were specific to the biosimilar's own expression system, complicating direct comparison 93.
The consistent adjudication logic is: locate the difference, tie it to (or exclude it from) mechanism of action, and confirm with functional potency and effector-function assays that the difference is not expected to be clinically meaningful. Truxima's review captured the endpoint most reviews reach, that the totality of evidence supported high similarity "except for minor components" 104.
The two comparators and the analytical bridge
Because many sponsors run pivotal clinical studies against the EU-approved reference product, the analytical package doubles as a "scientific bridge" among three products: the proposed biosimilar, US-licensed reference, and EU-approved reference. Reviews therefore report pairwise equivalence across all three arms. Mvasi ran ABP 215 vs US Avastin, ABP 215 vs EU bevacizumab, and EU bevacizumab vs US Avastin 141142; Cyltezo compared 13 biosimilar lots against 55 US and 86 EU Humira lots for TNF-alpha neutralization, and 13 vs 43 US and 53 EU lots for TNF-alpha binding, with all pairwise comparisons meeting equivalence 107108112114; and the Amjevita review noted the analytical similarity results supported a sufficiently robust bridge to justify EU-approved Humira as the clinical comparator 183184185. Trastuzumab-dkst similarly tested 22 US-Herceptin, 22 EU-Herceptin, and 9 biosimilar lots for HER2 binding by flow cytometry 136.
What this means for a 351(k) analytical strategy
Reading the approved record, several practical patterns hold across products:
- Expect FDA to manage the criticality ranking and to refuse downgrades of clearly high-risk attributes on uncertainty grounds 139183.
- Reserve Tier 1 equivalence testing for mechanism-of-action attributes (binding and potency, plus effector function for antibodies), and anchor the margin to reference-product lot variability 8797113137142.
- Justify Tier 2 quality-range multipliers by criticality; a tighter multiplier signals a more critical attribute 97.
- Anticipate minor differences in glycosylation, charge and size variants, and impurities, and pre-build the functional and effector-function data package that lets FDA conclude those differences are not clinically meaningful 105109125.
- Where a Tier 1 attribute misses, a mechanistic root-cause explanation plus supporting analytics can still support a high-similarity conclusion, as the etanercept-szzs disulfide/T7 case shows 8992.
- Plan the three-way analytical bridge if the clinical program uses a non-US comparator 141183.