Rhizome

Preclinical Toxicity Endpoints for Large-Molecule Drugs (FDA)

Chetan Mishra
Chetan Mishra
Sep 30, 2025

Defining an adequate nonclinical safety package is one of the earliest and most consequential decisions a development team makes for a biologic therapeutic. Regulators expect sponsors to apply a science- and risk-based rationale to endpoint selection — one that accounts for species specificity, immunogenicity, and the exaggerated pharmacology inherent to large molecules — rather than defaulting to the standard small-molecule paradigm. Gaps or misalignments in this package can delay IND acceptance, trigger clinical holds, or require costly additional studies late in development.

The analysis below maps the toxicity endpoint categories that FDA, operating under ICH S6(R1) and the broader ICH S-series framework, expects sponsors to address for biotechnology-derived pharmaceuticals. It covers general toxicity, safety pharmacology, toxicokinetics, reproductive and developmental toxicity, immunotoxicity, and carcinogenicity considerations, along with the study design adaptations specific to large molecules.

Want Rhizome's help on your own question? Try it for free.

Toxicity endpoints in preclinical testing of large-molecule (biologic) drugs for FDA approval

Large molecules (therapeutic proteins, monoclonal antibodies, fusion proteins, and related biotechnology-derived pharmaceuticals) are evaluated under a different nonclinical safety paradigm than small-molecule drugs. FDA reviews these programs against ICH S6(R1) and the broader ICH safety (S-series) framework, supplemented by product-class FDA guidances. The endpoints are largely the same categories a reviewer expects for any drug (general toxicity, safety pharmacology, toxicokinetics, reproductive toxicity), but the study designs, species choices, and some endpoint sets are adapted to the biology of large molecules: species specificity, immunogenicity, exaggerated pharmacology, and the fact that these products are not expected to interact with DNA. This overview summarizes the toxicity endpoints a sponsor is expected to generate.

Governing framework and its philosophy

ICH S6(R1) sets a flexible, case-by-case, science-driven approach rather than a fixed checklist 1. The stated objectives of the nonclinical program are to identify an initial safe human dose and dose-escalation scheme, to identify potential target organs of toxicity and whether effects are reversible, and to identify safety parameters for clinical monitoring 1. Standard rodent/dog designs used for small molecules are often not appropriate because of the species specificity, immunogenicity, and other unique properties of biopharmaceuticals 23. Studies are generally conducted under GLP, with any justified deviations documented and their significance assessed 2.

Species selection: the decision that drives every endpoint

For large molecules, endpoint interpretation is only meaningful in a pharmacologically relevant species, one in which the receptor or epitope is expressed so the product is pharmacologically active 1. Relevance is established by comparing target sequence homology across species, in vitro target binding affinity and receptor/ligand occupancy, and functional activity in species-specific cell systems or in vivo pharmacology, ideally supported by modulation of a pharmacodynamic marker 1. Using a non-relevant species can produce misleading toxicity findings 7.

Practical consequences for the study package:

  • Safety evaluation should typically include two relevant species, but one relevant species can suffice when justified; non-relevant species are discouraged 3.
  • When no relevant species exists, homologous (surrogate) proteins or transgenic animals expressing the human target may be used 3.
  • Tissue cross-reactivity (TCR) studies with human tissue panels are recommended for antibodies and antibody-like products to characterize binding and detect unexpected binding, but binding alone does not prove in vivo activity and must be read in the context of the full package 40. TCR has limited value for choosing species and mainly supplements knowledge of target distribution 1.
  • Both sexes are generally used unless otherwise justified 3.

Core general (repeat-dose) toxicity endpoints

Repeat-dose general toxicity studies are the backbone of the package. They should reflect the intended clinical route and regimen and, when feasible, incorporate toxicokinetics 1. The standard endpoints collected are:

  • Clinical signs (in-life observations) 34
  • Body weight and food consumption 35
  • Clinical pathology: hematology and clinical chemistry 34
  • Organ weights 36
  • Gross pathology / necropsy 36
  • Histopathology 3436
  • Ophthalmology, where appropriate 34
  • Systemic exposure (toxicokinetics) integrated into the study whenever possible 14

Dose selection should establish a dose-response relationship including, where possible, a toxic dose and a no-observed-adverse-effect level (NOAEL); the high dose should be scientifically justified and may be set on pharmacology/exposure margins rather than classical maximum-tolerated-dose logic 12. Single-dose studies remain useful for dose-response characterization and to select doses for repeat-dose studies, and can carry safety pharmacology parameters 1.

Recovery and reversibility

A recovery (non-dosing) period is generally included to assess whether effects are reversible, persistent, progressive, or delayed 37. Full reversibility does not need to be demonstrated; a trend toward reversibility plus a scientific assessment that full reversal is likely is generally sufficient 383912. Recovery arms are most important when reversibility cannot be predicted and the toxicity is severe or clinically relevant 3812.

Study duration

For most biotechnology-derived pharmaceuticals, repeat-dose animal dosing has generally been 1 to 3 months; short-term human use may be supported by studies up to 2 weeks, and chronic indications generally by 6-month studies, with longer or shorter durations scientifically justified 1. Where two relevant species exist, both are typically used for general toxicology up to 1 month; if findings are concordant or the mechanism is understood, longer studies in a single species usually suffice 6. Six-month chronic studies are generally adequate, as longer durations have not typically added information 12.

Safety pharmacology endpoints (ICH S7A core battery)

Safety pharmacology evaluates unwanted effects on vital organ systems, investigated before first-in-human dosing 21. The core battery covers three systems:

  • Cardiovascular: blood pressure, heart rate, and electrocardiogram; supplemental in vitro/in vivo/ex vivo assessment of repolarization and conductance; follow-up may add cardiac output, ventricular contractility, and vascular resistance 1720.
  • Respiratory: respiratory rate plus function measures such as tidal volume or hemoglobin oxygen saturation (clinical observation alone is inadequate); follow-up may add airway resistance, compliance, pulmonary arterial pressure, and blood gases/pH 1720.
  • Central nervous system: motor activity, behavior, coordination, sensory/motor reflexes, and body temperature; follow-up may add learning/memory, neurochemistry, and electrophysiology 1720.

For large molecules, a key efficiency applies: when a product achieves highly specific receptor targeting, these endpoints are often adequately evaluated within the toxicology and/or pharmacodynamic studies, so standalone safety pharmacology studies can be reduced or eliminated 2021. A novel therapeutic class, or a product lacking highly specific targeting, warrants a more extensive safety pharmacology evaluation 21.

Toxicokinetic endpoints (ICH S3A)

Toxicokinetics is integrated into toxicity studies to measure systemic exposure and relate it to dose and time 22. The measured endpoints are exposure parameters rather than toxicity findings: AUC, Cmax, and concentration at a specified time (C(time)), typically from plasma, whole blood, or serum, and sometimes tissue or metabolite concentrations or unbound concentration 222425. Exposure is estimated across an appropriate number of animals and dose groups to support risk assessment and to interpret differences across species, doses, and sexes 26, and it underpins species/regimen selection and later study design 2223.

Immunogenicity (anti-drug antibody) endpoints

Because many large molecules are immunogenic in animals, anti-drug antibodies (ADAs) are measured in repeat-dose studies to support interpretation 8. ADAs should be characterized by titer, number of responding animals, and neutralizing versus non-neutralizing status, and correlated with pharmacological and toxicological changes 8. Their impact on interpretation is assessed particularly when there is altered pharmacodynamic activity, unexpected exposure changes without a PD marker, or evidence of immune-mediated reactions such as immune complex disease, vasculitis, or anaphylaxis, with attention to effects on PK/PD, complement activation, new toxicities, and immune-complex pathology 89. Two reviewer caveats matter: antibody detection alone should not trigger early study termination unless the response neutralizes activity in a large proportion of animals 8, and animal antibody formation does not predict human immunogenicity 89. FDA's therapeutic protein immunogenicity guidance similarly frames animal immunogenicity data as useful for interpreting animal toxicology and designing later studies rather than for predicting the human response, with particular relevance when the product inhibits an evolutionarily conserved endogenous protein 41.

Immunotoxicity and immunomodulation endpoints

Routine tiered immunotoxicity batteries are not recommended for biopharmaceuticals; instead, screening signals in standard studies drive any follow-up mechanistic work 12. Endpoints and signals include:

  • Whether an effect reflects immunosuppression or immunostimulation 10
  • Hematologic changes; changes in immune organ weights/histology (thymus, spleen, lymph nodes, bone marrow); unexplained changes in serum globulins/immunoglobulins; increased infections; and increased tumors as a possible sign of immunosuppression 11
  • Where developmental immune concerns exist, functional assessments such as lymphocyte subset number/function, immunoglobulin class changes, and T-cell-dependent antibody response (TDAR) at appropriate developmental stages 13
  • Cytokine release syndrome as a specifically recognized adverse phenomenon to be captured and summarized 16

Reproductive and developmental toxicity (DART) endpoints

ICH S5(R3) defines the reproductive stages to be covered, from premating and fertility through embryo-fetal development, parturition/lactation, and postnatal development to sexual maturity, traditionally addressed by fertility and early embryonic development (FEED), embryo-fetal development (EFD), and pre-/postnatal development (PPND) studies 2728. For large molecules, S6(R1) and S5(R3) adapt these designs to species specificity, immunogenicity, mechanism of action, and long half-life 12:

  • If both rodent and non-rodent are relevant, EFD is typically assessed in two species; if the rodent is not relevant, a single relevant non-rodent species may be used 29.
  • When the nonhuman primate (NHP) is the only relevant species, an enhanced pre- and postnatal development (ePPND) study can replace separate EFD/PPND studies; it doses from about gestation day 20 to birth and assesses pregnancy outcome, offspring viability, external malformations, skeletal effects, and visceral morphology at necropsy 2931.
  • Because mating studies are impractical in NHPs, fertility is instead assessed via reproductive-tract evaluation in repeat-dose studies, with specialized endpoints such as menstrual cyclicity, sperm parameters, and reproductive hormones added when needed 30.
  • Placental transfer differences should be considered in design and interpretation 30; where no relevant species exists, surrogate molecules or transgenic models may be used, and if none are available a scientific justification for not conducting in vivo DART is required 29.

Genotoxicity and carcinogenicity: why the standard batteries do not apply

For large molecules the standard small-molecule genotoxicity and carcinogenicity batteries are generally not used 3:

  • Genotoxicity: the usual assays are not applicable and not needed, because these products are not expected to interact directly with DNA or chromosomal material; targeted studies may be warranted only for specific concerns such as an organic linker in a conjugated (ADC) product, and standard assays are not appropriate for assessing process contaminants 3.
  • Carcinogenicity: standard 2-year bioassays are generally inappropriate 3. A product-specific assessment of carcinogenic potential may still be needed depending on clinical treatment duration, patient population, and biological activity (for example growth factors or immunosuppressants). When warranted, this can include examining receptor expression in normal and malignant human cells, testing whether the product stimulates growth of receptor-bearing cells, using relevant animal models, adding sensitive indices of cellular proliferation to long-term repeat-dose studies, and, if justified, a single rodent species study when the product is biologically active and non-immunogenic in rodents 333.

Endpoints that feed the first-in-human starting dose

A central output of the toxicity program is a safe starting dose. Standard nonclinical approaches (for example a NOAEL-based margin) are generally appropriate for therapeutic proteins and monoclonal antibodies 45. For products with agonistic properties, notably bispecific antibodies, FDA recommends considering a minimally anticipated biological effect level (MABEL) approach, drawing on in vitro and in vivo pharmacology to set the initial dose 45. FDA does not, in these documents, give a single universal rule for choosing between NOAEL and MABEL; the choice is product- and mechanism-dependent 45.

Summary of toxicity endpoints by study type

Study typeKey endpointsNotes for large molecules
Repeat-dose general toxicityClinical signs; body weight; food consumption; hematology; clinical chemistry; organ weights; gross pathology; histopathology; ophthalmology; toxicokinetics 3435361Core of the package; conducted in relevant species, both sexes, with recovery arms for reversibility 337
Single-dose toxicityDose-response; may carry safety pharmacology parameters 1Supports dose selection for repeat-dose studies
Safety pharmacology (S7A)Cardiovascular (BP, HR, ECG); respiratory (rate, tidal volume, SpO2); CNS (activity, behavior, reflexes, temperature) 1720Often folded into toxicology/PD studies when targeting is highly specific 2021
Toxicokinetics (S3A)AUC, Cmax, C(time); tissue/metabolite concentrations 222425Exposure margins underpin interpretation and starting dose
ImmunogenicityADA titer, incidence, neutralizing status; correlation with PK/PD and pathology 89Interpretive tool; not predictive of human immunogenicity 89
Immunotoxicity (S8)Immune organ weights/histology; hematology; immunoglobulins; infection/tumor signals; TDAR; cytokine release 10111316No routine tiered battery; signal-driven 12
Reproductive/developmental (S5(R3), S6)Fertility; EFD; ePPND endpoints (viability, malformations, skeletal/visceral) 272831ePPND in NHP when NHP is the only relevant species 2931
Genotoxicity / carcinogenicityGenerally not assessed by standard batteries; product-specific proliferation/receptor assessments if warranted 333Large molecules do not interact directly with DNA 3

Scope notes and limitations

This overview reflects the harmonized ICH framework FDA applies to biotechnology-derived pharmaceuticals plus selected FDA product-class guidances, as captured in the cited documents. The governing principle is that the endpoint set and study designs are tailored case by case; specialized product classes carry additional endpoints not detailed here (for example, antibody-drug conjugates raise linker/payload genotoxicity and additional immunogenicity and drug-interaction considerations 48, and juvenile/pediatric programs add developmental endpoints under ICH S11 713). The FDA source set did not provide a single universal formula for selecting between NOAEL- and MABEL-based starting doses 45. For a specific molecule, a reviewer would refine the endpoint list against the product's modality, target biology, indication, patient population, and dosing duration.

Want Rhizome's help on your own question? Try it for free.