Rhizome

Human Factors Deficiencies in FDA Premarket Submissions: Recurring Patterns Across Device Types

Human factors and usability engineering findings are a leading driver of additional-information requests on medical device premarket submissions. For regulatory teams, an HFE deficiency late in review can mean repeating a validation study, restructuring a use-related risk analysis, or delaying clearance by months — so knowing where FDA reviewers most often find gaps is directly actionable when planning a submission.

The analysis below draws on FDA's human factors guidance and device- and biologic-level review practice to identify the deficiency patterns that recur across submissions: gaps in the chain from use-related risk analysis to critical tasks, unrepresentative validation studies, weak root-cause analysis of test findings, and reports that present data without reasoning. It then examines how these patterns differ by device type.

Want to ask Rhizome your own regulatory questions? Try it for free.

Human factors deficiencies in FDA premarket submissions: the recurring patterns and how they differ by device type

Human factors engineering (HFE) and usability engineering (UE) findings are among the most common reasons FDA issues an additional-information (deficiency) request on a medical device or combination-product submission. FDA's guidance does not publish a single ranked tally of "most common" deficiencies, but read together, the agency's human factors guidance and device- and biologic-level reviews expose a consistent set of failure modes. They cluster around four things: an incomplete chain from use-related risk analysis to critical tasks, validation studies that are not realistic or representative enough to generalize, weak root-cause analysis of what testing actually found, and reports that document data without documenting reasoning. This article summarizes those patterns and how they play out across device types.

What FDA means by a human factors deficiency

FDA evaluates the user interface, defined broadly to include the device hardware, displays, alarms, packaging, labels, and instructions for use (IFU), against the intended users, uses, and use environments. A deficiency is any gap that prevents a reviewer from concluding the interface is adequately designed for safe and effective use. The agency expects the human factors information in a submission to explain how HFE/UE was applied and summarize the evaluations performed, not simply provide raw data 64. Submissions that supply data without the analysis, conclusions, and residual-risk reasoning behind it are a frequent trigger for questions 6472.

The recurring deficiency categories

Across FDA's guidance, the same categories of problem surface repeatedly in HF/UE submissions and validation analyses.

  • Use errors, close calls, and use difficulties on critical tasks. FDA expects the analysis to focus on problems found during testing, including observed use errors and close calls on critical tasks plus participant-reported difficulties, and then to analyze each for severity and root cause 124126127131.
  • Poor or incomplete root-cause analysis. This is one of FDA's most emphasized expectations: reviewers want the analysis to determine which part of the user interface was involved and how the interaction led to the error, drawing on both observed behavior and subjective participant data 124125126127131. Failing to determine which interface element caused an error, or why it occurred from the participant's perspective, is a recurring shortcoming 125126128.
  • User interface inadequacies that confuse users. Confusing interactions and confusing GUI elements recur as examples; FDA notes participants may report an aspect of the interface as confusing or difficult even when no actual error occurred 124125126.
  • Inadequate feedback, cues, or device response. Examples include unexpected device operation and a mismatch between audible cues and hold-time requirements that causes confusion and use error 125126.
  • Problems with displays, labels, and information presentation. FDA specifically calls out difficulty reading a display and misinterpreting, not noticing, or not understanding a label, and treats labels and instructions as part of the user interface 126134.
  • Difficult alarm or alert comprehension and response. Difficulty hearing an alarm is a use difficulty, and tasks that require users to respond to alerts or alarms are treated as critical tasks because incorrect response can cause harm 126131142.
  • Awkward or difficult physical manipulation. Awkward manual manipulations recur as a subjective usability problem, and changes affecting the user's physical interaction can impact critical tasks 126130.
  • Unacceptable or unaddressed residual use-related risk. FDA expects submitters to assess whether design, labeling, or training needs modification to reduce use-related risk, and to provide a residual-risk analysis with a rationale for the controls that were accepted 72127.

Where the analysis breaks: use-related risk analysis and critical tasks

Many downstream problems trace back to a weak use-related risk analysis (URRA). FDA expects the URRA to be a comprehensive, systematic evaluation of all use steps, covering user tasks and knowledge tasks, identifying the errors users might commit and the potential clinical consequences, and incorporating known problems with similar products and their mitigations 153156160. A common shortcoming is not covering all of these elements.

Task analysis is the engine of that process: it should break device use into discrete tasks and analyze each for the interface components involved, possible use errors, and resulting harm 148151. An incomplete task analysis that does not adequately support identification of use-related hazards or critical tasks is a frequent deficiency 148151.

Critical (essential) tasks are those that, performed incorrectly or not at all, would or could cause serious harm; they should be identified from the URRA and categorized by severity 149154157. The recurring failures here are procedural rather than cosmetic: not clearly explaining the process used to identify critical tasks, not listing and describing them, or not tying them to severity of harm and use scenarios 149154157. In short, FDA expects a complete, traceable chain from task analysis and URRA to task categorization and critical-task identification; missing traceability, incomplete task coverage, or weak harm/severity rationale are the main shortcomings 149150153154156157.

Where the study fails: summative (validation) test design

When human factors validation is required, the most common design deficiencies FDA cites are about realism and representativeness.

  • Participants who do not represent intended users. The single most important consideration is that participants represent the population of intended users 128.
  • Sample size not justified by preliminary analyses. FDA expects validation sample size to be derived from preliminary analyses and evaluations rather than a one-size-fits-all number 128.
  • Test conditions that are not realistic enough to generalize. Simulated-use testing must be realistic enough to reflect actual use, with realism scaled to the risks tied to the intended use, users, environments, and interface 128169.
  • Use-environment factors left out. Environmental factors that affect interaction, such as dim lighting, competing alarms, distractions, and multitasking, should be built into the simulated environment 128.
  • Facilitator interference. Users should interact as independently and naturally as possible, without the facilitator influencing them 128.
  • "Think aloud" during validation. FDA states think-aloud is not acceptable in validation testing because it does not reflect actual use behavior 128.
  • Unrealistic training and labeling. Training in the summative study should approximate what actual users receive, and if labeling or help lines are available in real use they should be available in the study, with participants free to use them as they choose 128172173174. FDA also warns that knowledge gained through training decays over time, so training should not be relied on as the primary risk control 169.
  • Critical tasks omitted, or testing that is not sensitive enough. Validation should include all critical tasks and be sensitive enough to capture use problems arising from interface inadequacies even when users are unaware they erred 169173.
  • Rushed or incomplete testing. Testing should use the final user interface under expected use conditions so results generalize to actual use 169.

The documentation gap: report completeness and category rationale

A large share of additional-information requests are about what the report does not say. FDA organizes marketing-submission human factors content into three categories, and each has its own rationale requirement 65:

HF submission categoryWhat FDA expectsCommon deficiency
Category 1Justify that modifications do not affect HF considerations; describe any leveraged prior HFE/UE evaluations 6567No justification for the category, or reliance on prior evaluations that are not described or cross-referenced 6567
Category 2Provide a rationale for why there are no critical tasks, or why validation is not being submitted 6569No rationale for the absence of critical tasks or for omitting validation 65
Category 3A comprehensive HFE/UE report including human factors validation testing 6566Missing report elements (intended users/uses/environments/training, UI description, known use problems, preliminary analyses, URRA) or missing validation results 6566

At a minimum FDA expects a conclusion and high-level summary stating the interface has been found adequately designed for the intended users, uses, and environments, and identifying the HF submission category with supporting rationale 6465. Providing raw data instead of a summary of evaluations, processes, issues, resolutions, and conclusions is itself a documented problem area 64. FDA notes that including appropriate human factors information up front may reduce additional-information requests 68, and that sponsors can cross-reference previously submitted information rather than resubmit it 6567.

Residual risk, root cause, and labeling as a mitigation

FDA expects a clear line from hazards to specific risk controls, tests, manuals, or training, with verification that the controls are effective 184; a failure to show this traceability is a common deficiency 184. The agency also expects a qualitative analysis that aggregates observation, task-performance, and interview data to determine the root causes of use errors and to prioritize additional controls 126128.

Labeling and IFU are legitimate controls, but FDA is skeptical of over-reliance on them. Recurring labeling deficiencies include warnings or precautions that are not conspicuous, are overused ("overwarning"), are not placed near the relevant procedural step, or are otherwise not effective for the intended audience 186189192193. The design hierarchy matters: risk controls should prioritize eliminating or mitigating risk through device design where feasible, rather than relying primarily on labeling or training 96. Where residual use-related risk remains after validation, FDA expects it to be analyzed, stated as acceptable or not, and the interface further modified if warranted 91184185187.

How the patterns differ by device type

Highest-priority device types carry a default expectation

FDA takes a risk-based approach to whether a submission needs human factors data. For device types on its highest-priority list, a submission should include human factors data unless it does not change users, user tasks, user interface, or use environments relative to the predicate 133196. That priority list includes ablation generators; anesthesia machines; artificial pancreas systems; auto injectors; automated external defibrillators; duodenoscopes and gastroenterology-urology endoscopic ultrasound systems with elevator channels; hemodialysis and peritoneal dialysis systems; implanted and external infusion pumps; insulin delivery systems; home-use negative-pressure wound therapy; robotic catheter manipulation systems; robotic surgery devices; ventilators; and ventricular assist devices 196. For device types not on the list, human factors data should be included when the risk analysis shows that incorrect or omitted tasks could cause serious harm; FDA may also request data case-by-case for PMAs and De Novos, new or different user interfaces, different intended users, recalls or adverse events tied to use error, and high-risk modifications 133196.

The practical pattern: high-priority devices (infusion and insulin delivery, dialysis, ventilators, AEDs, robotic surgery, autoinjectors) draw the closest scrutiny of validation adequacy, while lower-risk and unchanged devices more often turn on whether the sponsor justified the absence of critical tasks or the leveraging of prior data.

Combination products and drug-delivery devices: the delivery task is the risk

For combination products, FDA expects the interface to support safe use of the whole product, including packaging, labels, IFU, and training, with explicit attention to medication errors and to user limitations such as reduced vision or hearing, low literacy, cognitive decline, or constrained home and emergency settings 81969799. Autoinjectors and pens generate characteristic use-related issues that map directly to critical tasks: inability to remove a cap (delaying administration), not fully activating the injector (missed or under-dose), and hold-time problems where high viscosity or a longer injection time makes completion harder and risks underdose 889293. Changes in injection time, angle, tissue plane, dose accuracy, activation force, or cap-removal force can all affect validation or performance 899095. When a product moves from a prefilled syringe to an autoinjector, or the indication, user population, injection site, or drug changes, FDA expects an assessment of whether the change affects critical tasks and may require additional validation 90102.

For generic drug-device combinations submitted in an ANDA, FDA expects applicants to minimize user-interface differences from the reference listed drug (RLD) early in development and to run threshold analyses: a labeling comparison, a comparative task analysis, and a physical comparison of the device constituent parts 103113114115. Whether comparative-use human factors study data are needed is decided case-by-case based on the threshold-analysis results and the specific differences identified 113114. The core concern is that patients and caregivers, who are less experienced than clinicians, face increased use-error risk when external critical design attributes differ, especially on substitution for the RLD 113114.

What the review record actually shows for combination-product BLAs

Combination-product reviews give the most concrete picture of how these deficiencies read in practice:

  • ALTUVIIIO: the validation study did not include a distinct caregiver user group despite caregivers being intended users, and did not separate injection-naive from injection-experienced users; FDA noted both as deviations but said they did not preclude review given similarity to a marketed product and hemophilia patients' familiarity with IV therapy 2. FDA also identified four use errors and three close calls on the critical task "user removes air from the syringe," flagging air embolism, no dose/underdose, and possible hemorrhage as consequences 3.
  • ESPEROCT: the study did not include an arm of untrained patients/caregivers; the sponsor justified this on the basis that hemophilia patients are trained before self-administration, and FDA accepted the rationale, treating it as a design limitation rather than an observed error 610. Overall FDA found the studies acceptable with no reported use errors, close calls, or need for administrator assistance 5.
  • ADZYNMA: the DMEPA reviewer found some draft IFU descriptions acceptable only if the intended users are health care professionals, a labeling/usability limitation tied to user qualification 78.
  • FESILTY: DMEPA concluded a validation study was not needed but asked that the URRA be updated to better capture the clinical impact of identified use errors or task failures 13.
  • PREVNAR 20 and MRESVIA: validation was either deemed unnecessary because HCPs are familiar with the tasks and packaging mirrors a marketed product, or was conducted to validate labeling and packaging without reported use errors 4912.

The through-line: the most frequent combination-product findings are not exotic hardware failures but user-group representativeness (were caregivers and naive users included?), critical-task use errors during preparation and injection, and IFU adequacy tied to who the intended user actually is 236713.

Devices cleared through 510(k)/De Novo

In device reviews, human factors work is typically described as a use-related risk analysis (URRA/UFMEA) to identify critical tasks, formative testing to iterate the interface, and summative validation in a simulated-use environment with representative end users 333435394345464748505152535455. The device categories that most often involve human factors validation are surgical devices (especially robotic or assisted surgical systems), clinical software and decision-support/viewer tools, diabetes and glucose-management and insulin-related devices, home monitoring/caregiver devices, and interventional/implant delivery systems 33353639434445464748505152535455. In many cleared submissions, sponsors document either that prior validation showed the interface change did not affect critical tasks 51, or that no critical tasks existed so a final validation study was not required 41, with residual issues addressed through training or labeling 3842454853.

Practical takeaways for submission strategy

The deficiency patterns point to a small number of high-leverage steps. Build a complete, traceable URRA and task analysis so every critical task is identifiable and tied to severity of harm 148149153154157. Design validation studies around realism and representativeness: representative users (including caregivers and naive users where relevant), a justified sample size, a realistic use environment, the final interface, no facilitator interference, and no think-aloud 128169. Analyze what testing found, not just that it occurred: determine the interface root cause of every use error and close call and decide whether residual risk is acceptable 124126128184. Prefer design controls over labeling and training, and where labeling is used, make it conspicuous, targeted, and placed at the relevant step 96186189192. Finally, write the report to FDA's expected structure with an explicit category determination and rationale, because a large share of additional-information requests are about missing reasoning rather than missing data 646568.