Risk analysis for software-containing devices forces a question that hardware reliability engineering never had to answer: how do you assign a probability of occurrence to a defect that is deterministic once the triggering conditions arise? FDA's answer shapes the structure of the risk file a submission is judged on — whether severity alone can carry a risk estimate, what a hazard analysis must show in place of a failure rate, and how mitigations are expected to be traced and verified. Teams that import a hardware FMEA approach into a software risk analysis often find it does not match what FDA expects on this point.
The analysis below traces FDA's stated position across the premarket software guidance and its interaction with ISO 14971's probability-times-severity definition of risk, including how the agency reconciles the two and what it expects in a risk management file when probability cannot be estimated. It then examines how the question has actually been handled in 510(k) summaries and De Novo decisions, drawing on the guidance record and decision documentation, with every point linked to its source.
Want Rhizome's help on your own question? Try it for free.
Software failure probability in device risk analysis: FDA's position under ISO 14971 and the premarket software guidance
The core premise: software failures are systematic, not random
FDA's foundational position is that software does not fail the way hardware does. Hardware components fail randomly and their failure rates can be characterized statistically; software failures are systematic, arising from latent design and coding defects that manifest deterministically under the right conditions. Because of that, FDA states that "for software, failures tend to be systematic in nature and therefore the probability of occurrence of a software failure cannot be determined using traditional statistical methods." 7 This single premise drives everything downstream in how FDA expects software risk to be treated.
The practical consequence is spelled out in the same guidance: where the overall probability of harm cannot be estimated, risk estimation should not rest on a manufactured probability figure. "While it may be possible to estimate the probability for other events in the sequence, if the overall probability of occurrence of harm cannot be estimated, the estimation of risk should be based on the severity of harm alone." 7 In other words, if you cannot credibly quantify how often the software will fail, you fall back to how bad it is when it does.
Reconciling this with ISO 14971
ISO 14971 defines risk as the combination of the probability of occurrence of harm and the severity of that harm, and risk estimation as the assignment of values to both. 66 FDA recognizes ISO 14971 as the framework for device risk management and expects manufacturers to work within it: identify hazards, estimate and evaluate the associated risks, implement controls, monitor their effectiveness, and judge the acceptability of individual and overall residual risk. 54 In premarket review, FDA may scrutinize the manufacturer's decisions on risk estimation, evaluation, acceptability, controls, and overall residual risk. 55
The tension is obvious: ISO 14971's risk model is two-dimensional (probability x severity), but FDA says the probability dimension is often unreliable or unavailable for software. FDA resolves this not by abandoning ISO 14971 but by acknowledging that probability is not always the meaningful or practicable factor. It already takes this view for adjacent problems, most explicitly use errors, where FDA notes that probability is very difficult to determine because many use errors cannot be anticipated until simulated and observed, so severity of potential harm becomes the more meaningful basis for deciding whether to eliminate or reduce harm. 58 A use-related risk analysis, in FDA's view, may focus on resulting harm and need not include estimated occurrence rates. 65 The same logic is applied to software. Where evidence genuinely exists, such as electromagnetic-interference event reports or experience with similar devices, FDA does expect that information to inform probability estimates. 57 The point is not that probability is banned, but that it must not be invented.
The worst-case rule: set software failure probability to 1
The 2023 Content of Premarket Submissions for Device Software Functions guidance (which superseded the 2005 software guidance) is explicit about the mechanics. FDA warns that "it is often difficult to adequately estimate the probability of software failures that could contribute to a hazardous situation," and that "applying unrealistically low probability estimates to software failures could result in unrealistic risk evaluation and subsequently lead to inappropriate risk control measures." 26 The recommended response is to stop trying to guess the number: "in some instances it may be prudent to focus on the identification of potential software functionality and failures that could result in hazardous situations instead of estimating probability." 26
When that worst-case approach is taken, FDA gives a concrete convention: "considering a worst case probability is appropriate, the probability for the software failure occurring should be set to 1." 26 This is a conservative classification device, not a claim that real-world software failure is certain. It forces the analysis onto the severity axis and the causal chain from failure to harm. FDA specifically directs manufacturers to document the "severity of the harm resulting from the hazardous situation," and to make the assessment before risk controls are applied. 26
Importantly, this is not an absolute rule that probability must always equal 1. The Off-The-Shelf Software guidance does not itself impose probability = 1 in every case; it defers to the companion software guidance and determines documentation based on device-level risk, including the probable risk from possible OTS software failures or flaws. 13914 FDA sets probability to 1 only in the circumstances where probability is not estimated and a worst-case posture is warranted. 26
Documentation Level as the operational risk threshold
The 2023 guidance replaced the old Minor/Moderate/Major "Level of Concern" terminology with a two-tier Documentation Level. 33 The distinction is itself severity-driven and applied before risk controls. Enhanced Documentation applies when a failure or a flaw of any device software function could create a hazardous situation with a "probable risk of death or serious injury" to a patient, user, or others in the use environment, assessed in the context of intended use and including foreseeable misuse and cybersecurity compromise; "probable" excludes purely hypothetical risks. 29 Basic Documentation applies when that Enhanced criterion is not met. 29 The determination reflects the device as a whole in the context of its intended use, not a software module viewed in isolation. 29 Older device-specific guidance still uses the Minor/Moderate/Major framing (for example, photobiomodulation software limited to on/off or timer functions being minor versus software controlling treatment parameters being moderate), with the rationale to be justified by the possible consequences of software failure. 30 Current updates generally map former Moderate examples to Basic and former Major examples to Enhanced, while stressing that the actual determination is device-specific. 3335
Alongside this, FDA expects a risk management file in the submission developed under an FDA-recognized version of ISO 14971. 41 The risk management plan should define individual risk-acceptability criteria, the need for risk reduction, and the method for evaluating overall residual risk after controls are implemented and verified, including how residual risk is weighed against the benefits of intended use. 41 For device software, the risk assessment should cover risk analysis, risk evaluation, risk control, and, where applicable, benefit-risk analysis, addressing the whole system and hardware environment when the software is integrated. 26
Production and quality-system software: a distinct standard
For software used in production or the quality management system, rather than software that is part of the device, FDA's Computer Software Assurance guidance draws a deliberate line between its risk-based analysis and ISO 14971 medical-device risk analysis. The analysis should consider "reasonably foreseeable (as opposed to likely)" failures rather than only probable ones, and a high process risk turns on whether failure "may result in a quality problem that foreseeably compromises safety, meaning a medical device risk." 59 The draft version was even more explicit that such software failures do not occur in a probabilistic manner where likelihood can be estimated from historical data or modeling; the analysis should instead address reasonably foreseeable failures and their resulting risks. 68 The through-line with the premarket guidance is the same: for software, "foreseeable" displaces "likely."
How 510(k) reviews have actually handled it
The 510(k) record shows no single universal convention. Summaries fall into three groups.
Summaries that estimate or categorize probability. Many device software risk analyses do assign a likelihood alongside severity, sometimes with detectability, and reassess after mitigations:
- IRMA Blood Analysis System (Diametrics Medical): the software hazard table rates inaccurate-algorithm hazards for calibration, sample calculations, system operation, incorrect IR reading, and temperature/barometric checks each as "Improbable" likelihood and "Marginal" severity after validation and QC, using likelihood bands from Frequent to Incredible and severity from Negligible to Catastrophic. 108
- ICEPlex C. difficile Assay Kit / ICEPlex System (PrimeraDx): one high-risk hazard was initially assigned a low probability of occurrence, reduced to very low after mitigation, with severity and mitigations documented for hardware and software hazards. 109
- BacT/ALERT VIRTUO Microbial Detection System (bioMerieux): each failure and cause was assessed before and after mitigation for severity, probability as appropriate, and detectability, with residual risks minor or moderate. 112
- cobas CT/NG v2.0 Test (Roche Molecular Systems): an FMEA ranked hazards by severity, occurrence (defined as potential failure frequency), and detectability, classifying risks Green/Yellow/Red with hardware, software, or user-process mitigations. 117
- Vitrea CT Multi-Chamber Cardiac Functional Analysis (Vital Images): added features were judged non-critical to the risk profile because the probability of harm after mitigations was "Improbable." 132
Summaries that assume failure and focus on severity. For software safety classification specifically, some sponsors apply the worst-case convention directly. The B.O.L.T patient monitoring system (Amzetta Technologies) states in its summary: "Probability of a software failure shall be assumed to be 1." 119 The decision flow then asks whether a hazardous situation can arise, whether the resulting risk is unacceptable, and what severity of injury is possible, distinguishing non-serious injury (Class B) from serious injury or death (Class C), crediting only risk controls external to the software system; the device's software was classified Class B (Moderate). 119
Summaries that reveal nothing. Many 510(k) summaries simply state that risk analysis and controls were completed without disclosing whether probability was quantitatively estimated, so the method cannot be inferred from those documents alone.
The net picture: 510(k) reviewers accept both a categorized probability-plus-severity FMEA 108109112117132 and the conservative probability-equals-1, severity-driven classification. 119 Which appears depends on the risk framework the manufacturer chose, not a fixed FDA mandate.
How higher-risk reviews (De Novo and level-of-concern) have handled it
De Novo classification decisions, which establish special controls for novel device types, treat software failure as a hazard-to-clinical-harm problem and operationalize the probability and severity judgments through special controls rather than a single numerical threshold.
- Quantitative FMEA scoring still appears: the Infrascanner Model 1000 FMEA assigned 1 to 10 ratings for occurrence, severity, and detection, multiplied into a Risk Priority Number, and required further reduction for RPN above 100, which its software failure modes did not exceed. 161
- Severity is framed by the downstream clinical consequence of a software or algorithm failure: delayed or incorrect treatment from erroneous output (Acumen HPI 165), fluid overload or missed treatment (AFM 168), incorrect diagnosis from inaccurate quantitative imaging (Caption Interpretation 166), and false prognostic outputs and delayed diagnosis (BrainSee 170). At the high end, Dexter L6 states that a software failure or latent flaw could directly cause severe injury or death, supporting a "Major" level-of-concern, Enhanced-documentation posture. 179
- Empirical failure rates complement the hazard analysis where they exist: Natural Cycles reported a method failure rate of 0.6 per 100 woman-years, reflecting a "green" (non-fertile) reading on a fertile day followed by pregnancy. 182
- Where residual probability is genuinely low, lower-tier controls can suffice: for Dexcom's STUDIO on the Cloud, incorrect analysis could lead to acute hypo- or hyperglycemia, hospitalization, death, or chronic complications, yet FDA found verification, validation, and design controls made malfunction risk "very low," and classified the device Class I under general controls. 181185
For Class II software De Novos, these judgments are enforced through special controls that consistently require software verification, validation, and hazard analysis/risk assessment 165167169170177184; algorithm and technical characterization of inputs, outputs, parameters, and patient population 165180; independent and representative performance testing including subgroup and data-quality characterization and indeterminate-output rates 164175180; usability and labeling that preserve clinician judgment and address overreliance 168177; and, where relevant, change-control and postmarket performance management 175178180. The common structure targets the causal chain from software or algorithm or input failure, to incorrect or absent output, to clinician use or misuse, to patient harm, precisely the sequence the premarket guidance says to analyze when probability is set aside.
What this means in practice
For a regulatory affairs team building a software risk file, the throughline is consistent across the guidance and the review record. Work within ISO 14971, but do not fabricate a software failure probability. Where the probability of harm cannot be credibly estimated, estimate risk on severity alone 7, and if you take the worst-case route, set the software failure probability to 1 and let severity and the failure-to-harm chain drive the analysis. 26 Expect the pre-control severity of a software failure, not a computed likelihood, to determine your Documentation Level and the depth of evidence FDA asks for. 29 A quantitative, categorized FMEA remains acceptable when the occurrence estimates are defensible 108117, but reviewers are equally comfortable with, and higher-risk decisions often rest on, a severity-first analysis backed by verification, validation, characterization, and special controls. 119165180