1,401 reports. That is how many FDA device reports Reuters found tied to AI-enabled devices between 2021 and October 2025 when it searched the agency's MAUDE database. Only 115 mentioned a problem with software, algorithms or programming.
And even those 115 are signals, not 115 confirmed AI failures.
That gap between a report and a cause is the real story behind the recent headlines about AI in operating rooms. When a navigation screen shows an instrument in the wrong place, the record captures the outcome. It rarely captures which part of the system broke.
This article builds on a video breakdown of the Reuters investigation published on February 9, 2026. I checked the headline TruDI figures and the company responses against other published coverage, and I read the npj Digital Medicine study directly. Every injury claim below is an allegation unless an FDA action is cited.
What the TruDI Reports Show and What They Don't
TruDI is a surgical navigation system used in sinus and skull-base procedures. It tracks an instrument and shows the surgeon where the tip sits relative to the patient's anatomy. The eyes, the base of the skull, the brain and major blood vessels all sit close by, so the margin for error is measured in millimeters.
Its maker, Acclarent, added machine learning features in 2021. TruSeg could automatically identify and segment anatomical structures from medical images. TruPath could calculate a route between two selected points.
Before those features arrived, the FDA had received seven unconfirmed malfunction reports and one report of patient injury. From late 2021 through the period Reuters examined, the agency received at least 100 malfunction or adverse event reports, including at least 10 reported injuries through November 2025.
That jump looks alarming. Timing does not prove cause. TruDI contains far more than its machine learning features, and one report cannot say which component, if any, was responsible. Integra, which later acquired the product line, told Reuters that reports of this kind only show a TruDI system was in use when an event occurred.
Two patients have sued. Erin Ralph alleges that during a June 2022 sinus procedure, the system displayed an inaccurate instrument location, her carotid artery was injured and she suffered a stroke. Donna Fernihough alleges a similar injury and a same-day stroke after a May 2023 procedure involving the same surgeon, Marc Dean, and the same system.
Acclarent denies the allegations in both cases. Integra says there is no credible evidence connecting TruDI's AI features to the alleged injuries. Reuters could not independently review the protected medical records in Ralph's case. The lawsuits describe catastrophic outcomes, but they do not establish that machine learning caused them.
Two Recalls, Neither Pointing at the AI Features
The FDA has found that the navigation system itself could show a location that did not match reality. In 2023 it announced a Class II recall covering TruDI software version 2.3.1 and certain surgical curettes. The actual position of the curette tip could differ from the position on the screen. The listed consequences included prolonged surgery, cerebrospinal fluid leaks, visual impairment and damage to structures around the base of the skull. The FDA named a software design change as the underlying cause.
In 2024, 265 TruDI suction instruments were recalled because incorrect calibration could cause the same mismatch. This time the FDA described a process control problem.
Neither recall was attributed by the FDA to TruSeg or TruPath. Neither established what happened during the two surgeries in the lawsuits.
What they do show is a chain. Tracking hardware, calibrated instruments, conventional software, patient imaging and, in newer versions, machine learning tools all feed the final picture. When that picture is wrong, finding the broken link can take far more information than a surgeon in the room will ever have.
What the MAUDE Database Can and Cannot Tell You
MAUDE is the FDA database that collects reports of medical device deaths, injuries and malfunctions. It sounds like the obvious place to look for AI failures. The FDA itself warns that these reports cannot reliably establish why an event happened, how often a device fails, or whether one device is more dangerous than another. Reports can be incomplete or duplicated, some incidents are never reported, and the number of devices in use is often unknown.
Back to the Reuters count: 1,401 reports from 2021 to October 2025 involving devices on the FDA's list of AI-enabled products. Just 115 mentioned a software, algorithm or programming problem.
A 2024 study in npj Digital Medicine by Handley and colleagues shows why even that filter is blunt. The researchers reviewed 429 reports tied to AI and machine learning enabled devices and sorted them by how likely AI was involved.
| Classification | Reports | Share of 429 |
|---|---|---|
| Potentially AI-related | 108 | 25.2% |
| Unlikely AI-related | 173 | 40.3% |
| Not enough information | 148 | 34.5% |
More than a third could not be classified at all. A clinician filing a report may know exactly what appeared on the screen, but not what happened inside the system beforehand. Was the model wrong? Was the input corrupted? Did conventional software fail? Was an instrument miscalibrated? The report records the outcome. It does not diagnose the machine.
Two More Cases Where the Cause Is Disputed
The pattern repeats outside surgery.
In June 2025, a voluntary report to the FDA alleged that an FDA-cleared prenatal ultrasound assistance tool linked fetal structures to the wrong body parts. In some instances it allegedly indicated a structure had been visualized when it had not. No patient injury was reported. The company that markets the software said the report did not demonstrate a safety problem, and the FDA did not request corrective action.
The risk here is subtle. An obvious error makes a clinician stop and look again. A believable one can do the opposite, because a false "already seen" signal tells the clinician to move on.
Then there is Medtronic's AccuRhythm AI, which filters alerts from implantable cardiac monitors. In Medtronic's validation data, it cut false atrial fibrillation alerts by 74.1% while preserving 99.3% of true AF alerts. For pause alerts, it cut false positives by 97.4% and preserved 100% of true pauses. Those are strong numbers.
Reuters still found at least 16 reports alleging that real rhythm events or pauses were missed. No injuries were reported. After reviewing the cases, Medtronic said only one involved a missed abnormal rhythm and that the others concerned how information was displayed or interpreted, not a failure of AccuRhythm.
One case shows the attribution problem clearly. An FDA report first said AccuRhythm had classified a real pause as false. Medtronic's investigation concluded that an earlier detection algorithm inside the device had already rejected the event, so AccuRhythm never received it. The monitor still produced the wrong outcome. The AI first blamed apparently never saw the event.
How AI Devices Reach Patients and Who Watches Them
Most AI medical devices do not go through an approval built from scratch. Under the 510(k) pathway, a manufacturer shows that a new device is substantially equivalent to a legally marketed device, called a predicate. The FDA can still require performance testing and clinical evidence when needed, and the pathway exists so every incremental improvement does not need a full approval.
AI strains the idea of "incremental." A Lancet Digital Health study traced the predicate networks of 285 AI devices cleared between 2019 and 2021. It found that 32.6% ultimately descended from first-generation predicates with no AI or machine learning components at all. Researchers call this predicate creep. It does not make a device unsafe, but it raises a different question: how much evidence shows that the AI in use today performs safely on real patients?
The scale keeps growing. By the FDA's 2026 update, roughly 1,524 AI-enabled devices had been identified in the United States, with about three-quarters in radiology.
Oversight capacity is part of the Reuters story too. Reuters interviewed five current and former FDA scientists. It reported that the division of imaging, diagnostics and software had grown to roughly 40 AI specialists, and that around 15 were laid off or left during 2025. The Digital Health Center of Excellence reportedly lost about a third of a staff of around 30, and former employees said workloads for remaining reviewers had nearly doubled. The FDA and the Department of Health and Human Services pushed back, saying patient safety remains the agency's highest priority, AI-enabled devices undergo rigorous review, and the FDA continues to recruit digital health expertise.
My Take
Start with what the numbers cannot do. 100 TruDI reports do not prove the AI caused harm. 115 software-related reports do not prove 115 AI failures. And 148 of 429 reviewed reports could not be classified at all.
The data cuts the other way too. Neither recall was tied to the machine learning features, but that does not clear them. The recalls cover a software design change and instrument calibration. Nobody in the public record has isolated the AI features, and the reporting system was not built to do it.
That matters well beyond medicine. If a system cannot log which component produced which output, nobody can answer the question after something goes wrong. I looked at a related gap, benchmark scores running ahead of real tasks, in AI Self-Improvement Has 5 Levels.
MAUDE is the weakest link in this chain.
- Reuters found 1,401 MAUDE reports on AI-enabled devices from 2021 to October 2025. Only 115 mentioned software, algorithm or programming problems.
- In a 2024 npj Digital Medicine review of 429 reports, 34.5% had too little information to judge whether AI was involved.
- The FDA's 2023 and 2024 TruDI recalls were not attributed to its machine learning features.
- The TruDI lawsuits are allegations. Acclarent denies them and Integra says there is no credible evidence linking the AI to the injuries.
FAQ
What is the MAUDE database?
MAUDE is the FDA database of reports about medical device deaths, injuries and malfunctions. The FDA warns that the reports cannot reliably establish why an event happened, how often a device fails, or whether one device is more dangerous than another.
Can a MAUDE report prove an AI medical device failed?
No. A report records an outcome, such as a wrong location on a screen. It usually cannot show whether the cause was the model, other software, hardware, calibration or something else. In the 2024 npj Digital Medicine review, 34.5% of 429 reports lacked enough information to tell.
Did the FDA link the TruDI recalls to its AI features?
No. The 2023 recall was tied to a software design change, and the 2024 recall to a calibration process control problem. The FDA did not attribute either to TruSeg or TruPath, the machine learning features added in 2021.
What is predicate creep in AI medical devices?
Predicate creep describes small changes accumulating across generations of 510(k) devices, even though each new device stays linked to an earlier one. A Lancet Digital Health study found that 32.6% of 285 cleared AI devices traced back to first-generation predicates with no AI or machine learning components.
Conclusion
Treat every figure here as a signal, not a verdict. The injury claims are allegations, the company responses are on the record, and the public data cannot yet separate an AI error from a hardware, calibration or software one. The lawsuits were still ongoing in the coverage I reviewed, and I have not seen a ruling on whether the AI features played any role.
Check the original Reuters investigation and the FDA recall database for anything that has changed since February 2026.
0 Comments