
Only three of the 1,357 U.S. Food and Drug Administration-authorized artificial intelligence (AI) medical devices have been evaluated for clinical effectiveness, according to a study published online Aug. 19 in PLOS Digital Health.
Rawan Abulibdeh, from the University of Toronto, and colleagues conducted a systematic analysis of all FDA-cleared AI/machine learning-enabled medical devices (through Dec. 5, 2025) using the FDA device database and the American College of Radiology Data Science Institute catalog, with searches linked to ClinicalTrials.gov and PubMed, to identify registered trials and publications.
The researchers found that of 1,357 cleared AI devices, only 2.5 percent were linked to registered prospective trials, 0.9 percent (12 studies) posted results, 0.9 percent had peer-reviewed publications, and 0.2 percent (three studies) evaluated patient-centered outcomes, including mortality, morbidity, or readmissions. Of those AI devices with identified studies, 62 percent had employed observational designs with small, homogeneous cohorts, limited subgroup analyses, and frequent exclusion of vulnerable populations. Rigorous evaluation was discouraged by structural barriers, including misaligned financial incentives, reliance on predicate-based regulatory pathways, and logistical challenges of multicenter trials.
“The market is crowded with AI tools, yet the evidence supporting patient-centered outcomes remains remarkably thin,” the authors write. “This evidence attrition highlights the profound gap between regulatory authorization and clinically meaningful validation.”
One author is a paid consultant for the Collaborative Institutional Training Initiative.
Regulatory authorization and real-world clinical effectiveness are not necessarily the same thing. A medical device might demonstrate technical performance or meet a regulatory pathway without having extensive evidence showing how it performs across different patient populations.
But why does this matter? An AI tool that hasn’t been trained or validated primarily on homogeneous populations may not perform identically across race, skin tone, age, socioeconomic status, or other patient characteristics. Inaccurate or biased data from AI tools could significantly affect patient care and health outcomes.
Previous studies that examined FDA-authorized AI/ML devices found that only 3.6 percent of authorization summaries reported race or ethnicity, while 99.1 percent didn’t report socioeconomic data.
For clinicians, it’s critical to know how these new AI tools are evaluated before implementing them in care.
As medical technology evolves, tools and devices must be evaluated across patients from diverse backgrounds. Black patients — in particular — have historically been underrepresented in medical research, creating concerns about whether medical technologies perform equally well across populations. If demographic information and subgroup performance aren’t reported, providers may have a limited ability to determine whether an AI tool has been adequately evaluated in Black patients.
These concerns aren’t to assume that AI is inherently dangerous for Black patients — it’s that insufficient validation of these medical technologies can make it harder to detect disparities before they become part of routine clinical use.
The latest study findings reveal a large evidence gap that poses an equity issue: an AI tool may be widely deployed to clinical practices before providers know whether its benefits and limitations are consistent across populations.

Before implementing AI-enabled medical devices into care, providers should ask the following questions:
FDA authorization should only be viewed as one part of the evidence lifecycle, not the end of the evaluation. AI systems can behave differently when deployed across different hospitals, patient populations, workflows, and datasets. Studies have shown ongoing challenges involving algorithmic bias, transparency, post-market surveillance, and continuously learning systems.
For clinicians serving historically underserved populations, post-market monitoring of AI-enabled medical devices can be especially important because performance disparities may not become apparent until an AI tool is used across a broader, more diverse patient population.
By subscribing, you consent to receive emails from BlackDoctor.pro You may unsubscribe at any time. Privacy Policy & Terms of Service.
Are you a healthcare professional? Register with us today!