Whether a diagnostic tool counts as a medical device depends on what the software claims to do, and that single decision shapes everything in AI diagnostics software development: the evidence you collect, the team you need and the date you can launch. The FDA had authorized over 1,600 AI-enabled medical devices as of September 2026, so the route is well travelled. Nobody ends up on it by accident, though.
Where AI diagnostics software development starts: the intended use
Start by writing down, in one sentence, what the tool will tell a clinician and what it will leave to them. The FDA regulates AI-enabled devices as medical devices, software included, rather than treating AI as a separate category, and some software functions sit outside the device definition under section 520(o) of the FD&C Act.
The line is drawn by claims, not by how clever the model is. A feature that surfaces information for a clinician to weigh is a different product from one that returns a diagnosis, and the FDA's clinical decision support software guidance, dated January 2026, is where you find out which side each feature falls on. It explains which decision support functions are excluded under section 520(o)(1)(E), which the FDA calls the Non-Device CDS criteria.
Read the full guidance, not a summary of it, before the roadmap is fixed. Moving one feature across that line late means redoing the evidence plan around it.
Which FDA pathway does AI diagnostic software follow?
Devices reach the US market through premarket clearance (510(k)), De Novo classification or premarket approval, depending on the device. Which one applies to you comes out of the intended use and the risk that use carries, so it's a regulatory conclusion to confirm with specialist counsel, not an engineering preference.
Two things follow for the build. The FDA also reviews modifications to AI-enabled software that could significantly affect safety or effectiveness, so a model you plan to retrain every quarter needs a change strategy designed in from the start. And the agency's lifecycle thinking covers development, validation, deployment, monitoring, maintenance and modification, which means the monitoring pipeline is part of the product rather than an ops afterthought.
The FDA's page also points to three guiding-principles documents, on Good Machine Learning Practice, predetermined change control plans for machine learning-enabled devices, and transparency for machine learning-enabled devices. Hand them to your data scientists early, because they're the FDA's own statement of how a machine learning team should work.
Why external validation decides whether the model survives contact with a hospital
A model that scores well on the data it was trained on will usually do worse somewhere else. A systematic review in Annals of Medicine and Surgery, published in October 2025, looked at six studies of radiology models on CT and MRI. Internal AUC ranged from 0.76 to 0.95, and on external data the median AUC fell by about 0.03.
The AUC hides the damage. Specificity dropped by as much as about 24 percentage points, and in one of the included studies it fell from 94% to 70% while sensitivity barely moved. In a clinic, that gap is a flood of false alarms that clinicians learn to ignore.
Treat the review with its own limits in mind. Six retrospective studies, CT and MRI only, none reporting prospective clinical deployment: it's a warning about a pattern, not a measurement of your model. Its authors still recommend mandatory external validation on diverse cohorts, multicenter training data, evaluation across age groups, institutions and equipment, and local validation before deployment.
The review also noted what seemed to help: training on data from several centers, and augmenting training data with GAN-generated images. In one included study, the external AUC was noticeably higher with augmentation than without. That's encouraging, but with so few studies it points to something worth testing on your own data rather than a technique to adopt on faith.
Our read is that this changes the first month of a project. Securing data from a second institution, with different scanners and a different patient mix, is a harder task than choosing an architecture, and it should start first.
Designing the workflow around the prediction

Photo by Christina Morillo on Pexels
A diagnostic output nobody sees at the right moment is worth nothing, so the interface carries as much risk as the model. The same review's authors argue that clinical tools should serve as decision support rather than standalone diagnostics, and that's a design brief as much as a regulatory position: show the finding, show what it was based on, and make it quick for a clinician to disagree.
Where the result lands matters. For a clinician, that's usually inside the record they already work in, which pulls in the questions covered in our guide to EHR software development, compliance and usability. For a patient, it may arrive through a portal, and building a patient portal has its own traps around what people can safely be shown without a clinician in the room.
We designed and built OptimalMD's website, members portal and mobile app end to end, from design through development to launch, and you can read how in the OptimalMD case study. That kind of patient-facing surface is where a diagnostic model's output succeeds or fails in practice.
Patient data raises the bar on everything else. Our overview of what HIPAA-compliant app development actually takes covers the obligations that apply the moment training data or predictions touch identifiable records.
A build order that avoids the expensive rework
The sequence matters more than the stack. This is the order we'd put the work in for a first diagnostic product, based on the points above.
- Write the intended use and check it against the FDA's decision support guidance.
- Line up data from at least two sites before choosing a model.
- Prototype the clinician-facing screen with real findings, including wrong ones, to see how people react to a false positive.
- Train, then test on data from a site the model never saw, and report specificity alongside AUC.
- Build monitoring and a change process before launch, since modifications that affect safety get reviewed.
Teams usually come to us at step two, with a model they like and no second dataset. If you're planning the product side, our apps and SaaS development and AI and automation work are where steps three and five get built, and the wider picture sits in our complete guide to healthcare software development.
Frequently asked questions
Can an AI diagnostic model be updated after FDA authorization?
Yes, within limits. A marketing submission may include a predetermined change control plan, and the FDA has issued final guidance on PCCPs for AI-enabled device software functions. Changes that could significantly affect safety or effectiveness are still reviewed.
Is there FDA guidance on the whole lifecycle of an AI diagnostic tool?
There's a draft, titled "Artificial Intelligence-Enabled Device Software Functions: Lifecycle Management and Marketing Submission Recommendations." It covers development, validation, deployment, monitoring, maintenance and modification, and a draft can change before it's final.
Should a diagnostic tool give a diagnosis or support a clinician's decision?
The authors of the 2025 radiology review recommend decision support rather than standalone diagnosis, plus local validation before deployment. Which one you can claim also affects your regulatory route, so settle it before design starts.
Cover photo by Daniil Komov on Pexels
Sources
- Artificial Intelligence in Software as a Medical Device — U.S. Food and Drug Administration
- Clinical Decision Support Software — U.S. Food and Drug Administration





























