Background. Clinical artificial intelligence (AI) is usually judged by diagnostic-accuracy metrics. Its public-health value, however, rests on whether a model improves decisions, reduces harm, and operates safely inside real health systems. Newly registered trials give an early reading of where the field is putting its prospective evidence.
Methods. We reviewed AI-related studies registered on ClinicalTrials.gov (NIH/NLM) between 1 May and 29 June 2026. Each record was classified along four axes: the model's functional role at the point of care, planned enrolment as a measure of scale, endpoint type, and public-health relevance. Five exemplar trials were chosen to show the direction of travel, and each was checked field by field against its full registry record through the ClinicalTrials.gov API v2.
Results. About 154 AI-related studies were registered in the window. Diagnostic imaging remained the largest functional category. A sizeable minority, though, embedded models inside the electronic health record (EHR), benchmarked general-purpose large language models (LLMs) prospectively against clinician reference standards, or extended AI into population-scale opportunistic screening. Planned enrolment ranged across three orders of magnitude, from a 690-patient LLM benchmarking study to a 200,000-participant pancreatic-cancer screening trial. Few trials carried hard clinical endpoints such as overall survival or 30-day mortality.
Conclusions. The registry points to an early but consistent shift away from accuracy demonstrations and toward evidence on outcomes, implementation, and equity. The trials of greatest public-health consequence are those that tie AI to hard endpoints, embed it in clinical workflow, or deploy it as screening infrastructure on imaging that already exists. Their scarcity is what makes them worth tracking.
The main evidentiary currency of clinical AI has been the diagnostic-accuracy metric: the area under the receiver-operating-characteristic curve (AUC), sensitivity, and specificity, each computed against a held-out test set. These quantities are necessary but not sufficient. A high AUC shows that a model can discriminate. It says little about whether deploying that model changes what clinicians do, whether patients end up better off, or whether the benefit reaches the people most likely to be missed. For public health, the binding questions sit downstream of accuracy, in the territory of decisions, harms, reach, and safe operation inside real health systems.[1]
Trial registries offer an early vantage point. Because studies are registered before they report, the registry shows the field's intended evidence: the endpoints investigators are willing to defend, the scale at which they plan to test, and the settings into which they plan to deploy. This information arrives months or years before any results. Reading newly registered AI trials as a set, rather than simply counting them, therefore gives a forward-looking signal of where prospective clinical AI evidence is heading.[1]
This review asks a focused question. Among AI-related studies registered in a recent two-month window, how are functional roles, scale, and endpoints distributed, and which trials carry the greatest public-health relevance? We organise the analysis around one hypothesis, that the field is starting to move from accuracy demonstrations toward outcome- and implementation-oriented evidence, and we test it against the registry record.
The data source was ClinicalTrials.gov, the NIH/National Library of Medicine registry of clinical studies, accessed programmatically through its public Application Programming Interface, version 2 (API v2).[1] The search window covered studies first posted between 1 May and 29 June 2026. Eligible records were studies whose condition, intervention, or design fields referenced artificial intelligence, machine learning, or deep learning, with recruitment status of recruiting or not-yet-recruiting.
Rather than tabulate trials by disease, we classified each record along four pre-specified axes chosen to expose public-health signal. Function captured what the model does at the point of care, sorted into seven categories: diagnosis and imaging; risk prediction and early warning; population and opportunistic screening; clinical decision support; patient-facing LLMs and communication; education and workflow; and therapeutic-adjacent or digital interventions. A single trial could fall into more than one category. Scale captured planned enrolment, used as a proxy for the capacity to generate population-level evidence. Endpoint type separated diagnostic-accuracy measures (AUC, sensitivity, specificity, agreement) from hard clinical endpoints (overall survival, mortality, length of stay, kidney injury). Public-health relevance flagged whether a study touched screening, antimicrobial stewardship, infectious-disease surveillance, equity, or health-system implementation.
Five exemplar trials were selected on purpose, not for novelty, but to show the direction of travel across the highest-relevance axes: population-scale opportunistic screening, antimicrobial stewardship under a hard safety endpoint, EHR-embedded implementation science, prospective benchmarking of clinicians against LLMs, and AI as infectious-disease surveillance infrastructure. Each exemplar was checked field by field against its full registry record retrieved through the API, confirming sponsor, country, study design, planned enrolment, recruitment status, and the pre-registered primary endpoint. The verified characteristics appear in Table 2. Where this review differs from informal summaries of these trials, the registry record governs.
This is a descriptive, registry-based review rather than a systematic review or meta-analysis. It applies no PRISMA screening, assesses no risk of bias, and reports no pooled effects. All enrolment figures, statuses, and sponsors are planned or current registry values that may change, and none is a result. The approach is meant to surface direction and emphasis in newly registered evidence, and its conclusions should be read in that light (Section 5).
About 154 AI-related studies were registered in the search window. The distribution bore out both halves of the framing hypothesis. Diagnostic imaging continued to dominate the count, while a smaller and qualitatively distinct group pushed toward outcomes, implementation, and scale.
Sorting by function rather than by disease shows where AI is merely reading and where it is starting to decide. Diagnostic and imaging applications, meaning models that read scans, slides, ultrasound, and endoscopy, formed the largest group. Risk-prediction and early-warning models, usually embedded in the EHR or the monitoring stream, were a growing category. The smallest but highest-stakes category was population and opportunistic screening: re-reading routine scans, or stratifying whole populations, to find disease that no one was actively looking for. Table 1 summarises the seven categories with representative registered trials.
| Functional category | What the model does | Relative volume | Representative NCT |
|---|---|---|---|
| Diagnosis & imaging | Reads scans, slides, ultrasound, endoscopy | Largest | NCT07647692 · lung nodule CT |
| Risk prediction & early warning | Forecasts deterioration (AKI, hypotension, dosing) | Growing | NCT07627607 · ICU hypotension |
| Population & opportunistic screening | Re-reads routine scans; stratifies populations | Small / highest-stakes | NCT06528223 · PANDA pancreatic |
| Clinical decision support | Advises the order, test, or stage at point of care | High-impact | NCT06163781 · ABC blood cultures |
| Patient-facing LLMs & communication | Translates, summarises, counsels | Emerging | NCT07666399 · health translation |
| Education & workflow | Trains clinicians; eases documentation | High volume | NCT07618975 · nursing process |
| Therapeutic-adjacent & digital | Personalises prevention, nutrition, rehab | Niche | NCT07622628 · AIM-MET microbiome |
Planned enrolment ranged across three orders of magnitude. The largest cohorts clustered in opportunistic screening and population surveillance, where the marginal cost of applying a model to imaging or records that already exist is close to zero. Scale is not a virtue on its own, but it is what separates a proof of concept from a study with the statistical power to move a guideline.
| NCT · Acronym | Sponsor · Country | Design | Planned N | Pre-registered primary endpoint | Status |
|---|---|---|---|---|---|
| NCT06528223 · PANDA | Zhejiang University · China | Interventional RCT (real-world validation) | 200,000 | Overall survival (to 3 yr) | Not yet |
| NCT07611695 · TB-ATLAS | Huashan Hospital · China (w/ HK PolyU) | Observational, retrospective-prospective cohort | 31,600 | AUROC, Easy- vs Hard-to-Treat stratification | Not yet |
| NCT07604662 · ML-AKI | UC San Francisco · USA (NIGMS) | Pragmatic 3-arm cluster-RCT | 25,518 | Change in serum creatinine, POD 1 to 2 | Not yet |
| NCT06163781 · ABC | Amsterdam UMC (VUmc) · Netherlands | RCT, non-inferiority | 7,584 | 30-day mortality | Recruiting |
| NCT07626060 · LLM-HEART | Marmara University · Türkiye | Prospective observational diagnostic accuracy | 690 | AUC for 30-day MACE prediction | Recruiting |
The most consequential finding concerns endpoints. Most registered studies were still diagnostic-accuracy work, with primary outcomes expressed as AUC, sensitivity, specificity, or agreement against a reference standard. Only a minority pre-registered hard clinical endpoints. Two exemplars show the difference inside the same registry: ABC defends a 30-day mortality non-inferiority endpoint,[3] and PANDA ties an opportunistic-screening pathway to overall survival,[2] whereas the LLM-benchmarking and TB-stratification studies, important as they are, report performance metrics such as AUC and AUROC as their primary outcomes.[5][6] That scarcity of hard outcomes is itself worth recording, because it makes the exceptions disproportionately valuable.
The five trials below are described from their verified registry records. Together they map the frontier where AI is being tested as public-health infrastructure rather than as a demonstration.
A deep-learning model re-reads routine non-contrast CT scans to flag pancreatic ductal adenocarcinoma that the original radiology report missed. Patients who are imaging-negative but PANDA-positive are recalled for gold-standard work-up, and the pre-registered primary endpoint is overall survival, assessed from diagnosis out to three years.[2]
A machine-learning model predicts whether a blood culture will return positive. In the intervention arm, when the predicted probability falls below 5%, the culture is cancelled. Earlier validation suggests the tool can cut blood-culture volume by at least 30%. The trial is powered for non-inferiority on 30-day mortality, with hospital admission, in-hospital mortality, and length of stay as key secondary outcomes.[3]
A pragmatic three-arm cluster-randomised trial randomises 75 to 100 attending anesthesiologists 1:1:1 to a hidden risk score (control), a visible preoperative AKI risk probability, or the visible score paired with an interruptive best-practice advisory. The pre-registered primary endpoint is the continuous change in serum creatinine from baseline to postoperative day 1 to 2.[4]
A prospective observational diagnostic-accuracy study pits two frontier language models (gpt-4o-2024-11-20 and claude-sonnet-4-6), accessed under deterministic settings, against a blinded three-expert consensus. The task is HEART-score calculation from free-text Turkish clinical notes and 30-day major-adverse-cardiac-event (MACE) prediction. The design follows STARD-AI 2025 reporting standards, and 690 patients are enrolled to yield 600 evaluable complete cases.[5]
A retrospective-prospective cohort builds and validates a modular AI clinical-decision-support system for whole-chain tuberculosis management, drawing on more than 30,000 retrospective patient records with prospective external validation in a cohort of at least 1,600. The pre-registered primary endpoint is the AUROC for separating "Easy-to-Treat" from "Hard-to-Treat" pulmonary tuberculosis.[6]
Read as a set, the newly registered trials support the framing hypothesis: the field is beginning to move from accuracy demonstrations toward outcome- and implementation-oriented evidence. Four patterns carry the signal.
Accuracy is increasingly treated as a starting point rather than a result. The most credible new trials take a high AUC as a precondition and then attach a downstream measure, such as survival, length of stay, or mortality, that reflects whether patients are actually better off. This is a quiet concession that test-set performance says little about clinical benefit.[2][3]
Opportunistic screening is reaching population scale. PANDA (planned 200,000) and a parallel routine-CT tumour cohort (planned 100,000) point to AI's largest public-health lever: extracting new diagnoses from imaging that already exists, at a marginal cost near zero. At this scale AI stops behaving like a clinic tool and starts behaving like screening infrastructure.[2]
LLMs are entering formal, prospective clinical benchmarking. General-purpose models are now tested against clinician gold standards on defined decisions with real outcomes, under reproducible, pre-registered protocols aligned to STARD-AI reporting standards. That is a marked shift from exam-style leaderboards toward the evidentiary bar that real triage requires. It is early and narrow, but it points the right way.[5]
Implementation, not model-building, is the decisive frontier. The ML-AKI cluster-RCT exemplifies a maturing recognition that a prediction model is inert until it is embedded in workflow and acted upon; its three-arm design is built to isolate which component of the human-and-model interface actually changes care.[4] Whether the large existing inventory of published prediction models ever earns a place in routine practice is an implementation question, and the registry shows that question finally being asked prospectively.
For public health, the central question is no longer whether an AI model can predict. The harder question is whether it improves decisions, reduces harm, reaches the right populations, and works safely inside real health systems.
The registry shows that shift beginning, away from performance claims and toward evidence on scale, endpoints, implementation, and equity. The trials that carry it are still a minority, which is exactly why they deserve close and continuing attention.
Several limitations qualify these findings. First, this is a descriptive registry review, not a systematic review: it applies no PRISMA-conformant screening, no risk-of-bias assessment, and no quantitative synthesis, and the functional taxonomy involves interpretive judgement that another reviewer might apply differently. Second, the approximate count of 154 studies depends on the specific search definition for "AI-related" and on registry indexing; it should be read as an order-of-magnitude estimate of activity, not a precise denominator, and the trials discussed are a purposive cross-section rather than a complete enumeration. Third, all enrolment figures, recruitment statuses, sponsors, and endpoints are planned or current registry values that may change before completion; none is a result, and registration does not guarantee that a study will be conducted, completed, or published. Fourth, registry records can contain errors or omissions; although exemplar trials were verified field-by-field against their full API records, non-exemplar trials were not exhaustively audited. Finally, a two-month window captures a snapshot of intent and cannot establish a trend on its own. The patterns described here are hypotheses about direction that future windows can confirm or revise.
Among AI-related studies registered between May and June 2026, diagnostic imaging still dominates by count, but the studies of greatest public-health consequence are those that tie AI to hard clinical endpoints, embed it within the electronic health record, benchmark language models prospectively against clinicians, or deploy it as population-scale screening infrastructure. These remain a minority of registrations. Their scarcity, set against the public-health magnitude of what they test, is the central finding of this review, and the reason newly registered AI trials warrant systematic, continuing surveillance rather than one-off enumeration.