Lakes LinkedPublic Health Watch
Lakes Linked · Field Notes in Public-Health Analytics Registry Review · Vol. 1, No. 2 · June 2026
Registry-Based Review

From Accuracy to Outcomes: A Structured Review of Newly Registered Artificial-Intelligence Clinical Trials (ClinicalTrials.gov, May to June 2026)

Sergey Soshnikov, MD, PhD
lakeslinked.com · Public-health analytics from open data.
Search window: 1 May to 29 June 2026 Compiled: 29 June 2026 Data source: ClinicalTrials.gov API v2
Abstract

Background. Clinical artificial intelligence (AI) is usually judged by diagnostic-accuracy metrics. Its public-health value, however, rests on whether a model improves decisions, reduces harm, and operates safely inside real health systems. Newly registered trials give an early reading of where the field is putting its prospective evidence.

Methods. We reviewed AI-related studies registered on ClinicalTrials.gov (NIH/NLM) between 1 May and 29 June 2026. Each record was classified along four axes: the model's functional role at the point of care, planned enrolment as a measure of scale, endpoint type, and public-health relevance. Five exemplar trials were chosen to show the direction of travel, and each was checked field by field against its full registry record through the ClinicalTrials.gov API v2.

Results. About 154 AI-related studies were registered in the window. Diagnostic imaging remained the largest functional category. A sizeable minority, though, embedded models inside the electronic health record (EHR), benchmarked general-purpose large language models (LLMs) prospectively against clinician reference standards, or extended AI into population-scale opportunistic screening. Planned enrolment ranged across three orders of magnitude, from a 690-patient LLM benchmarking study to a 200,000-participant pancreatic-cancer screening trial. Few trials carried hard clinical endpoints such as overall survival or 30-day mortality.

Conclusions. The registry points to an early but consistent shift away from accuracy demonstrations and toward evidence on outcomes, implementation, and equity. The trials of greatest public-health consequence are those that tie AI to hard endpoints, embed it in clinical workflow, or deploy it as screening infrastructure on imaging that already exists. Their scarcity is what makes them worth tracking.

Keywords: artificial intelligence; machine learning; large language models; clinical trials; ClinicalTrials.gov; opportunistic screening; clinical endpoints; implementation science; public health

1 · IntroductionIntroduction

The main evidentiary currency of clinical AI has been the diagnostic-accuracy metric: the area under the receiver-operating-characteristic curve (AUC), sensitivity, and specificity, each computed against a held-out test set. These quantities are necessary but not sufficient. A high AUC shows that a model can discriminate. It says little about whether deploying that model changes what clinicians do, whether patients end up better off, or whether the benefit reaches the people most likely to be missed. For public health, the binding questions sit downstream of accuracy, in the territory of decisions, harms, reach, and safe operation inside real health systems.[1]

Trial registries offer an early vantage point. Because studies are registered before they report, the registry shows the field's intended evidence: the endpoints investigators are willing to defend, the scale at which they plan to test, and the settings into which they plan to deploy. This information arrives months or years before any results. Reading newly registered AI trials as a set, rather than simply counting them, therefore gives a forward-looking signal of where prospective clinical AI evidence is heading.[1]

This review asks a focused question. Among AI-related studies registered in a recent two-month window, how are functional roles, scale, and endpoints distributed, and which trials carry the greatest public-health relevance? We organise the analysis around one hypothesis, that the field is starting to move from accuracy demonstrations toward outcome- and implementation-oriented evidence, and we test it against the registry record.

2 · MethodsMethods

2.1 Data source and search window

The data source was ClinicalTrials.gov, the NIH/National Library of Medicine registry of clinical studies, accessed programmatically through its public Application Programming Interface, version 2 (API v2).[1] The search window covered studies first posted between 1 May and 29 June 2026. Eligible records were studies whose condition, intervention, or design fields referenced artificial intelligence, machine learning, or deep learning, with recruitment status of recruiting or not-yet-recruiting.

2.2 Classification framework

Rather than tabulate trials by disease, we classified each record along four pre-specified axes chosen to expose public-health signal. Function captured what the model does at the point of care, sorted into seven categories: diagnosis and imaging; risk prediction and early warning; population and opportunistic screening; clinical decision support; patient-facing LLMs and communication; education and workflow; and therapeutic-adjacent or digital interventions. A single trial could fall into more than one category. Scale captured planned enrolment, used as a proxy for the capacity to generate population-level evidence. Endpoint type separated diagnostic-accuracy measures (AUC, sensitivity, specificity, agreement) from hard clinical endpoints (overall survival, mortality, length of stay, kidney injury). Public-health relevance flagged whether a study touched screening, antimicrobial stewardship, infectious-disease surveillance, equity, or health-system implementation.

2.3 Selection and verification of exemplar trials

Five exemplar trials were selected on purpose, not for novelty, but to show the direction of travel across the highest-relevance axes: population-scale opportunistic screening, antimicrobial stewardship under a hard safety endpoint, EHR-embedded implementation science, prospective benchmarking of clinicians against LLMs, and AI as infectious-disease surveillance infrastructure. Each exemplar was checked field by field against its full registry record retrieved through the API, confirming sponsor, country, study design, planned enrolment, recruitment status, and the pre-registered primary endpoint. The verified characteristics appear in Table 2. Where this review differs from informal summaries of these trials, the registry record governs.

2.4 Nature and limits of the design

This is a descriptive, registry-based review rather than a systematic review or meta-analysis. It applies no PRISMA screening, assesses no risk of bias, and reports no pooled effects. All enrolment figures, statuses, and sponsors are planned or current registry values that may change, and none is a result. The approach is meant to surface direction and emphasis in newly registered evidence, and its conclusions should be read in that light (Section 5).

3 · ResultsResults

About 154 AI-related studies were registered in the search window. The distribution bore out both halves of the framing hypothesis. Diagnostic imaging continued to dominate the count, while a smaller and qualitatively distinct group pushed toward outcomes, implementation, and scale.

3.1 Functional taxonomy

Sorting by function rather than by disease shows where AI is merely reading and where it is starting to decide. Diagnostic and imaging applications, meaning models that read scans, slides, ultrasound, and endoscopy, formed the largest group. Risk-prediction and early-warning models, usually embedded in the EHR or the monitoring stream, were a growing category. The smallest but highest-stakes category was population and opportunistic screening: re-reading routine scans, or stratifying whole populations, to find disease that no one was actively looking for. Table 1 summarises the seven categories with representative registered trials.

Table 1. Functional taxonomy of newly registered AI-related trials, with representative examples. Categories are not mutually exclusive; a single trial may span more than one.
Functional categoryWhat the model doesRelative volumeRepresentative NCT
Diagnosis & imagingReads scans, slides, ultrasound, endoscopyLargestNCT07647692 · lung nodule CT
Risk prediction & early warningForecasts deterioration (AKI, hypotension, dosing)GrowingNCT07627607 · ICU hypotension
Population & opportunistic screeningRe-reads routine scans; stratifies populationsSmall / highest-stakesNCT06528223 · PANDA pancreatic
Clinical decision supportAdvises the order, test, or stage at point of careHigh-impactNCT06163781 · ABC blood cultures
Patient-facing LLMs & communicationTranslates, summarises, counselsEmergingNCT07666399 · health translation
Education & workflowTrains clinicians; eases documentationHigh volumeNCT07618975 · nursing process
Therapeutic-adjacent & digitalPersonalises prevention, nutrition, rehabNicheNCT07622628 · AIM-MET microbiome

3.2 Distribution of scale

Planned enrolment ranged across three orders of magnitude. The largest cohorts clustered in opportunistic screening and population surveillance, where the marginal cost of applying a model to imaging or records that already exist is close to zero. Scale is not a virtue on its own, but it is what separates a proof of concept from a study with the statistical power to move a guideline.

Table 2. Exemplar trials, ordered by planned enrolment, with characteristics verified field by field against ClinicalTrials.gov API v2 records (retrieved 29 June 2026).
NCT · AcronymSponsor · CountryDesignPlanned NPre-registered primary endpointStatus
NCT06528223 · PANDA Zhejiang University · China Interventional RCT (real-world validation) 200,000 Overall survival (to 3 yr) Not yet
NCT07611695 · TB-ATLAS Huashan Hospital · China (w/ HK PolyU) Observational, retrospective-prospective cohort 31,600 AUROC, Easy- vs Hard-to-Treat stratification Not yet
NCT07604662 · ML-AKI UC San Francisco · USA (NIGMS) Pragmatic 3-arm cluster-RCT 25,518 Change in serum creatinine, POD 1 to 2 Not yet
NCT06163781 · ABC Amsterdam UMC (VUmc) · Netherlands RCT, non-inferiority 7,584 30-day mortality Recruiting
NCT07626060 · LLM-HEART Marmara University · Türkiye Prospective observational diagnostic accuracy 690 AUC for 30-day MACE prediction Recruiting

3.3 Endpoint type

The most consequential finding concerns endpoints. Most registered studies were still diagnostic-accuracy work, with primary outcomes expressed as AUC, sensitivity, specificity, or agreement against a reference standard. Only a minority pre-registered hard clinical endpoints. Two exemplars show the difference inside the same registry: ABC defends a 30-day mortality non-inferiority endpoint,[3] and PANDA ties an opportunistic-screening pathway to overall survival,[2] whereas the LLM-benchmarking and TB-stratification studies, important as they are, report performance metrics such as AUC and AUROC as their primary outcomes.[5][6] That scarcity of hard outcomes is itself worth recording, because it makes the exceptions disproportionately valuable.

3.4 Exemplar trials

The five trials below are described from their verified registry records. Together they map the frontier where AI is being tested as public-health infrastructure rather than as a demonstration.

NCT06528223 200,000 plannedInterventional RCTZhejiang University

PANDA: opportunistic pancreatic-cancer screening from ordinary CT

A deep-learning model re-reads routine non-contrast CT scans to flag pancreatic ductal adenocarcinoma that the original radiology report missed. Patients who are imaging-negative but PANDA-positive are recalled for gold-standard work-up, and the pre-registered primary endpoint is overall survival, assessed from diagnosis out to three years.[2]

Why it matters: opportunistic screening at population scale may be the most consequential public-health use of clinical AI. Pancreatic cancer is almost always found too late. Re-reading scans that people already have, and then anchoring the model to a survival endpoint rather than an AUC, turns the move from accuracy to outcomes into something concrete. If the approach holds up, the screening pathway costs little to deploy, because the imaging already exists.
NCT06163781 30-day mortalityRCT · 7,584Amsterdam UMC

ABC: machine learning to reduce unnecessary blood cultures in the ED

A machine-learning model predicts whether a blood culture will return positive. In the intervention arm, when the predicted probability falls below 5%, the culture is cancelled. Earlier validation suggests the tool can cut blood-culture volume by at least 30%. The trial is powered for non-inferiority on 30-day mortality, with hospital admission, in-hospital mortality, and length of stay as key secondary outcomes.[3]

Why it matters: antimicrobial stewardship is a core public-health priority, and this is one of the few AI trials willing to defend a hard safety endpoint instead of diagnostic accuracy. By cancelling low-yield cultures, it targets both antibiotic overuse and false-positive contamination, the everyday drivers of resistance. A non-inferiority result is what turns a prediction model into a defensible change in practice.
NCT07604662 25,518 planned3-arm cluster-RCTUCSF · NIGMS

ML-AKI: EHR-embedded acute-kidney-injury risk score after surgery

A pragmatic three-arm cluster-randomised trial randomises 75 to 100 attending anesthesiologists 1:1:1 to a hidden risk score (control), a visible preoperative AKI risk probability, or the visible score paired with an interruptive best-practice advisory. The pre-registered primary endpoint is the continuous change in serum creatinine from baseline to postoperative day 1 to 2.[4]

Why it matters: this is implementation science rather than model-building. A risk score helps only if it is embedded in the EHR, surfaced in the workflow, and acted on, and the three-arm design isolates which part of that chain changes outcomes. It is the question that decides whether thousands of published prediction models ever earn a place inside a real health system.
NCT07626060 LLM head-to-head690 · ED chest painMarmara University

LLM-HEART: GPT-4o vs Claude vs three-expert consensus on the HEART score

A prospective observational diagnostic-accuracy study pits two frontier language models (gpt-4o-2024-11-20 and claude-sonnet-4-6), accessed under deterministic settings, against a blinded three-expert consensus. The task is HEART-score calculation from free-text Turkish clinical notes and 30-day major-adverse-cardiac-event (MACE) prediction. The design follows STARD-AI 2025 reporting standards, and 690 patients are enrolled to yield 600 evaluable complete cases.[5]

Why it matters: general-purpose LLMs already sit in clinicians' pockets, so the relevant question is no longer whether they sound plausible but whether they hold up against a defined gold standard on a real triage decision with a 30-day outcome. Formal, prospective benchmarking of clinicians against AI, rather than leaderboard scores on exam questions, is how the field earns the right to put these tools near patients.
NCT07611695 31,600 plannedObservational cohortHuashan Hospital

TB-ATLAS: AI-driven tuberculosis stratification across a population

A retrospective-prospective cohort builds and validates a modular AI clinical-decision-support system for whole-chain tuberculosis management, drawing on more than 30,000 retrospective patient records with prospective external validation in a cohort of at least 1,600. The pre-registered primary endpoint is the AUROC for separating "Easy-to-Treat" from "Hard-to-Treat" pulmonary tuberculosis.[6]

Why it matters: tuberculosis remains one of the world's leading infectious killers, and case-finding, not treatment, is the binding constraint. Risk-stratifying tens of thousands of patients is AI used as surveillance infrastructure, steering scarce screening and treatment capacity toward the people most likely to be missed. This is the public-health lane where AI's marginal value is highest, because human screening cannot scale to meet the need.

4 · DiscussionDiscussion

Read as a set, the newly registered trials support the framing hypothesis: the field is beginning to move from accuracy demonstrations toward outcome- and implementation-oriented evidence. Four patterns carry the signal.

Accuracy is increasingly treated as a starting point rather than a result. The most credible new trials take a high AUC as a precondition and then attach a downstream measure, such as survival, length of stay, or mortality, that reflects whether patients are actually better off. This is a quiet concession that test-set performance says little about clinical benefit.[2][3]

Opportunistic screening is reaching population scale. PANDA (planned 200,000) and a parallel routine-CT tumour cohort (planned 100,000) point to AI's largest public-health lever: extracting new diagnoses from imaging that already exists, at a marginal cost near zero. At this scale AI stops behaving like a clinic tool and starts behaving like screening infrastructure.[2]

LLMs are entering formal, prospective clinical benchmarking. General-purpose models are now tested against clinician gold standards on defined decisions with real outcomes, under reproducible, pre-registered protocols aligned to STARD-AI reporting standards. That is a marked shift from exam-style leaderboards toward the evidentiary bar that real triage requires. It is early and narrow, but it points the right way.[5]

Implementation, not model-building, is the decisive frontier. The ML-AKI cluster-RCT exemplifies a maturing recognition that a prediction model is inert until it is embedded in workflow and acted upon; its three-arm design is built to isolate which component of the human-and-model interface actually changes care.[4] Whether the large existing inventory of published prediction models ever earns a place in routine practice is an implementation question, and the registry shows that question finally being asked prospectively.

Interpretation

For public health, the central question is no longer whether an AI model can predict. The harder question is whether it improves decisions, reduces harm, reaches the right populations, and works safely inside real health systems.

The registry shows that shift beginning, away from performance claims and toward evidence on scale, endpoints, implementation, and equity. The trials that carry it are still a minority, which is exactly why they deserve close and continuing attention.

5 · LimitationsLimitations

Several limitations qualify these findings. First, this is a descriptive registry review, not a systematic review: it applies no PRISMA-conformant screening, no risk-of-bias assessment, and no quantitative synthesis, and the functional taxonomy involves interpretive judgement that another reviewer might apply differently. Second, the approximate count of 154 studies depends on the specific search definition for "AI-related" and on registry indexing; it should be read as an order-of-magnitude estimate of activity, not a precise denominator, and the trials discussed are a purposive cross-section rather than a complete enumeration. Third, all enrolment figures, recruitment statuses, sponsors, and endpoints are planned or current registry values that may change before completion; none is a result, and registration does not guarantee that a study will be conducted, completed, or published. Fourth, registry records can contain errors or omissions; although exemplar trials were verified field-by-field against their full API records, non-exemplar trials were not exhaustively audited. Finally, a two-month window captures a snapshot of intent and cannot establish a trend on its own. The patterns described here are hypotheses about direction that future windows can confirm or revise.

6 · ConclusionConclusion

Among AI-related studies registered between May and June 2026, diagnostic imaging still dominates by count, but the studies of greatest public-health consequence are those that tie AI to hard clinical endpoints, embed it within the electronic health record, benchmark language models prospectively against clinicians, or deploy it as population-scale screening infrastructure. These remain a minority of registrations. Their scarcity, set against the public-health magnitude of what they test, is the central finding of this review, and the reason newly registered AI trials warrant systematic, continuing surveillance rather than one-off enumeration.

ReproducibilityData and methods summary

SourceClinicalTrials.gov API v2 (NIH / U.S. National Library of Medicine)
SearchAI / ML / deep-learning conditions or interventions · first posted 1 May to 29 June 2026 · recruiting or not-yet-recruiting
Records matched≈154 studies (search-definition dependent); a purposive cross-section is reported
VerificationFive exemplar trials confirmed field by field against full API records, 29 June 2026; every NCT links to its source
DesignDescriptive registry review; not a systematic review or meta-analysis
PreparationAI-assisted, human-reviewed

ReferencesReferences

  1. National Library of Medicine (US). ClinicalTrials.gov. Bethesda (MD): NIH/NLM. Registry of clinical studies; primary source for all trials cited (search: artificial intelligence / machine learning / deep learning, first posted 1 May 2026 onward). https://clinicaltrials.gov/
  2. Zhejiang University. Research of the Application of Artificial Intelligence Model "PANDA": A Multicenter, Prospective Randomized Controlled Clinical Trial. ClinicalTrials.gov identifier NCT06528223. Opportunistic pancreatic-cancer screening from non-contrast CT; primary endpoint overall survival. https://clinicaltrials.gov/study/NCT06528223
  3. Amsterdam UMC (location VUmc). Appropriate Use of Blood Cultures in the Emergency Department Through Machine Learning (ABC): a Randomized Controlled Trial. ClinicalTrials.gov identifier NCT06163781. Non-inferiority RCT; primary endpoint 30-day mortality. https://clinicaltrials.gov/study/NCT06163781
  4. University of California, San Francisco (with NIGMS). Prediction of Acute Kidney Injury After Surgery (ML-AKI): A Pragmatic Three-Arm Cluster-Randomized Trial. ClinicalTrials.gov identifier NCT07604662. EHR-embedded AKI risk score; primary endpoint change in serum creatinine, postoperative day 1 to 2. https://clinicaltrials.gov/study/NCT07604662
  5. Marmara University Pendik Training and Research Hospital. Diagnostic Accuracy of Large Language Models (GPT-4o and Claude) in HEART Score Calculation and 30-Day MACE Prediction in Emergency Department Chest Pain Patients: A Prospective Observational Validation Study Against Three-Expert Consensus (LLM-HEART). ClinicalTrials.gov identifier NCT07626060. https://clinicaltrials.gov/study/NCT07626060
  6. Huashan Hospital (with The Hong Kong Polytechnic University). Artificial Intelligence-driven Tuberculosis Landscape Analysis & Stratification Research (TB-ATLAS). ClinicalTrials.gov identifier NCT07611695. Retrospective-prospective cohort; primary endpoint AUROC for Easy- vs Hard-to-Treat stratification. https://clinicaltrials.gov/study/NCT07611695

DeclarationsDeclarations

Funding. None. This review was prepared without external funding as part of the Lakes Linked public-health analytics project.

Competing interests. The author declares no competing financial interests. The author has no relationship to the trials reviewed.

Data availability. All underlying data are publicly available on ClinicalTrials.gov; every trial cited is linked to its source record by NCT identifier.

Generative-AI disclosure. This review is AI-assisted (using tools such as Claude) and human-reviewed. Enrolment numbers, statuses, sponsors, and endpoints are planned or current registry values that may change; nothing reported here is a result. This is a research and public-health-education brief, not medical, regulatory, or investment advice. Confirm any trial detail against its source record before relying on it.

#ClinicalAI#PublicHealth#ClinicalTrials#ImplementationScience#HealthPolicy#AIinHealthcare#AntimicrobialStewardship#CancerScreening#HealthData#LLM

AboutAbout this review

Lakes Linked is a public-health analytics project that turns open data, such as ClinicalTrials.gov, CDC, US Census, and IHME, into briefs and dashboards that local decision-makers can use. This is Part 02 of a three-part digest on AI in public health; the literature review (Part 01) and news digest (Part 03) appear as separate articles. This journal-formatted version restructures the original Field Note into a registry-based review with verified exemplar data and numbered references.