NEJM AI’s cover photo
NEJM AI

NEJM AI

Book and Periodical Publishing

Waltham, Massachusetts 22,523 followers

AI is transforming clinical practice. Are you ready?

About us

NEJM AI, a new monthly journal from NEJM Group, is the first publication to engage both clinical and technology innovators in applying the rigorous research and publishing standards of the New England Journal of Medicine to evaluate the promises and pitfalls of clinical applications of AI. NEJM AI is leading the way in establishing a stronger evidence base for clinical AI while facilitating dialogue among all parties with a stake in these emerging technologies. We invite you to join your peers on this journey.

Website
https://ai.nejm.org/
Industry
Book and Periodical Publishing
Company size
201-500 employees
Headquarters
Waltham, Massachusetts
Founded
2023
Specialties
medical education and public health

Updates

  • NEJM AI reposted this

    Our piece is out in NEJM AI! More than 1,000 AI-enabled medical devices have been cleared by the FDA, but most are still evaluated using old data from a handful of hospitals. Unfortunately, these tools often fail when deployed in other clinics or as they see newer patient cases. And when they do, there's no alarm...just a confident wrong answer. We thus need a solution that protects patients without slowing innovation. In our new paper, we propose that the FDA require: - before clearance: multisite validation that reflects the diversity of our patients - after deployment: continuous monitoring so issues get flagged and addressed quickly Endlessly grateful for guidance and insights from Dr. Roxana Daneshjou, Dr. Kavita Patel, star student Alia Sajjadian, and the editorial team at NEJM AI. 📄: https://lnkd.in/gGX-giHi *Happy to send a PDF, please dm or comment below

    • No alternative text description for this image
  • Hematoxylin and eosin–stained tissue analysis has remained the global foundation of diagnostic pathology since the 19th century, and recent advances in #ArtificialIntelligence (AI) now enable the extraction of clinical and molecular insights from routine images.     Yet only 5 of the 1430 AI- and/or machine learning–enabled medical devices cleared by the U.S. Food and Drug Administration are modern pathology AI tools based on whole-slide images, whereas 1094 are classified in the radiology category. Furthermore, only approximately 4% of U.S. slide volume is currently read digitally.     Su and colleagues argue that this disconnect reflects a hierarchy of barriers to clinical impact: foundational infrastructure and economic hurdles, post-digitization challenges related to sample variability and clinician trust, an underdeveloped implementation and governance pathway, and persistent equity gaps in low- and middle-income settings.     Against this backdrop, two near- and medium-term translational pathways stand out: AI as a digital copilot for clinical evaluation and AI-enabled clinical trial design. Together, these approaches offer measurable workflow and trial-enabling benefits, with the potential to improve patient outcomes.     Read the Perspective “Bridging the Gap — Translating AI in Pathology into Clinical Impact” by F.-Y. Su (Fang-Yi Su) et al.: https://nejm.ai/44xQhI1    #AIinMedicine 

    • A quote from a Perspective published in NEJM AI reads as follows: "Standard pathology slides remain a largely untapped resource for precision medicine, but realizing their value requires more than algorithmic advances." The Perspective is titled “Bridging the Gap — Translating AI in Pathology into Clinical Impact” and the authors are F.-Y. Su et al. The background color is orange and includes the NEJM AI logo.
  • End point evaluation is central to the evidence generation process in a clinical trial, and #ArtificialIntelligence (AI) tools have the potential to improve the efficiency and accuracy of this process. In 2025, the European Medicines Agency and U.S. Food and Drug Administration independently endorsed AIM-NASH, an AI-based tool for histologic assessment in clinical trials for metabolic dysfunction–associated steatohepatitis (MASH). As the first ever regulatory certification of an AI-enabled end point, this represents a milestone for the emerging development and use of AI tools for drug development. A new article examines the evidence and policy decisions underlying the endorsement and explores the clinical and regulatory implications of AI-enabled end points for the transformation of clinical trials. Read the Perspective “Adoption of Artificial Intelligence–Based End Points — A New Era for Clinical Trials” by Kushal Kadakia, MD, MSc, and Harlan Krumholz, MD, SM: https://nejm.ai/4wkloTl #AIinMedicine

    • Table outlining potential applications of artificial intelligence for clinical trial endpoints. It includes focus areas such as histopathology, anatomical imaging, and physiological imaging. Descriptions and current examples are provided alongside AI-enabled future possibilities for each category.
  • Progress in #ArtificialIntelligence (AI)–based analysis of surgical videos has been constrained by reliance on manual frame-level annotations rather than patient-level outcomes. In addition, concerns about data privacy restrict the exchange of laparoscopic video data and, thereby, multicenter collaboration. To address these limitations, Saldanha and colleagues developed a pipeline that integrates weakly supervised deep learning with Swarm Learning, a decentralized machine learning approach that enables collaborative model training without data centralization. The authors evaluated the pipeline using a dataset of 397 laparoscopic appendectomy recordings from six international centers for two patient-level staging tasks: (1) laparoscopic grading of appendicitis and appendiceal perforation detection; and (2) histopathologic inflammation grading. They identified optimal modeling configurations (frame sampling rates and model architectures) using the binary perforation detection task, then compared Swarm Learning with single-center and centralized learning across the laparoscopic and histopathologic disease staging tasks. In addition, they surveyed participating centers to identify barriers to clinical implementation of our learning pipeline for surgical video analysis. For binary perforation detection, frame sampling at one frame per second and use of the SurgTempoNet architecture resulted in reliable classification performance, outperforming SurgFrameNet and Multiple Instance Learning. For both laparoscopic (area under the receiver operating curve [AUROC]: 0.818±0.092) and histopathologic disease staging (AUROC: 0.626±0.029), Swarm Learning consistently outperformed single-center training and achieved performance comparable to centralized learning on external validation (AUROC: 0.795±0.092 for laparoscopic grading; AUROC: 0.610±0.018 for histopathologic grading). The user survey identified hardware failure and limited integration of the decentralized learning pipeline with electronic patient records as key barriers to clinical implementation. Weakly supervised deep learning enables the prediction of patient-level labels directly from surgical video data. Swarm Learning facilitates privacy-preserving multicenter collaboration and achieves performance on par with centralized learning, highlighting its potential for advancing clinically relevant, collaborative AI development in surgical video analysis. Read the Original Article “Privacy-Preserving Surgical Video Analysis with Swarm Learning — Results from a Multinational Appendectomy Cohort” by O.L. Saldanha (Dr. Oliver Lester Saldanha) et al.: https://nejm.ai/44Ap0V9 #AIinMedicine

    • A digital illustration showing a medical diagram of human organs overlaid with surgical tools. At the bottom, there's a text box containing the article title: "Privacy-Preserving Surgical Video Analysis with Swarm Learning — Results from a Multinational Appendectomy Cohort" by O.L. Saldanha and Others.
  • Repetitive laboratory testing that is unlikely to yield clinically useful information is a common practice that burdens patients and increases health care costs. Education and feedback interventions have limited success, while general test ordering restrictions and electronic alerts impede appropriate clinical care. Liang and colleagues introduce and evaluate SmartAlert, a machine learning–driven clinical decision support (CDS) system integrated into the electronic health record that predicts stable laboratory results to reduce unnecessary repeat testing. A new Case Study describes the implementation process, challenges, and lessons learned from deploying SmartAlert targeting complete blood count (CBC) utilization in a randomized controlled pilot across 9270 admissions in eight acute care units across two hospitals between August 15, 2024, and March 15, 2025. The results show a significant decrease in the number of CBC results within 52 hours of SmartAlert display (1.54 vs. 1.82; P<0.01) without adverse effect on secondary safety outcomes, representing a 15% relative reduction in repetitive testing. Implementation lessons learned include interpretation of probabilistic model predictions in clinical contexts, stakeholder engagement to define acceptable model behavior, governance processes for deploying a complex model in a clinical environment, user interface design considerations, alignment with clinical operational priorities, and the value of qualitative feedback from end users. In conclusion, a machine learning–driven CDS system backed by a deliberate implementation and governance process can provide precision guidance on inpatient laboratory testing to safely reduce unnecessary repetitive testing. Read the Case Study “SmartAlert — Implementing Machine Learning–Driven Clinical Decision Support for Inpatient Laboratory Utilization Reduction” by A.S. Liang (April Liang) et al.: https://nejm.ai/4ezJxzb #ArtificialIntelligence #AIinMedicine

    • Table depicting demographics and outcomes of CBC Utilization SmartAlert Pilot. It displays data for treatment and control groups, including number of encounters, alerts, ages, and race percentages. Rates for ICU admission and mortality are included.
  • Brandon Rice, co-founder and CEO of Weave Bio, believes that improving health care sometimes means improving the systems behind it.     Through his work at Weave, he focuses on modernizing the regulatory infrastructure that helps bring new medicines to market.     In the latest episode of NEJM AI Grand Rounds, he discusses the challenge of organizing scientific knowledge, communicating with regulators, and navigating processes that can span more than a decade.     His vision is that #ArtificialIntelligence can reduce administrative burden while preserving the rigor required for patient safety and scientific progress.   Listen to the full episode with hosts Arjun (Raj) Manrai, and Andrew Beam, PhD: https://nejm.ai/ep44   #AIinMedicine

    • Promotional image for episode 44 of the NEJM AI Grand Rounds podcast featuring Weave’s Brandon Rice on rebuilding drug regulation with AI. Includes a photo of the guest.
  • What if better models reduce — not increase — the value of traditional oversight? In the latest episode of the NEJM AI Grand Rounds podcast, Dr. Karan Singhal discusses findings that challenge assumptions. Hear more from Dr. Singhal in the full episode hosted by Arjun (Raj) Manrai, PhD, and Andrew Beam, PhD: https://nejm.ai/ep43 #ArtificialIntelligence #AIinMedicine

  • Rare and undiagnosed genetic disorders affect millions of patients globally, and many patients endure years of inconclusive testing. Conventional genomic interpretation can be insufficiently sensitive and costly and is rarely repeated as knowledge evolves. Jaech and colleagues conducted a retrospective multicohort reanalysis using a large language model (LLM)–assisted workflow that ingests clinician notes, Human Phenotype Ontology (HPO) terms, and a filtered variant table to propose explanation-rich candidate hypotheses for expert adjudication under American College of Medical Genetics and Genomics and Association for Molecular Pathology criteria. A diagnosis was defined a priori as a variant classified as pathogenic or likely pathogenic, confirmed in a Clinical Laboratory Improvement Amendments–certified laboratory, and returned to families. Secondary outputs included “rediscoveries” of externally established diagnoses not yet available locally and hypothesis generation signals. Across four cohorts, new local diagnoses were made in 10 of 100 rare disease neurodevelopmental cases (10.0%, [exact binomial: 95% confidence interval (CI), 4.9 to 17.6]), 4 of 61 neuromuscular cases (6.6%, [CI, 1.8 to 16.0]), 2 of 200 cases of sudden unexpected death in pediatrics (1.0% [CI, 0.1 to 3.6]), and 2 of 15 early psychosis cases (13.3% [CI, 1.7 to 40.5]) for an overall diagnostic yield of 18 of 376 (4.8%, [CI, 2.9 to 7.5]). The authors identified seven rediscoveries in which pathogenic or likely pathogenic findings had been established externally but were not available in the local research record at the time of review. In one case, the model’s synthesis of genotype-quality patterns and phenotype concordance triaged a putative 22q11.2 deletion that was subsequently confirmed by whole-genome sequencing. The workflow also generated testable biological hypotheses, including a candidate association between the sphingosine-1-phosphate receptor 1 gene (S1PR1) and vitiligo. In retrospective reanalysis, an explanation-first LLM applied to routine HPO terms and variant tables produced clinically relevant gains in diagnostic yield, surfaced overlooked pathogenic findings, and generated biologically grounded hypotheses. These results motivate prospective multicenter evaluation with predefined end points, calibration reporting, and comparator baselines. Read the Case Study “LLM-Assisted Reanalysis of Unsolved Rare Disease Genomes Increases Diagnostic Yield” by A. Jaech et al.: https://nejm.ai/4eMh8pu 📖 Further reading in NEJM AI: Editorial by Sam Finlayson, MD, PhD, and Heidi Rehm, PhD: When Genomic Reanalysis Leaves the Laboratory — Clinical Genetics in the Age of Consumer AI https://nejm.ai/4vVYCAY #ArtificialIntelligence #AIinMedicine

    • Diagram showing protein structure predictions related to the pathological mechanism of the Sphingosine-1-Phosphate Receptor 1 (S1PR1) c.850_882del variant in a vitiligo patient. Image includes a protein sequence diagram, ribbon model structures, and molecular surface models.
  • Connecting free-text descriptions such as “3-mm nodule in the lower left lobe” to precise 3D segmentations remains an unsolved challenge in medical #ArtificialIntelligence (AI). Existing chest computed tomography (CT) datasets rely on structured labels or predefined categories, limiting their ability to represent the richness of clinical language and support grounded radiology report generation. Bridging this gap requires datasets that capture the expressiveness of free-text findings and link them to accurate annotations in volumetric imaging. Baharoon and colleagues introduce ReXGroundingCT, the first publicly available dataset linking free-text findings to 3D segmentations in chest CT scans. It includes 3142 noncontrast CT scans paired with standardized radiology reports from CT-RATE, constructed through a three-stage pipeline. First, GPT-4 extracted and standardized findings, descriptors, and metadata from Turkish reports machine-translated into English. Second, GPT4 omni (GPT-4o) categorized each finding into a hierarchical ontology of lung and pleural abnormalities. Third, 3D annotations were produced for all CT volumes: the training set underwent quality assurance by board-certified radiologists, and the validation and test sets were fully labeled by them. A complementary chain-of-thought dataset was also created, providing step-by-step anatomical reasoning for localizing findings within the CT volume, guided by organ-segmentation models. ReXGroundingCT provides 16,301 annotated entities across 8028 text-to-3D-segmentation pairs spanning diverse findings. About 79% of findings are focal abnormalities, while 21% are nonfocal. The dataset includes a public validation set of 50 cases and a private test set of 100 cases, annotated by board-certified radiologists. Model performance on the test set is hosted on a leaderboard at https://lnkd.in/gRJ_a3ZW. ReXGroundingCT is the first manually curated dataset linking free-text chest CT findings to 3D segmentation masks, providing a benchmark for developing and evaluating free-text medical segmentation models. It lays the foundation for the segmentation of free-text findings and generation of grounded radiology reports in CT imaging. The dataset is available at https://lnkd.in/gnRUbuuX. Read the Datasets, Benchmarks, and Protocols article “ReXGroundingCT: A 3D Chest CT Dataset for Segmentation of Findings from Free-Text Reports” by M. Baharoon et al.: https://nejm.ai/4giw8N6 #AIinMedicine

    • Figure 1 illustrates the ReXGroundingCT Dataset. Panel A displays lung CT scan findings. Panel B highlights a bar chart of findings per category across the dataset, grouped by typical pattern (nonfocal vs. focal). Panel C details an example of anatomical chain-of-thought reasoning.

Affiliated pages

Similar pages