arXivSub Start free trial

MICCAI 2026 Papers with AI Summaries

International Conference on Medical Image Computing and Computer-Assisted Intervention · 1165 papers

3D Cerebrovascular Shape Completion from Biplane Angiography and CTA Prior

Jehkul, Janik (Harvard Medical School, Brigham and Women's Hospital), Haouchine, Nazim (Harvard Medical School, Brigham and Women's Hospital)

RestorationConvolutional Neural NetworkTransformerDiffusion modelContrastive LearningGaussian SplattingImagePoint CloudBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Utilize dual-view DSA and CTA priors, employing a 3D Gaussian point cloud to complete the shape and recover the fine vascular structures missing in CTA.

3D Classification of Paramagnetic Rim Lesions in Multiple Sclerosis via Asymmetric QSM–FLAIR Modeling

Pignedoli, Veronica (University of Genova), Moro, Matteo (University of Genova)

ClassificationAnomaly DetectionConvolutional Neural NetworkTransformerSupervised Fine-TuningAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: This study proposes FRODO, a 3D deep learning framework based on QSM and FLAIR, for automatically distinguishing paramagnetic rim lesions (Rim+) from non-Rim lesions (Rim-) in multiple sclerosis.

A 3D Unrolling Framework for Joint Sodium MRI Reconstruction and Concentration Quantification

Cao, Yilin (ShanghaiTech University), Sun, Kaicong (ShanghaiTech University)

RestorationConvolutional Neural NetworkDiffusion modelScore-based ModelContrastive LearningBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Proposed a 3D deep transposed convolutional framework called SoReCon, which can end-to-end reconstruct sodium magnetic resonance imaging and quantitatively estimate sodium concentration from undersampled radial k-space.

A Clinical Guideline-Grounded Hybrid Agentic Framework for Holistic Epilepsy Management

Pham, Duy Khoa (Monash University), Mehta, Deval (Monash University)

ClassificationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningConvolutional Neural NetworkTransformerLarge Language ModelMixture of ExpertsVision Language ModelImageTextMultimodalityTabularBiomedical DataMagnetic Resonance ImagingElectronic Health RecordsRetrieval-Augmented Generation

🎯 What it does: Propose a hybrid multi-modal multi-task agent framework called EPI-GUIDE based on international epilepsy guidelines, for integrating multi-modal information such as EEG, MRI, and text to make comprehensive epilepsy management decisions.

A Confounder-aware Representation Learning Framework for Alzheimer’s Disease Classification

Chen, Yuanhao (Zhejiang University of Finance and Economics), Wang, Changmiao (Shenzhen Research Institute of Big Data)

ClassificationRepresentation LearningConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningImageTabularBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease

🎯 What it does: This paper proposes a hybrid modal framework based on causal inference for the early diagnosis of Alzheimer's disease;

A Counterfactual Framework for Directional Cell–Cell Interaction Analysis in Spatial Transcriptomics

Anzum, Humaira (University of Houston), Banerjee, Tania (University of Houston)

Drug DiscoveryGraph Neural NetworkGenerative Adversarial NetworkContrastive LearningBiomedical Data

🎯 What it does: Designed a directional cell-cell interaction framework based on adversarial inference, utilizing a neighborhood graph model to predict receptor cell states and quantify directional effects through hypothetical interventions.

A Deep Learning Surrogate Model for Microwave Thermal Ablation: From Preoperative Planning to Intraoperative Re-planning

Nahmed, Ilias (Université de Strasbourg), Cotin, Stéphane (Université de Strasbourg)

SegmentationData SynthesisOptimizationComputational EfficiencyConvolutional Neural NetworkTransformerSupervised Fine-TuningDiffusion modelContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound

🎯 What it does: This paper proposes a real-time deep learning surrogate model for predicting thermal injury and necrotic volume in microwave ablation (MWA) of the liver, and embeds it into an optimization loop for pre-surgical planning and intraoperative re-planning, to achieve rapid personalized treatment plans.

A Flexible Structure-Guided Feature Aggregation Paradigm for Efficient Vascular Representation

Qi, Yaolei (Southeast University), Yang, Guanyu (Southeast University)

SegmentationComputational EfficiencyRepresentation LearningConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningOptical FlowImagePoint CloudBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound

🎯 What it does: Propose VSP-Branch, a pluggable structural-guided feature aggregation module, which enhances the continuity and branch integrity of vascular representations by constructing learnable paths along the local vascular structure, aggregating cross-layer features, and adaptively injecting structural information through difficulty-gated mechanisms.

A Guideline-Aware AI Agent for Zero-Shot Target Volume Auto-Delineation

Kim, Yoon Jo (Oncosoft Inc.), Kim, Jin Sung (Oncosoft Inc.)

SegmentationExplainability and InterpretabilityConvolutional Neural NetworkTransformerLarge Language ModelAgentic AIImageTextBiomedical DataComputed TomographyRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed an AI agent framework called OncoAgent, which can zero-shot automatically convert text-based radiotherapy clinical guidelines into three-dimensional target volume volumes;

A Heterogeneous Prognosis Prediction Framework with Global Brain Connectivity and Local Image Features

Lin, Jingfeng, Tang, Zhenyu (Chinese PLA General Hospital)

ClassificationRepresentation LearningConvolutional Neural NetworkGraph Neural NetworkSupervised Fine-TuningContrastive LearningImageGraphBiomedical DataMagnetic Resonance ImagingDiffusion Tensor Imaging

🎯 What it does: This paper proposes a heterogeneous framework that achieves prognosis prediction for diffuse gliomas by iteratively combining graph neural networks based on DTI with local feature networks based on structural MRI.

A Hierarchical Multi-Task Framework for Dementia Diagnosis via Pathological Feature Learning from Multi-Organ Data

Zhao, Shilun (ShanghaiTech University), Shen, Dinggang (ShanghaiTech University)

ClassificationRepresentation LearningConvolutional Neural NetworkTransformerMixture of ExpertsContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyAlzheimer's DiseaseElectronic Health Records

🎯 What it does: Propose the MoMmNet framework, which utilizes multi-organ imaging (brain, heart, gut, liver, kidney) and clinical text for hierarchical multi-task reasoning to achieve multi-cause dementia diagnosis.

A Modality-Aware Mixture-of-Experts Framework for Clinical Multimodal Breast Ultrasound Diagnosis

Suo, Hao (Shanghai Jiao Tong University), Chen, Fang (Shanghai Jiao Tong University)

ClassificationAnomaly DetectionComputational EfficiencyRepresentation LearningTransformerMixture of ExpertsVision Language ModelContrastive LearningImageMultimodalityBiomedical DataUltrasound

🎯 What it does: Propose the MA-MoE model, which performs multi-modal fusion of paired functional ultrasound modalities such as B-mode with CDFI, elasticity, or CEUS, and construct a private benchmark dataset that covers real-world clinical heterogeneity including multiple devices, different clock positions, and long-tail distributions.

A Multi-center Benchmark for Abdominal Disease Diagnosis and Report Generation from Non-Contrast CT

Elbakry, Mariam (Ain Shams University), Elbatel, Marawan (Hong Kong University of Science and Technology)

ClassificationGenerationExplainability and InterpretabilityConvolutional Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningImageTextBiomedical DataComputed TomographyBenchmarkRetrieval-Augmented Generation

🎯 What it does: Designed and released a multi-center benchmark to evaluate the feasibility of generating multi-organ diagnostic reports from single-phase non-contrast CT.

A Multi-view, Hybrid, Hypergraph Learning Framework with Causal Perturbation for Early Alzheimer’s Disease Diagnosis

Jia, Yifan (ShanghaiTech University), Zhang, Han (ShanghaiTech University)

ClassificationExplainability and InterpretabilityConvolutional Neural NetworkGraph Neural NetworkTransformerSupervised Fine-TuningContrastive LearningGraphTime SeriesBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease

🎯 What it does: Proposes HyperBrainNet, a framework combining causal effective connectivity, high-order hybrid hypergraph networks, and multi-view adaptive fusion, for the diagnosis of early Alzheimer's disease.

A Multimodal 3D Foundation Model for Light Sheet Fluorescence Microscopy Enables Few-Shot Segmentation, Classification, and Deblurring

Scheinfeld, Adina (Weill Cornell Medicine), Paetzold, Johannes C. (Weill Cornell Medicine)

ClassificationRestorationSegmentationKnowledge DistillationRepresentation LearningConvolutional Neural NetworkTransformerSupervised Fine-TuningVision Language ModelContrastive LearningImageTextBiomedical Data

🎯 What it does: This paper constructs a multi-modal 3D foundation model, pre-trained using a large-scale unlabeled light sheet fluorescence microscopy (LSM) volume images and corresponding text descriptions, achieving efficient few-shot fine-tuning on downstream tasks such as segmentation, classification, and deblurring.

A Neurosymbolic Framework for Interpretable Skeleton-Based Seizure Detection via Concept-Driven Logical Reasoning

Ilyas, Talha (Monash University), Ge, Zongyuan (Monash University)

Anomaly DetectionExplainability and InterpretabilityGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringContrastive LearningVideoBiomedical DataElectrocardiogram

🎯 What it does: A neural-symbolic framework is constructed for epileptic seizure detection using patient skeleton sequences and interpretable concepts, providing three levels of auditable explanations;

A Real-World Evaluation of Failure Detection for Liver CT Segmentation

Bennett, Jeddy (University of Texas MD Anderson Cancer Center), Brock, Kristy K. (University of Texas MD Anderson Cancer Center)

SegmentationAnomaly DetectionConvolutional Neural NetworkContrastive LearningImageBiomedical DataComputed Tomography

🎯 What it does: Evaluate and compare multiple OOS detection methods in real clinical settings for failure detection in automatic liver CT segmentation models.

A Unified Few-Shot Framework for Multi-Atlas Neuroimaging Segmentation Leveraging Vision Foundation Models

Guo, Bocheng (University of Electronic Science and Technology of China), Zhang, Fan (Harvard Medical School)

SegmentationDomain AdaptationComputational EfficiencyConvolutional Neural NetworkTransformerPrompt EngineeringAuto EncoderContrastive LearningBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose a unified few-shot framework that transfers large visual foundation models to 3D multi-atlas neuroimaging segmentation, achieving efficient adaptation across atlases.

A Unified Framework for Joint Detection of Lacunes and Enlarged Perivascular Spaces

He, Lucas (University College London), Sudre, Carole H. (University College London)

Object DetectionSegmentationConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose a unified multi-task framework that jointly detects adenomas (EPVS) and small cavities (lacunes) in the brain and improves detection accuracy.

A Versatile Prompt-Driven Framework for Brain Segmentation, Parcellation, and Surface Reconstruction from Diverse Modalities

Lian, Zifeng (ShanghaiTech University), Shen, Dinggang (ShanghaiTech University)

RestorationSegmentationGenerationConvolutional Neural NetworkTransformerPrompt EngineeringMixture of ExpertsContrastive LearningMultimodalityBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Proposed a prompt-based unified framework called uBrain, which simultaneously achieves multi-modal brain voxel segmentation, ROI partitioning, and surface reconstruction.

AbdomenGen: Sequential Volume-Conditioned Diffusion Framework for Abdominal Anatomy Generation

Bhandari, Yubraj (Duke University), Lo, Joseph Y. (Duke University)

GenerationData SynthesisConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelContrastive LearningImageBiomedical DataComputed Tomography

🎯 What it does: Proposed a sequential volumetric conditional diffusion framework named AbdomenGen for controllable generation of computer tomography (CT) phantom containing 11 abdominal structures.

AC-MIL: Weakly-Supervised Atrial LGE-MRI Quality Assessment via Adversarial Concept Disentanglement

Sultan, K. M. Arefeen (University of Utah), Elhabian, Shireen Y. (University of Utah)

ClassificationExplainability and InterpretabilityConvolutional Neural NetworkTransformerGenerative Adversarial NetworkContrastive LearningBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose the AC-MIL framework to achieve weakly supervised atrial LGE-MRI quality assessment, decomposing overall quality into interpretable clinical concepts;

ACA: Post-hoc Adaptive Logit Alignment for 3D Medical Image Segmentation

Dong, Nanyu (Adelaide University), Liao, Zhibin (Adelaide University)

SegmentationContrastive LearningBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound

🎯 What it does: To address the insufficient generalization of pre-trained models for 3D medical image segmentation on unseen data, this paper proposes a post-adaptive clipping alignment (ACA) method. It enhances segmentation performance by achieving local distribution alignment within the model's uncertain logit interval through isotropic quantile mapping and isotropic regression.

Accurate Reconstruction of the Standard 12-Lead ECG from a Single Lead Based on Inherent Principles of ECG

Seo, Dong-hyuk (Hanyang University), Kim, Sang-Wook (Yonsei University)

RestorationConvolutional Neural NetworkTransformerAuto EncoderGenerative Adversarial NetworkTime SeriesBiomedical DataElectrocardiogram

🎯 What it does: This paper proposes a method for reconstructing complete 12-lead electrocardiograms (ECGs) based on a single-lead ECG.

ACMap: Across-scales Connectivity Mapping of Primate Brain

Lin, Runjia (Children's Hospital of Philadelphia), Huang, Hao (Children's Hospital of Philadelphia)

Diffusion modelScore-based ModelBiomedical DataDiffusion Tensor Imaging

🎯 What it does: Develop ACMap, which generates high-fidelity fiber trajectories from viral tracing data using the Fast Marching Method, completing a cross-scale (from micrometer to millimeter level) brain connectivity map;

Active Evaluation-Induced Graph Network Based Medical Hyperspectral Image Classification for Tumor Diagnosis

Wang, Meiling (Nanjing University of Posts and Telecommunications), Cao, Liying (China University of Mining and Technology)

ClassificationAnomaly DetectionGraph Neural NetworkGenerative Adversarial NetworkContrastive LearningImageBiomedical Data

🎯 What it does: Designed and implemented an Active Evaluation Induced Graph Network (AEGNet) for tumor diagnosis classification in medical hyperspectral images.

Active Source-free Domain Adaptation in Open-set Medical Image Segmentation via Decomposed Uncertainty and Prototype Discrepancy

Yang, Jin (Icahn School of Medicine at Mount Sinai), Yu, Xiaobing (Washington University School of Medicine in St. Louis)

SegmentationDomain AdaptationConvolutional Neural NetworkTransformerSupervised Fine-TuningAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: This paper proposes an active source-free open-set domain adaptation method (ASFOSDA) for medical image segmentation;

AdaBIMBA: Adaptive Token Compression for Long Endoscopy Video Understanding

Yung, Ka-Wai (Johnson & Johnson), Damasceno, Pablo F. (Johnson & Johnson)

Image TranslationAnomaly DetectionComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringVision Language ModelAuto EncoderContrastive LearningVideoTextBiomedical Data

🎯 What it does: Developed the AdaBIMBA adaptive token compression module to enable visual language model question answering for long-endoscopy videos;

ADAPT: Adaptive Profiling Transformers for Efficient Alzheimer’s Disease Diagnosis

Wang, Yifeng (Carnegie Mellon University), Wang, Haohan (University of Illinois Urbana-Champaign)

ClassificationComputational EfficiencyRepresentation LearningTransformerAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease

🎯 What it does: This study proposes an efficient 3.5D multi-view Transformer framework called ADAPT for diagnosing Alzheimer's disease from structural MRI.

Adapting Ordinal-Risk Alignment for Clinically Safe Diabetic Retinopathy Grading

Zhang, Yuhan, Ni, Dong (Shenzhen University)

ClassificationTransformerReinforcement LearningContrastive LearningImageBiomedical Data

🎯 What it does: Proposes a risk-sensitive reinforcement learning-based framework for diabetic retinopathy (DR) grading, named RAO-RL, which takes into account ordinal relationships and clinical safety.

Adapting Without Access: Black-Box Bayesian Adaptation for Trustworthy Cross-Center Medical Image Segmentation

Han, Xiaoxiang (Shanghai Jiao Tong University), Zhang, Qi (Tsinghua University)

SegmentationDomain AdaptationFederated LearningSafty and PrivacyBiomedical Data

🎯 What it does: Unable to identify paper content

Adaptive Frequency-Guided Parallel Mamba Network for Polyp Segmentation

Jiang, Aiwen (Jiangxi Normal University), Cui, Rui (Jiangxi Normal University)

SegmentationTransformerVision-Language-Action ModelDiffusion modelAuto EncoderContrastive LearningImageBiomedical DataBenchmark

🎯 What it does: Propose an adaptive frequency-guided parallel Mamba network (Polyp-AFMamba), achieving high-precision polyp segmentation by combining spatial structure with frequency domain information.

AdaSurvMamba: Dynamic Fusion and Semantic Scanning for Multimodal Survival Analysis

Zhong, Jialong (Dalian University of Technology), Lu, Huchuan (Dalian University of Technology)

Representation LearningTransformerContrastive Learning

🎯 What it does: Propose an AdaSurvMamba framework applicable to multi-modal survival analysis, combining WSI and genomic data to achieve accurate cancer prognosis prediction.

Addressing Gradient Conflicts in Multimodal Fundus Disease Recognition with Fusion-Guided Learning

Liu, Xiaozhou (Southwest University), Du, Zhiguo (Southwest University)

RecognitionOptimizationConvolutional Neural NetworkSupervised Fine-TuningContrastive LearningImageMultimodalityBiomedical DataAlzheimer's DiseaseReview/Survey Paper

🎯 What it does: Address gradient conflicts in multi-modal retinal image classification by introducing the fusion-guided gradient decoupling learning (FGDL) and cross-modal semantic alignment (CMSA) modules to improve dynamic optimization and enhance feature alignment;

Addressing Tissue and Appearance Heterogeneity in Text-Guided Few-Shot WSI Classification

Li, Yongcen (Dalian University of Technology), Xing, Xudong (Beijing Institute of Genomics)

ClassificationDomain AdaptationTransformerVision Language ModelContrastive LearningImageTextBiomedical DataComputed TomographyReview/Survey Paper

🎯 What it does: This paper proposes a Heterogeneity-Aware Text-guided MIL framework to address the domain drift problem caused by tissue composition and appearance heterogeneity in whole slide image classification under the extremely few-sample condition.

AEGIS: Anatomy-Embedded Group-Invariant Segmentation for Fair Medical Foundation Models

Wang, Sen (Shenzhen MSU-BIT University), Zhao, Weibing (Shenzhen MSU-BIT University)

SegmentationFederated LearningExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerSupervised Fine-TuningAuto EncoderContrastive LearningBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound

🎯 What it does: Addressing the fairness issues of medical foundation models (such as SAM and MedSegX) in the segmentation of head and neck squamous cell carcinoma (HNSCC), this paper proposes the AEGIS framework, achieving group-invariant segmentation under parameter-efficient fine-tuning (PEFT).

AffordTissue: Dense Affordance Prediction for Tool-Action Specific Tissue Interaction

Maksutova, Aiza (Johns Hopkins University), Unberath, Mathias (Johns Hopkins University)

SegmentationConvolutional Neural NetworkTransformerPrompt EngineeringVision Language ModelContrastive LearningVideoMultimodality

🎯 What it does: Propose a multimodal framework called AffordTissue for predicting tool-action-specific tissue manipulable regions during cholecystectomy and generating dense heatmaps.

Age Conditional Longitudinal Forecasting of Adolescent Functional Connectivity via Brownian Bridge Diffusion Models

Zhang, Rongye (Taiyuan University of Technology), Wen, Xin (Taiyuan University of Technology)

GenerationData SynthesisGraph Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderBiomedical DataMagnetic Resonance ImagingStochastic Differential Equation

🎯 What it does: Propose a framework that combines graph attention autoencoder with Brownian bridge diffusion for long-term evolution prediction of adolescent functional connectivity under age conditions.

AGE-MIL: Anchor-Guided Evidence Learning for Patient-Level Prediction

Niu, Jiawei (Xi'an Jiaotong University), Cai, Yi (Central South University)

ClassificationImage TranslationRetrievalAnomaly DetectionOptimizationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningData-Centric LearningConvolutional Neural NetworkTransformerMixture of ExpertsAuto EncoderContrastive LearningImageBiomedical DataElectronic Health RecordsReview/Survey PaperRetrieval-Augmented Generation

🎯 What it does: Proposed an Anchor-Guided Evidence MIL (AGE-MIL) framework for patient-level diagnosis and prognosis prediction.

Agent-Guided Relational Concept Discovery: Toward Interpretable Surgical Margin Assessment

Maghsoodi, Nooshin (Queen's University), Mousavi, Parvin (Queen's University)

Explainability and InterpretabilityDrug DiscoveryTransformerLarge Language ModelAgentic AIContrastive LearningBiomedical DataMagnetic Resonance ImagingRetrieval-Augmented Generation

🎯 What it does: Automatically discovers and explains interpretable concepts in REIMS data by embedding reasoning agents in the training loop, aiming to improve the accuracy and interpretability of surgical margin assessment.

AGGRNet: Selective Feature Extraction and Aggregation for Enhanced Medical Image Classification

Makwe, Ansh (Indian Institute of Technology Kanpur), Bagade, Priyanka (Indian Institute of Technology Kanpur)

ClassificationConvolutional Neural NetworkTransformerContrastive LearningImageBiomedical Data

🎯 What it does: This paper proposes the AGGRNet framework, which utilizes a feature extraction and aggregation module (FEA) to separate informative and non-informative features within the YOLOv11 backbone network, and captures global dependencies through cross-attention mechanisms, thereby improving the subcategory and severity classification of medical images.

AHPR-Net: Anatomy-aware Hierarchical Prompting Refinement Network for Spine Image Segmentation

Zhao, Junyong (Nanjing University of Aeronautics and Astronautics), Zhang, Daoqiang (Nanjing University of Aeronautics and Astronautics)

SegmentationConvolutional Neural NetworkTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Proposed an anatomy-aware hierarchical prompt refinement network (AHPR-Net) for spinal image segmentation.

AI-Driven Pulmonary Congestion Assessment for Lung Ultrasound via Segmentation-Guided Transformers

Fooladgar, Fahimeh, Kapur, Tina (Brigham and Women's Hospital)

ClassificationSegmentationAnomaly DetectionExplainability and InterpretabilityConvolutional Neural NetworkTransformerPrompt EngineeringMixture of ExpertsContrastive LearningImageBiomedical DataUltrasound

🎯 What it does: Propose a two-stage segmentation-guided transformer framework for automated severity scoring of lung B-lines, enabling consistent and reproducible assessment of pulmonary edema directly at the bedside in lung ultrasound.

ALFA: Biplanar X-Ray Reconstruction via Anatomy-Latent Field Adaptation

Yu, Tianqi (ShanghaiTech University), Zhang, Yuyao (Shanghai Jiao Tong University)

RestorationDomain AdaptationOptimizationDiffusion modelScore-based ModelFlow-based ModelRectified FlowNeural Radiance FieldAuto EncoderImageBiomedical DataComputed Tomography

🎯 What it does: Propose ALFA, a dual-plane X-ray 3D reconstruction framework that decouples the learning of anatomical priors from physical adaptation, leveraging implicit neural representations to construct a continuous anatomical space and achieving zero-shot geometric generalization through latent variable optimization at test time.

Aligning Prototypes and Updating Null Spaces for Continual WSI Learning

Li, Xianrui (City University of Hong Kong), Chan, Antoni B. (City University of Hong Kong)

ClassificationAnomaly DetectionFederated LearningSafty and PrivacyTransformerVision Language ModelAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasound

🎯 What it does: AlignMIL-NS is proposed in a continuous learning whole slide image (WSI) multi-instance learning (MIL) framework, enabling continuous learning without storing any patient data.

All-in-One Augmented Reality Guided Head and Neck Tumor Resection

Yang, Yue (Vanderbilt University), Wu, Jie Ying (Vanderbilt University)

Pose EstimationDepth EstimationAutonomous DrivingRobotic IntelligenceSimultaneous Localization and MappingOptical FlowPoint CloudMeshBiomedical DataMagnetic Resonance ImagingComputed TomographyReview/Survey Paper

🎯 What it does: Studied an全自动无标记AR system, using the depth sensor of HoloLens 2 to reposition the resected head and neck tumor specimens back to the resection bed, and to visualize the tumor margin in real-time at the surgical site.

ALSAnchorNet: Biologically Informed Multimodal MRI Fusion for Amyotrophic Lateral Sclerosis Diagnosis

Shen, Xiongri (Ant Group), Xie, Long (Ant Group)

ClassificationExplainability and InterpretabilityRepresentation LearningConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningMultimodalityBiomedical DataMagnetic Resonance ImagingDiffusion Tensor ImagingAlzheimer's Disease

🎯 What it does: Propose ALSAnchorNet, which integrates rs-fMRI, DTI, structural MRI, and brain lymphatic/cerebrospinal fluid biomarkers, utilizing a prefrontal anchor attention mechanism for multimodal fusion to diagnose ALS.

AMFG: Anatomy-Grounded Multi-finding Guidance for Training-Free Chest X-Ray Generation

Han, Yeon Gyu (Chungnam National University), Lee, Dongheon (Seoul National University)

GenerationData SynthesisConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningImageBiomedical DataComputed TomographyElectronic Health Records

🎯 What it does: Propose a training-agnostic decoding guidance framework AMFG, which generates accurate multi-lesion chest X-rays on frozen chest X-ray diffusion models through three guidance signals.

An Artifact-Based Agent Framework for Adaptive and Reproducible Medical Image Processing

Zuo, Lianrui (Vanderbilt University), Landman, Bennett A. (Vanderbilt University)

Image TranslationRestorationSegmentationReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Proposed an Agent framework based on artifact contracts, achieving adaptive configuration and reproducible execution for medical image processing;

Anatomically Accurate 3D Vessel Generation via 3-Phase Chebyshev Curve Diffusion

Mo, Jihwan (Korea Advanced Institute of Science and Technology), Chang, Dong Eui (Asan Medical Center)

GenerationData SynthesisTransformerPrompt EngineeringDiffusion modelScore-based ModelImageBiomedical DataComputed Tomography

🎯 What it does: A model is proposed that generates continuous, C∞ smooth vascular centerlines through the diffusion of three-phase Chebyshev curves, capable of recovering millimeter-level dimensions while maintaining the curve trajectory.

Anatomically Consistent TMJ Disc Segmentation via Semantic Anchoring and Clinical Priors

Ju, Dayun (Yonsei University), Hwang, Seong Jae (Yonsei University)

SegmentationConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Proposed a TMJ disc segmentation framework called TISC based on semantic anchoring and clinical priors, specifically addressing the segmentation challenges of this structure due to its small size, low contrast, and variable morphology.

Anatomy- and Site-Guided 3D Patch Diffusion for Robust MRI Harmonization

Hache, Barnabé (Inserm U1172), Lopes, Renaud (Inserm U1172)

Image HarmonizationConvolutional Neural NetworkTransformerDiffusion modelContrastive LearningBiomedical DataMagnetic Resonance Imaging

🎯 What it does: This paper proposes a condition diffusion model based on 3D patches, achieving non-destructive alignment and difference elimination for multi-site T1-weighted MRI, capable of unifying scan contrast while preserving fine anatomical details.

Anatomy-Anchored Self-Supervision: Distilling Vision Foundation Models for Invariant Ultrasound Representation

Zhu, Chunzheng, Li, Kenli (Hunan University)

ClassificationSegmentationDomain AdaptationRepresentation LearningTransformerPrompt EngineeringAuto EncoderContrastive LearningImageBiomedical DataUltrasound

🎯 What it does: Propose AnaUS, a self-supervised pre-training framework based on anatomical structures, which utilizes LP-SAM to automatically segment anatomical structures in ultrasound images, and conducts anatomical-level contrastive learning and core region prediction on this basis to enhance the representation capability of ultrasound images.

Anatomy-Aware Hierarchical Multiple Instance Learning for Interpretable COPD Diagnosis and Phenotype Analysis

Ahmadi, Raha (University of British Columbia), Tam, Roger (University of British Columbia)

ClassificationExplainability and InterpretabilityConvolutional Neural NetworkContrastive LearningImageBiomedical DataComputed Tomography

🎯 What it does: Proposed the AnatoMIL framework for COPD detection and grading from CT images.

Anatomy-Aware Prediction of Bronchoscopic Accessibility from 3D CT

Peng, Linkai (Northwestern University), Bagci, Ulas (Northwestern University)

ClassificationAnomaly DetectionConvolutional Neural NetworkRecurrent Neural NetworkGraph Neural NetworkMixture of ExpertsContrastive LearningImageBiomedical DataComputed Tomography

🎯 What it does: This paper formalizes pre-operative bronchoscope reachability prediction as a supervised learning task and proposes an anatomy-aware mixture-of-experts (Anatomy-Aware MoE) framework based on 3D CT.

Anatomy-Aware Reverse Symmetric Self-paced Learning for Prostate Cancer Detection

Zhang, Dan (University of Auckland), Reynolds, Hayley M. (University of Auckland)

SegmentationAnomaly DetectionConvolutional Neural NetworkTransformerContrastive LearningImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Proposed and verified a reverse symmetric self-paced learning (RSSPL) strategy, achieving hard-to-easy case scheduling in prostate cancer multi-parametric MRI detection through anatomy-aware difficulty estimation.

Anatomy-Aware Standard Plane Localization in 3D Ultrasound with Geometric Constraints

Nie, Jianlong (ShanghaiTech University), Shen, Dinggang (ShanghaiTech University)

SegmentationPose EstimationConvolutional Neural NetworkContrastive LearningImageBiomedical DataUltrasound

🎯 What it does: Propose a multi-standard plane localization framework for three-dimensional fetal brain ultrasound, combining anatomical structure segmentation with plane regression, and incorporating geometric constraints between planes;

Anatomy-Conditioned Domain Randomization for Zero-Shot 3D Cerebrovascular Segmentation

Pentassuglia, Matteo (EURECOM), Zuluaga, Maria A. (EURECOM)

SegmentationData SynthesisDomain AdaptationConvolutional Neural NetworkDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose a domain randomization framework based on whole-brain anatomical labels, using only synthetic images to train a 3D CNN, achieving zero-shot brain vessel segmentation on unseen TOF-MRA and CTA data;

Anatomy-Grounded Synthetic Coronary Angiography for Geometry-Informed Multi-view Matching

Lee, In Kyu (Medipixel, Inc.), Min, Jaesik (Medipixel, Inc.)

Image TranslationData SynthesisPose EstimationDepth EstimationTransformerDiffusion modelContrastive LearningOptical FlowImageBiomedical DataComputed Tomography

🎯 What it does: This paper generates digital reconstructed radiographs (DRR) from patient CT angiography data to construct a multi-view coronary artery correspondence dataset without manual annotations, and proposes a Geometry-Informed Matching Module (GIMM) to achieve fine matching.

Anatomy-Guided Residual Motion Diffusion for Controllable 4D Cardiac MRI Synthesis

Cao, Yiheng (Suzhou Institute of Biomedical Engineering and Technology, Chinese Academy of Science), Gao, Xin

SegmentationGenerationData SynthesisTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Proposed a 4D motion medical image generation framework based on semi-supervised VAE and two-stage residual latent diffusion models, which can simultaneously generate voxel sequences and corresponding segmentation masks, achieving controllable 4D cardiac MRI synthesis.

Anatomy-Structured Hierarchical MIL for Weakly-Supervised Thoracic Disease Detection in Chest X-Rays

Kim, Jeongin (Ewha Womans University), Noh, Junhyug (Ewha Womans University)

Object DetectionAnomaly DetectionExplainability and InterpretabilityConvolutional Neural NetworkTransformerContrastive LearningImageBiomedical DataComputed TomographyElectronic Health Records

🎯 What it does: Proposed an anatomy-based hierarchical multi-instance learning (ASH-MIL) framework to detect pulmonary diseases in chest X-rays using weak supervision.

Anatomy-Texture Aware Safe Gradient Guidance for Source-Free Domain Adaptive Echocardiography Video Segmentation

Lv, Jinrong (Southwest Jiaotong University), Jiang, Weili (Southwest Jiaotong University)

SegmentationDomain AdaptationConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningVideoBiomedical DataUltrasound

🎯 What it does: Propose a source domain-agnostic domain adaptation method for ultrasound video segmentation, adopting a decoupled anatomical structure and texture, along with a safe gradient guidance mechanism, to achieve robust adaptation to the target domain.

Angio-Stitch: An Unsupervised Coarse-to-Fine Framework for Sequential Stitching of Confocal Laser Endomicroscopy Images

Hu, Xiaoshi (Zhejiang University), Ye, Xuesong (Zhejiang University)

RestorationConvolutional Neural NetworkTransformerContrastive LearningOptical FlowImageSequentialBiomedical DataComputed TomographyUltrasound

🎯 What it does: Proposes an unsupervised coarse-to-fine framework called Angio-Stitch for sequentially merging liver CLE (confocal laser endomicroscopy) microvascular images, expanding the field of view while preserving fine structures.

Angular-Constrained Hyperbolic Learning for Hierarchical Multimodal Survival Prediction

Yang, Haotian (East China Normal University), Wang, Yan (East China Normal University)

Representation LearningData-Centric LearningTransformerAuto EncoderGenerative Adversarial NetworkContrastive LearningImageTabularBiomedical DataElectronic Health Records

🎯 What it does: Proposed the HyperGP framework, transforming the multimodal survival prediction of全景病理切片 (panoramic pathological slides) and genomic data into a hierarchical hyper-surface alignment problem.

Anomaly Detection in Fetal Echocardiography via Cross-Modal Translation and Region-Discriminated Error Calibration

Yang, Qianye (University of Oxford), Noble, J. Alison (University of Oxford)

Anomaly DetectionConvolutional Neural NetworkDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningVideoBiomedical DataUltrasound

🎯 What it does: Designed and verified a fetal ultrasound CHD abnormality detection framework based on cross-modal translation and regional calibration error.

AnomExpert: Identifying and Selecting Anatomical Planes for Prenatal Ultrasound Anomaly Diagnosis

Wang, Jian (Shenzhen University), Ni, Dong (Shenzhen University)

ClassificationAnomaly DetectionTransformerContrastive LearningImageBiomedical DataUltrasound

🎯 What it does: Developed the AnomExpert framework, which achieves prenatal ultrasound anomaly diagnosis using weak supervision (only case labels);

APEX-SAM: Anatomy-Aware Prompting with Expert Retrieval for Training-Free Medical Image Segmentation

Mao, Zhihao (China University of Geosciences (Wuhan)), Sun, Kun (China University of Geosciences (Wuhan))

SegmentationRetrievalDomain AdaptationTransformerPrompt EngineeringMixture of ExpertsContrastive LearningImageMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyRetrieval-Augmented Generation

🎯 What it does: Propose a training-free cross-domain few-shot medical image segmentation framework, APEX-SAM, which achieves instant segmentation of unknown anatomical structures by leveraging expert retrieval, anatomy-aware prompt mining, and multi-modal fusion.

ARCH-Net: Aligned 3D Teeth Reconstruction via Counterfactual-Enhanced Hierarchical Framework

Deng, Qingxin (Macao Polytechnic University), Zhang, Dian (Lingnan University)

RestorationGenerationTransformerDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningImagePoint CloudMesh

🎯 What it does: Proposed ARCH-Net, a three-stage alignment 3D tooth reconstruction framework that can generate high-precision 3D tooth models from a small number of uncalibrated intraoral photographs.

Are Foundational Models Less Biased Than Specialized Models? An Ultrasound Study

Fournel, Joris (Technical University of Denmark), Feragen, Aasa (University of Copenhagen)

ClassificationImage TranslationAnomaly DetectionConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningImageBiomedical DataUltrasound

🎯 What it does: Compared the bias performance of foundation models (FM) obtained via self-supervised pre-training (SSL) with that of traditional fully supervised training (CS) models on two prenatal ultrasound tasks (preterm birth prediction and fetal weight estimation).

Arrhythmia-Robust Cine-MRI via Latent Motion Artifact Characterization

Ning, Gaoning (Zhejiang University), Liu, Huafeng (Zhejiang University)

RestorationTransformerSupervised Fine-TuningDiffusion modelAuto EncoderContrastive LearningOptical FlowImageVideoBiomedical DataMagnetic Resonance ImagingElectrocardiogram

🎯 What it does: Propose a motion denoising framework based on latent dictionary encoding, LDE-MAR, for eliminating artifacts in cardiac cine-MRI caused by arrhythmia.

Artery and vein segmentation on brain CTA using CTP for contrast-phase augmentation

van Voorst, Henk (Stanford University), Heit, Jeremy (Stanford University)

SegmentationConvolutional Neural NetworkDiffusion modelContrastive LearningBiomedical DataComputed Tomography

🎯 What it does: We propose a deep learning method that utilizes CT perfusion (CTP) source data for contrast phase augmentation, achieving the segmentation of arteries and veins on cerebral CTA, and automatically calculating vascular quantitative scores based on this;

Artifact-aware Prompt Learning for Domain Generalization

Jia, Xibin (Beijing University of Technology), Fan, Chao (Beijing University of Technology)

ClassificationDomain AdaptationTransformerPrompt EngineeringContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound

🎯 What it does: This paper proposes a domain generalization framework called APDG based on visual Prompt learning, aiming to model and solve domain shift problems caused by imaging artifacts in skin lesion classification.

AS-TIME: Time-Conditioned Multimodal Modeling of Aortic Stenosis Progression

Kim, Diane (University Of British Columbia), Abolmaesumi, Purang (Vancouver General Hospital)

ClassificationExplainability and InterpretabilityTransformerMixture of ExpertsVideoTextBiomedical DataUltrasound

🎯 What it does: Proposed a time-conditioned multimodal model, AS-TIME, which predicts the risk of aortic stenosis (AS) progression using a single baseline echocardiogram and time interval ∆t.

ASTAR: Automated Induction of STAndardized Radiology Reporting Templates from Large-Scale Clinical Free-Text Corpora

Zhang, Xinfeng (Tsinghua University), Tian, Qiyuan (Sichuan University)

Explainability and InterpretabilityRepresentation LearningData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextBiomedical DataMagnetic Resonance ImagingRetrieval-Augmented Generation

🎯 What it does: This paper proposes the ASTAR framework, which automatically induces standardized report templates from large-scale radiology free-text reports using large language models.

Asym-MSF: Asymmetric Multi-scale Structural Fusion for Robust BCVA Prediction with Self-supervised Fidelity Constraints

Wang, Yifei (Wuhan University), Cao, Danmin (Aier Eye Hospital of Wuhan University)

ClassificationRepresentation LearningConvolutional Neural NetworkTransformerContrastive LearningImageTextMultimodalityBiomedical DataComputed TomographyElectronic Health Records

🎯 What it does: Propose an Asym-MSF framework based on OCT prioritization, an asymmetric multi-scale structure fusion framework, for predicting best corrected visual acuity after cataract surgery.

AsynDiff: Asynchronous-Timestep Diffusion with Noise-Aware Attention for Multi-sequence MRI Synthesis

Han, Luyi (Macao Polytechnic University), Tan, Tao (Macao Polytechnic University)

GenerationData SynthesisTransformerDiffusion modelScore-based ModelRectified FlowGenerative Adversarial NetworkImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: This paper proposes an asynchronous time-step diffusion framework, AsynDiff, for synthesizing missing sequences in multi-sequence MRI, and achieves cross-sequence information fusion through noise-aware attention.

AtomiMed: Hierarchical Atomic Fact-Checking for Universal Clinical-Aware Medical Report Evaluation

Wang, Yuan (Zhejiang University-University of Illinois Urbana-Champaign Institute, Zhejiang University), Zhang, Jianpeng (DAMO Academy, Alibaba Group)

Explainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelAgentic AIPrompt EngineeringMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasoundBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose a multi-modal medical report generation evaluation framework called AtomiMed, which decomposes reports into atomic clinical facts at the disease level and attribute level, and verifies clinical consistency through bidirectional Agentic Cross-Verification.

Attention is Matter for Inclusiveness: Generating Synthetic CT for Patients with Hip Implants

Zala, Nico Camillo (University Hospital Zurich), Dal Bello, Riccardo (University Hospital Zurich)

GenerationData SynthesisDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose a deep learning-based MR-only synthetic CT generation framework specifically for handling images of patients with metal implants (especially hip replacements and intramedullary nails), automatically extracting metal masks from MR images and improving the reconstruction quality of metal regions through attention mechanisms.

Attention-Based Prototype Calibration for Multi-Rater Few-Shot Medical Image Segmentation

Vu, Truong (Mohamed bin Zayed University of Artificial Intelligence), Xie, Yutong (Mohamed bin Zayed University of Artificial Intelligence)

SegmentationTransformerAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: This paper proposes a prototype calibration framework based on attention (JAPC), achieving multi-rater personalized prediction in few-shot medical image segmentation.

Attenuation-Resilient Alternating Optimization for Laparoscopic Liver Landmark Detection

Liu, Lanqing (Hong Kong Polytechnic University), Qin, Jing (Chinese University of Hong Kong)

SegmentationOptimizationConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningGaussian SplattingImageMultimodalityComputed TomographyReview/Survey Paper

🎯 What it does: Propose A2ONet, combining illumination compensation, frequency-domain directional filtering, and alternating segmentation-curve optimization to enhance the robustness of detection for laparoscopic liver surface markings.

Attribute-Aware Federated Learning with Pseudo-Bag Augmentation for Multi-Center WSI Classification

Deng, Yingjiao (East China Normal University), Li, Qingli (East China Normal University)

ClassificationDomain AdaptationFederated LearningData-Centric LearningContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasound

🎯 What it does: Designed the FedGMM-Bag framework, which generates pseudo bags using attribute-aware Gaussian Mixture Models (GMM) in a federated learning environment to classify whole slide images (WSI).

AutoLand: A Data-Efficient Automated Agentic Workflow for Universal Medical Landmark Detection

Li, Ziyi (Northeastern University), Qian, Wei (Northeastern University)

RecognitionPose EstimationData-Centric LearningTransformerAgentic AIPrompt EngineeringVision Language ModelImageMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasound

🎯 What it does: Propose AutoLand, an automated workflow based on Vision Language Models, which leverages physical fingerprints and a medical deep learning library to automatically generate feature point detection strategies suitable for new medical image datasets, significantly reducing the cost of manual parameter tuning.

Automated Disentangling Analysis of Skin Colour for Lesion Images

Yang, Wenbo (University of Waterloo), Wang, Zhou (University of Waterloo)

ClassificationImage TranslationRestorationAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical Data

🎯 What it does: Proposed an unsupervised skin color decoupling framework that can learn skin color captured in skin images (SCCI) and achieve controllable color conversion, enhancement, and normalization;

Automated Optical Density Normalization for Myelin Quantification: Cross-Modal Validation with 7T Ex Vivo MRI

Khodakarami, Zahra (University of Pennsylvania), Yushkevich, Paul A. (University of Pennsylvania)

Image TranslationImage HarmonizationRestorationSegmentationDomain AdaptationExplainability and InterpretabilityConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningOptical FlowImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Proposed an automated pipeline that utilizes a reverse attention mechanism and blind pigment deconvolution to perform optical density (OD) normalization on LFB-CV stained whole-slide images, obtaining quantified myelin levels; subsequently, the OD map was pixel-wise registered with the corresponding 7T ex vivo T2w MRI to validate the correlation between OD and MRI signals.

Automated Scope Pose Estimation in Real-World Robotic-Assisted Bronchoscopy with Integrated C-Arm Imaging

Zhang, Dingzhong (Johnson & Johnson MedTech Surgery), Matinfar, Babak (Johnson & Johnson MedTech Surgery)

SegmentationPose EstimationRobotic IntelligenceConvolutional Neural NetworkAuto EncoderContrastive LearningOptical FlowImageBiomedical DataComputed Tomography

🎯 What it does: This paper proposes a complete automated workflow for directly estimating the bronchoscope pose from CBCT images in robot-assisted bronchoscopy (RAB).

Automatic LV Localization and Short-Axis Plane Estimation from Arbitrary CMR Slice

Yu, Yi (Ohio State University), Xue, Yuan (Ohio State University)

Object DetectionPose EstimationConvolutional Neural NetworkSupervised Fine-TuningContrastive LearningBiomedical DataMagnetic Resonance Imaging

🎯 What it does: This paper proposes a method that simultaneously accomplishes left ventricle localization and three-dimensional SAX plane orientation estimation from a single arbitrary CMR slice.

B-spline Activations as Intrinsic Uncertainty Estimators in Medical Image Segmentation

Du, Wenju (Yangzhou University), Liu, Wei (Yangzhou University)

SegmentationExplainability and InterpretabilityComputational EfficiencyConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound

🎯 What it does: This paper introduces a residual KAN head in the output layer of a medical image segmentation model, leveraging the local support property of B-spline basis functions in the KAN layer to directly extract activation dispersion as an uncertainty estimate in a single forward pass.

B² Loss for Evidential Classification of Critical View of Safety in Laparoscopic Cholecystectomy

Nowak, Franciszek M. (UCL), Clarkson, Matthew J. (UCL)

ClassificationAnomaly DetectionSafty and PrivacyTransformerContrastive LearningVideoBiomedical Data

🎯 What it does: This paper proposes a Beta-Binomial evidence learning framework that probabilistically models the determination of critical safety views (CVS) during laparoscopic cholecystectomy. It directly utilizes the voting ratio information from multiple experts instead of hard labels for training, and introduces a video-level precision constraint evaluation.

Bayesian Temporal Pose Networks for Uncertainty-Calibrated Laparoscopic Tool Pose Tracking

Choudhry, Omar (University of Leeds), Jones, Dominic (University of Leeds)

Object TrackingPose EstimationRecurrent Neural NetworkTransformerVision-Language-Action ModelContrastive LearningOptical FlowVideoBiomedical DataMagnetic Resonance ImagingComputed TomographyBenchmark

🎯 What it does: Developed a 7-degree-of-freedom laparoscopic tool pose tracking framework based on Bayesian Time Pose Network (BTPN), which does not require geometric priors and can quantify uncertainty.

Bayesian Topological Analysis of Brain Networks

Zhu, Xukun (Duke University), Songdechakraiwut, Tananun

Mixture of ExpertsContrastive LearningBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease

🎯 What it does: The study proposes a Bayesian inference-based framework for comparing brain network topology, modeling the Wasserstein distance of the network using Gamma likelihood and Log-Normal prior, and estimating the posterior distribution through MCMC sampling;

BC-MultiSet: A Multi-target Dataset and Benchmark for Clinical Tasks in Breast Cancer

Illarionova, Svetlana (Applied AI Institute), Sharaev, Maxim (Applied AI Institute)

ClassificationSegmentationData SynthesisConvolutional Neural NetworkTransformerContrastive LearningImageBiomedical DataBenchmark

🎯 What it does: Created the BC-MultiSet dataset and conducted benchmark evaluations of multi-task deep learning models on it, exploring the association between cell morphology and clinical molecular markers.

BCER Agent: Reliable Long-Horizon MRI Workflow Execution via Compilation, Artifact Binding, and Bounded Local Recovery

Long, Ziyang (Cedars-Sinai Medical Center), Yang, Hsin-Jung (Cedars-Sinai Medical Center)

SegmentationGenerationData SynthesisTransformerLarge Language ModelAgentic AIPrompt EngineeringVision Language ModelImageBiomedical DataMagnetic Resonance ImagingBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Developed the BCER framework, which can reliably execute long-cycle MRI workflows, supporting the separation of high-level planning and compilable execution, symbolic binding, and local recovery, and was evaluated on a multi-organ multi-task benchmark.

BDFR-Net: Boundary-Aware Dual-Domain Feature Reconstruction Network for Fetal Ultrasound Image Segmentation

Zhang, Wenfeng (Chongqing Normal University), Bai, Jieyun (Jinan University)

SegmentationConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningImageBiomedical DataUltrasound

🎯 What it does: This paper proposes a dual-domain feature reconstruction network, BDFR-Net, for the segmentation of fetal head and pubic symphysis in ultrasound images.

Be Indiscrete: The Benefits of Learning Continuous Spine Degeneration Severity Scores

Monzon, Maria (ETH Zurich), Jamaludin, Amir (University of Oxford)

ClassificationRepresentation LearningConvolutional Neural NetworkContrastive LearningBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Proposed SpineRankNet, which learns continuous spinal degeneration severity scores using contrastive learning, and achieves multi-task prediction on a single 3D ResNet-18 architecture.

Benchmarking Foundation Models for Early Multi-retinopathy Detection at National Scale: Performance, Generalization, and Bias

Cheng, Wenquan (Tsinghua University), Jia, Huixun (Shanghai Jiao Tong University)

ClassificationAnomaly DetectionTransformerSupervised Fine-TuningVision Language ModelContrastive LearningImageBiomedical DataBenchmark

🎯 What it does: The study constructs a nationwide benchmark for the early detection of multiple retinal diseases from fundus photographs (OphFMbench), and evaluates the performance, generalization, and fairness of three categories of foundational models on it.

Benchmarking Pathology Foundation Models for Spatial Domain Understanding

Zhao, Bokai (University of Chinese Academy of Sciences), Jiang, Tianzi (University of Chinese Academy of Sciences)

Representation LearningGraph Neural NetworkTransformerAuto EncoderContrastive LearningImageTabularBiomedical DataBenchmark

🎯 What it does: Designed SpaPath-Bench, utilizing paired WSI and spatial transcriptomics data to evaluate the ability of pathological foundation models to understand the spatial domain.

Benchmarking Video Foundation Models for Remote Parkinson’s Disease Screening

Islam, Md Saiful (University of Rochester), Hoque, Ehsan (University of Cambridge)

ClassificationTransformerAuto EncoderContrastive LearningVideoBiomedical DataBenchmark

🎯 What it does: Constructed a large-scale benchmark, evaluating the performance of remote Parkinson's disease screening using seven video foundation models (VideoPrism, V-JEPA, V-JEPA-SSv2, ViViT, TimeSformer, VideoMAE, VideoMAEv2) on 32,847 videos from 1,888 subjects (727 with Parkinson's disease) performing 16 standard clinical tasks, systematically comparing the performance of different VFM architectures across tasks.

Better Said Than Seen: Exposing and Mitigating Modality Collapse with ICD-11-Grounded Evaluation

Hassan, Lara (Mohamed Bin Zayed University of Artificial Intelligence), Mahmoud, Abdulrahman (Mohamed Bin Zayed University of Artificial Intelligence)

ClassificationExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextMultimodalityBiomedical DataElectronic Health RecordsRetrieval-Augmented Generation

🎯 What it does: The study investigates the phenomenon where image information in multimodal large language models is suppressed by text priors in dermatology diagnosis, and proposes alleviating and evaluating this issue by replacing pixels with structured clinical vignettes and using an ICDLens assessment framework based on ICD-11.

Beyond Clean Test Sets: Spurious Correlations in Medical Vision-language Models and the Role of Concept Supervision

Corbetta, Valentina (Netherlands Cancer Institute), Wickstrøm, Kristoffer (Arctic University of Norway)

ClassificationImage TranslationRestorationAnomaly DetectionExplainability and InterpretabilityComputational EfficiencyRepresentation LearningAdversarial AttackData-Centric LearningTransformerSupervised Fine-TuningPrompt EngineeringVision Language ModelContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasoundReview/Survey PaperBenchmark

🎯 What it does: A set of adjustable parameter synthetic artifact and artifact removal/inversion dual protocols was proposed for systematically evaluating the robustness of medical vision-language models in diabetic retinopathy grading and breast imaging diagnosis, with experiments conducted on five different conceptual supervision levels of architectures.

Beyond Landmarks: Anchor-Guided Orientation Regression for Robust Spinopelvic Measurement

Hwang, Yunseob (Gwangju Institute of Science and Technology), Ro, Du Hyun (Gwangju Institute of Science and Technology)

Pose EstimationConvolutional Neural NetworkSupervised Fine-TuningContrastive LearningImage

🎯 What it does: This paper proposes a new method to address a specific problem, aiming to improve performance and efficiency.