arXivSub Start free trial

MICCAI 2026 Papers with Code

International Conference on Medical Image Computing and Computer-Assisted Intervention Β· 649 papers with a public code repository

3D Classification of Paramagnetic Rim Lesions in Multiple Sclerosis via Asymmetric QSM–FLAIR Modeling

Pignedoli, Veronica (University of Genova), Moro, Matteo (University of Genova)

CodeClassificationAnomaly DetectionConvolutional Neural NetworkTransformerSupervised Fine-TuningAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: This study proposes FRODO, a 3D deep learning framework based on QSM and FLAIR, for automatically distinguishing paramagnetic rim lesions (Rim+) from non-Rim lesions (Rim-) in multiple sclerosis.

A Confounder-aware Representation Learning Framework for Alzheimer’s Disease Classification

Chen, Yuanhao (Zhejiang University of Finance and Economics), Wang, Changmiao (Shenzhen Research Institute of Big Data)

CodeClassificationRepresentation LearningConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningImageTabularBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease

🎯 What it does: This paper proposes a hybrid modal framework based on causal inference for the early diagnosis of Alzheimer's disease;

A Counterfactual Framework for Directional Cell–Cell Interaction Analysis in Spatial Transcriptomics

Anzum, Humaira (University of Houston), Banerjee, Tania (University of Houston)

CodeDrug DiscoveryGraph Neural NetworkGenerative Adversarial NetworkContrastive LearningBiomedical Data

🎯 What it does: Designed a directional cell-cell interaction framework based on adversarial inference, utilizing a neighborhood graph model to predict receptor cell states and quantify directional effects through hypothetical interventions.

A Flexible Structure-Guided Feature Aggregation Paradigm for Efficient Vascular Representation

Qi, Yaolei (Southeast University), Yang, Guanyu (Southeast University)

CodeSegmentationComputational EfficiencyRepresentation LearningConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningOptical FlowImagePoint CloudBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound

🎯 What it does: Propose VSP-Branch, a pluggable structural-guided feature aggregation module, which enhances the continuity and branch integrity of vascular representations by constructing learnable paths along the local vascular structure, aggregating cross-layer features, and adaptively injecting structural information through difficulty-gated mechanisms.

A Guideline-Aware AI Agent for Zero-Shot Target Volume Auto-Delineation

Kim, Yoon Jo (Oncosoft Inc.), Kim, Jin Sung (Oncosoft Inc.)

CodeSegmentationExplainability and InterpretabilityConvolutional Neural NetworkTransformerLarge Language ModelAgentic AIImageTextBiomedical DataComputed TomographyRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed an AI agent framework called OncoAgent, which can zero-shot automatically convert text-based radiotherapy clinical guidelines into three-dimensional target volume volumes;

A Hierarchical Multi-Task Framework for Dementia Diagnosis via Pathological Feature Learning from Multi-Organ Data

Zhao, Shilun (ShanghaiTech University), Shen, Dinggang (ShanghaiTech University)

CodeClassificationRepresentation LearningConvolutional Neural NetworkTransformerMixture of ExpertsContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyAlzheimer's DiseaseElectronic Health Records

🎯 What it does: Propose the MoMmNet framework, which utilizes multi-organ imaging (brain, heart, gut, liver, kidney) and clinical text for hierarchical multi-task reasoning to achieve multi-cause dementia diagnosis.

A Multi-center Benchmark for Abdominal Disease Diagnosis and Report Generation from Non-Contrast CT

Elbakry, Mariam (Ain Shams University), Elbatel, Marawan (Hong Kong University of Science and Technology)

CodeClassificationGenerationExplainability and InterpretabilityConvolutional Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningImageTextBiomedical DataComputed TomographyBenchmarkRetrieval-Augmented Generation

🎯 What it does: Designed and released a multi-center benchmark to evaluate the feasibility of generating multi-organ diagnostic reports from single-phase non-contrast CT.

A Multi-view, Hybrid, Hypergraph Learning Framework with Causal Perturbation for Early Alzheimer’s Disease Diagnosis

Jia, Yifan (ShanghaiTech University), Zhang, Han (ShanghaiTech University)

CodeClassificationExplainability and InterpretabilityConvolutional Neural NetworkGraph Neural NetworkTransformerSupervised Fine-TuningContrastive LearningGraphTime SeriesBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease

🎯 What it does: Proposes HyperBrainNet, a framework combining causal effective connectivity, high-order hybrid hypergraph networks, and multi-view adaptive fusion, for the diagnosis of early Alzheimer's disease.

A Multimodal 3D Foundation Model for Light Sheet Fluorescence Microscopy Enables Few-Shot Segmentation, Classification, and Deblurring

Scheinfeld, Adina (Weill Cornell Medicine), Paetzold, Johannes C. (Weill Cornell Medicine)

CodeClassificationRestorationSegmentationKnowledge DistillationRepresentation LearningConvolutional Neural NetworkTransformerSupervised Fine-TuningVision Language ModelContrastive LearningImageTextBiomedical Data

🎯 What it does: This paper constructs a multi-modal 3D foundation model, pre-trained using a large-scale unlabeled light sheet fluorescence microscopy (LSM) volume images and corresponding text descriptions, achieving efficient few-shot fine-tuning on downstream tasks such as segmentation, classification, and deblurring.

A Real-World Evaluation of Failure Detection for Liver CT Segmentation

Bennett, Jeddy (University of Texas MD Anderson Cancer Center), Brock, Kristy K. (University of Texas MD Anderson Cancer Center)

CodeSegmentationAnomaly DetectionConvolutional Neural NetworkContrastive LearningImageBiomedical DataComputed Tomography

🎯 What it does: Evaluate and compare multiple OOS detection methods in real clinical settings for failure detection in automatic liver CT segmentation models.

A Unified Framework for Joint Detection of Lacunes and Enlarged Perivascular Spaces

He, Lucas (University College London), Sudre, Carole H. (University College London)

CodeObject DetectionSegmentationConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose a unified multi-task framework that jointly detects adenomas (EPVS) and small cavities (lacunes) in the brain and improves detection accuracy.

AC-MIL: Weakly-Supervised Atrial LGE-MRI Quality Assessment via Adversarial Concept Disentanglement

Sultan, K. M. Arefeen (University of Utah), Elhabian, Shireen Y. (University of Utah)

CodeClassificationExplainability and InterpretabilityConvolutional Neural NetworkTransformerGenerative Adversarial NetworkContrastive LearningBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose the AC-MIL framework to achieve weakly supervised atrial LGE-MRI quality assessment, decomposing overall quality into interpretable clinical concepts;

ACA: Post-hoc Adaptive Logit Alignment for 3D Medical Image Segmentation

Dong, Nanyu (Adelaide University), Liao, Zhibin (Adelaide University)

CodeSegmentationContrastive LearningBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound

🎯 What it does: To address the insufficient generalization of pre-trained models for 3D medical image segmentation on unseen data, this paper proposes a post-adaptive clipping alignment (ACA) method. It enhances segmentation performance by achieving local distribution alignment within the model's uncertain logit interval through isotropic quantile mapping and isotropic regression.

AdaSurvMamba: Dynamic Fusion and Semantic Scanning for Multimodal Survival Analysis

Zhong, Jialong (Dalian University of Technology), Lu, Huchuan (Dalian University of Technology)

CodeRepresentation LearningTransformerContrastive Learning

🎯 What it does: Propose an AdaSurvMamba framework applicable to multi-modal survival analysis, combining WSI and genomic data to achieve accurate cancer prognosis prediction.

Addressing Gradient Conflicts in Multimodal Fundus Disease Recognition with Fusion-Guided Learning

Liu, Xiaozhou (Southwest University), Du, Zhiguo (Southwest University)

CodeRecognitionOptimizationConvolutional Neural NetworkSupervised Fine-TuningContrastive LearningImageMultimodalityBiomedical DataAlzheimer's DiseaseReview/Survey Paper

🎯 What it does: Address gradient conflicts in multi-modal retinal image classification by introducing the fusion-guided gradient decoupling learning (FGDL) and cross-modal semantic alignment (CMSA) modules to improve dynamic optimization and enhance feature alignment;

Addressing Tissue and Appearance Heterogeneity in Text-Guided Few-Shot WSI Classification

Li, Yongcen (Dalian University of Technology), Xing, Xudong (Beijing Institute of Genomics)

CodeClassificationDomain AdaptationTransformerVision Language ModelContrastive LearningImageTextBiomedical DataComputed TomographyReview/Survey Paper

🎯 What it does: This paper proposes a Heterogeneity-Aware Text-guided MIL framework to address the domain drift problem caused by tissue composition and appearance heterogeneity in whole slide image classification under the extremely few-sample condition.

AEGIS: Anatomy-Embedded Group-Invariant Segmentation for Fair Medical Foundation Models

Wang, Sen (Shenzhen MSU-BIT University), Zhao, Weibing (Shenzhen MSU-BIT University)

CodeSegmentationFederated LearningExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerSupervised Fine-TuningAuto EncoderContrastive LearningBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound

🎯 What it does: Addressing the fairness issues of medical foundation models (such as SAM and MedSegX) in the segmentation of head and neck squamous cell carcinoma (HNSCC), this paper proposes the AEGIS framework, achieving group-invariant segmentation under parameter-efficient fine-tuning (PEFT).

Age Conditional Longitudinal Forecasting of Adolescent Functional Connectivity via Brownian Bridge Diffusion Models

Zhang, Rongye (Taiyuan University of Technology), Wen, Xin (Taiyuan University of Technology)

CodeGenerationData SynthesisGraph Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderBiomedical DataMagnetic Resonance ImagingStochastic Differential Equation

🎯 What it does: Propose a framework that combines graph attention autoencoder with Brownian bridge diffusion for long-term evolution prediction of adolescent functional connectivity under age conditions.

AGE-MIL: Anchor-Guided Evidence Learning for Patient-Level Prediction

Niu, Jiawei (Xi'an Jiaotong University), Cai, Yi (Central South University)

CodeClassificationImage TranslationRetrievalAnomaly DetectionOptimizationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningData-Centric LearningConvolutional Neural NetworkTransformerMixture of ExpertsAuto EncoderContrastive LearningImageBiomedical DataElectronic Health RecordsReview/Survey PaperRetrieval-Augmented Generation

🎯 What it does: Proposed an Anchor-Guided Evidence MIL (AGE-MIL) framework for patient-level diagnosis and prognosis prediction.

AI-Driven Pulmonary Congestion Assessment for Lung Ultrasound via Segmentation-Guided Transformers

Fooladgar, Fahimeh, Kapur, Tina (Brigham and Women's Hospital)

CodeClassificationSegmentationAnomaly DetectionExplainability and InterpretabilityConvolutional Neural NetworkTransformerPrompt EngineeringMixture of ExpertsContrastive LearningImageBiomedical DataUltrasound

🎯 What it does: Propose a two-stage segmentation-guided transformer framework for automated severity scoring of lung B-lines, enabling consistent and reproducible assessment of pulmonary edema directly at the bedside in lung ultrasound.

An Artifact-Based Agent Framework for Adaptive and Reproducible Medical Image Processing

Zuo, Lianrui (Vanderbilt University), Landman, Bennett A. (Vanderbilt University)

CodeImage TranslationRestorationSegmentationReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Proposed an Agent framework based on artifact contracts, achieving adaptive configuration and reproducible execution for medical image processing;

Anatomically Accurate 3D Vessel Generation via 3-Phase Chebyshev Curve Diffusion

Mo, Jihwan (Korea Advanced Institute of Science and Technology), Chang, Dong Eui (Asan Medical Center)

CodeGenerationData SynthesisTransformerPrompt EngineeringDiffusion modelScore-based ModelImageBiomedical DataComputed Tomography

🎯 What it does: A model is proposed that generates continuous, C∞ smooth vascular centerlines through the diffusion of three-phase Chebyshev curves, capable of recovering millimeter-level dimensions while maintaining the curve trajectory.

Anatomy-Aware Hierarchical Multiple Instance Learning for Interpretable COPD Diagnosis and Phenotype Analysis

Ahmadi, Raha (University of British Columbia), Tam, Roger (University of British Columbia)

CodeClassificationExplainability and InterpretabilityConvolutional Neural NetworkContrastive LearningImageBiomedical DataComputed Tomography

🎯 What it does: Proposed the AnatoMIL framework for COPD detection and grading from CT images.

Anatomy-Conditioned Domain Randomization for Zero-Shot 3D Cerebrovascular Segmentation

Pentassuglia, Matteo (EURECOM), Zuluaga, Maria A. (EURECOM)

CodeSegmentationData SynthesisDomain AdaptationConvolutional Neural NetworkDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose a domain randomization framework based on whole-brain anatomical labels, using only synthetic images to train a 3D CNN, achieving zero-shot brain vessel segmentation on unseen TOF-MRA and CTA data;

Anatomy-Grounded Synthetic Coronary Angiography for Geometry-Informed Multi-view Matching

Lee, In Kyu (Medipixel, Inc.), Min, Jaesik (Medipixel, Inc.)

CodeImage TranslationData SynthesisPose EstimationDepth EstimationTransformerDiffusion modelContrastive LearningOptical FlowImageBiomedical DataComputed Tomography

🎯 What it does: This paper generates digital reconstructed radiographs (DRR) from patient CT angiography data to construct a multi-view coronary artery correspondence dataset without manual annotations, and proposes a Geometry-Informed Matching Module (GIMM) to achieve fine matching.

Anatomy-Guided Residual Motion Diffusion for Controllable 4D Cardiac MRI Synthesis

Cao, Yiheng (Suzhou Institute of Biomedical Engineering and Technology, Chinese Academy of Science), Gao, Xin

CodeSegmentationGenerationData SynthesisTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Proposed a 4D motion medical image generation framework based on semi-supervised VAE and two-stage residual latent diffusion models, which can simultaneously generate voxel sequences and corresponding segmentation masks, achieving controllable 4D cardiac MRI synthesis.

Anatomy-Texture Aware Safe Gradient Guidance for Source-Free Domain Adaptive Echocardiography Video Segmentation

Lv, Jinrong (Southwest Jiaotong University), Jiang, Weili (Southwest Jiaotong University)

CodeSegmentationDomain AdaptationConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningVideoBiomedical DataUltrasound

🎯 What it does: Propose a source domain-agnostic domain adaptation method for ultrasound video segmentation, adopting a decoupled anatomical structure and texture, along with a safe gradient guidance mechanism, to achieve robust adaptation to the target domain.

Angio-Stitch: An Unsupervised Coarse-to-Fine Framework for Sequential Stitching of Confocal Laser Endomicroscopy Images

Hu, Xiaoshi (Zhejiang University), Ye, Xuesong (Zhejiang University)

CodeRestorationConvolutional Neural NetworkTransformerContrastive LearningOptical FlowImageSequentialBiomedical DataComputed TomographyUltrasound

🎯 What it does: Proposes an unsupervised coarse-to-fine framework called Angio-Stitch for sequentially merging liver CLE (confocal laser endomicroscopy) microvascular images, expanding the field of view while preserving fine structures.

Angular-Constrained Hyperbolic Learning for Hierarchical Multimodal Survival Prediction

Yang, Haotian (East China Normal University), Wang, Yan (East China Normal University)

CodeRepresentation LearningData-Centric LearningTransformerAuto EncoderGenerative Adversarial NetworkContrastive LearningImageTabularBiomedical DataElectronic Health Records

🎯 What it does: Proposed the HyperGP framework, transforming the multimodal survival prediction ofε…¨ζ™―η—…η†εˆ‡η‰‡ (panoramic pathological slides) and genomic data into a hierarchical hyper-surface alignment problem.

Anomaly Detection in Fetal Echocardiography via Cross-Modal Translation and Region-Discriminated Error Calibration

Yang, Qianye (University of Oxford), Noble, J. Alison (University of Oxford)

CodeAnomaly DetectionConvolutional Neural NetworkDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningVideoBiomedical DataUltrasound

🎯 What it does: Designed and verified a fetal ultrasound CHD abnormality detection framework based on cross-modal translation and regional calibration error.

AnomExpert: Identifying and Selecting Anatomical Planes for Prenatal Ultrasound Anomaly Diagnosis

Wang, Jian (Shenzhen University), Ni, Dong (Shenzhen University)

CodeClassificationAnomaly DetectionTransformerContrastive LearningImageBiomedical DataUltrasound

🎯 What it does: Developed the AnomExpert framework, which achieves prenatal ultrasound anomaly diagnosis using weak supervision (only case labels);

Arrhythmia-Robust Cine-MRI via Latent Motion Artifact Characterization

Ning, Gaoning (Zhejiang University), Liu, Huafeng (Zhejiang University)

CodeRestorationTransformerSupervised Fine-TuningDiffusion modelAuto EncoderContrastive LearningOptical FlowImageVideoBiomedical DataMagnetic Resonance ImagingElectrocardiogram

🎯 What it does: Propose a motion denoising framework based on latent dictionary encoding, LDE-MAR, for eliminating artifacts in cardiac cine-MRI caused by arrhythmia.

AS-TIME: Time-Conditioned Multimodal Modeling of Aortic Stenosis Progression

Kim, Diane (University Of British Columbia), Abolmaesumi, Purang (Vancouver General Hospital)

CodeClassificationExplainability and InterpretabilityTransformerMixture of ExpertsVideoTextBiomedical DataUltrasound

🎯 What it does: Proposed a time-conditioned multimodal model, AS-TIME, which predicts the risk of aortic stenosis (AS) progression using a single baseline echocardiogram and time interval βˆ†t.

ASTAR: Automated Induction of STAndardized Radiology Reporting Templates from Large-Scale Clinical Free-Text Corpora

Zhang, Xinfeng (Tsinghua University), Tian, Qiyuan (Sichuan University)

CodeExplainability and InterpretabilityRepresentation LearningData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextBiomedical DataMagnetic Resonance ImagingRetrieval-Augmented Generation

🎯 What it does: This paper proposes the ASTAR framework, which automatically induces standardized report templates from large-scale radiology free-text reports using large language models.

AsynDiff: Asynchronous-Timestep Diffusion with Noise-Aware Attention for Multi-sequence MRI Synthesis

Han, Luyi (Macao Polytechnic University), Tan, Tao (Macao Polytechnic University)

CodeGenerationData SynthesisTransformerDiffusion modelScore-based ModelRectified FlowGenerative Adversarial NetworkImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: This paper proposes an asynchronous time-step diffusion framework, AsynDiff, for synthesizing missing sequences in multi-sequence MRI, and achieves cross-sequence information fusion through noise-aware attention.

AtomiMed: Hierarchical Atomic Fact-Checking for Universal Clinical-Aware Medical Report Evaluation

Wang, Yuan (Zhejiang University-University of Illinois Urbana-Champaign Institute, Zhejiang University), Zhang, Jianpeng (DAMO Academy, Alibaba Group)

CodeExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelAgentic AIPrompt EngineeringMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasoundBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose a multi-modal medical report generation evaluation framework called AtomiMed, which decomposes reports into atomic clinical facts at the disease level and attribute level, and verifies clinical consistency through bidirectional Agentic Cross-Verification.

Attention is Matter for Inclusiveness: Generating Synthetic CT for Patients with Hip Implants

Zala, Nico Camillo (University Hospital Zurich), Dal Bello, Riccardo (University Hospital Zurich)

CodeGenerationData SynthesisDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose a deep learning-based MR-only synthetic CT generation framework specifically for handling images of patients with metal implants (especially hip replacements and intramedullary nails), automatically extracting metal masks from MR images and improving the reconstruction quality of metal regions through attention mechanisms.

Attention-Based Prototype Calibration for Multi-Rater Few-Shot Medical Image Segmentation

Vu, Truong (Mohamed bin Zayed University of Artificial Intelligence), Xie, Yutong (Mohamed bin Zayed University of Artificial Intelligence)

CodeSegmentationTransformerAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: This paper proposes a prototype calibration framework based on attention (JAPC), achieving multi-rater personalized prediction in few-shot medical image segmentation.

Attenuation-Resilient Alternating Optimization for Laparoscopic Liver Landmark Detection

Liu, Lanqing (Hong Kong Polytechnic University), Qin, Jing (Chinese University of Hong Kong)

CodeSegmentationOptimizationConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningGaussian SplattingImageMultimodalityComputed TomographyReview/Survey Paper

🎯 What it does: Propose A2ONet, combining illumination compensation, frequency-domain directional filtering, and alternating segmentation-curve optimization to enhance the robustness of detection for laparoscopic liver surface markings.

AutoLand: A Data-Efficient Automated Agentic Workflow for Universal Medical Landmark Detection

Li, Ziyi (Northeastern University), Qian, Wei (Northeastern University)

CodeRecognitionPose EstimationData-Centric LearningTransformerAgentic AIPrompt EngineeringVision Language ModelImageMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasound

🎯 What it does: Propose AutoLand, an automated workflow based on Vision Language Models, which leverages physical fingerprints and a medical deep learning library to automatically generate feature point detection strategies suitable for new medical image datasets, significantly reducing the cost of manual parameter tuning.

Automated Disentangling Analysis of Skin Colour for Lesion Images

Yang, Wenbo (University of Waterloo), Wang, Zhou (University of Waterloo)

CodeClassificationImage TranslationRestorationAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical Data

🎯 What it does: Proposed an unsupervised skin color decoupling framework that can learn skin color captured in skin images (SCCI) and achieve controllable color conversion, enhancement, and normalization;

Automatic LV Localization and Short-Axis Plane Estimation from Arbitrary CMR Slice

Yu, Yi (Ohio State University), Xue, Yuan (Ohio State University)

CodeObject DetectionPose EstimationConvolutional Neural NetworkSupervised Fine-TuningContrastive LearningBiomedical DataMagnetic Resonance Imaging

🎯 What it does: This paper proposes a method that simultaneously accomplishes left ventricle localization and three-dimensional SAX plane orientation estimation from a single arbitrary CMR slice.

B-spline Activations as Intrinsic Uncertainty Estimators in Medical Image Segmentation

Du, Wenju (Yangzhou University), Liu, Wei (Yangzhou University)

CodeSegmentationExplainability and InterpretabilityComputational EfficiencyConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound

🎯 What it does: This paper introduces a residual KAN head in the output layer of a medical image segmentation model, leveraging the local support property of B-spline basis functions in the KAN layer to directly extract activation dispersion as an uncertainty estimate in a single forward pass.

Bayesian Temporal Pose Networks for Uncertainty-Calibrated Laparoscopic Tool Pose Tracking

Choudhry, Omar (University of Leeds), Jones, Dominic (University of Leeds)

CodeObject TrackingPose EstimationRecurrent Neural NetworkTransformerVision-Language-Action ModelContrastive LearningOptical FlowVideoBiomedical DataMagnetic Resonance ImagingComputed TomographyBenchmark

🎯 What it does: Developed a 7-degree-of-freedom laparoscopic tool pose tracking framework based on Bayesian Time Pose Network (BTPN), which does not require geometric priors and can quantify uncertainty.

BC-MultiSet: A Multi-target Dataset and Benchmark for Clinical Tasks in Breast Cancer

Illarionova, Svetlana (Applied AI Institute), Sharaev, Maxim (Applied AI Institute)

CodeClassificationSegmentationData SynthesisConvolutional Neural NetworkTransformerContrastive LearningImageBiomedical DataBenchmark

🎯 What it does: Created the BC-MultiSet dataset and conducted benchmark evaluations of multi-task deep learning models on it, exploring the association between cell morphology and clinical molecular markers.

BCER Agent: Reliable Long-Horizon MRI Workflow Execution via Compilation, Artifact Binding, and Bounded Local Recovery

Long, Ziyang (Cedars-Sinai Medical Center), Yang, Hsin-Jung (Cedars-Sinai Medical Center)

CodeSegmentationGenerationData SynthesisTransformerLarge Language ModelAgentic AIPrompt EngineeringVision Language ModelImageBiomedical DataMagnetic Resonance ImagingBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Developed the BCER framework, which can reliably execute long-cycle MRI workflows, supporting the separation of high-level planning and compilable execution, symbolic binding, and local recovery, and was evaluated on a multi-organ multi-task benchmark.

Be Indiscrete: The Benefits of Learning Continuous Spine Degeneration Severity Scores

Monzon, Maria (ETH Zurich), Jamaludin, Amir (University of Oxford)

CodeClassificationRepresentation LearningConvolutional Neural NetworkContrastive LearningBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Proposed SpineRankNet, which learns continuous spinal degeneration severity scores using contrastive learning, and achieves multi-task prediction on a single 3D ResNet-18 architecture.

Benchmarking Video Foundation Models for Remote Parkinson’s Disease Screening

Islam, Md Saiful (University of Rochester), Hoque, Ehsan (University of Cambridge)

CodeClassificationTransformerAuto EncoderContrastive LearningVideoBiomedical DataBenchmark

🎯 What it does: Constructed a large-scale benchmark, evaluating the performance of remote Parkinson's disease screening using seven video foundation models (VideoPrism, V-JEPA, V-JEPA-SSv2, ViViT, TimeSformer, VideoMAE, VideoMAEv2) on 32,847 videos from 1,888 subjects (727 with Parkinson's disease) performing 16 standard clinical tasks, systematically comparing the performance of different VFM architectures across tasks.

Better Said Than Seen: Exposing and Mitigating Modality Collapse with ICD-11-Grounded Evaluation

Hassan, Lara (Mohamed Bin Zayed University of Artificial Intelligence), Mahmoud, Abdulrahman (Mohamed Bin Zayed University of Artificial Intelligence)

CodeClassificationExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextMultimodalityBiomedical DataElectronic Health RecordsRetrieval-Augmented Generation

🎯 What it does: The study investigates the phenomenon where image information in multimodal large language models is suppressed by text priors in dermatology diagnosis, and proposes alleviating and evaluating this issue by replacing pixels with structured clinical vignettes and using an ICDLens assessment framework based on ICD-11.

Beyond Stochastic Diffusion: Trajectory-Consistent Deterministic Flow Matching for Fluorescence Molecular Tomography

Xue, Qianqian (Shanxi University), Wang, Wenjian (Shanxi University)

CodeDrug DiscoveryTransformerDiffusion modelFlow-based ModelBiomedical DataComputed TomographyOrdinary Differential Equation

🎯 What it does: Propose a reconstruction framework for fluorescent molecular tomography (FMT) named TCDFlow-FMT based on deterministic flow matching, achieving fast and high-quality three-dimensional fluorescent source reconstruction.

Beyond the Batch: Momentum-Updated Virtual Cohorts for 3D PET-CT Prognosis

Liang, Xinglong (Netherlands Cancer Institute), Mann, Ritse (Netherlands Cancer Institute)

CodeAnomaly DetectionOptimizationRepresentation LearningConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningBiomedical DataComputed TomographyPositron Emission Tomography

🎯 What it does: Propose the Momentum-Contrastive Survival Framework (MCSF), which expands the risk set in 3D PET-CT survival prediction through a virtual cohort and improves optimization stability using a time-weighted contrastive loss.

Beyond Visual Cues: CoT-Enhanced Reasoning for Semi-supervised Medical Image Segmentation

Chen, Yuming (Southeast University), Zhou, Yi (Nanjing University of Science and Technology)

CodeSegmentationRetrievalExplainability and InterpretabilityConvolutional Neural NetworkTransformerLarge Language ModelVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingComputed TomographyRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This paper proposes a semi-supervised medical image segmentation framework called CERS based on Chain-of-Thought (CoT), which enhances segmentation performance by jointly retrieving generated diagnostic reasoning text and visual features.

Beyond Visual Forensics: Auditing Multimodal Robustness for Synthetic Medical Image Detection

Chiu, Ching-Hao (University of Notre Dame), Shi, Yiyu (University of Notre Dame)

CodeAnomaly DetectionExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextBiomedical DataComputed TomographyElectronic Health RecordsBenchmark

🎯 What it does: Proposed an evaluation benchmark for detecting the authenticity of multimodal medical images, by fixing the image and only modifying the accompanying structured metadata to detect text-induced decision bias.

BiC-ODE: Bidirectionally Coupled Neural ODEs for Rapid Laminar Surface Reconstruction of the Human Cortex from 5T MRI

Cao, Shui (ShanghaiTech University), Shen, Dinggang (ShanghaiTech University)

CodeSegmentationOptimizationComputational EfficiencyDiffusion modelScore-based ModelRectified FlowContrastive LearningImageBiomedical DataMagnetic Resonance ImagingStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Propose the BiC-ODE framework for fast and high-precision reconstruction of four surface layers (white matter, inner surface, outer surface, pial surface) in 5T FLAIR images.

Bidirectional Anatomy-Aware Post-Training for Longitudinal Chest X-Ray Progression Modeling

Chen, Yuming (Adelaide University), Xie, Yutong (Mohamed bin Zayed University of Artificial Intelligence)

CodeClassificationRepresentation LearningTransformerSupervised Fine-TuningContrastive LearningImageBiomedical DataComputed TomographyElectronic Health Records

🎯 What it does: Proposed a lightweight post-training strategy called BAAP to improve longitudinal chest radiograph progression prediction;

BIGUS: Beam Integrated Gaussian Scattering for Continuous Ultrasound View Synthesis

Abdelaziz, Youssif (Applied Innovation Center), Torki, Marwan (Applied Innovation Center)

CodeData SynthesisNeural Radiance FieldGaussian SplattingBiomedical DataUltrasound

🎯 What it does: Propose the BIGUS framework, which utilizes 3D Gaussian splats to achieve continuous ultrasound view synthesis and accurately simulate ultrasound scattering to preserve speckle details.

BiM-GeoAttn-Net: Linear-Time Depth Modeling with Geometry-Aware Attention for 3D Aortic Dissection CTA Segmentation

Zhang, Yuan (Sichuan Normal University), Mu, Nan (Sichuan Normal University)

CodeSegmentationConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningBiomedical DataComputed Tomography

🎯 What it does: Propose BiM-GeoAttn-Net, combining bidirectional deep Mamba with geometry-aware vessel attention for aortic dissection segmentation in 3D CTA.

BioFact-MoE: Biologically Factorized Mixture of Experts for Vision–Language Prognostic Modeling in Hepatocellular Carcinoma

Yang, Junlin (Yale University), Chapiro, Julius (Yale University)

CodeExplainability and InterpretabilityDrug DiscoveryTransformerLarge Language ModelMixture of ExpertsVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose the BioFact-MoE framework, utilizing a Mixture of Experts at the gene level for visual-language prognostic modeling in hepatocellular carcinoma.

BioFlow: A Biologically Valid Support-Preserving Flow for Histology-Conditioned Spatial Transcriptomics Prediction

Xu, Haoran (Sichuan University), Han, Xiao (Sichuan University)

CodeGenerationData SynthesisTransformerDiffusion modelFlow-based ModelImageBiomedical DataBenchmarkOrdinary Differential Equation

🎯 What it does: Proposed a flow matching model called BioFlow that maintains the non-negativity of gene expression during the generation process, specifically designed for predicting spatial transcriptomics data from tissue slice images.

BioGuide: Biomedically-Guided Segmentation for Medical Tasks

Alsahanova, Nadezhda (Applied AI Institute), Sharaev, Maxim (Applied AI Institute)

CodeSegmentationData-Centric LearningConvolutional Neural NetworkTransformerContrastive LearningImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose the BioGuide framework, which adds a lightweight MLP as an auxiliary task at the bottleneck of the encoder in medical image segmentation models, leveraging structured clinical descriptions (such as lesion location and radiological features) to guide feature learning.

Boosting Ultrasound Image Classification via Attribute-Guided Dual-Branch Framework

Zhao, Bo (Wuhan University), Du, Bo (Wuhan University)

CodeClassificationExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataUltrasound

🎯 What it does: Proposed a pluggable attribute-guided dual-branch framework to enhance the accuracy and interpretability of ultrasound image classification.

Born Different: A Multi-level Individualized Modeling Framework for Embryo Euploidy Prediction

Tan, Shuangyi (Chinese University of Hong Kong, Shenzhen), Li, Guanbin (Sun Yat-sen University)

CodeClassificationAnomaly DetectionConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningImageTabularTime SeriesBiomedical DataElectronic Health Records

🎯 What it does: Propose a multi-level personalized modeling framework for non-invasive embryo homology prediction

Boundary-Aware Multi-Granularity Learning for Depression Severity Estimation

Li, Zhihong (Yunnan University), Yang, Yun (Yunnan University)

CodeClassificationExplainability and InterpretabilityRepresentation LearningTransformerSupervised Fine-TuningContrastive LearningTextMultimodalityAudio

🎯 What it does: Propose a boundary-aware multi-grained learning framework that combines continuous regression with coarse-grained ordinal supervision to estimate depression severity.

Brain-Adapter: A Dual-Stream Vision-Language MIL Framework for Comprehensive 3D CT Diagnosis of Acute Intracranial Pathologies

Yi, Zhenyu (Shanghai Jiao Tong University), Zhang, Lichi (Shanghai Jiao Tong University)

CodeClassificationAnomaly DetectionTransformerSupervised Fine-TuningVision Language ModelContrastive LearningImageTextBiomedical DataComputed Tomography

🎯 What it does: Propose the Brain-Adapter dual-stream MIL framework, utilizing 2D vision-language models and diagnostic reports to achieve multi-label diagnosis of 3D CT pathology.

Brain-DiT: A Universal Multi-state fMRI Foundation Model with Metadata-Conditioned Pretraining

Xia, Junfeng (Southern University of Science and Technology), Liu, Quanying (Southern University of Science and Technology)

CodeGenerationData SynthesisRepresentation LearningTransformerDiffusion modelContrastive LearningBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Developed a diffusion Transformer base model called Brain-DiT, which is pre-trained using multi-state fMRI data (rest, task, natural stimulation, disease, sleep, etc.), and incorporates individual metadata for conditional learning during the pre-training phase;

BRAIN: Bi-directional Motion Reasoning with Dynamic Memory Pruning for Surgical Video Segmentation

Xu, Chuanzhen (Nanjing University), Shi, Yinghuan (Nanjing University)

CodeSegmentationTransformerContrastive LearningOptical FlowVideoBiomedical Data

🎯 What it does: Proposed the BRAIN framework to address segmentation errors caused by sudden instrument displacement and long-term disappearance in surgical videos.

BrainACU: Asynchronous Coordination Unit Modeling for Dynamic Brain Network Analysis

Guo, Guiliang (Northeastern University), Zaiane, Osmar R. (University of Alberta)

CodeClassificationAnomaly DetectionGraph Neural NetworkTransformerAuto EncoderContrastive LearningTime SeriesBiomedical DataMagnetic Resonance Imaging

🎯 What it does: A dynamic brain network asynchronous coordination unit framework called BrainACU for psychiatric disease diagnosis using rs-fMRI was studied.

BrainAnytime: Anatomy-Aware Cross-Modal Pretraining for Brain Image Analysis with Arbitrary Modality Availability

Yang, Guangqian (Hong Kong Polytechnic University), the Alzheimer’s Disease Neuroimaging Initiative

CodeClassificationSegmentationKnowledge DistillationRepresentation LearningTransformerAuto EncoderContrastive LearningMultimodalityBiomedical DataMagnetic Resonance ImagingPositron Emission TomographyAlzheimer's Disease

🎯 What it does: Developed a unified pre-training framework called BrainAnytime, which can perform brain image analysis under any available modality subset (multi-sequence MRI and amyloid-PET), supporting seamless switching from single-modality to multi-modality.

BrainHTF: Learning Causal Graph Representation of Brain Connectome with Hyperbolic Transformer

Sun, Qiyu (Nanjing University of Information Science and Technology), Wang, Mingliang (Nanjing University of Information Science and Technology)

CodeClassificationAnomaly DetectionRepresentation LearningGraph Neural NetworkTransformerAuto EncoderContrastive LearningGraphBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease

🎯 What it does: Propose the BrainHTF framework, which first splits the brain connectome into causal subgraphs and bias subgraphs using a mask generator, then learns their discrete representations on the Poincaré ball via hyperbolic Transformer, and applies causal interventions on the bias representations to discover transferable causal patterns for brain disease diagnosis.

BrainSTR: Spatio-Temporal Contrastive Learning for Interpretable Dynamic Brain Network Modeling

Guo, Guiliang, Zaiane, Osmar R. (Northeastern University)

CodeClassificationExplainability and InterpretabilityGraph Neural NetworkAuto EncoderContrastive LearningGraphBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease

🎯 What it does: Designed and implemented a dynamic brain network modeling framework called BrainSTR based on spatiotemporal contrastive learning, for the diagnosis and interpretability analysis of neuropsychiatric disorders.

BrainWeaver: Weaving Dynamic Brain Networks with ROI-Guided Attention for Cognitive Assessment

Liu, Tao (Hangzhou City University), Zhou, Binbin (Hangzhou City University)

CodeExplainability and InterpretabilityComputational EfficiencyGraph Neural NetworkTransformerContrastive LearningGraphBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Proposed the BrainWeaver framework, which utilizes ROI-guided dynamic multi-graph attention fusion, hierarchical graph representation, and individual-invariant contrastive learning to achieve interpretable prediction of cognitive functions.

BREIT: A Framework for Brain Stroke Reconstruction using Multi-frequency 3D EIT

Abdelmoumene, Djahid (CY Cergy Paris University), Daveau, Christian (CY Cergy Paris University)

CodeRestorationConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageMultimodalityBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Proposed the BREIT framework to generate frequency-dependent conductivity voxels from CT/MRI, providing a Python 3D CEM forward solver and 3D D-bar implementation, and developed the dFNO-bar learning-based reconstruction method based on this.

Bridging Heterogeneous Medical Datasets via Mixture-of-Specialists Adapters for Unified Medical Image Classification

Huang, Shixing (University of Sydney)

CodeClassificationTransformerSupervised Fine-TuningMixture of ExpertsContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound

🎯 What it does: Propose a parameter-efficient framework called MOSAIC, which freezes the ViT backbone and utilizes a modality-aware tokenizer and a hybrid expert adapter to unify 18 different modalities of medical imaging data (2D/3D) into a single model, addressing the problem of cross-modal feature entanglement.

Bridging Research and Practice: A Systematic Evaluation of Generalist and Dermatology-Specific Models in Clinical Skin Lesion Classification

dos Santos, Emanoel (Universidade Federal de Pernambuco), Ing Ren, Tsang (Universidade Federal de Pernambuco)

CodeClassificationAnomaly DetectionConvolutional Neural NetworkTransformerSupervised Fine-TuningPrompt EngineeringVision Language ModelContrastive LearningImageMultimodalityBiomedical DataElectronic Health RecordsBenchmark

🎯 What it does: Systematically evaluated the performance of various general-purpose and dermatology-specific vision and vision-language models in binary classification of malignant risk prediction across different skin disease image datasets.

Bridging Single Distortion Artifacts and Multifactorial Clinical Quality: Few-Shot Biparametric MRI Quality Assessment via Distortion-Trained Prototypical Networks

Tang, Yucheng (University College London), Hu, Yipeng (University College London)

CodeClassificationDomain AdaptationAnomaly DetectionMeta LearningConvolutional Neural NetworkContrastive LearningBiomedical DataMagnetic Resonance Imaging

🎯 What it does: A few-shot dual-parameter proton imaging quality assessment network is proposed, which can automatically identify DWI geometric distortions and transfer to clinical PI-QUAL scoring under conditions of limited annotation.

Bridging Vision Foundation Model Priors with CLIP for Spatial-aware Few-shot Anomaly Detection in Medical Images

Miao, Juzheng (Chinese University of Hong Kong), Heng, Pheng-Ann (Chinese University of Hong Kong)

CodeAnomaly DetectionTransformerVision Language ModelContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose a Spatial-FAD framework that combines the spatial prior of Vision Foundation Model (DINO) with CLIP for improving the spatial localization accuracy of anomaly detection in medical imaging with few samples.

CΒ²RM-Seg: Causal Counterfactual Reasoning with Structural-Semantic Priors for Weakly Supervised Histopathological Tissue Segmentation

Zhang, Hualong (Guilin University of Electronic Technology), Pan, Xipeng (Guilin University of Electronic Technology)

CodeSegmentationExplainability and InterpretabilityComputational EfficiencyConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningBiomedical DataMagnetic Resonance ImagingComputed TomographyReview/Survey Paper

🎯 What it does: Propose a weakly supervised organ segmentation framework, CRM-Seg, based on causal counterfactual reasoning and structural semantic priors. First, generate deconfounded pseudo labels through a causal module, and then perform segmentation using a dual-channel structural-semantic network, optimized by an uncertainty-gated margin loss.

CAG-WM: A Synthetic-Data-Driven Coronary World Model for Autonomous Guidewire Navigation

Cao, Yue (Tianjin University), Yang, Jiachen (Tianjin University)

CodeAutonomous DrivingRecurrent Neural NetworkTransformerReinforcement LearningDiffusion modelWorld ModelOptical FlowImagePoint CloudMeshBiomedical DataMagnetic Resonance ImagingComputed TomographyReview/Survey PaperStochastic Differential Equation

🎯 What it does: Propose CAG-WM, a coronary artery world model based on synthetic data, for autonomous guidewire navigation.

CalcSeg: Confidence-aware 3D Latent Context Curriculum Learning For Myocardial Scar Segmentation From Single-Stack LGE-CMRs

Jayakumar, Nivetha (University of Virginia), Zhang, Miaomiao (University of Virginia)

CodeSegmentationConvolutional Neural NetworkTransformerContrastive LearningBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose a method called CalcSeg for myocardial scar segmentation on single-stack LGE-CMR images.

CALHippo: Cell Segmentation for Neuronal Density Inference in the Human Hippocampus

Casari, Giovanni (University of Modena and Reggio Emilia), Grana, Costantino (University of Modena and Reggio Emilia)

CodeSegmentationGenerationData SynthesisConvolutional Neural NetworkMixture of ExpertsAuto EncoderContrastive LearningImagePoint CloudBiomedical DataMagnetic Resonance ImagingBenchmark

🎯 What it does: Built CALHippo, a multi-scale human hippocampal CA region cell type segmentation and density inference resource, covering all subregions from CA1 to CA4.

Calibrated Confidence Expression for Radiology Report Generation

Bani-Harouni, David (Technical University of Munich), Keicher, Matthias (Technical University of Munich)

CodeExplainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringVision Language ModelImageTextMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyElectronic Health Records

🎯 What it does: Propose ConRad, a reinforcement learning-based framework that enables medical large vision-language models to generate radiology reports along with calibrated self-confidence expressions;

CALM: Interpretable Cross-Modal Alignment for Biomarker Discovery from Unpaired Data

Wang, Jueqi (Boston University), Venkataraman, Archana (Boston University)

CodeClassificationExplainability and InterpretabilityDrug DiscoveryTransformerMixture of ExpertsAuto EncoderGenerative Adversarial NetworkContrastive LearningImageTabularBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease

🎯 What it does: Designed and implemented the CALM framework, which maps structural MRI and genomic data from completely non-overlapping populations into a shared latent space via linear projection, enabling the discovery of interpretable brain region-gene pathway associations without paired data, and using these associations for diagnostic prediction of autism spectrum disorder (ASD).

Can Agents Distinguish Visually Hard-to-Separate Diseases in a Zero-Shot Setting? A Pilot Study

Zhao, Zihao (University Hospital Aachen), Truhn, Daniel (University Hospital Aachen)

CodeClassificationAnomaly DetectionExplainability and InterpretabilityTransformerLarge Language ModelAgentic AIPrompt EngineeringVision Language ModelImageTextBiomedical DataElectronic Health RecordsBenchmarkChain-of-Thought

🎯 What it does: A zero-shot diagnostic benchmark for multi-modal large language models (MLLMs) is conducted on visually indistinguishable diseases (melanoma vs. atypical nevi, pulmonary edema vs. pneumonia), and a multi-agent contrastive reasoning framework called CARE is proposed.

Cancer-Type-Agnostic Pan-Cancer Gene Expression Prediction from Histopathological Images

Shen, Yiyang (Tsinghua University), Li, Xiu (Tsinghua University)

CodeImage TranslationRepresentation LearningData-Centric LearningDrug DiscoveryGraph Neural NetworkTransformerMixture of ExpertsAuto EncoderContrastive LearningImageBiomedical Data

🎯 What it does: This paper proposes a pan-cancer image gene expression prediction framework called PIGP, which can directly infer gene expression levels from H&E tissue sections without prior knowledge of cancer types.

CardioDiT: Latent Diffusion Transformers for 4D Cardiac MRI Synthesis

Seyfarth, Marvin (Heidelberg University), Engelhardt, Sandy (Heidelberg University)

CodeGenerationData SynthesisTransformerVision-Language-Action ModelDiffusion modelAuto EncoderBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Proposed CardioDiT, a 4D latent diffusion transformer for generating short-axis cardiac MRI (cine CMR) data from scratch, where the model simultaneously models spatial and temporal dimensions in the latent space;

CARE: Clinically Aligned Retrieval Evidence for Consistent Radiology Report Generation

Zhu, Jintao (Zhejiang Cas Angels Biotechnology Co Ltd), Ma, Jiaxin (Zhejiang Cas Angels Biotechnology Co Ltd)

CodeGenerationRetrievalDomain AdaptationTransformerLarge Language ModelContrastive LearningImageTextBiomedical DataComputed TomographyElectronic Health RecordsRetrieval-Augmented Generation

🎯 What it does: This paper proposes CARE, a radiology report generation framework based on retrieval augmentation, which maintains a shared retrieval memory queue based on a momentum encoder to achieve temporal consistency in retrieval supervision, and introduces a disease alignment loss to ensure consistency between the retrieved text and image diagnosis, thus improving the clinical accuracy and diversity of the generated reports.

CARE: Confidence-Aware Reasoning for Reliable Medical VQA

Du, Yuetian (Zhejiang University), Zhu, Qiang (Zhejiang University)

CodeExplainability and InterpretabilityComputational EfficiencyData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision Language ModelImageTextMultimodalityBiomedical DataElectronic Health RecordsBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose the CARE framework, which first automatically synthesizes structured Medical-CoT data, and then performs two-stage fine-tuning of the medical vision-and-language model using confidence-aware reinforcement learning (GRPO+CAR), to improve diagnostic accuracy and confidence calibration.

CARformer: Class-Aware Representation Learning for Small-Cohort, Imbalanced Psychiatric Disorder Classification in Brain MRI

Fu, Xingyue, Kim, Jinman (University of Sydney)

CodeClassificationRepresentation LearningTransformerSupervised Fine-TuningContrastive LearningImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: A CARformer framework is proposed and implemented for multi-class classification from structural MRI in the context of few-shot, class-imbalanced psychiatric disease diagnosis.

Case-Specific Priors-Guided Multimodal Fusion for Prognosis Prediction in Head and Neck Cancer

Lin, Yue (Zhongguancun Academy), Meng, Mingyuan (Zhongguancun Academy)

CodeClassificationAnomaly DetectionExplainability and InterpretabilityComputational EfficiencyRepresentation LearningDrug DiscoveryConvolutional Neural NetworkTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsContrastive LearningImageTextMultimodalityTabularBiomedical DataElectronic Health RecordsBenchmarkRetrieval-Augmented Generation

🎯 What it does: Utilize large language models to generate case-specific priors to guide multi-modal fusion for predicting 5-year overall survival and 2-year recurrence-free survival in patients with head and neck cancer.

CDFP-Net: Cross-Modal Dynamic Fusion with Diffusion Priors for PET-CT Tumor Segmentation

Wei, Minqin (Xinjiang University), Deng, Lei (Xinjiang University)

CodeSegmentationGraph Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningMultimodalityBiomedical DataComputed TomographyPositron Emission Tomography

🎯 What it does: This paper proposes CDFP-Net, a dual-stream framework for tumor segmentation in PET-CT images.

CDPM-Align: Multi-Scale Guidance-Aligned Diffusion Pretraining for Robust Few-Shot Anatomical Landmark Detection

Di Via, Roberto (University of Genoa), Pastore, Vito Paolo (University of Genoa)

CodePose EstimationRepresentation LearningConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelContrastive LearningImageBiomedical DataComputed Tomography

🎯 What it does: This paper proposes a pre-training framework called CDPM-Align, based on conditional diffusion models and multi-scale guided alignment, for achieving reliable anatomical landmark detection under conditions of very limited annotation.

Cell-Division-Aware Category-Enhanced Contrastive Learning Model for Embryonic Cleavage Stage Classification

Li, Yukun (Shenzhen University), Pei, Jihong

CodeClassificationConvolutional Neural NetworkTransformerContrastive LearningOptical FlowTime SeriesBiomedical Data

🎯 What it does: This paper proposes a classification framework for the embryo cleavage stage called 3CL-ECSC, based on cell division awareness and category-enhanced contrastive learning, for automatically identifying different developmental stages of embryos;

CerviThink: A Reinforced Visual Reasoning Framework for Cervical Cancer Cell Classification

Fei, Manman, Zhang, Lichi (Shanghai Jiao Tong University)

CodeClassificationReinforcement Learning from Human FeedbackTransformerReinforcement LearningVision Language ModelImageBiomedical DataRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This paper proposes a vision reasoning framework based on reinforcement learning called CerviThink, aimed at improving the classification accuracy of cervical cancer cell images.

CheXanatomy: Anatomy–Aware Vision–Language Modeling for Chest Radiographs

Gatidis, Sergios (Stanford University), Bluethgen, Christian (Stanford University)

CodeSegmentationData SynthesisTransformerSupervised Fine-TuningVision Language ModelDiffusion modelAuto EncoderImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: The authors introduce self-regressive token space dissection supervision into a pre-trained vision-language model (VLM), enabling the model to directly generate anatomical segmentation masks for chest X-rays without requiring an additional pixel-level decoder.

CHILD: Human-in-the-Loop OOD Detection for Safe Clinical Deployment

Ye, Jinlun (Sun Yat-sen University), Wang, Ruixuan (Sun Yat-sen University)

CodeAnomaly DetectionSafty and PrivacyContrastive LearningImageBiomedical DataElectronic Health RecordsRetrieval-Augmented Generation

🎯 What it does: Proposes a training-agnostic, sparse physician feedback-based clinical OOD detection framework called CHILD, achieving risk-aware sample selection and retrieval-based score calibration in streaming deployment environments.

ChronoSurv: A Clinical Pathway-Guided Graph Framework for Multimodal Survival Analysis

Miccinilli, Hugo (UniversitΓ© Paris-Saclay), Di Piazza, Theo (University of Lyon)

CodeGraph Neural NetworkTransformerMultimodalityGraphBiomedical DataElectronic Health RecordsBenchmark

🎯 What it does: Developed ChronoSurv, a heterogeneous hierarchical directed graph framework based on clinical pathways, for multimodal survival prediction in head and neck cancer.

CIGTSurv: Clinical Information Guided Tri-Modal Survival Prediction with Local Prototype Association and Global Feature Alignment

Dai, Jing (Dalian University of Technology), Xu, Hongming (Dalian University of Technology)

CodeClassificationRepresentation LearningDrug DiscoveryConvolutional Neural NetworkRecurrent Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelContrastive LearningImageTextTabularBiomedical DataElectronic Health Records

🎯 What it does: Propose the CIGTSurv framework, which uses clinical information guided fusion of three modalities (pathological images, gene expression, clinical tables) to predict the survival period of cancer patients.

CIPHER: Causal Intervention Pathways for Healthcare Equity and Robustness

Jia, Xinyu (Fudan University), Wang, Yuanyuan (Fudan University)

CodeGenerationData SynthesisFederated LearningSafty and PrivacyExplainability and InterpretabilityTransformerPrompt EngineeringDiffusion modelScore-based ModelImageBiomedical DataComputed TomographyElectronic Health Records

🎯 What it does: Proposed a diffusion generation framework based on causal intervention called CIPHER, aimed at eliminating performance differences among sensitive subgroups in medical image diagnosis

ClinCoT: Clinical-Aware Visual Chain-of-Thought for Medical Vision Language Models

Liu, Xiwei, Xie, Yutong (Mohamed bin Zayed University of Artificial Intelligence)

CodeOptimizationExplainability and InterpretabilityRepresentation LearningDrug DiscoveryTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision Language ModelImageTextBiomedical DataElectronic Health RecordsChain-of-Thought

🎯 What it does: Propose the ClinCoT framework, extending preference optimization from only improving at the final response level to the clinical visual chain-of-thought level; achieving visual-driven alignment between intermediate reasoning and final answers through automatically generating region proposals based on disease hypotheses, generating intermediate reasoning chains for each region, constructing preference pairs with multi-model consensus weighted scoring, and using margin-aware DPO along with iterative learning.

Clinically-Grounded Hierarchical Classification for Consistent Chest X-ray Interpretation

Moon, Jong Hak (Yeji X), Kim, Minjun (Yeji X)

CodeClassificationImage TranslationAnomaly DetectionTransformerPrompt EngineeringAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound

🎯 What it does: Proposes a clinically oriented hierarchical multi-label classification framework called CHASE, which achieves hierarchical diagnosis from coarse to fine based on chest X-rays.