arXivSub Start free trial

MICCAI 2026 Papers — Page 2

International Conference on Medical Image Computing and Computer-Assisted Intervention · 1165 papers

Beyond Stochastic Diffusion: Trajectory-Consistent Deterministic Flow Matching for Fluorescence Molecular Tomography

Xue, Qianqian (Shanxi University), Wang, Wenjian (Shanxi University)

Drug DiscoveryTransformerDiffusion modelFlow-based ModelBiomedical DataComputed TomographyOrdinary Differential Equation

🎯 What it does: Propose a reconstruction framework for fluorescent molecular tomography (FMT) named TCDFlow-FMT based on deterministic flow matching, achieving fast and high-quality three-dimensional fluorescent source reconstruction.

Beyond the Batch: Momentum-Updated Virtual Cohorts for 3D PET-CT Prognosis

Liang, Xinglong (Netherlands Cancer Institute), Mann, Ritse (Netherlands Cancer Institute)

Anomaly DetectionOptimizationRepresentation LearningConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningBiomedical DataComputed TomographyPositron Emission Tomography

🎯 What it does: Propose the Momentum-Contrastive Survival Framework (MCSF), which expands the risk set in 3D PET-CT survival prediction through a virtual cohort and improves optimization stability using a time-weighted contrastive loss.

Beyond Visual Cues: CoT-Enhanced Reasoning for Semi-supervised Medical Image Segmentation

Chen, Yuming (Southeast University), Zhou, Yi (Nanjing University of Science and Technology)

SegmentationRetrievalExplainability and InterpretabilityConvolutional Neural NetworkTransformerLarge Language ModelVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingComputed TomographyRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This paper proposes a semi-supervised medical image segmentation framework called CERS based on Chain-of-Thought (CoT), which enhances segmentation performance by jointly retrieving generated diagnostic reasoning text and visual features.

Beyond Visual Forensics: Auditing Multimodal Robustness for Synthetic Medical Image Detection

Chiu, Ching-Hao (University of Notre Dame), Shi, Yiyu (University of Notre Dame)

Anomaly DetectionExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextBiomedical DataComputed TomographyElectronic Health RecordsBenchmark

🎯 What it does: Proposed an evaluation benchmark for detecting the authenticity of multimodal medical images, by fixing the image and only modifying the accompanying structured metadata to detect text-induced decision bias.

BiC-ODE: Bidirectionally Coupled Neural ODEs for Rapid Laminar Surface Reconstruction of the Human Cortex from 5T MRI

Cao, Shui (ShanghaiTech University), Shen, Dinggang (ShanghaiTech University)

SegmentationOptimizationComputational EfficiencyDiffusion modelScore-based ModelRectified FlowContrastive LearningImageBiomedical DataMagnetic Resonance ImagingStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Propose the BiC-ODE framework for fast and high-precision reconstruction of four surface layers (white matter, inner surface, outer surface, pial surface) in 5T FLAIR images.

Bidirectional Anatomy-Aware Post-Training for Longitudinal Chest X-Ray Progression Modeling

Chen, Yuming (Adelaide University), Xie, Yutong (Mohamed bin Zayed University of Artificial Intelligence)

ClassificationRepresentation LearningTransformerSupervised Fine-TuningContrastive LearningImageBiomedical DataComputed TomographyElectronic Health Records

🎯 What it does: Proposed a lightweight post-training strategy called BAAP to improve longitudinal chest radiograph progression prediction;

BIGUS: Beam Integrated Gaussian Scattering for Continuous Ultrasound View Synthesis

Abdelaziz, Youssif (Applied Innovation Center), Torki, Marwan (Applied Innovation Center)

Data SynthesisNeural Radiance FieldGaussian SplattingBiomedical DataUltrasound

🎯 What it does: Propose the BIGUS framework, which utilizes 3D Gaussian splats to achieve continuous ultrasound view synthesis and accurately simulate ultrasound scattering to preserve speckle details.

BiM-GeoAttn-Net: Linear-Time Depth Modeling with Geometry-Aware Attention for 3D Aortic Dissection CTA Segmentation

Zhang, Yuan (Sichuan Normal University), Mu, Nan (Sichuan Normal University)

SegmentationConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningBiomedical DataComputed Tomography

🎯 What it does: Propose BiM-GeoAttn-Net, combining bidirectional deep Mamba with geometry-aware vessel attention for aortic dissection segmentation in 3D CTA.

BioFact-MoE: Biologically Factorized Mixture of Experts for Vision–Language Prognostic Modeling in Hepatocellular Carcinoma

Yang, Junlin (Yale University), Chapiro, Julius (Yale University)

Explainability and InterpretabilityDrug DiscoveryTransformerLarge Language ModelMixture of ExpertsVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose the BioFact-MoE framework, utilizing a Mixture of Experts at the gene level for visual-language prognostic modeling in hepatocellular carcinoma.

BioFlow: A Biologically Valid Support-Preserving Flow for Histology-Conditioned Spatial Transcriptomics Prediction

Xu, Haoran (Sichuan University), Han, Xiao (Sichuan University)

GenerationData SynthesisTransformerDiffusion modelFlow-based ModelImageBiomedical DataBenchmarkOrdinary Differential Equation

🎯 What it does: Proposed a flow matching model called BioFlow that maintains the non-negativity of gene expression during the generation process, specifically designed for predicting spatial transcriptomics data from tissue slice images.

BioGait-VLM: A Tri-Modal Vision–Language–Biomechanics Framework for Interpretable Clinical Gait Assessment

Chen, Erdong (Drexel University), Liu, Feng (Washington University)

ClassificationPose EstimationExplainability and InterpretabilityKnowledge DistillationTransformerSupervised Fine-TuningVision Language ModelVideoTextBiomedical Data

🎯 What it does: Proposed the BioGait-VLM tri-modal vision-language-biomechanics framework for interpretable clinical gait assessment.

BioGuide: Biomedically-Guided Segmentation for Medical Tasks

Alsahanova, Nadezhda (Applied AI Institute), Sharaev, Maxim (Applied AI Institute)

SegmentationData-Centric LearningConvolutional Neural NetworkTransformerContrastive LearningImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose the BioGuide framework, which adds a lightweight MLP as an auxiliary task at the bottleneck of the encoder in medical image segmentation models, leveraging structured clinical descriptions (such as lesion location and radiological features) to guide feature learning.

BiSA-DiT: A Bidirectional Spatially-Aware Diffusion Transformer for Sparse Tomographic Magnetic Particle Imaging

Zhang, Xun (Beihang University), Tian, Jie (Beihang University)

RestorationTransformerDiffusion modelBiomedical Data

🎯 What it does: This paper proposes a BiSA-DiT framework based on a bidirectional spatial-aware diffusion transformer for reconstructing three-dimensional volumes from sparse system matrices in magnetic particle imaging.

Boosting Ultrasound Image Classification via Attribute-Guided Dual-Branch Framework

Zhao, Bo (Wuhan University), Du, Bo (Wuhan University)

ClassificationExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataUltrasound

🎯 What it does: Proposed a pluggable attribute-guided dual-branch framework to enhance the accuracy and interpretability of ultrasound image classification.

Born Different: A Multi-level Individualized Modeling Framework for Embryo Euploidy Prediction

Tan, Shuangyi (Chinese University of Hong Kong, Shenzhen), Li, Guanbin (Sun Yat-sen University)

ClassificationAnomaly DetectionConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningImageTabularTime SeriesBiomedical DataElectronic Health Records

🎯 What it does: Propose a multi-level personalized modeling framework for non-invasive embryo homology prediction

Boundary-Aware Multi-Granularity Learning for Depression Severity Estimation

Li, Zhihong (Yunnan University), Yang, Yun (Yunnan University)

ClassificationExplainability and InterpretabilityRepresentation LearningTransformerSupervised Fine-TuningContrastive LearningTextMultimodalityAudio

🎯 What it does: Propose a boundary-aware multi-grained learning framework that combines continuous regression with coarse-grained ordinal supervision to estimate depression severity.

Brain-Adapter: A Dual-Stream Vision-Language MIL Framework for Comprehensive 3D CT Diagnosis of Acute Intracranial Pathologies

Yi, Zhenyu (Shanghai Jiao Tong University), Zhang, Lichi (Shanghai Jiao Tong University)

ClassificationAnomaly DetectionTransformerSupervised Fine-TuningVision Language ModelContrastive LearningImageTextBiomedical DataComputed Tomography

🎯 What it does: Propose the Brain-Adapter dual-stream MIL framework, utilizing 2D vision-language models and diagnostic reports to achieve multi-label diagnosis of 3D CT pathology.

Brain-DiT: A Universal Multi-state fMRI Foundation Model with Metadata-Conditioned Pretraining

Xia, Junfeng (Southern University of Science and Technology), Liu, Quanying (Southern University of Science and Technology)

GenerationData SynthesisRepresentation LearningTransformerDiffusion modelContrastive LearningBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Developed a diffusion Transformer base model called Brain-DiT, which is pre-trained using multi-state fMRI data (rest, task, natural stimulation, disease, sleep, etc.), and incorporates individual metadata for conditional learning during the pre-training phase;

Brain-Specialized Self-supervised Learning: Neuroanatomy-Aware Spatial Priors for 3D Brain MRI

Kim, Kyeong Ho (Hanyang University), Lee, Jong-Min (Hanyang University)

ClassificationTransformerAuto EncoderContrastive LearningBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease

🎯 What it does: This study proposes a self-supervised learning framework for 3D brain MRI, integrating latent space mask alignment, global contrastive learning, and spatial distance-based patch decorrelation regularization;

Brain-TM: Decoding Brain States from Functional Networks Using a Hierarchical Spatiotemporal Transformer-Mamba

Liu, Xuan (Northwestern Polytechnical University), Zhang, Shu (Northwestern Polytechnical University)

ClassificationExplainability and InterpretabilityGraph Neural NetworkTransformerContrastive LearningGraphTime SeriesBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose the Brain-TM framework, which decodes brain states using functional networks and provides an interpretable central hub mechanism.

BRAIN: Bi-directional Motion Reasoning with Dynamic Memory Pruning for Surgical Video Segmentation

Xu, Chuanzhen (Nanjing University), Shi, Yinghuan (Nanjing University)

SegmentationTransformerContrastive LearningOptical FlowVideoBiomedical Data

🎯 What it does: Proposed the BRAIN framework to address segmentation errors caused by sudden instrument displacement and long-term disappearance in surgical videos.

BrainACU: Asynchronous Coordination Unit Modeling for Dynamic Brain Network Analysis

Guo, Guiliang (Northeastern University), Zaiane, Osmar R. (University of Alberta)

ClassificationAnomaly DetectionGraph Neural NetworkTransformerAuto EncoderContrastive LearningTime SeriesBiomedical DataMagnetic Resonance Imaging

🎯 What it does: A dynamic brain network asynchronous coordination unit framework called BrainACU for psychiatric disease diagnosis using rs-fMRI was studied.

BrainAnytime: Anatomy-Aware Cross-Modal Pretraining for Brain Image Analysis with Arbitrary Modality Availability

Yang, Guangqian (Hong Kong Polytechnic University), the Alzheimer’s Disease Neuroimaging Initiative

ClassificationSegmentationKnowledge DistillationRepresentation LearningTransformerAuto EncoderContrastive LearningMultimodalityBiomedical DataMagnetic Resonance ImagingPositron Emission TomographyAlzheimer's Disease

🎯 What it does: Developed a unified pre-training framework called BrainAnytime, which can perform brain image analysis under any available modality subset (multi-sequence MRI and amyloid-PET), supporting seamless switching from single-modality to multi-modality.

BrainHTF: Learning Causal Graph Representation of Brain Connectome with Hyperbolic Transformer

Sun, Qiyu (Nanjing University of Information Science and Technology), Wang, Mingliang (Nanjing University of Information Science and Technology)

ClassificationAnomaly DetectionRepresentation LearningGraph Neural NetworkTransformerAuto EncoderContrastive LearningGraphBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease

🎯 What it does: Propose the BrainHTF framework, which first splits the brain connectome into causal subgraphs and bias subgraphs using a mask generator, then learns their discrete representations on the Poincaré ball via hyperbolic Transformer, and applies causal interventions on the bias representations to discover transferable causal patterns for brain disease diagnosis.

BrainKODE: Controlled Koopman Dynamics over Longitudinal Structure–Function Coupling for Early Prediction of Adolescent Substance Use

Mazumder, Badhan (Georgia State University), Ye, Dong Hye (Georgia State University)

ClassificationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerDiffusion modelScore-based ModelContrastive LearningGraphTabularTime SeriesBiomedical DataMagnetic Resonance ImagingDiffusion Tensor ImagingStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: This study proposes BrainKODE, a geometry-consistent deep framework that models longitudinal structure-function coupling as a controlled latent dynamical system to predict the risk of adolescent substance use initiation (SUI).

BrainSTR: Spatio-Temporal Contrastive Learning for Interpretable Dynamic Brain Network Modeling

Guo, Guiliang, Zaiane, Osmar R. (Northeastern University)

ClassificationExplainability and InterpretabilityGraph Neural NetworkAuto EncoderContrastive LearningGraphBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease

🎯 What it does: Designed and implemented a dynamic brain network modeling framework called BrainSTR based on spatiotemporal contrastive learning, for the diagnosis and interpretability analysis of neuropsychiatric disorders.

BrainWeaver: Weaving Dynamic Brain Networks with ROI-Guided Attention for Cognitive Assessment

Liu, Tao (Hangzhou City University), Zhou, Binbin (Hangzhou City University)

Explainability and InterpretabilityComputational EfficiencyGraph Neural NetworkTransformerContrastive LearningGraphBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Proposed the BrainWeaver framework, which utilizes ROI-guided dynamic multi-graph attention fusion, hierarchical graph representation, and individual-invariant contrastive learning to achieve interpretable prediction of cognitive functions.

Breaking Annotation Dependency in Coronary Stenosis Segmentation via Physics-Driven Sim2Real Learning

Zhang, Baochang (Technical University of Munich), Navab, Nassir (HELIOS Hospital)

SegmentationData SynthesisDomain AdaptationConvolutional Neural NetworkDiffusion modelScore-based ModelGenerative Adversarial NetworkContrastive LearningSimultaneous Localization and MappingOptical FlowImageBiomedical DataComputed TomographyPositron Emission TomographyPhysics RelatedStochastic Differential Equation

🎯 What it does: Propose a physics-driven sim-to-real framework, using controllable synthetic coronary artery stenosis models to train U-Net, and achieving coronary artery stenosis segmentation without lesion annotations through uncertainty-guided self-supervised refinement;

Breaking Teacher’s Mean Regression Bias: Boundary-Aware Contrastive Distillation for Endoscopic Super-Resolution

Zuo, Wenbin (Tianjin University), Wan, Liang (Tianjin University)

Super ResolutionKnowledge DistillationContrastive LearningImageBiomedical Data

🎯 What it does: This paper addresses the mean regression bias in endoscopic super-resolution by proposing the Boundary-Aware Contrastive Distillation (BACD) framework, which reshapes the distillation space to eliminate the smoothing bias of the teacher model by constructing extreme degradation samples as negative anchors and utilizing gradient-optimized inputs as positive anchors.

BREIT: A Framework for Brain Stroke Reconstruction using Multi-frequency 3D EIT

Abdelmoumene, Djahid (CY Cergy Paris University), Daveau, Christian (CY Cergy Paris University)

RestorationConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageMultimodalityBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Proposed the BREIT framework to generate frequency-dependent conductivity voxels from CT/MRI, providing a Python 3D CEM forward solver and 3D D-bar implementation, and developed the dFNO-bar learning-based reconstruction method based on this.

Bridging EEG to fMRI: Semi-Supervised Functional Connectivity Translation via Latent Alignment

Li, Xiaoye (University of Science and Technology of China), Shen, Dinggang (University of Science and Technology of China)

RetrievalConvolutional Neural NetworkTransformerContrastive LearningImageTextMultimodality

🎯 What it does: Proposed a deep learning-based image retrieval method

Bridging Heterogeneous Medical Datasets via Mixture-of-Specialists Adapters for Unified Medical Image Classification

Huang, Shixing (University of Sydney)

ClassificationTransformerSupervised Fine-TuningMixture of ExpertsContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound

🎯 What it does: Propose a parameter-efficient framework called MOSAIC, which freezes the ViT backbone and utilizes a modality-aware tokenizer and a hybrid expert adapter to unify 18 different modalities of medical imaging data (2D/3D) into a single model, addressing the problem of cross-modal feature entanglement.

Bridging Research and Practice: A Systematic Evaluation of Generalist and Dermatology-Specific Models in Clinical Skin Lesion Classification

dos Santos, Emanoel (Universidade Federal de Pernambuco), Ing Ren, Tsang (Universidade Federal de Pernambuco)

ClassificationAnomaly DetectionConvolutional Neural NetworkTransformerSupervised Fine-TuningPrompt EngineeringVision Language ModelContrastive LearningImageMultimodalityBiomedical DataElectronic Health RecordsBenchmark

🎯 What it does: Systematically evaluated the performance of various general-purpose and dermatology-specific vision and vision-language models in binary classification of malignant risk prediction across different skin disease image datasets.

Bridging Single Distortion Artifacts and Multifactorial Clinical Quality: Few-Shot Biparametric MRI Quality Assessment via Distortion-Trained Prototypical Networks

Tang, Yucheng (University College London), Hu, Yipeng (University College London)

ClassificationDomain AdaptationAnomaly DetectionMeta LearningConvolutional Neural NetworkContrastive LearningBiomedical DataMagnetic Resonance Imaging

🎯 What it does: A few-shot dual-parameter proton imaging quality assessment network is proposed, which can automatically identify DWI geometric distortions and transfer to clinical PI-QUAL scoring under conditions of limited annotation.

Bridging the Markerless Tracking Validation Gap: Intraoperative X-Ray Workflow Integration

Daly, Connor (Hamlyn Centre for Robotic Surgery Imperial College London), Rodriguez y Baena, Ferdinando (Hamlyn Centre for Robotic Surgery Imperial College London)

Pose EstimationSupervised Fine-TuningOptical FlowImagePoint CloudBiomedical DataComputed Tomography

🎯 What it does: A markerless tracking validation pipeline suitable for clinical surgery was developed, utilizing conventional C-arm fluoroscopy and RGB-D cameras to obtain the 6-degree-of-freedom pose of the vertebra, achieving non-invasive, radiation-free, and standard workflow-compatible ground truth annotation for the lower spine.

Bridging Vision Foundation Model Priors with CLIP for Spatial-aware Few-shot Anomaly Detection in Medical Images

Miao, Juzheng (Chinese University of Hong Kong), Heng, Pheng-Ann (Chinese University of Hong Kong)

Anomaly DetectionTransformerVision Language ModelContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose a Spatial-FAD framework that combines the spatial prior of Vision Foundation Model (DINO) with CLIP for improving the spatial localization accuracy of anomaly detection in medical imaging with few samples.

C²RM-Seg: Causal Counterfactual Reasoning with Structural-Semantic Priors for Weakly Supervised Histopathological Tissue Segmentation

Zhang, Hualong (Guilin University of Electronic Technology), Pan, Xipeng (Guilin University of Electronic Technology)

SegmentationExplainability and InterpretabilityComputational EfficiencyConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningBiomedical DataMagnetic Resonance ImagingComputed TomographyReview/Survey Paper

🎯 What it does: Propose a weakly supervised organ segmentation framework, CRM-Seg, based on causal counterfactual reasoning and structural semantic priors. First, generate deconfounded pseudo labels through a causal module, and then perform segmentation using a dual-channel structural-semantic network, optimized by an uncertainty-gated margin loss.

CAG-WM: A Synthetic-Data-Driven Coronary World Model for Autonomous Guidewire Navigation

Cao, Yue (Tianjin University), Yang, Jiachen (Tianjin University)

Autonomous DrivingRecurrent Neural NetworkTransformerReinforcement LearningDiffusion modelWorld ModelOptical FlowImagePoint CloudMeshBiomedical DataMagnetic Resonance ImagingComputed TomographyReview/Survey PaperStochastic Differential Equation

🎯 What it does: Propose CAG-WM, a coronary artery world model based on synthetic data, for autonomous guidewire navigation.

CalCErt: Bin-wise Certification of Confidence Calibration in Medical Image Classification

Fillioux, Leo (Université Paris-Saclay), Dolz, Jose (ETS Montréal)

ClassificationAnomaly DetectionExplainability and InterpretabilityConvolutional Neural NetworkTransformerContrastive LearningImageBiomedical DataBenchmark

🎯 What it does: Propose a post-processing method called CalCErt, which provides verifiable confidence calibration evaluation for any pre-trained differentiable classifier under a given ℓ₂ error radius.

CalcSeg: Confidence-aware 3D Latent Context Curriculum Learning For Myocardial Scar Segmentation From Single-Stack LGE-CMRs

Jayakumar, Nivetha (University of Virginia), Zhang, Miaomiao (University of Virginia)

SegmentationConvolutional Neural NetworkTransformerContrastive LearningBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose a method called CalcSeg for myocardial scar segmentation on single-stack LGE-CMR images.

CALHippo: Cell Segmentation for Neuronal Density Inference in the Human Hippocampus

Casari, Giovanni (University of Modena and Reggio Emilia), Grana, Costantino (University of Modena and Reggio Emilia)

SegmentationGenerationData SynthesisConvolutional Neural NetworkMixture of ExpertsAuto EncoderContrastive LearningImagePoint CloudBiomedical DataMagnetic Resonance ImagingBenchmark

🎯 What it does: Built CALHippo, a multi-scale human hippocampal CA region cell type segmentation and density inference resource, covering all subregions from CA1 to CA4.

Calibrated Confidence Expression for Radiology Report Generation

Bani-Harouni, David (Technical University of Munich), Keicher, Matthias (Technical University of Munich)

Explainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringVision Language ModelImageTextMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyElectronic Health Records

🎯 What it does: Propose ConRad, a reinforcement learning-based framework that enables medical large vision-language models to generate radiology reports along with calibrated self-confidence expressions;

CalibSleep: Cross-Modal Calibration and Rule-Aware Learning for Automatic Sleep Stage Classification

Chen, Xiang (Shanghai Jiao Tong University), Wang, Zheyuan (Shanghai Jiao Tong University)

ClassificationConvolutional Neural NetworkRecurrent Neural NetworkTransformerAuto EncoderContrastive LearningTime SeriesBiomedical DataElectrocardiogram

🎯 What it does: Propose a dual-stream network called CalibSleep, which jointly learns from time-domain EEG/EOG signals and their time-frequency representations, and achieves automatic sleep stage classification through cross-modal calibration and rule-aware classification heads.

CALM: Interpretable Cross-Modal Alignment for Biomarker Discovery from Unpaired Data

Wang, Jueqi (Boston University), Venkataraman, Archana (Boston University)

ClassificationExplainability and InterpretabilityDrug DiscoveryTransformerMixture of ExpertsAuto EncoderGenerative Adversarial NetworkContrastive LearningImageTabularBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease

🎯 What it does: Designed and implemented the CALM framework, which maps structural MRI and genomic data from completely non-overlapping populations into a shared latent space via linear projection, enabling the discovery of interpretable brain region-gene pathway associations without paired data, and using these associations for diagnostic prediction of autism spectrum disorder (ASD).

Can Agents Distinguish Visually Hard-to-Separate Diseases in a Zero-Shot Setting? A Pilot Study

Zhao, Zihao (University Hospital Aachen), Truhn, Daniel (University Hospital Aachen)

ClassificationAnomaly DetectionExplainability and InterpretabilityTransformerLarge Language ModelAgentic AIPrompt EngineeringVision Language ModelImageTextBiomedical DataElectronic Health RecordsBenchmarkChain-of-Thought

🎯 What it does: A zero-shot diagnostic benchmark for multi-modal large language models (MLLMs) is conducted on visually indistinguishable diseases (melanoma vs. atypical nevi, pulmonary edema vs. pneumonia), and a multi-agent contrastive reasoning framework called CARE is proposed.

Can Experts Adapt Without Training? On Test-Time Modality Generalization in MVLMs

Imam, Raza (Mohamed bin Zayed University of Artificial Intelligence), Yaqub, Mohammad (Mohamed bin Zayed University of Artificial Intelligence)

Domain AdaptationOptimizationComputational EfficiencyRepresentation LearningTransformerPrompt EngineeringMixture of ExpertsVision Language ModelContrastive LearningImageMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound

🎯 What it does: Proposed MoBE, a training-free optimization-based Mixture-of-Experts (MoE) framework, which enables test-time adaptive reasoning when medical vision-language models (MVLMs) encounter unknown modalities.

Cancer-Type-Agnostic Pan-Cancer Gene Expression Prediction from Histopathological Images

Shen, Yiyang (Tsinghua University), Li, Xiu (Tsinghua University)

Image TranslationRepresentation LearningData-Centric LearningDrug DiscoveryGraph Neural NetworkTransformerMixture of ExpertsAuto EncoderContrastive LearningImageBiomedical Data

🎯 What it does: This paper proposes a pan-cancer image gene expression prediction framework called PIGP, which can directly infer gene expression levels from H&E tissue sections without prior knowledge of cancer types.

CancerVerse: A Fully Open Longitudinal and Multimodal Dataset for Multicancer Screening

Li, Wenxuan (Johns Hopkins University), Zhou, Zongwei (Johns Hopkins Medicine)

SegmentationAnomaly DetectionConvolutional Neural NetworkTransformerLarge Language ModelContrastive LearningImageTextMultimodalityTabularBiomedical DataComputed TomographyElectronic Health Records

🎯 What it does: Proposed CancerVerse, an open dataset containing multi-cancer types, multi-timepoints, and cross-modalities (CT, radiology reports, pathology results, clinical variables) with voxel-level tumor annotations, and provides a validation cohort of healthy follow-ups.

Cardiac Motion Transfer with Adaptive Physics Constraint

Yu, Chengjin (Anhui University), Liu, Huafeng (Zhejiang University)

Image TranslationGenerationData SynthesisConvolutional Neural NetworkTransformerDiffusion modelGenerative Adversarial NetworkContrastive LearningOptical FlowImageVideoBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Transfer motion from cardiac cine-MRI to static LGE images to generate dynamic LGE-MRI for dynamic lesion assessment.

CardiacPULSE: Frequency-Aware Self-supervised Cardiac Phase Detection in Echocardiography

Guo, Jiarong (Hong Kong University of Science and Technology), Li, Xiaomeng (Hong Kong University of Science and Technology)

RecognitionConvolutional Neural NetworkContrastive LearningVideoBiomedical DataUltrasound

🎯 What it does: Proposed a lightweight, unsupervised cardiac phase detection framework called CardiacPULSE, which identifies ED/ES frames in echocardiography using frequency domain self-supervised learning.

CardioDiT: Latent Diffusion Transformers for 4D Cardiac MRI Synthesis

Seyfarth, Marvin (Heidelberg University), Engelhardt, Sandy (Heidelberg University)

GenerationData SynthesisTransformerVision-Language-Action ModelDiffusion modelAuto EncoderBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Proposed CardioDiT, a 4D latent diffusion transformer for generating short-axis cardiac MRI (cine CMR) data from scratch, where the model simultaneously models spatial and temporal dimensions in the latent space;

CardioGS: Deformation-Driven 4D Gaussian Splatting with Self-supervised Cardiac Clock Learning for Rotational DSA Reconstruction

Liu, Yilun (Northeastern University), Xu, Yan (Xi'an Jiaotong University)

Image TranslationRestorationGenerationData SynthesisConvolutional Neural NetworkTransformerSupervised Fine-TuningDiffusion modelContrastive LearningGaussian SplattingOptical FlowImageVideoBiomedical DataComputed TomographyElectronic Health RecordsElectrocardiogram

🎯 What it does: Propose CardioGS, which utilizes dual time coordinates (self-supervised cardiac clock and acquisition time) to separate the modeling of vascular geometric deformation and contrast agent permeation, and achieves 3D vascular reconstruction from sparse-view rotational DSA through 4D Gaussian Splatting.

CARE: Clinically Aligned Retrieval Evidence for Consistent Radiology Report Generation

Zhu, Jintao (Zhejiang Cas Angels Biotechnology Co Ltd), Ma, Jiaxin (Zhejiang Cas Angels Biotechnology Co Ltd)

GenerationRetrievalDomain AdaptationTransformerLarge Language ModelContrastive LearningImageTextBiomedical DataComputed TomographyElectronic Health RecordsRetrieval-Augmented Generation

🎯 What it does: This paper proposes CARE, a radiology report generation framework based on retrieval augmentation, which maintains a shared retrieval memory queue based on a momentum encoder to achieve temporal consistency in retrieval supervision, and introduces a disease alignment loss to ensure consistency between the retrieved text and image diagnosis, thus improving the clinical accuracy and diversity of the generated reports.

CARE: Confidence-Aware Reasoning for Reliable Medical VQA

Du, Yuetian (Zhejiang University), Zhu, Qiang (Zhejiang University)

Explainability and InterpretabilityComputational EfficiencyData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision Language ModelImageTextMultimodalityBiomedical DataElectronic Health RecordsBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose the CARE framework, which first automatically synthesizes structured Medical-CoT data, and then performs two-stage fine-tuning of the medical vision-and-language model using confidence-aware reinforcement learning (GRPO+CAR), to improve diagnostic accuracy and confidence calibration.

CARformer: Class-Aware Representation Learning for Small-Cohort, Imbalanced Psychiatric Disorder Classification in Brain MRI

Fu, Xingyue, Kim, Jinman (University of Sydney)

ClassificationRepresentation LearningTransformerSupervised Fine-TuningContrastive LearningImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: A CARformer framework is proposed and implemented for multi-class classification from structural MRI in the context of few-shot, class-imbalanced psychiatric disease diagnosis.

Case-Specific Priors-Guided Multimodal Fusion for Prognosis Prediction in Head and Neck Cancer

Lin, Yue (Zhongguancun Academy), Meng, Mingyuan (Zhongguancun Academy)

ClassificationAnomaly DetectionExplainability and InterpretabilityComputational EfficiencyRepresentation LearningDrug DiscoveryConvolutional Neural NetworkTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsContrastive LearningImageTextMultimodalityTabularBiomedical DataElectronic Health RecordsBenchmarkRetrieval-Augmented Generation

🎯 What it does: Utilize large language models to generate case-specific priors to guide multi-modal fusion for predicting 5-year overall survival and 2-year recurrence-free survival in patients with head and neck cancer.

CDFP-Net: Cross-Modal Dynamic Fusion with Diffusion Priors for PET-CT Tumor Segmentation

Wei, Minqin (Xinjiang University), Deng, Lei (Xinjiang University)

SegmentationGraph Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningMultimodalityBiomedical DataComputed TomographyPositron Emission Tomography

🎯 What it does: This paper proposes CDFP-Net, a dual-stream framework for tumor segmentation in PET-CT images.

CDPM-Align: Multi-Scale Guidance-Aligned Diffusion Pretraining for Robust Few-Shot Anatomical Landmark Detection

Di Via, Roberto (University of Genoa), Pastore, Vito Paolo (University of Genoa)

Pose EstimationRepresentation LearningConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelContrastive LearningImageBiomedical DataComputed Tomography

🎯 What it does: This paper proposes a pre-training framework called CDPM-Align, based on conditional diffusion models and multi-scale guided alignment, for achieving reliable anatomical landmark detection under conditions of very limited annotation.

Cell-Division-Aware Category-Enhanced Contrastive Learning Model for Embryonic Cleavage Stage Classification

Li, Yukun (Shenzhen University), Pei, Jihong

ClassificationConvolutional Neural NetworkTransformerContrastive LearningOptical FlowTime SeriesBiomedical Data

🎯 What it does: This paper proposes a classification framework for the embryo cleavage stage called 3CL-ECSC, based on cell division awareness and category-enhanced contrastive learning, for automatically identifying different developmental stages of embryos;

Cervical Vertebral Maturation Staging from a Continuous Perspective

Mahajan, Kush (Indian Institute of Technology Ropar), Gupta, Sukrit (Post Graduate Institute of Medical Education and Research)

RecognitionSegmentationConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: This study proposes the ORACLE framework, which utilizes continuous learning to automatically predict the maturation stage of the cervical spine, addressing the errors caused by traditional discrete grading and the clinically acceptable ±1 stage tolerance.

CerviThink: A Reinforced Visual Reasoning Framework for Cervical Cancer Cell Classification

Fei, Manman, Zhang, Lichi (Shanghai Jiao Tong University)

ClassificationReinforcement Learning from Human FeedbackTransformerReinforcement LearningVision Language ModelImageBiomedical DataRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This paper proposes a vision reasoning framework based on reinforcement learning called CerviThink, aimed at improving the classification accuracy of cervical cancer cell images.

CheXanatomy: Anatomy–Aware Vision–Language Modeling for Chest Radiographs

Gatidis, Sergios (Stanford University), Bluethgen, Christian (Stanford University)

SegmentationData SynthesisTransformerSupervised Fine-TuningVision Language ModelDiffusion modelAuto EncoderImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: The authors introduce self-regressive token space dissection supervision into a pre-trained vision-language model (VLM), enabling the model to directly generate anatomical segmentation masks for chest X-rays without requiring an additional pixel-level decoder.

CHILD: Human-in-the-Loop OOD Detection for Safe Clinical Deployment

Ye, Jinlun (Sun Yat-sen University), Wang, Ruixuan (Sun Yat-sen University)

Anomaly DetectionSafty and PrivacyContrastive LearningImageBiomedical DataElectronic Health RecordsRetrieval-Augmented Generation

🎯 What it does: Proposes a training-agnostic, sparse physician feedback-based clinical OOD detection framework called CHILD, achieving risk-aware sample selection and retrieval-based score calibration in streaming deployment environments.

ChronoSurv: A Clinical Pathway-Guided Graph Framework for Multimodal Survival Analysis

Miccinilli, Hugo (Université Paris-Saclay), Di Piazza, Theo (University of Lyon)

Graph Neural NetworkTransformerMultimodalityGraphBiomedical DataElectronic Health RecordsBenchmark

🎯 What it does: Developed ChronoSurv, a heterogeneous hierarchical directed graph framework based on clinical pathways, for multimodal survival prediction in head and neck cancer.

CIGTSurv: Clinical Information Guided Tri-Modal Survival Prediction with Local Prototype Association and Global Feature Alignment

Dai, Jing (Dalian University of Technology), Xu, Hongming (Dalian University of Technology)

ClassificationRepresentation LearningDrug DiscoveryConvolutional Neural NetworkRecurrent Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelContrastive LearningImageTextTabularBiomedical DataElectronic Health Records

🎯 What it does: Propose the CIGTSurv framework, which uses clinical information guided fusion of three modalities (pathological images, gene expression, clinical tables) to predict the survival period of cancer patients.

CiLA: Knowledge-Preserving Self-supervised Adaptation of CLIP for Fundus Imaging

Huang, Yijin (Southern University of Science and Technology), Tang, Xiaoying (University of British Columbia)

ClassificationDomain AdaptationTransformerPrompt EngineeringVision Language ModelAuto EncoderContrastive LearningImageBiomedical Data

🎯 What it does: This paper proposes the CiLA framework, which in the field of retinal fundus images combines the knowledge of the pre-trained CLIP model with clinical lesion attribute pseudo-labels to perform self-supervised domain adaptation;

CineMesh4D: Personalized 4D Whole Heart Reconstruction from Sparse Cine MRI

Liu, Xiaoyue (National University of Singapore), Li, Lei (National University of Singapore)

RestorationSegmentationGenerationConvolutional Neural NetworkGraph Neural NetworkTransformerSupervised Fine-TuningDiffusion modelAuto EncoderContrastive LearningImageVideoMeshBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose an end-to-end CineMesh4D approach that directly reconstructs 4D (3D+t) full-heart shape meshes from sparse multi-view cardiac cine MRI.

CIPHER: Causal Intervention Pathways for Healthcare Equity and Robustness

Jia, Xinyu (Fudan University), Wang, Yuanyuan (Fudan University)

GenerationData SynthesisFederated LearningSafty and PrivacyExplainability and InterpretabilityTransformerPrompt EngineeringDiffusion modelScore-based ModelImageBiomedical DataComputed TomographyElectronic Health Records

🎯 What it does: Proposed a diffusion generation framework based on causal intervention called CIPHER, aimed at eliminating performance differences among sensitive subgroups in medical image diagnosis

ClinCoT: Clinical-Aware Visual Chain-of-Thought for Medical Vision Language Models

Liu, Xiwei, Xie, Yutong (Mohamed bin Zayed University of Artificial Intelligence)

OptimizationExplainability and InterpretabilityRepresentation LearningDrug DiscoveryTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision Language ModelImageTextBiomedical DataElectronic Health RecordsChain-of-Thought

🎯 What it does: Propose the ClinCoT framework, extending preference optimization from only improving at the final response level to the clinical visual chain-of-thought level; achieving visual-driven alignment between intermediate reasoning and final answers through automatically generating region proposals based on disease hypotheses, generating intermediate reasoning chains for each region, constructing preference pairs with multi-model consensus weighted scoring, and using margin-aware DPO along with iterative learning.

Clinical-Injection Transformer with Domain-Adapted MAE for Lupus Nephritis Prognosis Prediction

Huang, Yuewen (Sun Yat-sen University), Lu, Yutong (Sun Yat-sen University)

ClassificationDomain AdaptationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningImageMultimodalityTabularBiomedical DataElectronic Health Records

🎯 What it does: This study proposes a multimodal computational pathology framework that uses PAS-stained renal biopsy sections and structured clinical data to predict the three-class treatment response in children with systemic lupus erythematosus nephritis (LN) (complete remission, partial remission, no response)

Clinically Structured Chain-of-Thought Reasoning for Intelligent Guidance in Coronary Artery Disease

Ma, Xinghua, Gao, Xin (King Abdullah University of Science and Technology)

Autonomous DrivingOptimizationFederated LearningSafty and PrivacyExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningAdversarial AttackHyperparameter SearchData-Centric LearningRobotic IntelligenceMeta LearningDrug DiscoveryReinforcement Learning from Human FeedbackNeural Architecture SearchProtein Structure PredictionGraph Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataComputed TomographyElectronic Health RecordsElectrocardiogramRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: A visual-language model called CS-CoT based on clinical structured chain-of-thought was constructed for intelligent diagnosis and management of coronary artery disease (CAD).

Clinically-Grounded Hierarchical Classification for Consistent Chest X-ray Interpretation

Moon, Jong Hak (Yeji X), Kim, Minjun (Yeji X)

ClassificationImage TranslationAnomaly DetectionTransformerPrompt EngineeringAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound

🎯 What it does: Proposes a clinically oriented hierarchical multi-label classification framework called CHASE, which achieves hierarchical diagnosis from coarse to fine based on chest X-rays.

ClinLoop: Human-AI Collaborative Agents for Multimodal Breast Cancer Diagnosis Under Incomplete Clinical Data

Xie, Hui (Macao Polytechnic University), Tan, Tao (University College Dublin)

ClassificationExplainability and InterpretabilityData-Centric LearningTransformerAgentic AIPrompt EngineeringDiffusion modelContrastive LearningMultimodalityTabularBiomedical DataMagnetic Resonance ImagingComputed TomographyElectronic Health RecordsRetrieval-Augmented Generation

🎯 What it does: Propose ClinLoop, which combines diagnostic agents with quality assurance agents to achieve a human-machine collaborative workflow for breast cancer staging under missing multimodal breast imaging data.

ClinRAG-GRAPH: Clinical-Prior Retrieval-Augmented Graph Model with Domain Adversarial Learning for Breast pCR Prediction

Duan, Yaofei (Radboud University Medical Center), Mann, Ritse (Radboud University Medical Center)

ClassificationDomain AdaptationExplainability and InterpretabilityGraph Neural NetworkTransformerLarge Language ModelContrastive LearningImageTextGraphBiomedical DataMagnetic Resonance ImagingRetrieval-Augmented Generation

🎯 What it does: This study proposes a framework called ClinRAG-GRAPH, which constructs a clinical prior graph by utilizing dynamic enhanced magnetic resonance imaging (DCE-MRI), clinical variables, and biopsy pathological biomarkers, and achieves multi-modal feature fusion through graph convolutional networks. Furthermore, domain adversarial learning and large language model retrieval augmented generation (RAG) mechanisms are introduced to realize cross-center prediction of pathological complete response (pCR) during the pre-processing phase.

CMD-AD: Cross-Modality Distillation from fMRI-EEG for Single-Modality Alzheimer’s Disease Diagnosis

Lin, Xin (ShanghaiTech University), Shen, Dinggang (ShanghaiTech University)

ClassificationDomain AdaptationKnowledge DistillationRepresentation LearningGraph Neural NetworkTransformerAuto EncoderContrastive LearningMultimodalityBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease

🎯 What it does: This paper proposes a three-stage cross-modal distillation framework, CMD-AD, for learning joint spatiotemporal features from fMRI and EEG and transferring them to single-modal Alzheimer's disease diagnosis.

CMSA-Net: Causal Multi-scale Aggregation with Adaptive Multi-source Reference for Video Polyp Segmentation

Wang, Tong (Southeast University), Xie, Yutong (Mohamed bin Zayed University of Artificial Intelligence)

SegmentationConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningVideoBiomedical Data

🎯 What it does: Propose a video polyp segmentation framework named CMSA-Net, combining causal multi-scale aggregation with dynamic multi-source reference, significantly enhancing the semantic discrimination and cross-frame consistency of low-contrast polyps.

Co-MoCo: Collaborative Learning Enabled Motion Compensation for Weight-Bearing CBCT

Lin, Tong (Shanghai Jiao Tong University), Zhang, Yikun (University of Science and Technology of China)

RecognitionConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningAudio

🎯 What it does: This paper proposes an end-to-end speech recognition framework based on multi-head attention, and evaluates its performance in multilingual environments.

Coarse-to-Fine Meta-Reweighting with Dynamic Retrieval for Adult-to-Pediatric Domain Adaptation in Tumor Segmentation

Al-Fakih, Abdulkhalek (Yonsei University), Al-masni, Mohammed A. (King Fahd University of Petroleum & Minerals)

SegmentationDomain AdaptationContrastive LearningImageBiomedical DataMagnetic Resonance ImagingRetrieval-Augmented Generation

🎯 What it does: To address the imbalance and domain differences in adult and pediatric brain tumor image data, a two-stage coarse-to-fine meta-reweighting framework is proposed. It first performs cross-domain binary segmentation for localization, and then refines the segmentation of pediatric-specific tumor sub-regions using gradient-aligned meta-reweighting.

COAST: Context-Aware Differential Learning for Gene Expression Prediction in Spatial Transcriptomics

Byeon, Keunho (Korea University), Kwak, Jin Tae (Korea University)

Image TranslationRepresentation LearningData-Centric LearningTransformerSupervised Fine-TuningContrastive LearningImageBiomedical Data

🎯 What it does: Propose the COAST method, which jointly encodes local and global context features with target image information in the Transformer, and combines absolute expression regression with differential regression as a joint objective, to predict spatial gene expression from H&E histogram images.

CoGaze: Closed-Loop Eye-Face Alignment for Multi-center Mobile Gaze-Based Cognitive Screening in Older Adults

Yu, Jiahui (Zhejiang University), Xu, Xin (Zhejiang University)

RecognitionPose EstimationFederated LearningComputational EfficiencyRepresentation LearningConvolutional Neural NetworkRecurrent Neural NetworkTransformerVision Language ModelDiffusion modelAuto EncoderContrastive LearningVideoBiomedical DataAlzheimer's Disease

🎯 What it does: Propose the CoGaze closed-loop eye-face alignment framework for eye tracking and cognitive screening on mobile devices for the elderly.

CoGE: Sim-to-Real Online Geometric Estimation for Monocular Colonoscopy

Shao, Liangjing (Chinese University of Hong Kong), Ren, Hongliang (Chinese University of Hong Kong)

Depth EstimationDomain AdaptationTransformerDiffusion modelAuto EncoderContrastive LearningVideoBiomedical Data

🎯 What it does: Propose the CoGE framework, which achieves online monocular colonoscope geometric estimation (depth and 3D reconstruction) using only simulated data.

Collaborative Multiscale Representation Learning for AMD Detection

Jang, Nayeon (Sungkyunkwan University), Choo, Hyunseung (Sungkyunkwan University)

ClassificationAnomaly DetectionRepresentation LearningConvolutional Neural NetworkTransformerMixture of ExpertsContrastive LearningImageBiomedical DataAlzheimer's Disease

🎯 What it does: An automatic detection framework for age-related macular degeneration (AMD), named CoMRep, is proposed, which combines dual-branch feature extraction from global retinal images and local macular regions of interest (ROI), cross-scale attention aggregation, and Mixture-of-Experts (MoE) aggregation, to achieve multi-scale and cross-location lesion representation;

Collaborative Topology and Connectivity Learning for EM Neuron Segmentation

Shi, Haoyuan (University of Science and Technology of China), Xiong, Zhiwei (University of Science and Technology of China)

SegmentationComputational EfficiencyRepresentation LearningConvolutional Neural NetworkTransformerContrastive LearningBiomedical DataMagnetic Resonance ImagingComputed TomographyBenchmark

🎯 What it does: Proposed a two-stage collaborative topology and connectivity learning framework, leveraging EM-specific semantic priors such as neuronal skeletons and distance transforms, significantly improving the topological accuracy of EM neuron segmentation.

Colon-Bench: An Agentic Workflow for Scalable Dense Lesion Annotation in Full-Procedure Colonoscopy Videos

Hamdi, Abdullah (King Abdullah University of Science and Technology), Gao, Xin

Object DetectionSegmentationAnomaly DetectionTransformerLarge Language ModelAgentic AIPrompt EngineeringVision Language ModelVideoTextMultimodalityBiomedical DataBenchmarkRetrieval-Augmented Generation

🎯 What it does: This paper proposes Colon-Bench, a densely annotated dataset of endoscopic videos based on a multi-stage agent workflow, and evaluates the performance of multi-modal large language models on tasks such as detection of colon lesions, video segmentation, and question answering on this dataset.

Community-Aware Dynamic Graph Learning for Brain Disorder Diagnosis

Li, Qiuyan (Dongguan University of Technology), Wang, Kai (Dongguan University of Technology)

ClassificationAnomaly DetectionComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerAuto EncoderContrastive LearningGraphBiomedical DataMagnetic Resonance ImagingComputed TomographyAlzheimer's Disease

🎯 What it does: Proposed Community-aware Dynamic Graph Neural Network (CaDGNN) for brain disease diagnosis, combining delay-aware dynamic brain functional connectivity with community structure.

Comparative Analysis of Automated Frame Selection Methods for Glaucoma Detection in Smartphone-Based Low-Quality Fundus Video

Abay, Solomon Gebru (KU Leuven), Geurts, Lucca (KU Leuven)

ClassificationAnomaly DetectionConvolutional Neural NetworkContrastive LearningOptical FlowVideoBiomedical DataAlzheimer's DiseaseElectronic Health Records

🎯 What it does: Propose an automatic frame selection framework based on optic disc localization and multiple quality metrics for glaucoma screening using ultra-low-cost smartphone retinal videos.

Compass: Prostate Cancer Detection Needs Multi-view Context

Wilson, Paul F. R. (Queen's University), Mousavi, Parvin (University of British Columbia)

ClassificationAnomaly DetectionConvolutional Neural NetworkTransformerContrastive LearningImageVideoBiomedical DataUltrasound

🎯 What it does: Propose Compass, a multi-view AI framework that integrates microwave ultrasound rotation scanning and biopsy images for prostate cancer risk assessment;

Computer-Assisted Intervention in Capsule Endoscopy: A Real-Time Edge-AI Auditing System

Bravo, Diego (Universidad Nacional de Colombia), Romero, Eduardo (Universidad Nacional de Colombia)

ClassificationRecognitionAnomaly DetectionConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningVideoBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasound

🎯 What it does: Designed and implemented a real-time, edge-side capsule endoscopy auditing system that can instantly identify anatomical regions and detect abnormalities during the capsule imaging process.

Concept-Guided Noisy Negative Suppression for Zero-Shot Classification and Grounding of Chest X-Ray Findings

Lian, Chenyu (Hong Kong Polytechnic University), Qin, Jing (Hong Kong Polytechnic University)

ClassificationObject DetectionRepresentation LearningTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataComputed TomographyElectronic Health Records

🎯 What it does: Propose the CoNNS framework, achieving chest X-ray vision-language alignment through concept-oriented noise negative sample suppression, thereby enhancing zero-shot classification and localization performance.

ConceptIL: Concept Incremental Learning for Interpretable Medical AI

Zhou, Hangqi (Fudan University), Zhuang, Xiahai (Fudan University)

ClassificationExplainability and InterpretabilityTransformerSupervised Fine-TuningContrastive LearningImageBiomedical Data

🎯 What it does: Propose the Concept Incremental Learning (ConceptIL) framework, enabling the model to progressively learn new concepts in continuous tasks with only partial concept annotations, and enhance the interpretability of medical decision-making through learned concepts.

Conditional Diffusion Prompting for Ambiguous Medical Image Segmentation

Zhao, Hongkai (Beijing University of Posts and Telecommunications), Lao, Qicheng (Beijing University of Posts and Telecommunications)

SegmentationTransformerDiffusion modelAuto EncoderImageBiomedical DataMagnetic Resonance ImagingComputed TomographyBenchmark

🎯 What it does: Propose the Conditional Diffusion Prompting (CDP) framework, which performs diffusion prompting in the dense prompt embedding space and regulates prompt uncertainty through image latent variables, achieving diverse and image-consistent segmentation of ambiguous boundaries in medical images.

Conditional Latent Diffusion Model with Fourier-Based Motion Modelling for Virtual Population Synthesis

Lan, Shaokun (University of Manchester), Frangi, Alejandro F. (University of Manchester)

GenerationData SynthesisTransformerDiffusion modelScore-based ModelAuto EncoderMeshBiomedical DataMagnetic Resonance ImagingElectronic Health Records

🎯 What it does: This paper proposes a four-dimensional (3D+t) cardiac mesh synthesis framework based on conditional latent diffusion models, named 4D F-MeshLDM, which achieves the generation of periodic mesh sequences by mapping cardiac motion to the Fourier coefficient space.

Confidence-Aware Supervision for Robust Multi-task Embryo Grading

Park, Bogyu (Kai Health), Kim, Hyung Min (Kai Health)

ClassificationImage TranslationData SynthesisConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningImageBiomedical DataReview/Survey Paper

🎯 What it does: This paper proposes a confidence-aware soft label learning framework for multi-task embryo grading, which adaptively softens the labels of high-confidence samples during training based on feature consistency and entropy stability, thereby improving the model's reliability across different grading dimensions.

Conformal 3D Lesion Segmentation with Balanced Risk Control

Tan, Binyu (University of Electronic Science and Technology of China), Shi, Xiaoshuang (University of Electronic Science and Technology of China)

SegmentationExplainability and InterpretabilityConvolutional Neural NetworkTransformerScore-based ModelContrastive LearningBiomedical DataMagnetic Resonance ImagingComputed TomographyBenchmark

🎯 What it does: Propose the CLS framework, which performs post-processing threshold calibration for 3D lesion segmentation to achieve statistical guarantees on the voxel-level missed detection rate.

Confusion-Aware Evidence Calibration for Few-Shot Whole Slide Image Classification

Li, Peng (Nanjing University of Aeronautics and Astronautics), Qin, Jing (Hong Kong Polytechnic University)

ClassificationImage TranslationData-Centric LearningMeta LearningTransformerPrompt EngineeringVision Language ModelContrastive LearningImageBiomedical DataComputed TomographyReview/Survey PaperBenchmark

🎯 What it does: This paper proposes a confusion-aware evidence calibration framework (CAEC) for few-shot whole slide image classification, which suppresses aggregated alignment noise through dynamic confusion memory and morphological evidence compensation.

CONNECT-4: Brain Connectivity-Guided Hyperedge Graph Fusion for Structural MRI to 4D Rest Functional MRI Synthesis

Hassan, Salma (Mohamed Bin Zayed University of Artificial Intelligence), Yaqub, Mohammad (Mohamed Bin Zayed University of Artificial Intelligence)

GenerationData SynthesisGraph Neural NetworkTransformerVision Language ModelDiffusion modelContrastive LearningImageGraphBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease

🎯 What it does: This paper proposes the CONNECT-4 method, which generates a complete 4D rs-fMRI spatiotemporal sequence from a single T1-weighted structural MRI to bridge the structural-functional gap;

ConnecToMind2: Inter-Subject fMRI Decoding via Whole-Brain Connectome-Guided Alignment

Bae, Gunwoo (Gwangju Institute of Science and Technology), Kim, Mansu (Gwangju Institute of Science and Technology)

Image TranslationRestorationGenerationData SynthesisTransformerPrompt EngineeringVision Language ModelDiffusion modelAuto EncoderContrastive LearningImageMultimodalityBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Proposed a whole-brain connectivity-guided fMRI-to-image decoding framework called ConnecToMind2, which uses region-level embeddings and structural connectivity priors to achieve cross-subject image reconstruction.

ConstTrack: Constellation-Guided Cell Tracking Under Lineage Constraints

Xu, Yiwen (University of New South Wales), Meijering, Erik (University of New South Wales)

Object TrackingConvolutional Neural NetworkTransformerContrastive LearningOptical FlowImageVideoBiomedical Data

🎯 What it does: This paper proposes a cell tracking and lineage reconstruction framework called ConstTrack, which is based on constellation features and can real-time predict cell identities and accurately assign cell division events in time-series microscopy images.

ContiStain: Cross-Domain Relation-Preserving Distillation for Continual Multi-domain Virtual IHC Staining

Chen, Fuqiang (Harbin Institute of Technology), Zhang, Yongbing (Tsinghua University)

Image TranslationDomain AdaptationKnowledge DistillationRepresentation LearningTransformerMixture of ExpertsGenerative Adversarial NetworkContrastive LearningImageBiomedical Data

🎯 What it does: For the task of continuous multi-domain virtual IHC staining in medical images, this paper proposes a continual learning framework called ContiStain, which maintains performance on previously learned domains as new biomarker data is gradually received.

Contrast-Invariant Reference-Based Slice Thickness Estimation in MRI via Spectral Matching

Remedios, Samuel W. (Johns Hopkins University), Dewey, Blake E.

Super ResolutionDiffusion modelAuto EncoderContrastive LearningBiomedical DataMagnetic Resonance Imaging

🎯 What it does: This paper proposes a contrast-invariant slice thickness estimation method called PRISM based on spectral matching, which matches the full voxel spectra of high-resolution reference images with low-resolution scanned images, thereby estimating 3D MRI slice thickness without the need for registration or pixel correspondence.