arXivSub Start free trial

MICCAI 2026 Papers with Code β€” Page 4

International Conference on Medical Image Computing and Computer-Assisted Intervention Β· 649 papers

LangOR: 3D Language Field Reconstruction for Operating Room

Li, Wei (Chinese University of Hong Kong), Heng, Pheng-Ann (Chinese University of Hong Kong)

CodeRecognitionSegmentationTransformerContrastive LearningGaussian SplattingImageVideoTextPoint CloudBenchmark

🎯 What it does: Propose an unsupervised 3D language field reconstruction framework, LangOR, for semantic understanding in operating room scenes.

Large-Scale Distributed GPU-Accelerated Respiratory Motion-Resolved Reconstruction of 3D Non-cartesian mGRE MRI

Zhang, Chao (Stony Brook University), Kee, Youngwook (Stony Brook University)

CodeOptimizationComputational EfficiencyBiomedical DataMagnetic Resonance Imaging

🎯 What it does: A hierarchical multi-node multi-GPU framework is proposed for respiratory motion-resolved reconstruction in three-dimensional non-Cartesian multi-echo gradient echo (mGRE) magnetic resonance imaging.

Large-Small Model Collaboration for Zero-Shot Surgical Phase Recognition

Zhang, Yiyi (Chinese University of Hong Kong), Heng, Pheng-Ann (Chinese University of Hong Kong)

CodeRecognitionConvolutional Neural NetworkTransformerPrompt EngineeringVision Language ModelContrastive LearningVideoText

🎯 What it does: Designed a large model-small model collaborative framework LaST to achieve zero-shot out-of-distribution surgery phase recognition, improving pseudo-label quality through iterative time refinement and cyclic replay.

Latent-to-Latent Flow for Volumetric Stochastic Segmentation

Todd, Omar (Imperial College London), Glocker, Ben (Imperial College London)

CodeSegmentationConvolutional Neural NetworkDiffusion modelScore-based ModelFlow-based ModelAuto EncoderBiomedical DataComputed TomographyStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: To address the diversity uncertainty in medical volume segmentation, this paper proposes and implements a flow matching method in the latent space (L2L-Flow), which significantly improves inference speed while maintaining clinically relevant uncertainty estimation.

Learning Beyond Pixels: Structured Knowledge Topology Consistency for Semi-supervised 2D/3D Medical Image Segmentation

Chen, Yu (Jinan University), Wei, Honghao (Jinan University)

CodeSegmentationContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: This paper proposes a semi-supervised 2D/3D medical image segmentation framework called SKTC-Net based on structural knowledge topological consistency.

Learning Class Difficulty via Dynamic Focal Attention for Histopathology Segmentation

Rathukohe Mudiyanselage, Lakmali Nadeesha Kumari (University of Kentucky), Cheung, Sen-Ching Samson (University of Kentucky)

CodeSegmentationTransformerContrastive LearningBiomedical DataBenchmark

🎯 What it does: Proposed and implemented Dynamic Focal Attention (DFA) in pathological image semantic segmentation, directly encoding class difficulty by introducing learnable class-level bias into the transformer cross-attention mechanism;

Learning Directional Semantic Transitions for Longitudinal Chest X-Ray Analysis

Hu, Zhangfeng (Rensselaer Polytechnic Institute), Yan, Pingkun (Massachusetts General Hospital, Harvard Medical School)

CodeClassificationExplainability and InterpretabilityRepresentation LearningTransformerVision Language ModelContrastive LearningImageTextBiomedical DataComputed TomographyElectronic Health Records

🎯 What it does: Proposes a visual-language pre-training framework called ProTrans, which learns directional semantic transitions in chest X-rays over time to enable longitudinal CXR disease progression analysis.

Learning Diverse and Realistic Polyp Data via Style-Fused Diffusion Models

Han, Longfei (Beijing Technology and Business University), Li, Haisheng (Beijing Technology and Business University)

CodeSegmentationGenerationData SynthesisConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningImageBiomedical Data

🎯 What it does: Propose a conditional diffusion model based on style fusion (SFD-Polyp), which injects the style of real images through the Style Integration Module (SIM), and generates diverse and anatomically accurate mask shapes during the inference phase using Morphology-Preserving Mask Synthesis (MPMS), thus synthesizing diverse and realistic polyp image-mask pairs to enhance the dataset.

Learning from Good Neighbors: Slice-wise Quality-aware CT-to-Diffusion MRI Synthesis

Na, You-Kyoung (Chonnam National University), Cho, Yeong-Jun (Chonnam National University)

CodeImage TranslationData SynthesisConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose a hierarchical quality-aware CT-to-MRI (DWI) synthesis framework called LGN. It first generates a draft MRI using a latent diffusion model, then predicts slice quality through a teacher-student network, and finally refines locally using cross-attention from high-quality neighboring slices, improving synthesis quality while maintaining three-dimensional anatomical continuity.

Learning Geometry-Aware Bundles with Spectral Heat Kernel Diffusion for fMRI-based Brain Disease Diagnosis

Yin, Feiyu (Fudan University), Yu, Jinhua (Fudan University)

CodeClassificationAnomaly DetectionRepresentation LearningGraph Neural NetworkDiffusion modelContrastive LearningGraphBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease

🎯 What it does: Proposed a dual-stream contrastive learning framework VB-GCL based on the vector bundle theory for disease diagnosis in fMRI neural networks

Learning Hierarchical Anatomy-Grounded Representation for Cross-Modal Registration

Mi, Jia (ShanghaiTech University), Shen, Dinggang (ShanghaiTech University)

CodeOptimizationRepresentation LearningConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Studied the HAGR framework, a hierarchical anatomy-driven cross-modal deformable registration representation learning framework

Learning Invariant Anatomical Representations from Multi-Sequence MRI for Brain Segmentation

Liang, Jianwen, Tang, Xiaoying (Southern University of Science and Technology)

CodeSegmentationDomain AdaptationRepresentation LearningConvolutional Neural NetworkTransformerPrompt EngineeringVision Language ModelAuto EncoderContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingRetrieval-Augmented Generation

🎯 What it does: Propose the Text-guided Invariant Anatomical Learning (TeInL) framework, which leverages self-supervised pre-training on multi-sequence MRI to decouple anatomical structures from sequence styles, and enhances brain tissue and lesion segmentation performance through text-image alignment learning to capture pathology-related features.

Learning Prototypes for Unsupervised Joint Generation of Conditional Templates and Atlases

Zhang, Jichang (Beijing Normal University), Li, Shuyu (Beijing Normal University)

CodeGenerationData SynthesisDiffusion modelAuto EncoderContrastive LearningBiomedical DataMagnetic Resonance Imaging

🎯 What it does: This paper proposes an unsupervised framework called LatentAtlas, which jointly generates conditional brain templates and atlases. It aligns public templates and atlases through shared latent space learning of prototypes, enabling the synchronous generation of templates and atlases under specific conditions (e.g., age).

Learning Quantitative Calibration of Computed Tomography for Bone Mineral Density Estimation Using Multi-Center Data

SchΓΆller, Job H. J. (Nara Institute of Science and Technology), Otake, Yoshito (Osaka University)

CodeImage TranslationRestorationSegmentationGenerationData SynthesisConvolutional Neural NetworkDiffusion modelContrastive LearningImageBiomedical DataComputed Tomography

🎯 What it does: Achieve calibration-free bone density estimation by generating voxel-level QCT density maps directly from conventional CT using a supervised 3D image-to-image translation framework.

Learning Role-Conditioned Alignment for Medical Image Referring Segmentation

Li, Kun (University of Liverpool), Zheng, Yalin (Chinese Academy of Sciences)

CodeSegmentationConvolutional Neural NetworkTransformerVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: This paper proposes a role-conditioned alignment-based model for medical image referential segmentation (RCA), which achieves more accurate segmentation by explicitly separating textual descriptions into three categories: lesion semantics, quantity, and location, and then performing fine-grained alignment with visual features;

Learning the Hierarchical Organization in Brain Network for Brain Disorder Diagnosis

Tang, Jingfeng (Northeastern University), Zaiane, Osmar R. (University of Alberta)

CodeClassificationExplainability and InterpretabilityGraph Neural NetworkTransformerContrastive LearningGraphBiomedical DataAlzheimer's Disease

🎯 What it does: Propose a framework called BrainHO that learns the hierarchical organization of brain networks for the diagnosis of brain disorders (ASD, MDD);

Learning to Distort: Weakly-Supervised Image Quality Transfer for Prostate DWI Correction

Tang, Yucheng (University College London), Hu, Yipeng (University College London)

CodeImage TranslationRestorationData SynthesisConvolutional Neural NetworkTransformerDiffusion modelGenerative Adversarial NetworkContrastive LearningImageBiomedical DataMagnetic Resonance ImagingDiffusion Tensor Imaging

🎯 What it does: Proposed a weakly supervised image quality transfer (IQT) framework, which first generates realistic distorted images from undistorted prostate DWI using quality prototype flow matching, and then uses these synthetic pairs to train a supervised correction network to achieve distortion correction for single-echo EPI DWI.

Learning to Read Where to Look: Disease-Aware Vision–Language Pretraining for 3D CT

Ging, Simon (University of Freiburg), Brox, Thomas (University of Freiburg)

CodeClassificationSegmentationRetrievalTransformerPrompt EngineeringVision Language ModelContrastive LearningGaussian SplattingBiomedical DataComputed Tomography

🎯 What it does: Built a 3D CT audio-visual language model called RadFinder based on 98k CT scan-report pairs, incorporating disease prompt supervision and slice-level localization supervision.

Learning to Segment Burns with Limited Labels: A Dual-Path Cascaded Network with Uncertainty-Guided Teacher Diffusion Exposure

Chauhan, Joohi (Motilal Nehru National Institute of Technology Allahabad), Singh, Prabal Pratap (Motilal Nehru National Institute of Technology Allahabad)

CodeSegmentationData SynthesisDomain AdaptationConvolutional Neural NetworkDiffusion modelContrastive LearningImageBiomedical Data

🎯 What it does: Propose a semi-supervised burn segmentation framework called UADA-MSPCNet, combining a dual-path cascade network, uncertainty-based teacher-student consistency learning, and diffusion generation enhancement on the teacher side.

Learning to Unify Deformable Shape and Texture Representations for Cardiac Video Classification

Hossain, Tonmoy (University of Virginia), Zhang, Miaomiao (University of Virginia)

CodeClassificationExplainability and InterpretabilityRepresentation LearningConvolutional Neural NetworkTransformerContrastive LearningOptical FlowVideoBiomedical DataMagnetic Resonance Imaging

🎯 What it does: This study proposes ShapeFuse, aiming to unify deformable shape representations and image texture representations into a shared latent space, and dynamically fuse shape and texture through bidirectional cross-modal temporal attention and adaptive gating for disease classification of cardiac MRI videos.

Learning Where to Look: A Reinforcement Learning Framework for Robust Micro-ultrasound Prostate Cancer Detection

Abootorabi, Mohammad Mahdi (University of British Columbia), Abolmaesumi, Purang (University of British Columbia)

CodeClassificationSegmentationConvolutional Neural NetworkTransformerReinforcement LearningPrompt EngineeringAuto EncoderContrastive LearningBiomedical DataUltrasound

🎯 What it does: Propose the Prost-RL framework, which learns the attention regions of micro-ultrasound images through reinforcement learning. By combining a foundational model encoder-decoder, it achieves 'learn attention first, then decode,' thereby improving the detection and localization of prostate cancer under weakly supervised conditions.

Learning Where to Look: Pathologist-Inspired Multi-field-of-View Evidence Retrieval for Cell Type Classification in H&E

Yuan, Ruizhi (University of Pittsburgh), Chen, Wei (University of Pittsburgh)

CodeClassificationExplainability and InterpretabilityComputational EfficiencyTransformerMixture of ExpertsContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyRetrieval-Augmented Generation

🎯 What it does: Proposed a hybrid expert framework called TRACE, which classifies cell types in H&E slices by leveraging token-level evidence retrieval from multi-scale perspectives.

Leuko, I Am Your Prototype: A Prototypical Detection Framework for Fine-Grained Laryngeal Lesion Stratification

Federici, Lorenzo (UniversitΓ  Politecnica delle Marche), Moccia, Sara (Istituto Italiano di Tecnologia)

CodeClassificationObject DetectionAnomaly DetectionConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningImageBiomedical DataComputed TomographyUltrasound

🎯 What it does: Proposed a dual-branch real-time detection framework that integrates prototype learning into the YOLO network for fine-grained hierarchical classification of laryngeal lesions in NBI endoscopy.

Leveraging Pathology Co-occurrence for Test-Time Adaptation in Chest X-Ray Diagnosis

Jeong, Woojin (Seoul National University), Lee, Jaewook (Seoul National University)

CodeClassificationDomain AdaptationConvolutional Neural NetworkSupervised Fine-TuningContrastive LearningImageBiomedical DataComputed Tomography

🎯 What it does: This study investigates a method that utilizes disease co-occurrence patterns for test-time adaptation in multi-label classification tasks on chest X-rays, aiming to alleviate performance degradation caused by domain shifts between different hospitals.

LinGuinE: Longitudinal Guidance Estimation for Volumetric Tumour Segmentation

Garibli, Nadine (AstraZeneca), Patwari, Mayank (AstraZeneca)

CodeObject TrackingSegmentationConvolutional Neural NetworkTransformerPrompt EngineeringDiffusion modelOptical FlowImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Proposed the LinGuinE framework, which utilizes prompts from radiologists at a single time point to achieve longitudinal tumor voxel segmentation and tracking through image registration and prompt segmentation models.

LiSAD: A Label-Efficient Implicit Model with Spatially Adaptive Directional Total Variation for Spatial Transcriptomics

Xia, Yinghao (Northwestern Polytechnical University), Jin, Qiangguo (Northwestern Polytechnical University)

CodeSegmentationRepresentation LearningData-Centric LearningGraph Neural NetworkNeural Radiance FieldAuto EncoderGenerative Adversarial NetworkContrastive LearningBiomedical DataBenchmark

🎯 What it does: Propose the LiSAD model, which achieves continuous representation of spatial transcriptomics through implicit neural representation (INR) and hypergraph encoding, and performs spatial domain partitioning based on this

LLM-Bootstrapped Targeted Finding Guidance for Factual MLLM-based Medical Report Generation

Yang, Cunyuan (Zhejiang University), Wang, Haishuai (Zhejiang University)

CodeClassificationGenerationExplainability and InterpretabilityData-Centric LearningConvolutional Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelImageTextMultimodalityBiomedical DataComputed TomographyElectronic Health Records

🎯 What it does: Proposed the Fact-Flow framework, which first uses a multi-label classifier to identify clinical findings in images, and then uses these findings as conditions to guide the MLLM to generate medical reports, thereby significantly improving the factual accuracy of the reports.

LLM-Steered Clinical Knowledge Discovery for CVD Outcome Modeling

Liang, Yuxuan (Rensselaer Polytechnic Institute), Yan, Pingkun (Rensselaer Polytechnic Institute)

CodeExplainability and InterpretabilityDrug DiscoveryGraph Neural NetworkTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTabularBiomedical DataElectronic Health Records

🎯 What it does: This paper proposes the EGL-LLM framework, which is based on evidence-driven learning (EGL) and LLM guidance, to discover auditable cardiovascular disease (CVD) knowledge graphs from clinical and imaging variables.

Localization-Grounded Supervision: Revisiting Vanilla SFT of Large Vision-Language Models for Medical Image Analysis

Shi, Yiming (Tsinghua University), Wu, Ji (Tsinghua University)

CodeExplainability and InterpretabilityTransformerSupervised Fine-TuningVision Language ModelImageBiomedical DataComputed TomographyReview/Survey Paper

🎯 What it does: This paper systematically analyzes the supervised fine-tuning of large vision-language models in medical image analysis, and proposes a location-based supervision mechanism called LGS to accelerate the alignment of fine-grained semantics and spatial information.

LoFi: Location-Aware Fine-Grained Representation Learning for Chest X-Ray

Kang, Myeongkyun (University of British Columbia), Li, Xiaoxiao (University of British Columbia)

CodeRetrievalRepresentation LearningTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataComputed TomographyElectronic Health RecordsRetrieval-Augmented Generation

🎯 What it does: Propose the LoFi method, jointly optimizing sigmoid, captioning, and position-aware captioning losses, leveraging a lightweight LLM to learn fine-grained representations of chest X-rays, and using them for retrieval-based context learning to achieve precise localization.

Lost in the Folds: When Cross-Validation Is Not a Deep Ensemble for Uncertainty Estimation

Kirscher, Tristan (University of Strasbourg), Maier-Hein, Klaus (German Cancer Research Center)

CodeSegmentationConvolutional Neural NetworkImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: This paper investigates the impact of two ensemble methods, cross-validation (CV) and deep ensembles (DE), on uncertainty estimation in medical image segmentation.

Lost in Volume: The CT-SpatialVQA Benchmark for Evaluating Semantic-Spatial Understanding of 3D Medical Vision–Language Models

Monon, Mashrafi (Mohamed Bin Zayed University of Artificial Intelligence), Yaqub, Mohammad (Mohamed Bin Zayed University of Artificial Intelligence)

CodeExplainability and InterpretabilityTransformerLarge Language ModelVision Language ModelContrastive LearningImageTextBiomedical DataComputed TomographyBenchmark

🎯 What it does: Proposed the CT-SpatialVQA benchmark for systematically evaluating the semantic spatial reasoning capabilities of 3D medical vision-language models on volumetric CT images.

LOT-Bridge: A Latent Optimal Transport Bridge Based on Signed Distance Fields for Automatic Skull Defect Reconstruction

Yan, Yanghui, Shui, Wuyang (Shanghai Jiao Tong University)

CodeRestorationGenerationConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningPoint CloudMeshBiomedical DataMagnetic Resonance ImagingComputed TomographyOrdinary Differential Equation

🎯 What it does: Propose a Latent Optimal Transport Bridge (LOT-Bridge) method based on Signed Distance Field (SDF) for automatic cranial defect reconstruction, reducing the number of generation iterations and improving reconstruction quality.

Low-Rank Text-Guided Spectral Learning for Semi-supervised MHSI Segmentation

Zhang, Siqi (East China Normal University), Li, Qingli (East China Normal University)

CodeSegmentationData-Centric LearningConvolutional Neural NetworkTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposes a low-rank text-guided spectral learning network called LoTS-Net for segmenting microscopic hyperspectral images (MHSI) under semi-supervised conditions.

Low-Rank Velocity Fields as a Structural Prior for Unsupervised 4D Medical Image Interpolation

Li, Haojin, Liu, Jiang (Southern University of Science and Technology)

CodeRestorationDiffusion modelScore-based ModelAuto EncoderContrastive LearningOptical FlowBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: A low-rank velocity field structural prior is proposed under endpoint unsupervised conditions for 4D medical image interpolation.

Low-Rank-Modulated Functa: Exploring the Latent Space of Implicit Neural Representations for Interpretable Ultrasound Video Analysis

Wolleb, Julia (Yale University), Papademetris, Xenophon (Yale University)

CodeCompressionExplainability and InterpretabilityNeural Radiance FieldAuto EncoderOptical FlowVideoBiomedical DataUltrasound

🎯 What it does: Propose the Low-Rank Modulated Functa network (LRM-Functa), achieving compression and interpretability analysis of ultrasound videos.

LumiState: Retinex-inspired Illumination Decoupling in 4D Gaussian Splatting for Endoscopic Reconstruction

Sun, Dai (University of Science and Technology of China), Zhou, Shaohua Kevin

CodeRestorationConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelNeural Radiance FieldAuto EncoderContrastive LearningGaussian SplattingImageVideoPoint Cloud

🎯 What it does: Proposes a method combining illumination decomposition based on Retinex theory with 4D Gaussian scattering (LumiState), achieving separation of variable near-field illumination and tissue reflectance in endoscope videos, thereby improving the geometric and visual quality of dynamic 3D reconstruction.

M3D-QAdapter: 3D Medical VQA with Lesion-Level Finding-Segmentation Alignment and Query-Driven Adaptive Token Reduction

Liu, Hong (Xiamen University), Wang, Liansheng (Xiamen University)

CodeClassificationSegmentationComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataComputed Tomography

🎯 What it does: M3D-QAdapter proposes a two-stage 3D medical vision question answering framework, first training a 3D ViT with lesion-level alignment, and then using a query-driven adaptive process to compress visual tokens to generate answers.

Making HE Histopathological Images More Colorful by Conditional Flow Matching

Ben Omrane, Mohamed Salim (UniversitΓ© Paris-Saclay), Pesquet, Jean-Christophe (UniversitΓ© Paris-Saclay)

CodeImage TranslationRestorationGenerationData SynthesisConvolutional Neural NetworkDiffusion modelScore-based ModelFlow-based ModelImageBiomedical Data

🎯 What it does: This paper proposes a conditional flow matching model based on the concentration domain, generating missing eosin components from HE-stained images to obtain complete HES-stained images.

Mamba Based Anisotropic Diffusion Model for Speckle Noise Reduction in Ultrasound Images

Pham, Nhat Duy (UniversitΓ© Sorbonne Paris Nord), Trinh, Dinh Hoan (Viettel AI)

CodeRestorationTransformerDiffusion modelImageBiomedical DataUltrasound

🎯 What it does: Propose a Mamba-based hybrid model for ultrasound speckle noise reduction, combining the diffusion process of partial differential equations with a learnable diffusion function to form an end-to-end trainable iterative denoising framework.

MammoFlow: Multiview Mammogram Synthesis with Anatomically Consistent Flow Matching

Du, Yuexi (Yale University), Dvornek, Nicha C. (Yale University)

CodeGenerationData SynthesisTransformerDiffusion modelScore-based ModelFlow-based ModelRectified FlowAuto EncoderOptical FlowImageBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound

🎯 What it does: Proposes a multi-view breast X-ray synthesis framework called MammoFlow based on flow matching, which can generate consistent CC and MLO views without requiring real reference images.

MARBLE: Lightweight EEG-to-fMRI Translation via Mamba-Attention and ROI-Conditioned Decoding

Kim, Emil (Pusan National University), Gahm, Jin Kyu (Pusan National University)

CodeImage TranslationExplainability and InterpretabilityComputational EfficiencyTransformerAuto EncoderContrastive LearningTime SeriesBiomedical DataMagnetic Resonance ImagingElectrocardiogram

🎯 What it does: Propose MARBLE, a lightweight EEG-to-fMRI translation framework capable of predicting whole-brain BOLD time series from EEG;

Marginal Constrained Morphological Prototype Learning for Patch Search in Whole Slide Images

Park, Sihyeon (Korea University), Kim, Bumsoo (Chung Ang University)

CodeClassificationRetrievalComputational EfficiencyRepresentation LearningConvolutional Neural NetworkTransformerMixture of ExpertsContrastive LearningImageBiomedical DataBenchmark

🎯 What it does: This paper proposes a selective slice-level prediction framework called MMPL based on global morphological prototypes. It first uses the Sinkhorn-Knopp algorithm and momentum feature queue to learn global prototypes, and then uses these prototypes to retrieve the top-k most diagnostically valuable slices from each slide for aggregation, thereby achieving end-to-end training and significantly reducing computational costs.

Mask to Concept: Auto-Promptable SAM3 via Efficient Test-Time Concept Embedding Search for Few-Shot Annotation

Zhou, Quan (Wuhan University of Technology), Wang, Zhiwei (Huazhong University of Science and Technology)

CodeSegmentationData-Centric LearningTransformerPrompt EngineeringContrastive LearningImageBiomedical Data

🎯 What it does: Propose the Mask to Concept (M2C) framework, leveraging the concept prompting capability of SAM3, automatically searching for visual concepts from a small number of labeled samples on the frozen SAM3 architecture, achieving medical image few-shot automatic annotation and realizing efficient human-machine collaborative closed-loop;

Maximizing Domain Generalization in Automated Fetal Brain Biometry

Li, Yijin (Tsinghua University), Tian, Qiyuan (Tsinghua University)

CodeDomain AdaptationConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Proposes BioTTA, a source-agnostic, unsupervised test-time adaptation framework for automatic fetal brain biometry.

MBFAD: Benchmarking Balance Function Assessment with Environmental Perturbation and Synergistic Tasks

Ge, Zhaoyang (Zhengzhou University), Xu, Mingliang (Zhengzhou University)

CodeClassificationAnomaly DetectionConvolutional Neural NetworkSupervised Fine-TuningContrastive LearningMultimodalityTime SeriesBenchmark

🎯 What it does: This paper proposes the MBFAD multimodal balanced function assessment dataset and the M²-Balance framework, aiming to achieve real-time evaluation of imbalance states under static, dynamic, and reactive activities through synchronized inertial and plantar pressure data.

MDL-DA Track: Self-supervised Sperm Motility Direction Learning and Density-Adaptive Association for Sperm Tracking

Zeng, Xiaoyu, Yang, Xuan (Shenzhen University)

CodeObject TrackingConvolutional Neural NetworkTransformerContrastive LearningOptical FlowVideoBiomedical Data

🎯 What it does: This paper proposes a multi-object tracking framework that combines self-supervised motion direction learning and density-adaptive association, for precise tracking of sperm in high-density microscopic videos.

Measuring Prediction Uncertainty in Neural Cellular Automata

Sadafi, Ario (Helmholtz Munich), Marr, Carsten (Helmholtz Munich)

CodeSegmentationExplainability and InterpretabilityImageBiomedical DataBenchmark

🎯 What it does: Studies the uncertainty estimation of neural cellular automata (NCA) in medical image segmentation, proposing a training-agnostic perturbation recovery method called Resilience.

Measuring What VLMs Don’t Say: Validation Metrics Hide Clinical Terminology Erasure in Radiology Report Generation

Parikh, Aditya (Technical University of Denmark), Frank, Stella (Technical University of Denmark)

CodeGenerationExplainability and InterpretabilityData-Centric LearningLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataElectronic Health RecordsRetrieval-Augmented Generation

🎯 What it does: This paper proposes two new evaluation metricsβ€”Clinical Association Displacement (CAD) and Weighted Association Elimination (WAE)β€”to detect lexical disappearance and potential biases in visual-language models when generating chest X-ray reports; meanwhile, experiments reveal the impact of decoding strategies (deterministic vs. random sampling) on clinical information retention and fairness.

MEC: A Multi-expert Consultation Framework for Synergizing Frozen Foundation Models in Whole Slide Image Analysis

Chen, Zongyi (Xiamen University), Wang, Liansheng (Hong Kong University of Science and Technology)

CodeClassificationAnomaly DetectionComputational EfficiencyRepresentation LearningTransformerMixture of ExpertsImageBiomedical Data

🎯 What it does: Propose a MEC multi-expert negotiation framework that utilizes multiple frozen foundation models combined with MIL for diagnosis in WSI analysis.

MEC: Medical Evidence Capsules for Retrieval-Augmented Generation in Medical Multimodal Question Answering

Xu, Zhenghua (Hebei University Of Technology), Tian, Tian (Hebei University Of Technology)

CodeRetrievalExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningDrug DiscoveryTransformerLarge Language ModelPrompt EngineeringVision Language ModelTextMultimodalityGraphTabularBiomedical DataElectronic Health RecordsBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose Medical Evidence Capsules (MEC), which unify heterogeneous evidence such as medical text, structured knowledge, model parameters, and imaging reports into a structured capsule format for retrieval-augmented generation (RAG) in medical multimodal question answering.

Med-CAP: Counterfactual Evidence and Adaptive Prior Suppression for Robust Medical Visual Question Answering

Huang, Zaiqiang (Tsinghua University), Wu, Xian (Tencent Jarvis Lab)

CodeDomain AdaptationExplainability and InterpretabilityRepresentation LearningTransformerSupervised Fine-TuningVision Language ModelContrastive LearningImageTextBiomedical DataBenchmark

🎯 What it does: This paper proposes Med-CAP, a framework that achieves robustness in medical visual question answering through contrastive evidence modeling and adaptive prior suppression.

Med3D-R1: Mitigating Narrative Bias and Enhancing Reasoning Consistency in 3D Medical Vision-Language Models

Lai, Haoran (University of Science and Technology of China), Zhou, Shaohua Kevin (University of Science and Technology of China)

CodeAnomaly DetectionExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerReinforcement LearningPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataComputed Tomography

🎯 What it does: Built and trained Med3D-R1, a vision-language model based on CT 3D medical imaging, for abnormal diagnosis and to improve reasoning consistency.

MedCenterDet: A Center-Based Object Detection Framework for 3D Medical Image

Zeng, Qiang (Lingnan University), Pan, Fei (Shenzhen University)

CodeObject DetectionConvolutional Neural NetworkTransformerGenerative Adversarial NetworkContrastive LearningGaussian SplattingOptical FlowImageBiomedical DataMagnetic Resonance ImagingComputed TomographyStochastic Differential Equation

🎯 What it does: Proposed an anchor-free 3D medical image object detection framework called MedCenterDet, which directly locates the target center and infers the size through center point heatmap regression.

MedEnv: Scaling Multimodal Virtual Medical Environments for Long-Horizon Diagnosis

Fan, Zhiting (Zhejiang University), Liu, Zuozhu (Zhejiang University)

CodeExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningData-Centric LearningDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackTransformerReinforcement LearningPrompt EngineeringVision-Language-Action ModelTextMultimodalityTabularTime SeriesBiomedical DataElectronic Health RecordsRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Built the MedEnv multimodal long-term diagnostic simulation environment, supporting multi-round Q&A and tool-assisted evidence acquisition

MedFuse-Seg: Multi-level Visual and Semantic Context Fusion for Segmentation-Based Medical Reasoning

Limaroon, Keetawan, Achakulvisut, Titipat (Mahidol University)

CodeImage TranslationSegmentationData SynthesisRepresentation LearningConvolutional Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasoundDiffusion Tensor Imaging

🎯 What it does: Propose the MedFuse-Seg model to achieve language-driven medical image segmentation through multi-level visual and semantic context fusion and reasoning-guided mask decoding.

Medical Knowledge-Guided Fusion of Holistic Hospital Data for Tumor Survival Analysis

Guan, Jinquan (South China University of Technology), Xie, Yutong (South China University of Technology)

CodeExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningDrug DiscoveryGraph Neural NetworkTransformerLarge Language ModelMixture of ExpertsContrastive LearningMultimodalityBiomedical DataElectronic Health Records

🎯 What it does: This paper proposes a medical knowledge-driven multi-modal fusion framework named MKGSurv, which integrates pan-hospital data from four disciplinesβ€”pathology, genetics, clinical, and treatmentβ€”to achieve tumor survival analysis.

MedPro: Dual-View Propagation and Prototypical Assessment for Training-Free Few-Shot Medical Segmentation

Zhang, Jingyi (Nanjing University of Science and Technology), Xie, Guo-Sen (Nanjing University of Science and Technology)

CodeSegmentationTransformerPrompt EngineeringDiffusion modelContrastive LearningBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose a training-free MedPro framework for few-shot medical image segmentation.

MedTri: Structured Medical Report Normalization for Enhanced Vision–Language Pretraining

Chu, Yuetan (King Abdullah University of Science and Technology), Gao, Xin (King Abdullah University of Science and Technology)

CodeClassificationFederated LearningExplainability and InterpretabilityComputational EfficiencyRepresentation LearningData-Centric LearningConvolutional Neural NetworkTransformerLarge Language ModelPrompt EngineeringVision Language ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageTextMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyElectronic Health Records

🎯 What it does: Propose the MedTri framework, which converts free-text medical reports into structured triplets [anatomical entity: radiological description + diagnostic category], to enhance the quality of text supervision in vision-language pre-training.

MedTriage-LM: Anatomically Grounded Visual Phenotype Synthesis for Interpretable ED Triage

Lu, Zhixiang (Xi'an Jiaotong-Liverpool University), Song, Sifan (Xi'an Jiaotong-Liverpool University)

CodeClassificationExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringVision Language ModelDiffusion modelTextMultimodalityTabularBiomedical DataElectronic Health RecordsRetrieval-Augmented Generation

🎯 What it does: This paper proposes MedTriage-LM, a multimodal large language model that integrates clinical tables, text, and synthetic visual phenotyping maps (VPM), for emergency triage instruction prediction and interpretable reasoning;

MedTriFlow: Efficient Resolution-Agnostic 3D Medical Image Generation with Implicit Triplane Representation

Xu, Chenfan (ShanghaiTech University), Cui, Zhiming (ShanghaiTech University)

CodeRestorationGenerationTransformerDiffusion modelScore-based ModelFlow-based ModelRectified FlowNeural Radiance FieldAuto EncoderBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose the MedTriFlow framework, which combines tri-plane representation with implicit anatomical fields to achieve 3D medical image generation from a compressed latent space to arbitrary resolutions, and enables efficient inference on medical image generation and reconstruction tasks.

MedTS-TTT: Test-Time Training for Medical Time Series Classification

Chen, Mingzhi (Peking University), Luo, Guibo (Peking University)

CodeClassificationAnomaly DetectionComputational EfficiencyMeta LearningConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningTime SeriesBiomedical DataElectrocardiogramReview/Survey Paper

🎯 What it does: Proposed the MedTS-TTT framework for test-time training in the classification of medical time series (EEG, ECG).

MedTSC-Net: Bridging Semantic and Scale Gaps in Text-Guided Medical Image Segmentation

Chen, Zhaomin (Wenzhou University), Chen, Huiling (Hangzhou Dianzi University)

CodeImage HarmonizationSegmentationConvolutional Neural NetworkTransformerVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose a text-guided medical image segmentation framework named MedTSC-Net, which is used for accurately segmenting lesions in COVID-19 chest imaging (X-ray and CT).

Memory-Guided Random-Direction Feature Disentanglement Model for Multi-phase MRI Translation

Xiao, Qianmu (Central South University), Zou, Beiji (Manchester Metropolitan University)

CodeImage TranslationRestorationSegmentationDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningOptical FlowImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Proposed a multi-phase CE-MRI translation framework (MG-RDFD) based on a shared VAE encoder and random direction feature disentanglement, achieving unified latent space separation of content and style, and realizing region-level style injection through a segmentation-guided memory pool.

MERIT: Multi-scale Mamba with Enhanced Information Retention for Treatment Response Prediction from Whole Slide Images

Hu, Taiyuan (Chinese Academy of Sciences), Yan, Rui (University of Science and Technology of China)

CodeClassificationImage TranslationDrug DiscoveryTransformerMixture of ExpertsContrastive LearningImageBiomedical DataComputed TomographyReview/Survey Paper

🎯 What it does: Designed and implemented a multi-scale Mamba framework called MERIT for predicting tumor treatment response from whole slide images (WSI).

Merlin Plus: A Large-Scale, Multi-cancer, Image-Mask-Report Dataset

Bassi, Pedro R. A. S. (Johns Hopkins University), Zhou, Zongwei (Johns Hopkins University)

CodeSegmentationConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningImageTextBiomedical DataComputed Tomography

🎯 What it does: This study first constructed the Merlin Plus dataset, providing 1,153 voxel-level tumor masks covering nine organs (spleen, bladder, gallbladder, stomach, duodenum, uterus, prostate, adrenal gland, esophagus) and supplemented with longitudinal metadata such as patient identifiers and scan times; subsequently, segmentation models such as MedFormer and R-Super were trained using these masks, and were compared and evaluated with the original Merlin dataset and publicly available multi-tumor segmentation models (Voxtell, ULS, FLARE).

MetaFormer: Efficient Metadata-Guided Transformer for Patient-Aware 12-Lead Electrocardiogram Reconstruction

Xie, Yanchong (Sun Yat-sen University), Huang, Kai (Sun Yat-sen University)

CodeRestorationTransformerAuto EncoderContrastive LearningTabularElectrocardiogram

🎯 What it does: This paper proposes MetaFormer, a Transformer architecture guided by patient metadata, which reconstructs single-lead ECG into complete 12-lead ECG.

MicroscopyCLIP: A Domain-Specific Vision-Language Model for Optical Microscopy

Yang, Zhuoqin (Shenzhen University), Shen, Linlin (Shenzhen University)

CodeClassificationImage TranslationDomain AdaptationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningVision Language ModelContrastive LearningImageTextBiomedical DataBenchmark

🎯 What it does: This paper proposes and trains a domain-specific vision-language model called MicroscopyCLIP for optical microscopy images.

Missingness-Aware Multimodal Learning for Four-Class Adrenal Tumor Classification

Chen, Dehua (Donghua University), An, Huimin (Shanghai Jiao Tong University)

CodeClassificationData-Centric LearningConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningImageTabularBiomedical DataComputed TomographyElectronic Health Records

🎯 What it does: Designed and implemented a missing-aware multimodal learning framework for four-class classification of adrenal tumors (Cushing's syndrome, primary aldosteronism, pheochromocytoma, and non-functional adenoma) using non-contrast CT images and conventional non-hormonal clinical variables.

MLLM-Enhanced Region-Aware Bidirectional Evidence-Based Model for Tongue Diagnosis

Du, Yiwei (Nanjing University), Shan, Caifeng (Nanjing University)

CodeClassificationExplainability and InterpretabilityRepresentation LearningConvolutional Neural NetworkTransformerLarge Language ModelMixture of ExpertsVision Language ModelContrastive LearningImageMultimodalityBiomedical Data

🎯 What it does: Propose a region-aware bidirectional evidence model integrated with a multimodal large language model for simultaneously predicting tongue coating syndrome patterns and visceral states.

MM-UNet: Meta Mamba UNet for Volumetric Medical Image Segmentation

Xie, Bin (Illinois Institute of Technology), Agam, Gady (Illinois Institute of Technology)

CodeSegmentationConvolutional Neural NetworkMixture of ExpertsBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose a unified U-shaped structure MM-UNet, combining convolutional networks with state space models (Mamba) for volumetric medical image segmentation.

MMIC-EndoDepth: Self-Supervised Endoscopic Depth Estimation with Multi-Mechanism Illumination Correction

Xu, Ziang (Chinese University of Hong Kong), Wang, Bing (Hong Kong Polytechnic University)

CodeDepth EstimationConvolutional Neural NetworkTransformerMixture of ExpertsAuto EncoderContrastive LearningVideoBiomedical DataBenchmark

🎯 What it does: Proposes the MMIC-EndoDepth framework, which estimates depth using self-supervised monocular endoscopic videos and improves reconstruction quality through multi-mechanism illumination correction.

MoCaf-Mamba: Modality Completion and Alignment in Feature Space for Missing-Modality Segmentation

Zhou, Yongsong (Shenzhen University), Shen, Linlin (Shenzhen University)

CodeSegmentationConvolutional Neural NetworkTransformerDiffusion modelContrastive LearningMultimodalityBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Proposes the MoCaf-Mamba framework to address missing modalities and residual geometric mismatches in multi-modal medical image segmentation.

MORI-Seg: Learning Morphological Geometry for Instance Segmentation Without Instance Annotations

Zhao, Leiyue, Deng, Ruining (Vanderbilt University)

CodeSegmentationConvolutional Neural NetworkTransformerDiffusion modelContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: In the absence of instance annotations, learn morphological geometric information through semantic segmentation results to achieve instance segmentation of kidney tissue;

Motion-Conditioned Multi-view Fusion for Myocardial Infarction Localization from Echocardiography

Yang, Guang (University of Oxford), Grau, Vicente (University of Oxford)

CodeSegmentationAnomaly DetectionConvolutional Neural NetworkTransformerContrastive LearningGaussian SplattingOptical FlowVideoBiomedical DataUltrasound

🎯 What it does: Propose a dual-view echocardiography myocardial infarction localization framework named MCF-Net, which combines extremely sparse annotated motion priors with a pre-trained EchoPrime base model;

MPFlow: Multi-modal Posterior-Guided Flow Matching for Zero-Shot MRI Reconstruction

Kim, Seunghoi (University College London), Alexander, Daniel C. (University College London)

CodeRestorationConvolutional Neural NetworkTransformerDiffusion modelFlow-based ModelRectified FlowContrastive LearningMultimodalityBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Proposed a zero-shot multi-modal MRI reconstruction framework called MPFlow, which reduces artifacts in reconstruction by guiding the reconstruction process with additional auxiliary modalities during inference, using a pre-trained unconditional flow model.

MRI2Rep: Autoregressive Structured Report Generation for 3D Liver MRI

Li, Xinran (Yale University), Staib, Lawrence H. (Yale University)

CodeImage TranslationGenerationRepresentation LearningConvolutional Neural NetworkTransformerLarge Language ModelVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingRetrieval-Augmented Generation

🎯 What it does: Developed MRI2Rep, an end-to-end framework for automatically generating structured reports from 3D liver MRI, which directly predicts diagnostic sequences from multi-phase volumetric images and renders them into standardized reports using templates.

MRSeg3D: A Perturbation-Resilient Model for 3D Medical Reasoning Segmentation

Hao, Qin (Xinjiang University), Ye, Xujiong (University of Exeter)

CodeSegmentationConvolutional Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelImageTextBiomedical DataComputed TomographyBenchmark

🎯 What it does: Proposed the MRSeg3D model and its three-level progressive perturbation benchmark to evaluate and enhance the robustness of 3D medical image reasoning segmentation under real language instructions.

MS-GDUN: Multi-scale Graph Deep Unfolding Network with Gradient-Threshold Learning for FMT Reconstruction

He, Xiaowei (Northwest University), Guo, Hongbo (Northwest University)

CodeOptimizationGraph Neural NetworkDiffusion modelAuto EncoderContrastive LearningGraphBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Proposed a physics-guided multi-scale graph deep unrolling network (MS-GDUN) for inversion reconstruction in fluorescence molecular tomography (FMT).

MSHF-Net: Multimodal Breast Cancer Molecular Subtype Prediction via Segmentation-Guided Hierarchical Fusion Network

Jiang, Xinjie (Hangzhou Dianzi University), Wang, Changmiao (Shenzhen Research Institute of Big Data)

CodeClassificationSegmentationAnomaly DetectionConvolutional Neural NetworkTransformerContrastive LearningImageMultimodalityTabularBiomedical DataMagnetic Resonance ImagingUltrasoundElectronic Health Records

🎯 What it does: A multi-modal segmentation-guided hierarchical fusion network (MSHF-Net) was constructed for non-invasive pre-surgical molecular subtype prediction of breast cancer, integrating mammography (MG), ultrasound (US), and clinical information.

MSQG-3DNet: Multi-scale Token Selection with Query-Decoder and Gated Fusion for 3D Breast Tumor Classification

Liu, Zefeng (Macao Polytechnic University), Tan, Tao (Macao Polytechnic University)

CodeClassificationConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningBiomedical DataUltrasound

🎯 What it does: A model named MSQG‑3DNet is proposed for classifying the lesion-level ROI in 3D automated breast ultrasound (ABUS) as benign or malignant.

MSS-Net: Learning Hierarchical Multi-scale and Multimodal Representations for Complex Cardiac Disease Classification

Chen, Yuling (Central South University), Zeng, Feng (Central South University)

CodeClassificationConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningTextTime SeriesElectrocardiogram

🎯 What it does: Proposed the MSS-Net framework for multi-label ECG classification in complex cardiac diseases.

MuellerPT: Decomposition Driven Pre-training for Dense Learning in Mueller Polarimetry

Tlemsani, Adam (Imperial College London), Elson, Daniel S. (Imperial College London)

CodeClassificationSegmentationConvolutional Neural NetworkSupervised Fine-TuningContrastive LearningImageBiomedical DataPhysics Related

🎯 What it does: A pre-training method called MuellerPT was developed, which learns dense representations by predicting the Lu-Chipman decomposition parameters of the Mueller matrix, and was evaluated on tasks such as segmentation of gray-white brain tissue and classification of colorectal cancer.

Multi-agent Test-Time Adaptation for Robust Medical Image-to-Image Translation

Iele, Irene (UniversitΓ  Campus Bio-Medico di Roma), Tortora, Matteo (University of Genoa)

CodeImage TranslationDomain AdaptationAgentic AIDiffusion modelFlow-based ModelAuto EncoderGenerative Adversarial NetworkImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Proposed a selection-based test-time adaptation framework based on multi-agent for image-to-image translation tasks in medical imaging.

Multi-Annotation Adaption: A Label-Informed Dynamic Framework for Medical Segmentation and Localization

Han, Luyi (Macao Polytechnic University), Mann, Ritse (Netherlands Cancer Institute)

CodeSegmentationOptimizationHyperparameter SearchData-Centric LearningConvolutional Neural NetworkTransformerContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound

🎯 What it does: Propose a multi-annotation dynamic framework called MALFOY, which encodes different annotation types through label weight encoding, achieving adaptive learning for medical image segmentation and localization.

Multi-Disease Diagnosis in Retinal Images via Patch-Level Reasoning and Selective Aggregation

Xie, Jianyang (University of Liverpool), Zheng, Yalin (University of Liverpool)

CodeClassificationAnomaly DetectionExplainability and InterpretabilityGraph Neural NetworkTransformerContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasoundFibre Orientation DistributionDiffusion Tensor ImagingAlzheimer's DiseaseElectronic Health RecordsElectrocardiogram

🎯 What it does: Proposes a multi-disease retinal image diagnosis framework that combines patch-level reasoning with selective aggregation to address lesion-disease mismatch caused by comorbidities.

Multi-parametric MRI for Contrast Agent Free Breast DCE-MRI Synthesis

Yang, Zhikai (KTH Royal Institute of Technology), Moreno, Rodrigo (KTH Royal Institute of Technology)

CodeImage TranslationData SynthesisConvolutional Neural NetworkDiffusion modelGenerative Adversarial NetworkImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Built a model based on conditional generative adversarial networks to synthesize multi-phase images of breast DCE-MRI from non-contrast-enhanced multi-parametric MRI (T1w, DWI, and ADC).

Multi-scale Dual-domain Fusion-Based Spatio-Temporal Graph Learning for fNIRS Signal Classification

Chu, Mengxiang (Northwest University), Guo, Hongbo (Northwest University)

CodeClassificationGraph Neural NetworkTransformerMixture of ExpertsContrastive LearningTime SeriesBiomedical Data

🎯 What it does: Propose a multi-scale dual-domain fusion spatiotemporal graph learning framework, MsDFSTGL, for multi-class classification of fNIRS signals.

Multi-Stage Dilated Convolution-Mamba Network for Surgical Action Triplet Recognition

Meng, Qingke (Jiangnan University), Zhou, Tao (Jiangnan University)

CodeRecognitionConvolutional Neural NetworkGraph Neural NetworkTransformerMixture of ExpertsContrastive LearningVideoTabularBenchmark

🎯 What it does: Propose a multi-stage sparse convolution-Mamba network (DCM-Net) for end-to-end identification of instrument-verb-target triplets in surgical videos.

Multi-stage NeRF for Efficient 3D Coronary Artery Reconstruction from Two Narrow-Angle Angiographic Projections

Meng, Deyu (University of Oxford), Banerjee, Abhirup (University of Oxford)

CodeOptimizationComputational EfficiencyDiffusion modelNeural Radiance FieldAuto EncoderContrastive LearningImageBiomedical DataComputed TomographyReview/Survey Paper

🎯 What it does: Reconstructs complete 3D vascular structures from two X-ray coronary angiography images with narrow viewing angles using a self-supervised multi-stage NeRF framework.

Multilayer Modularity-Aware Graph Autoencoder for Detecting Subpopulation Community Structure in Dynamic Brain Networks

See, Kai-Jun (Monash University Malaysia), Ting, Chee-Ming (Monash University Malaysia)

CodeAnomaly DetectionExplainability and InterpretabilityComputational EfficiencyRepresentation LearningGraph Neural NetworkAuto EncoderContrastive LearningGraphBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose a multi-layer modular amplitude perception graph autoencoder (MMGAE-WSBM), which decouples individual differences from the dynamic community structure of brain networks by identifying potential subgroups and using a contrastive coupling mechanism within the autoencoder.

Multimodal Contrastive Regression for Organ-Resolved Biological Age Prediction

Ecker, Veronika (University of Stuttgart), Yang, Bin (University Hospital of Tuebingen)

CodeConvolutional Neural NetworkTransformerContrastive LearningImageMultimodalityTabularBiomedical DataMagnetic Resonance Imaging

🎯 What it does: This paper proposes a multimodal contrastive regression framework that jointly uses MRI images and structured clinical variables to predict the biological age of various organs.

Multimodal Large Language Model-Driven Self-Verification Reasoning for Interpretable Diagnosis of Diffuse Cystic Lung Diseases

Jia, Qiwei (University of Science and Technology of China), Hu, Xiaowen (Central South University)

CodeAnomaly DetectionExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringImageMultimodalityBiomedical DataComputed TomographyRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed a self-verification reasoning framework based on a multi-modal large language model, utilizing global spatial awareness and multi-step self-verification to achieve interpretable diagnosis of diffuse cystic lung disease.

Multimodal-Conditioned Flow Matching for Pathology Nuclei Data Augmentation

Zhang, Yanan (Beihang University), Bai, Xiangzhi (Beihang University)

CodeSegmentationGenerationData SynthesisConvolutional Neural NetworkTransformerDiffusion modelFlow-based ModelContrastive LearningImageTextMultimodalityBiomedical DataOrdinary Differential Equation

🎯 What it does: A multi-modal conditional flow matching framework was developed to synthesize paired nuclear masks and pathological images for data augmentation.

Mutual Distillation of Dual-Foundation Models for Semi-supervised PET/CT Segmentation

Mao, Fuyou (Central South University), Tang, Yan (Central South University)

CodeSegmentationKnowledge DistillationConvolutional Neural NetworkTransformerPrompt EngineeringContrastive LearningGaussian SplattingBiomedical DataComputed TomographyPositron Emission Tomography

🎯 What it does: This paper proposes a mutual distillation framework (MuDuo) based on a dual baseline model to achieve semi-supervised PET/CT organ segmentation.

MUX-USCT: A Noise-Robust Neural Network for Ultrasound Computed Tomography

Yuan, Yuchen (George Mason University), Yang, Lei (University of North Carolina at Chapel Hill)

CodeRestorationTransformerAuto EncoderContrastive LearningBiomedical DataComputed TomographyUltrasound

🎯 What it does: Designed and implemented a deep learning architecture named MUX-USCT for ultrasound computed tomography (USCT) reconstruction under known acoustic acquisition geometry, achieving robustness to noise through adaptive MUX and attention mechanisms.

MVDA-Net: Multi-View Dual-Alignment Network for 3D Carotid Artery Segmentation

Wen, Yang (Shenzhen University), Sheng, Bin (Hong Kong University of Science and Technology)

CodeSegmentationConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningBiomedical DataComputed Tomography

🎯 What it does: Proposed MVDA-Net for precise segmentation of carotid arteries in 3D CTA images;

NeRD: Neuro-Symbolic Rule Distillation for Efficient Ontology-Grounded Chain-of-Thought in Medical Image Diagnosis

Yang, Hongxi (Monash University), Ge, Zongyuan (Monash University)

CodeClassificationExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningImageMultimodalityBiomedical DataChain-of-Thought

🎯 What it does: Propose the NeRD framework, which combines data-driven logical rules with multi-modal chain reasoning to generate efficient ontological diagnostic reasoning chains.

Network-Aware Bilinear Tokenization for Brain Functional Connectivity Representation Learning

Milecki, Leo, Zhao, Qingyu (Weill Cornell Medicine)

CodeRepresentation LearningTransformerAuto EncoderContrastive LearningBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Proposes the NERVE framework, which uses network-aware bilinear decomposition to tokenize brain functional connectivity matrices in blocks, and learns transferable features in MAE self-supervised learning.

NeuroFlow: Manifold-Aware Topological Flow Matching for EEG Decoding

Wang, Lei (South China University Of Technology), Xu, Yanwu (Pazhou Lab)

CodeGenerationRetrievalRepresentation LearningGraph Neural NetworkTransformerDiffusion modelFlow-based ModelAuto EncoderContrastive LearningImageMultimodalityAudio

🎯 What it does: Construct the NeuroFlow framework to achieve decoding and generation of visual content from EEG signals.