arXivSub Start free trial

MICCAI 2026 Papers — Page 7

International Conference on Medical Image Computing and Computer-Assisted Intervention · 1165 papers

LRMIL: Efficient Low-Resolution Multiple Instance Learning via High-Resolution Knowledge Distillation for Whole Slide Image Classification

Shin, Yonghan (Korea University), Jeong, Won-Ki (Korea University)

ClassificationComputational EfficiencyKnowledge DistillationTransformerContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasound

🎯 What it does: Propose a low-resolution multiple instance learning framework (LRMIL), which enables pathological image classification using only low-resolution slices through high-resolution knowledge distillation.

LumiState: Retinex-inspired Illumination Decoupling in 4D Gaussian Splatting for Endoscopic Reconstruction

Sun, Dai (University of Science and Technology of China), Zhou, Shaohua Kevin

RestorationConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelNeural Radiance FieldAuto EncoderContrastive LearningGaussian SplattingImageVideoPoint Cloud

🎯 What it does: Proposes a method combining illumination decomposition based on Retinex theory with 4D Gaussian scattering (LumiState), achieving separation of variable near-field illumination and tissue reflectance in endoscope videos, thereby improving the geometric and visual quality of dynamic 3D reconstruction.

M3D-QAdapter: 3D Medical VQA with Lesion-Level Finding-Segmentation Alignment and Query-Driven Adaptive Token Reduction

Liu, Hong (Xiamen University), Wang, Liansheng (Xiamen University)

ClassificationSegmentationComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataComputed Tomography

🎯 What it does: M3D-QAdapter proposes a two-stage 3D medical vision question answering framework, first training a 3D ViT with lesion-level alignment, and then using a query-driven adaptive process to compress visual tokens to generate answers.

MAGE: Color-Invariant and Spatial Knowledge Distillation for Gastric Neoplasm Classification

Jun, Jiho (Korea University), Uhm, Kwang-Hyun (Gachon University)

ClassificationDomain AdaptationExplainability and InterpretabilityKnowledge DistillationConvolutional Neural NetworkAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound

🎯 What it does: Proposed and implemented a dual-branch training framework named MAGE, which utilizes mask-based grayscale expert and spatial attention distillation to eliminate color bias and enhance the discriminability of morphological features between gastric adenomas and cancers during training, thereby achieving binary classification of gastric adenomas and cancers under full-frame WLI.

MAGEFormer: Learning Metric-Consistent Representations for Anisotropic CT Segmentation

Li, Jiaying (Durham University), Remagnino, Paolo (Durham University)

SegmentationTransformerContrastive LearningBiomedical DataComputed Tomography

🎯 What it does: This paper addresses the problem of uneven longitudinal resolution of clinical CT voxels by proposing MAGEFormer—a Transformer structure that performs geometric calibration in position encoding, cross-scale attention, and the inference phase, achieving metric-consistent representation learning for 3D medical images.

Making HE Histopathological Images More Colorful by Conditional Flow Matching

Ben Omrane, Mohamed Salim (Université Paris-Saclay), Pesquet, Jean-Christophe (Université Paris-Saclay)

Image TranslationRestorationGenerationData SynthesisConvolutional Neural NetworkDiffusion modelScore-based ModelFlow-based ModelImageBiomedical Data

🎯 What it does: This paper proposes a conditional flow matching model based on the concentration domain, generating missing eosin components from HE-stained images to obtain complete HES-stained images.

MamaDino: Mammography-Aware Multi-view Attentional DINO for 3-Year Breast Cancer Risk Prediction

Santeramo, Ruggiero (Fondazione Human Technopole), Jug, Florian (Fondazione Human Technopole)

ClassificationAnomaly DetectionConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose a hybrid visual model, MamaDino, that predicts the risk of breast cancer within three years using low-resolution (512×512) breast X-ray images.

Mamba Based Anisotropic Diffusion Model for Speckle Noise Reduction in Ultrasound Images

Pham, Nhat Duy (Université Sorbonne Paris Nord), Trinh, Dinh Hoan (Viettel AI)

RestorationTransformerDiffusion modelImageBiomedical DataUltrasound

🎯 What it does: Propose a Mamba-based hybrid model for ultrasound speckle noise reduction, combining the diffusion process of partial differential equations with a learnable diffusion function to form an end-to-end trainable iterative denoising framework.

MammoFlow: Multiview Mammogram Synthesis with Anatomically Consistent Flow Matching

Du, Yuexi (Yale University), Dvornek, Nicha C. (Yale University)

GenerationData SynthesisTransformerDiffusion modelScore-based ModelFlow-based ModelRectified FlowAuto EncoderOptical FlowImageBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound

🎯 What it does: Proposes a multi-view breast X-ray synthesis framework called MammoFlow based on flow matching, which can generate consistent CC and MLO views without requiring real reference images.

MAP-Diff: Multi-anchor Guided Diffusion for Progressive 3D Whole-Body Low-Dose PET Denoising

Jing, Peiyuan (Zurich University of Applied Sciences), Montoya-Zegarra, Javier A. (Imperial College London)

RestorationConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelBiomedical DataPositron Emission Tomography

🎯 What it does: Propose MAP-Diff, which utilizes a multi-anchor guided diffusion model to achieve step-by-step denoising of three-dimensional whole-body low-dose PET

MARBLE: Lightweight EEG-to-fMRI Translation via Mamba-Attention and ROI-Conditioned Decoding

Kim, Emil (Pusan National University), Gahm, Jin Kyu (Pusan National University)

Image TranslationExplainability and InterpretabilityComputational EfficiencyTransformerAuto EncoderContrastive LearningTime SeriesBiomedical DataMagnetic Resonance ImagingElectrocardiogram

🎯 What it does: Propose MARBLE, a lightweight EEG-to-fMRI translation framework capable of predicting whole-brain BOLD time series from EEG;

Marginal Constrained Morphological Prototype Learning for Patch Search in Whole Slide Images

Park, Sihyeon (Korea University), Kim, Bumsoo (Chung Ang University)

ClassificationRetrievalComputational EfficiencyRepresentation LearningConvolutional Neural NetworkTransformerMixture of ExpertsContrastive LearningImageBiomedical DataBenchmark

🎯 What it does: This paper proposes a selective slice-level prediction framework called MMPL based on global morphological prototypes. It first uses the Sinkhorn-Knopp algorithm and momentum feature queue to learn global prototypes, and then uses these prototypes to retrieve the top-k most diagnostically valuable slices from each slide for aggregation, thereby achieving end-to-end training and significantly reducing computational costs.

MaRS: Robust Out-of-Distribution Detection via Mahalanobis Residual Scoring

Di Salvo, Francesco (University of Bamberg), Ledig, Christian (University of Bamberg)

Anomaly DetectionTransformerAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound

🎯 What it does: A post-hoc, label-free anomaly detection method based on the Mahalanobis distance using residual features from an autoencoder (MaRS) is proposed for medical images, which can be applied on frozen baseline model features.

Mask to Concept: Auto-Promptable SAM3 via Efficient Test-Time Concept Embedding Search for Few-Shot Annotation

Zhou, Quan (Wuhan University of Technology), Wang, Zhiwei (Huazhong University of Science and Technology)

SegmentationData-Centric LearningTransformerPrompt EngineeringContrastive LearningImageBiomedical Data

🎯 What it does: Propose the Mask to Concept (M2C) framework, leveraging the concept prompting capability of SAM3, automatically searching for visual concepts from a small number of labeled samples on the frozen SAM3 architecture, achieving medical image few-shot automatic annotation and realizing efficient human-machine collaborative closed-loop;

Mask-Guided Attention Regulation for Anatomically Consistent Counterfactual CXR Synthesis

Zhang, Zichun (Tianjin University), Su, Yuting (Tianjin University)

GenerationData SynthesisTransformerPrompt EngineeringDiffusion modelImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Local controllable generation of pathologies in chest X-ray images was achieved by regulating the attention of diffusion models during inference, while maintaining anatomical consistency.

MASurv: Risk-Conditional Martingale Adversarial Training for Calibrated Multi-slide Survival Prediction in Whole-Slide Histopathology Images

Sabouri Rad, Meghdad (SUNY Upstate Medical University), Rodd, Bardia (SUNY Upstate Medical University)

Domain AdaptationAnomaly DetectionRepresentation LearningConvolutional Neural NetworkTransformerGenerative Adversarial NetworkContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyReview/Survey Paper

🎯 What it does: Under a multi-instance learning framework, the MASurv model is constructed by performing adversarial conditional moment test (ACMT) on the residual increments of the counting process for hierarchical calibration, and introducing consistency constraints between multi-slice patients.

Maximizing Domain Generalization in Automated Fetal Brain Biometry

Li, Yijin (Tsinghua University), Tian, Qiyuan (Tsinghua University)

Domain AdaptationConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Proposes BioTTA, a source-agnostic, unsupervised test-time adaptation framework for automatic fetal brain biometry.

MBFAD: Benchmarking Balance Function Assessment with Environmental Perturbation and Synergistic Tasks

Ge, Zhaoyang (Zhengzhou University), Xu, Mingliang (Zhengzhou University)

ClassificationAnomaly DetectionConvolutional Neural NetworkSupervised Fine-TuningContrastive LearningMultimodalityTime SeriesBenchmark

🎯 What it does: This paper proposes the MBFAD multimodal balanced function assessment dataset and the M²-Balance framework, aiming to achieve real-time evaluation of imbalance states under static, dynamic, and reactive activities through synchronized inertial and plantar pressure data.

MDL-DA Track: Self-supervised Sperm Motility Direction Learning and Density-Adaptive Association for Sperm Tracking

Zeng, Xiaoyu, Yang, Xuan (Shenzhen University)

Object TrackingConvolutional Neural NetworkTransformerContrastive LearningOptical FlowVideoBiomedical Data

🎯 What it does: This paper proposes a multi-object tracking framework that combines self-supervised motion direction learning and density-adaptive association, for precise tracking of sperm in high-density microscopic videos.

MDPNet: A Multi-Scale Deformable Prototype Network for Accurate and Interpretable Medical Image Classification

Yan, Qinglan, Ma, Changsheng (First Affiliated Hospital of Soochow University)

ClassificationExplainability and InterpretabilityConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningImageBiomedical DataComputed TomographyPositron Emission TomographyElectrocardiogram

🎯 What it does: Designed a medical image classification network called MDPNet based on multi-scale deformable prototypes, using prototype learning to interpret diagnostic results and improve accuracy.

MEASURE: Multi-Task Slice Selection and Regression for Neonatal Brain Biometry

Lee, Jiyang (Hanyang University), Kim, Seh Hyun (Seoul National University)

Convolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose the MEASURE framework, which decomposes the neonatal brain measurement task into two stages: ViT slice selection and multi-task regression, completing six biometric measurements according to the Kidokoro protocol without relying on voxel-level annotations.

Measuring Prediction Uncertainty in Neural Cellular Automata

Sadafi, Ario (Helmholtz Munich), Marr, Carsten (Helmholtz Munich)

SegmentationExplainability and InterpretabilityImageBiomedical DataBenchmark

🎯 What it does: Studies the uncertainty estimation of neural cellular automata (NCA) in medical image segmentation, proposing a training-agnostic perturbation recovery method called Resilience.

Measuring What VLMs Don’t Say: Validation Metrics Hide Clinical Terminology Erasure in Radiology Report Generation

Parikh, Aditya (Technical University of Denmark), Frank, Stella (Technical University of Denmark)

GenerationExplainability and InterpretabilityData-Centric LearningLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataElectronic Health RecordsRetrieval-Augmented Generation

🎯 What it does: This paper proposes two new evaluation metrics—Clinical Association Displacement (CAD) and Weighted Association Elimination (WAE)—to detect lexical disappearance and potential biases in visual-language models when generating chest X-ray reports; meanwhile, experiments reveal the impact of decoding strategies (deterministic vs. random sampling) on clinical information retention and fairness.

MEC: A Multi-expert Consultation Framework for Synergizing Frozen Foundation Models in Whole Slide Image Analysis

Chen, Zongyi (Xiamen University), Wang, Liansheng (Hong Kong University of Science and Technology)

ClassificationAnomaly DetectionComputational EfficiencyRepresentation LearningTransformerMixture of ExpertsImageBiomedical Data

🎯 What it does: Propose a MEC multi-expert negotiation framework that utilizes multiple frozen foundation models combined with MIL for diagnosis in WSI analysis.

MEC: Medical Evidence Capsules for Retrieval-Augmented Generation in Medical Multimodal Question Answering

Xu, Zhenghua (Hebei University Of Technology), Tian, Tian (Hebei University Of Technology)

RetrievalExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningDrug DiscoveryTransformerLarge Language ModelPrompt EngineeringVision Language ModelTextMultimodalityGraphTabularBiomedical DataElectronic Health RecordsBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose Medical Evidence Capsules (MEC), which unify heterogeneous evidence such as medical text, structured knowledge, model parameters, and imaging reports into a structured capsule format for retrieval-augmented generation (RAG) in medical multimodal question answering.

Med-CAP: Counterfactual Evidence and Adaptive Prior Suppression for Robust Medical Visual Question Answering

Huang, Zaiqiang (Tsinghua University), Wu, Xian (Tencent Jarvis Lab)

Domain AdaptationExplainability and InterpretabilityRepresentation LearningTransformerSupervised Fine-TuningVision Language ModelContrastive LearningImageTextBiomedical DataBenchmark

🎯 What it does: This paper proposes Med-CAP, a framework that achieves robustness in medical visual question answering through contrastive evidence modeling and adaptive prior suppression.

Med3D-R1: Mitigating Narrative Bias and Enhancing Reasoning Consistency in 3D Medical Vision-Language Models

Lai, Haoran (University of Science and Technology of China), Zhou, Shaohua Kevin (University of Science and Technology of China)

Anomaly DetectionExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerReinforcement LearningPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataComputed Tomography

🎯 What it does: Built and trained Med3D-R1, a vision-language model based on CT 3D medical imaging, for abnormal diagnosis and to improve reasoning consistency.

MedCenterDet: A Center-Based Object Detection Framework for 3D Medical Image

Zeng, Qiang (Lingnan University), Pan, Fei (Shenzhen University)

Object DetectionConvolutional Neural NetworkTransformerGenerative Adversarial NetworkContrastive LearningGaussian SplattingOptical FlowImageBiomedical DataMagnetic Resonance ImagingComputed TomographyStochastic Differential Equation

🎯 What it does: Proposed an anchor-free 3D medical image object detection framework called MedCenterDet, which directly locates the target center and infers the size through center point heatmap regression.

MedConcept: Unsupervised Concept Discovery for Interpretability in Medical VLMs

Haque, Md Rakibul (University of Utah), Elhabian, Shireen Y. (University of Utah)

Explainability and InterpretabilityTransformerLarge Language ModelVision Language ModelAuto EncoderContrastive LearningImageTextBiomedical DataComputed Tomography

🎯 What it does: Achieve unsupervised concept discovery in medical vision-language models (VLMs), and quantitatively verify the semantic consistency between concepts and radiology reports using large language models (LLMs).

MedEnv: Scaling Multimodal Virtual Medical Environments for Long-Horizon Diagnosis

Fan, Zhiting (Zhejiang University), Liu, Zuozhu (Zhejiang University)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningData-Centric LearningDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackTransformerReinforcement LearningPrompt EngineeringVision-Language-Action ModelTextMultimodalityTabularTime SeriesBiomedical DataElectronic Health RecordsRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Built the MedEnv multimodal long-term diagnostic simulation environment, supporting multi-round Q&A and tool-assisted evidence acquisition

MedExpMem: Adapting Experience Memory for Differential Diagnosis

Feng, Qianhan (Chinese University of Hong Kong), Dou, Qi (Chinese University of Hong Kong)

Recommendation SystemExplainability and InterpretabilityData-Centric LearningTransformerVision Language ModelImageTextBiomedical DataElectronic Health RecordsRetrieval-Augmented Generation

🎯 What it does: Proposes the MedExpMem framework, adding sustainable experiential memory to vision-language diagnostic models.

MedFM-Robust: Benchmarking Robustness of Medical Foundation Models

Cui, Xiangxiang (Beijing Normal University), Yin, Lu (University College London)

Image TranslationRestorationSegmentationAnomaly DetectionTransformerSupervised Fine-TuningPrompt EngineeringVision Language ModelVision-Language-Action ModelDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasoundReview/Survey PaperBenchmark

🎯 What it does: This paper proposes a robustness benchmark for medical foundation models (MedFM), systematically evaluating the impact of 40 modality-based perturbations (12 basic perturbations and 28 medical-specific perturbations) on visual language models and segmentation models across eight imaging modalities;

MedFuse-Seg: Multi-level Visual and Semantic Context Fusion for Segmentation-Based Medical Reasoning

Limaroon, Keetawan, Achakulvisut, Titipat (Mahidol University)

Image TranslationSegmentationData SynthesisRepresentation LearningConvolutional Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasoundDiffusion Tensor Imaging

🎯 What it does: Propose the MedFuse-Seg model to achieve language-driven medical image segmentation through multi-level visual and semantic context fusion and reasoning-guided mask decoding.

Medical Image Spatial Grounding with Semantic Sampling

Yu, Andrew Seohwan (Case Western Reserve University), Chaudhary, Vipin (Case Western Reserve University)

RecognitionSegmentationPrompt EngineeringVision Language ModelContrastive LearningImageMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyBenchmark

🎯 What it does: This paper proposes the MIS-Ground benchmark to systematically evaluate and challenge the spatial localization capabilities of vision-language models (VLMs) in 3D medical images, and implements a low-cost, plug-and-play during inference semantic sampling method called MIS-SemSam on this benchmark, significantly improving the model's localization accuracy.

Medical Knowledge-Guided Fusion of Holistic Hospital Data for Tumor Survival Analysis

Guan, Jinquan (South China University of Technology), Xie, Yutong (South China University of Technology)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningDrug DiscoveryGraph Neural NetworkTransformerLarge Language ModelMixture of ExpertsContrastive LearningMultimodalityBiomedical DataElectronic Health Records

🎯 What it does: This paper proposes a medical knowledge-driven multi-modal fusion framework named MKGSurv, which integrates pan-hospital data from four disciplines—pathology, genetics, clinical, and treatment—to achieve tumor survival analysis.

MedObvious: Exposing the Medical Moravec’s Paradox in VLMs via Clinical Triage

Khan, Ufaq (Mohamed bin Zayed University of Artificial Intelligence), Khan, Muhammad Haris (Mohamed bin Zayed University of Artificial Intelligence)

Anomaly DetectionSafty and PrivacyExplainability and InterpretabilityPrompt EngineeringVision Language ModelImageMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasoundBenchmark

🎯 What it does: Built and evaluated a pre-diagnostic visual consistency checking benchmark called MedObvious for medical vision-language models, covering 5 hierarchical levels and 5 evaluation formats, and conducted zero-shot testing on 17 VLMs.

MedPro: Dual-View Propagation and Prototypical Assessment for Training-Free Few-Shot Medical Segmentation

Zhang, Jingyi (Nanjing University of Science and Technology), Xie, Guo-Sen (Nanjing University of Science and Technology)

SegmentationTransformerPrompt EngineeringDiffusion modelContrastive LearningBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose a training-free MedPro framework for few-shot medical image segmentation.

MedPruner: Training-Free Hierarchical Token Pruning for Efficient 3D Medical Image Understanding in Vision-Language Models

Liu, Shengyuan (Chinese University of Hong Kong), Yuan, Yixuan (Chinese University of Hong Kong)

Computational EfficiencyPrompt EngineeringVision Language ModelContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose a training-free, model-agnostic hierarchical token pruning framework called MedPruner, which is used to efficiently process the input of 3D medical images in vision-language models.

MedQ-Engine: A Closed-Loop Data Engine for Evolving MLLMs in Medical Image Quality Assessment

Liu, Jiyao (Fudan University), Xu, Ningsheng (Fudan University)

Data-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelContrastive LearningImageMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyRetrieval-Augmented Generation

🎯 What it does: Developed and implemented MedQ-Engine, a closed-loop data engine designed to iteratively enhance the performance of multimodal large language models on the medical image quality assessment (Med-IQA) task.

MedRAG-SCA: A Retrieval-Augmented Self-correcting Agent for Clinically Compliant Cross-Modality MRI Synthesis

Yuan, Feng (Beijing Institute of Technology), Gao, Xin (Beijing Institute of Technology)

GenerationData SynthesisConvolutional Neural NetworkImageBiomedical DataMagnetic Resonance ImagingComputed TomographyRetrieval-Augmented Generation

🎯 What it does: This paper studies the use of deep learning methods for diagnosing medical images

MedTri: Structured Medical Report Normalization for Enhanced Vision–Language Pretraining

Chu, Yuetan (King Abdullah University of Science and Technology), Gao, Xin (King Abdullah University of Science and Technology)

ClassificationFederated LearningExplainability and InterpretabilityComputational EfficiencyRepresentation LearningData-Centric LearningConvolutional Neural NetworkTransformerLarge Language ModelPrompt EngineeringVision Language ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageTextMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyElectronic Health Records

🎯 What it does: Propose the MedTri framework, which converts free-text medical reports into structured triplets [anatomical entity: radiological description + diagnostic category], to enhance the quality of text supervision in vision-language pre-training.

MedTriage-LM: Anatomically Grounded Visual Phenotype Synthesis for Interpretable ED Triage

Lu, Zhixiang (Xi'an Jiaotong-Liverpool University), Song, Sifan (Xi'an Jiaotong-Liverpool University)

ClassificationExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringVision Language ModelDiffusion modelTextMultimodalityTabularBiomedical DataElectronic Health RecordsRetrieval-Augmented Generation

🎯 What it does: This paper proposes MedTriage-LM, a multimodal large language model that integrates clinical tables, text, and synthetic visual phenotyping maps (VPM), for emergency triage instruction prediction and interpretable reasoning;

MedTriFlow: Efficient Resolution-Agnostic 3D Medical Image Generation with Implicit Triplane Representation

Xu, Chenfan (ShanghaiTech University), Cui, Zhiming (ShanghaiTech University)

RestorationGenerationTransformerDiffusion modelScore-based ModelFlow-based ModelRectified FlowNeural Radiance FieldAuto EncoderBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose the MedTriFlow framework, which combines tri-plane representation with implicit anatomical fields to achieve 3D medical image generation from a compressed latent space to arbitrary resolutions, and enables efficient inference on medical image generation and reconstruction tasks.

MedTS-TTT: Test-Time Training for Medical Time Series Classification

Chen, Mingzhi (Peking University), Luo, Guibo (Peking University)

ClassificationAnomaly DetectionComputational EfficiencyMeta LearningConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningTime SeriesBiomedical DataElectrocardiogramReview/Survey Paper

🎯 What it does: Proposed the MedTS-TTT framework for test-time training in the classification of medical time series (EEG, ECG).

MedTSC-Net: Bridging Semantic and Scale Gaps in Text-Guided Medical Image Segmentation

Chen, Zhaomin (Wenzhou University), Chen, Huiling (Hangzhou Dianzi University)

Image HarmonizationSegmentationConvolutional Neural NetworkTransformerVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose a text-guided medical image segmentation framework named MedTSC-Net, which is used for accurately segmenting lesions in COVID-19 chest imaging (X-ray and CT).

MedVol-R1: Reward-Driven Evidence Grounding for Volumetric Reasoning Segmentation

Wang, Zichun (Beihang University), Wang, Zihua (Bytedance Inc.)

SegmentationExplainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringVision Language ModelImageBiomedical DataComputed TomographyBenchmark

🎯 What it does: Propose a Volumetric Reasoning Segmentation framework MedVol-R1 based on reinforcement learning, which first locates key 2D slices and bounding boxes in 3D medical images as verifiable evidence, and then expands them into 3D semantic segmentation masks through a frozen MedSAM2 model.

MedXEdit-GRPO: Reasoning-Aware Online RL for Counterfactual Medical Image Generation

Pan, Yaning (Fudan University), Zhang, Xiaobo (Fudan University)

GenerationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringVision Language ModelDiffusion modelImageMultimodalityBiomedical DataComputed TomographyElectronic Health Records

🎯 What it does: Propose MedXEdit-GRPO, a counterfactual editing framework for medical images based on reinforcement learning, which achieves instruction compliance and anatomical fidelity by constructing a reward model and optimizing the generation strategy using Group Relative Policy Optimization.

Memory-Guided Random-Direction Feature Disentanglement Model for Multi-phase MRI Translation

Xiao, Qianmu (Central South University), Zou, Beiji (Manchester Metropolitan University)

Image TranslationRestorationSegmentationDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningOptical FlowImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Proposed a multi-phase CE-MRI translation framework (MG-RDFD) based on a shared VAE encoder and random direction feature disentanglement, achieving unified latent space separation of content and style, and realizing region-level style injection through a segmentation-guided memory pool.

MERIT: Multi-scale Mamba with Enhanced Information Retention for Treatment Response Prediction from Whole Slide Images

Hu, Taiyuan (Chinese Academy of Sciences), Yan, Rui (University of Science and Technology of China)

ClassificationImage TranslationDrug DiscoveryTransformerMixture of ExpertsContrastive LearningImageBiomedical DataComputed TomographyReview/Survey Paper

🎯 What it does: Designed and implemented a multi-scale Mamba framework called MERIT for predicting tumor treatment response from whole slide images (WSI).

Merlin Plus: A Large-Scale, Multi-cancer, Image-Mask-Report Dataset

Bassi, Pedro R. A. S. (Johns Hopkins University), Zhou, Zongwei (Johns Hopkins University)

SegmentationConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningImageTextBiomedical DataComputed Tomography

🎯 What it does: This study first constructed the Merlin Plus dataset, providing 1,153 voxel-level tumor masks covering nine organs (spleen, bladder, gallbladder, stomach, duodenum, uterus, prostate, adrenal gland, esophagus) and supplemented with longitudinal metadata such as patient identifiers and scan times; subsequently, segmentation models such as MedFormer and R-Super were trained using these masks, and were compared and evaluated with the original Merlin dataset and publicly available multi-tumor segmentation models (Voxtell, ULS, FLARE).

MetaFormer: Efficient Metadata-Guided Transformer for Patient-Aware 12-Lead Electrocardiogram Reconstruction

Xie, Yanchong (Sun Yat-sen University), Huang, Kai (Sun Yat-sen University)

RestorationTransformerAuto EncoderContrastive LearningTabularElectrocardiogram

🎯 What it does: This paper proposes MetaFormer, a Transformer architecture guided by patient metadata, which reconstructs single-lead ECG into complete 12-lead ECG.

MetaVFM: Meta-Learned Efficient Adaptation of Vision Foundation Models for Medical Imaging

Mecharbat, Lotfi Abdelkrim (Mohamed bin Zayed University of Artificial Intelligence), Yaqub, Mohammad (Mohamed bin Zayed University of Artificial Intelligence)

ClassificationMeta LearningTransformerSupervised Fine-TuningPrompt EngineeringContrastive LearningImageBiomedical Data

🎯 What it does: Propose MetaVFM, a meta-learning-based framework that achieves efficient adaptation for task-agnostic search through a pre-constructed Adapter-Dataset Meta-Space.

Microglia-RTS: Instance Segmentation of Microglia in Hypoxic Ischemic Fetal Sheep Brain Histology

Loomes, Callan (University of Auckland), Abbasi, Hamid (University of Auckland)

ClassificationObject DetectionSegmentationConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningImageBiomedical DataAlzheimer's DiseaseReview/Survey Paper

🎯 What it does: This study proposes the Microglia-RTS two-stage pipeline, achieving instance segmentation and morphological classification of Iba1-stained microglia cells in fetal sheep HIE tissue.

MicroscopyCLIP: A Domain-Specific Vision-Language Model for Optical Microscopy

Yang, Zhuoqin (Shenzhen University), Shen, Linlin (Shenzhen University)

ClassificationImage TranslationDomain AdaptationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningVision Language ModelContrastive LearningImageTextBiomedical DataBenchmark

🎯 What it does: This paper proposes and trains a domain-specific vision-language model called MicroscopyCLIP for optical microscopy images.

MiGAD: Gut Microbiota-informed Genetic Surrogates for Multimodal Alzheimer’s Disease Diagnosis

Liang, Zhichao (Macao Polytechnic University), Shen, Dinggang (ShanghaiTech University)

ClassificationKnowledge DistillationRepresentation LearningData-Centric LearningTransformerContrastive LearningImageMultimodalityBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease

🎯 What it does: This study proposes the MiGAD framework, which integrates gut microbiota-related gene information with brain MRI and clinical features for early multimodal diagnosis of Alzheimer's disease.

MIL-PF: Multiple Instance Learning on Precomputed Features for Mammography Classification

Jovišić, Nikola (Institute for AI R&D of Serbia), Ćulibrk, Dubravko (University of Novi Sad)

ClassificationAnomaly DetectionComputational EfficiencyRepresentation LearningConvolutional Neural NetworkTransformerContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound

🎯 What it does: Propose the MIL-PF framework, which freezes a large pre-trained model to precompute features, and then uses a lightweight multi-instance learning head to complete the classification task on breast X-ray images.

MINT: Molecularly Informed Training with Spatial Transcriptomics Supervision for Pathology Foundation Models

Lee, Minsoo (LG AI Research), Jang, Jongseong (LG AI Research)

Knowledge DistillationRepresentation LearningData-Centric LearningDrug DiscoveryTransformerSupervised Fine-TuningAuto EncoderContrastive LearningMultimodalityBiomedical DataBenchmark

🎯 What it does: This study proposes the MINT framework, which fine-tunes a pre-trained pathological ViT using spatial transcriptomics as a supervisory signal, thereby improving the model's performance in molecular and morphological features.

Missing Modality-Aware Calibration for Trustworthy Brain Tumor Segmentation

Lee, Sol (Korea Advanced Institute of Science and Technology), Yi, Mun Yong (Korea Advanced Institute of Science and Technology)

SegmentationExplainability and InterpretabilityConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningMultimodalityBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose a post-processing calibration method called MMA-LTS for missing modalities in brain tumor segmentation, which can perform spatially adaptive calibration of segmentation confidence under different modality combinations;

Missingness-Aware Multimodal Learning for Four-Class Adrenal Tumor Classification

Chen, Dehua (Donghua University), An, Huimin (Shanghai Jiao Tong University)

ClassificationData-Centric LearningConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningImageTabularBiomedical DataComputed TomographyElectronic Health Records

🎯 What it does: Designed and implemented a missing-aware multimodal learning framework for four-class classification of adrenal tumors (Cushing's syndrome, primary aldosteronism, pheochromocytoma, and non-functional adenoma) using non-contrast CT images and conventional non-hormonal clinical variables.

Mitigating Grade Boundary Conflicts for Disease Severity Grading with Ambiguous Labels

Liao, Zehui (Northwestern Polytechnical University), Xia, Yong (Northwestern Polytechnical University)

ClassificationAnomaly DetectionData-Centric LearningConvolutional Neural NetworkSupervised Fine-TuningContrastive LearningBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose an HDU framework, which enhances the robustness of disease severity grading models by harmonizing the boundaries of fuzzy labels generated by multiple annotators and performing unimodal projection.

MLLM-Enhanced Region-Aware Bidirectional Evidence-Based Model for Tongue Diagnosis

Du, Yiwei (Nanjing University), Shan, Caifeng (Nanjing University)

ClassificationExplainability and InterpretabilityRepresentation LearningConvolutional Neural NetworkTransformerLarge Language ModelMixture of ExpertsVision Language ModelContrastive LearningImageMultimodalityBiomedical Data

🎯 What it does: Propose a region-aware bidirectional evidence model integrated with a multimodal large language model for simultaneously predicting tongue coating syndrome patterns and visceral states.

MM-UNet: Meta Mamba UNet for Volumetric Medical Image Segmentation

Xie, Bin (Illinois Institute of Technology), Agam, Gady (Illinois Institute of Technology)

SegmentationConvolutional Neural NetworkMixture of ExpertsBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose a unified U-shaped structure MM-UNet, combining convolutional networks with state space models (Mamba) for volumetric medical image segmentation.

MMIC-EndoDepth: Self-Supervised Endoscopic Depth Estimation with Multi-Mechanism Illumination Correction

Xu, Ziang (Chinese University of Hong Kong), Wang, Bing (Hong Kong Polytechnic University)

Depth EstimationConvolutional Neural NetworkTransformerMixture of ExpertsAuto EncoderContrastive LearningVideoBiomedical DataBenchmark

🎯 What it does: Proposes the MMIC-EndoDepth framework, which estimates depth using self-supervised monocular endoscopic videos and improves reconstruction quality through multi-mechanism illumination correction.

MoCaf-Mamba: Modality Completion and Alignment in Feature Space for Missing-Modality Segmentation

Zhou, Yongsong (Shenzhen University), Shen, Linlin (Shenzhen University)

SegmentationConvolutional Neural NetworkTransformerDiffusion modelContrastive LearningMultimodalityBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Proposes the MoCaf-Mamba framework to address missing modalities and residual geometric mismatches in multi-modal medical image segmentation.

Modality-Agnostic Domain Generalizable Medical Image Segmentation via Self-Coordinated Focusing

Tang, Wentao (Sichuan University), Tan, Rui (Sichuan University)

SegmentationDomain AdaptationConvolutional Neural NetworkContrastive LearningImageMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasound

🎯 What it does: Proposes SCFNet, a multi-modal medical image segmentation model that is generalizable to unseen domains, achieving input-adaptive feature perception and information fusion through dynamic focusing convolution (DFC) and interactive attention fusion decoding (IAFD).

Modeling Clinical Workflow for SYNTAX Scoring from Coronary Angiography Videos

Fu, Suzhong (Chinese University of Hong Kong), Li, Zhen (Chinese University of Hong Kong)

ClassificationRecognitionImage TranslationSegmentationData SynthesisConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningGaussian SplattingOptical FlowImageVideoBiomedical DataComputed TomographyElectronic Health Records

🎯 What it does: Designed and implemented a hierarchical model based on preserving vascular segment identity for predicting SYNTAX scores from coronary angiography videos.

MORI-Seg: Learning Morphological Geometry for Instance Segmentation Without Instance Annotations

Zhao, Leiyue, Deng, Ruining (Vanderbilt University)

SegmentationConvolutional Neural NetworkTransformerDiffusion modelContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: In the absence of instance annotations, learn morphological geometric information through semantic segmentation results to achieve instance segmentation of kidney tissue;

MoSE: Mixture-of-Scale-Experts for Medical Image Restoration

Li, Xingyu (Hong Kong University of Science and Technology (Guangzhou)), Zhu, Lei (Hong Kong University of Science and Technology (Guangzhou))

RestorationSuper ResolutionConvolutional Neural NetworkTransformerMixture of ExpertsAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission Tomography

🎯 What it does: This paper proposes a medical image restoration framework based on multi-scale expert mixture (MoSE), which can adaptively handle various degradation patterns while preserving anatomical structural details.

Motion-Conditioned Multi-view Fusion for Myocardial Infarction Localization from Echocardiography

Yang, Guang (University of Oxford), Grau, Vicente (University of Oxford)

SegmentationAnomaly DetectionConvolutional Neural NetworkTransformerContrastive LearningGaussian SplattingOptical FlowVideoBiomedical DataUltrasound

🎯 What it does: Propose a dual-view echocardiography myocardial infarction localization framework named MCF-Net, which combines extremely sparse annotated motion priors with a pre-trained EchoPrime base model;

MPFlow: Multi-modal Posterior-Guided Flow Matching for Zero-Shot MRI Reconstruction

Kim, Seunghoi (University College London), Alexander, Daniel C. (University College London)

RestorationConvolutional Neural NetworkTransformerDiffusion modelFlow-based ModelRectified FlowContrastive LearningMultimodalityBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Proposed a zero-shot multi-modal MRI reconstruction framework called MPFlow, which reduces artifacts in reconstruction by guiding the reconstruction process with additional auxiliary modalities during inference, using a pre-trained unconditional flow model.

MRI2Rep: Autoregressive Structured Report Generation for 3D Liver MRI

Li, Xinran (Yale University), Staib, Lawrence H. (Yale University)

Image TranslationGenerationRepresentation LearningConvolutional Neural NetworkTransformerLarge Language ModelVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingRetrieval-Augmented Generation

🎯 What it does: Developed MRI2Rep, an end-to-end framework for automatically generating structured reports from 3D liver MRI, which directly predicts diagnostic sequences from multi-phase volumetric images and renders them into standardized reports using templates.

MRSeg3D: A Perturbation-Resilient Model for 3D Medical Reasoning Segmentation

Hao, Qin (Xinjiang University), Ye, Xujiong (University of Exeter)

SegmentationConvolutional Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelImageTextBiomedical DataComputed TomographyBenchmark

🎯 What it does: Proposed the MRSeg3D model and its three-level progressive perturbation benchmark to evaluate and enhance the robustness of 3D medical image reasoning segmentation under real language instructions.

MS-GDUN: Multi-scale Graph Deep Unfolding Network with Gradient-Threshold Learning for FMT Reconstruction

He, Xiaowei (Northwest University), Guo, Hongbo (Northwest University)

OptimizationGraph Neural NetworkDiffusion modelAuto EncoderContrastive LearningGraphBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Proposed a physics-guided multi-scale graph deep unrolling network (MS-GDUN) for inversion reconstruction in fluorescence molecular tomography (FMT).

MSHF-Net: Multimodal Breast Cancer Molecular Subtype Prediction via Segmentation-Guided Hierarchical Fusion Network

Jiang, Xinjie (Hangzhou Dianzi University), Wang, Changmiao (Shenzhen Research Institute of Big Data)

ClassificationSegmentationAnomaly DetectionConvolutional Neural NetworkTransformerContrastive LearningImageMultimodalityTabularBiomedical DataMagnetic Resonance ImagingUltrasoundElectronic Health Records

🎯 What it does: A multi-modal segmentation-guided hierarchical fusion network (MSHF-Net) was constructed for non-invasive pre-surgical molecular subtype prediction of breast cancer, integrating mammography (MG), ultrasound (US), and clinical information.

mSpecFusion-Net: Intelligent Multimodal Smartphone Imaging for Mobile Differential Diagnosis of Psoriasis and Dermatitis and Deep Learning

Kim, Sewoong (Daegu Gyeongbuk Institute of Science and Technology), Hwang, Jae Youn (Daegu Gyeongbuk Institute of Science and Technology)

ClassificationAnomaly DetectionConvolutional Neural NetworkTransformerImageMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasoundFibre Orientation DistributionDiffusion Tensor ImagingAlzheimer's DiseaseElectronic Health RecordsElectrocardiogramReview/Survey PaperBenchmarkAgriculture RelatedFinance RelatedPhysics Related

🎯 What it does: Proposed a smartphone-based multimodal optical imaging and deep learning framework for differential diagnosis of psoriasis and seborrheic dermatitis in dermatological clinical settings.

MSQG-3DNet: Multi-scale Token Selection with Query-Decoder and Gated Fusion for 3D Breast Tumor Classification

Liu, Zefeng (Macao Polytechnic University), Tan, Tao (Macao Polytechnic University)

ClassificationConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningBiomedical DataUltrasound

🎯 What it does: A model named MSQG‑3DNet is proposed for classifying the lesion-level ROI in 3D automated breast ultrasound (ABUS) as benign or malignant.

MSS-Net: Learning Hierarchical Multi-scale and Multimodal Representations for Complex Cardiac Disease Classification

Chen, Yuling (Central South University), Zeng, Feng (Central South University)

ClassificationConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningTextTime SeriesElectrocardiogram

🎯 What it does: Proposed the MSS-Net framework for multi-label ECG classification in complex cardiac diseases.

MuellerPT: Decomposition Driven Pre-training for Dense Learning in Mueller Polarimetry

Tlemsani, Adam (Imperial College London), Elson, Daniel S. (Imperial College London)

ClassificationSegmentationConvolutional Neural NetworkSupervised Fine-TuningContrastive LearningImageBiomedical DataPhysics Related

🎯 What it does: A pre-training method called MuellerPT was developed, which learns dense representations by predicting the Lu-Chipman decomposition parameters of the Mueller matrix, and was evaluated on tasks such as segmentation of gray-white brain tissue and classification of colorectal cancer.

Multi-agent Test-Time Adaptation for Robust Medical Image-to-Image Translation

Iele, Irene (Università Campus Bio-Medico di Roma), Tortora, Matteo (University of Genoa)

Image TranslationDomain AdaptationAgentic AIDiffusion modelFlow-based ModelAuto EncoderGenerative Adversarial NetworkImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Proposed a selection-based test-time adaptation framework based on multi-agent for image-to-image translation tasks in medical imaging.

Multi-Annotation Adaption: A Label-Informed Dynamic Framework for Medical Segmentation and Localization

Han, Luyi (Macao Polytechnic University), Mann, Ritse (Netherlands Cancer Institute)

SegmentationOptimizationHyperparameter SearchData-Centric LearningConvolutional Neural NetworkTransformerContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound

🎯 What it does: Propose a multi-annotation dynamic framework called MALFOY, which encodes different annotation types through label weight encoding, achieving adaptive learning for medical image segmentation and localization.

Multi-camera AR Guidance System for Surgical Instrument Handling and Assembly: Investigating Workload and Efficiency

Li, Shiyu (Technical University of Munich), Roth, Daniel (Technical University of Munich)

Data SynthesisPose EstimationComputational EfficiencyRobotic IntelligenceConvolutional Neural NetworkTransformerSupervised Fine-TuningDiffusion modelContrastive LearningGaussian SplattingSimultaneous Localization and MappingOptical FlowImageVideoPoint CloudMesh

🎯 What it does: This paper designs and implements an AR guidance system based on multi-camera markerless 6D pose estimation for the handling and assembly of surgical instruments in the operating room.

Multi-depth Rubidium PET Graph Neural Network for Major Adverse Cardiac Event Prediction

Chevalley, Arthur (Centre Hospitalier Universitaire Vaudois), Depeursinge, Adrien (Centre Hospitalier Universitaire Vaudois)

ClassificationAnomaly DetectionHyperparameter SearchConvolutional Neural NetworkGraph Neural NetworkContrastive LearningImageBiomedical DataPositron Emission Tomography

🎯 What it does: Developed a complete automated 82Rb PET processing pipeline, utilizing myocardial perfusion information from multiple depths (endocardial, mid-wall, epicardial) to construct a graph neural network for predicting major adverse cardiac events (MACE)

Multi-Disease Diagnosis in Retinal Images via Patch-Level Reasoning and Selective Aggregation

Xie, Jianyang (University of Liverpool), Zheng, Yalin (University of Liverpool)

ClassificationAnomaly DetectionExplainability and InterpretabilityGraph Neural NetworkTransformerContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasoundFibre Orientation DistributionDiffusion Tensor ImagingAlzheimer's DiseaseElectronic Health RecordsElectrocardiogram

🎯 What it does: Proposes a multi-disease retinal image diagnosis framework that combines patch-level reasoning with selective aggregation to address lesion-disease mismatch caused by comorbidities.

Multi-Expert Representation Learning with Dynamic Routing for Brain Captioning

Guo, Zhilin (University of Cambridge), Oztireli, Cengiz (University of Tokyo)

Representation LearningTransformerMixture of ExpertsContrastive LearningImageTextBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Proposed a multi-expert routing and progressive alignment framework, MERPA, for decoding brain activation signals into natural language descriptions (brain transcription)

Multi-frame Restoration for 10 Hz Lissajous Confocal Laser Endomicroscopy

Lee, Minhee (POSTECH), Lee, Jaeho (POSTECH)

RestorationSegmentationRetrievalConvolutional Neural NetworkRecurrent Neural NetworkTransformerDiffusion modelScore-based ModelFlow-based ModelRectified FlowAuto EncoderContrastive LearningOptical FlowVideoBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound

🎯 What it does: Proposed the CLE image restoration task for 10 Hz Lissajous scanning, constructed the first publicly available paired dataset MaLissa, and proposed a lightweight recursive restoration framework named MIRA to fill in missing pixels and reduce motion artifacts.

Multi-Kernel Gated Decoder Adapters for Robust Multi-Task Thyroid Ultrasound under Cross-Center Shift

Sabouri, Maziar (University of British Columbia), Rahmim, Arman (University of British Columbia)

ClassificationSegmentationDomain AdaptationConvolutional Neural NetworkTransformerContrastive LearningImageBiomedical DataUltrasound

🎯 What it does: This paper proposes a lightweight decoder-side adapter, MKGA/ResMKGA, to achieve robustness in multi-task thyroid ultrasound under cross-center domain shifts, incorporating multi-scale kernel refinement, context gating, and residual fusion, addressing negative transfer between segmentation and malignancy risk assessment.

Multi-parametric MRI for Contrast Agent Free Breast DCE-MRI Synthesis

Yang, Zhikai (KTH Royal Institute of Technology), Moreno, Rodrigo (KTH Royal Institute of Technology)

Image TranslationData SynthesisConvolutional Neural NetworkDiffusion modelGenerative Adversarial NetworkImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Built a model based on conditional generative adversarial networks to synthesize multi-phase images of breast DCE-MRI from non-contrast-enhanced multi-parametric MRI (T1w, DWI, and ADC).

Multi-probe Trait–State Latent Modeling for Cross-Context Clinical Motor Assessment

Yu, Jiahui (Zhejiang University), Xu, Xin (Zhejiang University)

GenerationData SynthesisPose EstimationRetrievalTransformerMixture of ExpertsDiffusion modelAuto EncoderContrastive LearningVideoBiomedical DataElectronic Health Records

🎯 What it does: A multi-probe feature-state latent model was studied, using standardized tasks (single-leg standing, sit-to-stand, knee lift) as probes to infer shared motor traits and generate resting-state gait from these traits, to test consistency across contexts.

Multi-regional CT Radiomics Fusion to Quantify Pathology-Derived Pollutant Burden in Lung Cancer

Li, Yutong (University of Texas MD Anderson Cancer Center), Wu, Chengyue (University of Texas MD Anderson Cancer Center)

ClassificationAnomaly DetectionMixture of ExpertsContrastive LearningImageBiomedical DataComputed Tomography

🎯 What it does: A non-invasive lung pollution load prediction model based on multi-region CT radiomics, named CT-LPI, was developed, utilizing pathology-derived LPI as a reference to train region-specific classifiers and perform fusion.

Multi-scale Dual-domain Fusion-Based Spatio-Temporal Graph Learning for fNIRS Signal Classification

Chu, Mengxiang (Northwest University), Guo, Hongbo (Northwest University)

ClassificationGraph Neural NetworkTransformerMixture of ExpertsContrastive LearningTime SeriesBiomedical Data

🎯 What it does: Propose a multi-scale dual-domain fusion spatiotemporal graph learning framework, MsDFSTGL, for multi-class classification of fNIRS signals.

Multi-Stage Dilated Convolution-Mamba Network for Surgical Action Triplet Recognition

Meng, Qingke (Jiangnan University), Zhou, Tao (Jiangnan University)

RecognitionConvolutional Neural NetworkGraph Neural NetworkTransformerMixture of ExpertsContrastive LearningVideoTabularBenchmark

🎯 What it does: Propose a multi-stage sparse convolution-Mamba network (DCM-Net) for end-to-end identification of instrument-verb-target triplets in surgical videos.

Multi-stage NeRF for Efficient 3D Coronary Artery Reconstruction from Two Narrow-Angle Angiographic Projections

Meng, Deyu (University of Oxford), Banerjee, Abhirup (University of Oxford)

OptimizationComputational EfficiencyDiffusion modelNeural Radiance FieldAuto EncoderContrastive LearningImageBiomedical DataComputed TomographyReview/Survey Paper

🎯 What it does: Reconstructs complete 3D vascular structures from two X-ray coronary angiography images with narrow viewing angles using a self-supervised multi-stage NeRF framework.

Multi-view Consistency-Based Adaptation of SAM2 for 3D ABUS Tumor Segmentation

Kim, Soopil (Daegu Gyeongbuk Institute of Science and Technology), Park, Sang Hyun (Pohang University of Science and Technology)

SegmentationDomain AdaptationTransformerSupervised Fine-TuningPrompt EngineeringContrastive LearningBiomedical DataUltrasound

🎯 What it does: This paper proposes a weakly supervised learning framework based on SAM2, achieving label-efficient adaptation for 3D ABUS tumor segmentation through point annotations and multi-view consistency regularization.

Multilayer Modularity-Aware Graph Autoencoder for Detecting Subpopulation Community Structure in Dynamic Brain Networks

See, Kai-Jun (Monash University Malaysia), Ting, Chee-Ming (Monash University Malaysia)

Anomaly DetectionExplainability and InterpretabilityComputational EfficiencyRepresentation LearningGraph Neural NetworkAuto EncoderContrastive LearningGraphBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose a multi-layer modular amplitude perception graph autoencoder (MMGAE-WSBM), which decouples individual differences from the dynamic community structure of brain networks by identifying potential subgroups and using a contrastive coupling mechanism within the autoencoder.

Multimodal Contrastive Regression for Organ-Resolved Biological Age Prediction

Ecker, Veronika (University of Stuttgart), Yang, Bin (University Hospital of Tuebingen)

Convolutional Neural NetworkTransformerContrastive LearningImageMultimodalityTabularBiomedical DataMagnetic Resonance Imaging

🎯 What it does: This paper proposes a multimodal contrastive regression framework that jointly uses MRI images and structured clinical variables to predict the biological age of various organs.

Multimodal Large Language Model-Driven Self-Verification Reasoning for Interpretable Diagnosis of Diffuse Cystic Lung Diseases

Jia, Qiwei (University of Science and Technology of China), Hu, Xiaowen (Central South University)

Anomaly DetectionExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringImageMultimodalityBiomedical DataComputed TomographyRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed a self-verification reasoning framework based on a multi-modal large language model, utilizing global spatial awareness and multi-step self-verification to achieve interpretable diagnosis of diffuse cystic lung disease.

Multimodal-Conditioned Flow Matching for Pathology Nuclei Data Augmentation

Zhang, Yanan (Beihang University), Bai, Xiangzhi (Beihang University)

SegmentationGenerationData SynthesisConvolutional Neural NetworkTransformerDiffusion modelFlow-based ModelContrastive LearningImageTextMultimodalityBiomedical DataOrdinary Differential Equation

🎯 What it does: A multi-modal conditional flow matching framework was developed to synthesize paired nuclear masks and pathological images for data augmentation.

Mutual Distillation of Dual-Foundation Models for Semi-supervised PET/CT Segmentation

Mao, Fuyou (Central South University), Tang, Yan (Central South University)

SegmentationKnowledge DistillationConvolutional Neural NetworkTransformerPrompt EngineeringContrastive LearningGaussian SplattingBiomedical DataComputed TomographyPositron Emission Tomography

🎯 What it does: This paper proposes a mutual distillation framework (MuDuo) based on a dual baseline model to achieve semi-supervised PET/CT organ segmentation.

Mutually Promoted Medical Image Segmentation and Classification via Prompt-Driven Interaction

Deng, Zijian (Sichuan University), Yang, Hui (Sichuan University)

ClassificationSegmentationTransformerPrompt EngineeringVision Language ModelContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasound

🎯 What it does: Propose the BiMI framework, which combines the Segment Anything Model with a hierarchical classifier to achieve mutual promotion between segmentation and classification.

MUX-USCT: A Noise-Robust Neural Network for Ultrasound Computed Tomography

Yuan, Yuchen (George Mason University), Yang, Lei (University of North Carolina at Chapel Hill)

RestorationTransformerAuto EncoderContrastive LearningBiomedical DataComputed TomographyUltrasound

🎯 What it does: Designed and implemented a deep learning architecture named MUX-USCT for ultrasound computed tomography (USCT) reconstruction under known acoustic acquisition geometry, achieving robustness to noise through adaptive MUX and attention mechanisms.