MICCAI 2026 Papers — Page 3
International Conference on Medical Image Computing and Computer-Assisted Intervention · 1165 papers
Contrastive and Adaptive Multi-modal Masked Autoencoder for Spatial Transcriptomics
Kim, Joohyeok (Yonsei University), Hwang, Seong Jae (Yonsei University)
Representation LearningData-Centric LearningTransformerAuto EncoderContrastive LearningImageMultimodalityBiomedical Data
🎯 What it does: Propose CAMMST, a multi-modal framework based on masked autoencoders, to predict and interpolate the gene expression profiles of entire tissues using H&E images and a small number of gene expression anchors.
Contrastive Learning under Noisy Temporal Self-Supervision for Colonoscopy Videos
Parolari, Luca (University of Padova), Le Folgoc, Loïc (Institut Polytechnique de Paris)
Object DetectionRetrievalRepresentation LearningConvolutional Neural NetworkTransformerContrastive LearningVideoBiomedical Data
🎯 What it does: Propose a self-supervised contrastive learning framework based on temporal order for learning robust representations of polyp trajectories (tracklets) in colonoscopy videos.
Controllable Generation of Diverse Dermatological Imagery for Fair and Efficient Malignancy Classification
Carrión, Héctor (University of California Santa Cruz), Norouzi, Narges (University of California Berkeley)
ClassificationGenerationData SynthesisTransformerSupervised Fine-TuningPrompt EngineeringDiffusion modelImageBiomedical Data
🎯 What it does: Propose the cgDDI framework, which utilizes three types of synthetic skin images—healthy skin regeneration, non-parametric lesion mapping, and parametric semantic generation—to augment under-sampled datasets and improve the fairness and accuracy of malignant lesion classification.
Controllable Histopathology Image Synthesis with Training-Free Structural Initialization and Textural Modulation
Qiu, Yuheng (Harbin Institute of Technology (Shenzhen)), Cao, Jianfeng (Harbin Institute of Technology (Shenzhen))
SegmentationGenerationData SynthesisConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelImageBiomedical DataMagnetic Resonance ImagingComputed TomographyReview/Survey Paper
🎯 What it does: Propose the CHIS framework, which utilizes untrained structural initialization and texture modulation control for pathological image generation;
ConVL: Interpretable Concept-Guided Vision-Language MIL for Survival Analysis in Whole Slide Images
Li, Junjian (Central South University), Wang, Jianxin (Central South University)
ClassificationExplainability and InterpretabilityConvolutional Neural NetworkTransformerPrompt EngineeringVision Language ModelContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Proposed a interpretable concept-guided audio-visual language multi-instance learning framework, ConVL, for whole-slide image (WSI) survival analysis.
CoRe-DA: Contrastive Regression for Unsupervised Domain Adaptation in Surgical Skill Assessment
Anastasiou, Dimitrios (University College London), Mazomenos, Evangelos B. (University College London)
Domain AdaptationConvolutional Neural NetworkTransformerContrastive LearningVideoBiomedical DataBenchmark
🎯 What it does: Propose an unsupervised domain adaptation framework called CoRe-DA based on contrastive regression, for video-based surgical skill assessment, and build the first SSA regression domain adaptation benchmark.
CortexAdapt3D: Parameter-Efficient Fine-Tuning of General 3D Foundation Models for Cortical Surface Analysis
Li, Kehan (Xi'an Jiaotong University), Lian, Chunfeng (Xi'an Jiaotong University)
ClassificationSegmentationDomain AdaptationTransformerSupervised Fine-TuningContrastive LearningMeshBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Proposes CortexAdapt3D, a parameter-efficient fine-tuning framework tailored for cortical surface analysis.
CoSim: Unleashing Eye Movements for EEG-Free Emotion Recognition via Conditional Prompting and Similarity-Guided Augmentation
An, Xiaoling (Hebei University of Technology), Hao, Xiaoke (Nanjing University of Aeronautics and Astronautics)
RecognitionTransformerPrompt EngineeringGenerative Adversarial NetworkContrastive LearningMultimodalityTime Series
🎯 What it does: This paper proposes the CoSim framework, which enhances the performance of emotion recognition using only eye movement (EYE) signals by reconstructing pseudo EEG features from EYE through a conditional prompt generator (CPG), and by incorporating similar subjects' EEG-EYE data into training via similarity-guided knowledge augmentation (SKA).
Counterfactual Anatomy-Guided Spatial-Temporal Decoding for Annotation-Free Hallucination Mitigation in Medical VLMs
Lu, Yifan (Mohamed bin Zayed University of Artificial Intelligence), Razzak, Imran (Mohamed bin Zayed University of Artificial Intelligence)
Anomaly DetectionExplainability and InterpretabilityTransformerPrompt EngineeringVision Language ModelDiffusion modelContrastive LearningImageMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyReview/Survey Paper
🎯 What it does: Propose a fully unlabeled, counterfactual anatomy-guided spatial-temporal decoding (CAST) framework that is only used during inference, to alleviate hallucination phenomena in medical vision-language models (Med-VLM) during question answering.
Counterfactual Contrastive Analysis
He, Yunlong (Télécom Paris), Gori, Pietro (Télécom Paris)
GenerationData SynthesisExplainability and InterpretabilityTransformerDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission Tomography
🎯 What it does: Propose a classifier-free contrastive analysis framework that generates interpretable, high-quality visual counterfactual explanation (VCE) images by separating common and salient latent factors and refining them in the F-space of StyleGAN2.
Counterfactual Stress-Testing for Fixing Lung Cancer Segmentation Models
Santhirasekaram, Ainkaran (Imperial College London), Chen, Mitchell (Imperial College London)
SegmentationExplainability and InterpretabilityData-Centric LearningConvolutional Neural NetworkTransformerFlow-based ModelGenerative Adversarial NetworkContrastive LearningBiomedical DataComputed TomographyStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: This paper proposes a causal counterfactual-based stress testing framework to evaluate and improve the robustness of lung cancer CT segmentation models; by intervening on clinically interpretable radiomic parent attributes, identity-preserving control samples and corresponding pseudo labels are generated, which are then used to plot Counterfactual Robustness Curves (CRC), and targeted fine-tuning is performed using the samples marked as 'hard' in the CRC;
CPAgents: Agentic Composite Phenotype Generation for Cardiac Disease Association
Li, Zuoou (University College London), Qiao, Mengyun (University College London)
ClassificationExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelAgentic AIPrompt EngineeringTabularBiomedical DataMagnetic Resonance ImagingElectrocardiogramRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This study proposes the CPAgents framework, which enhances predictive performance in disease-wide association studies by automatically constructing and validating interpretable composite cardiac imaging phenotypes through an analytical cycle of analyze-propose-validate.
CPR: Chained Perceptual Refinement for Coarse-to-Fine Medical Image Classification
Lu, Si-Yuan (Nanjing University of Posts and Telecommunications), Huang, Tianjin (University of Exeter)
ClassificationConvolutional Neural NetworkTransformerReinforcement LearningContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasound
🎯 What it does: This study proposes a Chained Perception Refinement (CPR) framework, which starts from a low-resolution global view, dynamically predicts and extracts high-resolution local regions, and gradually fuses global and local information to achieve high-resolution medical image classification.
CPS4: Class Prompt Driven Semi-supervised Spine Segmentation with Class-Specific Consistency Constraint
Pan, Qingtao (Shandong University), Li, Shuo (Case Western Reserve University)
SegmentationConvolutional Neural NetworkTransformerPrompt EngineeringVision Language ModelContrastive LearningImageBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Developed the CPS4 framework, which utilizes class prompt-driven vision-language models to improve the quality of pseudo-labels in semi-supervised spinal segmentation, achieving high-precision segmentation through two-stage training.
CRIMSON: A Clinically-Grounded LLM-Based Metric for Generative Radiology Report Evaluation
Baharoon, Mohammed (Harvard Medical School), Rajpurkar, Pranav (Harvard Medical School)
Explainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBiomedical DataComputed TomographyBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposed and implemented a clinically oriented LLM evaluation framework named CRIMSON, which measures the diagnostic accuracy, context relevance, and patient safety of chest X-ray report generation models, and generates interpretable scores through comprehensive error classification weighted by clinical severity.
Cross-Cancer Expert-Routing Knowledge Transfer in Federated Prognosis Prediction
Wang, Shu (Central South University), Wang, Jianxin (Central South University)
Federated LearningKnowledge DistillationRepresentation LearningTransformerMixture of ExpertsContrastive LearningBiomedical Data
🎯 What it does: Propose the FedERF framework, which constructs a cross-cancer-type expert pool through federated low-rank decomposition and achieves dynamic routing on the target cancer type to improve prognosis prediction.
Cross-channel Consistent Mamba for Extreme Limited-View Photoacoustic Tomography Reconstruction
Xu, Shicheng (University of Science and Technology of China), Gao, Fei (University of Science and Technology of China)
RestorationTransformerAuto EncoderContrastive LearningBiomedical DataComputed Tomography
🎯 What it does: Propose a shared single-channel Mamba network (S²-Mamba) that can handle missing channels in extreme limited-view photoacoustic tomography, and introduce cross-channel consistency loss and adaptive gradient gating loss for training.
Cross-domain Controllable Generation Enables Zero-shot Fundus Image Analysis
Li, Kaiwen (Peking University), Lu, Yanye (Peking University)
Image TranslationSegmentationGenerationDomain AdaptationConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataComputed TomographyReview/Survey Paper
🎯 What it does: Developed a unified framework called UniVessel, which first generates vascular masks and bifurcation point labels by simulating vascular networks through modality-adapted vascular network simulation, and then uses a diffusion model that separates structure and style (via local phase bridging) to generate realistic fundus images that are consistent with the labels, achieving multi-modal zero-shot analysis (segmentation, registration, bifurcation detection).
Cross-Modal Alignment and Distillation for Retinal Imaging Dementia Screening
Xu, Zhuoting (ShanghaiTech University), Zhao, Yitian (ShanghaiTech University)
ClassificationDomain AdaptationKnowledge DistillationRepresentation LearningConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningGaussian SplattingMultimodalityBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease
🎯 What it does: This study proposes the EyeDEM framework, which utilizes paired ocular OCTA and brain MRI data. By leveraging cross-modal latent space decoupling, hierarchical alignment, and block knowledge distillation, it transfers brain structural priors to an ocular network, enabling dementia screening using only ocular images.
Cross-Modal Concept Transfer: From ECG Signals to Images for Explainable Disease Prediction
Lee, Chan (Pusan National University), Kwon, Sunyoung (Pusan National University)
ClassificationAnomaly DetectionExplainability and InterpretabilityConvolutional Neural NetworkTransformerPrompt EngineeringAuto EncoderContrastive LearningImageMultimodalityTime SeriesElectronic Health RecordsElectrocardiogramBenchmark
🎯 What it does: Proposes a cross-modal concept transfer (CMCT) framework, using quantitative measurements of ECG signals as a concept bottleneck to guide image encoders in learning features related to electrocardiography, thereby enabling ECG disease prediction based on images.
Cross-Modal Contrastive Learning of ECG and Angiography Representations for Severe Stenosis Classification
Cenikj, Nikola (Technical University of Munich), Müller, Philip (Technical University of Munich)
ClassificationAnomaly DetectionRepresentation LearningTransformerSupervised Fine-TuningContrastive LearningMultimodalityBiomedical DataElectrocardiogram
🎯 What it does: Propose the StenCE framework, which aligns ECG representations with coronary artery X-ray angiography (Angio) representations through cross-modal contrastive learning, enabling models using only ECG to identify severe coronary artery stenosis.
Cross-Modal Prediction of DNA Methylation from Histology Images
Raza, Manahil (University of Warwick), Rajpoot, Nasir (University of Oxford)
ClassificationImage TranslationRepresentation LearningDrug DiscoveryGraph Neural NetworkTransformerSupervised Fine-TuningAuto EncoderContrastive LearningImageGraphTabularBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyElectronic Health Records
🎯 What it does: Propose GraphCpG, which utilizes graph neural networks to predict DNA methylation β values at 10,000 CpG sites from H&E pathological slides, and uses the predicted methylation signals for clinical downstream tasks such as survival stratification, proliferation subtypes, and immune subtypes.
Cross-Surface Representation Learning for Accurate Gingiva Prediction in CBCT
Mei, Lanzhuju (ShanghaiTech University), Shen, Dinggang (ShanghaiTech University)
SegmentationRepresentation LearningConvolutional Neural NetworkContrastive LearningPoint CloudMeshBiomedical DataComputed Tomography
🎯 What it does: A CBCT gingival margin prediction method combining clinical scanning protocols and cross-surface representations is proposed. It utilizes physical isolation to construct a high-contrast air interface, and then transforms the 3D segmentation task into a 1D density regression along the tooth surface normal to achieve precise soft tissue boundary localization.
CrownFusion: 3D Dental Crown Generation Using Geometry Images and Latent Diffusion
Ye, Johan Ziruo (Technical University of Denmark), Søndergaard, Peter Lempel (Simon Fraser University)
GenerationData SynthesisTransformerDiffusion modelAuto EncoderPoint CloudMeshBiomedical DataBenchmark
🎯 What it does: Proposes CrownFusion, a geometry image-based latent diffusion model for generating 3D crown designs.
CSAM-HQ: A Multi-stage Refinement Framework for Surgical Instrument Segmentation based on SAM and Probabilistic Graphical Models
Shi, Xueyi (Guilin University of Electronic Technology), Luo, Huoling (Shenzhen University of Information Technology)
SegmentationConvolutional Neural NetworkTransformerPrompt EngineeringContrastive LearningImageBiomedical Data
🎯 What it does: To address issues such as ambiguous boundaries and topological fragmentation in the segmentation of minimally invasive surgical instruments, the CSAM-HQ framework is proposed, which incorporates HQ-Token and CRF for hierarchical refinement based on SAM.
CSGE: Break the SSL Bottleneck in Medical Image Segmentation via Collaborative Semantic-Geometric Experts
Wan, Yufei (Taiyuan University of Technology), Wu, Yongfei (Taiyuan University of Technology)
SegmentationConvolutional Neural NetworkTransformerMixture of ExpertsVision Language ModelAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound
🎯 What it does: Propose CSGe — a medical image semi-supervised segmentation framework that utilizes the collaborative evolution of the BiomedCLIP semantic expert and the SAM geometric expert.
CSWinUNETR: Segmentation of Thin Anatomical Structures in Medical Images
Moon, Junho (Hanyang University), Jang, Ikbeom (Hankuk University of Foreign Studies)
SegmentationConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyBenchmark
🎯 What it does: Proposed a novel segmentation model called CSWinUNETR for thin and curved anatomical structures, integrating CSWin self-attention, detail-enhanced multi-scale self-attention, and sparsely controlled dynamic serpentine convolution, achieving high-precision 2D/3D segmentation.
CT-Conditioned Diffusion Prior with Physics-Constrained Sampling for PET Super-Resolution
Yang, Liutao (Imperial College London), Yang, Guang (Zurich University of Applied Sciences)
Super ResolutionTransformerDiffusion modelScore-based ModelBiomedical DataComputed TomographyPositron Emission Tomography
🎯 What it does: Proposed a PET super-resolution method based on a CT-conditioned diffusion prior and physical constraint sampling, achieving inverse problem inference from low-quality PET to high-resolution PET.
CT2-DiT: Enabling Abdominal Non-contrast to Contrast-Enhanced CT Translation with Diffusion Transformers in the Low-Data Regime
Zhu, Han (Ritsumeikan University), Nishikawa, Ikuko (Ritsumeikan University)
Image TranslationData SynthesisTransformerDiffusion modelFlow-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningBiomedical DataComputed Tomography
🎯 What it does: A framework named CT‑DiT based on Diffusion Transformer (DiT) for translating abdominal non-contrast CT (NCCT) to contrast-enhanced CT (CECT) was studied, with a focus on liver lesions.
CTTok: Voxel-Abulary for Autoregressive 3D CT Volume Generation with Large Language Models
Wang, Jiayi (FAU Erlangen Nurnberg), Kainz, Bernhard (FAU Erlangen Nurnberg)
GenerationData SynthesisTransformerLarge Language ModelDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageTextBiomedical DataComputed Tomography
🎯 What it does: Propose a discrete vocabulary-based autoregressive model called CTTok, which uses an LLM to generate CT voxel sequences and employs a GAN to deblock the patch reconstruction, achieving fast synthesis of text-to-3D CT volumes.
Cubic Structure-Aware Neural Radiance Fields for Zero-Shot Super-Resolution of Susceptibility-Weighted Imaging
Bi, Hongwei (Ningbo University), Wu, Xiping (Ningbo Medical Center Lihuili Hospital)
SegmentationSuper ResolutionTransformerNeural Radiance FieldAuto EncoderBiomedical DataMagnetic Resonance Imaging
🎯 What it does: A zero-shot cubic structure-aware NeRF framework, CSANeRF, is proposed to improve super-resolution reconstruction of SWI images.
CURA: Calibrated Uncertainty with Retrieval and Agents for Trustworthy Multimodal Medical Decision Support
Zhang, Ruichen (University of North Carolina at Chapel Hill), Chen, Tianlong (University of North Carolina at Chapel Hill)
Recommendation SystemSafty and PrivacyExplainability and InterpretabilityDrug DiscoveryTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsTextMultimodalityBiomedical DataElectronic Health RecordsBenchmarkRetrieval-Augmented Generation
🎯 What it does: Designed the CURA framework, integrating knowledge graph retrieval, heterogeneous clinical deliberation, Bayesian aggregation, and reliability overlay to provide calibrated uncertainty and actionable safety coverage for multimodal medical decision-making.
Curia-2: Scaling Self-Supervised Learning for Radiology Foundation Models
Saporta, Antoine (Raidium), Manceron, Pierre (Raidium)
Image TranslationSegmentationAnomaly DetectionRepresentation LearningTransformerVision Language ModelAuto EncoderContrastive LearningGaussian SplattingImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: This paper proposes the Curia-2 framework, which improves upon the original Curia and enhances the self-supervised pre-training performance for CT and MRI, while extending the ViT model to the billion-parameter level.
Cut, Stitch, Adapt: Data-Efficient and Low-Calibration Speech BCIs
Naskar, Animan (Indian Institute of Technology Ropar), Bathula, Deepti (Indian Institute of Technology Ropar)
Data SynthesisDomain AdaptationComputational EfficiencyRecurrent Neural NetworkGenerative Adversarial NetworkContrastive LearningBiomedical DataElectrocardiogramAudio
🎯 What it does: A low-calibration, low-computation data augmentation and domain adaptation framework for speech brain-computer interfaces is proposed, which uses neural cutting and splicing (NCS) to synthesize diverse EEG samples and achieves generalization to future unseen recording sessions through adversarial domain adaptation (ADAN).
CW-B: Class Weighted Boosting Framework for Imbalance Resilient Multi Class Cardiac Phenotyping
Li, Sijia (Shanghai University of Engineering Science), Qiu, Xihe (Shanghai University of Engineering Science)
ClassificationExplainability and InterpretabilitySupervised Fine-TuningTabularBiomedical DataElectronic Health Records
🎯 What it does: To address class imbalance and missing records in real clinical data, this paper proposes a class-weighted boosting framework based on XGBoost (CW-B), used for five-class heart discharge diagnosis phenotyping prediction.
Cycle-Verified Gentle Teaching and Confusion Correction for Label-Scarce Cross-Modal Single-Cell Annotation
Lu, Zheng (East China Normal University), Wang, Yan (East China Normal University)
ClassificationDomain AdaptationKnowledge DistillationRepresentation LearningTransformerGenerative Adversarial NetworkContrastive LearningMultimodalityBiomedical Data
🎯 What it does: This paper proposes a semi-supervised framework called CADENCE for cross-modal single-cell annotation with only about 1% of source labels, mainly through a teacher-tutor-student structure, cyclic consistency gating, and prototype confusion pair correction strategies to avoid decision boundary collapse caused by pseudo-label errors.
D2C: Data-Efficient Skin Lesion Classification for Clinical Images by Representation Alignment Using Dermoscopic Resources
Zhang, Zichen (Leiden University Medical Center), Dzyubachyk, Oleh (Leiden University Medical Center)
ClassificationDomain AdaptationKnowledge DistillationRepresentation LearningConvolutional Neural NetworkAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound
🎯 What it does: A semi-supervised domain adaptation framework (D2C) is proposed, which constructs a feature library based on dermoscopic images to transfer high-quality diagnostic knowledge to clinical images, achieving skin lesion classification on clinical images.
Data-Driven Optimization of Prostate Sectorization from MRI
Boccarato, Florencia (Université Côte d'Azur), Delingette, Hervé (Université Côte d'Azur)
OptimizationConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Propose a differentiable geometric framework that maps PI-RADS regions to 3D MRI voxels by learning a global consensus template, enabling adaptive segmentation of regions such as anterior-posterior, base-mid-apex.
DBT-Bleed: Dual-Branch Temporal Modeling with Key-Frame Selection for Surgical Bleeding Detection
Mishra, Sudhanshu (National University of Singapore), Jin, Yueming (National Neuroscience Institute)
Anomaly DetectionRecurrent Neural NetworkTransformerPrompt EngineeringVision Language ModelContrastive LearningVideoMultimodality
🎯 What it does: This paper proposes a dual-branch multi-scale temporal modeling framework, DBT-Bleed, combined with hierarchy entropy-driven key frame selection (HiRED), for real-time detection of adverse bleeding events during surgery.
DCLDs-RAG: Evidence-Grounded Multimodal Retrieval-Augmented Generation for Diffuse Cystic Lung Diseases Diagnosis
Li, Haoqing (University of Science and Technology of China), Hu, Xiaowen (University of Science and Technology of China)
ClassificationRetrievalAnomaly DetectionGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageTextGraphBiomedical DataComputed TomographyRetrieval-Augmented Generation
🎯 What it does: Proposed a multi-modal framework called DCLDs-RAG based on retrieval augmented generation for precise diagnosis of rare pulmonary cystic diseases.
DD-INR: Dynamics-Driven Implicit Neural Representation for Accelerated Whole-Brain Functional MRI Reconstruction
Li, Qiaoxin (Inria), Ciuciu, Philippe (Inria)
RestorationDiffusion modelScore-based ModelFlow-based ModelRectified FlowNeural Radiance FieldAuto EncoderBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Propose a dynamics-driven implicit neural representation framework called DD-INR for accelerated fMRI reconstruction, which can recover brain functional signals from time-varying sampled data;
DDecomp: Structured Uncertainty Decomposition Using Masked Autoencoders for Anomaly Detection in Diffusion MRI
Fu, ZhiJia (University of Electronic Science and Technology of China), Zhang, Fan (Harvard Medical School)
Anomaly DetectionTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningBiomedical DataMagnetic Resonance ImagingDiffusion Tensor Imaging
🎯 What it does: A framework for unsupervised anomaly detection called DDecomp based on Masked Autoencoder is constructed, which can automatically identify brain lesions in dMRI.
DEC-Net: A Multi-Task Distillation Expert Collaboration Network for Glioma Grading and Molecular Subtyping Using Multimodal MRI
Liu, Fei (Xinjiang University), Li, Hong-Dong (Xinjiang University)
ClassificationExplainability and InterpretabilityKnowledge DistillationRepresentation LearningConvolutional Neural NetworkMixture of ExpertsMultimodalityBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Proposed a multi-task network DEC-Net that combines knowledge distillation, expert collaboration, and dependency constraints for non-invasive multi-modal MRI glioma grading and molecular subtype prediction.
Decoding the Alzheimer’s Continuum: Interpretable Multi-gate Routing for Diagnosis and Transition Prediction
Jiang, Yufeng (Hong Kong Polytechnic University), Cai, Jing (Hong Kong Polytechnic University)
ClassificationDomain AdaptationExplainability and InterpretabilityTransformerMixture of ExpertsContrastive LearningImageBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease
🎯 What it does: Propose a unified multi-task multi-gated mixture-of-experts framework, M³AD, which uses single-modal T1-weighted structural MRI (sMRI) to achieve three-class diagnosis (NC/MCI/AD) and predict disease stage transitions (stable/progressive/reversal)
Decouple and Reason: Anatomically Guided Two-Stage Grounding of Lung Lesions in 3D Chest CT from Free-Text Reports
Uhm, Kwang-Hyun (Gachon University), Ko, Sung-Jea (Gachon University)
Image TranslationSegmentationAnomaly DetectionConvolutional Neural NetworkTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataComputed Tomography
🎯 What it does: Propose a two-stage decoupled 3D thoracic CT pulmonary lesion visual localization framework, first performing class-agnostic lesion segmentation, then performing text-volume alignment.
Decoupled Ordinal Refinement with Geometric Alignment for Chest X-ray Severity Scoring
Jeong, Geon (Inha University), Lee, Hyun Gyu (Inha University)
ClassificationAnomaly DetectionTransformerScore-based ModelContrastive LearningImageBiomedical DataComputed Tomography
🎯 What it does: Studied multi-region severity scoring in chest X-rays, proposing DORGA, which addresses issues such as class imbalance, ordinal label discontinuity, and annotator discrepancies by separating the severity axis from the structural subspace and performing geometric alignment.
Decoupled Single-Mask Annotation Noise Detection via Cross-Sectional Patch Self-consistency
Zhu, Yinheng (Tsinghua University), Xu, Xiaowei (Guangdong Provincial People's Hospital)
SegmentationAnomaly DetectionConvolutional Neural NetworkAuto EncoderContrastive LearningImageBiomedical DataComputed TomographyRetrieval-Augmented Generation
🎯 What it does: Propose a noise localization framework based on cross-sectional self-consistency detection, which can automatically identify annotation errors and generate auditable quality maps from vascular CT data with only a single annotation.
Deep EM with Hierarchical Latent Label Modelling for Multi-Site Prostate Lesion Segmentation
Yan, Wen (University College London), Barratt, Dean C. (University College London)
SegmentationConvolutional Neural NetworkAuto EncoderContrastive LearningBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Propose a hierarchical Expectation-Maximization (HierEM) framework that models annotation noise in multi-center prostate lesion segmentation as observational errors of potential 'clean' lesion masks, and learns soft targets through the network.
Deep Probabilistic Cerebrovascular Atlas
Zhang, Xiaoming (EURECOM), Zuluaga, Maria A. (EURECOM)
RestorationAnomaly DetectionMixture of ExpertsDiffusion modelScore-based ModelNeural Radiance FieldAuto EncoderContrastive LearningGaussian SplattingImageBiomedical DataMagnetic Resonance Imaging
🎯 What it does: VITAL, an implicit neural network based on dynamic Mixture-of-Experts, was constructed to learn population-level, age- and gender-conditioned probabilistic vascular maps in continuous space, and can be used for visualization and anomaly analysis of vascular structures.
Deep-IMMA: an Interpretable Immune-Aware Multimodal Deep Representation Learning for Cancer Prognosis and Recurrence Prediction
Jung, Kyeong Joo (Ohio State University), Machiraju, Raghu (Ohio State University)
Explainability and InterpretabilityRepresentation LearningConvolutional Neural NetworkTransformerMixture of ExpertsVision Language ModelContrastive LearningImageMultimodalityBiomedical Data
🎯 What it does: Propose the Deep-IMMA two-stage multimodal framework, which achieves interpretable representations of immune-related features and predicts biochemical recurrence of prostate cancer through multitask learning on H&E images.
DEER: A Foundation Model with Dual-Branch CT-Enhanced Embeddings for Thoracoabdominal Radiographs
Sun, Yihua (Shanghai Jiao Tong University), Liao, Hongen (Shanghai Jiao Tong University)
ClassificationRecognitionRepresentation LearningConvolutional Neural NetworkTransformerVision Language ModelContrastive LearningImageTextMultimodalityBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Developed DEER, a dual-branch contrastive pre-training model that leverages chest and abdominal digital radiographs (DR) and CT localizers (LOC) along with their corresponding reports to construct a large-scale self-supervised medical foundation model for chest and abdominal radiology diagnosis.
Defining Robust Ultrasound Quality Metrics via an Ultrasound Foundation Model
Huang, Ziyang (Fudan University), Wang, Yuanyuan (Fudan University)
RestorationSuper ResolutionAnomaly DetectionTransformerDiffusion modelAuto EncoderContrastive LearningImageBiomedical DataUltrasound
🎯 What it does: Propose a framework for ultrasound image quality assessment based on TinyUSFM, including the full-reference metric TinyUSFM-uLPIPS and the no-reference metric TinyUSFM-NRQ, used to quantify the diagnostic utility of ultrasound reconstructed and generated images.
DeformTTA: Morphology-Aware Controllable Deformation for Test-Time Adaptation in Fetal Cardiac Ultrasound Image Segmentation
Li, Zhihao (Wuhan University), Du, Bo (Wuhan University)
SegmentationDomain AdaptationConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataUltrasound
🎯 What it does: Propose DeformTTA, a test-time adaptation framework for fetal cardiac ultrasound images, which utilizes morphological priors and controllable deformations to achieve domain drift adaptive segmentation.
DeGenseGS: Geometrically and Semantically Decoupled Surgical Scene Understanding in 4D Gaussian Splatting
Wang, Yimo (Southeast University), Jin, Yueming (National University of Singapore)
SegmentationComputational EfficiencyRepresentation LearningTransformerVision Language ModelDiffusion modelScore-based ModelContrastive LearningGaussian SplattingOptical FlowImageVideoTextBiomedical Data
🎯 What it does: Propose the DeGenseGS framework, achieving the decoupling of geometric deformation and semantic evolution in 4D Gaussian Splatting, enabling real-time, text-prompted surgical scene understanding.
Demographic-Conditioned State Space Models for ECG-Based Age Estimation
Bracke, Benjamin (University of Applied Sciences and Arts Dortmund), Friedrich, Christoph M. (University of Applied Sciences and Arts Dortmund)
TransformerMixture of ExpertsTabularTime SeriesElectrocardiogram
🎯 What it does: Propose a multi-modal architecture that combines the Mamba2 state space model with dual-axis attention aggregation to predict cardiac biological age using 12-lead ECG and gender information.
DenseTRF: Texture-Aware Unsupervised Representation Adaptation for Surgical Scene Dense Prediction
Liao, Guiqiu (University of Pennsylvania), Hashimoto, Daniel A. (University of Pennsylvania)
SegmentationDomain AdaptationTransformerMixture of ExpertsContrastive LearningVideoBiomedical Data
🎯 What it does: Proposed a texture-aware self-supervised distribution adaptation framework called DenseTRF based on slot attention, for dense prediction tasks in surgical videos;
DentalPSAM: Lifting SAM to Dental Plaque Segmentation with Hybrid 2D-3D Knowledge Fusion
Jiang, Haoning (University of Hong Kong), Qu, Liangqiong (University of Hong Kong)
SegmentationConvolutional Neural NetworkGraph Neural NetworkTransformerPrompt EngineeringDiffusion modelAuto EncoderContrastive LearningImagePoint CloudMeshBenchmark
🎯 What it does: Proposed DentalPSAM, a dual-branch 2D-3D model that integrates geometric and appearance information for plaque segmentation in intraoral scanning (IOS), and released the first open-source 3D plaque dataset, PlaqueIOS.
DentMamba: An Anatomy-Aware Global-Local Hybrid Network for Large-Scale Multi-class Dental Disease Detection
Zhang, Kang (Shandong University), Zhou, Yuanfeng (Shandong University)
Object DetectionConvolutional Neural NetworkTransformerMixture of ExpertsContrastive LearningImageBiomedical DataComputed TomographyReview/Survey Paper
🎯 What it does: Proposes the DentMamba two-stage dental disease detection framework, combining the anatomy-prior AnatoMamba module, MASAG bidirectional attention gating fusion, and dynamic scale-aware hybrid loss DSAHL, to achieve precise detection of multiple classes of lesions in panoramic X-ray images.
DeNuC: Decoupling Nuclei Detection and Classification in Histopathology
Yang, Zijiang (University of Science and Technology Beijing), Fu, Dongmei (Beijing Engineering Research Center of Industrial Spectrum Imaging)
ClassificationObject DetectionConvolutional Neural NetworkTransformerSupervised Fine-TuningVision-Language-Action ModelContrastive LearningBiomedical DataBenchmark
🎯 What it does: Proposes the DeNuC framework, which decouples nuclear detection from nuclear classification, allowing detection to be performed by a lightweight network and classification to be handled by a pathology foundation model, thereby improving the performance of nuclear detection and classification.
DepthPilot: From Controllability to Reliability in Colonoscopy Video Generation
Fu, Junhu (Fudan University), Li, Shuo
GenerationDepth EstimationTransformerSupervised Fine-TuningDiffusion modelScore-based ModelAuto EncoderContrastive LearningImageVideoBiomedical Data
🎯 What it does: Proposed a reliable colonoscopy video generation framework called DepthPilot, which utilizes depth priors to impose geometric constraints on the generation process and enhances nonlinear spatiotemporal modeling capabilities through an adaptive spline denoising module.
DermaFlux: Synthetic Skin Lesion Generation with Rectified Flows for Enhanced Image Classification
Galanakis, Stathis, Zafeiriou, Stefanos (Imperial College London)
ClassificationGenerationData SynthesisTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringDiffusion modelScore-based ModelRectified FlowImageTextBiomedical DataOrdinary Differential Equation
🎯 What it does: This paper proposes DermaFlux, a text-conditioned skin lesion generation framework based on rectified flow, which generates skin lesion images using structured descriptions generated by Llama 3.2, and significantly improves the accuracy and AUC of binary classification (benign vs. malignant) models through data augmentation with the generated images.
DermAgent: A Multi-Tool Agentic System for Evidence-Grounded and Traceable Reasoning in Dermatology
Liu, Yize (Monash University), Ge, Zongyuan (Monash University)
ClassificationRecommendation SystemAnomaly DetectionExplainability and InterpretabilityTransformerLarge Language ModelAgentic AIVision Language ModelImageTextMultimodalityBiomedical DataBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose DermAgent, a multi-tool collaborative dermatology image diagnosis agent system, which achieves step-by-step traceable reasoning through the Plan-Execute-Reflect framework.
DermDepth: Toward Monocular Metric Scale 3D Reconstruction Models for Dermatology
Carrión, Héctor (University of California Santa Cruz), Norouzi, Narges (University of California Berkeley)
Data SynthesisDepth EstimationTransformerSupervised Fine-TuningImageBiomedical Data
🎯 What it does: Achieve dense 3D reconstruction and surface normal estimation at a metric scale from single-view skin images, proposing the DermDepth model and the D-Synth synthetic dataset.
Det-Y: A Multi-center Dataset and Benchmark for Efficient Detection of Mycobacterium Tuberculosis in Sputum Smears
Xie, Zijian (Guangxi Medical University), Li, Yuexiang (Guangxi Medical University)
Object DetectionConvolutional Neural NetworkImageBiomedical DataBenchmark
🎯 What it does: Proposed the Det-Y multi-center annotated dataset and the lightweight YOLO-Y network for fast and accurate detection of tuberculosis bacilli in sputum smears.
Detail Consistent Stage-Wise Distillation for Efficient 3D MRI Segmentation
Fan, Mengchen (University of Alabama at Birmingham), Lan, Qizhen (UTHealth Houston)
SegmentationComputational EfficiencyKnowledge DistillationConvolutional Neural NetworkMultimodalityBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Proposes a training-time knowledge distillation framework called Detail Consistent Stage-Wise Distillation (DCD) for compressing 3D MRI segmentation models, reducing model size and inference latency while preserving fine structural details.
Detecting Clinical Hallucinations in LVLMs via Counterfactual Visual Grounding Uncertainty
Song, Xiao (Nanjing University), Shan, Caifeng (Nanjing University)
Anomaly DetectionExplainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelImageTextMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose a pluggable visual hallucination detection framework for LVLMs, which utilizes an external visual localizer to visually verify entities generated by LVLMs, and estimates visual uncertainty by applying counterfactual perturbations to entities, thereby determining hallucinations.
Deterministic Hallucination Detection in Medical VQA via Confidence-Evidence Bayesian Gain
Asadi, Mohammad (Stanford University), Adeli, Ehsan (Stanford University)
Anomaly DetectionExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataBenchmark
🎯 What it does: Proposed a deterministic hallucination detection method called CEBaG, based on the model's own log-probability, specifically for the medical VQA task;
Device-Constrained Real-Time EVD Catheter Segmentation
Seddiqi, Mustafa (Concordia University), Kersten-Oertel, Marta (Concordia University)
SegmentationData SynthesisComputational EfficiencyConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningGaussian SplattingImageVideoPoint CloudMeshMagnetic Resonance ImagingComputed TomographyUltrasound
🎯 What it does: This paper proposes a framework that can achieve real-time EVD catheter segmentation on mobile and XR devices, specifically designed for segmenting weak low-contrast structures;
DFuse: A Dual-Branch Adaptive Fusion Network for Multi-Channel MCG Denoising
Huang, Shuai, Jiang, Tianzi (Chinese Academy Of Sciences)
RestorationConvolutional Neural NetworkTransformerScore-based ModelAuto EncoderContrastive LearningTime SeriesBiomedical DataElectrocardiogramBenchmark
🎯 What it does: Proposed the DFuse dual-branch adaptive fusion network for denoising multi-channel TMR-based MCG.
DGFMIL: Fuzzy Multi-instance Learning with Cross-Granularity Calibration and Soft Semantic Routing for Histopathological Image Staging
Cai, Meiling (Taiyuan University of Technology), Qiang, Yan (North University of China)
ClassificationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingComputed TomographyReview/Survey Paper
🎯 What it does: Propose the DGFMIL framework to address uncertainty and multi-scale feature issues in whole slide image staging, achieving uncertainty awareness and semantic prior alignment through dual branches.
Di-PACT: Discrete Prognosis via Aligned Co-Topology of Pathways and Tumor Microenvironment
Cai, Zhiheng (Ocean University of China), Zhang, Shugang (Ocean University of China)
ClassificationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningDrug DiscoveryConvolutional Neural NetworkSpiking Neural NetworkTransformerMixture of ExpertsDiffusion modelScore-based ModelAuto EncoderContrastive LearningImageMultimodalityTabularBiomedical DataComputed TomographyAlzheimer's DiseaseElectronic Health RecordsReview/Survey PaperBenchmark
🎯 What it does: Proposed a bidirectional co-topology based multimodal discrete survival prediction framework, Di-PACT, which integrates whole-slide images and pathway-level transcriptomic data to achieve precise prognosis;
DiagMIL: Uncertainty-Gated Coarse-to-Fine Multi-resolution MIL for WSI Diagnosis
Jeong, Jiwon (Korea Advanced Institute of Science and Technology), Yi, Mun Yong (Korea Advanced Institute of Science and Technology)
ClassificationAnomaly DetectionComputational EfficiencyKnowledge DistillationRepresentation LearningGraph Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningGaussian SplattingImageBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasoundReview/Survey Paper
🎯 What it does: Propose a multi-resolution multi-instance learning framework called DiagMIL, which first performs a rough screening using low-resolution images, and only uses high-resolution images for fine-grained diagnosis when uncertain; meanwhile, the representation quality is enhanced through structured graph networks and cross-resolution alignment.
Diagnose Like a Doctor: Misdiagnosis-Aware Structured Memory for Continual Learning of Medical Vision-Language Models
Qin, Chendong (Shanghai Jiao Tong University), Yao, Yifei (Shanghai Jiao Tong University)
ClassificationExplainability and InterpretabilityKnowledge DistillationRepresentation LearningGraph Neural NetworkTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasoundFibre Orientation DistributionDiffusion Tensor ImagingRetrieval-Augmented Generation
🎯 What it does: Designed and implemented a misdiagnosis-aware hierarchical pitfall graph (Hierarchical Pitfall Graph) and retrieval-augmented generation (RAG) framework, combining dynamic retrieval from an offline multimodal database to address fine-grained feature confusion and catastrophic forgetting in medical vision-language models during continual learning.
Diagnosing 3D Volumes from Camera Video: A Semantic-Robust Framework Without Direct DICOM Access
Wu, Junlei (ShanghaiTech University), Wang, Qian (ShanghaiTech University)
ClassificationSegmentationAnomaly DetectionConvolutional Neural NetworkTransformerSupervised Fine-TuningMixture of ExpertsContrastive LearningVideoBiomedical DataComputed Tomography
🎯 What it does: Proposed a framework (ScreenCAD) that enables 3D medical image diagnosis directly from screen videos recorded by a camera without the need for a PACS interface.
Diagnostic Evidence-Based Prototype Learning via Dirichlet Process for Multimodal Bone Tumor Subtyping
Wu, Feng (University of Hong Kong), Yu, Lequan (University of Hong Kong)
ClassificationExplainability and InterpretabilityKnowledge DistillationRepresentation LearningConvolutional Neural NetworkTransformerContrastive LearningMultimodalityBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Propose a diagnostic evidence prototype learning framework based on Dirichlet Process (DPDEP), achieving multi-modal bone tumor subtyping, and resolving modal conflicts and missing data through evidence arbitration.
DiCoP-SAM: Disagreement-Band Guided Corrective Prompting for Collaborative Semi-Supervised Medical Image Segmentation
Wang, Xinyu (Fuzhou University), Lian, Sheng (Xiamen University)
SegmentationKnowledge DistillationData-Centric LearningTransformerPrompt EngineeringContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound
🎯 What it does: This paper proposes a two-stage collaborative framework, DiCoP-SAM, to improve the performance of medical image segmentation under the condition of limited labeled data.
Diff-SABR: Diffusion-Based Symmetry-Aware Reconstruction of Alveolar Bone Defects from CBCT
Choi, Ho Yoon (Seoul National University), Yi, Won-Jin (Seoul National University)
RestorationGenerationData SynthesisDomain AdaptationTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningImageBiomedical DataComputed Tomography
🎯 What it does: Developed a symmetry-aware framework based on diffusion models, named Diff-SABR, for automatically reconstructing alveolar bone defects in CBCT images.
Differential Privacy Representation Geometry for Medical Image Analysis
Tayebi Arasteh, Soroosh (RWTH Aachen University), Truhn, Daniel (RWTH Aachen University)
ClassificationSafty and PrivacyRepresentation LearningConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningImageBiomedical DataComputed TomographyPositron Emission TomographyUltrasound
🎯 What it does: Proposes the DP-RGMI framework, analyzing the geometric impact of differential privacy on the representation space of medical imaging models;
Diffusion-Guided Anatomical Position Encoding for Dense Longitudinal CT Correspondence
Zhuang, Mingrui (Central Hospital of Dalian University of Technology), Wang, Hongkai (Dalian University of Technology)
Image TranslationSegmentationConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningBiomedical DataComputed Tomography
🎯 What it does: Learn dense correspondences from unaligned and unlabeled CT scans by using intermediate features from a pre-trained diffusion model as anatomical priors.
Digital Twin of the Lung from Wearable Biosignals for Real-Time Respiratory Monitoring
Acharya, Partha (Indian Institute of Technology Kharagpur), Chakraborty, Suman (Indian Institute of Technology Kharagpur)
Image TranslationData SynthesisAnomaly DetectionComputational EfficiencyConvolutional Neural NetworkRecurrent Neural NetworkDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningOptical FlowMeshTime SeriesBiomedical DataComputed TomographyElectrocardiogram
🎯 What it does: This paper develops a digital twin framework of the lungs that takes a wearable single-lead ECG as input, extracts respiratory waveforms, and drives real-time animation of a 3D lung model, synchronously displaying lung expansion and contraction during patient breathing.
Dino U-Net: Exploiting High-Fidelity Dense Features from Foundation Models for Medical Image Segmentation
Li, Haoyue (University of Science and Technology of China), Gao, Xin (Chinese Academy of Sciences)
SegmentationConvolutional Neural NetworkTransformerMixture of ExpertsAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingUltrasound
🎯 What it does: This paper proposes Dino U-Net, an encoder-decoder structure that leverages the high-fidelity dense features of the DINOv3 foundation model for medical image segmentation.
DINO-3DRA: Leveraging 2D Foundation Model Semantics for 3D Cerebral Aneurysm Segmentation
Lu, Jiayang (University of Manchester), Sarrami-Foroushani, Ali (University of Manchester)
SegmentationConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Propose DINO-3DRA, a dual-pathway framework that injects frozen 2D DINOv3 semantic features into 3D U-Net to achieve segmentation of cerebral aneurysms in 3DRA.
DINO-Med3D: Bridging Dimension and Domain Gaps in Volumetric Segmentation via Progressive Adaptation
Hu, Haoyu (University of Chinese Academy of Sciences), Hou, Zeng-Guang (Institute of Automation Chinese Academy of Sciences)
SegmentationDomain AdaptationConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: This paper proposes a two-stage progressive method called DINO-Med3D, which transfers the 2D self-supervised pretraining model DINOv3 to 3D medical segmentation tasks.
Directed Ordinal Diffusion Regularization for Progression-Aware Diabetic Retinopathy Grading
Chen, Huangwei (Zhejiang University), Wu, Lei (Zhejiang University)
ClassificationExplainability and InterpretabilityTransformerDiffusion modelContrastive LearningImageBiomedical Data
🎯 What it does: This paper proposes a diabetic retinopathy (DR) grading method called D-ODR based on directed diffusion regularization, which directly models the unidirectional progression of the disease in the feature space.
Discovering Heterogeneous Neurodegenerative Disease Patterns From MRI Data for Improved Prediction
Zhang, Yuanwang (University of Pennsylvania), Fan, Yong (University of Pennsylvania)
ClassificationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningMixture of ExpertsAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease
🎯 What it does: A hybrid expert (MoE) framework combining prediction and subtyping discovery is proposed for identifying diverse subtypes of neurodegenerative diseases from MRI data and improving the accuracy of progression prediction.
Discriminative Directional Projection for Reverse Distillation in Breast Ultrasound Anomaly Detection
Xu, Shicheng (University of Science and Technology of China), Gao, Fei (University of Science and Technology of China)
Anomaly DetectionKnowledge DistillationConvolutional Neural NetworkContrastive LearningImageBiomedical DataUltrasound
🎯 What it does: This paper addresses abnormal detection in breast ultrasound by proposing a Discriminative Directional Projection Loss, improving the reverse distillation method to achieve more accurate pixel-level abnormal localization.
Discriminative Synergistic Temporal Diffusion for Ultrasound Video Thyroid Nodule Segmentation
Li, Jialu (Hong Kong University of Science and Technology (Guangzhou)), Zhu, Lei (Shenzhen Institute of Advanced Technology)
SegmentationTransformerDiffusion modelVideoBiomedical DataUltrasound
🎯 What it does: Propose a Discriminative Synergistic Temporal Diffusion (DSTD) framework for ultrasound video thyroid nodule segmentation, which can dynamically evaluate the contributions of motion and texture guidance during the denoising process.
Disease-Consistent Visual-Text Alignment and Fusion for Reliable Radiology Report Generation
Shao, Jianjun (Beijing Normal University), Tian, Yun (Beijing Normal University)
GenerationRetrievalRepresentation LearningTransformerVision Language ModelContrastive LearningImageTextBiomedical DataComputed TomographyElectronic Health RecordsRetrieval-Augmented Generation
🎯 What it does: Propose a disease-consistent visual-textual alignment and fusion framework (DCAF) for generating reliable radiology reports
Disentangling Clinic-Specific Acquisition Variability from Motion Dynamics in Intrapartum Ultrasound
Aytutuldu, Ilhan (Gebze Technical University), Akgul, Yusuf Sinan (Erzurum City Hospital)
Domain AdaptationRepresentation LearningTransformerAuto EncoderGenerative Adversarial NetworkContrastive LearningOptical FlowVideoTime SeriesBiomedical DataUltrasound
🎯 What it does: Propose a self-supervised framework that disentangles center-specific acquisition variations from physiological motion dynamics in multi-center prenatal ultrasound videos through temporal regularization, adversarial disentanglement, and hierarchical cross-attention.
Disentangling Prompt Dependence to Evaluate Segmentation Reliability in Gynecological MRI
Germani, Elodie (University of Rennes), Baxter, John S. H. (University of Rennes)
SegmentationExplainability and InterpretabilityTransformerPrompt EngineeringAuto EncoderBiomedical DataMagnetic Resonance Imaging
🎯 What it does: This paper proposes a framework to evaluate prompt dependency and measure segmentation reliability in prompt-based segmentation models for medical imaging.
Disentangling Uterine and Bladder Activity in Paired EHG-MRI Data Using Anatomy-Guided Multicluster Beamforming
Bustos-Vivas, Maria Camila (Uniklinikum Erlangen), Hutter, Jana (Uniklinikum Erlangen)
Diffusion modelScore-based ModelContrastive LearningOptical FlowImageMultimodalityBiomedical DataMagnetic Resonance ImagingStochastic Differential EquationOrdinary Differential EquationAudio
🎯 What it does: This paper proposes an anatomy-guided multi-cluster LCMV beamforming method, which utilizes simultaneously collected EHG and MRI data to separate the electrical activities of the uterus and bladder.
Displacement Preserving Relational Distillation for Robust Medical Segmentation
Ding, Zhicheng (Bowling Green State University), Lan, Qizhen (University of Houston - Clear Lake)
SegmentationKnowledge DistillationConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Proposes Displacement Preserving Relational Distillation (DPRD), enhancing the performance of lightweight models in 3D medical image segmentation through ROI-aware feature masks and displacement vector-based relational distillation.
Distilling Photon-Counting CT Into Routine Chest CT Through Clinically Validated Degradation Modeling
Liu, Junqi (Johns Hopkins University), Zhou, Zongwei (Johns Hopkins University)
RestorationSuper ResolutionKnowledge DistillationTransformerDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataComputed Tomography
🎯 What it does: This paper proposes the SUMI method, which simulates degradation on high-quality PCCT images through clinical validation, and then learns an inverse degradation network in the latent space of a pre-trained CT autoencoder to enhance conventional EICT scans to the level of PCCT.
Distilling Temporal Coherence Into 2D Networks for Transrectal Ultrasound Prostate Video Segmentation
Kim, Dong Yeong (Seoul National University), Kim, Young-Gon (Seoul National University Hospital)
SegmentationKnowledge DistillationConvolutional Neural NetworkContrastive LearningOptical FlowVideoBiomedical DataUltrasound
🎯 What it does: This paper proposes a learning framework that distills temporal consistency into a 2D network, achieving real-time prostate TRUS video segmentation.
DIVER-Surv: Diffusion and Virtual-Node Graph Fusion for Cross-Cohort Glioma Survival
Cho, Minji (Sungkyunkwan University), Park, Hyunjin (Sungkyunkwan University)
ClassificationDomain AdaptationRepresentation LearningConvolutional Neural NetworkGraph Neural NetworkTransformerLarge Language ModelVision-Language-Action ModelDiffusion modelContrastive LearningImageTextBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Propose the DIVER-Surv model, which achieves cross-cohort survival prediction for brain gliomas by fusing diffusion-guided 3D mpMRI representations, LLM-encoded clinical text, and virtual node attention graphs.
Diverse Normal Prototypes-Guided Contrastive Reconstruction for Medical Anomaly Detection
Li, Luhu, Fu, Shujun (Shandong University)
Domain AdaptationAnomaly DetectionTransformerAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyBenchmark
🎯 What it does: Proposes the DNP-ConFormer framework, achieving unsupervised anomaly detection in medical images through a trainable encoder, momentum teacher, and a decoder guided by diverse normal prototypes.
DK-Net: Semi-supervised Wavelet-KAN Landmark Localization for Angle of Progression Measurement in Intrapartum Transperineal Ultrasound
Deng, Bo, Li, Shuo (Jinan University)
Image TranslationRestorationPose EstimationConvolutional Neural NetworkTransformerSupervised Fine-TuningAuto EncoderContrastive LearningImageBiomedical DataUltrasound
🎯 What it does: Developed a semi-supervised Wavelet-KAN network called DK-Net to locate three key points (the two pubic symphysis and the fetal head tangent point) in obstetric transvaginal ultrasound images and calculate the angle of descent (AoP) of the fetal head.
DLCE-MIL: Depth-aware Local Context Enhanced Multiple Instance Learning for Patient-Level IVCM Fungal Keratitis Classification
Men, Xiaoyan (Shandong Normal University), Chen, Yiqiang (Institute of Computing Technology Chinese Academy of Sciences)
ClassificationConvolutional Neural NetworkTransformerContrastive LearningImageBiomedical DataComputed TomographyReview/Survey Paper
🎯 What it does: Propose the DLCE-MIL framework for patient-level subtype classification of fungal keratitis (Fusarium vs. Aspergillus) based on IVCM images;
Do Not Break the Vessels: Structure-Preserving Mean Flow for Vascular Image Translation
Sun, Changjin (Southeast University), Zhou, Guangquan (Suzhou Microclear Medical Instruments Co.)
Image TranslationRestorationDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningOptical FlowImageBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound
🎯 What it does: A Structural Preserving Mean Flow (SPMF) framework is proposed, which utilizes orthogonal constraints and prototype-guided structural refinement to achieve cross-modal translation from NIRII to 2PF vascular images, ensuring the vascular topology remains invariant during the translation process.
DosePlug: Observation-Driven Dose Modeling for Robust Low-Dose Reconstruction and Segmentation
Ki, Do Hyun (Chonnam National University), Yoo, Seok Bong (Chonnam National University)
RestorationSegmentationConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelFlow-based ModelAuto EncoderContrastive LearningImageBiomedical DataComputed TomographyPositron Emission Tomography
🎯 What it does: Proposes DosePlug, a lightweight observation-driven dose estimation module, which can achieve adaptive reconstruction and segmentation of low-dose CT/PET images without dose metadata.