International Conference on Medical Image Computing and Computer-Assisted Intervention Β· 649 papers
ClinRAG-GRAPH: Clinical-Prior Retrieval-Augmented Graph Model with Domain Adversarial Learning for Breast pCR Prediction
Duan, Yaofei (Radboud University Medical Center), Mann, Ritse (Radboud University Medical Center)
CodeClassificationDomain AdaptationExplainability and InterpretabilityGraph Neural NetworkTransformerLarge Language ModelContrastive LearningImageTextGraphBiomedical DataMagnetic Resonance ImagingRetrieval-Augmented Generation
π― What it does: This study proposes a framework called ClinRAG-GRAPH, which constructs a clinical prior graph by utilizing dynamic enhanced magnetic resonance imaging (DCE-MRI), clinical variables, and biopsy pathological biomarkers, and achieves multi-modal feature fusion through graph convolutional networks. Furthermore, domain adversarial learning and large language model retrieval augmented generation (RAG) mechanisms are introduced to realize cross-center prediction of pathological complete response (pCR) during the pre-processing phase.
CMSA-Net: Causal Multi-scale Aggregation with Adaptive Multi-source Reference for Video Polyp Segmentation
Wang, Tong (Southeast University), Xie, Yutong (Mohamed bin Zayed University of Artificial Intelligence)
CodeSegmentationConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningVideoBiomedical Data
π― What it does: Propose a video polyp segmentation framework named CMSA-Net, combining causal multi-scale aggregation with dynamic multi-source reference, significantly enhancing the semantic discrimination and cross-frame consistency of low-contrast polyps.
π― What it does: To address the imbalance and domain differences in adult and pediatric brain tumor image data, a two-stage coarse-to-fine meta-reweighting framework is proposed. It first performs cross-domain binary segmentation for localization, and then refines the segmentation of pediatric-specific tumor sub-regions using gradient-aligned meta-reweighting.
π― What it does: Propose the CoGaze closed-loop eye-face alignment framework for eye tracking and cognitive screening on mobile devices for the elderly.
CodeClassificationAnomaly DetectionRepresentation LearningConvolutional Neural NetworkTransformerMixture of ExpertsContrastive LearningImageBiomedical DataAlzheimer's Disease
π― What it does: An automatic detection framework for age-related macular degeneration (AMD), named CoMRep, is proposed, which combines dual-branch feature extraction from global retinal images and local macular regions of interest (ROI), cross-scale attention aggregation, and Mixture-of-Experts (MoE) aggregation, to achieve multi-scale and cross-location lesion representation;
π― What it does: Proposed a two-stage collaborative topology and connectivity learning framework, leveraging EM-specific semantic priors such as neuronal skeletons and distance transforms, significantly improving the topological accuracy of EM neuron segmentation.
π― What it does: Proposed Community-aware Dynamic Graph Neural Network (CaDGNN) for brain disease diagnosis, combining delay-aware dynamic brain functional connectivity with community structure.
π― What it does: Propose Compass, a multi-view AI framework that integrates microwave ultrasound rotation scanning and biopsy images for prostate cancer risk assessment;
π― What it does: Designed and implemented a real-time, edge-side capsule endoscopy auditing system that can instantly identify anatomical regions and detect abnormalities during the capsule imaging process.
Concept-Guided Noisy Negative Suppression for Zero-Shot Classification and Grounding of Chest X-Ray Findings
Lian, Chenyu (Hong Kong Polytechnic University), Qin, Jing (Hong Kong Polytechnic University)
CodeClassificationObject DetectionRepresentation LearningTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataComputed TomographyElectronic Health Records
π― What it does: Propose the CoNNS framework, achieving chest X-ray vision-language alignment through concept-oriented noise negative sample suppression, thereby enhancing zero-shot classification and localization performance.
π― What it does: Propose the Conditional Diffusion Prompting (CDP) framework, which performs diffusion prompting in the dense prompt embedding space and regulates prompt uncertainty through image latent variables, achieving diverse and image-consistent segmentation of ambiguous boundaries in medical images.
π― What it does: Propose the CLS framework, which performs post-processing threshold calibration for 3D lesion segmentation to achieve statistical guarantees on the voxel-level missed detection rate.
π― What it does: Proposed a whole-brain connectivity-guided fMRI-to-image decoding framework called ConnecToMind2, which uses region-level embeddings and structural connectivity priors to achieve cross-subject image reconstruction.
ConstTrack: Constellation-Guided Cell Tracking Under Lineage Constraints
Xu, Yiwen (University of New South Wales), Meijering, Erik (University of New South Wales)
CodeObject TrackingConvolutional Neural NetworkTransformerContrastive LearningOptical FlowImageVideoBiomedical Data
π― What it does: This paper proposes a cell tracking and lineage reconstruction framework called ConstTrack, which is based on constellation features and can real-time predict cell identities and accurately assign cell division events in time-series microscopy images.
ContiStain: Cross-Domain Relation-Preserving Distillation for Continual Multi-domain Virtual IHC Staining
Chen, Fuqiang (Harbin Institute of Technology), Zhang, Yongbing (Tsinghua University)
CodeImage TranslationDomain AdaptationKnowledge DistillationRepresentation LearningTransformerMixture of ExpertsGenerative Adversarial NetworkContrastive LearningImageBiomedical Data
π― What it does: For the task of continuous multi-domain virtual IHC staining in medical images, this paper proposes a continual learning framework called ContiStain, which maintains performance on previously learned domains as new biomarker data is gradually received.
π― What it does: This paper proposes a contrast-invariant slice thickness estimation method called PRISM based on spectral matching, which matches the full voxel spectra of high-resolution reference images with low-resolution scanned images, thereby estimating 3D MRI slice thickness without the need for registration or pixel correspondence.
Contrastive and Adaptive Multi-modal Masked Autoencoder for Spatial Transcriptomics
Kim, Joohyeok (Yonsei University), Hwang, Seong Jae (Yonsei University)
CodeRepresentation LearningData-Centric LearningTransformerAuto EncoderContrastive LearningImageMultimodalityBiomedical Data
π― What it does: Propose CAMMST, a multi-modal framework based on masked autoencoders, to predict and interpolate the gene expression profiles of entire tissues using H&E images and a small number of gene expression anchors.
CodeClassificationGenerationData SynthesisTransformerSupervised Fine-TuningPrompt EngineeringDiffusion modelImageBiomedical Data
π― What it does: Propose the cgDDI framework, which utilizes three types of synthetic skin imagesβhealthy skin regeneration, non-parametric lesion mapping, and parametric semantic generationβto augment under-sampled datasets and improve the fairness and accuracy of malignant lesion classification.
π― What it does: Propose the CHIS framework, which utilizes untrained structural initialization and texture modulation control for pathological image generation;
ConVL: Interpretable Concept-Guided Vision-Language MIL for Survival Analysis in Whole Slide Images
Li, Junjian (Central South University), Wang, Jianxin (Central South University)
CodeClassificationExplainability and InterpretabilityConvolutional Neural NetworkTransformerPrompt EngineeringVision Language ModelContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
π― What it does: Proposed a interpretable concept-guided audio-visual language multi-instance learning framework, ConVL, for whole-slide image (WSI) survival analysis.
CoSim: Unleashing Eye Movements for EEG-Free Emotion Recognition via Conditional Prompting and Similarity-Guided Augmentation
An, Xiaoling (Hebei University of Technology), Hao, Xiaoke (Nanjing University of Aeronautics and Astronautics)
CodeRecognitionTransformerPrompt EngineeringGenerative Adversarial NetworkContrastive LearningMultimodalityTime Series
π― What it does: This paper proposes the CoSim framework, which enhances the performance of emotion recognition using only eye movement (EYE) signals by reconstructing pseudo EEG features from EYE through a conditional prompt generator (CPG), and by incorporating similar subjects' EEG-EYE data into training via similarity-guided knowledge augmentation (SKA).
π― What it does: Propose a classifier-free contrastive analysis framework that generates interpretable, high-quality visual counterfactual explanation (VCE) images by separating common and salient latent factors and refining them in the F-space of StyleGAN2.
CPAgents: Agentic Composite Phenotype Generation for Cardiac Disease Association
Li, Zuoou (University College London), Qiao, Mengyun (University College London)
CodeClassificationExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelAgentic AIPrompt EngineeringTabularBiomedical DataMagnetic Resonance ImagingElectrocardiogramRetrieval-Augmented GenerationChain-of-Thought
π― What it does: This study proposes the CPAgents framework, which enhances predictive performance in disease-wide association studies by automatically constructing and validating interpretable composite cardiac imaging phenotypes through an analytical cycle of analyze-propose-validate.
CPS4: Class Prompt Driven Semi-supervised Spine Segmentation with Class-Specific Consistency Constraint
Pan, Qingtao (Shandong University), Li, Shuo (Case Western Reserve University)
CodeSegmentationConvolutional Neural NetworkTransformerPrompt EngineeringVision Language ModelContrastive LearningImageBiomedical DataMagnetic Resonance Imaging
π― What it does: Developed the CPS4 framework, which utilizes class prompt-driven vision-language models to improve the quality of pseudo-labels in semi-supervised spinal segmentation, achieving high-precision segmentation through two-stage training.
CRIMSON: A Clinically-Grounded LLM-Based Metric for Generative Radiology Report Evaluation
Baharoon, Mohammed (Harvard Medical School), Rajpurkar, Pranav (Harvard Medical School)
CodeExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBiomedical DataComputed TomographyBenchmarkRetrieval-Augmented Generation
π― What it does: Proposed and implemented a clinically oriented LLM evaluation framework named CRIMSON, which measures the diagnostic accuracy, context relevance, and patient safety of chest X-ray report generation models, and generates interpretable scores through comprehensive error classification weighted by clinical severity.
Cross-Cancer Expert-Routing Knowledge Transfer in Federated Prognosis Prediction
Wang, Shu (Central South University), Wang, Jianxin (Central South University)
CodeFederated LearningKnowledge DistillationRepresentation LearningTransformerMixture of ExpertsContrastive LearningBiomedical Data
π― What it does: Propose the FedERF framework, which constructs a cross-cancer-type expert pool through federated low-rank decomposition and achieves dynamic routing on the target cancer type to improve prognosis prediction.
π― What it does: Developed a unified framework called UniVessel, which first generates vascular masks and bifurcation point labels by simulating vascular networks through modality-adapted vascular network simulation, and then uses a diffusion model that separates structure and style (via local phase bridging) to generate realistic fundus images that are consistent with the labels, achieving multi-modal zero-shot analysis (segmentation, registration, bifurcation detection).
Cross-Modal Concept Transfer: From ECG Signals to Images for Explainable Disease Prediction
Lee, Chan (Pusan National University), Kwon, Sunyoung (Pusan National University)
CodeClassificationAnomaly DetectionExplainability and InterpretabilityConvolutional Neural NetworkTransformerPrompt EngineeringAuto EncoderContrastive LearningImageMultimodalityTime SeriesElectronic Health RecordsElectrocardiogramBenchmark
π― What it does: Proposes a cross-modal concept transfer (CMCT) framework, using quantitative measurements of ECG signals as a concept bottleneck to guide image encoders in learning features related to electrocardiography, thereby enabling ECG disease prediction based on images.
π― What it does: Propose the StenCE framework, which aligns ECG representations with coronary artery X-ray angiography (Angio) representations through cross-modal contrastive learning, enabling models using only ECG to identify severe coronary artery stenosis.
CSAM-HQ: A Multi-stage Refinement Framework for Surgical Instrument Segmentation based on SAM and Probabilistic Graphical Models
Shi, Xueyi (Guilin University of Electronic Technology), Luo, Huoling (Shenzhen University of Information Technology)
CodeSegmentationConvolutional Neural NetworkTransformerPrompt EngineeringContrastive LearningImageBiomedical Data
π― What it does: To address issues such as ambiguous boundaries and topological fragmentation in the segmentation of minimally invasive surgical instruments, the CSAM-HQ framework is proposed, which incorporates HQ-Token and CRF for hierarchical refinement based on SAM.
π― What it does: Proposed a novel segmentation model called CSWinUNETR for thin and curved anatomical structures, integrating CSWin self-attention, detail-enhanced multi-scale self-attention, and sparsely controlled dynamic serpentine convolution, achieving high-precision 2D/3D segmentation.
π― What it does: Proposed a PET super-resolution method based on a CT-conditioned diffusion prior and physical constraint sampling, achieving inverse problem inference from low-quality PET to high-resolution PET.
π― What it does: Propose a discrete vocabulary-based autoregressive model called CTTok, which uses an LLM to generate CT voxel sequences and employs a GAN to deblock the patch reconstruction, achieving fast synthesis of text-to-3D CT volumes.
CURA: Calibrated Uncertainty with Retrieval and Agents for Trustworthy Multimodal Medical Decision Support
Zhang, Ruichen (University of North Carolina at Chapel Hill), Chen, Tianlong (University of North Carolina at Chapel Hill)
CodeRecommendation SystemSafty and PrivacyExplainability and InterpretabilityDrug DiscoveryTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsTextMultimodalityBiomedical DataElectronic Health RecordsBenchmarkRetrieval-Augmented Generation
π― What it does: Designed the CURA framework, integrating knowledge graph retrieval, heterogeneous clinical deliberation, Bayesian aggregation, and reliability overlay to provide calibrated uncertainty and actionable safety coverage for multimodal medical decision-making.
π― What it does: A low-calibration, low-computation data augmentation and domain adaptation framework for speech brain-computer interfaces is proposed, which uses neural cutting and splicing (NCS) to synthesize diverse EEG samples and achieves generalization to future unseen recording sessions through adversarial domain adaptation (ADAN).
DBT-Bleed: Dual-Branch Temporal Modeling with Key-Frame Selection for Surgical Bleeding Detection
Mishra, Sudhanshu (National University of Singapore), Jin, Yueming (National Neuroscience Institute)
CodeAnomaly DetectionRecurrent Neural NetworkTransformerPrompt EngineeringVision Language ModelContrastive LearningVideoMultimodality
π― What it does: This paper proposes a dual-branch multi-scale temporal modeling framework, DBT-Bleed, combined with hierarchy entropy-driven key frame selection (HiRED), for real-time detection of adverse bleeding events during surgery.
Li, Haoqing (University of Science and Technology of China), Hu, Xiaowen (University of Science and Technology of China)
CodeClassificationRetrievalAnomaly DetectionGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageTextGraphBiomedical DataComputed TomographyRetrieval-Augmented Generation
π― What it does: Proposed a multi-modal framework called DCLDs-RAG based on retrieval augmented generation for precise diagnosis of rare pulmonary cystic diseases.
π― What it does: Propose a dynamics-driven implicit neural representation framework called DD-INR for accelerated fMRI reconstruction, which can recover brain functional signals from time-varying sampled data;
π― What it does: Propose a two-stage decoupled 3D thoracic CT pulmonary lesion visual localization framework, first performing class-agnostic lesion segmentation, then performing text-volume alignment.
π― What it does: Propose a hierarchical Expectation-Maximization (HierEM) framework that models annotation noise in multi-center prostate lesion segmentation as observational errors of potential 'clean' lesion masks, and learns soft targets through the network.
Deep-IMMA: an Interpretable Immune-Aware Multimodal Deep Representation Learning for Cancer Prognosis and Recurrence Prediction
Jung, Kyeong Joo (Ohio State University), Machiraju, Raghu (Ohio State University)
CodeExplainability and InterpretabilityRepresentation LearningConvolutional Neural NetworkTransformerMixture of ExpertsVision Language ModelContrastive LearningImageMultimodalityBiomedical Data
π― What it does: Propose the Deep-IMMA two-stage multimodal framework, which achieves interpretable representations of immune-related features and predicts biochemical recurrence of prostate cancer through multitask learning on H&E images.
π― What it does: Developed DEER, a dual-branch contrastive pre-training model that leverages chest and abdominal digital radiographs (DR) and CT localizers (LOC) along with their corresponding reports to construct a large-scale self-supervised medical foundation model for chest and abdominal radiology diagnosis.
π― What it does: Propose a framework for ultrasound image quality assessment based on TinyUSFM, including the full-reference metric TinyUSFM-uLPIPS and the no-reference metric TinyUSFM-NRQ, used to quantify the diagnostic utility of ultrasound reconstructed and generated images.
Demographic-Conditioned State Space Models for ECG-Based Age Estimation
Bracke, Benjamin (University of Applied Sciences and Arts Dortmund), Friedrich, Christoph M. (University of Applied Sciences and Arts Dortmund)
CodeTransformerMixture of ExpertsTabularTime SeriesElectrocardiogram
π― What it does: Propose a multi-modal architecture that combines the Mamba2 state space model with dual-axis attention aggregation to predict cardiac biological age using 12-lead ECG and gender information.
DenseTRF: Texture-Aware Unsupervised Representation Adaptation for Surgical Scene Dense Prediction
Liao, Guiqiu (University of Pennsylvania), Hashimoto, Daniel A. (University of Pennsylvania)
CodeSegmentationDomain AdaptationTransformerMixture of ExpertsContrastive LearningVideoBiomedical Data
π― What it does: Proposed a texture-aware self-supervised distribution adaptation framework called DenseTRF based on slot attention, for dense prediction tasks in surgical videos;
DentMamba: An Anatomy-Aware Global-Local Hybrid Network for Large-Scale Multi-class Dental Disease Detection
Zhang, Kang (Shandong University), Zhou, Yuanfeng (Shandong University)
CodeObject DetectionConvolutional Neural NetworkTransformerMixture of ExpertsContrastive LearningImageBiomedical DataComputed TomographyReview/Survey Paper
π― What it does: Proposes the DentMamba two-stage dental disease detection framework, combining the anatomy-prior AnatoMamba module, MASAG bidirectional attention gating fusion, and dynamic scale-aware hybrid loss DSAHL, to achieve precise detection of multiple classes of lesions in panoramic X-ray images.
π― What it does: Proposes the DeNuC framework, which decouples nuclear detection from nuclear classification, allowing detection to be performed by a lightweight network and classification to be handled by a pathology foundation model, thereby improving the performance of nuclear detection and classification.
DepthPilot: From Controllability to Reliability in Colonoscopy Video Generation
Fu, Junhu (Fudan University), Li, Shuo
CodeGenerationDepth EstimationTransformerSupervised Fine-TuningDiffusion modelScore-based ModelAuto EncoderContrastive LearningImageVideoBiomedical Data
π― What it does: Proposed a reliable colonoscopy video generation framework called DepthPilot, which utilizes depth priors to impose geometric constraints on the generation process and enhances nonlinear spatiotemporal modeling capabilities through an adaptive spline denoising module.
CodeClassificationRecommendation SystemAnomaly DetectionExplainability and InterpretabilityTransformerLarge Language ModelAgentic AIVision Language ModelImageTextMultimodalityBiomedical DataBenchmarkRetrieval-Augmented GenerationChain-of-Thought
π― What it does: Propose DermAgent, a multi-tool collaborative dermatology image diagnosis agent system, which achieves step-by-step traceable reasoning through the Plan-Execute-Reflect framework.
CodeData SynthesisDepth EstimationTransformerSupervised Fine-TuningImageBiomedical Data
π― What it does: Achieve dense 3D reconstruction and surface normal estimation at a metric scale from single-view skin images, proposing the DermDepth model and the D-Synth synthetic dataset.
π― What it does: Proposed the Det-Y multi-center annotated dataset and the lightweight YOLO-Y network for fast and accurate detection of tuberculosis bacilli in sputum smears.
π― What it does: Proposes a training-time knowledge distillation framework called Detail Consistent Stage-Wise Distillation (DCD) for compressing 3D MRI segmentation models, reducing model size and inference latency while preserving fine structural details.
Deterministic Hallucination Detection in Medical VQA via Confidence-Evidence Bayesian Gain
Asadi, Mohammad (Stanford University), Adeli, Ehsan (Stanford University)
CodeAnomaly DetectionExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataBenchmark
π― What it does: Proposed a deterministic hallucination detection method called CEBaG, based on the model's own log-probability, specifically for the medical VQA task;
π― What it does: Propose a multi-resolution multi-instance learning framework called DiagMIL, which first performs a rough screening using low-resolution images, and only uses high-resolution images for fine-grained diagnosis when uncertain; meanwhile, the representation quality is enhanced through structured graph networks and cross-resolution alignment.
π― What it does: Propose a diagnostic evidence prototype learning framework based on Dirichlet Process (DPDEP), achieving multi-modal bone tumor subtyping, and resolving modal conflicts and missing data through evidence arbitration.
π― What it does: Proposes the DP-RGMI framework, analyzing the geometric impact of differential privacy on the representation space of medical imaging models;
π― What it does: Learn dense correspondences from unaligned and unlabeled CT scans by using intermediate features from a pre-trained diffusion model as anatomical priors.
π― What it does: This paper develops a digital twin framework of the lungs that takes a wearable single-lead ECG as input, extracts respiratory waveforms, and drives real-time animation of a 3D lung model, synchronously displaying lung expansion and contraction during patient breathing.
Dino U-Net: Exploiting High-Fidelity Dense Features from Foundation Models for Medical Image Segmentation
Li, Haoyue (University of Science and Technology of China), Gao, Xin (Chinese Academy of Sciences)
CodeSegmentationConvolutional Neural NetworkTransformerMixture of ExpertsAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingUltrasound
π― What it does: This paper proposes Dino U-Net, an encoder-decoder structure that leverages the high-fidelity dense features of the DINOv3 foundation model for medical image segmentation.
π― What it does: Propose DINO-3DRA, a dual-pathway framework that injects frozen 2D DINOv3 semantic features into 3D U-Net to achieve segmentation of cerebral aneurysms in 3DRA.
π― What it does: This paper proposes a two-stage progressive method called DINO-Med3D, which transfers the 2D self-supervised pretraining model DINOv3 to 3D medical segmentation tasks.
Directed Ordinal Diffusion Regularization for Progression-Aware Diabetic Retinopathy Grading
Chen, Huangwei (Zhejiang University), Wu, Lei (Zhejiang University)
CodeClassificationExplainability and InterpretabilityTransformerDiffusion modelContrastive LearningImageBiomedical Data
π― What it does: This paper proposes a diabetic retinopathy (DR) grading method called D-ODR based on directed diffusion regularization, which directly models the unidirectional progression of the disease in the feature space.
Discovering Heterogeneous Neurodegenerative Disease Patterns From MRI Data for Improved Prediction
Zhang, Yuanwang (University of Pennsylvania), Fan, Yong (University of Pennsylvania)
CodeClassificationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningMixture of ExpertsAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease
π― What it does: A hybrid expert (MoE) framework combining prediction and subtyping discovery is proposed for identifying diverse subtypes of neurodegenerative diseases from MRI data and improving the accuracy of progression prediction.
Disentangling Prompt Dependence to Evaluate Segmentation Reliability in Gynecological MRI
Germani, Elodie (University of Rennes), Baxter, John S. H. (University of Rennes)
CodeSegmentationExplainability and InterpretabilityTransformerPrompt EngineeringAuto EncoderBiomedical DataMagnetic Resonance Imaging
π― What it does: This paper proposes a framework to evaluate prompt dependency and measure segmentation reliability in prompt-based segmentation models for medical imaging.
π― What it does: Proposes Displacement Preserving Relational Distillation (DPRD), enhancing the performance of lightweight models in 3D medical image segmentation through ROI-aware feature masks and displacement vector-based relational distillation.
π― What it does: This paper proposes the SUMI method, which simulates degradation on high-quality PCCT images through clinical validation, and then learns an inverse degradation network in the latent space of a pre-trained CT autoencoder to enhance conventional EICT scans to the level of PCCT.
π― What it does: This paper proposes a learning framework that distills temporal consistency into a 2D network, achieving real-time prostate TRUS video segmentation.
π― What it does: Propose the DIVER-Surv model, which achieves cross-cohort survival prediction for brain gliomas by fusing diffusion-guided 3D mpMRI representations, LLM-encoded clinical text, and virtual node attention graphs.
π― What it does: Proposes the DNP-ConFormer framework, achieving unsupervised anomaly detection in medical images through a trainable encoder, momentum teacher, and a decoder guided by diverse normal prototypes.
π― What it does: Developed a semi-supervised Wavelet-KAN network called DK-Net to locate three key points (the two pubic symphysis and the fetal head tangent point) in obstetric transvaginal ultrasound images and calculate the angle of descent (AoP) of the fetal head.
π― What it does: Propose a 2D-3D ultrasound registration framework called DreamReg based on a world model, which achieves real-time pose estimation through continuous belief updates.
DRGFuse: Alignment-Guided Dual-Layer Reliability-Gated CTβWSI Fusion for Preoperative TRG0 Vs TRG1β3 Prediction in Esophageal Cancer
Liu, Zhenbing (Guilin University of Electronic Technology), Lu, Haoxiang (Guilin University of Electronic Technology)
CodeClassificationSegmentationExplainability and InterpretabilityComputational EfficiencyData-Centric LearningConvolutional Neural NetworkTransformerMixture of ExpertsDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageMultimodalityBiomedical DataComputed TomographyReview/Survey Paper
π― What it does: This paper proposes a dual reliability-gated fusion framework, DRGFuse, for predicting tumor regression grades TRG0 and TRG1-3 in esophageal cancer using preoperative CT and biopsy whole slide images (WSI).
π― What it does: Propose a voxel-level spherical harmonic regression network based on diffusion tensor prior (DTI-SHNet), achieving the synthesis of multi-shell diffusion MRI data from single-shell sampling;
Dual-Adaptive SAM3: Hierarchical Routing over Low-Rank Expert Layers for Parameter-Efficient Medical Image Segmentation
Chen, Ying, Liu, Yang (Nanyang Technological University)
CodeSegmentationExplainability and InterpretabilityComputational EfficiencyTransformerPrompt EngineeringMixture of ExpertsContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
π― What it does: Propose the Dual-Adaptive SAM3 framework, which inserts a hierarchical Mixture-of-Experts into the fusion module of SAM3, combining dynamic expert routing and low-rank expert decomposition to achieve efficient and interpretable adaptation for medical image segmentation.
π― What it does: This paper proposes DMC-Net, a dual-domain meta-conditional network, for segmenting fetal heads and the pubic symphysis in obstetric ultrasound images to achieve accurate progress angle estimation.
CodeDrug DiscoveryGraph Neural NetworkTransformerSupervised Fine-TuningContrastive LearningGraphSequentialBiomedical Data
π― What it does: Proposed a dual-modal neural network called MGK-DTI, which integrates sequence and structural information of drugs and targets for drug-target interaction prediction.
π― What it does: Proposes a framework utilizing zero-cost pseudo masks generated by SAM3 for dual-space active learning in the cold-start scenario of medical image segmentation.
π― What it does: Propose a dual-teacher knowledge distillation framework that combines an internal EMA teacher with an external SAM-Med3D teacher, using learnable prototypes to guide semi-supervised 3D medical segmentation.
π― What it does: Propose a dynamic collaborative continuous test-time adaptation framework, DyCo-CTA, for online updating of 3D vascular segmentation models under cross-center domain shift, reducing topological destruction and pseudo branches.
Dynamic Sub-domain Modeling for Robust Medical Image Segmentation
Lee, Kyungsu (Jeonbuk National University), Woo, Jonghye (Massachusetts General Hospital and Harvard Medical School)
CodeSegmentationConvolutional Neural NetworkTransformerPrompt EngineeringMixture of ExpertsContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasoundReview/Survey Paper
π― What it does: Achieve robustness in medical image segmentation without gradient updates through dynamic subdomain modeling, significantly improving boundary accuracy and internal variability;
π― What it does: Propose an MRI reconstruction framework called ECHO based on zero-shot self-supervised learning, which directly optimizes network parameters on a single undersampled acquisition, avoiding the training-inference mismatch in traditional methods.
Echo2ECG: Enhancing ECG Representations with Cardiac Morphology from Multi-view Echos
Liman, Michelle Espranita (Technical University of Munich), MΓΌller, Philip (Technical University of Munich)
CodeClassificationRetrievalRepresentation LearningTransformerContrastive LearningMultimodalityUltrasoundElectronic Health RecordsElectrocardiogram
π― What it does: Propose a multi-modal self-supervised framework called Echo2ECG, which aligns the cardiac morphology information from multi-view echocardiography (Echo) to electrocardiogram (ECG), thereby generating ECG representations rich in structural information for subsequent structural cardiac phenotyping and cross-modal retrieval.
π― What it does: This study proposes Echo4DIR, an implicit framework for real-time reconstruction of 4D cardiac geometry from sparse 2D ultrasound views.
π― What it does: Propose an EchoTracker2 model that tracks myocardial points using only the refinement phase, focusing on local motion to achieve pixel-level accurate tracking.
EdgeCVS: Democratization of surgical AI with a Distilled Edge-Deployable Critical View of Safety (CVS) model
Yamlahi, Amine (German Cancer Research Center), Maier-Hein, Lena (German Cancer Research Center)
CodeClassificationSegmentationComputational EfficiencyKnowledge DistillationRepresentation LearningConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningVideoBiomedical Data
π― What it does: The study proposes the EdgeCVS framework, which compresses a 305M parameter teacher model into a 5M parameter edge model through knowledge distillation, spatial priors, and stage-limited pseudo-labels, achieving real-time evaluation of critical safety views (CVS) during laparoscopic cholecystectomy.
EEGRFusion: Uncertainty-Aware EEG-to-Image Evidence for Adjunct Bedside Assessment of CMD in Disorders of Consciousness
Hong, Chenyuan (South China University of Technology), Xu, Yanwu (Pazhou Lab)
CodeRestorationGenerationRetrievalExplainability and InterpretabilityTransformerVision Language ModelDiffusion modelScore-based ModelRectified FlowContrastive LearningImageTime SeriesBiomedical DataElectrocardiogram
π― What it does: Propose the EEGRFusion framework, which retrieves and generates images consistent with the stimulus using EEG signals in the CLIP visual-semantic space, providing interpretable and uncertainty-aware visual evidence for bedside assessment of patients with disorders of consciousness.
π― What it does: Propose a CT reconstruction framework based on flow matching, FMCT, and its efficient version EFMCT, which reduces the number of network evaluations by reusing the velocity field, achieving reconstruction quality comparable to diffusion models in sparse-view CT reconstruction while significantly improving inference efficiency.
Eigen-Directed Random Projection for Medical Image Class Incremental Learning
Wu, Yongyi (Xi'an Jiaotong University), Wang, Hong (Xi'an Jiaotong University)
CodeClassificationComputational EfficiencyRepresentation LearningData-Centric LearningTransformerContrastive LearningImageBiomedical Data
π― What it does: Propose a training-agnostic, dual-view incremental learning framework named Eigen-Directed Random Projection (EDRP), which utilizes data-driven principal component projection instead of traditional random projection, significantly enhancing the discriminability of fine-grained medical image categories, and updates the classifier with closed-form Ridge regression on pre-trained features.
π― What it does: This paper proposes an event-level neural ordinary differential equation (ENC-ODE) model based on diagnostic conditions to predict the future evolution of multimodal brain biomarkers in Alzheimer's disease patients in continuous time.
CodeRecognitionSegmentationRetrievalAnomaly DetectionTransformerPrompt EngineeringVision Language ModelAuto EncoderContrastive LearningImageVideoTextBiomedical DataElectronic Health Records
π― What it does: Propose EndoVLM β a vision-language pre-training model specifically designed for gastrointestinal endoscopy, capable of aligning clinical reports with unordered image sets;
CodeData SynthesisAnomaly DetectionExplainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringVision Language ModelImageBiomedical DataMagnetic Resonance ImagingRetrieval-Augmented GenerationChain-of-Thought
π― What it does: Propose the BrReMark framework, which adopts a two-round 'annotation-reflection' interaction, enabling brain MRI abnormality detection, description, and diagnosis based on auditable ROI annotations, generating verifiable diagnostic chains.
Enhancing Pathological VLMs with Cross-scale Reasoning
Phan, Chi (National University of Singapore), Jin, Yueming (National University of Singapore)
CodeClassificationImage TranslationRestorationAnomaly DetectionExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsVision Language ModelContrastive LearningImageTextBiomedical DataElectronic Health RecordsBenchmark
π― What it does: Proposes a training and evaluation paradigm for cross-scale pathological vision-language models, constructs the Scale-VQA dataset, and trains the ScaleReasoner-R1 model capable of reasoning at multiple magnification levels.
Evidential Fusion Network for Multimodal Survival Prediction Under Missing Modalities
Xing, Yucheng (National University of Singapore), Feng, Mengling (National University of Singapore)
CodeTransformerMultimodalityBiomedical Data
π― What it does: Propose a multi-modal survival prediction model EMMS based on evidence fusion, which can still provide reliable predictions and uncertainty estimates when some modalities are missing.
Evidential Perfusion Physics-Informed Neural Networks with Residual Uncertainty Quantification
Lee, Junhyeok (Seoul National University), Choi, Kyu Sung (Seoul National University)
CodeAnomaly DetectionOptimizationConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelContrastive LearningImageBiomedical DataComputed TomographyPhysics Related
π― What it does: Propose an Evidence-based Physics-informed Neural Network (EPPINN) based on deep learning to solve the deconvolution problem in computed tomography perfusion (CTP) images of acute ischemic stroke, and provide confidence estimates for each voxel.
π― What it does: By modeling the retrieval process as a Markov Decision Process, a self-evolving retrieval agent is designed to dynamically delete, insert, and terminate retrieved cases to construct a reference set with higher diagnostic homogeneity, thereby improving the diagnostic accuracy for rare retinal diseases.
π― What it does: Propose an end-to-end framework based on internal geometry detection and graph neural networks for automatically detecting and segmenting cerebral aneurysms using 3D point clouds.
π― What it does: This paper proposes a verified tracking workflow that combines prompts generated during registration with verification or correction by radiologists, and then utilizes longitudinal information from baseline scans for lesion segmentation.