MICCAI 2026 Papers — Page 12
International Conference on Medical Image Computing and Computer-Assisted Intervention · 1165 papers
Uncertainty-Aware Hypergraph Consistency Learning for Semi-supervised Medical Image Segmentation
Xie, Kexin (Guangdong University of Technology), Luo, Guibo (Peking University)
SegmentationGraph Neural NetworkAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound
🎯 What it does: Propose a semi-supervised medical image segmentation framework (UAHC), which enhances segmentation performance under limited annotation by combining uncertainty guidance with hypergraph consistency learning.
Uncertainty-Aware Spatio-Semantic Contextual Prompts for Multimodal Medical Segmentation
Chattopadhyay, Soumitri (University of California San Diego), Niethammer, Marc (University of California San Diego)
SegmentationConvolutional Neural NetworkTransformerPrompt EngineeringContrastive LearningMultimodalityBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Propose a training-agnostic multi-modal 3D medical image segmentation framework that generates dense prompts by fusing spatial and semantic context to achieve cross-modal, low-sample segmentation.
Uncertainty-Aware Spatiotemporal Graph Learning for Individualized Cortical Parcellation and Somato-Cognitive Integration in the Developing Brain
Xue, Yuxin (University of Electronic Science and Technology of China), Zhang, Yu (Shanghai Jiao Tong University)
SegmentationExplainability and InterpretabilityRepresentation LearningRecurrent Neural NetworkGraph Neural NetworkContrastive LearningGraphTime SeriesBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Utilize an uncertainty-aware spatiotemporal graph learning framework to perform individualized cortical parcellation on adolescent natural experimental movie-watching fMRI data;
Uncertainty-Guided Conservative Propagation for Coronary Artery Segmentation
Huang, Huan (Kennesaw State University), Zhao, Chen (Kennesaw State University)
SegmentationConvolutional Neural NetworkTransformerDiffusion modelContrastive LearningBiomedical DataComputed Tomography
🎯 What it does: Propose an uncertainty-guided conservative propagation framework (UGCP), treating coronary artery segmentation inference as a state evolution process with limited steps, exchanging neighborhood information in the logit space to enhance structural consistency.
Uncovering Interpretable Point-Wise Image Correspondences from Deep Segmentation Features
Duval, Eli (KU Leuven), Maes, Frederik (KU Leuven)
SegmentationExplainability and InterpretabilityConvolutional Neural NetworkContrastive LearningOptical FlowImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Propose an untrained framework that estimates point-to-point correspondences between different medical images by leveraging multi-scale features from pre-trained segmentation networks.
Understanding Model Behavior in Monocular Polyp Sizing
Xiong, Xinqi (University of North Carolina at Chapel Hill), Sengupta, Roni (University of California San Diego)
ClassificationSegmentationDepth EstimationExplainability and InterpretabilityConvolutional Neural NetworkTransformerPrompt EngineeringContrastive LearningImageBiomedical DataComputed TomographyReview/Survey Paper
🎯 What it does: Conducting a diagnostic audit on the classification of polyp size from monocular colonoscopy, systematically evaluating the performance of multiple models, input modalities, and cross-center datasets
Uni-Brain: Anatomy-Informed Diffusion for Multi-Tracer Synthesis and Counterfactual Prognosis from Cross-Sectional Data
Yu, Hongjie (ShanghaiTech University), Shen, Dinggang (ShanghaiTech University)
GenerationData SynthesisTransformerDiffusion modelScore-based ModelAuto EncoderImageBiomedical DataMagnetic Resonance ImagingPositron Emission TomographyAlzheimer's Disease
🎯 What it does: This paper proposes Uni-Brain, a unified diffusion framework that generates multi-tracer PET images and performs causal prediction based on MRI-derived anatomical conditions and clinical conditions.
UniCT: A Unified Joint Multi-Task Framework for 3D Chest CT Abnormality Analysis
Di Piazza, Theo (INSA Lyon), Boussel, Loic (Philips Clinical Informatics)
ClassificationSegmentationGenerationConvolutional Neural NetworkTransformerVision Language ModelContrastive LearningImageBiomedical DataComputed TomographyRetrieval-Augmented Generation
🎯 What it does: Propose UniCT, a unified end-to-end multi-task framework that jointly performs multi-label abnormal classification, segmentation, and report generation for 3D chest CT.
UniFCHarm: Unified Prompt-Guided Harmonization of fMRI Functional Connectivity Across Scanners and Atlases
Lin, Xin (ShanghaiTech University), Shen, Dinggang (ShanghaiTech University)
Image HarmonizationDomain AdaptationRepresentation LearningGraph Neural NetworkTransformerPrompt EngineeringDiffusion modelAuto EncoderContrastive LearningGraphBiomedical DataMagnetic Resonance Imaging
🎯 What it does: A unified prompt-guided fMRI functional connectivity harmonization method, UniFCHarm, was studied, which can achieve functional connectivity harmonization under multi-scanner and multi-spatial-template conditions.
Unified Multimodal Model for Brain MRI Imputation and Understanding
Song, Zhiyun (Imperial College London), Bai, Wenjia (Imperial College London)
RecognitionImage TranslationRestorationAnomaly DetectionTransformerMixture of ExpertsVision Language ModelFlow-based ModelAuto EncoderImageBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Propose a unified multimodal model, UniBrain, for imputing and diagnosing missing modalities in brain MRI
UniField: A Unified Field-Aware MRI Enhancement Framework
Lin, Yiyang (Chinese University of Hong Kong), Yuan, Yixuan (Chinese University of Hong Kong)
RestorationSuper ResolutionTransformerDiffusion modelAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Propose UniField, a unified MRI field strength enhancement framework that integrates multi-modal (T1, T2, FLAIR) and cross-field enhancement tasks (64mT→3T, 3T→7T), and leverages a pre-trained 3D video super-resolution prior to achieve structural prior knowledge;
UniLiver: Gradient-Conditioned Unified Model for Multi-target Hepatic Segmentation
Qiu, Yue (Chinese University of Hong Kong), Fu, Chi-Wing (Chinese University of Hong Kong)
SegmentationConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningBiomedical DataMagnetic Resonance ImagingComputed TomographyBenchmark
🎯 What it does: This paper proposes UniLiver, a unified multi-task liver segmentation model that can simultaneously segment liver vasculature, Couinaud segments, and tumors.
UniSpine-GS: An Efficient Physics-Aware Gaussian Framework for Cross-Modality Multi-view Spine Image Synthesis
Chen, Qiuhua (Wuhan University), Du, Bo (Wuhan University)
GenerationData SynthesisDiffusion modelNeural Radiance FieldGaussian SplattingImageMultimodalityBiomedical DataComputed TomographyUltrasound
🎯 What it does: Propose a physics-aware 3D Gaussian model called UniSpine-GS for efficient synthesis of cross-modal multi-view spinal images.
Unleashing Video Language Models for Fine-grained HRCT Report Generation
Fang, Yingying (Imperial College London), Yang, Guang (Imperial College London)
GenerationAnomaly DetectionOptimizationRepresentation LearningData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelImageTextBiomedical DataComputed TomographyChain-of-Thought
🎯 What it does: This paper studies the migration of general video-language models to high-resolution CT (HRCT) report generation, and proposes the AbSteering framework, enabling the model to first identify abnormalities before generating reports;
Unlocking the Power of Medical Tabular Data via Semantic-Aware Multimodal Pre-training
Liu, Yingsheng (Monash University), Yu, Zhen (University of Queensland)
ClassificationAnomaly DetectionRepresentation LearningData-Centric LearningConvolutional Neural NetworkTransformerPrompt EngineeringMixture of ExpertsVision Language ModelAuto EncoderContrastive LearningImageMultimodalityTabularBiomedical DataElectronic Health Records
🎯 What it does: Proposed a semantic-aware multimodal pre-training framework named AID, specifically designed to model the two-dimensional hierarchical structure of medical tabular data, significantly enhancing the robustness and generalization ability of image-table joint representations.
Unsupervised and Segmentation-Free Phase Unwrapping for 4D Flow MRI Using Deep Image Prior
Ren, Yuyang (ShanghaiTech University), Hu, Peng (ShanghaiTech University)
RestorationConvolutional Neural NetworkDiffusion modelOptical FlowBiomedical DataMagnetic Resonance Imaging
🎯 What it does: This paper proposes an unsupervised, manual segmentation-free active aortic 4D flow MRI phase unwrapping method called PUDIP-Flow.
Unsupervised Brain Anomaly Detection as a Bayesian Inverse Problem with Diffusion Prior
Roy, Hugues (Sorbonne Université), Burgos, Ninon (Sorbonne Université)
Anomaly DetectionTransformerDiffusion modelScore-based ModelAuto EncoderBiomedical DataMagnetic Resonance ImagingPositron Emission TomographyAlzheimer's DiseaseStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: This paper models unsupervised brain anomaly detection as a Bayesian inverse problem with a diffusion prior, and introduces a latent anomaly mask to jointly infer pseudo-healthy images and anomalous regions.
Unsupervised Domain Adaptation for Pelvic Landmark Localization in CT and MR Using Landmark-Conditioned Synthesis
Ye, Kangqing (Shanghai Jiao Tong University), Zheng, Guoyan (Shanghai Jiao Tong University)
Pose EstimationDomain AdaptationDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: This paper proposes a 3D pelvic landmark localization framework based on unsupervised domain adaptation, which utilizes landmark-conditioned image synthesis and pseudo-label sets to enhance cross-modal performance between CT and MR.
Unveiling Brain-Body Axis Interactions in Psychiatric-Endocrine Disorder with Deep Graph Causal Neural Network
Wan, Zhonghua (Nanjing University of Science and Technology), Wu, Ye (Nanjing University of Science and Technology)
Explainability and InterpretabilityGraph Neural NetworkAuto EncoderContrastive LearningGraphTabularBiomedical DataMagnetic Resonance ImagingDiffusion Tensor ImagingAlzheimer's DiseaseElectronic Health Records
🎯 What it does: Built a causal graph framework based on graph neural networks to explore the brain-body axis interactions between psychiatric and endocrine diseases.
URLCF: Universal Representation Learning for Cross-Domain Few-Shot Autism Spectrum Disorder Detection
Mao, Zhenlin (Nanjing University of Information Science and Technology), Wang, Mingliang (Nanjing University of Information Science and Technology)
ClassificationDomain AdaptationKnowledge DistillationRepresentation LearningMeta LearningGraph Neural NetworkTransformerSupervised Fine-TuningAuto EncoderGenerative Adversarial NetworkContrastive LearningBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Proposed the URLCF framework, which leverages multi-source domain teacher model distillation and mask generation alignment to learn general representations, and then performs few-shot fine-tuning through a lightweight module to achieve cross-domain autism spectrum disorder detection.
UT-MIL: Uncertainty-Rectified and TME-Decoupling Dual-Stream Aggregation for Robust WSI Survival Prediction
Hu, Taiyuan (Chinese Academy of Sciences), Jiang, Jinrong (University of Chinese Academy of Sciences)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerContrastive LearningBiomedical DataBenchmark
🎯 What it does: Proposes the UT-MIL framework, combining uncertainty calibration with a dual-stream aggregation for tumor microenvironment (TME) decoupling, for robust and interpretable survival prediction on whole-slide images (WSI).
VA-Adapter: Adapting Ultrasound Foundation Model to Echocardiography Probe Guidance
Wang, Teng (Tsinghua University), Huang, Gao (Tsinghua University)
Image TranslationDomain AdaptationComputational EfficiencyRecurrent Neural NetworkTransformerSupervised Fine-TuningVision-Language-Action ModelContrastive LearningImageBiomedical DataUltrasoundElectrocardiogram
🎯 What it does: Propose Vision-Action Adapter (VA-Adapter), which utilizes a cardiac ultrasound base model and incorporates action information to achieve real-time guidance for the cardiac ultrasound probe;
Validation of simulation-based evaluation by inclusion of unlabeled data and meta-ablation
Ploner, Stefan B. (FAU Erlangen-Nürnberg), Maier, Andreas (Tufts University)
Data SynthesisDomain AdaptationOptimizationDiffusion modelAuto EncoderContrastive LearningOptical FlowImageBiomedical DataComputed Tomography
🎯 What it does: This paper proposes an evaluation framework based on the 'meta-ablation' method for assessing unlabeled data, used to verify the authenticity of simulated data, and reveals and closes the reality gap by mixing some real data into simulated data.
VarDeFlow: Variational Deformation Learning with Flow Matching for Multi-organ Cross-Modality Medical Image Synthesis
Singh, Peeyush Kumar (Indian Institute of Technology), Gupta, Pankaj (Post-Graduate Institute of Medical Education and Research)
GenerationData SynthesisConvolutional Neural NetworkTransformerDiffusion modelFlow-based ModelAuto EncoderImageBiomedical DataMagnetic Resonance ImagingComputed TomographyOrdinary Differential Equation
🎯 What it does: This paper proposes a unified framework called VarDeFlow, which combines variational deformation learning with flow matching techniques to achieve multi-organ cross-modal medical image synthesis without requiring pre-registration or registration information from the target modality.
VCDA-Net: Graph-Guided Voxel-Cluster Density Attention for Global-Regional Brain Age Gap Estimation
Nguyen, Quoc-Tien (Chonnam National University), Yang, Hyung-Jeong (Chonnam National University)
ClassificationExplainability and InterpretabilityRepresentation LearningConvolutional Neural NetworkGraph Neural NetworkTransformerAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease
🎯 What it does: Proposed a multi-scale brain age prediction framework named VCDA-Net, integrating voxel cluster density attention and graph neural networks to estimate global and local brain age gaps.
VDSB-GWSyn: Diffusion Schrödinger Bridge for Controllable and Anatomically Feasible Guidewire Synthesis in Coronary Angiography
Tang, Haoyuan (Tianjin University), Yang, Jiachen (Tianjin University)
Image TranslationGenerationData SynthesisPose EstimationConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataComputed TomographyElectronic Health RecordsReview/Survey PaperStochastic Differential Equation
🎯 What it does: Propose the VDSB-GWSyn framework, which utilizes the Diffusion Schrödinger Bridge to controllably synthesize guidewires in the coronary artery vessel context, generating high-quality guidewire images with real endpoint coordinates, and using these synthetic data to improve the performance of guidewire endpoint localization.
VecHeart: Holistic Four-Chamber Cardiac Anatomy Modeling via Hybrid VecSets
Chen, Yihong (EPFL), Fua, Pascal (EPFL)
RestorationGenerationTransformerDiffusion modelAuto EncoderContrastive LearningPoint CloudMeshBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound
🎯 What it does: This study proposes the VecHeart framework, which can uniformly reconstruct and generate four-chamber heart structures and supports the recovery of complete heart geometry from sparse or missing data;
VEGA: Vision Enhanced Generative Augmentation for Cross-Subject Electrocardiogram Generalization
Shi, Tianyi (Beijing University Of Posts And Telecommunications), Zhao, Zhicheng (Beijing University Of Posts And Telecommunications)
GenerationData SynthesisDomain AdaptationAnomaly DetectionTransformerSupervised Fine-TuningDiffusion modelAuto EncoderContrastive LearningImageTime SeriesElectrocardiogram
🎯 What it does: Propose a generation enhancement method based on a visual pre-trained model (VEGA), which converts one-dimensional electrocardiogram (ECG) signals into two-dimensional images. It utilizes a visual MAE for self-supervised reconstruction, thereby generating diverse ECG samples to improve cross-subject generalization performance.
VesselSim: learning 3D blood vessel segmentation without expert annotations
Rainville, Erin (Concordia University), Xiao, Yiming (Concordia University)
SegmentationData SynthesisDomain AdaptationConvolutional Neural NetworkDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Propose a two-stage framework, first generating 16,500 synthetic 3D vascular images using random geometry-driven vascular simulation and domain randomization, then training a 3D U-Net on these synthetic data; during inference, test-time adaptation is achieved through a self-supervised masked reconstruction decoder without any real annotations;
VigilSAM3: Reinforcement-Optimized Memory Gating for Ultra-Long Surgical Video Segmentation
Yao, Xincheng (Hong Kong University of Science and Technology), Zhu, Lei (Hong Kong University of Science and Technology)
SegmentationTransformerReinforcement LearningContrastive LearningVideoBiomedical Data
🎯 What it does: Proposed a memory-gated framework based on reinforcement learning to achieve real-time vascular structure segmentation in ultra-long surgical videos, and suppressed temporal drift through hierarchical memory firewall and dynamic scheduling mechanism.
VIHD: Visual Intervention-Based Hallucination Detection for Medical Visual Question Answering
Chen, Jiayi (Monash University), Cai, Jianfei (Monash University)
Anomaly DetectionExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataBenchmark
🎯 What it does: Developed a no-training, vision-based intervention medical visual question answering hallucination detection framework called VIHD.
ViPSAM: Visual Prompting Medical Image Segmentation Using Segment Anything Model
Lee, San (Sungkyunkwan University), Kim, Boah (Sungkyunkwan University)
SegmentationConvolutional Neural NetworkTransformerPrompt EngineeringContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Improving liver lesion segmentation on low-contrast CT using visual prompts from contrast-enhanced MRI.
Virtual 3D H&E Staining from Phase-Contrast Back-Illumination Interference Tomography
Song, Anthony A. (Johns Hopkins University), Durr, Nicholas J. (Johns Hopkins Hospital)
Image TranslationSegmentationGenerationData SynthesisTransformerDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataComputed Tomography
🎯 What it does: Constructed HistoBIT3D, the first voxel-level paired 3D Back-illumination Interference Tomography (BIT) and fluorescent nucleus image dataset, and proposed a virtual H&E staining framework based on Vision Transformer CycleGAN.
Vis2Reg: Visibility-Aware Landmark-Free Geometric 3D–2D Registration for Liver Laparoscopy
Feng, Jiaming (University of Leeds), Ali, Sharib (University of Leeds)
Pose EstimationDepth EstimationGraph Neural NetworkTransformerDiffusion modelNeural Radiance FieldAuto EncoderContrastive LearningGaussian SplattingOptical FlowImagePoint CloudBiomedical DataComputed TomographyUltrasoundReview/Survey PaperBenchmark
🎯 What it does: Proposed a visualization-aware, unlabeled point, geometry self-supervised 3D-2D liver registration framework that utilizes visible region constraints to deformations, achieving real-time AR guidance;
VisiFAT: Estimating Visceral Fat from Dual-View DXA Scans Using Cross-View Attention
Maqsood, Arooba (Edith Cowan University), Gilani, Syed Zulqarnain (Edith Cowan University)
RestorationRepresentation LearningConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: A multi-view cross-view attention framework named VisiFAT is proposed to accurately estimate visceral fat volume by integrating lateral and anterior-posterior DXA scans along with demographic information such as gender, age, height, and weight.
Vision-Based Force Estimation Using Brain Surface Deformation Models
Heraoui, Chakib (Université TÉLUQ), Gueziri, Houssem-Eddine (Neurosurgical Simulation and Artificial Intelligence Learning Centre)
Data SynthesisConvolutional Neural NetworkDiffusion modelContrastive LearningVideoBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasoundFibre Orientation DistributionDiffusion Tensor Imaging
🎯 What it does: Propose a deep learning framework based on monocular surgical microscope videos, which directly estimates tool-brain tissue contact force by utilizing brain surface deformation.
Visual-Query Curriculum Learning for Rare Fetal Biometry Plane Detection in Blind-Sweep Ultrasound Video
Kwon, Jong (University of Oxford), Noble, J. Alison (University of Oxford)
Object DetectionDomain AdaptationAnomaly DetectionConvolutional Neural NetworkTransformerContrastive LearningImageVideoBiomedical DataUltrasound
🎯 What it does: Proposed a Spatio-Temporal Transformer model based on visual queries for detecting biometric planes (HC, AC, FL) in blind-sweep prenatal ultrasound videos, achieving automatic assessment of scan quality and rapid resampling;
VL-RewardGen: Vision–Language Reward Driven Skin Lesion Image Generation
Chu, Yusung (Yonsei University), Yang, Sejung (Yonsei University)
ClassificationGenerationData SynthesisConvolutional Neural NetworkTransformerSupervised Fine-TuningReinforcement LearningPrompt EngineeringVision Language ModelDiffusion modelImageTextBiomedical Data
🎯 What it does: Proposed a control-based skin lesion image generation framework called VL-RewardGen, which is based on visual-language rewards
VM-NeXT UNet: Synergizing ConvNeXT and Visual Mamba for Robust Medical Image Segmentation
Cao, Boxuan (Xi'an Jiaotong University), Jiang, Peilin (Xi'an Jiaotong University)
SegmentationConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Propose a dual-encoder U-Net structure that integrates ConvNeXT with Visual Mamba for medical image segmentation
VOGeo-Gaze: Calibration-Free, Geometry-Aware Deep Learning for Real-Time Gaze Tracking in Clinical Video-Oculography
Zhao, Jingkang (LMU University Hospital), Wuehr, Max (LMU University Hospital)
Pose EstimationOptimizationComputational EfficiencyConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningOptical FlowImageVideoBiomedical DataElectrocardiogramReview/Survey Paper
🎯 What it does: Propose a calibration-free, geometry-aware deep learning framework VOGeo-Gaze for real-time eye tracking, which can recover ocular parameters and infer gaze direction from geometric information obtained through image segmentation.
VolDiT: Controllable Volumetric Medical Image Synthesis with Diffusion Transformers
Seyfarth, Marvin (Heidelberg University), Engelhardt, Sandy (Heidelberg University)
SegmentationGenerationData SynthesisTransformerDiffusion modelAuto EncoderBiomedical DataComputed Tomography
🎯 What it does: Proposed and implemented VolDiT — a fully Transformer-based 3D diffusion model for unsupervised generation and conditional generation of volumetric medical images.
Volume-Fused Super-Resolution OCT Reveals Cellular-Level Structures in the Outer Segment Across Wider Fields Without Adaptive Optics
Ploner, Stefan B. (FAU Erlangen-Nürnberg), Maier, Andreas (Massachussets Institute of Technology)
Super ResolutionConvolutional Neural NetworkOptical FlowBiomedical Data
🎯 What it does: By multi-frame visible light OCT convolution fusion and 3D super-resolution reconstruction, achieving visualization of the photoreceptor outer segment (OS) cone cell-level structures within a wide field of view (3×3 mm) without adaptive optics (AO) equipment.
VoxShield: Protecting 3D Medical Datasets from Unauthorized Training via Frequency-Aware Inter-Slice Disruption
Liu, Xinyao (Westlake University), Zheng, Yefeng (Westlake University)
SegmentationSafty and PrivacyConvolutional Neural NetworkTransformerGenerative Adversarial NetworkContrastive LearningBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Propose a data protection method for 3D medical image segmentation called VoxShield, which generates unlearnable examples (UE) to prevent unauthorized model training;
VRUGA: Visual and Reasoning Uncertainty Guided Active Learning for Medical MLLM Post-training
Xiao, Boyi (University of Science and Technology of China), Zhang, Shaoting (Shanghai Artificial Intelligence Laboratory)
Explainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelReinforcement LearningPrompt EngineeringVision Language ModelImageTextMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyBenchmarkChain-of-Thought
🎯 What it does: Proposed a proactive learning framework named VRUGA to maximize the utilization efficiency of expert annotation budgets during post-training of medical multimodal large language models (MLLMs); high-risk samples are filtered through dual uncertainty quantification (visual stability and reasoning trajectory consistency); logical-aware diversity sampling (LADS) is used to select diverse and informative samples from the candidate pool; and cross-threshold pseudo-labeling (CTPL) is adopted during post-training with reinforcement learning to balance expert correction and self-supervised anchors, stabilizing policy convergence.
WaCT-US: Cyclic-Alignment Scaled Contour Tubes for Efficient, Calibrated Uncertainty in Echocardiography
Kazemi Esfeh, Mohammad Mahdi (University of British Columbia), Abolmaesumi, Purang (University of British Columbia)
SegmentationConvolutional Neural NetworkDiffusion modelScore-based ModelContrastive LearningImageBiomedical DataUltrasound
🎯 What it does: Propose a framework called WaCT-US, which is based on closed curves and provides infinitely calibrable uncertainty. It can directly generate a set of probabilistic predictions surrounding the prediction boundary (contour tube) after a single forward pass, and eliminates index ambiguity by utilizing the cyclic alignment of closed curves.
Wasserstein-Aligned Localisation for VLM-Based Distributional OOD Detection in Medical Imaging
Kainz, Bernhard (FAU Erlangen-Nürnberg), Bercea, Cosmin (Technical University Munich)
Image TranslationAnomaly DetectionTransformerPrompt EngineeringVision Language ModelContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyRetrieval-Augmented Generation
🎯 What it does: Proposes an untrained framework called WALDO for zero-shot anomaly localization using vision-language models (VLMs), primarily by identifying lesions through comparison with healthy reference images.
WaveDE: Wavelet-Routed Dual-Expert Network for Fundus Photograph Enhancement
Nie, Haitao, Xie, Bin
RestorationConvolutional Neural NetworkMixture of ExpertsImage
🎯 What it does: Unable to determine
WaveDiT: Distribution-Aware Wavelet Flow Matching for Efficient 3D Brain MRI Synthesis
Danese, Danilo (Politecnico di Bari), Di Noia, Tommaso (Politecnico di Bari)
GenerationData SynthesisTransformerDiffusion modelFlow-based ModelRectified FlowContrastive LearningBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Propose a conditional flow matching framework called WaveDiT based on the Haar 3D discrete wavelet transform for high-resolution 3D brain MRI synthesis, and implement distribution-aware uncertainty scheduling through the Morpheus module, combined with hierarchical spatiotemporal attention for efficient sampling.
Wavelet-Enhanced Coarse-to-Fine Radiology Report Generation with Large Language Models
Yang, Shuanghong (Chongqing Normal University), Zhang, Wenfeng (Chongqing Normal University)
GenerationRetrievalExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextBiomedical DataComputed TomographyElectronic Health RecordsRetrieval-Augmented Generation
🎯 What it does: Proposes the Wavelet-Enhanced Coarse-to-Fine Radiology Report Generation (WERG) framework based on a multi-modal large language model, which generates more detailed and clinically consistent chest X-ray reports by leveraging patients' prior images, retrieved text segments, and enhancing diagnostic-guided visual feature calibration with frequency domain details.
Wavelet-Guided Spectral–Spatial Learning for Medical Hyperspectral Image Segmentation
Ji, Junya (University of Exeter), Ye, Xujiong (University of Exeter)
SegmentationConvolutional Neural NetworkTransformerBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Proposed a wavelet decomposition-based waveguide spectral-spatial learning framework for pixel-level segmentation of medical hyperspectral images;
wBCAM: Windowed Bi-directional Cross-Attention with Mamba for Enhanced 3D Vascular Reconstruction in Free-Hand Photoacoustic and Ultrasound Imaging
Lee, SiYeoul (Pusan National University), Kim, MinWoo (Pusan National University)
RestorationTransformerContrastive LearningOptical FlowVideoBiomedical DataPositron Emission TomographyUltrasound
🎯 What it does: This paper proposes a sensor-free freehand 3D PAUS imaging framework called wBCAM, which achieves high-resolution 3D vascular reconstruction by estimating scanning motion through windowed bidirectional cross-attention and Mamba state space modules.
Weakly-Supervised Coronary Artery Segmentation from DSA Sequence via Motion-Aware Modeling
Wu, Han (ShanghaiTech University), Shen, Dinggang (ShanghaiTech University)
SegmentationConvolutional Neural NetworkRecurrent Neural NetworkTransformerContrastive LearningOptical FlowVideoBiomedical DataComputed TomographyBenchmark
🎯 What it does: Proposed the MACOS framework, which utilizes long-term digital subtraction angiography (DSA) sequences to achieve weakly supervised segmentation of coronary arteries using only keyframe annotations.
Weakly-Supervised Temporal Decomposition with Graph Neural Networks for Parkinsonian Turning Assessment
Zhang, Jieming (Sungkyunkwan University), Park, Hogun (Sungkyunkwan University)
ClassificationPose EstimationExplainability and InterpretabilityGraph Neural NetworkContrastive LearningVideoBiomedical DataReview/Survey Paper
🎯 What it does: This study proposes a weakly supervised temporal decomposition framework that utilizes graph neural networks to automatically learn the temporal segmentation of turning movements from video-level diagnostic labels, and combines attention aggregation for PD (Parkinson's disease) turning assessment.
When Brains Disagree: Biological Ambiguity Underlies the Challenge of Amyloid PET Synthesis from Structural MRI
Baron, Louise E. G. (University College London), Zhang, Hui (AINOSTICS Ltd)
Image TranslationData SynthesisTransformerDiffusion modelGenerative Adversarial NetworkImageMultimodalityBiomedical DataMagnetic Resonance ImagingPositron Emission TomographyAlzheimer's Disease
🎯 What it does: Investigate the biological ambiguity in the synthesis from MRI to Alzheimer's alpha PET through controlled experiments, and explore whether multimodal inputs can resolve this issue.
When Does Uncertainty Become Meaningful in Medical Vision-Language Models? Alignment Enables Disease-wise Uncertainty and Risk-Aware Prediction
Kim, Tae Hun (Inha University), Lee, Hyun Gyu (Inha University)
ClassificationAnomaly DetectionExplainability and InterpretabilityTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataComputed TomographyReview/Survey Paper
🎯 What it does: In chest X-ray multi-label diagnosis, Bi-EDL is proposed by combining bidirectional multi-choice training with evidential deep learning to generate disease-level uncertainty and achieve risk-aware prediction.
When Tissue Speaks: Vibroacoustic Monitoring of Tissue Response During Electrosurgery
Esmaeili, Nazila (University Medical Center Göttingen), Illanes, Alfredo (Fraunhofer Research Institution for Individualized and Cell-based Medical Engineering)
ClassificationRecognitionTime SeriesBiomedical Data
🎯 What it does: By attaching a vibration acoustic sensor to a conventional bipolar electrosurgical forceps handle, structural vibration signals are collected in real-time to monitor tissue response during electrocoagulation.
When Views Disagree: Conflict-Aware Evidential Inference for Multi-view Mammography
Ahmed, Md Rayhan (University of British Columbia), Lasserre, Patricia (University of British Columbia)
ClassificationAnomaly DetectionConvolutional Neural NetworkTransformerContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound
🎯 What it does: Designed and implemented a multi-view mammography image classification framework based on evidence reasoning, CEI-Net, which can fuse diagnostic information from two views in an uncertainty-driven manner when views are inconsistent.
When, Where, and How: Adaptive Binning for Tabular Self-supervised Learning
Kim, Daehwan (Hanyang University), Jang, Ikbeom (Hankuk University of Foreign Studies)
Representation LearningData-Centric LearningAuto EncoderContrastive LearningTabularBiomedical DataElectronic Health Records
🎯 What it does: Propose a self-supervised learning framework on medical tabular data, utilizing an adaptive binning mechanism to enhance representation learning.
Where Is the Needle Tip? VA Sensing Reveals Hidden Needle-Tissue Transitions During US Guidance
Urrutia, Robin (Otto von Guericke University Magdeburg), Illanes, Alfredo (Otto von Guericke University Magdeburg)
Anomaly DetectionAutonomous DrivingRobotic IntelligenceDiffusion modelContrastive LearningOptical FlowVideoBiomedical DataUltrasoundAudio
🎯 What it does: Investigate the mechanism changes when the needle tip is invisible under ultrasound guidance by capturing the changes through proximal vibration acoustic (VA) sensors when the needle is inserted into different tissue interfaces;
WHO-CXRBench: A CXR Benchmark for Diagnostic Rule Retrieval in CLIP-Style MedVLMs
Liu, Yixiang (Southern University of Science and Technology), Tang, Xiaoying (Monash University)
RetrievalExplainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelVision Language ModelContrastive LearningImageTextBiomedical DataBenchmarkRetrieval-Augmented Generation
🎯 What it does: Constructed the WHO-CXRBench evaluation benchmark based on the WHO chest X-ray diagnostic guidelines to assess the ability of CLIP-style medical vision-language models in retrieving fine-grained diagnostic rules from chest X-rays.
Wrist Camera Pose-Guided Multi-view Fusion for Occlusion Reconstruction in Robot-Assisted Microsurgical Anastomosis
Zhou, Xinyao (Shanghai Jiao Tong University), Guo, Yao (Shanghai Jiao Tong University)
SegmentationRobotic IntelligenceTransformerAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataReview/Survey Paper
🎯 What it does: Propose a multi-view fusion framework guided by the pose of the wrist camera, aimed at reconstructing and segmenting suture thread occlusions in robot-assisted microsurgical anastomosis.
X-Edit: Exact, Explicit, and Explainable Null-Space Editing for Medical Vision Transformers
Liu, Yuanye (Fudan University), Zhuang, Xiahai (Johns Hopkins University)
ClassificationExplainability and InterpretabilityComputational EfficiencyTransformerSupervised Fine-TuningContrastive LearningImageBiomedical Data
🎯 What it does: Proposes a model editing framework called X-Edit based on zero-space projection, used to correct misjudgments in medical vision Transformers without causing catastrophic forgetting.
X2Bone: Reconstructing 3D Bone Structures from 2D Biplanar X-Rays via Cross-RWKV
Pan, Zhaohong (Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences), Liang, Xiaokun (Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences)
Image TranslationRestorationGenerationPose EstimationDepth EstimationConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningImageBiomedical DataComputed Tomography
🎯 What it does: Developed an end-to-end X2Bone framework for reconstructing patient-specific 3D skeletal structures from dual-plane X-rays.
XTinyU-Net: Training-Free U-Net Scaling via Initialization-Time Sensitivity
Kimbowa, Alvin (University of British Columbia), Hacihaliloglu, Ilker (University of British Columbia)
SegmentationConvolutional Neural NetworkAuto EncoderContrastive LearningImageMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasound
🎯 What it does: Propose a training-agnostic U-Net width scaling method that selects the minimal stable network by utilizing Jacobian sensitivity at initialization.
Zero-Shot CT Super-Resolution Using Diffusion-Based 2D Projection Priors and Signed 3D Gaussians
Noh, Jeonghyun (Korea University), Jeong, Won-Ki (Korea University)
Super ResolutionConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelGaussian SplattingImageBiomedical DataComputed Tomography
🎯 What it does: Propose a zero-shot 3D CT super-resolution framework: first, use a diffusion model to reconstruct high-resolution projections from low-resolution projections, then use negative α mixed Gaussian splatting (NAB-GS) to map the residual to the volume space, achieving HR volume recovery.