MICCAI 2026 Papers — Page 9
International Conference on Medical Image Computing and Computer-Assisted Intervention · 1165 papers
POT-SAM3: Prompt-Only Tuning for One-Shot Electron Microscopy Segmentation
Gu, Xuncheng (Beijing University of Posts and Telecommunications), Xiao, Li (Beijing University of Posts and Telecommunications)
SegmentationTransformerSupervised Fine-TuningPrompt EngineeringContrastive LearningBiomedical Data
🎯 What it does: Propose a Prompt-Only Tuning framework (POT-SAM3), which learns a few concept prompt tokens on a single annotated slice, enabling the frozen SAM3 model to achieve prompt-free detection and cross-slice propagation instance segmentation across the entire electron microscopy sequence.
PRA-PoE: Robust Multimodal Alzheimer’s Disease Classification under Arbitrary Modality Missingness
Yang, Guangqian (Hong Kong Polytechnic University), Wang, Shujun (Hong Kong Polytechnic University)
ClassificationRepresentation LearningData-Centric LearningConvolutional Neural NetworkTransformerMixture of ExpertsContrastive LearningMultimodalityTabularBiomedical DataMagnetic Resonance ImagingPositron Emission TomographyAlzheimer's DiseaseElectronic Health Records
🎯 What it does: Propose the PRA-PoE framework to achieve multi-modal classification of Alzheimer's disease under arbitrary missing modalities
Precision Recall Controllable Radiology Report Generation via Hybrid Natural Language and Clinical Reward Learning
Chen, Ling (Ohio State University), Wu, Dufan (Ohio State University)
GenerationReinforcement Learning from Human FeedbackTransformerReinforcement LearningVision Language ModelImageTextBiomedical DataComputed TomographyElectronic Health Records
🎯 What it does: Proposed a controllable reinforcement learning framework for generating radiology reports from chest X-ray images, achieving a trade-off between clinical precision and recall through control parameters.
Predicting Radiologist Expertise from 3D Gaze Patterns During CT Interpretation
Khaertdinova, Leila (University of Copenhagen), Ibragimov, Bulat (University of Copenhagen)
ClassificationExplainability and InterpretabilityConvolutional Neural NetworkTransformerVision Language ModelContrastive LearningImageBiomedical DataComputed Tomography
🎯 What it does: Construct a professionalism classification model based on Transformer using eye-tracking during the CT interpretation process, integrating 3D fixation information to predict the experience level of radiologists.
Preoperative Simulation of Personalized Breast Reconstruction via Implant Prior Controllable Diffusion
Wang, Yizhi (Zhejiang University), Jin, Yaochu (Westlake University)
Image TranslationRestorationSegmentationGenerationConvolutional Neural NetworkTransformerPrompt EngineeringDiffusion modelAuto EncoderContrastive LearningImageBiomedical DataComputed Tomography
🎯 What it does: Proposed a pre-surgical breast reconstruction simulation framework based on a controllable diffusion model, which can generate post-operative reconstruction results from pre-operative CT scans and support editing of implant parameters;
Pressure-Tuned Auditor for Sycophancy-Resistant Breast Ultrasound VLM Diagnosis
Zhang, Hongze (Shanghai University of Engineering Science), Qiu, Xihe (Shanghai University of Engineering Science)
Anomaly DetectionExplainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelImageMultimodalityBiomedical DataUltrasoundChain-of-Thought
🎯 What it does: Propose a two-agent pressure-tuned auditor (PTA) to detect and correct sycophancy errors in breast ultrasound vision-language models (VLMs) when they receive biased prompts, while enhancing interpretability through visual localization evidence.
PriLoRA: Prior-Conditioned Low-Rank Adapters for Medical Vision Models
Kazerouni, Amirhossein (University of Toronto), Taati, Babak (University of Toronto)
ClassificationRestorationSegmentationGenerationTransformerPrompt EngineeringContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Propose Prior-Conditioned Low-Rank Adapters (PriLoRA), a parameter-efficient fine-tuning method for medical vision models that dynamically adjusts low-rank sinusoidal updates using input priors.
Prior-Anchored Debiasing for Long-Tailed Multi-Organ Pathology Report Generation
Yang, Feng (City University of Hong Kong), Chen, Ping (University of Massachusetts Boston)
GenerationData-Centric LearningTransformerVision Language ModelContrastive LearningImageTextBiomedical DataElectronic Health RecordsRetrieval-Augmented Generation
🎯 What it does: This paper proposes the PriOrGen framework to address the long-tail distribution bias in multi-organ pathology report generation.
Prior-Guided Hierarchical Vector Quantized-VAE for Digital Subtraction Angiography Generation
Liu, Yunbi, Liu, Qingshan (Nanjing University of Posts and Telecommunications)
RestorationGenerationTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Proposed and implemented a Prior-Guided Hierarchical Vector Quantized Variational Autoencoder (PH-VQVAE), which generates digital subtraction angiography images without motion artifacts and with complete vascular connectivity by achieving 'soft subtraction' of contrast images and pre-contrast masks through a dual-branch encoder and windowed cross-attention.
PRISM: Differentiable Analysis-by-Synthesis for Fixel Recovery in Diffusion MRI
Abouagour, Mohamed (Indiana University Bloomington), Garyfallidis, Eleftherios
OptimizationDiffusion modelAuto EncoderContrastive LearningGaussian SplattingBiomedical DataMagnetic Resonance ImagingDiffusion Tensor Imaging
🎯 What it does: A differentiable analysis-synthesis framework called PRISM is constructed to end-to-end optimize multi-component microstructure models in 3D space while simultaneously correcting intensity distortions, thereby recovering fiber peak (fixel) information in diffusion magnetic resonance imaging.
Privacy-Preserving Video-Based Facial Palsy Recognition via Dynamic Facial Action Modeling
Chang, Yuance (Xi'an Jiaotong University), Xi, Wei (Xi'an Jiaotong University)
RecognitionSafty and PrivacyConvolutional Neural NetworkRecurrent Neural NetworkTransformerAuto EncoderContrastive LearningImageVideo
🎯 What it does: Designed and implemented a video-based privacy-preserving facial palsy recognition framework named DP-Face, and publicly released the largest Palsy-330 video dataset.
Proactive Domain Unification for Robust Echocardiography Segmentation
Pang, Xintao (Macao Polytechnic University), Tan, Tao (Macao Polytechnic University)
SegmentationDomain AdaptationTransformerDiffusion modelAuto EncoderContrastive LearningImageBiomedical DataUltrasound
🎯 What it does: Propose a proactive domain unification framework (PDU) during inference, which maps cardiac ultrasound images from different centers and devices into a source-domain-aligned style space to enhance cross-center segmentation performance.
Probabilistic Multi-rater Segmentation via Style-Aware Boundary Conditioning
Karimijafarbigloo, Sanaz (University of Regensburg), Merhof, Dorit (University of Regensburg)
SegmentationConvolutional Neural NetworkDiffusion modelScore-based ModelContrastive LearningBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Propose a lightweight probabilistic multi-evaluator segmentation framework, which models experts' boundary preferences and generates diverse segmentation results that align with individual preferences through style-conditioned personalization and differential boundary loss.
Probe-EM: Targeted Neuron Tracing via Training-Free Semantic Verification
Jiang, Liuyun (Chinese Academy of Sciences), Han, Hua (Chinese Academy of Sciences)
SegmentationExplainability and InterpretabilityComputational EfficiencyTransformerSupervised Fine-TuningVision Language ModelDiffusion modelContrastive LearningPoint CloudBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasoundFibre Orientation DistributionDiffusion Tensor ImagingElectronic Health Records
🎯 What it does: A training-agnostic directional neuronal tracing framework called Probe-EM was developed, utilizing skeleton-guided heuristic spatial search and dimension-aware semantic verification based on the Foundation model NeuroSAM 2, achieving high-precision neuronal morphology reconstruction and accelerated manual correction.
Progressive Self-supervised Learning with Individualized Community Assignment for Brain Network Analysis
Chen, Hairui (Harbin Institute of Technology at Shenzhen), Ma, Ting (Harbin Institute of Technology at Shenzhen)
ClassificationAnomaly DetectionRepresentation LearningGraph Neural NetworkTransformerContrastive LearningGraphBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease
🎯 What it does: This paper proposes the BrainPICM framework, which performs representation learning and disease diagnosis on functional magnetic resonance brain networks through advanced self-supervised learning combined with personalized community assignment.
ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
Chhetri, Aavash (NepAl Applied Mathematics and Informatics Institute for research), Bhattarai, Binod (West Virginia University)
Data SynthesisFederated LearningTransformerMixture of ExpertsContrastive LearningImageTextMultimodalityElectronic Health Records
🎯 What it does: A multi-modal federated learning framework called ProMoE-FL is studied for synthesizing missing modality features under missing modality scenarios, enabling feature synthesis of missing modalities between different hospitals without public data sharing.
Prompt and Probe: Data-Efficient Adaptation of Black-Box Foundation Models via Unified Active Visual Prompting
Mahapatra, Dwarikanath (Khalifa University), Razzak, Imran (MBZUAI)
ClassificationSegmentationComputational EfficiencyData-Centric LearningTransformerPrompt EngineeringContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound
🎯 What it does: A black-box adaptation framework called Prompt and Probe is constructed through a visual prompt decoder and multi-scale probing, enabling frozen medical foundation models to efficiently adapt and actively select annotated samples for classification and segmentation tasks.
Prompt Group-Aware Training for Robust Text-Guided Nuclei Segmentation
Wu, Yonghuang (Shanghai University), Yu, Jinhua (Shanghai University)
Protein Structure PredictionGraph Neural NetworkPrompt EngineeringContrastive LearningGraphBiomedical Data
🎯 What it does: This paper proposes a protein interaction prediction model based on a hierarchical graph attention network.
PromptGate: Client-Adaptive Vision–Language Gating for Open-Set Federated Active Learning
Nesturi, Adea (University of Bonn), Albarqouni, Shadi (University of Bonn)
Anomaly DetectionOptimizationFederated LearningTransformerPrompt EngineeringVision Language ModelContrastive LearningImageBiomedical Data
🎯 What it does: Propose PromptGate, a dynamic vision-language gating framework for open-set federated active learning;
PromptMedCT: A Training-Free LLM-Guided Framework for Volumetric CT Data Realization with Quantifiable Disease Manifestations and Multi-Tier Annotations
Owais, Muhammad (Khalifa University), Hussain, Irfan (Khalifa University)
Image TranslationSegmentationGenerationData SynthesisTransformerLarge Language ModelPrompt EngineeringDiffusion modelImageBiomedical DataComputed TomographyRetrieval-Augmented Generation
🎯 What it does: Propose the PromptMedCT framework, which utilizes training-free LLM-guided diffusion models to generate lung CT images with multi-level annotations (descriptions, region masks, bounding boxes).
Propagating Structural Guidance: Synthesizing Fluorescein Angiography from Fundus Images and Sparse OCT Scans
Ma, Tengfei (Southeast University), Fan, Wen (Nanjing University of Science and Technology)
Image TranslationGenerationData SynthesisTransformerDiffusion modelGenerative Adversarial NetworkContrastive LearningImageMultimodalityBiomedical DataComputed TomographyReview/Survey Paper
🎯 What it does: Construct a tri-modal fundus image dataset and develop an FFA synthesis framework guided by OCT structure.
ProSMA-UNet: Decoder Conditioning for Proximal-Sparse Skip Feature Selection
Cheng, Chun-Wun (University of Cambridge), Aviles-Rivero, Angelica I. (Tsinghua University)
SegmentationConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningImagePoint CloudBiomedical DataComputed TomographyPositron Emission TomographyUltrasound
🎯 What it does: Propose ProSMA-UNet, which explicitly selects features in the skip connections of U-Net through a decoder-conditioned sparse multi-scale attention gate (Proximal-Sparse Gate), significantly suppressing noise and background information.
ProSyn-Net: A Reliable Prototypical Synergy Network for Imbalanced Dermatologic Multimodal Ultrasound Diagnosis
Lin, Jicheng (Tongji University), Luo, Ye (Tongji University)
ClassificationExplainability and InterpretabilityComputational EfficiencyConvolutional Neural NetworkTransformerMixture of ExpertsContrastive LearningImageMultimodalityBiomedical DataUltrasound
🎯 What it does: Propose the ProSyn-Net model, which achieves reliable diagnosis of dermatological ultrasound images through dynamic cross-modal attention fusion, nonlinear prototype mapping, and uncertainty-aware ensemble.
Proto-CAP: Prototype Memory Fusion and Curriculum-Adaptive Loss for Pediatric Myopia Visual Question Answering
Yan, Xu (Nankai University), Li, Tao (Tianjin Medical University Eye Hospital)
RecognitionData-Centric LearningConvolutional Neural NetworkTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBiomedical DataBenchmark
🎯 What it does: Constructed the first multi-modal medical visual question answering dataset (PM-VQA) specifically focused on childhood myopia, and proposed the Proto-CAP framework to address the semantic gap and overfitting issues in few-shot medical VQA.
Protocol-Manifold Coverage Learning (PMCL): Source-Free Active Adaptation for Multi-Center Nasopharyngeal Carcinoma MRI Segmentation
Lu, Jiawen (Fuzhou University), Wu, Xiangjun (Fuzhou University)
SegmentationDomain AdaptationConvolutional Neural NetworkSupervised Fine-TuningContrastive LearningImageBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Propose the Protocol‑Manifold Coverage Learning (PMCL) framework to achieve source-agnostic active adaptation, addressing the problem of tumor boundary drift caused by protocol mixing in cross-center CE-T1 MRI.
PROTON: Prototype-Based Test-Time Online OOD Detection for Medical VLMs
Das, Abhijit (Mohamed bin Zayed University of Artificial Intelligence), Razzak, Imran (Khalifa University)
Domain AdaptationAnomaly DetectionComputational EfficiencyRepresentation LearningPrompt EngineeringVision Language ModelContrastive LearningImageMultimodalityBiomedical DataElectronic Health RecordsRetrieval-Augmented Generation
🎯 What it does: This paper proposes PROTON, an online out-of-distribution (OOD) detection framework based on prototypes, which can improve the OOD detection performance of medical vision-language models (VLMs) under different types of shifts without modifying the model, training data, or prompt engineering. It achieves this by maintaining an online prototype library and adaptively fusing it with Maximum Concept Matching (MCM) scores based on variance.
ProtoReg: Prototype Guidance and Dual-Level Uncertainty-Weighted Consistency for Semi-supervised Pulmonary Perfusion Regression
Zhuang, Yitao (Hong Kong Polytechnic University), Ren, Ge (Hong Kong Polytechnic University)
Image TranslationRestorationDomain AdaptationConvolutional Neural NetworkTransformerAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataComputed TomographyPositron Emission Tomography
🎯 What it does: Propose a semi-supervised framework ProtoReg for predicting lung perfusion maps from non-contrast CT;
Prototype Instance-Semantic Disentanglement with Low-Rank Regularized Subspace Clustering for WSIs Explainable Recognition
Li, Chentao (Columbia University), Huang, Pan (Hong Kong Polytechnic University)
RecognitionExplainability and InterpretabilityContrastive LearningBiomedical Data
🎯 What it does: Proposed an end-to-end prototype instance-semantic disentangled framework called PID-LRSC, which uses low-rank regularized subspace clustering to eliminate instance and semantic confusion in whole slide image (WSI) multi-instance learning, thereby improving diagnostic performance and interpretability.
Prototype Learning for Visual Field Estimation from Fundus Photography in High Myopia
Li, Guoliang, Liang, Dong (Hong Kong Polytechnic University)
Explainability and InterpretabilityRepresentation LearningConvolutional Neural NetworkAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Predicting Visual Field Sensitivity in High Myopia Based on Fundus Photographs
Prototype Memory-Guided Training-Free Anomaly Classification and Localization in Prenatal Ultrasound
Liang, Huanwen (Shenzhen University), Ni, Dong (Shenzhen University)
ClassificationSegmentationAnomaly DetectionTransformerVision-Language-Action ModelAuto EncoderContrastive LearningImageBiomedical DataUltrasound
🎯 What it does: The paper proposes a training-free multi-class prenatal ultrasound anomaly classification and localization framework that relies only on a small number of reference images.
Prototype-based Physiological Transfer Enables NCCT-only Hyperacute Stroke Tissue-Window Segmentation Under Missing Perfusion
Dan, Ying (Chinese University of Hong Kong), Tong, Raymond Kai-yu (Chinese University of Hong Kong)
SegmentationConvolutional Neural NetworkTransformerSupervised Fine-TuningPrompt EngineeringContrastive LearningImageBiomedical DataComputed Tomography
🎯 What it does: Propose a prototype-based physiological information transfer framework, ProPhyT, for the segmentation of the core and penumbra in acute ischemic stroke using only NCCT images.
Prototype-Driven Robust Domain Generalization in Corneal Endothelial Cell Segmentation
Li, Hongshuo (Ningbo Institute of Materials Technology and Engineering, Chinese Academy of Sciences), Zhao, Yitian (ShanghaiTech University)
SegmentationDomain AdaptationConvolutional Neural NetworkContrastive LearningBiomedical DataComputed TomographyUltrasound
🎯 What it does: Designed and implemented a domain adaptation framework based on prototype learning, named ProtoCEC, for iris endothelial cell segmentation, significantly improving segmentation robustness in various out-of-domain scenarios such as cross-centers, cross-diseases, cross-devices, and cross-species.
Proxy-CAC: Learning Coronary Calcium Segmentation from Volume-Level Scores
Sato Imuro, Sandra Emi (Rice University), Sabharwal, Ashutosh (Houston Methodist)
SegmentationConvolutional Neural NetworkScore-based ModelContrastive LearningBiomedical DataComputed Tomography
🎯 What it does: Propose the Proxy-CAC method, which uses volume-level Agatston score to train a 3D U-Net, achieving segmentation of coronary artery calcification and prediction of the score.
PseudoRET: A General Retinal Segmentation Framework with Multi-target Pseudo Embeddings
Wang, Zhonghua (Monash University), Ge, Zongyuan (Airdoc LLC)
SegmentationConvolutional Neural NetworkTransformerContrastive LearningImageBiomedical DataBenchmark
🎯 What it does: Proposes PseudoRET, a semi-supervised multi-task retinal segmentation framework based on pseudo-label memory, capable of unifying fragmented, task-specific fundus image datasets.
PSP: Harnessing Position and Shape Priors for Cross-Domain Few-Shot Medical Image Segmentation
Xu, Bin (Nanjing University of Science and Technology), Zhang, Haofeng (Nanjing University of Science and Technology)
SegmentationDomain AdaptationConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: This paper proposes a cross-domain few-shot medical image segmentation framework called PSP, which utilizes position and shape priors to counteract texture differences across different modalities, significantly improving segmentation performance.
Pulmonary Embolism Risk Stratification from CTPA and Medical Records: Vascular Graphs Are Not All You Need
Painchaud, Nathan (INSA-Lyon), Merveille, Odyssée (École Polytechnique)
ClassificationAnomaly DetectionGraph Neural NetworkTransformerSupervised Fine-TuningContrastive LearningImageGraphTabularBiomedical DataComputed TomographyElectronic Health Records
🎯 What it does: Propose an automated pipeline based on CTPA and medical records, using cardiac biomarkers and pulmonary vascular graphs for risk stratification prediction of pulmonary embolism.
QCAgent: An Agentic Framework for Quality-Controllable Pathology Report Generation from Whole Slide Image
Wang, Rundong (University of Science and Technology of China), Zhou, Shaohua Kevin (University of Science and Technology of China)
GenerationRetrievalExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelAgentic AIVision Language ModelImageTextBiomedical DataElectronic Health RecordsRetrieval-Augmented Generation
🎯 What it does: Propose a quality-controllable whole slide image (WSI) pathology report generation framework called QCAgent, which automatically generates verifiable and complete diagnostic reports through a closed-loop audit-retrieval-revision process.
QG-MIL: A Gated Transformer Aggregator for Domain-Agnostic Multiple Instance Learning in Medical Imaging
Zedda, Luca (University of Cagliari), Loddo, Andrea (University of Cagliari)
Anomaly DetectionTransformerImageBiomedical Data
🎯 What it does: Proposes QG-MIL — a gated Transformer aggregator that addresses the attention focusing problem in multi-instance learning.
Quality-Guided Semi-supervised Learning for Medical Image Segmentation
Abhishek, Kumar (Simon Fraser University), Hamarneh, Ghassan (Simon Fraser University)
SegmentationConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningImageBiomedical Data
🎯 What it does: Propose a semi-supervised medical image segmentation framework based on quality prediction, which uses a pre-trained quality assessment network to guide the quality of unlabeled data.
Quantification of 3D Musculoskeletal CT Image-Based Biomarkers in the Lumbar Spine: Application in Metastatic Prostate Cancer
Tabatabaei, Saleh (Sunnybrook Research Institute), Hardisty, Michael (Sunnybrook Research Institute)
RecognitionSegmentationConvolutional Neural NetworkGraph Neural NetworkImageBiomedical DataComputed Tomography
🎯 What it does: This study developed a deep learning-based automated pipeline that utilizes sequential abdominal CT scans to quantitatively assess spinal bone mineral density and the volume and density of the psoas major muscles in the lumbar region, enabling longitudinal monitoring of skeletal muscle function in patients with prostate cancer;
Quantification of Uncertainty with Adversarial Models in Medical Image Segmentation
Jebril, Hana (Medical University of Vienna), Bogunović, Hrvoje (Medical University of Vienna)
SegmentationExplainability and InterpretabilityAdversarial AttackConvolutional Neural NetworkMixture of ExpertsAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: This paper proposes a post-hoc adversarial search-based framework for uncertainty quantification in medical image segmentation, named QUAM-SM, which can identify 'fragile' regions where model predictions at the pixel level are easily flipped by adversarial perturbations, and generate reliable pixel-level uncertainty maps based on this.
Quantify, Collect, and Correct: Evidence-Driven Multi-agent Framework for Critical Evidence Refinement in Emergency Rooms
Cao, Haoyu (Harbin Institute of Technology), Wang, Wei (Harbin Institute of Technology)
OptimizationExplainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AITextBiomedical DataElectronic Health RecordsRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Studied a closed-loop multi-agent framework called QCC, which prioritizes evidence as the primary optimization objective, aiming to improve the quality of evidence collection and decision-making in interactive diagnosis within emergency room scenarios.
Question-Driven Decision Tree via Knowledge Distillation for Interpretable Prediction
Yao, Xuancheng (Shanghai Jiao Tong University), Hong, Yi (Shanghai Jiao Tong University)
Explainability and InterpretabilityKnowledge DistillationConvolutional Neural NetworkTransformerContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyElectronic Health Records
🎯 What it does: This study proposes a question-driven decision tree (QDT) framework that leverages knowledge distillation to transfer the visual reasoning capabilities of a high-capacity teacher model to an interpretable student model constrained by logic, achieving auditable diagnostic pathways.
R-MSA: Rhythm-Conditioned Surgical Video Skill Assessment via Morphology Synchronization and Uncertainty Calibration
Ke, Xiao (Fuzhou University), Xu, Huangbiao (Fuzhou University)
Explainability and InterpretabilityComputational EfficiencyRobotic IntelligenceConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningGaussian SplattingVideoBiomedical DataBenchmark
🎯 What it does: Propose the R-MSA framework, which realizes video-based surgical skill assessment using rhythm-conditioned encoding, morphological synchronization decoding, and uncertainty calibration trees.
R2P-SAM: Root-to-Prompt Hair Instance Segmentation with Geometric Priors under Point Supervision
Hu, Haigen (Zhejiang University of Technology), Su, Yiping (First People's Hospital)
SegmentationConvolutional Neural NetworkTransformerPrompt EngineeringContrastive LearningImage
🎯 What it does: Proposes R2P-SAM, a sparse point-supervised hair instance segmentation framework based on root-point prompting, achieving fine, thin, and curly hair instance segmentation with only point-level annotations through the root-to-prompt generation task and geometric prior calibration.
RACA: Rule-Aligned Collaborative Agents for Evidence-Grounded Breast Ultrasound Classification
Chen, Lingyu (Nanjing University of Aeronautics and Astronautics), Meng, Qingjie (University of Birmingham)
ClassificationExplainability and InterpretabilityAgentic AIImageBiomedical DataUltrasound
🎯 What it does: A rule-aligned collaborative agent framework named RACA is proposed for evidence-based breast ultrasound classification, aiming to address the shortcomings of existing methods in clinical diagnostic rule alignment and decision traceability.
RaD-Seg: Region-Aware Disentangled Learning for Plan-Guided Longitudinal Tumor Segmentation in Head-and-Neck Adaptive Radiotherapy
Khor, Hee Guan (Tsinghua University), Liao, Hongen (Tsinghua University)
SegmentationConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance Imaging
🎯 What it does: This paper proposes a RaD-Seg framework for plan-based longitudinal tumor segmentation in adaptive radiotherapy for head and neck regions.
RADAR: A Multimodal Benchmark for 3D Image-Based Radiology Report Review
Sun, Zhaoyi (University of Washington), Ben Abacha, Asma (Microsoft Health AI)
TransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextMultimodalityBiomedical DataComputed TomographyBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose the RADAR multimodal benchmark, using real pre- and final abdominal CT report differences to evaluate agreement, severity, and types of edits in image-supported report editing.
RadHiera: Semantic Hierarchical Reinforcement Learning for Medical Report Generation
Du, Bodong (Hong Kong University of Science and Technology), Li, Xiaomeng (Hong Kong University of Science and Technology)
GenerationTransformerLarge Language ModelReinforcement LearningVision Language ModelImageTextBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound
🎯 What it does: Proposes the RadHiera framework, which generates radiology reports that are more aligned with medical structures using hierarchical reinforcement learning.
RADIANT-PET: Reasoning-Augmented PET/CT Lesion Segmentation with Large Language Models and Reinforcement Learning
Wang, Jiasheng (Ohio State University Comprehensive Cancer Center), Zhu, Simeng (Ohio State University Comprehensive Cancer Center)
SegmentationTransformerLarge Language ModelReinforcement LearningPrompt EngineeringImageTextBiomedical DataComputed TomographyPositron Emission TomographyRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose an RADIANT-PET framework that combines high-sensitivity voxel-level segmentation with lesion-level judgment based on large language models (LLMs), achieving fine segmentation of PET/CT lesions.
Radiogenomics-Driven Hierarchical Multimodal Stacking for Non-invasive Triage of High-Risk PKD1-Truncating Genotypes in ADPKD
Valiyev, Oybek (Kumoh National Institute of Technology), Kim, Youngwoo (Kumoh National Institute of Technology)
ClassificationExplainability and InterpretabilityConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningImageTabularBiomedical DataMagnetic Resonance Imaging
🎯 What it does: A two-stage hierarchical cascade framework was constructed, combining 3D MRI radiomics with clinical variables to achieve non-invasive stratification screening for high-risk PKD1 truncating genotypes in ADPKD;
RadSLDP: Selective Local Differential Privacy for Radiology Vision-Language Models
Zhao, Konghao (University of Southern California), Liu, Ruishan (Virginia Polytechnic Institute and State University)
Safty and PrivacyTransformerLarge Language ModelVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingComputed TomographyElectronic Health Records
🎯 What it does: Implementing selective local differential privacy perturbation on non-visual clinical context (NVCC) word embeddings for radiology vision-language models in an untrusted training environment.
RaLMPH: Reliability-Aware Learning for Multi-Pathologist Harmonization in Whole-Slide Image Classification
Hong, Sungrae (Korea Advanced Institute of Science and Technology), Yi, Mun Yong (Korea Advanced Institute of Science and Technology)
ClassificationImage HarmonizationAnomaly DetectionTransformerSupervised Fine-TuningMixture of ExpertsContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyReview/Survey Paper
🎯 What it does: Propose a reliability-aware MIL framework named RaLMPH with multi-pathologist fusion for multi-label integration and learning in WSI
RAM-Missing: Retrieval-Augmented Missing-Aware Fusion for Robust Lung Cancer Subtyping
Li, Fulin (Ocean University of China), Li, Jinxing (Ocean University of China)
ClassificationRetrievalTransformerMixture of ExpertsContrastive LearningImageMultimodalityBiomedical DataComputed TomographyRetrieval-Augmented Generation
🎯 What it does: Proposes RAM-Missing, a retrieval-enhanced, missing-aware fusion framework, to address the practical scenario of severe CT pattern missing in lung cancer subtype classification.
RAPO: Risk-Aware Anatomical Prior Optimization via Reinforcement Learning for CAC Detection in Rheumatoid Arthritis
Xu, Jiashu (Hong Kong University of Science and Technology Guangzhou), Zhu, Lei (Hong Kong University of Science and Technology Guangzhou)
Object DetectionSegmentationOptimizationReinforcement Learning from Human FeedbackTransformerSupervised Fine-TuningReinforcement LearningVision Language ModelContrastive LearningImageBiomedical DataComputed TomographyAlzheimer's Disease
🎯 What it does: Constructed the first coronary artery calcification (CAC) detection dataset specifically for patients with rheumatoid arthritis (RA), and proposed a reinforcement learning framework named RAPO for precise CAC localization and classification on non-contrast chest CT scans with low contrast and small lesions.
RareGCD: Toward Rare Disease Discovery via Generalized Category Discovery
Ma, Yuan (Japan Advanced Institute of Science and Technology), Ju, Lie (University College London)
ClassificationAnomaly DetectionRepresentation LearningDrug DiscoveryConvolutional Neural NetworkTransformerContrastive LearningImageBiomedical DataBenchmark
🎯 What it does: Propose the RareGCD framework, which utilizes the prior of known class labels to guide prototype-based sample association, combined with density-aware prototype contrastive learning and decision boundary refinement, achieving simultaneous identification of known classes and clustering of unknown rare diseases in long-tailed medical image data.
RASP: Bridging the Long-Tail Gap in Surgical Video Understanding via Retrieval-Augmented Perception
Luo, Yuxiang (Hong Kong Polytechnic University), Chen, Zhen (Sichuan University)
RecognitionTransformerVision Language ModelContrastive LearningVideoTextRetrieval-Augmented Generation
🎯 What it does: Developed the RASP framework, utilizing retrieval augmentation to address the long-tail problem in the unified understanding of surgical videos.
RatSeizure: A Benchmark and Saliency-Context Transformer for Rat Seizure Localization
Tsai, Ting Yu (State University of New York), Chang, Ming-Ching (GE HealthCare Technology and Innovation Center)
RecognitionObject DetectionAnomaly DetectionConvolutional Neural NetworkTransformerVision-Language-Action ModelDiffusion modelScore-based ModelContrastive LearningVideoBiomedical DataBenchmark
🎯 What it does: Proposed the RatSeizure dataset and the RaSeformer model for precise detection and localization of rat seizure behaviors.
RDE-Seg: Role-Disentangled Experts with Residual Routing and Anatomy Constraints for DSA Guidewire Segmentation
Zhang, Yining, Zhao, Jianhui (Wuhan University)
SegmentationTransformerPrompt EngineeringMixture of ExpertsContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasound
🎯 What it does: Proposed an RDE framework based on role-decoupled experts and residual routing for joint segmentation of guidewires and vessels in DSA images, and implemented this framework on SAM.
Real-Time 6D Ultrasound Probe Pose Tracking in Freehand Scanning
Son, Jungjae (Korea Advanced Institute of Science and Technology), Bae, Hyeon-Min (Korea Advanced Institute of Science and Technology)
Object TrackingPose EstimationConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningOptical FlowVideoPoint CloudBiomedical DataUltrasound
🎯 What it does: This paper proposes a real-time 6D ultrasound probe pose tracking framework called UPTrack based on a single camera, which can estimate the six-degree-of-freedom pose of the probe in real-time during handheld scanning.
Real-Time Hardware-Free HIFU Interference Suppression via Teacher-Student Diffusion Framework
Cai, Dejia (Hong Kong University of Science and Technology), Chen, Hao (Hong Kong University of Science and Technology)
RestorationKnowledge DistillationConvolutional Neural NetworkTransformerDiffusion modelImageBiomedical DataUltrasound
🎯 What it does: Proposed a real-time HIFU interference suppression framework based on the image domain, named mHC-Diff, which utilizes a teacher-student diffusion model to achieve interference suppression without hardware synchronization;
Real-Time Prediction of Impending Severe Fetal Distress Using Deep Learning on Fetal Heart Rate Monitoring
Seo, Donghyeon, Lee, Seung Mi (Seoul National University)
Anomaly DetectionComputational EfficiencyConvolutional Neural NetworkRecurrent Neural NetworkTime SeriesBiomedical DataElectrocardiogram
🎯 What it does: A deep learning model was developed that can predict severe fetal distress 5 minutes in advance in fetal heart rate monitoring.
Really Need Recursion? Revisiting One-Pass Registration with Multi-Level Multi-Space Contextual Guidance
Lin, Yue (Zhongguancun Academy), Meng, Mingyuan (Zhongguancun Academy)
SegmentationOptimizationComputational EfficiencyConvolutional Neural NetworkAuto EncoderContrastive LearningOptical FlowImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Propose a first-order multi-level multi-space context-guided registration network, M2CG, aiming to verify the necessity of recursion in deformable medical image registration;
Reasoning Text-to-Video Retrieval for Operating Room Clips via Action-Driven Digital Twins
Shen, Yiqing (Johns Hopkins University), Unberath, Mathias (Johns Hopkins University)
RetrievalTransformerLarge Language ModelVision Language ModelContrastive LearningVideoTextMultimodalityRetrieval-Augmented Generation
🎯 What it does: Designed the OR3 method, which uses action-driven digital twins (ActDT) to perform structured representation of operating room video clips, and achieves precise retrieval of implicit queries through LLM imagination retrieval and evidence-driven improvements.
Reasoning Trace Divergence: An Empirical Signal for Trustworthy Black-Box MLLMs in Histopathology Classification
Kazmina, Anastasiia (Korea Advanced Institute of Science and Technology), Yi, Mun Yong (Mohamed bin Zayed University of Artificial Intelligence)
ClassificationExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataChain-of-Thought
🎯 What it does: Proposed and evaluated a novel empirical uncertainty signal called Reasoning Trace Divergence (RTD) for reliability assessment of black-box multimodal large language models (MLLMs) in histopathological slide classification tasks.
Reconstructing Isotropic 3D Cervical MRI from Anisotropic Clinical Scans via Anatomy-Style Decoupled INRs
Zhang, Qi (Shanghai Jiao Tong University), Sun, Jianqi (Shanghai Jiao Tong University)
RestorationSuper ResolutionConvolutional Neural NetworkDiffusion modelNeural Radiance FieldAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Use a decoupled hybrid implicit neural representation (Hybrid INR) framework to reconstruct clinical multi-view non-uniform 2D spinal MRI into isometric 3D images.
Recover Semantics First, Generate Better: Improved Latent Modeling for 3D MRI Reconstruction and Cross-Contrast Synthesis
Chen, Yonghao (Hong Kong University of Science and Technology (Guangzhou)), Zhu, Lei (Hong Kong University of Science and Technology (Guangzhou))
RestorationGenerationData SynthesisTransformerDiffusion modelAuto EncoderContrastive LearningBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Propose a semantics-priority latent modeling framework, including Latent Harmonization Encoder, Semantic Recovery Block, and Anatomy-aware Frequency Loss, for 3D MRI reconstruction and cross-contrast synthesis.
ReDIVE: A Two-Stage Hybrid Network for Keratitis Diagnosis via Retinex-based Enhancement and Depth-Informed Vertical Embedding
Liao, Anxu (Institute of Computing Technology Chinese Academy of Sciences), Chen, Yiqiang (Institute of Computing Technology Chinese Academy of Sciences)
ClassificationConvolutional Neural NetworkTransformerContrastive LearningImageBiomedical Data
🎯 What it does: Propose ReDIVE, a two-stage hybrid network, which performs Retinex preprocessing and depth information embedding on IVCM image sequences to achieve the diagnosis of infectious keratitis.
Reducing Expert Annotation Effort for CTA Vessel Delineation with Synthetic Pretraining
Bracci, Jacopo (Friedrich-Alexander-Universität Erlangen-Nürnberg), Maier, Andreas (Siemens Healthineers AG)
SegmentationData SynthesisConvolutional Neural NetworkSupervised Fine-TuningContrastive LearningBiomedical DataComputed Tomography
🎯 What it does: This study proposes a two-stage deep learning pipeline that uses synthetic pre-training combined with a small amount of manually corrected CTA vascular lumen and wall segmentation.
Reference-Guided Gradient Balancing Loss for Long-Tailed Multi-label Medical Image Classification
Lin, Ying-Chih (National Yang Ming Chiao Tung University), Chen, Yong-Sheng (National Yang Ming Chiao Tung University)
ClassificationConvolutional Neural NetworkSupervised Fine-TuningContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasoundElectronic Health Records
🎯 What it does: Propose the Reference-Guided Gradient Balancing (RGB) loss to address the long-tailed distribution problem in multi-label classification of medical images.
Refining 3D Medical Segmentation with Verbal Instruction
Xie, Kangxian (University at Buffalo), Gao, Mingchen (University at Buffalo)
SegmentationData SynthesisConvolutional Neural NetworkTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningTextPoint CloudBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Constructed a dataset called CoWTalk and proposed an iterative refinement model for 3D medical segmentation based on language instructions, which can gradually correct initial segmentation results according to verbal correction instructions from radiologists.
Reformulating Chest X-Ray Report Generation as Fine-Grained Visual Question Answering with Structured Clinical Report Templates
Xiao, Ziqin (Northwestern Polytechnical University), Xia, Yong (Wenzhou Medical University)
GenerationData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelImageTextBiomedical DataComputed TomographyElectronic Health Records
🎯 What it does: Reformulate the chest X-ray report generation task as a fine-grained template-based visual question answering task, decomposing reports into clinical attributes using a tree-structured template, and constructing dense supervision from raw text through a two-stage process
RefTr: Recurrent Refinement of Confluent Trajectories for 3D Tubular Tree Centerlines
Naeem, Roman (Chalmers University of Technology), Kahl, Fredrik (Chalmers University of Technology)
SegmentationComputational EfficiencyRepresentation LearningTransformerReinforcement LearningAuto EncoderContrastive LearningImageBiomedical DataComputed Tomography
🎯 What it does: Propose RefTr, a Producer-Refiner framework based on Transformer, which extracts the centerline maps of tree-like lumens such as blood vessels or airways from 3D CT images by recursively refining confluent trajectories.
Reliability-Aware Cross-Prompt Aggregation for Propagation-Based Segmentation of Arbitrary 3D Medical Objects
Li, Haoshen (Peking University), Zhang, Li (Peking University)
SegmentationTransformerPrompt EngineeringContrastive LearningMultimodalityBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: This paper proposes an untrained reliability-aware cross-prompt aggregation framework to improve propagation-based 3D medical image segmentation, reducing error accumulation and enhancing adaptability to different anatomical structures.
ReLiF-3D: Prior-Guided Semi-supervised 3D MRI Segmentation via Robust Bias-Consistent Paired Views
Jangid, Kunal (Indian Institute of Science Education & Research), Kurmi, Vinod (Indian Institute of Science Education & Research)
SegmentationConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelGenerative Adversarial NetworkContrastive LearningBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Propose a semi-supervised 3D MRI segmentation framework ReLiF-3D based on the frozen SAM-Med3D prior, using a lightweight U-Net for training and generating view pairs that are consistent with bias through Smooth Orthogonal Bias Field (SOBF). It combines Confidence-Gated Consistency (CGCR) and Lesion-Aware Representation Consistency (LARC) to suppress the mispropagation of pseudo labels.
ReMeDI: Refined Memory for Disambiguation of Identities with SAM3 in Surgical Segmentation
Bundele, Valay (University of Tübingen), Lensch, Hendrik P. A. (University of Tübingen)
Object TrackingSegmentationTransformerContrastive LearningOptical FlowVideoBiomedical Data
🎯 What it does: Proposes ReMeDI-SAM3, a training-agnostic extension that improves SAM3's tool segmentation in surgical videos, addressing issues of occlusion, long-term tracking, and identity recovery.
ReMiX-Seg: Latent Modality Completion Meets Expert Routing for Robust Glioma Segmentation
Nghiem, Van Quang (National Taiwan University), Lin, Che (National Taiwan University)
SegmentationConvolutional Neural NetworkTransformerMixture of ExpertsAuto EncoderMultimodalityBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Propose the ReMiX-Seg framework for glioma segmentation under any missing magnetic resonance modality;
Remixing Embeddings for Patient Augmentation in Data Scarce Multiple Instance Learning
Dasdelen, Muhammed Furkan (Helmholtz Munich), Sadafi, Ario (Helmholtz Munich)
Data SynthesisAnomaly DetectionRepresentation LearningTransformerContrastive LearningGaussian SplattingImageTabularBiomedical DataAlzheimer's DiseaseElectronic Health RecordsReview/Survey Paper
🎯 What it does: Propose the RECIPE method, which uses an unsupervised Gaussian Mixture Model (GMM) to cluster instances in the embedding space, statistically captures the 'recipe' of different diseases, and then samples new patient bags (patients) from the existing instance pool based on these recipes for offline data augmentation.
RePCM: Region-Specific and Phenotype-Adaptive Bi-ventricular Cardiac Motion Synthesis
Yang, Xuan (National University of Singapore), Li, Lei (National University of Singapore)
GenerationData SynthesisTransformerMixture of ExpertsDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningMeshBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Propose a full-cycle cardiac motion synthesis framework named RePCM based on single-frame end-diastolic meshes, used to generate three-dimensional time series of human ventricles.
ReportX: The BraTS Clinical Report Dataset
Marchesini, Kevin (University of Modena and Reggio Emilia), Bolelli, Federico (University of Modena and Reggio Emilia)
SegmentationConvolutional Neural NetworkTransformerPrompt EngineeringVision Language ModelAuto EncoderContrastive LearningImageTextBiomedical DataMagnetic Resonance Imaging
🎯 What it does: A paired dataset called ReportX was constructed on the BraTS-GLI2023 dataset, combining structured clinical reports written by expert physicians with automatically generated quantitative fields, and these reports were used as auxiliary semantic supervision for 3D brain tumor segmentation.
Representation Learning with Distance-Residualized Functional Connectivity for Cortical Parcellation from rs-fMRI
Zhu, Jianfei (Harbin Institute of Technology), Yi, Chunzhi (Harbin Institute of Technology)
Representation LearningGraph Neural NetworkAuto EncoderContrastive LearningBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Propose a self-supervised functional connectivity representation learning framework, utilizing distance residualization, variational autoencoder, neighborhood contrastive learning, and Laplacian position encoding to learn integrated cortical functional parcellation from rs-fMRI data.
ReScan-IA: A Spatially-Adaptive Diffusion Framework for Controllable 3D Intracranial Aneurysm Inpainting
Chuang, Tzu I (Charité - Universitätsmedizin Berlin), Hilbert, Adam (Charité - Universitätsmedizin Berlin)
RestorationSegmentationGenerationData SynthesisConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningImageBiomedical DataComputed Tomography
🎯 What it does: This paper proposes a spatially adaptive 3D diffusion model, ReScan-IA, for controllably synthesizing/filling intracranial aneurysms in CTA imaging, simulating the appearance of aneurysms during repeated scans.
Research Design Considerations for Empirical User Studies in MICCAI
Cho, Sue Min (Johns Hopkins University), Unberath, Mathias (Johns Hopkins University)
Biomedical DataReview/Survey Paper
🎯 What it does: This paper summarizes and proposes design considerations and best practices for controlled user studies in the field of medical image computing and computer-aided surgery, including guidance on study objectives, technical system design, tasks and interfaces, subject recruitment, and measurement evaluation.
Residual Diffusion Bridge in Wavelet Space for Medical Image Translation
Yang, Xiao (Newcastle University), Zhang, Jingjing (Newcastle University)
Image TranslationRestorationExplainability and InterpretabilityConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Propose a Wavelet Residual Bridge (WRB) framework for cross-modal MRI translation, first mapping low-frequency structures through dual-tree complex wavelet transform, then performing random sampling in the high-frequency texture subspace using Gaussian bridge, and finally reconstructing the target image.
Residual-Aware Structured Pruning with Dual-Stream Alignment for Adapter-Tuned SAM
Luo, Xin (University of Science and Technology of China), Zhou, Shaohua Kevin (University of Science and Technology of China)
SegmentationComputational EfficiencyKnowledge DistillationTransformerAuto EncoderBiomedical DataReview/Survey Paper
🎯 What it does: Proposes a structured pruning framework called RASP-SAM for Adapter-tuned SAM (e.g., SAM-Med2D), which can significantly compress the model while preserving pre-trained and downstream knowledge.
Resolution-Conditioned Representation Learning for Unified Brain MRI Modeling
Dong, Ke (ShanghaiTech University), Sun, Kaicong (ShanghaiTech University)
ClassificationSegmentationRepresentation LearningConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningMultimodalityBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Proposed and implemented a resolution-conditioned representation learning framework that utilizes self-supervised reconstruction and multi-modal metadata alignment to train feature representations that can uniformly process brain MRI across different slice thicknesses, with applications in brain tissue segmentation, ASD/MDD diagnosis, etc.
Response-Aware Multimodal Learning for Post-treatment Visual Acuity Forecasting
Bui, Phuoc-Nguyen, Choo, Hyunseung
Explainability and InterpretabilityComputational EfficiencyData-Centric LearningTabularElectronic Health Records
🎯 What it does: This paper conducts a multiple linear regression analysis on a public dataset to explore the relationships between variables;
Rest2Visual: Predicting Visually Evoked fMRI from Resting-State Scans
Zhou, Chuyang (University of Sydney), Xu, Chang (University of Sydney)
RestorationGenerationData SynthesisConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Predict the corresponding visually induced fMRI activation maps by conditioning individual resting-state fMRI on visual stimulus images.
Rethinking RAG: Mitigating Text-Driven Hallucinations via Feature-Level Retrieval Integration in Medical LVLMs
Kim, Yujoong (Pohang University of Science and Technology), Park, Sang Hyun (Daegu Gyeongbuk Institute of Science and Technology)
RetrievalExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringVision Language ModelAuto EncoderContrastive LearningImageTextBiomedical DataBenchmarkRetrieval-Augmented Generation
🎯 What it does: This paper proposes a framework that directly injects retrieved image-text pairs into the visual encoder of a large medical audio-visual model to reduce text-driven hallucinations.
Rethinking the Adaptation of Vision Foundation Models for Efficient Cell Segmentation
Xu, Qing (University of Nottingham), Chen, Zhen (University of Lincoln)
SegmentationConvolutional Neural NetworkTransformerPrompt EngineeringAuto EncoderContrastive LearningImageBiomedical Data
🎯 What it does: Propose the EffiCell-Seg framework, which utilizes a frozen Vision Foundation Model and performs lightweight learning only on structural prompts and the decoder to achieve efficient cell segmentation.
Retinopathy of Prematurity Staging via Prototype-Evidential Learning
Wu, Donghan (Ningbo Institute of Materials Technology and Engineering, Chinese Academy of Sciences), Zhao, Yitian (Ningbo Institute of Materials Technology and Engineering, Chinese Academy of Sciences)
ClassificationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasoundFibre Orientation DistributionDiffusion Tensor ImagingAlzheimer's DiseaseElectronic Health RecordsElectrocardiogram
🎯 What it does: This paper proposes ProtoEvi-ROP, a dual-branch network combining prototype learning with evidence theory, for early-stage ROP (retinopathy of prematurity) staging diagnosis.
Retrieval-Augmented Anatomical Guidance for Text-to-CT Generation
Molino, Daniele (Università Campus Bio-Medico di Roma), Guarrasi, Valerio (Università Campus Bio-Medico di Roma)
GenerationData SynthesisRetrievalTransformerVision Language ModelDiffusion modelAuto EncoderMultimodalityBiomedical DataComputed TomographyRetrieval-Augmented Generation
🎯 What it does: This paper proposes a retrieval-enhanced text-to-CT generation framework, which utilizes the case structure retrieved based on semantic similarity as a proxy to inject into ControlNet to achieve anatomy-guided generation.
Retrieval-Based Mixture-of-Experts for Patient-Specific Cancer Survival Prediction with Incomplete Multimodal Data
Lim, Minjoo (Korea University), Kam, Tae-Eui (Korea University)
RetrievalExplainability and InterpretabilityComputational EfficiencyDrug DiscoveryTransformerMixture of ExpertsContrastive LearningMultimodalityBiomedical DataElectronic Health RecordsRetrieval-Augmented Generation
🎯 What it does: Proposed a Retrieval-based Mixture-of-Experts (RMoE) model for patient-specific cancer survival prediction in the presence of missing multimodal (WSI and gene expression) data;
Retrieval-Guided Semi-supervised Federated Learning for Medical Image Segmentation
Park, Heejung (Daegu Gyeongbuk Institute of Science and Technology), Park, Sang Hyun (Pohang Unversity of Science and Technology)
SegmentationDomain AdaptationFederated LearningConvolutional Neural NetworkTransformerContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasoundRetrieval-Augmented Generation
🎯 What it does: Proposed a retrieval-reference-guided semi-supervised federated learning framework, R2S2Fed, for medical image segmentation.
REVEAL++: Differentiable Phenotypic Grouping for Vision–Language Retinal Modeling of Alzheimer’s Disease Risk
Meidinger, Ethan (University of Virginia), Fang, Ruogu (University of Florida)
ClassificationExplainability and InterpretabilityRepresentation LearningConvolutional Neural NetworkTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataAlzheimer's DiseaseElectronic Health RecordsRetrieval-Augmented Generation
🎯 What it does: This paper proposes REVEAL++, a differentiable phenotype-weighted audio-visual language alignment framework, for predicting Alzheimer's disease risk through retinal images and clinical risk narratives.
Revisiting Integration of Image and Metadata for DICOM Series Classification: Cross-Attention and Dictionary Learning
Truong, Tuan (Bayer AG), Lenga, Matthias (Bayer AG)
ClassificationConvolutional Neural NetworkTransformerContrastive LearningMultimodalityBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Propose an end-to-end multi-modal framework for automatically identifying DICOM image series, jointly modeling image content and acquisition metadata to address issues of missing, diverse, and heterogeneous data.
Reward-Guided Distillation: A Progressive Pseudo-Bag Purification Framework for WSI Multiple Instance Learning
Jia, Qi (Dalian University of Technology), Fan, Xin (Dalian University of Technology)
ClassificationExplainability and InterpretabilityKnowledge DistillationRepresentation LearningConvolutional Neural NetworkTransformerReinforcement LearningContrastive LearningImageBiomedical DataReview/Survey PaperBenchmark
🎯 What it does: Proposes a hierarchical pseudo-bag purification framework, utilizing Progressive Shapley ranking to construct pseudo-bags and filtering noise through the Reward-Guided Selection Module, thereby enhancing sparse classification performance in weakly supervised multiple instance learning (WSI-MIL).
Riemannian Batch Normalization on Correlation Manifolds for EEG Decoding
Yang, Jiarui, Wu, Xiao-Jun (Jiangnan University)
ClassificationRepresentation LearningBiomedical Data
🎯 What it does: A Correlation Batch Normalization (CorBN) layer is proposed for the Riemannian geometry of full-rank correlation matrices in EEG decoding, achieving geometrically consistent normalization on the correlation matrix manifold.
Right Answer, Wrong Evidence: Revealing Visual Confabulation in Multi-Image Medical VLMs with the Evidence Attribution Score
Kuppani, Chidwipak (Indian Institute of Information Technology Sri City), Halavar, Bheemappa (Indian Institute of Information Technology Sri City)
Explainability and InterpretabilityPrompt EngineeringVision Language ModelImageTextBiomedical DataBenchmarkChain-of-Thought
🎯 What it does: Propose the Evidence Attribution Score (EAS), which measures whether the model truly relies on the correct visual evidence when providing the correct answer, through leave-one-out ablation and multiple prompting on multi-image medical question answering.
Robotic Ultrasound Makes CBCT Alive
Li, Feng (Technical University of Munich), Bi, Yuan (Technical University of Munich)
Robotic IntelligenceConvolutional Neural NetworkTransformerSupervised Fine-TuningAuto EncoderGenerative Adversarial NetworkContrastive LearningOptical FlowImageVideoBiomedical DataComputed TomographyUltrasound
🎯 What it does: Designed and implemented a real-time CBCT update framework based on robotic ultrasound, which estimates soft tissue deformation using ultrasound sequences and applies this deformation in real-time to static CBCT slices, achieving dynamic navigation without repeated radiation exposure.