arXivSub Start free trial

MICCAI 2026 Papers with Code — Page 5

International Conference on Medical Image Computing and Computer-Assisted Intervention · 649 papers

NoiBox: Safeguard Frontier via Reverse Adversarial Diffusion for Noisy-Box-Supervised 3D Tumor Segmentation

Ji, Kailun (Wuhan University of Technology), Zhou, Quan (Zhongnan Hospital of Wuhan University)

CodeSegmentationTransformerDiffusion modelGenerative Adversarial NetworkContrastive LearningBiomedical DataComputed Tomography

🎯 What it does: This paper proposes a method for 3D tumor segmentation using rough 3D bounding boxes (noisy boxes), constructing a safe frontier soft supervision target to avoid missed and mislabeled annotations caused by hard box constraints.

Non-linear INR-Based Motion Modeling for 4D Radiotherapy

Gebauer, Johannes B. (University Medical Center Hamburg-Eppendorf), Werner, René (University Medical Center Hamburg-Eppendorf)

CodeOptimizationRepresentation LearningNeural Radiance FieldContrastive LearningOptical FlowImageBiomedical DataComputed Tomography

🎯 What it does: Propose an implicit neural representation (INR) model based on SIREN to achieve patient-specific nonlinear respiratory motion modeling;

Non-parametric Prototypes Enable Semantic Graph Learning for Whole-Slide Pathology

Gao, Zixuan (East China Normal University), Li, Qingli (East China Normal University)

CodeClassificationAnomaly DetectionComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerVision Language ModelContrastive LearningImageTextBiomedical Data

🎯 What it does: Proposed a non-parametric prototype-guided semantic graph learning network (ProtoSG-Net) for classifying whole slide images (WSI);

Non-Stationary Spectral State-Space Networks with Region-Wise Frequency Routing for Robust Fundus Vessel Segmentation

Viriyasaranon, Thanaporn (Ewha Womans University), Choi, Jang-Hwan (Ewha Womans University)

CodeSegmentationDomain AdaptationRecurrent Neural NetworkTransformerMixture of ExpertsVision-Language-Action ModelAuto EncoderContrastive LearningImageBiomedical Data

🎯 What it does: Proposed an NS Net network based on explicit frequency-domain hybrid experts for robust retinal vessel segmentation and artery-vein separation.

NR-Align: Non-rigid Alignment for Non-simultaneous Two-View 3D Coronary Reconstruction in Complex Cardiac Interventions

Wang, Pengbo (Beijing University of Posts and Telecommunications), Di, Chunxia (Beijing University of Posts and Telecommunications)

CodeSegmentationData SynthesisConvolutional Neural NetworkDiffusion modelAuto EncoderContrastive LearningOptical FlowImageBiomedical DataComputed TomographyElectrocardiogram

🎯 What it does: Propose the NR-Align module, performing non-rigid alignment on asynchronous dual-view DSA backprojection volumes, thereby improving the connectivity and topological consistency of 3D coronary artery reconstruction.

Omni-Sleep: A Sleep Foundation Model via Hierarchical Contrastive Learning of CNS–ANS Dynamics

Hou, Zhoujie (Southern University of Science and Technology), Liu, Quanying (Southern University of Science and Technology)

CodeClassificationRepresentation LearningConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningMultimodalityTime SeriesBiomedical DataElectrocardiogram

🎯 What it does: This paper proposes Omni-Sleep, a sleep foundation model based on the hierarchical alignment of the central nervous system (CNS) and the autonomic nervous system (ANS), which performs unsupervised pre-training on multi-modal PSG using hierarchical contrastive learning and long-time-slot latent mask modeling.

On the Behavior of Calibration Error Under Rater Disagreement

Kumar, Harshit (Whiterabbit.ai), Matthews, Thomas Paul (Whiterabbit.ai)

CodeClassificationAnomaly DetectionExplainability and InterpretabilityBiomedical DataReview/Survey Paper

🎯 What it does: This paper investigates the problem of the deviation between calibration error (ECE) and majority voting labels in medical AI evaluation under multiple annotator uncertainty, and proposes two calibration schemes (MR-ECE and AdjustedECE) to restore oracle guarantee.

Ontology-Grounded Structured Prediction for Dental CBCT Reporting

Lumetti, Luca (University of Modena and Reggio Emilia), Bolelli, Federico (University of Modena and Reggio Emilia)

CodeObject DetectionSegmentationRepresentation LearningConvolutional Neural NetworkTransformerContrastive LearningImageTextBiomedical DataComputed TomographyRetrieval-Augmented Generation

🎯 What it does: Constructed an ontology-based structured report resource based on dental CBCT, providing 893 clinical reports, 13 types of findings and their attributes annotated with RDF/OWL, and proposed a three-stage structured prediction framework.

Open-PMC-18M: A High-Fidelity Large Scale Medical Dataset for Multimodal Representation Learning

Baghbanzadeh, Negin (Vector Institute), Afkanpour, Arash

CodeClassificationRetrievalRepresentation LearningGraph Neural NetworkTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBiomedical DataBenchmarkRetrieval-Augmented Generation

🎯 What it does: Constructed a high-fidelity medical vision-language dataset, Open-PMC-18M, consisting of 18 million subfigure-subtitle + context abstract pairs, and proposed a scalable subfigure detection and text enhancement pipeline.

OpenRC: An Open-Source Robotic Colonoscopy Framework for Multimodal Data Acquisition and Autonomy Research

Kapuria, Siddhartha (University of Texas at Austin), Alambeigi, Farshid (Technical University of Munich)

CodeRobotic IntelligenceSimultaneous Localization and MappingOptical FlowVideoMultimodalityBiomedical Data

🎯 What it does: Proposed and implemented a low-cost, scalable open-source robotic colonoscopy platform called OpenRC, which provides multi-degree-of-freedom mechanical actuation to traditional colonoscopes while preserving clinical workflows, and simultaneously records video, operator commands, mechanical status, and EM tracking of the end-effector pose;

OPGAgent: An Agent for Auditable Dental Panoramic X-ray Interpretation

Yu, Zhaolin (Monash University), Ge, Zongyuan (Monash University)

CodeAnomaly DetectionExplainability and InterpretabilityTransformerLarge Language ModelAgentic AIMixture of ExpertsVision Language ModelImageTextBiomedical DataComputed TomographyBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This paper proposes OPGAgent, a multi-tool intelligent agent for auditable全景牙片 (panoramic radiograph) interpretation, achieving multi-task diagnosis through hierarchical evidence collection, a specialized toolset, and a consensus sub-agent.

OphthFlowBench: A Unified Ophthalmology Multimodal Benchmark with Workflow-Aligned Evaluation

Zheng, Qiaojian (Shenzhen University), Shen, Linlin (Shenzhen Second People's Hospital)

CodeClassificationRecognitionAnomaly DetectionExplainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision Language ModelImageTextMultimodalityBiomedical DataBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Constructed OphthFlowBench — a unified multi-modal ophthalmology evaluation benchmark, and trained a specialized ophthalmology multi-modal large language model, Oph-RL, based on this benchmark.

Opportunistic Cardiac Health Assessment: Estimating Phenotypes from Localizer MRI Through Multi-modal Representations

Zeybek, Busra Nur (Technical University of Munich), Kafali, Sevgi Gokce (Technical University of Munich)

CodeRepresentation LearningData-Centric LearningTransformerAuto EncoderContrastive LearningImageMultimodalityTabularBiomedical DataMagnetic Resonance ImagingElectronic Health RecordsElectrocardiogram

🎯 What it does: This paper proposes C-TRIP, a multimodal framework that jointly learns a latent space using fast, low-quality local MRI, electrocardiogram (ECG), and clinical tabular data, and predicts cardiac phenotypes during inference using only local MRI.

OPS: One-Shot Point Prompted Semantics-Aware Learning for Medical Image Registration

Xie, Housheng (Shanghai Jiao Tong University), Zheng, Guoyan (Shanghai Jiao Tong University)

CodeSegmentationTransformerPrompt EngineeringDiffusion modelContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Proposed the OPS framework, which utilizes single-point prompts and the pre-trained DINOv3 semantic clustering capability to generate dense semantic similarity maps across the entire dataset, and embeds them into an unsupervised medical image registration network to achieve semantic-aware registration.

Ordinal Diffusion Models for Color Fundus Images

Schmidt, Gustav (University of Tübingen), Müller, Sarah (University of Tübingen)

CodeGenerationData SynthesisTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningImageBiomedical Data

🎯 What it does: Proposed an ordinal diffusion model capable of generating realistic color fundus images based on diabetic retinopathy (DR) grading.

Ordinal Priors for Colonoscopy Temporal Segmentation

Jang, Seunghyun (Seoul National University), Park, Chang Min (Seoul National University Hospital)

CodeSegmentationComputational EfficiencyConvolutional Neural NetworkTransformerContrastive LearningVideoBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Proposed a colonoscopy temporal segmentation framework that considers anatomical order, combining sequence-level ordinal regression loss with an ordinal constrained decoder.

Organ-Invariant Tumor Representation Learning via Disentangled Adversarial MAE for Multi-organ Ultrasound Imaging

Xiong, Xiangyu (Macao Polytechnic University), Tan, Tao (Macao Polytechnic University)

CodeClassificationSegmentationRepresentation LearningTransformerPrompt EngineeringAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataUltrasound

🎯 What it does: Proposed a Decoupled Adversarial Masked Autoencoder (DA-MAE), which learns organ-agnostic tumor representations in multi-organ ultrasound images through a dual-branch encoder, and uses text prompts to assist decoding.

OrthoFlow: Decoupled Landmark Prediction and Image Synthesis for Orthodontic Treatment Outcome Visualization

Wang, Leyuan (ShanghaiTech University), Cui, Zhiming (ShanghaiTech University)

CodeImage TranslationGenerationPose EstimationGraph Neural NetworkTransformerPrompt EngineeringDiffusion modelScore-based ModelFlow-based ModelImageText

🎯 What it does: Designed the OrthoFlow two-stage decoupled framework for predicting geometric deformation and image synthesis of lateral cephalograms after orthodontic treatment.

OSCAR: Occupancy-Based Shape Completion via Acoustic Neural Implicit Representations

Wysocki, Magdalena (Technical University of Munich), Navab, Nassir (Technical University of Munich)

CodeRestorationDiffusion modelScore-based ModelNeural Radiance FieldAuto EncoderGenerative Adversarial NetworkContrastive LearningOptical FlowImageBiomedical DataComputed TomographyUltrasound

🎯 What it does: Propose the OSCAR framework for spine shape completion based on ultrasound, which directly recovers complete 3D geometry from B-mode images by using an implicit network that couples acoustic information with the occupancy field.

OTCHA: Optimal Transport-Driven Confidence-Aware Latent Hub Alignment for Multi-view Medical Image Classification

Yang, Jiwoong (Hanyang University), Jang, Ikbeom (Hankuk University of Foreign Studies)

CodeClassificationImage TranslationDomain AdaptationOptimizationKnowledge DistillationRepresentation LearningConvolutional Neural NetworkTransformerContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound

🎯 What it does: Propose the OTCHA module, which refines the patch tokens from each perspective first through optimized transport matching and then fuses information on a shared latent hub before fusion in multi-view medical image classification.

PancCADx: A Multimodal Framework for Pancreatic Cancer Diagnosis

Hu, Shan (Wuhan University), Wang, Zhongyuan (Wuhan University)

CodeClassificationExplainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningVision Language ModelImageMultimodalityTabularBiomedical DataUltrasoundChain-of-Thought

🎯 What it does: A explainable multimodal framework called PancCADx was constructed, using Qwen3‑VL‑8B‑Thinking combined with chain-of-thought (CoT) reasoning to diagnose pancreatic cancer and integrate it with clinical data.

Parameter-Efficient Self-Supervised Adaptation for EEG-FM under Fixed Computational Budgets

Dani, Meghal (University of Tübingen), Liebe, Stefanie (University of Tübingen)

CodeDomain AdaptationAnomaly DetectionComputational EfficiencyRepresentation LearningTransformerSupervised Fine-TuningContrastive LearningTime SeriesBiomedical DataElectrocardiogram

🎯 What it does: This paper studies self-supervised learning (SSL) adaptation methods for updating only 9–29% of the parameters of the EEG foundation model (EEG-FM) in resource-constrained clinical environments, and improves the performance of downstream tasks by freezing the encoder and training only the final layer.

Patch-to-Global: Random Patch Diffusion for Globally Consistent Megapixel Artifact Inpainting in Whole Slide Images

Lee, Hyeseong (Korea University), Ahn, Sangjeong (Seoul National University)

CodeRestorationConvolutional Neural NetworkTransformerDiffusion modelContrastive LearningBiomedical Data

🎯 What it does: Proposes the RestorePath framework, which addresses large-scale (megapixel-level) artifacts in whole digital pathology slides by using random patch diffusion to achieve globally consistent restoration, avoiding high-confidence erroneous predictions by the model on artifact-containing images.

Pathologist Attention–Aligned Report Generation for Prostate Histopathology

Xue, Ruoyu (Stony Brook University), Samaras, Dimitris (Stony Brook University)

CodeGenerationExplainability and InterpretabilityTransformerSupervised Fine-TuningVision Language ModelGaussian SplattingImageTextBiomedical DataElectronic Health Records

🎯 What it does: This paper proposes incorporating pathologist visual attention supervision into the pathological report generation model, and improves the generation quality through attention alignment loss.

Patient-Conditioned Vision–Language Priors with Stable Mixture-of-Experts for Label-Efficient Alzheimer’s MRI Staging

Ding, Shan (Southeast University), Shu, Huazhong (Southeast University)

CodeClassificationRepresentation LearningData-Centric LearningTransformerMixture of ExpertsVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease

🎯 What it does: This study proposes a semi-supervised framework based on patient-conditioned vision-language prior, scale-preserving mixture-of-experts classifier, and EMA teacher for Alzheimer's disease MRI staging under few-label conditions.

PC-Seg: Progressive Cross-View Consistency for 3D OCT Segmentation from Sparse 2D Annotations

Konno, Tsubasa (Tohoku University), Aoki, Takafumi (Tohoku University)

CodeSegmentationKnowledge DistillationConvolutional Neural NetworkTransformerContrastive LearningBiomedical Data

🎯 What it does: Proposed the PC-Seg framework, which trains a 3D OCT segmentation model step-by-step through cross-view consistency using sparse 2D annotations.

PET-Adapter: Test-Time Domain Adaptation for Full and Limited-Angle PET Image Reconstruction

Yilmaz, Rüveyda (RWTH Aachen University), Schulz, Volkmar (RWTH Aachen University)

CodeRestorationDomain AdaptationTransformerDiffusion modelScore-based ModelContrastive LearningImageBiomedical DataMagnetic Resonance ImagingPositron Emission Tomography

🎯 What it does: Propose a PET-Adapter framework that can reconstruct clinical PET data through unsupervised test-time adaptation, based solely on a diffusion model pre-trained on a simulated phantom.

PhaseAT: Fourier Phase Adversarial Training for Medical Image Domain Generalization

Sharshar, Ahmed (Mohamed bin Zayed University of Artificial Intelligence), Guizani, Mohsen (Mohamed bin Zayed University of Artificial Intelligence)

CodeClassificationDomain AdaptationConvolutional Neural NetworkGenerative Adversarial NetworkContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasound

🎯 What it does: Propose PhaseAT, which enhances the domain generalization ability of medical images by utilizing phase adversarial training in the Fourier frequency domain.

PhaseGen: A Diffusion-Based Approach for Complex-Valued MRI Data Generation

Rempe, Moritz (University Hospital Essen), Kleesiek, Jens (University Hospital Essen)

CodeRestorationGenerationData SynthesisConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose a diffusion model called PhaseGen conditioned on amplitude, which generates phase from amplitude-only MRI data to synthesize complete complex k-space data, and uses these synthetic data for pre-training downstream tasks.

PhyDiCT: Plug-and-Play CT Reconstruction from Sparse X-Rays via Differentiable Rendering and Strong Priors

Dai, Weicheng (Boston University), Batmanghelich, Kayhan (Boston University)

CodeRestorationTransformerPrompt EngineeringVision Language ModelDiffusion modelScore-based ModelImageTextBiomedical DataComputed Tomography

🎯 What it does: Train freely to reconstruct 3D lung CT images from a few X-rays through differentiable rendering and strong text-guided diffusion prior.

PhyRadGS: Physics-Aware Radiative Gaussian Splatting for Fly-Scanning CBCT Reconstruction

Li, Anwei (CVTE Research), Wang, Rongqiu (CVTE Research)

CodeDiffusion modelScore-based ModelContrastive LearningGaussian SplattingImageBiomedical DataComputed Tomography

🎯 What it does: Propose a physics-aware 3D Gaussian scattering framework called PhyRadGS for reconstruction in flying scan CBCT, capable of generating clear medical images under continuous rotation and rolling shutter conditions.

Physical-Driven Unified Implicit Regularization Network with Scan Parameter Prompts for Accelerated Multi-parametric MR Imaging

Yang, Yan (Xi'an Jiaotong University), Sun, Jian (Shenzhen Institute of Advanced Technology)

CodeRestorationOptimizationComputational EfficiencyConvolutional Neural NetworkTransformerPrompt EngineeringDiffusion modelBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose a physics-driven unified implicit regularization network, PDUIR-Net, which can simultaneously complete multi-contrast image reconstruction and quantitative parameter mapping under undersampled MR data.

Physics-Based iOCT Sonification for Real-Time Interaction Awareness in Subretinal Injection

Reyes Vargas, Luis D. (Technische Universität München), Matinfar, Sasan (Technische Universität Dresden)

CodeSegmentationOptimizationComputational EfficiencyRobotic IntelligenceConvolutional Neural NetworkDiffusion modelContrastive LearningOptical FlowImageBiomedical DataPhysics RelatedAudio

🎯 What it does: Developed a real-time iOCT acoustic mapping system that uses a physics-inspired mass-spring-damper model to convert retinal layer structure, tip motion, and tissue deformation caused by injection into audible feedback, helping surgeons precisely locate and monitor bubble formation during subretinal injection.

Physics-Grounded Weakly Supervised Histopathological Tissue Segmentation via Frozen Tri-Domain Prototypes

Zhang, Shuyu (Jiangnan University), Pan, Xiang (Jiangnan University)

CodeSegmentationTransformerMixture of ExpertsContrastive LearningBiomedical DataBenchmarkPhysics Related

🎯 What it does: Propose the PhysPro weakly supervised organ segmentation framework, which utilizes physics-prior triple-domain static prototypes to stabilize feature learning, decouples spatial-frequency features, and achieves triple-state adaptive training through differential supervision.

Physics-Inspired Continuous Transformer for Fast QDSA Reconstruction

Liu, Yang (Southern University of Science and Technology), Si, Weixin (Shenzhen University of Advanced Technology)

CodeImage TranslationRestorationExplainability and InterpretabilityComputational EfficiencyTransformerDiffusion modelScore-based ModelRectified FlowAuto EncoderContrastive LearningImageTime SeriesBiomedical DataMagnetic Resonance ImagingPhysics Related

🎯 What it does: This study proposes PIC-Former, a physics-informed continuous transformer, for reconstructing continuous time density curves (TDC) from irregular DSA samples, enabling fast and robust generation of hemodynamic parameters (such as TTP, CBV, MTT) and QDSA velocity estimation, supporting intraoperative hemodynamic assessment;

Physiology-Guided Cross-View Spatial-Temporal Aligning for Early Prediction of pCR in Breast Cancer from Longitudinal Multimodal MRI

Li, Xin, Wang, Hongkai (Dalian University of Technology)

CodeClassificationImage TranslationAnomaly DetectionTransformerSupervised Fine-TuningContrastive LearningImageMultimodalityBiomedical DataMagnetic Resonance Imaging

🎯 What it does: This paper proposes the CALM framework, which uses multimodal longitudinal MRI (DCE-MRI, DWI, T2WI) to predict whether patients with breast cancer can achieve pathological complete response (pCR) early in neoadjuvant chemotherapy.

PhysOCT: Physics-Guided Retinal Modeling for OCT Image Synthesis

Wang, Yutong (Nanjing University of Science and Technology), Chen, Qiang (Nanjing University of Science and Technology)

CodeGenerationData SynthesisConvolutional Neural NetworkDiffusion modelNeural Radiance FieldOptical FlowImageBiomedical DataComputed Tomography

🎯 What it does: This paper proposes a physics-guided OCT image synthesis framework called PhysOCT, which enhances anatomical consistency in generated images by simultaneously considering lesion evolution and retinal layer deformation in conditional generation.

PINet: A Multimodal Network for Pontine Infarction Segmentation and Early Neurological Deterioration Prediction

Liu, Jinshuo, Pan, Yi (Shenzhen Institute of Advanced Technology)

CodeClassificationSegmentationConvolutional Neural NetworkTransformerMixture of ExpertsDiffusion modelAuto EncoderContrastive LearningImageTabularBiomedical DataMagnetic Resonance ImagingDiffusion Tensor Imaging

🎯 What it does: This study proposes a multi-modal multi-task network called PINet, which can simultaneously perform segmentation of pontine infarction lesions and predict early neurologic deterioration (END);

PINNOCHIO: Physics-Informed Neural Network for Coupled Hyperelastic Interface-Volume Simulation in Orthognathic Surgery

Lee, Jungwook (Rensselaer Polytechnic Institute), Yan, Pingkun (Houston Methodist Research Institute)

CodeGraph Neural NetworkDiffusion modelPoint CloudMeshBiomedical DataComputed TomographyReview/Survey PaperPhysics Related

🎯 What it does: Developed a physics-informed neural network framework called PINNOCHIO for predicting facial soft tissue deformation caused by bone repositioning in orthognathic surgery.

PlaneCycle: Training-Free 2D-to-3D Lifting of Foundation Models Without Adapters

Yu, Yinghong (ELLIS Institute Finland), Yang, Jiancheng (Aalto University)

CodeClassificationSegmentationComputational EfficiencyRepresentation LearningTransformerPrompt EngineeringVision Language ModelContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasound

🎯 What it does: The study proposes a training-free and adapter-free PlaneCycle operator that can seamlessly upgrade pre-trained 2D foundational models into 3D networks without changing the original parameters.

PoseBridgeNet: Learning a Pose-Evolving Schrödinger Bridge for Multi-Organs Segmentation in Prenatal Volumetric Ultrasound

Gan, Shushen, Gao, Yi (Shenzhen University)

CodeSegmentationPose EstimationConvolutional Neural NetworkTransformerMixture of ExpertsAuto EncoderContrastive LearningImageBiomedical DataUltrasoundStochastic Differential Equation

🎯 What it does: Propose PoseBridgeNet for the segmentation of multiple organs in the fetal abdomen using 3D ultrasound during the mid-pregnancy period, leveraging the feature evolution caused by pose variations to enhance segmentation accuracy.

Posterior-Aware Motor Phenotyping with Multimodal Imaging Validation in Parkinson’s Disease

Tirhekar, Harsh (The University of Texas at Austin), Bajaj, Chandrajit (The University of Texas at Austin)

CodeClassificationFederated LearningExplainability and InterpretabilityRepresentation LearningHyperparameter SearchData-Centric LearningDrug DiscoveryContrastive LearningGaussian SplattingMultimodalityTabularBiomedical DataMagnetic Resonance ImagingPositron Emission Tomography

🎯 What it does: This paper uses Bayesian Gaussian Mixture Model (BGMM) to cluster 29,366 PPMI visit-level MDS-UPDRS-III assessments, proposes a three-way posterior confidence classification, and verifies it through multi-modal imaging with DaTSCAN SPECT and structural MRI. Finally, the model is migrated to the BioFIND dataset for external validation.

POT-SAM3: Prompt-Only Tuning for One-Shot Electron Microscopy Segmentation

Gu, Xuncheng (Beijing University of Posts and Telecommunications), Xiao, Li (Beijing University of Posts and Telecommunications)

CodeSegmentationTransformerSupervised Fine-TuningPrompt EngineeringContrastive LearningBiomedical Data

🎯 What it does: Propose a Prompt-Only Tuning framework (POT-SAM3), which learns a few concept prompt tokens on a single annotated slice, enabling the frozen SAM3 model to achieve prompt-free detection and cross-slice propagation instance segmentation across the entire electron microscopy sequence.

PRA-PoE: Robust Multimodal Alzheimer’s Disease Classification under Arbitrary Modality Missingness

Yang, Guangqian (Hong Kong Polytechnic University), Wang, Shujun (Hong Kong Polytechnic University)

CodeClassificationRepresentation LearningData-Centric LearningConvolutional Neural NetworkTransformerMixture of ExpertsContrastive LearningMultimodalityTabularBiomedical DataMagnetic Resonance ImagingPositron Emission TomographyAlzheimer's DiseaseElectronic Health Records

🎯 What it does: Propose the PRA-PoE framework to achieve multi-modal classification of Alzheimer's disease under arbitrary missing modalities

Precision Recall Controllable Radiology Report Generation via Hybrid Natural Language and Clinical Reward Learning

Chen, Ling (Ohio State University), Wu, Dufan (Ohio State University)

CodeGenerationReinforcement Learning from Human FeedbackTransformerReinforcement LearningVision Language ModelImageTextBiomedical DataComputed TomographyElectronic Health Records

🎯 What it does: Proposed a controllable reinforcement learning framework for generating radiology reports from chest X-ray images, achieving a trade-off between clinical precision and recall through control parameters.

Predicting Radiologist Expertise from 3D Gaze Patterns During CT Interpretation

Khaertdinova, Leila (University of Copenhagen), Ibragimov, Bulat (University of Copenhagen)

CodeClassificationExplainability and InterpretabilityConvolutional Neural NetworkTransformerVision Language ModelContrastive LearningImageBiomedical DataComputed Tomography

🎯 What it does: Construct a professionalism classification model based on Transformer using eye-tracking during the CT interpretation process, integrating 3D fixation information to predict the experience level of radiologists.

Preoperative Simulation of Personalized Breast Reconstruction via Implant Prior Controllable Diffusion

Wang, Yizhi (Zhejiang University), Jin, Yaochu (Westlake University)

CodeImage TranslationRestorationSegmentationGenerationConvolutional Neural NetworkTransformerPrompt EngineeringDiffusion modelAuto EncoderContrastive LearningImageBiomedical DataComputed Tomography

🎯 What it does: Proposed a pre-surgical breast reconstruction simulation framework based on a controllable diffusion model, which can generate post-operative reconstruction results from pre-operative CT scans and support editing of implant parameters;

PriLoRA: Prior-Conditioned Low-Rank Adapters for Medical Vision Models

Kazerouni, Amirhossein (University of Toronto), Taati, Babak (University of Toronto)

CodeClassificationRestorationSegmentationGenerationTransformerPrompt EngineeringContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose Prior-Conditioned Low-Rank Adapters (PriLoRA), a parameter-efficient fine-tuning method for medical vision models that dynamically adjusts low-rank sinusoidal updates using input priors.

Prior-Anchored Debiasing for Long-Tailed Multi-Organ Pathology Report Generation

Yang, Feng (City University of Hong Kong), Chen, Ping (University of Massachusetts Boston)

CodeGenerationData-Centric LearningTransformerVision Language ModelContrastive LearningImageTextBiomedical DataElectronic Health RecordsRetrieval-Augmented Generation

🎯 What it does: This paper proposes the PriOrGen framework to address the long-tail distribution bias in multi-organ pathology report generation.

PRISM: Differentiable Analysis-by-Synthesis for Fixel Recovery in Diffusion MRI

Abouagour, Mohamed (Indiana University Bloomington), Garyfallidis, Eleftherios

CodeOptimizationDiffusion modelAuto EncoderContrastive LearningGaussian SplattingBiomedical DataMagnetic Resonance ImagingDiffusion Tensor Imaging

🎯 What it does: A differentiable analysis-synthesis framework called PRISM is constructed to end-to-end optimize multi-component microstructure models in 3D space while simultaneously correcting intensity distortions, thereby recovering fiber peak (fixel) information in diffusion magnetic resonance imaging.

Privacy-Preserving Video-Based Facial Palsy Recognition via Dynamic Facial Action Modeling

Chang, Yuance (Xi'an Jiaotong University), Xi, Wei (Xi'an Jiaotong University)

CodeRecognitionSafty and PrivacyConvolutional Neural NetworkRecurrent Neural NetworkTransformerAuto EncoderContrastive LearningImageVideo

🎯 What it does: Designed and implemented a video-based privacy-preserving facial palsy recognition framework named DP-Face, and publicly released the largest Palsy-330 video dataset.

Proactive Domain Unification for Robust Echocardiography Segmentation

Pang, Xintao (Macao Polytechnic University), Tan, Tao (Macao Polytechnic University)

CodeSegmentationDomain AdaptationTransformerDiffusion modelAuto EncoderContrastive LearningImageBiomedical DataUltrasound

🎯 What it does: Propose a proactive domain unification framework (PDU) during inference, which maps cardiac ultrasound images from different centers and devices into a source-domain-aligned style space to enhance cross-center segmentation performance.

Probabilistic Multi-rater Segmentation via Style-Aware Boundary Conditioning

Karimijafarbigloo, Sanaz (University of Regensburg), Merhof, Dorit (University of Regensburg)

CodeSegmentationConvolutional Neural NetworkDiffusion modelScore-based ModelContrastive LearningBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose a lightweight probabilistic multi-evaluator segmentation framework, which models experts' boundary preferences and generates diverse segmentation results that align with individual preferences through style-conditioned personalization and differential boundary loss.

Progressive Self-supervised Learning with Individualized Community Assignment for Brain Network Analysis

Chen, Hairui (Harbin Institute of Technology at Shenzhen), Ma, Ting (Harbin Institute of Technology at Shenzhen)

CodeClassificationAnomaly DetectionRepresentation LearningGraph Neural NetworkTransformerContrastive LearningGraphBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease

🎯 What it does: This paper proposes the BrainPICM framework, which performs representation learning and disease diagnosis on functional magnetic resonance brain networks through advanced self-supervised learning combined with personalized community assignment.

ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities

Chhetri, Aavash (NepAl Applied Mathematics and Informatics Institute for research), Bhattarai, Binod (West Virginia University)

CodeData SynthesisFederated LearningTransformerMixture of ExpertsContrastive LearningImageTextMultimodalityElectronic Health Records

🎯 What it does: A multi-modal federated learning framework called ProMoE-FL is studied for synthesizing missing modality features under missing modality scenarios, enabling feature synthesis of missing modalities between different hospitals without public data sharing.

PromptGate: Client-Adaptive Vision–Language Gating for Open-Set Federated Active Learning

Nesturi, Adea (University of Bonn), Albarqouni, Shadi (University of Bonn)

CodeAnomaly DetectionOptimizationFederated LearningTransformerPrompt EngineeringVision Language ModelContrastive LearningImageBiomedical Data

🎯 What it does: Propose PromptGate, a dynamic vision-language gating framework for open-set federated active learning;

Propagating Structural Guidance: Synthesizing Fluorescein Angiography from Fundus Images and Sparse OCT Scans

Ma, Tengfei (Southeast University), Fan, Wen (Nanjing University of Science and Technology)

CodeImage TranslationGenerationData SynthesisTransformerDiffusion modelGenerative Adversarial NetworkContrastive LearningImageMultimodalityBiomedical DataComputed TomographyReview/Survey Paper

🎯 What it does: Construct a tri-modal fundus image dataset and develop an FFA synthesis framework guided by OCT structure.

ProSyn-Net: A Reliable Prototypical Synergy Network for Imbalanced Dermatologic Multimodal Ultrasound Diagnosis

Lin, Jicheng (Tongji University), Luo, Ye (Tongji University)

CodeClassificationExplainability and InterpretabilityComputational EfficiencyConvolutional Neural NetworkTransformerMixture of ExpertsContrastive LearningImageMultimodalityBiomedical DataUltrasound

🎯 What it does: Propose the ProSyn-Net model, which achieves reliable diagnosis of dermatological ultrasound images through dynamic cross-modal attention fusion, nonlinear prototype mapping, and uncertainty-aware ensemble.

Proto-CAP: Prototype Memory Fusion and Curriculum-Adaptive Loss for Pediatric Myopia Visual Question Answering

Yan, Xu (Nankai University), Li, Tao (Tianjin Medical University Eye Hospital)

CodeRecognitionData-Centric LearningConvolutional Neural NetworkTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBiomedical DataBenchmark

🎯 What it does: Constructed the first multi-modal medical visual question answering dataset (PM-VQA) specifically focused on childhood myopia, and proposed the Proto-CAP framework to address the semantic gap and overfitting issues in few-shot medical VQA.

Prototype Instance-Semantic Disentanglement with Low-Rank Regularized Subspace Clustering for WSIs Explainable Recognition

Li, Chentao (Columbia University), Huang, Pan (Hong Kong Polytechnic University)

CodeRecognitionExplainability and InterpretabilityContrastive LearningBiomedical Data

🎯 What it does: Proposed an end-to-end prototype instance-semantic disentangled framework called PID-LRSC, which uses low-rank regularized subspace clustering to eliminate instance and semantic confusion in whole slide image (WSI) multi-instance learning, thereby improving diagnostic performance and interpretability.

Prototype Learning for Visual Field Estimation from Fundus Photography in High Myopia

Li, Guoliang, Liang, Dong (Hong Kong Polytechnic University)

CodeExplainability and InterpretabilityRepresentation LearningConvolutional Neural NetworkAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Predicting Visual Field Sensitivity in High Myopia Based on Fundus Photographs

Prototype-based Physiological Transfer Enables NCCT-only Hyperacute Stroke Tissue-Window Segmentation Under Missing Perfusion

Dan, Ying (Chinese University of Hong Kong), Tong, Raymond Kai-yu (Chinese University of Hong Kong)

CodeSegmentationConvolutional Neural NetworkTransformerSupervised Fine-TuningPrompt EngineeringContrastive LearningImageBiomedical DataComputed Tomography

🎯 What it does: Propose a prototype-based physiological information transfer framework, ProPhyT, for the segmentation of the core and penumbra in acute ischemic stroke using only NCCT images.

PseudoRET: A General Retinal Segmentation Framework with Multi-target Pseudo Embeddings

Wang, Zhonghua (Monash University), Ge, Zongyuan (Airdoc LLC)

CodeSegmentationConvolutional Neural NetworkTransformerContrastive LearningImageBiomedical DataBenchmark

🎯 What it does: Proposes PseudoRET, a semi-supervised multi-task retinal segmentation framework based on pseudo-label memory, capable of unifying fragmented, task-specific fundus image datasets.

PSP: Harnessing Position and Shape Priors for Cross-Domain Few-Shot Medical Image Segmentation

Xu, Bin (Nanjing University of Science and Technology), Zhang, Haofeng (Nanjing University of Science and Technology)

CodeSegmentationDomain AdaptationConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: This paper proposes a cross-domain few-shot medical image segmentation framework called PSP, which utilizes position and shape priors to counteract texture differences across different modalities, significantly improving segmentation performance.

Pulmonary Embolism Risk Stratification from CTPA and Medical Records: Vascular Graphs Are Not All You Need

Painchaud, Nathan (INSA-Lyon), Merveille, Odyssée (École Polytechnique)

CodeClassificationAnomaly DetectionGraph Neural NetworkTransformerSupervised Fine-TuningContrastive LearningImageGraphTabularBiomedical DataComputed TomographyElectronic Health Records

🎯 What it does: Propose an automated pipeline based on CTPA and medical records, using cardiac biomarkers and pulmonary vascular graphs for risk stratification prediction of pulmonary embolism.

QCAgent: An Agentic Framework for Quality-Controllable Pathology Report Generation from Whole Slide Image

Wang, Rundong (University of Science and Technology of China), Zhou, Shaohua Kevin (University of Science and Technology of China)

CodeGenerationRetrievalExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelAgentic AIVision Language ModelImageTextBiomedical DataElectronic Health RecordsRetrieval-Augmented Generation

🎯 What it does: Propose a quality-controllable whole slide image (WSI) pathology report generation framework called QCAgent, which automatically generates verifiable and complete diagnostic reports through a closed-loop audit-retrieval-revision process.

QG-MIL: A Gated Transformer Aggregator for Domain-Agnostic Multiple Instance Learning in Medical Imaging

Zedda, Luca (University of Cagliari), Loddo, Andrea (University of Cagliari)

CodeAnomaly DetectionTransformerImageBiomedical Data

🎯 What it does: Proposes QG-MIL — a gated Transformer aggregator that addresses the attention focusing problem in multi-instance learning.

Quality-Guided Semi-supervised Learning for Medical Image Segmentation

Abhishek, Kumar (Simon Fraser University), Hamarneh, Ghassan (Simon Fraser University)

CodeSegmentationConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningImageBiomedical Data

🎯 What it does: Propose a semi-supervised medical image segmentation framework based on quality prediction, which uses a pre-trained quality assessment network to guide the quality of unlabeled data.

Quantification of Uncertainty with Adversarial Models in Medical Image Segmentation

Jebril, Hana (Medical University of Vienna), Bogunović, Hrvoje (Medical University of Vienna)

CodeSegmentationExplainability and InterpretabilityAdversarial AttackConvolutional Neural NetworkMixture of ExpertsAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: This paper proposes a post-hoc adversarial search-based framework for uncertainty quantification in medical image segmentation, named QUAM-SM, which can identify 'fragile' regions where model predictions at the pixel level are easily flipped by adversarial perturbations, and generate reliable pixel-level uncertainty maps based on this.

Quantify, Collect, and Correct: Evidence-Driven Multi-agent Framework for Critical Evidence Refinement in Emergency Rooms

Cao, Haoyu (Harbin Institute of Technology), Wang, Wei (Harbin Institute of Technology)

CodeOptimizationExplainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AITextBiomedical DataElectronic Health RecordsRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Studied a closed-loop multi-agent framework called QCC, which prioritizes evidence as the primary optimization objective, aiming to improve the quality of evidence collection and decision-making in interactive diagnosis within emergency room scenarios.

RACA: Rule-Aligned Collaborative Agents for Evidence-Grounded Breast Ultrasound Classification

Chen, Lingyu (Nanjing University of Aeronautics and Astronautics), Meng, Qingjie (University of Birmingham)

CodeClassificationExplainability and InterpretabilityAgentic AIImageBiomedical DataUltrasound

🎯 What it does: A rule-aligned collaborative agent framework named RACA is proposed for evidence-based breast ultrasound classification, aiming to address the shortcomings of existing methods in clinical diagnostic rule alignment and decision traceability.

RADAR: A Multimodal Benchmark for 3D Image-Based Radiology Report Review

Sun, Zhaoyi (University of Washington), Ben Abacha, Asma (Microsoft Health AI)

CodeTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextMultimodalityBiomedical DataComputed TomographyBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose the RADAR multimodal benchmark, using real pre- and final abdominal CT report differences to evaluate agreement, severity, and types of edits in image-supported report editing.

RadHiera: Semantic Hierarchical Reinforcement Learning for Medical Report Generation

Du, Bodong (Hong Kong University of Science and Technology), Li, Xiaomeng (Hong Kong University of Science and Technology)

CodeGenerationTransformerLarge Language ModelReinforcement LearningVision Language ModelImageTextBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound

🎯 What it does: Proposes the RadHiera framework, which generates radiology reports that are more aligned with medical structures using hierarchical reinforcement learning.

RADIANT-PET: Reasoning-Augmented PET/CT Lesion Segmentation with Large Language Models and Reinforcement Learning

Wang, Jiasheng (Ohio State University Comprehensive Cancer Center), Zhu, Simeng (Ohio State University Comprehensive Cancer Center)

CodeSegmentationTransformerLarge Language ModelReinforcement LearningPrompt EngineeringImageTextBiomedical DataComputed TomographyPositron Emission TomographyRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose an RADIANT-PET framework that combines high-sensitivity voxel-level segmentation with lesion-level judgment based on large language models (LLMs), achieving fine segmentation of PET/CT lesions.

RadSLDP: Selective Local Differential Privacy for Radiology Vision-Language Models

Zhao, Konghao (University of Southern California), Liu, Ruishan (Virginia Polytechnic Institute and State University)

CodeSafty and PrivacyTransformerLarge Language ModelVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingComputed TomographyElectronic Health Records

🎯 What it does: Implementing selective local differential privacy perturbation on non-visual clinical context (NVCC) word embeddings for radiology vision-language models in an untrusted training environment.

RAM-Missing: Retrieval-Augmented Missing-Aware Fusion for Robust Lung Cancer Subtyping

Li, Fulin (Ocean University of China), Li, Jinxing (Ocean University of China)

CodeClassificationRetrievalTransformerMixture of ExpertsContrastive LearningImageMultimodalityBiomedical DataComputed TomographyRetrieval-Augmented Generation

🎯 What it does: Proposes RAM-Missing, a retrieval-enhanced, missing-aware fusion framework, to address the practical scenario of severe CT pattern missing in lung cancer subtype classification.

RAPO: Risk-Aware Anatomical Prior Optimization via Reinforcement Learning for CAC Detection in Rheumatoid Arthritis

Xu, Jiashu (Hong Kong University of Science and Technology Guangzhou), Zhu, Lei (Hong Kong University of Science and Technology Guangzhou)

CodeObject DetectionSegmentationOptimizationReinforcement Learning from Human FeedbackTransformerSupervised Fine-TuningReinforcement LearningVision Language ModelContrastive LearningImageBiomedical DataComputed TomographyAlzheimer's Disease

🎯 What it does: Constructed the first coronary artery calcification (CAC) detection dataset specifically for patients with rheumatoid arthritis (RA), and proposed a reinforcement learning framework named RAPO for precise CAC localization and classification on non-contrast chest CT scans with low contrast and small lesions.

RareGCD: Toward Rare Disease Discovery via Generalized Category Discovery

Ma, Yuan (Japan Advanced Institute of Science and Technology), Ju, Lie (University College London)

CodeClassificationAnomaly DetectionRepresentation LearningDrug DiscoveryConvolutional Neural NetworkTransformerContrastive LearningImageBiomedical DataBenchmark

🎯 What it does: Propose the RareGCD framework, which utilizes the prior of known class labels to guide prototype-based sample association, combined with density-aware prototype contrastive learning and decision boundary refinement, achieving simultaneous identification of known classes and clustering of unknown rare diseases in long-tailed medical image data.

RASP: Bridging the Long-Tail Gap in Surgical Video Understanding via Retrieval-Augmented Perception

Luo, Yuxiang (Hong Kong Polytechnic University), Chen, Zhen (Sichuan University)

CodeRecognitionTransformerVision Language ModelContrastive LearningVideoTextRetrieval-Augmented Generation

🎯 What it does: Developed the RASP framework, utilizing retrieval augmentation to address the long-tail problem in the unified understanding of surgical videos.

RatSeizure: A Benchmark and Saliency-Context Transformer for Rat Seizure Localization

Tsai, Ting Yu (State University of New York), Chang, Ming-Ching (GE HealthCare Technology and Innovation Center)

CodeRecognitionObject DetectionAnomaly DetectionConvolutional Neural NetworkTransformerVision-Language-Action ModelDiffusion modelScore-based ModelContrastive LearningVideoBiomedical DataBenchmark

🎯 What it does: Proposed the RatSeizure dataset and the RaSeformer model for precise detection and localization of rat seizure behaviors.

RDE-Seg: Role-Disentangled Experts with Residual Routing and Anatomy Constraints for DSA Guidewire Segmentation

Zhang, Yining, Zhao, Jianhui (Wuhan University)

CodeSegmentationTransformerPrompt EngineeringMixture of ExpertsContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasound

🎯 What it does: Proposed an RDE framework based on role-decoupled experts and residual routing for joint segmentation of guidewires and vessels in DSA images, and implemented this framework on SAM.

Reasoning Trace Divergence: An Empirical Signal for Trustworthy Black-Box MLLMs in Histopathology Classification

Kazmina, Anastasiia (Korea Advanced Institute of Science and Technology), Yi, Mun Yong (Mohamed bin Zayed University of Artificial Intelligence)

CodeClassificationExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataChain-of-Thought

🎯 What it does: Proposed and evaluated a novel empirical uncertainty signal called Reasoning Trace Divergence (RTD) for reliability assessment of black-box multimodal large language models (MLLMs) in histopathological slide classification tasks.

Reconstructing Isotropic 3D Cervical MRI from Anisotropic Clinical Scans via Anatomy-Style Decoupled INRs

Zhang, Qi (Shanghai Jiao Tong University), Sun, Jianqi (Shanghai Jiao Tong University)

CodeRestorationSuper ResolutionConvolutional Neural NetworkDiffusion modelNeural Radiance FieldAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Use a decoupled hybrid implicit neural representation (Hybrid INR) framework to reconstruct clinical multi-view non-uniform 2D spinal MRI into isometric 3D images.

Recover Semantics First, Generate Better: Improved Latent Modeling for 3D MRI Reconstruction and Cross-Contrast Synthesis

Chen, Yonghao (Hong Kong University of Science and Technology (Guangzhou)), Zhu, Lei (Hong Kong University of Science and Technology (Guangzhou))

CodeRestorationGenerationData SynthesisTransformerDiffusion modelAuto EncoderContrastive LearningBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose a semantics-priority latent modeling framework, including Latent Harmonization Encoder, Semantic Recovery Block, and Anatomy-aware Frequency Loss, for 3D MRI reconstruction and cross-contrast synthesis.

Reference-Guided Gradient Balancing Loss for Long-Tailed Multi-label Medical Image Classification

Lin, Ying-Chih (National Yang Ming Chiao Tung University), Chen, Yong-Sheng (National Yang Ming Chiao Tung University)

CodeClassificationConvolutional Neural NetworkSupervised Fine-TuningContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasoundElectronic Health Records

🎯 What it does: Propose the Reference-Guided Gradient Balancing (RGB) loss to address the long-tailed distribution problem in multi-label classification of medical images.

Refining 3D Medical Segmentation with Verbal Instruction

Xie, Kangxian (University at Buffalo), Gao, Mingchen (University at Buffalo)

CodeSegmentationData SynthesisConvolutional Neural NetworkTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningTextPoint CloudBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Constructed a dataset called CoWTalk and proposed an iterative refinement model for 3D medical segmentation based on language instructions, which can gradually correct initial segmentation results according to verbal correction instructions from radiologists.

RefTr: Recurrent Refinement of Confluent Trajectories for 3D Tubular Tree Centerlines

Naeem, Roman (Chalmers University of Technology), Kahl, Fredrik (Chalmers University of Technology)

CodeSegmentationComputational EfficiencyRepresentation LearningTransformerReinforcement LearningAuto EncoderContrastive LearningImageBiomedical DataComputed Tomography

🎯 What it does: Propose RefTr, a Producer-Refiner framework based on Transformer, which extracts the centerline maps of tree-like lumens such as blood vessels or airways from 3D CT images by recursively refining confluent trajectories.

ReMeDI: Refined Memory for Disambiguation of Identities with SAM3 in Surgical Segmentation

Bundele, Valay (University of Tübingen), Lensch, Hendrik P. A. (University of Tübingen)

CodeObject TrackingSegmentationTransformerContrastive LearningOptical FlowVideoBiomedical Data

🎯 What it does: Proposes ReMeDI-SAM3, a training-agnostic extension that improves SAM3's tool segmentation in surgical videos, addressing issues of occlusion, long-term tracking, and identity recovery.

ReMiX-Seg: Latent Modality Completion Meets Expert Routing for Robust Glioma Segmentation

Nghiem, Van Quang (National Taiwan University), Lin, Che (National Taiwan University)

CodeSegmentationConvolutional Neural NetworkTransformerMixture of ExpertsAuto EncoderMultimodalityBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose the ReMiX-Seg framework for glioma segmentation under any missing magnetic resonance modality;

RePCM: Region-Specific and Phenotype-Adaptive Bi-ventricular Cardiac Motion Synthesis

Yang, Xuan (National University of Singapore), Li, Lei (National University of Singapore)

CodeGenerationData SynthesisTransformerMixture of ExpertsDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningMeshBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose a full-cycle cardiac motion synthesis framework named RePCM based on single-frame end-diastolic meshes, used to generate three-dimensional time series of human ventricles.

ReportX: The BraTS Clinical Report Dataset

Marchesini, Kevin (University of Modena and Reggio Emilia), Bolelli, Federico (University of Modena and Reggio Emilia)

CodeSegmentationConvolutional Neural NetworkTransformerPrompt EngineeringVision Language ModelAuto EncoderContrastive LearningImageTextBiomedical DataMagnetic Resonance Imaging

🎯 What it does: A paired dataset called ReportX was constructed on the BraTS-GLI2023 dataset, combining structured clinical reports written by expert physicians with automatically generated quantitative fields, and these reports were used as auxiliary semantic supervision for 3D brain tumor segmentation.

ReScan-IA: A Spatially-Adaptive Diffusion Framework for Controllable 3D Intracranial Aneurysm Inpainting

Chuang, Tzu I (Charité - Universitätsmedizin Berlin), Hilbert, Adam (Charité - Universitätsmedizin Berlin)

CodeRestorationSegmentationGenerationData SynthesisConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningImageBiomedical DataComputed Tomography

🎯 What it does: This paper proposes a spatially adaptive 3D diffusion model, ReScan-IA, for controllably synthesizing/filling intracranial aneurysms in CTA imaging, simulating the appearance of aneurysms during repeated scans.

Residual Diffusion Bridge in Wavelet Space for Medical Image Translation

Yang, Xiao (Newcastle University), Zhang, Jingjing (Newcastle University)

CodeImage TranslationRestorationExplainability and InterpretabilityConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose a Wavelet Residual Bridge (WRB) framework for cross-modal MRI translation, first mapping low-frequency structures through dual-tree complex wavelet transform, then performing random sampling in the high-frequency texture subspace using Gaussian bridge, and finally reconstructing the target image.

Rethinking the Adaptation of Vision Foundation Models for Efficient Cell Segmentation

Xu, Qing (University of Nottingham), Chen, Zhen (University of Lincoln)

CodeSegmentationConvolutional Neural NetworkTransformerPrompt EngineeringAuto EncoderContrastive LearningImageBiomedical Data

🎯 What it does: Propose the EffiCell-Seg framework, which utilizes a frozen Vision Foundation Model and performs lightweight learning only on structural prompts and the decoder to achieve efficient cell segmentation.

REVEAL++: Differentiable Phenotypic Grouping for Vision–Language Retinal Modeling of Alzheimer’s Disease Risk

Meidinger, Ethan (University of Virginia), Fang, Ruogu (University of Florida)

CodeClassificationExplainability and InterpretabilityRepresentation LearningConvolutional Neural NetworkTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataAlzheimer's DiseaseElectronic Health RecordsRetrieval-Augmented Generation

🎯 What it does: This paper proposes REVEAL++, a differentiable phenotype-weighted audio-visual language alignment framework, for predicting Alzheimer's disease risk through retinal images and clinical risk narratives.

Revisiting Integration of Image and Metadata for DICOM Series Classification: Cross-Attention and Dictionary Learning

Truong, Tuan (Bayer AG), Lenga, Matthias (Bayer AG)

CodeClassificationConvolutional Neural NetworkTransformerContrastive LearningMultimodalityBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose an end-to-end multi-modal framework for automatically identifying DICOM image series, jointly modeling image content and acquisition metadata to address issues of missing, diverse, and heterogeneous data.

Reward-Guided Distillation: A Progressive Pseudo-Bag Purification Framework for WSI Multiple Instance Learning

Jia, Qi (Dalian University of Technology), Fan, Xin (Dalian University of Technology)

CodeClassificationExplainability and InterpretabilityKnowledge DistillationRepresentation LearningConvolutional Neural NetworkTransformerReinforcement LearningContrastive LearningImageBiomedical DataReview/Survey PaperBenchmark

🎯 What it does: Proposes a hierarchical pseudo-bag purification framework, utilizing Progressive Shapley ranking to construct pseudo-bags and filtering noise through the Reward-Guided Selection Module, thereby enhancing sparse classification performance in weakly supervised multiple instance learning (WSI-MIL).

Riemannian Batch Normalization on Correlation Manifolds for EEG Decoding

Yang, Jiarui, Wu, Xiao-Jun (Jiangnan University)

CodeClassificationRepresentation LearningBiomedical Data

🎯 What it does: A Correlation Batch Normalization (CorBN) layer is proposed for the Riemannian geometry of full-rank correlation matrices in EEG decoding, achieving geometrically consistent normalization on the correlation matrix manifold.

Robust Tooth Segmentation Under Orthodontic CBCT: A Metal Artifact-Aware Approach

Jo, SeungKwan (Soongsil University), Chung, Minyoung (Osstem Implant)

CodeSegmentationConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningBiomedical DataComputed Tomography

🎯 What it does: This paper proposes a two-stage dental instance segmentation framework, which utilizes metal artifacts in orthodontic CBCT as localization cues. Subsequently, it suppresses the impact of artifacts on boundaries, achieving robust dental detection and segmentation.

RPG-SAM: Reliability-Weighted Prototypes and Geometric Adaptive Threshold Selection for Training-Free One-Shot Polyp Segmentation

Lin, Weikun (East China Normal University), Wang, Yan (East China Normal University)

CodeSegmentationTransformerPrompt EngineeringAuto EncoderContrastive LearningImageBiomedical Data

🎯 What it does: Proposes a training-free one-time polyp segmentation framework called RPG-SAM, leveraging the base model SAM2 for high-quality segmentation.