arXivSub Start free trial

MICCAI 2026 Papers with Code — Page 6

International Conference on Medical Image Computing and Computer-Assisted Intervention · 649 papers

RT-Super: Learning Tumor Segmentation from Longitudinal Images and Reports

Bassi, Pedro R. A. S. (Johns Hopkins University), Zhou, Zongwei (Johns Hopkins University)

CodeSegmentationKnowledge DistillationRepresentation LearningData-Centric LearningConvolutional Neural NetworkTransformerLarge Language ModelDiffusion modelContrastive LearningImageTextBiomedical DataComputed TomographyRetrieval-Augmented Generation

🎯 What it does: Leverage the large amount of available imaging reports, longitudinal scans, and multi-phase CT from hospitals to build a teacher-student self-distillation architecture, learning multi-tumor (esophagus, spleen, uterus) segmentation without relying on a large number of manual masks;

S-CEReBrO: Breaking the Memory Barrier in Continuous EEG Monitoring

Bucagu, Glenn Anta (ETH Zurich), Benini, Luca (University of Bologna)

CodeAnomaly DetectionComputational EfficiencyRepresentation LearningTransformerAuto EncoderContrastive LearningTime SeriesBiomedical Data

🎯 What it does: Proposed S-CEReBrO, a Transformer-based continuous EEG foundation model that utilizes windowed alternating attention to achieve constant KV cache memory, and was pre-trained with masked autoencoding on 25,000 hours of EEG corpus.

S-GRPO: Structural Group Relative Policy Optimization for Medical Image Segmentation

Duan, Yiru (Nanjing University of Science and Technology), Chen, Qiang (Nanjing University of Science and Technology)

CodeSegmentationKnowledge DistillationConvolutional Neural NetworkTransformerReinforcement LearningMixture of ExpertsImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: This paper proposes a structured parallel sampling reinforcement learning framework called S-GRPO, aimed at improving pixel accuracy and global anatomical consistency in medical image segmentation.

S2DLight: Single Step Diffusion Network via Degradation-Aware Adaptation for Surgical Endoscopic Image Low-Light Enhancement

Fei, Song (Hong Kong University of Science and Technology), Zhu, Lei (Hong Kong University of Science and Technology)

CodeRestorationConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningImageBiomedical Data

🎯 What it does: Propose the S2DLight single-step diffusion framework to achieve efficient enhancement of low-light images in surgical endoscopy;

SA-CFM: Structure-Aware Continuous Flow Matching for Normative Modeling of EEG Microstate

Hu, Shiang (Anhui University), Lv, Zhao (Anhui University)

CodeClassificationAnomaly DetectionRepresentation LearningTransformerFlow-based ModelAuto EncoderContrastive LearningTime SeriesBiomedical DataAlzheimer's DiseaseElectrocardiogramOrdinary Differential Equation

🎯 What it does: Propose the Structure-Aware Continuous Flow Matching (SA-CFM) model, which provides a standardized modeling approach for EEG microstate potentials, generating reproducible individual deviation representations.

SAFE-Diff: Scale-Aware Attention and Feature-Dispersive Diffusion with Uncertainty Estimation for Contrast-Enhanced Breast MRI Synthesis

Zhang, Tianyu (Radboud University Medical Center), Mann, Ritse (Netherlands Cancer Institute)

CodeGenerationData SynthesisTransformerDiffusion modelScore-based ModelBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose SAFE-Diff, which synthesizes high-fidelity contrast-enhanced images on non-contrast breast MRI using a multi-scale attention diffusion model.

SAFE-Diff: Structurally Anchored Diffusion for Anatomically Faithful CT Image Super-Resolution

Sneha (Indian Institute of Technology Jodhpur), Paul, Angshuman (Indian Institute of Technology Jodhpur)

CodeSuper ResolutionConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelImageBiomedical DataComputed Tomography

🎯 What it does: Propose a two-stage CT super-resolution framework called SAFE-Diff: the first stage uses a residual prediction network to recover structure; the second stage uses truncated diffusion to refine details, and retains low-frequency structure and high-frequency details through SWT/ISWT frequency domain fusion.

SAGEAgent: A Self-evolving Agent for Cost-Aware Modality Acquisition in Multimodal Survival Prediction

Qu, Chongyu (Vanderbilt University), Huo, Yuankai (Vanderbilt University)

CodeExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelReinforcement LearningMultimodalityBiomedical DataElectronic Health RecordsRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose a self-evolving LLM agent called SAGEAgent for cost-aware sequential modal acquisition in multimodal survival prediction.

SAM2 as a Key Bridge Between 2D Slices and 3D Volumes for Sparsely-supervised Medical Image Segmentation

Yu, Peng (East China Normal University), Wang, Yan (East China Normal University)

CodeSegmentationTransformerVision Language ModelAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: This paper studies how to achieve complete 3D medical image segmentation using only three orthogonal central 2D slices with annotations, leveraging SAM2 for sparse annotation scenarios, and proposes a unified framework that integrates multi-view propagation, cross-view consistency weighted fusion, and dynamic pseudo-label refinement.

SAR-Net: Sample-Agnostic Retrieval Network for Robust Brain Tumor Segmentation with Missing Modalities

Zheng, Shenhai (Chongqing University of Posts and Telecommunications), Li, Weisheng (Chongqing University of Posts and Telecommunications)

CodeSegmentationRetrievalRepresentation LearningConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningMultimodalityBiomedical DataMagnetic Resonance ImagingRetrieval-Augmented Generation

🎯 What it does: Proposed a Sample-Agnostic Retrieval Network (SAR-Net) that completes missing modalities and performs brain tumor segmentation through a retrieval-based approach.

Scalable Training of Spatially Grounded 2D Vision–Language Models for Radiology

Salcan, Yusuf (University of Freiburg), Brox, Thomas (University of Freiburg)

CodeClassificationRecognitionObject DetectionSegmentationAnomaly DetectionTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingComputed TomographyRetrieval-Augmented Generation

🎯 What it does: Trained a CT/MRI VLM called RadGrounder for visual localization, supporting report generation, VQA, and localization.

SCKAN: Structural Consensus-Based KAN Prototype Learning for Semi-supervised Pancreas Segmentation

Liu, Yuqi, Li, Shuo (Tongji University)

CodeSegmentationTransformerAuto EncoderContrastive LearningImageBiomedical DataComputed Tomography

🎯 What it does: Proposes a structure consensus-based KAN prototype learning framework, SCKAN, to address the problem of supervisory bias in pancreatic segmentation.

Seamless Whole Slide Label-Free Virtual Staining

Kwark, Dou Hoon (University of Illinois Urbana-Champaign), Bhargava, Rohit (University of Illinois Urbana-Champaign)

CodeImage TranslationRestorationSegmentationGenerationConvolutional Neural NetworkTransformerDiffusion modelGenerative Adversarial NetworkImageBiomedical DataMagnetic Resonance ImagingComputed TomographyRetrieval-Augmented Generation

🎯 What it does: This work proposes Consistency Memory Bank (COMB), a framework for full-slide label-free virtual staining, aiming to eliminate artifacts such as stitching and color drift caused by block-based inference.

SeedPro: Place Brachytherapy Seeds Like Expert Clinicians via Hierarchical Reinforcement Agents

Li, Haitao (Shanghai Jiao Tong University), Chen, Xiaojun (Shanghai Jiao Tong University)

CodeRecurrent Neural NetworkTransformerReinforcement LearningAuto EncoderImageBiomedical DataComputed Tomography

🎯 What it does: Developed SeedPro, an automatic radioactive seed pre-planning framework based on hierarchical reinforcement learning, capable of generating radioactive seed layout plans comparable to those of experts.

SegDINO: Introducing Multi-Scale Structure into DINO for Efficient Medical Image Segmentation

Yang, Sicheng (Hong Kong University of Science and Technology (Guangzhou)), Zhu, Lei (Sun Yat-sen University)

CodeSegmentationConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningImageBiomedical DataComputed TomographyUltrasoundBenchmark

🎯 What it does: This study proposes the SegDINO framework, which transfers the pre-trained DINOv3 visual model to medical image segmentation tasks, achieving efficient segmentation through lightweight multi-scale modeling;

Segment Anything Model with Symmetric Spectral Uncertainty Gating for Generalizable Medical Image Segmentation

Li, He (University of Electronic Science and Technology of China), Zhang, Shaoting (University of Electronic Science and Technology of China)

CodeSegmentationDomain AdaptationTransformerSupervised Fine-TuningDiffusion modelContrastive LearningImageBiomedical DataMagnetic Resonance ImagingBenchmark

🎯 What it does: Propose a single-source domain generalization medical image segmentation framework called SSUG based on the Segment Anything Model (SAM), which utilizes three modules: Spectral Manifold Extrapolation, Uncertainty-aware Symmetric Channel Gating, and Hierarchical Decoupled Consistency, to jointly suppress style sensitivity and extract domain-invariant anatomical features.

Selective Rank-1 Orthogonality Regularization: Rethinking Stability-Plasticity Balance in Continual Medical Image Classification

Zhang, Jingyang (Southeast University), Ma, Lei (Tongji University)

CodeClassificationTransformerSupervised Fine-TuningPrompt EngineeringContrastive LearningImageBiomedical DataBenchmark

🎯 What it does: Propose Selective Rank-1 Orthogonality Regularization (SR1OR), which addresses the stability-plasticity balance problem in medical image continual learning by imposing selective orthogonality constraints on unit rank-1 subspaces within the LoRA low-rank update space.

Self-auditing Parameter-Efficient Fine-Tuning for Few-Shot 3D Medical Image Segmentation

Ly, Son Thai (University of Houston), Nguyen, Hien V. (University of Houston)

CodeSegmentationComputational EfficiencyData-Centric LearningTransformerSupervised Fine-TuningContrastive LearningPoint CloudBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose SEA-PEFT, a parameter-efficient fine-tuning method that automatically configures adapters through an online self-audit process, specifically designed for few-shot 3D medical image segmentation.

Self-Supervised Implicit CEST Reconstruction via Physics-Informed Lorentz Encoding

Li, Dexuan (East China Normal University), Yang, Guang (East China Normal University)

CodeOptimizationComputational EfficiencyRepresentation LearningTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Proposes a physics-informed implicit neural representation framework called Lorentz Encoding (LE), for self-supervised reconstruction of high-resolution CEST Z spectra under sparse sampling, and generates continuous, interpretable metabolic maps through physical constraints.

Self-supervised Learning on Lingual Ultrasound Video Encodes Clinically Meaningful Articulatory Structure of Rhotic Production in Residual Speech Sound Disorder

Kim, Taehyo, Shu, Hai (New York University)

CodeClassificationRepresentation LearningTransformerAuto EncoderContrastive LearningVideoBiomedical DataUltrasound

🎯 What it does: This paper proposes a joint embedding framework based on self-supervised learning, which learns the tongue position structure of American English rhotic pronunciation using tongue ultrasound videos, and reduces speaker differences through speaker adversarial regularization.

Self-supervised Spline Autoencoder for Smooth Latent Dynamics in Echocardiography

Kwan, Leong Chit Jeff (Hong Kong University of Science and Technology), Chung, Albert C. S. (Hong Kong University of Science and Technology)

CodeAnomaly DetectionRepresentation LearningDiffusion modelAuto EncoderContrastive LearningOptical FlowVideoBiomedical DataUltrasound

🎯 What it does: This study proposes a self-supervised Spline Autoencoder, which learns low-rank latent trajectories of cardiac motion through smooth curve fitting, and achieves unlabeled cardiac cycle and ED/ES event localization by utilizing spectral embedding and phase projection.

Self-supervised Temporal Regularization for Landmark-Based Cardiac Segmentation with Automatic AHA Regional Mapping

Montalvo-García, David (Universidad Politécnica de Madrid), Ferrante, Enzo (Universidad de Buenos Aires)

CodeSegmentationConvolutional Neural NetworkGraph Neural NetworkAuto EncoderContrastive LearningOptical FlowImageVideoBiomedical DataUltrasound

🎯 What it does: Building upon the previously learned implicit correspondence cardiac graph network Mask‑HybridGNet Dual, this study introduces self-supervised temporal regularization (velocity and acceleration penalty) as a post-training phase to enhance temporal consistency in ultrasound sequence segmentation, and utilizes the learned correspondence to achieve automatic AHA 17-segment standard mapping, completing regional motion analysis.

Semantic Feature Modulation for Mammographic Lesion Classification

Mahpod, Shahar (Ariel University), Ben Artzi, Gil (Ariel University)

CodeClassificationAnomaly DetectionConvolutional Neural NetworkTransformerContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Introduce a semantic feature modulation framework based on BI-RADS descriptors in breast X-ray image classification to achieve dynamic fusion of images and clinical semantics.

Sequence-Conditioned Flow-Based Models for Digital Phantom Generation in MRI

Belousova, Kseniya (ITMO University), Al-Haidri, Walid (ITMO University)

CodeGenerationData SynthesisConvolutional Neural NetworkDiffusion modelScore-based ModelFlow-based ModelImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: This study proposes a self-supervised hierarchical flow model for estimating corresponding quantitative T1, T2, and PD maps from T1, T2, and PD-weighted MRI images, and for synthesizing new weighted images using the MR signal model.

Set-Inclusive Uncertainty Modeling for Robust Brain Tumor Segmentation

Baek, Seunghun (Pohang University of Science and Technology), Kim, Won Hwa (Pohang University of Science and Technology)

CodeSegmentationConvolutional Neural NetworkContrastive LearningGaussian SplattingMultimodalityBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose a probabilistic representation framework that models the task representation of each modality subset as a Gaussian distribution, utilizing the set inclusion hierarchy to guide mean alignment and variance to reflect information missing, thus achieving brain tumor segmentation under missing modalities;

SF-DINO: Spatial-Frequency Adapted Foundation Model for Multiple Myeloma Diagnosis

Ye, Zhaoyi (Wuhan University), Lei, Cheng (Wuhan University)

CodeClassificationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningDrug DiscoveryConvolutional Neural NetworkTransformerPrompt EngineeringContrastive LearningImageBiomedical Data

🎯 What it does: Propose the SF-DINO model, which uses a spatial morphology adapter and frequency-selective hash attention for fine-grained classification of multiple myeloma cells;

Shape-Adaptive Guidance Signal for Interactive Cortical Sulcal Labeling

Son, Jiwon (POSTECH), Lyu, Ilwoo (POSTECH)

CodeSegmentationConvolutional Neural NetworkDiffusion modelContrastive LearningBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose an interactive cortical sulcus annotation framework based on shape-adaptive guiding signals, using spherical CNN to perform binary segmentation of the left posterior frontal sulcus.

Shape-Scale Aware Diffusion Models for Weakly Supervised Medical Image Anomaly Detection

Zheng, Shenhai (Chongqing University of Posts and Telecommunications), Li, Laquan (Chongqing College of Artificial Intelligence)

CodeAnomaly DetectionConvolutional Neural NetworkDiffusion modelScore-based ModelContrastive LearningBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Proposed a direction-guided shape-scale-aware diffusion model for weakly supervised medical image anomaly detection.

ShapeEFM: A Shape-Based Pretrained Foundation Model on ECG Morphology Comprehension

Liu, Huan (Ant Group), Lu, Le (Zhejiang University School of Medicine)

CodeExplainability and InterpretabilityRepresentation LearningTransformerSupervised Fine-TuningAuto EncoderContrastive LearningTime SeriesBiomedical DataElectrocardiogram

🎯 What it does: Propose ShapeEFM, a shape-based pre-trained ECG foundation model, which decomposes ECG waveforms into Cardio-Morphemes, treating them as 'words,' and uses Transformer for self-supervised learning to enhance ECG morphology understanding and interpretability.

ShapKO: Shapley-Adaptive Modality Knockout for Robust Multimodal Learning

Nizam, Nusrat Binta (Cornell University), Sabuncu, Mert R. (Weill Cornell Medicine)

CodeExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerVision Language ModelAuto EncoderContrastive LearningImageTextMultimodalityTabularBiomedical DataElectronic Health Records

🎯 What it does: By dynamically estimating modal importance using Shapley values during training, adaptively adjusting the elimination probability for each modality, thereby enhancing the robustness of multi-modal models under missing modality scenarios.

Shared Gaussian Geometry for Brain MRI Spatial Resolution Harmonization Without External Training

Gao, Yifan (Cedars-Sinai Medical Center), Li, Debiao (Cedars-Sinai Medical Center)

CodeImage HarmonizationRestorationSuper ResolutionTransformerDiffusion modelAuto EncoderContrastive LearningGaussian SplattingImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose a shared Gaussian geometric framework to achieve spatial resolution harmonization of multi-contrast brain MRI without the need for external training sets.

SHINE: An Entropy-Guided Digital Twin Framework for 3D Myocardial Scar Reconstruction from Sparse CMR

Zong, Dexiang, Zhong, Liang (National Heart Centre Singapore)

CodeSegmentationGenerationData SynthesisConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: A continuous function learning framework based on sparse LGE-MRI was constructed to generate a three-dimensional digital twin of myocardial scar.

Single-Stage Hierarchical Rectification for Weakly Supervised Histopathology Segmentation

Nguyen, Trong Duc (VinUniversity), Pham, Huy Hieu (Posts and Telecommunications Institute of Technology)

CodeSegmentationConvolutional Neural NetworkRectified FlowContrastive LearningBiomedical Data

🎯 What it does: Propose a single-stage hierarchical correction framework, SSHR, which achieves weakly supervised histopathological semantic segmentation by correcting intermediate features during the forward pass, avoiding multi-stage CAM generation and recursive pseudo-labeling.

Single-Subject Multi-view MRI Super-Resolution via Implicit Neural Representations

Kim, Heejong (Weill Cornell Medicine), Sabuncu, Mert R. (Weill Cornell Medicine)

CodeSuper ResolutionNeural Radiance FieldAuto EncoderOptical FlowBiomedical DataMagnetic Resonance Imaging

🎯 What it does: A single-organ multi-view MRI super-resolution framework based on implicit neural representations was studied, which can reconstruct isotropic high-resolution images from multi-directional anisotropic scans without the need for pre-registration.

SIRA: Reasoning-Aware Surgical Instrument Segmentation via Query-Anchored Alignment

Zhang, Zhibo (Huazhong University of Science and Technology), Yan, Zengqiang (Huazhong University of Science and Technology)

CodeSegmentationTransformerLarge Language ModelVision Language ModelContrastive LearningImageTextMultimodalityBiomedical DataBenchmark

🎯 What it does: Proposed a query-conditioned surgical instrument segmentation framework called SIRA, achieving multi-modal semantic alignment segmentation tasks on the newly constructed SurgRS dataset.

Situational context embedding and knowledge-based grounding for LLMs in surgical training assistance

Bahari Malayeri, Ali (University of Zurich), Fürnstahl, Philipp (Zurich University of Applied Sciences)

CodeExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningRobotic IntelligenceDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringTextBiomedical DataReview/Survey PaperRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed and implemented a tutoring system for total hip arthroplasty (THA) training simulator based on large language models, which provides context-aware answers by embedding real-time situational information from the simulator and combining it with knowledge base retrieval.

SkepticalNet: A Safety-Aware Interactive Segmentation Framework for Neuro-Oncology

Kyriazis, Efstathios (Foundation for Research and Technology Hellas), Marias, Kostas (Foundation for Research and Technology Hellas)

CodeSegmentationConvolutional Neural NetworkPrompt EngineeringContrastive LearningMultimodalityBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Proposes a multi-modal 3D interactive segmentation framework called SkepticalNet for precise segmentation of neural tumors;

Skin2Mind: A Multimodal Framework for Mental Health Risk Screening via Facial Images and Structured Skin Reports

Liu, Jiahe (Monash University), Ge, Zongyuan (Monash University)

CodeClassificationExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelVision Language ModelContrastive LearningImageTextMultimodalityElectronic Health Records

🎯 What it does: Proposed and implemented Skin2Mind, a multimodal framework based on multispectral facial skin images and structured skin reports, for predicting individual-level DSM-5 level mental health risks.

SLICE-MIL: Counterfactual Semantic Anchoring of Disentangled Evidence for Child-Pugh Grading

Deng, Yihan (University of Science and Technology of China), Zheng, Jian (University of Science and Technology of China)

CodeClassificationExplainability and InterpretabilityTransformerContrastive LearningImageBiomedical DataComputed Tomography

🎯 What it does: Proposed a SLICE-MIL framework based on weakly supervised multiple instance learning, using anatomy-guided dual-query attention and adversarial semantic anchoring to estimate Child-Pugh grades from portal-phase CT imaging.

Slim nnU-Net: Revisiting nnU-Net Self Configuration from an Optimization Angle

Zhao, Yidong (Delft University of Technology), Tao, Qian (Delft University of Technology)

CodeSegmentationOptimizationComputational EfficiencyConvolutional Neural NetworkAuto EncoderImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: This paper reformulates the training of nnU-Net as a proximal optimization problem with a hybrid L1+Group L2 sparse regularization, and achieves extremely sparse Slim nnU-Net by performing channel-wise pruning during training.

Smartphone-Based ASD Screening via Instance-Adaptive Cross-Modal Fusion and Contrastive Learning

Xu, Jie (Northwestern Polytechnical University), Xia, Chen (Northwestern Polytechnical University)

CodeClassificationAnomaly DetectionConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningVideoMultimodality

🎯 What it does: Implementing an autism spectrum disorder screening based on eye movement and head pose on smartphones, proposing the AMF-Net framework.

SonoCLIP: Mask-Guided Region-Aware Vision–Language Pretraining for Fetal Ultrasound Analysis

Su, Hang (Wuhan University), Du, Bo (Wuhan University)

CodeRecognitionSegmentationTransformerVision Language ModelContrastive LearningImageTextBiomedical DataUltrasound

🎯 What it does: This paper proposes SonoCLIP, a fetal ultrasound foundation model based on vision-language alignment, which can achieve region-controllable representation learning at both global and local levels.

Sparse Autoencoders for Interpretable Medical Image Representation Learning

Wesp, Philipp (Stanford University), Gatidis, Sergios (Stanford University)

CodeRetrievalExplainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelVision Language ModelAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: This paper proposes using a sparse autoencoder (Matryoshka SAE + BatchTopK) to transform dense embeddings from medical imaging foundational models (BiomedParse, DINOv3) into interpretable sparse features, and validates their interpretability and performance through various retrieval and evaluation methods;

Sparsely Supervised Surgical Video Segmentation with Reliable Asymmetric Dual Memory

Sapkota, Nishchal (University of Notre Dame), Chen, Danny Z. (University of Notre Dame)

CodeSegmentationConvolutional Neural NetworkTransformerVision-Language-Action ModelContrastive LearningVideo

🎯 What it does: Proposes D-GEM, a reliable asynchronous dual-memory framework for semantic segmentation of surgical videos under sparse supervision.

SPEAR: Error-Guided Sparse Refinement for Scalable Deformable Medical Image Registration

Bai, Hao (Shanghai Jiao Tong University), Hong, Yi (Shanghai Jiao Tong University)

CodeImage TranslationSegmentationComputational EfficiencyConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningOptical FlowImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose the SPEAR framework, which achieves high-resolution 3D medical image deformation registration through error-guided sparse refinement.

Spectral-Semantic Consensus for Few-Shot Whole Slide Image Classification

Du, Hanyu (Wuhan University), Xu, Yongchao (Wuhan University)

CodeClassificationAnomaly DetectionExplainability and InterpretabilityComputational EfficiencyRepresentation LearningMeta LearningTransformerPrompt EngineeringVision Language ModelAuto EncoderContrastive LearningImageTextBiomedical DataBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposes a Spectral-Semantic Consensus (SSC) framework based on text priors for few-shot whole-slide image classification, automatically selecting diagnostic information-rich patches and aggregating them.

Speech-Guided Multimodal Learning for Vocal Tract Segmentation in Real-Time MRI

Liu, Daiqi (Friedrich-Alexander-Universität Erlangen-Nürnberg), Pérez-Toro, Paula Andrea (Friedrich-Alexander-Universität Erlangen-Nürnberg)

CodeSegmentationData-Centric LearningConvolutional Neural NetworkTransformerContrastive LearningImageMultimodalityBiomedical DataMagnetic Resonance ImagingAudio

🎯 What it does: This study proposes a three-stage multimodal learning framework for precise segmentation of vocal tract organs using speech-guided real-time MRI.

Spinverse: Differentiable Physics for Permeability-Aware Microstructure Reconstruction from Diffusion MRI

Khole, Prathamesh Pradeep (University of California Santa Cruz), Marinescu, Razvan (University of California Santa Cruz)

CodeOptimizationGraph Neural NetworkDiffusion modelScore-based ModelGenerative Adversarial NetworkContrastive LearningMeshBiomedical DataMagnetic Resonance ImagingDiffusion Tensor Imaging

🎯 What it does: Propose the Spinverse method, which optimizes surface permeability on a fixed tetrahedral mesh using a differentiable Bloch-Torrey finite element simulator to inversely reconstruct the microstructural boundaries corresponding to diffusion MRI signals.

SRMA-Mamba: Spatial Reverse Mamba Attention Network for Pathological Liver Segmentation in MRI Volumes

Zeng, Jun, Jha, Debesh (University Of South Dakota)

CodeSegmentationConvolutional Neural NetworkTransformerMixture of ExpertsDiffusion modelContrastive LearningImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Proposed and implemented a three-dimensional convolution + Mamba network called SRMA-Mamba for precise segmentation of liver cirrhosis lesions in magnetic resonance imaging (MRI) volumes.

SRPAN: Implicit Neural Slice-to-Volume Reconstruction for Super-Resolution Pancreatic MRI

Olofsson, Fredrik (Leiden University Medical Center), Dijkstra, Jouke (Leiden University Medical Center)

CodeSuper ResolutionDiffusion modelScore-based ModelNeural Radiance FieldAuto EncoderBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose an implicit neural slice-to-volume reconstruction method called SRPAN for low-redundancy pancreatic MRI data;

SSPT: Spiking Serialized Point Transformer for Couinaud Segmentation in 3D Medical Data

Li, Yan (Chinese University of Hong Kong), Heng, Pheng-Ann (Chinese University of Hong Kong)

CodeSegmentationComputational EfficiencySpiking Neural NetworkTransformerAuto EncoderContrastive LearningPoint CloudBiomedical DataComputed Tomography

🎯 What it does: Proposed a Spiking Serialized Point Transformer (SSPT) based on spiking neural networks for Couinaud segmentation of 3D medical point clouds.

Stabilizing Temporal Inference Dynamics for Online Surgical Phase Recognition

Liu, Yang (King's College London), Ourselin, Sébastien (University Of Electronic Science And Technology Of China)

CodeRecognitionRecurrent Neural NetworkTransformerSupervised Fine-TuningContrastive LearningVideoBiomedical DataBenchmark

🎯 What it does: This paper proposes a unified Train-Inference-Evaluation framework specifically designed to enhance the temporal stability of models for surgical phase recognition (SPR) during the online surgery phase, avoiding the cascading errors caused by short-term jitter and misjudgment;

StrokeTimer: Robust Representation Learning for Ischemic Stroke Onset-Time Estimation from Non-Contrast CT

Wang, Weiru (Eindhoven University of Technology), Su, Ruisheng (Maastricht University)

CodeClassificationRepresentation LearningConvolutional Neural NetworkAuto EncoderContrastive LearningImageBiomedical DataComputed Tomography

🎯 What it does: This study proposes an end-to-end framework called StrokeTimer, which can automatically estimate the onset time window of ischemic stroke in non-enhanced CT images (<4.5 h, 4.5–6 h, >6 h).

Structural Congruence Matters: Biological Pretraining for Vascular Graph Extraction

Scavone, Alessandro (EURECOM), Zuluaga, Maria A. (EURECOM)

CodeImage TranslationSegmentationDomain AdaptationGraph Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningImageGraphBiomedical DataAgriculture Related

🎯 What it does: A 'nature-to-nature' pre-training paradigm is proposed, using the structural pre-training of plant (grapevine) skeleton graphs for end-to-end extraction of medical images to vascular graph structures.

Structure-Adaptive Sparse Diffusion in Voxel Space for 3D Medical Image Enhancement

Jiang, Hongxu (University of Florida), Shao, Wei (University of Florida)

CodeRestorationSuper ResolutionConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelContrastive LearningBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission Tomography

🎯 What it does: A sparse voxel space diffusion framework is proposed for denoising and super-resolution of 3D medical images. By directly predicting clean images in the voxel space, incorporating velocity supervision, sparse time step scheduling, and structure-aware trajectory modulation, the method achieves results with only 5 time steps, a 10-fold increase in training speed, while maintaining fine details and structural integrity.

Structure-Contrast Disentangled INRs for Accelerated Multi-contrast MRI Reconstruction

Vavasour, Zach (University of Toronto), Chiew, Mark (University of Toronto)

CodeRestorationNeural Radiance FieldAuto EncoderContrastive LearningBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Proposes a structure-contrast disentangled implicit neural representation, DISINR, for multi-contrast MRI accelerated reconstruction;

Subject- and Task-Aware EEG Foundation Model

An, Sion (Pohang University of Science and Technology), Park, Sang Hyun (Pohang University of Science and Technology)

CodeClassificationRepresentation LearningMeta LearningTransformerSupervised Fine-TuningContrastive LearningTime SeriesBiomedical DataElectrocardiogram

🎯 What it does: Propose the STEM model, combining self-supervised contrastive learning with metric-based meta-learning, explicitly decoupling EEG features into subject-related and task-related components, thereby fully utilizing subject identity and task label information during the pre-training phase.

SurfMark3D: Surface-Based Graph Convolution Framework for Distal Femoral Landmark Localisation in Total Knee Arthroplasty

B., Dharshan (Indian Institute of Technology Madras), Sivaprakasam, Mohanasankar (Indian Institute of Technology Madras)

CodePose EstimationGraph Neural NetworkDiffusion modelAuto EncoderContrastive LearningGaussian SplattingPoint CloudMeshBiomedical DataMagnetic Resonance ImagingComputed TomographyBenchmark

🎯 What it does: Proposed a graph convolutional framework based on surface point clouds, named SurfMark3D, for anatomical landmark localization of the distal femur during total knee arthroplasty.

Surgical Video Temporal Grounding

Liu, Qixuan (Chinese University of Hong Kong), Heng, Pheng-Ann (Chinese University of Hong Kong)

CodeRecognitionSegmentationTransformerSupervised Fine-TuningPrompt EngineeringVision Language ModelVideoMultimodalityBenchmark

🎯 What it does: Proposed the SurgVTG task for temporal localization in surgical videos, constructed the HMSSurgVTG benchmark, and provided corresponding annotations.

SurgOnAir: Hierarchy-Aware Real-Time Surgical Video Commentary

He, Jingyi (TU Munich), Bi, Yuan (TU Munich)

CodeGenerationComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringVision Language ModelVideoTextRetrieval-Augmented GenerationAudio

🎯 What it does: Propose the SurgOnAir model to achieve real-time hierarchical narration of surgical videos.

SurgRFO: Foundation Model Based Compositional Synthesis of Critical Retained Foreign Objects in Intraoperative Chest X-Rays

Hu, Yuanyun (Tsinghua University), Bai, Harrison (Johns Hopkins University)

CodeObject DetectionGenerationData SynthesisConvolutional Neural NetworkTransformerSupervised Fine-TuningPrompt EngineeringDiffusion modelImageBiomedical DataComputed TomographyBenchmark

🎯 What it does: Designed and implemented a two-stage synthesis framework called SurgRFO, which first generates surgery chest X-ray backgrounds without RFO using a latent diffusion model, and then samples local RFOs with a lightweight generator and generates realistic composite images through conditional Poisson fusion, used for data augmentation to improve detection performance.

SurgVLA-Bench: Towards Evaluating Vision-Language-Action Models for Laparoscopic Surgical Robotics

Sun, Jiashuo (National Engineering Research Center of Robot Visual Perception and Control Technology), Liu, Min (National Engineering Research Center of Robot Visual Perception and Control Technology)

CodeRobotic IntelligenceTransformerSupervised Fine-TuningReinforcement LearningVision-Language-Action ModelImageVideoTextSequentialBiomedical DataBenchmark

🎯 What it does: Propose SurgVLA-Bench, constructing a surgical vision-language-action (VLA) evaluation benchmark based on the SurRoL simulation platform, which includes a hierarchical task system and corresponding datasets ranging from atomic actions to complete surgical procedures.

SurgZSD: Zero-Shot Surgical Action Triplet Detection via Attribute Composition and Clinical Feasibility Inference

Deng, Zuxing (Hefei University of Technology), Zhou, Wenrui (Hefei University of Technology)

CodeObject DetectionTransformerSupervised Fine-TuningPrompt EngineeringVision Language ModelImageBiomedical DataBenchmarkRetrieval-Augmented Generation

🎯 What it does: Developed the SurgZSD zero-shot surgical action triple detection framework, which utilizes medical prior knowledge to achieve joint reasoning of attribute combination and clinical feasibility constraints.

SutureFormer: Learning Surgical Trajectories via Goal-Conditioned Offline RL in Pixel Space

Liu, Huanrong (University of Macau), Li, Qingbiao (University of Macau)

CodeRobotic IntelligenceConvolutional Neural NetworkTransformerReinforcement LearningDiffusion modelVideoBiomedical Data

🎯 What it does: Learn a surgical suturing trajectory prediction model called SutureFormer in the pixel space through offline reinforcement learning, enabling the prediction of future needle tip trajectories based solely on endoscopic videos and sparse keyframes.

SwallowReg: SAM-Assisted Deformable Registration with Adaptive Global-Local Features for Cine-MRI Swallowing Function Quantification

Tang, Zhiwen (Nanjing University of Posts and Telecommunications), Yu, Han (Nanjing Medical University)

CodeImage TranslationRestorationSegmentationAnomaly DetectionConvolutional Neural NetworkTransformerDiffusion modelContrastive LearningOptical FlowImageVideoBiomedical DataMagnetic Resonance Imaging

🎯 What it does: This study proposes an end-to-end SwallowReg framework for joint segmentation and registration in dynamic Cine-MRI of patients with tongue cancer after head and neck resection, thereby enabling quantitative evaluation of swallowing function.

Synergistic Perception-Reasoning Governance: Grounding Medical MLLMs with Verifiable Anatomical Evidence

Hao, Rui, Zeng, Zhigang (Huazhong University Of Science And Technology)

CodeExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextBiomedical DataMagnetic Resonance ImagingComputed TomographyElectronic Health RecordsRetrieval-Augmented Generation

🎯 What it does: A training-agnostic, anatomy ROI evidence injection-based multimodal large model hallucination suppression framework is proposed, combining visual activation re-adjustment with text structural injection, and introducing a task-aware dynamic router.

SynHydro: Biomechanics-Driven Domain Randomization for Hydrocephalus-Agnostic, Generalizable Tissue-Ventricle Segmentation Across Age, Modality, and Resolution

Ren, Zehua (Xi'an Jiaotong University), Lian, Chunfeng (Xi'an Jiaotong University)

CodeSegmentationData SynthesisDomain AdaptationConvolutional Neural NetworkDiffusion modelGenerative Adversarial NetworkContrastive LearningImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Simulate hydrocephalus using biomechanically driven deformations based on healthy brain tissue labels, then generate diverse multimodal, low-resolution, noisy, and artifact-containing MRI data through domain randomization, and train nnU-Net to achieve joint segmentation of brain tissue and ventricles in pediatric hydrocephalus patients.

SYNPRED: A Synergistic Approach to Multimodal Learning for Clinical Prediction

Janíčková, Ivana (Medical University of Vienna), Langs, Georg (Medical University of Vienna)

CodeClassificationExplainability and InterpretabilityRepresentation LearningData-Centric LearningTransformerSupervised Fine-TuningMixture of ExpertsAuto EncoderGenerative Adversarial NetworkContrastive LearningImageMultimodalityTabularBiomedical DataMagnetic Resonance ImagingElectronic Health Records

🎯 What it does: Propose a multi-modal variational autoencoder (VAE) framework called SYNPRED, which integrates medical imaging and RNA sequencing data, to enable clinical prediction when multi-modal data are missing.

TAIR: Text-Guided Adaptive Prompt Refinement for Coarse-to-Fine all-in-One Medical Image Restoration

Cui, Jiaqi, Wang, Yan (Sichuan University)

CodeRestorationSuper ResolutionTransformerPrompt EngineeringVision Language ModelContrastive LearningImageMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission Tomography

🎯 What it does: Propose TAIR, a text-guided adaptive prompt refinement framework, achieving full-scenario medical image restoration in coarse-to-fine stages.

Task-Performance Routing for Multi-teacher Distillation in CT Medical Image Segmentation

Yang, Jiaye, Wang, Peng (University Of Electronic Science And Technology Of China)

CodeSegmentationKnowledge DistillationTransformerMixture of ExpertsBiomedical DataComputed Tomography

🎯 What it does: CT organ segmentation based on multi-teacher knowledge distillation, proposing the Task-Performance Routing (TPR) framework, which dynamically selects the best teacher and fuses knowledge for each region in space.

TAVR-VLM: Risk-Conditioned Causal Grounding for Hallucination-Resistant Report Generation

Lu, Zhixiang (Xi'an Jiaotong-Liverpool University), Wang, Jinfeng (Xi'an Jiaotong-Liverpool University)

CodeGenerationAnomaly DetectionExplainability and InterpretabilityConvolutional Neural NetworkTransformerLarge Language ModelVision Language ModelContrastive LearningImageTextMultimodalityTabularBiomedical DataComputed TomographyUltrasoundElectronic Health RecordsRetrieval-Augmented Generation

🎯 What it does: This paper proposes a framework named TAVR‑VLM for generating preoperative reports for transcatheter aortic valve replacement (TAVR), and suppresses diagnostic hallucinations through risk-conditioned causal grounding.

TCCT: Trajectory-Conditioned CBCT Reconstruction for Sinusoidal Non-circular Orbits with a Fourier Neural Operator

Ye, Chengze (Friedrich-Alexander-Universität Erlangen-Nürnberg), Maier, Andreas (Friedrich-Alexander-Universität Erlangen-Nürnberg)

CodeRestorationConvolutional Neural NetworkDiffusion modelAuto EncoderContrastive LearningBiomedical DataComputed Tomography

🎯 What it does: Proposed the TCCT framework, which utilizes the orbit-conditioned spectral neural operator to predict redundant weights and embeds them into a differentiable Shift-Variant FBP, achieving fast CBCT reconstruction for sinusoidal non-circular orbits.

TCFD-Net: Temporal-Calibrated Feature Disentanglement Network for Dual-Timepoint CT-Based Immunotherapy Response Prediction in Advanced Non-small Cell Lung Cancer

Fu, Yao (Beihang University), Mu, Wei (Beihang University)

CodeClassificationTransformerContrastive LearningImageBiomedical DataComputed Tomography

🎯 What it does: Developed TCFD-Net based on dual-time-point CT images to predict early immunotherapy response in patients with non-small cell lung cancer (NSCLC).

TEDi: Temporal Memory-Enhanced and Denoising Transformer for Surgical Instrument Segmentation

Yuan, Jiahong (Tsinghua University), Zhou, Haoyin (Harvard Medical School)

CodeSegmentationTransformerContrastive LearningImageVideoBenchmark

🎯 What it does: Proposed the TEDi framework, combining memory search enhancement and temporal consistency denoising Transformer to improve surgical instrument segmentation, addressing cross-frame semantic consistency and class confusion issues.

TeDyS: Temporal Dynamics with Spatial Context for Longitudinal Progression Prediction

Hwang, Jeonghyun (Ewha Womans University), Choi, Jang-Hwan (Ewha Womans University)

CodeClassificationAnomaly DetectionTransformerImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose the TeDyS framework, which utilizes time-dynamic conditional queries for progress prediction in long-term medical imaging.

Temporal Phase-Difference Guided Spatiotemporal Learning for DSA Vessel Segmentation

Liu, Kun (Beijing University of Posts and Telecommunications), Yang, Huihua (Beijing University of Posts and Telecommunications)

CodeSegmentationConvolutional Neural NetworkRecurrent Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningVideoBiomedical Data

🎯 What it does: Propose a spatiotemporal segmentation framework combining forward phase difference projection maximum intensity projection (FPD‑MaxIP) and a Mamba-based temporal encoder for precise segmentation of vessels in digital subtraction angiography (DSA).

Test-Time Adaptation for ECG Classification via SQI-Gated Self-training and Beat-Rhythm Consistency

Jiang, Wenhan (Westlake University), Zheng, Yefeng (Westlake University)

CodeClassificationDomain AdaptationConvolutional Neural NetworkSupervised Fine-TuningContrastive LearningTime SeriesBiomedical DataElectrocardiogram

🎯 What it does: Achieve test-time adaptation in electrocardiogram classification under unseen domains.

Test-Time Adaptation for Rare Surgical Phase Recognition: Bridging the Coverage-Gap Paradox

Park, Ho-min (Ghent University Global Campus), Vankerschaver, Joris (Ghent University Hospital)

CodeRecognitionDomain AdaptationRecurrent Neural NetworkTransformerSupervised Fine-TuningContrastive LearningVideoBiomedical Data

🎯 What it does: This paper addresses the coverage gap issue in the identification of rare surgical phases by proposing a lightweight test-time adaptation framework, which significantly improves the accuracy of rare phase identification through adaptive threshold pseudo-label updating and temporal smoothing.

Test-Time Adaptation in Optical Coherence Tomography Using Trajectory-Aligned Time-Independent Flow

Hucke, Veit (Medical University of Vienna), Bogunović, Hrvoje (Medical University of Vienna)

CodeRestorationSegmentationDomain AdaptationConvolutional Neural NetworkSupervised Fine-TuningDiffusion modelFlow-based ModelImageBiomedical DataOrdinary Differential Equation

🎯 What it does: Proposes a test-time adaptation framework called TTA-Flow based on flow matching and histogram alignment, which transforms noisy images from low-cost OCT devices into high-quality training domain images.

Test-time Adaptation of Pelvic Bone Segmentation Models via Dynamic Reliability-Guided

Ren, Ling, Zheng, Kai (Nanjing University of Posts and Telecommunications)

CodeSegmentationDomain AdaptationConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningImageBiomedical DataComputed Tomography

🎯 What it does: An online adaptive framework for pelvic bone segmentation in medical CT imaging.

Text as Illumination: Spatial Contrastive Retinex Learning for Language-Guided Medical Image Segmentation

Shi, Jian (Dalian University of Technology), Lu, Huchuan (Dalian University of Technology)

CodeSegmentationTransformerVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose a Retinex-inspired network called TIRNet that utilizes text embeddings as semantic illumination to achieve language-guided medical image segmentation.

Text-Deficient Multimodal Stroke Segmentation with Lesion-Grounded Self-Retrieval-Augmented Generation

Eum, Heeseong (Seoul National University), Choi, Kyu Sung (Seoul National University)

CodeSegmentationGenerationRetrievalConvolutional Neural NetworkTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextBiomedical DataMagnetic Resonance ImagingRetrieval-Augmented Generation

🎯 What it does: The study proposes LeG-RAG, a lesion-aligned retrieval-based generation method for missing report generation in the text-missing scenario of acute ischemic stroke MRI, and implements report-conditioned multimodal segmentation using LLMSwin.

Text-Guided Multi-frequency Latent Diffusion for Medical Image Segmentation

Gao, Qiang (Monash University), Chen, Cunjian (Monash University)

CodeSegmentationTransformerPrompt EngineeringVision Language ModelDiffusion modelScore-based ModelAuto EncoderContrastive LearningImageTextBiomedical Data

🎯 What it does: Proposed a text-guided multi-frequency latent diffusion framework (TMF-Seg) for medical image segmentation.

TGBD-Net: Tumor-Guided Bridging Distillation for Joint EGFR Mutation and Survival Prediction in Lung Cancer from CT Imaging

Yang, Huihui (Beihang University), Tian, Jie (Beihang University)

CodeClassificationKnowledge DistillationRepresentation LearningTransformerContrastive LearningImageBiomedical DataComputed Tomography

🎯 What it does: Propose TGBD-Net for jointly predicting EGFR mutation status and progression-free survival (PFS) in lung cancer CT images, achieving stable multi-task learning in a heterogeneous supervision environment.

TGH-DB: Template-Guided Heteroscedastic Diffusion Bridge for Brain MRI-to-PET Synthesis

Pang, Haowen (Beijing Institute of Technology), Qiu, Anqi (Hong Kong Polytechnic University)

CodeGenerationData SynthesisTransformerDiffusion modelMultimodalityBiomedical DataMagnetic Resonance ImagingPositron Emission TomographyAlzheimer's Disease

🎯 What it does: Propose a Template-Guided Heteroscedastic Diffusion Bridge (TGH-DB) based on population-level amyloid PET templates, which can generate structured and physiologically reasonable PET images from multi-modal MRI.

The Learnability Gap in Medical Latent Diffusion

Dombrowski, Mischa (FAU Erlangen Nurnberg), Kainz, Bernhard (FAU Erlangen Nurnberg)

CodeClassificationGenerationKnowledge DistillationRepresentation LearningConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyElectronic Health Records

🎯 What it does: This paper systematically investigates the 'learnability gap' in potential diffusion models for medical image generation, and proposes a noise-conditioned latent classifier (FiLM layer + image space distillation) to diagnose and partially alleviate this gap.

The Paper Has a GitHub, the GitHub Has a README, the README Has Nothing: Reproducibility Signals for Review Support

Bolelli, Federico (University of Modena and Reggio Emilia), Grana, Costantino (University of Modena and Reggio Emilia)

CodeExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelPrompt EngineeringImageTextTabularReview/Survey PaperBenchmarkRetrieval-Augmented Generation

🎯 What it does: This paper proposes an interpretable decision support tool called paper-snitch based on LLM, which is used in the peer review process of medical imaging papers to automatically parse PDFs, extract and statically check code repositories, and generate evidence-supported reproducibility assessment reports based on the MICCAI reproducibility policy.

Think Global, Look Focal: Dual-Stream Retrieval-Augmented Pathology Report Generation for Whole Slide Images

Xiong, Liang (Chongqing Normal University), Cui, ShaoGuo (Chongqing Normal University)

CodeGenerationRetrievalTransformerPrompt EngineeringVision Language ModelImageTextBiomedical DataRetrieval-Augmented Generation

🎯 What it does: This paper proposes a dual-stream retrieval-enhanced pathological report generation framework, DR-Gen, for automatically generating pathological reports.

Through the Schrödinger Bridge: Benchmarking Antemortem Image Restoration from Postmortem Autolysis to Enhance Forensic Diagnostics

Hao, Shuang (Xi'an Jiaotong University), Lian, Chunfeng (Xi'an Jiaotong University)

CodeImage TranslationRestorationTransformerDiffusion modelScore-based ModelGenerative Adversarial NetworkContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyBenchmarkStochastic Differential Equation

🎯 What it does: Propose the task of postmortem self-autopsy image restoration in forensic histopathology, construct the first homologous but unpaired AutoPath dataset, and use the Schrödinger Bridge generative model to achieve image restoration, followed by the design of an evaluation framework based on diagnostic distribution consistency.

Time Matters: Rethinking Diffusion and Flow Models in One-Step Medical Image Translation

Mei, Siyuan (Friedrich-Alexander-Universität Erlangen-Nürnberg), Maier, Andreas (Siemens Healthineers)

CodeImage TranslationGenerationConvolutional Neural NetworkTransformerDiffusion modelFlow-based ModelImageBiomedical DataMagnetic Resonance ImagingPositron Emission TomographyAlzheimer's DiseaseOrdinary Differential Equation

🎯 What it does: Proposes a first-order deterministic generative framework JiR that only utilizes the temporal condition t for medical image translation (MR→PET, T1w→T1ce).

Tooth Alignment via Virtual Trajectories with Clinical Constraints

Jo, SeungKwan (Soongsil University), Chung, Minyoung (Osstem Implant)

CodePose EstimationOptimizationGraph Neural NetworkTransformerDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningPoint CloudBiomedical DataBenchmark

🎯 What it does: Designed a teeth alignment framework based on virtual progressive trajectories, utilizing pre-treatment and post-treatment oral scan data to progressively predict K-step SE(3) transformations, and applying clinical constraints at each step to simulate the incremental process of real orthodontic treatment.

ToothFairy3: Scaling CBCT Maxillofacial Segmentation to 77 Classes with U-Mamba2

Lumetti, Luca (University of Modena and Reggio Emilia), Bolelli, Federico (University of Modena and Reggio Emilia)

CodeSegmentationConvolutional Neural NetworkTransformerBiomedical DataComputed Tomography

🎯 What it does: Proposed a large-scale CBCT facial bone and tooth segmentation dataset called ToothFairy3 (77 classes, 582 scans) and designed an efficient U-Mamba2 architecture.

Topology-Driven Transferability Estimation of Medical Foundation Models for Segmentation

Tang, Jiaqi (Peking University), Chen, Qingchao (Peking University)

CodeSegmentationDomain AdaptationConvolutional Neural NetworkGraph Neural NetworkTransformerContrastive LearningImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: A topology-based transferability estimation framework is proposed to predict the performance of models in medical image segmentation tasks without fine-tuning.

TopoOR: A Unified Topological Scene Representation for the Operating Room

Wang, Tony Danjun (Technical University of Munich), Bastian, Lennart (Technical University of Munich)

CodeRepresentation LearningRobotic IntelligenceGraph Neural NetworkTransformerVision-Language-Action ModelContrastive LearningImageTextMultimodalityAudio

🎯 What it does: Propose TopoOR, which uses composite complexes and high-order attention networks to uniformly model and reason about multi-modal entities and multi-agent relationships in the operating room.

TotalFM: An Organ-Separated 3D-CT Foundation Model Leveraging Large-Scale Routine Clinical Radiology Data

Yamamoto, Kohei (Jichi Medical University), Kikuchi, Tomohiro (Jichi Medical University)

CodeSegmentationAnomaly DetectionConvolutional Neural NetworkTransformerLarge Language ModelVision Language ModelContrastive LearningImageTextBiomedical DataComputed Tomography

🎯 What it does: Propose TotalFM, a 3D-CT foundation model that utilizes large-scale clinical CT data, through automated conversion of raw CT images and reports into organ-level image-text pairs for self-supervised training.

Toward Synergistic Learning for Liver Vessel and Couinaud Segmentations

Qiu, Yue (Chinese University of Hong Kong), Fu, Chi-Wing (Chinese University of Hong Kong)

CodeSegmentationConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelFlow-based ModelRectified FlowNeural Radiance FieldAuto EncoderGenerative Adversarial NetworkContrastive LearningBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Proposes a collaborative learning framework based on vascular skeletons to achieve joint segmentation of liver vessels and Couinaud segments.

Toward Thyroid Cancer BRAF V600E Mutation Prediction via Multimodal Large Language Model: Dataset and Model Development

Hao, Pengfei, Zhu, Lei (Hong Kong University Of Science And Technology (Guangzhou))

CodeClassificationData-Centric LearningDrug DiscoveryTransformerLarge Language ModelSupervised Fine-TuningVision Language ModelContrastive LearningImageTextMultimodalityBiomedical DataUltrasoundElectronic Health Records

🎯 What it does: This paper constructs the first multi-modal thyroid cancer BRAF V600E mutation prediction dataset, MM-BRAF, and proposes a BRAF-MLLM framework based on a multi-modal large language model to achieve mutation prediction.

Towards Direction-Equilibrated Segmentation in 3D Medical Images

Zhu, Zhiqin, Cong, Baisen (State University of New York)

CodeSegmentationTransformerBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasound

🎯 What it does: A direction-aware balanced Mamba network (DEM-Mamba) is proposed to achieve 3D medical image segmentation.

Towards the Digital Dental Models: High-Fidelity CBCT-to-IOS Fusion via Anatomy-Aware Geometric Transformers

Zhou, Hanqing (Beijing University of Posts and Telecommunications), Xiao, Li (Beijing University of Posts and Telecommunications)

CodeRestorationConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningPoint CloudMeshBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: This paper proposes an anatomy-aware geometric transformer (AAGT) to achieve automatic high-precision registration between CBCT and IOS.

Towards Unified Image–Video Modeling Across B-Mode and Contrast-Enhanced Ultrasound for Multi-task Analysis

Huang, Qiang (Macao Polytechnic University), Tan, Tao (Macao Polytechnic University)

CodeClassificationSegmentationTransformerPrompt EngineeringContrastive LearningImageVideoBiomedical DataUltrasound

🎯 What it does: Proposes a unified multi-task framework called UIV-US, which can simultaneously perform image and video segmentation and classification tasks under B-mode and contrast-enhanced ultrasound (CEUS);