MICCAI 2026 Papers — Page 10
International Conference on Medical Image Computing and Computer-Assisted Intervention · 1165 papers
Robust automatic brain vessel segmentation in 3D CTA scans using dynamic 4D-CTA data
Ceballos Arroyo, Alberto Mario (Northeastern University), Jiang, Huaizu (Northeastern University)
SegmentationConvolutional Neural NetworkAuto EncoderContrastive LearningOptical FlowImageVideoBiomedical DataComputed Tomography
🎯 What it does: Based on dynamic 4D-CTA, a semi-automatic annotation workflow is proposed and the DynaVessel dataset is constructed, followed by using nnUNet for arterial and venous segmentation of cerebral vasculature.
Robust Colonoscopy Video Polyp Segmentation via Quality-Aware Memory and Prototype Alignment
Xiong, Sijing (Chongqing University), Liu, Hongyu (Chongqing University)
SegmentationConvolutional Neural NetworkRecurrent Neural NetworkTransformerAuto EncoderVideoBiomedical Data
🎯 What it does: Proposes a query-enhanced cascaded memory framework for multiple polyp segmentation in colonoscopy videos with distortions such as motion blur, specular reflection, and occlusion.
Robust Mitigation of Age-Dependent Confounding Effects via Sample-Difficulty Decorrelation
Kurian, Nikhil Cherian (Adelaide University), Palmer, Lyle J. (Adelaide University)
ClassificationDomain AdaptationAnomaly DetectionContrastive LearningImageBiomedical DataElectronic Health Records
🎯 What it does: This paper proposes a regularization framework that disentangles sample difficulty from age to mitigate age confounding effects in medical image classification;
Robust Tooth Segmentation Under Orthodontic CBCT: A Metal Artifact-Aware Approach
Jo, SeungKwan (Soongsil University), Chung, Minyoung (Osstem Implant)
SegmentationConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningBiomedical DataComputed Tomography
🎯 What it does: This paper proposes a two-stage dental instance segmentation framework, which utilizes metal artifacts in orthodontic CBCT as localization cues. Subsequently, it suppresses the impact of artifacts on boundaries, achieving robust dental detection and segmentation.
ROTE: Routing–Timescale Framework for Balanced Expert Allocation in MoE-Based Continual Segmentation
Zhu, Zhanshi (Harbin Institute of Technology), Li, Shuo (Case Western Reserve University)
SegmentationComputational EfficiencyKnowledge DistillationTransformerMixture of ExpertsContrastive LearningImageMultimodalityBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: This paper proposes ROTE, a continual medical image segmentation framework built upon a frozen SAM base model, which employs a Mixture-of-Experts (MoE) decoder to specifically address the issues of expert allocation imbalance and catastrophic forgetting.
RPG-SAM: Reliability-Weighted Prototypes and Geometric Adaptive Threshold Selection for Training-Free One-Shot Polyp Segmentation
Lin, Weikun (East China Normal University), Wang, Yan (East China Normal University)
SegmentationTransformerPrompt EngineeringAuto EncoderContrastive LearningImageBiomedical Data
🎯 What it does: Proposes a training-free one-time polyp segmentation framework called RPG-SAM, leveraging the base model SAM2 for high-quality segmentation.
RT-Super: Learning Tumor Segmentation from Longitudinal Images and Reports
Bassi, Pedro R. A. S. (Johns Hopkins University), Zhou, Zongwei (Johns Hopkins University)
SegmentationKnowledge DistillationRepresentation LearningData-Centric LearningConvolutional Neural NetworkTransformerLarge Language ModelDiffusion modelContrastive LearningImageTextBiomedical DataComputed TomographyRetrieval-Augmented Generation
🎯 What it does: Leverage the large amount of available imaging reports, longitudinal scans, and multi-phase CT from hospitals to build a teacher-student self-distillation architecture, learning multi-tumor (esophagus, spleen, uterus) segmentation without relying on a large number of manual masks;
S-CEReBrO: Breaking the Memory Barrier in Continuous EEG Monitoring
Bucagu, Glenn Anta (ETH Zurich), Benini, Luca (University of Bologna)
Anomaly DetectionComputational EfficiencyRepresentation LearningTransformerAuto EncoderContrastive LearningTime SeriesBiomedical Data
🎯 What it does: Proposed S-CEReBrO, a Transformer-based continuous EEG foundation model that utilizes windowed alternating attention to achieve constant KV cache memory, and was pre-trained with masked autoencoding on 25,000 hours of EEG corpus.
S-GRPO: Structural Group Relative Policy Optimization for Medical Image Segmentation
Duan, Yiru (Nanjing University of Science and Technology), Chen, Qiang (Nanjing University of Science and Technology)
SegmentationKnowledge DistillationConvolutional Neural NetworkTransformerReinforcement LearningMixture of ExpertsImageBiomedical DataMagnetic Resonance Imaging
🎯 What it does: This paper proposes a structured parallel sampling reinforcement learning framework called S-GRPO, aimed at improving pixel accuracy and global anatomical consistency in medical image segmentation.
S2DLight: Single Step Diffusion Network via Degradation-Aware Adaptation for Surgical Endoscopic Image Low-Light Enhancement
Fei, Song (Hong Kong University of Science and Technology), Zhu, Lei (Hong Kong University of Science and Technology)
RestorationConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningImageBiomedical Data
🎯 What it does: Propose the S2DLight single-step diffusion framework to achieve efficient enhancement of low-light images in surgical endoscopy;
S2GAIN: Spatio-Spectral Guided Affinity Imputation Network for Alzheimer’s Disease Diagnosis with Incomplete Modality
Wu, Yanan (Ningbo University), Wu, Yuehan (Ningbo University)
ClassificationExplainability and InterpretabilityRepresentation LearningData-Centric LearningConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageMultimodalityBiomedical DataMagnetic Resonance ImagingPositron Emission TomographyAlzheimer's Disease
🎯 What it does: Proposed the S GAIN network, achieving early diagnosis of Alzheimer's disease under missing PET modality through cross-modal interaction backbone, multi-stage lightweight imputer, and spatial-frequency joint alignment, affinity-guided reliability weighting, and image-domain reconstruction regularization.
S3H-Net: Injecting Hybrid Semantic-Spatial Guidance via Superpixel Hypergraphs for Medical Image Segmentation
Yang, Zongjian (Heilongjiang University), Ma, Jiquan (Heilongjiang University)
SegmentationDepth EstimationConvolutional Neural NetworkGraph Neural NetworkTransformerDiffusion modelScore-based ModelFlow-based ModelRectified FlowNeural Radiance FieldAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasound
🎯 What it does: Propose the S H-Net network based on a hypergraph of superpixel semantic space to address the challenges of complex target geometry and ambiguous boundaries in medical image segmentation.
SA-CFM: Structure-Aware Continuous Flow Matching for Normative Modeling of EEG Microstate
Hu, Shiang (Anhui University), Lv, Zhao (Anhui University)
ClassificationAnomaly DetectionRepresentation LearningTransformerFlow-based ModelAuto EncoderContrastive LearningTime SeriesBiomedical DataAlzheimer's DiseaseElectrocardiogramOrdinary Differential Equation
🎯 What it does: Propose the Structure-Aware Continuous Flow Matching (SA-CFM) model, which provides a standardized modeling approach for EEG microstate potentials, generating reproducible individual deviation representations.
SAFE-Diff: Scale-Aware Attention and Feature-Dispersive Diffusion with Uncertainty Estimation for Contrast-Enhanced Breast MRI Synthesis
Zhang, Tianyu (Radboud University Medical Center), Mann, Ritse (Netherlands Cancer Institute)
GenerationData SynthesisTransformerDiffusion modelScore-based ModelBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Propose SAFE-Diff, which synthesizes high-fidelity contrast-enhanced images on non-contrast breast MRI using a multi-scale attention diffusion model.
SAFE-Diff: Structurally Anchored Diffusion for Anatomically Faithful CT Image Super-Resolution
Sneha (Indian Institute of Technology Jodhpur), Paul, Angshuman (Indian Institute of Technology Jodhpur)
Super ResolutionConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelImageBiomedical DataComputed Tomography
🎯 What it does: Propose a two-stage CT super-resolution framework called SAFE-Diff: the first stage uses a residual prediction network to recover structure; the second stage uses truncated diffusion to refine details, and retains low-frequency structure and high-frequency details through SWT/ISWT frequency domain fusion.
SAGEAgent: A Self-evolving Agent for Cost-Aware Modality Acquisition in Multimodal Survival Prediction
Qu, Chongyu (Vanderbilt University), Huo, Yuankai (Vanderbilt University)
Explainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelReinforcement LearningMultimodalityBiomedical DataElectronic Health RecordsRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose a self-evolving LLM agent called SAGEAgent for cost-aware sequential modal acquisition in multimodal survival prediction.
SAM-Cluster and SAM-LP: Structure-Aware Evaluation Metrics for Selecting WSI Feature Extractors Without MIL Training
Kim, Juhyeon (Korea Advanced Institute of Science and Technology), Yi, Mun Yong (Korea Advanced Institute of Science and Technology)
ClassificationImage TranslationSegmentationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningData-Centric LearningTransformerPrompt EngineeringDiffusion modelAuto EncoderContrastive LearningImageBiomedical DataBenchmark
🎯 What it does: Propose a method to pre-evaluate whole slide image (WSI) feature extractors (FE) without requiring complete multiple instance learning (MIL) training, utilizing the SAM model to generate structured masks and designing two structure-aware evaluation metrics (SAM-Cluster and SAM-LP) to rank candidate FEs.
SAM2 as a Key Bridge Between 2D Slices and 3D Volumes for Sparsely-supervised Medical Image Segmentation
Yu, Peng (East China Normal University), Wang, Yan (East China Normal University)
SegmentationTransformerVision Language ModelAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: This paper studies how to achieve complete 3D medical image segmentation using only three orthogonal central 2D slices with annotations, leveraging SAM2 for sparse annotation scenarios, and proposes a unified framework that integrates multi-view propagation, cross-view consistency weighted fusion, and dynamic pseudo-label refinement.
SAMoCo: A Learning-Initialized and Physics-Refined Framework for Sampling-Agnostic Motion Correction in MRI
Liu, Weijia (Shanghai Jiao Tong University), Jiang, Dengrong (Shanghai Jiao Tong University)
RestorationConvolutional Neural NetworkRecurrent Neural NetworkContrastive LearningOptical FlowBiomedical DataMagnetic Resonance Imaging
🎯 What it does: This paper proposes a hybrid framework (SAMoCo) that combines deep learning initialization with physics-driven self-focusing for motion correction in 2D multi-ray MRI.
SANT-CBM: Structurally-Aware and Noise-Tolerant Semi-supervised Concept Bottleneck Models
Liu, Hongwei (Shanghai Normal University), Lai, Songning (HKUST)
ClassificationExplainability and InterpretabilityData-Centric LearningConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningGaussian SplattingImageBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound
🎯 What it does: Proposes a structure-aware and noise-tolerant semi-supervised concept bottleneck model, SANT-CBM, to address the issues of scarce concept labels and noisy pseudo-labels in medical image diagnosis.
SAR-Net: Sample-Agnostic Retrieval Network for Robust Brain Tumor Segmentation with Missing Modalities
Zheng, Shenhai (Chongqing University of Posts and Telecommunications), Li, Weisheng (Chongqing University of Posts and Telecommunications)
SegmentationRetrievalRepresentation LearningConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningMultimodalityBiomedical DataMagnetic Resonance ImagingRetrieval-Augmented Generation
🎯 What it does: Proposed a Sample-Agnostic Retrieval Network (SAR-Net) that completes missing modalities and performs brain tumor segmentation through a retrieval-based approach.
SAVer: Structure-Aware Validation of Standard Fetal Head Planes via Clinically Grounded Anatomical Completeness
Amba Prasanna, Anusha (Indian Institute Of Technology Madras), Sivaprakasam, Mohanasankar (Indian Institute Of Technology Madras)
ClassificationSegmentationAnomaly DetectionConvolutional Neural NetworkTransformerPrompt EngineeringScore-based ModelContrastive LearningImageBiomedical DataUltrasound
🎯 What it does: This paper proposes a structure-aware standard fetal head plane verification framework (SAVER), which simultaneously performs multi-structure segmentation and plane classification, and introduces an anatomical reasoning engine to determine whether the image meets clinical standards.
SAW: Toward a Surgical Action World Model via Controllable and Scalable Video Generation
Rapuri, Sampath (Johns Hopkins University), Unberath, Mathias (Johns Hopkins University)
GenerationData SynthesisTransformerPrompt EngineeringVision-Language-Action ModelDiffusion modelWorld ModelImageVideoTextBiomedical Data
🎯 What it does: Proposed the SAW model, achieving controllable surgical video generation based on lightweight conditions (language prompts, reference frames, tissue availability masks, tool trajectories).
Scalable Training of Spatially Grounded 2D Vision–Language Models for Radiology
Salcan, Yusuf (University of Freiburg), Brox, Thomas (University of Freiburg)
ClassificationRecognitionObject DetectionSegmentationAnomaly DetectionTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingComputed TomographyRetrieval-Augmented Generation
🎯 What it does: Trained a CT/MRI VLM called RadGrounder for visual localization, supporting report generation, VQA, and localization.
SCISSR: Scribble-Conditioned Interactive Surgical Segmentation and Refinement
Ping, Haonan (Shanghai Jiao Tong University), Ban, Yutong (Shanghai Jiao Tong University)
SegmentationConvolutional Neural NetworkTransformerSupervised Fine-TuningPrompt EngineeringAuto EncoderImageBiomedical DataComputed TomographyUltrasound
🎯 What it does: Proposes the SCISSR framework, which achieves interactive surgical image segmentation through hand-drawn sketch prompts, supporting multi-round revisions.
SCKAN: Structural Consensus-Based KAN Prototype Learning for Semi-supervised Pancreas Segmentation
Liu, Yuqi, Li, Shuo (Tongji University)
SegmentationTransformerAuto EncoderContrastive LearningImageBiomedical DataComputed Tomography
🎯 What it does: Proposes a structure consensus-based KAN prototype learning framework, SCKAN, to address the problem of supervisory bias in pancreatic segmentation.
ScribbleDose: Scribble-Guided Dose Prediction in Radiotherapy
Zhang, Zhenxi (Hong Kong Polytechnic University), Ren, Ge (Hong Kong Polytechnic University)
Image TranslationRestorationSegmentationGenerationConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningImageBiomedical DataComputed TomographyRetrieval-Augmented Generation
🎯 What it does: Studied the use of sparse scribbles as structural hints for radiation therapy dose prediction.
SD-SAM: Synergizing SAM and DINOv3 for Parameter-Efficient Polyp Segmentation and a Challenging Benchmark Dataset
Yang, Haoran (Qufu Normal University), Liu, Jianlei (Qufu Normal University)
SegmentationComputational EfficiencyKnowledge DistillationConvolutional Neural NetworkTransformerSupervised Fine-TuningScore-based ModelContrastive LearningGaussian SplattingImageBiomedical DataBenchmark
🎯 What it does: A two-stage parameter-efficient multi-scale multi-polar segmentation framework called SD-SAM is proposed: the first stage uses high-confidence Gaussian heatmap regression to locate the polyp center points and generate prompts; the second stage implements DINOv3-guided hierarchical fine-tuning on the SAM encoder, with shallow layers using cosine similarity distillation to enhance edge perception, and deep layers injecting semantic priors through QKV feature fusion, with overall parameters less than 3.5M.
Seamless Whole Slide Label-Free Virtual Staining
Kwark, Dou Hoon (University of Illinois Urbana-Champaign), Bhargava, Rohit (University of Illinois Urbana-Champaign)
Image TranslationRestorationSegmentationGenerationConvolutional Neural NetworkTransformerDiffusion modelGenerative Adversarial NetworkImageBiomedical DataMagnetic Resonance ImagingComputed TomographyRetrieval-Augmented Generation
🎯 What it does: This work proposes Consistency Memory Bank (COMB), a framework for full-slide label-free virtual staining, aiming to eliminate artifacts such as stitching and color drift caused by block-based inference.
SeedPro: Place Brachytherapy Seeds Like Expert Clinicians via Hierarchical Reinforcement Agents
Li, Haitao (Shanghai Jiao Tong University), Chen, Xiaojun (Shanghai Jiao Tong University)
Recurrent Neural NetworkTransformerReinforcement LearningAuto EncoderImageBiomedical DataComputed Tomography
🎯 What it does: Developed SeedPro, an automatic radioactive seed pre-planning framework based on hierarchical reinforcement learning, capable of generating radioactive seed layout plans comparable to those of experts.
Seeing Through Multiple Views: Parameter-Efficient Fine-Tuning via Selective Neurons for Consistent Radiology Report Generation
Chen, Yucheng (Nanyang Technological University), Yeo, Si Yong (Nanyang Technological University)
GenerationExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelSupervised Fine-TuningVision Language ModelImageTextBiomedical DataMagnetic Resonance ImagingComputed TomographyElectronic Health Records
🎯 What it does: Proposed a parameter-efficient framework called View-PNDF for multi-view chest X-ray report generation by detecting and fine-tuning view-specific pattern neurons.
Seeing Through Smoke: Surgical Desmoking for Improved Visual Perception
Lu, Jingpei (Intuitive Surgical), Mohareri, Omid (Intuitive Surgical)
Image TranslationRestorationSegmentationDepth EstimationTransformerDiffusion modelContrastive LearningImageBiomedical DataComputed Tomography
🎯 What it does: Proposed a Transformer-based surgical smoke removal model and constructed a large-scale synthetic and real paired dataset to achieve high-quality smoke removal;
SegDINO: Introducing Multi-Scale Structure into DINO for Efficient Medical Image Segmentation
Yang, Sicheng (Hong Kong University of Science and Technology (Guangzhou)), Zhu, Lei (Sun Yat-sen University)
SegmentationConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningImageBiomedical DataComputed TomographyUltrasoundBenchmark
🎯 What it does: This study proposes the SegDINO framework, which transfers the pre-trained DINOv3 visual model to medical image segmentation tasks, achieving efficient segmentation through lightweight multi-scale modeling;
Segment Anything Model with Symmetric Spectral Uncertainty Gating for Generalizable Medical Image Segmentation
Li, He (University of Electronic Science and Technology of China), Zhang, Shaoting (University of Electronic Science and Technology of China)
SegmentationDomain AdaptationTransformerSupervised Fine-TuningDiffusion modelContrastive LearningImageBiomedical DataMagnetic Resonance ImagingBenchmark
🎯 What it does: Propose a single-source domain generalization medical image segmentation framework called SSUG based on the Segment Anything Model (SAM), which utilizes three modules: Spectral Manifold Extrapolation, Uncertainty-aware Symmetric Channel Gating, and Hierarchical Decoupled Consistency, to jointly suppress style sensitivity and extract domain-invariant anatomical features.
Selective Confidence Fusion for Specular Highlight Detection in Medical Images
Hu, Daosong (Sun Yat Sen University), Huang, Kai (Sun Yat Sen University)
SegmentationAnomaly DetectionConvolutional Neural NetworkDiffusion modelAuto EncoderContrastive LearningImageVideoBiomedical Data
🎯 What it does: This paper proposes an adaptive confidence fusion framework tailored for endoscopic images, aimed at accurately detecting and removing specular highlights.
Selective Rank-1 Orthogonality Regularization: Rethinking Stability-Plasticity Balance in Continual Medical Image Classification
Zhang, Jingyang (Southeast University), Ma, Lei (Tongji University)
ClassificationTransformerSupervised Fine-TuningPrompt EngineeringContrastive LearningImageBiomedical DataBenchmark
🎯 What it does: Propose Selective Rank-1 Orthogonality Regularization (SR1OR), which addresses the stability-plasticity balance problem in medical image continual learning by imposing selective orthogonality constraints on unit rank-1 subspaces within the LoRA low-rank update space.
Self-auditing Parameter-Efficient Fine-Tuning for Few-Shot 3D Medical Image Segmentation
Ly, Son Thai (University of Houston), Nguyen, Hien V. (University of Houston)
SegmentationComputational EfficiencyData-Centric LearningTransformerSupervised Fine-TuningContrastive LearningPoint CloudBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Propose SEA-PEFT, a parameter-efficient fine-tuning method that automatically configures adapters through an online self-audit process, specifically designed for few-shot 3D medical image segmentation.
Self-Supervised Implicit CEST Reconstruction via Physics-Informed Lorentz Encoding
Li, Dexuan (East China Normal University), Yang, Guang (East China Normal University)
OptimizationComputational EfficiencyRepresentation LearningTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Proposes a physics-informed implicit neural representation framework called Lorentz Encoding (LE), for self-supervised reconstruction of high-resolution CEST Z spectra under sparse sampling, and generates continuous, interpretable metabolic maps through physical constraints.
Self-supervised Learning on Lingual Ultrasound Video Encodes Clinically Meaningful Articulatory Structure of Rhotic Production in Residual Speech Sound Disorder
Kim, Taehyo, Shu, Hai (New York University)
ClassificationRepresentation LearningTransformerAuto EncoderContrastive LearningVideoBiomedical DataUltrasound
🎯 What it does: This paper proposes a joint embedding framework based on self-supervised learning, which learns the tongue position structure of American English rhotic pronunciation using tongue ultrasound videos, and reduces speaker differences through speaker adversarial regularization.
Self-supervised Longitudinal Neuroimaging Augmentation for Continuous Staging of Alzheimer’s Disease
Liu, Haoyuan (Indiana University Indianapolis), Yan, Jingwen (Indiana University Indianapolis)
GenerationData SynthesisOptimizationExplainability and InterpretabilityHyperparameter SearchGenerative Adversarial NetworkContrastive LearningBiomedical DataPositron Emission TomographyAlzheimer's Disease
🎯 What it does: Propose a self-supervised residual GAN framework that generates feasible follow-up Amyloid PET images on single-visit subjects by leveraging the biological prior of unidirectional amyloid deposition, and use these synthetic samples to enhance the SLOPE continuous staging model.
Self-supervised Spline Autoencoder for Smooth Latent Dynamics in Echocardiography
Kwan, Leong Chit Jeff (Hong Kong University of Science and Technology), Chung, Albert C. S. (Hong Kong University of Science and Technology)
Anomaly DetectionRepresentation LearningDiffusion modelAuto EncoderContrastive LearningOptical FlowVideoBiomedical DataUltrasound
🎯 What it does: This study proposes a self-supervised Spline Autoencoder, which learns low-rank latent trajectories of cardiac motion through smooth curve fitting, and achieves unlabeled cardiac cycle and ED/ES event localization by utilizing spectral embedding and phase projection.
Self-supervised Temporal Regularization for Landmark-Based Cardiac Segmentation with Automatic AHA Regional Mapping
Montalvo-García, David (Universidad Politécnica de Madrid), Ferrante, Enzo (Universidad de Buenos Aires)
SegmentationConvolutional Neural NetworkGraph Neural NetworkAuto EncoderContrastive LearningOptical FlowImageVideoBiomedical DataUltrasound
🎯 What it does: Building upon the previously learned implicit correspondence cardiac graph network Mask‑HybridGNet Dual, this study introduces self-supervised temporal regularization (velocity and acceleration penalty) as a post-training phase to enhance temporal consistency in ultrasound sequence segmentation, and utilizes the learned correspondence to achieve automatic AHA 17-segment standard mapping, completing regional motion analysis.
Semantic Feature Modulation for Mammographic Lesion Classification
Mahpod, Shahar (Ariel University), Ben Artzi, Gil (Ariel University)
ClassificationAnomaly DetectionConvolutional Neural NetworkTransformerContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Introduce a semantic feature modulation framework based on BI-RADS descriptors in breast X-ray image classification to achieve dynamic fusion of images and clinical semantics.
Semantic-Aware Organ-Level Esophageal Tumor Synthesis via Latent Rectified Flow
Yu, Qinji (DAMO Academy, Alibaba Group), Zhang, Ling (Fudan University)
SegmentationGenerationData SynthesisConvolutional Neural NetworkTransformerPrompt EngineeringDiffusion modelScore-based ModelFlow-based ModelRectified FlowAuto EncoderContrastive LearningImageBiomedical DataComputed Tomography
🎯 What it does: Propose a semantic-aware organ-level generative framework called SOFT, aimed at synthesizing complete tumors and the resulting organ shape changes while keeping the background unchanged, and automatically generating corresponding labels.
Semantic-Consistent Dual-Image Vision-Language Alignment via Modality Routing for 3D PET/CT
Lu, Lin (Tsinghua University), Zhang, Hui (Tsinghua University)
RetrievalAnomaly DetectionTransformerVision Language ModelContrastive LearningMultimodalityBiomedical DataComputed TomographyPositron Emission Tomography
🎯 What it does: Proposed the PETCT‑DiVLA framework, achieving 3D PET/CT dual-branch vision-language alignment, automatically distinguishing PET and CT information and performing global and local multi-grained matching.
Semantically Consistent Whole-Body PET/CT Report Generation with Local Lesion-Level Guidance
Yu, Mingyang (ShanghaiTech University), Shen, Dinggang (ShanghaiTech University)
GenerationTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringContrastive LearningImageTextBiomedical DataComputed TomographyPositron Emission TomographyRetrieval-Augmented Generation
🎯 What it does: Propose a full-body PET/CT report generation framework that compresses a variable number of lesions into fixed entities using local lesion prompts and generates complete reports.
Sequence-Conditioned Flow-Based Models for Digital Phantom Generation in MRI
Belousova, Kseniya (ITMO University), Al-Haidri, Walid (ITMO University)
GenerationData SynthesisConvolutional Neural NetworkDiffusion modelScore-based ModelFlow-based ModelImageBiomedical DataMagnetic Resonance Imaging
🎯 What it does: This study proposes a self-supervised hierarchical flow model for estimating corresponding quantitative T1, T2, and PD maps from T1, T2, and PD-weighted MRI images, and for synthesizing new weighted images using the MR signal model.
Set-Based Groupwise Registration for Variable-Length, Variable-Contrast Cardiac MRI
Zhang, Yi (Delft University of Technology), Tao, Qian (Delft University of Technology)
Image TranslationRestorationConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningOptical FlowImageBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Propose a set-based inter-group registration framework called Any Reg, for motion correction of cardiac MRI sequences with variable length and contrast.
Set-Inclusive Uncertainty Modeling for Robust Brain Tumor Segmentation
Baek, Seunghun (Pohang University of Science and Technology), Kim, Won Hwa (Pohang University of Science and Technology)
SegmentationConvolutional Neural NetworkContrastive LearningGaussian SplattingMultimodalityBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Propose a probabilistic representation framework that models the task representation of each modality subset as a Gaussian distribution, utilizing the set inclusion hierarchy to guide mean alignment and variance to reflect information missing, thus achieving brain tumor segmentation under missing modalities;
SF-DINO: Spatial-Frequency Adapted Foundation Model for Multiple Myeloma Diagnosis
Ye, Zhaoyi (Wuhan University), Lei, Cheng (Wuhan University)
ClassificationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningDrug DiscoveryConvolutional Neural NetworkTransformerPrompt EngineeringContrastive LearningImageBiomedical Data
🎯 What it does: Propose the SF-DINO model, which uses a spatial morphology adapter and frequency-selective hash attention for fine-grained classification of multiple myeloma cells;
Shape-Adaptive Guidance Signal for Interactive Cortical Sulcal Labeling
Son, Jiwon (POSTECH), Lyu, Ilwoo (POSTECH)
SegmentationConvolutional Neural NetworkDiffusion modelContrastive LearningBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Propose an interactive cortical sulcus annotation framework based on shape-adaptive guiding signals, using spherical CNN to perform binary segmentation of the left posterior frontal sulcus.
Shape-Aware Registration Without Pretraining: CT-Ultrasound Bone Surface Alignment via Instance-Level Shape Completion
Wu, Luohong (University of Zurich), Fürnstahl, Philipp (University of Zurich)
RestorationOptimizationComputational EfficiencyTransformerDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningPoint CloudBiomedical DataComputed TomographyUltrasound
🎯 What it does: Propose an end-to-end CT-US bone surface registration framework called ComReg, which simultaneously completes the bone surface reconstruction and registration by utilizing instance-level shape completion.
Shape-Scale Aware Diffusion Models for Weakly Supervised Medical Image Anomaly Detection
Zheng, Shenhai (Chongqing University of Posts and Telecommunications), Li, Laquan (Chongqing College of Artificial Intelligence)
Anomaly DetectionConvolutional Neural NetworkDiffusion modelScore-based ModelContrastive LearningBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Proposed a direction-guided shape-scale-aware diffusion model for weakly supervised medical image anomaly detection.
ShapeEFM: A Shape-Based Pretrained Foundation Model on ECG Morphology Comprehension
Liu, Huan (Ant Group), Lu, Le (Zhejiang University School of Medicine)
Explainability and InterpretabilityRepresentation LearningTransformerSupervised Fine-TuningAuto EncoderContrastive LearningTime SeriesBiomedical DataElectrocardiogram
🎯 What it does: Propose ShapeEFM, a shape-based pre-trained ECG foundation model, which decomposes ECG waveforms into Cardio-Morphemes, treating them as 'words,' and uses Transformer for self-supervised learning to enhance ECG morphology understanding and interpretability.
ShapKO: Shapley-Adaptive Modality Knockout for Robust Multimodal Learning
Nizam, Nusrat Binta (Cornell University), Sabuncu, Mert R. (Weill Cornell Medicine)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerVision Language ModelAuto EncoderContrastive LearningImageTextMultimodalityTabularBiomedical DataElectronic Health Records
🎯 What it does: By dynamically estimating modal importance using Shapley values during training, adaptively adjusting the elimination probability for each modality, thereby enhancing the robustness of multi-modal models under missing modality scenarios.
Shared Gaussian Geometry for Brain MRI Spatial Resolution Harmonization Without External Training
Gao, Yifan (Cedars-Sinai Medical Center), Li, Debiao (Cedars-Sinai Medical Center)
Image HarmonizationRestorationSuper ResolutionTransformerDiffusion modelAuto EncoderContrastive LearningGaussian SplattingImageBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Propose a shared Gaussian geometric framework to achieve spatial resolution harmonization of multi-contrast brain MRI without the need for external training sets.
SHINE: An Entropy-Guided Digital Twin Framework for 3D Myocardial Scar Reconstruction from Sparse CMR
Zong, Dexiang, Zhong, Liang (National Heart Centre Singapore)
SegmentationGenerationData SynthesisConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataMagnetic Resonance Imaging
🎯 What it does: A continuous function learning framework based on sparse LGE-MRI was constructed to generate a three-dimensional digital twin of myocardial scar.
SHOT-CCR: Biologically Guided Adversarial Training for Test-Time Adaptation in Cellular Morphology
Dee, William (Queen Mary University of London), Slabaugh, Gregory (Queen Mary University of London)
ClassificationDomain AdaptationConvolutional Neural NetworkGenerative Adversarial NetworkContrastive LearningBiomedical Data
🎯 What it does: Propose a biology-guided test-time adaptation framework called SHOT-CCR based on cell count gradient reversal, which is used to eliminate batch effects in Cell Painting data and improve the accuracy of gene perturbation classification.
SIGMA-VAE: Sketched Isotropic Gaussian Multi-View Autoencoder for Normative Modeling
Zabihi, Mariam (University Hospital Tübingen), Wolfers, Thomas
Anomaly DetectionRepresentation LearningTransformerMixture of ExpertsScore-based ModelAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease
🎯 What it does: Proposed the SIGMA-VAE model for normative modeling on large-scale structural MRI data and quantifying individual brain abnormalities.
Single-Stage Hierarchical Rectification for Weakly Supervised Histopathology Segmentation
Nguyen, Trong Duc (VinUniversity), Pham, Huy Hieu (Posts and Telecommunications Institute of Technology)
SegmentationConvolutional Neural NetworkRectified FlowContrastive LearningBiomedical Data
🎯 What it does: Propose a single-stage hierarchical correction framework, SSHR, which achieves weakly supervised histopathological semantic segmentation by correcting intermediate features during the forward pass, avoiding multi-stage CAM generation and recursive pseudo-labeling.
Single-Subject Multi-view MRI Super-Resolution via Implicit Neural Representations
Kim, Heejong (Weill Cornell Medicine), Sabuncu, Mert R. (Weill Cornell Medicine)
Super ResolutionNeural Radiance FieldAuto EncoderOptical FlowBiomedical DataMagnetic Resonance Imaging
🎯 What it does: A single-organ multi-view MRI super-resolution framework based on implicit neural representations was studied, which can reconstruct isotropic high-resolution images from multi-directional anisotropic scans without the need for pre-registration.
SIRA: Reasoning-Aware Surgical Instrument Segmentation via Query-Anchored Alignment
Zhang, Zhibo (Huazhong University of Science and Technology), Yan, Zengqiang (Huazhong University of Science and Technology)
SegmentationTransformerLarge Language ModelVision Language ModelContrastive LearningImageTextMultimodalityBiomedical DataBenchmark
🎯 What it does: Proposed a query-conditioned surgical instrument segmentation framework called SIRA, achieving multi-modal semantic alignment segmentation tasks on the newly constructed SurgRS dataset.
Situational context embedding and knowledge-based grounding for LLMs in surgical training assistance
Bahari Malayeri, Ali (University of Zurich), Fürnstahl, Philipp (Zurich University of Applied Sciences)
Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningRobotic IntelligenceDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringTextBiomedical DataReview/Survey PaperRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposed and implemented a tutoring system for total hip arthroplasty (THA) training simulator based on large language models, which provides context-aware answers by embedding real-time situational information from the simulator and combining it with knowledge base retrieval.
SkepticalNet: A Safety-Aware Interactive Segmentation Framework for Neuro-Oncology
Kyriazis, Efstathios (Foundation for Research and Technology Hellas), Marias, Kostas (Foundation for Research and Technology Hellas)
SegmentationConvolutional Neural NetworkPrompt EngineeringContrastive LearningMultimodalityBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Proposes a multi-modal 3D interactive segmentation framework called SkepticalNet for precise segmentation of neural tumors;
Skill-Evolving Grounded Reasoning for Free-Text Promptable 3D Medical Image Segmentation
Zhang, Tongrui (Fudan University), Shan, Hongming (Fudan University)
SegmentationExplainability and InterpretabilityKnowledge DistillationData-Centric LearningTransformerSupervised Fine-TuningPrompt EngineeringVision Language ModelImageTextMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This paper proposes SEER, a framework that addresses the inconsistency between linguistic diversity and anatomical accuracy in free-text prompting for 3D medical image segmentation by explicitly reasoning about 'skills' and dynamically evolving skills.
Skin2Mind: A Multimodal Framework for Mental Health Risk Screening via Facial Images and Structured Skin Reports
Liu, Jiahe (Monash University), Ge, Zongyuan (Monash University)
ClassificationExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelVision Language ModelContrastive LearningImageTextMultimodalityElectronic Health Records
🎯 What it does: Proposed and implemented Skin2Mind, a multimodal framework based on multispectral facial skin images and structured skin reports, for predicting individual-level DSM-5 level mental health risks.
SLICE-MIL: Counterfactual Semantic Anchoring of Disentangled Evidence for Child-Pugh Grading
Deng, Yihan (University of Science and Technology of China), Zheng, Jian (University of Science and Technology of China)
ClassificationExplainability and InterpretabilityTransformerContrastive LearningImageBiomedical DataComputed Tomography
🎯 What it does: Proposed a SLICE-MIL framework based on weakly supervised multiple instance learning, using anatomy-guided dual-query attention and adversarial semantic anchoring to estimate Child-Pugh grades from portal-phase CT imaging.
SlideGuard: WSI Artifact Detection Benchmark
Kaczmarek, Gabriela (IDEAS Research Institute), Swiderska-Chadaj, Zaneta (IDEAS Research Institute)
ClassificationObject DetectionSegmentationAnomaly DetectionConvolutional Neural NetworkTransformerSupervised Fine-TuningAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataBenchmark
🎯 What it does: Proposed the SlideGuard benchmark to evaluate the performance of AI-based quality control tools for multi-source, multi-species whole slide images (WSI) in artifact detection.
Slim nnU-Net: Revisiting nnU-Net Self Configuration from an Optimization Angle
Zhao, Yidong (Delft University of Technology), Tao, Qian (Delft University of Technology)
SegmentationOptimizationComputational EfficiencyConvolutional Neural NetworkAuto EncoderImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: This paper reformulates the training of nnU-Net as a proximal optimization problem with a hybrid L1+Group L2 sparse regularization, and achieves extremely sparse Slim nnU-Net by performing channel-wise pruning during training.
SMART-KNEE: Sparse-Morphology Based Anatomical Reconstruction Using Topology-Aware Networks for Imageless Total Knee Arthroplasty
Rajasekar, Durga (Indian Institute of Technology Madras), Sivaprakasam, Mohanasankar (Indian Institute of Technology Madras)
RestorationGenerationGraph Neural NetworkDiffusion modelAuto EncoderGenerative Adversarial NetworkPoint CloudMeshGraphBiomedical DataComputed TomographyReview/Survey Paper
🎯 What it does: Proposed a two-stage framework named SMART‑KNEE, which can predict the complete femoral and tibial landmark sets using only four anatomical landmarks acquired intraoperatively, and fuse them with local surface point clouds to achieve patient-specific bone surface reconstruction without pre-imaging data.
Smartphone-Based ASD Screening via Instance-Adaptive Cross-Modal Fusion and Contrastive Learning
Xu, Jie (Northwestern Polytechnical University), Xia, Chen (Northwestern Polytechnical University)
ClassificationAnomaly DetectionConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningVideoMultimodality
🎯 What it does: Implementing an autism spectrum disorder screening based on eye movement and head pose on smartphones, proposing the AMF-Net framework.
SONIC: Sonographer-Inspired Conceptual Explanation for Neural Networks
Wang, Yingni (Tsinghua University), Liao, Hongen (Shanghai Jiaotong University)
RecognitionExplainability and InterpretabilityConvolutional Neural NetworkGraph Neural NetworkAuto EncoderContrastive LearningImageBiomedical DataUltrasound
🎯 What it does: Propose SONIC, an interpretable framework based on medical anatomical concepts, used to identify fetal ultrasound standard planes.
SonoCLIP: Mask-Guided Region-Aware Vision–Language Pretraining for Fetal Ultrasound Analysis
Su, Hang (Wuhan University), Du, Bo (Wuhan University)
RecognitionSegmentationTransformerVision Language ModelContrastive LearningImageTextBiomedical DataUltrasound
🎯 What it does: This paper proposes SonoCLIP, a fetal ultrasound foundation model based on vision-language alignment, which can achieve region-controllable representation learning at both global and local levels.
Source-Free Domain Adaptation for Medical Image Segmentation via LLM-Agent Collaboration
Zhang, Tianyu (Hangzhou Dianzi University), Wang, Shuai (Hangzhou Dianzi University)
SegmentationDomain AdaptationTransformerLarge Language ModelAgentic AIPrompt EngineeringVision Language ModelContrastive LearningImageMultimodalityBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Propose the RaMA framework, which generates spatial constraints using a multimodal LLM and combines them with a multi-agent SAM model to produce consistent pseudo-labels, achieving reliable self-training in source-free domain adaptation tasks.
Sparse Autoencoders for Interpretable Medical Image Representation Learning
Wesp, Philipp (Stanford University), Gatidis, Sergios (Stanford University)
RetrievalExplainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelVision Language ModelAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: This paper proposes using a sparse autoencoder (Matryoshka SAE + BatchTopK) to transform dense embeddings from medical imaging foundational models (BiomedParse, DINOv3) into interpretable sparse features, and validates their interpretability and performance through various retrieval and evaluation methods;
Sparse Local Latents for Explainable Zero-Shot Reasoning in Medical Vision-Language Models
Goetze, Martin (University Hospital Aachen), Truhn, Daniel (University Hospital Aachen)
ClassificationExplainability and InterpretabilityKnowledge DistillationTransformerPrompt EngineeringVision Language ModelAuto EncoderContrastive LearningImageTextBiomedical DataComputed Tomography
🎯 What it does: This paper extracts locally text-aligned features from MedSigLIP through distillation and token alignment, and uses a TopK sparse autoencoder to learn sparse latents, constructing an interpretable zero-shot reasoning probe while maintaining text alignment capabilities;
Sparsely Supervised Surgical Video Segmentation with Reliable Asymmetric Dual Memory
Sapkota, Nishchal (University of Notre Dame), Chen, Danny Z. (University of Notre Dame)
SegmentationConvolutional Neural NetworkTransformerVision-Language-Action ModelContrastive LearningVideo
🎯 What it does: Proposes D-GEM, a reliable asynchronous dual-memory framework for semantic segmentation of surgical videos under sparse supervision.
Spatially-Aware Graph Geometric Representation for Lung Nodules Detection
Guayacán, Luis (Universidad Industrial de Santander), Martinez, Fabio (Universidad Industrial de Santander)
Object DetectionRepresentation LearningGraph Neural NetworkDiffusion modelContrastive LearningImageGraphBiomedical DataComputed Tomography
🎯 What it does: Propose a geometry graph model based on superpixels, using second-order covariance embedding and differentiable pooling (DiffPool) for nodule detection in lung CT images.
Spatiotemporal Modeling of Longitudinal MRI with Causal Graph Reasoning for Breast Cancer pCR Prediction
Lin, Haiwei (Shenzhen University), Huang, Bingsheng (Shenzhen University)
ClassificationExplainability and InterpretabilityGraph Neural NetworkTransformerVision Language ModelScore-based ModelContrastive LearningImageTabularBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Predict the probability of pathologic complete response (pCR) to preoperative chemotherapy in breast cancer, integrating multi-center longitudinal DCE-MRI and clinical indicators, by proposing a spatiotemporal attention and causal graph reasoning framework based on a medical foundation model.
Spatiotemporal Modeling of Tumor Dynamics via Optimal Transport for Enhanced Longitudinal Analysis
Yang, Jiannan (Stony Brook University), Veeraraghavan, Harini (Memorial Sloan Kettering Cancer Center)
Image TranslationSegmentationGenerationData SynthesisAnomaly DetectionConvolutional Neural NetworkTransformerDiffusion modelFlow-based ModelContrastive LearningOptical FlowImagePoint CloudBiomedical DataComputed TomographyOrdinary Differential Equation
🎯 What it does: Proposed a framework called OxFlow that combines optimal transport (OT) mask interpolation with flow matching (FM) generative models to synthesize intermediate-time lung cancer CT images, enabling visualization and prediction of longitudinal tumor dynamics.
SPEAR: Error-Guided Sparse Refinement for Scalable Deformable Medical Image Registration
Bai, Hao (Shanghai Jiao Tong University), Hong, Yi (Shanghai Jiao Tong University)
Image TranslationSegmentationComputational EfficiencyConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningOptical FlowImageBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Propose the SPEAR framework, which achieves high-resolution 3D medical image deformation registration through error-guided sparse refinement.
SpecF2M: A Spectral-Aware Multi-task Network Estimating Axial Length and Refractive Error from Pediatric Fundus Photographs
He, Mengxian (Chinese University of Hong Kong), Yuan, Wu (Chinese University of Hong Kong)
RecognitionImage TranslationRestorationConvolutional Neural NetworkTransformerMixture of ExpertsAuto EncoderContrastive LearningImageBiomedical DataAlzheimer's DiseaseElectronic Health Records
🎯 What it does: Propose a spectral-aware multi-task network (SpecF2M) that simultaneously estimates axial length (AL) and the spherical (SPH) and cylindrical (CYL) components of spherical equivalent refractive error (SER) from 45° posterior pole retinal photographs of children.
Spectral-Semantic Consensus for Few-Shot Whole Slide Image Classification
Du, Hanyu (Wuhan University), Xu, Yongchao (Wuhan University)
ClassificationAnomaly DetectionExplainability and InterpretabilityComputational EfficiencyRepresentation LearningMeta LearningTransformerPrompt EngineeringVision Language ModelAuto EncoderContrastive LearningImageTextBiomedical DataBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposes a Spectral-Semantic Consensus (SSC) framework based on text priors for few-shot whole-slide image classification, automatically selecting diagnostic information-rich patches and aggregating them.
Spectral-Temporal State Space Modeling on Functional Brain Networks
Sim, Jaeyoon (Pohang University of Science and Technology), Kim, Won Hwa (Pohang University of Science and Technology)
ClassificationExplainability and InterpretabilityGraph Neural NetworkTransformerAuto EncoderContrastive LearningGraphTime SeriesBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease
🎯 What it does: Proposed a Spectral-Temporal Mamba model that combines graph spectral frequency domain information with state space models to achieve window-free, frequency-driven spatiotemporal modeling.
Speech-Guided Multimodal Learning for Vocal Tract Segmentation in Real-Time MRI
Liu, Daiqi (Friedrich-Alexander-Universität Erlangen-Nürnberg), Pérez-Toro, Paula Andrea (Friedrich-Alexander-Universität Erlangen-Nürnberg)
SegmentationData-Centric LearningConvolutional Neural NetworkTransformerContrastive LearningImageMultimodalityBiomedical DataMagnetic Resonance ImagingAudio
🎯 What it does: This study proposes a three-stage multimodal learning framework for precise segmentation of vocal tract organs using speech-guided real-time MRI.
SPHARM-Mamba: Rotation-Invariant Multiscale Modeling for Brain Age Prediction
Choi, Junho (POSTECH), Lyu, Ilwoo (POSTECH)
ClassificationComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerAuto EncoderContrastive LearningMeshBiomedical DataMagnetic Resonance Imaging
🎯 What it does: This paper proposes a Mamba network based on spherical harmonics for rotation-invariant brain age prediction without the need for surface registration.
Spherical Clinical Embedding with Adaptive Multi-expert Fusion for AAA Rupture Classification
Roby, Merjulah (University of Texas at San Antonio), Finol, Ender A. (Northwestern University)
ClassificationRepresentation LearningData-Centric LearningConvolutional Neural NetworkTransformerSupervised Fine-TuningMixture of ExpertsContrastive LearningImageTabularBiomedical DataComputed TomographyAlzheimer's DiseaseElectronic Health Records
🎯 What it does: Proposed an adaptive multi-expert fusion framework (XMoE) for binary classification prediction of abdominal aortic aneurysm (AAA) rupture risk.
SpikeEEGformer: A Spike-Driven Transformer for Generalized EEG Recognition
Luo, Jiacheng, Wang, Ben (Hangzhou Normal University)
ClassificationRecognitionSpiking Neural NetworkTransformerSupervised Fine-TuningContrastive LearningTime SeriesBiomedical Data
🎯 What it does: Proposes SpikeEEGformer, a directly trained spike-driven Transformer for classifying EEG signals.
SpikeMamba: Spike-Driven State Space Models for Energy-Efficient Biomedical Sequence Modeling
Lee, Si Yong (Stony Brook University), Yang, Yoon Seok (State University of New York Korea)
Anomaly DetectionComputational EfficiencySpiking Neural NetworkTransformerTime SeriesBiomedical DataElectrocardiogram
🎯 What it does: Convert the Mamba state space model into the spike domain, using LIF neurons and event-driven gating to achieve sparse computation, thereby significantly reducing energy consumption while maintaining accuracy at the level of artificial neural networks (ANNs).
Spinverse: Differentiable Physics for Permeability-Aware Microstructure Reconstruction from Diffusion MRI
Khole, Prathamesh Pradeep (University of California Santa Cruz), Marinescu, Razvan (University of California Santa Cruz)
OptimizationGraph Neural NetworkDiffusion modelScore-based ModelGenerative Adversarial NetworkContrastive LearningMeshBiomedical DataMagnetic Resonance ImagingDiffusion Tensor Imaging
🎯 What it does: Propose the Spinverse method, which optimizes surface permeability on a fixed tetrahedral mesh using a differentiable Bloch-Torrey finite element simulator to inversely reconstruct the microstructural boundaries corresponding to diffusion MRI signals.
SRMA-Mamba: Spatial Reverse Mamba Attention Network for Pathological Liver Segmentation in MRI Volumes
Zeng, Jun, Jha, Debesh (University Of South Dakota)
SegmentationConvolutional Neural NetworkTransformerMixture of ExpertsDiffusion modelContrastive LearningImageBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Proposed and implemented a three-dimensional convolution + Mamba network called SRMA-Mamba for precise segmentation of liver cirrhosis lesions in magnetic resonance imaging (MRI) volumes.
SRPAN: Implicit Neural Slice-to-Volume Reconstruction for Super-Resolution Pancreatic MRI
Olofsson, Fredrik (Leiden University Medical Center), Dijkstra, Jouke (Leiden University Medical Center)
Super ResolutionDiffusion modelScore-based ModelNeural Radiance FieldAuto EncoderBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Propose an implicit neural slice-to-volume reconstruction method called SRPAN for low-redundancy pancreatic MRI data;
SS-IoSR: Self-supervised Intraoral Scans Repair
Farhat, Manel (Digital Research Center of Sfax), Smaoui, Oussama (Biotech Dental Group)
RestorationTransformerDiffusion modelScore-based ModelRectified FlowNeural Radiance FieldAuto EncoderContrastive LearningGaussian SplattingPoint CloudComputed Tomography
🎯 What it does: A self-supervised hybrid framework named SS-IoSR is studied for repairing geometric defects in endoscopic scans.
SSPT: Spiking Serialized Point Transformer for Couinaud Segmentation in 3D Medical Data
Li, Yan (Chinese University of Hong Kong), Heng, Pheng-Ann (Chinese University of Hong Kong)
SegmentationComputational EfficiencySpiking Neural NetworkTransformerAuto EncoderContrastive LearningPoint CloudBiomedical DataComputed Tomography
🎯 What it does: Proposed a Spiking Serialized Point Transformer (SSPT) based on spiking neural networks for Couinaud segmentation of 3D medical point clouds.
Stabilizing Temporal Inference Dynamics for Online Surgical Phase Recognition
Liu, Yang (King's College London), Ourselin, Sébastien (University Of Electronic Science And Technology Of China)
RecognitionRecurrent Neural NetworkTransformerSupervised Fine-TuningContrastive LearningVideoBiomedical DataBenchmark
🎯 What it does: This paper proposes a unified Train-Inference-Evaluation framework specifically designed to enhance the temporal stability of models for surgical phase recognition (SPR) during the online surgery phase, avoiding the cascading errors caused by short-term jitter and misjudgment;
Stage-Aware Adaptive Reaction–Diffusion Framework for Modeling Alzheimer’s Disease Progression with Tau and Amyloid-β PET Scans
Qin, Hongyi (University of Liverpool), El-bouri, Wahbi K. (University of Liverpool)
Recurrent Neural NetworkGraph Neural NetworkDiffusion modelAuto EncoderImageBiomedical DataPositron Emission TomographyAlzheimer's DiseaseStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Developed a stage-aware adaptive reaction-diffusion neural ODE framework for predicting regional progression of tau and amyloidβ PET in Alzheimer's disease.
STAMP: Subtype-Aware Spatio-Temporal Atlases for Personalized Prediction of Longitudinal Alzheimer’s Disease Progression
Kalkhof, John (Inria Center at University Côte d'Azur), Lorenzi, Marco (Inria Center at University Côte d'Azur)
Recurrent Neural NetworkGraph Neural NetworkTransformerContrastive LearningImageBiomedical DataMagnetic Resonance ImagingAlzheimer's DiseaseStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: This paper proposes the STAMP model, which integrates disease subtyping with continuous-time spatiotemporal graph learning, achieving personalized prediction of Alzheimer's disease progression based on 3D T1 MRI.
Steer the Sampling, Not the Kernel Grid: Geometry-Guided Sampling Operator for Volumetric Segmentation
Wang, Sizhe (Monash University), Chen, Zhaolin (Monash University)
SegmentationConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound
🎯 What it does: Propose a geometry-guided sampling operator, GeoSample, for 3D medical image segmentation, replacing the traditional stride-1 refinement and downsampling processes, by utilizing local geometric information to guide feature sampling and aggregation.
StrokeTimer: Robust Representation Learning for Ischemic Stroke Onset-Time Estimation from Non-Contrast CT
Wang, Weiru (Eindhoven University of Technology), Su, Ruisheng (Maastricht University)
ClassificationRepresentation LearningConvolutional Neural NetworkAuto EncoderContrastive LearningImageBiomedical DataComputed Tomography
🎯 What it does: This study proposes an end-to-end framework called StrokeTimer, which can automatically estimate the onset time window of ischemic stroke in non-enhanced CT images (<4.5 h, 4.5–6 h, >6 h).
Structural Congruence Matters: Biological Pretraining for Vascular Graph Extraction
Scavone, Alessandro (EURECOM), Zuluaga, Maria A. (EURECOM)
Image TranslationSegmentationDomain AdaptationGraph Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningImageGraphBiomedical DataAgriculture Related
🎯 What it does: A 'nature-to-nature' pre-training paradigm is proposed, using the structural pre-training of plant (grapevine) skeleton graphs for end-to-end extraction of medical images to vascular graph structures.