arXivSub Start free trial

MICCAI 2026 Papers — Page 11

International Conference on Medical Image Computing and Computer-Assisted Intervention · 1165 papers

Structural-Semantic Aware Information Reduction for Asymmetric EEG-Visual Alignment

Chen, Hongan (Nanjing University), Fang, Yuqi (Nanjing University)

RetrievalDomain AdaptationRepresentation LearningConvolutional Neural NetworkGraph Neural NetworkTransformerPrompt EngineeringVision Language ModelAuto EncoderContrastive LearningImageMultimodalityBiomedical Data

🎯 What it does: The paper proposes a structure-semantic aware information reduction framework (SAIR), achieving zero-shot EEG-visual alignment through neuroscience-inspired image dimensionality reduction and parallel EEG encoding.

Structure-Adaptive Sparse Diffusion in Voxel Space for 3D Medical Image Enhancement

Jiang, Hongxu (University of Florida), Shao, Wei (University of Florida)

RestorationSuper ResolutionConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelContrastive LearningBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission Tomography

🎯 What it does: A sparse voxel space diffusion framework is proposed for denoising and super-resolution of 3D medical images. By directly predicting clean images in the voxel space, incorporating velocity supervision, sparse time step scheduling, and structure-aware trajectory modulation, the method achieves results with only 5 time steps, a 10-fold increase in training speed, while maintaining fine details and structural integrity.

Structure-Aware Sensorless 3D Ultrasound and Photoacoustic Reconstruction

Huang, Yi (Harbin Institute of Technology), Gao, Fei (Harbin Institute of Technology)

ClassificationConvolutional Neural NetworkTransformerContrastive LearningImage

🎯 What it does: The paper proposes a new deep learning model for image classification tasks.

Structure-Contrast Disentangled INRs for Accelerated Multi-contrast MRI Reconstruction

Vavasour, Zach (University of Toronto), Chiew, Mark (University of Toronto)

RestorationNeural Radiance FieldAuto EncoderContrastive LearningBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Proposes a structure-contrast disentangled implicit neural representation, DISINR, for multi-contrast MRI accelerated reconstruction;

Sub-Resolution Monocyte Detection in Time-Lapse MRI: A Benchmark

Rexeisen, Robin (University of Münster), Jiang, Xiaoyi (University of Münster)

Object DetectionDomain AdaptationConvolutional Neural NetworkSupervised Fine-TuningContrastive LearningTime SeriesBiomedical DataMagnetic Resonance ImagingBenchmark

🎯 What it does: This paper creates the first publicly available time-series MRI dataset for detecting single cells (monocytes) with sizes ranging from 2 to 6 pixels, and provides a benchmark evaluation scheme.

Subject- and Task-Aware EEG Foundation Model

An, Sion (Pohang University of Science and Technology), Park, Sang Hyun (Pohang University of Science and Technology)

ClassificationRepresentation LearningMeta LearningTransformerSupervised Fine-TuningContrastive LearningTime SeriesBiomedical DataElectrocardiogram

🎯 What it does: Propose the STEM model, combining self-supervised contrastive learning with metric-based meta-learning, explicitly decoupling EEG features into subject-related and task-related components, thereby fully utilizing subject identity and task label information during the pre-training phase.

Subject-Specific Low-Field MRI Synthesis via a Neural Operator

Gao, Ziqi (Yale University), Constable, R. Todd (Yale University)

Image TranslationData SynthesisConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Designed and trained a subject-specific synthesis framework called H2LO for high-field to low-field MRI, enabling the generation of low-field images from high-field scans, supporting low-field image enhancement and device evaluation;

Sulcus-Aware Hippocampal Surface Modeling with Training-Time Sulcus-Guided Learning

Kim, Yuyeon (Pohang University of Science and Technology), Lyu, Ilwoo (Pohang University of Science and Technology)

RestorationSegmentationConvolutional Neural NetworkDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImagePoint CloudBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease

🎯 What it does: This study proposes a hippocampal surface reconstruction framework based on template deformation, which can accurately preserve the fine hippocampal sulci while maintaining topological consistency and point correspondence.

SuRe-Flow: Surface-to-CT Synthesis via Retrieval-Augmented Rectified Flow for CT-Free Superficial Radiotherapy

Liu, Cong (Shanghai Business School), Xie, Kai (Nanjing Medical University)

Image TranslationData SynthesisConvolutional Neural NetworkTransformerDiffusion modelRectified FlowContrastive LearningImageBiomedical DataComputed TomographyRetrieval-Augmented Generation

🎯 What it does: Proposed a retrieval-enhanced Rectified Flow model, SuRe-Flow, which can directly generate CT volumes from optical surface scans, enabling a CT-free surface radiotherapy workflow.

SurfMark3D: Surface-Based Graph Convolution Framework for Distal Femoral Landmark Localisation in Total Knee Arthroplasty

B., Dharshan (Indian Institute of Technology Madras), Sivaprakasam, Mohanasankar (Indian Institute of Technology Madras)

Pose EstimationGraph Neural NetworkDiffusion modelAuto EncoderContrastive LearningGaussian SplattingPoint CloudMeshBiomedical DataMagnetic Resonance ImagingComputed TomographyBenchmark

🎯 What it does: Proposed a graph convolutional framework based on surface point clouds, named SurfMark3D, for anatomical landmark localization of the distal femur during total knee arthroplasty.

Surgical Anatomy Recognition with Context Learning using Foundation Representations

de Jong, Ronald L. P. D. (Eindhoven University of Technology), van der Sommen, Fons (University Medical Center Utrecht)

RecognitionSegmentationConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningVideoBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound

🎯 What it does: This study first constructed the ATLAS-120k medical surgical video slice semantic segmentation dataset, and proposed a temporal semantic segmentation model called ATLAS, which combines base model embeddings with procedure and phase context queries, for anatomical structure identification in minimally invasive surgery.

Surgical Video Temporal Grounding

Liu, Qixuan (Chinese University of Hong Kong), Heng, Pheng-Ann (Chinese University of Hong Kong)

RecognitionSegmentationTransformerSupervised Fine-TuningPrompt EngineeringVision Language ModelVideoMultimodalityBenchmark

🎯 What it does: Proposed the SurgVTG task for temporal localization in surgical videos, constructed the HMSSurgVTG benchmark, and provided corresponding annotations.

SurgLQA: Scalable Long-Horizon Surgical Video Question Answering

Guo, Diandian (Chinese University of Hong Kong), Heng, Pheng-Ann (Chinese University of Hong Kong)

RetrievalCompressionRepresentation LearningConvolutional Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVideoTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose SurgLQA, a scalable video question answering framework designed for long-duration surgical videos, capable of achieving precise temporal modeling and sparse event localization throughout the entire surgical process.

SurgMark: Structure-Aware Watermarking for 4D Gaussian Splatting in Surgical Scenes

Jeong, Minjae (Pohang University of Science and Technology), Kim, Won Hwa (Pohang University of Science and Technology)

SegmentationSafty and PrivacyTransformerAuto EncoderContrastive LearningGaussian SplattingVideoBiomedical DataComputed TomographyUltrasound

🎯 What it does: Proposed a structure-aware watermarking framework called SurgMark for 4D Gaussian Splatting models in surgical scenarios, which can embed invisible watermarks without affecting rendering quality or segmentation performance.

SurgOnAir: Hierarchy-Aware Real-Time Surgical Video Commentary

He, Jingyi (TU Munich), Bi, Yuan (TU Munich)

GenerationComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringVision Language ModelVideoTextRetrieval-Augmented GenerationAudio

🎯 What it does: Propose the SurgOnAir model to achieve real-time hierarchical narration of surgical videos.

SurgRFO: Foundation Model Based Compositional Synthesis of Critical Retained Foreign Objects in Intraoperative Chest X-Rays

Hu, Yuanyun (Tsinghua University), Bai, Harrison (Johns Hopkins University)

Object DetectionGenerationData SynthesisConvolutional Neural NetworkTransformerSupervised Fine-TuningPrompt EngineeringDiffusion modelImageBiomedical DataComputed TomographyBenchmark

🎯 What it does: Designed and implemented a two-stage synthesis framework called SurgRFO, which first generates surgery chest X-ray backgrounds without RFO using a latent diffusion model, and then samples local RFOs with a lightweight generator and generates realistic composite images through conditional Poisson fusion, used for data augmentation to improve detection performance.

SurgVLA-Bench: Towards Evaluating Vision-Language-Action Models for Laparoscopic Surgical Robotics

Sun, Jiashuo (National Engineering Research Center of Robot Visual Perception and Control Technology), Liu, Min (National Engineering Research Center of Robot Visual Perception and Control Technology)

Robotic IntelligenceTransformerSupervised Fine-TuningReinforcement LearningVision-Language-Action ModelImageVideoTextSequentialBiomedical DataBenchmark

🎯 What it does: Propose SurgVLA-Bench, constructing a surgical vision-language-action (VLA) evaluation benchmark based on the SurRoL simulation platform, which includes a hierarchical task system and corresponding datasets ranging from atomic actions to complete surgical procedures.

SurgZSD: Zero-Shot Surgical Action Triplet Detection via Attribute Composition and Clinical Feasibility Inference

Deng, Zuxing (Hefei University of Technology), Zhou, Wenrui (Hefei University of Technology)

Object DetectionTransformerSupervised Fine-TuningPrompt EngineeringVision Language ModelImageBiomedical DataBenchmarkRetrieval-Augmented Generation

🎯 What it does: Developed the SurgZSD zero-shot surgical action triple detection framework, which utilizes medical prior knowledge to achieve joint reasoning of attribute combination and clinical feasibility constraints.

SurvTTA: Order-Aware Test-Time Adaptation for Heterogeneous Domain Shifts in Multi-modal Survival Analysis

Zhao, Rongchang (Central South University), Li, Shuo (Central South University)

Domain AdaptationGraph Neural NetworkTransformerSupervised Fine-TuningContrastive LearningMultimodalityBiomedical DataComputed TomographyPositron Emission Tomography

🎯 What it does: Proposes SurvTTA, a multi-modal survival analysis framework based on test-time adaptation, aimed at addressing cross-center heterogeneous domain shift issues.

SutureFormer: Learning Surgical Trajectories via Goal-Conditioned Offline RL in Pixel Space

Liu, Huanrong (University of Macau), Li, Qingbiao (University of Macau)

Robotic IntelligenceConvolutional Neural NetworkTransformerReinforcement LearningDiffusion modelVideoBiomedical Data

🎯 What it does: Learn a surgical suturing trajectory prediction model called SutureFormer in the pixel space through offline reinforcement learning, enabling the prediction of future needle tip trajectories based solely on endoscopic videos and sparse keyframes.

SwallowReg: SAM-Assisted Deformable Registration with Adaptive Global-Local Features for Cine-MRI Swallowing Function Quantification

Tang, Zhiwen (Nanjing University of Posts and Telecommunications), Yu, Han (Nanjing Medical University)

Image TranslationRestorationSegmentationAnomaly DetectionConvolutional Neural NetworkTransformerDiffusion modelContrastive LearningOptical FlowImageVideoBiomedical DataMagnetic Resonance Imaging

🎯 What it does: This study proposes an end-to-end SwallowReg framework for joint segmentation and registration in dynamic Cine-MRI of patients with tongue cancer after head and neck resection, thereby enabling quantitative evaluation of swallowing function.

SWoMo: Neuro-Symbolic World Model for Cataract Surgery Simulation

Sivakumar, Ssharvien Kumar (Technical University Darmstadt), Mukhopadhyay, Anirban (Carl Zeiss AG)

GenerationData SynthesisRobotic IntelligenceConvolutional Neural NetworkRecurrent Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningWorld ModelOptical FlowVideoBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound

🎯 What it does: Propose a neuro-symbolic world model (SWoMo) for cataract surgery simulation, generating motion and interaction through symbolic scene graphs and rule-based physical simulation, and then using diffusion models to generate high-quality visual representations.

SymmDiff: Symmetry-Guided Diffusion for Fracture Localization in Pelvic Radiographs

Rahman, Abdul (Korea Institute of Energy Technology), Lee, Bumshik (Chosun University)

Anomaly DetectionConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningImageBiomedical DataComputed Tomography

🎯 What it does: Propose SymmDiff, an unsupervised anomaly detection framework that incorporates contralateral symmetric information into latent diffusion models for pelvic X-ray fracture localization.

Synergistic Information Disentanglement for Omni-Modal Slide Representation Learning in Computational Pathology

Liu, Mingxin (Nanjing University of Information Science and Technology), Xu, Jun (Nanjing University of Information Science and Technology)

Representation LearningTransformerContrastive LearningMultimodalityBiomedical Data

🎯 What it does: Propose a self-supervised framework called Phi-Omni that leverages synergistic information disentanglement to learn multi-modal representations from whole slides.

Synergistic Perception-Reasoning Governance: Grounding Medical MLLMs with Verifiable Anatomical Evidence

Hao, Rui, Zeng, Zhigang (Huazhong University Of Science And Technology)

Explainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextBiomedical DataMagnetic Resonance ImagingComputed TomographyElectronic Health RecordsRetrieval-Augmented Generation

🎯 What it does: A training-agnostic, anatomy ROI evidence injection-based multimodal large model hallucination suppression framework is proposed, combining visual activation re-adjustment with text structural injection, and introducing a task-aware dynamic router.

SynHydro: Biomechanics-Driven Domain Randomization for Hydrocephalus-Agnostic, Generalizable Tissue-Ventricle Segmentation Across Age, Modality, and Resolution

Ren, Zehua (Xi'an Jiaotong University), Lian, Chunfeng (Xi'an Jiaotong University)

SegmentationData SynthesisDomain AdaptationConvolutional Neural NetworkDiffusion modelGenerative Adversarial NetworkContrastive LearningImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Simulate hydrocephalus using biomechanically driven deformations based on healthy brain tissue labels, then generate diverse multimodal, low-resolution, noisy, and artifact-containing MRI data through domain randomization, and train nnU-Net to achieve joint segmentation of brain tissue and ventricles in pediatric hydrocephalus patients.

SynMamba: Synergizing Visual Mamba with Synthesized Clinical Semantics for Boundary-Aware Segmentation

Wei, Xuan (Xiamen University), Hong, Qingqi (Xiamen University)

SegmentationData SynthesisConvolutional Neural NetworkTransformerVision Language ModelContrastive LearningImageTextBiomedical DataComputed Tomography

🎯 What it does: Propose SynMamba dual-branch network, integrating the global modeling of Visual Mamba and the local edge detection of CNN to achieve precise medical image segmentation.

SYNPRED: A Synergistic Approach to Multimodal Learning for Clinical Prediction

Janíčková, Ivana (Medical University of Vienna), Langs, Georg (Medical University of Vienna)

ClassificationExplainability and InterpretabilityRepresentation LearningData-Centric LearningTransformerSupervised Fine-TuningMixture of ExpertsAuto EncoderGenerative Adversarial NetworkContrastive LearningImageMultimodalityTabularBiomedical DataMagnetic Resonance ImagingElectronic Health Records

🎯 What it does: Propose a multi-modal variational autoencoder (VAE) framework called SYNPRED, which integrates medical imaging and RNA sequencing data, to enable clinical prediction when multi-modal data are missing.

TAIR: Text-Guided Adaptive Prompt Refinement for Coarse-to-Fine all-in-One Medical Image Restoration

Cui, Jiaqi, Wang, Yan (Sichuan University)

RestorationSuper ResolutionTransformerPrompt EngineeringVision Language ModelContrastive LearningImageMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission Tomography

🎯 What it does: Propose TAIR, a text-guided adaptive prompt refinement framework, achieving full-scenario medical image restoration in coarse-to-fine stages.

Task-Performance Routing for Multi-teacher Distillation in CT Medical Image Segmentation

Yang, Jiaye, Wang, Peng (University Of Electronic Science And Technology Of China)

SegmentationKnowledge DistillationTransformerMixture of ExpertsBiomedical DataComputed Tomography

🎯 What it does: CT organ segmentation based on multi-teacher knowledge distillation, proposing the Task-Performance Routing (TPR) framework, which dynamically selects the best teacher and fuses knowledge for each region in space.

TAVR-VLM: Risk-Conditioned Causal Grounding for Hallucination-Resistant Report Generation

Lu, Zhixiang (Xi'an Jiaotong-Liverpool University), Wang, Jinfeng (Xi'an Jiaotong-Liverpool University)

GenerationAnomaly DetectionExplainability and InterpretabilityConvolutional Neural NetworkTransformerLarge Language ModelVision Language ModelContrastive LearningImageTextMultimodalityTabularBiomedical DataComputed TomographyUltrasoundElectronic Health RecordsRetrieval-Augmented Generation

🎯 What it does: This paper proposes a framework named TAVR‑VLM for generating preoperative reports for transcatheter aortic valve replacement (TAVR), and suppresses diagnostic hallucinations through risk-conditioned causal grounding.

TC-SSA: Token Compression via Semantic Slot Aggregation for Gigapixel Pathology Reasoning

Chen, Zhuo (Shenzhen University of Advanced Technology), Xu, Lijian (Shenzhen University of Advanced Technology)

CompressionComputational EfficiencyRepresentation LearningTransformerVision Language ModelContrastive LearningImageBiomedical DataRetrieval-Augmented Generation

🎯 What it does: Propose the TC-SSA module, which performs semantic slot aggregation and compression on WSI visual tokens while preserving diagnostic key information.

TCCT: Trajectory-Conditioned CBCT Reconstruction for Sinusoidal Non-circular Orbits with a Fourier Neural Operator

Ye, Chengze (Friedrich-Alexander-Universität Erlangen-Nürnberg), Maier, Andreas (Friedrich-Alexander-Universität Erlangen-Nürnberg)

RestorationConvolutional Neural NetworkDiffusion modelAuto EncoderContrastive LearningBiomedical DataComputed Tomography

🎯 What it does: Proposed the TCCT framework, which utilizes the orbit-conditioned spectral neural operator to predict redundant weights and embeds them into a differentiable Shift-Variant FBP, achieving fast CBCT reconstruction for sinusoidal non-circular orbits.

TCFD-Net: Temporal-Calibrated Feature Disentanglement Network for Dual-Timepoint CT-Based Immunotherapy Response Prediction in Advanced Non-small Cell Lung Cancer

Fu, Yao (Beihang University), Mu, Wei (Beihang University)

ClassificationTransformerContrastive LearningImageBiomedical DataComputed Tomography

🎯 What it does: Developed TCFD-Net based on dual-time-point CT images to predict early immunotherapy response in patients with non-small cell lung cancer (NSCLC).

TEDi: Temporal Memory-Enhanced and Denoising Transformer for Surgical Instrument Segmentation

Yuan, Jiahong (Tsinghua University), Zhou, Haoyin (Harvard Medical School)

SegmentationTransformerContrastive LearningImageVideoBenchmark

🎯 What it does: Proposed the TEDi framework, combining memory search enhancement and temporal consistency denoising Transformer to improve surgical instrument segmentation, addressing cross-frame semantic consistency and class confusion issues.

TeDyS: Temporal Dynamics with Spatial Context for Longitudinal Progression Prediction

Hwang, Jeonghyun (Ewha Womans University), Choi, Jang-Hwan (Ewha Womans University)

ClassificationAnomaly DetectionTransformerImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose the TeDyS framework, which utilizes time-dynamic conditional queries for progress prediction in long-term medical imaging.

Temporal Phase-Difference Guided Spatiotemporal Learning for DSA Vessel Segmentation

Liu, Kun (Beijing University of Posts and Telecommunications), Yang, Huihua (Beijing University of Posts and Telecommunications)

SegmentationConvolutional Neural NetworkRecurrent Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningVideoBiomedical Data

🎯 What it does: Propose a spatiotemporal segmentation framework combining forward phase difference projection maximum intensity projection (FPD‑MaxIP) and a Mamba-based temporal encoder for precise segmentation of vessels in digital subtraction angiography (DSA).

Test-Time Adaptation for ECG Classification via SQI-Gated Self-training and Beat-Rhythm Consistency

Jiang, Wenhan (Westlake University), Zheng, Yefeng (Westlake University)

ClassificationDomain AdaptationConvolutional Neural NetworkSupervised Fine-TuningContrastive LearningTime SeriesBiomedical DataElectrocardiogram

🎯 What it does: Achieve test-time adaptation in electrocardiogram classification under unseen domains.

Test-Time Adaptation for Rare Surgical Phase Recognition: Bridging the Coverage-Gap Paradox

Park, Ho-min (Ghent University Global Campus), Vankerschaver, Joris (Ghent University Hospital)

RecognitionDomain AdaptationRecurrent Neural NetworkTransformerSupervised Fine-TuningContrastive LearningVideoBiomedical Data

🎯 What it does: This paper addresses the coverage gap issue in the identification of rare surgical phases by proposing a lightweight test-time adaptation framework, which significantly improves the accuracy of rare phase identification through adaptive threshold pseudo-label updating and temporal smoothing.

Test-Time Adaptation in Optical Coherence Tomography Using Trajectory-Aligned Time-Independent Flow

Hucke, Veit (Medical University of Vienna), Bogunović, Hrvoje (Medical University of Vienna)

RestorationSegmentationDomain AdaptationConvolutional Neural NetworkSupervised Fine-TuningDiffusion modelFlow-based ModelImageBiomedical DataOrdinary Differential Equation

🎯 What it does: Proposes a test-time adaptation framework called TTA-Flow based on flow matching and histogram alignment, which transforms noisy images from low-cost OCT devices into high-quality training domain images.

Test-time Adaptation of Pelvic Bone Segmentation Models via Dynamic Reliability-Guided

Ren, Ling, Zheng, Kai (Nanjing University of Posts and Telecommunications)

SegmentationDomain AdaptationConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningImageBiomedical DataComputed Tomography

🎯 What it does: An online adaptive framework for pelvic bone segmentation in medical CT imaging.

Testable Concept Bottlenecks for Pretreatment pCR Prediction in Breast DCE-MRI: Faithfulness and Reliability

Tan, Jiyang (Harbin Institute of Technology), Zhang, Tianyu (Netherlands Cancer Institute)

ClassificationExplainability and InterpretabilityConvolutional Neural NetworkTransformerPrompt EngineeringVision Language ModelContrastive LearningImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: This paper proposes a projection-based concept bottleneck model to predict pathological complete response during the preprocessing stage of breast DCE-MRI, and provides visualizable and intervenable BI-RADS concept evidence.

Text as a Compass: Semantic-Navigated 3D Bone Collapse Prediction via Orthogonal Micro-volume Primitives

Zeng, Qingyuan (Hong Kong University of Science and Technology), Chen, Jintai (Hong Kong University of Science and Technology)

ClassificationAnomaly DetectionExplainability and InterpretabilityConvolutional Neural NetworkTransformerVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Proposes a text-guided 3D fracture prediction framework called COMPASS, which utilizes orthogonal micro-volume primitives (OMVP) and causal counterfactual validation to capture fine-grained lesion features, thereby improving the accuracy of predicting early collapse in avascular necrosis of the femoral head.

Text as Illumination: Spatial Contrastive Retinex Learning for Language-Guided Medical Image Segmentation

Shi, Jian (Dalian University of Technology), Lu, Huchuan (Dalian University of Technology)

SegmentationTransformerVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose a Retinex-inspired network called TIRNet that utilizes text embeddings as semantic illumination to achieve language-guided medical image segmentation.

Text as the Sieve: Information Bottleneck Guided Prototype Sieving for Text-Guided Medical Image Segmentation

Long, Jiake (Sichuan University), Wang, Yan (Sichuan University)

SegmentationExplainability and InterpretabilityTransformerVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose a framework called TeSi based on the information bottleneck for early visual prototype screening, using text guidance to segment medical images.

Text Guidance and Distance Gating Improve Detection and Segmentation of Focal Cortical Dysplasia

Mikhelson, German (University of Bonn), Schultz, Thomas

SegmentationAnomaly DetectionGraph Neural NetworkTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance Imaging

🎯 What it does: A framework for FCD detection and segmentation based on graph neural networks was constructed, fusing text prompts and spatial distance gating on the cortical surface graph.

Text-Deficient Multimodal Stroke Segmentation with Lesion-Grounded Self-Retrieval-Augmented Generation

Eum, Heeseong (Seoul National University), Choi, Kyu Sung (Seoul National University)

SegmentationGenerationRetrievalConvolutional Neural NetworkTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextBiomedical DataMagnetic Resonance ImagingRetrieval-Augmented Generation

🎯 What it does: The study proposes LeG-RAG, a lesion-aligned retrieval-based generation method for missing report generation in the text-missing scenario of acute ischemic stroke MRI, and implements report-conditioned multimodal segmentation using LLMSwin.

Text-Guided Multi-frequency Latent Diffusion for Medical Image Segmentation

Gao, Qiang (Monash University), Chen, Cunjian (Monash University)

SegmentationTransformerPrompt EngineeringVision Language ModelDiffusion modelScore-based ModelAuto EncoderContrastive LearningImageTextBiomedical Data

🎯 What it does: Proposed a text-guided multi-frequency latent diffusion framework (TMF-Seg) for medical image segmentation.

TG-OT: Topology-Guided CCTA-IVUS Registration via Optimal Transport Matching

van Herten, Rudolf L. M. (Weill Cornell Medicine), Išgum, Ivana (Amsterdam UMC, University of Amsterdam)

OptimizationConvolutional Neural NetworkDiffusion modelContrastive LearningOptical FlowBiomedical DataComputed TomographyUltrasound

🎯 What it does: Developed an全自动 coronary CT angiography (CCTA) and intravascular ultrasound (IVUS) registration framework called TG-OT, which can complete the fusion of two modalities without prior segmentation or manual intervention.

TGBD-Net: Tumor-Guided Bridging Distillation for Joint EGFR Mutation and Survival Prediction in Lung Cancer from CT Imaging

Yang, Huihui (Beihang University), Tian, Jie (Beihang University)

ClassificationKnowledge DistillationRepresentation LearningTransformerContrastive LearningImageBiomedical DataComputed Tomography

🎯 What it does: Propose TGBD-Net for jointly predicting EGFR mutation status and progression-free survival (PFS) in lung cancer CT images, achieving stable multi-task learning in a heterogeneous supervision environment.

TGH-DB: Template-Guided Heteroscedastic Diffusion Bridge for Brain MRI-to-PET Synthesis

Pang, Haowen (Beijing Institute of Technology), Qiu, Anqi (Hong Kong Polytechnic University)

GenerationData SynthesisTransformerDiffusion modelMultimodalityBiomedical DataMagnetic Resonance ImagingPositron Emission TomographyAlzheimer's Disease

🎯 What it does: Propose a Template-Guided Heteroscedastic Diffusion Bridge (TGH-DB) based on population-level amyloid PET templates, which can generate structured and physiologically reasonable PET images from multi-modal MRI.

The Learnability Gap in Medical Latent Diffusion

Dombrowski, Mischa (FAU Erlangen Nurnberg), Kainz, Bernhard (FAU Erlangen Nurnberg)

ClassificationGenerationKnowledge DistillationRepresentation LearningConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyElectronic Health Records

🎯 What it does: This paper systematically investigates the 'learnability gap' in potential diffusion models for medical image generation, and proposes a noise-conditioned latent classifier (FiLM layer + image space distillation) to diagnose and partially alleviate this gap.

The Neuro-Developing Engine: a Surrogate Longitudinal Model for Infant Brain Dynamics Via Causal Distillation

Hu, Dan (University of North Carolina at Chapel Hill), Li, Gang (University of North Carolina at Chapel Hill)

Explainability and InterpretabilityKnowledge DistillationRepresentation LearningConvolutional Neural NetworkRecurrent Neural NetworkGraph Neural NetworkDiffusion modelScore-based ModelGenerative Adversarial NetworkContrastive LearningTime SeriesBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose the Neuro-Developing Engine (NDE), a personalized neural simulation model based on rs-fMRI, capable of continuously simulating infant brain functional dynamics and performing causal interventions.

The Paper Has a GitHub, the GitHub Has a README, the README Has Nothing: Reproducibility Signals for Review Support

Bolelli, Federico (University of Modena and Reggio Emilia), Grana, Costantino (University of Modena and Reggio Emilia)

Explainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelPrompt EngineeringImageTextTabularReview/Survey PaperBenchmarkRetrieval-Augmented Generation

🎯 What it does: This paper proposes an interpretable decision support tool called paper-snitch based on LLM, which is used in the peer review process of medical imaging papers to automatically parse PDFs, extract and statically check code repositories, and generate evidence-supported reproducibility assessment reports based on the MICCAI reproducibility policy.

Think Global, Look Focal: Dual-Stream Retrieval-Augmented Pathology Report Generation for Whole Slide Images

Xiong, Liang (Chongqing Normal University), Cui, ShaoGuo (Chongqing Normal University)

GenerationRetrievalTransformerPrompt EngineeringVision Language ModelImageTextBiomedical DataRetrieval-Augmented Generation

🎯 What it does: This paper proposes a dual-stream retrieval-enhanced pathological report generation framework, DR-Gen, for automatically generating pathological reports.

Through the Schrödinger Bridge: Benchmarking Antemortem Image Restoration from Postmortem Autolysis to Enhance Forensic Diagnostics

Hao, Shuang (Xi'an Jiaotong University), Lian, Chunfeng (Xi'an Jiaotong University)

Image TranslationRestorationTransformerDiffusion modelScore-based ModelGenerative Adversarial NetworkContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyBenchmarkStochastic Differential Equation

🎯 What it does: Propose the task of postmortem self-autopsy image restoration in forensic histopathology, construct the first homologous but unpaired AutoPath dataset, and use the Schrödinger Bridge generative model to achieve image restoration, followed by the design of an evaluation framework based on diagnostic distribution consistency.

Through-Plane Attention for Anisotropic CT Super-Resolution Improves Quantitative Lung Cancer Analysis

Boubnovski Martell, Marc (Imperial College London Hammersmith Campus), Aboagye, Eric O. (Imperial College London Hammersmith Campus)

RestorationSegmentationSuper ResolutionConvolutional Neural NetworkTransformerContrastive LearningImageBiomedical DataComputed Tomography

🎯 What it does: Propose TVSRN-V2, which utilizes a voxel-level super-resolution method combining Swin Transformer V2 with planar attention, to reconstruct thick-slice CT scans, thereby improving the structural continuity for quantitative analysis of lung cancer.

Time Matters: Rethinking Diffusion and Flow Models in One-Step Medical Image Translation

Mei, Siyuan (Friedrich-Alexander-Universität Erlangen-Nürnberg), Maier, Andreas (Siemens Healthineers)

Image TranslationGenerationConvolutional Neural NetworkTransformerDiffusion modelFlow-based ModelImageBiomedical DataMagnetic Resonance ImagingPositron Emission TomographyAlzheimer's DiseaseOrdinary Differential Equation

🎯 What it does: Proposes a first-order deterministic generative framework JiR that only utilizes the temporal condition t for medical image translation (MR→PET, T1w→T1ce).

Time-Grounded Clinical Prediction from Longitudinal Multimodal Patient Histories

Yang, Wanqi (Mohamed bin Zayed University of Artificial Intelligence), Xie, Yutong (University of Technology Sydney)

ClassificationRecommendation SystemExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelImageTextMultimodalityTime SeriesBiomedical DataElectronic Health RecordsElectrocardiogramBenchmark

🎯 What it does: Proposed the MIMIC-MTE dataset and designed a multi-timepoint clinical prediction task based on it;

TinyH2GNet: Exploring Lightweight Models for Gene Expression Prediction from H&E Images

Kaur, Maninder (Indian Institute of Technology Ropar), Gupta, Sukrit (Hamad Bin Khalifa University)

Computational EfficiencyKnowledge DistillationDrug DiscoveryConvolutional Neural NetworkTransformerContrastive LearningImageBiomedical Data

🎯 What it does: Use a lightweight convolutional network to predict spatial gene expression corresponding to H&E tissue section images

TissueCodePilot: A Code-Action Agent for AI-Assisted Spatial Tissue Analysis

Vo, Hung Q. (University of Houston), Nguyen, Hien V. (University of Houston)

Drug DiscoveryAI Code AssistantTransformerLarge Language ModelAgentic AIPrompt EngineeringVision-Language-Action ModelImageTextBiomedical Data

🎯 What it does: Built an environment-interactive code acting agent called TissueCodePilot, which can automatically generate, execute, and debug Python code from minimal natural language queries to complete spatial tissue organization analysis tasks.

TOF-GR: Unsupervised PET Reconstruction via Time-of-Flight Gaussian Representation and Kernel-Based Prior Modeling

Zhou, Zilan (Zhejiang University), Zhu, Wentao (Zhejiang University)

RestorationGaussian SplattingBiomedical DataPositron Emission Tomography

🎯 What it does: Propose the TOF-GR framework, which utilizes 3D Gaussian expansion and the TOF system matrix to achieve unsupervised PET reconstruction.

TOFlow: Learning-Based Isotropic Reconstruction and Flow Mapping from Multi-orientation TOF-MRA

Hasin, Daria (Tel Aviv Sourasky Medical Center), Ben Bashat, Dafna (Tel Aviv Sourasky Medical Center)

RestorationSuper ResolutionDiffusion modelFlow-based ModelNeural Radiance FieldAuto EncoderContrastive LearningOptical FlowBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Developed a multi-directional TOF-MRA reconstruction and blood flow mapping framework based on neural fields, named TOFlow, which can generate isotropic high-resolution cerebral vascular images and estimate voxel-level blood flow direction.

Tooth Alignment via Virtual Trajectories with Clinical Constraints

Jo, SeungKwan (Soongsil University), Chung, Minyoung (Osstem Implant)

Pose EstimationOptimizationGraph Neural NetworkTransformerDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningPoint CloudBiomedical DataBenchmark

🎯 What it does: Designed a teeth alignment framework based on virtual progressive trajectories, utilizing pre-treatment and post-treatment oral scan data to progressively predict K-step SE(3) transformations, and applying clinical constraints at each step to simulate the incremental process of real orthodontic treatment.

ToothFairy3: Scaling CBCT Maxillofacial Segmentation to 77 Classes with U-Mamba2

Lumetti, Luca (University of Modena and Reggio Emilia), Bolelli, Federico (University of Modena and Reggio Emilia)

SegmentationConvolutional Neural NetworkTransformerBiomedical DataComputed Tomography

🎯 What it does: Proposed a large-scale CBCT facial bone and tooth segmentation dataset called ToothFairy3 (77 classes, 582 scans) and designed an efficient U-Mamba2 architecture.

Topology-Driven Transferability Estimation of Medical Foundation Models for Segmentation

Tang, Jiaqi (Peking University), Chen, Qingchao (Peking University)

SegmentationDomain AdaptationConvolutional Neural NetworkGraph Neural NetworkTransformerContrastive LearningImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: A topology-based transferability estimation framework is proposed to predict the performance of models in medical image segmentation tasks without fine-tuning.

TopoOR: A Unified Topological Scene Representation for the Operating Room

Wang, Tony Danjun (Technical University of Munich), Bastian, Lennart (Technical University of Munich)

Representation LearningRobotic IntelligenceGraph Neural NetworkTransformerVision-Language-Action ModelContrastive LearningImageTextMultimodalityAudio

🎯 What it does: Propose TopoOR, which uses composite complexes and high-order attention networks to uniformly model and reason about multi-modal entities and multi-agent relationships in the operating room.

TotalFM: An Organ-Separated 3D-CT Foundation Model Leveraging Large-Scale Routine Clinical Radiology Data

Yamamoto, Kohei (Jichi Medical University), Kikuchi, Tomohiro (Jichi Medical University)

SegmentationAnomaly DetectionConvolutional Neural NetworkTransformerLarge Language ModelVision Language ModelContrastive LearningImageTextBiomedical DataComputed Tomography

🎯 What it does: Propose TotalFM, a 3D-CT foundation model that utilizes large-scale clinical CT data, through automated conversion of raw CT images and reports into organ-level image-text pairs for self-supervised training.

Toward Markerless Video-Based Tremor Analysis: Objective Quantification of Pathological Tremor in Mouse Preclinical Models

Koshimoto, Yota (Keio University), Isogawa, Mariko (Keio University)

Pose EstimationAnomaly DetectionConvolutional Neural NetworkTransformerContrastive LearningOptical FlowVideoTime SeriesBiomedical Data

🎯 What it does: Proposes a non-invasive method for quantifying tremors in mice based on lateral RGB video, markerless pose estimation, and peak significance detection

Toward Synergistic Learning for Liver Vessel and Couinaud Segmentations

Qiu, Yue (Chinese University of Hong Kong), Fu, Chi-Wing (Chinese University of Hong Kong)

SegmentationConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelFlow-based ModelRectified FlowNeural Radiance FieldAuto EncoderGenerative Adversarial NetworkContrastive LearningBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Proposes a collaborative learning framework based on vascular skeletons to achieve joint segmentation of liver vessels and Couinaud segments.

Toward Thyroid Cancer BRAF V600E Mutation Prediction via Multimodal Large Language Model: Dataset and Model Development

Hao, Pengfei, Zhu, Lei (Hong Kong University Of Science And Technology (Guangzhou))

ClassificationData-Centric LearningDrug DiscoveryTransformerLarge Language ModelSupervised Fine-TuningVision Language ModelContrastive LearningImageTextMultimodalityBiomedical DataUltrasoundElectronic Health Records

🎯 What it does: This paper constructs the first multi-modal thyroid cancer BRAF V600E mutation prediction dataset, MM-BRAF, and proposes a BRAF-MLLM framework based on a multi-modal large language model to achieve mutation prediction.

Towards Class-Aware Semantic Decoupling in Multi-modal Disease Diagnosis with Missing Modalities

Zhou, Feixiang (University of Liverpool), Zheng, Yalin (University of Liverpool)

ClassificationDomain AdaptationAnomaly DetectionRepresentation LearningTransformerVision Language ModelAuto EncoderContrastive LearningImageTextMultimodalityTime SeriesBiomedical DataElectronic Health Records

🎯 What it does: This paper proposes the Class‑Aware Semantic Decoupling (CSD) framework to address the problem of missing modalities in multi-modal medical diagnosis;

Towards Direction-Equilibrated Segmentation in 3D Medical Images

Zhu, Zhiqin, Cong, Baisen (State University of New York)

SegmentationTransformerBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasound

🎯 What it does: A direction-aware balanced Mamba network (DEM-Mamba) is proposed to achieve 3D medical image segmentation.

Towards the Digital Dental Models: High-Fidelity CBCT-to-IOS Fusion via Anatomy-Aware Geometric Transformers

Zhou, Hanqing (Beijing University of Posts and Telecommunications), Xiao, Li (Beijing University of Posts and Telecommunications)

RestorationConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningPoint CloudMeshBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: This paper proposes an anatomy-aware geometric transformer (AAGT) to achieve automatic high-precision registration between CBCT and IOS.

Towards Unified Image–Video Modeling Across B-Mode and Contrast-Enhanced Ultrasound for Multi-task Analysis

Huang, Qiang (Macao Polytechnic University), Tan, Tao (Macao Polytechnic University)

ClassificationSegmentationTransformerPrompt EngineeringContrastive LearningImageVideoBiomedical DataUltrasound

🎯 What it does: Proposes a unified multi-task framework called UIV-US, which can simultaneously perform image and video segmentation and classification tasks under B-mode and contrast-enhanced ultrasound (CEUS);

Towards Unified Surgical Scene Understanding: Bridging Reasoning and Grounding via MLLMs

Huang, Jincai (Southern University of Science and Technology), Si, Weixin (Nanfang Hospital)

RecognitionSegmentationTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageVideoTextBenchmark

🎯 What it does: Propose the SurgMLLM framework to achieve unified processing of procedural phase recognition, IVT triplet reasoning, and pixel-level entity localization in surgical videos.

Towards Versatile and Safe Multi-electrode Path Planning for SEEG Implantation

Qu, Ziqiao (Southern University of Science and Technology), Si, Weixin (Shenzhen University of Advanced Technology)

Reinforcement LearningImageMeshBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: A reinforcement learning-based multi-electrode SEEG implantation path planning method is studied, which can determine the target position and approach of each electrode at once, and generate safe straight trajectories under strict anatomical and spacing constraints.

TRACE-PSG: Traceable Sleep Diagnosis Support with Pointer-Grounded Evidence

Wu, Sipeng (University of Hong Kong), Meng, Gaofeng (Shandong University)

Explainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextBiomedical DataElectronic Health RecordsRetrieval-Augmented Generation

🎯 What it does: Developed a two-stage traceable sleep diagnosis support system, TRACE-PSG, which generates auditable structured JSON outputs by utilizing multi-page PSG reports and clinical text.

TraceCXR: Verifiable Chest X-ray Reports via Evidence-Addressable Graph Prompting

Dong, Hang (Wuhan University), Du, Bo (Wuhan University)

GenerationExplainability and InterpretabilityGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringContrastive LearningImageTextGraphBiomedical DataComputed TomographyElectronic Health RecordsRetrieval-Augmented Generation

🎯 What it does: Proposed TraceCXR, a traceable framework for generating chest X-ray reports, which can align the generated clinical descriptions with image evidence, knowledge graph nodes, and local visual evidence.

Tracer-Adaptive Expert Learning for Generalizable PET Lesion Segmentation

Shu, Yiran (ShanghaiTech University), Shen, Dinggang (ShanghaiTech University)

SegmentationConvolutional Neural NetworkTransformerMixture of ExpertsAuto EncoderContrastive LearningBiomedical DataComputed TomographyPositron Emission Tomography

🎯 What it does: Propose a prototype-driven Mixture-of-Experts framework called ProMoE for lesion segmentation in multi-tracer PET images.

TracerAD: Training-Free Few-Shot 3D Anomaly Detection for Novel PET Tracers

Huang, Haolin (ShanghaiTech University), Wang, Qian (Fudan University)

Anomaly DetectionExplainability and InterpretabilityRepresentation LearningGraph Neural NetworkTransformerContrastive LearningImageBiomedical DataPositron Emission TomographyAlzheimer's Disease

🎯 What it does: Propose a training-free, few-shot 3D PET anomaly detection framework called TracerAD;

Track2Map: Online Deformable SLAM with Motion-Aware Pose Optimization in Robotic Surgery

Song, Tianyi (University College London), Vasconcelos, Francisco (University College London)

OptimizationRobotic IntelligenceTransformerContrastive LearningGaussian SplattingSimultaneous Localization and MappingOptical FlowVideoBiomedical DataBenchmark

🎯 What it does: Implement online deformable SLAM in robot-assisted surgery videos, jointly optimizing camera pose and 3D Gaussian light scattering model.

Trackerless Ultrasound Pose Estimation via Correspondence-Aware Tokenization and Flow-Matching Transformers

Kim, Sang-yun (KAIST), Bae, Hyeon-Min (KAIST)

Pose EstimationTransformerDiffusion modelContrastive LearningOptical FlowVideoBiomedical DataUltrasoundOrdinary Differential Equation

🎯 What it does: Proposes a free-hand 3D ultrasound pose estimation method without tracking, based on explicit correspondence and flow-matching Transformer

Tractography-Driven Synthetic Data Generation for Fiber Bundle Segmentation in Tracer Histology

Bintsi, Kyriaki-Margarita (Massachusetts General Hospital and Harvard Medical School), Yendiki, Anastasia (Massachusetts General Hospital and Harvard Medical School)

SegmentationData SynthesisConvolutional Neural NetworkDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataMagnetic Resonance ImagingDiffusion Tensor Imaging

🎯 What it does: Using ex vivo dMRI tractography as a generative prior, synthetic 2D image-mask pairs were generated through domain randomization, and combined with a small number of real annotated fiber bundle images to train a 2D U-Net, achieving automated segmentation of fiber bundles in macaque tracing-stained images.

Training LLMs with Reinforcement Learning over Digital Twin Representations for Reasoning-Intensive Surgical VideoQA

Shen, Yiqing (Johns Hopkins University), Unberath, Mathias (Johns Hopkins University)

Reinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringVision Language ModelVideoTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Training large language models through reinforcement learning to perform reasoning on digital twin representations for multi-step reasoning question answering on surgical videos.

Trajectory-Aware Cross-Paradigm Transformers for Efficient System Matrix Calibration in Magnetic Particle Imaging

Li, Jintao (Northwest University), Guo, Hongbo (Concordia University)

Super ResolutionOptimizationComputational EfficiencyConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningBiomedical Data

🎯 What it does: Propose a cross-paradigm transformer named TACPT-Net for efficiently and accurately calibrating and upscaling the system matrix of magnetic particle imaging (MPI).

Tri-KA: Tri-level knowledge anchoring test-time adaptation for source-free cross-site MRI rectal cancer segmentation

Bo, Wang (University of Science and Technology of China), Zhou, Shaohua Kevin (University of Science and Technology of China)

SegmentationDomain AdaptationSafty and PrivacyConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: This paper proposes a three-layer knowledge anchoring (Tri-KA) framework for test-time adaptation (TTA) in cross-site rectal cancer MRI segmentation, without requiring source data or privacy leakage.

TRIAGE-MIL: Multi-axis Instance Selection and Semantic Hypergraph Modeling for Survival Prediction from Whole-Slide Images

Subramanian, Barathi (Stanford University), Shen, Jeanne (Stanford University)

ClassificationImage TranslationAnomaly DetectionComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasoundElectronic Health Records

🎯 What it does: This paper proposes the TRIAGE-MIL framework, which combines unsupervised multi-axis hierarchical sampling (MASS) and a semantic hierarchical hypergraph model to achieve multi-instance learning (MIL) on whole slide images (WSI) for survival prediction.

TriFlow: Triplane Latent Conditional Flow Matching for Efficient 3D Skull Shape Completion

Liu, Zhenhong (Beijing Normal University), Wang, Xingce (Beijing Normal University)

RestorationGenerationTransformerDiffusion modelFlow-based ModelAuto EncoderMeshBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose TriFlow, a conditional flow matching model trained in a three-plane latent space, for efficient and realistic 3D cranial defect restoration.

TRUST: Efficient Abdominal Trauma Recognition via Image-to-Ultrasound-Video Transfer Learning

Wang, Enguang (Southeast University), Zhou, Guangquan (Southeast University)

RecognitionDomain AdaptationAnomaly DetectionConvolutional Neural NetworkTransformerVision Language ModelContrastive LearningImageVideoMultimodalityUltrasound

🎯 What it does: Propose the TRUST framework, which utilizes parameter-efficient image-to-video transfer learning (PEIVTL) to achieve rapid identification and localization of abdominal ultrasound trauma.

TrustSyn: Reliable and Divergent Synergy for Mixed-Domain Fundus Segmentation

Qiao, Baojun (Henan University), Norouzifard, Mohammad (Yoobee College of Creative Innovation)

SegmentationDomain AdaptationKnowledge DistillationConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound

🎯 What it does: Propose the TrustSyn framework, achieving hybrid-domain semi-supervised retinal image segmentation through teacher-student collaboration;

Trustworthy Endoscopic Super-Resolution

Silva-Rodríguez, Julio (ETH Zurich), Konukoglu, Ender (ETH Zurich)

Super ResolutionExplainability and InterpretabilityComputational EfficiencyConvolutional Neural NetworkTransformerDiffusion modelContrastive LearningImageVideoBiomedical DataFibre Orientation DistributionReview/Survey Paper

🎯 What it does: In the endoscopic super-resolution task, this paper proposes a framework that combines a lightweight error prediction network with conformal failure masks to identify regions of reconstruction failure and enable real-time deployment.

TSegAgent: Zero-Shot Tooth Segmentation via Geometry-Aware Vision-Language Agents

Zhuang, Shaojie, Zhou, Yuanfeng (Shandong University)

RecognitionSegmentationTransformerAgentic AIPrompt EngineeringVision Language ModelDiffusion modelPoint CloudMeshBiomedical DataComputed TomographyRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose TSegAgent, a method for zero-shot tooth segmentation and identification through a geometry-aware vision-language agent;

Tumor-Aware Direct Preference Optimization for Longitudinal Post-operative MRI Generation

Kang, Bogyeong (Korea University), Kam, Tae-Eui (Korea University)

GenerationData SynthesisTransformerPrompt EngineeringDiffusion modelImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Developed the TADPO framework, which generates post-surgical MRI images that align with clinical preferences using pre-surgical MRI, and improves the generation accuracy through tumor-aware direct preference optimization.

TumorFlow: Physics-Guided Longitudinal MRI Synthesis of Glioblastoma Growth

Biller, Valentin (Technical University of Munich), Weidner, Jonas (Munich Center for Machine Learning)

GenerationData SynthesisConvolutional Neural NetworkTransformerDiffusion modelFlow-based ModelAuto EncoderImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: A physics-based generative framework was constructed to synthesize multi-timepoint glioblastoma MRI in 3D space using continuous tumor concentration fields and biophysical diffusion models.

Two-Stage Cross-Domain Cervical Abnormality Screening with Cytopathological Image Synthesis and Knowledge Distillation

Li, Jincheng (Nantong University), Zhao, Lili (Shanghai Jiao Tong University)

Image TranslationObject DetectionDomain AdaptationKnowledge DistillationConvolutional Neural NetworkTransformerDiffusion modelGenerative Adversarial NetworkContrastive LearningImageBiomedical Data

🎯 What it does: A two-stage cross-domain cervical cell abnormal screening framework is studied, which first generates an intermediate domain through SC-UNSB, and then improves detection performance by using dual-layer feature alignment knowledge distillation.

U-ReLSE: Voxel-Wise Evidential Uncertainty-Regularized MIL for Vertebral Metastasis Detection with Hybrid Supervision

Koike, Tatsuki (FUJIFILM Corporation), Miyake, Mototaka (National Cancer Center Hospital)

SegmentationAnomaly DetectionConvolutional Neural NetworkMixture of ExpertsContrastive LearningImageBiomedical DataComputed Tomography

🎯 What it does: Detection and localization of spinal bone metastases in CT scans using a hybrid supervised 3D CNN combined with Dirichlet uncertainty weighted Log-Sum-Exp pooling.

Ultrafast Echocardiography Based on Dynamic Subspace-Guided and Temporal Adaptive Diffusion

Lu, Jingfeng (Sichuan University), Zhang, Yi (Polytechnique Montréal)

RestorationConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelOptical FlowImageBiomedical DataUltrasound

🎯 What it does: This paper proposes a fast cardiac ultrasound reconstruction framework called ST‑CODE based on diffusion models, which achieves high-quality and temporally consistent image reconstruction by utilizing dynamic subspace guidance and time-related sampling.

UltraStar: Semantic-Aware Star Graph Modeling for Echocardiography Navigation

Wang, Teng (Tsinghua University), Huang, Gao (Beijing Academy of Artificial Intelligence)

Robotic IntelligenceGraph Neural NetworkTransformerContrastive LearningImageBiomedical DataUltrasound

🎯 What it does: Propose the UltraStar framework, transforming ultrasound probe navigation from path regression to global localization based on star maps.

Uncertainty-Aware Bayesian Prompt Adaptation for Robust Cross-Modality Medical Segmentation

Hong, Sakang (Jeonbuk National University), Lee, Kyungsu (Jeonbuk National University)

SegmentationDomain AdaptationTransformerPrompt EngineeringContrastive LearningMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound

🎯 What it does: A medical image segmentation model based on SAM2, which proposes the BayesPrompt framework to achieve cross-modal adaptive segmentation under limited target modality annotations.