MICCAI 2026 Papers — Page 6
International Conference on Medical Image Computing and Computer-Assisted Intervention · 1165 papers
Interactive MWA Planning via Probabilistic Inverse Inference
Dettori, Francesco (Université de Strasbourg), Cotin, Stéphane (Université de Strasbourg)
OptimizationExplainability and InterpretabilityComputational EfficiencyConvolutional Neural NetworkDiffusion modelScore-based ModelContrastive LearningImagePoint CloudBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound
🎯 What it does: This study proposes an interactive microwave ablation (MWA) planning framework based on deep learning, which uses inverse reasoning to directly predict operational parameters such as generator power, dwell time, and insertion depth from tumor targets and blood flow distribution, and realizes real-time safety assessment and multi-antenna sequential optimization during clinical path planning.
Interpretable Local-to-Global Estimation of Brain Aging Speed From Morphological Changes Using Longitudinal Structural MRI Data
Zhang, Yuanwang (University of Pennsylvania), Fan, Yong (University of Pennsylvania)
Explainability and InterpretabilityRepresentation LearningConvolutional Neural NetworkDiffusion modelContrastive LearningBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease
🎯 What it does: What was done: Propose a framework for local–global brain aging speed estimation based on Jacobian discriminative maps derived from deformation registration, first calculating local speeds in blocks and then aggregating them into a global speed through learnable weights.
Interpretable Longitudinal Disease Forecasting with Graph-Based Multimodal Encoding and Population-Aware Language Generation
Fan, Yuheng (University of Liverpool), Zhao, He (University of Liverpool)
Explainability and InterpretabilityDrug DiscoveryGraph Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningMultimodalityTime SeriesBiomedical DataMagnetic Resonance ImagingAlzheimer's DiseaseElectronic Health RecordsRetrieval-Augmented Generation
🎯 What it does: Propose the ENIGMA framework, which utilizes multimodal time series data through graph neural network encoding and dual population memory modules, jointly with large language models, to achieve prediction of Alzheimer's disease and interpretable text generation.
Interpretable Medical Image Diagnosis via VLM-based Concept Alignment and Graph Reasoning
Hu, Yilan (Shenzhen University), Huang, Bingsheng (Shenzhen University)
ClassificationExplainability and InterpretabilityGraph Neural NetworkVision Language ModelContrastive LearningImageBiomedical DataComputed TomographyUltrasound
🎯 What it does: Proposed an interpretable medical image diagnosis framework called ConceptAlignE-G, which integrates the concept alignment of vision-language models with graph convolutional reasoning to achieve decision-making based on clinical imaging biomarkers;
Interpretable Probabilistic Medical Image Segmentation via Gaussian Process with Explicit Modelling of Annotation Bias and Variability
Li, Qi (University College London), Hu, Yipeng (King's College London)
SegmentationExplainability and InterpretabilityConvolutional Neural NetworkContrastive LearningBiomedical DataUltrasound
🎯 What it does: Propose a multi-annotator medical image segmentation framework that explicitly models annotation bias and variation in the logit space, using sparse variational Gaussian process (SVGP) to learn the image reference logit, and capturing systematic bias and random error from different annotators through annotation-specific mean and variance.
Interpretable Structure-Function Coupling via State-Space Models for Brain Connectome Analysis
Niwarthana, Amashi (Nanyang Technological University), Rajapakse, Jagath C. (Nanyang Technological University)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerAuto EncoderContrastive LearningGraphBiomedical DataMagnetic Resonance ImagingDiffusion Tensor ImagingAlzheimer's Disease
🎯 What it does: Proposes a multi-modal structure-functional coupling method based on the state space model, which can achieve many-to-many structure-functional interactions and introduces hierarchical variants to leverage the brain's modular structure.
Intraoperative Fluoroscopy-Volume Registration Using Virtual-Supervision Skeleton Attention
Zhang, Shurui (Xiamen University), Luo, Xiongbiao (Xiamen University)
Pose EstimationOptimizationData-Centric LearningConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningOptical FlowImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Propose a framework for intraoperative 2D-3D MRI/CT and fluoroscopy X-ray registration based on a virtual supervised pose regression network (VPRN) and pose optimization, which can achieve accurate registration without manual annotations.
Intraoperative X-Ray-Guided Tumor Tracking via Dynamic 3D/2D Registration with Task-Driven 3D Gaussian Priors
Geng, Haixiao, Yang, Jian (Chinese PLA General Hospital First Medical Center)
Image TranslationSegmentationPose EstimationDepth EstimationOptimizationComputational EfficiencyConvolutional Neural NetworkTransformerDiffusion modelNeural Radiance FieldAuto EncoderContrastive LearningGaussian SplattingOptical FlowImageBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasound
🎯 What it does: Propose a two-stage dynamic 3D/2D registration framework, which constructs task-driven label-aware 3D Gaussian priors using pre-operative CT, and achieves real-time 3D localization and deformation estimation of the liver and tumor through Gaussian displacement updates guided by intraoperative X-ray sequences and 2D motion cues.
IntraStyler: Intra-Domain Style Synthesis for Cross-Modality MRI Domain Adaptation
Liu, Han (Siemens Healthineers), Oguz, Ipek (Vanderbilt University)
SegmentationDomain AdaptationTransformerDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Proposed a cross-modal MRI domain adaptation method called IntraStyler, which utilizes unpaired image translation to achieve fine-grained style adaptive synthesis within the target domain, thereby enhancing the generalization ability of downstream segmentation models.
InvDetect: Unsupervised Medical Anomaly Detection in the Noise Latent Space of DDIM
Ma, Xinyu (McMaster University), Chu, Lingyang (McMaster University)
Anomaly DetectionTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasoundDiffusion Tensor Imaging
🎯 What it does: Proposes an unsupervised medical anomaly detection framework named InvDetect, which uses a DDIM model trained only on normal images to perform reverse inference, constructing a noise latent space and performing anomaly discrimination via one-class SVM in this space, while introducing spatial coherence post-processing to obtain the final anomaly segmentation.
Inverse Protocol Prediction from Spheroid Microscopy Imaging via Morphology-Aware Structured Learning
Chauhan, Joohi (Motilal Nehru National Institute of Technology), Srivastava, Ayush (Motilal Nehru National Institute of Technology)
ClassificationDomain AdaptationExplainability and InterpretabilityConvolutional Neural NetworkTransformerContrastive LearningImageBiomedical Data
🎯 What it does: This paper proposes the task of directly inferring experimental culture protocols (Inverse Protocol Prediction) from a single bright-field spirochete image, aiming to utilize morphological information to verify and detect the reliability of experimental records.
Inversion from ECG Foundation Model
Lee, Junseok (Korea University), Yoo, Chuck (Korea University)
Safty and PrivacyRepresentation LearningAdversarial AttackConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningTime SeriesBiomedical DataElectrocardiogram
🎯 What it does: Proposes the SHINE framework, which uses a limited number (≤1000) of queries to invert the embeddings of ECG foundation models (FMs), generating high-fidelity ECG signals and revealing privacy risks caused by embedding leakage.
IOSVLM: A 3D Vision-Language Model for Unified Dental Diagnosis from Intraoral Scans
Xiong, Huimin (Zhejiang University), Liu, Zuozhu (Zhejiang University)
ClassificationGenerationAnomaly DetectionTransformerLarge Language ModelSupervised Fine-TuningVision Language ModelContrastive LearningTextMultimodalityPoint CloudBiomedical Data
🎯 What it does: Designed an end-to-end 3D vision-language model, IOSVLM, which utilizes raw 3D intraoral scanning point clouds to achieve unified diagnosis for multiple diseases and generative visual question answering, and constructed a large-scale IOSVQA dataset.
IRIS: An Intelligent Vision-Language System for Ocular Surface Diseases via Topic Tree and Scene-Driven VQA Generation
Wei, Hao (Chinese University of Hong Kong), Yuan, Wu (Chinese University of Hong Kong)
Explainability and InterpretabilityComputational EfficiencyData-Centric LearningDrug DiscoveryTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataAlzheimer's DiseaseElectronic Health RecordsBenchmarkRetrieval-Augmented Generation
🎯 What it does: Constructed a visual question-answering dataset for ocular surface diseases called IRIS-120K, and performed low-rank adaptation fine-tuning on a 4B parameter visual-language model on this dataset, ultimately obtaining the IRIS-4B model, which can be deployed on mobile devices.
JANUS: Anatomy-Conditioned Gating for Robust CT Triage Under Distribution Shift
Dahal, Lavsen (Duke University), Lo, Joseph Y. (Duke University)
ClassificationDomain AdaptationAnomaly DetectionConvolutional Neural NetworkTransformerContrastive LearningImageBiomedical DataComputed Tomography
🎯 What it does: Propose a dual-stream physiological conditioning network named JANUS, which integrates macro radiomic quantitative features with visual representations to achieve more robust multi-label classification under distribution shift in the CT triage task.
Joint Biomarker and Survival Prediction via Concept-Conditioned Multimodal Slot Factorization
Huang, Zuqi (Shanghai Jiao Tong University), Li, Zhongyu (Shanghai Jiao Tong University)
Drug DiscoveryRecurrent Neural NetworkTransformerPrompt EngineeringMixture of ExpertsVision Language ModelContrastive LearningImageTextMultimodalityBiomedical Data
🎯 What it does: Propose a concept-conditioned slot factorization framework called PathoSlot, which simultaneously performs biomarker prediction and survival risk estimation on multimodal pathological data (whole slide images and pathological reports).
Joint Decomposition and Learned Deformation for Part-to-Whole MRI to Histology Registration
Athalye, Chinmayee (University of Pennsylvania), Yushkevich, Paul A. (University of Pennsylvania)
Image TranslationOptimizationDiffusion modelNeural Radiance FieldAuto EncoderGenerative Adversarial NetworkContrastive LearningOptical FlowImageBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Jointly utilize implicit neural representations to achieve 3D-to-2D registration from MRI to single tissue slices, including slice localization, fragment stitching, and deformation modeling.
Joint Imaging–ROI Representation Learning via Cross-View Contrastive Alignment for Brain Disorder Classification
Liang, Wei (Lehigh University), He, Lifang (Lehigh University)
ClassificationRepresentation LearningConvolutional Neural NetworkGraph Neural NetworkTransformerContrastive LearningBiomedical DataMagnetic Resonance Imaging
🎯 What it does: This paper proposes a cross-view contrastive learning framework for jointly learning global image and local ROI graph representations to achieve brain disease classification.
Joint Multi-modal Learning for Susceptibility-Induced Distortion Correction in Diffusion MRI
Feng, Jianhui (Fudan University), Qiao, Yuchuan (Fudan University)
RestorationConvolutional Neural NetworkAuto EncoderContrastive LearningMultimodalityBiomedical DataMagnetic Resonance ImagingFibre Orientation DistributionDiffusion Tensor Imaging
🎯 What it does: Proposes a multi-modal unsupervised learning framework, MMDC-Net, for jointly correcting magnetic susceptibility distortions in diffusion MRI using B0 and FOD images.
Joint Segmentation and Graph-Based Skeletal Representation Learning with Geometric Priors
Wu, Yinhao (University of Texas at Arlington), Huang, Junzhou (University of Texas at Arlington)
SegmentationRepresentation LearningConvolutional Neural NetworkGraph Neural NetworkDiffusion modelScore-based ModelFlow-based ModelRectified FlowNeural Radiance FieldAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease
🎯 What it does: Propose an end-to-end deep learning framework that can directly predict the s-rep (skeleton model) of the hippocampus and voxel segmentation from 3D medical images.
Joint-Level Motor Imagery Decoding in fNIRS via a Unified Pairwise Machine Learning Framework
Nosheen, Tayyaba (Silesian University of Technology), Malik, Aamir Saeed (National University of Sciences and Technology)
ClassificationRecognitionHyperparameter SearchSupervised Fine-TuningTime SeriesBiomedical Data
🎯 What it does: A unified binary classification machine learning framework (TSCIML-MI) was constructed and evaluated for fine-grained decoding of imagined movements of eight joints in the upper limbs from fNIRS signals.
Just 32 Tokens: Medical Referring Image Segmentation via Next-Token Mask Prediction
Chen, Xinyu (University of Sydney), Li, Yonghui (University of Sydney)
SegmentationComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageBiomedical DataMagnetic Resonance ImagingComputed TomographyBenchmarkChain-of-Thought
🎯 What it does: This paper proposes the 32TokenMRISeg framework, which utilizes a multimodal large language model to achieve medical reference image segmentation through 32 discrete mask tokens.
K-MaT: Knowledge-Anchored Manifold Transport for Cross-Modal Prompt Learning in Medical Imaging
Zeng, Jiajun (University of Bonn), Albarqouni, Shadi (University of Bonn)
ClassificationDomain AdaptationRepresentation LearningTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound
🎯 What it does: Developed a knowledge-anchored manifold transport (K-MaT) framework to achieve cross-modal prompt learning for zero-shot low-end imaging after training on high-end medical imaging.
KAN-AINet: Kolmogorov-Arnold Network with Adaptive Illumination Modulation for Generalizable Polyp Segmentation
Timklaypachara, Watcharapong (Mahidol University), Achakulvisut, Titipat (Mahidol University)
SegmentationConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningImageBiomedical Data
🎯 What it does: Propose a KAN-AINet model based on ConvNeXt-U-Net, utilizing the Kolmogorov-Arnold network (KAN) to achieve adaptive illumination modulation and boundary attention, thereby enhancing the cross-dataset robustness of colon polyp segmentation.
KANEx: Translating Kolmogorov-Arnold Networks’ Interpretability to Medical Explainability
Shailya, Krithi (Indian Institute of Technology Madras), Ravindran, Balaraman (Indian Institute of Technology Madras)
ClassificationExplainability and InterpretabilityConvolutional Neural NetworkTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataComputed TomographyElectronic Health Records
🎯 What it does: Proposes the KANEx framework, combining Kolmogorov-Arnold networks (KAN) with Vision-Language models (VLM), to simultaneously generate interpretable heatmaps and textual reports in multi-label diagnosis of chest X-rays.
KANResDiff: Learning Local Residual Diffusion via Kolmogorov-Arnold Network for Ambiguous Medical Image Segmentation
Li, Fanding (Beijing Institute Of Technology), Li, Shuo (Beijing Institute Of Technology)
SegmentationGenerationTransformerDiffusion modelScore-based ModelImageStochastic Differential Equation
🎯 What it does: This paper proposes a new loss function based on RSB for training score matching models.
KEPIL: Knowledge-Enhanced Prompt-Image Learning for Prompt-Robust Disease Detection
Luo, Haozhe (University of Bern), Reyes, Mauricio (University of Bern)
ClassificationSegmentationAnomaly DetectionTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataComputed TomographyElectronic Health Records
🎯 What it does: Proposed the KEPIL framework, achieving prompt-robust zero-shot disease detection through knowledge-enhanced dynamic prompts and semantic contrastive learning.
Keypoint Correspondence Learning for CBCT–IOS Registration
Park, Jeonglok (Soongsil University), Chung, Minyoung (Soongsil University)
Pose EstimationTransformerContrastive LearningOptical FlowPoint CloudMeshBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: This paper proposes an end-to-end keypoint correspondence learning framework for rigid registration between Cone-Beam CT (CBCT) and intraoral scanning (IOS), directly estimating keypoint correspondences in the feature space and refining the transformation through iterative SVD, avoiding the traditional process of segmentation to surface registration.
KGT-Reg: Knowledge-Guided Multimodal Framework for Medical Image Registration
Dai, Yuhe (National University of Singapore), Huang, Weimin (Agency for Science, Technology and Research)
Representation LearningGraph Neural NetworkTransformerVision Language ModelDiffusion modelContrastive LearningOptical FlowImageTextGraphBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Propose the KGT-Reg framework, integrating medical image registration with knowledge graphs (KG) and text semantics to improve anatomical consistency and registration accuracy.
KiD-Seg: Kinetic-Disentangled Contrastive Learning with Anchor Guidance for Incomplete Multi-Modal Liver Tumor Segmentation
Khor, Hee Guan (Ant Group), Lu, Le (Ant Group)
SegmentationKnowledge DistillationConvolutional Neural NetworkAuto EncoderContrastive LearningMultimodalityBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Propose the KiD-Seg framework to achieve liver tumor segmentation based on T2WI anchors, supporting arbitrary missing multi-modal MRI at the group level.
KIDA: Kinematic-Intent Dual-path Alignment for Surgical Error Detection
Luo, Yuxuan (Huazhong University of Science and Technology), Wang, Zhiwei (Huazhong University of Science and Technology)
Anomaly DetectionRobotic IntelligenceRecurrent Neural NetworkTransformerVision-Language-Action ModelContrastive LearningOptical FlowVideoTextBiomedical Data
🎯 What it does: Proposes a dual-path alignment framework called KIDA for detecting subtle errors in robotic-assisted surgery by aligning actual execution with clinical standard intentions.
KiMoCo-Net: Dual-Domain Learning for MRI Motion Artifact Correction
Fekry, Marina Maher (Cairo University), Al-masni, Mohammed A. (King Fahd University of Petroleum & Minerals)
RestorationConvolutional Neural NetworkGenerative Adversarial NetworkContrastive LearningBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Proposed a KiMoCo-Net, which first restores phase and amplitude consistency in the k-space complex domain, and then refines it in the image domain to complete MRI motion artifact correction.
Kinetics-Informed Neural Representation for Dual-Tracer Dynamic PET Separation
Wang, Xingyue (ShanghaiTech University), Cui, Zhiming (Shanghai United Imaging Healthcare Co., Ltd.)
OptimizationRepresentation LearningDiffusion modelScore-based ModelNeural Radiance FieldAuto EncoderContrastive LearningBiomedical DataPositron Emission Tomography
🎯 What it does: Proposed a dual-tracer PET signal separation method based on implicit neural representations (INR), which directly predicts dynamic parameters for each voxel through coordinate encoding + MLP, and utilizes a chamber-chamber model to reconstruct single-tracer imaging.
KJD-25K: A 3D Multi-modal Dataset and Benchmark for Knee Joint Diagnosis
Li, Yihui, Niu, Ben (Shenzhen Second Peoples Hospital)
TransformerLarge Language ModelSupervised Fine-TuningVision Language ModelImageTextMultimodalityBiomedical DataMagnetic Resonance ImagingBenchmarkRetrieval-Augmented Generation
🎯 What it does: The paper constructs a large-scale multi-modal dataset named KJD-25K, containing 25,000 pairs of knee MRI volumes and reports, and conducts benchmark experiments on report generation using eight state-of-the-art multi-modal large language models on this dataset.
KMP-MIL: Knowledge Memory Pool Multiple Instance Learning with Foundation Model for Continual Whole Slide Image Classification
Yuan, Lingling (Northeastern University), Li, Chen (University of Lübeck)
ClassificationImage TranslationComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerVision Language ModelContrastive LearningImageBiomedical DataRetrieval-Augmented Generation
🎯 What it does: Proposed a replay-free multi-instance learning framework, KMP-MIL, for continual learning in whole slide image classification.
Knowledge Distillation-Driven Quality Transfer with Importance-Guided Supervision for Fundus Image Enhancement
Hou, Qingshan (Taiyuan University of Technology), Liu, Yangyang (Jiangsu Second Normal University)
RestorationDomain AdaptationKnowledge DistillationConvolutional Neural NetworkDiffusion modelGenerative Adversarial NetworkContrastive LearningImageBiomedical Data
🎯 What it does: Propose a semi-supervised dual-branch framework that achieves global quality transfer through unsupervised knowledge distillation and prioritizes the repair of clinically critical structures via pixel-level importance-guided supervision, thereby improving retinal image quality.
Knowledge-Enhanced Representation Learning with Retrieval-Augmented Multimodal Fusion for Survival Prediction
Zhang, Zeyu (University of British Columbia), Bashashati, Ali (University of British Columbia)
RetrievalRepresentation LearningTransformerVision Language ModelContrastive LearningImageMultimodalityTabularBiomedical DataElectronic Health RecordsRetrieval-Augmented Generation
🎯 What it does: Developed the KERA framework, combining knowledge-enhanced pre-training with retrieval-enhanced multi-modal fusion, achieving survival prediction through the integration of whole-slide images and transcriptomics.
KOAL: Knowledge-Driven Prostate Cancer Grading with Ordinal-Aware Learning
Guo, Zheng (Sichuan University), Wang, Yan (Sichuan University)
ClassificationConvolutional Neural NetworkTransformerLarge Language ModelPrompt EngineeringContrastive LearningImageTextMultimodalityBiomedical DataMagnetic Resonance ImagingRetrieval-Augmented Generation
🎯 What it does: Based on multi-modal mpMRI, combining clinical variables and report priors, the KOAL framework is proposed for lesion-level Gleason Grade Group prediction.
KS-SAM: Kolmogorov-Arnold-driven Shape Alignment for Single-Domain Generalization in 3D Medical Image Segmentation
Wu, Rengmin (Shenzhen University), Lei, Baiying (Qingyuan People's Hospital)
SegmentationDomain AdaptationKnowledge DistillationConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyBenchmark
🎯 What it does: Designed and implemented the KS-SAM framework, achieving geometric-driven adaptation in single-domain generalization for 3D medical images through the Kolmogorov-Arnold network.
L-TGVN: Leveraging Longitudinal Priors for Personalized Rapid MRI
Atalık, Arda (NYU Center for Data Science), Sodickson, Daniel K. (Center for Advanced Imaging Innovation and Research, NYU Grossman School of Medicine)
RestorationConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Proposed a longitudinal trust-guided variational network called L-TGVN for reconstructing MRI images at extremely high acceleration rates, while utilizing the patient's most recent follow-up image as side information.
LAD: Learnable Anisotropic Deformation for 3D Medical Image Analysis
Liu, Chengcai (Beihang University), Wang, Shuo (Beihang University)
ClassificationSegmentationTransformerContrastive LearningBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Propose a learnable anisotropic deformation framework called LAD, enabling 3D Transformer to adapt to the spatial inhomogeneity of medical images.
LaGuadia: Language-Guided Adaptive Distillation from Pathology Foundation Models
Kim, Gangsu (Korea University), Jeong, Won-Ki (Korea University)
Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerSupervised Fine-TuningPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataElectronic Health RecordsBenchmark
🎯 What it does: Proposes the LaGuadia framework, which dynamically integrates the expertise of multiple pathological foundation models (PFMs) into a lightweight (87M parameters) WSI encoder through language-guided adaptive knowledge distillation.
Landmark-Free Assessment of Lower-Limb Alignment with Implicit Neural Shape Functions from Knee Radiographs
Hu, Zhisen (University of Manchester), Tiulpin, Aleksei (University of Oulu)
Pose EstimationOptimizationConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Developed an automatic method for evaluating lower limb alignment in knee X-rays based on implicit neural shape functions, directly regressing alignment angles from the latent space and eliminating the traditional anatomical landmarking step.
LangOR: 3D Language Field Reconstruction for Operating Room
Li, Wei (Chinese University of Hong Kong), Heng, Pheng-Ann (Chinese University of Hong Kong)
RecognitionSegmentationTransformerContrastive LearningGaussian SplattingImageVideoTextPoint CloudBenchmark
🎯 What it does: Propose an unsupervised 3D language field reconstruction framework, LangOR, for semantic understanding in operating room scenes.
Large-Scale Distributed GPU-Accelerated Respiratory Motion-Resolved Reconstruction of 3D Non-cartesian mGRE MRI
Zhang, Chao (Stony Brook University), Kee, Youngwook (Stony Brook University)
OptimizationComputational EfficiencyBiomedical DataMagnetic Resonance Imaging
🎯 What it does: A hierarchical multi-node multi-GPU framework is proposed for respiratory motion-resolved reconstruction in three-dimensional non-Cartesian multi-echo gradient echo (mGRE) magnetic resonance imaging.
Large-Small Model Collaboration for Zero-Shot Surgical Phase Recognition
Zhang, Yiyi (Chinese University of Hong Kong), Heng, Pheng-Ann (Chinese University of Hong Kong)
RecognitionConvolutional Neural NetworkTransformerPrompt EngineeringVision Language ModelContrastive LearningVideoText
🎯 What it does: Designed a large model-small model collaborative framework LaST to achieve zero-shot out-of-distribution surgery phase recognition, improving pseudo-label quality through iterative time refinement and cyclic replay.
Latent Lessons from Histopathology: A Registration-Free Cross-Modal Supervision Framework for mpMRI-Based ISUP Grading
Schouten, Daan, Litjens, Geert
ClassificationConvolutional Neural NetworkImage
🎯 What it does: The paper explores a new deep learning model aimed at improving the accuracy of image classification.
Latent Motion Spectral Regularity for Congenital Heart Defect Screening in Fetal Echocardiography
Yang, Yingyu (University of Oxford), Noble, J. Alison (University of Oxford)
Anomaly DetectionAuto EncoderContrastive LearningOptical FlowVideoBiomedical DataUltrasound
🎯 What it does: This paper proposes an abnormal detection framework based on normal fetal four-chamber view videos for the automatic screening of congenital heart defects (CHD).
Latent Pathology–Age Decoupling in Functional Connectivity Graphs for Multi-task Schizophrenia Assessment
Yan, Yige (Nanyang Technological University), Rajapakse, Jagath C. (Nanyang Technological University)
Anomaly DetectionRepresentation LearningGraph Neural NetworkAuto EncoderContrastive LearningGraphBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease
🎯 What it does: Propose the LPAD framework, which utilizes a dual-stream graph neural network to achieve decoupling between disease pathology and age in the latent space, and reintroduces age as a secondary context through a clinical information gating module to enhance the prediction of clinical symptoms and cognitive deficits in schizophrenia.
Latent-CURE: Interpretable Breast Cancer Diagnosis via Dual-Asymmetric Chain-of-Thought
Zhao, Weiyi (Shanghai University of Engineering Science), Qiu, Xihe (Shanghai University of Engineering Science)
ClassificationAnomaly DetectionExplainability and InterpretabilityConvolutional Neural NetworkTransformerPrompt EngineeringContrastive LearningImageBiomedical DataUltrasoundRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose the Latent-CURE framework, achieving interpretable and deterministic diagnosis in breast ultrasound by constructing implicit chain-of-thought reasoning (Feature-aware CoT) in the latent space and employing dual asymmetric optimization.
Latent-to-Latent Flow for Volumetric Stochastic Segmentation
Todd, Omar (Imperial College London), Glocker, Ben (Imperial College London)
SegmentationConvolutional Neural NetworkDiffusion modelScore-based ModelFlow-based ModelAuto EncoderBiomedical DataComputed TomographyStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: To address the diversity uncertainty in medical volume segmentation, this paper proposes and implements a flow matching method in the latent space (L2L-Flow), which significantly improves inference speed while maintaining clinically relevant uncertainty estimation.
Learnable Gaussian Prototype Guided Ordinal Learning for Fine-Grained WSI Grading
Deng, Yingjiao (East China Normal University), Li, Qingli (East China Normal University)
ClassificationExplainability and InterpretabilityComputational EfficiencyTransformerAuto EncoderContrastive LearningGaussian SplattingImageBiomedical DataElectronic Health Records
🎯 What it does: Propose a Gaussian Prototype-based Ordinal Multi-Instance Learning framework (GPO-MIL), which constructs a prototype space consistent with disease severity and uses instance responsibility to guide attention for WSI-level grading;
Learnable In-Context Templates for Medical Image Segmentation via In-Context Learning
Bao, Xueqi (Beijing University of Posts and Telecommunications), Lao, Qicheng (Beijing University of Posts and Telecommunications)
SegmentationComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerPrompt EngineeringContrastive LearningImageBiomedical Data
🎯 What it does: This study proposes learnable context templates, achieving efficient zero-shot, retrieval-free adaptation for medical image segmentation tasks through template distillation within visual foundation models.
Learning Beyond Pixels: Structured Knowledge Topology Consistency for Semi-supervised 2D/3D Medical Image Segmentation
Chen, Yu (Jinan University), Wei, Honghao (Jinan University)
SegmentationContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: This paper proposes a semi-supervised 2D/3D medical image segmentation framework called SKTC-Net based on structural knowledge topological consistency.
Learning Class Difficulty via Dynamic Focal Attention for Histopathology Segmentation
Rathukohe Mudiyanselage, Lakmali Nadeesha Kumari (University of Kentucky), Cheung, Sen-Ching Samson (University of Kentucky)
SegmentationTransformerContrastive LearningBiomedical DataBenchmark
🎯 What it does: Proposed and implemented Dynamic Focal Attention (DFA) in pathological image semantic segmentation, directly encoding class difficulty by introducing learnable class-level bias into the transformer cross-attention mechanism;
Learning Directional Semantic Transitions for Longitudinal Chest X-Ray Analysis
Hu, Zhangfeng (Rensselaer Polytechnic Institute), Yan, Pingkun (Massachusetts General Hospital, Harvard Medical School)
ClassificationExplainability and InterpretabilityRepresentation LearningTransformerVision Language ModelContrastive LearningImageTextBiomedical DataComputed TomographyElectronic Health Records
🎯 What it does: Proposes a visual-language pre-training framework called ProTrans, which learns directional semantic transitions in chest X-rays over time to enable longitudinal CXR disease progression analysis.
Learning Diverse and Realistic Polyp Data via Style-Fused Diffusion Models
Han, Longfei (Beijing Technology and Business University), Li, Haisheng (Beijing Technology and Business University)
SegmentationGenerationData SynthesisConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningImageBiomedical Data
🎯 What it does: Propose a conditional diffusion model based on style fusion (SFD-Polyp), which injects the style of real images through the Style Integration Module (SIM), and generates diverse and anatomically accurate mask shapes during the inference phase using Morphology-Preserving Mask Synthesis (MPMS), thus synthesizing diverse and realistic polyp image-mask pairs to enhance the dataset.
Learning Dual Prior Presentation for Opportunistic Screening of Visceral Artery Aneurysms on Non-Contrast CT
Zhang, Jianfeng (DAMO Academy, Alibaba Group), Xu, Minfeng (DAMO Academy, Alibaba Group)
SegmentationAnomaly DetectionConvolutional Neural NetworkTransformerSupervised Fine-TuningAuto EncoderContrastive LearningImageBiomedical DataComputed Tomography
🎯 What it does: This paper proposes a dual prior learning framework called DeepVAA for opportunistic screening of visceral aneurysms (VAA) on non-enhanced CT (NCT).
Learning from Good Neighbors: Slice-wise Quality-aware CT-to-Diffusion MRI Synthesis
Na, You-Kyoung (Chonnam National University), Cho, Yeong-Jun (Chonnam National University)
Image TranslationData SynthesisConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Propose a hierarchical quality-aware CT-to-MRI (DWI) synthesis framework called LGN. It first generates a draft MRI using a latent diffusion model, then predicts slice quality through a teacher-student network, and finally refines locally using cross-attention from high-quality neighboring slices, improving synthesis quality while maintaining three-dimensional anatomical continuity.
Learning Geometry-Aware Bundles with Spectral Heat Kernel Diffusion for fMRI-based Brain Disease Diagnosis
Yin, Feiyu (Fudan University), Yu, Jinhua (Fudan University)
ClassificationAnomaly DetectionRepresentation LearningGraph Neural NetworkDiffusion modelContrastive LearningGraphBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease
🎯 What it does: Proposed a dual-stream contrastive learning framework VB-GCL based on the vector bundle theory for disease diagnosis in fMRI neural networks
Learning Hierarchical Anatomy-Grounded Representation for Cross-Modal Registration
Mi, Jia (ShanghaiTech University), Shen, Dinggang (ShanghaiTech University)
OptimizationRepresentation LearningConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Studied the HAGR framework, a hierarchical anatomy-driven cross-modal deformable registration representation learning framework
Learning Invariant Anatomical Representations from Multi-Sequence MRI for Brain Segmentation
Liang, Jianwen, Tang, Xiaoying (Southern University of Science and Technology)
SegmentationDomain AdaptationRepresentation LearningConvolutional Neural NetworkTransformerPrompt EngineeringVision Language ModelAuto EncoderContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingRetrieval-Augmented Generation
🎯 What it does: Propose the Text-guided Invariant Anatomical Learning (TeInL) framework, which leverages self-supervised pre-training on multi-sequence MRI to decouple anatomical structures from sequence styles, and enhances brain tissue and lesion segmentation performance through text-image alignment learning to capture pathology-related features.
Learning Modality-Invariant Structure and Distribution Alignment for Multimodal Medical Image Registration
Xiao, Siqi (Shenzhen University), Wang, Yi (Shenzhen University)
Convolutional Neural NetworkTransformerContrastive LearningMultimodalityBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Propose a cross-modal medical image registration framework that jointly learns structure-agnostic representations and alignment
Learning Prototypes for Unsupervised Joint Generation of Conditional Templates and Atlases
Zhang, Jichang (Beijing Normal University), Li, Shuyu (Beijing Normal University)
GenerationData SynthesisDiffusion modelAuto EncoderContrastive LearningBiomedical DataMagnetic Resonance Imaging
🎯 What it does: This paper proposes an unsupervised framework called LatentAtlas, which jointly generates conditional brain templates and atlases. It aligns public templates and atlases through shared latent space learning of prototypes, enabling the synchronous generation of templates and atlases under specific conditions (e.g., age).
Learning Quantitative Calibration of Computed Tomography for Bone Mineral Density Estimation Using Multi-Center Data
Schöller, Job H. J. (Nara Institute of Science and Technology), Otake, Yoshito (Osaka University)
Image TranslationRestorationSegmentationGenerationData SynthesisConvolutional Neural NetworkDiffusion modelContrastive LearningImageBiomedical DataComputed Tomography
🎯 What it does: Achieve calibration-free bone density estimation by generating voxel-level QCT density maps directly from conventional CT using a supervised 3D image-to-image translation framework.
Learning Robust Medical Image Segmentation Under Mixed-Quality Annotations
Jeong, Minjae (Pohang University of Science and Technology), Kim, Won Hwa (Pohang University of Science and Technology)
SegmentationData-Centric LearningConvolutional Neural NetworkDiffusion modelImageBiomedical Data
🎯 What it does: This paper proposes an iterative training framework named RGSS, aimed at learning robust models in medical image segmentation tasks with mixed quality annotations;
Learning Role-Conditioned Alignment for Medical Image Referring Segmentation
Li, Kun (University of Liverpool), Zheng, Yalin (Chinese Academy of Sciences)
SegmentationConvolutional Neural NetworkTransformerVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: This paper proposes a role-conditioned alignment-based model for medical image referential segmentation (RCA), which achieves more accurate segmentation by explicitly separating textual descriptions into three categories: lesion semantics, quantity, and location, and then performing fine-grained alignment with visual features;
Learning Task-Specific Anatomical Priors via Context Distillation for In-Context Medical Image Segmentation
Yang, Guoqing (Northwestern Polytechnical University), Xia, Yong (Northwestern Polytechnical University)
SegmentationKnowledge DistillationMeta LearningTransformerContrastive LearningBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound
🎯 What it does: Proposed a Context Distillation via Masked Image Modeling (CD-MIM) framework, which learns demonstration samples in medical image segmentation as compressed task-specific anatomical prior templates, to improve segmentation accuracy in in-context learning (ICL) and reduce redundancy;
Learning the Hierarchical Organization in Brain Network for Brain Disorder Diagnosis
Tang, Jingfeng (Northeastern University), Zaiane, Osmar R. (University of Alberta)
ClassificationExplainability and InterpretabilityGraph Neural NetworkTransformerContrastive LearningGraphBiomedical DataAlzheimer's Disease
🎯 What it does: Propose a framework called BrainHO that learns the hierarchical organization of brain networks for the diagnosis of brain disorders (ASD, MDD);
Learning to Distort: Weakly-Supervised Image Quality Transfer for Prostate DWI Correction
Tang, Yucheng (University College London), Hu, Yipeng (University College London)
Image TranslationRestorationData SynthesisConvolutional Neural NetworkTransformerDiffusion modelGenerative Adversarial NetworkContrastive LearningImageBiomedical DataMagnetic Resonance ImagingDiffusion Tensor Imaging
🎯 What it does: Proposed a weakly supervised image quality transfer (IQT) framework, which first generates realistic distorted images from undistorted prostate DWI using quality prototype flow matching, and then uses these synthetic pairs to train a supervised correction network to achieve distortion correction for single-echo EPI DWI.
Learning to Optimize Radiotherapy Plans via Fluence Maps Diffusion Model Generation and LSTM-Based Optimization
Poles, Isabella (Politecnico di Milano), Comaniciu, Dorin (Politecnico di Milano)
OptimizationRecurrent Neural NetworkTransformerDiffusion modelScore-based ModelGenerative Adversarial NetworkContrastive LearningBiomedical DataComputed TomographyReview/Survey Paper
🎯 What it does: This paper proposes an end-to-end VMAT radiotherapy planning method. It first uses a first-order diffusion model to generate feasible beam intensity maps, and then employs an LSTM to learn an optimizer that rapidly iteratively improves the beam intensity while maintaining dose consistency.
Learning to Read Where to Look: Disease-Aware Vision–Language Pretraining for 3D CT
Ging, Simon (University of Freiburg), Brox, Thomas (University of Freiburg)
ClassificationSegmentationRetrievalTransformerPrompt EngineeringVision Language ModelContrastive LearningGaussian SplattingBiomedical DataComputed Tomography
🎯 What it does: Built a 3D CT audio-visual language model called RadFinder based on 98k CT scan-report pairs, incorporating disease prompt supervision and slice-level localization supervision.
Learning to Segment Burns with Limited Labels: A Dual-Path Cascaded Network with Uncertainty-Guided Teacher Diffusion Exposure
Chauhan, Joohi (Motilal Nehru National Institute of Technology Allahabad), Singh, Prabal Pratap (Motilal Nehru National Institute of Technology Allahabad)
SegmentationData SynthesisDomain AdaptationConvolutional Neural NetworkDiffusion modelContrastive LearningImageBiomedical Data
🎯 What it does: Propose a semi-supervised burn segmentation framework called UADA-MSPCNet, combining a dual-path cascade network, uncertainty-based teacher-student consistency learning, and diffusion generation enhancement on the teacher side.
Learning to Trim: End-to-End Causal Graph Pruning with Dynamic Anatomical Feature Banks for Medical VQA
Xu, Zibo (Tianjin University), Su, Yuting (Tianjin University)
TransformerVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasound
🎯 What it does: This paper proposes an end-to-end causal pruning framework called LCT, which suppresses dataset bias in medical VQA by utilizing a dynamic anatomical feature bank (DAFB) and a learnable pruning module (Causal Trimming).
Learning to Unify Deformable Shape and Texture Representations for Cardiac Video Classification
Hossain, Tonmoy (University of Virginia), Zhang, Miaomiao (University of Virginia)
ClassificationExplainability and InterpretabilityRepresentation LearningConvolutional Neural NetworkTransformerContrastive LearningOptical FlowVideoBiomedical DataMagnetic Resonance Imaging
🎯 What it does: This study proposes ShapeFuse, aiming to unify deformable shape representations and image texture representations into a shared latent space, and dynamically fuse shape and texture through bidirectional cross-modal temporal attention and adaptive gating for disease classification of cardiac MRI videos.
Learning What Can Move: Patient-Specific Motion Subspaces from a Single Scan
Zhu, Yinheng (University of Texas Southwestern Medical Center), Lu, Weiguo (University of Texas Southwestern Medical Center)
RestorationPose EstimationOptimizationComputational EfficiencyConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningOptical FlowImagePoint CloudMeshBiomedical DataComputed Tomography
🎯 What it does: Predict patient-specific motion subspaces from a single static CT image, and rapidly reconstruct complete 3D motion fields under different real-time observations (3D images, 2D images, 2D segmentation masks, surface points) using a closed-form coefficient solver.
Learning Where to Look: A Reinforcement Learning Framework for Robust Micro-ultrasound Prostate Cancer Detection
Abootorabi, Mohammad Mahdi (University of British Columbia), Abolmaesumi, Purang (University of British Columbia)
ClassificationSegmentationConvolutional Neural NetworkTransformerReinforcement LearningPrompt EngineeringAuto EncoderContrastive LearningBiomedical DataUltrasound
🎯 What it does: Propose the Prost-RL framework, which learns the attention regions of micro-ultrasound images through reinforcement learning. By combining a foundational model encoder-decoder, it achieves 'learn attention first, then decode,' thereby improving the detection and localization of prostate cancer under weakly supervised conditions.
Learning Where to Look: Pathologist-Inspired Multi-field-of-View Evidence Retrieval for Cell Type Classification in H&E
Yuan, Ruizhi (University of Pittsburgh), Chen, Wei (University of Pittsburgh)
ClassificationExplainability and InterpretabilityComputational EfficiencyTransformerMixture of ExpertsContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyRetrieval-Augmented Generation
🎯 What it does: Proposed a hybrid expert framework called TRACE, which classifies cell types in H&E slices by leveraging token-level evidence retrieval from multi-scale perspectives.
LEGEND: A Language-Aligned EEG Foundation Model with Flow Matching Latent Denoising
Gui, Yiyu (Peking University), Luo, Guibo (Peking University)
ClassificationAnomaly DetectionRepresentation LearningConvolutional Neural NetworkRecurrent Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningFlow-based ModelContrastive LearningTextTime SeriesBiomedical DataAlzheimer's DiseaseElectrocardiogramReview/Survey Paper
🎯 What it does: Propose LEGEND, a foundational model that aligns EEG signals with natural language instructions, and introduce flow matching latent denoising to enhance structural learning.
LeGend: Lesion-Guided Longitudinal CXR Generation via Gaussian-Biased Causal Attention
Wang, Yiran (University of Sydney), Zhou, Luping (University of Sydney)
GenerationData SynthesisTransformerVision Language ModelDiffusion modelGaussian SplattingImageBiomedical DataComputed TomographyPositron Emission TomographyUltrasoundElectronic Health Records
🎯 What it does: Studied the issue of insufficient lesion focus in autoregressive models when generating follow-up chest X-ray images, and proposed a lesion-guided generation framework based on Gaussian-biased causal attention.
Lesion-Centric Structured Report Generation for 3D CT with Multimodal Large Language Models via Hierarchical Evidence Tokens
Zhao, Xiuheng (Tsinghua University), Li, Miao (Tsinghua University)
GenerationExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageMultimodalityBiomedical DataComputed Tomography
🎯 What it does: A hierarchical evidence marking mechanism is proposed for generating structured reports of lesions in 3D CT images;
Leuko, I Am Your Prototype: A Prototypical Detection Framework for Fine-Grained Laryngeal Lesion Stratification
Federici, Lorenzo (Università Politecnica delle Marche), Moccia, Sara (Istituto Italiano di Tecnologia)
ClassificationObject DetectionAnomaly DetectionConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningImageBiomedical DataComputed TomographyUltrasound
🎯 What it does: Proposed a dual-branch real-time detection framework that integrates prototype learning into the YOLO network for fine-grained hierarchical classification of laryngeal lesions in NBI endoscopy.
Leveraging Pathology Co-occurrence for Test-Time Adaptation in Chest X-Ray Diagnosis
Jeong, Woojin (Seoul National University), Lee, Jaewook (Seoul National University)
ClassificationDomain AdaptationConvolutional Neural NetworkSupervised Fine-TuningContrastive LearningImageBiomedical DataComputed Tomography
🎯 What it does: This study investigates a method that utilizes disease co-occurrence patterns for test-time adaptation in multi-label classification tasks on chest X-rays, aiming to alleviate performance degradation caused by domain shifts between different hospitals.
LinGuinE: Longitudinal Guidance Estimation for Volumetric Tumour Segmentation
Garibli, Nadine (AstraZeneca), Patwari, Mayank (AstraZeneca)
Object TrackingSegmentationConvolutional Neural NetworkTransformerPrompt EngineeringDiffusion modelOptical FlowImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Proposed the LinGuinE framework, which utilizes prompts from radiologists at a single time point to achieve longitudinal tumor voxel segmentation and tracking through image registration and prompt segmentation models.
LiSAD: A Label-Efficient Implicit Model with Spatially Adaptive Directional Total Variation for Spatial Transcriptomics
Xia, Yinghao (Northwestern Polytechnical University), Jin, Qiangguo (Northwestern Polytechnical University)
SegmentationRepresentation LearningData-Centric LearningGraph Neural NetworkNeural Radiance FieldAuto EncoderGenerative Adversarial NetworkContrastive LearningBiomedical DataBenchmark
🎯 What it does: Propose the LiSAD model, which achieves continuous representation of spatial transcriptomics through implicit neural representation (INR) and hypergraph encoding, and performs spatial domain partitioning based on this
LLM-Bootstrapped Targeted Finding Guidance for Factual MLLM-based Medical Report Generation
Yang, Cunyuan (Zhejiang University), Wang, Haishuai (Zhejiang University)
ClassificationGenerationExplainability and InterpretabilityData-Centric LearningConvolutional Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelImageTextMultimodalityBiomedical DataComputed TomographyElectronic Health Records
🎯 What it does: Proposed the Fact-Flow framework, which first uses a multi-label classifier to identify clinical findings in images, and then uses these findings as conditions to guide the MLLM to generate medical reports, thereby significantly improving the factual accuracy of the reports.
LLM-Steered Clinical Knowledge Discovery for CVD Outcome Modeling
Liang, Yuxuan (Rensselaer Polytechnic Institute), Yan, Pingkun (Rensselaer Polytechnic Institute)
Explainability and InterpretabilityDrug DiscoveryGraph Neural NetworkTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTabularBiomedical DataElectronic Health Records
🎯 What it does: This paper proposes the EGL-LLM framework, which is based on evidence-driven learning (EGL) and LLM guidance, to discover auditable cardiovascular disease (CVD) knowledge graphs from clinical and imaging variables.
Localization-Grounded Supervision: Revisiting Vanilla SFT of Large Vision-Language Models for Medical Image Analysis
Shi, Yiming (Tsinghua University), Wu, Ji (Tsinghua University)
Explainability and InterpretabilityTransformerSupervised Fine-TuningVision Language ModelImageBiomedical DataComputed TomographyReview/Survey Paper
🎯 What it does: This paper systematically analyzes the supervised fine-tuning of large vision-language models in medical image analysis, and proposes a location-based supervision mechanism called LGS to accelerate the alignment of fine-grained semantics and spatial information.
LoFi: Location-Aware Fine-Grained Representation Learning for Chest X-Ray
Kang, Myeongkyun (University of British Columbia), Li, Xiaoxiao (University of British Columbia)
RetrievalRepresentation LearningTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataComputed TomographyElectronic Health RecordsRetrieval-Augmented Generation
🎯 What it does: Propose the LoFi method, jointly optimizing sigmoid, captioning, and position-aware captioning losses, leveraging a lightweight LLM to learn fine-grained representations of chest X-rays, and using them for retrieval-based context learning to achieve precise localization.
LoHi-Net: A Low-to-High Hierarchical Wavelet Network for Laparoscopic Surgical Smoke Removal
Wu, Shiwei, Yan, Hui (Nanjing University of Science and Technology)
RestorationConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageVideoBenchmark
🎯 What it does: Proposed LoHi-Net, a low-to-high frequency network based on wavelet transform, for smoke removal in laparoscopic surgery.
Longitudinal Multi-view Modeling for Breast Cancer Risk Prediction
Thrun, Solveig (UiT Arctic University of Norway), Kampffmeyer, Michael (UiT Arctic University of Norway)
ClassificationAnomaly DetectionRepresentation LearningConvolutional Neural NetworkTransformerContrastive LearningImageBiomedical DataComputed TomographyReview/Survey Paper
🎯 What it does: Propose LMV-Net, which jointly utilizes longitudinal mammography images from the CC and MLO views for breast cancer risk prediction, achieving explicit feature space alignment and dual-stream attention fusion.
Longitudinal Reliability of Cross-Sectional Normative Models Under Measurement Error
Ji, Wenxing (Newcastle University), Wang, Yujiang (Newcastle University)
Explainability and InterpretabilityComputational EfficiencyBiomedical DataMagnetic Resonance ImagingReview/Survey Paper
🎯 What it does: Proposed and validated the Longitudinal Deviation Index (LDI), used to measure the reliability of individual residual changes over time compared to measurement error when using cross-sectional normative models across time.
LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making
Chen, Yanzhen (Zhejiang University), Liu, Zuozhu (Zhejiang University)
TransformerLarge Language ModelAgentic AITabularTime SeriesBiomedical DataElectronic Health RecordsBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose LongMedBench, a long-term EHR benchmark based on MIMIC-IV, for evaluating the performance of medical LLM agents in long-term clinical decision-making across multiple visits.
Lost in the Folds: When Cross-Validation Is Not a Deep Ensemble for Uncertainty Estimation
Kirscher, Tristan (University of Strasbourg), Maier-Hein, Klaus (German Cancer Research Center)
SegmentationConvolutional Neural NetworkImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: This paper investigates the impact of two ensemble methods, cross-validation (CV) and deep ensembles (DE), on uncertainty estimation in medical image segmentation.
Lost in Volume: The CT-SpatialVQA Benchmark for Evaluating Semantic-Spatial Understanding of 3D Medical Vision–Language Models
Monon, Mashrafi (Mohamed Bin Zayed University of Artificial Intelligence), Yaqub, Mohammad (Mohamed Bin Zayed University of Artificial Intelligence)
Explainability and InterpretabilityTransformerLarge Language ModelVision Language ModelContrastive LearningImageTextBiomedical DataComputed TomographyBenchmark
🎯 What it does: Proposed the CT-SpatialVQA benchmark for systematically evaluating the semantic spatial reasoning capabilities of 3D medical vision-language models on volumetric CT images.
LOT-Bridge: A Latent Optimal Transport Bridge Based on Signed Distance Fields for Automatic Skull Defect Reconstruction
Yan, Yanghui, Shui, Wuyang (Shanghai Jiao Tong University)
RestorationGenerationConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningPoint CloudMeshBiomedical DataMagnetic Resonance ImagingComputed TomographyOrdinary Differential Equation
🎯 What it does: Propose a Latent Optimal Transport Bridge (LOT-Bridge) method based on Signed Distance Field (SDF) for automatic cranial defect reconstruction, reducing the number of generation iterations and improving reconstruction quality.
Low-Rank Text-Guided Spectral Learning for Semi-supervised MHSI Segmentation
Zhang, Siqi (East China Normal University), Li, Qingli (East China Normal University)
SegmentationData-Centric LearningConvolutional Neural NetworkTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposes a low-rank text-guided spectral learning network called LoTS-Net for segmenting microscopic hyperspectral images (MHSI) under semi-supervised conditions.
Low-Rank Velocity Fields as a Structural Prior for Unsupervised 4D Medical Image Interpolation
Li, Haojin, Liu, Jiang (Southern University of Science and Technology)
RestorationDiffusion modelScore-based ModelAuto EncoderContrastive LearningOptical FlowBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: A low-rank velocity field structural prior is proposed under endpoint unsupervised conditions for 4D medical image interpolation.
Low-Rank-Modulated Functa: Exploring the Latent Space of Implicit Neural Representations for Interpretable Ultrasound Video Analysis
Wolleb, Julia (Yale University), Papademetris, Xenophon (Yale University)
CompressionExplainability and InterpretabilityNeural Radiance FieldAuto EncoderOptical FlowVideoBiomedical DataUltrasound
🎯 What it does: Propose the Low-Rank Modulated Functa network (LRM-Functa), achieving compression and interpretability analysis of ultrasound videos.
Low-Resource Cross-Modality Nucleus Detection via One-Sided Domain Adaptation
Xing, Fuyong (University of Colorado Anschutz Medical Campus), Ren, Siyang (University of Colorado Anschutz Medical Campus)
Object DetectionDomain AdaptationDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageMultimodalityBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: This study proposes an unsupervised domain adaptation framework combining one-sided GAN and mask contrastive learning (MCL) for cell nucleus detection in low-resource cross-modal microscopy images;