🎯 What it does: Leverage the large amount of available imaging reports, longitudinal scans, and multi-phase CT from hospitals to build a teacher-student self-distillation architecture, learning multi-tumor (esophagus, spleen, uterus) segmentation without relying on a large number of manual masks;
S-CEReBrO: Breaking the Memory Barrier in Continuous EEG Monitoring
Bucagu, Glenn Anta (ETH Zurich), Benini, Luca (University of Bologna)
CodeAnomaly DetectionComputational EfficiencyRepresentation LearningTransformerAuto EncoderContrastive LearningTime SeriesBiomedical Data
🎯 What it does: Proposed S-CEReBrO, a Transformer-based continuous EEG foundation model that utilizes windowed alternating attention to achieve constant KV cache memory, and was pre-trained with masked autoencoding on 25,000 hours of EEG corpus.
S-GRPO: Structural Group Relative Policy Optimization for Medical Image Segmentation
Duan, Yiru (Nanjing University of Science and Technology), Chen, Qiang (Nanjing University of Science and Technology)
CodeSegmentationKnowledge DistillationConvolutional Neural NetworkTransformerReinforcement LearningMixture of ExpertsImageBiomedical DataMagnetic Resonance Imaging
🎯 What it does: This paper proposes a structured parallel sampling reinforcement learning framework called S-GRPO, aimed at improving pixel accuracy and global anatomical consistency in medical image segmentation.
🎯 What it does: Propose the Structure-Aware Continuous Flow Matching (SA-CFM) model, which provides a standardized modeling approach for EEG microstate potentials, generating reproducible individual deviation representations.
🎯 What it does: Propose SAFE-Diff, which synthesizes high-fidelity contrast-enhanced images on non-contrast breast MRI using a multi-scale attention diffusion model.
🎯 What it does: Propose a two-stage CT super-resolution framework called SAFE-Diff: the first stage uses a residual prediction network to recover structure; the second stage uses truncated diffusion to refine details, and retains low-frequency structure and high-frequency details through SWT/ISWT frequency domain fusion.
CodeExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelReinforcement LearningMultimodalityBiomedical DataElectronic Health RecordsRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose a self-evolving LLM agent called SAGEAgent for cost-aware sequential modal acquisition in multimodal survival prediction.
SAM2 as a Key Bridge Between 2D Slices and 3D Volumes for Sparsely-supervised Medical Image Segmentation
Yu, Peng (East China Normal University), Wang, Yan (East China Normal University)
CodeSegmentationTransformerVision Language ModelAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: This paper studies how to achieve complete 3D medical image segmentation using only three orthogonal central 2D slices with annotations, leveraging SAM2 for sparse annotation scenarios, and proposes a unified framework that integrates multi-view propagation, cross-view consistency weighted fusion, and dynamic pseudo-label refinement.
🎯 What it does: Proposed a Sample-Agnostic Retrieval Network (SAR-Net) that completes missing modalities and performs brain tumor segmentation through a retrieval-based approach.
🎯 What it does: Proposes a structure consensus-based KAN prototype learning framework, SCKAN, to address the problem of supervisory bias in pancreatic segmentation.
🎯 What it does: This work proposes Consistency Memory Bank (COMB), a framework for full-slide label-free virtual staining, aiming to eliminate artifacts such as stitching and color drift caused by block-based inference.
🎯 What it does: Developed SeedPro, an automatic radioactive seed pre-planning framework based on hierarchical reinforcement learning, capable of generating radioactive seed layout plans comparable to those of experts.
🎯 What it does: This study proposes the SegDINO framework, which transfers the pre-trained DINOv3 visual model to medical image segmentation tasks, achieving efficient segmentation through lightweight multi-scale modeling;
🎯 What it does: Propose a single-source domain generalization medical image segmentation framework called SSUG based on the Segment Anything Model (SAM), which utilizes three modules: Spectral Manifold Extrapolation, Uncertainty-aware Symmetric Channel Gating, and Hierarchical Decoupled Consistency, to jointly suppress style sensitivity and extract domain-invariant anatomical features.
🎯 What it does: Propose Selective Rank-1 Orthogonality Regularization (SR1OR), which addresses the stability-plasticity balance problem in medical image continual learning by imposing selective orthogonality constraints on unit rank-1 subspaces within the LoRA low-rank update space.
🎯 What it does: Propose SEA-PEFT, a parameter-efficient fine-tuning method that automatically configures adapters through an online self-audit process, specifically designed for few-shot 3D medical image segmentation.
🎯 What it does: Proposes a physics-informed implicit neural representation framework called Lorentz Encoding (LE), for self-supervised reconstruction of high-resolution CEST Z spectra under sparse sampling, and generates continuous, interpretable metabolic maps through physical constraints.
Self-supervised Learning on Lingual Ultrasound Video Encodes Clinically Meaningful Articulatory Structure of Rhotic Production in Residual Speech Sound Disorder
🎯 What it does: This paper proposes a joint embedding framework based on self-supervised learning, which learns the tongue position structure of American English rhotic pronunciation using tongue ultrasound videos, and reduces speaker differences through speaker adversarial regularization.
🎯 What it does: This study proposes a self-supervised Spline Autoencoder, which learns low-rank latent trajectories of cardiac motion through smooth curve fitting, and achieves unlabeled cardiac cycle and ED/ES event localization by utilizing spectral embedding and phase projection.
🎯 What it does: Building upon the previously learned implicit correspondence cardiac graph network Mask‑HybridGNet Dual, this study introduces self-supervised temporal regularization (velocity and acceleration penalty) as a post-training phase to enhance temporal consistency in ultrasound sequence segmentation, and utilizes the learned correspondence to achieve automatic AHA 17-segment standard mapping, completing regional motion analysis.
🎯 What it does: Introduce a semantic feature modulation framework based on BI-RADS descriptors in breast X-ray image classification to achieve dynamic fusion of images and clinical semantics.
🎯 What it does: This study proposes a self-supervised hierarchical flow model for estimating corresponding quantitative T1, T2, and PD maps from T1, T2, and PD-weighted MRI images, and for synthesizing new weighted images using the MR signal model.
🎯 What it does: Propose a probabilistic representation framework that models the task representation of each modality subset as a Gaussian distribution, utilizing the set inclusion hierarchy to guide mean alignment and variance to reflect information missing, thus achieving brain tumor segmentation under missing modalities;
SF-DINO: Spatial-Frequency Adapted Foundation Model for Multiple Myeloma Diagnosis
Ye, Zhaoyi (Wuhan University), Lei, Cheng (Wuhan University)
CodeClassificationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningDrug DiscoveryConvolutional Neural NetworkTransformerPrompt EngineeringContrastive LearningImageBiomedical Data
🎯 What it does: Propose the SF-DINO model, which uses a spatial morphology adapter and frequency-selective hash attention for fine-grained classification of multiple myeloma cells;
🎯 What it does: Propose an interactive cortical sulcus annotation framework based on shape-adaptive guiding signals, using spherical CNN to perform binary segmentation of the left posterior frontal sulcus.
ShapeEFM: A Shape-Based Pretrained Foundation Model on ECG Morphology Comprehension
Liu, Huan (Ant Group), Lu, Le (Zhejiang University School of Medicine)
CodeExplainability and InterpretabilityRepresentation LearningTransformerSupervised Fine-TuningAuto EncoderContrastive LearningTime SeriesBiomedical DataElectrocardiogram
🎯 What it does: Propose ShapeEFM, a shape-based pre-trained ECG foundation model, which decomposes ECG waveforms into Cardio-Morphemes, treating them as 'words,' and uses Transformer for self-supervised learning to enhance ECG morphology understanding and interpretability.
ShapKO: Shapley-Adaptive Modality Knockout for Robust Multimodal Learning
Nizam, Nusrat Binta (Cornell University), Sabuncu, Mert R. (Weill Cornell Medicine)
CodeExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerVision Language ModelAuto EncoderContrastive LearningImageTextMultimodalityTabularBiomedical DataElectronic Health Records
🎯 What it does: By dynamically estimating modal importance using Shapley values during training, adaptively adjusting the elimination probability for each modality, thereby enhancing the robustness of multi-modal models under missing modality scenarios.
🎯 What it does: Propose a shared Gaussian geometric framework to achieve spatial resolution harmonization of multi-contrast brain MRI without the need for external training sets.
🎯 What it does: A continuous function learning framework based on sparse LGE-MRI was constructed to generate a three-dimensional digital twin of myocardial scar.
Single-Stage Hierarchical Rectification for Weakly Supervised Histopathology Segmentation
Nguyen, Trong Duc (VinUniversity), Pham, Huy Hieu (Posts and Telecommunications Institute of Technology)
CodeSegmentationConvolutional Neural NetworkRectified FlowContrastive LearningBiomedical Data
🎯 What it does: Propose a single-stage hierarchical correction framework, SSHR, which achieves weakly supervised histopathological semantic segmentation by correcting intermediate features during the forward pass, avoiding multi-stage CAM generation and recursive pseudo-labeling.
🎯 What it does: A single-organ multi-view MRI super-resolution framework based on implicit neural representations was studied, which can reconstruct isotropic high-resolution images from multi-directional anisotropic scans without the need for pre-registration.
SIRA: Reasoning-Aware Surgical Instrument Segmentation via Query-Anchored Alignment
Zhang, Zhibo (Huazhong University of Science and Technology), Yan, Zengqiang (Huazhong University of Science and Technology)
CodeSegmentationTransformerLarge Language ModelVision Language ModelContrastive LearningImageTextMultimodalityBiomedical DataBenchmark
🎯 What it does: Proposed a query-conditioned surgical instrument segmentation framework called SIRA, achieving multi-modal semantic alignment segmentation tasks on the newly constructed SurgRS dataset.
Situational context embedding and knowledge-based grounding for LLMs in surgical training assistance
Bahari Malayeri, Ali (University of Zurich), Fürnstahl, Philipp (Zurich University of Applied Sciences)
CodeExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningRobotic IntelligenceDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringTextBiomedical DataReview/Survey PaperRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposed and implemented a tutoring system for total hip arthroplasty (THA) training simulator based on large language models, which provides context-aware answers by embedding real-time situational information from the simulator and combining it with knowledge base retrieval.
CodeClassificationExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelVision Language ModelContrastive LearningImageTextMultimodalityElectronic Health Records
🎯 What it does: Proposed and implemented Skin2Mind, a multimodal framework based on multispectral facial skin images and structured skin reports, for predicting individual-level DSM-5 level mental health risks.
SLICE-MIL: Counterfactual Semantic Anchoring of Disentangled Evidence for Child-Pugh Grading
Deng, Yihan (University of Science and Technology of China), Zheng, Jian (University of Science and Technology of China)
CodeClassificationExplainability and InterpretabilityTransformerContrastive LearningImageBiomedical DataComputed Tomography
🎯 What it does: Proposed a SLICE-MIL framework based on weakly supervised multiple instance learning, using anatomy-guided dual-query attention and adversarial semantic anchoring to estimate Child-Pugh grades from portal-phase CT imaging.
🎯 What it does: This paper reformulates the training of nnU-Net as a proximal optimization problem with a hybrid L1+Group L2 sparse regularization, and achieves extremely sparse Slim nnU-Net by performing channel-wise pruning during training.
🎯 What it does: Implementing an autism spectrum disorder screening based on eye movement and head pose on smartphones, proposing the AMF-Net framework.
SonoCLIP: Mask-Guided Region-Aware Vision–Language Pretraining for Fetal Ultrasound Analysis
Su, Hang (Wuhan University), Du, Bo (Wuhan University)
CodeRecognitionSegmentationTransformerVision Language ModelContrastive LearningImageTextBiomedical DataUltrasound
🎯 What it does: This paper proposes SonoCLIP, a fetal ultrasound foundation model based on vision-language alignment, which can achieve region-controllable representation learning at both global and local levels.
Sparse Autoencoders for Interpretable Medical Image Representation Learning
Wesp, Philipp (Stanford University), Gatidis, Sergios (Stanford University)
CodeRetrievalExplainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelVision Language ModelAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: This paper proposes using a sparse autoencoder (Matryoshka SAE + BatchTopK) to transform dense embeddings from medical imaging foundational models (BiomedParse, DINOv3) into interpretable sparse features, and validates their interpretability and performance through various retrieval and evaluation methods;
🎯 What it does: Propose the SPEAR framework, which achieves high-resolution 3D medical image deformation registration through error-guided sparse refinement.
Spectral-Semantic Consensus for Few-Shot Whole Slide Image Classification
Du, Hanyu (Wuhan University), Xu, Yongchao (Wuhan University)
CodeClassificationAnomaly DetectionExplainability and InterpretabilityComputational EfficiencyRepresentation LearningMeta LearningTransformerPrompt EngineeringVision Language ModelAuto EncoderContrastive LearningImageTextBiomedical DataBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposes a Spectral-Semantic Consensus (SSC) framework based on text priors for few-shot whole-slide image classification, automatically selecting diagnostic information-rich patches and aggregating them.
🎯 What it does: This study proposes a three-stage multimodal learning framework for precise segmentation of vocal tract organs using speech-guided real-time MRI.
🎯 What it does: Propose the Spinverse method, which optimizes surface permeability on a fixed tetrahedral mesh using a differentiable Bloch-Torrey finite element simulator to inversely reconstruct the microstructural boundaries corresponding to diffusion MRI signals.
SRMA-Mamba: Spatial Reverse Mamba Attention Network for Pathological Liver Segmentation in MRI Volumes
Zeng, Jun, Jha, Debesh (University Of South Dakota)
CodeSegmentationConvolutional Neural NetworkTransformerMixture of ExpertsDiffusion modelContrastive LearningImageBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Proposed and implemented a three-dimensional convolution + Mamba network called SRMA-Mamba for precise segmentation of liver cirrhosis lesions in magnetic resonance imaging (MRI) volumes.
🎯 What it does: Proposed a Spiking Serialized Point Transformer (SSPT) based on spiking neural networks for Couinaud segmentation of 3D medical point clouds.
🎯 What it does: This paper proposes a unified Train-Inference-Evaluation framework specifically designed to enhance the temporal stability of models for surgical phase recognition (SPR) during the online surgery phase, avoiding the cascading errors caused by short-term jitter and misjudgment;
🎯 What it does: This study proposes an end-to-end framework called StrokeTimer, which can automatically estimate the onset time window of ischemic stroke in non-enhanced CT images (<4.5 h, 4.5–6 h, >6 h).
Structural Congruence Matters: Biological Pretraining for Vascular Graph Extraction
Scavone, Alessandro (EURECOM), Zuluaga, Maria A. (EURECOM)
CodeImage TranslationSegmentationDomain AdaptationGraph Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningImageGraphBiomedical DataAgriculture Related
🎯 What it does: A 'nature-to-nature' pre-training paradigm is proposed, using the structural pre-training of plant (grapevine) skeleton graphs for end-to-end extraction of medical images to vascular graph structures.
🎯 What it does: A sparse voxel space diffusion framework is proposed for denoising and super-resolution of 3D medical images. By directly predicting clean images in the voxel space, incorporating velocity supervision, sparse time step scheduling, and structure-aware trajectory modulation, the method achieves results with only 5 time steps, a 10-fold increase in training speed, while maintaining fine details and structural integrity.
🎯 What it does: Propose the STEM model, combining self-supervised contrastive learning with metric-based meta-learning, explicitly decoupling EEG features into subject-related and task-related components, thereby fully utilizing subject identity and task label information during the pre-training phase.
🎯 What it does: Proposed a graph convolutional framework based on surface point clouds, named SurfMark3D, for anatomical landmark localization of the distal femur during total knee arthroplasty.
Liu, Qixuan (Chinese University of Hong Kong), Heng, Pheng-Ann (Chinese University of Hong Kong)
CodeRecognitionSegmentationTransformerSupervised Fine-TuningPrompt EngineeringVision Language ModelVideoMultimodalityBenchmark
🎯 What it does: Proposed the SurgVTG task for temporal localization in surgical videos, constructed the HMSSurgVTG benchmark, and provided corresponding annotations.
SurgOnAir: Hierarchy-Aware Real-Time Surgical Video Commentary
He, Jingyi (TU Munich), Bi, Yuan (TU Munich)
CodeGenerationComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringVision Language ModelVideoTextRetrieval-Augmented GenerationAudio
🎯 What it does: Propose the SurgOnAir model to achieve real-time hierarchical narration of surgical videos.
🎯 What it does: Designed and implemented a two-stage synthesis framework called SurgRFO, which first generates surgery chest X-ray backgrounds without RFO using a latent diffusion model, and then samples local RFOs with a lightweight generator and generates realistic composite images through conditional Poisson fusion, used for data augmentation to improve detection performance.
SurgVLA-Bench: Towards Evaluating Vision-Language-Action Models for Laparoscopic Surgical Robotics
Sun, Jiashuo (National Engineering Research Center of Robot Visual Perception and Control Technology), Liu, Min (National Engineering Research Center of Robot Visual Perception and Control Technology)
🎯 What it does: Propose SurgVLA-Bench, constructing a surgical vision-language-action (VLA) evaluation benchmark based on the SurRoL simulation platform, which includes a hierarchical task system and corresponding datasets ranging from atomic actions to complete surgical procedures.
SurgZSD: Zero-Shot Surgical Action Triplet Detection via Attribute Composition and Clinical Feasibility Inference
Deng, Zuxing (Hefei University of Technology), Zhou, Wenrui (Hefei University of Technology)
CodeObject DetectionTransformerSupervised Fine-TuningPrompt EngineeringVision Language ModelImageBiomedical DataBenchmarkRetrieval-Augmented Generation
🎯 What it does: Developed the SurgZSD zero-shot surgical action triple detection framework, which utilizes medical prior knowledge to achieve joint reasoning of attribute combination and clinical feasibility constraints.
SutureFormer: Learning Surgical Trajectories via Goal-Conditioned Offline RL in Pixel Space
Liu, Huanrong (University of Macau), Li, Qingbiao (University of Macau)
CodeRobotic IntelligenceConvolutional Neural NetworkTransformerReinforcement LearningDiffusion modelVideoBiomedical Data
🎯 What it does: Learn a surgical suturing trajectory prediction model called SutureFormer in the pixel space through offline reinforcement learning, enabling the prediction of future needle tip trajectories based solely on endoscopic videos and sparse keyframes.
🎯 What it does: This study proposes an end-to-end SwallowReg framework for joint segmentation and registration in dynamic Cine-MRI of patients with tongue cancer after head and neck resection, thereby enabling quantitative evaluation of swallowing function.
Synergistic Perception-Reasoning Governance: Grounding Medical MLLMs with Verifiable Anatomical Evidence
Hao, Rui, Zeng, Zhigang (Huazhong University Of Science And Technology)
CodeExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextBiomedical DataMagnetic Resonance ImagingComputed TomographyElectronic Health RecordsRetrieval-Augmented Generation
🎯 What it does: A training-agnostic, anatomy ROI evidence injection-based multimodal large model hallucination suppression framework is proposed, combining visual activation re-adjustment with text structural injection, and introducing a task-aware dynamic router.
SynHydro: Biomechanics-Driven Domain Randomization for Hydrocephalus-Agnostic, Generalizable Tissue-Ventricle Segmentation Across Age, Modality, and Resolution
🎯 What it does: Simulate hydrocephalus using biomechanically driven deformations based on healthy brain tissue labels, then generate diverse multimodal, low-resolution, noisy, and artifact-containing MRI data through domain randomization, and train nnU-Net to achieve joint segmentation of brain tissue and ventricles in pediatric hydrocephalus patients.
SYNPRED: A Synergistic Approach to Multimodal Learning for Clinical Prediction
Janíčková, Ivana (Medical University of Vienna), Langs, Georg (Medical University of Vienna)
CodeClassificationExplainability and InterpretabilityRepresentation LearningData-Centric LearningTransformerSupervised Fine-TuningMixture of ExpertsAuto EncoderGenerative Adversarial NetworkContrastive LearningImageMultimodalityTabularBiomedical DataMagnetic Resonance ImagingElectronic Health Records
🎯 What it does: Propose a multi-modal variational autoencoder (VAE) framework called SYNPRED, which integrates medical imaging and RNA sequencing data, to enable clinical prediction when multi-modal data are missing.
🎯 What it does: Propose TAIR, a text-guided adaptive prompt refinement framework, achieving full-scenario medical image restoration in coarse-to-fine stages.
Task-Performance Routing for Multi-teacher Distillation in CT Medical Image Segmentation
Yang, Jiaye, Wang, Peng (University Of Electronic Science And Technology Of China)
CodeSegmentationKnowledge DistillationTransformerMixture of ExpertsBiomedical DataComputed Tomography
🎯 What it does: CT organ segmentation based on multi-teacher knowledge distillation, proposing the Task-Performance Routing (TPR) framework, which dynamically selects the best teacher and fuses knowledge for each region in space.
CodeGenerationAnomaly DetectionExplainability and InterpretabilityConvolutional Neural NetworkTransformerLarge Language ModelVision Language ModelContrastive LearningImageTextMultimodalityTabularBiomedical DataComputed TomographyUltrasoundElectronic Health RecordsRetrieval-Augmented Generation
🎯 What it does: This paper proposes a framework named TAVR‑VLM for generating preoperative reports for transcatheter aortic valve replacement (TAVR), and suppresses diagnostic hallucinations through risk-conditioned causal grounding.
🎯 What it does: Proposed the TCCT framework, which utilizes the orbit-conditioned spectral neural operator to predict redundant weights and embeds them into a differentiable Shift-Variant FBP, achieving fast CBCT reconstruction for sinusoidal non-circular orbits.
🎯 What it does: Developed TCFD-Net based on dual-time-point CT images to predict early immunotherapy response in patients with non-small cell lung cancer (NSCLC).
🎯 What it does: Proposed the TEDi framework, combining memory search enhancement and temporal consistency denoising Transformer to improve surgical instrument segmentation, addressing cross-frame semantic consistency and class confusion issues.
Temporal Phase-Difference Guided Spatiotemporal Learning for DSA Vessel Segmentation
Liu, Kun (Beijing University of Posts and Telecommunications), Yang, Huihua (Beijing University of Posts and Telecommunications)
CodeSegmentationConvolutional Neural NetworkRecurrent Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningVideoBiomedical Data
🎯 What it does: Propose a spatiotemporal segmentation framework combining forward phase difference projection maximum intensity projection (FPD‑MaxIP) and a Mamba-based temporal encoder for precise segmentation of vessels in digital subtraction angiography (DSA).
Test-Time Adaptation for Rare Surgical Phase Recognition: Bridging the Coverage-Gap Paradox
Park, Ho-min (Ghent University Global Campus), Vankerschaver, Joris (Ghent University Hospital)
CodeRecognitionDomain AdaptationRecurrent Neural NetworkTransformerSupervised Fine-TuningContrastive LearningVideoBiomedical Data
🎯 What it does: This paper addresses the coverage gap issue in the identification of rare surgical phases by proposing a lightweight test-time adaptation framework, which significantly improves the accuracy of rare phase identification through adaptive threshold pseudo-label updating and temporal smoothing.
🎯 What it does: Proposes a test-time adaptation framework called TTA-Flow based on flow matching and histogram alignment, which transforms noisy images from low-cost OCT devices into high-quality training domain images.
Text as Illumination: Spatial Contrastive Retinex Learning for Language-Guided Medical Image Segmentation
Shi, Jian (Dalian University of Technology), Lu, Huchuan (Dalian University of Technology)
CodeSegmentationTransformerVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Propose a Retinex-inspired network called TIRNet that utilizes text embeddings as semantic illumination to achieve language-guided medical image segmentation.
Text-Deficient Multimodal Stroke Segmentation with Lesion-Grounded Self-Retrieval-Augmented Generation
Eum, Heeseong (Seoul National University), Choi, Kyu Sung (Seoul National University)
CodeSegmentationGenerationRetrievalConvolutional Neural NetworkTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextBiomedical DataMagnetic Resonance ImagingRetrieval-Augmented Generation
🎯 What it does: The study proposes LeG-RAG, a lesion-aligned retrieval-based generation method for missing report generation in the text-missing scenario of acute ischemic stroke MRI, and implements report-conditioned multimodal segmentation using LLMSwin.
CodeSegmentationTransformerPrompt EngineeringVision Language ModelDiffusion modelScore-based ModelAuto EncoderContrastive LearningImageTextBiomedical Data
🎯 What it does: Proposed a text-guided multi-frequency latent diffusion framework (TMF-Seg) for medical image segmentation.
🎯 What it does: Propose TGBD-Net for jointly predicting EGFR mutation status and progression-free survival (PFS) in lung cancer CT images, achieving stable multi-task learning in a heterogeneous supervision environment.
🎯 What it does: Propose a Template-Guided Heteroscedastic Diffusion Bridge (TGH-DB) based on population-level amyloid PET templates, which can generate structured and physiologically reasonable PET images from multi-modal MRI.
Dombrowski, Mischa (FAU Erlangen Nurnberg), Kainz, Bernhard (FAU Erlangen Nurnberg)
CodeClassificationGenerationKnowledge DistillationRepresentation LearningConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyElectronic Health Records
🎯 What it does: This paper systematically investigates the 'learnability gap' in potential diffusion models for medical image generation, and proposes a noise-conditioned latent classifier (FiLM layer + image space distillation) to diagnose and partially alleviate this gap.
The Paper Has a GitHub, the GitHub Has a README, the README Has Nothing: Reproducibility Signals for Review Support
Bolelli, Federico (University of Modena and Reggio Emilia), Grana, Costantino (University of Modena and Reggio Emilia)
CodeExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelPrompt EngineeringImageTextTabularReview/Survey PaperBenchmarkRetrieval-Augmented Generation
🎯 What it does: This paper proposes an interpretable decision support tool called paper-snitch based on LLM, which is used in the peer review process of medical imaging papers to automatically parse PDFs, extract and statically check code repositories, and generate evidence-supported reproducibility assessment reports based on the MICCAI reproducibility policy.
Xiong, Liang (Chongqing Normal University), Cui, ShaoGuo (Chongqing Normal University)
CodeGenerationRetrievalTransformerPrompt EngineeringVision Language ModelImageTextBiomedical DataRetrieval-Augmented Generation
🎯 What it does: This paper proposes a dual-stream retrieval-enhanced pathological report generation framework, DR-Gen, for automatically generating pathological reports.
🎯 What it does: Propose the task of postmortem self-autopsy image restoration in forensic histopathology, construct the first homologous but unpaired AutoPath dataset, and use the Schrödinger Bridge generative model to achieve image restoration, followed by the design of an evaluation framework based on diagnostic distribution consistency.
🎯 What it does: Proposes a first-order deterministic generative framework JiR that only utilizes the temporal condition t for medical image translation (MR→PET, T1w→T1ce).
🎯 What it does: Designed a teeth alignment framework based on virtual progressive trajectories, utilizing pre-treatment and post-treatment oral scan data to progressively predict K-step SE(3) transformations, and applying clinical constraints at each step to simulate the incremental process of real orthodontic treatment.
🎯 What it does: Proposed a large-scale CBCT facial bone and tooth segmentation dataset called ToothFairy3 (77 classes, 582 scans) and designed an efficient U-Mamba2 architecture.
🎯 What it does: A topology-based transferability estimation framework is proposed to predict the performance of models in medical image segmentation tasks without fine-tuning.
🎯 What it does: Propose TopoOR, which uses composite complexes and high-order attention networks to uniformly model and reason about multi-modal entities and multi-agent relationships in the operating room.
TotalFM: An Organ-Separated 3D-CT Foundation Model Leveraging Large-Scale Routine Clinical Radiology Data
Yamamoto, Kohei (Jichi Medical University), Kikuchi, Tomohiro (Jichi Medical University)
CodeSegmentationAnomaly DetectionConvolutional Neural NetworkTransformerLarge Language ModelVision Language ModelContrastive LearningImageTextBiomedical DataComputed Tomography
🎯 What it does: Propose TotalFM, a 3D-CT foundation model that utilizes large-scale clinical CT data, through automated conversion of raw CT images and reports into organ-level image-text pairs for self-supervised training.
🎯 What it does: Proposes a collaborative learning framework based on vascular skeletons to achieve joint segmentation of liver vessels and Couinaud segments.
Toward Thyroid Cancer BRAF V600E Mutation Prediction via Multimodal Large Language Model: Dataset and Model Development
Hao, Pengfei, Zhu, Lei (Hong Kong University Of Science And Technology (Guangzhou))
CodeClassificationData-Centric LearningDrug DiscoveryTransformerLarge Language ModelSupervised Fine-TuningVision Language ModelContrastive LearningImageTextMultimodalityBiomedical DataUltrasoundElectronic Health Records
🎯 What it does: This paper constructs the first multi-modal thyroid cancer BRAF V600E mutation prediction dataset, MM-BRAF, and proposes a BRAF-MLLM framework based on a multi-modal large language model to achieve mutation prediction.
🎯 What it does: This paper proposes an anatomy-aware geometric transformer (AAGT) to achieve automatic high-precision registration between CBCT and IOS.
🎯 What it does: Proposes a unified multi-task framework called UIV-US, which can simultaneously perform image and video segmentation and classification tasks under B-mode and contrast-enhanced ultrasound (CEUS);