π― What it does: A hierarchical multi-node multi-GPU framework is proposed for respiratory motion-resolved reconstruction in three-dimensional non-Cartesian multi-echo gradient echo (mGRE) magnetic resonance imaging.
Large-Small Model Collaboration for Zero-Shot Surgical Phase Recognition
Zhang, Yiyi (Chinese University of Hong Kong), Heng, Pheng-Ann (Chinese University of Hong Kong)
CodeRecognitionConvolutional Neural NetworkTransformerPrompt EngineeringVision Language ModelContrastive LearningVideoText
π― What it does: Designed a large model-small model collaborative framework LaST to achieve zero-shot out-of-distribution surgery phase recognition, improving pseudo-label quality through iterative time refinement and cyclic replay.
π― What it does: To address the diversity uncertainty in medical volume segmentation, this paper proposes and implements a flow matching method in the latent space (L2L-Flow), which significantly improves inference speed while maintaining clinically relevant uncertainty estimation.
π― What it does: This paper proposes a semi-supervised 2D/3D medical image segmentation framework called SKTC-Net based on structural knowledge topological consistency.
π― What it does: Proposed and implemented Dynamic Focal Attention (DFA) in pathological image semantic segmentation, directly encoding class difficulty by introducing learnable class-level bias into the transformer cross-attention mechanism;
Learning Directional Semantic Transitions for Longitudinal Chest X-Ray Analysis
Hu, Zhangfeng (Rensselaer Polytechnic Institute), Yan, Pingkun (Massachusetts General Hospital, Harvard Medical School)
CodeClassificationExplainability and InterpretabilityRepresentation LearningTransformerVision Language ModelContrastive LearningImageTextBiomedical DataComputed TomographyElectronic Health Records
π― What it does: Proposes a visual-language pre-training framework called ProTrans, which learns directional semantic transitions in chest X-rays over time to enable longitudinal CXR disease progression analysis.
Learning Diverse and Realistic Polyp Data via Style-Fused Diffusion Models
Han, Longfei (Beijing Technology and Business University), Li, Haisheng (Beijing Technology and Business University)
CodeSegmentationGenerationData SynthesisConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningImageBiomedical Data
π― What it does: Propose a conditional diffusion model based on style fusion (SFD-Polyp), which injects the style of real images through the Style Integration Module (SIM), and generates diverse and anatomically accurate mask shapes during the inference phase using Morphology-Preserving Mask Synthesis (MPMS), thus synthesizing diverse and realistic polyp image-mask pairs to enhance the dataset.
π― What it does: Propose a hierarchical quality-aware CT-to-MRI (DWI) synthesis framework called LGN. It first generates a draft MRI using a latent diffusion model, then predicts slice quality through a teacher-student network, and finally refines locally using cross-attention from high-quality neighboring slices, improving synthesis quality while maintaining three-dimensional anatomical continuity.
π― What it does: Proposed a dual-stream contrastive learning framework VB-GCL based on the vector bundle theory for disease diagnosis in fMRI neural networks
π― What it does: Propose the Text-guided Invariant Anatomical Learning (TeInL) framework, which leverages self-supervised pre-training on multi-sequence MRI to decouple anatomical structures from sequence styles, and enhances brain tissue and lesion segmentation performance through text-image alignment learning to capture pathology-related features.
π― What it does: This paper proposes an unsupervised framework called LatentAtlas, which jointly generates conditional brain templates and atlases. It aligns public templates and atlases through shared latent space learning of prototypes, enabling the synchronous generation of templates and atlases under specific conditions (e.g., age).
π― What it does: Achieve calibration-free bone density estimation by generating voxel-level QCT density maps directly from conventional CT using a supervised 3D image-to-image translation framework.
Learning Role-Conditioned Alignment for Medical Image Referring Segmentation
Li, Kun (University of Liverpool), Zheng, Yalin (Chinese Academy of Sciences)
CodeSegmentationConvolutional Neural NetworkTransformerVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingComputed Tomography
π― What it does: This paper proposes a role-conditioned alignment-based model for medical image referential segmentation (RCA), which achieves more accurate segmentation by explicitly separating textual descriptions into three categories: lesion semantics, quantity, and location, and then performing fine-grained alignment with visual features;
Learning the Hierarchical Organization in Brain Network for Brain Disorder Diagnosis
Tang, Jingfeng (Northeastern University), Zaiane, Osmar R. (University of Alberta)
CodeClassificationExplainability and InterpretabilityGraph Neural NetworkTransformerContrastive LearningGraphBiomedical DataAlzheimer's Disease
π― What it does: Propose a framework called BrainHO that learns the hierarchical organization of brain networks for the diagnosis of brain disorders (ASD, MDD);
π― What it does: Proposed a weakly supervised image quality transfer (IQT) framework, which first generates realistic distorted images from undistorted prostate DWI using quality prototype flow matching, and then uses these synthetic pairs to train a supervised correction network to achieve distortion correction for single-echo EPI DWI.
Learning to Read Where to Look: Disease-Aware VisionβLanguage Pretraining for 3D CT
Ging, Simon (University of Freiburg), Brox, Thomas (University of Freiburg)
CodeClassificationSegmentationRetrievalTransformerPrompt EngineeringVision Language ModelContrastive LearningGaussian SplattingBiomedical DataComputed Tomography
π― What it does: Built a 3D CT audio-visual language model called RadFinder based on 98k CT scan-report pairs, incorporating disease prompt supervision and slice-level localization supervision.
Learning to Segment Burns with Limited Labels: A Dual-Path Cascaded Network with Uncertainty-Guided Teacher Diffusion Exposure
Chauhan, Joohi (Motilal Nehru National Institute of Technology Allahabad), Singh, Prabal Pratap (Motilal Nehru National Institute of Technology Allahabad)
CodeSegmentationData SynthesisDomain AdaptationConvolutional Neural NetworkDiffusion modelContrastive LearningImageBiomedical Data
π― What it does: Propose a semi-supervised burn segmentation framework called UADA-MSPCNet, combining a dual-path cascade network, uncertainty-based teacher-student consistency learning, and diffusion generation enhancement on the teacher side.
π― What it does: This study proposes ShapeFuse, aiming to unify deformable shape representations and image texture representations into a shared latent space, and dynamically fuse shape and texture through bidirectional cross-modal temporal attention and adaptive gating for disease classification of cardiac MRI videos.
π― What it does: Propose the Prost-RL framework, which learns the attention regions of micro-ultrasound images through reinforcement learning. By combining a foundational model encoder-decoder, it achieves 'learn attention first, then decode,' thereby improving the detection and localization of prostate cancer under weakly supervised conditions.
Learning Where to Look: Pathologist-Inspired Multi-field-of-View Evidence Retrieval for Cell Type Classification in H&E
Yuan, Ruizhi (University of Pittsburgh), Chen, Wei (University of Pittsburgh)
CodeClassificationExplainability and InterpretabilityComputational EfficiencyTransformerMixture of ExpertsContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyRetrieval-Augmented Generation
π― What it does: Proposed a hybrid expert framework called TRACE, which classifies cell types in H&E slices by leveraging token-level evidence retrieval from multi-scale perspectives.
π― What it does: Proposed a dual-branch real-time detection framework that integrates prototype learning into the YOLO network for fine-grained hierarchical classification of laryngeal lesions in NBI endoscopy.
π― What it does: This study investigates a method that utilizes disease co-occurrence patterns for test-time adaptation in multi-label classification tasks on chest X-rays, aiming to alleviate performance degradation caused by domain shifts between different hospitals.
π― What it does: Proposed the LinGuinE framework, which utilizes prompts from radiologists at a single time point to achieve longitudinal tumor voxel segmentation and tracking through image registration and prompt segmentation models.
π― What it does: Propose the LiSAD model, which achieves continuous representation of spatial transcriptomics through implicit neural representation (INR) and hypergraph encoding, and performs spatial domain partitioning based on this
CodeClassificationGenerationExplainability and InterpretabilityData-Centric LearningConvolutional Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelImageTextMultimodalityBiomedical DataComputed TomographyElectronic Health Records
π― What it does: Proposed the Fact-Flow framework, which first uses a multi-label classifier to identify clinical findings in images, and then uses these findings as conditions to guide the MLLM to generate medical reports, thereby significantly improving the factual accuracy of the reports.
CodeExplainability and InterpretabilityDrug DiscoveryGraph Neural NetworkTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTabularBiomedical DataElectronic Health Records
π― What it does: This paper proposes the EGL-LLM framework, which is based on evidence-driven learning (EGL) and LLM guidance, to discover auditable cardiovascular disease (CVD) knowledge graphs from clinical and imaging variables.
Localization-Grounded Supervision: Revisiting Vanilla SFT of Large Vision-Language Models for Medical Image Analysis
Shi, Yiming (Tsinghua University), Wu, Ji (Tsinghua University)
CodeExplainability and InterpretabilityTransformerSupervised Fine-TuningVision Language ModelImageBiomedical DataComputed TomographyReview/Survey Paper
π― What it does: This paper systematically analyzes the supervised fine-tuning of large vision-language models in medical image analysis, and proposes a location-based supervision mechanism called LGS to accelerate the alignment of fine-grained semantics and spatial information.
LoFi: Location-Aware Fine-Grained Representation Learning for Chest X-Ray
Kang, Myeongkyun (University of British Columbia), Li, Xiaoxiao (University of British Columbia)
CodeRetrievalRepresentation LearningTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataComputed TomographyElectronic Health RecordsRetrieval-Augmented Generation
π― What it does: Propose the LoFi method, jointly optimizing sigmoid, captioning, and position-aware captioning losses, leveraging a lightweight LLM to learn fine-grained representations of chest X-rays, and using them for retrieval-based context learning to achieve precise localization.
π― What it does: This paper investigates the impact of two ensemble methods, cross-validation (CV) and deep ensembles (DE), on uncertainty estimation in medical image segmentation.
Lost in Volume: The CT-SpatialVQA Benchmark for Evaluating Semantic-Spatial Understanding of 3D Medical VisionβLanguage Models
Monon, Mashrafi (Mohamed Bin Zayed University of Artificial Intelligence), Yaqub, Mohammad (Mohamed Bin Zayed University of Artificial Intelligence)
CodeExplainability and InterpretabilityTransformerLarge Language ModelVision Language ModelContrastive LearningImageTextBiomedical DataComputed TomographyBenchmark
π― What it does: Proposed the CT-SpatialVQA benchmark for systematically evaluating the semantic spatial reasoning capabilities of 3D medical vision-language models on volumetric CT images.
π― What it does: Propose a Latent Optimal Transport Bridge (LOT-Bridge) method based on Signed Distance Field (SDF) for automatic cranial defect reconstruction, reducing the number of generation iterations and improving reconstruction quality.
Low-Rank Text-Guided Spectral Learning for Semi-supervised MHSI Segmentation
Zhang, Siqi (East China Normal University), Li, Qingli (East China Normal University)
CodeSegmentationData-Centric LearningConvolutional Neural NetworkTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBenchmarkRetrieval-Augmented Generation
π― What it does: Proposes a low-rank text-guided spectral learning network called LoTS-Net for segmenting microscopic hyperspectral images (MHSI) under semi-supervised conditions.
Low-Rank-Modulated Functa: Exploring the Latent Space of Implicit Neural Representations for Interpretable Ultrasound Video Analysis
Wolleb, Julia (Yale University), Papademetris, Xenophon (Yale University)
CodeCompressionExplainability and InterpretabilityNeural Radiance FieldAuto EncoderOptical FlowVideoBiomedical DataUltrasound
π― What it does: Propose the Low-Rank Modulated Functa network (LRM-Functa), achieving compression and interpretability analysis of ultrasound videos.
π― What it does: Proposes a method combining illumination decomposition based on Retinex theory with 4D Gaussian scattering (LumiState), achieving separation of variable near-field illumination and tissue reflectance in endoscope videos, thereby improving the geometric and visual quality of dynamic 3D reconstruction.
M3D-QAdapter: 3D Medical VQA with Lesion-Level Finding-Segmentation Alignment and Query-Driven Adaptive Token Reduction
Liu, Hong (Xiamen University), Wang, Liansheng (Xiamen University)
CodeClassificationSegmentationComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataComputed Tomography
π― What it does: M3D-QAdapter proposes a two-stage 3D medical vision question answering framework, first training a 3D ViT with lesion-level alignment, and then using a query-driven adaptive process to compress visual tokens to generate answers.
CodeImage TranslationRestorationGenerationData SynthesisConvolutional Neural NetworkDiffusion modelScore-based ModelFlow-based ModelImageBiomedical Data
π― What it does: This paper proposes a conditional flow matching model based on the concentration domain, generating missing eosin components from HE-stained images to obtain complete HES-stained images.
π― What it does: Propose a Mamba-based hybrid model for ultrasound speckle noise reduction, combining the diffusion process of partial differential equations with a learnable diffusion function to form an end-to-end trainable iterative denoising framework.
π― What it does: Proposes a multi-view breast X-ray synthesis framework called MammoFlow based on flow matching, which can generate consistent CC and MLO views without requiring real reference images.
Marginal Constrained Morphological Prototype Learning for Patch Search in Whole Slide Images
Park, Sihyeon (Korea University), Kim, Bumsoo (Chung Ang University)
CodeClassificationRetrievalComputational EfficiencyRepresentation LearningConvolutional Neural NetworkTransformerMixture of ExpertsContrastive LearningImageBiomedical DataBenchmark
π― What it does: This paper proposes a selective slice-level prediction framework called MMPL based on global morphological prototypes. It first uses the Sinkhorn-Knopp algorithm and momentum feature queue to learn global prototypes, and then uses these prototypes to retrieve the top-k most diagnostically valuable slices from each slide for aggregation, thereby achieving end-to-end training and significantly reducing computational costs.
Mask to Concept: Auto-Promptable SAM3 via Efficient Test-Time Concept Embedding Search for Few-Shot Annotation
Zhou, Quan (Wuhan University of Technology), Wang, Zhiwei (Huazhong University of Science and Technology)
CodeSegmentationData-Centric LearningTransformerPrompt EngineeringContrastive LearningImageBiomedical Data
π― What it does: Propose the Mask to Concept (M2C) framework, leveraging the concept prompting capability of SAM3, automatically searching for visual concepts from a small number of labeled samples on the frozen SAM3 architecture, achieving medical image few-shot automatic annotation and realizing efficient human-machine collaborative closed-loop;
π― What it does: This paper proposes the MBFAD multimodal balanced function assessment dataset and the MΒ²-Balance framework, aiming to achieve real-time evaluation of imbalance states under static, dynamic, and reactive activities through synchronized inertial and plantar pressure data.
MDL-DA Track: Self-supervised Sperm Motility Direction Learning and Density-Adaptive Association for Sperm Tracking
Zeng, Xiaoyu, Yang, Xuan (Shenzhen University)
CodeObject TrackingConvolutional Neural NetworkTransformerContrastive LearningOptical FlowVideoBiomedical Data
π― What it does: This paper proposes a multi-object tracking framework that combines self-supervised motion direction learning and density-adaptive association, for precise tracking of sperm in high-density microscopic videos.
CodeSegmentationExplainability and InterpretabilityImageBiomedical DataBenchmark
π― What it does: Studies the uncertainty estimation of neural cellular automata (NCA) in medical image segmentation, proposing a training-agnostic perturbation recovery method called Resilience.
Measuring What VLMs Donβt Say: Validation Metrics Hide Clinical Terminology Erasure in Radiology Report Generation
Parikh, Aditya (Technical University of Denmark), Frank, Stella (Technical University of Denmark)
CodeGenerationExplainability and InterpretabilityData-Centric LearningLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataElectronic Health RecordsRetrieval-Augmented Generation
π― What it does: This paper proposes two new evaluation metricsβClinical Association Displacement (CAD) and Weighted Association Elimination (WAE)βto detect lexical disappearance and potential biases in visual-language models when generating chest X-ray reports; meanwhile, experiments reveal the impact of decoding strategies (deterministic vs. random sampling) on clinical information retention and fairness.
MEC: A Multi-expert Consultation Framework for Synergizing Frozen Foundation Models in Whole Slide Image Analysis
Chen, Zongyi (Xiamen University), Wang, Liansheng (Hong Kong University of Science and Technology)
CodeClassificationAnomaly DetectionComputational EfficiencyRepresentation LearningTransformerMixture of ExpertsImageBiomedical Data
π― What it does: Propose a MEC multi-expert negotiation framework that utilizes multiple frozen foundation models combined with MIL for diagnosis in WSI analysis.
MEC: Medical Evidence Capsules for Retrieval-Augmented Generation in Medical Multimodal Question Answering
Xu, Zhenghua (Hebei University Of Technology), Tian, Tian (Hebei University Of Technology)
CodeRetrievalExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningDrug DiscoveryTransformerLarge Language ModelPrompt EngineeringVision Language ModelTextMultimodalityGraphTabularBiomedical DataElectronic Health RecordsBenchmarkRetrieval-Augmented Generation
π― What it does: Propose Medical Evidence Capsules (MEC), which unify heterogeneous evidence such as medical text, structured knowledge, model parameters, and imaging reports into a structured capsule format for retrieval-augmented generation (RAG) in medical multimodal question answering.
CodeDomain AdaptationExplainability and InterpretabilityRepresentation LearningTransformerSupervised Fine-TuningVision Language ModelContrastive LearningImageTextBiomedical DataBenchmark
π― What it does: This paper proposes Med-CAP, a framework that achieves robustness in medical visual question answering through contrastive evidence modeling and adaptive prior suppression.
Med3D-R1: Mitigating Narrative Bias and Enhancing Reasoning Consistency in 3D Medical Vision-Language Models
Lai, Haoran (University of Science and Technology of China), Zhou, Shaohua Kevin (University of Science and Technology of China)
CodeAnomaly DetectionExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerReinforcement LearningPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataComputed Tomography
π― What it does: Built and trained Med3D-R1, a vision-language model based on CT 3D medical imaging, for abnormal diagnosis and to improve reasoning consistency.
π― What it does: Proposed an anchor-free 3D medical image object detection framework called MedCenterDet, which directly locates the target center and infers the size through center point heatmap regression.
CodeExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningData-Centric LearningDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackTransformerReinforcement LearningPrompt EngineeringVision-Language-Action ModelTextMultimodalityTabularTime SeriesBiomedical DataElectronic Health RecordsRetrieval-Augmented GenerationChain-of-Thought
π― What it does: Built the MedEnv multimodal long-term diagnostic simulation environment, supporting multi-round Q&A and tool-assisted evidence acquisition
π― What it does: Propose the MedFuse-Seg model to achieve language-driven medical image segmentation through multi-level visual and semantic context fusion and reasoning-guided mask decoding.
Medical Knowledge-Guided Fusion of Holistic Hospital Data for Tumor Survival Analysis
Guan, Jinquan (South China University of Technology), Xie, Yutong (South China University of Technology)
CodeExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningDrug DiscoveryGraph Neural NetworkTransformerLarge Language ModelMixture of ExpertsContrastive LearningMultimodalityBiomedical DataElectronic Health Records
π― What it does: This paper proposes a medical knowledge-driven multi-modal fusion framework named MKGSurv, which integrates pan-hospital data from four disciplinesβpathology, genetics, clinical, and treatmentβto achieve tumor survival analysis.
MedTri: Structured Medical Report Normalization for Enhanced VisionβLanguage Pretraining
Chu, Yuetan (King Abdullah University of Science and Technology), Gao, Xin (King Abdullah University of Science and Technology)
CodeClassificationFederated LearningExplainability and InterpretabilityComputational EfficiencyRepresentation LearningData-Centric LearningConvolutional Neural NetworkTransformerLarge Language ModelPrompt EngineeringVision Language ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageTextMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyElectronic Health Records
π― What it does: Propose the MedTri framework, which converts free-text medical reports into structured triplets [anatomical entity: radiological description + diagnostic category], to enhance the quality of text supervision in vision-language pre-training.
CodeClassificationExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringVision Language ModelDiffusion modelTextMultimodalityTabularBiomedical DataElectronic Health RecordsRetrieval-Augmented Generation
π― What it does: This paper proposes MedTriage-LM, a multimodal large language model that integrates clinical tables, text, and synthetic visual phenotyping maps (VPM), for emergency triage instruction prediction and interpretable reasoning;
π― What it does: Propose the MedTriFlow framework, which combines tri-plane representation with implicit anatomical fields to achieve 3D medical image generation from a compressed latent space to arbitrary resolutions, and enables efficient inference on medical image generation and reconstruction tasks.
π― What it does: Propose a text-guided medical image segmentation framework named MedTSC-Net, which is used for accurately segmenting lesions in COVID-19 chest imaging (X-ray and CT).
π― What it does: Proposed a multi-phase CE-MRI translation framework (MG-RDFD) based on a shared VAE encoder and random direction feature disentanglement, achieving unified latent space separation of content and style, and realizing region-level style injection through a segmentation-guided memory pool.
MERIT: Multi-scale Mamba with Enhanced Information Retention for Treatment Response Prediction from Whole Slide Images
Hu, Taiyuan (Chinese Academy of Sciences), Yan, Rui (University of Science and Technology of China)
CodeClassificationImage TranslationDrug DiscoveryTransformerMixture of ExpertsContrastive LearningImageBiomedical DataComputed TomographyReview/Survey Paper
π― What it does: Designed and implemented a multi-scale Mamba framework called MERIT for predicting tumor treatment response from whole slide images (WSI).
π― What it does: This study first constructed the Merlin Plus dataset, providing 1,153 voxel-level tumor masks covering nine organs (spleen, bladder, gallbladder, stomach, duodenum, uterus, prostate, adrenal gland, esophagus) and supplemented with longitudinal metadata such as patient identifiers and scan times; subsequently, segmentation models such as MedFormer and R-Super were trained using these masks, and were compared and evaluated with the original Merlin dataset and publicly available multi-tumor segmentation models (Voxtell, ULS, FLARE).
π― What it does: This paper proposes MetaFormer, a Transformer architecture guided by patient metadata, which reconstructs single-lead ECG into complete 12-lead ECG.
Missingness-Aware Multimodal Learning for Four-Class Adrenal Tumor Classification
Chen, Dehua (Donghua University), An, Huimin (Shanghai Jiao Tong University)
CodeClassificationData-Centric LearningConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningImageTabularBiomedical DataComputed TomographyElectronic Health Records
π― What it does: Designed and implemented a missing-aware multimodal learning framework for four-class classification of adrenal tumors (Cushing's syndrome, primary aldosteronism, pheochromocytoma, and non-functional adenoma) using non-contrast CT images and conventional non-hormonal clinical variables.
MLLM-Enhanced Region-Aware Bidirectional Evidence-Based Model for Tongue Diagnosis
Du, Yiwei (Nanjing University), Shan, Caifeng (Nanjing University)
CodeClassificationExplainability and InterpretabilityRepresentation LearningConvolutional Neural NetworkTransformerLarge Language ModelMixture of ExpertsVision Language ModelContrastive LearningImageMultimodalityBiomedical Data
π― What it does: Propose a region-aware bidirectional evidence model integrated with a multimodal large language model for simultaneously predicting tongue coating syndrome patterns and visceral states.
MM-UNet: Meta Mamba UNet for Volumetric Medical Image Segmentation
Xie, Bin (Illinois Institute of Technology), Agam, Gady (Illinois Institute of Technology)
CodeSegmentationConvolutional Neural NetworkMixture of ExpertsBiomedical DataMagnetic Resonance ImagingComputed Tomography
π― What it does: Propose a unified U-shaped structure MM-UNet, combining convolutional networks with state space models (Mamba) for volumetric medical image segmentation.
MMIC-EndoDepth: Self-Supervised Endoscopic Depth Estimation with Multi-Mechanism Illumination Correction
Xu, Ziang (Chinese University of Hong Kong), Wang, Bing (Hong Kong Polytechnic University)
CodeDepth EstimationConvolutional Neural NetworkTransformerMixture of ExpertsAuto EncoderContrastive LearningVideoBiomedical DataBenchmark
π― What it does: Proposes the MMIC-EndoDepth framework, which estimates depth using self-supervised monocular endoscopic videos and improves reconstruction quality through multi-mechanism illumination correction.
π― What it does: Proposes the MoCaf-Mamba framework to address missing modalities and residual geometric mismatches in multi-modal medical image segmentation.
π― What it does: In the absence of instance annotations, learn morphological geometric information through semantic segmentation results to achieve instance segmentation of kidney tissue;
π― What it does: Propose a dual-view echocardiography myocardial infarction localization framework named MCF-Net, which combines extremely sparse annotated motion priors with a pre-trained EchoPrime base model;
π― What it does: Proposed a zero-shot multi-modal MRI reconstruction framework called MPFlow, which reduces artifacts in reconstruction by guiding the reconstruction process with additional auxiliary modalities during inference, using a pre-trained unconditional flow model.
MRI2Rep: Autoregressive Structured Report Generation for 3D Liver MRI
Li, Xinran (Yale University), Staib, Lawrence H. (Yale University)
CodeImage TranslationGenerationRepresentation LearningConvolutional Neural NetworkTransformerLarge Language ModelVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingRetrieval-Augmented Generation
π― What it does: Developed MRI2Rep, an end-to-end framework for automatically generating structured reports from 3D liver MRI, which directly predicts diagnostic sequences from multi-phase volumetric images and renders them into standardized reports using templates.
MRSeg3D: A Perturbation-Resilient Model for 3D Medical Reasoning Segmentation
Hao, Qin (Xinjiang University), Ye, Xujiong (University of Exeter)
CodeSegmentationConvolutional Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelImageTextBiomedical DataComputed TomographyBenchmark
π― What it does: Proposed the MRSeg3D model and its three-level progressive perturbation benchmark to evaluate and enhance the robustness of 3D medical image reasoning segmentation under real language instructions.
π― What it does: Proposed a physics-guided multi-scale graph deep unrolling network (MS-GDUN) for inversion reconstruction in fluorescence molecular tomography (FMT).
MSHF-Net: Multimodal Breast Cancer Molecular Subtype Prediction via Segmentation-Guided Hierarchical Fusion Network
Jiang, Xinjie (Hangzhou Dianzi University), Wang, Changmiao (Shenzhen Research Institute of Big Data)
CodeClassificationSegmentationAnomaly DetectionConvolutional Neural NetworkTransformerContrastive LearningImageMultimodalityTabularBiomedical DataMagnetic Resonance ImagingUltrasoundElectronic Health Records
π― What it does: A multi-modal segmentation-guided hierarchical fusion network (MSHF-Net) was constructed for non-invasive pre-surgical molecular subtype prediction of breast cancer, integrating mammography (MG), ultrasound (US), and clinical information.
π― What it does: A model named MSQGβ3DNet is proposed for classifying the lesion-level ROI in 3D automated breast ultrasound (ABUS) as benign or malignant.
MuellerPT: Decomposition Driven Pre-training for Dense Learning in Mueller Polarimetry
Tlemsani, Adam (Imperial College London), Elson, Daniel S. (Imperial College London)
CodeClassificationSegmentationConvolutional Neural NetworkSupervised Fine-TuningContrastive LearningImageBiomedical DataPhysics Related
π― What it does: A pre-training method called MuellerPT was developed, which learns dense representations by predicting the Lu-Chipman decomposition parameters of the Mueller matrix, and was evaluated on tasks such as segmentation of gray-white brain tissue and classification of colorectal cancer.
π― What it does: Proposed a selection-based test-time adaptation framework based on multi-agent for image-to-image translation tasks in medical imaging.
π― What it does: Propose a multi-annotation dynamic framework called MALFOY, which encodes different annotation types through label weight encoding, achieving adaptive learning for medical image segmentation and localization.
Multi-Disease Diagnosis in Retinal Images via Patch-Level Reasoning and Selective Aggregation
Xie, Jianyang (University of Liverpool), Zheng, Yalin (University of Liverpool)
CodeClassificationAnomaly DetectionExplainability and InterpretabilityGraph Neural NetworkTransformerContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasoundFibre Orientation DistributionDiffusion Tensor ImagingAlzheimer's DiseaseElectronic Health RecordsElectrocardiogram
π― What it does: Proposes a multi-disease retinal image diagnosis framework that combines patch-level reasoning with selective aggregation to address lesion-disease mismatch caused by comorbidities.
π― What it does: Built a model based on conditional generative adversarial networks to synthesize multi-phase images of breast DCE-MRI from non-contrast-enhanced multi-parametric MRI (T1w, DWI, and ADC).
CodeClassificationGraph Neural NetworkTransformerMixture of ExpertsContrastive LearningTime SeriesBiomedical Data
π― What it does: Propose a multi-scale dual-domain fusion spatiotemporal graph learning framework, MsDFSTGL, for multi-class classification of fNIRS signals.
Multi-Stage Dilated Convolution-Mamba Network for Surgical Action Triplet Recognition
Meng, Qingke (Jiangnan University), Zhou, Tao (Jiangnan University)
CodeRecognitionConvolutional Neural NetworkGraph Neural NetworkTransformerMixture of ExpertsContrastive LearningVideoTabularBenchmark
π― What it does: Propose a multi-stage sparse convolution-Mamba network (DCM-Net) for end-to-end identification of instrument-verb-target triplets in surgical videos.
Multi-stage NeRF for Efficient 3D Coronary Artery Reconstruction from Two Narrow-Angle Angiographic Projections
Meng, Deyu (University of Oxford), Banerjee, Abhirup (University of Oxford)
CodeOptimizationComputational EfficiencyDiffusion modelNeural Radiance FieldAuto EncoderContrastive LearningImageBiomedical DataComputed TomographyReview/Survey Paper
π― What it does: Reconstructs complete 3D vascular structures from two X-ray coronary angiography images with narrow viewing angles using a self-supervised multi-stage NeRF framework.
π― What it does: Propose a multi-layer modular amplitude perception graph autoencoder (MMGAE-WSBM), which decouples individual differences from the dynamic community structure of brain networks by identifying potential subgroups and using a contrastive coupling mechanism within the autoencoder.
π― What it does: This paper proposes a multimodal contrastive regression framework that jointly uses MRI images and structured clinical variables to predict the biological age of various organs.
Multimodal Large Language Model-Driven Self-Verification Reasoning for Interpretable Diagnosis of Diffuse Cystic Lung Diseases
Jia, Qiwei (University of Science and Technology of China), Hu, Xiaowen (Central South University)
CodeAnomaly DetectionExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringImageMultimodalityBiomedical DataComputed TomographyRetrieval-Augmented GenerationChain-of-Thought
π― What it does: Proposed a self-verification reasoning framework based on a multi-modal large language model, utilizing global spatial awareness and multi-step self-verification to achieve interpretable diagnosis of diffuse cystic lung disease.
π― What it does: A multi-modal conditional flow matching framework was developed to synthesize paired nuclear masks and pathological images for data augmentation.
π― What it does: This paper proposes a mutual distillation framework (MuDuo) based on a dual baseline model to achieve semi-supervised PET/CT organ segmentation.
π― What it does: Designed and implemented a deep learning architecture named MUX-USCT for ultrasound computed tomography (USCT) reconstruction under known acoustic acquisition geometry, achieving robustness to noise through adaptive MUX and attention mechanisms.
CodeClassificationExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningImageMultimodalityBiomedical DataChain-of-Thought
π― What it does: Propose the NeRD framework, which combines data-driven logical rules with multi-modal chain reasoning to generate efficient ontological diagnostic reasoning chains.
π― What it does: Proposes the NERVE framework, which uses network-aware bilinear decomposition to tokenize brain functional connectivity matrices in blocks, and learns transferable features in MAE self-supervised learning.