International Conference on Medical Image Computing and Computer-Assisted Intervention Β· 649 papers
Exploring Subnetwork Interactions in Heterogeneous Brain Network via Prior-Informed Graph Learning
Liu, Siyu (Northeastern University), Zaiane, Osmar R. (University of Alberta)
CodeClassificationExplainability and InterpretabilityDrug DiscoveryGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBiomedical DataAlzheimer's DiseaseElectronic Health Records
π― What it does: Propose KD-Brain, a graph learning framework based on prior information, for modeling interactions among subnetworks in heterogeneous brain networks and discovering interpretable functional pathways in mental disorder diagnosis.
Fair Curriculum Learning for Concept Bottleneck Models in Dermatology
Cockayne, Matthew J. (Keele University), Al-Bander, Baidaa (Keele University)
CodeClassificationFederated LearningExplainability and InterpretabilityData-Centric LearningTransformerSupervised Fine-TuningGenerative Adversarial NetworkContrastive LearningImageBiomedical Data
π― What it does: This paper proposes a Fair Curriculum Concept Bottleneck Model (Fair Curriculum CBM) based on the concept of fair curriculum, which reduces skin tone disparities in dermatological diagnosis through phased fair objectives training.
FairGE: Gated Expert Routing for Intersectional Fairness in Medical Foundation Models
Shao, Yuchen (South China University of Technology), Chen, Qi (South China University of Technology)
CodeClassificationFederated LearningSafty and PrivacyExplainability and InterpretabilityComputational EfficiencyRepresentation LearningAdversarial AttackTransformerMixture of ExpertsGenerative Adversarial NetworkContrastive LearningImageBiomedical DataElectronic Health Records
π― What it does: The paper proposes a cross-fair adaptation framework called FairGE for medical foundational models, achieving fair improvements for cross-subgroup sensitive attributes through multi-attribute gated Mixture-of-Experts and adversarial feature purification.
π― What it does: We propose the FairPneum framework, specifically designed to enhance fairness in the tasks of pneumonia diagnosis and lesion segmentation.
FARSIGHT: A Dynamic Multi Modal Medical Image Analysis Framework Powered by a Prior Knowledge Guided Vision Language Model
Yang, Jinghan (Chinese Academy of Sciences), Wang, Kun (Chinese Academy of Sciences)
CodeImage TranslationImage HarmonizationSegmentationAnomaly DetectionExplainability and InterpretabilityConvolutional Neural NetworkTransformerPrompt EngineeringMixture of ExpertsVision Language ModelContrastive LearningImageMultimodalityBiomedical DataUltrasound
π― What it does: Propose the FARSIGHT framework, achieving an end-to-end pipeline from raw medical image archives to patient-level diagnosis, including unsupervised MAGOS data governance and multi-modal multi-instance learning HMMIL.
π― What it does: This paper proposes a fast, data-efficient mesh-free neural surrogate based on SE(3)-equivariant neural fields for simulating embolism dynamics in ischemic stroke.
π― What it does: After airway segmentation in chest CT, the FDRAS framework is proposed, which first predicts structural mismatches through error maps, and then performs localization repair based on these error maps.
π― What it does: Propose a breast cancer classification framework for dynamic contrast-enhanced breast magnetic resonance imaging (DCE-MRI), which first locates the breast region of interest (ROI) through 3D segmentation, then selects key phases in the latent space dynamically, and finally improves sensitivity to small lesions using large-scale anatomical pre-training.
π― What it does: The study explores how to utilize multi-party model checkpoints for unlabeled OOD performance estimation in a federated learning environment, proposing the FedAgree method and designing six checkpoint combination strategies.
π― What it does: Propose the FedMoSA framework, which utilizes dedicated and shared adapters on top of SAM to achieve federated learning for medical image segmentation, addressing the problem of multi-modal heterogeneity.
π― What it does: A framework called FedProIn is proposed for federated medical imaging learning, which mitigates client drift by using learnable multi-prototypes and influence aggregation, thereby improving the model's robustness on heterogeneous data.
CodeClassificationImage TranslationRestorationSegmentationRecommendation SystemAnomaly DetectionConvolutional Neural NetworkTransformerLarge Language ModelAgentic AIVision Language ModelContrastive LearningImageVideoTextBiomedical DataUltrasound
π― What it does: This paper proposes the FetalAgents multi-agent system, achieving end-to-end automation from planar classification, segmentation, and measurement of fetal ultrasound images to video summarization and structured report generation.
Fibers of Asymmetric Similarity: A Framework for Clinical and Imaging Data
MΓΌller, Johanna P. (Friedrich-Alexander University Erlangen-NΓΌrnberg), Kainz, Bernhard (Friedrich-Alexander University Erlangen-NΓΌrnberg)
CodeClassificationAnomaly DetectionExplainability and InterpretabilityData-Centric LearningTransformerContrastive LearningImageTabularBiomedical DataUltrasoundElectronic Health Records
π― What it does: A unified framework based on asymmetric Tversky similarity was studied, which handles clinical tabular data and frozen image encoders, addressing class imbalance and missing data issues in rare disease screening.
FOCUS: Towards Fetal Obstetric Corrective UltraSound Guidance
Lamdouar, Hala (University of Oxford), Noble, J. Alison (University of Oxford)
CodeRecognitionGenerationData SynthesisConvolutional Neural NetworkTransformerPrompt EngineeringVision Language ModelContrastive LearningVideoTextMultimodalityUltrasound
π― What it does: Proposed and implemented the FOCUS framework, which utilizes multimodal information from video and text to generate targeted corrective feedback in real-time for obstetric ultrasound learners, helping them improve technical details such as probe posture and parameter settings.
π― What it does: A framework for selecting key anatomical frames based on foundational models is proposed, using blind sweep ultrasound videos to estimate fetal birth weight 48 hours before delivery.
π― What it does: Propose a frequency and geometry guided graph clustering framework FG3-Cluster, which achieves skin lesion segmentation using weak labels;
CodeClassificationConvolutional Neural NetworkContrastive LearningBiomedical Data
π― What it does: This paper proposes a frequency-aware multi-margin loss (FMML) and a long-term tail rank deviation (LTRD) metric to address the long-tailed imbalance problem in medical ordinal classification.
π― What it does: This paper proposes a frequency-aware neural architecture search (NAS) framework, utilizing bidirectional Mamba modules for cross-band modeling in EEG emotion recognition, and independently searching for variable topologies and internal temporal configurations within each frequency band.
π― What it does: This paper proposes a post-training quantization framework based on frequency awareness, specifically designed for medical image denoising;
Frequency-Aware Prototype Learning in VLM for Intracranial Aneurysm HR-VWI Segmentation
Wang, Yixin (Southern University of Science and Technology), Si, Weixin (Shenzhen University of Advanced Technology)
CodeSegmentationTransformerPrompt EngineeringVision Language ModelDiffusion modelContrastive LearningImageTextBiomedical DataMagnetic Resonance Imaging
π― What it does: This paper proposes a visual-language model based on frequency-aware prototype learning for brain aneurysm segmentation in high-resolution vascular wall imaging (HR-VWI).
π― What it does: Proposed a 3D latent diffusion model based on large-scale chest CT pre-training, capable of synthesizing future IPF progression CT images from a single baseline scan.
From Geometry to Clinic: A Physics-Aware Evaluation Framework and Benchmark for Automated Dental Crown Design
Wu, Jiamin (University of Hong Kong), Tsoi, James Kit Hon (Fuzhou University)
CodeGenerationTransformerDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningPoint CloudMeshBiomedical DataBenchmarkPhysics Related
π― What it does: Proposes CrownBench, a physics-aware evaluation framework and benchmark for assessing automated crown design from a clinical perspective.
From Perception to Anticipation: Forecasting VesselβInstrument Interactions in Endoscopic Surgery under Unreliable Observations
Chen, Yueyao (Chinese University of Hong Kong), Dou, Qi (Chinese University of Hong Kong)
CodeClassificationRecognitionSegmentationExplainability and InterpretabilityComputational EfficiencyRecurrent Neural NetworkTransformerPrompt EngineeringVision Language ModelAuto EncoderContrastive LearningVideoBiomedical DataElectronic Health Records
π― What it does: Proposes a task of predicting the imminent interaction status (no interaction, approaching, contact) between vessels and instruments based on short-term videos and segmentation masks under endoscopy, and achieves real-time prediction.
π― What it does: Modeling preterm birth prediction as a multiple instance learning (MIL) problem, using multiple transvaginal ultrasound (TVUS) images from each pregnant woman. The Gaussian Mixture Model (GMM) aggregator maps the distribution of image features into a fixed-length representation, which is then used to predict the risk of preterm birth.
From Scanning Guidelines to Action: A Robotic Ultrasound Agent with LLM-Based Reasoning
Bi, Yuan (Technical University of Munich), Navab, Nassir (Technical University of Munich)
CodeRobotic IntelligenceReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AITextBiomedical DataUltrasoundRetrieval-Augmented Generation
π― What it does: Achieve autonomous operation of robotic ultrasound based on scanning guidelines by using a large language model as an autonomous planner combined with a set of tools.
π― What it does: This paper proposes an explicit spatial modeling framework called ExpGS, based on a fixed virtual dual-view projection paradigm, to reconstruct three-dimensional skeletal structures from dual-plane X-ray images, combining three-dimensional Gaussian splatting with Transformer to achieve reconstruction.
π― What it does: A coarse-to-fine residual pyramid decoder directly applicable to 3D deformation registration is constructed by leveraging a frozen large-scale pre-trained 3D segmentation encoder as an anatomical feature space. More robust correspondence learning is achieved through feature space similarity loss and the DPI-Fuse interaction module.
FunPiQ: A New Benchmark for Pixel-Level Quality Assessment in Fundus Images
Wang, Pengwei (Medical University of Vienna), BogunoviΔ, Hrvoje (Medical University of Vienna)
CodeClassificationAnomaly DetectionExplainability and InterpretabilityConvolutional Neural NetworkContrastive LearningImageBiomedical DataBenchmark
π― What it does: This paper develops the FunPiQ dataset, which uses pixel-level annotations to evaluate the quality of fundus images, and proposes the EFIQA-CP method to achieve interpretable quality prediction.
π― What it does: Propose an adaptive cardiac cine MRI reconstruction framework based on Gabor primitives, achieving scan-specific reconstruction from undersampled k-space data.
π― What it does: Proposes the FedGMA framework to address cross-cancer domain generalization issues in federated learning when different centers use heterogeneous label spaces, achieving feature and label decoupling through generated anchors and geometric consistency regularization.
Dasdelen, Muhammed Furkan (Helmholtz Munich), Marr, Carsten (Helmholtz Munich)
CodeClassificationRetrievalRepresentation LearningTransformerAuto EncoderContrastive LearningImageMultimodalityTabularBiomedical Data
π― What it does: Proposes the GenBloom framework, which aligns single-cell blood smear images with chromosomal aberration and mutation data through multi-modal alignment to learn patient-level encodings and improve hematology diagnostic performance.
π― What it does: Under the extremely limited data environment with only 62 expert-annotated subjects, they utilized a large-scale unannotated UK Biobank 30k sample for self-supervised pre-training. Subsequently, they integrated three-channel semantic inputs into the pre-trained spherical encoder through a flexible topological prior injector, ultimately achieving high-precision annotation of more than 60 sulci in the entire right hemisphere of the brain (average Dice 0.77).
Geometry-Aware Manifold Trajectory Modeling for Task fMRI Activation Mapping
Li, Yueran (Harbin Institute of Technology), Su, Jingyong (Sichuan University)
CodeExplainability and InterpretabilityComputational EfficiencyRepresentation LearningBiomedical DataMagnetic Resonance Imaging
π― What it does: This paper proposes a geometry-aware manifold trajectory modeling framework for task fMRI activation mapping without the HRF assumption.
π― What it does: This paper proposes a geometry-aware point-to-voxel fusion framework called GeoP2VNet, aimed at improving the segmentation accuracy of intracranial aneurysms in CTA images.
Geometry-Aware Superpixel Graph Transformer with Metadata for Skin Lesion Classification
Azeem, Muhammad (Edge Hill University), Behera, Ardhendu (Edge Hill University)
CodeClassificationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningConvolutional Neural NetworkGraph Neural NetworkTransformerDiffusion modelScore-based ModelFlow-based ModelRectified FlowAuto EncoderContrastive LearningImageGraphBiomedical DataReview/Survey Paper
π― What it does: Proposed a Transformer framework based on geometry-aware superpixel graphs (GeoMeta-GT), which segments skin lesion images into superpixel nodes and embeds patient metadata as dedicated nodes into the graph, achieving region-level relationship modeling and multimodal fusion.
π― What it does: Proposed the GFR-MIL framework, which utilizes a three-stage iterative closed-loop, Glance-Focus-Reflect, to achieve bidirectional interaction between low-resolution global information and high-resolution details, and combines a recycling mechanism and memory mask to enhance the performance of whole slide image (WSI) multi-instance learning.
π― What it does: Proposed a gene-guided self-supervised autoencoder (GG-AE), which first learns a compact genetic latent space via a variational autoencoder (VAE) of genes, and then uses a dual-latent-space non-negative decoder to decompose brain volume degeneration images into a genetically aligned subspace and a residual subspace, thereby achieving disentanglement between genetically driven brain atrophy patterns and non-genetic changes.
π― What it does: Proposed a novel medical image segmentation framework called Glasses, which utilizes the frozen DINOv3 and combines it with a multi-scale Deformable Attention Adapter and an edge-aware convolutional encoder.
CodeSegmentationGraph Neural NetworkTransformerVision Language ModelAuto EncoderContrastive LearningImageBiomedical DataComputed Tomography
π― What it does: Propose a graph-structure-based lesion description modeling and verification framework called GLeVE, which can directly align and precisely locate lesion descriptions from radiology reports in three-dimensional CT images.
Global-Local Structure-Aware Alignment for Automated LGE MRI Report Generation
Zhou, Chenyang (Beijing Normal University), Duan, Jinming (University of Manchester)
CodeGenerationGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingRetrieval-Augmented Generation
π― What it does: This paper proposes the GLSA-RG framework to achieve automatic report generation based on LGE MRI, utilizing structural graph representation and structural alignment to capture regional-level enhancement patterns and cross-slice continuity.
π― What it does: This paper proposes a morphology analysis method for aortic valve based on graph neural networks (GNN), which constructs a graph structure using the candidate distribution of detected key anatomical landmarks and performs adaptive correction of morphological measurements through GNN.
π― What it does: This paper proposes the HadBalance framework, which combines the Hadwiger shape prior with conflict-aware target balancing for general medical image segmentation.
Hallucination Detection and Correction in Medical VLMs via Counter-Evidence Verification
Zhou, Nan (Sichuan University), Fu, Huazhu (Agency for Science, Technology and Research)
CodeAnomaly DetectionExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningDrug DiscoveryPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataElectronic Health RecordsReview/Survey PaperRetrieval-Augmented GenerationChain-of-Thought
π― What it does: Proposes a training-agnostic, plug-and-play CoEV framework that utilizes adversarial masking of visual regions to detect and correct hallucinations in medical vision-language models
CodeClassificationSegmentationTransformerMixture of ExpertsContrastive LearningBiomedical DataUltrasound
π― What it does: This paper proposes a Hierarchical Affinity-Aware Routing Mixture-of-Experts (HAR-MoE) model that jointly learns pixel-level segmentation and image-level classification tasks for multi-organ ultrasound images.
π― What it does: A spatiotemporal heatmap feature fusion network (STHFFN) for the auxiliary diagnosis of ADHD children based on functional near-infrared spectroscopy (fNIRS) is proposed. It converts discrete HbO2/HbR signals into continuous heatmaps and jointly models brain region activation, hemoglobin coupling, and long-term attention instability using parallel spatiotemporal attention, bidirectional hemoglobin complementary attention, and memory pool mechanisms.
π― What it does: This paper proposes the Hypothesis-Driven Test-Time Adaptation (HD-TTA) framework, aiming to achieve safe and interpretable brain tumor segmentation. Instead of blindly performing global optimization, the framework adapts the model during testing by selecting geometric hypotheses (compression and expansion) and using an unsupervised selector for self-optimization.
π― What it does: Designed HeartFormer, a dual-structure semantic-aware Transformer, for reconstructing complete four-chamber cardiac point clouds from conventional short-axis and long-axis Cine MRI.
π― What it does: Proposed the HeartVolMesh framework, achieving end-to-end generation from CTA images to tetrahedral cardiac volume meshes with consistent vertex correspondence.
π― What it does: Proposed a medical image segmentation network called HEDS-Net, which combines global sequence modeling and local detail enhancement, and introduces axial bridging and progressive weighted deep supervision within a U-shaped structure.
CodeDiffusion modelAuto EncoderContrastive LearningBiomedical DataMagnetic Resonance ImagingPhysics Related
π― What it does: A physics-based brain perfusion digital twin, HemoPIC, was constructed to directly estimate perfusion indicators such as blood flow, blood volume, and transit time from dynamic perfusion imaging, without the need for manually selecting an arterial input function.
Hepato-LLaVA: An Expert MLLM with Sparse Topo-Pack Attention for Hepatocellular Pathology Analysis on Whole Slide Images
Yang, Yuxuan, Ma, Zhanyu
CodeRecognitionTransformerLarge Language ModelPrompt EngineeringVision Language ModelAuto EncoderContrastive LearningImageTextMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyBenchmark
π― What it does: Propose Hepato-LLaVA, a multimodal large language model specialized for hepatocellular carcinoma (HCC) pathological images, and construct the HepatoPathoVQA multi-scale question-answering dataset.
π― What it does: A whole-body PET-guided MRI translation model based on the differential SchrΓΆdinger bridge (HA-DSB) was studied, achieving the synthesis of high-resolution T2 sequences from rapidly acquired LAVA sequences.
π― What it does: Proposed the HiEDL framework, which decomposes tumor segmentation into a two-level hierarchical task of organ-tumor, models predictive uncertainty using evidential deep learning, and achieves uncertainty-aware 3D CT tumor segmentation by combining distance map boundary loss
CodeClassificationPose EstimationAnomaly DetectionRepresentation LearningConvolutional Neural NetworkTransformerMixture of ExpertsContrastive LearningImageVideoMultimodalityBiomedical DataAlzheimer's DiseaseElectronic Health Records
π― What it does: Propose a hierarchical cross-modal fusion framework that integrates three complementary modalitiesβocular imaging, behavior, and gaitβto achieve high-precision non-invasive aging assessment.
π― What it does: This study investigates semi-supervised multi-modal medical image segmentation, proposing the HDCL framework, which leverages hierarchical decoupled representations and prediction consistency to improve segmentation performance.
Hierarchical Memory for Radiology Report Generation with Rich Context
Jiao, Qingyue (University of Notre Dame), Shi, Yiyu (University of Notre Dame)
CodeGenerationRetrievalRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningImageTextBiomedical DataComputed TomographyElectronic Health RecordsRetrieval-Augmented Generation
π― What it does: Proposed a hierarchical memory-based radiology report generation framework called HM-RRG, which utilizes cross-patient retrieved reports and longitudinal historical information from the same patient;
CodeClassificationAnomaly DetectionExplainability and InterpretabilityTransformerVision Language ModelContrastive LearningVideoBiomedical DataUltrasound
π― What it does: Proposed a hierarchical multi-video multi-instance learning (HMV-MIL) method for directly performing patient-level tuberculosis (TB) detection from multi-video ultrasound images.
π― What it does: This paper proposes a hierarchical region-aware multi-granularity Mamba (HiRAM) model for the white matter lesion segmentation task, and integrates it into the SAM decoder to achieve fine-grained processing of internal, boundary, and background region features.
CodeSegmentationConvolutional Neural NetworkTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance Imaging
π― What it does: Propose the TextSCP framework, which uses hierarchical text guidance to achieve fine segmentation of three sub-regions of brain tumors (WT, TC, ET).
HiFi-Rep: Leveraging High-Resolution Vision Foundation Models with Tri-Planar Slice Context for CT Report Generation
Limaroon, Keetawan (King Mongkut's University of Technology Thonburi), Tarnpradab, Sansiri (King Mongkut's University of Technology Thonburi)
CodeGenerationTransformerLarge Language ModelPrompt EngineeringContrastive LearningImageTextBiomedical DataComputed TomographyRetrieval-Augmented Generation
π― What it does: Propose the HiFi-Rep framework, which utilizes high-resolution DINOv3 features and HiFi-S2D aggregation, combined with the MedGemma LLM to achieve automatic generation of CT reports.
π― What it does: Propose the HGβSemiSeg framework, which improves the quality of pseudo labels in semi-supervised ultrasound image segmentation through high-frequency information enhancement.
HiPath: Hierarchical Vision-Language Alignment for Structured Pathology Report Prediction
Yuan, Ruicheng (Hunan University), Yang, Guang (Imperial College London)
CodeExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningVision Language ModelContrastive LearningImageTextBiomedical DataElectronic Health Records
π― What it does: Designed and trained HiPath, a lightweight vision-language model for predicting structured pathological reports.
How Good Are Foundation Models in Longitudinal MRI Disease Progression Reasoning?
Al Ghallabi, Wafa (Mohamed bin Zayed University of Artificial Intelligence), Khan, Fahad Shahbaz (Mohamed bin Zayed University of Artificial Intelligence)
CodeTransformerPrompt EngineeringVision Language ModelImageMultimodalityTime SeriesBiomedical DataMagnetic Resonance ImagingBenchmarkRetrieval-Augmented GenerationChain-of-Thought
π― What it does: This paper designs a Time-Aware Multi-View MRI Benchmark, integrating multi-disease, longitudinal, and multi-view MRI data with structured question answering to evaluate the ability of vision-language models in temporal and spatial reasoning.
π― What it does: A hybrid Transformer-Mamba structure called TranSamba was studied for weakly supervised volumetric medical segmentation based on planar-level labels.
π― What it does: This paper proposes a framework based on hyperbolic contrastive learning and inference loss, utilizing pathological images to predict gene expression in spatial transcriptomics.
HypOProto: Hyperbolic Ordinal Prototypes for Left Ventricular Filling Pressure Classification
Wu, Victoria (University of British Columbia), Tsang, Teresa S. M. (University of British Columbia)
CodeClassificationExplainability and InterpretabilityTransformerSupervised Fine-TuningContrastive LearningImageVideoBiomedical DataUltrasound
π― What it does: Implement an interpretable classification of left ventricular filling pressure (LVFP) on B-mode ultrasound images using a frozen DINOv3 base model, proposing the HypOProto method, which maps prototypes into a hyperbolic (Lorentzian) space, organizing them sequentially according to clinical E/e' thresholds;
IΒ²-Med: Interpretable Medical Inference Through Visual-Guided Dynamic Logits Calibration
Sajid, Md (Indian Institute of Technology Indore), Tanveer, Mohammad (Indian Institute of Technology Indore)
CodeExplainability and InterpretabilityComputational EfficiencyTransformerReinforcement LearningPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound
π― What it does: Propose a vision-guided dynamic logits calibration framework called I-Med, which combines structured reasoning based on reinforcement learning with visual inspection during inference, thereby improving the interpretability and visual authenticity of medical vision-language models without altering model parameters.
Image-mediated fMRI-to-caption generation with visual pathway tokens and hyperbolic alignment
Choi, Hyoungshin (Sungkyunkwan University), Park, Hyunjin (Sungkyunkwan University)
CodeGenerationKnowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringVision Language ModelDiffusion modelGenerative Adversarial NetworkContrastive LearningImageTextMultimodalityBiomedical DataMagnetic Resonance ImagingRetrieval-Augmented Generation
π― What it does: Propose a two-stage brain-image-text generation framework, which first maps fMRI signals to image embeddings, and then generates natural language captions using LLaMA3.
π― What it does: Propose a weakly supervised cell instance segmentation framework called ImPartial, which can achieve cell segmentation on multi-channel images with only a small number of scribble annotations.
Improving Factuality of 3D Brain MRI Report Generation with Paired Image-domain Retrieval and Text-domain Augmentation
Lee, Junhyeok (Seoul National University), Ye, Jong Chul (KAIST)
CodeGenerationData SynthesisRetrievalTransformerLarge Language ModelPrompt EngineeringAuto EncoderImageTextBiomedical DataMagnetic Resonance ImagingRetrieval-Augmented Generation
π― What it does: Proposed a framework named PIRTA, which generates radiology reports with higher factual accuracy by utilizing image-domain retrieval and text-domain enhancement from 3D brain MRI.
Inception-Powered Unrolling Reconstruction for Accelerated Dynamic Cardiac Late Gadolinium Enhancement MRI
Zhong, Ya (Chongqing University of Technology), Zhu, Yanjie (Paul C. Lauterbur Research Center for Biomedical Imaging, Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences)
π― What it does: To address the motion artifact problem in late gadolinium enhancement (LGE) cardiac imaging, this paper proposes to reconstruct it as a dynamic multi-frame accelerated reconstruction task, achieving motion robustness improvement without additional scanning time by segmenting the acquired k-space into multiple time frames.
CodeSegmentationKnowledge DistillationConvolutional Neural NetworkTransformerContrastive LearningImageBiomedical Data
π― What it does: This paper proposes a framework called LiteΟ, which significantly improves segmentation performance by injecting structural and semantic priors from foundation models (SAM, DINOv2, OneFormer) into a lightweight multi-adenoma segmentation baseline.
π― What it does: In the multi-slice alignment and fusion of spatial transcriptomics, the INST-Align framework is proposed, which achieves unsupervised joint alignment and reconstruction using a shared implicit neural expression field and deformation network.
π― What it does: What was done: Propose a framework for localβglobal brain aging speed estimation based on Jacobian discriminative maps derived from deformation registration, first calculating local speeds in blocks and then aggregating them into a global speed through learnable weights.
Interpretable Longitudinal Disease Forecasting with Graph-Based Multimodal Encoding and Population-Aware Language Generation
Fan, Yuheng (University of Liverpool), Zhao, He (University of Liverpool)
CodeExplainability and InterpretabilityDrug DiscoveryGraph Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningMultimodalityTime SeriesBiomedical DataMagnetic Resonance ImagingAlzheimer's DiseaseElectronic Health RecordsRetrieval-Augmented Generation
π― What it does: Propose the ENIGMA framework, which utilizes multimodal time series data through graph neural network encoding and dual population memory modules, jointly with large language models, to achieve prediction of Alzheimer's disease and interpretable text generation.
CodeClassificationExplainability and InterpretabilityGraph Neural NetworkVision Language ModelContrastive LearningImageBiomedical DataComputed TomographyUltrasound
π― What it does: Proposed an interpretable medical image diagnosis framework called ConceptAlignE-G, which integrates the concept alignment of vision-language models with graph convolutional reasoning to achieve decision-making based on clinical imaging biomarkers;
Interpretable Probabilistic Medical Image Segmentation via Gaussian Process with Explicit Modelling of Annotation Bias and Variability
Li, Qi (University College London), Hu, Yipeng (King's College London)
CodeSegmentationExplainability and InterpretabilityConvolutional Neural NetworkContrastive LearningBiomedical DataUltrasound
π― What it does: Propose a multi-annotator medical image segmentation framework that explicitly models annotation bias and variation in the logit space, using sparse variational Gaussian process (SVGP) to learn the image reference logit, and capturing systematic bias and random error from different annotators through annotation-specific mean and variance.
π― What it does: Proposed a cross-modal MRI domain adaptation method called IntraStyler, which utilizes unpaired image translation to achieve fine-grained style adaptive synthesis within the target domain, thereby enhancing the generalization ability of downstream segmentation models.
π― What it does: Proposes an unsupervised medical anomaly detection framework named InvDetect, which uses a DDIM model trained only on normal images to perform reverse inference, constructing a noise latent space and performing anomaly discrimination via one-class SVM in this space, while introducing spatial coherence post-processing to obtain the final anomaly segmentation.
π― What it does: Propose a dual-stream physiological conditioning network named JANUS, which integrates macro radiomic quantitative features with visual representations to achieve more robust multi-label classification under distribution shift in the CT triage task.
CodeDrug DiscoveryRecurrent Neural NetworkTransformerPrompt EngineeringMixture of ExpertsVision Language ModelContrastive LearningImageTextMultimodalityBiomedical Data
π― What it does: Propose a concept-conditioned slot factorization framework called PathoSlot, which simultaneously performs biomarker prediction and survival risk estimation on multimodal pathological data (whole slide images and pathological reports).
π― What it does: This paper proposes a cross-view contrastive learning framework for jointly learning global image and local ROI graph representations to achieve brain disease classification.
π― What it does: Propose an end-to-end deep learning framework that can directly predict the s-rep (skeleton model) of the hippocampus and voxel segmentation from 3D medical images.
CodeSegmentationConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningImageBiomedical Data
π― What it does: Propose a KAN-AINet model based on ConvNeXt-U-Net, utilizing the Kolmogorov-Arnold network (KAN) to achieve adaptive illumination modulation and boundary attention, thereby enhancing the cross-dataset robustness of colon polyp segmentation.
KANEx: Translating Kolmogorov-Arnold Networksβ Interpretability to Medical Explainability
Shailya, Krithi (Indian Institute of Technology Madras), Ravindran, Balaraman (Indian Institute of Technology Madras)
CodeClassificationExplainability and InterpretabilityConvolutional Neural NetworkTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataComputed TomographyElectronic Health Records
π― What it does: Proposes the KANEx framework, combining Kolmogorov-Arnold networks (KAN) with Vision-Language models (VLM), to simultaneously generate interpretable heatmaps and textual reports in multi-label diagnosis of chest X-rays.
KEPIL: Knowledge-Enhanced Prompt-Image Learning for Prompt-Robust Disease Detection
Luo, Haozhe (University of Bern), Reyes, Mauricio (University of Bern)
CodeClassificationSegmentationAnomaly DetectionTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataComputed TomographyElectronic Health Records
π― What it does: Proposed the KEPIL framework, achieving prompt-robust zero-shot disease detection through knowledge-enhanced dynamic prompts and semantic contrastive learning.
KIDA: Kinematic-Intent Dual-path Alignment for Surgical Error Detection
Luo, Yuxuan (Huazhong University of Science and Technology), Wang, Zhiwei (Huazhong University of Science and Technology)
CodeAnomaly DetectionRobotic IntelligenceRecurrent Neural NetworkTransformerVision-Language-Action ModelContrastive LearningOptical FlowVideoTextBiomedical Data
π― What it does: Proposes a dual-path alignment framework called KIDA for detecting subtle errors in robotic-assisted surgery by aligning actual execution with clinical standard intentions.
π― What it does: Proposed a KiMoCo-Net, which first restores phase and amplitude consistency in the k-space complex domain, and then refines it in the image domain to complete MRI motion artifact correction.
Knowledge-Enhanced Representation Learning with Retrieval-Augmented Multimodal Fusion for Survival Prediction
Zhang, Zeyu (University of British Columbia), Bashashati, Ali (University of British Columbia)
CodeRetrievalRepresentation LearningTransformerVision Language ModelContrastive LearningImageMultimodalityTabularBiomedical DataElectronic Health RecordsRetrieval-Augmented Generation
π― What it does: Developed the KERA framework, combining knowledge-enhanced pre-training with retrieval-enhanced multi-modal fusion, achieving survival prediction through the integration of whole-slide images and transcriptomics.
π― What it does: Based on multi-modal mpMRI, combining clinical variables and report priors, the KOAL framework is proposed for lesion-level Gleason Grade Group prediction.
π― What it does: Propose a learnable anisotropic deformation framework called LAD, enabling 3D Transformer to adapt to the spatial inhomogeneity of medical images.
LaGuadia: Language-Guided Adaptive Distillation from Pathology Foundation Models
Kim, Gangsu (Korea University), Jeong, Won-Ki (Korea University)
CodeExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerSupervised Fine-TuningPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataElectronic Health RecordsBenchmark
π― What it does: Proposes the LaGuadia framework, which dynamically integrates the expertise of multiple pathological foundation models (PFMs) into a lightweight (87M parameters) WSI encoder through language-guided adaptive knowledge distillation.