These 649 MICCAI 2026 papers come with a code repository. Each shows an AI one-line summary below β get the verified repo link + the full 6-part summary (innovation, method, data, results, limitations) and search every MICCAI 2026 paper, free trial on arXivSub.
3D Classification of Paramagnetic Rim Lesions in Multiple Sclerosis via Asymmetric QSMβFLAIR Modeling
Pignedoli, Veronica (University of Genova), Moro, Matteo (University of Genova)
π― What it does: This study proposes FRODO, a 3D deep learning framework based on QSM and FLAIR, for automatically distinguishing paramagnetic rim lesions (Rim+) from non-Rim lesions (Rim-) in multiple sclerosis.
A Counterfactual Framework for Directional CellβCell Interaction Analysis in Spatial Transcriptomics
Anzum, Humaira (University of Houston), Banerjee, Tania (University of Houston)
CodeDrug DiscoveryGraph Neural NetworkGenerative Adversarial NetworkContrastive LearningBiomedical Data
π― What it does: Designed a directional cell-cell interaction framework based on adversarial inference, utilizing a neighborhood graph model to predict receptor cell states and quantify directional effects through hypothetical interventions.
π― What it does: Propose VSP-Branch, a pluggable structural-guided feature aggregation module, which enhances the continuity and branch integrity of vascular representations by constructing learnable paths along the local vascular structure, aggregating cross-layer features, and adaptively injecting structural information through difficulty-gated mechanisms.
A Guideline-Aware AI Agent for Zero-Shot Target Volume Auto-Delineation
Kim, Yoon Jo (Oncosoft Inc.), Kim, Jin Sung (Oncosoft Inc.)
CodeSegmentationExplainability and InterpretabilityConvolutional Neural NetworkTransformerLarge Language ModelAgentic AIImageTextBiomedical DataComputed TomographyRetrieval-Augmented GenerationChain-of-Thought
π― What it does: Proposed an AI agent framework called OncoAgent, which can zero-shot automatically convert text-based radiotherapy clinical guidelines into three-dimensional target volume volumes;
CodeClassificationRepresentation LearningConvolutional Neural NetworkTransformerMixture of ExpertsContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyAlzheimer's DiseaseElectronic Health Records
π― What it does: Propose the MoMmNet framework, which utilizes multi-organ imaging (brain, heart, gut, liver, kidney) and clinical text for hierarchical multi-task reasoning to achieve multi-cause dementia diagnosis.
A Multi-center Benchmark for Abdominal Disease Diagnosis and Report Generation from Non-Contrast CT
Elbakry, Mariam (Ain Shams University), Elbatel, Marawan (Hong Kong University of Science and Technology)
CodeClassificationGenerationExplainability and InterpretabilityConvolutional Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningImageTextBiomedical DataComputed TomographyBenchmarkRetrieval-Augmented Generation
π― What it does: Designed and released a multi-center benchmark to evaluate the feasibility of generating multi-organ diagnostic reports from single-phase non-contrast CT.
π― What it does: Proposes HyperBrainNet, a framework combining causal effective connectivity, high-order hybrid hypergraph networks, and multi-view adaptive fusion, for the diagnosis of early Alzheimer's disease.
A Multimodal 3D Foundation Model for Light Sheet Fluorescence Microscopy Enables Few-Shot Segmentation, Classification, and Deblurring
Scheinfeld, Adina (Weill Cornell Medicine), Paetzold, Johannes C. (Weill Cornell Medicine)
CodeClassificationRestorationSegmentationKnowledge DistillationRepresentation LearningConvolutional Neural NetworkTransformerSupervised Fine-TuningVision Language ModelContrastive LearningImageTextBiomedical Data
π― What it does: This paper constructs a multi-modal 3D foundation model, pre-trained using a large-scale unlabeled light sheet fluorescence microscopy (LSM) volume images and corresponding text descriptions, achieving efficient few-shot fine-tuning on downstream tasks such as segmentation, classification, and deblurring.
π― What it does: Evaluate and compare multiple OOS detection methods in real clinical settings for failure detection in automatic liver CT segmentation models.
π― What it does: Propose a unified multi-task framework that jointly detects adenomas (EPVS) and small cavities (lacunes) in the brain and improves detection accuracy.
π― What it does: Propose the AC-MIL framework to achieve weakly supervised atrial LGE-MRI quality assessment, decomposing overall quality into interpretable clinical concepts;
π― What it does: To address the insufficient generalization of pre-trained models for 3D medical image segmentation on unseen data, this paper proposes a post-adaptive clipping alignment (ACA) method. It enhances segmentation performance by achieving local distribution alignment within the model's uncertain logit interval through isotropic quantile mapping and isotropic regression.
π― What it does: Propose an AdaSurvMamba framework applicable to multi-modal survival analysis, combining WSI and genomic data to achieve accurate cancer prognosis prediction.
Addressing Gradient Conflicts in Multimodal Fundus Disease Recognition with Fusion-Guided Learning
Liu, Xiaozhou (Southwest University), Du, Zhiguo (Southwest University)
CodeRecognitionOptimizationConvolutional Neural NetworkSupervised Fine-TuningContrastive LearningImageMultimodalityBiomedical DataAlzheimer's DiseaseReview/Survey Paper
π― What it does: Address gradient conflicts in multi-modal retinal image classification by introducing the fusion-guided gradient decoupling learning (FGDL) and cross-modal semantic alignment (CMSA) modules to improve dynamic optimization and enhance feature alignment;
Addressing Tissue and Appearance Heterogeneity in Text-Guided Few-Shot WSI Classification
Li, Yongcen (Dalian University of Technology), Xing, Xudong (Beijing Institute of Genomics)
CodeClassificationDomain AdaptationTransformerVision Language ModelContrastive LearningImageTextBiomedical DataComputed TomographyReview/Survey Paper
π― What it does: This paper proposes a Heterogeneity-Aware Text-guided MIL framework to address the domain drift problem caused by tissue composition and appearance heterogeneity in whole slide image classification under the extremely few-sample condition.
π― What it does: Addressing the fairness issues of medical foundation models (such as SAM and MedSegX) in the segmentation of head and neck squamous cell carcinoma (HNSCC), this paper proposes the AEGIS framework, achieving group-invariant segmentation under parameter-efficient fine-tuning (PEFT).
π― What it does: Propose a framework that combines graph attention autoencoder with Brownian bridge diffusion for long-term evolution prediction of adolescent functional connectivity under age conditions.
AI-Driven Pulmonary Congestion Assessment for Lung Ultrasound via Segmentation-Guided Transformers
Fooladgar, Fahimeh, Kapur, Tina (Brigham and Women's Hospital)
CodeClassificationSegmentationAnomaly DetectionExplainability and InterpretabilityConvolutional Neural NetworkTransformerPrompt EngineeringMixture of ExpertsContrastive LearningImageBiomedical DataUltrasound
π― What it does: Propose a two-stage segmentation-guided transformer framework for automated severity scoring of lung B-lines, enabling consistent and reproducible assessment of pulmonary edema directly at the bedside in lung ultrasound.
An Artifact-Based Agent Framework for Adaptive and Reproducible Medical Image Processing
Zuo, Lianrui (Vanderbilt University), Landman, Bennett A. (Vanderbilt University)
CodeImage TranslationRestorationSegmentationReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
π― What it does: Proposed an Agent framework based on artifact contracts, achieving adaptive configuration and reproducible execution for medical image processing;
π― What it does: A model is proposed that generates continuous, Cβ smooth vascular centerlines through the diffusion of three-phase Chebyshev curves, capable of recovering millimeter-level dimensions while maintaining the curve trajectory.
π― What it does: Propose a domain randomization framework based on whole-brain anatomical labels, using only synthetic images to train a 3D CNN, achieving zero-shot brain vessel segmentation on unseen TOF-MRA and CTA data;
π― What it does: This paper generates digital reconstructed radiographs (DRR) from patient CT angiography data to construct a multi-view coronary artery correspondence dataset without manual annotations, and proposes a Geometry-Informed Matching Module (GIMM) to achieve fine matching.
π― What it does: Proposed a 4D motion medical image generation framework based on semi-supervised VAE and two-stage residual latent diffusion models, which can simultaneously generate voxel sequences and corresponding segmentation masks, achieving controllable 4D cardiac MRI synthesis.
π― What it does: Propose a source domain-agnostic domain adaptation method for ultrasound video segmentation, adopting a decoupled anatomical structure and texture, along with a safe gradient guidance mechanism, to achieve robust adaptation to the target domain.
π― What it does: Proposes an unsupervised coarse-to-fine framework called Angio-Stitch for sequentially merging liver CLE (confocal laser endomicroscopy) microvascular images, expanding the field of view while preserving fine structures.
Angular-Constrained Hyperbolic Learning for Hierarchical Multimodal Survival Prediction
Yang, Haotian (East China Normal University), Wang, Yan (East China Normal University)
CodeRepresentation LearningData-Centric LearningTransformerAuto EncoderGenerative Adversarial NetworkContrastive LearningImageTabularBiomedical DataElectronic Health Records
π― What it does: Proposed the HyperGP framework, transforming the multimodal survival prediction ofε ¨ζ―η ηεη (panoramic pathological slides) and genomic data into a hierarchical hyper-surface alignment problem.
π― What it does: Designed and verified a fetal ultrasound CHD abnormality detection framework based on cross-modal translation and regional calibration error.
π― What it does: Propose a motion denoising framework based on latent dictionary encoding, LDE-MAR, for eliminating artifacts in cardiac cine-MRI caused by arrhythmia.
AS-TIME: Time-Conditioned Multimodal Modeling of Aortic Stenosis Progression
Kim, Diane (University Of British Columbia), Abolmaesumi, Purang (Vancouver General Hospital)
CodeClassificationExplainability and InterpretabilityTransformerMixture of ExpertsVideoTextBiomedical DataUltrasound
π― What it does: Proposed a time-conditioned multimodal model, AS-TIME, which predicts the risk of aortic stenosis (AS) progression using a single baseline echocardiogram and time interval βt.
CodeExplainability and InterpretabilityRepresentation LearningData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextBiomedical DataMagnetic Resonance ImagingRetrieval-Augmented Generation
π― What it does: This paper proposes the ASTAR framework, which automatically induces standardized report templates from large-scale radiology free-text reports using large language models.
π― What it does: This paper proposes an asynchronous time-step diffusion framework, AsynDiff, for synthesizing missing sequences in multi-sequence MRI, and achieves cross-sequence information fusion through noise-aware attention.
AtomiMed: Hierarchical Atomic Fact-Checking for Universal Clinical-Aware Medical Report Evaluation
Wang, Yuan (Zhejiang University-University of Illinois Urbana-Champaign Institute, Zhejiang University), Zhang, Jianpeng (DAMO Academy, Alibaba Group)
CodeExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelAgentic AIPrompt EngineeringMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasoundBenchmarkRetrieval-Augmented Generation
π― What it does: Propose a multi-modal medical report generation evaluation framework called AtomiMed, which decomposes reports into atomic clinical facts at the disease level and attribute level, and verifies clinical consistency through bidirectional Agentic Cross-Verification.
π― What it does: Propose a deep learning-based MR-only synthetic CT generation framework specifically for handling images of patients with metal implants (especially hip replacements and intramedullary nails), automatically extracting metal masks from MR images and improving the reconstruction quality of metal regions through attention mechanisms.
π― What it does: This paper proposes a prototype calibration framework based on attention (JAPC), achieving multi-rater personalized prediction in few-shot medical image segmentation.
Attenuation-Resilient Alternating Optimization for Laparoscopic Liver Landmark Detection
Liu, Lanqing (Hong Kong Polytechnic University), Qin, Jing (Chinese University of Hong Kong)
CodeSegmentationOptimizationConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningGaussian SplattingImageMultimodalityComputed TomographyReview/Survey Paper
π― What it does: Propose A2ONet, combining illumination compensation, frequency-domain directional filtering, and alternating segmentation-curve optimization to enhance the robustness of detection for laparoscopic liver surface markings.
π― What it does: Propose AutoLand, an automated workflow based on Vision Language Models, which leverages physical fingerprints and a medical deep learning library to automatically generate feature point detection strategies suitable for new medical image datasets, significantly reducing the cost of manual parameter tuning.
Automated Disentangling Analysis of Skin Colour for Lesion Images
Yang, Wenbo (University of Waterloo), Wang, Zhou (University of Waterloo)
CodeClassificationImage TranslationRestorationAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical Data
π― What it does: Proposed an unsupervised skin color decoupling framework that can learn skin color captured in skin images (SCCI) and achieve controllable color conversion, enhancement, and normalization;
π― What it does: This paper proposes a method that simultaneously accomplishes left ventricle localization and three-dimensional SAX plane orientation estimation from a single arbitrary CMR slice.
π― What it does: This paper introduces a residual KAN head in the output layer of a medical image segmentation model, leveraging the local support property of B-spline basis functions in the KAN layer to directly extract activation dispersion as an uncertainty estimate in a single forward pass.
π― What it does: Developed a 7-degree-of-freedom laparoscopic tool pose tracking framework based on Bayesian Time Pose Network (BTPN), which does not require geometric priors and can quantify uncertainty.
π― What it does: Created the BC-MultiSet dataset and conducted benchmark evaluations of multi-task deep learning models on it, exploring the association between cell morphology and clinical molecular markers.
BCER Agent: Reliable Long-Horizon MRI Workflow Execution via Compilation, Artifact Binding, and Bounded Local Recovery
Long, Ziyang (Cedars-Sinai Medical Center), Yang, Hsin-Jung (Cedars-Sinai Medical Center)
CodeSegmentationGenerationData SynthesisTransformerLarge Language ModelAgentic AIPrompt EngineeringVision Language ModelImageBiomedical DataMagnetic Resonance ImagingBenchmarkRetrieval-Augmented GenerationChain-of-Thought
π― What it does: Developed the BCER framework, which can reliably execute long-cycle MRI workflows, supporting the separation of high-level planning and compilable execution, symbolic binding, and local recovery, and was evaluated on a multi-organ multi-task benchmark.
π― What it does: Proposed SpineRankNet, which learns continuous spinal degeneration severity scores using contrastive learning, and achieves multi-task prediction on a single 3D ResNet-18 architecture.
π― What it does: Constructed a large-scale benchmark, evaluating the performance of remote Parkinson's disease screening using seven video foundation models (VideoPrism, V-JEPA, V-JEPA-SSv2, ViViT, TimeSformer, VideoMAE, VideoMAEv2) on 32,847 videos from 1,888 subjects (727 with Parkinson's disease) performing 16 standard clinical tasks, systematically comparing the performance of different VFM architectures across tasks.
Better Said Than Seen: Exposing and Mitigating Modality Collapse with ICD-11-Grounded Evaluation
Hassan, Lara (Mohamed Bin Zayed University of Artificial Intelligence), Mahmoud, Abdulrahman (Mohamed Bin Zayed University of Artificial Intelligence)
CodeClassificationExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextMultimodalityBiomedical DataElectronic Health RecordsRetrieval-Augmented Generation
π― What it does: The study investigates the phenomenon where image information in multimodal large language models is suppressed by text priors in dermatology diagnosis, and proposes alleviating and evaluating this issue by replacing pixels with structured clinical vignettes and using an ICDLens assessment framework based on ICD-11.
π― What it does: Propose a reconstruction framework for fluorescent molecular tomography (FMT) named TCDFlow-FMT based on deterministic flow matching, achieving fast and high-quality three-dimensional fluorescent source reconstruction.
π― What it does: Propose the Momentum-Contrastive Survival Framework (MCSF), which expands the risk set in 3D PET-CT survival prediction through a virtual cohort and improves optimization stability using a time-weighted contrastive loss.
Beyond Visual Cues: CoT-Enhanced Reasoning for Semi-supervised Medical Image Segmentation
Chen, Yuming (Southeast University), Zhou, Yi (Nanjing University of Science and Technology)
CodeSegmentationRetrievalExplainability and InterpretabilityConvolutional Neural NetworkTransformerLarge Language ModelVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingComputed TomographyRetrieval-Augmented GenerationChain-of-Thought
π― What it does: This paper proposes a semi-supervised medical image segmentation framework called CERS based on Chain-of-Thought (CoT), which enhances segmentation performance by jointly retrieving generated diagnostic reasoning text and visual features.
Beyond Visual Forensics: Auditing Multimodal Robustness for Synthetic Medical Image Detection
Chiu, Ching-Hao (University of Notre Dame), Shi, Yiyu (University of Notre Dame)
CodeAnomaly DetectionExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextBiomedical DataComputed TomographyElectronic Health RecordsBenchmark
π― What it does: Proposed an evaluation benchmark for detecting the authenticity of multimodal medical images, by fixing the image and only modifying the accompanying structured metadata to detect text-induced decision bias.
π― What it does: Propose the BiC-ODE framework for fast and high-precision reconstruction of four surface layers (white matter, inner surface, outer surface, pial surface) in 5T FLAIR images.
Bidirectional Anatomy-Aware Post-Training for Longitudinal Chest X-Ray Progression Modeling
Chen, Yuming (Adelaide University), Xie, Yutong (Mohamed bin Zayed University of Artificial Intelligence)
CodeClassificationRepresentation LearningTransformerSupervised Fine-TuningContrastive LearningImageBiomedical DataComputed TomographyElectronic Health Records
π― What it does: Proposed a lightweight post-training strategy called BAAP to improve longitudinal chest radiograph progression prediction;
π― What it does: Propose the BIGUS framework, which utilizes 3D Gaussian splats to achieve continuous ultrasound view synthesis and accurately simulate ultrasound scattering to preserve speckle details.
π― What it does: Propose BiM-GeoAttn-Net, combining bidirectional deep Mamba with geometry-aware vessel attention for aortic dissection segmentation in 3D CTA.
BioFact-MoE: Biologically Factorized Mixture of Experts for VisionβLanguage Prognostic Modeling in Hepatocellular Carcinoma
Yang, Junlin (Yale University), Chapiro, Julius (Yale University)
CodeExplainability and InterpretabilityDrug DiscoveryTransformerLarge Language ModelMixture of ExpertsVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance Imaging
π― What it does: Propose the BioFact-MoE framework, utilizing a Mixture of Experts at the gene level for visual-language prognostic modeling in hepatocellular carcinoma.
π― What it does: Proposed a flow matching model called BioFlow that maintains the non-negativity of gene expression during the generation process, specifically designed for predicting spatial transcriptomics data from tissue slice images.
π― What it does: Propose the BioGuide framework, which adds a lightweight MLP as an auxiliary task at the bottleneck of the encoder in medical image segmentation models, leveraging structured clinical descriptions (such as lesion location and radiological features) to guide feature learning.
Boosting Ultrasound Image Classification via Attribute-Guided Dual-Branch Framework
Zhao, Bo (Wuhan University), Du, Bo (Wuhan University)
CodeClassificationExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataUltrasound
π― What it does: Proposed a pluggable attribute-guided dual-branch framework to enhance the accuracy and interpretability of ultrasound image classification.
CodeClassificationExplainability and InterpretabilityRepresentation LearningTransformerSupervised Fine-TuningContrastive LearningTextMultimodalityAudio
π― What it does: Propose a boundary-aware multi-grained learning framework that combines continuous regression with coarse-grained ordinal supervision to estimate depression severity.
CodeClassificationAnomaly DetectionTransformerSupervised Fine-TuningVision Language ModelContrastive LearningImageTextBiomedical DataComputed Tomography
π― What it does: Propose the Brain-Adapter dual-stream MIL framework, utilizing 2D vision-language models and diagnostic reports to achieve multi-label diagnosis of 3D CT pathology.
π― What it does: Developed a diffusion Transformer base model called Brain-DiT, which is pre-trained using multi-state fMRI data (rest, task, natural stimulation, disease, sleep, etc.), and incorporates individual metadata for conditional learning during the pre-training phase;
CodeSegmentationTransformerContrastive LearningOptical FlowVideoBiomedical Data
π― What it does: Proposed the BRAIN framework to address segmentation errors caused by sudden instrument displacement and long-term disappearance in surgical videos.
π― What it does: A dynamic brain network asynchronous coordination unit framework called BrainACU for psychiatric disease diagnosis using rs-fMRI was studied.
π― What it does: Developed a unified pre-training framework called BrainAnytime, which can perform brain image analysis under any available modality subset (multi-sequence MRI and amyloid-PET), supporting seamless switching from single-modality to multi-modality.
π― What it does: Designed and implemented a dynamic brain network modeling framework called BrainSTR based on spatiotemporal contrastive learning, for the diagnosis and interpretability analysis of neuropsychiatric disorders.
BrainWeaver: Weaving Dynamic Brain Networks with ROI-Guided Attention for Cognitive Assessment
Liu, Tao (Hangzhou City University), Zhou, Binbin (Hangzhou City University)
CodeExplainability and InterpretabilityComputational EfficiencyGraph Neural NetworkTransformerContrastive LearningGraphBiomedical DataMagnetic Resonance Imaging
π― What it does: Proposed the BrainWeaver framework, which utilizes ROI-guided dynamic multi-graph attention fusion, hierarchical graph representation, and individual-invariant contrastive learning to achieve interpretable prediction of cognitive functions.
π― What it does: Proposed the BREIT framework to generate frequency-dependent conductivity voxels from CT/MRI, providing a Python 3D CEM forward solver and 3D D-bar implementation, and developed the dFNO-bar learning-based reconstruction method based on this.
Bridging Heterogeneous Medical Datasets via Mixture-of-Specialists Adapters for Unified Medical Image Classification
Huang, Shixing (University of Sydney)
CodeClassificationTransformerSupervised Fine-TuningMixture of ExpertsContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound
π― What it does: Propose a parameter-efficient framework called MOSAIC, which freezes the ViT backbone and utilizes a modality-aware tokenizer and a hybrid expert adapter to unify 18 different modalities of medical imaging data (2D/3D) into a single model, addressing the problem of cross-modal feature entanglement.
Bridging Research and Practice: A Systematic Evaluation of Generalist and Dermatology-Specific Models in Clinical Skin Lesion Classification
dos Santos, Emanoel (Universidade Federal de Pernambuco), Ing Ren, Tsang (Universidade Federal de Pernambuco)
CodeClassificationAnomaly DetectionConvolutional Neural NetworkTransformerSupervised Fine-TuningPrompt EngineeringVision Language ModelContrastive LearningImageMultimodalityBiomedical DataElectronic Health RecordsBenchmark
π― What it does: Systematically evaluated the performance of various general-purpose and dermatology-specific vision and vision-language models in binary classification of malignant risk prediction across different skin disease image datasets.
π― What it does: A few-shot dual-parameter proton imaging quality assessment network is proposed, which can automatically identify DWI geometric distortions and transfer to clinical PI-QUAL scoring under conditions of limited annotation.
Bridging Vision Foundation Model Priors with CLIP for Spatial-aware Few-shot Anomaly Detection in Medical Images
Miao, Juzheng (Chinese University of Hong Kong), Heng, Pheng-Ann (Chinese University of Hong Kong)
CodeAnomaly DetectionTransformerVision Language ModelContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
π― What it does: Propose a Spatial-FAD framework that combines the spatial prior of Vision Foundation Model (DINO) with CLIP for improving the spatial localization accuracy of anomaly detection in medical imaging with few samples.
CΒ²RM-Seg: Causal Counterfactual Reasoning with Structural-Semantic Priors for Weakly Supervised Histopathological Tissue Segmentation
Zhang, Hualong (Guilin University of Electronic Technology), Pan, Xipeng (Guilin University of Electronic Technology)
CodeSegmentationExplainability and InterpretabilityComputational EfficiencyConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningBiomedical DataMagnetic Resonance ImagingComputed TomographyReview/Survey Paper
π― What it does: Propose a weakly supervised organ segmentation framework, CRM-Seg, based on causal counterfactual reasoning and structural semantic priors. First, generate deconfounded pseudo labels through a causal module, and then perform segmentation using a dual-channel structural-semantic network, optimized by an uncertainty-gated margin loss.
π― What it does: Built CALHippo, a multi-scale human hippocampal CA region cell type segmentation and density inference resource, covering all subregions from CA1 to CA4.
Calibrated Confidence Expression for Radiology Report Generation
Bani-Harouni, David (Technical University of Munich), Keicher, Matthias (Technical University of Munich)
CodeExplainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringVision Language ModelImageTextMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyElectronic Health Records
π― What it does: Propose ConRad, a reinforcement learning-based framework that enables medical large vision-language models to generate radiology reports along with calibrated self-confidence expressions;
CodeClassificationExplainability and InterpretabilityDrug DiscoveryTransformerMixture of ExpertsAuto EncoderGenerative Adversarial NetworkContrastive LearningImageTabularBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease
π― What it does: Designed and implemented the CALM framework, which maps structural MRI and genomic data from completely non-overlapping populations into a shared latent space via linear projection, enabling the discovery of interpretable brain region-gene pathway associations without paired data, and using these associations for diagnostic prediction of autism spectrum disorder (ASD).
Can Agents Distinguish Visually Hard-to-Separate Diseases in a Zero-Shot Setting? A Pilot Study
Zhao, Zihao (University Hospital Aachen), Truhn, Daniel (University Hospital Aachen)
CodeClassificationAnomaly DetectionExplainability and InterpretabilityTransformerLarge Language ModelAgentic AIPrompt EngineeringVision Language ModelImageTextBiomedical DataElectronic Health RecordsBenchmarkChain-of-Thought
π― What it does: A zero-shot diagnostic benchmark for multi-modal large language models (MLLMs) is conducted on visually indistinguishable diseases (melanoma vs. atypical nevi, pulmonary edema vs. pneumonia), and a multi-agent contrastive reasoning framework called CARE is proposed.
CodeImage TranslationRepresentation LearningData-Centric LearningDrug DiscoveryGraph Neural NetworkTransformerMixture of ExpertsAuto EncoderContrastive LearningImageBiomedical Data
π― What it does: This paper proposes a pan-cancer image gene expression prediction framework called PIGP, which can directly infer gene expression levels from H&E tissue sections without prior knowledge of cancer types.
π― What it does: Proposed CardioDiT, a 4D latent diffusion transformer for generating short-axis cardiac MRI (cine CMR) data from scratch, where the model simultaneously models spatial and temporal dimensions in the latent space;
CARE: Clinically Aligned Retrieval Evidence for Consistent Radiology Report Generation
Zhu, Jintao (Zhejiang Cas Angels Biotechnology Co Ltd), Ma, Jiaxin (Zhejiang Cas Angels Biotechnology Co Ltd)
CodeGenerationRetrievalDomain AdaptationTransformerLarge Language ModelContrastive LearningImageTextBiomedical DataComputed TomographyElectronic Health RecordsRetrieval-Augmented Generation
π― What it does: This paper proposes CARE, a radiology report generation framework based on retrieval augmentation, which maintains a shared retrieval memory queue based on a momentum encoder to achieve temporal consistency in retrieval supervision, and introduces a disease alignment loss to ensure consistency between the retrieved text and image diagnosis, thus improving the clinical accuracy and diversity of the generated reports.
CARE: Confidence-Aware Reasoning for Reliable Medical VQA
Du, Yuetian (Zhejiang University), Zhu, Qiang (Zhejiang University)
CodeExplainability and InterpretabilityComputational EfficiencyData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision Language ModelImageTextMultimodalityBiomedical DataElectronic Health RecordsBenchmarkRetrieval-Augmented GenerationChain-of-Thought
π― What it does: Propose the CARE framework, which first automatically synthesizes structured Medical-CoT data, and then performs two-stage fine-tuning of the medical vision-and-language model using confidence-aware reinforcement learning (GRPO+CAR), to improve diagnostic accuracy and confidence calibration.
π― What it does: A CARformer framework is proposed and implemented for multi-class classification from structural MRI in the context of few-shot, class-imbalanced psychiatric disease diagnosis.
CodeClassificationAnomaly DetectionExplainability and InterpretabilityComputational EfficiencyRepresentation LearningDrug DiscoveryConvolutional Neural NetworkTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsContrastive LearningImageTextMultimodalityTabularBiomedical DataElectronic Health RecordsBenchmarkRetrieval-Augmented Generation
π― What it does: Utilize large language models to generate case-specific priors to guide multi-modal fusion for predicting 5-year overall survival and 2-year recurrence-free survival in patients with head and neck cancer.
π― What it does: This paper proposes a pre-training framework called CDPM-Align, based on conditional diffusion models and multi-scale guided alignment, for achieving reliable anatomical landmark detection under conditions of very limited annotation.
Cell-Division-Aware Category-Enhanced Contrastive Learning Model for Embryonic Cleavage Stage Classification
Li, Yukun (Shenzhen University), Pei, Jihong
CodeClassificationConvolutional Neural NetworkTransformerContrastive LearningOptical FlowTime SeriesBiomedical Data
π― What it does: This paper proposes a classification framework for the embryo cleavage stage called 3CL-ECSC, based on cell division awareness and category-enhanced contrastive learning, for automatically identifying different developmental stages of embryos;
CodeClassificationReinforcement Learning from Human FeedbackTransformerReinforcement LearningVision Language ModelImageBiomedical DataRetrieval-Augmented GenerationChain-of-Thought
π― What it does: This paper proposes a vision reasoning framework based on reinforcement learning called CerviThink, aimed at improving the classification accuracy of cervical cancer cell images.
π― What it does: The authors introduce self-regressive token space dissection supervision into a pre-trained vision-language model (VLM), enabling the model to directly generate anatomical segmentation masks for chest X-rays without requiring an additional pixel-level decoder.
CHILD: Human-in-the-Loop OOD Detection for Safe Clinical Deployment
Ye, Jinlun (Sun Yat-sen University), Wang, Ruixuan (Sun Yat-sen University)
CodeAnomaly DetectionSafty and PrivacyContrastive LearningImageBiomedical DataElectronic Health RecordsRetrieval-Augmented Generation
π― What it does: Proposes a training-agnostic, sparse physician feedback-based clinical OOD detection framework called CHILD, achieving risk-aware sample selection and retrieval-based score calibration in streaming deployment environments.
CodeGraph Neural NetworkTransformerMultimodalityGraphBiomedical DataElectronic Health RecordsBenchmark
π― What it does: Developed ChronoSurv, a heterogeneous hierarchical directed graph framework based on clinical pathways, for multimodal survival prediction in head and neck cancer.
CIGTSurv: Clinical Information Guided Tri-Modal Survival Prediction with Local Prototype Association and Global Feature Alignment
Dai, Jing (Dalian University of Technology), Xu, Hongming (Dalian University of Technology)
CodeClassificationRepresentation LearningDrug DiscoveryConvolutional Neural NetworkRecurrent Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelContrastive LearningImageTextTabularBiomedical DataElectronic Health Records
π― What it does: Propose the CIGTSurv framework, which uses clinical information guided fusion of three modalities (pathological images, gene expression, clinical tables) to predict the survival period of cancer patients.
CodeGenerationData SynthesisFederated LearningSafty and PrivacyExplainability and InterpretabilityTransformerPrompt EngineeringDiffusion modelScore-based ModelImageBiomedical DataComputed TomographyElectronic Health Records
π― What it does: Proposed a diffusion generation framework based on causal intervention called CIPHER, aimed at eliminating performance differences among sensitive subgroups in medical image diagnosis
ClinCoT: Clinical-Aware Visual Chain-of-Thought for Medical Vision Language Models
Liu, Xiwei, Xie, Yutong (Mohamed bin Zayed University of Artificial Intelligence)
CodeOptimizationExplainability and InterpretabilityRepresentation LearningDrug DiscoveryTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision Language ModelImageTextBiomedical DataElectronic Health RecordsChain-of-Thought
π― What it does: Propose the ClinCoT framework, extending preference optimization from only improving at the final response level to the clinical visual chain-of-thought level; achieving visual-driven alignment between intermediate reasoning and final answers through automatically generating region proposals based on disease hypotheses, generating intermediate reasoning chains for each region, constructing preference pairs with multi-model consensus weighted scoring, and using margin-aware DPO along with iterative learning.
π― What it does: Proposes a clinically oriented hierarchical multi-label classification framework called CHASE, which achieves hierarchical diagnosis from coarse to fine based on chest X-rays.