MICCAI 2026 Papers — Page 4
International Conference on Medical Image Computing and Computer-Assisted Intervention · 1165 papers
DPMix-GAN: Dual-Domain Perception and Dynamic Patch-Wise Mixing GAN for 3D CT Synthesis in MR-only Radiotherapy
Li, Jiapeng (Hebei Medical University), Feng, Hongying (China Three Gorges University)
GenerationData SynthesisTransformerDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Proposed and implemented a MR-CT synthesis framework based on DPMix-GAN, capable of generating synthetic CT images with high skeletal boundary clarity for MR-only radiotherapy.
DreamReg: Belief-Driven World Model for 2D–3D Ultrasound Registration
Kang, Luoyao (Chinese University of Hong Kong), Cheng, Shing Shin (Chinese University of Hong Kong)
Pose EstimationConvolutional Neural NetworkRecurrent Neural NetworkWorld ModelOptical FlowBiomedical DataUltrasound
🎯 What it does: Propose a 2D-3D ultrasound registration framework called DreamReg based on a world model, which achieves real-time pose estimation through continuous belief updates.
DRGFuse: Alignment-Guided Dual-Layer Reliability-Gated CT–WSI Fusion for Preoperative TRG0 Vs TRG1–3 Prediction in Esophageal Cancer
Liu, Zhenbing (Guilin University of Electronic Technology), Lu, Haoxiang (Guilin University of Electronic Technology)
ClassificationSegmentationExplainability and InterpretabilityComputational EfficiencyData-Centric LearningConvolutional Neural NetworkTransformerMixture of ExpertsDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageMultimodalityBiomedical DataComputed TomographyReview/Survey Paper
🎯 What it does: This paper proposes a dual reliability-gated fusion framework, DRGFuse, for predicting tumor regression grades TRG0 and TRG1-3 in esophageal cancer using preoperative CT and biopsy whole slide images (WSI).
DSCE-Net: Dual-Scale Conditioned Evolving Network with Balanced Optimization for All-in-One Medical Image Restoration
Hu, Xulin (Guizhou University), Wang, Lihui (Guizhou University)
RestorationSuper ResolutionConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningImageMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission Tomography
🎯 What it does: Propose a fully functional medical image restoration network called DSCE-Net, which can restore multi-modal and multi-task images under a single model.
DTI-Guided Volumetric Spherical Harmonics Regression for Single-to-Multi-shell dMRI Synthesis
Li, Binghua (Juntendo University), Aoki, Shigeki (Juntendo University)
Data SynthesisConvolutional Neural NetworkDiffusion modelScore-based ModelBiomedical DataMagnetic Resonance ImagingDiffusion Tensor Imaging
🎯 What it does: Propose a voxel-level spherical harmonic regression network based on diffusion tensor prior (DTI-SHNet), achieving the synthesis of multi-shell diffusion MRI data from single-shell sampling;
Dual Agreement Consistency Learning for Semi-supervised Fetal Ultrasound Segmentation
Wang, Fangyijie (University College Dublin), Curran, Kathleen M. (University College Dublin)
SegmentationConvolutional Neural NetworkTransformerContrastive LearningImageBiomedical DataUltrasound
🎯 What it does: Proposed a dual consistency learning framework (DACL) for fetal ultrasound image segmentation with very few annotations.
Dual-Adaptive SAM3: Hierarchical Routing over Low-Rank Expert Layers for Parameter-Efficient Medical Image Segmentation
Chen, Ying, Liu, Yang (Nanyang Technological University)
SegmentationExplainability and InterpretabilityComputational EfficiencyTransformerPrompt EngineeringMixture of ExpertsContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Propose the Dual-Adaptive SAM3 framework, which inserts a hierarchical Mixture-of-Experts into the fusion module of SAM3, combining dynamic expert routing and low-rank expert decomposition to achieve efficient and interpretable adaptation for medical image segmentation.
Dual-Domain Cross-Modal Decoding for Clinical Text-Guided Medical Image Segmentation
Rahman, Md Maklachur (Texas A&M University), Hammond, Tracy (Texas A&M University)
SegmentationData-Centric LearningConvolutional Neural NetworkTransformerVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Proposed a dual-domain cross-modal decoder (DD-CMD) for lung infection segmentation guided by clinical text.
Dual-Domain Meta Conditioning Network for Fetal Head and Pubic Symphysis Segmentation in Ultrasound Images Analysis
Hu, Xinglong (Wuhan University), Lei, Cheng (Wuhan University)
SegmentationConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningImageBiomedical DataUltrasound
🎯 What it does: This paper proposes DMC-Net, a dual-domain meta-conditional network, for segmenting fetal heads and the pubic symphysis in obstetric ultrasound images to achieve accurate progress angle estimation.
Dual-Modality Neural Network Integrating Sequence and Structural Information for Drug-Target Interaction Prediction
Chen, Dongjie (Foshan University), Yang, Zhihui (Foshan University)
Drug DiscoveryGraph Neural NetworkTransformerSupervised Fine-TuningContrastive LearningGraphSequentialBiomedical Data
🎯 What it does: Proposed a dual-modal neural network called MGK-DTI, which integrates sequence and structural information of drugs and targets for drug-target interaction prediction.
Dual-Registration Reciprocal Super-Resolution for Prior-Guided Real-Time 4D-MRI
Wang, Yinghui (Hong Kong Polytechnic University), Cai, Jing (Northeastern University)
RestorationSuper ResolutionConvolutional Neural NetworkRecurrent Neural NetworkAuto EncoderGenerative Adversarial NetworkContrastive LearningOptical FlowImageBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Utilizing bidirectional registration and super-resolution techniques, low-quality real-time 4D-MRI is reconstructed into high-quality images, with the integration of anatomical priors from diagnostic T1-weighted images and pseudo-labels from high-resolution scans during apnea throughout the process.
Dual-Space Cold-Start Active Learning Guided by SAM3 for Medical Image Segmentation
Ye, Ping, Wang, Guotai (University Of Electronic Science And Technology Of China)
SegmentationTransformerAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Proposes a framework utilizing zero-cost pseudo masks generated by SAM3 for dual-space active learning in the cold-start scenario of medical image segmentation.
Dual-Stage Course-to-Fine Tooth Alignment Framework with Decoupled Position, Orientation and Shape Features
Jin, Xiaoxian (Dalian University of Technology), Wang, Hongkai (Dalian University of Technology)
Pose EstimationData-Centric LearningRecurrent Neural NetworkTransformerDiffusion modelScore-based ModelContrastive LearningPoint CloudBiomedical DataComputed Tomography
🎯 What it does: Propose a two-stage coarse-to-fine teeth alignment framework that separates and progressively optimizes the position, orientation, and shape.
Dual-Stream Diffusion with Structure-Aware Fusion for Joint Synthesis of Coherent Medical Image-Mask Pairs
Li, Yu (Ritsumeikan University), Chen, Yen-Wei (Ritsumeikan University)
SegmentationGenerationData SynthesisTransformerDiffusion modelScore-based ModelBiomedical Data
🎯 What it does: Propose a Dual-Stream Diffusion framework that jointly generates medical images and corresponding segmentation masks using structural stream and appearance stream, respectively.
Dual-Teacher Knowledge Distillation with SAM and Learnable Prototype Guidance for Semi-supervised 3D Medical Segmentation
Huo, Shuaiguang (Heilongjiang University), Xi, Heran (Heilongjiang University)
SegmentationKnowledge DistillationTransformerContrastive LearningBiomedical DataComputed TomographyBenchmark
🎯 What it does: Propose a dual-teacher knowledge distillation framework that combines an internal EMA teacher with an external SAM-Med3D teacher, using learnable prototypes to guide semi-supervised 3D medical segmentation.
DualCaMIL: Dual-Level Causal Multi-instance Learning for Patient-Level Diagnosis in Reflectance Confocal Microscopy
Wang, Changxin (Jiangnan University), Zou, Yunmin (Jiangnan University)
ClassificationAnomaly DetectionExplainability and InterpretabilityConvolutional Neural NetworkTransformerAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical Data
🎯 What it does: Propose the DualCaMIL framework, which decomposes RCM multi-instance diagnosis into micro cross-modal causal disentanglement and macro patient-level causal modeling to address imaging and patient-level biases.
DualGGR-MoME: A Mixture of Multimodal Experts with Dual-Branch Gating and Global-Guided Routing for Cancer Survival Prediction
Tan, Kaiwen (Kunming University of Science and Technology), Yu, Zhengtao (Kunming University of Science and Technology)
ClassificationAnomaly DetectionDrug DiscoveryConvolutional Neural NetworkSpiking Neural NetworkTransformerMixture of ExpertsContrastive LearningImageMultimodalityTabularBiomedical DataElectronic Health Records
🎯 What it does: Propose DualGGR-MoME, a dual-branch gated and globally guided routing mixture-of-experts model for cancer survival prediction.
DUCX: Decomposing Unfairness in Tool-Using Chest X-Ray Agents
Xu, Zikang (Hefei Comprehensive National Science Center), Li, Xiaoxiao (University of British Columbia)
Safty and PrivacyExplainability and InterpretabilityTransformerLarge Language ModelAgentic AIImageTextBiomedical DataComputed TomographyBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Systematically evaluate the fairness of a tool-based chest X-ray image QA agent system, and propose a staged fairness decomposition framework.
Dynamic Collaborative Continual Test-Time Adaptation for 3D Vessel Segmentation
Li, Xiang (Nanjing University), Shan, Caifeng (Nanjing University)
SegmentationDomain AdaptationConvolutional Neural NetworkTransformerAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Propose a dynamic collaborative continuous test-time adaptation framework, DyCo-CTA, for online updating of 3D vascular segmentation models under cross-center domain shift, reducing topological destruction and pseudo branches.
Dynamic Spatial-Temporal Mapping for EEG-to-fMRI Synthesis via Gaussian-Guided Heterogeneous Graph Neural Network
Zheng, Jialun (Hong Kong Polytechnic University), Saxena, Divya (Indian Institute of Technology Jodhpur)
Data SynthesisGraph Neural NetworkTransformerDiffusion modelContrastive LearningBiomedical DataMagnetic Resonance ImagingElectrocardiogram
🎯 What it does: Propose the DSTP framework, which utilizes dynamic heterogeneous graph neural networks and learnable Gaussian masks to achieve dynamic spatiotemporal mapping from EEG to fMRI, thereby generating high spatial resolution fMRI.
Dynamic Sub-domain Modeling for Robust Medical Image Segmentation
Lee, Kyungsu (Jeonbuk National University), Woo, Jonghye (Massachusetts General Hospital and Harvard Medical School)
SegmentationConvolutional Neural NetworkTransformerPrompt EngineeringMixture of ExpertsContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasoundReview/Survey Paper
🎯 What it does: Achieve robustness in medical image segmentation without gradient updates through dynamic subdomain modeling, significantly improving boundary accuracy and internal variability;
DynFS-MoE: Dynamic Functional-Structural Mixture-of-Experts for Post-Traumatic Epilepsy Diagnosis
Ding, Jun-En (Rutgers University), Liu, Feng (Rutgers University)
ClassificationAnomaly DetectionExplainability and InterpretabilityComputational EfficiencyGraph Neural NetworkMixture of ExpertsContrastive LearningMultimodalityBiomedical DataMagnetic Resonance ImagingDiffusion Tensor Imaging
🎯 What it does: This study proposes a multi-modal fusion framework based on dynamic mixture of experts (DynFS-MoE) for the early diagnosis of epilepsy after traumatic brain injury.
E-MRL: Cross-View Aligned Evidence-Driven Multimodal Reinforcement Learning for Reliable 3D Tumor Analysis
Li, Sijing (Zhejiang University), Zhang, Ling (DAMO Academy, Alibaba Group)
Explainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision Language ModelImageBiomedical DataComputed Tomography
🎯 What it does: This paper proposes the E-MRL framework, which combines cross-perspective aligned evidence-driven multi-modal reinforcement learning for 3D CT tumor analysis; the framework decomposes report generation into three steps: global diagnosis, key slice localization, and local verification, and introduces cross-perspective consistency into the reinforcement learning reward to achieve verifiability of visual evidence;
EA-LDM: Explicit Alignment Latent Diffusion Model for Patient-Specific CT Reconstruction from Bi-planar X-Rays
Xiao, Jiewen (Dalian University of Technology), Fan, Xin (Dalian University of Technology)
OptimizationConvolutional Neural NetworkTransformerDiffusion modelNeural Radiance FieldAuto EncoderGenerative Adversarial NetworkContrastive LearningOptical FlowImageBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasound
🎯 What it does: Achieve high-quality patient-specific CT volume reconstruction from biplanar X-ray using explicit alignment of latent diffusion models (EA-LDM) and patient-specific optimization (PSO);
ECFR: Entropy-Consistency Flow Rectification in Test-Time Adaptation for Medical Foundation Models
Li, Shangkun (Shanghai Jiao Tong University), Zhang, Hanxiao (Shanghai Jiao Tong University)
Domain AdaptationAnomaly DetectionTransformerSupervised Fine-TuningFlow-based ModelContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasoundReview/Survey PaperStochastic Differential Equation
🎯 What it does: Propose the ECFR method, achieving safe and real-time adaptation of medical foundation models through bidirectional optimization based on entropy and consistency quadrants.
ECG-CoCa: Physiology-Informed Contrastive and Generative Framework for ECG Interpretation
Liu, Xinyue (Beihang University), Liu, Qingjie (Beihang University)
GenerationRetrievalExplainability and InterpretabilityRepresentation LearningConvolutional Neural NetworkTransformerPrompt EngineeringGenerative Adversarial NetworkContrastive LearningTextMultimodalityTime SeriesBiomedical DataElectronic Health RecordsElectrocardiogramRetrieval-Augmented Generation
🎯 What it does: This paper proposes ECG‑CoCa, a unified framework that integrates contrastive learning and generative learning, directly interpreting raw ECG signals and generating diagnostic reports.
Echo-SCAR: Temporal Feature Learning for Myocardial Scar Detection from Routine Echocardiography
Shehzad, Mansoor (Rice University), Sabharwal, Ashutosh (Rice University)
ClassificationSegmentationExplainability and InterpretabilityRepresentation LearningConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingUltrasound
🎯 What it does: Developed a multi-stage pipeline supervised by CMR to perform segment-level detection of myocardial fibrosis using conventional echocardiography;
ECHO: Estimated Composite Hybrid Observation for Zero-Shot Self-Supervised MRI Reconstruction
Shang, Wenlei (ShanghaiTech University), Zhou, Zijian (ShanghaiTech University)
RestorationConvolutional Neural NetworkAuto EncoderContrastive LearningBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Propose an MRI reconstruction framework called ECHO based on zero-shot self-supervised learning, which directly optimizes network parameters on a single undersampled acquisition, avoiding the training-inference mismatch in traditional methods.
Echo2ECG: Enhancing ECG Representations with Cardiac Morphology from Multi-view Echos
Liman, Michelle Espranita (Technical University of Munich), Müller, Philip (Technical University of Munich)
ClassificationRetrievalRepresentation LearningTransformerContrastive LearningMultimodalityUltrasoundElectronic Health RecordsElectrocardiogram
🎯 What it does: Propose a multi-modal self-supervised framework called Echo2ECG, which aligns the cardiac morphology information from multi-view echocardiography (Echo) to electrocardiogram (ECG), thereby generating ECG representations rich in structural information for subsequent structural cardiac phenotyping and cross-modal retrieval.
Echo4DIR: 4D Implicit Heart Reconstruction from 2D Echocardiography Videos
Liu, Yanan (National University of Singapore), Li, Lei (Yunnan University)
Image TranslationGenerationDepth EstimationRecurrent Neural NetworkTransformerDiffusion modelNeural Radiance FieldAuto EncoderOptical FlowImageVideoBiomedical DataUltrasound
🎯 What it does: This study proposes Echo4DIR, an implicit framework for real-time reconstruction of 4D cardiac geometry from sparse 2D ultrasound views.
EchoFID: A Multi-teacher Distilled Evaluation Network for Cardiac Ultrasound Image Generation
Wu, Junde (University of Oxford), Grau, Vicente (University of Oxford)
GenerationKnowledge DistillationRepresentation LearningConvolutional Neural NetworkTransformerDiffusion modelGenerative Adversarial NetworkContrastive LearningImageBiomedical DataUltrasound
🎯 What it does: Propose EchoFID, an evaluation network based on multi-teacher distillation, used to measure the quality of cardiac ultrasound image generation models.
EchoLVFM: One-Step Video Generation via Latent Flow Matching for Echocardiogram Synthesis
Oladokun, Emmanuel (University of Oxford), Grau, Vicente (University of Oxford)
GenerationData SynthesisTransformerDiffusion modelScore-based ModelFlow-based ModelAuto EncoderOptical FlowVideoBiomedical DataUltrasound
🎯 What it does: Propose EchoLVFM, a one-shot latent video stream matching framework for controllable cardiac ultrasound video generation.
EchoPilot: Training-Free Ultrasound Video Segmentation via Scale-Space Semantic Prompting and Reliability-Gated Memory
Xiao, Ruiqiang (Hong Kong University Of Science And Technology Guangzhou), Zhu, Lei (Hong Kong University Of Science And Technology Guangzhou)
SegmentationTransformerPrompt EngineeringVision Language ModelContrastive LearningVideoBiomedical DataUltrasound
🎯 What it does: Proposed the EchoPilot framework, achieving ultrasound video segmentation without training, requiring only a single click and the name of the anatomical category.
EchoTracker2: Enhancing Myocardial Point Tracking by Modeling Local Motion
Azad, Md Abulkalam (Norwegian University of Science and Technology), Østvik, Andreas (Norwegian University of Science and Technology)
Object TrackingConvolutional Neural NetworkTransformerContrastive LearningOptical FlowVideoBiomedical DataUltrasound
🎯 What it does: Propose an EchoTracker2 model that tracks myocardial points using only the refinement phase, focusing on local motion to achieve pixel-level accurate tracking.
EchoVQA: Enabling Conversational Assistance for Point-of-Care Cardiac Ultrasound
Bellos, Filippos (University of Michigan), Corso, Jason J. (University of Michigan)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataUltrasound
🎯 What it does: Proposed the EchoVQA large-scale echocardiography visual question answering dataset, and introduced a parameter-efficient multi-layer prompting learning method;
ECLIPSE: EHR-Constrained Localized Inference for Pan-Tumor Segmentation
Meng, Runqi (ShanghaiTech University), Shen, Dinggang (ShanghaiTech University)
SegmentationConvolutional Neural NetworkTransformerPrompt EngineeringVision Language ModelImageTabularBiomedical DataComputed TomographyElectronic Health Records
🎯 What it does: This study proposes the ECLIPSE framework, which utilizes tumor burden information from electronic health records to perform localized segmentation of cross-organ tumors.
EdgeCVS: Democratization of surgical AI with a Distilled Edge-Deployable Critical View of Safety (CVS) model
Yamlahi, Amine (German Cancer Research Center), Maier-Hein, Lena (German Cancer Research Center)
ClassificationSegmentationComputational EfficiencyKnowledge DistillationRepresentation LearningConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningVideoBiomedical Data
🎯 What it does: The study proposes the EdgeCVS framework, which compresses a 305M parameter teacher model into a 5M parameter edge model through knowledge distillation, spatial priors, and stage-limited pseudo-labels, achieving real-time evaluation of critical safety views (CVS) during laparoscopic cholecystectomy.
EEGRFusion: Uncertainty-Aware EEG-to-Image Evidence for Adjunct Bedside Assessment of CMD in Disorders of Consciousness
Hong, Chenyuan (South China University of Technology), Xu, Yanwu (Pazhou Lab)
RestorationGenerationRetrievalExplainability and InterpretabilityTransformerVision Language ModelDiffusion modelScore-based ModelRectified FlowContrastive LearningImageTime SeriesBiomedical DataElectrocardiogram
🎯 What it does: Propose the EEGRFusion framework, which retrieves and generates images consistent with the stimulus using EEG signals in the CLIP visual-semantic space, providing interpretable and uncertainty-aware visual evidence for bedside assessment of patients with disorders of consciousness.
Efficient Conformal Volumetry for Template-Based Segmentation
Cheung, Matt Y. (Rice University), Balakrishnan, Guha (Rice University)
SegmentationExplainability and InterpretabilityComputational EfficiencyContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Propose ConVOLT, a conformal prediction (CP) framework that leverages deformation field features, within a template-based segmentation pipeline, for generating effective confidence intervals for volumetric metrics;
Efficient Flow Matching for Sparse-View CT Reconstruction
Shi, Jiayang (Centrum Wiskunde en Informatica), Batenburg, K. Joost (Leiden University)
RestorationConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelFlow-based ModelRectified FlowBiomedical DataComputed TomographyOrdinary Differential Equation
🎯 What it does: Propose a CT reconstruction framework based on flow matching, FMCT, and its efficient version EFMCT, which reduces the number of network evaluations by reusing the velocity field, achieving reconstruction quality comparable to diffusion models in sparse-view CT reconstruction while significantly improving inference efficiency.
Eigen-Directed Random Projection for Medical Image Class Incremental Learning
Wu, Yongyi (Xi'an Jiaotong University), Wang, Hong (Xi'an Jiaotong University)
ClassificationComputational EfficiencyRepresentation LearningData-Centric LearningTransformerContrastive LearningImageBiomedical Data
🎯 What it does: Propose a training-agnostic, dual-view incremental learning framework named Eigen-Directed Random Projection (EDRP), which utilizes data-driven principal component projection instead of traditional random projection, significantly enhancing the discriminability of fine-grained medical image categories, and updates the classifier with closed-form Ridge regression on pre-trained features.
EmERGE: A Multimodal LLM-Based Evidence-Mediated ECG Report Generation and Evaluation Approach
Yao, Keyu (Sun Yat-sen University), Shen, Gang (Sun Yat-sen University)
ClassificationExplainability and InterpretabilityComputational EfficiencyData-Centric LearningConvolutional Neural NetworkRecurrent Neural NetworkLarge Language ModelContrastive LearningImageTextBiomedical DataElectrocardiogram
🎯 What it does: The paper explores a new method to improve the performance of machine learning models, especially when dealing with complex datasets.
ENC-ODE: Event-level Neurodegenerative Modeling in Continuous Time with Neural ODEs
Song, Yujee (Pohang University of Science and Technology), Kim, Won Hwa (University of North Carolina at Chapel Hill)
Recurrent Neural NetworkTransformerAuto EncoderMultimodalityTabularTime SeriesBiomedical DataAlzheimer's DiseaseElectronic Health RecordsOrdinary Differential Equation
🎯 What it does: This paper proposes an event-level neural ordinary differential equation (ENC-ODE) model based on diagnostic conditions to predict the future evolution of multimodal brain biomarkers in Alzheimer's disease patients in continuous time.
End2Reg: Learning Task-Specific Segmentation for Markerless Registration in Spine Surgery
Pettinari, Lorenzo (University of Basel), Licci, Maria (University Children's Hospital)
SegmentationPose EstimationOptimizationGraph Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningPoint CloudBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Propose End2Reg, an end-to-end deep learning framework that jointly learns segmentation and registration, achieving marker-free RGB-D registration for spinal surgery without requiring manual segmentation labels.
EndoDA3: Unified Depth and Pose Estimation for Endoscopic Video
Tan, Jie, Lu, Li (University of Shanghai for Science and Technology)
Pose EstimationDepth EstimationTransformerSupervised Fine-TuningContrastive LearningSimultaneous Localization and MappingOptical FlowVideoBiomedical Data
🎯 What it does: Propose the EndoDA3 framework to achieve unified depth and pose estimation for monocular endoscopic videos.
EndoGSim: Physics-Aware 4D Dynamic Endoscopic Scene Simulations via MLLM-Guided Gaussian Splatting
Liu, Changjing (Chinese University of Hong Kong), Ren, Hongliang (Chinese University of Hong Kong)
Data SynthesisOptimizationComputational EfficiencyRepresentation LearningTransformerLarge Language ModelDiffusion modelNeural Radiance FieldAuto EncoderGenerative Adversarial NetworkContrastive LearningGaussian SplattingOptical FlowImageVideoMultimodalityBiomedical DataPhysics Related
🎯 What it does: A unified 4D Gaussian Splatting framework is constructed by automatically initializing based on a multimodal large language model and finely optimizing through differentiable material point methods, achieving physics-aware scene reconstruction and dynamic simulation from endoscopic videos.
EndoVLM: An Endoscopy Vision-Language Pre-training Model via Anatomy-Guided Sparsity and Progressive Alignment
Yi, Zhenyu (Alibaba Group), Xia, Yingda (Alibaba Group)
RecognitionSegmentationRetrievalAnomaly DetectionTransformerPrompt EngineeringVision Language ModelAuto EncoderContrastive LearningImageVideoTextBiomedical DataElectronic Health Records
🎯 What it does: Propose EndoVLM — a vision-language pre-training model specifically designed for gastrointestinal endoscopy, capable of aligning clinical reports with unordered image sets;
EndoX: A GPU-Accelerated Endoscopic Perception Simulation Framework
Lam, Kwan Tung (NVIDIA AI Technology Center, NVIDIA Corporation), Cheung, Ka Chun (NVIDIA AI Technology Center, NVIDIA Corporation)
SegmentationData SynthesisPose EstimationDepth EstimationDiffusion modelNeural Radiance FieldGaussian SplattingOptical FlowImageVideoMeshBiomedical DataComputed Tomography
🎯 What it does: Developed EndoX, a GPU-based endoscopic perception simulation framework that can automatically generate anatomical meshes from CT data and render RGB and multiple geometric annotations (depth, normal, optical flow, occlusion, overlay) in one go.
Enhancing Brain MRI Anomaly Detection and Reasoning with ROI Rethink and Synthetic Data
Li, Shangkun (Fudan University), Wang, Yuanyuan (Fudan University)
Data SynthesisAnomaly DetectionExplainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringVision Language ModelImageBiomedical DataMagnetic Resonance ImagingRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose the BrReMark framework, which adopts a two-round 'annotation-reflection' interaction, enabling brain MRI abnormality detection, description, and diagnosis based on auditable ROI annotations, generating verifiable diagnostic chains.
Enhancing Pathological VLMs with Cross-scale Reasoning
Phan, Chi (National University of Singapore), Jin, Yueming (National University of Singapore)
ClassificationImage TranslationRestorationAnomaly DetectionExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsVision Language ModelContrastive LearningImageTextBiomedical DataElectronic Health RecordsBenchmark
🎯 What it does: Proposes a training and evaluation paradigm for cross-scale pathological vision-language models, constructs the Scale-VQA dataset, and trains the ScaleReasoner-R1 model capable of reasoning at multiple magnification levels.
Equity-Aware Self-supervised Multi-domain Connectome Representation Learning for Major Depressive Disorder Diagnosis
Barman, Jyotismita (Indian Institute of Technology Delhi), Kumar, Sandeep (Indian Institute of Technology Delhi)
ClassificationDomain AdaptationFederated LearningExplainability and InterpretabilityComputational EfficiencyRepresentation LearningData-Centric LearningGraph Neural NetworkAuto EncoderContrastive LearningGraphTime SeriesBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease
🎯 What it does: Propose a self-supervised multi-domain brain connectivity network learning framework named RESOLVE, which constructs a dual graph network using temporal domain correlation and frequency domain coherence, and achieves fair and reliable MDD diagnosis through aligned latent space.
ESRT: Endoscopic Stereo Reconstruction Transformer
Chang, Jiun-Feng (National Yang Ming Chiao Tung University), Huang, Chun-Rong (Industrial Technology Research Institute)
Depth EstimationKnowledge DistillationTransformerContrastive LearningImagePoint Cloud
🎯 What it does: Propose a Transformer network called ESRT for endoscopic stereo images to achieve precise depth reconstruction.
Estimated Age-Guided Cerebral Microbleed Segmentation
Kwon, Junmo (Sungkyunkwan University), Park, Hyunjin (Sungkyunkwan University)
SegmentationConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Propose the AgeCMBNet framework, which utilizes age estimation based on T1-weighted MRI for the segmentation of cerebral microbleeds.
Evi-Steer: Learning to Steer Biomedical Vision-Language Models through Efficient and Generalizable Evidential Tuning
Koleilat, Taha (Concordia University), Xiao, Yiming (Concordia University)
ClassificationDomain AdaptationExplainability and InterpretabilityComputational EfficiencyTransformerSupervised Fine-TuningPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound
🎯 What it does: Propose the Evi-Steer framework, which utilizes low-dimensional cross-modal fine-tuning and evidence reasoning to perform parameter-efficient, uncertainty-aware fine-tuning of BiomedCLIP in the activation space, achieving multi-modal medical image classification with few data and domain generalization.
Evidential Fusion Network for Multimodal Survival Prediction Under Missing Modalities
Xing, Yucheng (National University of Singapore), Feng, Mengling (National University of Singapore)
TransformerMultimodalityBiomedical Data
🎯 What it does: Propose a multi-modal survival prediction model EMMS based on evidence fusion, which can still provide reliable predictions and uncertainty estimates when some modalities are missing.
Evidential Neural Networks for Uncertainty-Aware Alzheimer’s Disease Screening from Resting-State EEG
Ujjain, Siddhant (Indian Institute of Technology Delhi), Gandhi, Tapan K.
ClassificationAnomaly DetectionExplainability and InterpretabilityComputational EfficiencyConvolutional Neural NetworkScore-based ModelContrastive LearningBiomedical DataAlzheimer's DiseaseAudio
🎯 What it does: An evidence neural network based on EEGNet performs binary classification of resting-state EEG between AD+FTD and normal cognition (CN), providing class probabilities and uncertainty in a single forward pass.
Evidential Perfusion Physics-Informed Neural Networks with Residual Uncertainty Quantification
Lee, Junhyeok (Seoul National University), Choi, Kyu Sung (Seoul National University)
Anomaly DetectionOptimizationConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelContrastive LearningImageBiomedical DataComputed TomographyPhysics Related
🎯 What it does: Propose an Evidence-based Physics-informed Neural Network (EPPINN) based on deep learning to solve the deconvolution problem in computed tomography perfusion (CTP) images of acute ischemic stroke, and provide confidence estimates for each voxel.
Evo-RAD: Navigating Rare Retinal Disease Diagnosis via Self-Evolving Agentic Retrieval
Xia, Wangding (Hong Kong Polytechnic University), Wang, Shujun (Hong Kong Polytechnic University)
ClassificationRetrievalRepresentation LearningGraph Neural NetworkTransformerReinforcement LearningAgentic AIPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataRetrieval-Augmented Generation
🎯 What it does: By modeling the retrieval process as a Markov Decision Process, a self-evolving retrieval agent is designed to dynamically delete, insert, and terminate retrieved cases to construct a reference set with higher diagnostic homogeneity, thereby improving the diagnostic accuracy for rare retinal diseases.
EvoMed: Self-Evolving Medical VLM via Training-free Continued Learning
Li, Jiaming (ShanghaiTech University), Wang, Qian (ShanghaiTech University)
Recommendation SystemExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningDrug DiscoveryReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBiomedical DataComputed TomographyBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose EvoMed, a training-free, parameter-freezing medical vision-language model framework, which enhances medical VQA performance through self-cycling reasoning, criticism, summarization, and knowledge base maintenance, achieving performance improvements during the self-evolution phase with only 500 samples.
Expert-like Bone Ultrasound Segmentation through Expert-in-the-loop Mask-conditioned Progressive Learning
Tavangar, Arash (McGill University), Hooshiar, Amir (McGill University)
SegmentationConvolutional Neural NetworkSupervised Fine-TuningContrastive LearningImageBiomedical DataUltrasound
🎯 What it does: Proposed the ExiL expert cyclic mask conditional progressive learning framework for automatically completing and refining ultrasound bone segmentation.
Explainability in Multimodal Deep Transformation Models for Stroke Outcome Prediction
Herzog, Lisa, Sick, Beate (Zurich University of Applied Sciences)
ClassificationExplainability and InterpretabilityConvolutional Neural NetworkTransformerImageTabularBiomedical DataDiffusion Tensor ImagingElectronic Health Records
🎯 What it does: A multi-modal deep transformation model (DTM) was constructed and evaluated, combining 3D convolutional neural networks with clinical tabular features, for predicting functional independence in stroke patients three months later. Occlusion and Grad-CAM methods were adapted to DTM to generate visual interpretability maps.
Exploiting Interior Geometry for Intracranial Aneurysm Detection and Segmentation on 3D Point Clouds via GNNs
Layachi, Mohamed Amine (Ibn Tofail University), Autrusseau, Florent (University of Nantes)
SegmentationExplainability and InterpretabilityGraph Neural NetworkSupervised Fine-TuningContrastive LearningPoint CloudBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Propose an end-to-end framework based on internal geometry detection and graph neural networks for automatically detecting and segmenting cerebral aneurysms using 3D point clouds.
Exploiting Longitudinal Context in Clinician-Verified Interactive Lesion Tracking
Kirchhoff, Yannick (German Cancer Research Center Heidelberg), Maier-Hein, Klaus (German Cancer Research Center Heidelberg)
SegmentationData SynthesisConvolutional Neural NetworkTransformerPrompt EngineeringDiffusion modelAuto EncoderContrastive LearningImageBiomedical DataComputed TomographyPositron Emission TomographyBenchmark
🎯 What it does: This paper proposes a verified tracking workflow that combines prompts generated during registration with verification or correction by radiologists, and then utilizes longitudinal information from baseline scans for lesion segmentation.
Exploring Efficient Weakly Supervised Ultrasound Video Object Segmentation with VLM-Guided Structure Advantage Learning
Li, Jialu (Hong Kong University of Science and Technology (Guangzhou)), Zhu, Lei (Hong Kong University of Science and Technology (Guangzhou))
SegmentationTransformerPrompt EngineeringVision Language ModelDiffusion modelContrastive LearningVideoBiomedical DataUltrasound
🎯 What it does: This paper proposes a weakly supervised ultrasound video object segmentation method guided by a vision-language model (VLM), aiming to significantly reduce the cost of manual annotation.
Exploring Subnetwork Interactions in Heterogeneous Brain Network via Prior-Informed Graph Learning
Liu, Siyu (Northeastern University), Zaiane, Osmar R. (University of Alberta)
ClassificationExplainability and InterpretabilityDrug DiscoveryGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBiomedical DataAlzheimer's DiseaseElectronic Health Records
🎯 What it does: Propose KD-Brain, a graph learning framework based on prior information, for modeling interactions among subnetworks in heterogeneous brain networks and discovering interpretable functional pathways in mental disorder diagnosis.
FA-Diff: Leveraging Frequency-Aware RWKV Diffusion for OCT to Correlation-Stability Map Translation
Zhang, Chuanhao (Tsinghua University), Liao, Hongen (Tsinghua University)
Image TranslationRestorationTransformerDiffusion modelAuto EncoderBiomedical DataComputed Tomography
🎯 What it does: Proposed a frequency-aware RWKV diffusion framework, FA-Diff, for translating optical coherence tomography (OCT) images into corresponding stiffness (CS) elasticity images, achieving virtual elasticity imaging without hardware compression.
Fair Curriculum Learning for Concept Bottleneck Models in Dermatology
Cockayne, Matthew J. (Keele University), Al-Bander, Baidaa (Keele University)
ClassificationFederated LearningExplainability and InterpretabilityData-Centric LearningTransformerSupervised Fine-TuningGenerative Adversarial NetworkContrastive LearningImageBiomedical Data
🎯 What it does: This paper proposes a Fair Curriculum Concept Bottleneck Model (Fair Curriculum CBM) based on the concept of fair curriculum, which reduces skin tone disparities in dermatological diagnosis through phased fair objectives training.
FairGE: Gated Expert Routing for Intersectional Fairness in Medical Foundation Models
Shao, Yuchen (South China University of Technology), Chen, Qi (South China University of Technology)
ClassificationFederated LearningSafty and PrivacyExplainability and InterpretabilityComputational EfficiencyRepresentation LearningAdversarial AttackTransformerMixture of ExpertsGenerative Adversarial NetworkContrastive LearningImageBiomedical DataElectronic Health Records
🎯 What it does: The paper proposes a cross-fair adaptation framework called FairGE for medical foundational models, achieving fair improvements for cross-subgroup sensitive attributes through multi-attribute gated Mixture-of-Experts and adversarial feature purification.
Fairness Beyond Demographics: Optimizing Performance Across Appearance-Based Hidden Cohorts in Medical Imaging
Masroor, Milad (University of Surrey), Carneiro, Gustavo (University of Surrey)
ClassificationOptimizationConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningImageBiomedical DataBenchmark
🎯 What it does: This paper proposes Label-Independent Hidden Cohort Fairness Training (LHCF), which identifies potential subpopulations by clustering image appearances and optimizes fairness on these hidden cohorts, while also constructing the HIDFairBench benchmark to evaluate multi-attribute fairness.
FairPneum: Improving Fairness in Pneumonia Diagnosis and Lesion Segmentation
Chu, Yuetan (King Abdullah University of Science and Technology), Gao, Xin (King Abdullah University of Science and Technology)
ClassificationSegmentationAnomaly DetectionConvolutional Neural NetworkTransformerGenerative Adversarial NetworkContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: We propose the FairPneum framework, specifically designed to enhance fairness in the tasks of pneumonia diagnosis and lesion segmentation.
Faithful, Interpretable Chest X-ray Diagnosis with Artifact-Free B-cos Networks
Arya, Shreyash (Max-Planck-Institute for Informatics), Keuper, Margret (University of Mannheim)
ClassificationAnomaly DetectionExplainability and InterpretabilityConvolutional Neural NetworkAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Proposed an anti-aliasing B-cos network for chest X-ray diagnosis and generating artifact-free interpretable contribution maps.
FARSIGHT: A Dynamic Multi Modal Medical Image Analysis Framework Powered by a Prior Knowledge Guided Vision Language Model
Yang, Jinghan (Chinese Academy of Sciences), Wang, Kun (Chinese Academy of Sciences)
Image TranslationImage HarmonizationSegmentationAnomaly DetectionExplainability and InterpretabilityConvolutional Neural NetworkTransformerPrompt EngineeringMixture of ExpertsVision Language ModelContrastive LearningImageMultimodalityBiomedical DataUltrasound
🎯 What it does: Propose the FARSIGHT framework, achieving an end-to-end pipeline from raw medical image archives to patient-level diagnosis, including unsupervised MAGOS data governance and multi-modal multi-instance learning HMMIL.
Fast 3D Cell Volume Reconstruction from Multi-focus Images by Cross-modality Knowledge Distillation
Lubis, Imam Khairi (Kumamoto University), Nagahara, Hajime (Osaka University)
SegmentationKnowledge DistillationConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningImageBiomedical Data
🎯 What it does: A framework for image segmentation based on Deep Image Prior (DIP) is proposed. Under an unsupervised setting, this framework directly learns pixel-level segmentation results from raw images through an improved network structure and loss function.
Fast Few-Shot Embolization Simulations in Ischemic Stroke Using Equivariant Neural Fields
Kuipers, Thijs P. (Amsterdam UMC), Bekkers, Erik J. (University of Amsterdam)
OptimizationComputational EfficiencyRobotic IntelligenceMeta LearningDrug DiscoveryGraph Neural NetworkTransformerDiffusion modelNeural Radiance FieldContrastive LearningMeshBiomedical DataMagnetic Resonance ImagingComputed TomographyReview/Survey PaperStochastic Differential Equation
🎯 What it does: This paper proposes a fast, data-efficient mesh-free neural surrogate based on SE(3)-equivariant neural fields for simulating embolism dynamics in ischemic stroke.
FDMMBone: Flow-Based Conditional Deformable Mesh Modeling for Accurate 3D Bone Defect Completion
Liu, Zhenhong (Beijing Normal University), Shui, Wuyang (Alder Hey Children's NHS Foundation Trust)
RestorationGenerationData SynthesisConvolutional Neural NetworkTransformerDiffusion modelFlow-based ModelAuto EncoderGenerative Adversarial NetworkPoint CloudMeshBiomedical DataMagnetic Resonance ImagingComputed TomographyOrdinary Differential Equation
🎯 What it does: Propose a conditional 3D bone defect completion framework based on flow matching, FDMMBone, which can generate high-quality, topologically consistent, and aligned with the defect boundary triangular mesh implants.
FDRAS: Failure Diagnosis and Repair for Airway Segmentation
Zhang, Francis Xiatian (University of Edinburgh), Khadem, Mohsen (University of Edinburgh)
SegmentationAnomaly DetectionConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningImageBiomedical DataComputed Tomography
🎯 What it does: After airway segmentation in chest CT, the FDRAS framework is proposed, which first predicts structural mismatches through error maps, and then performs localization repair based on these error maps.
Feature Space Guidance for Breast Cancer Classification in DCE-MRI
Hamm, Benjamin (German Cancer Research Center (DKFZ)), Maier-Hein, Klaus (German Cancer Research Center (DKFZ))
ClassificationSegmentationAnomaly DetectionConvolutional Neural NetworkTransformerSupervised Fine-TuningAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Propose a breast cancer classification framework for dynamic contrast-enhanced breast magnetic resonance imaging (DCE-MRI), which first locates the breast region of interest (ROI) through 3D segmentation, then selects key phases in the latent space dynamically, and finally improves sensitivity to small lesions using large-scale anatomical pre-training.
FedAgree: Label-Free Performance Estimation Under Distribution Shift for Federated Medical Imaging Analysis
Serra, Giuseppe (German Cancer Research Center), Buettner, Florian (German Cancer Research Center)
Domain AdaptationAnomaly DetectionFederated LearningConvolutional Neural NetworkTransformerContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasound
🎯 What it does: The study explores how to utilize multi-party model checkpoints for unlabeled OOD performance estimation in a federated learning environment, proposing the FedAgree method and designing six checkpoint combination strategies.
FedCHDP: Federated Concept-Guided Dual-Modal Prompting for CHD Diagnosis
Gu, Ruilin (Wuhan University), Du, Bo (Wuhan University)
ClassificationAnomaly DetectionFederated LearningRepresentation LearningData-Centric LearningTransformerSupervised Fine-TuningPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataUltrasound
🎯 What it does: Proposed a FedCHDP federated concept-guided dual-modal prompting framework for the screening of congenital heart disease (CHD) using first-trimester fetal cardiac ultrasound images;
Federated Medical Image Segmentation under Modality Heterogeneity via Specialized Adapters
Jamir, Obed (Indian Institute of Technology Jodhpur), Paul, Angshuman (Indian Institute of Technology Jodhpur)
SegmentationFederated LearningTransformerPrompt EngineeringContrastive LearningImageMultimodalityBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Propose the FedMoSA framework, which utilizes dedicated and shared adapters on top of SAM to achieve federated learning for medical image segmentation, addressing the problem of multi-modal heterogeneity.
FedProIn: Mitigating Client Drift for Learnable Prototypes in Federated Medical Imaging
Kumar, Harsh (Indian Institute of Science Bengaluru), Sundaresan, Vaanathi (Indian Institute of Science Bengaluru)
ClassificationAnomaly DetectionFederated LearningConvolutional Neural NetworkContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: A framework called FedProIn is proposed for federated medical imaging learning, which mitigates client drift by using learnable multi-prototypes and influence aggregation, thereby improving the model's robustness on heterogeneous data.
FedSwitch: From Standard to Balanced Aggregation via Private Histogram Estimation in Federated Medical Image Classification
Cacace, Paolo (CERN), Serio, Luigi (CERN)
ClassificationFederated LearningSafty and PrivacyAgentic AIContrastive LearningImageBiomedical DataBenchmark
🎯 What it does: Propose FedSwitch, which enables switching from standard FedAvg to distribution-aware aggregation in federated learning through private histogram estimation, to address label imbalance issues in medical image classification.
FermatSyn: SAM2-Enhanced Bidirectional Mamba with Isotropic Spiral Scanning for Multi-Modal Medical Image Synthesis
Yuan, Feng (Xi'an Jiaotong University), Gao, Xin (Xi'an Jiaotong University)
Data SynthesisConvolutional Neural NetworkContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: This paper proposes a new deep learning model for image processing tasks.
FetalAgents: A Multi-Agent System for Fetal Ultrasound Image and Video Analysis
Hu, Xiaotian (Tsinghua University), Tian, Qiyuan (Sichuan University)
ClassificationImage TranslationRestorationSegmentationRecommendation SystemAnomaly DetectionConvolutional Neural NetworkTransformerLarge Language ModelAgentic AIVision Language ModelContrastive LearningImageVideoTextBiomedical DataUltrasound
🎯 What it does: This paper proposes the FetalAgents multi-agent system, achieving end-to-end automation from planar classification, segmentation, and measurement of fetal ultrasound images to video summarization and structured report generation.
Fibers of Asymmetric Similarity: A Framework for Clinical and Imaging Data
Müller, Johanna P. (Friedrich-Alexander University Erlangen-Nürnberg), Kainz, Bernhard (Friedrich-Alexander University Erlangen-Nürnberg)
ClassificationAnomaly DetectionExplainability and InterpretabilityData-Centric LearningTransformerContrastive LearningImageTabularBiomedical DataUltrasoundElectronic Health Records
🎯 What it does: A unified framework based on asymmetric Tversky similarity was studied, which handles clinical tabular data and frozen image encoders, addressing class imbalance and missing data issues in rare disease screening.
Fine-Grained Cerebrovascular Parsing in DSA via Structurally-Grounded Semantic Disentanglement
Zhu, Kai (Wuhan University of Science and Technology), Zhao, Yitian (Ningbo Institute of Materials Technology and Engineering, Chinese Academy of Sciences)
SegmentationConvolutional Neural NetworkTransformerContrastive LearningImageBiomedical Data
🎯 What it does: This paper achieves fine-grained cerebral vascular parsing in digital subtraction angiography (DSA) images, generating pixel-level annotations for six categories of blood vessels.
FlowReg: Reconstruction-Guided Flow Matching for Universal Spinal CT/X-Ray Registration
Shen, Ao (Hohai University), Zhou, Shaohua Kevin (Hohai University)
Image TranslationPose EstimationConvolutional Neural NetworkTransformerDiffusion modelFlow-based ModelContrastive LearningOptical FlowImageMultimodalityBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Proposes FlowReg, an end-to-end flow matching framework that utilizes 3D reconstruction as a spatial prior, enabling single-segment and multi-segment CT/X-ray registration.
FluxCut: Shifting from Global Reconstruction to Local Geometry for Unsupervised Small Lesion Segmentation
Deng, Zhiwei (University of Southern California), Shi, Yonggang (University of Southern California)
SegmentationAnomaly DetectionConvolutional Neural NetworkGraph Neural NetworkDiffusion modelAuto EncoderOptical FlowImageBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease
🎯 What it does: Propose FluxCut, a framework that shifts unsupervised small lesion segmentation from global reconstruction to local geometric modeling.
FOCUS: Towards Fetal Obstetric Corrective UltraSound Guidance
Lamdouar, Hala (University of Oxford), Noble, J. Alison (University of Oxford)
RecognitionGenerationData SynthesisConvolutional Neural NetworkTransformerPrompt EngineeringVision Language ModelContrastive LearningVideoTextMultimodalityUltrasound
🎯 What it does: Proposed and implemented the FOCUS framework, which utilizes multimodal information from video and text to generate targeted corrective feedback in real-time for obstetric ultrasound learners, helping them improve technical details such as probe posture and parameter settings.
Footprint-Guided Exemplar-Free Continual Histopathology Report Generation
Kumari, Pratibha (University of Regensburg), Merhof, Dorit (University of Regensburg)
GenerationDomain AdaptationTransformerPrompt EngineeringVision Language ModelImageTextBiomedical DataRetrieval-Augmented Generation
🎯 What it does: Proposes a whole slide image (WSI) to pathology report generation framework for continual learning, which does not require storing any original WSI or patch examples during training, and can retain memory of previous domains when new domains arrive.
Foundation Model-Driven Key Anatomy Frame Selection for Blind-Sweep Ultrasound Fetal Birth Weight Estimation
Ou, Le (Shenzhen University), Ni, Dong (Shenzhen University)
Image TranslationAnomaly DetectionComputational EfficiencyRepresentation LearningConvolutional Neural NetworkTransformerVision Language ModelContrastive LearningVideoBiomedical DataUltrasound
🎯 What it does: A framework for selecting key anatomical frames based on foundational models is proposed, using blind sweep ultrasound videos to estimate fetal birth weight 48 hours before delivery.
FrameONE: Hierarchical Motion Modeling for Universal Multi-view Echocardiographic Keyframe Detection
Chen, Rusi (Shenzhen University), Ni, Dong (Shenzhen University)
ClassificationRecognitionPose EstimationConvolutional Neural NetworkRecurrent Neural NetworkTransformerContrastive LearningOptical FlowVideoBiomedical DataUltrasound
🎯 What it does: Proposed a unified multi-view cardiac ultrasound keyframe (ES/ED) detection framework called FrameONE;
FreeBridge: Variational Schrödinger Bridges for Cellular Transition Dynamics
Wang, Xurui (Stony Brook University), You, Chenyu (University Health Network)
GenerationData SynthesisDrug DiscoveryConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningImageBiomedical DataStochastic Differential Equation
🎯 What it does: Propose the FreeBridge method, which learns continuous transition dynamics of chemical or genetic perturbations at the single-cell level using variational Schrödinger bridges, supervised only by endpoint distributions.
Frequency and Geometry Guided Graph Clustering for Weakly Supervised Skin Lesion Segmentation
Deng, Zhaoxin (Henan Normal University), Shen, Hualei (Henan Normal University)
SegmentationGraph Neural NetworkDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageGraph
🎯 What it does: Propose a frequency and geometry guided graph clustering framework FG3-Cluster, which achieves skin lesion segmentation using weak labels;
Frequency-Aware Multi-Margin Loss for Imbalanced Medical Ordinal Classification
Yin, Haojie (Duke Kunshan University), Huang, Kaizhu (Duke Kunshan University)
ClassificationConvolutional Neural NetworkContrastive LearningBiomedical Data
🎯 What it does: This paper proposes a frequency-aware multi-margin loss (FMML) and a long-term tail rank deviation (LTRD) metric to address the long-tailed imbalance problem in medical ordinal classification.
Frequency-Aware Neural Architecture Search with Bidirectional Mamba for EEG Emotion Recognition
Xu, Shasha (Taiyuan University of Technology), Wen, Xin (Taiyuan University of Technology)
RecognitionNeural Architecture SearchRecurrent Neural NetworkGraph Neural NetworkTransformerReinforcement LearningContrastive LearningTime SeriesSequentialBiomedical DataElectrocardiogramReview/Survey Paper
🎯 What it does: This paper proposes a frequency-aware neural architecture search (NAS) framework, utilizing bidirectional Mamba modules for cross-band modeling in EEG emotion recognition, and independently searching for variable topologies and internal temporal configurations within each frequency band.
Frequency-Aware Post-Training Quantization for Medical Image Denoising
Ko, Uni (Ewha Womans University), Choi, Jang-Hwan (Ewha Womans University)
RestorationAuto EncoderContrastive LearningImageBiomedical DataComputed Tomography
🎯 What it does: This paper proposes a post-training quantization framework based on frequency awareness, specifically designed for medical image denoising;
Frequency-Aware Prototype Learning in VLM for Intracranial Aneurysm HR-VWI Segmentation
Wang, Yixin (Southern University of Science and Technology), Si, Weixin (Shenzhen University of Advanced Technology)
SegmentationTransformerPrompt EngineeringVision Language ModelDiffusion modelContrastive LearningImageTextBiomedical DataMagnetic Resonance Imaging
🎯 What it does: This paper proposes a visual-language model based on frequency-aware prototype learning for brain aneurysm segmentation in high-resolution vascular wall imaging (HR-VWI).
From Baseline to Future CT: Large-Scale Diffusion Pretraining for Single-Scan IPF Progression Prediction
McConnell, Niccolò (University College London), Jacob, Joseph (University College London)
GenerationData SynthesisConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataComputed Tomography
🎯 What it does: Proposed a 3D latent diffusion model based on large-scale chest CT pre-training, capable of synthesizing future IPF progression CT images from a single baseline scan.
From Exams to Trajectories: Neural Controlled Differential Equations for Longitudinal Mammography Vision-Language Pretraining
Sun, Zijun (UiT Arctic University of Norway), Kampffmeyer, Michael (Norwegian Computing Center)
ClassificationImage TranslationRestorationAnomaly DetectionRepresentation LearningData-Centric LearningTransformerVision Language ModelVision-Language-Action ModelNeural Radiance FieldAuto EncoderContrastive LearningImageTextMultimodalityTime SeriesBiomedical DataComputed TomographyElectronic Health RecordsStochastic Differential Equation
🎯 What it does: Propose the Dynamo framework, which performs multimodal pre-training on the longitudinal history of mammography images, combining continuous time models and vision-language alignment.