arXivSub Start free trial

MICCAI 2026 Papers — Page 8

International Conference on Medical Image Computing and Computer-Assisted Intervention · 1165 papers

MVDA-Net: Multi-View Dual-Alignment Network for 3D Carotid Artery Segmentation

Wen, Yang (Shenzhen University), Sheng, Bin (Hong Kong University of Science and Technology)

SegmentationConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningBiomedical DataComputed Tomography

🎯 What it does: Proposed MVDA-Net for precise segmentation of carotid arteries in 3D CTA images;

N6-TreeMamba: Radiology Knowledge-Guided 6-Neighborhood Tree Scanning for 3D MRI-Based Stroke Complication Prediction

Zhao, Chenyang (Shanghai Institute of Technology), Hu, Chuanfei (Wenzhou Medical University)

ClassificationAnomaly DetectionComputational EfficiencyTransformerDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose the N6-TreeMamba framework for predicting pneumonia and hemorrhagic transformation risks in patients with acute ischemic stroke.

NASD-Diff: A Noise-Assisted Structural Difference Diffusion Model for CT Angiography Generation from NCCT

Xie, Zeng, Zhang, Rong (Ningbo University)

GenerationData SynthesisConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderGenerative Adversarial NetworkImageBiomedical DataComputed Tomography

🎯 What it does: Propose a noise-assisted structural difference diffusion model (NASD-Diff) to generate CTA images from non-contrast CT (NCCT);

NEARL: Interacted Query Adaptation with Orthogonal Regularization for Medical Vision-Language Understanding

Peng, Zelin (Shanghai Jiao Tong University), Shen, Wei (Chinese Academy Of Science)

ClassificationDomain AdaptationComputational EfficiencyRepresentation LearningTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingComputed TomographyAlzheimer's Disease

🎯 What it does: Proposes NEARL, a parameter-efficient medical vision-language understanding framework that leverages the pre-trained CLIP model;

NeRD: Neuro-Symbolic Rule Distillation for Efficient Ontology-Grounded Chain-of-Thought in Medical Image Diagnosis

Yang, Hongxi (Monash University), Ge, Zongyuan (Monash University)

ClassificationExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningImageMultimodalityBiomedical DataChain-of-Thought

🎯 What it does: Propose the NeRD framework, which combines data-driven logical rules with multi-modal chain reasoning to generate efficient ontological diagnostic reasoning chains.

NetCorr: A Resection-Induced Network Perturbation Framework for Predicting Seizure Freedom in Drug-Resistant Epilepsy

Monsoor, Tonmoy, Roychowdhury, Vwani (UCLA)

ClassificationExplainability and InterpretabilityComputational EfficiencyDrug DiscoveryGraph Neural NetworkSupervised Fine-TuningAuto EncoderContrastive LearningGaussian SplattingOptical FlowGraphTabularTime SeriesBiomedical DataElectronic Health RecordsElectrocardiogram

🎯 What it does: Developed the NetCorr framework, which predicts the probability of seizure freedom after surgery for drug-resistant epilepsy by constructing patient-level 2D predictive features. This is achieved by coupling channel-level spike-HFO with regression of synchronized network features R², and by considering surgery as a network perturbation.

Network-Aware Bilinear Tokenization for Brain Functional Connectivity Representation Learning

Milecki, Leo, Zhao, Qingyu (Weill Cornell Medicine)

Representation LearningTransformerAuto EncoderContrastive LearningBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Proposes the NERVE framework, which uses network-aware bilinear decomposition to tokenize brain functional connectivity matrices in blocks, and learns transferable features in MAE self-supervised learning.

NeuroAlign: A Unified Plug-and-Play Enhancer for Visual and Linguistic Brain Decoding

Li, Jinke (Lanzhou University), Fu, Yu (Lanzhou University)

RestorationGenerationData SynthesisRepresentation LearningTransformerVision Language ModelDiffusion modelContrastive LearningImageTextMultimodalityBiomedical DataMagnetic Resonance ImagingRetrieval-Augmented Generation

🎯 What it does: This paper proposes a general plugin called NeuroAlign, which aligns the fMRI encoder with the generative decoder, significantly improving the geometric and semantic quality of visual reconstruction and text decoding.

NeuroFlow: Manifold-Aware Topological Flow Matching for EEG Decoding

Wang, Lei (South China University Of Technology), Xu, Yanwu (Pazhou Lab)

GenerationRetrievalRepresentation LearningGraph Neural NetworkTransformerDiffusion modelFlow-based ModelAuto EncoderContrastive LearningImageMultimodalityAudio

🎯 What it does: Construct the NeuroFlow framework to achieve decoding and generation of visual content from EEG signals.

NeuroSonic: Conditional Flow Matching for EEG-to-Speech Reconstruction

Gao, Wenhao (Stony Brook University), You, Chenyu (Stony Brook University)

RestorationGenerationData SynthesisTransformerDiffusion modelFlow-based ModelMultimodalityBiomedical DataElectrocardiogramOrdinary Differential EquationAudio

🎯 What it does: Designed a EEG-to-speech reconstruction framework called NeuroSonic based on conditional flow matching, which maps noisy audio states to clean speech by learning a deterministic probability-flow velocity field.

NeuroSymb-MRG: Differentiable Abductive Reasoning with Active Uncertainty Minimization for Radiology Report Generation

Fu, Rong (University of Macau), Fong, Simon (University of Macau)

GenerationExplainability and InterpretabilityData-Centric LearningConvolutional Neural NetworkTransformerLarge Language ModelPrompt EngineeringDiffusion modelImageTextBiomedical DataMagnetic Resonance ImagingComputed TomographyElectronic Health RecordsRetrieval-Augmented Generation

🎯 What it does: A framework named NeuroSymb-MRG is built, which automatically generates structured and interpretable radiology reports using differentiable neuro-symbolic reasoning, retrieval-augmented generation, and active uncertainty minimization.

NoiBox: Safeguard Frontier via Reverse Adversarial Diffusion for Noisy-Box-Supervised 3D Tumor Segmentation

Ji, Kailun (Wuhan University of Technology), Zhou, Quan (Zhongnan Hospital of Wuhan University)

SegmentationTransformerDiffusion modelGenerative Adversarial NetworkContrastive LearningBiomedical DataComputed Tomography

🎯 What it does: This paper proposes a method for 3D tumor segmentation using rough 3D bounding boxes (noisy boxes), constructing a safe frontier soft supervision target to avoid missed and mislabeled annotations caused by hard box constraints.

Noise-Aware Importance–Uncertainty Disentangled Multimodal Learning for Robust Cancer Survival Prediction

Ming, Wenlong (Nanjing University of Information Science and Technology), Wang, Xiangxue (Nanjing University of Information Science and Technology)

ClassificationExplainability and InterpretabilityDrug DiscoveryGraph Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageMultimodalityTabularBiomedical DataAlzheimer's DiseaseElectronic Health RecordsBenchmark

🎯 What it does: A multi-modal survival prediction framework is constructed that can simultaneously utilize tumor whole slide images and mRNA expression data, specifically considering the technical noise and asymmetry in importance between the two modalities.

Non-intrusive Body Composition Assessment from Full-Body mmWave Scans

Senne, Miriam (Technical University of Munich), Navab, Nassir (Technical University of Munich)

Data SynthesisDomain AdaptationOptimizationRepresentation LearningTransformerDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImagePoint CloudMeshBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: This paper proposes to use millimeter-wave radar scanning to acquire full-body point clouds, and to achieve regression prediction of visceral fat volume (VAT), body fat percentage (BFP), and multiple anthropometric indicators through a multi-task learning model.

Non-linear INR-Based Motion Modeling for 4D Radiotherapy

Gebauer, Johannes B. (University Medical Center Hamburg-Eppendorf), Werner, René (University Medical Center Hamburg-Eppendorf)

OptimizationRepresentation LearningNeural Radiance FieldContrastive LearningOptical FlowImageBiomedical DataComputed Tomography

🎯 What it does: Propose an implicit neural representation (INR) model based on SIREN to achieve patient-specific nonlinear respiratory motion modeling;

Non-parametric Prototypes Enable Semantic Graph Learning for Whole-Slide Pathology

Gao, Zixuan (East China Normal University), Li, Qingli (East China Normal University)

ClassificationAnomaly DetectionComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerVision Language ModelContrastive LearningImageTextBiomedical Data

🎯 What it does: Proposed a non-parametric prototype-guided semantic graph learning network (ProtoSG-Net) for classifying whole slide images (WSI);

Non-Stationary Spectral State-Space Networks with Region-Wise Frequency Routing for Robust Fundus Vessel Segmentation

Viriyasaranon, Thanaporn (Ewha Womans University), Choi, Jang-Hwan (Ewha Womans University)

SegmentationDomain AdaptationRecurrent Neural NetworkTransformerMixture of ExpertsVision-Language-Action ModelAuto EncoderContrastive LearningImageBiomedical Data

🎯 What it does: Proposed an NS Net network based on explicit frequency-domain hybrid experts for robust retinal vessel segmentation and artery-vein separation.

NR-Align: Non-rigid Alignment for Non-simultaneous Two-View 3D Coronary Reconstruction in Complex Cardiac Interventions

Wang, Pengbo (Beijing University of Posts and Telecommunications), Di, Chunxia (Beijing University of Posts and Telecommunications)

SegmentationData SynthesisConvolutional Neural NetworkDiffusion modelAuto EncoderContrastive LearningOptical FlowImageBiomedical DataComputed TomographyElectrocardiogram

🎯 What it does: Propose the NR-Align module, performing non-rigid alignment on asynchronous dual-view DSA backprojection volumes, thereby improving the connectivity and topological consistency of 3D coronary artery reconstruction.

NutriDiff: Phenotype-Consistent Diffusion Augmentation for Malnutrition Screening

Yang, Hanqiu (Jilin University), He, Lili (Jilin University)

ClassificationImage TranslationData SynthesisAnomaly DetectionTransformerSupervised Fine-TuningPrompt EngineeringDiffusion modelContrastive LearningImageBiomedical DataAlzheimer's Disease

🎯 What it does: Propose the NutriDiff framework for few-shot enhancement and diagnosis of facial images of elderly patients with malnutrition.

Occlusion-aware Disparity Estimation via Depth Prior Embedded Cost Volume for Stereo Endoscopic Images

Liu, Ziteng, Fu, Yili (Harbin Institute of Technology)

Depth EstimationConvolutional Neural NetworkTransformerContrastive LearningOptical FlowImagePoint CloudBenchmark

🎯 What it does: Proposes DEDENet, a stereo disparity estimation network that integrates monocular depth priors, specifically designed to address occlusion problems in endoscopic images.

OCS-MAMBA: Object-Centric, Logic-Guided Approach for Surgical Action Understanding

Shuvo, Md Rezowan H. F. (Robert Gordon University), Elyan, Eyad

RecognitionSegmentationAnomaly DetectionExplainability and InterpretabilityComputational EfficiencyRepresentation LearningRecurrent Neural NetworkGraph Neural NetworkTransformerVision-Language-Action ModelContrastive LearningImageVideoGraphBiomedical DataBenchmarkRetrieval-Augmented Generation

🎯 What it does: This paper proposes a three-stage, object-centric, logic-constrained surgical action understanding framework called OCS-Mamba.

Oculo: A Multilabel Dataset for the Identification of Ocular Abnormalities from Ultrasound Images

Kumari, Sneha (Microsoft Research), Jain, Mohit (Microsoft Research)

ClassificationAnomaly DetectionConvolutional Neural NetworkTransformerSupervised Fine-TuningImageBiomedical DataUltrasound

🎯 What it does: This paper constructs the public Oculo ophthalmic ultrasound multi-label dataset and systematically benchmarks various deep learning models and domain-specific foundation models on this dataset.

Omni-Sleep: A Sleep Foundation Model via Hierarchical Contrastive Learning of CNS–ANS Dynamics

Hou, Zhoujie (Southern University of Science and Technology), Liu, Quanying (Southern University of Science and Technology)

ClassificationRepresentation LearningConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningMultimodalityTime SeriesBiomedical DataElectrocardiogram

🎯 What it does: This paper proposes Omni-Sleep, a sleep foundation model based on the hierarchical alignment of the central nervous system (CNS) and the autonomic nervous system (ANS), which performs unsupervised pre-training on multi-modal PSG using hierarchical contrastive learning and long-time-slot latent mask modeling.

On the Behavior of Calibration Error Under Rater Disagreement

Kumar, Harshit (Whiterabbit.ai), Matthews, Thomas Paul (Whiterabbit.ai)

ClassificationAnomaly DetectionExplainability and InterpretabilityBiomedical DataReview/Survey Paper

🎯 What it does: This paper investigates the problem of the deviation between calibration error (ECE) and majority voting labels in medical AI evaluation under multiple annotator uncertainty, and proposes two calibration schemes (MR-ECE and AdjustedECE) to restore oracle guarantee.

On the Robustness of Temporal Vision-Language Models for Surgical Endoscopy Videos

Rashid, Darakshan (Mohamed bin Zayed University of Artificial Intelligence), Xie, Yutong (Mohamed bin Zayed University of Artificial Intelligence)

ClassificationRestorationRepresentation LearningTransformerSupervised Fine-TuningPrompt EngineeringVision Language ModelContrastive LearningImageVideoTextBiomedical DataBenchmark

🎯 What it does: Investigate the robustness of temporal vision-language models for clinical endoscopy videos under noise/distortion, propose the Endo-C6 degradation benchmark, and develop the lightweight adaptation model RobustEndoCLIP.

On-Manifold Variational Learning with Heat-Kernel Priors

Xing, Jiarui (Yale), Wang, Jian (Yale)

GenerationAnomaly DetectionRepresentation LearningGraph Neural NetworkDiffusion modelAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose a manifold-anchored variational learning framework based on the heat kernel prior, for generating subgroup prototypes on the data manifold and achieving generative modeling in unsupervised representation learning of medical images.

One Sequence to Segment Them All: Efficient Data Augmentation for CT and MRI Cross-Domain 3D Spine Segmentation

Molinier, Nathan (Polytechnique Montréal), Cohen-Adad, Julien (Technical University of Munich)

SegmentationDomain AdaptationComputational EfficiencyConvolutional Neural NetworkContrastive LearningMultimodalityBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: A GPU-accelerated strong data augmentation scheme was developed under the nnUNet framework, training CT and MRI spinal segmentation models with single-sequence data, and evaluated on multi-modal, multi-sequence discrete domains.

One-Shot Data Selection for Medical Image Classification via Graph Coverage

Rustamov, Zahiriddin (United Arab Emirates University), Zaki, Nazar (United Arab Emirates University)

ClassificationData-Centric LearningGraph Neural NetworkImageGraphBiomedical Data

🎯 What it does: Proposed a graph-based single-pass data selection method for medical image classification, which prunes the labeled training pool by using frozen base model embeddings.

Oneclass4ASD: Autism Spectrum Disorder Screening via One-class Classifier

Chen, Fan (Beijing Jiaotong University), Guan, Qingji (Beijing Jiaotong University)

ClassificationAnomaly DetectionConvolutional Neural NetworkAuto EncoderContrastive LearningImageBiomedical Data

🎯 What it does: This paper proposes a one-class classification framework called Oneclass4ASD, which uses only eye movement data from healthy subjects to screen for autism spectrum disorder.

Ontology-Grounded Structured Prediction for Dental CBCT Reporting

Lumetti, Luca (University of Modena and Reggio Emilia), Bolelli, Federico (University of Modena and Reggio Emilia)

Object DetectionSegmentationRepresentation LearningConvolutional Neural NetworkTransformerContrastive LearningImageTextBiomedical DataComputed TomographyRetrieval-Augmented Generation

🎯 What it does: Constructed an ontology-based structured report resource based on dental CBCT, providing 893 clinical reports, 13 types of findings and their attributes annotated with RDF/OWL, and proposed a three-stage structured prediction framework.

Open World MRI Reconstruction with Bias-Calibrated Adaptation

Liu, Jiyao (Fudan University), Xu, Ningsheng (Fudan University)

RestorationTransformerDiffusion modelScore-based ModelBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose the BiasRecon framework, achieving adaptive calibration and regularization adjustment for open-world MRI reconstruction;

Open-PMC-18M: A High-Fidelity Large Scale Medical Dataset for Multimodal Representation Learning

Baghbanzadeh, Negin (Vector Institute), Afkanpour, Arash

ClassificationRetrievalRepresentation LearningGraph Neural NetworkTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBiomedical DataBenchmarkRetrieval-Augmented Generation

🎯 What it does: Constructed a high-fidelity medical vision-language dataset, Open-PMC-18M, consisting of 18 million subfigure-subtitle + context abstract pairs, and proposed a scalable subfigure detection and text enhancement pipeline.

OpenRC: An Open-Source Robotic Colonoscopy Framework for Multimodal Data Acquisition and Autonomy Research

Kapuria, Siddhartha (University of Texas at Austin), Alambeigi, Farshid (Technical University of Munich)

Robotic IntelligenceSimultaneous Localization and MappingOptical FlowVideoMultimodalityBiomedical Data

🎯 What it does: Proposed and implemented a low-cost, scalable open-source robotic colonoscopy platform called OpenRC, which provides multi-degree-of-freedom mechanical actuation to traditional colonoscopes while preserving clinical workflows, and simultaneously records video, operator commands, mechanical status, and EM tracking of the end-effector pose;

OPGAgent: An Agent for Auditable Dental Panoramic X-ray Interpretation

Yu, Zhaolin (Monash University), Ge, Zongyuan (Monash University)

Anomaly DetectionExplainability and InterpretabilityTransformerLarge Language ModelAgentic AIMixture of ExpertsVision Language ModelImageTextBiomedical DataComputed TomographyBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This paper proposes OPGAgent, a multi-tool intelligent agent for auditable全景牙片 (panoramic radiograph) interpretation, achieving multi-task diagnosis through hierarchical evidence collection, a specialized toolset, and a consensus sub-agent.

OphthFlowBench: A Unified Ophthalmology Multimodal Benchmark with Workflow-Aligned Evaluation

Zheng, Qiaojian (Shenzhen University), Shen, Linlin (Shenzhen Second People's Hospital)

ClassificationRecognitionAnomaly DetectionExplainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision Language ModelImageTextMultimodalityBiomedical DataBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Constructed OphthFlowBench — a unified multi-modal ophthalmology evaluation benchmark, and trained a specialized ophthalmology multi-modal large language model, Oph-RL, based on this benchmark.

Opportunistic Cardiac Health Assessment: Estimating Phenotypes from Localizer MRI Through Multi-modal Representations

Zeybek, Busra Nur (Technical University of Munich), Kafali, Sevgi Gokce (Technical University of Munich)

Representation LearningData-Centric LearningTransformerAuto EncoderContrastive LearningImageMultimodalityTabularBiomedical DataMagnetic Resonance ImagingElectronic Health RecordsElectrocardiogram

🎯 What it does: This paper proposes C-TRIP, a multimodal framework that jointly learns a latent space using fast, low-quality local MRI, electrocardiogram (ECG), and clinical tabular data, and predicts cardiac phenotypes during inference using only local MRI.

OPS: One-Shot Point Prompted Semantics-Aware Learning for Medical Image Registration

Xie, Housheng (Shanghai Jiao Tong University), Zheng, Guoyan (Shanghai Jiao Tong University)

SegmentationTransformerPrompt EngineeringDiffusion modelContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Proposed the OPS framework, which utilizes single-point prompts and the pre-trained DINOv3 semantic clustering capability to generate dense semantic similarity maps across the entire dataset, and embeds them into an unsupervised medical image registration network to achieve semantic-aware registration.

Optimal Steps for Fast Diffeomorphic Shape Registration

Bigo-Balland, Hadrien (Inria, Université Paris Cité, Inserm, HeKA), Feydy, Jean (Inria, Université Paris Cité, Inserm, HeKA)

OptimizationComputational EfficiencyData-Centric LearningDiffusion modelScore-based ModelGaussian SplattingOptical FlowPoint CloudMeshBiomedical DataComputed Tomography

🎯 What it does: Propose a learning-free, fast, and topology-preserving point cloud/surface shape registration method that can automatically achieve high-precision registration at the clinical scale.

OPVD: On-Policy Chain-of-Visual-Thought Distillation for Medical Reasoning

Zhang, Ningyue (Hong Kong University of Science and Technology), Yang, Qiushi (City University of Hong Kong)

Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision Language ModelContrastive LearningImageTextMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasoundChain-of-Thought

🎯 What it does: Propose an On-Policy Chain-of-Visual-Thought Distillation (OPVD) framework, which uses the same model as both teacher and student, generating dense token-level supervision through visual and textual interaction, thereby enhancing the reasoning ability of medical vision-language models.

OR-Action: Multi-Role Video Understanding with Fine-Grained Actions

Tristram, Felix (Technical University of Munich), Navab, Nassir (Technical University of Munich)

RecognitionGraph Neural NetworkTransformerVision-Language-Action ModelContrastive LearningVideoGraphBenchmark

🎯 What it does: This paper proposes the OR-Action benchmark, which converts the publicly available EgoExOR scene graph annotations into fine-grained, multi-role actions using rules, and evaluates the performance of scene graphs and visual models on action recognition in external operating room videos based on this; subsequently, a purely visual temporal model is designed, and a multi-view to single-view feature alignment strategy is introduced to improve action recognition performance under a single view.

Ordinal Diffusion Models for Color Fundus Images

Schmidt, Gustav (University of Tübingen), Müller, Sarah (University of Tübingen)

GenerationData SynthesisTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningImageBiomedical Data

🎯 What it does: Proposed an ordinal diffusion model capable of generating realistic color fundus images based on diabetic retinopathy (DR) grading.

Ordinal Priors for Colonoscopy Temporal Segmentation

Jang, Seunghyun (Seoul National University), Park, Chang Min (Seoul National University Hospital)

SegmentationComputational EfficiencyConvolutional Neural NetworkTransformerContrastive LearningVideoBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Proposed a colonoscopy temporal segmentation framework that considers anatomical order, combining sequence-level ordinal regression loss with an ordinal constrained decoder.

Organ-Invariant Tumor Representation Learning via Disentangled Adversarial MAE for Multi-organ Ultrasound Imaging

Xiong, Xiangyu (Macao Polytechnic University), Tan, Tao (Macao Polytechnic University)

ClassificationSegmentationRepresentation LearningTransformerPrompt EngineeringAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataUltrasound

🎯 What it does: Proposed a Decoupled Adversarial Masked Autoencoder (DA-MAE), which learns organ-agnostic tumor representations in multi-organ ultrasound images through a dual-branch encoder, and uses text prompts to assist decoding.

OrthoFlow: Decoupled Landmark Prediction and Image Synthesis for Orthodontic Treatment Outcome Visualization

Wang, Leyuan (ShanghaiTech University), Cui, Zhiming (ShanghaiTech University)

Image TranslationGenerationPose EstimationGraph Neural NetworkTransformerPrompt EngineeringDiffusion modelScore-based ModelFlow-based ModelImageText

🎯 What it does: Designed the OrthoFlow two-stage decoupled framework for predicting geometric deformation and image synthesis of lateral cephalograms after orthodontic treatment.

OrthoMorphNet: Alignment-Guided Hierarchical Morphological Synthesis of Post-orthodontic Dentition and Gingiva

Han, Ji Yong (Seoul National University), Yi, Won-Jin (Tech University of Korea)

GenerationData SynthesisTransformerDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningPoint CloudMeshComputed TomographyPositron Emission TomographyReview/Survey Paper

🎯 What it does: Proposed OrthoMorphNet, a hierarchical framework capable of simultaneously predicting post-orthodontic tooth alignment and gingival morphology.

OSCAR: Occupancy-Based Shape Completion via Acoustic Neural Implicit Representations

Wysocki, Magdalena (Technical University of Munich), Navab, Nassir (Technical University of Munich)

RestorationDiffusion modelScore-based ModelNeural Radiance FieldAuto EncoderGenerative Adversarial NetworkContrastive LearningOptical FlowImageBiomedical DataComputed TomographyUltrasound

🎯 What it does: Propose the OSCAR framework for spine shape completion based on ultrasound, which directly recovers complete 3D geometry from B-mode images by using an implicit network that couples acoustic information with the occupancy field.

OsteoFlow: Lyapunov-Guided Flow Distillation for Predicting Bone Remodeling after Mandibular Reconstruction

Aftabi, Hamidreza (University of British Columbia), Hardisty, Michael (University of Toronto)

GenerationData SynthesisKnowledge DistillationConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelFlow-based ModelRectified FlowOptical FlowImageBiomedical DataComputed TomographyStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Designed a flow-based generative framework called OsteoFlow, which predicts CT images from postoperative day 5 to one year after mandible reconstruction surgery under low-sample conditions using Lyapunov-guided trajectory distillation.

OT-Coupled Flow: Towards Realistic TEE Simulation via Unpaired Cross-Modality Translation with Anatomical Consistency

Li, Yunhao (Hong Kong Polytechnic University), Qin, Jing (University of Hong Kong)

Image TranslationData SynthesisTransformerDiffusion modelFlow-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningOptical FlowImageBiomedical DataComputed TomographyUltrasound

🎯 What it does: Generate realistic TEE images from CT slices in an unsupervised manner.

OTCHA: Optimal Transport-Driven Confidence-Aware Latent Hub Alignment for Multi-view Medical Image Classification

Yang, Jiwoong (Hanyang University), Jang, Ikbeom (Hankuk University of Foreign Studies)

ClassificationImage TranslationDomain AdaptationOptimizationKnowledge DistillationRepresentation LearningConvolutional Neural NetworkTransformerContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound

🎯 What it does: Propose the OTCHA module, which refines the patch tokens from each perspective first through optimized transport matching and then fuses information on a shared latent hub before fusion in multi-view medical image classification.

OTFusion: Optimal Transport-Driven 3D Multi-modal Medical Image Fusion for Enhanced Cross-Modality Representation

Wei, Xinjian (Nankai University), Xu, Jing (Nankai University)

Image HarmonizationOptimizationRepresentation LearningConvolutional Neural NetworkTransformerDiffusion modelContrastive LearningOptical FlowMultimodalityBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose a 3D multi-modal medical image fusion framework based on optimal transport (OT) called OTFusion, which adopts a distribution alignment approach to achieve cross-modal information fusion;

P2R-Net: Physics-to-Reality SM Calibration Network with Sparse Measurement Guidance for High-Precision MPI Reconstruction

Zhang, Lizhi (Northwest University), He, Xiaowei (Shenzhen University of Advanced Technology)

Super ResolutionConvolutional Neural NetworkTransformerSupervised Fine-TuningDiffusion modelAuto EncoderContrastive LearningBiomedical Data

🎯 What it does: Propose a Physics-to-Reality (P2R) dual-branch network, using sparse measured data to guide the physical model in generating the complete system matrix (SM), thereby achieving high-precision calibration and reconstruction of the system matrix in magnetic particle imaging (MPI);

PALETTE: Pathology-Aware Latent Embedding for Targeted Tissue Expression

Ma, Rongze (Northwestern Polytechnical University), Xia, Yong (Northwestern Polytechnical University)

Image TranslationGenerationData SynthesisTransformerDiffusion modelAuto EncoderContrastive LearningImageBiomedical Data

🎯 What it does: Propose the PALETTE framework, which divides the virtual IHC transformation from H&E into two steps: semantic encoding and molecular synthesis. It utilizes a frozen PathFM to extract high-capacity continuous representations and generates discrete IHC markers through Cond-VAR in a coarse-to-fine autoregressive manner.

PancCADx: A Multimodal Framework for Pancreatic Cancer Diagnosis

Hu, Shan (Wuhan University), Wang, Zhongyuan (Wuhan University)

ClassificationExplainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningVision Language ModelImageMultimodalityTabularBiomedical DataUltrasoundChain-of-Thought

🎯 What it does: A explainable multimodal framework called PancCADx was constructed, using Qwen3‑VL‑8B‑Thinking combined with chain-of-thought (CoT) reasoning to diagnose pancreatic cancer and integrate it with clinical data.

PANet: Probability-Anchored Network for Ultrasound Bone Segmentation

Yan, Wenqing (Tsinghua University), Wang, Guangzhi (Tsinghua University)

SegmentationConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningImageBiomedical DataUltrasound

🎯 What it does: Propose a probabilistic anchoring network (PANet) that utilizes a probabilistic spatial intermediate representation for ultrasound bone surface segmentation.

Parameter-Efficient Self-Supervised Adaptation for EEG-FM under Fixed Computational Budgets

Dani, Meghal (University of Tübingen), Liebe, Stefanie (University of Tübingen)

Domain AdaptationAnomaly DetectionComputational EfficiencyRepresentation LearningTransformerSupervised Fine-TuningContrastive LearningTime SeriesBiomedical DataElectrocardiogram

🎯 What it does: This paper studies self-supervised learning (SSL) adaptation methods for updating only 9–29% of the parameters of the EEG foundation model (EEG-FM) in resource-constrained clinical environments, and improves the performance of downstream tasks by freezing the encoder and training only the final layer.

Patch-to-Global: Random Patch Diffusion for Globally Consistent Megapixel Artifact Inpainting in Whole Slide Images

Lee, Hyeseong (Korea University), Ahn, Sangjeong (Seoul National University)

RestorationConvolutional Neural NetworkTransformerDiffusion modelContrastive LearningBiomedical Data

🎯 What it does: Proposes the RestorePath framework, which addresses large-scale (megapixel-level) artifacts in whole digital pathology slides by using random patch diffusion to achieve globally consistent restoration, avoiding high-confidence erroneous predictions by the model on artifact-containing images.

Pathologist Attention–Aligned Report Generation for Prostate Histopathology

Xue, Ruoyu (Stony Brook University), Samaras, Dimitris (Stony Brook University)

GenerationExplainability and InterpretabilityTransformerSupervised Fine-TuningVision Language ModelGaussian SplattingImageTextBiomedical DataElectronic Health Records

🎯 What it does: This paper proposes incorporating pathologist visual attention supervision into the pathological report generation model, and improves the generation quality through attention alignment loss.

PathoMamba: Piecewise-Diffeomorphic Registration via Stiffness-Modulated State Space Dynamics

Asimeng, Ernest (Jiangsu University), Liu, Zhe (Jiangsu University)

Image TranslationSegmentationOptimizationComputational EfficiencyTransformerAuto EncoderContrastive LearningOptical FlowImageBiomedical DataMagnetic Resonance ImagingStochastic Differential Equation

🎯 What it does: Proposes a segmentation-aware longitudinal image registration framework for brain tumors called PathoMamba, based on a stiffness-modulated state space model, which can simultaneously preserve the rigidity of healthy tissues and the plasticity of tumors while ensuring topological safety.

Patient-Conditioned Vision–Language Priors with Stable Mixture-of-Experts for Label-Efficient Alzheimer’s MRI Staging

Ding, Shan (Southeast University), Shu, Huazhong (Southeast University)

ClassificationRepresentation LearningData-Centric LearningTransformerMixture of ExpertsVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease

🎯 What it does: This study proposes a semi-supervised framework based on patient-conditioned vision-language prior, scale-preserving mixture-of-experts classifier, and EMA teacher for Alzheimer's disease MRI staging under few-label conditions.

PAVL: Enhancing Feature Prototype Alignment with Vision–Language Model for Semi-supervised Medical Image Segmentation

Kuang, Junchang (Guilin University of Electronic Technology), Pan, Xipeng (Guilin University of Electronic Technology)

SegmentationRepresentation LearningData-Centric LearningConvolutional Neural NetworkTransformerVision Language ModelAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose a semi-supervised medical image segmentation framework called PAVL based on CLIP, achieving feature alignment and semantic enhancement through unified representation matching, prototype alignment, and embedding fusion.

PC-Seg: Progressive Cross-View Consistency for 3D OCT Segmentation from Sparse 2D Annotations

Konno, Tsubasa (Tohoku University), Aoki, Takafumi (Tohoku University)

SegmentationKnowledge DistillationConvolutional Neural NetworkTransformerContrastive LearningBiomedical Data

🎯 What it does: Proposed the PC-Seg framework, which trains a 3D OCT segmentation model step-by-step through cross-view consistency using sparse 2D annotations.

PC-VLG: Position-Correlated Vision–Language Graph Alignment for Semi-supervised Medical Image Segmentation

Qiu, Luyi (Nanyang Technological University), Kong, Adams Wai-Kin (Nanyang Technological University)

SegmentationGraph Neural NetworkTransformerVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose the PC-VLG framework, achieving location-aware visual-language graph alignment in semi-supervised medical image segmentation, enhancing anatomical consistency and boundary accuracy.

PCRP: Progressive Coarse-to-Fine Refinement for Pathology Grounding

Peng, Liang (Wuhan University), Dong, Xingping (Wuhan University)

RecognitionImage TranslationSegmentationConvolutional Neural NetworkTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical Data

🎯 What it does: Designed and implemented a progressive refinement framework called PCRP for precisely locating lesion regions in pathological images based on natural language descriptions;

Peeling an Onion: Layer-Wise Doppler-Backprojected Ultrasound Microvascular Vector Flow Imaging

Fei, Zetao (Xiamen University), Chen, Yinran (Xiamen University)

OptimizationOptical FlowBiomedical DataUltrasound

🎯 What it does: This paper proposes a hierarchical peeling method based on microvascular morphology and multi-angle Doppler signals, which utilizes the Hessian matrix to estimate blood flow direction and obtains microvascular vector flow fields through weighted global optimization back-projection;

Percept-Aware Surgical Planning for Visual Cortical Prostheses with Vascular Avoidance

Pogoncheff, Galen, Beyeler, Michael (University Of California Santa Barbara)

OptimizationSafty and PrivacyContrastive LearningGaussian SplattingImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Proposes a perception-based, differentiable framework for surgical planning to optimize visual cortex electrode placement in three-dimensional brain structures while satisfying safety constraints.

Perfusion Aware Infarct Core and Penumbra Segmentation on NCCT with Region Constraints

Zhang, Donghao (Huazhong Science of Science and Technology), Qiu, Wu (Huazhong Science of Science and Technology)

SegmentationExplainability and InterpretabilityComputational EfficiencyConvolutional Neural NetworkTransformerDiffusion modelContrastive LearningImageBiomedical DataComputed Tomography

🎯 What it does: This study proposes a collaborative learning framework, CoMPASS-Net, which utilizes a foundational model to jointly segment the infarct core and ischemic penumbra from NCCT images, and enhances segmentation consistency and interpretability through structural and physiological constraints.

PET-Adapter: Test-Time Domain Adaptation for Full and Limited-Angle PET Image Reconstruction

Yilmaz, Rüveyda (RWTH Aachen University), Schulz, Volkmar (RWTH Aachen University)

RestorationDomain AdaptationTransformerDiffusion modelScore-based ModelContrastive LearningImageBiomedical DataMagnetic Resonance ImagingPositron Emission Tomography

🎯 What it does: Propose a PET-Adapter framework that can reconstruct clinical PET data through unsupervised test-time adaptation, based solely on a diffusion model pre-trained on a simulated phantom.

PETFlux: Taming Natural Foundation Models for Accurate MRI-to-PET Synthesis

Yin, Yuan (ShanghaiTech University), Wang, Qian (ShanghaiTech University)

Image TranslationData SynthesisTransformerLarge Language ModelPrompt EngineeringDiffusion modelRectified FlowImageBiomedical DataMagnetic Resonance ImagingPositron Emission TomographyAlzheimer's Disease

🎯 What it does: This paper proposes PETFlux, a framework for transferring natural image pre-trained models to MRI-to-PET generation.

PhaseAT: Fourier Phase Adversarial Training for Medical Image Domain Generalization

Sharshar, Ahmed (Mohamed bin Zayed University of Artificial Intelligence), Guizani, Mohsen (Mohamed bin Zayed University of Artificial Intelligence)

ClassificationDomain AdaptationConvolutional Neural NetworkGenerative Adversarial NetworkContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasound

🎯 What it does: Propose PhaseAT, which enhances the domain generalization ability of medical images by utilizing phase adversarial training in the Fourier frequency domain.

PhaseGen: A Diffusion-Based Approach for Complex-Valued MRI Data Generation

Rempe, Moritz (University Hospital Essen), Kleesiek, Jens (University Hospital Essen)

RestorationGenerationData SynthesisConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose a diffusion model called PhaseGen conditioned on amplitude, which generates phase from amplitude-only MRI data to synthesize complete complex k-space data, and uses these synthetic data for pre-training downstream tasks.

PhaseMamba: Two-Stage Learning from Short Clips to Full Procedures for Online Surgical Phase Recognition

Mohamed, Shaheer (Queensland University of Technology), Fookes, Clinton (Queensland University of Technology)

RecognitionTransformerContrastive LearningVideoBiomedical DataBenchmark

🎯 What it does: Proposes PhaseMamba, a two-stage online surgical phase recognition framework.

PhyDiCT: Plug-and-Play CT Reconstruction from Sparse X-Rays via Differentiable Rendering and Strong Priors

Dai, Weicheng (Boston University), Batmanghelich, Kayhan (Boston University)

RestorationTransformerPrompt EngineeringVision Language ModelDiffusion modelScore-based ModelImageTextBiomedical DataComputed Tomography

🎯 What it does: Train freely to reconstruct 3D lung CT images from a few X-rays through differentiable rendering and strong text-guided diffusion prior.

PhyRadGS: Physics-Aware Radiative Gaussian Splatting for Fly-Scanning CBCT Reconstruction

Li, Anwei (CVTE Research), Wang, Rongqiu (CVTE Research)

Diffusion modelScore-based ModelContrastive LearningGaussian SplattingImageBiomedical DataComputed Tomography

🎯 What it does: Propose a physics-aware 3D Gaussian scattering framework called PhyRadGS for reconstruction in flying scan CBCT, capable of generating clear medical images under continuous rotation and rolling shutter conditions.

Physical-Driven Unified Implicit Regularization Network with Scan Parameter Prompts for Accelerated Multi-parametric MR Imaging

Yang, Yan (Xi'an Jiaotong University), Sun, Jian (Shenzhen Institute of Advanced Technology)

RestorationOptimizationComputational EfficiencyConvolutional Neural NetworkTransformerPrompt EngineeringDiffusion modelBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose a physics-driven unified implicit regularization network, PDUIR-Net, which can simultaneously complete multi-contrast image reconstruction and quantitative parameter mapping under undersampled MR data.

Physics-Based iOCT Sonification for Real-Time Interaction Awareness in Subretinal Injection

Reyes Vargas, Luis D. (Technische Universität München), Matinfar, Sasan (Technische Universität Dresden)

SegmentationOptimizationComputational EfficiencyRobotic IntelligenceConvolutional Neural NetworkDiffusion modelContrastive LearningOptical FlowImageBiomedical DataPhysics RelatedAudio

🎯 What it does: Developed a real-time iOCT acoustic mapping system that uses a physics-inspired mass-spring-damper model to convert retinal layer structure, tip motion, and tissue deformation caused by injection into audible feedback, helping surgeons precisely locate and monitor bubble formation during subretinal injection.

Physics-Driven Cold Diffusion with Severity-Aware Sampling and Adaptive Degradation Estimation for MRI Motion Artifact Removal

Li, Chuanpu (Southern Medical University), Yang, Wei (Southern Medical University)

RestorationDiffusion modelScore-based ModelBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Proposed a physics-driven noise-free cold diffusion framework for motion artifact removal in magnetic resonance imaging

Physics-Grounded Weakly Supervised Histopathological Tissue Segmentation via Frozen Tri-Domain Prototypes

Zhang, Shuyu (Jiangnan University), Pan, Xiang (Jiangnan University)

SegmentationTransformerMixture of ExpertsContrastive LearningBiomedical DataBenchmarkPhysics Related

🎯 What it does: Propose the PhysPro weakly supervised organ segmentation framework, which utilizes physics-prior triple-domain static prototypes to stabilize feature learning, decouples spatial-frequency features, and achieves triple-state adaptive training through differential supervision.

Physics-Informed Surrogate Model Using Graph Neural Network for Cardiac Electrophysiology

Tran, Vu Anh (National University of Singapore), Li, Lei (National University of Singapore)

Explainability and InterpretabilityComputational EfficiencyRecurrent Neural NetworkGraph Neural NetworkMeshGraphTime SeriesBiomedical DataMagnetic Resonance ImagingFibre Orientation DistributionPhysics RelatedStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: A physics-informed graph neural network was developed for fast and high-fidelity prediction of cardiac electrophysiological dynamics, achieving real-time surrogate modeling.

Physics-Inspired Continuous Transformer for Fast QDSA Reconstruction

Liu, Yang (Southern University of Science and Technology), Si, Weixin (Shenzhen University of Advanced Technology)

Image TranslationRestorationExplainability and InterpretabilityComputational EfficiencyTransformerDiffusion modelScore-based ModelRectified FlowAuto EncoderContrastive LearningImageTime SeriesBiomedical DataMagnetic Resonance ImagingPhysics Related

🎯 What it does: This study proposes PIC-Former, a physics-informed continuous transformer, for reconstructing continuous time density curves (TDC) from irregular DSA samples, enabling fast and robust generation of hemodynamic parameters (such as TTP, CBV, MTT) and QDSA velocity estimation, supporting intraoperative hemodynamic assessment;

Physio-Anatomical Prior-Guided Probabilistic Field Learning for Ultrasound-Based Neuro-Fascial Interface Landmark Localization

Dai, Guowei (Xidian University), Chen, Hu (Xi'an Jiaotong University)

Image TranslationConvolutional Neural NetworkGenerative Adversarial NetworkImage

🎯 What it does: A systematic evaluation and comparison of multiple image style transfer algorithms were conducted.

Physiologically Consistent Missing Channel Reconstruction and Cascaded MoE for Impairment Categorization from Visual Electrophysiological Signals

Yao, Chenglin (Southern University of Science and Technology), Liu, Jiang (Southern University of Science and Technology)

ClassificationConvolutional Neural NetworkTransformerMixture of ExpertsGenerative Adversarial NetworkContrastive LearningMultimodalityTime SeriesBiomedical DataElectrocardiogram

🎯 What it does: This paper proposes a missing channel reconstruction and cascaded Mixture-of-Experts (CMoE) framework tailored for visual electrophysiological signals, aiming to achieve more robust grading of visual dysfunction.

Physiology and Anatomy Aware Inverse Inference of Myocardial Infarction for Cardiac Digital Twin

Wang, Mengxiao (Shanghai Jiao Tong University), Li, Lei (Shanghai Jiao Tong University)

SegmentationData SynthesisExplainability and InterpretabilityConvolutional Neural NetworkRecurrent Neural NetworkGraph Neural NetworkDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImagePoint CloudMeshBiomedical DataMagnetic Resonance ImagingElectrocardiogram

🎯 What it does: This paper proposes a non-invasive myocardial infarction (MI) localization framework based on cardiac digital twins, which uses three-dimensional myocardial geometry and multi-lead ECG for inverse inference.

Physiology-Aligned Hierarchical Network for Heart Murmur Detection

Lin, Wenjun, Chui, Chee Kong (National University Of Singapore)

ClassificationAnomaly DetectionConvolutional Neural NetworkRecurrent Neural NetworkTransformerContrastive LearningBiomedical DataElectrocardiogramAudio

🎯 What it does: Proposed the Physiology‑Aligned Hierarchical Network (PAHNet), which achieves detection and rough localization of heart murmurs by performing feature encoding and fusion at multiple levels such as phonocardiogram phase, cardiac cycle, and auscultation location.

Physiology-Guided Cross-View Spatial-Temporal Aligning for Early Prediction of pCR in Breast Cancer from Longitudinal Multimodal MRI

Li, Xin, Wang, Hongkai (Dalian University of Technology)

ClassificationImage TranslationAnomaly DetectionTransformerSupervised Fine-TuningContrastive LearningImageMultimodalityBiomedical DataMagnetic Resonance Imaging

🎯 What it does: This paper proposes the CALM framework, which uses multimodal longitudinal MRI (DCE-MRI, DWI, T2WI) to predict whether patients with breast cancer can achieve pathological complete response (pCR) early in neoadjuvant chemotherapy.

PhysioSplat: Physics-Informed Dynamic Gaussian Splatting for Surgical Scene Reconstruction

Basak, Hritam (Stony Brook University), Yin, Zhaozheng (Stony Brook University)

Depth EstimationRobotic IntelligenceConvolutional Neural NetworkTransformerSupervised Fine-TuningDiffusion modelNeural Radiance FieldAuto EncoderGenerative Adversarial NetworkContrastive LearningGaussian SplattingOptical FlowImageVideoPoint CloudBiomedical DataBenchmarkPhysics Related

🎯 What it does: Propose PhysioSplat, a physics-informed constrained dynamic Gaussian splatting (4DGS) method for high-precision reconstruction in surgical scenarios, capable of separating tools from tissues, applying elastic regularization, and modeling specular highlights on wet surfaces.

PhysOCT: Physics-Guided Retinal Modeling for OCT Image Synthesis

Wang, Yutong (Nanjing University of Science and Technology), Chen, Qiang (Nanjing University of Science and Technology)

GenerationData SynthesisConvolutional Neural NetworkDiffusion modelNeural Radiance FieldOptical FlowImageBiomedical DataComputed Tomography

🎯 What it does: This paper proposes a physics-guided OCT image synthesis framework called PhysOCT, which enhances anatomical consistency in generated images by simultaneously considering lesion evolution and retinal layer deformation in conditional generation.

PICA: Physics-Guided Implicit Neural Fields with Boundary Conditioning for Intracranial Aneurysm Hemodynamics

Hu, Mengfan (ShanghaiTech University), Zhang, Zeng (ShanghaiTech University)

Convolutional Neural NetworkTransformerDiffusion modelScore-based ModelFlow-based ModelRectified FlowNeural Radiance FieldContrastive LearningBiomedical DataComputed TomographyBenchmarkPhysics Related

🎯 What it does: Proposes PICA, a physics-guided implicit neural field, for rapidly predicting 3D blood flow fields from patient-specific vascular geometries and inlet boundary conditions.

Piecewise Dynamic Diffusion Regularization for Reconstruction of Cardiac Cine MRI

Fürnrohr, Florian (Technical University of Munich), Heckel, Reinhard (Technical University of Munich)

RestorationConvolutional Neural NetworkTransformerDiffusion modelBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose the Piecewise Dynamic Diffusion Regularization (PDDR) method, which realizes real-time cardiac cine MRI reconstruction by using a dynamic diffusion model as a block-wise regularizer.

PiMoE: Physics-Informed Mixture-of-Experts for Accelerated MRI Reconstruction

Kim, Kyuri (Seoul National University), Ye, Sung-Joon (Seoul National University)

RestorationConvolutional Neural NetworkMixture of ExpertsBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose the PiMoE framework, which utilizes physics-informed Mixture-of-Experts to adaptively model accelerated MRI reconstruction;

PINet: A Multimodal Network for Pontine Infarction Segmentation and Early Neurological Deterioration Prediction

Liu, Jinshuo, Pan, Yi (Shenzhen Institute of Advanced Technology)

ClassificationSegmentationConvolutional Neural NetworkTransformerMixture of ExpertsDiffusion modelAuto EncoderContrastive LearningImageTabularBiomedical DataMagnetic Resonance ImagingDiffusion Tensor Imaging

🎯 What it does: This study proposes a multi-modal multi-task network called PINet, which can simultaneously perform segmentation of pontine infarction lesions and predict early neurologic deterioration (END);

PINNOCHIO: Physics-Informed Neural Network for Coupled Hyperelastic Interface-Volume Simulation in Orthognathic Surgery

Lee, Jungwook (Rensselaer Polytechnic Institute), Yan, Pingkun (Houston Methodist Research Institute)

Graph Neural NetworkDiffusion modelPoint CloudMeshBiomedical DataComputed TomographyReview/Survey PaperPhysics Related

🎯 What it does: Developed a physics-informed neural network framework called PINNOCHIO for predicting facial soft tissue deformation caused by bone repositioning in orthognathic surgery.

Pix2Rep-v2: Data-Efficient Representation Learning for Dense Medical Imaging Applications

Sifaoui, Sofiane (Institut Polytechnique de Paris), Le Folgoc, Loïc (Institut Polytechnique de Paris)

SegmentationRepresentation LearningConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose Pix2Rep-v2, a sparse and efficient self-supervised dense representation learning framework for medical imaging.

Plane-Wise Retrieval Memory and Structural Evidence Fusion for Tetralogy of Fallot Video Diagnosis

Lv, Xingguo (Hunan University), Li, Kenli (Hunan University)

ClassificationDomain AdaptationExplainability and InterpretabilityTransformerVision Language ModelContrastive LearningVideoBiomedical DataUltrasoundRetrieval-Augmented Generation

🎯 What it does: Propose a fetal cardiac ultrasound video diagnostic framework based on view evolution for detecting Tetralogy of Fallot.

PlaneCycle: Training-Free 2D-to-3D Lifting of Foundation Models Without Adapters

Yu, Yinghong (ELLIS Institute Finland), Yang, Jiancheng (Aalto University)

ClassificationSegmentationComputational EfficiencyRepresentation LearningTransformerPrompt EngineeringVision Language ModelContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasound

🎯 What it does: The study proposes a training-free and adapter-free PlaneCycle operator that can seamlessly upgrade pre-trained 2D foundational models into 3D networks without changing the original parameters.

Planning the Perfect Burn: Multi-objective PSO for Multi-needle Thermal Ablation

Mehtali, Jonas, Essert, Caroline

Machine LearningOptimizationConvolutional Neural NetworkImage

🎯 What it does: This paper proposes a new method to address a specific problem, aiming to improve performance and efficiency.

PoCoLT: Position-Aware Cross-Modal Learning with LLM-Decoupled Hierarchical Reports for TMD Diagnosis

Zeng, Xinyi (Sichuan University), Wang, Yan (Sichuan University)

ClassificationAnomaly DetectionRepresentation LearningConvolutional Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose the PoCoLT framework, achieving TMD diagnosis through hierarchical reports generated by combining dual oral posture MRI with LLM decomposition.

PolypSteer: Counterfactual Endoscopic Synthesis via Training-Free Activation Steering

Pham, Trong Thang (University of Arkansas), Le, Ngan (University of Arkansas)

Image TranslationGenerationData SynthesisTransformerPrompt EngineeringDiffusion modelContrastive LearningImage

🎯 What it does: Proposes PolypSteer, a training-agnostic activation scheduling framework for generating pathological and non-pathological endoscopic images under the same anatomical structure.

PoseBridgeNet: Learning a Pose-Evolving Schrödinger Bridge for Multi-Organs Segmentation in Prenatal Volumetric Ultrasound

Gan, Shushen, Gao, Yi (Shenzhen University)

SegmentationPose EstimationConvolutional Neural NetworkTransformerMixture of ExpertsAuto EncoderContrastive LearningImageBiomedical DataUltrasoundStochastic Differential Equation

🎯 What it does: Propose PoseBridgeNet for the segmentation of multiple organs in the fetal abdomen using 3D ultrasound during the mid-pregnancy period, leveraging the feature evolution caused by pose variations to enhance segmentation accuracy.

Positive Semi-definite Group-Aware Instance Disentangled Learning for WSI Representation

Li, Chentao, Qin, Jing (Hong Kong Polytechnic University)

Explainability and InterpretabilityRepresentation LearningTransformerAuto EncoderContrastive LearningGaussian SplattingImageBiomedical DataReview/Survey Paper

🎯 What it does: Propose a three-stage interpretable WSI representation learning framework named PG-CIDL. First, it maps instances into a latent space and clusters them into three groups: tumor, microenvironment, and background via Positive Semi-Definite Latent Factor Grouping (PSD-LFG). Subsequently, it evaluates the causal effects of each group on prediction using Cluster-Reasoning Instance Disentangling (CID) based on causal inference. Finally, it generates decoupled WSI representations by reweighting instances according to their effect weights. The whole process achieves end-to-end optimization.

Posterior-Aware Motor Phenotyping with Multimodal Imaging Validation in Parkinson’s Disease

Tirhekar, Harsh (The University of Texas at Austin), Bajaj, Chandrajit (The University of Texas at Austin)

ClassificationFederated LearningExplainability and InterpretabilityRepresentation LearningHyperparameter SearchData-Centric LearningDrug DiscoveryContrastive LearningGaussian SplattingMultimodalityTabularBiomedical DataMagnetic Resonance ImagingPositron Emission Tomography

🎯 What it does: This paper uses Bayesian Gaussian Mixture Model (BGMM) to cluster 29,366 PPMI visit-level MDS-UPDRS-III assessments, proposes a three-way posterior confidence classification, and verifies it through multi-modal imaging with DaTSCAN SPECT and structural MRI. Finally, the model is migrated to the BioFIND dataset for external validation.