arXivSub Start free trial

MICCAI 2026 Papers — Page 5

International Conference on Medical Image Computing and Computer-Assisted Intervention · 1165 papers

From Geometry to Clinic: A Physics-Aware Evaluation Framework and Benchmark for Automated Dental Crown Design

Wu, Jiamin (University of Hong Kong), Tsoi, James Kit Hon (Fuzhou University)

GenerationTransformerDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningPoint CloudMeshBiomedical DataBenchmarkPhysics Related

🎯 What it does: Proposes CrownBench, a physics-aware evaluation framework and benchmark for assessing automated crown design from a clinical perspective.

From Patches to Patients: A Study of the Tile-to-Slide Performance Transferability in Digital Pathology

Boutaj, Sofiène (Université Paris-Saclay), Marza, Pierre (Université Paris-Saclay)

ClassificationAnomaly DetectionRepresentation LearningConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningImageBiomedical DataBenchmark

🎯 What it does: Investigated the performance transferability of foundation models (FM) from patch-level to slide-level in digital pathology, evaluated the performance of 19 FMs on 16 patch-level tasks and 42 slide-level tasks, and verified that patch-level linear probing can serve as an efficient proxy for slide-level evaluation.

From Perception to Anticipation: Forecasting Vessel–Instrument Interactions in Endoscopic Surgery under Unreliable Observations

Chen, Yueyao (Chinese University of Hong Kong), Dou, Qi (Chinese University of Hong Kong)

ClassificationRecognitionSegmentationExplainability and InterpretabilityComputational EfficiencyRecurrent Neural NetworkTransformerPrompt EngineeringVision Language ModelAuto EncoderContrastive LearningVideoBiomedical DataElectronic Health Records

🎯 What it does: Proposes a task of predicting the imminent interaction status (no interaction, approaching, contact) between vessels and instruments based on short-term videos and segmentation masks under endoscopy, and achieves real-time prediction.

From Point Estimates to Distributions: GMM Pooling for MIL in Preterm Birth Prediction

Alasmawi, Hussain (Mohamed bin Zayed University), Yaqub, Mohammad (Mohamed bin Zayed University)

ClassificationImage TranslationAnomaly DetectionConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningImageBiomedical DataUltrasoundBenchmark

🎯 What it does: Modeling preterm birth prediction as a multiple instance learning (MIL) problem, using multiple transvaginal ultrasound (TVUS) images from each pregnant woman. The Gaussian Mixture Model (GMM) aggregator maps the distribution of image features into a fixed-length representation, which is then used to predict the risk of preterm birth.

From Scanning Guidelines to Action: A Robotic Ultrasound Agent with LLM-Based Reasoning

Bi, Yuan (Technical University of Munich), Navab, Nassir (Technical University of Munich)

Robotic IntelligenceReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AITextBiomedical DataUltrasoundRetrieval-Augmented Generation

🎯 What it does: Achieve autonomous operation of robotic ultrasound based on scanning guidelines by using a large language model as an autonomous planner combined with a set of tools.

From Sparsity to Geometry: Spatial Modeling for 3D Reconstruction from Biplanar Bone X-rays

Long, Nuo’er, Chen, Peikai (University of Hong Kong)

RestorationGenerationDepth EstimationConvolutional Neural NetworkTransformerDiffusion modelContrastive LearningGaussian SplattingImageBiomedical DataComputed TomographyPositron Emission Tomography

🎯 What it does: This paper proposes an explicit spatial modeling framework called ExpGS, based on a fixed virtual dual-view projection paradigm, to reconstruct three-dimensional skeletal structures from dual-plane X-ray images, combining three-dimensional Gaussian splatting with Transformer to achieve reconstruction.

From Trajectories to Phenotypes: Disease Progression as Structural Priors for Multi-organ Imaging Representation Learning

Wang, Zian, Wang, Chengyan (Fudan University)

Knowledge DistillationRepresentation LearningTransformerLarge Language ModelContrastive LearningImageTextBiomedical DataElectronic Health Records

🎯 What it does: Proposed a knowledge distillation framework based on disease trajectories, transferring the trajectory structure knowledge learned by a large-scale diagnostic sequence generative Transformer to a multi-organ imaging-derived phenotypes (IDP) encoder, to improve disease risk and duration prediction.

FSE-Reg: Enhancing 3D Deformable Registration with Frozen Large-Scale Pre-trained Segmentation Encoders

Kang, Hao (East China Normal University), Wen, Ying (DAMO Academy, Alibaba Group)

SegmentationConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningOptical FlowBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: A coarse-to-fine residual pyramid decoder directly applicable to 3D deformation registration is constructed by leveraging a frozen large-scale pre-trained 3D segmentation encoder as an anatomical feature space. More robust correspondence learning is achieved through feature space similarity loss and the DPI-Fuse interaction module.

FunPiQ: A New Benchmark for Pixel-Level Quality Assessment in Fundus Images

Wang, Pengwei (Medical University of Vienna), Bogunović, Hrvoje (Medical University of Vienna)

ClassificationAnomaly DetectionExplainability and InterpretabilityConvolutional Neural NetworkContrastive LearningImageBiomedical DataBenchmark

🎯 What it does: This paper develops the FunPiQ dataset, which uses pixel-level annotations to evaluate the quality of fundus images, and proposes the EFIQA-CP method to achieve interpretable quality prediction.

Fusion-E2Pulse: A Multimodal Event-RGB Fusion Network for Non-contact Pulse Wave Reconstruction

Feng, Qian (Taiyuan University of Technology), Li, Yidi (Taiyuan University of Technology)

RestorationConvolutional Neural NetworkTransformerSupervised Fine-TuningAuto EncoderGenerative Adversarial NetworkContrastive LearningVideoMultimodalityBiomedical Data

🎯 What it does: Designed and implemented a multi-modal network called Fusion-E2Pulse, which integrates event camera and RGB video for high-fidelity non-contact pulse wave reconstruction.

G3R: Gaussian-Based Geometry-Guided Reconstruction for Fetal Brain MRI

Cai, Zhibao (South China University of Technology), Yang, Chaoxiang

RestorationConvolutional Neural NetworkTransformerScore-based ModelNeural Radiance FieldAuto EncoderContrastive LearningGaussian SplattingImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose a 3D fetal brain MRI reconstruction framework based on 3D Gaussian splines (G3R), achieving structurally stable and high-quality reconstructions by introducing geometry-guided priors.

Gabor Primitives for Accelerated Cardiac Cine MRI Reconstruction

Huang, Wenqi (Technical University of Munich), Rueckert, Daniel (Technical University of Munich)

RestorationBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose an adaptive cardiac cine MRI reconstruction framework based on Gabor primitives, achieving scan-specific reconstruction from undersampled k-space data.

GCM-Net: Anatomy-Aware Gaussian-Contrastive Multi-layer Fusion Network for Abdominal Ultrasound Standard Plane Classification

Wang, Anqi (Tsinghua University), Ning, Guochen (Tsinghua University)

ClassificationConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningImageVideoBiomedical DataUltrasound

🎯 What it does: Propose a network called GCM-Net for automatic classification of standard planes in abdominal ultrasound images.

Gene-Selective Morphological Conditioning with Heteroscedastic Flow Matching for Spatial Transcriptomics

Namgung, Hyun (Pohang University of Science & Technology), Park, Sang Hyun (Pohang University of Science & Technology)

GenerationData SynthesisExplainability and InterpretabilityTransformerDiffusion modelScore-based ModelFlow-based ModelBiomedical DataStochastic Differential Equation

🎯 What it does: Utilize gene binding modules and regulators to enable each gene to focus only on the morphological subspace relevant to it, and combine heteroscedastic flow matching to adaptively generate spatial transcriptomics.

GeneRAG: A Retrieval-Augmented Framework for Spatially Resolved Gene Expression Prediction

Kim, Hyeongsub (Seoul National University), Kim, Kyungsu (Seoul National University)

GenerationData SynthesisRetrievalTransformerContrastive LearningBiomedical DataRetrieval-Augmented Generation

🎯 What it does: The GeneRAG framework proposes a retrieval-augmented generation method in spatial transcriptomics, utilizing dual-constrained retrieval based on morphological features and gene expression to predict complete gene expression profiles, particularly capable of predicting genes not seen during training.

Generative Anchor-Guided Federated Domain Generalization for Heterogeneous Pan Cancer Image Analysis

Jin, Qiangguo (Northwestern Polytechnical University), Cong, Cong (Hainan University)

Domain AdaptationFederated LearningTransformerGenerative Adversarial NetworkContrastive LearningImageBiomedical DataBenchmark

🎯 What it does: Proposes the FedGMA framework to address cross-cancer domain generalization issues in federated learning when different centers use heterogeneous label spaces, achieving feature and label decoupling through generated anchors and geometric consistency regularization.

Genetically Aligned Patient Representations Improve Hematological Diagnosis

Dasdelen, Muhammed Furkan (Helmholtz Munich), Marr, Carsten (Helmholtz Munich)

ClassificationRetrievalRepresentation LearningTransformerAuto EncoderContrastive LearningImageMultimodalityTabularBiomedical Data

🎯 What it does: Proposes the GenBloom framework, which aligns single-cell blood smear images with chromosomal aberration and mutation data through multi-modal alignment to learn patient-level encodings and improve hematology diagnostic performance.

GeoCAN: Nonlinear Causal-Geometric Learning for Echocardiography Quality Assessment

Li, Yiran (China University of Petroleum), Li, Shuo (Case Western Reserve University)

TransformerContrastive LearningImageBiomedical DataUltrasound

🎯 What it does: A cardiac ultrasound image quality assessment network named GeoCAN based on nonlinear causal-geometric learning was developed.

Geometric-to-Semantic Spherical Transfer Learning for Cortical Sulci Labeling

Tounsi, Saeb (Paris-Saclay University), Mangin, Jean-François (Paris-Saclay University)

SegmentationConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningMeshBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Under the extremely limited data environment with only 62 expert-annotated subjects, they utilized a large-scale unannotated UK Biobank 30k sample for self-supervised pre-training. Subsequently, they integrated three-channel semantic inputs into the pre-trained spherical encoder through a flexible topological prior injector, ultimately achieving high-precision annotation of more than 60 sulci in the entire right hemisphere of the brain (average Dice 0.77).

Geometry-Aware Manifold Trajectory Modeling for Task fMRI Activation Mapping

Li, Yueran (Harbin Institute of Technology), Su, Jingyong (Sichuan University)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningBiomedical DataMagnetic Resonance Imaging

🎯 What it does: This paper proposes a geometry-aware manifold trajectory modeling framework for task fMRI activation mapping without the HRF assumption.

Geometry-Aware Point-to-Voxel Fusion via Gaussian Splatting for Intracranial Aneurysm Segmentation

Qin, Tian (University of Sydney), Luo, Tao (Beijing University of Posts and Telecommunications)

SegmentationTransformerAuto EncoderContrastive LearningGaussian SplattingPoint CloudBiomedical DataComputed Tomography

🎯 What it does: This paper proposes a geometry-aware point-to-voxel fusion framework called GeoP2VNet, aimed at improving the segmentation accuracy of intracranial aneurysms in CTA images.

Geometry-Aware Self-supervised Learning for Whole-Heart 3D+t Representation

Hu, Liwei (Imperial College London), Yang, Guang (Imperial College London)

SegmentationTransformerAuto EncoderContrastive LearningBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose a geometry-aware self-supervised learning framework for learning full-heart 3D+t representations from sparse multi-plane cardiac magnetic resonance imaging.

Geometry-Aware Superpixel Graph Transformer with Metadata for Skin Lesion Classification

Azeem, Muhammad (Edge Hill University), Behera, Ardhendu (Edge Hill University)

ClassificationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningConvolutional Neural NetworkGraph Neural NetworkTransformerDiffusion modelScore-based ModelFlow-based ModelRectified FlowAuto EncoderContrastive LearningImageGraphBiomedical DataReview/Survey Paper

🎯 What it does: Proposed a Transformer framework based on geometry-aware superpixel graphs (GeoMeta-GT), which segments skin lesion images into superpixel nodes and embeds patient metadata as dedicated nodes into the graph, achieving region-level relationship modeling and multimodal fusion.

GES-Net: A Gesture, Error, and Smoothness Aware Network for Surgical Skill Assessment

He, Runlong (University College London), Mazomenos, Evangelos B. (University College London)

ClassificationRecognitionAnomaly DetectionRobotic IntelligenceConvolutional Neural NetworkTransformerPrompt EngineeringMixture of ExpertsVision Language ModelContrastive LearningVideoMultimodalityBiomedical DataBenchmark

🎯 What it does: GES-Net proposes a multi-task learning framework that jointly evaluates robotic-assisted surgical techniques by integrating three modules: gesture recognition, error detection, and operational smoothness.

GFR-MIL: Glance–Focus–Reflect Based Multiple Instance Learning for Multi-Scale Whole Slide Image Analysis

Qi, Mingxin (Beihang University), Mu, Wei (Beihang University)

ClassificationImage TranslationAnomaly DetectionComputational EfficiencyRepresentation LearningTransformerContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Proposed the GFR-MIL framework, which utilizes a three-stage iterative closed-loop, Glance-Focus-Reflect, to achieve bidirectional interaction between low-resolution global information and high-resolution details, and combines a recycling mechanism and memory mask to enhance the performance of whole slide image (WSI) multi-instance learning.

GG-AE: Genetically-Guided Autoencoder for Disentangling Structural Networks of Brain Atrophy

De Pian, Marilena (University of Pennsylvania), Davatzikos, Christos (University of Pennsylvania)

Explainability and InterpretabilityRepresentation LearningGraph Neural NetworkAuto EncoderContrastive LearningBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease

🎯 What it does: Proposed a gene-guided self-supervised autoencoder (GG-AE), which first learns a compact genetic latent space via a variational autoencoder (VAE) of genes, and then uses a dual-latent-space non-negative decoder to decompose brain volume degeneration images into a genetically aligned subspace and a residual subspace, thereby achieving disentanglement between genetically driven brain atrophy patterns and non-genetic changes.

GGFI-Net: Genomics-Guided Cross-Modal Imputation for Multimodal Survival Analysis with Missing Modality

Liu, Jiaxuan (Politecnico di Milano), Corino, Valentina D. A. (Centro Cardiologico Monzino IRCCS)

Knowledge DistillationRepresentation LearningData-Centric LearningTransformerAuto EncoderContrastive LearningImageMultimodalityTabularBiomedical DataComputed Tomography

🎯 What it does: Propose the GGFI-Net framework, which utilizes genome-guided cross-modal feature interpolation and knowledge distillation to achieve multi-modal survival prediction, maintaining high predictive performance even when imaging modalities are missing.

Glasses: Adapter-Enhanced DINOv3 with Boundary-Aware Multi-Scale Representations for Medical Image Segmentation

Zhu, Jiahao (Northwestern Polytechnical University), Xia, Yong (Northwestern Polytechnical University)

SegmentationConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningImageBiomedical DataComputed Tomography

🎯 What it does: Proposed a novel medical image segmentation framework called Glasses, which utilizes the frozen DINOv3 and combines it with a multi-scale Deformable Attention Adapter and an edge-aware convolutional encoder.

GLeVE: Graph-Guided Lesion Grounding with Proposal Verification in 3D CT

Jiang, Shuo (Hangzhou Dianzi University), Chen, Yifei (Hangzhou Dianzi University)

SegmentationGraph Neural NetworkTransformerVision Language ModelAuto EncoderContrastive LearningImageBiomedical DataComputed Tomography

🎯 What it does: Propose a graph-structure-based lesion description modeling and verification framework called GLeVE, which can directly align and precisely locate lesion descriptions from radiology reports in three-dimensional CT images.

Global-Local Structure-Aware Alignment for Automated LGE MRI Report Generation

Zhou, Chenyang (Beijing Normal University), Duan, Jinming (University of Manchester)

GenerationGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingRetrieval-Augmented Generation

🎯 What it does: This paper proposes the GLSA-RG framework to achieve automatic report generation based on LGE MRI, utilizing structural graph representation and structural alignment to capture regional-level enhancement patterns and cross-slice continuity.

Good Enough? An Investigation on the Impact of Label Quality in Large-Scale Medical Datasets

Jaus, Alexander (Karlsruhe Institute of Technology), Stiefelhagen, Rainer (Karlsruhe Institute of Technology)

SegmentationConvolutional Neural NetworkTransformerSupervised Fine-TuningBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: This paper investigates the impact of label quality on model performance on large-scale medical image segmentation datasets, including two scenarios: direct training and pre-training followed by fine-tuning.

Graph Neural Networks for Morphometric Characterization of the Aortic Valve

Ibragimov, Kamil (University of Ljubljana), Vrtovec, Tomaz (University of Ljubljana)

Graph Neural NetworkSupervised Fine-TuningAuto EncoderContrastive LearningImageBiomedical DataComputed Tomography

🎯 What it does: This paper proposes a morphology analysis method for aortic valve based on graph neural networks (GNN), which constructs a graph structure using the candidate distribution of detected key anatomical landmarks and performs adaptive correction of morphological measurements through GNN.

Graph-of-Differences: Anatomy-Structured Difference Alignment for Medical Image Re-identification

Wasalathilaka, Nichula (University of Peradeniya), Mahapatra, Dwarikanath (Khalifa University)

RecognitionExplainability and InterpretabilityGraph Neural NetworkTransformerAuto EncoderContrastive LearningImageGraphBiomedical DataComputed TomographyElectronic Health Records

🎯 What it does: Propose the Graph-of-Differences (GoD) model, which utilizes anatomical structure graphs for patient medical image re-identification, and uses the differences of named anatomical nodes in comparison.

Group Equivariant Diffusion for Anomaly Detection in Computational Cytology

Chatterjee, Swarnadip (Uppsala University), Mukhopadhyay, Anirban (Technical University of Darmstadt)

Anomaly DetectionConvolutional Neural NetworkDiffusion modelScore-based ModelAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasoundDiffusion Tensor ImagingAlzheimer's DiseaseElectronic Health RecordsElectrocardiogram

🎯 What it does: Proposed a D4-equivariant diffusion model that enforces symmetry constraints on cellular-level medical images for unsupervised anomaly detection.

Group-Conditioned Representation Modulation for Fair Skin Disease Diagnosis

Xu, Gelei (University of Notre Dame), Shi, Yiyu (University of Notre Dame)

ClassificationDomain AdaptationFederated LearningExplainability and InterpretabilityRepresentation LearningConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningImageBiomedical DataBenchmark

🎯 What it does: What was done: Proposed a group-specific intervention mask framework (GIM), which improves fairness in dermatological diagnosis by generating binary channel masks on a shared backbone network to modulate activations for different populations.

GTRE: A Physiological-Semantic Foundation Model for Generalizable EEG Analysis via Preference-Guided Alignment

Azfar, Mohd, Khan, Izhar Dad

ClassificationRetrievalTransformerSupervised Fine-TuningContrastive LearningTextMultimodality

🎯 What it does: This paper proposes a cross-modal semantic trajectory model that can simultaneously align and learn representations of EEG signals and text during different pre-training stages.

GU-SAM2: A Semi-supervised Medical Image Segmentation Framework via Gated Feature Fusion of MIM-Pretrained U-Net and SAM 2

Liu, Yinjun, Wang, Hongkai (Dalian University Of Technology)

SegmentationConvolutional Neural NetworkTransformerContrastive LearningBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose the GU-SAM2 framework, which first pretrains the U-Net using masked image modeling on unlabeled data, and then injects medical domain knowledge into SAM2 through a gated feature fusion module, achieving semi-supervised medical image segmentation.

H2GT: Hub-enhanced Hierarchical Graph Transformer for Explainable Diagnosis of ASD with Functional Brain Networks

Xiu, Yiqi, Zhao, Shijie (Northwestern Polytechnical University)

ClassificationExplainability and InterpretabilityGraph Neural NetworkTransformerContrastive LearningGraphBiomedical DataMagnetic Resonance Imaging

🎯 What it does: For the diagnosis of autism spectrum disorder (ASD), this paper proposes and implements a Hub-enhanced Hierarchical Graph Transformer (H GT) model based on functional brain networks (FBN), which can simultaneously learn local node interactions and global module collaboration.

H2M-Net: Hierarchical Hyperbolic Memory Network for Pathological Report Generation

Yao, Shuilian (Dalian University of Technology), Fan, Xin (Dalian University of Technology)

GenerationRepresentation LearningTransformerVision Language ModelDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageTextBiomedical DataElectronic Health Records

🎯 What it does: A framework for automatically generating pathology reports based on a hierarchical hyperbolic space memory network (H-M-Net) was constructed, embedding whole slide images (WSI) and corresponding text in the Lorentz model, and generating layer by layer from bottom to top through a cross-layer memory module.

HadBalance: A Plug-and-Play Unified Global Geometric Prior Framework for Generalizable Biomedical Segmentation

Gao, Zhuangzhi (University of Liverpool), Zheng, Yalin (University of Liverpool)

SegmentationContrastive LearningOptical FlowBiomedical DataUltrasoundStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: This paper proposes the HadBalance framework, which combines the Hadwiger shape prior with conflict-aware target balancing for general medical image segmentation.

Hallucination Detection and Correction in Medical VLMs via Counter-Evidence Verification

Zhou, Nan (Sichuan University), Fu, Huazhu (Agency for Science, Technology and Research)

Anomaly DetectionExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningDrug DiscoveryPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataElectronic Health RecordsReview/Survey PaperRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposes a training-agnostic, plug-and-play CoEV framework that utilizes adversarial masking of visual regions to detect and correct hallucinations in medical vision-language models

HAM-BGT: Harmonic-Aware Multimodal Brain Graph Transformer for Brain Disease Diagnosis

Xu, Heming (Xi'an Jiaotong University), Du, Shaoyi (Xi'an Jiaotong University)

ClassificationAnomaly DetectionGraph Neural NetworkTransformerContrastive LearningGraphBiomedical DataDiffusion Tensor ImagingAlzheimer's Disease

🎯 What it does: Proposed a multimodal brain graph convolutional transformer based on connectome harmonics (HAM-BGT), which integrates functional and structural brain networks for brain disease diagnosis through spectral graph filtering encoding and harmonic-aware filters.

HAR-MoE: Hierarchical Affinity-aware Routing Mixture-of-Experts for Unified Multi-Organ Ultrasound Analysis

Mu, Yajun (Southwest Jiaotong University), Kang, Qingbo (Sichuan University)

ClassificationSegmentationTransformerMixture of ExpertsContrastive LearningBiomedical DataUltrasound

🎯 What it does: This paper proposes a Hierarchical Affinity-Aware Routing Mixture-of-Experts (HAR-MoE) model that jointly learns pixel-level segmentation and image-level classification tasks for multi-organ ultrasound images.

Harmless Copyright Protection for CT Dataset

Xu, Fengzhi (Sichuan University), Zhang, Yi (Sichuan University)

Safty and PrivacyConvolutional Neural NetworkAuto EncoderContrastive LearningBiomedical DataComputed Tomography

🎯 What it does: Proposed a harmless copyright protection paradigm and a learnable measurement domain harmless backdoor watermark (LMHBW), which embeds a watermark in the CT measurement domain and verifies dataset ownership using the prediction accuracy of watermark samples;

Harmonic Kernel Mixing for Interpretable Oculomotor Event Segmentation in Video-Nystagmography

Chaturvedi, Kunal (University of Technology Sydney), Prasad, Mukesh (University of Technology Sydney)

SegmentationExplainability and InterpretabilityConvolutional Neural NetworkTransformerTime SeriesBiomedical DataElectrocardiogram

🎯 What it does: Designed NyST-Net, which combines an interpretable spectral frontend based on harmonic kernel mixing (HKM) and a multi-scale temporal network, to achieve precise segmentation of slow phases, fast phases, no eye movement, and artifacts in video nystagmography (VNG), and to reconstruct slow phase velocity (SPV) curves using the segmentation results.

Harnessing Adversarial Distillation to Customise Debiased, Disease-Specific Pathology Foundation Models for Breast Cancer

Chen, Zhiwei (South China University of Technology), Yang, Kaixiang (South China University of Technology)

ClassificationAnomaly DetectionComputational EfficiencyKnowledge DistillationRepresentation LearningConvolutional Neural NetworkTransformerAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical Data

🎯 What it does: Construct a lightweight, bias-free breast cancer-specific pathological foundation model through multi-teacher knowledge distillation and adversarial distillation

HbO2 and HbR Spatiotemporal Heatmap Feature Fusion Network for ADHD Diagnosis in Children

Chu, Mengxiang (Northwest University), Yu, Jingjing (Shaanxi Normal University)

ClassificationAnomaly DetectionRecurrent Neural NetworkTransformerContrastive LearningTime SeriesSequentialBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease

🎯 What it does: A spatiotemporal heatmap feature fusion network (STHFFN) for the auxiliary diagnosis of ADHD children based on functional near-infrared spectroscopy (fNIRS) is proposed. It converts discrete HbO2/HbR signals into continuous heatmaps and jointly models brain region activation, hemoglobin coupling, and long-term attention instability using parallel spatiotemporal attention, bidirectional hemoglobin complementary attention, and memory pool mechanisms.

HD-TTA: Hypothesis-Driven Test-Time Adaptation for Safer Brain Tumor Segmentation

Jhawar, Kartik (Nanyang Technological University), Wang, Lipo (Nanyang Technological University)

SegmentationDomain AdaptationExplainability and InterpretabilityComputational EfficiencyConvolutional Neural NetworkAuto EncoderContrastive LearningBiomedical DataMagnetic Resonance Imaging

🎯 What it does: This paper proposes the Hypothesis-Driven Test-Time Adaptation (HD-TTA) framework, aiming to achieve safe and interpretable brain tumor segmentation. Instead of blindly performing global optimization, the framework adapts the model during testing by selecting geometric hypotheses (compression and expansion) and using an unsupervised selector for self-optimization.

HeartFormer: Semantic-Aware Dual-Structure Transformers for 3D Four-Chamber Cardiac Point Cloud Reconstruction

Ma, Zhengda (University of Oxford), Banerjee, Abhirup (University of Oxford)

RestorationGenerationTransformerDiffusion modelContrastive LearningPoint CloudBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Designed HeartFormer, a dual-structure semantic-aware Transformer, for reconstructing complete four-chamber cardiac point clouds from conventional short-axis and long-axis Cine MRI.

HeartVolMesh: Cardiac Volumetric Mesh Reconstruction via Covariance-Guided Graph Deformation

Lin, Fengming (University of Manchester), Frangi, Alejandro F. (University of Manchester)

SegmentationGenerationData SynthesisConvolutional Neural NetworkGraph Neural NetworkDiffusion modelScore-based ModelContrastive LearningGaussian SplattingImageMeshBiomedical DataComputed Tomography

🎯 What it does: Proposed the HeartVolMesh framework, achieving end-to-end generation from CTA images to tetrahedral cardiac volume meshes with consistent vertex correspondence.

HEDS-Net: A Hybrid State-Space Architecture with Axial Bridge and Progressive Weighting for Medical Image Segmentation

Fan, Jiaying (Guilin University of Electronic Technology), Li, Yuexiang

SegmentationDepth EstimationConvolutional Neural NetworkTransformerContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound

🎯 What it does: Proposed a medical image segmentation network called HEDS-Net, which combines global sequence modeling and local detail enhancement, and introduces axial bridging and progressive weighted deep supervision within a U-shaped structure.

HemoPIC: A Physics-Informed Cerebral Hemodynamics Digital Twin for Brain Perfusion

Lee, Yi-Chen (Johns Hopkins University), Liu, Peirong (Johns Hopkins University)

Diffusion modelAuto EncoderContrastive LearningBiomedical DataMagnetic Resonance ImagingPhysics Related

🎯 What it does: A physics-based brain perfusion digital twin, HemoPIC, was constructed to directly estimate perfusion indicators such as blood flow, blood volume, and transit time from dynamic perfusion imaging, without the need for manually selecting an arterial input function.

Hepato-LLaVA: An Expert MLLM with Sparse Topo-Pack Attention for Hepatocellular Pathology Analysis on Whole Slide Images

Yang, Yuxuan, Ma, Zhanyu

RecognitionTransformerLarge Language ModelPrompt EngineeringVision Language ModelAuto EncoderContrastive LearningImageTextMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyBenchmark

🎯 What it does: Propose Hepato-LLaVA, a multimodal large language model specialized for hepatocellular carcinoma (HCC) pathological images, and construct the HepatoPathoVQA multi-scale question-answering dataset.

HERO: Hypothesis-Driven Evidence Retrieval from Omics for Multi-task Breast Cancer Analysis

Li, Xiangyu (Tianjin University), Su, Ran (Tianjin University)

ClassificationExplainability and InterpretabilityComputational EfficiencyDrug DiscoveryTransformerSupervised Fine-TuningPrompt EngineeringVision Language ModelContrastive LearningMultimodalityTabularBiomedical DataElectronic Health RecordsRetrieval-Augmented Generation

🎯 What it does: By mapping DNA methylation and miRNA data into sparse morphological intention vectors, utilizing TF-IDF keyword retrieval combined with a vision-language model, and employing cosine consistency gates and defect-driven repair mechanisms, auditable multi-task breast cancer prediction is achieved.

Heterogeneity-Adaptive Diffusion Schrödinger Bridge for PET-Guided Whole-Body MRI Translation

Wang, Chengbo (University of Sydney), Wang, Xiuying (University of Sydney)

Image TranslationRestorationSuper ResolutionConvolutional Neural NetworkTransformerVision Language ModelDiffusion modelScore-based ModelContrastive LearningImageBiomedical DataMagnetic Resonance ImagingPositron Emission TomographyStochastic Differential Equation

🎯 What it does: A whole-body PET-guided MRI translation model based on the differential Schrödinger bridge (HA-DSB) was studied, achieving the synthesis of high-resolution T2 sequences from rapidly acquired LAVA sequences.

Hi-FL: Hierarchical Federated Learning Adaptation of Vision-Language Models for Multi-Institutional Surgical Phase Recognition

Alekseenko, Julia (University of Strasbourg), Padoy, Nicolas (University of Strasbourg)

RecognitionFederated LearningTransformerSupervised Fine-TuningVision Language ModelContrastive LearningVideoText

🎯 What it does: Surgical phase recognition using hierarchically adapted vision-language models across multiple institutions.

HiEDL: Hierarchical Evidential Deep Learning for Uncertainty-Aware Tumor Segmentation from 3D CT via Boundary Regularization

Park, Younghyun (Yonsei University), Yang, Sejung (Yonsei University)

SegmentationConvolutional Neural NetworkTransformerBiomedical DataComputed Tomography

🎯 What it does: Proposed the HiEDL framework, which decomposes tumor segmentation into a two-level hierarchical task of organ-tumor, models predictive uncertainty using evidential deep learning, and achieves uncertainty-aware 3D CT tumor segmentation by combining distance map boundary loss

Hierarchical Cautious Optimization for Semi-supervised Medical Image Segmentation

Olender, Aaron (Hebrew University of Jerusalem), Joskowicz, Leo (Hebrew University of Jerusalem)

SegmentationOptimizationMixture of ExpertsContrastive LearningBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: In semi-supervised learning for medical image segmentation, a novel optimizer called Hierarchical Cautious Optimization (HCO) is proposed. It determines a trustworthy direction using gradients from labeled data in momentum updates, allowing only unlabeled gradients that align with this direction to participate, thereby improving learning stability and performance.

Hierarchical Cross-Modal Fusion with Modality-Specific Training Paradigms for High-Precision Non-invasive Aging Assessment

Wu, Yihang (Zhejiang University), Huang, Shuaihan (Zhejiang University)

ClassificationPose EstimationAnomaly DetectionRepresentation LearningConvolutional Neural NetworkTransformerMixture of ExpertsContrastive LearningImageVideoMultimodalityBiomedical DataAlzheimer's DiseaseElectronic Health Records

🎯 What it does: Propose a hierarchical cross-modal fusion framework that integrates three complementary modalities—ocular imaging, behavior, and gait—to achieve high-precision non-invasive aging assessment.

Hierarchical Disentangled Consistency Learning for Semi-supervised Multi-modal Medical Image Segmentation

Zhou, Feixiang (University of Liverpool), Zheng, Yalin (University of Liverpool)

SegmentationConvolutional Neural NetworkTransformerAuto EncoderGenerative Adversarial NetworkContrastive LearningMultimodalityBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: This study investigates semi-supervised multi-modal medical image segmentation, proposing the HDCL framework, which leverages hierarchical decoupled representations and prediction consistency to improve segmentation performance.

Hierarchical Hyperbolic Self-Attention Network for Medical Image Segmentation

Jin, Shaocheng (Jiangnan University), Zhou, Tao (Jiangnan University)

SegmentationConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical Data

🎯 What it does: Proposed H₂SA-Net, a medical image segmentation framework that integrates waveform transformation with hierarchical hyperspherical self-attention.

Hierarchical Memory for Radiology Report Generation with Rich Context

Jiao, Qingyue (University of Notre Dame), Shi, Yiyu (University of Notre Dame)

GenerationRetrievalRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningImageTextBiomedical DataComputed TomographyElectronic Health RecordsRetrieval-Augmented Generation

🎯 What it does: Proposed a hierarchical memory-based radiology report generation framework called HM-RRG, which utilizes cross-patient retrieved reports and longitudinal historical information from the same patient;

Hierarchical Multi-Video Multiple Instance Learning: AI-Interpreted Point-of-Care Ultrasound for TB Detection

Brokowski, Trevor (Yale University), Hartley, Mary-Anne (Yale University)

ClassificationAnomaly DetectionExplainability and InterpretabilityTransformerVision Language ModelContrastive LearningVideoBiomedical DataUltrasound

🎯 What it does: Proposed a hierarchical multi-video multi-instance learning (HMV-MIL) method for directly performing patient-level tuberculosis (TB) detection from multi-video ultrasound images.

Hierarchical Region-Aware Multi-Granularity Mamba for White Matter Lesion Segmentation

Lee, Dahye (DEEPNOID Inc.), Oh, Kwanseok (Korea University)

SegmentationConvolutional Neural NetworkTransformerSupervised Fine-TuningDiffusion modelContrastive LearningBiomedical DataMagnetic Resonance Imaging

🎯 What it does: This paper proposes a hierarchical region-aware multi-granularity Mamba (HiRAM) model for the white matter lesion segmentation task, and integrates it into the SAM decoder to achieve fine-grained processing of internal, boundary, and background region features.

Hierarchical Text-Guided Brain Tumor Segmentation via Sub-Region-Aware Prompts

Mohammadi, Bahram (Macquarie University), Qi, Yuankai (Macquarie University)

SegmentationConvolutional Neural NetworkTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose the TextSCP framework, which uses hierarchical text guidance to achieve fine segmentation of three sub-regions of brain tumors (WT, TC, ET).

HiFi-Rep: Leveraging High-Resolution Vision Foundation Models with Tri-Planar Slice Context for CT Report Generation

Limaroon, Keetawan (King Mongkut's University of Technology Thonburi), Tarnpradab, Sansiri (King Mongkut's University of Technology Thonburi)

GenerationTransformerLarge Language ModelPrompt EngineeringContrastive LearningImageTextBiomedical DataComputed TomographyRetrieval-Augmented Generation

🎯 What it does: Propose the HiFi-Rep framework, which utilizes high-resolution DINOv3 features and HiFi-S2D aggregation, combined with the MedGemma LLM to achieve automatic generation of CT reports.

High-Capacity Robust Medical Image Exfiltration via Neural Network Weight Replacement

Thellier, Elie (Université Côte d'Azur, Inria), Delingette, Hervé (Université Côte d'Azur, Inria)

Safty and PrivacyDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasound

🎯 What it does: High-capacity medical image steganography is achieved by embedding continuous latent codes into the weight initialization of neural networks, allowing the model to recover images after exporting while maintaining the original task performance.

High-Frequency Guided Feature Refinement for Semi-supervised Ultrasound Image Segmentation

Wu, Xiaming (Wuhan Institute of Technology), Xu, Guoping (Hefei University of Technology)

SegmentationConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningImageBiomedical DataUltrasound

🎯 What it does: Propose the HG‑SemiSeg framework, which improves the quality of pseudo labels in semi-supervised ultrasound image segmentation through high-frequency information enhancement.

HiLoGraph: Hierarchical-Longitudinal Brain Network Representation Learning

Chen, Tong (University of Texas at Arlington), Zhu, Dajiang (University of Georgia)

Anomaly DetectionRepresentation LearningGraph Neural NetworkAuto EncoderContrastive LearningGraphBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease

🎯 What it does: Propose a self-supervised multi-scale hierarchical graph neural network called HiLoGraph, which learns a shared latent space of brain networks using longitudinal structural magnetic resonance images from healthy subjects, to establish a normalized benchmark and detect Alzheimer's disease and Lewy body dementia.

HiPath: Hierarchical Vision-Language Alignment for Structured Pathology Report Prediction

Yuan, Ruicheng (Hunan University), Yang, Guang (Imperial College London)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningVision Language ModelContrastive LearningImageTextBiomedical DataElectronic Health Records

🎯 What it does: Designed and trained HiPath, a lightweight vision-language model for predicting structured pathological reports.

HiProS: A Hierarchical Prompt Selection for Stage-wise Vision–Language Alignment in Medical Image Segmentation

Kim, Da-Hee (Kyungpook National University), Lee, Dong-Gyu (Kyungpook National University)

SegmentationTransformerPrompt EngineeringVision Language ModelImageTextBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Proposed a hierarchical prompt selection framework called HiProS to achieve staged visual-language alignment in medical image segmentation.

HK-Fuse: Hilbert-Interleaved Kimi Delta Attention for Incomplete-Modality Brain Tumor Segmentation

Zhang, Weizhi (Nanjing University), Shi, Yinghuan (Nanjing University)

SegmentationTransformerMultimodalityBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose a Hilbert-Interleaved Kimi Fusion framework named HK-Fuse for brain tumor segmentation with incomplete modalities.

HomoSeg: A Weak-Label-Guided Semi-supervised Framework for Tumor Microenvironment Segmentation

Zhao, Haoyun, Tao, Dapeng (Yunnan University)

SegmentationData SynthesisKnowledge DistillationConvolutional Neural NetworkTransformerAuto EncoderGenerative Adversarial NetworkContrastive LearningBiomedical DataReview/Survey PaperBenchmark

🎯 What it does: Proposes HomoSeg, a weakly labeled guided semi-supervised framework for pixel-level segmentation of the tumor microenvironment.

HoughSAM: Geometry-Guided Zero-Shot Segmentation of Main Incision in Cataract Surgery Using SAM2 for Skill Assessment

Dai, Weina (Johns Hopkins University), Vedula, Swaroop (Johns Hopkins University)

SegmentationAnomaly DetectionComputational EfficiencyTransformerPrompt EngineeringAuto EncoderContrastive LearningOptical FlowVideoBiomedical DataMagnetic Resonance ImagingComputed TomographyReview/Survey Paper

🎯 What it does: Designed and implemented an untrained zero-shot segmentation framework called HoughSAM, which uses the Hough transform to automatically generate prompt points, combines SAM2 to segment the main incision in cataract surgery, and extracts incision geometric features for technical evaluation.

How Good Are Foundation Models in Longitudinal MRI Disease Progression Reasoning?

Al Ghallabi, Wafa (Mohamed bin Zayed University of Artificial Intelligence), Khan, Fahad Shahbaz (Mohamed bin Zayed University of Artificial Intelligence)

TransformerPrompt EngineeringVision Language ModelImageMultimodalityTime SeriesBiomedical DataMagnetic Resonance ImagingBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This paper designs a Time-Aware Multi-View MRI Benchmark, integrating multi-disease, longitudinal, and multi-view MRI data with structured question answering to evaluate the ability of vision-language models in temporal and spatial reasoning.

Human-AI Collaboration for 2D/3D Registration Quality Assurance

Cho, Sue Min (Johns Hopkins University), Unberath, Mathias (Johns Hopkins University)

ClassificationExplainability and InterpretabilityConvolutional Neural NetworkContrastive LearningOptical FlowImageBiomedical DataComputed TomographyPositron Emission TomographyUltrasound

🎯 What it does: Designed and implemented a user study to evaluate the effectiveness of human-AI (including explainable AI) collaboration in 2D/3D image registration quality assurance tasks. The study used a learning-based registration quality judgment model, incorporated Grad-CAM heatmaps and confidence indicators in the interface, and compared three decision-making modes: Human-only, Human+AI, and Human+XAI.

Human-In-The-Loop Multi-agent Ventilator Decision Support with Contextual Bandit Preference Learning

Li, Sijia (Shanghai University of Engineering Science), Qiu, Xihe (Shanghai University of Engineering Science)

Autonomous DrivingOptimizationFederated LearningSafty and PrivacyExplainability and InterpretabilityComputational EfficiencyRobotic IntelligenceMeta LearningDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsVision Language ModelTabularTime SeriesBiomedical DataElectronic Health RecordsChain-of-Thought

🎯 What it does: Developed a human-machine collaborative multi-agent ventilator decision support system (VDSS), which utilizes context bandit learning to achieve online adaptation to doctors' preferences.

Hybrid Transformer-Mamba for Weakly Supervised Volumetric Medical Segmentation

Lyu, Yiheng (University Of Western Australia), Dwivedi, Girish

SegmentationTransformerContrastive LearningBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound

🎯 What it does: A hybrid Transformer-Mamba structure called TranSamba was studied for weakly supervised volumetric medical segmentation based on planar-level labels.

Hyperbolic Contrastive Learning with Entailment for Spatial Transcriptomics

Vega, Daniela (Universidad de los Andes), Arbeláez, Pablo (Universidad de los Andes)

Representation LearningData-Centric LearningTransformerContrastive LearningImageMultimodalityBiomedical DataBenchmarkRetrieval-Augmented Generation

🎯 What it does: This paper proposes a framework based on hyperbolic contrastive learning and inference loss, utilizing pathological images to predict gene expression in spatial transcriptomics.

Hyperbolic Vision-Language Interaction for Semi-supervised Medical Image Segmentation

He, Qiuchi (Jiangnan University), Zhou, Tao (Jiangnan University)

SegmentationTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose a visual-language interaction framework based on hyperbolic geometry, named HyperSemi, for semi-supervised medical image segmentation;

HyperVLP: Enhancing Hierarchical Surgical Video-Language Pre-training in Hyperbolic Space

Hu, Yaojun (Zhejiang University), Padoy, Nicolas (Zhejiang University)

RetrievalRepresentation LearningTransformerVision Language ModelContrastive LearningVideoTextMultimodality

🎯 What it does: Propose a hierarchical surgical video-language pre-training framework based on hyperbolic space, named HyperVLP.

HypOProto: Hyperbolic Ordinal Prototypes for Left Ventricular Filling Pressure Classification

Wu, Victoria (University of British Columbia), Tsang, Teresa S. M. (University of British Columbia)

ClassificationExplainability and InterpretabilityTransformerSupervised Fine-TuningContrastive LearningImageVideoBiomedical DataUltrasound

🎯 What it does: Implement an interpretable classification of left ventricular filling pressure (LVFP) on B-mode ultrasound images using a frozen DINOv3 base model, proposing the HypOProto method, which maps prototypes into a hyperbolic (Lorentzian) space, organizing them sequentially according to clinical E/e' thresholds;

I²-Med: Interpretable Medical Inference Through Visual-Guided Dynamic Logits Calibration

Sajid, Md (Indian Institute of Technology Indore), Tanveer, Mohammad (Indian Institute of Technology Indore)

Explainability and InterpretabilityComputational EfficiencyTransformerReinforcement LearningPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound

🎯 What it does: Propose a vision-guided dynamic logits calibration framework called I-Med, which combines structured reasoning based on reinforcement learning with visual inspection during inference, thereby improving the interpretability and visual authenticity of medical vision-language models without altering model parameters.

IAVS: A Multi-center Dataset and Applicability Evaluation System for Computational Fluid Dynamics-Oriented Intracranial Aneurysm Segmentation

Xiao, Feiyang (Shanghai Academy of Artificial Intelligence for Science), Cheng, Yuan (Artificial Intelligence Innovation and Incubation Institute, Fudan University)

SegmentationConvolutional Neural NetworkDiffusion modelContrastive LearningImageMeshBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Constructed the IAVS multi-center intracranial aneurysm and vascular segmentation dataset, and designed an automatic CFD applicability assessment system and CFD-AS metric to evaluate the usability of segmentation results in clinical CFD modeling.

ICH-DIANet: A Modality-Disentangled and Attention-Expert Interaction Network for Brain Injury Assessment in Intracerebral Hemorrhage

Liu, Chengqing (Nanchang University), Shu, Xujun (Nanjing University)

ClassificationAnomaly DetectionConvolutional Neural NetworkTransformerSupervised Fine-TuningMixture of ExpertsContrastive LearningImageTextBiomedical DataComputed TomographyElectronic Health Records

🎯 What it does: Proposed ICH-DIANet, a brain injury assessment network that integrates 3D CT images with clinical text;

ICHOR: A Robust Representation Learning Approach for ASL CBF Maps with Self-supervised Masked Autoencoders

Beltran-Urbano, Xavier (University of Pennsylvania), Dolui, Sudipto (University of Pennsylvania)

ClassificationRepresentation LearningTransformerSupervised Fine-TuningAuto EncoderContrastive LearningBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease

🎯 What it does: Proposed a self-supervised pre-training framework called ICHOR based on 3D Masked Autoencoders for representation learning in ASL CBF images, and fine-tuned it for downstream tasks using LoRA.

IgG4-QPC: Quantitative Pathology-Consistent Virtual Staining from H&E to IgG4

Luo, Mengyue (Sichuan University), Zhang, Haixian (Sichuan University)

Image TranslationSegmentationGenerationData SynthesisConvolutional Neural NetworkGenerative Adversarial NetworkContrastive LearningImageBiomedical DataPositron Emission TomographyReview/Survey Paper

🎯 What it does: This paper proposes a model called IgG4-QPC for virtual staining of H&E images to IgG4 IHC, aiming to achieve discrete instance-level quantitative consistency of IgG4-positive plasma cells under conditions of non-specific staining (NSS) noise.

Image-mediated fMRI-to-caption generation with visual pathway tokens and hyperbolic alignment

Choi, Hyoungshin (Sungkyunkwan University), Park, Hyunjin (Sungkyunkwan University)

GenerationKnowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringVision Language ModelDiffusion modelGenerative Adversarial NetworkContrastive LearningImageTextMultimodalityBiomedical DataMagnetic Resonance ImagingRetrieval-Augmented Generation

🎯 What it does: Propose a two-stage brain-image-text generation framework, which first maps fMRI signals to image embeddings, and then generates natural language captions using LLaMA3.

Imbalance-Aware Distributional Alignment of Heterogeneous Modalities for HER2 Status Prediction

Huang, Ximeng (Xidian University), Zhang, Jianjia (Xidian University)

ClassificationImage TranslationImage HarmonizationRestorationObject DetectionObject TrackingSegmentationGenerationData SynthesisPose EstimationDepth EstimationSuper ResolutionRetrievalCompressionDomain AdaptationRecommendation SystemAnomaly DetectionAutonomous DrivingOptimizationFederated LearningSafty and PrivacyExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningAdversarial AttackHyperparameter SearchData-Centric LearningRobotic IntelligenceMeta LearningDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackNeural Architecture SearchProtein Structure PredictionConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasoundDiffusion Tensor ImagingAlzheimer's DiseaseElectronic Health Records

🎯 What it does: This paper proposes a novel deep learning model that integrates convolutional networks with self-attention mechanisms to address image classification/segmentation (or time series prediction) tasks, and validates its effectiveness through experiments.

ImPartial: Multi-channel Whole-Cell Segmentation Using Partial Annotations

Shrivastava, Gunjan (Memorial Sloan Kettering Cancer Center), Nadeem, Saad (Memorial Sloan Kettering Cancer Center)

SegmentationConvolutional Neural NetworkDiffusion modelAuto EncoderContrastive LearningImageBiomedical DataBenchmark

🎯 What it does: Propose a weakly supervised cell instance segmentation framework called ImPartial, which can achieve cell segmentation on multi-channel images with only a small number of scribble annotations.

Improved Vascular Segmentation via Binary Flow Matching with Straight-Path Regularization

Kim, Yongjun (Korea Institute of Science and Technology), Ryu, Kanghyun (Korea University)

SegmentationConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelFlow-based ModelRectified FlowImageBiomedical DataComputed Tomography

🎯 What it does: This paper proposes a binary flow matching (BFM) method for iterative segmentation of fine vessels in X-ray coronary images.

Improving Factuality of 3D Brain MRI Report Generation with Paired Image-domain Retrieval and Text-domain Augmentation

Lee, Junhyeok (Seoul National University), Ye, Jong Chul (KAIST)

GenerationData SynthesisRetrievalTransformerLarge Language ModelPrompt EngineeringAuto EncoderImageTextBiomedical DataMagnetic Resonance ImagingRetrieval-Augmented Generation

🎯 What it does: Proposed a framework named PIRTA, which generates radiology reports with higher factual accuracy by utilizing image-domain retrieval and text-domain enhancement from 3D brain MRI.

Inception-Powered Unrolling Reconstruction for Accelerated Dynamic Cardiac Late Gadolinium Enhancement MRI

Zhong, Ya (Chongqing University of Technology), Zhu, Yanjie (Paul C. Lauterbur Research Center for Biomedical Imaging, Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences)

RestorationConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningBiomedical DataMagnetic Resonance Imaging

🎯 What it does: To address the motion artifact problem in late gadolinium enhancement (LGE) cardiac imaging, this paper proposes to reconstruct it as a dynamic multi-frame accelerated reconstruction task, achieving motion robustness improvement without additional scanning time by segmenting the acquired k-space into multiple time frames.

Induce to Empower: Improving Lightweight Baselines via Foundation Model Induction for Generalized Polyp Segmentation

Agnihotri, Shivanshu, Jha, Debesh (Khalifa University)

SegmentationKnowledge DistillationConvolutional Neural NetworkTransformerContrastive LearningImageBiomedical Data

🎯 What it does: This paper proposes a framework called Liteπ, which significantly improves segmentation performance by injecting structural and semantic priors from foundation models (SAM, DINOv2, OneFormer) into a lightweight multi-adenoma segmentation baseline.

Inflated Performance in fMRI-Based ASD Diagnosis: A Reproducibility Study of Data Leakage in Feature Selection Pipelines

Disoki, Mohamad Bashar, Sonbol, Riad

ClassificationAnomaly DetectionOptimizationComputational EfficiencyData-Centric LearningConvolutional Neural NetworkRecurrent Neural NetworkTransformerSupervised Fine-TuningContrastive LearningBiomedical DataMagnetic Resonance Imaging

🎯 What it does: The paper explores a new method to improve the efficiency and accuracy of data processing.

Information Maximization for Long-Tailed Semi-Supervised Domain Generalization

Fillioux, Leo (Université Paris-Saclay), Dolz, Jose (ETS Montréal)

ClassificationDomain AdaptationContrastive LearningImageBiomedical Data

🎯 What it does: Proposes an objective function called IMaX based on information maximization (InfoMax) to address the long-tailed class distribution problem in semi-supervised domain generalization (SSDG).

Injecting a Low-Rank Background Prior: A Plugin for Medical Image Segmentation

Chen, Minxin (Hong Kong Polytechnic University), Cheung, James Chung-Wai (Hong Kong Polytechnic University)

SegmentationConvolutional Neural NetworkAuto EncoderContrastive LearningImageBiomedical DataComputed TomographyUltrasound

🎯 What it does: This paper proposes a low-rank background prior plugin that learns the background subspace in the bottleneck feature space, extracts reconstruction residuals, and injects residual energy into the segmentation network to improve lesion segmentation performance.

InSPIRE: Multiparameter Inversion for Ultrasound Computed Tomography via Sequential Position-Based Implicit Representation

Zeng, Xiaolu (Huazhong University of Science and Technology), Yuchi, Ming (Huazhong University of Science and Technology)

OptimizationConvolutional Neural NetworkAuto EncoderBiomedical DataComputed TomographyUltrasound

🎯 What it does: Propose the InSPIRE framework to achieve unsupervised multi-parameter (sound speed and acoustic impedance) USCT inversion

INST-Align: Implicit Neural Alignment for Spatial Transcriptomics via Canonical Expression Fields

Han, Bonian (New Jersey Institute of Technology), Wei, Zhi (New Jersey Institute of Technology)

Representation LearningData-Centric LearningNeural Radiance FieldAuto EncoderGenerative Adversarial NetworkContrastive LearningOptical FlowBiomedical Data

🎯 What it does: In the multi-slice alignment and fusion of spatial transcriptomics, the INST-Align framework is proposed, which achieves unsupervised joint alignment and reconstruction using a shared implicit neural expression field and deformation network.

Interaction-Aware Contrastive Multi-task Learning for Joint Segmentation and Malignancy Classification in Thyroid Ultrasound

Lin, Chaochao (Beijing Institute of Technology), Werghi, Naoufel (Khalifa University)

ClassificationSegmentationConvolutional Neural NetworkContrastive LearningImageBiomedical DataUltrasound

🎯 What it does: Proposed an IC-MTL framework for joint segmentation of nodules and classification of malignancy in thyroid ultrasound images.