arXivSub Start free trial

MICCAI 2026 Papers with Code β€” Page 7

International Conference on Medical Image Computing and Computer-Assisted Intervention Β· 649 papers

Towards Unified Surgical Scene Understanding: Bridging Reasoning and Grounding via MLLMs

Huang, Jincai (Southern University of Science and Technology), Si, Weixin (Nanfang Hospital)

CodeRecognitionSegmentationTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageVideoTextBenchmark

🎯 What it does: Propose the SurgMLLM framework to achieve unified processing of procedural phase recognition, IVT triplet reasoning, and pixel-level entity localization in surgical videos.

Tracer-Adaptive Expert Learning for Generalizable PET Lesion Segmentation

Shu, Yiran (ShanghaiTech University), Shen, Dinggang (ShanghaiTech University)

CodeSegmentationConvolutional Neural NetworkTransformerMixture of ExpertsAuto EncoderContrastive LearningBiomedical DataComputed TomographyPositron Emission Tomography

🎯 What it does: Propose a prototype-driven Mixture-of-Experts framework called ProMoE for lesion segmentation in multi-tracer PET images.

Tractography-Driven Synthetic Data Generation for Fiber Bundle Segmentation in Tracer Histology

Bintsi, Kyriaki-Margarita (Massachusetts General Hospital and Harvard Medical School), Yendiki, Anastasia (Massachusetts General Hospital and Harvard Medical School)

CodeSegmentationData SynthesisConvolutional Neural NetworkDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataMagnetic Resonance ImagingDiffusion Tensor Imaging

🎯 What it does: Using ex vivo dMRI tractography as a generative prior, synthetic 2D image-mask pairs were generated through domain randomization, and combined with a small number of real annotated fiber bundle images to train a 2D U-Net, achieving automated segmentation of fiber bundles in macaque tracing-stained images.

Trajectory-Aware Cross-Paradigm Transformers for Efficient System Matrix Calibration in Magnetic Particle Imaging

Li, Jintao (Northwest University), Guo, Hongbo (Concordia University)

CodeSuper ResolutionOptimizationComputational EfficiencyConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningBiomedical Data

🎯 What it does: Propose a cross-paradigm transformer named TACPT-Net for efficiently and accurately calibrating and upscaling the system matrix of magnetic particle imaging (MPI).

Tri-KA: Tri-level knowledge anchoring test-time adaptation for source-free cross-site MRI rectal cancer segmentation

Bo, Wang (University of Science and Technology of China), Zhou, Shaohua Kevin (University of Science and Technology of China)

CodeSegmentationDomain AdaptationSafty and PrivacyConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: This paper proposes a three-layer knowledge anchoring (Tri-KA) framework for test-time adaptation (TTA) in cross-site rectal cancer MRI segmentation, without requiring source data or privacy leakage.

TRIAGE-MIL: Multi-axis Instance Selection and Semantic Hypergraph Modeling for Survival Prediction from Whole-Slide Images

Subramanian, Barathi (Stanford University), Shen, Jeanne (Stanford University)

CodeClassificationImage TranslationAnomaly DetectionComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasoundElectronic Health Records

🎯 What it does: This paper proposes the TRIAGE-MIL framework, which combines unsupervised multi-axis hierarchical sampling (MASS) and a semantic hierarchical hypergraph model to achieve multi-instance learning (MIL) on whole slide images (WSI) for survival prediction.

TrustSyn: Reliable and Divergent Synergy for Mixed-Domain Fundus Segmentation

Qiao, Baojun (Henan University), Norouzifard, Mohammad (Yoobee College of Creative Innovation)

CodeSegmentationDomain AdaptationKnowledge DistillationConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound

🎯 What it does: Propose the TrustSyn framework, achieving hybrid-domain semi-supervised retinal image segmentation through teacher-student collaboration;

TSegAgent: Zero-Shot Tooth Segmentation via Geometry-Aware Vision-Language Agents

Zhuang, Shaojie, Zhou, Yuanfeng (Shandong University)

CodeRecognitionSegmentationTransformerAgentic AIPrompt EngineeringVision Language ModelDiffusion modelPoint CloudMeshBiomedical DataComputed TomographyRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose TSegAgent, a method for zero-shot tooth segmentation and identification through a geometry-aware vision-language agent;

UltraStar: Semantic-Aware Star Graph Modeling for Echocardiography Navigation

Wang, Teng (Tsinghua University), Huang, Gao (Beijing Academy of Artificial Intelligence)

CodeRobotic IntelligenceGraph Neural NetworkTransformerContrastive LearningImageBiomedical DataUltrasound

🎯 What it does: Propose the UltraStar framework, transforming ultrasound probe navigation from path regression to global localization based on star maps.

Uncertainty-Aware Bayesian Prompt Adaptation for Robust Cross-Modality Medical Segmentation

Hong, Sakang (Jeonbuk National University), Lee, Kyungsu (Jeonbuk National University)

CodeSegmentationDomain AdaptationTransformerPrompt EngineeringContrastive LearningMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound

🎯 What it does: A medical image segmentation model based on SAM2, which proposes the BayesPrompt framework to achieve cross-modal adaptive segmentation under limited target modality annotations.

Uncertainty-Aware Hypergraph Consistency Learning for Semi-supervised Medical Image Segmentation

Xie, Kexin (Guangdong University of Technology), Luo, Guibo (Peking University)

CodeSegmentationGraph Neural NetworkAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound

🎯 What it does: Propose a semi-supervised medical image segmentation framework (UAHC), which enhances segmentation performance under limited annotation by combining uncertainty guidance with hypergraph consistency learning.

Uncertainty-Aware Spatio-Semantic Contextual Prompts for Multimodal Medical Segmentation

Chattopadhyay, Soumitri (University of California San Diego), Niethammer, Marc (University of California San Diego)

CodeSegmentationConvolutional Neural NetworkTransformerPrompt EngineeringContrastive LearningMultimodalityBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose a training-agnostic multi-modal 3D medical image segmentation framework that generates dense prompts by fusing spatial and semantic context to achieve cross-modal, low-sample segmentation.

Uncertainty-Guided Conservative Propagation for Coronary Artery Segmentation

Huang, Huan (Kennesaw State University), Zhao, Chen (Kennesaw State University)

CodeSegmentationConvolutional Neural NetworkTransformerDiffusion modelContrastive LearningBiomedical DataComputed Tomography

🎯 What it does: Propose an uncertainty-guided conservative propagation framework (UGCP), treating coronary artery segmentation inference as a state evolution process with limited steps, exchanging neighborhood information in the logit space to enhance structural consistency.

Understanding Model Behavior in Monocular Polyp Sizing

Xiong, Xinqi (University of North Carolina at Chapel Hill), Sengupta, Roni (University of California San Diego)

CodeClassificationSegmentationDepth EstimationExplainability and InterpretabilityConvolutional Neural NetworkTransformerPrompt EngineeringContrastive LearningImageBiomedical DataComputed TomographyReview/Survey Paper

🎯 What it does: Conducting a diagnostic audit on the classification of polyp size from monocular colonoscopy, systematically evaluating the performance of multiple models, input modalities, and cross-center datasets

Uni-Brain: Anatomy-Informed Diffusion for Multi-Tracer Synthesis and Counterfactual Prognosis from Cross-Sectional Data

Yu, Hongjie (ShanghaiTech University), Shen, Dinggang (ShanghaiTech University)

CodeGenerationData SynthesisTransformerDiffusion modelScore-based ModelAuto EncoderImageBiomedical DataMagnetic Resonance ImagingPositron Emission TomographyAlzheimer's Disease

🎯 What it does: This paper proposes Uni-Brain, a unified diffusion framework that generates multi-tracer PET images and performs causal prediction based on MRI-derived anatomical conditions and clinical conditions.

UniField: A Unified Field-Aware MRI Enhancement Framework

Lin, Yiyang (Chinese University of Hong Kong), Yuan, Yixuan (Chinese University of Hong Kong)

CodeRestorationSuper ResolutionTransformerDiffusion modelAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance ImagingStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Propose UniField, a unified MRI field strength enhancement framework that integrates multi-modal (T1, T2, FLAIR) and cross-field enhancement tasks (64mTβ†’3T, 3Tβ†’7T), and leverages a pre-trained 3D video super-resolution prior to achieve structural prior knowledge;

UniLiver: Gradient-Conditioned Unified Model for Multi-target Hepatic Segmentation

Qiu, Yue (Chinese University of Hong Kong), Fu, Chi-Wing (Chinese University of Hong Kong)

CodeSegmentationConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningBiomedical DataMagnetic Resonance ImagingComputed TomographyBenchmark

🎯 What it does: This paper proposes UniLiver, a unified multi-task liver segmentation model that can simultaneously segment liver vasculature, Couinaud segments, and tumors.

UniSpine-GS: An Efficient Physics-Aware Gaussian Framework for Cross-Modality Multi-view Spine Image Synthesis

Chen, Qiuhua (Wuhan University), Du, Bo (Wuhan University)

CodeGenerationData SynthesisDiffusion modelNeural Radiance FieldGaussian SplattingImageMultimodalityBiomedical DataComputed TomographyUltrasound

🎯 What it does: Propose a physics-aware 3D Gaussian model called UniSpine-GS for efficient synthesis of cross-modal multi-view spinal images.

Unleashing Video Language Models for Fine-grained HRCT Report Generation

Fang, Yingying (Imperial College London), Yang, Guang (Imperial College London)

CodeGenerationAnomaly DetectionOptimizationRepresentation LearningData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelImageTextBiomedical DataComputed TomographyChain-of-Thought

🎯 What it does: This paper studies the migration of general video-language models to high-resolution CT (HRCT) report generation, and proposes the AbSteering framework, enabling the model to first identify abnormalities before generating reports;

Unlocking the Power of Medical Tabular Data via Semantic-Aware Multimodal Pre-training

Liu, Yingsheng (Monash University), Yu, Zhen (University of Queensland)

CodeClassificationAnomaly DetectionRepresentation LearningData-Centric LearningConvolutional Neural NetworkTransformerPrompt EngineeringMixture of ExpertsVision Language ModelAuto EncoderContrastive LearningImageMultimodalityTabularBiomedical DataElectronic Health Records

🎯 What it does: Proposed a semantic-aware multimodal pre-training framework named AID, specifically designed to model the two-dimensional hierarchical structure of medical tabular data, significantly enhancing the robustness and generalization ability of image-table joint representations.

Unsupervised Brain Anomaly Detection as a Bayesian Inverse Problem with Diffusion Prior

Roy, Hugues (Sorbonne UniversitΓ©), Burgos, Ninon (Sorbonne UniversitΓ©)

CodeAnomaly DetectionTransformerDiffusion modelScore-based ModelAuto EncoderBiomedical DataMagnetic Resonance ImagingPositron Emission TomographyAlzheimer's DiseaseStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: This paper models unsupervised brain anomaly detection as a Bayesian inverse problem with a diffusion prior, and introduces a latent anomaly mask to jointly infer pseudo-healthy images and anomalous regions.

Unsupervised Domain Adaptation for Pelvic Landmark Localization in CT and MR Using Landmark-Conditioned Synthesis

Ye, Kangqing (Shanghai Jiao Tong University), Zheng, Guoyan (Shanghai Jiao Tong University)

CodePose EstimationDomain AdaptationDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: This paper proposes a 3D pelvic landmark localization framework based on unsupervised domain adaptation, which utilizes landmark-conditioned image synthesis and pseudo-label sets to enhance cross-modal performance between CT and MR.

Unveiling Brain-Body Axis Interactions in Psychiatric-Endocrine Disorder with Deep Graph Causal Neural Network

Wan, Zhonghua (Nanjing University of Science and Technology), Wu, Ye (Nanjing University of Science and Technology)

CodeExplainability and InterpretabilityGraph Neural NetworkAuto EncoderContrastive LearningGraphTabularBiomedical DataMagnetic Resonance ImagingDiffusion Tensor ImagingAlzheimer's DiseaseElectronic Health Records

🎯 What it does: Built a causal graph framework based on graph neural networks to explore the brain-body axis interactions between psychiatric and endocrine diseases.

URLCF: Universal Representation Learning for Cross-Domain Few-Shot Autism Spectrum Disorder Detection

Mao, Zhenlin (Nanjing University of Information Science and Technology), Wang, Mingliang (Nanjing University of Information Science and Technology)

CodeClassificationDomain AdaptationKnowledge DistillationRepresentation LearningMeta LearningGraph Neural NetworkTransformerSupervised Fine-TuningAuto EncoderGenerative Adversarial NetworkContrastive LearningBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Proposed the URLCF framework, which leverages multi-source domain teacher model distillation and mask generation alignment to learn general representations, and then performs few-shot fine-tuning through a lightweight module to achieve cross-domain autism spectrum disorder detection.

UT-MIL: Uncertainty-Rectified and TME-Decoupling Dual-Stream Aggregation for Robust WSI Survival Prediction

Hu, Taiyuan (Chinese Academy of Sciences), Jiang, Jinrong (University of Chinese Academy of Sciences)

CodeExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerContrastive LearningBiomedical DataBenchmark

🎯 What it does: Proposes the UT-MIL framework, combining uncertainty calibration with a dual-stream aggregation for tumor microenvironment (TME) decoupling, for robust and interpretable survival prediction on whole-slide images (WSI).

VA-Adapter: Adapting Ultrasound Foundation Model to Echocardiography Probe Guidance

Wang, Teng (Tsinghua University), Huang, Gao (Tsinghua University)

CodeImage TranslationDomain AdaptationComputational EfficiencyRecurrent Neural NetworkTransformerSupervised Fine-TuningVision-Language-Action ModelContrastive LearningImageBiomedical DataUltrasoundElectrocardiogram

🎯 What it does: Propose Vision-Action Adapter (VA-Adapter), which utilizes a cardiac ultrasound base model and incorporates action information to achieve real-time guidance for the cardiac ultrasound probe;

VecHeart: Holistic Four-Chamber Cardiac Anatomy Modeling via Hybrid VecSets

Chen, Yihong (EPFL), Fua, Pascal (EPFL)

CodeRestorationGenerationTransformerDiffusion modelAuto EncoderContrastive LearningPoint CloudMeshBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound

🎯 What it does: This study proposes the VecHeart framework, which can uniformly reconstruct and generate four-chamber heart structures and supports the recovery of complete heart geometry from sparse or missing data;

VEGA: Vision Enhanced Generative Augmentation for Cross-Subject Electrocardiogram Generalization

Shi, Tianyi (Beijing University Of Posts And Telecommunications), Zhao, Zhicheng (Beijing University Of Posts And Telecommunications)

CodeGenerationData SynthesisDomain AdaptationAnomaly DetectionTransformerSupervised Fine-TuningDiffusion modelAuto EncoderContrastive LearningImageTime SeriesElectrocardiogram

🎯 What it does: Propose a generation enhancement method based on a visual pre-trained model (VEGA), which converts one-dimensional electrocardiogram (ECG) signals into two-dimensional images. It utilizes a visual MAE for self-supervised reconstruction, thereby generating diverse ECG samples to improve cross-subject generalization performance.

VIHD: Visual Intervention-Based Hallucination Detection for Medical Visual Question Answering

Chen, Jiayi (Monash University), Cai, Jianfei (Monash University)

CodeAnomaly DetectionExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataBenchmark

🎯 What it does: Developed a no-training, vision-based intervention medical visual question answering hallucination detection framework called VIHD.

ViPSAM: Visual Prompting Medical Image Segmentation Using Segment Anything Model

Lee, San (Sungkyunkwan University), Kim, Boah (Sungkyunkwan University)

CodeSegmentationConvolutional Neural NetworkTransformerPrompt EngineeringContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Improving liver lesion segmentation on low-contrast CT using visual prompts from contrast-enhanced MRI.

Virtual 3D H&E Staining from Phase-Contrast Back-Illumination Interference Tomography

Song, Anthony A. (Johns Hopkins University), Durr, Nicholas J. (Johns Hopkins Hospital)

CodeImage TranslationSegmentationGenerationData SynthesisTransformerDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataComputed Tomography

🎯 What it does: Constructed HistoBIT3D, the first voxel-level paired 3D Back-illumination Interference Tomography (BIT) and fluorescent nucleus image dataset, and proposed a virtual H&E staining framework based on Vision Transformer CycleGAN.

Vis2Reg: Visibility-Aware Landmark-Free Geometric 3D–2D Registration for Liver Laparoscopy

Feng, Jiaming (University of Leeds), Ali, Sharib (University of Leeds)

CodePose EstimationDepth EstimationGraph Neural NetworkTransformerDiffusion modelNeural Radiance FieldAuto EncoderContrastive LearningGaussian SplattingOptical FlowImagePoint CloudBiomedical DataComputed TomographyUltrasoundReview/Survey PaperBenchmark

🎯 What it does: Proposed a visualization-aware, unlabeled point, geometry self-supervised 3D-2D liver registration framework that utilizes visible region constraints to deformations, achieving real-time AR guidance;

Vision-Based Force Estimation Using Brain Surface Deformation Models

Heraoui, Chakib (Université TÉLUQ), Gueziri, Houssem-Eddine (Neurosurgical Simulation and Artificial Intelligence Learning Centre)

CodeData SynthesisConvolutional Neural NetworkDiffusion modelContrastive LearningVideoBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasoundFibre Orientation DistributionDiffusion Tensor Imaging

🎯 What it does: Propose a deep learning framework based on monocular surgical microscope videos, which directly estimates tool-brain tissue contact force by utilizing brain surface deformation.

Visual-Query Curriculum Learning for Rare Fetal Biometry Plane Detection in Blind-Sweep Ultrasound Video

Kwon, Jong (University of Oxford), Noble, J. Alison (University of Oxford)

CodeObject DetectionDomain AdaptationAnomaly DetectionConvolutional Neural NetworkTransformerContrastive LearningImageVideoBiomedical DataUltrasound

🎯 What it does: Proposed a Spatio-Temporal Transformer model based on visual queries for detecting biometric planes (HC, AC, FL) in blind-sweep prenatal ultrasound videos, achieving automatic assessment of scan quality and rapid resampling;

VL-RewardGen: Vision–Language Reward Driven Skin Lesion Image Generation

Chu, Yusung (Yonsei University), Yang, Sejung (Yonsei University)

CodeClassificationGenerationData SynthesisConvolutional Neural NetworkTransformerSupervised Fine-TuningReinforcement LearningPrompt EngineeringVision Language ModelDiffusion modelImageTextBiomedical Data

🎯 What it does: Proposed a control-based skin lesion image generation framework called VL-RewardGen, which is based on visual-language rewards

VOGeo-Gaze: Calibration-Free, Geometry-Aware Deep Learning for Real-Time Gaze Tracking in Clinical Video-Oculography

Zhao, Jingkang (LMU University Hospital), Wuehr, Max (LMU University Hospital)

CodePose EstimationOptimizationComputational EfficiencyConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningOptical FlowImageVideoBiomedical DataElectrocardiogramReview/Survey Paper

🎯 What it does: Propose a calibration-free, geometry-aware deep learning framework VOGeo-Gaze for real-time eye tracking, which can recover ocular parameters and infer gaze direction from geometric information obtained through image segmentation.

VolDiT: Controllable Volumetric Medical Image Synthesis with Diffusion Transformers

Seyfarth, Marvin (Heidelberg University), Engelhardt, Sandy (Heidelberg University)

CodeSegmentationGenerationData SynthesisTransformerDiffusion modelAuto EncoderBiomedical DataComputed Tomography

🎯 What it does: Proposed and implemented VolDiT β€” a fully Transformer-based 3D diffusion model for unsupervised generation and conditional generation of volumetric medical images.

VoxShield: Protecting 3D Medical Datasets from Unauthorized Training via Frequency-Aware Inter-Slice Disruption

Liu, Xinyao (Westlake University), Zheng, Yefeng (Westlake University)

CodeSegmentationSafty and PrivacyConvolutional Neural NetworkTransformerGenerative Adversarial NetworkContrastive LearningBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose a data protection method for 3D medical image segmentation called VoxShield, which generates unlearnable examples (UE) to prevent unauthorized model training;

WaCT-US: Cyclic-Alignment Scaled Contour Tubes for Efficient, Calibrated Uncertainty in Echocardiography

Kazemi Esfeh, Mohammad Mahdi (University of British Columbia), Abolmaesumi, Purang (University of British Columbia)

CodeSegmentationConvolutional Neural NetworkDiffusion modelScore-based ModelContrastive LearningImageBiomedical DataUltrasound

🎯 What it does: Propose a framework called WaCT-US, which is based on closed curves and provides infinitely calibrable uncertainty. It can directly generate a set of probabilistic predictions surrounding the prediction boundary (contour tube) after a single forward pass, and eliminates index ambiguity by utilizing the cyclic alignment of closed curves.

Wasserstein-Aligned Localisation for VLM-Based Distributional OOD Detection in Medical Imaging

Kainz, Bernhard (FAU Erlangen-NΓΌrnberg), Bercea, Cosmin (Technical University Munich)

CodeImage TranslationAnomaly DetectionTransformerPrompt EngineeringVision Language ModelContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyRetrieval-Augmented Generation

🎯 What it does: Proposes an untrained framework called WALDO for zero-shot anomaly localization using vision-language models (VLMs), primarily by identifying lesions through comparison with healthy reference images.

WaveDiT: Distribution-Aware Wavelet Flow Matching for Efficient 3D Brain MRI Synthesis

Danese, Danilo (Politecnico di Bari), Di Noia, Tommaso (Politecnico di Bari)

CodeGenerationData SynthesisTransformerDiffusion modelFlow-based ModelRectified FlowContrastive LearningBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose a conditional flow matching framework called WaveDiT based on the Haar 3D discrete wavelet transform for high-resolution 3D brain MRI synthesis, and implement distribution-aware uncertainty scheduling through the Morpheus module, combined with hierarchical spatiotemporal attention for efficient sampling.

Wavelet-Guided Spectral–Spatial Learning for Medical Hyperspectral Image Segmentation

Ji, Junya (University of Exeter), Ye, Xujiong (University of Exeter)

CodeSegmentationConvolutional Neural NetworkTransformerBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Proposed a wavelet decomposition-based waveguide spectral-spatial learning framework for pixel-level segmentation of medical hyperspectral images;

wBCAM: Windowed Bi-directional Cross-Attention with Mamba for Enhanced 3D Vascular Reconstruction in Free-Hand Photoacoustic and Ultrasound Imaging

Lee, SiYeoul (Pusan National University), Kim, MinWoo (Pusan National University)

CodeRestorationTransformerContrastive LearningOptical FlowVideoBiomedical DataPositron Emission TomographyUltrasound

🎯 What it does: This paper proposes a sensor-free freehand 3D PAUS imaging framework called wBCAM, which achieves high-resolution 3D vascular reconstruction by estimating scanning motion through windowed bidirectional cross-attention and Mamba state space modules.

Weakly-Supervised Coronary Artery Segmentation from DSA Sequence via Motion-Aware Modeling

Wu, Han (ShanghaiTech University), Shen, Dinggang (ShanghaiTech University)

CodeSegmentationConvolutional Neural NetworkRecurrent Neural NetworkTransformerContrastive LearningOptical FlowVideoBiomedical DataComputed TomographyBenchmark

🎯 What it does: Proposed the MACOS framework, which utilizes long-term digital subtraction angiography (DSA) sequences to achieve weakly supervised segmentation of coronary arteries using only keyframe annotations.

When Views Disagree: Conflict-Aware Evidential Inference for Multi-view Mammography

Ahmed, Md Rayhan (University of British Columbia), Lasserre, Patricia (University of British Columbia)

CodeClassificationAnomaly DetectionConvolutional Neural NetworkTransformerContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound

🎯 What it does: Designed and implemented a multi-view mammography image classification framework based on evidence reasoning, CEI-Net, which can fuse diagnostic information from two views in an uncertainty-driven manner when views are inconsistent.

When, Where, and How: Adaptive Binning for Tabular Self-supervised Learning

Kim, Daehwan (Hanyang University), Jang, Ikbeom (Hankuk University of Foreign Studies)

CodeRepresentation LearningData-Centric LearningAuto EncoderContrastive LearningTabularBiomedical DataElectronic Health Records

🎯 What it does: Propose a self-supervised learning framework on medical tabular data, utilizing an adaptive binning mechanism to enhance representation learning.

Wrist Camera Pose-Guided Multi-view Fusion for Occlusion Reconstruction in Robot-Assisted Microsurgical Anastomosis

Zhou, Xinyao (Shanghai Jiao Tong University), Guo, Yao (Shanghai Jiao Tong University)

CodeSegmentationRobotic IntelligenceTransformerAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataReview/Survey Paper

🎯 What it does: Propose a multi-view fusion framework guided by the pose of the wrist camera, aimed at reconstructing and segmenting suture thread occlusions in robot-assisted microsurgical anastomosis.

X-Edit: Exact, Explicit, and Explainable Null-Space Editing for Medical Vision Transformers

Liu, Yuanye (Fudan University), Zhuang, Xiahai (Johns Hopkins University)

CodeClassificationExplainability and InterpretabilityComputational EfficiencyTransformerSupervised Fine-TuningContrastive LearningImageBiomedical Data

🎯 What it does: Proposes a model editing framework called X-Edit based on zero-space projection, used to correct misjudgments in medical vision Transformers without causing catastrophic forgetting.

X2Bone: Reconstructing 3D Bone Structures from 2D Biplanar X-Rays via Cross-RWKV

Pan, Zhaohong (Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences), Liang, Xiaokun (Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences)

CodeImage TranslationRestorationGenerationPose EstimationDepth EstimationConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningImageBiomedical DataComputed Tomography

🎯 What it does: Developed an end-to-end X2Bone framework for reconstructing patient-specific 3D skeletal structures from dual-plane X-rays.