International Conference on Medical Image Computing and Computer-Assisted Intervention Β· 649 papers
Towards Unified Surgical Scene Understanding: Bridging Reasoning and Grounding via MLLMs
Huang, Jincai (Southern University of Science and Technology), Si, Weixin (Nanfang Hospital)
CodeRecognitionSegmentationTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageVideoTextBenchmark
π― What it does: Propose the SurgMLLM framework to achieve unified processing of procedural phase recognition, IVT triplet reasoning, and pixel-level entity localization in surgical videos.
Tractography-Driven Synthetic Data Generation for Fiber Bundle Segmentation in Tracer Histology
Bintsi, Kyriaki-Margarita (Massachusetts General Hospital and Harvard Medical School), Yendiki, Anastasia (Massachusetts General Hospital and Harvard Medical School)
π― What it does: Using ex vivo dMRI tractography as a generative prior, synthetic 2D image-mask pairs were generated through domain randomization, and combined with a small number of real annotated fiber bundle images to train a 2D U-Net, achieving automated segmentation of fiber bundles in macaque tracing-stained images.
CodeSuper ResolutionOptimizationComputational EfficiencyConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningBiomedical Data
π― What it does: Propose a cross-paradigm transformer named TACPT-Net for efficiently and accurately calibrating and upscaling the system matrix of magnetic particle imaging (MPI).
π― What it does: This paper proposes a three-layer knowledge anchoring (Tri-KA) framework for test-time adaptation (TTA) in cross-site rectal cancer MRI segmentation, without requiring source data or privacy leakage.
CodeClassificationImage TranslationAnomaly DetectionComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasoundElectronic Health Records
π― What it does: This paper proposes the TRIAGE-MIL framework, which combines unsupervised multi-axis hierarchical sampling (MASS) and a semantic hierarchical hypergraph model to achieve multi-instance learning (MIL) on whole slide images (WSI) for survival prediction.
π― What it does: Propose the TrustSyn framework, achieving hybrid-domain semi-supervised retinal image segmentation through teacher-student collaboration;
π― What it does: Propose the UltraStar framework, transforming ultrasound probe navigation from path regression to global localization based on star maps.
π― What it does: A medical image segmentation model based on SAM2, which proposes the BayesPrompt framework to achieve cross-modal adaptive segmentation under limited target modality annotations.
π― What it does: Propose a semi-supervised medical image segmentation framework (UAHC), which enhances segmentation performance under limited annotation by combining uncertainty guidance with hypergraph consistency learning.
π― What it does: Propose a training-agnostic multi-modal 3D medical image segmentation framework that generates dense prompts by fusing spatial and semantic context to achieve cross-modal, low-sample segmentation.
π― What it does: Propose an uncertainty-guided conservative propagation framework (UGCP), treating coronary artery segmentation inference as a state evolution process with limited steps, exchanging neighborhood information in the logit space to enhance structural consistency.
Understanding Model Behavior in Monocular Polyp Sizing
Xiong, Xinqi (University of North Carolina at Chapel Hill), Sengupta, Roni (University of California San Diego)
CodeClassificationSegmentationDepth EstimationExplainability and InterpretabilityConvolutional Neural NetworkTransformerPrompt EngineeringContrastive LearningImageBiomedical DataComputed TomographyReview/Survey Paper
π― What it does: Conducting a diagnostic audit on the classification of polyp size from monocular colonoscopy, systematically evaluating the performance of multiple models, input modalities, and cross-center datasets
π― What it does: This paper proposes Uni-Brain, a unified diffusion framework that generates multi-tracer PET images and performs causal prediction based on MRI-derived anatomical conditions and clinical conditions.
π― What it does: Propose UniField, a unified MRI field strength enhancement framework that integrates multi-modal (T1, T2, FLAIR) and cross-field enhancement tasks (64mTβ3T, 3Tβ7T), and leverages a pre-trained 3D video super-resolution prior to achieve structural prior knowledge;
π― What it does: This paper proposes UniLiver, a unified multi-task liver segmentation model that can simultaneously segment liver vasculature, Couinaud segments, and tumors.
Unleashing Video Language Models for Fine-grained HRCT Report Generation
Fang, Yingying (Imperial College London), Yang, Guang (Imperial College London)
CodeGenerationAnomaly DetectionOptimizationRepresentation LearningData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelImageTextBiomedical DataComputed TomographyChain-of-Thought
π― What it does: This paper studies the migration of general video-language models to high-resolution CT (HRCT) report generation, and proposes the AbSteering framework, enabling the model to first identify abnormalities before generating reports;
Unlocking the Power of Medical Tabular Data via Semantic-Aware Multimodal Pre-training
Liu, Yingsheng (Monash University), Yu, Zhen (University of Queensland)
CodeClassificationAnomaly DetectionRepresentation LearningData-Centric LearningConvolutional Neural NetworkTransformerPrompt EngineeringMixture of ExpertsVision Language ModelAuto EncoderContrastive LearningImageMultimodalityTabularBiomedical DataElectronic Health Records
π― What it does: Proposed a semantic-aware multimodal pre-training framework named AID, specifically designed to model the two-dimensional hierarchical structure of medical tabular data, significantly enhancing the robustness and generalization ability of image-table joint representations.
π― What it does: This paper models unsupervised brain anomaly detection as a Bayesian inverse problem with a diffusion prior, and introduces a latent anomaly mask to jointly infer pseudo-healthy images and anomalous regions.
π― What it does: This paper proposes a 3D pelvic landmark localization framework based on unsupervised domain adaptation, which utilizes landmark-conditioned image synthesis and pseudo-label sets to enhance cross-modal performance between CT and MR.
Unveiling Brain-Body Axis Interactions in Psychiatric-Endocrine Disorder with Deep Graph Causal Neural Network
Wan, Zhonghua (Nanjing University of Science and Technology), Wu, Ye (Nanjing University of Science and Technology)
CodeExplainability and InterpretabilityGraph Neural NetworkAuto EncoderContrastive LearningGraphTabularBiomedical DataMagnetic Resonance ImagingDiffusion Tensor ImagingAlzheimer's DiseaseElectronic Health Records
π― What it does: Built a causal graph framework based on graph neural networks to explore the brain-body axis interactions between psychiatric and endocrine diseases.
π― What it does: Proposed the URLCF framework, which leverages multi-source domain teacher model distillation and mask generation alignment to learn general representations, and then performs few-shot fine-tuning through a lightweight module to achieve cross-domain autism spectrum disorder detection.
UT-MIL: Uncertainty-Rectified and TME-Decoupling Dual-Stream Aggregation for Robust WSI Survival Prediction
Hu, Taiyuan (Chinese Academy of Sciences), Jiang, Jinrong (University of Chinese Academy of Sciences)
CodeExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerContrastive LearningBiomedical DataBenchmark
π― What it does: Proposes the UT-MIL framework, combining uncertainty calibration with a dual-stream aggregation for tumor microenvironment (TME) decoupling, for robust and interpretable survival prediction on whole-slide images (WSI).
π― What it does: Propose Vision-Action Adapter (VA-Adapter), which utilizes a cardiac ultrasound base model and incorporates action information to achieve real-time guidance for the cardiac ultrasound probe;
π― What it does: This study proposes the VecHeart framework, which can uniformly reconstruct and generate four-chamber heart structures and supports the recovery of complete heart geometry from sparse or missing data;
π― What it does: Propose a generation enhancement method based on a visual pre-trained model (VEGA), which converts one-dimensional electrocardiogram (ECG) signals into two-dimensional images. It utilizes a visual MAE for self-supervised reconstruction, thereby generating diverse ECG samples to improve cross-subject generalization performance.
CodeAnomaly DetectionExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataBenchmark
π― What it does: Developed a no-training, vision-based intervention medical visual question answering hallucination detection framework called VIHD.
π― What it does: Constructed HistoBIT3D, the first voxel-level paired 3D Back-illumination Interference Tomography (BIT) and fluorescent nucleus image dataset, and proposed a virtual H&E staining framework based on Vision Transformer CycleGAN.
π― What it does: Proposed a visualization-aware, unlabeled point, geometry self-supervised 3D-2D liver registration framework that utilizes visible region constraints to deformations, achieving real-time AR guidance;
π― What it does: Propose a deep learning framework based on monocular surgical microscope videos, which directly estimates tool-brain tissue contact force by utilizing brain surface deformation.
π― What it does: Proposed a Spatio-Temporal Transformer model based on visual queries for detecting biometric planes (HC, AC, FL) in blind-sweep prenatal ultrasound videos, achieving automatic assessment of scan quality and rapid resampling;
VOGeo-Gaze: Calibration-Free, Geometry-Aware Deep Learning for Real-Time Gaze Tracking in Clinical Video-Oculography
Zhao, Jingkang (LMU University Hospital), Wuehr, Max (LMU University Hospital)
CodePose EstimationOptimizationComputational EfficiencyConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningOptical FlowImageVideoBiomedical DataElectrocardiogramReview/Survey Paper
π― What it does: Propose a calibration-free, geometry-aware deep learning framework VOGeo-Gaze for real-time eye tracking, which can recover ocular parameters and infer gaze direction from geometric information obtained through image segmentation.
π― What it does: Proposed and implemented VolDiT β a fully Transformer-based 3D diffusion model for unsupervised generation and conditional generation of volumetric medical images.
π― What it does: Propose a data protection method for 3D medical image segmentation called VoxShield, which generates unlearnable examples (UE) to prevent unauthorized model training;
π― What it does: Propose a framework called WaCT-US, which is based on closed curves and provides infinitely calibrable uncertainty. It can directly generate a set of probabilistic predictions surrounding the prediction boundary (contour tube) after a single forward pass, and eliminates index ambiguity by utilizing the cyclic alignment of closed curves.
π― What it does: Proposes an untrained framework called WALDO for zero-shot anomaly localization using vision-language models (VLMs), primarily by identifying lesions through comparison with healthy reference images.
π― What it does: Propose a conditional flow matching framework called WaveDiT based on the Haar 3D discrete wavelet transform for high-resolution 3D brain MRI synthesis, and implement distribution-aware uncertainty scheduling through the Morpheus module, combined with hierarchical spatiotemporal attention for efficient sampling.
π― What it does: Proposed a wavelet decomposition-based waveguide spectral-spatial learning framework for pixel-level segmentation of medical hyperspectral images;
π― What it does: This paper proposes a sensor-free freehand 3D PAUS imaging framework called wBCAM, which achieves high-resolution 3D vascular reconstruction by estimating scanning motion through windowed bidirectional cross-attention and Mamba state space modules.
π― What it does: Proposed the MACOS framework, which utilizes long-term digital subtraction angiography (DSA) sequences to achieve weakly supervised segmentation of coronary arteries using only keyframe annotations.
π― What it does: Designed and implemented a multi-view mammography image classification framework based on evidence reasoning, CEI-Net, which can fuse diagnostic information from two views in an uncertainty-driven manner when views are inconsistent.
When, Where, and How: Adaptive Binning for Tabular Self-supervised Learning
Kim, Daehwan (Hanyang University), Jang, Ikbeom (Hankuk University of Foreign Studies)
CodeRepresentation LearningData-Centric LearningAuto EncoderContrastive LearningTabularBiomedical DataElectronic Health Records
π― What it does: Propose a self-supervised learning framework on medical tabular data, utilizing an adaptive binning mechanism to enhance representation learning.
CodeSegmentationRobotic IntelligenceTransformerAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataReview/Survey Paper
π― What it does: Propose a multi-view fusion framework guided by the pose of the wrist camera, aimed at reconstructing and segmenting suture thread occlusions in robot-assisted microsurgical anastomosis.
CodeClassificationExplainability and InterpretabilityComputational EfficiencyTransformerSupervised Fine-TuningContrastive LearningImageBiomedical Data
π― What it does: Proposes a model editing framework called X-Edit based on zero-space projection, used to correct misjudgments in medical vision Transformers without causing catastrophic forgetting.
X2Bone: Reconstructing 3D Bone Structures from 2D Biplanar X-Rays via Cross-RWKV
Pan, Zhaohong (Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences), Liang, Xiaokun (Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences)