ICML 2026 Papers — Page 63
International Conference on Machine Learning · 6554 papers
Unifying Stacking and Cascading for Efficient Ensemble Inference
Ashwin Gerard Colaco (University of California Irvine), Unnat Jain (University of California Irvine)
ClassificationRecommendation SystemComputational EfficiencyReinforcement LearningMixture of ExpertsImageTextMultimodalityTabularAudio
🎯 What it does: Propose LazyStack, a method for efficient ensemble inference through progressively aggregating model predictions in a black-box environment
Unifying Value Alignment and Assignment in Cross-Domain Offline Reinforcement Learning with Heterogeneous Datasets
Zhongjian Qiao (City University of Hong Kong), Shuang Qiu (City University of Hong Kong)
Recurrent Neural NetworkReinforcement LearningContrastive LearningSequential
🎯 What it does: Studied heterogeneous offline reinforcement learning scenarios where the source domain is diverse and the behavior policies differ, and proposed the V2A framework to address the value mismatch problem.
UniJEPA: Enhancing Robot Policy via Unified Continuous and Discrete Representation Learning
Jianke Zhang (Tsinghua University), Jianyu Chen (Tsinghua University)
Representation LearningRobotic IntelligenceReinforcement Learning from Human FeedbackTransformerSupervised Fine-TuningMixture of ExpertsVision Language ModelVision-Language-Action ModelFlow-based ModelContrastive LearningVideoTextMultimodality
🎯 What it does: Propose UniJEPA, a VLA framework that unifies discrete language understanding with continuous visual prediction, achieving generalized learning of robot control strategies through two-stage pre-training and fine-tuning.
UniMapping: Unified SLAM Framework for Map-Centric Embodied Perception
Xiaze Zhang (Fudan University), Rui Feng (Fudan University)
Autonomous DrivingOptimizationComputational EfficiencyRepresentation LearningConvolutional Neural NetworkTransformerNeural Radiance FieldContrastive LearningSimultaneous Localization and MappingOptical FlowImageMultimodalityPoint Cloud
🎯 What it does: Propose UniMapping, a unified SLAM framework that integrates multi-modal observations (camera, LiDAR, RGB-D) into persistent neural descriptor maps, and directly supports downstream perception tasks (object detection, semantic segmentation).
UniMedVL: Unifying Medical Multimodal Understanding and Generation through Observation-Knowledge-Analysis
Junzhi Ning (Shanghai Artificial Intelligence Laboratory), Junjun He (Shanghai Artificial Intelligence Laboratory)
RecognitionImage TranslationRestorationGenerationData SynthesisTransformerVision Language ModelDiffusion modelScore-based ModelFlow-based ModelRectified FlowAuto EncoderContrastive LearningImageTextMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasoundBenchmark
🎯 What it does: Proposed a unified medical multimodal model called UniMedVL, which can simultaneously perform medical image understanding and image generation under the same set of parameters.
UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and Texture
Shuo Cao (University of Science and Technology of China), Yihao Liu (Shanghai Artificial Intelligence Laboratory)
Image TranslationGenerationTransformerSupervised Fine-TuningReinforcement LearningVision Language ModelContrastive LearningImageTextMultimodalityBenchmark
🎯 What it does: Proposed the UniPercept framework, constructing a unified perceptual-level image understanding benchmark (UniPercept-Bench), covering four areas: aesthetics, quality, structure, and texture. Based on this, the UniPercept model was trained (through domain-adaptive pre-training + task-aligned RL), and it was used as a reward model for text-to-image generation.
UniRRM: Unified Reasoning Reward Models Across Languages and Evaluation Paradigms
Peng Lai, Guanhua Chen (Southern University of Science and Technology)
Representation LearningData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposed a unified reasoning reward model, UniRRM, and constructed a large-scale preference dataset called MixReward covering 103 languages and 6 domains, used for unified multi-paradigm (pair-wise, list-wise, point-wise) evaluation.
UniRTL: Unifying Code and Graph for Robust RTL Representation Learning
Yi Liu (Chinese University of Hong Kong), Qiang Xu (Chinese University of Hong Kong)
Knowledge DistillationRepresentation LearningGraph Neural NetworkTransformerLarge Language ModelContrastive LearningTextMultimodalityGraph
🎯 What it does: Propose UniRTL, a multi-modal pre-training framework that unifies RTL code with complete control data flow graphs (CDFG) to obtain more robust RTL representations.
UniScale: Adaptive Unified Inference Scaling via Online Joint Optimization of Model Routing and Test-Time Scaling
Kaiyu Huang (Tongji University), Qingjiang Shi (Tongji University)
OptimizationComputational EfficiencyTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsText
🎯 What it does: When deploying large language models in the real world, this paper proposes a unified inference scaling (UIS) framework that integrates model routing with inference-time scaling (TTS) into a single decision space, and implements adaptive selection within this space using online contextual multi-armed bandit (LinUCB). The algorithm implementation includes semantic representation, cost modeling, path-aware early stopping, and dense validation feedback. Experiments are conducted on multi-difficulty mathematical reasoning datasets.
Unison: Benchmarking Unified Multimodal Models via Synergistic Understanding and Generation
Jinyu Liu (Fudan University), Yu-Gang Jiang (Fudan University)
GenerationData SynthesisTransformerLarge Language ModelPrompt EngineeringVision Language ModelDiffusion modelImageTextMultimodalityBenchmark
🎯 What it does: Design and release the Unison benchmark to evaluate the four-dimensional performance of unified multimodal models in collaborative capabilities of understanding and generation.
UniSparse: Combining Weight Pruning and Spike Sparsification in Spiking Neural Networks
Xinyu Shi (Peking University), Zhaofei Yu (Peking University)
ClassificationComputational EfficiencySpiking Neural NetworkAuto EncoderContrastive LearningImageVideo
🎯 What it does: Propose UniSparse, which unifies weight pruning and pulse sparsification techniques to improve the energy efficiency of SNNs.
UniSVQ: 2-bit Unified Scalar-Vector Quantization
Haoyu Wang (Tsinghua University), Maosong Sun (Tsinghua University)
Computational EfficiencyRepresentation LearningTransformerLarge Language ModelAuto EncoderText
🎯 What it does: Propose and implement UniSVQ, a unified 2-bit scalar-vector quantization framework that balances the efficiency of SQ and the accuracy of VQ.
Unitary Convolutions for Message-passing and Positional Encodings on Directed Graphs
Lukas Fesser (Harvard University), Melanie Weber (Harvard University)
ClassificationRepresentation LearningGraph Neural NetworkTransformerContrastive LearningGraph
🎯 What it does: Designed a unit matrix convolution network called Dune suitable for directed graphs and graphs with edge features, addressing the over-smoothing and gradient vanishing problems that traditional GNNs encounter in directed graphs.
Universal Algorithm-Implicit Learning
Stefano Woerner (University of Tübingen), Christian F. Baumgartner (University of Tübingen)
ClassificationDomain AdaptationComputational EfficiencyRepresentation LearningMeta LearningTransformerPrompt EngineeringContrastive LearningImageTextMultimodalityBiomedical DataRetrieval-Augmented GenerationAudio
🎯 What it does: A new meta-learning theoretical framework is proposed, and under this framework, a Transformer-based algorithm implicit meta-learner called TAIL is designed to achieve generalization across domains, modalities, and label spaces in few-shot scenarios.
Universal Approximation with Softmax Attention
Jerry Yao-Chieh Hu (Northwestern University), Han Liu (Northwestern University)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformer
🎯 What it does: Prove that using only soft max attention (without feed-forward networks) can approximate any continuous sequence-to-sequence function;
Universal Learning of Nonlinear Dynamics
Evan Dogariu (New York University), Elad Hazan (Princeton University)
OptimizationComputational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackContrastive LearningTime SeriesSequentialPhysics RelatedStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Propose a spectral filtering algorithm that does not require identifying the system model and directly uses the observed sequence to perform single-step prediction for edge-stable nonlinear dynamics.
Universal Multiclass Transductive Online Learning
Steve Hanneke (Purdue University), Hongao Wang (Purdue University)
ClassificationOptimizationFederated LearningMeta LearningReinforcement Learning from Human FeedbackSupervised Fine-TuningReinforcement LearningMixture of ExpertsContrastive LearningReview/Survey Paper
🎯 What it does: Studied the general learnability of multi-class transferable online learning, and presented three possible learning rates: constant, logarithmic, and linear; simultaneously proposed an algorithm with an upper bound of approximately √T in the agnostic scenario.
Universal One-third Time Scaling in Learning Peaked Distributions
Yizhou Liu (Massachusetts Institute of Technology), Jeff Gore (Massachusetts Institute of Technology)
OptimizationComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelText
🎯 What it does: Studied that during the learning of peak distribution, softmax and cross-entropy loss themselves lead to power-law decay in the training process, and proposed and verified a general 1/3 time scaling law.
Universal Reasoner: A Single, Composable Plug-and-Play Reasoner for Frozen LLMs
Jaemin Kim (Korea Advanced Institute of Science and Technology), Jong Chul Ye (Korea Advanced Institute of Science and Technology)
Computational EfficiencyKnowledge DistillationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextMultimodalityRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposes Universal Reasoner (UniR), a lightweight and composable reasoning module that can be added to the logits of a frozen large language model (LLM);
Universal Redundancies in Time Series Foundation Models
Anthony Bao (UT Austin), William Gilpin (UT Austin)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerSupervised Fine-TuningPrompt EngineeringAuto EncoderContrastive LearningTabularTime SeriesBenchmark
🎯 What it does: Investigated and demonstrated that multiple time series foundation models (TSFM) exhibit widespread redundancy in their intermediate layers, revealing redundant components through large-scale ablation experiments, residual flow analysis, and direct logit attribution; simultaneously proposed a theoretical framework based on kernel regression to explain phenomena such as context parrots and seasonal bias, and developed visualization and ablation toolkits.
Universal Representation of Generalized Convex Functions and their Gradients
Moeen Nehzati (New York University)
OptimizationRepresentation LearningDiffusion modelScore-based ModelFlow-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningTabularBenchmark
🎯 What it does: A differentiable, generalized convex function (GCF) with a convex parameter space and a parameterization method for its gradient is proposed. It is proven that under moderate regularity conditions, this method can densely approximate all GCFs and their gradients. Subsequently, this method is applied to optimal transport and multi-item mechanism design problems, transforming the original bilevel or minimax problems into single-level problems solvable via first-order optimization.
Universal Skeleton Understanding via Differentiable Rendering and MLLMs
Ziyi Wang (Peking University), Mengyuan Liu (Peking University)
Data SynthesisPose EstimationKnowledge DistillationRepresentation LearningTransformerLarge Language ModelVision Language ModelContrastive LearningGaussian SplattingImageVideoSequential
🎯 What it does: Propose SkeletonLLM, which converts any skeleton sequence into visual images directly processable by MLLM through a differentiable renderer called DrAction, achieving cross-format skeleton understanding;
Unlearning in Diffusion Models: A Unified Framework with KL Divergence and Likelihood Constraints
Shervin Khalafi (University of Pennsylvania), Dongsheng Ding (University of Tennessee)
GenerationData SynthesisOptimizationSafty and PrivacyDiffusion modelScore-based ModelImageText
🎯 What it does: A unified constrained optimization framework is proposed for diffusion models, enabling the models to retain their original practicality while forgetting specified data or concepts.
Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
Xiaoyu Xu (Hong Kong Polytechnic University), Haibo Hu (Hong Kong Polytechnic University)
Federated LearningSafty and PrivacyExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringContrastive LearningText
🎯 What it does: This paper studies the unlearning process of large language models (LLMs) and explores whether the information is truly deleted, rather than merely suppressed;
Unlearning Isn’t Forgetting: Revealing Hidden Leakage in Class Unlearning Evaluations
Ali Ebrahimpour-Boroojeny (University of Illinois at Urbana-Champaign), Hari Sundaram (University of Illinois at Urbana-Champaign)
ClassificationSafty and PrivacyAdversarial AttackConvolutional Neural NetworkContrastive LearningImage
🎯 What it does: Proposed a class-level forgetting method based on tilted reweighting (TREW) and designed a new class member inference attack (CMIA) to evaluate the forgetting effect.
Unlearning with Asymmetric Sources: Improved Unlearning-Utility Trade-off with Public Data
Ahmed Mehdi Inane (Universite de Montreal), Ioannis Mitliagkas (Google DeepMind)
Safty and PrivacyImageTextStochastic Differential Equation
🎯 What it does: Proposes the Asymmetric Langevin Unlearning (ALU) framework, which utilizes public data to reduce noise requirements, thereby improving the privacy-utility balance of machine learning models when meeting data deletion requests.
Unlearning’s Blind Spots: Over‑Unlearning and Prototypical Relearning Attack
SeungBum Ha (Ulsan National Institute of Science and Technology), Sung Whan Yoon (Ulsan National Institute of Science and Technology)
ClassificationSafty and PrivacyKnowledge DistillationAdversarial AttackConvolutional Neural NetworkTransformerPrompt EngineeringDiffusion modelContrastive LearningImageTextMultimodality
🎯 What it does: Propose and evaluate two blind spots in machine unlearning—over-unlearning and prototype relearning attacks—and present a general framework called Spotter that can simultaneously suppress both blind spots.
Unleashing Implicit Rewards: Prefix-Value Learning for Distribution-Level Optimization
Shiping Gao (Sun Yat-sen University), Lifu Huang (University of California Davis)
OptimizationTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextSequential
🎯 What it does: Propose the Implicit Prefix-Value Reward Model (IPVRM) and combine it with Distribution-Level RL (DistRL) to improve the quality of process rewards based on terminal correctness labels, and achieve more efficient sample utilization in RL.
Unleashing the Representational Power of Fourier Shapes for Attacking Infrared Object Detection
Yixing Yong (Xi'an Jiaotong University), Fan Li (Xi'an Jiaotong University)
Object DetectionAdversarial AttackDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImage
🎯 What it does: In infrared target detection, the authors design and optimize a learnable Fourier shape attack, generating shapes that can be physically cut into thermal insulation materials, thereby making the target ignored by the detector in infrared images.
Unlocking Cross-Modal Biosignal Synthesis: A Temporally-Aware VAE-Diffusion Model
Chenyang Xu (Xidian University), Hao Wang (Xidian University)
GenerationData SynthesisTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningTime SeriesBiomedical DataElectrocardiogramAudio
🎯 What it does: Studied a cross-modal generation architecture combining VAE and diffusion models to synthesize corresponding PCG (heart sound) waveforms from common ECG signals.
Unlocking Noise-Resistant Vision: Key Architectural Secrets for Robust Models Against Gaussian Noise
Bum Jun Kim (University of Tokyo), Yutaka Matsuo (University of Tokyo)
ClassificationImage TranslationRestorationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningConvolutional Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningDiffusion modelScore-based ModelAuto EncoderContrastive LearningGaussian SplattingImage
🎯 What it does: This paper systematically evaluates and analyzes 1,174 pre-trained visual models and conducts control experiments on ResNet, revealing which architectural choices and input pipelines can enhance model robustness to additive Gaussian noise.
Unlocking Speech–Text Compositional Powers: Instruction-Following Speech Language Models without Instruction Tuning
Congrui Du (University of California Santa Barbara), Shiyu Chang (University of California Santa Barbara)
RecognitionGenerationExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsContrastive LearningTextMultimodalityRetrieval-Augmented GenerationAudio
🎯 What it does: Proposed SPEECHCOMBINE, a speech-language model that performs text and voice instructions without instruction fine-tuning, achieved through weighted merging.
Unlocking the Potential of Continual Model Merging: An ODE Perspective
Lihong Lin (Northeastern University), Haidong Kang (Northeastern University)
OptimizationComputational EfficiencyRepresentation LearningTransformerContrastive LearningImageMultimodalityBenchmarkStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Proposes an ODE-driven continuous model merging framework, ODE-M, treating each model merge as a continuous trajectory in the parameter space rather than a one-time update.
Unlocking Zero-Shot Geospatial Reasoning via Indirect Rewards
Chenhui Xu (University at Buffalo), Jinjun Xiong (University at Buffalo)
Representation LearningData-Centric LearningTransformerLarge Language ModelReinforcement LearningVision Language ModelImageTextMultimodalityBenchmarkChain-of-Thought
🎯 What it does: In the absence of scarce task annotations, Geo-R1 is proposed to train visual-language models for zero-shot geospatial reasoning by using indirectly verifiable rewards from cross-perspective (ground view and aerial view) pairings.
UnMaskFork: Test-Time Scaling for Masked Diffusion via Deterministic Action Branching
Kou Misaki (Sakana AI), Takuya Akiba (Sakana AI)
GenerationComputational EfficiencyAI Code AssistantTransformerReinforcement LearningPrompt EngineeringMixture of ExpertsDiffusion modelTextSequential
🎯 What it does: Propose a test-time scaling framework called UnMaskFork (UMF) for masked diffusion language models (MDLMs), which leverages Monte Carlo Tree Search (MCTS) to explore different decoding paths during inference.
Unpaired Visual Editing with Self-Consistent Flow Matching
Yoad Tewel (NVIDIA), Lior Wolf (Tel Aviv University)
Image TranslationRestorationGenerationDomain AdaptationTransformerPrompt EngineeringDiffusion modelScore-based ModelFlow-based ModelContrastive LearningOptical FlowImageVideoText
🎯 What it does: Train a flow matching editing model on unpaired data using instruction following information from a frozen model and cyclic consistency, achieving image and video editing.
Unraveling Syntax: Language Modeling and the Substructure of Grammars
Laura Ying Schulz (ETH Zürich), Tomaso Poggio (Massachusetts Institute of Technology)
Computational EfficiencyRepresentation LearningTransformerLarge Language ModelText
🎯 What it does: Study the learning dynamics of neural language models on substructures of context-free grammars (CFG), and prove that the loss function can be recursively decomposed into contributions from sub-grammars.
Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents
Xu Li (Northeastern University), Weiyan Shi (Northeastern University)
Safty and PrivacyTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: The paper constructs the MT-AgentRisk benchmark and proposes a training-agnostic, self-inspection defensive mechanism called ToolShield, aimed at evaluating and enhancing the safety of multi-turn tool-using agents.
Unsat Core Prediction through Polarity-Aware Representation Learning over Clause-Literal Hypergraphs
Zhenchao Sun (Beihang University), Chongyang Tao (Beihang University)
OptimizationRepresentation LearningGraph Neural NetworkTransformerSupervised Fine-TuningContrastive LearningGraph
🎯 What it does: This paper proposes a polarity-aware hypergraph representation learning framework for SAT formulas, aimed at predicting unsatisfiable core variables.
Unsupervised Camouflaged Object Detection with Dual-Eigenvector Spectral Pseudo-Labeling and Contrastive Refinement
Pingzhu Liu (Tsinghua University), Xiu Li (Tsinghua University)
Object DetectionSegmentationConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningImage
🎯 What it does: Propose an unsupervised hidden object detection framework called DualUCOD, which achieves high-quality segmentation through spectral pseudo-label generation and contrastive learning.
Unsupervised Diffusion Solver for Combinatorial Optimization via Combinatorial Adjoint Matching
Shengyu Feng (Carnegie Mellon University), Yiming Yang (Carnegie Mellon University)
OptimizationReinforcement LearningDiffusion modelScore-based ModelGraphTabular
🎯 What it does: Proposed an unsupervised discrete diffusion solver called Combinatorial Adjoint Matching (CAM), which utilizes discrete adjoint dynamics to propagate gradients along combinatorial optimization trajectories.
Unsupervised Disentanglement Without Compromises : How Functional Orthogonality Enforces Identifiability
Mathieu Cyrille Simon (UCLouvain), Christophe De Vleeschouwer (UCLouvain)
Representation LearningFlow-based ModelContrastive Learning
🎯 What it does: Proposed and verified an unsupervised decomposable representation learning framework based on functional orthogonality (Jacobian orthogonality), demonstrating that model identifiability can be achieved without requiring statistical independence or causal assumptions.
Unsupervised Hierarchical Skill Discovery
Damion Harvey (University of Witwatersrand), Steven James (University of Witwatersrand)
TransformerReinforcement LearningAuto EncoderGenerative Adversarial NetworkContrastive LearningWorld ModelOptical FlowVideoSequential
🎯 What it does: Propose a fully unsupervised framework HiSD, combining temporal action segmentation with syntax compression, which can automatically extract reusable multi-level skill hierarchies from unlabelled observed trajectories.
Unsupervised Neural Langevin Sampler for Mixed Integer Linear Programming
Yixin Huang (Carnegie Mellon University), Yiming Yang (Carnegie Mellon University)
OptimizationGraph Neural NetworkScore-based ModelContrastive LearningTabularBenchmarkStochastic Differential Equation
🎯 What it does: Proposed an unsupervised neural Langevin sampler for solving mixed integer linear programming problems.
Unsupervised Partner Design Enables Robust Ad-hoc Teamwork
Constantin Ruhdorfer (University of Stuttgart), Andreas Bulling (University of Stuttgart)
Reinforcement LearningGenerative Adversarial NetworkContrastive LearningTabularTime SeriesSequential
🎯 What it does: Propose Unsupervised Partner Design (UPD), an adaptive training framework that does not require explicit partner groups, instead generating partners online and selecting them based on learnability criteria;
Unsupervised Process-Aware Coreset Selection for In-Context Learning
Wei Zheng (Shandong University), YUQING SUN
OptimizationRepresentation LearningData-Centric LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextRetrieval-Augmented Generation
🎯 What it does: Proposes an unsupervised process-aware core set selection method (PaCS-ICL) for generating high-quality few-shot context demonstrations for large language models under extremely low labeling budgets.
Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking
Ravi Ghadia (Together AI), Max Ryabinin (Together AI)
Computational EfficiencyTransformerLarge Language ModelText
🎯 What it does: Propose UPipe, a context parallel technique that executes self-attention layers by splitting across heads;
Unveiling And Addressing Dimensional Collapse In Vector Quantization Models Via Codebook Regularization
Fang Zhang (University of Science and Technology of China), Linli Xu (University of Science and Technology of China)
GenerationCompressionDiffusion modelAuto EncoderContrastive LearningImage
🎯 What it does: Investigate and address the codebook dimension collapse problem in vector quantization models
Unveiling Multi-regime Patterns in SciML: Distinct Failure Modes and Regime-specific Optimization
Yuxin Wang (Dartmouth College), Yaoqing Yang (Dartmouth College)
OptimizationExplainability and InterpretabilityComputational EfficiencyContrastive LearningTabularTime SeriesSequentialPhysics RelatedStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: This study systematically investigates three training scenarios (well-trained, under-trained, overfitted) and corresponding failure modes of scientific machine learning (SciML) models (such as PINNs, FNOs, PINO, NODEs, PINODEs) under different physical parameters, training supervision, and optimization strategies, by constructing a polymorphic diagnostic framework based on training and testing errors, training dynamics, and the geometric properties of the loss landscape.
Unveiling Prior-Data Fitted Networks on Causal Effect Estimation: Pre-Training or Fine-Tuning?
Haotian Wang (National University of Defense Technology), Zhouchen Lin (Peking University)
Explainability and InterpretabilityRepresentation LearningTransformerSupervised Fine-TuningContrastive LearningTabularReview/Survey PaperBenchmark
🎯 What it does: Analyzes the unified pre-training limits of Prior-Data Fitted Networks (PFN) in causal effect estimation, proving that a single observational structural causal model leads to an exponential intervention distribution space, resulting in 'prior uncoverage' that causes posterior inconsistency and estimation bias, and proposes two targeted fine-tuning strategies (Pointwise Intervention Fine-tuning PWF and Meta-Sampling Fine-tuning MSF) to restore local and global generalization.
Unveiling the Entropy Dynamics of Chain-of-Thought Reasoning
Ting Xu (Chinese University of Hong Kong), Jianye HAO
Explainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelTextBenchmarkChain-of-Thought
🎯 What it does: Investigate the entropy dynamics of Chain-of-Thought reasoning, revealing a two-stage structure from the uncertain region to the confidence region, and propose a real-time change-point detection based on CUSUM to achieve early exit and test-time scaling.
Unveiling the Potential of Quantization with MXFP4: Strategies for Quantization Error Reduction
Jatin Chhugani (Meta Platforms Inc), Changkyu Kim (Meta Platforms Inc)
Computational EfficiencyKnowledge DistillationTransformerLarge Language ModelPrompt EngineeringText
🎯 What it does: To address the insufficient accuracy of 4-bit MXFP4 quantization, two software techniques, Overflow-Aware Scaling (OAS) and Macro Block Scaling (MBS), are proposed, significantly improving quantization quality.
Unveiling the Role of Data Uncertainty in Tabular Deep Learning
Nikolay Kartashev (HSE University), Artem Babenko (Yandex)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningData-Centric LearningMixture of ExpertsAuto EncoderContrastive LearningTabularBenchmarkRetrieval-Augmented Generation
🎯 What it does: Systematically analyze various tabular deep learning methods from the perspective of data uncertainty and propose a new numerical feature embedding scheme.
Unveiling the Structure of Do-Calculus Reasoning via Derivation Graphs
Clément Yvernes (Univ Grenoble Alpes), Eric Gaussier (Univ Grenoble Alpes)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningAuto EncoderGenerative Adversarial NetworkContrastive LearningGraphTabularTime SeriesBenchmarkPhysics Related
🎯 What it does: Propose and study the 'derivation graph' to represent and organize all equivalent transformations of Pearl's do-calculus rules on causal expressions, proving that any equivalent expression can be obtained through two R2 and two R3 rules, and provide a graphical criterion for determining equivalence.
Unveiling the Visual Counting Bottleneck in Vision-Language Models
Xingzhou Pang (ETH Zürich), Mrinmaya Sachan (ETH Zürich)
Explainability and InterpretabilityRepresentation LearningPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodality
🎯 What it does: Studied the systematic generalization bottleneck of large-scale vision-language models in visual counting tasks, and diagnosed the issue by decomposing counting into three stages: visual individuation, numerical perception, and symbol mapping.
UOTIP: Unbalanced Optimal Transport Map for Unpaired Inverse Problems
Donggyu Lee (Seoul National University), Jaewoong Choi (Sungkyunkwan University)
RestorationOptimizationConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelImage
🎯 What it does: Proposes UOTIP, an unpaired inverse problem solver based on Unbalanced Optimal Transport.
Updating Parametric Knowledge with Context Distillation Retains Post-Training Capabilities
Shankar Padmanabhan (Cornell University), Tanya Goyal (Cornell University)
Knowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningTextBiomedical DataBenchmarkRetrieval-Augmented Generation
🎯 What it does: This paper studies how to perform continuous knowledge adaptation on LLMs that have already completed post-training, while retaining existing capabilities during the update process.
Upper-Linearizability of Online Non-Monotone DR-Submodular Maximization over Down-Closed Convex Sets
Yiyang Lu (Purdue University), Vaneet Aggarwal (Purdue University)
Optimization
🎯 What it does: Propose an online maximization algorithm for non-monotone DR-submodular functions over the down-closed convex set, breaking the 1/e approximation lower bound and achieving O(√T) static, adaptive, and dynamic regret.
UrbanFusion: Stochastic Multimodal Fusion for Contrastive Learning of Robust Spatial Representations
Dominik J. Mühlematter (ETH Zurich), Nina Wiedemann (ETH Zurich)
Federated LearningRepresentation LearningData-Centric LearningTransformerVision Language ModelAuto EncoderContrastive LearningImageTextMultimodality
🎯 What it does: Learned city location representations that generalize under missing modalities by randomly masking multimodal inputs and using contrastive and reconstruction losses.
UrbanMLLM: Joint Learning of Cross-view Imagery for Urban Understanding
Xin Zhang (Tsinghua University), Yong Li (Tsinghua University)
Recommendation SystemAnomaly DetectionAutonomous DrivingOptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningDrug DiscoveryReinforcement Learning from Human FeedbackNeural Architecture SearchProtein Structure PredictionTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsVision Language ModelDiffusion modelScore-based ModelFlow-based ModelRectified FlowNeural Radiance FieldAuto EncoderGenerative Adversarial NetworkContrastive LearningImageTextMultimodalityBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose a unified multi-modal large language model called UrbanMLLM, which jointly learns from satellite images and street view images to achieve a more comprehensive understanding of cities.
URS: A Unified Neural Routing Solver for Cross-Problem Zero-Shot Generalization
Changliang Zhou (Southern University of Science and Technology), Qingfu Zhang (City University of Hong Kong)
OptimizationTransformerReinforcement LearningMixture of ExpertsGraphTabular
🎯 What it does: Proposed a unified neural routing solver (URS) that achieves zero-shot generalization across problems, capable of solving 110 different vehicle routing problem (VRP) variants on a single model;
Use What You Know: Causal Foundation Models with Partial Graphs
Arik Reuter (University of Cambridge), Bernhard Schölkopf (Max Planck Institute for Intelligent Systems)
Federated LearningExplainability and InterpretabilityComputational EfficiencyRepresentation LearningData-Centric LearningGraph Neural NetworkTransformerPrompt EngineeringContrastive LearningGraphTabularTime SeriesSequentialBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposes a method to inject partial causal graph (ancestry matrix) information into the causal foundation model (CFM), enabling it to utilize known causal structures when inferring causal effects.
USE: A Unified Self-Ensembling Framework for Test-Time Prompt Tuning
Siru Jiang (University of Chinese Academy of Sciences), Tieniu Tan (University of Chinese Academy of Sciences)
ClassificationDomain AdaptationOptimizationComputational EfficiencyTransformerPrompt EngineeringVision Language ModelAuto EncoderContrastive LearningImageTextMultimodality
🎯 What it does: Propose a unified self-integration framework called USE, which employs a self-integration strategy (SE) to perform self-supervised optimization on text prompts during testing, and uniformly applies the same strategy during inference to obtain more reliable pseudo-labels;
User-Aware Active Knowledge Acquisition for Emotional Support Dialogue
Mufan Xu (Harbin Institute of Technology), Min Zhang (Harbin Institute of Technology)
Recommendation SystemExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposes the User-Aware Active Knowledge Acquisition (UKA) framework, which uses gradient-agnostic active learning methods to actively acquire and utilize emotional knowledge in emotional support dialogues.
Utility Boundary of Dataset Distillation: Scaling and Coverage Laws
Zhengquan Luo (Mohamed bin Zayed University of Artificial Intelligence), zhiqiang xu
OptimizationKnowledge DistillationData-Centric LearningConvolutional Neural NetworkContrastive LearningImage
🎯 What it does: Proposed a configuration-dynamic-error (CDE) analysis framework to unify different matching-based dataset distillation methods and analyze the robustness of distillation under different configurations.
Utility-Diversity Aware Online Batch Selection for LLM Supervised Fine-tuning
Heming Zou (Tsinghua University), Xiangyang Ji (Tsinghua University)
Computational EfficiencyData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningText
🎯 What it does: A framework for online batch selection (UDS) oriented towards supervised fine-tuning of large language models was researched and implemented, aiming to improve efficiency and effectiveness by dynamically evaluating and filtering training samples.
Utonia: Toward One Encoder for All Point Clouds
Yujia Zhang (University of Hong Kong), Hengshuang Zhao (University of Hong Kong)
SegmentationAutonomous DrivingKnowledge DistillationRobotic IntelligenceTransformerAuto EncoderContrastive LearningPoint Cloud
🎯 What it does: This paper proposes Utonia, a single self-supervised point cloud Transformer encoder that can be jointly trained and shared across multiple point cloud domains (indoor, outdoor, object, remote sensing, video lifted point cloud).
V-ABS: Action-Observer Driven Beam Search for Dynamic Visual Reasoning
Zhiwei Ning (Shanghai Jiao Tong University), Wei Liu (Shanghai Jiao Tong University)
Computational EfficiencyRepresentation LearningRobotic IntelligenceTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision-Language-Action ModelImageTextMultimodality
🎯 What it does: Propose the V-ABS framework, achieving a closed-loop Beam Search of think-execute-observe for dynamic visual reasoning;
V-LynX: Token Interface Alignment for Video+X LLMs
Jungin Park (Yonsei University), Kwanghoon Sohn (Yonsei University)
Domain AdaptationComputational EfficiencyRepresentation LearningTransformerLarge Language ModelVision Language ModelContrastive LearningVideoTextMultimodalityAudio
🎯 What it does: Propose the V-LynX framework, which leverages the internal token interface of video LLMs to map new modalities into the existing video embedding space, achieving a lightweight, zero-shot multi-modal extension.
V1: Unifying Generation and Self-Verification for Parallel Reasoners
Harman Singh (University Of California Berkeley), Kurt Keutzer (University Of California Berkeley)
GenerationComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringContrastive LearningTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: This study proposes a parallel reasoning framework called V1, which unifies generation and self-verification, and improves the test-time extrapolation performance by ranking candidate answers for comparison.
Value Aggregation with Uncertainty in Online Decentralized MARL
Ziyue Chu (University of Birmingham), Leonardo Stella (University of Birmingham)
Federated LearningReinforcement LearningGraphTabular
🎯 What it does: An adaptive global consensus (AGC) mechanism is designed for online decentralized multi-agent reinforcement learning (MARL), used for value aggregation based on uncertainty to eliminate learning rollback.
Value-as-Return: A Two-Stage Framework to Align on the Optimal Score Function
Shikun Sun (Tsinghua University), Jia Jia (Tsinghua University)
GenerationOptimizationReinforcement Learning from Human FeedbackTransformerSupervised Fine-TuningReinforcement LearningDiffusion modelScore-based ModelImageTextMultimodalityStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Proposes a two-stage reinforcement learning framework VRPO, which first learns a paragraph-level value distribution function, and then uses this value to perform stage-wise DPO fine-tuning on a diffusion model, ultimately achieving precise alignment with image rewards.
VALUEFLOW: Toward Pluralistic and Steerable Value-based Alignment in Large Language Models
Woojin Kim (Seoul National University), Jaeyoung Do (Seoul National University)
Explainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringContrastive LearningText
🎯 What it does: Propose the VALUEFLOW framework to achieve controllable generation with adjustable intensity for value extraction and evaluation;
VAnim: Rendering-Aware Sparse State Modeling for Structure-Preserving Vector Animation
Guotao Liang (Beihang University), Qian Yu (University of Hong Kong)
GenerationData SynthesisTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision Language ModelImageVideoTextMultimodality
🎯 What it does: Designed and implemented VAnim, an open-domain text-to-SVG animation generation framework based on large language models, viewing animations as sparse state updates on a persistent SVG DOM tree.
Var-JEPA: A Variational Formulation of the Joint-Embedding Predictive Architecture – Bridging Predictive and Generative Self-Supervised Learning
Moritz Gögl, Christopher Yau (University of Oxford)
Representation LearningData-Centric LearningTransformerAuto EncoderContrastive LearningTabular
🎯 What it does: This paper reinterprets the Joint-Embedding Predictive Architecture (JEPA) as a variational inference framework, proposing Variational JEPA (Var-JEPA), which naturally avoids representation collapse through a single ELBO and can generate uncertainty estimates; subsequently, this framework is applied to heterogeneous tabular data, constructing Var-T-JEPA.
Variable Clustering via Distributionally Robust Nodewise Regression
Kaizheng Wang (Columbia University), XUNYU ZHOU
OptimizationFederated LearningComputational EfficiencyRepresentation LearningContrastive LearningImageGraphTabularFinance Related
🎯 What it does: A distributed robust node regression (DRO) method is proposed for variable clustering, and it is combined with the multi-factor block model and subspace clustering.
Variable-Length Tokenization via Learnable Global Merging for Diffusion Transformers
Dong Hoon Lee (KAIST), Seunghoon Hong (KAIST)
GenerationTransformerDiffusion modelAuto EncoderContrastive LearningImage
🎯 What it does: Propose a variable-length tokenizer based on learnable global merging (Learnable Global Merging, LGM), and combine it with a diffusion Transformer to generate high-quality images at different compression rates.
Variance Driven Exploration: A Provable and Efficient Methodology for Pure Exploration in Highly Stochastic Environments
Khang Luong (Hanoi University of Science and Technology), Tuan Quang Dam (Quantum AI & Cyber Security Institute, FPT Corporation)
Reinforcement LearningScore-based ModelGaussian SplattingTabularTime SeriesSequentialBenchmarkStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Proposes the Variance Driven Exploration (VarDE) method for sample allocation in pure exploration scenarios, with the core idea of minimizing the uncertainty of the final decision by multiplying the sensitivity weight of the decision function with the local variance;
Variance-Reduced $(\varepsilon, \delta)-$Unlearning using Forget Set Gradients
Martin Van Waerebeke (INRIA), El-Mahdi El-Mhamdi (Ecole polytechnique)
OptimizationSafty and PrivacyTabular
🎯 What it does: Proposed a first-order (ε,δ)-machine unlearning algorithm called Variance-Reduced Unlearning (VRU), which can directly utilize the gradients of the forgotten samples for updates, and accelerate model retraining while maintaining differential privacy guarantees.
Variational Adapter for Cross-modal Similarity Representation
WenZhang Wei (Wuhan University), Huayi Wu (Wuhan University)
RetrievalRepresentation LearningTransformerAuto EncoderContrastive LearningImageTextMultimodality
🎯 What it does: Designed and verified a variational adapter VACSR to model cross-modal similarity as a latent probability distribution, to alleviate the problem of false negative samples caused by binary annotations.
Variational Bayesian Flow Network for Graph Generation
Yida Xiong (Wuhan University), Wenbin Hu (Wuhan University)
GenerationDrug DiscoveryGraph Neural NetworkDiffusion modelFlow-based ModelGraph
🎯 What it does: Proposes a Variational Bayesian Flow Network (VBFN) for discrete generation of graph structures, utilizing structured precision to achieve coupled reasoning of nodes and edges.
Variational Entropic Optimal Transport
Roman Dyachenko (Higher School of Economics), Alexander Korotin (Applied AI Institute)
Image TranslationOptimizationComputational EfficiencyDiffusion modelScore-based ModelNeural Radiance FieldAuto EncoderContrastive LearningImagePoint CloudTabularStochastic Differential Equation
🎯 What it does: Proposed a variational entropy regularized optimal transport (VarEOT) based on a weak dual form of variational rewriting, achieving efficient solution without simulated training.
Variational Flow Maps: Make Some Noise for One-Step Conditional Generation
Abbas Mammadov (University of Oxford), Julius Berner (NVIDIA)
RestorationGenerationTransformerDiffusion modelScore-based ModelFlow-based ModelAuto EncoderContrastive LearningImageStochastic Differential Equation
🎯 What it does: Propose the Variational Flow Maps (VFM) framework, which maps noise to the data space by learning a noise sampler conditioned on observations, enabling one-shot conditional generation.
Variational Inference for Uncertain Optimal Transport via Sinkhorn Parametrization
Ananyapam De (Technische Universität Clausthal), Benjamin Säfken
OptimizationComputational EfficiencyRepresentation LearningScore-based ModelFlow-based ModelContrastive LearningImagePoint CloudTabularBiomedical DataStochastic Differential Equation
🎯 What it does: Proposes an scalable variational inference framework SPVI, which learns the posterior distribution within the transportation polytope using Sinkhorn parameterization.
Variational inference via Gaussian interacting particles in the Bures-Wasserstein geometry
Giacomo Borghi (Heriot-Watt University), Jose A. Carrillo
OptimizationScore-based ModelGaussian SplattingTabularBenchmarkStochastic Differential Equation
🎯 What it does: Propose a high-dimensional Gaussian Consensus-Based Optimization (Gaussian CBO) algorithm based on the linearized Bures-Wasserstein space, used for solving Gaussian variational inference problems, i.e., finding the best approximation of the target distribution within the Gaussian measure space through a stochastic consensus-driven particle system.
Variational Learning for Insertion-based Generation
Yangtian Zhang (Yale University), Jiaxin Shi (Meta Superintelligence Labs)
GenerationDrug DiscoveryTransformerReinforcement LearningDiffusion modelTextSequentialBiomedical Data
🎯 What it does: Propose a variable-length generation model called Insertion Process (IP), which learns the insertion order to achieve non-unidirectional sequence generation.
Variational Learning of Disentangled Representations
Yuli Slavutsky (Columbia University), Bianca Dumitrascu (Columbia University)
Representation LearningAuto EncoderGenerative Adversarial NetworkContrastive LearningImageTabularBiomedical Data
🎯 What it does: Propose DisCoVR, a framework based on variational inference for learning separable shared (z) and condition-specific (w) representations from multi-condition data.
Variational Routing: A Scalable Bayesian Framework for Calibrated Mixture-of-Experts Transformers
Albus Li, Matthew Robert Wicker
Explainability and InterpretabilityComputational EfficiencyTransformerMixture of ExpertsText
🎯 What it does: Introduce Variational Routing (VMoER) in Mixture-of-Experts Transformer, achieving calibrated uncertainty inference by performing Bayesian uncertainty modeling in the expert routing layer through variational inference.
Variational Speculative Decoding: Rethinking Draft Training from Token Likelihood to Sequence Acceptance
Xiandong Zou (Singapore Management University), Pan Zhou (Singapore Management University)
Computational EfficiencyAI Code AssistantTransformerLarge Language ModelTextMultimodality
🎯 What it does: Propose the Variational Speculative Decoding (VSD) framework, redesigning the training of the draft model to shift from a single token likelihood objective to variational inference based on sequence acceptance rate.
VBA: Vector Bundle Attention for Intrinsically Geometric Representation Learning
Shenglei Fang (Cardiff University), You Zhou (Cardiff University)
Representation LearningTransformerContrastive LearningPoint CloudTabularBiomedical Data
🎯 What it does: A Transformer model based on the vector bundle theory was constructed, embedding geometric information into the attention mechanism to achieve feature alignment and similarity calculation with geometric consistency.
VCG-Bench: Towards A Unified Visual-Centric Benchmark for Structured Generation and Editing
Xiaoyan Su (Hong Kong University of Science and Technology), Xiaowen Chu (Hong Kong University of Science and Technology)
GenerationData SynthesisLarge Language ModelPrompt EngineeringVision Language ModelImageTextGraphTabularBenchmarkRetrieval-Augmented Generation
🎯 What it does: Establish a unified visual centralization benchmark, VCG-Bench, to evaluate the ability of vision-language models in chart generation and editable code conversion.
VDW-GNNs: Vector diffusion wavelets for geometric graph neural networks
David R Johnson (Boise State University), Michael Perlmutter (Boise State University)
Computational EfficiencyRepresentation LearningData-Centric LearningGraph Neural NetworkDiffusion modelAuto EncoderContrastive LearningPoint CloudGraphTime Series
🎯 What it does: Proposes Vector Diffusion Wavelets (VDW) and integrates them into graph neural networks, forming VDW-GNN, for processing geometric graphs with vector features in Euclidean space.
VecDesigner: Exploring Visual Guidance and Structural Consistency for Semantic Typography
Liu Yu (East China Normal University), Liang He (East China Normal University)
GenerationData SynthesisTransformerPrompt EngineeringDiffusion modelScore-based ModelContrastive LearningImageTextMultimodality
🎯 What it does: Proposed a vector semantic layout method called VecDesigner, which can deform the glyph according to the semantics of the input word while maintaining the readability of the characters.
VecMol: Vector-Field Representations for 3D Molecule Generation
Yuchen Hua (Peking University), Muhan Zhang (Peking University)
GenerationRepresentation LearningDrug DiscoveryGraph Neural NetworkDiffusion modelNeural Radiance FieldAuto EncoderContrastive LearningPoint CloudGraphBiomedical DataOrdinary Differential Equation
🎯 What it does: This paper proposes a new 3D molecular generation framework called VecMol, which uses a continuous vector field as the molecular representation. It maps molecules to a fixed-dimensional latent space through a neural field autoencoder, and then generates molecules using a latent diffusion probabilistic model in this latent space. Finally, atomic coordinates and bond structures are recovered from the vector field through gradient ascent and clustering.
Vector Linking via Cross-Model Local Isometric Consistency
Ziying Chen (University of Edinburgh), Tianjian Yang (University of Edinburgh)
RetrievalRepresentation LearningTransformerAuto EncoderContrastive LearningOptical FlowTextBenchmarkStochastic Differential Equation
🎯 What it does: Studies how to achieve cross-model vector linking between embedding clouds generated by two different black-box contrastive learning encoders, using only a small number of known corresponding seeds.
VectorWorld: Efficient Streaming World Model via Diffusion Flow on Vector Graphs
Chaokang Jiang (Bosch), Li Sun (Bosch)
GenerationData SynthesisAutonomous DrivingComputational EfficiencyGraph Neural NetworkTransformerDiffusion modelFlow-based ModelAuto EncoderWorld ModelGraphTabularSequential
🎯 What it does: Propose VectorWorld, a system capable of real-time vector map world modeling and dynamic scene rendering in closed-loop evaluation.
Veda: Scalable Video Diffusion via Distilled Sparse Attention
Shihao Han (ByteDance), Xiaojuan Qi (University of Hong Kong)
GenerationComputational EfficiencyKnowledge DistillationTransformerDiffusion modelScore-based ModelContrastive LearningVideo
🎯 What it does: Proposed the Veda framework, achieving efficient generation in video diffusion Transformers through sparse attention
Vegas: Self-Speculative Decoding with Verification-Guided Sparse Attention
Yikang Yue (University of Illinois Urbana Champaign), Jian Huang (University of Illinois Urbana Champaign)
Computational EfficiencyTransformerLarge Language ModelText
🎯 What it does: To address the memory bottleneck of KV cache in long-context LLM inference, the Vegas scheme is proposed, which synergistically designs the draft phase and the verification phase of inference.
VELR: Efficient Video Reward Feedback via Ensemble Latent Reward Models
Liyu Zhang (Zhejiang University), Jiming Chen (Zhejiang University)
Computational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackConvolutional Neural NetworkTransformerReinforcement LearningMixture of ExpertsDiffusion modelContrastive LearningVideoText
🎯 What it does: Propose an efficient video reward feedback learning framework, VELR, which utilizes an integrated latent reward model (LRM) to address the memory bottleneck of traditional video reward models in ReFL;
VENOMREC: Cross-Modal Interactive Poisoning for Targeted Promotion in Multimodal LLM Recommender Systems
Guowei Guan (Nanyang Technological University), Wei Yang Bryan Lim (Nanyang Technological University)
Recommendation SystemAdversarial AttackTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodality
🎯 What it does: A cross-modal interactive poisoning attack called VENOMREC is proposed for multi-modal large language model recommendation systems. It utilizes synchronized visual and textual perturbations to guide the fused representation towards high-exposure semantic directions, achieving targeted promotion of target products.