arXivSub Start free trial

ICML 2026 Papers — Page 31

International Conference on Machine Learning · 6554 papers

Learn from Your Mistakes: Tree-like Self-Play for Secure Code LLMs

Wenqi Chen (University of Electronic Science and Technology of China), Zhengsu Chen (Beihang University)

Safty and PrivacyAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningContrastive LearningText

🎯 What it does: Propose the Tree-like Self-Play (TSP) framework, which enhances the safety of code generation by conducting self-play at critical risk nodes in the code generation tree.

Learn to change the world: Multi-level reinforcement learning with model-changing actions

Ziqing Lu (University of Iowa), Weiyu Xu (University of Iowa)

Reinforcement LearningTabularTime Series

🎯 What it does: This paper proposes a reinforcement learning framework that can actively change the environment dynamics, defining a multi-layer configurable time-varying Markov decision process (MCTVMDP), and decomposing it into an upper configuration layer and a lower execution layer, constructing two special subproblems: bi-layer configurable MDP and time-varying configurable MDP.

Learn to Merge: Meta-Learning for Adaptive Multi-Task Model Merging

Jun Chen (Shenzhen University), Ziyue Qiao (Great Bay University)

ClassificationMeta LearningTransformerMixture of ExpertsImageText

🎯 What it does: This paper proposes the MetaMerging framework, which utilizes meta-learning to adaptively optimize model merging coefficients, constructing a unified multi-task model suitable for subsequent adapter training.

Learn to Think: Improving Multimodal Reasoning through Vision-Aware Self-Improvement Training

Qihuang Zhong (Wuhan University), Dacheng Tao (Nanyang Technological University)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelImageTextMultimodalityChain-of-Thought

🎯 What it does: Enhance the reasoning ability of multimodal large language models through self-improvement training, and propose the VISTA framework

Learn-to-learn on Arbitrary Textual Conditioning: A Hypernetwork-Driven Meta-gated LLM

Luo Ji (Geely AI Lab, Geely Auto Group), Hongyan Li (Geely AI Lab, Geely Auto Group)

Meta LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringText

🎯 What it does: Propose MeGan, which utilizes a hypernetwork to generate adaptive β-SwiGLU gating, injecting meta-learning control signals into the FFN of LLMs, enabling fast adaptation under any text conditions.

Learnability-Driven Knowledge Assimilation for Class-Incremental Semantic Segmentation

Xinyue Zhang (Huazhong University of Science and Technology), Luxin Yan (Huazhong University of Science and Technology)

SegmentationKnowledge DistillationTransformerSupervised Fine-TuningContrastive LearningImage

🎯 What it does: Propose a learnable-driven knowledge assimilation framework called LDKA, targeting the optimization bottleneck in low boundary regions during class-incremental semantic segmentation.

Learnability-Informed Fine-Tuning of Diffusion Language Models

Shubham Parashar (Texas A&M University), Shuiwang Ji (Texas A&M University)

Computational EfficiencyRepresentation LearningData-Centric LearningTransformerSupervised Fine-TuningDiffusion modelAuto EncoderText

🎯 What it does: Propose a learnable mask fine-tuning method called LIFT, which adaptively masks training for token learning difficulty and timing in diffusion language models.

Learnable Kernel Density Estimation for Graphs and Its Application to Graph-Level Anomaly Detection

Xudong Wang (Chinese University of Hong Kong, Shenzhen), Jicong Fan (Chinese University of Hong Kong, Shenzhen)

Anomaly DetectionRepresentation LearningGraph Neural NetworkContrastive LearningGraph

🎯 What it does: Propose a learnable graph kernel density estimation framework, LGKDE, for learning the probability density of graphs and achieving graph-level anomaly detection.

LearniBridge: Learnable Calibration of Feature Caching for Diffusion Models Acceleration

Xuyue Huang (Tsinghua University), Xiao-Ping Zhang (Tsinghua University)

GenerationComputational EfficiencyTransformerDiffusion modelImageVideo

🎯 What it does: Propose LearniBridge, a caching acceleration method that utilizes LoRA to perform learnable feature calibration on the last block of the Diffusion Transformer, significantly reducing the computational requirements during inference.

Learning $U$-Statistics with Active Inference

Xiaoning Wang (Nankai University), Changliang Zou (Nankai University)

OptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyData-Centric LearningReinforcement LearningContrastive LearningTabularBiomedical DataElectronic Health Records

🎯 What it does: Under the scenario of limited labeling budgets, this paper designs a sampling strategy for U-statistics through an active inference framework to achieve efficient estimation and effective inference;

Learning 3D-Gaussian Simulators from RGB Videos

Mikel Zhobro (University of Tübingen), Georg Martius (University of Tübingen)

GenerationData SynthesisOptimizationComputational EfficiencyRepresentation LearningTransformerDiffusion modelNeural Radiance FieldAuto EncoderContrastive LearningGaussian SplattingImageVideoPoint Cloud

🎯 What it does: This paper proposes an end-to-end differentiable 3D Gaussian simulator, 3DGSim, which can directly learn physical interactions from multi-view RGB videos, and achieve future frame prediction and visualization through inverse rendering, a Transformer-based dynamics model, and Gaussian polishing rendering;

Learning a Generative Meta-Model of LLM Activations

Grace Luo (UC Berkeley), Jacob Steinhardt (UC Berkeley)

GenerationData SynthesisExplainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelDiffusion modelFlow-based ModelAuto EncoderText

🎯 What it does: This study trains a diffusion model (GLP) to learn the distribution of internal activations in large language models (LLMs), and uses this model as a prior for on-manifold adjustment and interpretation tasks of activations.

Learning a Zeroth-Order Optimizer for Fine-Tuning LLMs

Kairun Zhang (University of Illinois Urbana Champaign), Huan Zhang (University of Illinois Urbana Champaign)

OptimizationComputational EfficiencyRepresentation LearningMeta LearningTransformerLarge Language ModelSupervised Fine-TuningText

🎯 What it does: Designed and trained a learning-based zeroth-order optimizer, ZO Fine-tuner, for efficiently fine-tuning large-scale language models.

Learning Adaptive Perturbation-Conditioned Contexts for Robust Transcriptional Response Prediction

Yinhua Piao (KAIST), Sungsoo Ahn (KAIST)

Drug DiscoveryGraph Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningDiffusion modelAuto EncoderContrastive LearningTextGraphBiomedical Data

🎯 What it does: To predict transcriptional responses in single-cell CRISPR knockout experiments, the ADAPERT model is proposed. It extracts subgraphs related to specific perturbations from a global knowledge graph through differential differentiable node selection, and combines an adaptive learning scheme (reconstruction loss, non-differential gene robust loss, alignment loss) to suppress mean collapse and improve the recovery of sparse gene responses.

Learning Adaptive Topology with FiLM-Guided Distillation for Tertiary Structure-Based RNA Design

Zixun Zhang (Shenzhen Future Network of Intelligence Institute), Zhen Li (Chinese University of Hong Kong)

Knowledge DistillationRepresentation LearningProtein Structure PredictionGraph Neural NetworkTransformerDiffusion modelContrastive LearningGraphBiomedical Data

🎯 What it does: Propose the ATL-FGD framework for tertiary structure-driven RNA design, achieving more accurate sequence recovery through learnable topology and knowledge distillation.

Learning Anisotropic Value Geometry with Finsler Reinforcement Learning

Jumman Hossain (University of Maryland, Baltimore County), Nirmalya Roy (University of Maryland, Baltimore County)

OptimizationRobotic IntelligenceReinforcement LearningTabularTime Series

🎯 What it does: Propose Finslerian Reinforcement Learning (FiRL), combining direction-sensitive Finsler cost with CVaR risk objective, to improve energy consumption and safety of legged robots in inclined or windy environments.

Learning Attribute–Affordance Hierarchies in Hyperbolic Space for Open-Vocabulary 3D Object Affordance Grounding

Yuxuan Wang (Xidian University), Cheng Deng (Xidian University)

RecognitionObject DetectionRepresentation LearningGraph Neural NetworkTransformerPrompt EngineeringVision-Language-Action ModelContrastive LearningImageMultimodalityPoint Cloud

🎯 What it does: This paper proposes an AAH framework for open-vocabulary 3D object affordance localization, which can model and map the hierarchical relationship between local attributes and affordance regions into hyperbolic space to improve localization accuracy.

Learning Biophysical Models of Large-Scale Multineuronal Data To Enable Precise Neurostimulation

Amrith Lotlikar (Stanford University), Subhasish Mitra (Stanford University)

OptimizationComputational EfficiencyDrug DiscoveryReinforcement Learning from Human FeedbackSpiking Neural NetworkDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningTime SeriesBiomedical DataElectrocardiogramPhysics RelatedStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Extract electrophysiological features using high-density multi-electrode arrays, and rapidly infer physiological parameters of retinal ganglion cells (RGCs) under multi-electrode stimulation by combining a differentiable multi-compartment Hodgkin-Huxley (HH) model with Simulated Bayesian Inference (SBI), thereby predicting firing responses.

Learning Cardiac Latent Representations in Vectorcardiogram Space

Bosong Huang (Griffith University), Shirui Pan (Griffith University)

Anomaly DetectionRepresentation LearningRecurrent Neural NetworkScore-based ModelFlow-based ModelRectified FlowAuto EncoderContrastive LearningGaussian SplattingTime SeriesElectrocardiogram

🎯 What it does: Proposes a self-supervised representation learning framework called LVCG in the vector electrocardiogram (VCG) space, utilizing geometric projection to transform multi-lead ECG into perspective-invariant latent representations.

Learning Coherent Representations: A Topological Approach to Interpretability

Sigurd Gaukstad (NTNU), Benjamin Adric Dunn

Explainability and InterpretabilityRepresentation LearningTransformerPrompt EngineeringVision-Language-Action ModelDiffusion modelAuto EncoderContrastive LearningImageText

🎯 What it does: This paper proposes a topology-based 'coherence' regularization, aiming to maintain topological consistency between the sample space and feature space of neural networks, thus obtaining interpretable representations.

Learning Compressed Shape-Aware Molecular Representations for Virtual Screening

Robin Winter (Pfizer), Djork-Arné Clevert (Pfizer)

RetrievalCompressionRepresentation LearningDrug DiscoveryGraph Neural NetworkAuto EncoderContrastive LearningGraphBiomedical Data

🎯 What it does: Utilize the SAND framework to learn compressed shape-aware representations based on 2D molecular graphs, enabling shape similarity retrieval from a 1-billion-level molecular library;

Learning Context-Conditioned Predicate Semantics via Prototype Feedback

NamGyu Jung (Gachon University), Chang Choi (Gachon University)

RecognitionObject DetectionGenerationRecurrent Neural NetworkTransformerVision Language ModelContrastive LearningImageText

🎯 What it does: The paper proposes the AlignG model, which achieves contextualized learning of predicate semantics in scene graph generation through a prototype feedback mechanism.

Learning Coupled Continuous-Time Latent Dynamics from Irregular Events

Jiankai Zuo (Suzhou University of Science and Technology), Yaying Zhang (Tongji University)

Recommendation SystemOptimizationComputational EfficiencyRecurrent Neural NetworkTransformerDiffusion modelAuto EncoderTabularTime SeriesSequentialStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Proposes the Coupled Continuous-Time Latent Dynamics (CoCLD) framework, jointly modeling the continuous dynamics of individuals and groups in time-irregular events.

Learning Credal Ensembles via Distributionally Robust Optimization

Kaizheng Wang (KU Leuven), Hans Hallez (KU Leuven)

ClassificationDomain AdaptationAnomaly DetectionConvolutional Neural NetworkContrastive LearningImageBiomedical DataBenchmark

🎯 What it does: Proposes a trustworthy ensemble model called CreDRO based on distributionally robust optimization (DRO), which is used to more accurately quantify the epistemic uncertainty (EU) of neural networks and improve robustness against distribution shifts.

Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic

Shuo Liu (Northeastern University), Christopher Amato (Northeastern University)

Federated LearningTransformerLarge Language ModelReinforcement LearningAgentic AIText

🎯 What it does: This study proposes two multi-agent actor-critic (MAAC) methods — CoLLM-CC (Central Critic) and CoLLM-DC (Distributed Critic) — for optimizing decentralized LLM collaboration;

Learning Discrete Diffusion on Graphs via Free-Energy Gradient Flows

Dario Rancati (Institute of Science and Technology Austria), Francesco Locatello (Institute of Science and Technology Austria)

GenerationOptimizationRepresentation LearningGraph Neural NetworkDiffusion modelScore-based ModelContrastive LearningGraphBiomedical DataStochastic Differential Equation

🎯 What it does: Propose a method for learning probabilistic flow dynamics on discrete graph spaces by utilizing discrete gradient flows and free energy gradient flow frameworks;

Learning Discriminative and Generalizable Anomaly Detector for Dynamic Graph with Limited Supervision

Yuxing Tian (Université de Montréal), Jian-Yun Nie (Université de Montréal)

Anomaly DetectionGraph Neural NetworkFlow-based ModelContrastive LearningGraph

🎯 What it does: A dynamic graph anomaly detection framework called SDGAD is designed to learn discriminative boundaries under label-free or few-label conditions.

Learning Disentangled Multi-Agent World Model for Decentralized Control

Di Xue (Nanjing University), Yang Yu (Nanjing University)

Recurrent Neural NetworkTransformerReinforcement LearningWorld ModelImageTabular

🎯 What it does: A learnable discrete world model called DMAWM is proposed and implemented for distributed control in multi-agent reinforcement learning, which can generate decoupled agent states in the latent space and use them to generate imagined trajectories for training distributed policies.

Learning Dynamic Stability Landscapes in Synchronization Networks

Christian Nauck (Potsdam Institute for Climate Impact Research), Frank Hellmann (Potsdam Institute for Climate Impact Research)

Anomaly DetectionOptimizationExplainability and InterpretabilityConvolutional Neural NetworkGraph Neural NetworkSupervised Fine-TuningContrastive LearningGraphTabularPhysics Related

🎯 What it does: This paper proposes to use graph neural networks (GNN) to directly predict the synchronization stability landscape (basin landscape) for each node from the network topology, rather than the traditional single scalar stability index;

Learning Dynamics of Zeroth-Order Optimization: A Kernel Perspective

Zhe Li (Rochester Institute of Technology), Haibo Yang (Rochester Institute of Technology)

OptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerSupervised Fine-TuningContrastive LearningImageText

🎯 What it does: This paper analyzes and explains the learning dynamics of zeroth-order stochastic gradient descent (ZO-SGD) in fine-tuning large models by introducing the empirical neural tangent kernel (eNTK) and the Johnson–Lindenstrauss (JL) theory.

Learning Efficient Guardrails for Compliance

Xiaofei Wen (University of CaliforniaDavis), Muhao Chen (University of CaliforniaDavis)

Anomaly DetectionComputational EfficiencyData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningTextSequentialBenchmarkRetrieval-Augmented Generation

🎯 What it does: Constructed a POLICYGUARDBENCH benchmark with 60k scale to evaluate policy compliance of autonomous network agents in long-path tasks, and trained a lightweight POLICYGUARD-4B model;

Learning Fingerprints for Medical Time Series with Redundancy-Constrained Information Maximization

Huayu Li (University of Arizona), Ao Li (University of Arizona)

Explainability and InterpretabilityRepresentation LearningData-Centric LearningTransformerAuto EncoderContrastive LearningTime SeriesBiomedical DataElectrocardiogram

🎯 What it does: This paper proposes a self-supervised pre-training framework called TS-Fingerprint, which can compress variable-length medical time series into a fixed number k of de-redundant Fingerprint Tokens;

Learning Flexible Generalization in Video Quality Assessment by Bringing Device and Viewing Condition Distributions

Nikolay Safonov (MSU Institute for Artificial Intelligence), Dmitriy S. Vatolin

Domain AdaptationComputational EfficiencyData-Centric LearningConvolutional Neural NetworkTransformerSupervised Fine-TuningAuto EncoderContrastive LearningImageVideo

🎯 What it does: Constructed a subjective comparative evaluation dataset covering more than 300 Android devices, containing display parameters and environmental information, and used the Blade-Chest model to aggregate preference votes. Subsequently, a lightweight conditional adaptation network was trained to improve the prediction of existing VQA metrics across different devices and environments.

Learning from Comparison: Constrained Projection Policy Optimization for Pareto-Front Improvement

Jintao Li (Shenzhen Institute for Advanced Study, University of Electronic Science and Technology of China), Yun Li (Shenzhen Institute for Advanced Study, University of Electronic Science and Technology of China)

OptimizationGraph Neural NetworkReinforcement LearningContrastive LearningGraphTabularBenchmark

🎯 What it does: Propose a new Constrained Projection (CoPro) framework that utilizes group-wise comparison information to generate non-negative weights through KL-regularized moment-constrained projection, directly learning the optimal strategy from hard constraints and multi-objective feedback;

Learning from Fine-Grained Visual Discrepancies: Mitigating Multimodal Hallucinations via In-Context Visual Contrastive Optimization

Haolin Deng (Hong Kong University of Science and Technology), Xuming Hu (Hong Kong University of Science and Technology)

OptimizationExplainability and InterpretabilityKnowledge DistillationRepresentation LearningData-Centric LearningTransformerPrompt EngineeringVision Language ModelDiffusion modelGenerative Adversarial NetworkContrastive LearningImageTextMultimodalityRetrieval-Augmented Generation

🎯 What it does: Propose a framework called IC-VCO for visual contrastive optimization in a shared multi-graph context, aiming to eliminate the partition function mismatch issue in traditional visual DPO and achieve more refined visual alignment through visual contrastive visualization of positive and negative samples;

Learning Gaussian Graphical Models from a Glauber Trajectory Without Mixing

Eric Shen (Massachusetts Institute of Technology), Ankur Moitra (Massachusetts Institute of Technology)

OptimizationComputational EfficiencyRepresentation LearningContrastive LearningGraphStochastic Differential Equation

🎯 What it does: Learn the structure of a sparse Gaussian graphical model on a single Glauber dynamics trajectory without relying on the chain's mixing time.

Learning Gaussian Mixture-distributed Prototypes for 3D Scene Graph Generation from RGB-D Sequences

Rongxing Ding (Nanjing University of Science and Technology), Xiangbo Shu (Nanjing University of Science and Technology)

GenerationRepresentation LearningRecurrent Neural NetworkGraph Neural NetworkTransformerMixture of ExpertsDiffusion modelContrastive LearningGaussian SplattingImagePoint Cloud

🎯 What it does: Studied how to improve 3D scene graph generation (3DSGG) based on RGB-D sequences by utilizing Gaussian Mixture Distribution Prototypes (GMP), through constructing multi-component prototypes in the category space to simultaneously capture intra-class diversity and inter-class similarity;

Learning General Causal Structures with Hidden Dynamic Process for Climate Analysis

Minghao Fu (Mohamed bin Zayed University of Artificial Intelligence), Kun Zhang (Mohamed bin Zayed University of Artificial Intelligence)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerScore-based ModelFlow-based ModelAuto EncoderContrastive LearningTabularTime SeriesPhysics Related

🎯 What it does: Propose a unified framework called CaDRe, which can simultaneously learn the causal structure of hidden dynamic processes and the causal relationships between observed variables from observed time-series data.

Learning Generalizable Skill Policy with Data-Efficient Unsupervised RL

Jongchan Park (Sungkyunkwan University), Yusung Kim (Sungkyunkwan University)

TransformerReinforcement LearningAuto EncoderContrastive LearningImageTabularTime Series

🎯 What it does: This paper proposes the GenDa framework, which addresses the issues of semantic drift and overfitting to global context in existing off-policy URL through two modules: skill remarking and mutual information bottleneck (CIB), significantly improving the efficiency and generality of pre-training.

Learning Generalized Label Distributions

Haitao Wu (Nanjing University of Science and Technology), Xiuyi Jia (Nanjing University of Science and Technology)

Anomaly DetectionRepresentation LearningData-Centric LearningImageMultimodalityTabular

🎯 What it does: Proposed a general label distribution representation — General Label Distribution (GLD), which can fully restore the original data, maintain the order consistency between samples, and map to multiple label representations, such as multi-label and label ranking, without information loss;

Learning Generalized Trackers with Elastic Token Budgets

Yinchao Ma (University Of Science And Technology Of China), Tianzhu Zhang (University Of Science And Technology Of China)

Object TrackingComputational EfficiencyTransformerReinforcement LearningContrastive LearningVideo

🎯 What it does: Through the post-training framework ETBTrack, Transformer-based visual trackers can achieve robust tracking under an elastic token budget.

Learning Global Representation from Queries for Vectorized HD Map Construction

Shoumeng Qiu (Fudan University), Jian Pu (Fudan University)

Autonomous DrivingOptimizationRepresentation LearningConvolutional Neural NetworkTransformerContrastive LearningSimultaneous Localization and MappingImagePoint Cloud

🎯 What it does: This paper proposes an online vectorized high-precision map construction framework called MapGR, which improves map quality by learning global representations of queries and guiding local queries.

Learning Graph Foundation Models on Riemannian Graph-of-Graphs

Haokun Liu (University of Science and Technology of China), Xike Xie (University of Science and Technology of China)

Representation LearningGraph Neural NetworkMixture of ExpertsContrastive LearningGaussian SplattingGraph

🎯 What it does: Proposed a Riemannian geometry-based Graph-of-Graphs (GoG) foundational model called R-GFM, which can capture multi-scale structural information and adapt to different graph domains during the pre-training phase through adaptive hop-number subgraph sampling, similarity-driven GoG construction, and dynamic Mixture-of-Experts (MoE) routing.

Learning GUI Grounding with Spatial Reasoning from Visual Feedback

Yu Zhao (University of Edinburgh), Robert Sim (Microsoft)

Computational EfficiencyRepresentation LearningTransformerSupervised Fine-TuningReinforcement LearningAgentic AIVision Language ModelImageMultimodalityChain-of-Thought

🎯 What it does: Propose an interactive GUI localization method, where the model gradually searches for target UI elements by moving a virtual cursor and provides visual feedback at each step;

Learning Hamiltonian Dynamics at Scale: A Differential-Geometric Approach

Katharina Friedl (KTH Royal Institute of Technology), Danica Kragic (KTH Royal Institute of Technology)

Convolutional Neural NetworkRecurrent Neural NetworkGraph Neural NetworkAuto EncoderContrastive LearningTabularTime SeriesSequentialPhysics RelatedStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Propose a high-dimensional Hamiltonian system learning framework named RO-HNN, which can preserve the symplectic structure in a low-dimensional latent space and learn energy-conserving or dissipative dynamics;

Learning Hamiltonian Flow Maps: Mean Flow Consistency for Large-Timestep Molecular Dynamics

Winfried Ripken (Technical University Berlin), Klaus Robert Muller

Drug DiscoveryTransformerFlow-based ModelGraphTabularPhysics RelatedStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: This paper proposes and trains a Hamiltonian Flow Map (HFM) without trajectory learning, which learns the average displacement field over any time interval by sampling phase space states and their instantaneous force labels at once, thereby achieving stable numerical integration with large time steps.

Learning High-Dimensional Parity Functions with Product Networks using Gradient Descent

Guillaume Larue (Orange Research), Rekaya-Ben Othman

OptimizationComputational EfficiencyRepresentation LearningTabular

🎯 What it does: Propose a high-dimensional XOR function learning method under sparse input based on product networks,

Learning High-Frequency Continuous Action Chunks in Latent Space

Kunyun Wang (Shanghai Jiao Tong University), Wenchao Ding (Fudan University)

Robotic IntelligenceTransformerReinforcement LearningFlow-based ModelRectified FlowAuto EncoderTime SeriesSequential

🎯 What it does: In this study, the authors propose using a variational autoencoder (VAE) in the latent space to learn high-frequency continuous action blocks, and design a 'Reuse-then-Refine' (RTR) strategy to maintain continuity between action blocks under asynchronous inference, thereby enabling smooth and continuous execution by robots under high-frequency control.

Learning Human-Robot Collaboration via Heterogeneous-Agent Lyapunov Policy Optimization

Hao Zhang (University of Texas at Arlington), Eric H. Tseng

OptimizationRobotic IntelligenceReinforcement LearningPoint CloudTabularTime Series

🎯 What it does: This paper proposes a heterogeneous multi-agent reinforcement learning framework called HALO for human-robot collaboration, achieving stable convergence of decoupled gradient updates by introducing Lyapunov stability guarantees in the policy parameter space.

Learning in Bayesian Stackelberg Games With Unknown Follower's Types

Matteo Bollini (Politecnico di Milano), Alberto Marchesi (Politecnico di Milano)

OptimizationReinforcement Learning

🎯 What it does: Studies an algorithm for online learning in Bayesian Stackelberg games, where the leader learns with followers of unknown types, aiming to minimize the leader's regret.

Learning in Structured Stackelberg Games

Maria Florina Balcan, Keegan Harris (University of California, Berkeley)

OptimizationReinforcement LearningContrastive Learning

🎯 What it does: Studied the learning problem in Stackelberg games with contextual information, proposed a structured Stackelberg game model, and provided theoretical upper bounds for instance-optimal online and distributed learning.

Learning in the Fisher Subspace: A Guided Initialization for LoRA Fine-Tuning

Zhi-Quan Feng (National Cheng Kung University), Hung-Yu Kao (National Tsing Hua University)

Computational EfficiencyRepresentation LearningData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningImageText

🎯 What it does: Propose a LoRA initialization method based on Fisher energy (FILet), which selects low-sensitivity directions by estimating Fisher information under the data distribution, thereby better performing parameter-efficient fine-tuning of large models.

Learning Interpretable Options by Identifying Reward Diffusion Bottlenecks in Reinforcement Learning

Yiming Fei (Zhejiang University), Huajin Tang (Zhejiang University)

OptimizationExplainability and InterpretabilityReinforcement LearningContrastive LearningImageTabularTime Series

🎯 What it does: This paper proposes a bottleneck metric called Value Power Strength (VPS) based on the value function, which identifies reward diffusion bottlenecks. By using VPS as a potential function, interpretable options are learned, thereby improving exploration efficiency in sparse reward environments.

Learning Junta Distributions, Quantum Junta States, and QAC$^0$ Circuits

Jinge Bao (University of Edinburgh), Francisco Escudero Gutiérrez (Qusoft and CWI)

OptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyRepresentation LearningDiffusion modelScore-based ModelFlow-based ModelRectified FlowNeural Radiance FieldAuto EncoderContrastive LearningGaussian SplattingPhysics Related

🎯 What it does: Studied the learning and testing of k-junta distributions, quantum k-junta states, and QAC0 circuits.

Learning Latent Action World Models in the Wild

Quentin Garrido (FAIR at Meta), Michael Rabbat (FAIR at Meta)

Autonomous DrivingRepresentation LearningRobotic IntelligenceTransformerReinforcement LearningVision-Language-Action ModelAuto EncoderContrastive LearningWorld ModelOptical FlowVideoChain-of-Thought

🎯 What it does: Studied latent action world models trained on real-world natural videos (in-the-wild), and explored the impact of different information regularization strategies (sparsity, noise injection, quantization) on model performance.

Learning Locally, Revising Globally: Global Reviser for Federated Learning with Noisy Labels

Yuxin Tian (Sichuan University), Jiancheng Lv (Sichuan University)

ClassificationOptimizationFederated LearningSafty and PrivacyKnowledge DistillationData-Centric LearningSupervised Fine-TuningContrastive LearningImageTabular

🎯 What it does: Propose a new federated learning method called FedGR, which utilizes the slow memory characteristics of the global model under noisy labels to actively correct noisy labels and enhance model robustness through three modules.

Learning Long Range Spatio-Temporal Representations over Continuous Time Dynamic Graphs with State Space Models

Ayushman Raghuvanshi (Indian Institute of Science), Mahesh Chandran (Fujitsu Research India)

Representation LearningGraph Neural NetworkTransformerGraphTime SeriesSequential

🎯 What it does: Propose a state-space model based on continuous-time dynamic graphs (CTDG) (CTDG-SSM), which utilizes topology-aware HiPPO (CTT-HiPPO) together with graph Laplacian polynomial filters to achieve efficient memory and update of long-term and multi-hop spatial information.

Learning Manifold and Itô Dynamics with Branched Neural Rough Differential Equations

Luke Thompson (University of Sydney), Andi Han (University of Sydney)

OptimizationComputational EfficiencyRepresentation LearningContrastive LearningTabularTime SeriesSequentialPhysics RelatedStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Propose Branched Neural Rough Differential Equations (B-NRDE), extending NRDE to Itô integration and manifold dynamics through the Hopf algebra framework, achieving more accurate continuous-time modeling.

Learning Manifold Data with Flow Matching

Sophia Pi (Northwestern University), Han Liu (Northwestern University)

GenerationData SynthesisTransformerScore-based ModelFlow-based ModelStochastic Differential Equation

🎯 What it does: Theoretical analysis of flow matching generative models under the low-dimensional subspace assumption, proposing velocity decomposition and providing sample complexity and distribution convergence rate

Learning Molecular Semantic Invariant Representation with Prototype Constraint

Zhiqiang Li (Shanxi University), Jiye Liang (Shanxi University)

Domain AdaptationRepresentation LearningAdversarial AttackDrug DiscoveryGraph Neural NetworkMixture of ExpertsAuto EncoderContrastive LearningGraphBiomedical Data

🎯 What it does: Proposes a framework called MoSIR, which decomposes molecular embeddings into semantic-invariant components and environment-related residuals by utilizing a learnable prototype dictionary, thereby enhancing the generalization ability of molecular property prediction in out-of-distribution (OOD) scenarios.

Learning More from Less: Unlocking Internal Representations for Benchmark Compression

Yueqi Zhang (Beijing Institute of Technology), Kan Li (Beijing Institute of Technology)

CompressionRepresentation LearningTransformerLarge Language ModelContrastive LearningTextMultimodalityBenchmark

🎯 What it does: A framework called REPCORE is studied, which utilizes the internal hidden states of large language models for benchmark compression.

Learning Multi-Agent Coordination via Sheaf-ADMM

Jeffrey Seely (Sakana AI), Llion Jones (Sakana AI)

OptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackGraph Neural NetworkImageGraphTabular

🎯 What it does: Propose a differentiable multi-agent coordination framework called Sheaf-ADMM, which enables agents with local views to collaboratively solve global problems on low-dimensional projections through cellular sheaf constraints and ADMM iterations.

Learning Multi-Scale Hypergraph for High-Order Brain Connectivity Analysis

Jaeyoon Sim (Pohang University of Science and Technology), Won Hwa Kim (Pohang University of Science and Technology)

ClassificationExplainability and InterpretabilityRepresentation LearningGraph Neural NetworkTransformerSupervised Fine-TuningContrastive LearningGraphBiomedical DataAlzheimer's DiseaseBenchmark

🎯 What it does: Proposed the MuHL framework, which utilizes learnable multi-scale graph wavelet transforms and dynamically constructed hyperedges to capture high-order interactions in brain networks, thereby achieving disease staging and classification for neurodegenerative diseases.

Learning Multi-Timescale Abstractions for Hierarchical Combinatorial Planning

Vivienne Huiling Wang (Aalto University), Joni Pajarinen (Aalto University)

OptimizationGraph Neural NetworkReinforcement LearningWorld ModelTabularTime SeriesSequentialBenchmark

🎯 What it does: Propose a model-based hierarchical reinforcement learning framework called LMTA, which utilizes a multi-timescale SMDP world model and potential space tree search planning, combined with a subgoal-conditioned budget strategy, to address the problem of sequentialized stochastic combinatorial optimization (SSCO).

Learning Normalized Energy Models for Linear Inverse Problems

Nicolas Zilberstein (Rice University), Florentin Guth (Flatiron Institute)

RestorationConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelImage

🎯 What it does: Propose a method that can learn normalized energy models (Normalized Energy Models) to solve linear inverse problems such as image deblurring and inpainting, and can be applied to various denoising conditions without additional training.

Learning Partial Concept Classes and Universal Rates Under Massart Noise

Ariel Avital (BenGurion University), Steve Hanneke (Purdue University)

Classification

🎯 What it does: The study addresses the binary classification problem under Massart noise conditions, extending the theoretical framework of unified learning rates and partial concept classes.

Learning Permutation Distributions via Reflected Diffusion on Ranks

Sizhuang He (Yale University), David van Dijk (Yale University)

OptimizationTransformerDiffusion modelScore-based ModelImageSequentialBenchmark

🎯 What it does: This paper proposes a new discrete diffusion framework called Soft-Rank Diffusion, which utilizes a continuous latent space of soft ranks for reflective diffusion and employs a contextualized generalized Plackett-Luce (cGPL) denoiser during the reverse process to learn permutation distributions.

Learning Permutation from Structure Without Supervision

Ran Eisenberg (Bar-Ilan University), Ofir Lindenbaum (Bar-Ilan University)

OptimizationContrastive LearningImageTabular

🎯 What it does: This paper proposes an unsupervised learning method to discover hidden permutations by directly optimizing the reordered data using structural loss, without requiring true permutation labels.

Learning Permutation-invariant Macroscopic Dynamics

Zhichao Han (National University of Singapore), Qianxiao Li (National University of Singapore)

Graph Neural NetworkFlow-based ModelAuto EncoderVideoPoint CloudTabularPhysics RelatedStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Study the macroscopic dynamics of disordered microscopic systems, construct a distribution-aware autoencoder to learn permutation-invariant closed variables, and use these to predict macroscopic evolution.

Learning Protein Structure-Function Relationships through Knowledge-guided Representation Decomposition

Mingqing Wang (Tsinghua University), Zhixiang Ren (Pengcheng Laboratory)

Explainability and InterpretabilityRepresentation LearningProtein Structure PredictionTransformerAuto EncoderContrastive LearningBiomedical Data

🎯 What it does: Propose the ProtDiS framework, which utilizes knowledge-guided representation decomposition to split pre-trained protein microenvironment embeddings into independent channels aligned with biophysical attributes, thereby improving the interpretability and predictive performance of structure-function relationships.

Learning Query-Aware Budget-Tier Routing for Runtime Agent Memory

Haozhen Zhang (Nanyang Technological University), Wenya Wang (Nanyang Technological University)

Computational EfficiencyTransformerLarge Language ModelReinforcement LearningAgentic AITextRetrieval-Augmented Generation

🎯 What it does: Propose BudgetMem, a modular framework that decomposes runtime memory extraction into adjustable budget levels, and learns a lightweight router to make decisions across different budget levels, achieving controllable balance between memory extraction cost and quality in LLM agents.

Learning Randomized Reductions

Ferhat Erata (Yale University), Ruzica Piskac (Yale University)

OptimizationData-Centric LearningAI Code AssistantLarge Language ModelSupervised Fine-TuningAgentic AIPrompt EngineeringTextTabularBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposes a framework for automatically learning random self-reductions (RSR) from known programs and implements a system called Bitween.

Learning Rate Annealing Improves Tuning Robustness in Stochastic Optimization

Amit Attia (Tel Aviv University), Tomer Koren (Tel Aviv University)

OptimizationHyperparameter SearchImageTabular

🎯 What it does: Studied the impact of using learning rate annealing scheduling in random optimization on the robustness of hyperparameter tuning, proposed theoretical proofs and provided upper bounds in convex optimization scenarios, and then verified the robustness of annealing scheduling during coarse grid search on two tasks: synthetic logistic regression and CIFAR‑10.

Learning Rate Scaling across LoRA Ranks and Transfer to Full Finetuning

Nan Chen (Johns Hopkins University), Soufiane Hayou (Johns Hopkins University)

OptimizationFederated LearningRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningDiffusion modelContrastive LearningImageTextMultimodality

🎯 What it does: Propose the Maximal‑Update Adaptation (µA) theoretical framework to analyze the scaling relationship between learning rate, model width, and adapter rank during LoRA fine-tuning, and provide two scaling patterns.

Learning Reward–Cost Balance in Safe RL via Score-Based World Models

Yuetian Wang (Shanghai Jiao Tong University), Chunping Qiu (Intelligent Game and Decision Lab)

Autonomous DrivingKnowledge DistillationRecurrent Neural NetworkTransformerReinforcement LearningScore-based ModelContrastive LearningWorld ModelImageVideo

🎯 What it does: Propose the USB-RL framework, which learns and utilizes unsupervised reward-cost balance scores in model-based safe reinforcement learning to improve the trade-off between safety and performance.

Learning Rewrite-Invariant Reasoning with Targeted Alternation Training

Mousa Arraf (Technion Israel Institute of Technology), Kira Radinsky

Explainability and InterpretabilityComputational EfficiencyData-Centric LearningGraph Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextChain-of-Thought

🎯 What it does: By sampling and aggregating multiple reasoning trajectories of large language models on semantically preserving rewriting tasks, we construct a reasoning graph for each problem; using this graph, we identify the 'Solution Boundary Cut' (SBC), which marks the transition from recoverable to unrecoverable states, and generate a small number of targeted positive and negative example pairs for model fine-tuning or in-context learning, thereby improving the model's reasoning robustness under semantically preserving rewriting.

Learning Robust Multi-Agent Policies via Selective Adversarial Fault Induction

David Henry Mguni, Yaodong Yang (Peking University)

Recurrent Neural NetworkTransformerReinforcement LearningGenerative Adversarial NetworkContrastive LearningTabularTime SeriesBenchmark

🎯 What it does: Propose a pluggable multi-agent reinforcement learning robust framework called MARTA, which enhances the system's fault tolerance by selectively inducing failures in agents during critical states through the Switcher-Adversary mechanism during training.

Learning Self-Correction in Vision–Language Models via Rollout Augmentation

Yi Ding (Purdue University), Ruqi Zhang (Purdue University)

Explainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerSupervised Fine-TuningReinforcement LearningPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmarkChain-of-Thought

🎯 What it does: Proposed the Octopus framework, which constructs dense self-correction samples by self-correcting pairing and recombination of existing RL rollouts, and trains controllable self-correction capabilities on Vision-Language Models.

Learning Self-Interpretation from Interpretability Artifacts: Training Lightweight Adapters on Vector-Label Pairs

Keenan Pepper (AE Studio), Diogo S de Lucena

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelSupervised Fine-TuningAuto EncoderContrastive LearningText

🎯 What it does: Train a lightweight adapter to enable frozen language models to generate natural language explanations of their internal states through inserted vectors.

Learning Sparse Visual Representations via Spatial-Semantic Factorization

Theodore Zhao, Mu Wei (Microsoft)

ClassificationSegmentationCompressionRepresentation LearningTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningOptical FlowImage

🎯 What it does: STELLAR learns a sparse low-rank representation by decomposing image representations into a small number of semantic concepts (What) and their spatial distribution (Where), achieving high-quality image reconstruction while obtaining strong semantic representations.

Learning Stochastic Bridges for Video Object Removal via Video-to-Video Translation

Zijie Lou (Meitu Inc.), Ting Liu (Meitu Inc.)

Image TranslationRestorationGenerationTransformerDiffusion modelScore-based ModelAuto EncoderVideoBenchmarkStochastic Differential Equation

🎯 What it does: This study proposes a video object removal method based on a stochastic bridge model, treating the task as video-to-video translation.

Learning Structured Reasoning via Tractable Trajectory Control

Po-Nien Kung (University of California Los Angeles), Kai-Wei Chang (University of California Los Angeles)

OptimizationExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringVision Language ModelTextMultimodalityChain-of-Thought

🎯 What it does: Propose the Ctrl-R framework, achieving structured reasoning learning for language models through computable trajectory control;

Learning syntax without semantics: Disentangled tiny language models

Ezra Winston (Carnegie Melon University), J Zico Kolter

Computational EfficiencyKnowledge DistillationRepresentation LearningData-Centric LearningTransformerLarge Language ModelPrompt EngineeringText

🎯 What it does: By training extremely small language models with constrained paraphrasing (SAMBAL) on text, they learn syntactic structures while suppressing the influence of semantics and world knowledge.

Learning Task-Sufficient World Models by Synergizing Agentic Exploration and Structured Modeling

Fan Feng (University of California San Diego), Kun Zhang (Mohamed bin Zayed University of Artificial Intelligence)

Robotic IntelligenceTransformerReinforcement LearningAgentic AIContrastive LearningWorld ModelTabularTime SeriesSequentialBenchmark

🎯 What it does: Propose an MIST-WM framework that obtains task-specific, minimal, and sufficient latent representations through agent active exploration and structured world model learning.

Learning Taxonomic Trees with Hierarchical Representation Regularization for Large Multimodal Models

Hulingxiao He (Peking University), Yuxin Peng (Peking University)

RecognitionRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningVision Language ModelContrastive LearningImageTextMultimodality

🎯 What it does: Propose HiR 2, a parameter-free hierarchical representation regularization method, which utilizes non-parametric cross-attention to construct semantic visual trees from intermediate layers of LMMs, and enhances hierarchical visual recognition (HVR) consistency through hyperbolic implication loss and spherical angular dispersion loss.

Learning the Best Under Constraints: A Duality-Based Framework

Mingjie Hu (Fudan University), Jianqiang Hu (Fudan University)

OptimizationReinforcement Learning from Human FeedbackContrastive LearningTabular

🎯 What it does: Studied the problem of actively selecting covariates in the fixed confidence setting for the constrained linear best arm identification problem, proposed instance-dependent lower bounds, relaxed lower bounds, and their dual forms, and designed an efficient dual decomposition algorithm.

Learning the ESG Geometry with Domain Aware Language Models

Kunal Pradeep Pimparkhede (Indian Institute of Technology), Mahesh Mohan M R (Indian Institute of Technology)

OptimizationTransformerLarge Language ModelSupervised Fine-TuningTextMultimodalityTime SeriesFinance RelatedRetrieval-Augmented Generation

🎯 What it does: Propose a domain-aware LLM framework that can simultaneously predict multi-modal time series including ESG risk, financial returns, and news sentiment, and achieve asset selection through trajectory-level representations.

Learning the Interaction Prior for Protein-Protein Interaction Prediction: A Model-Agnostic Approach

Ziqi Gao (Tsinghua University), Jia Li (Hong Kong University of Science and Technology Guangzhou)

Drug DiscoveryProtein Structure PredictionGraph Neural NetworkTransformerPrompt EngineeringContrastive LearningGraphBiomedical Data

🎯 What it does: In protein interaction prediction, a graph prompting learning framework called L3-PPI based on the L3 path is proposed as a model-agnostic classification head.

Learning the Minimum Action Distance

Lorenzo Steccanella (Universitat Pompeu Fabra), Anders Jonsson (Universitat Pompeu Fabra)

Representation LearningRecurrent Neural NetworkTransformerReinforcement LearningContrastive LearningSequentialBenchmark

🎯 What it does: This paper proposes an offline state representation learning framework that learns the Minimum Action Distance (MAD) using only state trajectory data (without requiring reward or action information) and uses this distance for goal-directed reinforcement learning and reward shaping.

Learning the Neighborhood: Contrast-Free Multimodal Self-Supervised Molecular Graph Pretraining

Boshra Ariguib (University of Stuttgart), Andrei Manolache (University of Stuttgart)

Drug DiscoveryGraph Neural NetworkTransformerSupervised Fine-TuningAuto EncoderContrastive LearningMultimodalityGraph

🎯 What it does: This paper proposes the C‑FREE (Contrast‑Free Representation Learning on Ego‑nets) framework, which pretrains molecular graph neural networks using a contrast-free, self-supervised subgraph prediction approach under a multimodal setting (2D graph + 3D conformation), and then fine-tunes them on downstream tasks.

Learning Tight Rejection Boundaries without Negatives for Strict One-Class Audio Deepfake Detection

Yuze Zhao (Harbin Institute of Technology), Wei Jiang (Harbin Institute of Technology)

Anomaly DetectionTransformerAuto EncoderContrastive LearningAudio

🎯 What it does: Propose a strictly one-class learning audio deepfake detection method called CA-SOADD, which utilizes the distribution shift view without negative samples as a boundary detector, and achieves a clear delineation of the compact distribution of real audio and the rejection boundary through triple-objective central anchor point learning.

Learning to Approximate Uniform Facility Location via Graph Neural Networks

Chendi Qian (RWTHAachen University), Christian Sohler (University of Cologne)

OptimizationGraph Neural NetworkReinforcement LearningContrastive LearningGraph

🎯 What it does: A fully differentiable message passing graph neural network (MPNN) is proposed to approximate the unified facility location (UniFL) problem, and the network is trained using unsupervised learning to obtain good solutions.

Learning to Bet for Horizon-Aware Anytime-Valid Testing

Ege Onur Taga (University of Michigan), Shubhanshu Shekhar (University of Michigan)

OptimizationRecurrent Neural NetworkTransformerReinforcement LearningTabularTime SeriesBiomedical DataFinance Related

🎯 What it does: Propose a horizon-aware validation and confidence sequence that is effective at any time under a hard deadline N, and view it as a finite-horizon optimal control problem through the betting/e-process framework.

Learning to Correct: Reinforcement Learning for Multi-Attempt Chain-of-Thought

Muhammed Emrullah Ildiz, Samet Oymak (University of Michigan)

Reinforcement LearningTextChain-of-Thought

🎯 What it does: This paper studies the use of reinforcement learning in multi-attempt chain-of-thought (Multi‑Attempt CoT) to maximize the success probability of Verification@K.

Learning to Decode Against Compositional Hallucination in Video Multimodal Large Language Models

Wenbin Xing (Sun Yat-sen University), Junchi Yan (Shanghai Jiao Tong University)

ClassificationAnomaly DetectionExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelReinforcement LearningPrompt EngineeringVision Language ModelContrastive LearningOptical FlowVideoTextMultimodalityBenchmark

🎯 What it does: Proposed the OmniVCHall benchmark and the TriCD three-path contrastive decoding framework for detecting and mitigating compositional hallucinations in video multimodal large language models;

Learning to Discover at Test Time

Mert Yuksekgonul (Stanford University), Yu Sun (Stanford University)

OptimizationDrug DiscoveryAI Code AssistantNeural Architecture SearchProtein Structure PredictionTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextTabularTime SeriesBiomedical Data

🎯 What it does: Utilizing test-time reinforcement learning to continuously train large language models in order to discover novel optimal solutions for single scientific problems.

Learning to Emulate Chaos: Adversarial Optimal Transport Regularization

Gabriel Melo (Institut Polytechnique de Paris), Peter Y. Lu (Tufts University)

Data SynthesisOptimizationExplainability and InterpretabilityConvolutional Neural NetworkRecurrent Neural NetworkTransformerDiffusion modelScore-based ModelGenerative Adversarial NetworkContrastive LearningTime SeriesSequentialPhysics Related

🎯 What it does: Using adversarial optimal transport regularization, jointly learn interpretable summary statistics and high-quality simulators, approximating the statistical properties of chaotic systems with only a single noisy trajectory.

Learning to Evict from Key-Value Cache

Luca Moschella (Apple), Ozan Sener (Apple)

OptimizationComputational EfficiencyTransformerLarge Language ModelReinforcement LearningTextSequential

🎯 What it does: Designed and trained an offline RL framework called KVP to predict the future usage value of tokens in the KV cache, thereby achieving efficient cache eviction.

Learning to Execute Graph Algorithms Exactly with Graph Neural Networks

Muhammad Fetrat Qharabagh (University of Waterloo), Kimon Fountoulakis (University of Waterloo)

OptimizationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningGraph Neural NetworkPrompt EngineeringMixture of ExpertsContrastive LearningGraphTabular

🎯 What it does: This paper proposes a graph neural network (GNN) framework based on the integration of multi-layer perceptrons (MLPs). It first trains the MLP on a binary instruction set without a graph to achieve local node updates, and then during the inference phase, embeds the trained local updates into the GNN to enable precise execution of graph algorithms. The neural tangent kernel (NTK) theory is used to prove that under the conditions of infinite width, finite degree, and finite precision, the framework can achieve high probability execution of any LOCAL model graph algorithm (such as message flooding, BFS, DFS, Bellman-Ford).

Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization

Xingyuan Hua (Tsinghua University), Ju Ren (Tsinghua University)

OptimizationTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AITextMultimodality

🎯 What it does: Propose an exploration-aware reinforcement learning framework, EAPO, tailored for large language model agents, enabling agents to proactively explore the environment and record memories when needed, thereby improving decision quality.