arXivSub Start free trial

ICML 2026 Papers — Page 3

International Conference on Machine Learning · 6554 papers

Active Tabular Augmentation via Policy-Guided Diffusion Inpainting

Zheyu Zhang (Technical University Of Munich), Gjergji Kasneci (Technical University Of Munich)

GenerationData SynthesisData-Centric LearningTransformerReinforcement LearningDiffusion modelTabularElectronic Health Records

🎯 What it does: Propose a strategy-guided diffusion inpainting (TAP) framework to achieve active augmentation of tabular data, which can dynamically decide what kind of samples to generate and when to inject them during the learning process.

Active Timepoint Selection for Learning Measure-Valued Trajectories

Nicolas Huynh (University of Cambridge), Mihaela van der Schaar (University of Cambridge)

OptimizationComputational EfficiencyRepresentation LearningData-Centric LearningDiffusion modelScore-based ModelGaussian SplattingTabularTime SeriesBiomedical DataBenchmarkStochastic Differential Equation

🎯 What it does: An active learning framework is proposed to select time points that most effectively reduce trajectory uncertainty under a limited budget when measuring sparse distribution trajectories (probability paths);

ACTIVE-o3 : Empowering MLLMs with Active Perception via Pure Reinforcement Learning

Muzhi Zhu (Zhejiang University), Chunhua Shen (Ant Group)

Autonomous DrivingOptimizationComputational EfficiencyRobotic IntelligenceTransformerLarge Language ModelReinforcement LearningPrompt EngineeringVision Language ModelVision-Language-Action ModelImageMultimodality

🎯 What it does: Proposed the ACTIVE-o3 framework, which utilizes pure reinforcement learning to enable multi-modal large language models to have active perception capabilities, allowing them to autonomously generate meaningful region proposals and perform tasks.

ActiveScope: Actively Seeking and Correcting Perception for MLLMs

Yajing Wang (University of Chinese Academy of Sciences), Qingming Huang (University of Chinese Academy of Sciences)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBenchmarkChain-of-Thought

🎯 What it does: Propose the training-free ActiveScope framework, which enhances the performance of multi-modal large language models in high-resolution fine-grained visual understanding through active localization and self-correction.

ActiveUltraFeedback: Efficient Preference Data Generation using Active Learning

Davit Melikidze (ETH Zurich), Andreas Krause (ETH Zurich)

Computational EfficiencyData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringText

🎯 What it does: Proposes ACTIVEULTRAFEEDBACK, a modular active learning pipeline designed to efficiently generate preference data required for aligned models.

AD-BTS: Adaptive Dual-Branch Token Sparsification via Spatial Information Density

Xinpei Gao (University of Science and Technology of China), S Kevin Zhou (University of Science and Technology of China)

Computational EfficiencyRepresentation LearningTransformerLarge Language ModelVision Language ModelAuto EncoderContrastive LearningImageVideoTextMultimodality

🎯 What it does: Propose an adaptive dual-branch token sparsification framework for efficient inference in high-resolution vision encoders of multi-modal large language models.

AD-MIR: Bridging the Gap from Perception to Persuasion in Advertising Video Understanding via Structured Reasoning

Binxiao Xu (Peking University), Wentao Zhang (Peking University)

Recommendation SystemExplainability and InterpretabilityTransformerPrompt EngineeringVision Language ModelVideoTextMultimodalityRetrieval-Augmented GenerationChain-of-ThoughtAudio

🎯 What it does: Construct the AD-MIR framework, achieving advertising video intent understanding through structured memory and reasoning.

AdaEraser: Training-Free Object Removal via Adaptive Attention Suppression

Dingming Liu (Peking University)

Image TranslationImage HarmonizationRestorationTransformerPrompt EngineeringDiffusion modelImage

🎯 What it does: Propose an object removal method called AdaEraser that requires no training, utilizing the self-attention dynamic suppression of pre-trained diffusion models to achieve object erasure and generate reasonable backgrounds in blank areas.

AdaGC: Enhancing LLM Pretraining Stability via Adaptive Gradient Clipping

Guoxia Wang (Baidu Inc.), Li Shen (Sun Yat-sen University)

OptimizationTransformerLarge Language ModelText

🎯 What it does: Propose AdaGC, an adaptive tensor-level gradient clipping method, to eliminate loss spikes during the pre-training of large-scale language models.

AdaHC: Accelerating Multi-Token Prediction with Adaptive Head Chunking with Pipeline Parallelism

Yan Wang (University of Chinese Academy of Sciences), Weile Jia (University of Chinese Academy of Sciences)

Computational EfficiencyTransformerLarge Language ModelText

🎯 What it does: Propose the AdaHC framework, which utilizes adaptive head chunking and pipeline parallelism to accelerate the multi-token prediction (MTP) module in LLM training.

Adalina: Adaptive Linear Approximation for the Shapley Value and Beyond

Weida Li (National University of Singapore), Bryan Kian Hsiang Low (National University of Singapore)

Explainability and InterpretabilityComputational EfficiencyImageTabular

🎯 What it does: Proposed an adaptive linear approximation algorithm called Adalina, which achieves linear time and linear space approximation for half-values (including Shapley values, Banzhaf values, etc.) under the Θ(n) space constraint;

AdaMEM: Test-Time Adaptive Memory for Language Agents

Yunxiang Zhang (University of Michigan), Lu Wang (University of Michigan)

Autonomous DrivingRobotic IntelligenceReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringTextSequentialRetrieval-Augmented Generation

🎯 What it does: This paper proposes the ADAMEM framework, which enables language agents to achieve adaptive behaviors with parameter-free updates during reasoning by combining long-term trajectories stored offline with dynamically generated short-term strategies.

AdaMeZO: Adam-style Zeroth-Order Optimizer for LLM Fine-tuning Without Maintaining the Moments

Zhijie Cai (Shenzhen International Center for Industrial and Applied Mathematics), Guangxu Zhu (Shenzhen International Center for Industrial and Applied Mathematics)

OptimizationFederated LearningComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningScore-based ModelContrastive LearningText

🎯 What it does: Propose AdaMeZO, a zeroth-order optimizer that does not require storing first- or second-order moments, and can use Adam-style preconditioned updates during LLM fine-tuning.

AdamO: A Collapse-Suppressed Optimizer for Offline RL

Nan Qiao (Tsinghua University), Ju Ren (Tsinghua University)

OptimizationReinforcement LearningTabularTime SeriesBenchmark

🎯 What it does: Proposes AdamO, an Adam optimizer that suppresses TD error collapse in offline reinforcement learning through decoupled orthogonality correction.

AdaNav: Adaptive Reasoning with Uncertainty for Vision-Language Navigation

Xin Ding (University of Science and Technology of China), Ting Cao (Institute for AI Industry Research, Tsinghua University)

Autonomous DrivingExplainability and InterpretabilityComputational EfficiencyTransformerReinforcement LearningVision Language ModelVision-Language-Action ModelImageTextMultimodality

🎯 What it does: Propose AdaNav, an adaptive reasoning framework based on uncertainty, integrating a lightweight UAR Block and Heuristic-to-RL training, enabling visual-language navigation agents to automatically determine when and how to reason under limited samples.

Adapting Noise to Data: Generative Flows from Learned 1D Processes

Jannis Chemseddine (TU Berlin), Gabriele Steidl (TU Berlin)

GenerationData SynthesisFlow-based ModelImageTabularTime SeriesOrdinary Differential Equation

🎯 What it does: This paper proposes constructing a data-adaptive noise distribution by learning a one-dimensional quantile function, and applies it to flow matching models to address the challenges faced by traditional Gaussian-based noise in learning heavy-tailed or multi-modal distributions.

Adapting to Evolving Graphs: A Scalable Framework for Dynamic Coarsening

Abhishek Gupta (IIT Delhi), Sandeep Kumar (IIT Delhi)

OptimizationComputational EfficiencyRepresentation LearningGraph Neural NetworkContrastive LearningGraphTime Series

🎯 What it does: Proposes a unified framework for incrementally sparsifying discrete-time evolving graphs, avoiding the need to recompute the mapping matrix from scratch each time an update occurs.

Adaptive Bandit Algorithms for Contextual Matching Markets

Shiyun Lin (Peking University), Nadav Merlis (Technion Israel Institute of Technology)

OptimizationReinforcement LearningTabular

🎯 What it does: Studied the contextualized linear multi-armed bandit problem in matching markets, and proposed an adaptive learning algorithm for both stochastic and adversarial contexts;

Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent

Hiroki Naganuma (Mila), Hao-Jun Michael Shi (Meta Platforms)

OptimizationImageText

🎯 What it does: Propose an adaptive batch size method based on gradient noise scale (GNS) for non-Euclidean optimizers such as signSGD/Signum and specSGD/Muon.

Adaptive Code Watermarking Through Reinforcement Learning

Zhimeng Guo (Penn State University), Minhao Cheng (Penn State University)

AI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringText

🎯 What it does: Proposes CodeTracer, an adaptive framework for dynamically embedding watermarks into code generated by large language models;

Adaptive Coding Emerges in Stabilized Supralinear Networks Trained with Local Plasticity

Haoyu Albert Wang (Fudan University), Yuguo Yu (Fudan University)

ClassificationComputational EfficiencyRepresentation LearningAuto EncoderContrastive LearningImage

🎯 What it does: Train stable superlinear networks (SSN) using local plasticity rules to learn natural image features and spontaneously switch from population coding to sparse coding strategies under different contrast and noise conditions.

Adaptive Contracts for Cost-Effective AI Delegation

Eden Saig (Technion Israel Institute Of Technology), Jamie Tucker-Foltz (Yale)

OptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackLarge Language ModelPrompt EngineeringContrastive LearningTextTabularBenchmark

🎯 What it does: This paper studies the introduction of adaptive contracts in AI delegation scenarios, allowing the choice of whether to pay for more expensive fine evaluations after an initial rough assessment, in order to reduce costs and improve utility.

Adaptive DNA Sequence Modeling via Synergistic Plasticity Units

Binghao Liu (DAMO Academy, Alibaba Group), Fei Gu (DAMO Academy, Alibaba Group)

Drug DiscoveryProtein Structure PredictionConvolutional Neural NetworkTransformerSupervised Fine-TuningMixture of ExpertsContrastive LearningBiomedical DataBenchmark

🎯 What it does: This paper proposes an expandable Synergistic Plasticity Unit (SPU) for DNA sequence modeling. The SPU integrates local motifs, global dependencies, and frequency-domain periodic signals through multi-layer plasticity mechanisms, constructing an efficient DNA foundation model (SPU-DNA).

Adaptive Estimation and Inference in Semi-parametric Heterogeneous Clustered Multitask Learning via Neyman Orthogonality

Hanxiao Chen (Boston University), Debarghya Mukherjee (Boston University)

OptimizationFederated LearningExplainability and InterpretabilityMeta LearningMixture of ExpertsContrastive LearningTabularTime SeriesBenchmarkFinance Related

🎯 What it does: Propose a semi-parametric heterogeneous multi-task learning framework, utilizing Neyman orthogonal loss and adaptive paired fusion penalty. First, task-local initial estimates are used to quantify task similarity, and then precise estimation of the target parameters is achieved through an aggregation phase that enables clustering recovery.

Adaptive Generation of Bias-Eliciting Questions for LLMs

Robin Staab (ETH Zurich), Martin Vechev (ETH Zurich)

GenerationData SynthesisExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose a framework for automatically generating realistic open-ended questions based on contrastive variations to trigger biased behaviors in LLMs;

Adaptive Memory Retention in Dynamic Graphs

Fabrizio De Castelli (University of Pisa), Davide Bacciu (University of Pisa)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningRecurrent Neural NetworkGraph Neural NetworkTransformerContrastive LearningGraphTime SeriesSequentialStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Proposed the LAMP model based on the differential equations of neural impulses for long-range memory retention in dynamic graphs.

Adaptive Momentum and Nonlinear Damping for Neural Network Training

Aikaterini Karoni (University of Bristol), Gabriel Stoltz (Institut Polytechnique de Paris)

OptimizationHyperparameter SearchTransformerContrastive LearningImageTextStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Proposed an adaptive momentum and nonlinear damping optimizer based on continuous-time dynamics, mainly including Individual Kinetic Friction Adaptive Descent (iKFAD), Cubically Damped mSGD (CD), and Cubically Damped Adam (CADAM).

Adaptive Multi-Round Allocation with Stochastic Arrivals

Yuqi Pan (Harvard University), Cheryl Johnson (World Health Organization)

OptimizationReinforcement LearningTabularBiomedical Data

🎯 What it does: This paper studies a multi-round resource allocation problem with a limited budget for adaptive network recruitment, and proposes a dynamic programming method based on greedy allocation and population hierarchical value functions.

Adaptive Multiscale Binary Expansion Tests for Independence

Yang Yang (University of Illinois Chicago), Ping-Shou Zhong (University of Illinois Chicago)

Anomaly DetectionExplainability and InterpretabilityComputational EfficiencyContrastive LearningTabularBiomedical DataBenchmark

🎯 What it does: This paper proposes a novel multi-scale independence test method based on binary expansion (CoBET, dCoBET, wa-dCoBET), which can test the independence between multivariate random variables without relying on kernel functions.

Adaptive Node Feature Selection for Graph Neural Networks

Madeline Navarro (Rice University), Santiago Segarra (Rice University)

ClassificationGraph Neural NetworkSupervised Fine-TuningContrastive LearningGraph

🎯 What it does: Proposes an adaptive node feature selection method that evaluates feature importance through feature permutation and dynamically removes unnecessary features during training.

Adaptive Personalized Federated Learning via Multi-task Averaging of Kernel Mean Embeddings

Jean-Baptiste Fermanian (University of Montpellier), Aurélien Bellet (University of Montpellier)

Domain AdaptationOptimizationFederated LearningComputational EfficiencyImageTabular

🎯 What it does: Propose an adaptive personalized federated learning method that improves the model performance at the target end by learning the collaboration weights of each participant and minimizing the weighted empirical risk;

Adaptive Physics Transformer with Fused Global-Local Attention for Subsurface Energy Systems

Xin Ju (Stanford University), Gege Wen (Imperial College London)

OptimizationComputational EfficiencyGraph Neural NetworkTransformerSupervised Fine-TuningDiffusion modelAuto EncoderContrastive LearningGraphTabularTime SeriesPhysics Related

🎯 What it does: Propose Adaptive Physics Transformer (APT), a neural operator that can be directly trained on adaptive mesh data while simultaneously integrating global Perceiver and local GNO, for solving numerical simulation problems in multi-scale subsurface energy systems.

Adaptive Policy Backbone via Shared Network

Bumgeun Park (Korea Advanced Institute of Science and Technology), Donghwan Lee (Korea Advanced Institute of Science and Technology)

Domain AdaptationMeta LearningSupervised Fine-TuningReinforcement LearningTabularTime Series

🎯 What it does: Studied an adaptation strategy (APB) that freezes the shared backbone network and only fine-tunes the front and back linear layers for task adaptation and OOD generalization in reinforcement learning.

Adaptive Preconditioners Trigger Loss Spikes in Adam

Zhiwei Bai (Shanghai Jiao Tong University), Zhi-Qin John Xu (Shanghai Jiao Tong University)

OptimizationConvolutional Neural NetworkRecurrent Neural NetworkGraph Neural NetworkTransformerText

🎯 What it does: This paper studies the 'loss spike' phenomenon that occurs in the Adam optimizer and explains it from the perspective of internal mechanisms.

Adaptive Probe-based Steering for Robust LLM Jailbreaking

Junxi Chen (Sun Yat Sen University), Xiaohua Xie (Sun Yat Sen University)

Explainability and InterpretabilityComputational EfficiencyAdversarial AttackTransformerLarge Language ModelPrompt EngineeringContrastive LearningText

🎯 What it does: This paper proposes an adaptive probe-based steering vector to break aligned LLMs, significantly improving the effectiveness of the attack without additional contrastive prompts or tedious manual parameter tuning.

Adaptive Protein Tokenization

Rohit Dilip (California Institute Of Technology), David Van Valen (California Institute Of Technology)

Protein Structure PredictionTransformerDiffusion modelScore-based ModelFlow-based ModelAuto EncoderContrastive LearningBiomedical Data

🎯 What it does: Designed and implemented a global adaptive protein tokenizer (APT), combining a diffusion decoder to achieve protein structure generation, reconstruction, and representation learning.

Adaptive Quasimetric Mapping : Principled Topological Abstraction for Robust Offline Goal-Conditioned Navigation

Anthony Kobanda (Inria), Rémy Portelas (Ubisoft La Forge)

Autonomous DrivingOptimizationComputational EfficiencyGraph Neural NetworkReinforcement LearningContrastive LearningGraphTabularTime Series

🎯 What it does: This paper studies a path planning framework called Adaptive Quasimetric Mapping (AQM) for offline goal-conditioned reinforcement learning. It achieves adaptive planning and zero-shot replanning by learning a time-to-arrival quasimetric and constructing a sparse keypoint graph.

Adaptive Querying with AI Persona Priors

Kaizheng Wang (Columbia University), assaf zeevi

Recommendation SystemExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelMixture of ExpertsText

🎯 What it does: Under a limited questioning budget, adaptive querying is achieved by leveraging AI role priors to learn user-specific quantities.

Adaptive Recurrent Message Passing for Test Time Computing on Graphs

Junshu Sun (Chinese Academy of Sciences), Shuhui Wang (Chinese Academy of Sciences)

ClassificationRecommendation SystemFederated LearningComputational EfficiencyRepresentation LearningRecurrent Neural NetworkGraph Neural NetworkContrastive LearningGraph

🎯 What it does: Propose an adjustable iterative recursive graph model AdaR, achieving adaptive reasoning in graph learning.

Adaptive Reinforcement Learning for Unobservable Random Delays

John Wikman (KTH Royal Institute of Technology), David Broman (KTH Royal Institute of Technology)

Recurrent Neural NetworkGraph Neural NetworkReinforcement LearningAuto EncoderWorld ModelTime SeriesSequential

🎯 What it does: In reinforcement learning environments with unobservable random delays, this paper proposes the Interaction Layer framework and implements a model-based Actor-Critic with Delay Adaptation (ACDA) algorithm, significantly improving the agent's learning performance under delay conditions.

Adaptive Residual-Update Steering for Low-Overhead Hallucination Mitigation in Large Vision-Language Models

Zhengtao Zou (Aalto University), Pekka Marttinen (Aalto University)

Explainability and InterpretabilityComputational EfficiencyTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityRetrieval-Augmented Generation

🎯 What it does: Proposes the RUDDER framework, which generates visual evidence direction (CARD) by extracting residual updates from the pre-fill phase during LVLM inference, and adaptively injects this direction into the decoding stage via a Beta Gate, reducing visual confusion;

Adaptive Sharpness-Aware Minimization with a Polyak-type Step size: A Theory-Grounded Scheduler

Dimitris Oikonomou (Johns Hopkins University), Nicolas Loizou (Johns Hopkins University)

OptimizationImage

🎯 What it does: Proposed an adaptive learning rate scheduler based on the Polyak step size, modifying Unnormalized Sharpness-Aware Minimization (USAM) and its regularized form SAM to achieve robust optimization without manual parameter tuning.

Adaptive Symmetry Discovery for Dynamical System Identification

Behrooz Tahmasebi (Harvard University), Melanie Weber (Harvard University)

OptimizationComputational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackContrastive LearningTime SeriesSequentialPhysics RelatedStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Propose a method that, under the premise of observing only a single trajectory, can both discover unknown discrete group symmetries and perform parameter identification for equivariant dynamical systems.

Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks

Peiyu Li (University of Notre Dame), Nitesh V Chawla (University of Notre Dame)

TransformerLarge Language ModelTextBenchmark

🎯 What it does: Propose ATLAS — an adaptive testing framework based on IRT, designed for efficiently evaluating the capabilities of large language models (LLMs).

Adaptive Time Series Reasoning via Segment Selection

Shvat Messica (Harvard Medical School), Marinka Zitnik (Harvard Medical School)

Explainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextTime SeriesBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Solve time series reasoning tasks through a controller-reasoner architecture that adaptively selects and reasons about time periods, allowing the model to actively retrieve relevant time periods and generate answers during inference.

Adaptive Token Refinement in Long-Tailed Large Vision-Language Models Fine-Tuning

Wenjun Miao (Beihang University), Zheng Wei (Tencent PCG)

ClassificationRecognitionRecommendation SystemComputational EfficiencyData-Centric LearningTransformerSupervised Fine-TuningPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodality

🎯 What it does: To address the fine-tuning of large vision-language models (LVLMs) under long-tailed distributions, this paper proposes the Adaptive Token Refinement (ATR) framework. It dynamically filters high-confidence tokens and reweights low-confidence tokens at the output end using a boundary adaptive loss (BAL). At the input end, it randomly masks low-entropy visual tokens via visual token masking (VTM), thereby suppressing overfitting on head classes and enhancing learning on tail classes.

Adaptive Utilization of Low-Rank Adaptation via Conditioned Gating

Guang Yang (Beijing Jiaotong University), Kaiyu Huang (Beijing Jiaotong University)

Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningText

🎯 What it does: Introduce an adaptive utilization mechanism into the LoRA PEFT framework, utilizing token-level conditional gating and sentence-level aggregation to dynamically allocate low-rank updates, and using EMA prior to stabilize training.

Adaptive Visual Autoregressive Acceleration via Dual-Linkage Entropy Analysis

Yu Zhang (Tongji University), Longbing Cao (Macquarie University)

GenerationComputational EfficiencyTransformerDiffusion modelImage

🎯 What it does: Proposes NOVA, a training-agnostic self-adaptive acceleration framework that dynamically prunes low-entropy tokens in visual autoregressive (VAR) models through entropy analysis, considering both scale and hierarchical levels;

Adaptive Volumetric Mechanical Property Fields Invariant to Resolution

Rishit Dagli (NVIDIA), Maria Shugrina

GenerationData SynthesisTransformerVision Language ModelDiffusion modelNeural Radiance FieldAuto EncoderContrastive LearningImagePoint CloudMeshPhysics Related

🎯 What it does: Predict dense spatially varying mechanical properties (Young's modulus, Poisson's ratio, density) for 3D assets and generate voxelized material fields that can be directly used for physics simulations.

Adaptively Grouped Contextual Bandits for Heterogeneous Human-AI Decision Making with Conformal Prediction Sets

Yanchen Wu (Tsinghua University), Bo Li (Tsinghua University)

OptimizationExplainability and InterpretabilityComputational EfficiencyReinforcement LearningContrastive LearningTextTabular

🎯 What it does: This paper addresses the problem of personalized selection of significance level α in human-machine collaboration, modeling it as a contextual bandits problem and proposing an Adaptive Grouping Contextual Bandits (AGCB) framework, which leverages the assumptions of continuity and monotonicity to achieve efficient learning.

Adaptively Robust Resettable Streaming

Edith Cohen (Google Research), Uri Stemmer (Tel Aviv University)

Federated LearningSafty and PrivacyAdversarial Attack

🎯 What it does: This paper proposes a streaming sketch that is robust against adaptive attacks in resettable streaming models, capable of accurately estimating cardinality, summation, and a class of sublinear Bernstein statistics within polynomially small (polylog) space.

AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally

Shaowen Wang (Tsinghua University), Jian Li (Tsinghua University)

RetrievalComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextBenchmark

🎯 What it does: Propose AdaRoPE, which learns rotation frequency and attention scaling for each attention head to improve the performance of Transformers on long contexts.

AdaS: Adaptive Gradient Descent for Spiking Transformers

Zijian Zhou (University of Electronic Science and Technology of China), Haizhou Li (Shenzhen Loop Area Institute)

OptimizationSpiking Neural NetworkTransformerImageVideoTextBiomedical Data

🎯 What it does: Proposed AdaS, an adaptive gradient descent optimizer specifically designed for Spiking Transformers, aiming to alleviate the excessive noise problem during the training process with surrogate gradients.

AdaSCALE: Adaptive Scaling for OOD Detection

Sudarshan Regmi (Dartmouth College)

ClassificationAnomaly DetectionExplainability and InterpretabilityComputational EfficiencyConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelFlow-based ModelRectified FlowAuto EncoderGenerative Adversarial NetworkContrastive LearningImage

🎯 What it does: Proposed a post-processing method called AdaSCALE, which dynamically adjusts the scaling threshold of activation or logit by estimating the offset of activation under minor perturbations for each sample, thereby significantly improving the OOD detection capability.

AdaSplash-2: Faster Differentiable Sparse Attention

Nuno Gonçalves, Marcos Vinicius Treviso

Computational EfficiencyTransformerLarge Language ModelTextBenchmark

🎯 What it does: Proposed a sparse attention mechanism called ADASPLASH-2 based on α-entmax, significantly accelerating the forward and backward computations of Transformers.

Addressing Instrument-Outcome Confounding in Mendelian Randomization through Representation Learning

Shimeng Huang (Institute of Science and Technology Austria), Francesco Locatello (Institute of Science and Technology Austria)

Representation LearningDrug DiscoveryMixture of ExpertsAuto EncoderContrastive LearningTabularBiomedical DataElectronic Health Records

🎯 What it does: Use multi-environment data to recover invariant latent components in genetic instrumental variables through representation learning methods, eliminating environmental confounding between instrumental variables and outcomes, thus enabling reliable estimation of causal effects.

Addressing Semantic Blind Spots in Text-to-SQL via Component Pre-generation and AST Matching Rewards

Xingyu Ma (School of Cyber Science and Engineering Huazhong University of Science and Technology), Jinqiao Wang (Institute of Automation Chinese Academy of Sciences)

AI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextTabular

🎯 What it does: Proposes pre-generating SQL components and using the maximum connected subtree matching reward based on the AST to alleviate the 'semantic blind spot' problem in the text-to-SQL task, where the model cannot utilize the semantics of the components to be generated during the generation process.

ADEPT: RL-Aligned Agentic Decoding of Emotion via Evidence Probing Tools — From Consensus Learning to Ambiguity-Driven Emotion Reasoning

Esther Sun (Carnegie Mellon University), Carlos Busso (Carnegie Mellon University)

RecognitionTransformerReinforcement LearningAgentic AIPrompt EngineeringRetrieval-Augmented GenerationAudio

🎯 What it does: Proposes the ADEPT framework, transforming speech emotion recognition into a multi-round active querying and evidence retrieval process, generating auditable emotional reasoning paths;

ADHD Disease Detection Based on Short- and Long-Term Brain Function Encoding and Memory Graph Network

Dongxun Jiang (Tongji University), Dongdong Zhang (Tongji University)

ClassificationAnomaly DetectionRecurrent Neural NetworkGraph Neural NetworkTransformerContrastive LearningGraphBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease

🎯 What it does: Proposed an ADHD detection model named SLT-BFGN that integrates short-term brain functional reorganization with long-term structural features;

AdLift: Lifting Adversarial Perturbations to Safeguard 3D Gaussian Splatting Assets Against Instruction-Driven Editing

Ziming Hong (University of Sydney), Tongliang Liu (University of Sydney)

Safty and PrivacyAdversarial AttackTransformerDiffusion modelGenerative Adversarial NetworkContrastive LearningGaussian SplattingImagePoint Cloud

🎯 What it does: Active protection of 3D Gaussian Splatting assets to prevent illegal edits driven by instructions

Advancing Analytic Class-Incremental Learning through Vision-Language Calibration

Binyu Zhao (Harbin Institute of Technology), Ivor Tsang (Agency for Science, Technology and Research)

ClassificationComputational EfficiencyRepresentation LearningMeta LearningTransformerReinforcement LearningPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBenchmark

🎯 What it does: This study proposes a framework called VILA, which is based on pre-trained models and is designed for analytical class-incremental learning, achieving fast and efficient continuous learning through a dual-branch visual-language calibration mechanism.

Advancing LLM Reasoning with Natural Language and Numerical Feedback

Xiaoying Zhang (Chinese University of Hong Kong), Helen M. Meng (Chinese University of Hong Kong)

OptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningTextChain-of-Thought

🎯 What it does: Developed an online reinforcement learning framework called Critique-GRPO, which can simultaneously utilize numerical rewards and natural language critiques to promote self-improvement of large language models in reasoning tasks.

Advancing SVD-based LLM Compression via Layer-Wise Error Model Search

Moritz Thoma (Technical University of Munich), Ulf Schlichtmann (Technical University of Munich)

CompressionComputational EfficiencyKnowledge DistillationHyperparameter SearchTransformerLarge Language ModelText

🎯 What it does: This paper proposes two techniques, hierarchical error model search (LEMS) and token-based Fisher approximation (KFAC-SVD), for performing SVD compression on large language models;

Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation

Xixiang He (National University of Defense Technology), Qingyong Hu (Intelligent Game and Decision Lab)

OptimizationReinforcement Learning from Human FeedbackReinforcement LearningPrompt EngineeringTextTabularBenchmarkChain-of-Thought

🎯 What it does: Studied the advantage collapse problem that occurs under the RLVR framework in GRPO, and proposed the ACR diagnostic metric and the AVSPO repair method.

Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion Models

Shuchen Xue (University of Chinese Academy of Sciences), Zhi-Ming Ma (University of Chinese Academy of Sciences)

GenerationReinforcement Learning from Human FeedbackTransformerSupervised Fine-TuningReinforcement LearningPrompt EngineeringDiffusion modelScore-based ModelFlow-based ModelAuto EncoderImageStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: This paper proposes the Advantage Weighted Matching (AWM) method, aligning the objectives of reinforcement learning with diffusion model pre-training, and performs variance analysis on existing methods such as DDPO.

AdverMCTS: Combating Pseudo-Correctness in Code Generation via Adversarial Monte Carlo Tree Search

Qingyao Li (Shanghai Jiao Tong University), Bo An (Nanyang Technological University)

Explainability and InterpretabilityComputational EfficiencyAdversarial AttackAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsTextSequentialRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose the ADVERMCTS framework, which addresses the issue of pseudo-correctness in code generation through dual MCTS adversarial search by Solver and Attacker.

Adversarial Attack and Defense for Denoising Diffusion Sampling

Zhao-Rong Lai (Guangzhou University), Jian Weng (Guangzhou University)

GenerationData SynthesisOptimizationAdversarial AttackDiffusion modelScore-based ModelContrastive LearningImagePoint CloudTabularStochastic Differential Equation

🎯 What it does: This study investigates adversarial attacks and defenses in the Denoising Diffusion Sampling (DDS) process, proposing an attack method that injects perturbations into the sampling stage, and designing a defense framework called ADDDS based on Local Variation (LV) regularization. It also provides a conjugate gradient solving algorithm combined with zero-order Monte Carlo (ZOD-MC).

Adversarial Attacks and Robust Training for Hypergraph Neural Networks

Naheed Anjum Arafat (Howard University), Danda B. Rawat (Howard University)

Adversarial AttackGraph Neural NetworkContrastive LearningGraph

🎯 What it does: Proposes a gray-box attack framework MeLA (Meta-Laplacian Attack) for hypergraph neural networks and its corresponding robust training method MeLA-D, which can simultaneously perform low-budget perturbations on the hypergraph structure and node features, and uses the Laplacian operator as the meta-objective.

Adversarial Dual On-Policy Distillation from Expressive Teacher

Zhenglin Wan (National University of Singapore), Yang You (National University of Singapore)

Knowledge DistillationRobotic IntelligenceReinforcement Learning from Human FeedbackSpiking Neural NetworkTransformerReinforcement LearningDiffusion modelScore-based ModelFlow-based ModelRectified FlowGenerative Adversarial NetworkContrastive LearningTabularTime SeriesSequential

🎯 What it does: Propose a new learning from demonstration method called FA-OPD, which uses Flow Matching (FM) as a co-trainable teacher during the learning process. It employs a dual-channel (reward and action) self-supervised training on a student (lightweight MLP) to accomplish robotic control tasks without requiring real environment rewards.

Adversarial Flow Models

Shanchuan Lin (ByteDance Seed), Haoqi Fan (ByteDance Seed)

GenerationTransformerDiffusion modelScore-based ModelFlow-based ModelRectified FlowGenerative Adversarial NetworkImage

🎯 What it does: Proposed Adversarial Flow Models, combining adversarial training and flow matching on the Transformer architecture to achieve single-step or multi-step generation, and supporting end-to-end training of very deep models.

Adversarial Latent Embedding Repair for LLM Continual Learning

Xilin Xia (University of Science and Technology of China), Feng Wu (University of Science and Technology of China)

Knowledge DistillationRepresentation LearningAdversarial AttackTransformerLarge Language ModelSupervised Fine-TuningGenerative Adversarial NetworkContrastive LearningTextBenchmarkFinance Related

🎯 What it does: Propose the ALER framework, which first performs adversarial latent embedding search to locate the most forgettable regions of the model, and then performs online distillation repair on these embeddings to achieve continuous learning in LLMs;

Adversarial Reinforcement Learning for Robust Diffusion Large Language Model Unlearning

Zhiwei Zhang (Pennsylvania State University), Suhang Wang (Pennsylvania State University)

Safty and PrivacyExplainability and InterpretabilityAdversarial AttackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringDiffusion modelText

🎯 What it does: A systematic study on unlearning in diffusion-based language models (DLMs), and proposes an adversarial reinforcement learning framework tailored for DLMs, enhancing the robustness of unlearning in adversarial contexts.

Adversarial Robustness of Implicit Neural Representation-Based Classifiers

Jayoung Kim (Korea Advanced Institute of Science & Technology), Sanghyun Hong (Oregon State University)

ClassificationAdversarial AttackPrompt EngineeringDiffusion modelScore-based ModelNeural Radiance FieldGenerative Adversarial NetworkContrastive LearningImagePoint Cloud

🎯 What it does: This paper systematically evaluates the robustness of classifiers based on implicit neural representations (INR) under adversarial attacks, and proposes an attack framework based on a surrogate model.

Adversarial Training for Process Reward Models

Gurusha Juneja (University of California, Santa Barbara), William Yang Wang (University of California, Santa Barbara)

OptimizationExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerReinforcement LearningPrompt EngineeringGenerative Adversarial NetworkContrastive LearningTextBenchmarkChain-of-Thought

🎯 What it does: Propose an adversarial training framework APRM for Process Reward Models (PRM), enabling the generator to learn to produce misleading reasoning steps to enhance the robustness and generalization ability of PRM.

Adversarial Vulnerability from Interference Between Features in Superposition

Edward Stevinson (Imperial College London), Tolga Birdal (Imperial College London)

Representation LearningAdversarial AttackTransformerContrastive LearningImage

🎯 What it does: Studied how feature interference caused by superposition in neural networks leads to adversarial attacks, and verified this mechanism through theoretical and experimental analysis.

Adversarially Robust Approximate Furthest Neighbor

Kiarash Banihashem (University of Maryland), Sandeep Silwal (University of Wisconsin Madison)

OptimizationComputational EfficiencyAdversarial AttackDiffusion modelScore-based ModelAuto EncoderContrastive LearningGaussian Splatting

🎯 What it does: Propose an adversarial robust approximate farthest neighbor query data structure that supports adaptive queries and achieves sublinear query time identical to the optimal non-adversarial algorithm.

Adversarially Robust Control of Conditional Value-at-Risk via Rockafellar-Uryasev Conformal Inference

Catherine Chen (Stanford University), Lihua Lei (Stanford University)

OptimizationSafty and PrivacyReinforcement Learning from Human FeedbackReinforcement LearningContrastive LearningTextTabularTime SeriesFinance Related

🎯 What it does: Proposes an online, distribution-agnostic Conditional Value-at-Risk (CVaR) control framework that transforms tail risk control into a two-layer online optimization problem by leveraging the Rockafellar-Uryasev (RU) variational representation;

AdvEvo-MARL: Shaping Internalized Safety through Adversarial Co-Evolution in Multi-Agent Reinforcement Learning

Zhenyu Pan (Northwestern University), Han Liu (Northwestern University)

Safty and PrivacyTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningText

🎯 What it does: Propose AdvEvo-MARL, a multi-agent reinforcement learning framework that internalizes security awareness through co-evolutionary training between attackers and defenders;

AES: Curing Optimizer Blindness in Long-Tailed Recognition via State-Aware Correction

Fanfu Wang (University of Science and Technology of China), Yang Wang (University of Science and Technology of China)

ClassificationRecognitionOptimizationSupervised Fine-TuningContrastive LearningImage

🎯 What it does: Propose the AES framework to address the optimizer blind spots in long-tailed classification, dynamically correcting biases in supervision, gradient, and inference processes through adaptive residual supervision, entropy-aware PCGrad, and sample-level conflict arbitrated fusion.

AesFormer: Transform Everyday Photos into Beautiful Memories

Tianxiang Du (Peking University), Yuxin Peng (Peking University)

Image TranslationRestorationGenerationTransformerLarge Language ModelReinforcement LearningVision Language ModelVision-Language-Action ModelFlow-based ModelImageVideoText

🎯 What it does: Propose a two-stage framework called AesFormer, which achieves the aesthetic photo reconstruction task by first planning aesthetic actions and then executing structural edits.

AffIn-Space: Learning Affine-Invariant Representations for 3D Spatial Understanding with MLLMs

Zhenyu Lu (Peng Cheng Laboratory), Yaowei Wang (Harbin Institute of Technology)

Representation LearningRobotic IntelligenceReinforcement Learning from Human FeedbackTransformerLarge Language ModelVision Language ModelAuto EncoderContrastive LearningImageVideoTextMultimodalityPoint Cloud

🎯 What it does: Propose a framework named AFFIN-SPACE, which learns strict affine-invariant representations in multi-modal large language models, thereby achieving robust understanding of 3D space.

Affine-Equivariant Kernel Space Encoding for NeRF Editing

Mikołaj Zieliński (Poznan University of Technology), Przemysław Spurek

Computational EfficiencyKnowledge DistillationRepresentation LearningNeural Radiance FieldContrastive LearningGaussian SplattingOptical FlowImagePoint Cloud

🎯 What it does: This paper proposes an affine equivariant Gaussian kernel space encoding (EKS) for achieving editable and physics-driven neural radiance field rendering.

Affine-Scaled Attention: Towards Flexible and Stable Transformer Attention

Jeongin Bae (NAVER Cloud), Dongsoo Lee (NAVER Cloud)

Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelContrastive LearningText

🎯 What it does: This paper proposes a new Transformer attention mechanism called Affine-Scaled Attention, which adds input-dependent scaling factors and biases to the standard softmax weights, thereby relaxing normalization constraints and enhancing the flexibility and stability of attention.

AG-REPA: Causal Layer Selection for Representation Alignment in Audio Flow Matching

Pengfei ZHANG, Li Liu (Hong Kong University of Science and Technology)

GenerationComputational EfficiencyRepresentation LearningTransformerScore-based ModelFlow-based ModelRectified FlowContrastive LearningAudio

🎯 What it does: Propose a mid-layer alignment method based on causal attribution, AG-REPA, to improve the training efficiency and generation quality of audio stream matching models.

Agent JIT Compilation for Latency-Optimizing Web Agent Planning and Scheduling

Caleb Winston (Stanford University), Christos Kozyrakis (Stanford University)

OptimizationComputational EfficiencyAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextTabularBenchmark

🎯 What it does: This paper proposes an Agent JIT compilation framework that compiles natural language tasks into executable code, reducing LLM calls and significantly improving the execution speed and accuracy of web automation tasks.

Agent Learning via Early Experience

Kai Zhang (Meta Superintelligence Labs), Yifan Wu (Meta Superintelligence Labs)

TransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AIWorld ModelTextBenchmarkChain-of-Thought

🎯 What it does: Propose the 'Early Experience' paradigm, utilizing the future states generated by the agent's own actions as a reward-free supervisory signal to further enhance the learning effectiveness of language agents.

Agent Primitives: Reuseable Latent Building Blocks for Multi-Agent Systems

Haibo Jin (University of Illinois at Urbana-Champaign), Haohan Wang (University of Illinois at Urbana-Champaign)

Autonomous DrivingOptimizationComputational EfficiencyAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextSequentialBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposes Agent Primitives as reusable potential building blocks, using KV Cache for implicit communication to construct an automated multi-agent system.

Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning

Zhaoyang Wang (University of North Carolina at Chapel Hill), Yuxiong He (Snowflake)

Data SynthesisAI Code AssistantTransformerLarge Language ModelReinforcement LearningPrompt EngineeringWorld ModelTextTabularBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Developed Agent World Model (AWM) — a full-process automated pipeline based on LLMs, capable of generating 1,000 executable, database-driven agentic environments from scene descriptions, and providing complete tool interfaces and verification mechanisms for large-scale reinforcement learning training.

Agent-Omit: Adaptive Context Omission for Efficient LLM Agents

Yansong Ning (Hong Kong University of Science and Technology), Hao Liu (Hong Kong University of Science and Technology)

Computational EfficiencyKnowledge DistillationData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose the Agent-Omit framework, which achieves efficient reasoning for LLM agents by adaptively omitting redundant thoughts and observations in multi-round interactions.

Agent0-VL: Exploring Self-Evolving Agent for Tool-Integrated Vision-Language Reasoning

Jiaqi Liu (University Of North Carolina At Chapel Hill), Huaxiu Yao (University Of North Carolina At Chapel Hill)

Autonomous DrivingOptimizationRobotic IntelligenceMeta LearningReinforcement Learning from Human FeedbackTransformerSupervised Fine-TuningReinforcement LearningAgentic AIVision Language ModelImageTextMultimodalityBenchmarkChain-of-Thought

🎯 What it does: Propose Agent0-VL, a self-evolving vision-language reasoning agent that integrates tool usage. The internal loop consists of Solver (reasoning) and Verifier (self-assessment and self-correction), enabling continuous improvement through tool verification and repair without relying on external rewards.

AgentConductor: Topology Evolution for Multi-Agent Competition-Level Code Generation

Siyu Wang (Shanghai Jiao Tong University), Xinping Guan (Shanghai Jiao Tong University)

OptimizationAI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AIPrompt EngineeringTextGraph

🎯 What it does: Propose AgentConductor, a multi-agent system based on reinforcement learning, which utilizes an LLM-led orchestrator to end-to-end dynamically generate hierarchical DAG interaction topologies that match task difficulty, for competition-level code generation;

AgentExpt: Automating AI Experiment Design with LLM-based Resource Retrieval Agent

Yu Li (Tsinghua University), Yong Li (Tsinghua University)

Recommendation SystemExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This paper constructs the AgentExpt framework, which utilizes a large-scale paper-baseline-dataset knowledge base to automatically recommend experimental baselines and datasets.

AgentHijack: Benchmarking Computer Use Agent Robustness to Common Environment Corruptions

Jingwei Sun (Hong Kong Baptist University), Bo Han (Hong Kong Baptist University)

Anomaly DetectionRobotic IntelligenceTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringVision-Language-Action ModelImageTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose the AgentHijack benchmark, using nine configurable common environmental disruptions (pop-ups, resolution changes, annotations, subtitles, multiple windows, accidental touches, minimization, network errors, verification) to evaluate the robustness of multi-modal large language model-driven computer usage agents, and build the AgentHijack-Agent framework based on an action generator and a 'bystander' module to enhance robustness.

Agentic Confidence Calibration

Jiaxin Zhang (Salesforce AI Research), Chien-Sheng Wu (Salesforce AI Research)

Explainability and InterpretabilityData-Centric LearningRecurrent Neural NetworkTransformerLarge Language ModelAgentic AIAuto EncoderContrastive LearningGaussian SplattingTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Studied agentic confidence calibration, proposing the Holistic Trajectory Calibration (HTC) framework, which uses process-level features to calibrate the confidence of large language model agents.

Agentic Framework for Epidemiological Modeling

Rituparna Datta (University of Virginia), Anil Vullikanti (University of Virginia)

OptimizationExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelAgentic AITextTabularTime SeriesBiomedical DataBenchmarkRetrieval-Augmented Generation

🎯 What it does: Built a proxy-based framework called EPIAGENT for automatically generating, calibrating, and verifying epidemiological simulators from natural language scenario descriptions.

Agentic Model Predictive Questioning Control in Visual Design

Kuang-Da Wang (National Yang Ming Chiao Tung University), Shingo Takamatsu (Sony Group Corporation)

GenerationOptimizationTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringImageTextMultimodalityRetrieval-Augmented Generation

🎯 What it does: This paper proposes an agent-based model predictive query control (A-MPQC) for visual design, which enhances alignment between design and user intent and reduces cognitive load through multi-round clarifications under a fixed query budget.

Agentic Monte Carlo: Simulating Reinforcement Learning for Black-Box Agents

Dae Yon Hwang (Layer 6 AI), Brendan Leigh Ross (Layer 6 AI)

TransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringTextBenchmark

🎯 What it does: A framework called Agentic Monte Carlo (AMC) is proposed to optimize the behavior of black-box large language models (LLMs) that are only accessible through APIs, without requiring gradient information during testing.

Agentic Proposing: Enhancing Large language Model Reasoning via Compositional Skill Synthesis

Zhengbo Jiao (Alibaba Group), Linfeng Zhang (Shanghai Jiao Tong University)

Data SynthesisOptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningAdversarial AttackHyperparameter SearchData-Centric LearningRobotic IntelligenceMeta LearningDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackNeural Architecture SearchProtein Structure PredictionTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AIPrompt EngineeringMixture of ExpertsDiffusion modelScore-based ModelRectified FlowNeural Radiance FieldAuto EncoderGenerative Adversarial NetworkContrastive LearningGaussian SplattingSimultaneous Localization and MappingWorld ModelOptical FlowTextSequentialChain-of-ThoughtStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Proposes a framework called Agentic Proposing, which automatically generates high-difficulty, verifiable training data by treating question generation as a goal-driven sequence decision process, leveraging composable reasoning skills and closed-loop internal reflection with tool calls.

AgentLAB: Benchmarking LLM Agents against Long-Horizon Attacks

Tanqiu Jiang (Stony Brook University), Ting Wang (Stony Brook University)

Safty and PrivacyAdversarial AttackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed the AgentLAB benchmark, specifically designed to evaluate the security of large language model (LLM) agents in multi-round, long-term attack scenarios.

AgentNoiseBench: Benchmarking Robustness of Tool-Using LLM Agents Under Noisy Condition

Ruipeng Wang (University of Science and Technology of China), Tat-Seng Chua (National University of Singapore)

Explainability and InterpretabilityRobotic IntelligenceReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringContrastive LearningTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose the AgentNoiseBench framework to systematically evaluate the robustness of LLM agents in noisy environments;