arXivSub Start free trial

ICML 2026 Papers — Page 28

International Conference on Machine Learning · 6554 papers

Implicit Turn-Wise Policy Optimization for Proactive User-LLM Interaction

Haoyu Peter Wang (Georgia Tech), Pan Li (Georgia Tech)

OptimizationExplainability and InterpretabilityTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextSequentialRetrieval-Augmented Generation

🎯 What it does: Proposes the Implicit Turn-Wise Policy Optimization (ITPO) framework, which generates fine-grained turn-wise rewards for multi-turn human-machine interaction using an implicit process reward model, and improves training stability through a normalization mechanism, achieving end-to-end reinforcement learning optimization.

Imposing Boundary Conditions on Neural Operators via Learned Function Extensions

Sepehr Mousavi (ETH Zurich), Laura De Lorenzis (ETH Zurich)

Graph Neural NetworkTransformerContrastive LearningGraphBenchmarkPhysics Related

🎯 What it does: Proposes a framework that maps boundary conditions to a global pseudo-extension, enabling neural operators to directly utilize complex inhomogeneous boundary information to predict PDE solutions.

ImpQuant: Fine-Grained Importance-Aware Quantization for Large Vision-Language Models

Jundong Zhou (Shanghai Jiao Tong University), Nanyang Ye (Shanghai Jiao Tong University)

Computational EfficiencyRepresentation LearningTransformerVision Language ModelMultimodality

🎯 What it does: Propose an importance-aware post-training quantization framework called ImpQuant, which optimizes the accuracy degradation problem of vision-language models during low-bit deployment.

Improved Algorithms for Nash Welfare in Linear Bandits

Dhruv Sarkar (Indian Institute of Technology Kharagpur), Sayak Ray Chowdhury (Indian Institute of Technology Kanpur)

OptimizationReinforcement LearningTabular

🎯 What it does: This paper proposes an optimal algorithm for fairness in linear bandits, introducing the Nash regret and a more general p-means regret metric.

Improved Analysis of the Accelerated Noisy Power Method with Applications to Decentralized PCA

Pierre Aguié (Inria, DI ENS, PSL Research University), Laurent Massoulié (Inria, DI ENS, PSL Research University)

OptimizationFederated LearningGraphTabularBenchmark

🎯 What it does: An improved accelerated noise power method (ANPM) is proposed and applied to decentralized principal component analysis (PCA), with theoretical analysis showing that it maintains accelerated convergence under more relaxed noise conditions. Subsequently, the ADePM algorithm is designed to achieve decentralized accelerated PCA.

Improved Bounds for Private and Robust Alignment

Wenqian Weng (Wayne State University), Xingyu Zhou (Wayne State University)

OptimizationSafty and PrivacyReinforcement Learning from Human FeedbackLarge Language ModelReinforcement LearningContrastive LearningText

🎯 What it does: This paper presents a theoretical analysis of language model alignment under privacy protection (local differential privacy) and robustness (Huber corruption) conditions, and provides upper bounds on the optimal sample complexity for both offline and online scenarios.

Improved Bounds for Reward-Agnostic and Reward-Free Exploration

Oran Ridel (Tel Aviv University), Alon Cohen (Tel Aviv University)

Reinforcement Learning

🎯 What it does: This paper proposes a new reward-agnostic and reward-oblivious exploration algorithm, significantly reducing the dependence on the accuracy parameter ϵ, and provides the optimal lower bound for reward-agnostic exploration under time-inhomogeneous MDPs, improving the theoretical description of the global optimal sample complexity.

Improved Convergence Analysis of Topology Dependence in Decentralized SGD

Yuki Takezawa (Toyota Motor Corporation), Sebastian U Stich

OptimizationFederated LearningImageTabular

🎯 What it does: The paper improves the convergence analysis of decentralized SGD (Decentralized SGD), revealing the impact of network topology on convergence rate;

Improved Dimension Dependence for Bandit Convex Optimization with Gradient Variation

Hang Yu (Nanjing University), Peng Zhao (Nanjing University)

Optimization

🎯 What it does: This paper studies bandit convex optimization with gradient variation (BCO), proposing an improved method for analyzing discontinuous gradient variation, significantly enhancing dimension dependence.

Improved Distribution Estimation in $\ell_\infty$

Doron Cohen (Ben-Gurion University of the Negev), Yonatan Livshitz (Ben-Gurion University of the Negev)

OptimizationData-Centric LearningLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsContrastive LearningRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Optimal expected error and high-probability error bounds are provided for estimates under the ℓ∞ norm with discrete distributions, and the worst-case instances are identified through theoretical analysis.

Improved Dynamic Algorithm for Non-monotone Submodular Maximization under Cardinality Constraint

Kiarash Banihashem (University of Maryland), Morteza Monemizadeh (TU Eindhoven)

OptimizationReinforcement LearningContrastive Learning

🎯 What it does: Proposes an algorithm for non-monotonic submodular maximization under constraint k in a fully dynamic environment

Improved Scaling Laws via Weak-to-Strong Generalization in Random Features Ridge Regression

Diyuan Wu (Institute of Science and Technology Austria), Marco Mondelli (Institute of Science and Technology Austria)

OptimizationKnowledge DistillationRepresentation LearningContrastive LearningImageTabular

🎯 What it does: This paper studies the training of a strong student model by using weak teacher models to generate labels within the Random Feature Ridge Regression (RFRR) framework, and proves that the student can achieve better scaling law performance than the teacher under appropriate regularization and overparameterization conditions.

Improved Stochastic Optimization of LogSumExp

Egor Gladin (HSE University), Pavel Dvurechensky (WIAS)

OptimizationContrastive LearningImageTabularStochastic Differential Equation

🎯 What it does: Propose a LogSumExp approximation based on Safe KL divergence, which maintains convexity and smoothness and provides unbiased stochastic gradients, solving the numerical instability and gradient bias problems caused by traditional exponential summation;

Improving Adversarial Robustness of Attribution via Implicit Regularization

Amir Mehrpanah (Kth Royal Institute Of Technology), Hossein Azizpour (Kth Royal Institute Of Technology)

Explainability and InterpretabilityAdversarial AttackConvolutional Neural NetworkTransformerContrastive LearningImage

🎯 What it does: By analyzing the implicit regularization in the learning dynamics of SGD, this paper investigates methods to enhance the robustness of gradient-based explanation methods, and reveals the issue that the robustness is limited due to the entropy constraint on softmax attention.

Improving Backward Conformal Prediction via Non-Conformity Score Transformation

Junxian Liu (Southern University of Science and Technology), Hongxin Wei (Southern University of Science and Technology)

ClassificationConvolutional Neural NetworkScore-based ModelImage

🎯 What it does: Proposed a data-dependent non-consistency score transformation called ST-BCP, improving the coverage error and control of the prediction set size in backward consistency prediction;

Improving Classifier-Free Guidance of Flow Matching via Manifold Projection

Jian-Feng Cai (Hong Kong University of Science and Technology), Chao Wang (Southern University of Science and Technology)

GenerationOptimizationComputational EfficiencyTransformerDiffusion modelScore-based ModelFlow-based ModelRectified FlowImageText

🎯 What it does: Proposes a theoretical explanation of classifier-free guidance (CFG) in flow matching models from an optimization perspective, and designs algorithms to improve CFG sampling through manifold projection, named CFG-MP and its accelerated version CFG-MP+.

Improving CLIP Adaptation by Breaking Tail Alignment for Source-Free Cross-Domain Few-Shot Learning

Shuai Yi (Huazhong University of Science and Technology), Ruixuan Li (Huazhong University of Science and Technology)

ClassificationDomain AdaptationTransformerPrompt EngineeringContrastive LearningImageText

🎯 What it does: In cross-domain few-shot learning tasks without source domain data, we propose adaptive alignment of different image patches in the CLIP vision transformer: bringing the head tokens rich in semantic information closer, and pushing away the tail tokens with insufficient semantic information, to improve the classification performance in the target domain.

Improving Diffusion Planners by Self-Supervised Action Gating with Energies

Yuan Lu (University College London), Dongsheng Li (Microsoft Research)

TransformerReinforcement LearningDiffusion modelContrastive LearningTabularTime SeriesSequentialBenchmark

🎯 What it does: Propose SAGE (Self-supervised Action Gating with Energies), a plug-and-play solution for diffusion planning in offline reinforcement learning, which re-ranks sampled candidate trajectories using self-supervised energy scoring to reduce failures caused by dynamic inconsistency.

Improving Explicit Dynamic Gaussian Splatting Optimization via Update Mixture

Renjie Ding (Hunan University), Xiang Chen (Hunan University)

OptimizationGaussian SplattingVideo

🎯 What it does: An improvement is proposed for the explicit dynamic Gaussian splatting (Dynamic GS) optimization method, introducing an updated hybrid strategy to alleviate generalization degradation in large motion scenarios.

Improving LLM-Based Recommenders with Conservative Generative Flow Networks

Xuan Yu (University of Science and Technology of China), Yang Wang (University of Science and Technology of China)

Recommendation SystemTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningFlow-based ModelText

🎯 What it does: This paper studies the use of generative flow networks in offline LLM recommendation scenarios to address distribution mismatch issues, and proposes the CFlower framework, which penalizes forward traffic not supported by logs through conservative sub-trajectory balancing regularization.

Improving ML Attacks on LWE with Data Repetition and Stepwise Regression

Alberto Alfarano (Axiom Math), Kristin E. Lauter

Adversarial AttackTransformerSupervised Fine-TuningContrastive LearningTabularPhysics Related

🎯 What it does: This paper proposes an improved machine learning attack method, using a larger training set, repeated examples, and stepwise regression to attack LWE secrecy.

Improving Sampling for Masked Diffusion Models via Information Gain

Kaisen Yang (Tsinghua University), Alex Lamb (Tsinghua University)

GenerationData SynthesisOptimizationComputational EfficiencyTransformerPrompt EngineeringDiffusion modelScore-based ModelImageTextMultimodalityChain-of-Thought

🎯 What it does: Propose an information gain sampler (Info-Gain Sampler) that provides a training-agnostic global planning decoding strategy for unsupervised masked diffusion models (MDMs);

Improving the Performance and Learning Stability of Parallelizable RNNs Designed for Ultra-Low Power Applications

Julien Brandoit (University of Liège), Guillaume Drion (University of Liège)

Computational EfficiencyRepresentation LearningRecurrent Neural NetworkContrastive LearningTextTime SeriesSequentialAudio

🎯 What it does: Proposed the Cumulative Memory Recurrent Unit (CMRU) and its relaxed version α CMRU, which solves the gradient blocking problem in BMRU during state updates, achieving a parallelizable trained persistent memory RNN.

Improving the Robustness-Utility Trade-off in Decentralized Learning over Sparse Networks

Yangnan Li (Hong Kong University of Science and Technology), Shenghui Song (Hong Kong University of Science and Technology)

OptimizationFederated LearningImage

🎯 What it does: Improved the trade-off between robustness and efficiency in distributed learning under sparse networks, proposing Scaled Dual Ascent (SDA) within the augmented Lagrangian framework and its variants BRED and BRED-M, achieving linear acceleration with significantly reduced transient complexity.

Improving the Sensitivity of Backdoor Detectors via Class Subspace Orthogonalization

Guangmingmei Yang (Penn State University), George Kesidis (Penn State University)

Anomaly DetectionAdversarial AttackConvolutional Neural NetworkSupervised Fine-TuningContrastive LearningImageBenchmark

🎯 What it does: Propose a post-training backdoor detection method called Class Subspace Orthogonalization (CSO), which enhances detection sensitivity by suppressing intra-class inherent features and emphasizing the direction of the backdoor trigger.

Improving Topic Modeling by Distilling Soft Labels from Language Models

Raymond Li (University of British Columbia), Giuseppe Carenini (University of British Columbia)

Knowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningText

🎯 What it does: Train neural topic models by using temperature-scaled soft labels generated by small language models (SLMs) under topic prompts as the reconstruction target of the topic model;

Improving Visual Token Reduction via Rectifying Distortions for Efficient Multimodal LLM Inference

Hyeonwoo Cho (Yonsei University), Bumsub Ham (Yonsei University)

Computational EfficiencyRepresentation LearningTransformerLarge Language ModelVision Language ModelContrastive LearningImageTextMultimodality

🎯 What it does: Propose the RESTORE framework to correct spatial and attention distortions that occur during the visual token reduction (VTR) process in multimodal large language models;

Improving Zero-Shot Offline RL via Behavioral Task Sampling

Nazim Bendib (Sorbonne Universite), Olivier Sigaud (Sorbonne Universite)

TransformerReinforcement LearningAuto EncoderGenerative Adversarial NetworkContrastive LearningTabularTime Series

🎯 What it does: Propose a behavior task distribution (BTD) sampling method based on offline data extraction to improve the task sampling strategy in offline zero-shot reinforcement learning.

ImpText: A Benchmark and Tool-Augmented Framework for Implicit Text Reasoning

Litao Guo (Hong Kong University of Science and Technology), Ying-Cong Chen (Hong Kong University of Science and Technology)

Data-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision Language ModelImageTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: Define the implicit text reasoning task, construct the ImpText-Bench benchmark, and propose the tool-enhanced framework ImpText-Reader.

In-Context Generation with Regional Constraints for Instructional Video Editing

Zhongwei Zhang (University of Science and Technology of China), Tao Mei (HiDream.ai Inc)

RestorationGenerationTransformerSupervised Fine-TuningPrompt EngineeringDiffusion modelVideoText

🎯 What it does: Proposed a source-target video joint denoising framework called ReCo based on width concatenation, which utilizes natural language instructions for instructive video editing and ensures precise localization of the editing region and background preservation through regional constraints.

In-Context Learning as Rate–Distortion Optimization

Jiayu Zhang (Peking University), Canran Xiao (Sun Yat-sen University)

CompressionOptimizationData-Centric LearningMeta LearningTransformerPrompt EngineeringContrastive LearningTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: This paper proposes an RDCO method, a training-free context construction approach based on rate-distortion optimization, for efficiently selecting and compressing demonstration examples under a limited token budget.

In-Context Learning Is Provably Bayesian Inference: A Generalization Theory for Meta-Learning

Tomoya Wakayama (RIKEN Center for Advanced Intelligence Project), Taiji Suzuki (RIKEN Center for Advanced Intelligence Project)

Meta LearningTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningText

🎯 What it does: Studied the finite-sample statistical theory under a meta-learning framework, proving that in-context learning (ICL) is equivalent to Bayesian inference, and provided a decomposition of the ICL risk into Bayesian gap and posterior variance.

In-Context Universal Approximation, Compositional Generalization, and Algorithm Emulation

Jerry Yao-Chieh Hu (Northwestern University), Han Liu (Northwestern University)

Meta LearningTransformerPrompt Engineering

🎯 What it does: Prove that any sequence-to-sequence function satisfying L-Lipschitz continuity can be approximated by a softmax Transformer with fixed weights through different prompts.

In-Training Defenses Against Emergent Misalignment in Language Models

David Kaczér (University of Bonn), Florian Mai (University of Bonn)

Safty and PrivacyExplainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringMixture of ExpertsTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Studied and systematically evaluated various training-time regularization methods to prevent the emergence of 'Emergent Misalignment (EM)' during the fine-tuning of large language models (LLMs).

Incentivized Exploration with Stochastic Covariates: A Two-Stage Mechanism Design for Recommender System

Yuantong Li (Meta Platforms Inc), Xiaowu Dai (University of California, Los Angeles)

Recommendation SystemReinforcement LearningTabularBiomedical Data

🎯 What it does: Designed a two-stage mechanism that balances incentive exploration and linear contextual bandits, achieving a recommendation system that satisfies dynamic Bayesian incentive compatibility.

Incentivizing Truthfulness and Collaborative Fairness in Bayesian Learning

Rachael Hwee Ling Sim (National University of Singapore), Bryan Kian Hsiang Low (National University of Singapore)

Federated LearningExplainability and InterpretabilityData-Centric LearningContrastive LearningImageTabularBiomedical Data

🎯 What it does: Designed a mechanism that enables data sources in collaborative machine learning to receive fair rewards and be motivated to submit truthful data.

Incomplete Multi-View Clustering via Neighborhood-Conditioned Diffusion

Qian Guo (Taiyuan University of Science and Technology), Jianjian Ding (Taiyuan University of Science and Technology)

RestorationRepresentation LearningTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningImageGraph

🎯 What it does: A missing multi-view clustering framework based on neighborhood conditional diffusion, named IMVC-NCD, is proposed. It utilizes view-specific autoencoders to learn low-dimensional latent representations, and encodes cross-view neighbor information into a unified conditional vector through a neighborhood condition construction module. Stable latent recovery is then performed within a diffusion model, ultimately achieving better clustering results.

Inconsistency-Aware Minimization: Improving Generalization with Unlabeled Data

Hee-Sung Kim (Hanyang University), Sungyoon Lee (Hanyang University)

ClassificationOptimizationData-Centric LearningConvolutional Neural NetworkContrastive LearningImage

🎯 What it does: Propose a new information geometric quantity — local inconsistency — to measure the sensitivity of neural network outputs to parameter perturbations, and develop an inconsistency-aware minimization (IAM) training method based on this.

Incorporating Importance Weighting in Optimal Transport Based Domain Alignment

Okan Koç (RIKEN), Masashi Sugiyama (RIKEN)

Domain AdaptationOptimizationConvolutional Neural NetworkTransformerContrastive LearningImage

🎯 What it does: To address the Wasserstein boundary in unsupervised domain adaptation, the authors propose two assumptions: Gradual Shift (GS) and Probabilistic Margin (PM). Based on these assumptions, they design a Wasserstein Regularized Risk (W2R) objective with importance weighting, thereby achieving a tighter upper bound and improving optimization.

Incremental BPE Tokenization

Shenghu Jiang (Beihang University), Ruihao Gong (Beihang University)

Computational EfficiencyAI Code AssistantLarge Language ModelTextBenchmark

🎯 What it does: Proposes an incremental Byte Pair Encoding (BPE) tokenization algorithm that can update tokenization results in worst-case O(log n) time per byte, and supports streaming input and immediate output;

Incremental Learning of Sparse Attention Patterns in Transformers

Oğuz Kaan Yüksel (EPFL)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerSequentialStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Studied how Transformer learns sparse attention patterns in high-order Markov chain tasks step by step, observing a phased learning process transitioning from competition to cooperation.

Incremental Transformer Neural Processes

Philip Mortimer (University of Cambridge), Richard E. Turner (University of Cambridge)

Computational EfficiencyRepresentation LearningData-Centric LearningTransformerReinforcement LearningPrompt EngineeringContrastive LearningTabularTime SeriesSequential

🎯 What it does: Proposes an efficient incremental update model called Transformer Neural Process (incTNP) that can operate on real-time streaming data;

Independent Component Discovery in Temporal Count Data

Alexandre Chaussard (CNRS), Sylvain Le Corff (CNRS)

Data SynthesisAnomaly DetectionExplainability and InterpretabilityRepresentation LearningAuto EncoderContrastive LearningTabularTime SeriesBiomedical DataBenchmark

🎯 What it does: Proposed an adaptive dynamic ICA model based on Poisson log-normal for analyzing time count data.

INDEXGUARD: Index-only Backdoor Vetting for Secure Federated PEFT of Large Language Models

Javad Dogani (IMDEA Networks Institute), Nikolaos Laoutaris (IMDEA Networks Institute)

Anomaly DetectionFederated LearningSafty and PrivacyTransformerLarge Language ModelPrompt EngineeringContrastive LearningText

🎯 What it does: Developed IndexGuard, an index-only pre-aggregation backdoor detection mechanism for federated parameter-efficient fine-tuning (PEFT), which detects backdoors while meeting the privacy requirements of secure aggregation.

IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference

Xintong Yang (Hong Kong University of Science and Technology), Sirui Han (Hong Kong University of Science and Technology)

Computational EfficiencyKnowledge DistillationTransformerLarge Language ModelTextBenchmark

🎯 What it does: In LLM inference with long contexts, a learnable KV cache replacement strategy is used to compress memory.

Inducing Overthink: Hierarchical Genetic Algorithm-based DoS Attack on Black-Box Large Language Reasoning Models

Shuqiang Wang (Zhejiang University), Zhixuan Chu (Zhejiang University)

OptimizationComputational EfficiencyAdversarial AttackReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringTextBenchmark

🎯 What it does: The study uses hierarchical genetic algorithms to induce 'overthinking' attacks in black-box large-scale reasoning models, aiming to exhaust computational resources.

Induction Heads Interpolate N-Grams

Francesco D'Angelo (EPFL), Nicolas Flammarion (EPFL)

Explainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelContrastive LearningText

🎯 What it does: Investigate and demonstrate that transformer's induction heads can achieve n-gram smoothing through soft context matching and BOS pseudo-counts, thereby performing soft count estimation in Markov chain tasks;

Induction Meets Biology: Mechanisms of Repeat Detection in Protein Language Models

Gal Pomerants (Technion Israel Institute of Technology), Yonatan Belinkov (Technion Israel Institute of Technology)

Explainability and InterpretabilityProtein Structure PredictionTransformerLarge Language ModelContrastive LearningBiomedical Data

🎯 What it does: By reverse engineering the internal mechanisms of PLMs, this study investigates and explains how protein language models identify and complete exact or approximate repetitive sequences in the masked language modeling task.

INDUCTION: Finite-Structure Concept Synthesis in First-Order Logic

Serafim Batzoglou (Independent Researcher)

Data SynthesisLarge Language ModelPrompt EngineeringTextBenchmark

🎯 What it does: Designed and evaluated the mechanically verifiable finite structure concept synthesis benchmark INDUCTION, covering three task patterns: full observation, contrastive induction, and existence completion.

InertialAR: Autoregressive 3D Molecule Generation with Inertial Frames

Haorui Li (Chinese University of Hong Kong), Shengchao Liu (Chinese University of Hong Kong)

GenerationData SynthesisDrug DiscoveryTransformerDiffusion modelGenerative Adversarial NetworkContrastive LearningGraphTabular

🎯 What it does: Proposed a self-attention 3D molecular generation model named InertialAR;

INFER: Learning Implicit Neural Frequency Response Fields for Confined Acoustic Environments

Harshvardhan C. Takawale (University of Maryland), C. Phillip Brown (Dolby Laboratories, Inc.)

Neural Radiance FieldMeshAudio

🎯 What it does: Studied a frequency-domain based neural implicit frequency response field (INFER) for high-precision modeling of acoustic transmission in narrow and constrained environments.

Inference from Quantized Data via Normal Variance-Mean Mixtures

Chenyu Gao (ShanghaiTech University), Ziping Zhao (ShanghaiTech University)

Data SynthesisOptimizationComputational EfficiencyRepresentation LearningContrastive LearningMultimodalityTabularTime SeriesReview/Survey PaperBenchmark

🎯 What it does: This paper proposes an algorithm based on the Expectation Conditional Maximization (ECM) for maximum likelihood estimation of parameters in the normal variance-mean mixture (NVMM) model under multi-bit quantized observations, and generalizes this framework to structured learning tasks such as quantized linear regression, matrix completion, compressed sensing, and covariance estimation.

Inference of Online Newton Methods with Nesterov's Accelerated Sketching

Haoxuan Wang (Georgia Institute of Technology), Sen Na (Georgia Institute of Technology)

OptimizationComputational EfficiencyData-Centric LearningTabularTime Series

🎯 What it does: Proposed an online Newton method based on Nesterov accelerated projection (accelerated sketch-and-project) for streaming data; approximated the Newton direction by using the mean of the Hessian, maintaining a time complexity of O(d²); also provided the global almost sure convergence, asymptotic normality, and covariance matrix that can be estimated online for this method;

Inference Time Optimization with Confidence Dynamics

Yu Wang (Accenture), Wei Wei (Accenture)

OptimizationComputational EfficiencyTransformerLarge Language ModelTextBenchmarkChain-of-Thought

🎯 What it does: This paper studies the confidence dynamics during the reasoning process of large language models (LLMs), proposing a new voting method called Confidence Dynamic Gain (CDG) voting, aiming to optimize reasoning time.

Inference-Aware Meta-Alignment of LLMs via Non-Linear GRPO

Shokichi Takakura (LY Corporation), Taiji Suzuki (University of Tokyo)

OptimizationMeta LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringContrastive LearningText

🎯 What it does: Propose the IAMA framework, a Meta-Alignment approach that enables adaptation to multiple reward functions during inference, training the base model to quickly align with various human preferences using different alignment algorithms during inference.

Inference-time Alignment with Rewards in Besov Spaces: Provable Advantages of Feature Learning and Multi-Step Policy Updates

Naoki Nishikawa (University of Tokyo), Taiji Suzuki (University of Tokyo)

OptimizationReinforcement Learning

🎯 What it does: The study addresses the Inference-time Alignment problem in Besov spaces, proving that neural networks outperform linear estimators in reward function estimation and policy updates, and proposes a multi-step training strategy to further reduce regret.

Inference-Time Conformal Reasoning with Valid Factuality Control for Large Language Models

Ting Wang (University of Illinois Urbana Champaign), Huan Zhang (University of Illinois Urbana Champaign)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelTextGraphBenchmarkChain-of-Thought

🎯 What it does: Proposes an Inference-Time Conformal Reasoning (ITCR) framework that applies conformal prediction in real-time during multi-step reasoning, achieving factual control over the reasoning graph and determining when to stop expanding.

Inference-time optimization for experiment-grounded protein ensemble generation

Sai Advaith Maddipatla (Institute of Science and Technology), Alexander Bronstein

OptimizationProtein Structure PredictionTransformerDiffusion modelScore-based ModelBiomedical Data

🎯 What it does: Generate protein conformation ensembles consistent with experimental data (such as NMR NOE, X-ray density, and ipTM signals) by optimizing the Pairformer embeddings with gradient optimization during the inference stage of AlphaFold3.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training

Pengkai Wang (Hong Kong Polytechnic University), Hongxia Yang (Hong Kong Polytechnic University)

Explainability and InterpretabilityComputational EfficiencyData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBiomedical DataElectronic Health RecordsBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose an incremental reinforcement learning framework called ORBIT based on evaluation criteria, aimed at improving the performance of large language models in open-ended medical dialogues.

Infinite Mask Diffusion for Few-Step Distillation

Jaehoon Yoo (Korea Advanced Institute of Science and Technology), Seunghoon Hong (Korea Advanced Institute of Science and Technology)

GenerationKnowledge DistillationTransformerDiffusion modelScore-based ModelText

🎯 What it does: Propose the Infinite Mask Diffusion Model (IMDM), which introduces an infinite-state random mask on the basis of the Masked Diffusion Model (MDM), thereby eliminating the factorization error lower bound of MDM in few-step inference.

Infinite-Dimensional Generative Diffusions via Doob’s h-Transform

Thorben Pieper-Sethmacher (Nanyang Technological University), Daniel Paulin (Nanyang Technological University)

GenerationData SynthesisDiffusion modelScore-based ModelImagePoint CloudTabularTime SeriesPhysics RelatedStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Proposes constructing generative diffusion models in infinite-dimensional spaces through Doob's h-transform, and provides rigorous existence and approximation theories.

Infinite-Precision Autoregressive Modeling for Vector Graphics and Layouts

Yeonsang Shin (Seoul National University), Bohyung Han (Seoul National University)

GenerationData SynthesisTransformerDiffusion modelScore-based ModelAuto EncoderImageMeshSequentialBenchmark

🎯 What it does: Propose a self-attention model called IPAM, capable of jointly processing discrete and continuous values in sequences, achieving variable-length, infinite-precision vector graphics and layout generation.

Infinite-World: Scaling Interactive World Models to 1000-Frame Horizons via Pose-Free Hierarchical Memory

Ruiqi Wu (NKIARI), Ming-Ming Cheng (NKIARI)

GenerationCompressionRepresentation LearningReinforcement Learning from Human FeedbackRecurrent Neural NetworkTransformerSupervised Fine-TuningDiffusion modelAuto EncoderContrastive LearningWorld ModelImageVideo

🎯 What it does: Propose Infinite-World, an interactive world model that maintains consistency over long sequences of more than 1000 frames, enabling high-quality view switching and loop closure in real-world videos.

Influence-Disentangled Federated Training: Learning Models That Are Easy to Unlearn

Canran Xiao (Sun Yat-sen University), Liwei Hou (Airon Tech)

OptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencySupervised Fine-TuningAuto EncoderContrastive LearningImage

🎯 What it does: This paper proposes a framework that records and separates the influence of each client during the federated learning training process, enabling efficient and stable client-level unlearning by simply performing a single subtraction and short-term repair when a deletion request is made.

Influence-Guided Symbolic Regression: Scientific Discovery via LLM-Driven Equation Search with Granular Feedback

Evgeny Saveliev, Mihaela van der Schaar (University of Cambridge)

Explainability and InterpretabilityDrug DiscoveryTransformerLarge Language ModelMixture of ExpertsTabularBiomedical DataBenchmarkPhysics Related

🎯 What it does: Utilize large language models (LLMs) to generate candidate basis functions, and then perform fine-grained evaluation and pruning based on the impact score (∆) of each term, thereby automatically discovering interpretable scientific equations.

InfoAtlas: A Foundation Model for Zero-Shot Statistical Dependence Estimate

Zhengyang Hu (University of Hong Kong), Yanchao Yang (University of Hong Kong)

Data SynthesisComputational EfficiencyRepresentation LearningMeta LearningTransformerMixture of ExpertsDiffusion modelScore-based ModelFlow-based ModelContrastive LearningImageTextMultimodalityPoint CloudTabularBenchmark

🎯 What it does: Propose InfoAtlas, a foundational model that achieves zero-shot mutual information (MI) estimation through pre-training, capable of directly providing the statistical dependency strength between multi-dimensional random variables in a single forward pass.

InfoDLM: an Information-Adaptive Framework for Discrete Diffusion Language Model Pretraining

Shirou Jing (Rice University), Tony Geng (Rice University)

GenerationComputational EfficiencyRepresentation LearningTransformerLarge Language ModelReinforcement LearningDiffusion modelText

🎯 What it does: Proposed an adaptive masking strategy based on information gain, called InfoDLM, to optimize the pretraining of discrete diffusion language models.

InfoFlow KV: Information-Flow-Aware KV Recomputation for Long Context

Xin Teng (New York University), Shenji Wan (New York University)

Recommendation SystemComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextMultimodalityRetrieval-Augmented Generation

🎯 What it does: Propose an information-flow-aware KV recomputation method, which utilizes attention weights under query conditions to select tokens that need to be recomputed in the global context, and performs position reconstruction and optional block sorting on retrieved document blocks, thereby improving the accuracy and efficiency of long-text reasoning.

InfoGeo: Information-Theoretic Object-Centric Learning for Cross-View Generalizable UAV Geo-Localization

Hongyang ZHANG, Man On Pun (Chinese University of Hong Kong (Shenzhen))

RetrievalDomain AdaptationRepresentation LearningTransformerAuto EncoderContrastive LearningImagePoint Cloud

🎯 What it does: Designed an InfoGeo framework based on the information bottleneck theory for cross-view geolocation under UAV perspectives.

InfoGlobe: Local-and-Global Information-Preserving Statistical Manifold Learning for Single-Cell Transcriptomics

Cheng Wang (Shanghai Jiao Tong University), Hongyi Xin (Shanghai Jiao Tong University)

Explainability and InterpretabilityRepresentation LearningDrug DiscoveryFlow-based ModelAuto EncoderContrastive LearningBiomedical Data

🎯 What it does: Propose the InfoGlobe method, which models single-cell RNA-seq counts as a multinomial distribution, and achieves geometry-preserving decomposition from high-dimensional gene space to low-dimensional functional groups (factors) space through information geometry (Fisher-Rao metric).

InfoLaw: Information Scaling Laws for Large Language Models with Quality-Weighted Mixture Data and Repetition

Weidong Zhou (ByteDance), Taifeng Wang (ByteDance)

Computational EfficiencyData-Centric LearningTransformerLarge Language ModelMixture of ExpertsText

🎯 What it does: Proposed an information scaling law called InfoLaw, aiming to predict model performance by considering the impact of data quality weighted mixing and repetition on the pretraining of large language models (LLMs).

InfoPO: Information-Driven Policy Optimization for User-Centric Agents

Fanqi Kong (Peking University), Bang Liu (Université de Montréal)

OptimizationTransformerReinforcement LearningPrompt EngineeringContrastive LearningTextSequentialBenchmarkRetrieval-Augmented Generation

🎯 What it does: This paper proposes InfoPO, a reinforcement learning framework tailored for user-centric multi-turn interactions, achieving finer-grained credit assignment by calculating contrastive information gain rewards at each step.

Information dynamics and Memory in Neural Networks through Fisher Information Diffusion

Haodong Qin (University of California San Diego), Tatyana Sharpee

ClassificationOptimizationRepresentation LearningRecurrent Neural NetworkDiffusion modelContrastive LearningImageTextSequential

🎯 What it does: This paper proposes a theoretical framework based on Fisher information diffusion, analyzing how information flows and is retained over time among different subgroups in modular recurrent networks, and designs a Fisher-optimal initialization scheme based on this framework.

Information Flow Reveals When to Trust Language Models

Rui Xu (Hong Kong University of Science and Technology), Sihong Xie (Hong Kong University of Science and Technology)

RetrievalExplainability and InterpretabilityTransformerLarge Language ModelContrastive LearningTextRetrieval-Augmented Generation

🎯 What it does: Proposed an uncertainty quantification framework for retrieval-enhanced language models based on information flow, utilizing a hierarchical contribution matrix to track the influence of contextual words on the final prediction;

Information Geometry Loss for Time Series Forecasting

Jiayu Fang (University of Sydney), Junbin Gao (University of Sydney)

OptimizationTime SeriesBenchmark

🎯 What it does: Proposed an information geometry-based loss function for time series forecasting called InfoGeo Loss.

Information-Geometric Adaptive Sampling for Graph Diffusion

Yuhui Lu (Lanzhou University), Kun Zhan (Lanzhou University)

GenerationDrug DiscoveryGraph Neural NetworkDiffusion modelScore-based ModelGraphTabularStochastic Differential Equation

🎯 What it does: Proposes a drift variation score (DVS) adaptive sampling method based on information geometry for reverse sampling in graph diffusion models.

Information-Theoretic Disentangled Latent Modeling with Conditional Diffusion for Incomplete Multi-View Clustering

Wenlan Chen (Central South University), Fei Guo (Shandong Normal University)

Data SynthesisRepresentation LearningTransformerDiffusion modelAuto EncoderContrastive LearningMultimodality

🎯 What it does: Propose an information theory-based disentangled latent modeling framework called IDCD, and use conditional diffusion models to consistently generate missing views, achieving incomplete multi-view clustering;

Information-Theoretic Generalization Bounds for VAEs: A Role of Encoder and Latent Variable

Futoshi Futami (University of Osaka), Masahiro Fujisawa (University of Osaka)

Information TheoryGenerationRepresentation LearningAuto EncoderImage

🎯 What it does: Proposes the first set of information-theoretic generalization bounds for standard continuous latent variable variational autoencoders (VAEs), and extends this framework to hierarchical VAEs, further providing an upper bound on the 2-Wasserstein distance for generative quality.

Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access

Daniel Ebi (Karlsruhe Institute of Technology), Gaspard Lambrechts (McGill University)

Reinforcement LearningTabularTime SeriesSequential

🎯 What it does: This paper proposes the 'Informational Asymmetric Actor-Critic' framework, which allows the critic to use arbitrary state-related privileged signals during training while maintaining unbiased policy gradients.

InfraRL: A Benchmark for Constrained Resource Allocation in Large-Scale Infrastructure Asset Management

Yantian Wang (Tongji University), Bo Jin (Tongji University)

OptimizationReinforcement LearningTabularBenchmark

🎯 What it does: Propose the InfraRL benchmark for offline constrained resource allocation, based on the National Bridge Inventory (NBI) to construct a large-scale multi-agent bridge maintenance task;

InftyThink+: Effective and Efficient Infinite-Horizon Reasoning via Reinforcement Learning

Yuchen Yan (Zhejiang University), Yongliang Shen (Zhejiang University)

OptimizationComputational EfficiencyTransformerSupervised Fine-TuningReinforcement LearningTextBenchmarkChain-of-Thought

🎯 What it does: This paper proposes a reinforcement learning framework called InftyThink+, which can adaptively decide when to compress and summarize, how to compress, and how to continue reasoning during the iterative reasoning process;

InfVSR: Toward Consistency-Driven Streaming Generative Video Super-Resolution

Ziqing Zhang, Yulun Zhang (Shanghai Jiao Tong University)

Super ResolutionTransformerDiffusion modelScore-based ModelGenerative Adversarial NetworkContrastive LearningOptical FlowVideo

🎯 What it does: Proposed a generative video super-resolution framework called InfVSR that can process infinitely long videos in a streaming manner.

Initialization is Half the Battle: Generating Diverse Images from a Guidance Potential Posterior

Xiang Li (National University of Singapore), Kenji Kawaguchi (National University of Singapore)

GenerationData SynthesisTransformerDiffusion modelScore-based ModelFlow-based ModelRectified FlowImageTextMultimodalityStochastic Differential Equation

🎯 What it does: Propose a diversity initialization method called DivIn based on guided potential posterior, which samples from the initial noise space of diffusion models and flow matching models using Langevin dynamics, significantly enhancing generation diversity without sacrificing image quality.

Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression

Yao DU, Xiaomeng Li (Hong Kong University of Science and Technology)

Data-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningMultimodalityTabularBenchmark

🎯 What it does: A distribution-aware reinforcement learning framework (CCC-GRPO) is proposed to improve numerical prediction of multi-modal large language models in deep imbalanced regression tasks.

Inner-layer Token Self-modulation as Another Scaling Axis for LLMs

Yebin Yang (Shanghai Jiao Tong University), Junchi Yan (Shanghai Jiao Tong University)

Computational EfficiencyRepresentation LearningTransformerLarge Language ModelMixture of ExpertsTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: Designed and implemented two module based on token index parameters — ReToken (for dense Transformer) and MoRT (for sparse Mixture-of-Experts), which enhance model capacity and performance by performing lightweight Hadamard multiplication modulation on the FFN residual of Transformer layers through token-specific modulation vectors retrieved from the embedding table, without significantly increasing FLOPs.

InnoEval: On Research Idea Evaluation as a Knowledge-Grounded, Multi-Perspective Reasoning Problem

Shuofei Qiao (Zhejiang University), Emine Yilmaz (University College London)

Recommendation SystemFederated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextReview/Survey PaperBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed and implemented the InnoEval framework, which is used for systematic evaluation of scientific research innovation ideas in terms of knowledge rooting, multi-dimensional, and multi-perspective aspects;

Innovation: An Almost Characterization of Hallucination

Nishant P. Das (Tata Institute of Fundamental Research), Piyush Srivastava (Tata Institute of Fundamental Research)

Explainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextReview/Survey PaperRetrieval-Augmented Generation

🎯 What it does: This paper studies the hallucination problem in LLMs within the statistical framework of Kalai-Vempala, proposing a new metric called 'innovation' (the probability assigned by the model to unseen sentences) and proving that innovation is almost a necessary and sufficient condition for hallucinations under the K-sparse and conventional fact assumptions; further, using the innovation rate, it provides Markov-style and high-confidence lower bounds for hallucination rates, and associates these lower bounds with missing mass, thus eliminating dependence on the size of the corpus; finally, experiments on n-gram models and public evaluation datasets verify a strong correlation between innovation rate and hallucination rate.

Insertion Based Sequence Generation with Learnable Order Dynamics

Dhruvesh Patel (University of Massachusetts Amherst), Andrew McCallum (University of Massachusetts Amherst)

GenerationDrug DiscoveryTransformerReinforcement LearningDiffusion modelFlow-based ModelGraphSequential

🎯 What it does: Propose Lo FlexMDM, an insertion-based masked diffusion model that can learn data-dependent insertion and unmasked rates to generate variable-length sequences.

Inside the Visual Mind: Neuroscience-Motivated Concept Circuits for Interpreting and Steering Vision Transformers

Tang Li (University of Delaware), Xi Peng (University of Virginia)

Explainability and InterpretabilityRepresentation LearningTransformerVision Language ModelAuto EncoderContrastive LearningImageTextMultimodality

🎯 What it does: Propose the ViSAE toolbox, which utilizes sparse autoencoders (SAE) to provide mechanism-level interpretation of Vision Transformers at the concept level, and offers automatic reading, tracking, and editing of concept circuits, enabling interpretable auditing and precise regulation of ViT behavior.

Instance-Dependent Continuous-Time Reinforcement Learning via Maximum Likelihood Estimation

Runze Zhao (Indiana University Bloomington), Dongruo Zhou (Indiana University Bloomington)

Reinforcement LearningStochastic Differential Equation

🎯 What it does: This paper proposes a continuous-time reinforcement learning algorithm called CT-MLE based on maximum likelihood estimation, which uses the estimation of the state marginal density to guide policy learning and is compatible with any function approximator;

Instance-Level Costs for Nuanced Classifier Evaluation

Kabir Kang (Georgia Institute of Technology), Stephen Mussmann (Georgia Institute of Technology)

ClassificationExplainability and InterpretabilityData-Centric LearningSupervised Fine-TuningContrastive LearningImageTextTabular

🎯 What it does: Proposed an evaluation metric called NEC, which measures classification error based on instance misclassification cost, and explored how to derive cost from annotator disagreement, threshold distance, or confidence.

Instance-Specific Approximation Ratios for Correlation Clustering and Max-Cut

Sebastian Lüderssen (TU Wien), Stefan Neumann (TU Wien)

OptimizationGraph

🎯 What it does: The study investigates instance-specific approximation ratios, proposing an efficient algorithm to compute LP lower bounds for related clustering (CC) and maximum cut (MAX-CUT) problems on sparse graphs in near-linear time, and using these lower bounds to evaluate the performance of heuristic algorithms on real-world data.

InstEmb: Instruction-Following Embeddings through Glimpses of the Future

Tianhao Gao (JD.com), Qixia Jiang (JD.com)

RetrievalKnowledge DistillationRepresentation LearningTransformerPrompt EngineeringContrastive LearningText

🎯 What it does: Propose InstEmb, an instruction-following embedding framework that captures intrinsic input semantics and output-aware semantics simultaneously by utilizing a learnable look-ahead token and self-supervised distillation.

Instruction Decomposition and Action Alignment for Vision-Language Navigation

Zihao Xin (Nanjing University of Aeronautics and Astronautics), Sheng-Jun Huang

Autonomous DrivingComputational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringVision Language ModelVision-Language-Action ModelContrastive LearningImageVideoTextMultimodalityChain-of-Thought

🎯 What it does: Proposes IDEAL-VLN, decomposing visual language navigation into semantic anchoring (thinking) and action alignment, and achieving instruction decoding and action generation through a think-before-acting mechanism and a hierarchical correction mechanism;

Instruction Lens Score: Your Instruction Contributes a Powerful Object Hallucination Detector for Multimodal Large Language Models

Runhe Lai (Sun Yat-sen University), Ruixuan Wang (Sun Yat-sen University)

Anomaly DetectionExplainability and InterpretabilityTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodality

🎯 What it does: This paper studies and proposes an instruction-embedding-based, no-training, plug-and-play method for detecting object hallucinations — Instruction Lens Score (InsLen).

INT vs. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization Formats

Mengzhao Chen (University of Hong Kong), Ping Luo (University of Hong Kong)

Computational EfficiencyRepresentation LearningTransformerLarge Language ModelContrastive LearningTextBenchmark

🎯 What it does: This paper systematically evaluates the performance of low-bit integers (INT) and floating-point (FP) quantization at different fine-grained (block) levels, providing a theoretical framework, experimental results, and hardware cost analysis.

Integrated Episodic and Semantic Memory via Modulating Transformer FeedForward Layers

Yiqun Yao (Beijing Academy of Artificial Intelligence), Yequan Wang (Beijing Academy of Artificial Intelligence)

Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningTextRetrieval-Augmented Generation

🎯 What it does: Propose HyperMem, a memory framework that leverages hypernetworks to directly map context to Transformer FFN parameters, achieving the unification of short-term (episodic) and long-term (semantic) memory.

Intentional Updates for Streaming Reinforcement Learning

Arsalan Sharifnassab (Openmind Research Institute), Richard S Sutton (Openmind Research Institute)

Reinforcement LearningTabularTime SeriesSequentialBenchmark

🎯 What it does: This paper proposes an 'Intentional Update' framework, which automatically calculates the step size using a specified target function change (such as a fixed proportion reduction of the TD error or a fixed increment in the policy's log probability), thereby achieving stable training with single-sample, no-replay in streaming deep reinforcement learning; meanwhile, three instances are provided: Intentional TD, Intentional Q-Learning, and Intentional Policy Gradient, combining techniques such as eligibility traces, RMSProp-style diagonal scaling, and δ clipping to realize the complete algorithm.

IntentRL: Training Proactive User-intent Agents for Open-ended Deep Research via Reinforcement Learning

Haohao Luo (Sun Yat-sen University), Ying Shen (Sun Yat-sen University)

Autonomous DrivingOptimizationRobotic IntelligenceDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This paper trains an active user intent clarification agent, which first engages in multi-round interactions with users to clarify the implicit intent of deep research tasks, and then passes the refined query to the DR agent for long-term research, thereby improving task efficiency and report quality.

InteractBench: Benchmarking LLMs on Competitive Programming under Unrevealed Information

Jiaze Li (Huazhong University of Science and Technology), Xianjun Deng (Huazhong University of Science and Technology)

AI Code AssistantTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This work proposes and constructs InteractBench, a benchmark containing 322 interactive competitive programming problems, aimed at evaluating the interactive programming capabilities of LLMs when no prior information is available.