ICML 2026 Papers — Page 20
International Conference on Machine Learning · 6554 papers
EvoMAS: Evolutionary Generation of Multi-Agent Systems
Yuntong Hu (Emory University), Stefano Soatto (Amazon Web Services)
Autonomous DrivingOptimizationDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringTextTabularBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposes EvoMAS, an evolution-based method leveraging LLMs to automatically generate multi-agent systems (MAS) within structured configuration spaces and continuously optimize configurations through execution feedback.
EvoMAS: Heuristics in the Loop—Evolving Smarter Agentic Workflows
Yangbo Wei (Shanghai Jiao Tong University), WEI W. XING
Autonomous DrivingOptimizationRobotic IntelligenceDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackGraph Neural NetworkTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringMixture of ExpertsDiffusion modelScore-based ModelContrastive LearningTextGraphTabularTime SeriesSequentialBenchmark
🎯 What it does: Propose a bio-inspired evolutionary framework called EvoMAS to automatically design and optimize the workflow of multi-agent systems, integrating role-level evolution, dynamically diversified evolutionary strategies, curriculum learning, and a meta-controller called Cyber Creator, achieving adaptive and scalable agent collaboration processes;
EvReflection: Event-Driven Micro-Dynamics for Reflection Removal
Jiaxiao Wang (University of Science and Technology of China), Xiaoyan Sun (University of Science and Technology of China)
RestorationSpiking Neural NetworkTransformerDiffusion modelContrastive LearningOptical FlowImageVideo
🎯 What it does: Propose EvReflection, an end-to-end network that removes reflections by leveraging micro-dynamics from event cameras.
Exact and Approximate Algorithms for Polytree Learning
Juha Harviainen (University of Helsinki), Manuel Sorge (TU Wien)
OptimizationComputational EfficiencyRepresentation LearningReview/Survey Paper
🎯 What it does: This paper studies exact and approximate algorithms for learning optimal polytrees under a given local scoring function, with a particular focus on in-degree upper bounds, properties of the scoring function, and approximation error.
Exact Functional ANOVA Decomposition for Categorical Inputs Models
Baptiste Ferrere (EDF R&D), Joseph Muré (EDF R&D)
Explainability and InterpretabilityComputational EfficiencyTabular
🎯 What it does: Proposes an exact functional ANOVA decomposition method for discrete (categorical) inputs, providing closed-form expressions that can handle arbitrary dependency structures and non-rectangular supports.
Exact Unlearning in Reinforcement Learning
Thanh Nguyen-Tang (New Jersey Institute of Technology), Raman Arora (Johns Hopkins University)
Safty and PrivacyReinforcement LearningPrompt EngineeringDiffusion modelScore-based ModelFlow-based ModelRectified FlowNeural Radiance FieldAuto EncoderGenerative Adversarial NetworkContrastive LearningGaussian SplattingTabular
🎯 What it does: Proposes a theoretical framework for achieving exact unlearning in reinforcement learning (Tabular MDP), and provides specific algorithms to implement this framework;
Exactly Computing do-Shapley Values
R. Teal Witter (Claremont Mckenna College), Lucas Rosenblatt (Williams College)
Explainability and InterpretabilityComputational EfficiencyScore-based ModelContrastive LearningTabularBenchmark
🎯 What it does: The paper proposes a method to precisely calculate do-Shapley values by utilizing irreducible sets in structural causal models.
Excited Pfaffians: Generalized Neural Wave Functions Across Structure and State
Nicholas Gao (Technical University of Munich), Stephan Günnemann
Drug DiscoveryGraph Neural NetworkDiffusion modelScore-based ModelFlow-based ModelNeural Radiance FieldAuto EncoderGenerative Adversarial NetworkContrastive LearningOptical FlowGraphTabularTime SeriesSequentialPhysics Related
🎯 What it does: A novel neural network wave function model for multi-excited states, called Excited Pfaffian, is proposed within the variational Monte Carlo framework, combined with multi-state importance sampling (MSIS) to achieve nearly constant-scale calculations for excited states.
ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation
Yiran Wu (Pennsylvania State University), Anand Mudgerikar (Microsoft Security AI Research)
Explainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringAuto EncoderTextGraphTabularBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This work constructs ExCyTIn-Bench, a benchmark for evaluating large language model (LLM) agents in network threat investigation tasks; by creating an interactive MySQL environment, automatically generated question-answer pairs, and fine-grained progress rewards, the performance of agents in practical security log querying and reasoning is assessed.
Executable Agentic Memory for GUI Agent
Zerui Qin (Tsinghua University), Ju Ren (Tsinghua University)
Explainability and InterpretabilityComputational EfficiencyGraph Neural NetworkTransformerReinforcement LearningAgentic AIPrompt EngineeringMixture of ExpertsTextGraphRetrieval-Augmented Generation
🎯 What it does: Construct an Executable Agentic Memory (EAM) that encodes GUI interaction logic into a knowledge graph, and retrieves executable paths on the graph through MCTS guided by a Q-model, achieving efficient long-term GUI automation.
ExpAlign: Expectation-Guided Vision–Language Alignment for Open-Vocabulary Grounding
Junyi Hu (Tsinghua University), Yi ZHANG
RecognitionObject DetectionSegmentationTransformerVision Language ModelContrastive LearningImageTextMultimodality
🎯 What it does: Proposed a vision-language alignment framework called ExpAlign based on expectation guidance for open-vocabulary localization tasks under weak supervision.
Expand Neurons, Not Parameters
Linghao Kong (Massachusetts Institute of Technology), Nir N Shavit
ClassificationComputational EfficiencyRepresentation LearningContrastive LearningImageText
🎯 What it does: By increasing the number of neurons under a fixed non-zero parameter budget, reducing feature collisions and ambiguity, thus improving model performance.
Expandable, Compressible, Mineable: Open-World Thermal Infrared Image Restoration
Pu Li (Kunming University of Science and Technology), Jie Wen (Harbin Institute of Technology)
RestorationConvolutional Neural NetworkPrompt EngineeringAuto EncoderContrastive LearningImage
🎯 What it does: Propose ECMRNet, a scalable, compressible, and mineable continual learning framework for thermal infrared images that can adapt to new degradations in open-world environments.
Expanding the AI Evaluation Toolbox with Statistical Models
Drew Keller (National Institute of Standards and Technology), A. Stevie Bergman (National Institute of Standards and Technology)
Explainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelScore-based ModelFlow-based ModelRectified FlowGaussian SplattingTextBenchmarkStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: This study introduces a statistical model into the benchmark evaluation of large language models (LLMs), formally distinguishing between benchmark accuracy and generalization accuracy, and improving uncertainty estimation using a generalized linear mixed model (GLMM).
Expanding the Capabilities of Reinforcement Learning via Text Feedback
Yuda Song (Carnegie Mellon University), Andrea Zanette (Carnegie Mellon University)
Knowledge DistillationRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringContrastive LearningTextBenchmark
🎯 What it does: This paper proposes a new framework for reinforcement learning using text feedback—RL from Text Feedback (RLTF), and designs two methods: Self Distillation (RLTF-SD) and Feedback Modeling (RLTF-FM) to enhance the performance of single-turn LLMs.
Expanding the Chaos: Neural Operator for Stochastic (Partial) Differential Equations
Dai Shi (University of Cambridge), Junbin Gao (University of Sydney)
Graph Neural NetworkDiffusion modelScore-based ModelImageGraphTabularTime SeriesBiomedical DataBenchmarkFinance RelatedPhysics RelatedStochastic Differential Equation
🎯 What it does: Propose a neural operator (NO) framework based on Wiener-chaos expansion, which enables learning the solution operator for SDEs and SPDEs in a single step, projecting the stochastic driver onto orthogonal Wick-Hermite bases, and then using the NO model to learn the deterministic propagator.
Expectation Alignment of Language Models for Real-World User Expectations
Miaomiao Li (Chinese University of Hong Kong), Kam-Fai Wong (Chinese University of Hong Kong)
Recommendation SystemExplainability and InterpretabilityComputational EfficiencyRepresentation LearningData-Centric LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextSequentialBenchmarkRetrieval-Augmented Generation
🎯 What it does: A user expectation extraction method based on real multi-turn dialogues was constructed, the EXPECTBENCH benchmark was created, and the LENS was proposed to improve the response consistency of LLMs through latent expectation representations.
Expectation Consistency Loss: Rethink Confidence Calibration under Covariate Shift
Jinzong Dong (Central South University), Bo Yang (Central South University)
ClassificationDomain AdaptationConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelContrastive LearningImage
🎯 What it does: Propose the expected consistency condition under covariate shift and design an unsupervised domain adaptation loss ECL based on it, used to calibrate the confidence of classification models, supporting standard, class-level, and maximum confidence calibration.
Expected Return Causes Outcome-Level Mode Collapse in Reinforcement Learning and How to Fix It with Inverse Probability Scaling
ABHIJEET SINHA, Dianbo Liu (National University of Singapore)
Drug DiscoveryReinforcement LearningTextMultimodalityBiomedical Data
🎯 What it does: This paper analyzes the phenomenon of result-level mode collapse caused by the expected return objective in reinforcement learning, proposes the inverse probability scaling (IPS) method to eliminate the probability amplification effect, rewrites the reward signal, and implements IPS-GRPO under the group policy gradient framework; subsequently, its effectiveness is verified on various multi-modal tasks.
Expected Returns and Policy Inconsistency-Aware Offline Federated Deep Reinforcement Learning
Meng XU, Jianping Wang (City University of Hong Kong)
Federated LearningReinforcement LearningTabularTime SeriesSequential
🎯 What it does: Proposed a general offline federated deep reinforcement learning framework that utilizes policy inconsistency and Q-values to jointly evaluate the importance of client models, subsequently generating weights through soft-max normalization and attenuating the impact of weak global models during the local training phase.
Experience Augmented Policy Optimization for LLM Reasoning
Jinda Lu (University of Science and Technology of China), Jingren Zhou (Independent Researcher)
OptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Improve the reasoning performance of LLMs by dynamically injecting experience from previously optimized models during the reinforcement learning process, avoiding the full sampling and experience mismatch issues of traditional RLVR methods.
Experience is the Best Teacher: Motivating Effective Exploration in Reinforcement Learning for LLMs
Wenjian Zhang (Computer Network Information Center Chinese Academy of Sciences), Baisheng Lai (Computer Network Information Center Chinese Academy of Sciences)
Reinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringText
🎯 What it does: Propose a reinforcement learning framework called HeRL based on retrospective experience, which utilizes failure trajectories and unmet rubrics as guidance to drive LLMs to generate higher quality answers during the exploration phase and further conduct reinforcement learning.
Experience-Evolving Multi-Turn Tool-Use Agent with Hybrid Episodic–Procedural Memory
Sijia Li (Hong Kong University of Science and Technology), Rui Wang (Microsoft Research Asia)
Autonomous DrivingOptimizationRobotic IntelligenceTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextSequentialBenchmarkRetrieval-Augmented Generation
🎯 What it does: A hybrid episodic-procedural memory strategy (H-EPM) was designed for multi-round tool-using agents, achieving experience retrieval during reasoning and memory guidance during reinforcement learning by simultaneously storing tool call sequences and context summaries in an experience graph.
Expert-guided Clinical Text Augmentation via Query-Based Model Collaboration
Dongkyu Cho (New York University), Rumi Chunara (New York University)
Data SynthesisExplainability and InterpretabilityKnowledge DistillationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBiomedical DataElectronic Health RecordsRetrieval-Augmented Generation
🎯 What it does: Propose a query-based model collaboration framework guided by experts, using a lightweight expert model to constrain the clinical text generated by LLMs, achieving safe and information-preserving text enhancement.
Expert-level Leaf Cell Layout Generation via Preference-Optimized LLM
Yaohui Han (Chinese University of Hong Kong), Tsung-Yi Ho (Chinese University of Hong Kong)
OptimizationReinforcement Learning from Human FeedbackGraph Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningGraph
🎯 What it does: Propose the GenLeaf framework, which utilizes LLM to automatically generate layout scripts for leaf cells and drives placement and routing (PnR) through these scripts. The generated layouts achieve PPA (power, performance, area) metrics comparable to or exceeding those of layouts manually designed by human experts.
ExpertWeaver: Unlocking the Inherent MoE in Dense LLMs with GLU Activation Patterns
Ziyu Zhao (Zhejiang University), Yu Cheng (Chinese University of Hong Kong)
Computational EfficiencyKnowledge DistillationTransformerLarge Language ModelMixture of ExpertsText
🎯 What it does: Training-free sparse MoE conversion for pre-trained dense LLMs
Explainable Federated Learning via Global–Local Attribution Alignment
Dawood Wasif (Virginia Tech), Jin-Hee Cho (Virginia Tech)
OptimizationFederated LearningSafty and PrivacyExplainability and InterpretabilityKnowledge DistillationAuto EncoderContrastive LearningGaussian SplattingImageTextTabularBiomedical DataAlzheimer's DiseaseBenchmarkFinance Related
🎯 What it does: Propose the xFedAlign framework to achieve collaborative task learning and interpretability in federated learning, without sharing raw data or gradients, by only exchanging sparse, noisy top-k attribution summaries, achieving locally trustworthy and globally consistent explanations.
Explainable Forensics of Manipulated Segments in Untrimmed Long Videos
Yue Feng (Nanjing University of Aeronautics and Astronautics), Jie Qin (Nanjing University of Aeronautics and Astronautics)
Anomaly DetectionExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningVideoMultimodality
🎯 What it does: Propose the task of paragraph detection and explanation for long video AI generation, and construct the TASLE large-scale long video dataset.
ExPLAIND: Unifying Model, Data, and Training Attribution to Study Model Behavior
Florian Eichin (LMU Munich), Michael A. Hedderich (LMU Munich)
OptimizationExplainability and InterpretabilityConvolutional Neural NetworkTransformerLarge Language ModelContrastive LearningImageText
🎯 What it does: Proposed the ExPLAIND framework to uniformly attribute the influence of models, data, and training processes to model behavior, and verified its effectiveness in CNNs, Transformers (Grokking), and LLM pretraining.
Explaining Concept Shift with Interpretable Feature Attribution
Ruiqi Lyu (Carnegie Mellon University), Bryan Wilder (Carnegie Mellon University)
Domain AdaptationExplainability and InterpretabilityTabularBiomedical DataElectronic Health Records
🎯 What it does: Studies concept shift (differences in conditional distributions between source and target) and proposes the SGShift method, using the source model as a baseline and only learning sparse corrections to explain and repair performance degradation.
Explaining Data Mixing Scaling Laws
Rui Dai (Beijing Institute of Technology), SHURAN ZHENG
Explainability and InterpretabilityData-Centric LearningLarge Language ModelText
🎯 What it does: Proposes a unified theoretical framework to explain the scalar laws of multi-domain data mixing and predict the optimal mixing ratio.
Explanations are a Means to an End: Decision Theoretic Explanation Evaluation
Ziyang Guo (Northwestern University), Jessica Hullman (Northwestern University)
ClassificationAnomaly DetectionExplainability and InterpretabilityLarge Language ModelContrastive LearningImageTextTabularBiomedical DataElectronic Health RecordsReview/Survey Paper
🎯 What it does: Proposes a decision theory framework that treats explanations as information signals and quantifies their potential improvement on decision performance, defining three estimable indicators: theoretical value, human complementary value, and behavioral value, and providing a complete validation workflow.
Explicit representation of germline and non-germline residues improves antibody language modeling
Jeonghyeon Kim (Duke University), Philip Romero (Duke University)
Drug DiscoveryProtein Structure PredictionTransformerLarge Language ModelPrompt EngineeringContrastive LearningBiomedical Data
🎯 What it does: Propose the PRISM model, which improves antibody language models by explicitly distinguishing between germline (G) and non-germline (NGL) residues.
Explicitly Modeling Censoring Produces Superior Survival Predictors
Shi-ang Qi (University of Alberta), Russell Greiner (University of Alberta)
Computational EfficiencyRepresentation LearningGraph Neural NetworkTransformerContrastive LearningTabularTime SeriesBiomedical DataElectronic Health RecordsBenchmark
🎯 What it does: Propose a survival prediction framework that explicitly models the censoring process, treating event time and censoring time as two processes sharing parameters.
Exploiting weight-space symmetries for approximating curvature
Artem Artemev (MediaTek), Alberto Bernacchia (MediaTek)
OptimizationConvolutional Neural NetworkTransformerImageText
🎯 What it does: Structural Hessian approximation is achieved by orbit averaging a single gradient using weight space symmetry, enabling efficient curvature estimation.
Exploration Hacking: Can LLMs Learn to Resist RL Training?
Eyon Jang (MATS), David Lindner (Google DeepMind)
Drug DiscoveryAI Code AssistantNeural Architecture SearchTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextTabularTime SeriesBiomedical DataBenchmarkChain-of-Thought
🎯 What it does: This paper investigates the 'exploration hacking' phenomenon where LLMs hinder training in RL by reducing exploration behavior, and constructs model organisms to simulate this behavior.
Exploration-free Algorithms for Multi-group Mean Estimation
Ziyi Wei (Virginia Tech), Xiaocheng Li (Imperial College London)
OptimizationFederated LearningReinforcement LearningTabular
🎯 What it does: This paper studies the problem of multi-group mean estimation, aiming to allocate a limited sampling budget among multiple groups to achieve uniformly accurate mean estimation.
Exploring 3D Dataset Pruning
Xiaohan Zhao (Mohamed bin Zayed University of Artificial Intelligence), Zhiqiang Shen (Mohamed bin Zayed University of Artificial Intelligence)
ClassificationKnowledge DistillationRepresentation LearningData-Centric LearningContrastive LearningPoint CloudMesh
🎯 What it does: This paper proposes a pruning framework called 3D-Pruner for 3D datasets, addressing the conflict between overall accuracy (OA) and mean accuracy (mAcc) under long-tailed distributions.
Exploring Accurate and Transparent Domain Adaptation in Predictive Healthcare via Concept-Grounded Orthogonal Inference
Pengfei Hu (Stevens Institute of Technology), Yue Ning (Stevens Institute of Technology)
Domain AdaptationExplainability and InterpretabilityRepresentation LearningSupervised Fine-TuningAuto EncoderContrastive LearningTabularElectronic Health Records
🎯 What it does: Propose the ExtraCare framework, which decomposes patient representations into label-related immutable subspaces and domain-related mutable subspaces to achieve precise and interpretable domain adaptation;
Exploring and Exploiting Stability in Latent Flow Matching
Rania Briq (Forschungszentrum Jülich), Stefan Kesselheim (Forschungszentrum Jülich)
GenerationComputational EfficiencyData-Centric LearningTransformerDiffusion modelFlow-based ModelAuto EncoderContrastive LearningImageOrdinary Differential Equation
🎯 What it does: Investigate and utilize the stability of the Latent Flow Matching (LFM) model under perturbations in data subsets, model capacity, and training configurations, proposing data pruning methods and a two-stage inference acceleration technique from coarse to fine.
Exploring Data-Free LoRA Transferability for Video Diffusion Models
Yuchen Wang (Hong Kong University of Science and Technology Guangzhou), Zeke Xie (Hong Kong University of Science and Technology Guangzhou)
GenerationKnowledge DistillationTransformerDiffusion modelContrastive LearningVideo
🎯 What it does: Studied the compatibility issues that arise when migrating LoRA to distilled models in video diffusion models (VDM), and proposed a data-agnostic Cluster-Aware Spectral Arbitration (CASA) method to achieve training-free migration of LoRA.
Exploring More to Solve More: Boosting Diversity in Text Diffusion Models via Entropy-Based Guidance
Jingwei Zhang (Chinese University of Hong Kong), Farzan Farnia (Chinese University of Hong Kong)
GenerationData SynthesisComputational EfficiencyTransformerPrompt EngineeringDiffusion modelScore-based ModelContrastive LearningText
🎯 What it does: Propose a training-agnostic semantic-aware kernel entropy guidance method (SAKE) to balance diversity and quality in text diffusion models
Exploring Motif-based Heterogeneous Graph Learning for ReDoS Detection
Hong Huang (Chinese Academy of Sciences), Guiyi He (Chinese Academy of Sciences)
Anomaly DetectionComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerSupervised Fine-TuningContrastive LearningTextGraph
🎯 What it does: Developed a ReDoS detection framework RMGNN based on graph neural networks, which utilizes heterogeneous graphs (HRG) of regular expressions and counts of related substructures for prediction.
Exploring Nonlinear Pathway in Parameter Space for Machine Unlearning
Yingdan Shi (Illinois Institute of Technology), Ren Wang (Illinois Institute of Technology)
ClassificationSafty and PrivacyDiffusion modelScore-based ModelContrastive LearningImage
🎯 What it does: Proposes a nonlinear path exploration framework called MCU based on pattern connectivity, achieving efficient amnesia in machine learning models.
Expo-GS: Exposure-Aware Signed Distance Function in Gaussian Splatting for High Dynamic Range
Chaoda Song (Case Western Reserve University), Vipin Chaudhary (Case Western Reserve University)
GenerationData SynthesisDiffusion modelScore-based ModelNeural Radiance FieldGaussian SplattingImagePoint Cloud
🎯 What it does: Propose the Expo-GS framework, combining exposure-aware Signed Distance Function (Expo-SDF) with 3D Gaussian Splatting, to achieve high dynamic range novel view synthesis, capable of simultaneously optimizing the radiance field and geometric field.
Exposing Hidden Biases in Text-to-Image Models via Automated Prompt Search
Manos Plitsis (National and Kapodistrian University of Athens), Yannis Panagakis (National and Kapodistrian University of Athens)
GenerationExplainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringDiffusion modelImageText
🎯 What it does: Propose an automated prompt search framework (BGPS) based on large language models and internal attribute classifiers, which can generate prompts that maintain text naturalness while significantly amplifying implicit social biases in text generation models.
Exposing Vulnerabilities in Explanation for Time Series Classifiers via Dual-Target Attacks
Bohan Wang (Emory University), Wei Jin (Emory University)
Explainability and InterpretabilityAdversarial AttackTime SeriesBiomedical DataElectrocardiogramBenchmark
🎯 What it does: This paper proposes the TSEF dual-objective attack framework, which can induce time series classifiers to produce target predictions while maintaining the consistency of explanations.
Expressive Graph Neural Networks via Equivariant Use of Noise
Xiyuan Wang (Peking University), Muhan Zhang (Peking University)
Computational EfficiencyRepresentation LearningDrug DiscoveryGraph Neural NetworkContrastive LearningGraph
🎯 What it does: Propose the Equivariant Noise GNN (ENGNN) framework, which enhances the expressive power of graph neural networks by injecting noise onto nodes and processing the noise with equivariance.
Expressivity-Efficiency Tradeoffs for Hybrid Sequence Models
John Cooper (University of Wisconsin), Frederic Sala (University of Wisconsin)
Computational EfficiencyRepresentation LearningTransformerLarge Language ModelMixture of ExpertsSequential
🎯 What it does: This paper studies hybrid sequence models that combine Transformer and state space models (SSMs), and theoretically and experimentally demonstrates their superiority over single-architecture models on certain core tasks (such as selective copy and associative memory).
ExpWeaver: LLM Agents Learn from Experience via Latent RAG
Tao Feng (University of Illinois Urbana Champaign), Jiaxuan You (University of Illinois Urbana Champaign)
RetrievalRecommendation SystemData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposes the ExpWeaver framework, enabling LLM agents to retrieve and utilize past interaction experiences in the latent space, eliminating the need for traditional RAG modules;
Extending Fair Null-Space Projections for Continuous Attributes to Kernel Methods
Felix Störck (Bielefeld University), Barbara Hammer (Bielefeld University)
OptimizationComputational EfficiencyRepresentation LearningTabular
🎯 What it does: This paper addresses the problem of fair regression for continuous attributes by proposing a model-free technique that extends the null-space projection to kernel methods, directly removing predictive information about protected attributes in the empirical feature space;
Extending Prediction-Powered Inference through Conformal Prediction
Daniel Csillag (Fundação Getúlio Vargas), Guilherme Tegoni Goedert (Fundação Getúlio Vargas)
Explainability and InterpretabilityData-Centric LearningTabularTime SeriesBiomedical DataBenchmark
🎯 What it does: Proposed a general framework that combines prediction-driven inference with adaptive prediction (conformal prediction). By using prediction sets trained on a calibration set, it fills in missing data and corrects bias, thereby constructing valid confidence intervals and online tests for tasks such as mean estimation, Z/M estimation, and e-value inference.
Extra-Merge: Tracing the Rank-1 Subspace of Model Merging in Language Model Pre-Training
WenJie Zhou, Xueqi Cheng (University of Chinese Academy of Sciences)
OptimizationComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelContrastive LearningText
🎯 What it does: Geometric analysis of the trajectory of merged checkpoints during the pre-training process of large language models reveals that merged checkpoints cluster in a one-dimensional Rank-1 subspace. Based on this finding, the paper proposes an 'Extra-Merge' linear extrapolation method that does not require gradient updates.
Extracting alignment data in open models
Federico Barbero (Google DeepMind), Jamie Hayes (Google DeepMind)
Knowledge DistillationReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextRetrieval-Augmented Generation
🎯 What it does: This paper demonstrates that open-weight large language models can be induced to output alignment (SFT/RL) data by using chat templates, and utilizes high-quality embedding retrieval methods to measure the model's approximate semantic memory.
ExVerus: Verus Proof Repair via Counterexample Reasoning
Jun Yang (University of Chicago), Kexin Pei (University of Chicago)
OptimizationExplainability and InterpretabilityAI Code AssistantTransformerLarge Language ModelPrompt EngineeringTextBenchmark
🎯 What it does: Propose the EXVERUS framework, which utilizes LLM to generate, verify, and leverage source-level counterexamples during the proof repair process in Verus.
Eyes-on-Me: Scalable RAG Poisoning through Transferable Attention-Steering Attractors
Yen-Shan Chen (CyCraft AI Lab), Yun-Nung Chen (National Taiwan University)
Adversarial AttackTransformerPrompt EngineeringTextRetrieval-Augmented Generation
🎯 What it does: Propose an scalable RAG data poisoning framework called EYES-ON-ME, which utilizes reusable attention attractors and insertable Focus regions to induce the retriever and generator to produce malicious outputs.
FAB: A First-Order AB-based Gradient Algorithm for Distributed Bilevel Optimization over Time-Varying Directed Graphs
Yaoshuai Ma (Peng Cheng Laboratory), Jin Zhang (Peng Cheng Laboratory)
OptimizationFederated LearningReinforcement LearningImageTextTabularTime Series
🎯 What it does: This paper proposes FAB, a full-gradient push-pull algorithm AB/Push-Pull first implemented on time-varying directed graphs, for distributed two-layer optimization.
FACT: Fuzzy Alignment with Comorbidity Topology for Reliable Multi-Label Medical Image Diagnosis
Yingyu Chen (Sichuan University), Yi Zhang (Sichuan University)
ClassificationRecognitionAnomaly DetectionExplainability and InterpretabilityConvolutional Neural NetworkGraph Neural NetworkTransformerAuto EncoderContrastive LearningImageGraphBiomedical DataMagnetic Resonance ImagingComputed TomographyElectronic Health Records
🎯 What it does: Propose the FACT framework, which treats multi-label medical image diagnosis as a fuzzy alignment problem between atomic visual evidence and disease semantic anchors, addressing the limitations of hard segmentation caused by visual ambiguity and disease associations.
FactGuard: Agentic Video Misinformation Detection via Reinforcement Learning
Zehao Li (Institute of Computing Technology, Chinese Academy of Sciences), Zhaoqi Wang
Explainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningAgentic AIVision Language ModelVideoMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposed a FactGuard agentic framework based on a multi-modal large language model for video rumor detection, which employs iterative reasoning and tool calls to obtain external evidence and gradually refine judgments.
Factor-Wise Homogeneity of Slot-Attention for Continual Object-Centric Learning
Ilmin Kang (Gwangju Institute of Science and Technology), Kangil Kim (Gwangju Institute of Science and Technology)
Representation LearningTransformerDiffusion modelContrastive LearningImage
🎯 What it does: This paper studies object-centric learning in a continual learning scenario, proposing Slot Attention to naturally form 'Factor-Wise Homogeneity' in the latent space, and utilizing this property to learn new tasks while preserving knowledge of old tasks.
Factored Causal Representation Learning for Robust Reward Modeling in RLHF
Yupei Yang (Shanghai Jiao Tong University), Lei Xu (Shanghai Jiao Tong University)
OptimizationExplainability and InterpretabilityRepresentation LearningReinforcement Learning from Human FeedbackTransformerSupervised Fine-TuningReinforcement LearningAuto EncoderGenerative Adversarial NetworkContrastive LearningText
🎯 What it does: This paper proposes a causal decomposition-based reward model framework called CausalRM, which splits the context embedding into causal and non-causal parts. It specifically uses the causal component for reward prediction, and further suppresses reward leakage related to rewards by backpropagating adversarial gradients through the non-causal component, thereby inhibiting reward deception in RLHF.
Factored Classifier-Free Guidance
Tian Xia (Imperial College London), Ben Glocker (Imperial College London)
GenerationData SynthesisConvolutional Neural NetworkTransformerPrompt EngineeringDiffusion modelScore-based ModelImageBiomedical DataComputed TomographyElectronic Health RecordsReview/Survey Paper
🎯 What it does: To address the adversarial generation task of differential generation, this paper proposes a group-based classifier-free guidance method (FCFG), which assigns different weights to attribute groups (intervened and non-intervened) during the inference phase, in order to eliminate the attribute amplification problem caused by traditional CFG.
Factored Gossip DiLoCo: Reducing Blocking Communication within DiLoCo
Chamin P Hewa Koneputugodage, Alexander Long (Pluralis Research)
OptimizationFederated LearningComputational EfficiencyTransformerLarge Language ModelMixture of ExpertsContrastive LearningText
🎯 What it does: This paper proposes Factored Gossip DiLoCo, which significantly reduces blocking communication and improves computational utilization while maintaining optimization stability by splitting the original DiLoCo external synchronization into two steps: non-blocking Mix1 that can overlap with computation and blocking Mix2.
Factored Latent Action World Models
Zizhao Wang (University Of Texas At Austin), Peter Stone (University Of Texas At Austin)
GenerationAutonomous DrivingTransformerDiffusion modelAuto EncoderContrastive LearningWorld ModelImageVideo
🎯 What it does: This paper proposes a factorized latent action model called FLAM, which achieves more accurate world model learning and controllable video generation in multi-entity scenarios by decomposing the scene into independent factors (slots) and learning independent latent actions for each factor.
Factored Value Functions for Graph-Based Multi-Agent Reinforcement Learning
Ahmed Rashwan (University of Bath), Lisa Maria Kreusser
Graph Neural NetworkReinforcement LearningDiffusion modelGraph
🎯 What it does: Propose Diffusion Value Function (DVF) and the Diffusion A2C (DA2C) algorithm based on it, and design Learned DropEdge GNN (LD-GNN) to achieve distributed communication and decision-making
Factorized Scheduling Principle: Learning Interpretable and Transferable Policies via Structured Additive Functions
Hong Je-Gal (Sejong University), Hyun-Suk Lee (Sejong University)
OptimizationExplainability and InterpretabilitySpiking Neural NetworkTransformerReinforcement LearningTabularTime Series
🎯 What it does: Propose the Factorized Scheduling Principle (FSP) framework, which learns scheduling policies through interpretable transferable priority functions.
Facts in Stats: Impacts of Pretraining Diversity on Language Model Generalization
Tina Behnia (University of British Columbia), Christos Thrampoulidis (University of British Columbia)
Data SynthesisExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningText
🎯 What it does: Propose a controllable synthetic testing platform to study the impact of pre-training diversity on the generalization of language models along two learning streams: statistical and factual, and systematically analyze the effects of different context structures and diversity levels on ID and OOD performance.
FAFO: Lossy KV Cache Compression for Lossless Inference Acceleration via Draftless Fumble Decoding
Hoang Anh Duy Le (Rice University), Xia Hu (Rice University)
Computational EfficiencyAI Code AssistantTransformerLarge Language ModelPrompt EngineeringAuto EncoderTextBenchmark
🎯 What it does: This paper proposes the FAFO framework, which achieves lossless inference acceleration while compressing the KV cache.
FAIL: Flow Matching Adversarial Imitation Learning for Image Generation
Yeyao Ma (Shanghai Jiao Tong University), Weidi Xie (Shanghai Jiao Tong University)
GenerationAdversarial AttackReinforcement Learning from Human FeedbackTransformerSupervised Fine-TuningReinforcement LearningDiffusion modelScore-based ModelFlow-based ModelContrastive LearningImageVideoTextBenchmark
🎯 What it does: View the post-training of flow matching models as adversarial imitation learning, and propose the FAIL framework, which does not require explicit rewards or preference pairs, and uses a discriminator to minimize the gap between the policy and the expert distribution.
Failure is Feedback: History-Aware Backtracking for Agentic Traversal in Multimodal Graphs
Joohyung Yun (POSTECH), Wook-Shin Han (POSTECH)
RetrievalGraph Neural NetworkTransformerLarge Language ModelAgentic AIPrompt EngineeringMultimodalityGraphRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose a proxy-based multi-modal graph retrieval method that integrates history-aware backtracking and cost-aware strategy progression based on failure feedback;
Failure-Driven Workflow Refinement
Jusheng Zhang (Sun Yat-sen University), Keze Wang (Sun Yat-sen University)
OptimizationAI Code AssistantTransformerLarge Language ModelPrompt EngineeringTextGraphBenchmark
🎯 What it does: Proposes a workflow optimization framework called CE-Graph, which is based on failure distribution, utilizing failure signature space, clustering, and a 'proposal-validation' iterative loop to improve LLM workflows.
Fair Classification with Efficient and Post-hoc Controllable Fairness-Accuracy Trade-off
Maaya Sakata (University of Tsukuba), Kazuto Fukuchi (University of Tsukuba)
ClassificationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningSupervised Fine-TuningContrastive LearningImageTabular
🎯 What it does: Propose a novel fair classification algorithm called GFB (Guidance to Fairest-Boundary), which learns feature representations during training so that the model after post-processing can maintain efficient performance at different fairness-accuracy trade-off points, and can adjust the trade-off parameters without retraining after deployment.
Fair Dataset Distillation via Cross-Group Barycenter Alignment
Mohammad Hossein Moslemi (Western University), Boyu Wang (Western University)
ClassificationFederated LearningExplainability and InterpretabilityKnowledge DistillationData-Centric LearningConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageTabular
🎯 What it does: This paper proposes a fair dataset distillation framework called COBRA, which significantly reduces the fairness gap between subgroups by performing balanced cross-group barycenter alignment of representations for different subgroups during the distillation process, while maintaining overall performance.
Fair Decisions from Calibrated Scores: Achieving Optimal Classification While Satisfying Sufficiency
Etam Benger (Hebrew University of Jerusalem), Katrina Ligett (Hebrew University of Jerusalem)
ClassificationOptimizationFederated LearningSafty and PrivacyExplainability and InterpretabilityScore-based ModelTabularFinance Related
🎯 What it does: For binary classification problems with limited scores and group calibration, this paper derives the exact solution for the optimal (possibly randomized) classifier that satisfies the sufficiency constraint, and proposes a post-processing algorithm based on geometric features, which can achieve the optimal fair decision using only scores and group information.
Fair Transit Stop Placement: A Clustering Perspective and Beyond
Haris Aziz (University of New South Wales), Jeremy Vollen (Northwestern University)
OptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyPoint CloudGraphTabularBenchmark
🎯 What it does: Study the bus stop placement problem in general metric spaces, propose fairness constraints (core, just representative) and design corresponding algorithms.
FAIR-Calib: Frontier-Aware Instability-Reweighted Calibration for Post-Training Quantization of Diffusion Large Language Models
Haoyu Huang (Beihang University), Baochang Zhang (Beihang University)
Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelDiffusion modelAuto EncoderContrastive LearningTextBenchmark
🎯 What it does: Propose a post-training quantization method called FAIR‑Calib specifically designed for diffusion-based large language models, addressing issues caused by quantization such as instability in writing frontiers and error amplification.
Fair-FedMOE: Group-Fair One-Shot Federated Learning via Prototype-Guided Experts for Medical Imaging Analysis
Lingzhao Meng (China University of Petroleum (East China)), Tao Chen (China University of Petroleum (East China))
ClassificationImage TranslationAnomaly DetectionFederated LearningSupervised Fine-TuningMixture of ExpertsContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasoundElectronic Health Records
🎯 What it does: Designed and implemented a system called Fair-FedMOE, which achieves group fairness using a medical imaging foundation model within the one-shot federated learning (One-Shot Federated Learning) framework.
FairGB: A Fair Granular-Ball Generation Method for Data Classification
Qifen Yang (Jinan University), Lin Cui (Jinan University)
ClassificationFederated LearningSafty and PrivacyExplainability and InterpretabilityData-Centric LearningDiffusion modelScore-based ModelGenerative Adversarial NetworkContrastive LearningTabularBenchmark
🎯 What it does: Designed the FairGB framework to achieve fair classification based on Granular Balls;
FairJudge : An Adaptive, Debiased, and Consistent LLM-as-a-Judge
Bo Yang (Zhejiang University), Shijian Li
Federated LearningExplainability and InterpretabilityData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningTextMultimodalityBenchmark
🎯 What it does: Propose FairJudge, an adaptive, bias-free, and consistent LLM evaluation model that significantly improves the reliability and fairness of evaluations.
FairMerging: Rethinking Model Merging through the Lens of Fairness
Bing Liu (Huazhong University of Science and Technology), Xianjun Deng (Huazhong University of Science and Technology)
ClassificationFederated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationConvolutional Neural NetworkTransformerContrastive LearningImage
🎯 What it does: Studied the impact of model merging on subgroup fairness and proposed a two-stage fairness-aware merging method called FairMerging.
Fairness in Aggregation: Optimal Top-$k$ and Improved Full Ranking
Diptarka Chakraborty (National University Of Singapore), Alvin Hong Yao Yan (National University Of Singapore)
Recommendation SystemOptimizationTabular
🎯 What it does: This paper studies the problem of fair ranking aggregation under the Spearman footrule distance, providing a polynomial optimal algorithm for top-k fair ranking, and proposing a 2-approximation algorithm for full ranking.
FairRARI: A Plug and Play Framework for Fairness-Aware PageRank
Emmanouil Kariotakis (KU Leuven), Aritra Konar (KU Leuven)
OptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyRepresentation LearningGraph Neural NetworkGraph
🎯 What it does: Propose a framework named FairRARI, which utilizes a variational formulation of PageRank to achieve pluggable solutions for group fairness across multiple groups;
FairSSL: Fair Multimodal Self-Supervised Learning
Jiaee Cheong (Harvard University), Sinan Kalkan (METU)
Federated LearningSafty and PrivacyExplainability and InterpretabilityRepresentation LearningAdversarial AttackTransformerAuto EncoderContrastive LearningImageTextMultimodalityTabularBiomedical DataElectronic Health RecordsAudio
🎯 What it does: Proposed a fair self-supervised learning framework called FairSSL, designed for heterogeneous, variable-length multimodal data, aiming to improve fairness while maintaining performance on downstream tasks.
Faithful Mobile GUI Agents with Guided Advantage Estimator
Haowen Hu (Shanghai Jiao Tong University), Zhuosheng Zhang (Shanghai Jiao Tong University)
Explainability and InterpretabilityComputational EfficiencyAdversarial AttackReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AIPrompt EngineeringVision Language ModelImageTextMultimodality
🎯 What it does: Propose Faithful-Agent, which is trained through two stages (SFT + GRPO + GuAE) to train a mobile GUI agent, enabling it to stop or retreat when facing missing or conflicting evidence, thus avoiding fundamental execution.
Faithful Relational Reasoning with Region-based Embeddings: Expressivity of Convex Coordinate-wise Models
Victor Charpenay (Mines Saint-Etienne), Steven Schockaert (Cardiff University)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningGraph Neural NetworkContrastive LearningGraphBenchmark
🎯 What it does: Studied the expressiveness and limitations of convex coordinate-based region embeddings in expressing closed path rules
FakeWorld 1.0: An Omni-modal Benchmark for Fake Media and Content
Yifeng Gao (Fudan University), Yu-Gang Jiang (Fudan University)
Explainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringImageVideoTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-ThoughtAudio
🎯 What it does: Propose FakeWorld 1.0, a multimodal deception detection benchmark, and design the OmniCheck agent framework for joint authenticity and truth verification.
Falling Trees: A Model Class for Interpretable Risk Prioritization
Varun Babbar (Duke University), Cynthia Rudin (Duke University)
ClassificationExplainability and InterpretabilityTabularBiomedical DataElectronic Health Records
🎯 What it does: Proposed the falling trees model class and designed the GRAVITree algorithm to learn its Rashomon set, achieving interpretable risk priority prediction;
Fantastic Reasoning Behaviors and Where to Find Them: Unsupervised Discovery of the Reasoning Process
Zhenyu Zhang (University of Texas at Austin), Lun Wang (Google DeepMind)
Explainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelPrompt EngineeringAuto EncoderTextChain-of-Thought
🎯 What it does: Researchers proposed an unsupervised framework called RISE, which uses sparse autoencoders to learn 'reasoning vectors' from the chain-of-thought activations of LLMs, and achieves controllable intervention in the model's reasoning behavior by injecting these vectors.
FaPS: A General and Fast Training Method for Diffusion Models
Xianglu Wang (University of Science and Technology of China), Hu Ding (University of Science and Technology of China)
GenerationComputational EfficiencyTransformerReinforcement LearningDiffusion modelContrastive LearningImage
🎯 What it does: Propose the FaPS method, which accelerates diffusion model training through frequency-aware patch selection.
Fast and Accurate Causal Parallel Decoding using Jacobi Forcing
Lanxiang Hu (University of California San Diego), Hao Zhang (University of California San Diego)
Computational EfficiencyKnowledge DistillationAI Code AssistantTransformerLarge Language ModelText
🎯 What it does: Based on autoregressive (AR) large language models (LLMs), the Jacobi Forcing training method is proposed, directly transforming AR models into efficient parallel multi-token decoders;
Fast and Expressive Multi-Byte Prediction with Probabilistic Circuits
Andreas Grivas (University of Edinburgh), Antonio Vergari (University of Edinburgh)
GenerationComputational EfficiencyTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningText
🎯 What it does: Introduce the MTPC framework for multi-byte prediction in byte-level large language models (LLMs), which uses probabilistic circuits (PC) to model the joint distribution of future byte windows and combines with autoregressive (AR) models using speculative decoding to ensure generation quality.
Fast and Highly Expressive Policy Learning for Offline Reinforcement Learning via Bootstrapped Flow Q-Learning
Thanh Xuan Nguyen (KAIST), Chang D. Yoo (KAIST)
Reinforcement LearningFlow-based ModelTabularTime SeriesBenchmark
🎯 What it does: Proposed a single-step action generation strategy called BFQ based on Bootstrapped Flow for offline reinforcement learning.
Fast and Near-Optimal Algorithms for Private Hypothesis Selection
Hilal Asi (Apple), Hongjie Chen (ETH Zurich)
Safty and PrivacyComputational Efficiency
🎯 What it does: A two-stage private hypothesis selection algorithm is proposed, which, under the premise of differential privacy, uses a loss function composed of a small number of strong hypotheses, achieving near-optimal sample complexity and nearly linear time complexity.
Fast and Scalable Analytical Diffusion
Xinyi Shang (University College London), Zhiqiang Shen (Mohamed bin Zayed University of Artificial Intelligence)
GenerationData SynthesisComputational EfficiencyDiffusion modelScore-based ModelImageRetrieval-Augmented Generation
🎯 What it does: This paper proposes a training-agnostic dynamic golden subset retrieval framework called GOLDDIFF, aimed at accelerating the inference of analytical diffusion models.
Fast Byte Latent Transformer
Julie Kallini (FAIR at Meta), Srini Iyer (FAIR at Meta)
GenerationData SynthesisComputational EfficiencyTransformerLarge Language ModelDiffusion modelAuto EncoderText
🎯 What it does: Proposes multiple fast generation methods based on the byte-level BLT model, including block-level diffusion BLT-D, self-speculation-based BLT-S, and diffusion + verification BLT-DV;
Fast Estimation for Forest Matrix of Signed Graphs
Haoxin Sun (Fudan University), Zhongzhi Zhang (Fudan University)
OptimizationComputational EfficiencyGraph Neural NetworkGraph
🎯 What it does: This paper studies the fast estimation of the forest matrix of signed graphs, proposes the signed forest matrix theorem, designs a generation algorithm called GSCF based on positive cycle-erasing random walks, and introduces two low-variance sampling estimation methods, FMDE and FMDE+, as well as an opinion estimation algorithm called FJOE applied to the signed Friedkin-Johnsen model.
Fast k-means Seeding Under The Manifold Hypothesis
Poojan Chetan Shah (Indian Institute of Technology Delhi), Ragesh Jaiswal (Indian Institute of Technology Delhi)
OptimizationComputational EfficiencyRepresentation LearningContrastive LearningImageTextMultimodalityTabular
🎯 What it does: This paper proposes a fast k-means random seeding algorithm called Qkmeans based on the manifold assumption, which improves the sampling process of k-means++ by using rejection sampling.
Fast kernel methods: Sobolev, physics-informed, and additive models
Nathan Doumèche (Sorbonne University), Claire Boyer (Universite Paris-Saclay)
OptimizationComputational EfficiencyTabularTime SeriesPhysics Related
🎯 What it does: This paper proposes a scalable kernel regression framework with O(n log n) complexity, achieved by leveraging Fourier representation and the non-uniform fast Fourier transform (NUFFT).
Fast KV Compaction via Attention Matching
Adam Zweiger (Massachusetts Institute of Technology), Yoon Kim (Massachusetts Institute of Technology)
CompressionComputational EfficiencyTransformerLarge Language ModelTextMultimodalityBenchmarkRetrieval-Augmented Generation
🎯 What it does: This study proposes a fast KV cache compression method based on Attention Matching, which directly approximates the original attention output and attention quality by optimizing the compressed key-value pairs, maintaining the model's performance in long contexts.
Fast Mixing Steady-State Control in Markov Decision Processes
Federico Corso (Politecnico di Milano), Alberto Maria Metelli (Politecnico di Milano)
OptimizationReinforcement LearningTabular
🎯 What it does: This paper aims to transform the steady-state control problem in control theory into a Markov decision process (MDP) framework, proposing the Fast Mixing Steady-State (FMSS) control problem, with the goal of synthesizing a Markov policy that induces the target steady-state distribution at the fastest convergence speed.