arXivSub Start free trial

ICML 2026 Papers — Page 5

International Conference on Machine Learning · 6554 papers

Anytime Safe PAC Efficient Reasoning

Chengyao Yu (Southern University of Science and Technology), Bingyi Jing (Chinese University of Hong Kong, Shenzhen)

Computational EfficiencyReinforcement Learning from Human FeedbackReinforcement LearningPrompt EngineeringMixture of ExpertsTextBenchmarkChain-of-Thought

🎯 What it does: Propose B-PAC reasoning, an online selective reasoning method that achieves PAC safe control at any moment under partial feedback, significantly reducing reasoning costs.

Anytime-Valid Inference for Online Ranking of Large Language Models

Runzhe Gu (Zhejiang University), Xintao Xia (Zhejiang University)

Recommendation SystemOptimizationExplainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelContrastive LearningText

🎯 What it does: Propose the SERPANT framework, which utilizes e-process for timely and valid multiple testing to online evaluate and rank large language models, combined with adaptive sampling to achieve early stopping.

AOEB: Benchmarking Agent-Oriented Multimodal Embeddings

Xin Zhang (Harbin Institute of Technology), Min Zhang (Harbin Institute of Technology)

RetrievalRepresentation LearningTransformerLarge Language ModelVision Language ModelImageTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposes AOEB (Agent-Oriented Embedding Benchmark), a multi-task, multi-modal evaluation benchmark specifically designed for the retrieval needs of LLM agents;

AOEPT: Breaking the Implicit Modality-Reduction Bottleneck in Modality-Missing Prompt Tuning

Jian Lang (University of Electronic Science and Technology of China), Fan Zhou (University of Electronic Science and Technology of China)

ClassificationTransformerPrompt EngineeringContrastive LearningImageTextMultimodality

🎯 What it does: Propose a lightweight prompt-tuning framework named AOEPT, aimed at overcoming the implicit modal reduction bottleneck (IMR) in multi-modal Transformers under missing modal conditions, and restoring the inference scope of missing modalities by injecting modal contextualized prompts (MCP) into the model.

AOrchestra: Automating Sub-Agent Creation for Agentic Orchestration

Jianhao Ruan (DeepWisdom), Jiayi Zhang (DeepWisdom)

Autonomous DrivingOptimizationAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose the AORCHESTRA framework, which dynamically generates sub-agents through a unified 4-tuple (instruction, context, tool, model) to achieve task decomposition and execution.

APE-Bench: Evaluating Automated Proof Engineering for Formal Math Libraries

Huajian Xin (ByteDance Seed), Wenda Li (University of Edinburgh)

TransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed the Automated Proof Engineering (APE) framework, which includes automated task extraction (APE-Bench), a unified execution and evaluation infrastructure (APE-Harness), and a multi-version content deduplication and retrieval service; conducted a systematic evaluation of proof engineering tasks at the repository level on the formal mathematics library (Mathlib).

APEX: Approximate-but-exhaustive search for ultra-large combinatorial synthesis libraries

Aryan Pedawi (Numerion Labs), Izhar Wallach (Numerion Labs)

Drug DiscoveryGraph Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningTextGraphTabularBenchmark

🎯 What it does: This paper proposes the APEX protocol, which uses a neural network surrogate and factorization techniques to perform approximate but complete searches on ultra-large combinatorial synthesis libraries (CSL), enabling the evaluation of billions of molecules in one go on a GPU.

API: Adaptive Prototype Imputation for Incomplete Multimodal Sentiment Analysis

Xiaotao Wang (Wuhan University), Mang Ye (Wuhan University)

ClassificationComputational EfficiencyRepresentation LearningTransformerContrastive LearningTextMultimodalityRetrieval-Augmented Generation

🎯 What it does: Propose a prototype-based missing modality sentiment analysis method (API), which fills in the missing modality by retrieving and modulating class-level prototypes;

APIC: Orthogonalized Neuro-Symbolic Modeling for Nonlinear Dissipative Dynamics

YanHui Zhu, Yinhao Li (Liaoning Technical University)

OptimizationConvolutional Neural NetworkTransformerImagePoint CloudTabularTime SeriesBenchmarkPhysics Related

🎯 What it does: Proposed the Adaptive Physics-Informed Computing (APIC) architecture, which integrates physical priors with deep learning to achieve nonlinear dissipative dynamics modeling.

Approximate Equivariance via Projection-Based Regularisation

Torben Berndt (Heidelberg Institute for Theoretical Studies), Jan Stühmer (Heidelberg Institute for Theoretical Studies)

ClassificationRestorationRepresentation LearningConvolutional Neural NetworkContrastive LearningImagePoint CloudBiomedical DataComputed TomographyReview/Survey Paper

🎯 What it does: This paper proposes a projection-based regularization framework for learning approximate equivariant neural networks. By performing an orthogonal decomposition of the network weights, the equivariant and non-equivariant components are regularized separately, thus gradually guiding the network toward equivariance while maintaining model flexibility.

Approximate Nearest Neighbor Search for Modern AI: A Projection-Augmented Graph Approach

Kejing Lu (University of Yamanashi), Jianbin Qin (Shenzhen University)

RetrievalComputational EfficiencyGraph Neural NetworkContrastive LearningImageTextMultimodalityGraphTabularAudio

🎯 What it does: Proposes a new approximate nearest neighbor search framework called Projection-Augmented Graph (PAG), which combines projection techniques with graph indexing to reduce unnecessary exact distance computations.

Approximate Proportionality in Online Fair Division

Davin Choo (Harvard University), Nicholas Teh (University of Oxford)

OptimizationFederated LearningReinforcement Learning from Human FeedbackReview/Survey Paper

🎯 What it does: Studied the online fair division problem, where indivisible goods arrive sequentially and must be allocated immediately. Addressed the issue of approximating proportional fairness (PROP1) in this setting.

Approximating Drift-Diffusion Models for User Decisions under Nudging and External Information

Gustavo Grivol (New York University), Alexander Tuzhilin (New York University)

Recommendation SystemOptimizationExplainability and InterpretabilityDiffusion modelScore-based ModelContrastive LearningTime SeriesSequentialBenchmarkStochastic Differential Equation

🎯 What it does: This paper proposes an extended drift-diffusion model (ExtDDM), providing a closed-form approximation for the first passage time of a single-threshold drift-diffusion process under time-varying drift, and using this approximation to derive theoretical conditions for the optimal intervention (nudging) timing in high-threshold scenarios.

Approximating f -Divergences with Rank Statistics

Viktor Stein (Technical University of Munich), José Manuel de Frutos (Universidad Carlos III)

GenerationData SynthesisOptimizationRepresentation LearningContrastive LearningImagePoint CloudGraphTabularTime SeriesReview/Survey PaperBenchmark

🎯 What it does: Propose an f-divergence approximation method based on rank statistics, which uses discrete rank histograms (with resolution K) to replace explicit density ratio estimation, directly calculating the divergence from the rank information of samples.

Approximation Bounds for Transformer Networks with Application to Regression

Yuling Jiao (Wuhan University), Bokai Yan (Hong Kong University Of Science And Technology)

TransformerTime SeriesSequentialReview/Survey Paper

🎯 What it does: The study investigates and provides Lp approximation bounds of the standard soft-maximized Transformer on Hölder and Sobolev function classes, and gives the excess risk convergence rate of sliding window empirical risk minimization under β-mixing dependent sequences.

Approximation Error Upper and Lower Bounds for Hölder Class with Transformers

Xin He (Wuhan University), Jerry Zhijian Yang (Wuhan University)

TransformerReview/Survey Paper

🎯 What it does: Study the upper and lower bounds of the approximation error of the standard Transformer in approximating Hölder class functions, and provide the precise relationship between the approximation error and the number of model blocks.

Approximation of Log-Partition Function in Policy Mirror Descent Induces Implicit Regularization for LLM Post-Training

Zhenghao Xu (Georgia Institute of Technology), Tuo Zhao (Georgia Institute of Technology)

OptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningContrastive LearningText

🎯 What it does: This paper studies the PMD-MEAN algorithm used in the post-training of large language models, derives its implicit regularization caused by the logarithmic partition function approximation, and provides theoretical and experimental verification.

Approximation Preserving Coresets

Milind Prabhu (University of Michigan), Sudarshan Shyam (Aarhus University)

OptimizationComputational EfficiencyData-Centric LearningImageTabular

🎯 What it does: Proposed and constructed the 'Approximate Preservation Core Set' (APC), which allows the use of a smaller core set at the cost of retaining only an approximate solution;

Approximation Theory for Lipschitz Continuous Transformers

Takashi Furuya (Doshisha University), Carola-Bibiane Schönlieb (University of Cambridge)

OptimizationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerAuto EncoderContrastive LearningStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: A class of Transformer architectures that maintain 1-Lipschitz continuity in context is proposed, and it is proven that they are universally approximating for all continuous mappings satisfying the same Lipschitz constraint on compact domains.

AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use

Junzhi Chen (New York University), Ashish Sabharwal (Allen Institute for AI)

Autonomous DrivingRobotic IntelligenceDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringWorld ModelTextTabularBenchmarkRetrieval-Augmented Generation

🎯 What it does: This paper proposes the AppWorld-UL benchmark, systematically converting existing one-way tool usage tasks into tasks requiring user interaction, and evaluates the user interaction capabilities of LLM agents through three types of interactions (clarification, infeasibility communication, confirmation) and their combinations.

Arboreal Neural Network

Wubin Yan (Du Xiaoman Technology (Beijing) Co., Ltd.), Dongliang Xu (Du Xiaoman Technology (Beijing) Co., Ltd.)

ClassificationOptimizationExplainability and InterpretabilityData-Centric LearningSupervised Fine-TuningAuto EncoderContrastive LearningTabularFinance Related

🎯 What it does: Proposes Arboreal Neural Network (ArbNN), a neural symbolic framework that compiles decision trees into differentiable ArborCell and supports bidirectional reversibility.

ArborKV: Structure-Aware KV Cache Management for Scaling Tree-based LLM Reasoning

Yeqiu Chen (University of Science and Technology of China), Lei Liu (University of Science and Technology of China)

OptimizationComputational EfficiencyTransformerLarge Language ModelTextChain-of-Thought

🎯 What it does: Design a structure-aware KV cache management framework called ArborKV in Tree-of-Thoughts reasoning, which can significantly compress KV memory usage while maintaining reasoning quality.

ARC-Decode: Accelerated Decoding with Risk-Bounded Acceptance

Ying Li (Westlake University), Huan Wang (Westlake University)

Computational EfficiencyAI Code AssistantTransformerLarge Language ModelPrompt EngineeringScore-based ModelTextRetrieval-Augmented Generation

🎯 What it does: Propose ARC-Decode, an improvement to Speculative Decoding that is training-agnostic, allowing safe relaxation of the draft acceptance rule in sampling mode;

ArcDAE: Asymmetric Rectified Contrastive Diffusion Autoencoder for Unified Representation Learning

Ge Gao (Nanjing University), Shuo Chen (Nanjing University)

GenerationRepresentation LearningTransformerDiffusion modelScore-based ModelRectified FlowAuto EncoderContrastive LearningImage

🎯 What it does: Propose ArcDAE — an asynchronous corrected contrastive diffusion autoencoder — to address the issues of information splitting and information overload in diffusion bridges;

Architecture Matters for Multi-Agent Security

Ben Hagag (Carnegie Mellon University), Sarah Scheffler

Safty and PrivacyTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmark

🎯 What it does: Experimentally evaluate the impact of multi-agent system architectures on security, systematically studying the effects of role configuration, communication topology, and memory visibility on attack success rates and task performance.

ArcVQ-VAE: A Spherical Vector Quantization Framework with ArcCosine Additive Margin

Jaeyung Kim (Chung-Ang University), YoungJoon Yoo (Chung-Ang University)

GenerationRepresentation LearningDiffusion modelAuto EncoderContrastive LearningImage

🎯 What it does: Propose ArcVQ‑VAE, a vector quantization framework that introduces a spherical angular interval prior based on VQ‑VAE, improving the codebook utilization and reconstruction/generation quality.

Are Common Substructures Transferable? Riemannian Graph Foundation Model with Neural Vector Bundles

Li Sun (Beijing University of Posts and Telecommunications), Philip S. Yu (University of Illinois Chicago)

Domain AdaptationExplainability and InterpretabilityRepresentation LearningGraph Neural NetworkSupervised Fine-TuningContrastive LearningGraph

🎯 What it does: Propose GAUGE, a pretrainable graph model that learns the intrinsic geometry of graphs using Neural Vector Bundle and evaluates the transferability of substructures through Dirichlet loss.

Are First-Order Diffusion Samplers Really Slower? A Fast Forward-Value Approach

Yuchen Jiao (Chinese University of Hong Kong), Gen Li (Chinese University of Hong Kong)

GenerationDiffusion modelImageStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Propose a training-agnostic, first-order forward value sampler, F-DPMSolver, which improves sampling quality by using a single look-ahead prediction at each step.

Are Large Reasoning Models Interruptible?

Tsung-Han Wu (UC Berkeley), Joseph E. Gonzalez (UC Berkeley)

AI Code AssistantTransformerLarge Language ModelPrompt EngineeringTextSequentialBenchmark

🎯 What it does: This paper proposes a benchmark for evaluating the robustness of large inference models in dynamic interruption scenarios (time constraints and update-driven), and systematically analyzes the failure modes of the models.

Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations

Dani Roytburg (Carnegie Mellon University), Narmeen Fatimah Oozeer

Explainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextChain-of-Thought

🎯 What it does: This paper investigates self-preference in large language model evaluators, proposing a quality baseline for evaluators to separate evaluation uncertainty from narcissistic bias, and re-replicating and correcting previous experimental results.

Are Object-Centric Representations Better at Compositional Generalization?

Ferdinand Kapl (Technical University of Munich), Andrea Dittadi (Technical University of Munich)

Data SynthesisComputational EfficiencyRepresentation LearningTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: Studied the compositional generalization ability of object-centric representations in visual question answering (VQA) tasks, and constructed three controllable synthetic visual worlds (CLEVRTex, Super-CLEVR, MOVi-C) as benchmarks;

Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM Reasoning

Qinghe Ma (Nanjing University), Yinghuan Shi (Nanjing University)

Computational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringVision Language ModelImageTextMultimodalityChain-of-Thought

🎯 What it does: This paper proposes AutoTool, which can adaptively decide whether to call tools in multi-modal LLMs to improve reasoning efficiency and accuracy.

Are Two Datasets Close Enough With Statistical Significance? A Kernel Distributional Closeness Testing Approach

Zhijian Zhou (University of Melbourne), Feng Liu (University of Melbourne)

Anomaly DetectionRepresentation LearningData-Centric LearningContrastive LearningImageMultimodalityTabularTime Series

🎯 What it does: Proposes a distribution closeness test (DCT) framework based on kernel methods, using a new metric called normalized maximum mean discrepancy (NAMMD) to determine whether two distributions are close within a given tolerance.

Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination

Chufan Shi (University of Southern California), Xuezhe Ma (University of Southern California)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmark

🎯 What it does: This paper proposes a diagnostic framework called VISUALSWAP, combined with a custom VS-BENCH benchmark (800 image-question-answer triplets), to test whether visual language models truly re-examine images when engaging in self-reflection.

Are We Overconfident in Models and Results for Semi-Supervised 3D Medical Image Segmentation?

Jun Li (Southwest Jiaotong University), Ziwei Qin (Southwest Jiaotong University)

SegmentationConvolutional Neural NetworkContrastive LearningBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose a triple-space calibrated semi-supervised 3D medical image segmentation framework (TCSeg), which reduces pseudo-label confirmation bias by separating confidence from uncertainty.

Are Your Agents Upward Deceivers?

Dadi Guo (Shanghai Artificial Intelligence Laboratory), Xia Hu (Shanghai Artificial Intelligence Laboratory)

Recommendation SystemAnomaly DetectionAutonomous DrivingOptimizationFederated LearningExplainability and InterpretabilityDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextTabularTime SeriesSequentialBiomedical DataReview/Survey PaperBenchmarkFinance RelatedRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Studied the 'upward deception' phenomenon exhibited by large language model (LLM) agents when facing environmental constraints, where agents report success or fabricate information even after task failure.

AREA: Attribute Extraction and Aggregation for CLIP-Based Class-Incremental Learning

Zhen-Hao Xie (Nanjing University), Da-Wei Zhou (Nanjing University)

ClassificationDomain AdaptationComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerSupervised Fine-TuningPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodality

🎯 What it does: This paper proposes the AREA method, which addresses the catastrophic forgetting problem in continual learning by stabilizing attribute extraction and aggregation based on CLIP.

AReaL-DTA: Dynamic Tree Attention for Efficient Reinforcement Learning of Large Language Models

Jiarui Zhang (Hong Kong University of Science and Technology), Binhang Yuan (Hong Kong University of Science and Technology)

Computational EfficiencyTransformerLarge Language ModelReinforcement LearningPrompt EngineeringText

🎯 What it does: Systematically leverage prefix sharing during RL fine-tuning through Dynamic Tree Attention (DTA) and load balancing distribution strategies, significantly improving computational efficiency and GPU memory utilization in large language model (LLM) policy training.

ArenaRL: Scaling RL for Open-Ended Agents via Tournament-based Relative Ranking

Qiang Zhang (University of Science and Technology of China), Jiawei Liu (University of Science and Technology of China)

TransformerLarge Language ModelReinforcement LearningPrompt EngineeringContrastive LearningTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose ArenaRL, a reinforcement learning framework for open-ended tasks, which abandons traditional point-to-point scalar rewards and instead performs process-aware pairwise evaluation among generated trajectories within the same group, and uses a tournament for relative ranking, thereby providing more stable advantage signals; simultaneously, a complete open-ended agent evaluation benchmark, Open-Travel and Open-DeepResearch, is constructed.

Ariadne's Thread of LipSync: Unraveling Forgeries via Inconsistency between Lip Motions and Head Poses

Tianyi She (University of Science and Technology of China), Kejiang Chen (University of Science and Technology of China)

Anomaly DetectionConvolutional Neural NetworkRecurrent Neural NetworkTransformerContrastive LearningVideoMultimodalityAudio

🎯 What it does: Propose a LipDA framework based on the inconsistency between lip movements and head movements, for simultaneously detecting LipSync forgeries and attributing them to specific generative models.

ARLArena: A Unified Framework for Stable Agentic Reinforcement Learning

Xiaoxuan Wang (University of California Los Angeles), Wei Wang (University of California Los Angeles)

TransformerSupervised Fine-TuningReinforcement LearningAgentic AIPrompt EngineeringTextSequential

🎯 What it does: Proposed the ARLArena framework for building a stable Agentic RL training environment, and designed the SAMPO algorithm to achieve unified and stable training.

Artemis: Structured Visual Reasoning for Perception Policy Learning

Wei Tang (Nanjing University of Science and Technology), Zechao Li (Nanjing University of Science and Technology)

Object DetectionAutonomous DrivingOptimizationTransformerSupervised Fine-TuningReinforcement LearningVision Language ModelContrastive LearningImageMultimodality

🎯 What it does: Proposed a visual perception strategy learning framework called Artemis, which employs structured visual reasoning, outputting (label, bounding box) pairs at intermediate steps to achieve verifiable spatial reasoning; and uses verifiable rewards (GRPO) in reinforcement learning to guide the model's learning.

Artificial Hippocampus Networks for Efficient Long-Context Modeling

Yunhao Fang (ByteDance Seed), Lai Wei (ByteDance Seed)

Computational EfficiencyKnowledge DistillationRepresentation LearningRecurrent Neural NetworkTransformerLarge Language ModelText

🎯 What it does: Propose an Artificial Hippocampus Network (AHN) framework, which compresses long-sequence information into a fixed-size long-term memory outside the sliding window KV cache of Transformer, to achieve efficient long-context modeling;

ASAP: Exploiting the Satisficing Generalization Edge in Neural Combinatorial Optimization

Han Fang (Shanghai Jiao Tong University), Yutong Ban (Shanghai Jiao Tong University)

OptimizationMeta LearningTransformerGraph

🎯 What it does: Proposes the ASAP framework, which improves the robustness of neural combinatorial optimization under distribution shift by splitting the decision process into two stages: proposal and selection.

ASIR: Steganography for Diffusion Models via Antipodal Sampling and Iterative Recovery

Yaofei Wang (Hefei University of Technology), Donghui Hu (Hefei University of Technology)

Safty and PrivacyConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelImage

🎯 What it does: This paper proposes the ASIR framework, which achieves untrained, verifiably secure steganography in diffusion models by embedding messages during the reverse sampling process;

Ask Less, See More: Communication-Conditioned Token Pruning for Vehicle-to-Vehicle Cooperative Autonomous Driving with Multimodal Large Language Models

Shiqi Sun (Northwestern Polytechnical University), Chenglie Du (Northwestern Polytechnical University)

Autonomous DrivingComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningMultimodalityPoint CloudRetrieval-Augmented Generation

🎯 What it does: A dual-stage communication-conditioned Token pruning framework named V2V-CCM is proposed, which utilizes QSM and SCM messages to guide the selection of LiDAR visual Tokens, achieving efficient inference of large language models in cooperative autonomous driving between vehicles.

Asking What Matters: Reward-Driven Clarification for Software Engineering Tasks

Sanidhya Vijayvargiya (Carnegie Mellon University), Graham Neubig (Carnegie Mellon University)

AI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: This paper systematically studies the effectiveness of clarifying questions in software engineering tasks, quantifies the impact of missing information on task success and question answerability, and trains a clarification module called CLARITI.

ASRU: Activation Steering Meets Reinforcement Unlearning for Multimodal Large Language Models

Jiahui Guang (Harbin Institute of Technology), Zhaoquan Gu (Harbin Institute of Technology)

Safty and PrivacyExplainability and InterpretabilityTransformerLarge Language ModelReinforcement LearningPrompt EngineeringVision Language ModelImageTextMultimodality

🎯 What it does: Propose a controllable multimodal large model learning framework ASRU, which achieves fine-grained control over forgotten knowledge through activation steering and reinforcement learning, maintaining generation quality and model utility.

Assistive Prompt Mediation: Evaluating Language Models Under Accessibility Constraints

Priyaranjan Pattnayak (Oracle America Inc), Ishan Banerjee (Indian Statistical Institute)

Recommendation SystemExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose Assistive Prompt Mediation (APM), a framework for evaluating large language models as assistants under restricted input conditions (such as dyslexia, motor impairments, speech recognition errors, etc.), focusing on recovering the user's latent intent and minimizing cognitive load and hallucination risks without allowing clarifications.

ASTRA: Communication-Efficient Acceleration for Multi-Device Transformer Inference

Xiao Liu (University of Massachusetts Amherst), Hui Guan (University of Massachusetts Amherst)

Federated LearningComputational EfficiencyTransformerLarge Language ModelAuto EncoderContrastive LearningImageText

🎯 What it does: This paper proposes the ASTRA framework, which utilizes sequence parallelism and mixed-precision attention to achieve multi-device Transformer inference, significantly reducing cross-device communication and improving inference speed in low-bandwidth environments.

Asymmetric conformal prediction with penalized kernel sum-of-squares

Louis Allain (Univ Rennes, Ensai, CNRS, CREST - UMR 9194), Brian Staber (Safran Tech, Digital Sciences & Technologies)

OptimizationComputational EfficiencyData-Centric LearningTabularTime Series

🎯 What it does: Proposed a new framework that can simultaneously consider symmetric and asymmetric prediction intervals in distribution-free conformal prediction;

Asymmetric Contrastive Objectives for Efficient Phenotypic Screening

Luke Nightingale (Francis Crick Institute), Michael Howell (Francis Crick Institute)

Computational EfficiencyRepresentation LearningDrug DiscoveryConvolutional Neural NetworkTransformerSupervised Fine-TuningAuto EncoderContrastive LearningImageBiomedical Data

🎯 What it does: Proposed a heterogenous target based on contrastive learning, introducing experimental metadata as learnable class vectors into the contrastive loss, and designing a Spherical Phenotype Clustering (SPC) mechanism to learn more discriminative image representations in high-throughput cell imaging phenotype screening.

Asymmetric Multi-View Clustering with Hyperbolic Uncertainty Modeling

Yiming Wang (Nanjing University of Posts and Telecommunications), Fu Xiao (Nanjing University of Posts and Telecommunications)

Representation LearningAuto EncoderContrastive LearningMultimodality

🎯 What it does: Propose a non-symmetric multi-view clustering framework HAMC based on hyperbolic geometry, which maps view features to the Poincaré ball, uses radius as a proxy for confidence, and enhances clustering robustness through asymmetric view alignment and confidence screening for global clustering learning.

Asymmetric Perturbation in Solving Bilinear Saddle-Point Optimization

Kenshi Abe (CyberAgent), Atsushi Iwasaki (University of ElectroCommunications)

OptimizationReinforcement LearningTabularSequentialBenchmark

🎯 What it does: Proposed a method that applies asymmetric perturbation only to one side's payoff function to solve bilinear saddle-point optimization problems and achieve fast convergence.

Asymmetric Prompt Weighting for Reinforcement Learning with Verifiable Rewards

Reinhard Heckel (Technical University of Munich), Christos Thrampoulidis (University of British Columbia)

TransformerReinforcement LearningPrompt EngineeringText

🎯 What it does: This paper proposes an Asymmetric Prompt Weighting scheme for reinforcement learning in verifiable rewards, aiming to enhance gradient signals for prompts with low success rates, thereby accelerating learning;

ASyMOB: Algebraic Symbolic Mathematical Operations Benchmark

Michael Shalyt (Technion Israel Institute of Technology), Ido Kaminer (Technion Israel Institute of Technology)

TransformerLarge Language ModelPrompt EngineeringTextBenchmarkPhysics RelatedRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposes a high-resolution symbolic mathematics operation benchmark called ASyMOB, containing 35,368 university-level symbolic integration, limit, differential equation, series, and hypergeometric problems, and generates a large number of variants through systematic symbolic, numerical, and equivalent transformations, specifically designed to evaluate the true reasoning ability of LLMs in symbolic reasoning rather than memorization patterns.

Asymptotic Optimality of the High-Dimensional Gaussian Mechanism and Improved Low-Dimensional Mechanisms for Differential Privacy

Yu Wei (Georgia Institute of Technology), Antigoni Polychroniadou (JPMorgan AI Research & AlgoCRYPT CoE)

Safty and PrivacyGaussian Splatting

🎯 What it does: This paper studies the asymptotic optimality of the high-dimensional Gaussian mechanism and the application of improved low-dimensional mechanisms in differential privacy, proposing a new spherical generalized gamma differential privacy mechanism. It proves that the privacy-utility trade-off of the Gaussian mechanism is optimal in high-dimensional settings and identifies mechanisms that outperform the Gaussian and ℓ2 mechanisms in certain low-dimensional settings.

Asymptotic Theory of Iterated Empirical Risk Minimization, with Applications to Active Learning

Hugo Cui (Université Paris-Saclay), Yue M. Lu (Harvard John A. Paulson School of Engineering and Applied Sciences)

OptimizationData-Centric LearningImageBiomedical DataReview/Survey Paper

🎯 What it does: This paper studies the theoretical and practical aspects of performing two rounds of empirical risk minimization (iterated ERM) on the same dataset, and applies this framework to pool-based active learning.

Asymptotic Universal Alignment: A New Alignment Framework via Test-Time Scaling

Yang Cai (Yale University), Weiqiang Zheng (Yale University)

OptimizationRepresentation LearningReinforcement Learning from Human FeedbackReinforcement LearningPrompt EngineeringMixture of ExpertsContrastive LearningTextReview/Survey PaperChain-of-Thought

🎯 What it does: Proposes a theoretical framework for achieving universal alignment through test-time scaling, and provides the optimal convergence rate.

Asymptotically Fast Clebsch-Gordan Tensor Products with Vector Spherical Harmonics

YuQing Xie (Massachusetts Institute of Technology), Tess Smidt (Massachusetts Institute of Technology)

ClassificationComputational EfficiencyRepresentation LearningGraph Neural NetworkDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningPoint CloudMeshPhysics Related

🎯 What it does: Studied acceleration methods for Clebsch-Gordan tensor products (CGTP) in E(3) equivariant neural networks, proposing a vector signal tensor product (VSTP) based on vector spherical harmonics to achieve complete and asymptotically efficient CGTP computation.

Asymptotically Optimal Sequential Testing with Markovian Data

Alhad Sethi (Indian Institute of Science), P. N. Karthik (Indian Institute of Technology)

Anomaly DetectionOptimizationReinforcement LearningMixture of ExpertsContrastive LearningSequential

🎯 What it does: The study investigates a one-sided α-regular sequential test for a composite null hypothesis and a composite alternative hypothesis under finite-state Markov chain data, and provides instance-dependent non-asymptotic lower bounds as well as asymptotically optimal tests that match these bounds.

AsyncSpade: Efficient Test-Time Scaling with Asynchronous Sparse Decoding

Shuqing Luo (University of North Carolina at Chapel Hill), Tianlong Chen (University of North Carolina at Chapel Hill)

Computational EfficiencyTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsTextBenchmark

🎯 What it does: AsyncSpade achieves efficient inference-time scaling by asynchronously decoupling KV cache selection from forward inference, predicting the next query state and performing sparse decoding in parallel during inference.

At the Edge of Understanding: Sparse Autoencoders Trace The Limits of Transformer Generalization

Praneet Suresh (Mila-Quebec AI Institute), Danilo Bzdok (Mila-Quebec AI Institute)

Explainability and InterpretabilityRepresentation LearningAdversarial AttackTransformerLarge Language ModelSupervised Fine-TuningAuto EncoderContrastive LearningTextAudio

🎯 What it does: Utilize sparse autoencoders (SAE) to conduct interpretability analysis on the internal representations of large language models, construct energy scores to detect out-of-distribution (OOD) inputs, and perform sample-priority fine-tuning based on these scores to enhance model robustness and defend against jailbreak attacks.

AtelierEval: Agentic Evaluation of Humans & LLMs as Text-to-Image Prompters

Hanjun Luo (New York University Abu Dhabi), Hanan Salam (New York University Abu Dhabi)

GenerationData SynthesisTransformerLarge Language ModelAgentic AIPrompt EngineeringImageTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: Established a unified evaluation benchmark called AtelierEval to quantitatively measure the prompting ability of humans and multimodal large language models (MLLMs) in text-to-image (T2I) prompting engineering, and proposed an automatic evaluator called AtelierJudge;

ATLAS: Learning to Optimally Memorize the Context at Test Time

Ali Behrouz (Google Research), Vahab Mirrokni (Google Research)

OptimizationComputational EfficiencyRepresentation LearningRecurrent Neural NetworkTransformerLarge Language ModelMixture of ExpertsContrastive LearningTextRetrieval-Augmented Generation

🎯 What it does: Propose the Atlas model, combining sliding window learning rules, higher-order feature mapping, and the MuON optimizer to build a high-capacity long-term memory module;

AtomWorld: A Benchmark for Evaluating Spatial Reasoning in Large Language Models on Material Structures

Taoyuze Lv (University of Science and Technology of China), Tong Xie (University of New South Wales)

Drug DiscoveryTransformerLarge Language ModelPrompt EngineeringTextGraphTabularBenchmarkPhysics RelatedRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose the AtomWorld benchmark to evaluate the ability of large language models in reasoning and manipulation within the material structure space;

Attacking Gray-Box Large Vision-Language Models with Adaptive SVD-Structured Adversarial Alignment

Daizong Liu (Wuhan University), Dengpan Ye (Guangzhou University)

Adversarial AttackTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodality

🎯 What it does: A gray-box attack framework is designed, which can make large audio-visual language models produce target answers by applying minor perturbations to the visual encoder and leveraging the target text semantics for guidance.

Attacks on Machine-Text Detectors Retain Stylistic Fingerprints

Rafael Alberto Rivera Soto (Johns Hopkins University), Nicholas Andrews (Johns Hopkins University)

Representation LearningAdversarial AttackTransformerLarge Language ModelContrastive LearningTextRetrieval-Augmented Generation

🎯 What it does: Investigate the limits of machine text detectors when facing various evasion attacks, demonstrate that the style feature space is robust against attacks, and propose a style-aware rewriting method that can simultaneously optimize detectability and author-specific writing style at the single-document level, thereby breaking through the detection of traditional and style detectors.

Attend to Anything: Foundation Model for Unified Human Attention Modeling

Wenzhuo Zhao (Sichuan University), Qijun Zhao (Sichuan University)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerVision Language ModelDiffusion modelAuto EncoderContrastive LearningImageVideoMultimodalityBenchmarkStochastic Differential EquationAudio

🎯 What it does: Propose the Attend to Anything Model (AAM), unifying attention (salience) modeling across images, videos, and audio-visual modalities, forming a cross-modal, cross-scenario foundational model.

Attention Hijacking: Backdooring Text Dataset Distillation via Semantic Anchors

Hang Ren (Sichuan University), Junqing Le (Chongqing University)

Knowledge DistillationAdversarial AttackTransformerPrompt EngineeringContrastive LearningText

🎯 What it does: Studied security vulnerabilities in text dataset distillation, proposing the Attention Hijacking (AH) attack, which embeds a backdoor without loss during distillation by hardening attention labels.

Attention Illuminates LLM Reasoning: The Uncovered Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization

Yang Li (Shanghai Jiao Tong University), Junchi Yan (Shanghai Jiao Tong University)

OptimizationExplainability and InterpretabilityTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkChain-of-Thought

🎯 What it does: This paper reveals the 'pre-planning-anchoring' rhythm in the reasoning process of large language models by analyzing their attention dynamics, and based on this, designs a fine-grained reinforcement learning credit assignment strategy;

Attention Implements the Fisher Geometry of Exponential Families

Bodie Rubacher (Independent Researcher)

OptimizationFederated LearningExplainability and InterpretabilityRepresentation LearningTransformerScore-based ModelContrastive LearningGaussian SplattingTabularTime SeriesSequentialStochastic Differential Equation

🎯 What it does: This paper proves that under a finite discrete exponential family observation model, a single attention head can accurately realize the Bayesian posterior and posterior mean, and explains when sharing a quadratic metric is feasible and when multiple heads with local curvature spectra are needed. It also extends this framework to theoretical and experimental analysis on the context estimation (ICE) task.

Attention Projection Mixing with Exogenous Anchors

Jonathan Su (Independent Researcher)

Computational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsContrastive LearningText

🎯 What it does: This paper proposes a new Transformer architecture called ExoFormer, which significantly improves the model's training stability and data efficiency by decoupling the conflict between token identification and feature transformation through placing the attention projection's anchor outside of the layers.

Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse

Zizhuo Fu (Peking University), Meng Li (Peking University)

Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelMixture of ExpertsTextBenchmark

🎯 What it does: This paper investigates and clarifies the relationship between Vanilla, Sink, and Gated attention mechanisms, proving that attention sink inherently forms a built-in Mixture-of-Experts (MoE) structure, and proposes a sink-aware auxiliary load balancing loss to address the head collapse problem;

Attention Sinks as Internal Signals for Hallucination Detection in Large Language Models

Jakub Binkowski (Wroclaw University of Science and Technology), Tomasz Jan Kajdanowicz

Anomaly DetectionExplainability and InterpretabilityTransformerLarge Language ModelText

🎯 What it does: Propose SinkProbe, a method that detects hallucinatory outputs generated by large language models using only the attention sink scores from the Transformer decoder.

Attention Sinks in Diffusion Transformers: A Causal Analysis

FANGZHENG WU, Brian Summa (Tulane University)

GenerationExplainability and InterpretabilityTransformerDiffusion modelScore-based ModelImageText

🎯 What it does: Studied the attention sink in diffusion Transformers and verified its function through causal intervention.

Attention Sparsity is Input-Stable: Training-Free Sparse Attention for Video Generation via Offline Sparsity Profiling and Online QK Co-Clustering

Jiayi Luo (Beihang University), Jianxin Li (Beihang University)

GenerationComputational EfficiencyTransformerDiffusion modelScore-based ModelImageVideoText

🎯 What it does: This paper proposes a training-free sparse attention framework called SVOO to accelerate the inference of video generation models.

Attention with Routed-Memory for Learnable Sparse Control

QIUHAO Zeng, Boyu Wang (Western University)

Computational EfficiencyRepresentation LearningTransformerLarge Language ModelMixture of ExpertsTextRetrieval-Augmented Generation

🎯 What it does: Proposes Attention with Routed Memory (ARM) — a differentiable, fixed-size KV cache structure that utilizes a hierarchical router to achieve learnable sparse control.

Attention's forward pass and Frank-Wolfe

Albert Alcalde (Friedrich-Alexander University), Domènec Ruiz-Balet (Universitat de Barcelona)

OptimizationExplainability and InterpretabilityComputational EfficiencyTransformerOrdinary Differential Equation

🎯 What it does: This paper reveals that the update rule of self-attention mechanisms in the hard maximization limit at extremely low temperatures (β→+∞) is equivalent to the Frank-Wolfe step, and investigates the dynamics of positive definite and negative definite key-query matrices. Subsequently, it compares this limit behavior with soft attention at finite temperatures, proving that the latter maintains dynamics similar to the limit dynamics on an exponential time scale (dynamic metastability).

Attentive Multi-Layer Fusion for Vision Transformers

Laure Ciernik (Technische Universität Berlin), Lukas Muttenthaler (Helmholtz Munich)

ClassificationRepresentation LearningTransformerSupervised Fine-TuningContrastive LearningImageMultimodalityBenchmark

🎯 What it does: Studied an attention-based multi-layer fusion method (ALF) that dynamically fuses CLS and average pooling (AP) features from all layers in a visual Transformer, and performed linear probing on frozen pre-trained models.

Attn-QAT: 4-Bit Attention With Quantization-Aware Training

Peiyuan Zhang (University Of California San Diego), Hao Zhang (University Of California San Diego)

Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelDiffusion modelVideoText

🎯 What it does: Studied and implemented 4-bit quantization-aware training (Attn-QAT) for high-quality attention computation on FP4 GPUs.

Attributed Network Alignment: Statistical Limits and Efficient Algorithm

Dong Huang (Tsinghua University), Pengkun Yang (Tsinghua University)

OptimizationRepresentation LearningGraph Neural NetworkContrastive LearningMultimodalityGraphTabular

🎯 What it does: Studies how to recover the hidden vertex correspondence between two related graphs under the scenario where both weighted edges and node features are simultaneously observed.

Attribution-Guided and Coverage-Maximized Pruning for Structural MoE Compression

Yifu Ding (Beihang University), Dacheng Tao (Nanyang Technological University)

CompressionExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelMixture of ExpertsContrastive LearningText

🎯 What it does: A channel-level structured pruning framework is proposed for Mixture-of-Experts (MoE) language models, which utilizes attribution-guided loss approximation to estimate expert importance, followed by a coverage maximization strategy to allocate pruning ratios, and employs alignment-aware reallocation to ensure compatibility with low-bit quantization.

AudioChat: Unified Audio Storytelling, Editing, and Understanding with Transfusion Forcing

William Chen (Adobe Research), Zeyu Jin (Adobe Research)

GenerationData SynthesisTransformerLarge Language ModelPrompt EngineeringDiffusion modelTextMultimodalityRetrieval-Augmented GenerationChain-of-ThoughtAudio

🎯 What it does: Developed AudioChat, a unified audio foundation model capable of generating, editing, and understanding complex multi-source audio stories, with the aid of AudioCopilot to generate synthetic training data.

AudioMosaic: Contrastive Masked Audio Representation Learning

Hanxun Huang (University of Melbourne), Sarah Monazam Erfani

Representation LearningTransformerLarge Language ModelAuto EncoderContrastive LearningMultimodalityAudio

🎯 What it does: AudioMosaic constructs positive sample pairs by applying structured time-frequency masking on spectrograms and trains a Transformer encoder using contrastive learning to learn general sentence-level audio representations.

Auditing Sybil: Explaining Deep Lung Cancer Risk Prediction Through Generative Interventional Attributions

Bartlomiej Sobieski (University of Warsaw), Przemyslaw Biecek (University of Warsaw)

Anomaly DetectionExplainability and InterpretabilityDiffusion modelScore-based ModelAuto EncoderImageBiomedical DataComputed Tomography

🎯 What it does: Propose the S(H)NAP framework to perform causal intervention-based auditing on the Sybil lung cancer risk prediction model, using 3D diffusion bridge to generate lung nodule insertion/deletion, and construct SHNAP and SNAP explanation methods.

AugMask: Training Diffusion Models on Incomplete Tabular Data via Stochastic Augmentation and Masking

Jungkyu Kim (Yonsei University), Kibok Lee (Yonsei University)

GenerationData SynthesisRecommendation SystemAnomaly DetectionOptimizationFederated LearningComputational EfficiencyData-Centric LearningTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningTabularTime SeriesSequentialBiomedical DataElectronic Health RecordsReview/Survey PaperBenchmarkFinance RelatedStochastic Differential Equation

🎯 What it does: Propose a training framework named AugMask, which utilizes random conditional augmentation and supervision only on observed coordinates, enabling traditional score-based diffusion models to be directly trained and generate complete samples from missing tabular data.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving

Ying Wang (Zhejiang University), Wenzhi CHEN

OptimizationComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose AugServe, an efficient serving framework for enhanced large language model inference, capable of dynamically scheduling requests and adaptively batching token budgets;

AURA: Visually Interpretable Affective Understanding via Robust Archetypes

Guanyu Hu (Xi'an Jiaotong University), Xinyu Yang (Xi'an Jiaotong University)

ClassificationRecognitionExplainability and InterpretabilityTransformerPrompt EngineeringVision Language ModelContrastive LearningImage

🎯 What it does: Propose the AURA framework, which replaces text prompts with visual archetypes to achieve interpretability and efficiency in sentiment analysis, supporting emotion recognition, facial action unit detection, and emotional dimension (Valence–Arousal) regression.

AuTAgent: A Reinforcement Learning Framework for Tool-Augmented Audio Reasoning

Siqian Tong (Institute of Acoustics Chinese Academy of Sciences), HAO Chengpeng (Institute of Acoustics Chinese Academy of Sciences)

TransformerReinforcement LearningAgentic AIPrompt EngineeringMultimodalityAudio

🎯 What it does: Propose AuTAgent, a reinforcement learning framework that learns when and which external tools to invoke in audio reasoning to improve the model's reasoning accuracy.

Auto-regressive In-context Demonstration Selection

Yunzhe Qi (University of Illinois Urbana Champaign), Jingrui He (University of Illinois Urbana Champaign)

Computational EfficiencyRepresentation LearningData-Centric LearningMeta LearningTransformerLarge Language ModelReinforcement LearningPrompt EngineeringContrastive LearningTextRetrieval-Augmented Generation

🎯 What it does: Propose a demonstration selection framework called AUTOSELECT based on autoregressive decision-making, which can progressively generate high-quality demonstration sequences under given queries, thereby enhancing the few-shot reasoning performance of LLMs.

AutoBaxBuilder: Bootstrapping Code Security Benchmarking

Tobias von Arx (Eth Zurich), Martin Vechev (Eth Zurich)

Safty and PrivacyAI Code AssistantTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose AUTOBAXBUILDER, an LLM-based automated pipeline for generating Web backend security benchmark scenarios, functional tests, and end-to-end exploit scripts from scratch, and subsequently building the AUTOBAXBENCH benchmark based on this pipeline.

Autobidding Auctions with LLM-Powered Creatives

Bingzhe Wang (Renmin University of China), Qi Qi (Renmin University of China)

Recommendation SystemOptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningTextTabularFinance Related

🎯 What it does: Propose a Platform-Investment Mechanism (PIM), integrating the enhancement of LLM-generated ideas with automatic bidding, constructing a dynamic Stackelberg game model, enabling the platform to make decisions and optimize LLM inference costs in ad bidding under budget constraints.

AutoControl Arena: Synthesizing Executable Test Environments for Frontier AI Risk Evaluation

Changyi Li (Fudan University), Min Yang (Fudan University)

Safty and PrivacyExplainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringGenerative Adversarial NetworkWorld ModelTextTabularBenchmarkChain-of-Thought

🎯 What it does: Propose AutoControl Arena, an automated framework that combines executable code with narratives generated by LLMs to build scalable and high-fidelity testing environments for evaluating frontier AI risks.

AutoMat: Physics-Guided Agentic Reasoning for Solving Ill-Posed Inverse Microscopy Problems

Yaotian Yang (Tsinghua University), Fei Wei (Tsinghua University)

Image TranslationGenerationData SynthesisConvolutional Neural NetworkGraph Neural NetworkTransformerAgentic AIMixture of ExpertsContrastive LearningImageTabularBiomedical DataMagnetic Resonance ImagingComputed TomographyBenchmarkPhysics RelatedRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposes AutoMat, a physics-guided agent reasoning system for STEM images, capable of automatically generating crystal CIF structures and predicting formation energies from single noisy STEM projections.

Automata-Conditioned Cooperative Multi-Agent Reinforcement Learning

Beyazit Yalcinkaya (University of California, Berkeley), Sanjit A. Seshia (University of California, Berkeley)

OptimizationFederated LearningRepresentation LearningRecurrent Neural NetworkGraph Neural NetworkReinforcement LearningAuto EncoderContrastive LearningGraphSequential

🎯 What it does: Propose the ACC-MARL framework in multi-task, multi-agent reinforcement learning, using DFA to represent tasks, and achieve solutions for historical dependency, credit assignment, and representation bottleneck through DFA progress, potential reward shaping, and pre-trained RAD embeddings, thereby learning decentralized policies that can perform optimal task allocation at test time.

Automated Formal Proofs of Combinatorial Identities via Wilf–Zeilberger Guidance and LLMs

Beibei Xiong (East China Normal University), Lihong Zhi (University of Chinese Academy of Sciences)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextBenchmarkChain-of-Thought

🎯 What it does: Proposed a neuro-symbolic framework called WZ-LLM, combining the Wilf-Zeilberger (WZ) method with large language models (LLMs), for automatically formalizing proofs of combinatorial identities on Lean 4.

Automatic Construction of Clinical Scoring Systems with LLM Agents

Silas Ruhrberg Estévez (University of Cambridge), Mihaela van der Schaar (University of Cambridge)

OptimizationExplainability and InterpretabilityTransformerLarge Language ModelAgentic AIPrompt EngineeringTabularBiomedical DataElectronic Health RecordsRetrieval-Augmented Generation

🎯 What it does: Propose an automated framework called AgentScore, based on large language models, for constructing clinically executable scoring systems that can be manually performed. It generates unit-weighted rule lists while satisfying clinical workflow constraints such as deployability, interpretability, and memorability.

Automatic Layer Selection for Hallucination Detection

Xinpeng Wang (University of Virginia), Zhe Zeng (University of Virginia)

Anomaly DetectionExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelAuto EncoderContrastive LearningText

🎯 What it does: This paper studies automatically selecting suitable layers for hallucination detection under the hidden state probing framework, and proposes two new methods, FEPoID and FST;

Automatic Pruning Discovery for Large Language Models

Haidong Kang (Northeastern University), Hao Wang (Xidian University)

OptimizationComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Automatically prune models without expert knowledge by generating and optimizing sparsification rules using large language models themselves.