ACL 2026 Papers — Page 7
Annual Meeting of the Association for Computational Linguistics · 2296 papers
DUD: Decoupled Update Dynamics for Reliable Uncertainty Quantification in Large Language Models
Yixin Bu (Nanjing University of Aeronautics and Astronautics), Piji Li (Nanjing University of Aeronautics and Astronautics)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextRetrieval-Augmented Generation
🎯 What it does: Propose the DUD framework, which separates the FFN and attention modules of LLMs through noise interference and causal intervention, and constructs dual-flow dynamic features using recovery scores to measure uncertainty.
DVCQR: Dual-View Conversational Query Rewriting with Stage-wise Reinforcement Learning
Chenyi Li (Central China Normal University), Zaixiang Wang (Central China Normal University)
RetrievalTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextRetrieval-Augmented Generation
🎯 What it does: Propose the DVCQR framework, which generates sparse view and dense view dialogue query rewrites through structured reasoning in a single step.
DVI-DTM: Dual-View Representation Learning for Interpretable Short Text Dynamic Topic Modeling
Di Liu (Beijing University of Posts and Telecommunications), Bin Wu (Beijing University of Posts and Telecommunications)
Explainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelContrastive LearningText
🎯 What it does: Propose a dual-perspective representation learning and explainability-enhanced dynamic topic model for short texts, DVI-DTM, to address semantic ambiguity and interpretability ambiguity.
DVMap: Fine-Grained Pluralistic Value Alignment via High-Consensus Demographic-Value Mapping
Pengyun Zhu (Tianjin University), Deyi Xiong (Tianjin University)
OptimizationExplainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextTabularBenchmarkChain-of-Thought
🎯 What it does: By constructing a high-consensus demography-value mapping corpus and combining structured thinking prompts with group relative policy optimization, the DVMap framework is proposed to achieve fine-grained multi-value alignment for large language models.
DyBBT: Dynamic Balance via Bandit-inspired Targeting for Dialog Policy with Cognitive Dual Systems
Shuyu Zhang (Shanghai Jiao Tong University), Bin Li (SIAT Chinese Academy Of Sciences)
Knowledge DistillationRepresentation LearningTransformerReinforcement LearningPrompt EngineeringMixture of ExpertsTextBenchmark
🎯 What it does: Propose the DyBBT framework, which balances dynamic exploration and exploitation through the cognitive state space C and a meta-controller.
Dynamic Emotion and Personality Profiling for Multimodal Deception Detection
Li Zheng (Wuhan University), Donghong Ji (Wuhan University)
ClassificationRecognitionAnomaly DetectionRecurrent Neural NetworkTransformerLarge Language ModelPrompt EngineeringContrastive LearningVideoTextMultimodalityAudio
🎯 What it does: This paper proposes a multimodal annotation scheme for dynamic emotional and personality labels, and constructs the DDEP dataset based on this. Subsequently, a Rel-DDEP reliability-weighted fusion framework is designed for joint detection of lying, emotion, and personality.
Dynamic Generation of Multi LLM Agents Communication Topologies with Graph Diffusion Models
Eric Hanchen Jiang (University of California Los Angeles), Ying Nian Wu (University of California Los Angeles)
OptimizationAI Code AssistantReinforcement Learning from Human FeedbackGraph Neural NetworkTransformerLarge Language ModelDiffusion modelTextGraphBenchmarkChain-of-Thought
🎯 What it does: Designing task-adaptive communication topologies for multi-agent systems driven by multilingual models
Dynamic Infilling Anchors for Format-Constrained Generation in Diffusion Large Language Models
Boyan Han (AGI Lab Westlake University), Chi Zhang (AGI Lab Westlake University)
GenerationTransformerLarge Language ModelDiffusion modelText
🎯 What it does: Propose a no-training, two-stage method called Dynamic Infilling Anchors (DIA) for achieving format-constrained text generation in diffusion large language models.
Dynamic Long Context Reasoning over Compressed Memory via End-to-End Reinforcement Learning
Zhuoen Chen (Harbin Institute of Technology), Min Zhang (Harbin Institute of Technology)
Computational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextRetrieval-Augmented Generation
🎯 What it does: Propose the LycheeMemory framework, which splits long texts into blocks, uses a Compressor to compress the blocks into KV-cache style memory; dynamically filters relevant blocks through a Gate, and iteratively updates the working memory with a Reasoner to complete inference.
Dynamics of Cognitive Heterogeneity: Investigating Behavioral Biases in Multi-Stage Supply Chains with LLM-Based Simulation
Jiuyun Jiang (Harbin Institute of Technology), Guang Xiao (Hong Kong Polytechnic University)
OptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelTabularSequentialBenchmarkChain-of-Thought
🎯 What it does: Utilize large language models (LLMs) to simulate multi-stage decision-making in dynamic supply chain simulations (Beer Distribution Game), examining the impact of cognitive heterogeneity on behavioral deviations;
E-ViC: Reasoning Beyond Text via Embodied Visual Chain for Spatial Intelligence
Junbo Qi (Waseda University), Xiaozhu Ju (X-Humanoid)
Robotic IntelligenceReinforcement Learning from Human FeedbackTransformerReinforcement LearningAgentic AIVision Language ModelImageTextMultimodalityChain-of-Thought
🎯 What it does: Propose the E-ViC framework, which transforms visual reasoning from a text chain to a visual chain, utilizing executable visual tools (such as zooming and point marking) to implement a 'look-and-confirm' loop;
E2EDev: Benchmarking Large Language Models in End-to-End Software Development Task
Jingyao Liu (Sichuan University), Yang Deng (Singapore Management University)
AI Code AssistantTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmark
🎯 What it does: Constructed the E2EDev benchmark based on BDD to evaluate the performance of large language models in end-to-end software development.
EA-Agent: A Structured Multi-Step Reasoning Agent for Entity Alignment
Yixuan Nan (Institute of Information Engineering, Chinese Academy of Sciences), Yanan Cao (Institute of Information Engineering, Chinese Academy of Sciences)
OptimizationComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelAgentic AITextGraphBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposes EA-Agent, a structured multi-step reasoning agent based on tool planning and execution, for entity alignment.
EASE: Entity-Aware Sub-table Generation for Real-world Multi-table QA
Myunghoon Kang (Korea University), Heuiseok Lim (Korea University)
Explainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextTabularBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose a multi-table question answering framework called EASE based on entity-aware subtable generation, and construct a Noisy Multi-table QA dataset containing noisy tables;
EchoVLM: Dynamic Mixture-of-Experts Vision-Language Model for Universal Ultrasound Intelligence
Chaoyin She (Northwestern Polytechnical University), Qinghua Huang (Northwestern Polytechnical University)
TransformerSupervised Fine-TuningPrompt EngineeringMixture of ExpertsVision Language ModelImageTextBiomedical DataUltrasound
🎯 What it does: Designed and trained a 10B-parameter vision-language model called EchoVLM for ultrasound medicine, supporting report generation, diagnostic prediction, and VQA.
EconProver: Towards More Economical Test-Time Scaling for Automated Theorem Proving
Mukai Li (Tencent), Dong Yu (Tencent)
OptimizationComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsTextBenchmarkChain-of-Thought
🎯 What it does: By constructing the EconRL framework, combining dynamic link thinking switching and multi-head parallel reinforcement learning, the token and sampling costs of the ATP model during testing are reduced, while maintaining the original performance.
EdgeFormer: Latency-Aware Collaborative Multi-Head Attention of Transformer Inference in Edge Networks
Yiming Yao (Beihang University), Tao Ren (Beihang University)
Federated LearningComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsContrastive LearningTextRetrieval-Augmented Generation
🎯 What it does: Propose the EdgeFormer framework to achieve distributed Transformer inference, dynamically allocate model blocks, and improve parallelism through collaborative multi-head attention.
Edit-Aware Reward Modeling for Chinese Grammatical Error Correction
Yilin Li (Peking University), Xiaojun Wan (Peking University)
Reinforcement Learning from Human FeedbackTransformerSupervised Fine-TuningReinforcement LearningPrompt EngineeringText
🎯 What it does: Proposed an edit-aware reward model called EARM, and combined it with the GRPO reinforcement learning framework for the Chinese grammar error correction task.
Editing the Moving World: Model Editing for Video LLMs
Qian Zhang (Harbin Institute of Technology), Dianbo Sui (Nanyang Technological University)
Explainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelVideoTextMultimodalityBenchmark
🎯 What it does: Constructed the VMEB benchmark to evaluate the performance of six model editing methods on three Vid-LLMs;
EDSD: Entropy-Driven Design for Faster Speculative Decoding
Longkai Cheng (Ant Group), Haixiang Hu (Ant Group)
Computational EfficiencyAI Code AssistantTransformerLarge Language ModelPrompt EngineeringText
🎯 What it does: Proposes the EDSD framework, which utilizes entropy-driven draft model training and architecture design to achieve more efficient speculative decoding, thereby accelerating the inference of large language models.
EduBench: A Comprehensive Benchmarking Dataset for Evaluating Large Language Models in Diverse Educational Scenarios
Bin Xu (Beijing Institute Of Technology), Heyan Huang (Beijing Institute Of Technology)
Data SynthesisKnowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextMultimodalityBenchmarkRetrieval-Augmented Generation
🎯 What it does: This study created the EduBench educational scenario benchmark dataset, covering 9 major educational scenarios, over 4,000 educational contexts, and 18,821 question samples, and designed a 12-dimensional evaluation scale, combining manual annotation with LLM evaluation; further trained a small-scale model and achieved performance comparable to large models through multi-source distillation.
EDUMATH: Generating Standards-aligned Educational Math Word Problems
Bryan R Christ (University of Virginia), Thomas Hartvigsen (University of Virginia)
GenerationData SynthesisTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextRetrieval-Augmented Generation
🎯 What it does: Use large language models (LLM) to automatically generate educational math word problems (MWP) that comply with American standards, and construct a standardized dataset annotated by teachers (STEM) for training and evaluation.
Effects of Collaboration on the Performance of Interactive Theme Discovery Systems
Alvin Po-Chun Chen (University of Colorado Boulder), Maria Leonor Pacheco (University of Colorado Boulder)
Explainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextReview/Survey PaperRetrieval-Augmented Generation
🎯 What it does: Compare the impact of synchronous and asynchronous collaboration modes on the quality of coding results from three types of NLP-assisted qualitative analysis tools (topic models, relational models, LLMs).
Efficient Hyperparameter Optimization for LLM Reinforcement Learning
Minping Chen (Hong Kong University of Science and Technology), Zeyi Wen (Hong Kong University of Science and Technology)
OptimizationHyperparameter SearchTransformerReinforcement LearningText
🎯 What it does: Propose a joint fidelity hyperparameter optimization (JF-HPO) framework that simultaneously adjusts model scale and training budget in large language model reinforcement learning to improve hyperparameter search efficiency.
Efficient KL Divergence Estimation via Truncated Top-K Integration for Large Language Models
Xinyuan Wang (Fudan University), Xipeng Qiu (Fudan University)
Computational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningScore-based ModelText
🎯 What it does: Proposed a TIKE method for estimating KL divergence based on Top-k truncation and importance sampling, addressing the memory bottleneck in KL computation during LLM RLHF training.
Efficient Learned Data Compression via Dual-Stream Feature Decoupling
Huidong Ma (Nankai University), Wentong Cai (Nanyang Technological University)
CompressionConvolutional Neural NetworkTransformerPrompt EngineeringMixture of ExpertsDiffusion modelFlow-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageVideoTextRetrieval-Augmented GenerationAudio
🎯 What it does: Proposes FADE, a dual-stream network that separates features into micro-syntax and macro-semantic components, accompanied by a lossless compression framework with hierarchical gating refinement and parallel pipeline.
Efficient Low-Resource Language Adaptation via Multi-Source Dynamic Logit Fusion
Chen Zhang (Peking University), Yansong Feng (Peking University)
Domain AdaptationComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsText
🎯 What it does: Proposes a test-time logit fusion framework called TRIMIX, aiming to adapt large language models (LLMs) to low-resource languages (LRLs) by dynamically balancing capabilities from different sources.
Efficient Multi-Agent System Training with Data Influence-Oriented Tree Search
Wentao Shi (University of Science and Technology of China), Chenyan Xiong (Carnegie Mellon University)
OptimizationFederated LearningData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelTextChain-of-Thought
🎯 What it does: This paper proposes the DITS framework, which utilizes influence scores to guide self-supervised data synthesis and selection in multi-agent systems, thereby improving model training performance.
Efficient Paths and Dense Rewards: Probabilistic Flow Reasoning for Large Language Models
Yan Liu (Meituan), Yangdong Deng (Tsinghua University)
Explainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelReinforcement LearningFlow-based ModelTextBenchmarkChain-of-Thought
🎯 What it does: A framework named CoT-Flow, which treats chain-of-thought reasoning as a continuous probabilistic flow, was constructed, along with the introduction of the information gain metric Probabilistic Flow Progress (PFP).
Efficient Prior-Guided Reasoning for Robust Retrieval-Augmented Generation under Conflicts
Xiaowei Yuan (Chinese Academy of Sciences), Kang Liu (Chinese Academy of Sciences)
RetrievalExplainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringDiffusion modelScore-based ModelTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This paper proposes Prior-Guided Reasoning and the BrPr framework, aiming to enhance the robustness of retrieval-augmented generation (RAG) when retrieval results contain conflicts or noise.
Efficient Process Reward Modeling via Contrastive Mutual Information
Nakyung Lee (Seoul National University), Jungwoo Lee (Seoul National University)
Computational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringContrastive LearningTextBenchmark
🎯 What it does: This paper proposes an automatic step reward annotation method based on contrastive point mutual information (CPMI) to construct step-level supervision datasets for process reward models (PRMs), which are then validated across various reasoning tasks.
Efficient Provably Secure Linguistic Steganography via Range Coding
Ruiyi Yan (Kyoto University), Yugo Murawaki (Kyoto University)
CompressionSafty and PrivacyTransformerLarge Language ModelText
🎯 What it does: A provably secure linguistic steganography method utilizing range coding is studied.
Efficient Self-Evaluation for Diffusion Language Models via Sequence Regeneration
Linhao Zhong (Zhejiang University), Chunhua Shen (Zhejiang University)
GenerationExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelDiffusion modelScore-based ModelContrastive LearningText
🎯 What it does: Proposed a self-evaluation method called DiSE based on Diffusion LLM, which quantifies output quality by utilizing the probability of re-generating tokens on complete sequences, and achieves conditional likelihood estimation, confidence quantification, and variable-length generation without training.
EfficientLLM: Unified Pruning-Aware Pretraining for Auto-Designed Compact Language Models
Xingrun Xing (Beijing Academy of Artificial Intelligence), Jiajun Zhang (University of Chinese Academy of Sciences)
Computational EfficiencyKnowledge DistillationNeural Architecture SearchTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringContrastive LearningText
🎯 What it does: Propose a unified pruning-aware pre-training framework (EfficientLLM), which prunes large language models during the pre-training phase and automatically designs the structure of small models, thus achieving significant compression while maintaining performance.
Efficiently Learning To Reason or Not to Reason: Root-token Policy Optimization for Adaptive Thinking
Taehyeon Kim (LG AI Research), Moontae Lee (LG AI Research)
Computational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningTextChain-of-Thought
🎯 What it does: This paper proposes the Root-token Policy Optimization (RPO) framework, which enables adaptive reasoning by leveraging the decision-making process of large reasoning models (LRM) to update only the root token after generating the initial thinking (thought) label. This allows the model to engage in long-chain reasoning for difficult problems and directly answer simple questions.
EGSS: Entropy-guided Stepwise Scaling for Reliable Software Engineering
Chenhui Mao (Ant Group), Yong Li (Ant Group)
OptimizationComputational EfficiencyAI Code AssistantTransformerLarge Language ModelAgentic AIPrompt EngineeringMixture of ExpertsTextRetrieval-Augmented Generation
🎯 What it does: Propose the Entropy-Guided Stepwise Scaling (EGSS) framework, which utilizes tool entropy to dynamically focus on inference computation, and enhances the reliability of software engineering tasks through Test Consolidation Augmentation (TCA) that integrates multi-trajectory debugging signals.
EIFFEL: a novel benchmark to measure bias of English heavy training on French idiomatic expressions
Charlotte Noel (LINAGORA), Julie Hunter (LINAGORA)
TransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: This paper proposes and implements the EIFFE L benchmark to evaluate the ability of large language models in understanding and using French idioms, and conducts assessments on various mainstream and French-focused multilingual models on this benchmark; meanwhile, a series of models with 1B parameters are trained to explore the impact of the French/English training ratio on model performance.
Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks
Wenqi Zhang (Zhejiang University), Yueting Zhuang (Zhejiang University)
Robotic IntelligenceReinforcement Learning from Human FeedbackGraph Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringVision Language ModelVision-Language-Action ModelImageTextMultimodalitySequentialRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Built and trained an embodied language-visual model called Embodied-Reasoner, capable of actively observing, reasoning, planning, searching, and executing actions in interactive environments.
Embracing Anisotropy: Turning Massive Activations into Interpretable Control Knobs for Large Language Models
Youngji Roh (Yonsei University), Jaehyung Kim (Yonsei University)
Domain AdaptationSafty and PrivacyExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringText
🎯 What it does: Studied the heterogeneity of internal activations in large language models, finding that a small number of highly activated dimensions are domain-specific functional units. These 'domain-critical dimensions' were identified without training by analyzing activation magnitudes, and then a method called Critical Dimension Steering was proposed to regulate only these dimensions for more precise model control.
EMCEE: Improving Multilingual Capability of LLMs via Bridging Knowledge and Reasoning with Extracted Synthetic Multilingual Context
Hamin Koo (Yonsei University), Jaehyung Kim (Yonsei University)
Data SynthesisExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringAuto EncoderTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose the EMCEE framework, which leverages synthetic multilingual contexts generated by the LLM itself and integrates them with inference results to improve the quality of LLM responses to non-English queries.
Emergent Misalignment via In-Context Learning: Narrow in-context examples can produce broadly misaligned LLMs
Nikita Afonin (AIRI), Mikhail Seleznyov (AIRI)
Safty and PrivacyExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringTextFinance RelatedChain-of-Thought
🎯 What it does: This paper systematically evaluates whether using examples from narrow domains in the context learning of multiple state-of-the-art LLMs leads to models producing misleading or harmful responses on irrelevant inputs, referred to as the 'Emergent Mismatch' (EM) phenomenon.
emg2speech: synthesizing speech from electromyography using self-supervised speech models
Harshavardhana T Gowda, Lee M. Miller (University of California, Davis)
GenerationData SynthesisConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningTime SeriesBiomedical DataAudio
🎯 What it does: This study proposes a neuromuscular speech interface that directly converts facial and jaw muscle EMG signals into audio using a self-supervised speech model.
EmoHarbor: Evaluating Personalized Emotional Support by Simulating the User’s Internal World
Jing Ye (State Key Laboratory of Multimodal Artificial Intelligence Systems Institute of Automation Chinese Academy of Sciences), Chengqing Zong (State Key Laboratory of Multimodal Artificial Intelligence Systems Institute of Automation Chinese Academy of Sciences)
Recommendation SystemExplainability and InterpretabilityTransformerLarge Language ModelAgentic AITextTabularBenchmarkChain-of-Thought
🎯 What it does: This paper proposes the EmoHarbor framework, which uses the User-as-a-Judge mode to simulate the user's inner world through a chain-of-agent approach, constructs a benchmark with 100 multidimensional user profiles, and evaluates the personalized emotional support performance of 20 LLMs.
EmoMAS: Emotion-Aware Multi-Agent System for High-Stakes Edge-Deployable Negotiation with Bayesian Orchestration
Yunbo Long (University of Cambridge), Liming Xu (University of Cambridge)
OptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsTextTabularBenchmark
🎯 What it does: Propose EmoMAS, a Bayesian multi-agent emotion-driven negotiation framework that can realize high-risk, emotion-sensitive negotiations on edge devices.
EmoS: A High-Fidelity Multimodal Benchmark for Fine-grained Streaming Emotional Understanding
Pengze Guo (University of Macau), Derek F. Wong (University of Macau)
RecognitionTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningVideoTextMultimodalityBenchmarkAudio
🎯 What it does: Proposed and constructed a high-fidelity bilingual emotion recognition benchmark called EmoS, addressing the shortcomings of existing datasets in ecological validity, signal clarity, and fine-grained annotation, and introduced a continuous streaming monologue subset to capture emotional evolution.
Emotion-Wheel-Guided Audio-Referred Text Representation for Multimodal Emotion Recognition in Conversation
Eunseon Seong (Hanyang University), Dong-Kyu Chae (Hanyang University)
RecognitionRecurrent Neural NetworkTransformerSupervised Fine-TuningContrastive LearningTextMultimodalityAudio
🎯 What it does: Propose the EMART framework, which achieves dialogue emotion recognition (ERC) using audio-guided text representations, enhancing performance through cross-modal fusion of audio and text and contrastive learning of emotional structures.
Empathy in Diversity: Personalized Depression and Anxiety Therapy via Dialogue State Tracking and Patient-Aware Planning
Xinwei Yang (Sichuan University), Yao Song (Sichuan University)
Recommendation SystemDrug DiscoveryReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningMixture of ExpertsDiffusion modelScore-based ModelTextSequentialBiomedical DataElectronic Health RecordsBenchmarkRetrieval-Augmented Generation
🎯 What it does: Studied the use of large language models (LLM) for personalized psychological treatment of depression and anxiety, and proposed the MindEval evaluation protocol, the MindData large-scale real clinical dialogue dataset, and the MindApt system.
Empirical Analysis of Decoding Biases in Masked Diffusion Models
Pengcheng Huang (Alibaba Group), Maosong Sun (Tsinghua University)
GenerationAI Code AssistantTransformerPrompt EngineeringDiffusion modelScore-based ModelTextBenchmarkChain-of-Thought
🎯 What it does: Analyze the decoding bias of Masked Diffusion Models (MDM) and propose the UNCODE calibration framework to improve generation quality.
Empowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task Planning
Tianyi Men (Key Laboratory of Cognition and Decision Intelligence for Complex Systems Institute of Automation Chinese Academy of Sciences), Jun Zhao (Key Laboratory of Cognition and Decision Intelligence for Complex Systems Institute of Automation Chinese Academy of Sciences)
Autonomous DrivingFederated LearningComputational EfficiencyRepresentation LearningRobotic IntelligenceTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AIPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBenchmark
🎯 What it does: Proposes the PEEU method based on autonomous experience exploration and retrospection, enhancing the cross-site generalization ability of small-scale multi-modal language models in GUI task planning.
Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy Optimization
Yifeng Ding (University of Illinois Urbana Champaign), Anoop Deoras (Aws Ai Labs)
OptimizationAI Code AssistantTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposes Group Turn Policy Optimization (GTPO), a reinforcement learning method specifically designed for multi-round tool-integrated reasoning (TIR), enabling large language models to better utilize tools and generate code in each reasoning round.
Empowering Tabular Data Preparation with Language Models: Why and How?
Mengshi Chen (Shanghai Jiao Tong University), Wenjie Zhang (University of New South Wales)
Data-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringMixture of ExpertsTabularReview/Survey PaperRetrieval-Augmented Generation
🎯 What it does: Systematically reviews and summarizes the application of large language models (LLM/SLM) in four core phases of table data preparation (acquisition, integration, cleaning, transformation), as well as feasible workflows.
Enabling Agents to Communicate Entirely in Latent Space
Zhuoyun Du (Zhejiang University), Haochao Ying (Zhejiang University)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelAuto EncoderContrastive LearningTextChain-of-Thought
🎯 What it does: Propose the Interlat framework, enabling full communication between LLM agents in the latent space, directly transmitting the final hidden layer states of the model and performing compression;
End-to-End Optimization of LLM-Driven Multi-Agent Search Systems via Heterogeneous-Group-Based Reinforcement Learning
Guanzhong Chen (Xiaomi Inc.), Zenglin Xu (Fudan University)
Autonomous DrivingOptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningAdversarial AttackHyperparameter SearchData-Centric LearningRobotic IntelligenceMeta LearningDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackNeural Architecture SearchProtein Structure PredictionTransformerLarge Language ModelReinforcement LearningAgentic AITextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposes MHGOPO — a critic-free reinforcement learning framework for end-to-end optimization of multi-agent LLM search systems (MASS).
Enhanced Reasoning for Biomedical Document-Level Relation Extraction via a Novel Cascade Language Model Framework
Haohua Song (Shenzhen University), Zexuan Zhu (Shenzhen University)
Drug DiscoveryTransformerLarge Language ModelPrompt EngineeringTextBiomedical DataRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposed a hierarchical detection-rethink framework called CoRE for solving the biomedical document-level relation extraction task. The framework first uses a pre-trained language model (PLM) to efficiently predict entity pairs and provide confidence scores; low-confidence samples are then passed to a large language model (LLM), which further infers with the help of retrieved examples and iterative reasoning mechanisms.
Enhancing Agentic Textual Graph Retrieval with Synthetic Stepwise Supervision
Ge Chang (Tsinghua University), Yunxin Liu (Tsinghua University)
Data SynthesisRetrievalGraph Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AIPrompt EngineeringTextGraphRetrieval-Augmented Generation
🎯 What it does: Propose the Graph‑S³ framework, which leverages LLM-driven interactive graph retrieval and improves multi-hop graph question answering performance through synthesized step-by-step supervision.
Enhancing Lexical Relation Mining with Structured Sememe Knowledge
Hansi Wang (Key Laboratory of Computational Linguistics, Ministry of Education, Peking University), Yang Liu (School of Computer Science, Peking University)
Knowledge DistillationRepresentation LearningGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextGraphBenchmark
🎯 What it does: This paper studies how to leverage structured word meaning knowledge (Sememe tree) to improve the lexical relation mining task, which includes two subtasks: lexical relation classification (LRC) and lexical entailment (LE);
Enhancing Linguistic Competence of Language Models through Pre-training with Language Learning Tasks
Atsuki Yamaguchi (University of Sheffield), Nikolaos Aletras (University of Sheffield)
Representation LearningData-Centric LearningTransformerLarge Language ModelContrastive LearningTextBenchmark
🎯 What it does: Propose the L2T pre-training framework, adding 14 language learning tasks alongside the CLM objective to enhance the model's linguistic capabilities.
Enhancing LLM-based Search Agents via Contribution Weighted Group Relative Policy Optimization
Junzhe Wang (Fudan University), Qi Zhang (Fudan University)
RetrievalOptimizationTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextRetrieval-Augmented Generation
🎯 What it does: A method called Contribution-Weighted GRPO (CW-GRPO) is proposed, which integrates process supervision into the Group Relative Policy Optimization (GRPO) framework. It evaluates the contribution of each retrieval round through an LLM evaluator, thereby achieving fine-grained credit allocation while maintaining the stability of GRPO.
Enhancing Multimodal Large Language Models for Ancient Chinese Character Evolution Analysis via Glyph-Driven Fine-Tuning
Rui Song (Jilin University), Hao Xu (Jilin University)
RecognitionTransformerLarge Language ModelSupervised Fine-TuningVision Language ModelContrastive LearningImageTextMultimodalityBenchmark
🎯 What it does: A benchmark for evaluating the evolution of ancient characters in multi-modal large language models (MLLMs) was constructed, and model capabilities were assessed through tasks of character shape comparison and evolutionary reasoning. A stage-wise fine-tuning framework called GEVO, based on character shape comparison, was proposed to enhance the model's evolutionary reasoning and character shape recognition performance.
Enhancing Reinforcement Learning for Radiology Report Generation with Evidence-aware Rewards and Self-correcting Preference Learning
Qin Zhou (Key Laboratory of Smart Manufacturing in Energy Chemical Process Ministry of Education), Zhe Wang (Key Laboratory of Smart Manufacturing in Energy Chemical Process Ministry of Education)
GenerationData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringVision Language ModelImageTextBiomedical DataMagnetic Resonance ImagingComputed TomographyElectronic Health RecordsRetrieval-Augmented Generation
🎯 What it does: This paper proposes a reinforcement learning framework (ESC-RL) that combines group-wise evidence-aligned rewards and self-correcting preference learning for generating clinically credible chest X-ray reports.
Enhancing the Transferability of Jailbreak Attacks on Large Language Models via Exploiting Reparameterization Invariance
Ao Wang (China University of Petroleum East China), Weifeng Liu (China University of Petroleum East China)
Adversarial AttackTransformerLarge Language ModelPrompt EngineeringText
🎯 What it does: This study proposes a token-level jailbreaking attack framework called RIGJ based on natural gradients, aiming to enhance cross-model transferability on large language models.
Enhancing Two Steps Textual Anomaly Detection through Anisotropy Mitigation
Pierre Fihey (Institut Polytechnique de Paris), Pavlo Mozharovskyi (Institut Polytechnique de Paris)
Anomaly DetectionTransformerContrastive LearningTextBenchmark
🎯 What it does: This paper explores the use of similarity-trained sentence embedding models in a two-stage text anomaly detection framework, and introduces whitening post-processing to alleviate the anisotropy of embeddings.
EpiCaR: Knowing What You Don’t Know Matters for Better Reasoning in LLMs
Jewon Yeom (Seoul National University), Taesup Kim (Seoul National University)
Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Constructed the EPICAR framework, which internalizes model confidence calibration through dual-objective training (inference and self-assessment) in an iterative self-training loop, avoiding overconfidence caused by positive feedback.
EQUIP: EQUivariant preserving In-Place updates for Efficient Token Pruning
Arun Ramachandran (AMDIndia Private Limited), Prakash Raghavendra (AMDIndia Private Limited)
Computational EfficiencyTransformerLarge Language ModelText
🎯 What it does: Propose the EQUIP method, which replaces the eviction and insertion operations in token pruning with in-place replacement, thereby reducing the overhead of KV cache copying and RoPE re-calculation, while maintaining the equivariance of attention computation;
EquivPruner: Boosting Efficiency and Quality in LLM-Based Search via Action Pruning
Jiawei Liu (University of Science and Technology of China), Defu Lian (University of Science and Technology of China)
Computational EfficiencyAI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningTextRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Significantly reduce token consumption and improve inference accuracy by identifying and pruning semantically equivalent operations in the search tree during large language model inference.
Erasing Without Remembering: Implicit Knowledge Forgetting in Large Language Models
Huazheng Wang (Beijing University of Posts and Telecommunications), Dacheng Tao (Nanyang Technological University)
Safty and PrivacyExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringContrastive LearningText
🎯 What it does: This study investigates the issue of implicit knowledge forgetting in large language models, finding that existing forgetting methods have poor generalization capabilities for related semantic variations (such as rewrites, relation reversals, etc.); it proposes PERMU, a forgetting framework based on probabilistic perturbation, which injects noise into the most perturbation-sensitive tokens using adversarial samples and model sensitivity metric MSM, thereby suppressing the probability of fact-related tokens without significantly compromising model utility.
ERCThinker: Fast-Slow Thinking for Emotion Recognition in Conversation
Yumeng Fu (Harbin Institute of Technology), Bingquan Liu (Harbin Institute of Technology)
RecognitionTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AITextMultimodalityChain-of-Thought
🎯 What it does: This paper proposes the ERCThinker framework, which combines fast and slow thinking modes and achieves emotional reasoning and prediction through adaptive strategy switching and Agent-as-Judge rewards.
ES4R: Speech Encoding Based on Prepositive Affective Modeling for Empathetic Response Generation
Zhuoyue Gao (Northeastern University), Yifei Zhang (Northeastern University)
GenerationConvolutional Neural NetworkRecurrent Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningAuto EncoderTextMultimodalityRetrieval-Augmented GenerationAudio
🎯 What it does: This paper proposes the ES4R framework for generating empathetic text and speech responses from voice inputs, achieving a complete process in three stages (empathy understanding, empathy generation, and speech synthesis);
Escaping the Echo Trap: On Credit Assignment Failure in Multi-turn LLM Self-Reflection
Linxuan Du (Soochow University), Min Zhang (Soochow University)
OptimizationExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkChain-of-Thought
🎯 What it does: Propose a tree-structured GRPO extension called TRAE for multi-round self-reflection, addressing behavior collapse caused by Echo Trap.
Estimating the Black-box LLM Uncertainty with Distribution-Aligned Adversarial Distillation
Huizi Cui, Changqing Zhang (Tianjin University)
Explainability and InterpretabilityKnowledge DistillationTransformerLarge Language ModelSupervised Fine-TuningGenerative Adversarial NetworkContrastive LearningText
🎯 What it does: Proposes the Distribution‑Aligned Adversarial Distillation (DisAAD) framework, which performs adversarial distillation within the high-probability output regions of a black-box LLM using a lightweight proxy model, enabling real-time estimation of uncertainty with just a single input-output pair;
ET-Agent: Incentivizing Effective Tool-Integrated Reasoning Agent via Behavior Calibration
Yifei Chen (Renmin University of China), Zhicheng Dou (Renmin University of China)
Explainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningTextRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This paper proposes the ET-Agent framework, which improves the behavioral efficiency and accuracy of LLMs in tool-integrated reasoning through two stages: self-evolving data flywheel and behavioral calibration training.
EthicMind: A Risk-Aware Framework for Ethical-Emotional Alignment in Multi-Turn Dialogue
Jiawen Deng (University of Electronic Science and Technology of China), Fuji Ren (University of Electronic Science and Technology of China)
Safty and PrivacyExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation
🎯 What it does: Propose the ETHICMIND framework to achieve dynamic alignment of ethics and emotion in multi-turn dialogues;
ETR: Entropy Trend Reward for Efficient Chain-of-Thought Reasoning
Xuan Xiong (University of Toronto), Yang Wang (Concordia University)
Computational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextChain-of-Thought
🎯 What it does: This paper proposes an entropy trend-based reward mechanism (ETR) to encourage gradual reduction of uncertainty in chain-of-thought (CoT) reasoning, thereby achieving efficient, short-chain reasoning while maintaining answer accuracy.
Evaluating Answer Leakage Robustness of LLM Tutors against Adversarial Student Attacks
Jin Zhao, Tanja Käser (EPFL)
Explainability and InterpretabilityAdversarial AttackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This paper constructs a multi-turn dialogue evaluation to assess the answer leakage robustness of LLM mentors when facing adversarial student attacks.
Evaluating Language Model Pluralism through In-the-wild Crowd Discussions
Gagan Mundada (University of California San Diego), Zhouhang Xie (Allen Institute for Artificial Intelligence)
Explainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose the PLURALEVAL framework for evaluating the diversity of views generated by LLMs through open-ended generation, and construct the WILDSCOPE dataset.
Evaluating Legal Reasoning Traces with Legal Issue Tree Rubrics
Jinu Lee (University of Illinois Urbana Champaign), Julia Hockenmaier (University of Illinois Urbana Champaign)
Autonomous DrivingOptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningData-Centric LearningRobotic IntelligenceMeta LearningDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackNeural Architecture SearchProtein Structure PredictionTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Constructed a large-scale (24K entries) Korean legal judgment prediction dataset called LEGIT, and utilized legal question trees from court judgments as a fine-grained evaluation framework for assessing legal reasoning trajectories.
Evaluating LLMs on Large-Scale Graph Property Estimation via Random Walks
Sunil Kumar Maurya (University of Tokyo), Xin Liu (AIST)
Computational EfficiencyRepresentation LearningGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringGraphBenchmark
🎯 What it does: Proposes the EstGraph benchmark, which uses random walk statistics to evaluate the reasoning ability of LLMs in large-scale graph attribute estimation tasks.
Evaluating Robustness of Large Language Models Against Multilingual Typographical Errors
Raoyuan Zhao (LMU Munich), Michael A. Hedderich (LMU Munich)
TransformerLarge Language ModelPrompt EngineeringTextMultimodalityBenchmark
🎯 What it does: Proposed the MULTYPO multilingual typing error generation algorithm, and systematically evaluated its robustness in multilingual and multitask environments across 18 open-source LLMs
Evaluating Structure-Aware Retrieval and Safety in Statute-Centric Legal QA
Kyubyung Chae (Seoul National University), Taesup Kim (Seoul National University)
RetrievalSafty and PrivacyGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextGraphBenchmarkRetrieval-Augmented Generation
🎯 What it does: Constructed a structured retrieval and safety evaluation benchmark (SEARCHFIRESAFETY) based on South Korean fire safety regulations, used to test legal question-answering systems in hierarchical legal clause retrieval and safety under incomplete contexts.
Evaluating Temporal Consistency in Multi-Turn Language Models
Yash Kumar Atri (University of Virginia), Thomas Hartvigsen
TransformerLarge Language ModelPrompt EngineeringTextSequentialBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Construct and evaluate the ChronoScope benchmark to detect the ability of language models to maintain, cover, and transfer time ranges in multi-turn dialogues, and systematically assess the performance of various existing models on this task.
Evaluating the Expressive Appropriateness of Speech in Rich Contexts
Tianrui Wang (Tianjin University), Jianwu Dang (Tianjin University)
ClassificationKnowledge DistillationTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningTextChain-of-ThoughtAudio
🎯 What it does: Proposed and implemented the CEAEval framework for evaluating the appropriateness of speech expressions in rich contexts; constructed the CEAEval-D dataset and developed the CEAEval-M evaluation model;
Evaluating the Impact of Verbal Multiword Expressions on Machine Translation
Linfeng Liu (University of Cincinnati), Tianyu Jiang (University of Cincinnati)
TransformerLarge Language ModelSupervised Fine-TuningTextReview/Survey PaperBenchmark
🎯 What it does: Researchers systematically quantify the negative impact of Verbal Multi-word Expressions (VMWEs) on translation quality by conducting automatic and manual quality assessments on sentences with and without VMWEs across eight mainstream MT systems and seven target languages.
Evaluating Visual Narrative Coherence in Story Visualization via Diversified Storylines
Minha Jhang (Seoul National University), Kyomin Jung (Seoul National University)
GenerationData SynthesisTransformerVision Language ModelDiffusion modelContrastive LearningImageTextMultimodality
🎯 What it does: This paper proposes the Visual Context‑Aware Metric for Story Visualization (VCMS) evaluation framework and designs a hybrid dataset generation pipeline based on a discrete diffusion language model to simultaneously examine the match between text and images and the coherence between images.
Evaluation Pitfalls and Challenges in Multimedia Event Extraction
Philipp Seeberger (Technische Hochschule Nürnberg Georg Simon Ohm), Korbinian Riedhammer (Technische Hochschule Nürnberg Georg Simon Ohm)
MultimodalityBenchmark
🎯 What it does: Systematically analyze the pitfalls in multi-modal event extraction evaluation and propose a rigorous evaluation framework
EventWeave: A Dynamic Framework for Capturing Core and Supporting Events in Dialogue Systems
Zhengyi Zhao (Chinese University of Hong Kong), Kam-Fai Wong (Chinese University of Hong Kong)
GenerationComputational EfficiencyGraph Neural NetworkTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextGraphRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose the EventWeave framework, which constructs dynamic event graphs, distinguishes between core events and supporting events, and generates more natural dialogue responses by selecting the most relevant events with multi-head attention.
EverMemOS: A Self-Organizing Memory Operating System for Structured Long-Horizon Reasoning
Chuanrui Hu (EverMind), Yafeng Deng (EverMind)
RetrievalExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposes EverMemOS, a lifecycle-based memory operating system that converts the conversation history of LLMs into sustainable memory structures to support long-sequence reasoning.
Every Response Counts: Quantifying Uncertainty of LLM-based Multi-Agent Systems through Tensor Decomposition
Tiejin Chen (Arizona State University), Hua Wei (Arizona State University)
OptimizationExplainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Quantify uncertainty in multi-agent systems (MAS) built with large language models (LLM), proposing the MATU framework.
Every Token Counts: Generalizing 16M Ultra-Long Context in Large Language Models
Xiang Hu (Tencent), Jianguo Li (Ant Group)
Computational EfficiencyRepresentation LearningData-Centric LearningAI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringMixture of ExpertsTextSequentialRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: The HSA-UltraLong model was studied and implemented, achieving length generalization from a 16K pre-training window to a 16M context by utilizing Hierarchical Sparse Attention (HSA) and Sliding Window Attention (SWA).
Everyone is unique: Towards Behaviorally Heterogeneous Negotiation Dialogue Systems for Debt Collection
Yuhang Yang (Zhejiang University), Zhixin Zhang (Ant Group)
TransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkFinance RelatedRetrieval-Augmented Generation
🎯 What it does: This paper presents research on a dialogue system for debt collection, constructing the first publicly available benchmark with rich personas (DebtBench) and training a specialized debt collection agent (DebtGPT) based on this benchmark.
Evidence-Augmented Policy Optimization with Reward Co-Evolution for Long-Context Reasoning
Xin Guan (Southeast University), Jiuxin Cao (Hong Kong University of Science and Technology)
OptimizationTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposed the Evidence-Augmented Policy Optimization (EAPO) framework, which utilizes reinforcement learning to optimize the evidence retrieval and citation process in long-text reasoning through dense process rewards.
EVM-QuestBench: An Execution-Grounded Benchmark for Natural-Language Transaction Code Generation
Pei Yang (Gradient), Tianyu Shi (Soochow University)
AI Code AssistantTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose EVM-QuestBench, an execution-driven benchmark based on EVM-compatible chains, used to evaluate the ability of large language models to generate executable transaction scripts from natural language instructions.
Evo-Attacker: Memory-Augmented Reinforcement Learning for Long-Horizon Tool Attacks on LLM-MAS
Bingyu Yan (Beihang University), Litian Zhang (Beijing University of Posts and Telecommunications China Academy of Information and Communications Technology)
Adversarial AttackAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningTextSequentialBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose Evo-Attacker, which utilizes a memory-enhanced reinforcement learning framework to conduct long-term, dynamic attacks on the tool channels in LLM-multi-agent systems;
Evo-PI: Aligning Medical Reasoning via Evolving Principle-Guided Supervision
Xianda Zheng (University of Auckland), Shangyang Li (Beijing University of Posts and Telecommunications)
Explainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringVision Language ModelMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasound
🎯 What it does: Propose the Evo-PI framework, which uses evolvable reasoning principles as linguistic supervision signals in the medical vision-and-language question answering task, combining reinforcement learning to train a multimodal language model;
Evolutionary Guided Decoding: Iterative Value Refinement for LLMs
Zhenhua Liu (Shanghai Artificial Intelligence Laboratory), Jing Shao (Shanghai Artificial Intelligence Laboratory)
Computational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringText
🎯 What it does: Proposes the Iterative Value Refinement (IVR) framework, which improves the accuracy of the value function through two iterative steps: Value Exploration and Iterative Self-Refinement, thereby achieving higher quality guided decoding;
Evolutionary Negative Module Pruning for Better LoRA Merging
Anda Cao (Zhejiang University), Jie Song (Zhejiang University)
OptimizationComputational EfficiencyRepresentation LearningTransformerLarge Language ModelContrastive LearningImageText
🎯 What it does: The paper proposes a negative module pruning method called ENMP for merging LoRA models, which can remove LoRA layers that have a negative impact on global performance before merging.
Evolutionary Strategies at Scale lead to Catastrophic Forgetting
Immanuel Abdi (University of California Berkeley), Gopala Anumanchipalli (University of California Berkeley)
OptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningHyperparameter SearchMeta LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningText
🎯 What it does: Evaluate the performance of Evolution Strategies (ES) in fine-tuning large-scale language models and its forgetting behavior during continual learning, and conduct comparative experiments with Group Relative Policy Optimization (GRPO).
Evolving Agents
Leonardo Ranaldi (University of Edinburgh)
Computational EfficiencyRepresentation LearningRobotic IntelligenceTransformerLarge Language ModelReinforcement LearningAgentic AIWorld ModelTextMultimodalityRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This paper proposes the EVA framework, which integrates three modules: Perceptor, Actor, and Controller. It utilizes semi-structured quasi-symbolic abstraction to abstract and evolve the experiences of LLM agents, thereby achieving unified collaboration among reasoning, memory, and control.
Evolving Beyond Snapshots: Harmonizing Structure and Sequence via Entity State Tuning for Temporal Knowledge Graph Forecasting
Siyuan Li (Dalian University of Technology), Te Sun (Shanghai Jiao Tong University)
Recommendation SystemFederated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningAdversarial AttackDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackNeural Architecture SearchRecurrent Neural NetworkGraph Neural NetworkTransformerPrompt EngineeringMixture of ExpertsAuto EncoderGenerative Adversarial NetworkContrastive LearningGraphTime SeriesSequentialBenchmark
🎯 What it does: This paper proposes the Entity State Tuning (EST) framework, modeling entities as state containers that continuously evolve, leveraging global state buffering and closed-loop updates to balance structural reasoning and time series learning;
Evolving Sparsity: Leveraging Token Importance Dynamics for Efficient LLM Decoding with Sparse Attention
Ruizi Han (Harbin Institute of Technology (Shenzhen)), Liqiang Nie (Harbin Institute of Technology (Shenzhen))
Computational EfficiencyTransformerLarge Language ModelText
🎯 What it does: Propose the EvoSparse framework, which improves sparse attention in LLM long-text reasoning by using cross-step accumulation and cross-layer propagation to dynamically capture token importance, significantly enhancing efficiency and performance.
EvolvR: Self-Evolving Pairwise Reasoning for Story Evaluation to Enhance Generation
Xinda Wang (Peking University), Feng Xiao (Alibaba Group)
GenerationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningAgentic AITextChain-of-Thought
🎯 What it does: Proposed the EvolvR framework, which provides high-quality story evaluation and reward models for open-source models through self-evolving pairwise evaluation and multi-persona Chain-of-Thought generation.