arXivSub Start free trial

ACL 2026 Papers — Page 13

Annual Meeting of the Association for Computational Linguistics · 2296 papers

LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification

Penghui Yang (Nanyang Technological University), Bo An (Nanyang Technological University)

GenerationComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringText

🎯 What it does: Proposed the LONGSPEC framework for long-text reasoning scenarios, achieving lossless acceleration of long-context reasoning.

LongTutor: Benchmarking Large Language Models for Long-term Personalized Tutoring

Ning Li (University Of Science And Technology Of China), Enhong Chen (University Of Science And Technology Of China)

TransformerLarge Language ModelPrompt EngineeringTextSequentialBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose the LongTutor benchmark to evaluate the capabilities of large language models in long-term personalized tutoring.

LongVideoAgent: Multi-Agent Reasoning with Long Videos

Runtao Liu (Hong Kong University Of Science And Technology), Qifeng Chen (Hong Kong University Of Science And Technology)

Explainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringVision Language ModelVideoText

🎯 What it does: Proposed LongVideoAgent, a multi-agent framework where the main LLM collaborates with localization and visual agents to accomplish long video question answering.

Look Less, Reason More: Rollout-Guided Adaptive Pixel-Space Reasoning

Xuchen Li (Institute of Automation, Chinese Academy of Sciences), Wentao Zhang (Zhongguancun Academy)

Explainability and InterpretabilityComputational EfficiencyTransformerSupervised Fine-TuningReinforcement LearningPrompt EngineeringVision Language ModelImageTextMultimodalityRetrieval-Augmented Generation

🎯 What it does: Proposed an adaptive pixel space reasoning framework that allows vision-language models (VLMs) to dynamically decide whether to perform pixel-level operations based on input queries;

Look Light, Think Heavy: What Multimodal Chain-of-Thought Reasoning Can and Cannot Do

Zhuoran Jin (Institute of Automation, Chinese Academy of Sciences), Jun Zhao (Institute of Automation, Chinese Academy of Sciences)

Explainability and InterpretabilityRepresentation LearningTransformerPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmarkChain-of-Thought

🎯 What it does: This paper systematically evaluates 12 types of multimodal tasks (covering two dimensions: perception and reasoning), comparing the performance differences between direct answering and chain-of-thought (CoT) reasoning, and contrasts the performance of 14 general-purpose models and 8 reasoning models during the CoT process.

Look Within or Beyond? A Theoretical Comparison Between Parameter-Efficient and Full Fine-Tuning

YongKang Liu, Hinrich Schuetze

Federated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningTextReview/Survey PaperBenchmark

🎯 What it does: Studied and compared the theoretical and empirical performance of parameter-efficient fine-tuning (PEFT) and full-parameter fine-tuning (FFT), focusing on representation capability, robustness, and the impact on data/parameter gains.

Looking at Radiology Report Generation through a Causal Lens: A Survey

Satyam Kumar (Indian Institute of Technology Bombay), Pushpak Bhattacharyya (Indian Institute of Technology Bombay)

GenerationExplainability and InterpretabilityPrompt EngineeringContrastive LearningTextBiomedical DataElectronic Health RecordsReview/Survey Paper

🎯 What it does: Systematically reviews the sources of bias in the automatic radiology report generation (RRG) process, proposes methods for modeling, mitigating, and evaluating bias from a causal inference perspective, and analyzes the limitations of commonly used datasets and evaluation metrics.

Looking Beyond the One: Operationalizing and Eliciting Visual Ambiguity in VLLMs

Yuchong Chen (Soochow University), Yu Hong (Soochow University)

Explainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringVision Language ModelDiffusion modelContrastive LearningImageTextMultimodality

🎯 What it does: This paper quantifies the ambiguity of visual questions from a region-based perspective and systematically evaluates the performance of current vision-language large models (VLLMs) in multi-answer scenarios, investigating whether they implicitly contain ambiguity information and proposing a gating mechanism to dynamically switch between multiple focal answers based on hidden states.

LoopTool: Closing the Data–Training Loop for Robust LLM Tool Calls

Kangning Zhang (Shanghai Jiao Tong University), Yong Yu (Shanghai Jiao Tong University)

Autonomous DrivingOptimizationData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringDiffusion modelTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: LoopTool proposes a closed-loop adaptive data generation and model training framework to enhance the tool calling capabilities of large language models.

LoPT: Lossless Parallel Tokenization Acceleration for Long Context Inference of Large Language Model

Wei Shao (Huawei), Yuwei Fan (Huawei)

Computational EfficiencyTransformerLarge Language ModelPrompt EngineeringText

🎯 What it does: Proposes the LoPT (Lossless Parallel Tokenization) framework, which utilizes multi-process parallel tokenization and character position matching to achieve faster tokenization of long texts while ensuring the output is completely consistent with serial tokenization.

LoRA on the Go: Instance-level Dynamic LoRA Selection and Merging

Seungeon Lee (MPI-SWS), Krishna P. Gummadi (MPI-SWS)

Computational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsTextMultimodalityBenchmark

🎯 What it does: Propose the LOGO framework, which dynamically selects and fuses LoRA adapters for each instance during inference without requiring training.

Losing our Tail, Again: (Un)Natural Selection & Multilingual LLMs

Eva Vanmassenhove (Tilburg University)

TransformerLarge Language ModelTextMultimodalityReview/Survey Paper

🎯 What it does: This paper, through a review of the training mechanisms of multilingual large language models (LLMs), model self-destruction (model collapse), and related research in linguistics, computer vision, and machine translation, proposes the role of LLMs in 'natural selection' and discusses their potential threat to the dilution of language tails (rare words, low-probability structures) and the loss of linguistic diversity. It also calls for considering linguistic diversity and expressive richness in the evaluation and training of LLMs.

Lost in Execution: On the Multilingual Robustness of Tool Calling in Large Language Models

Zheng Luo (University of Southern California), Xiyang Hu (Arizona State University)

Explainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This paper constructs a multilingual tool calling diagnostic benchmark, MLCL, and systematically evaluates the robustness of tool calling in Chinese, Hindi, and Igbo.

Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluations

Preethi Seshadri (University of California Irvine), Seraphina Goldfarb-Tarrant (Cohere)

Recommendation SystemExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This paper systematically evaluates the reliability of using LLMs as user simulators in agent assessment, exploring their robustness, effectiveness, and fairness; through experiments with real users in the United States, India, Kenya, and Nigeria, it compares the interaction results between different user LLMs and real users; further analyzing the differences between simulated users and real users in terms of interaction structure, error patterns, and preferences.

Lost in the Mix: Evaluating LLM Understanding of Code-Switched Text

Amr Mohamed (MBZUAI), Guokan Shang (MBZUAI)

Explainability and InterpretabilityRepresentation LearningAI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextMultimodality

🎯 What it does: Systematically evaluate the understanding ability of large language models (LLMs) on code-switching (CSW) texts, construct a controllable code-switching test set based on linguistic theories, and compare the performance of different models and switching methods.

Lost in Translation, and Found: Detecting and Interpreting Translation Effects

Shira Wein (University of South Florida), Maria Leonor Pacheco (Proof School)

ClassificationExplainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningText

🎯 What it does: This paper performs binary classification to determine whether English text is translated text, constructing a high-precision model for detecting translated discourse.

LOTUS: Evolving Multimodal Unlearning via Hyperbolic Entailment and Lorentz Transport

Zekun Wang, Liang Yang (Dalian University Of Technology)

Safty and PrivacyExplainability and InterpretabilityTransformerVision Language ModelContrastive LearningImageTextMultimodalityBenchmark

🎯 What it does: Propose the LOTUS framework, which achieves precise zero-shot learning in multi-modal large language models by performing semantic slicing of specific visual concepts through the reverse meaning cone loss in the Lorentz hyperspace, and aligning with the safety rejection prior using Lorentz transport.

LoVeC: Reinforcement Learning for Better Verbalized Confidence in Long-Form Generation

Caiqi Zhang, Andreas Vlachos (University Of Cambridge)

GenerationReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningTextRetrieval-Augmented Generation

🎯 What it does: Train large language models using reinforcement learning to append a numerical confidence score after each sentence when generating long text, enabling online calibration of the authenticity of generated content.

Luring as a Proxy: Evaluating Corpus Transferability for Cybergrooming Detection

Shiying Fan (Fraunhofer SIT), Martin Steinebach (Fraunhofer SIT)

ClassificationDomain AdaptationTransformerSupervised Fine-TuningPrompt EngineeringContrastive LearningText

🎯 What it does: Evaluate the cross-domain transfer feasibility of public lure corpus in online sexual exploitation detection, and construct a framework for feature alignment and model generalization evaluation.

LVLMs and Humans Ground Differently in Referential Communication

Peter Zeng (Stony Brook University), Owen Rambow (Stony Brook University)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextMultimodality

🎯 What it does: In an experiment, researchers had humans engage in referential dialogues with large vision-language models (LVLMs) and among humans to complete multi-round object matching tasks, investigating the establishment of a common ground.

Lying with Truths: Open-Channel Multi-Agent Collusion for Belief Manipulation via Generative Montage

Jinwei Hu (University of Liverpool), Xiaowei Huang (University of Liverpool)

Adversarial AttackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose a method called cognitive collusion attack, which constructs misleading narratives through publicly available real fragments to induce LLM agents to generate false beliefs and spread rumors;

M^2PO: Multi-Perspective Multi-Pair Preference Optimization for Machine Translation

Hao Wang (Alibaba Group), Weihua Luo (Alibaba Group)

OptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringText

🎯 What it does: A multi-perspective multi-pair preference optimization framework named M2PO is proposed, aiming to improve human preference alignment in machine translation, addressing the blind spots of existing quality estimation models in partial errors (such as partial hallucinations and omissions).

M^3-VQA: A Benchmark for Multimodal, Multi-Entity, Multi-Hop Visual Question Answering

Jiatong Ma (Chinese Academy of Sciences), Jing Liu (Chinese Academy of Sciences)

TransformerLarge Language ModelVision Language ModelImageTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Constructed a multi-modal, multi-entity, multi-hop visual question answering benchmark called M3-VQA, and evaluated the performance of multi-modal large language models on this task.

MAB-DQA: Addressing Query Aspect Importance in Document Question Answering with Multi-Armed Bandits

Yixin Xiang (Nanjing University of Science and Technology), Jinhui Tang (Nanjing Forestry University)

GenerationRetrievalOptimizationGraph Neural NetworkTransformerReinforcement LearningAgentic AIPrompt EngineeringVision Language ModelTextMultimodalityGraphTabularBenchmarkRetrieval-Augmented Generation

🎯 What it does: This paper proposes MAB-DQA, a multi-modal document question answering framework based on the multi-armed bandit, which can explicitly model and dynamically allocate the importance of implicit aspects in queries, thereby improving retrieval effectiveness and answer generation quality.

Macaron: Controlled, Human-Written Benchmark for Multilingual and Multicultural Reasoning via Template-Filling

Alaa Elsetohy (MBZUAI), Alham Fikri Aji (Capital One)

Large Language ModelPrompt EngineeringTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: Constructed Macaron, a multilingual multicultural benchmark that separates reasoning types, cultural dimensions, and languages through templates.

MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference

Bo Li (Tsinghua University), Shaolin Zhu (Tianjin University)

Computational EfficiencyTransformerMixture of ExpertsVision Language ModelImageTextMultimodality

🎯 What it does: This paper proposes the MACS framework, aiming to address the 'dragging effect' in expert parallel inference of multi-modal MoE models, achieving efficient inference through information-aware and modality-dynamic adaptive capacity allocation.

MADE: A Living Benchmark for Multi-Label Text Classification with Uncertainty Quantification of Medical Device Adverse Events

Raunak Agarwal (Fraunhofer Heinrich Hertz Institute), Jackie Ma (Fraunhofer Heinrich Hertz Institute)

ClassificationTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextBiomedical DataBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed and made publicly available a sustainable and continuously updated multi-label text classification benchmark called MADE, specifically designed for medical device adverse event reports;

MaDS: Long-Horizon GUI Automation via Synergizing Dual-Layer Memory and Multi-Round Debate

Pengchen Chen (Zhejiang University), Wei Xiang (Zhejiang University)

Autonomous DrivingFederated LearningComputational EfficiencyRepresentation LearningRobotic IntelligenceReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIVision Language ModelVision-Language-Action ModelImageTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose a MaDS framework for long-sequence GUI automation, combining dual-layer memory with multi-round debates to achieve pre-task verification and experience cycles.

MAESTRO: Meta-learning Adaptive Estimation of Scalarization Trade-offs for Reward Optimization

Yang Zhao (Harbin Institute of Technology), Ting Liu (Harbin Institute of Technology)

OptimizationMeta LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextChain-of-Thought

🎯 What it does: Developed and implemented the MAESTRO framework, achieving multi-objective optimization in the alignment task of open-domain large language models (LLMs) through dynamic adaptive reward weighting.

MAGIC: Deep Geometric Evolution with Structural Consensus for Temporal Knowledge Graph Reasoning

Chengao Liu, Jianbin Jiao (University Of Chinese Academy Of Sciences)

Federated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningMeta LearningDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackNeural Architecture SearchProtein Structure PredictionGraph Neural NetworkContrastive LearningGraphTime SeriesStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Propose the MAGIC model, which utilizes multi-geometric space deep evolution and structural consensus for temporal knowledge graph reasoning

MagicBench: Diagnosing Visual Agency Loss and Semantic Dependency in Multimodal LLMs

Tang Da Huang (Xidian University), Xianpeng Guo (Xidian University)

Explainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningVideoTextMultimodalityBenchmark

🎯 What it does: Studied the performance of multimodal large language models in semantic adversarial scenarios, and proposed the MagicBench benchmark to evaluate physical consistency in visual and language interaction.

MAGMA: A Multi-Graph based Agentic Memory Architecture for AI Agents

Dongming Jiang (University of Texas at Dallas), Bingzhe Li (University of Texas at Dallas)

RetrievalExplainability and InterpretabilityComputational EfficiencyGraph Neural NetworkTransformerLarge Language ModelAgentic AIPrompt EngineeringTextGraphBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposes the MAGMA multi-graph agent memory architecture, achieving structured retrieval and reasoning of external memory through four types of graphs: semantic, temporal, causal, and entity.

MAGNET: Towards Adaptive GUI Agents with Memory-Driven Knowledge Evolution

Libo Sun (Fudan University), Zhongyu Wei (University of Southern California)

Robotic IntelligenceReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIVision Language ModelImageTextMultimodalityRetrieval-Augmented Generation

🎯 What it does: Propose the MAGNET framework, which utilizes a dual-memory architecture to achieve adaptive mobile GUI agents, overcoming issues of interface appearance drift and workflow drift.

Make LLMs See Like Investigators, Not Just Think More: The Role of Structured Analysis in Investigative Reasoning

Jaewook Lee (Electronics and Telecommunications Research Institute), Jong-hun Shin (Electronics and Telecommunications Research Institute)

Explainability and InterpretabilityTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkChain-of-Thought

🎯 What it does: Propose and verify the method PRISM, which migrates structured analysis techniques from the criminal investigation field (including information filtering, M.O.M.A structured hypothesis, and ACH hypothesis testing) into large language models (LLMs), to enhance the reasoning performance of identifying the perpetrator under complex narratives.

Make Mechanistic Interpretability Auditable: A Call to Develop Guidelines via Continuous Collaborative Reviewing

Michael Lan (Martian), Amir Abdullah

Safty and PrivacyExplainability and InterpretabilityAgentic AIPrompt EngineeringTextReview/Survey PaperRetrieval-Augmented Generation

🎯 What it does: Propose a Mechanism Interpretability (MI) audit framework, which includes a continuous collaborative review platform, community-developed 'living document' guidelines, and source-code-based audit systems.

Making Large Language Models Efficient Dense Retrievers

Yibin Lei (University of Amsterdam), Andrew Yates (Johns Hopkins University)

RetrievalComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelSupervised Fine-TuningAuto EncoderContrastive LearningText

🎯 What it does: This paper proposes the EffiR framework, which performs sparse compression on large language models (LLMs) to build efficient dense retrievers. It first coarsely removes redundant MLP layers, then further reduces model width through adaptive width compression (self-slimming), and finally performs contrastive learning fine-tuning for retrieval tasks.

MalruleLib: Large-Scale Executable Misconception Reasoning with Step Traces for Modeling Student Thinking in Mathematics

Xinghe Chen (Rice University), Shashank Sonkar (University of Central Florida)

Large Language ModelPrompt EngineeringTextBenchmarkChain-of-Thought

🎯 What it does: Proposed the MALRULELIB framework, which converts 101 learning science-based mathematical misconceptions (malrule) into executable programs, and generates dual-path (correct and incorrect) step-by-step problem-solving trajectories through 498 parameterized templates, thereby constructing a large-scale cross-template student error reasoning benchmark;

Mango: Multi-Agent Web Navigation via Global-View Optimization

Weixi Tong (Purdue University), Tianyi Zhang (Purdue University)

Autonomous DrivingOptimizationTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringMixture of ExpertsTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose the MANGO framework, which utilizes the global structure of websites and multi-armed bandits to achieve efficient web navigation.

Map of Encoders – Mapping Sentence Encoders using Quantum Relative Entropy

Gaifan Zhang (University of Liverpool), Danushka Bollegala (University of Liverpool)

Representation LearningTransformerContrastive LearningTextBenchmark

🎯 What it does: Propose a sentence encoder mapping method based on quantum relative entropy (QRE), constructing a 2D visualization map of 1101 sentence encoders.

Mapping the Circumplex of Affect: Geometric Analysis of Emotion Representations via Hyperspherical Contrastive Learning

Yusuke Yamauchi (University of Tokyo), Akiko Aizawa (National Institute of Informatics)

Explainability and InterpretabilityRepresentation LearningTransformerContrastive LearningText

🎯 What it does: By contrastive learning on emotion labels, inducing circular structures on hyperspheres for emotion representations, and comparing with traditional contrastive learning methods;

MARCH: Multi-Agent Radiology Clinical Hierarchy for CT Report Generation

Yi Lin (Weill Cornell Medicine), Yifan Peng (Weill Cornell Medicine)

GenerationRetrievalExplainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningVision Language ModelImageTextMultimodalityBiomedical DataComputed TomographyRetrieval-Augmented Generation

🎯 What it does: A multi-agent framework called MARCH is proposed, which generates 3D CT reports by imitating the hierarchical structure of residents, researchers, and attending physicians in radiology, and introduces retrieval-enhanced and iterative consensus steps.

MARCH: Multi-Agent Reinforced Check for Hallucination

Zhuo Li (Qwen Large Model Application Team, Alibaba), Guanjun Jiang (Qwen Large Model Application Team, Alibaba)

GenerationData SynthesisExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringMixture of ExpertsTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposed a retrieval-augmented generation framework called MARCH based on multi-agent reinforcement learning, aimed at eliminating hallucinations in large language models during retrieval-augmented generation (RAG) tasks.

MARD: Module-Aware Reasoning Distillation for Language Models with Adaptive Supervision

Wenqi Yang (Huazhong University of Science and Technology), Yushen Fang (Huazhong University of Science and Technology)

Computational EfficiencyKnowledge DistillationRepresentation LearningMeta LearningTransformerLarge Language ModelSupervised Fine-TuningMixture of ExpertsDiffusion modelTextBenchmarkChain-of-Thought

🎯 What it does: Proposed a module-aware reasoning distillation framework called MARD, which injects supervision into key submodules of the Transformer (FFN down projection and self-attention output projection) using lightweight adapters, and achieves efficient transfer of multi-step reasoning capabilities while freezing the backbone.

Market-Bench: Benchmarking Large Language Models on Economic and Trade Competition

Yushuo Zheng (Shanghai Jiao Tong University), Guangtao Zhai (Shanghai Jiao Tong University)

Reinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkFinance Related

🎯 What it does: Propose Market-Bench, a closed-loop multi-agent supply chain economic simulation environment, for evaluating the economic decision-making capabilities of large language models in procurement auctions, pricing, and marketing language generation.

Markovian Linguistic-Temporal Bridge: Unlocking the Potential of LLMs for Time Series Forecasting

Siming Sun (Zhejiang University), Qinmin Yang (Zhejiang University)

Knowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextTime SeriesBenchmark

🎯 What it does: Proposed the MGSAA framework, which extracts Markovian structures from LLMs and uses them as cross-modal structural priors. Subsequently, it performs structured alignment with state constraints on time series, ultimately mapping time series into token sequences that are isomorphic to the intrinsic structure of LLMs, thereby activating the inference capability of frozen LLMs for time series forecasting.

MARS-RA: Rank Aggregation for Credit Assignment via Multimodal Comparisons in Embodied Multi-Agent Cooperation

Dawei Wang (Newcastle University), Richard Davison (Newcastle University)

TransformerLarge Language ModelReinforcement LearningVision-Language-Action ModelMultimodalityBenchmark

🎯 What it does: Proposes the MARS-RA framework, which transforms the credit assignment problem in multi-agent collaboration into a rank aggregation problem based on pairwise comparisons generated by large-scale multi-modal models, and achieves dynamic and robust credit assignment through potential function reward shaping.

MARS^2: Scaling Multi-Agent Tree Search via Reinforcement Learning for Code Generation

Pengfei Li (Shanghai Artificial Intelligence Laboratory), Bowen Zhou (Shanghai Artificial Intelligence Laboratory)

AI Code AssistantTransformerReinforcement LearningText

🎯 What it does: Designed and implemented MARS 2, a multi-agent tree search reinforcement learning framework, to improve performance in code generation tasks.

MAS-Bench: A Unified Benchmark for Shortcut-Augmented Hybrid Mobile GUI Agents

Pengxiang Zhao (Zhejiang University), Yong Liu (Zhejiang University)

TransformerLarge Language ModelAgentic AITextSequentialBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose MAS-Bench, a unified benchmark for evaluating hybrid operations of mobile GUI and shortcuts;

Mask-to-Correct^+: Leveraging Retriever Diversity for Masking-guided Faithful Fact Correction

Payel Santra (Indian Association for the Cultivation of Science), Partha Basuchowdhuri (Indian Association for the Cultivation of Science)

RetrievalExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: Proposes the Mask-to-Correct framework, which is training and annotation-free, leveraging diverse masks and retrieval-augmented generation for fact correction.

Masked by Consensus: Disentangling Privileged Knowledge in LLM Correctness

Tomer Ashuach (Technion Israel Institute of Technology), Liat Ein-Dor (IBM Research)

Explainability and InterpretabilityKnowledge DistillationTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: Investigated whether large language models possess internal knowledge that is only visible to themselves when predicting the correctness of their own answers.

MASPO: Unifying Gradient Utilization, Probability Mass, and Signal Reliability for Robust and Sample-Efficient LLM Reasoning

Xiaoliang Fu (Meituan), Xunliang Cai (Meituan)

OptimizationComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsTextBenchmark

🎯 What it does: A unified RLVR algorithm called MASPO is studied, which addresses three major bottlenecks: gradient utilization, probability mass sensitivity, and signal reliability.

Massively Multilingual Joint Segmentation and Glossing

Michael Ginn (University of Colorado Boulder), Alexis Palmer (University of Colorado Boulder)

SegmentationTransformerLarge Language ModelReinforcement LearningPrompt EngineeringContrastive LearningText

🎯 What it does: Proposed and trained a multilingual joint segmentation and interlinear glossing model called POLYGLOSS, which can output morphological segmentation of words and corresponding glosses in one go.

MATCH: Modulating Attention via In-Context Retrieval for Long-Context Transformers

Linrui Ma (Huawei), Yufei Cui (Huawei)

RetrievalComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: Combine sparse attention models with an external retrieval module, dynamically introducing key positions retrieved into the attention mask to enhance long-context reasoning capabilities.

MatchTIR: Fine-Grained Supervision for Tool-Integrated Reasoning via Bipartite Matching

Changle Qu (Renmin University of China), Dawei Yin (Baidu Inc)

Reinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmark

🎯 What it does: Propose a fine-grained tool call reward allocation framework based on bilateral matching, called MatchTIR, to enhance the training of tool-integrated reasoning (TIR) models.

MathCanvas: Intrinsic Visual Chain-of-Thought for Multimodal Mathematical Reasoning

Weikang Shi (Chinese University of Hong Kong), Hongsheng Li (Chinese University of Hong Kong)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningVision Language ModelDiffusion modelRectified FlowAuto EncoderImageTextMultimodalityBenchmarkChain-of-Thought

🎯 What it does: Designed and implemented the MathCanvas framework, enabling large multimodal models to introspectively generate and edit visual thoughts (VCoT) to complete mathematical reasoning tasks.

Mathematical Proof as a Litmus Test: Revealing Failure Modes of Advanced Large Reasoning Models

Dadi Guo (Hong Kong University of Science and Technology), Yi R. Fung (Hong Kong University of Science and Technology)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose the RFMDataset and use LLM-as-a-judge to evaluate the performance of large inference models on natural language mathematical proof tasks.

MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems

Shuhang Chen (Zhejiang University), Yi Yang (Zhejiang University)

RecognitionExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmarkChain-of-Thought

🎯 What it does: Proposes the FlowVerse benchmark for fine-grained evaluation of perception and reasoning abilities in visual math problems, and designs a multi-module solution called MathFlow that separates perception and reasoning. A specialized model, MathFlow-P-7B, is trained to improve the extraction and description of graphics.

MathSight: A Benchmark Exploring Have Vision-Language Models Really Seen in University-Level Mathematical Reasoning?

Yuandong Wang (Capital Normal University), Zhenzhou Shao (Capital Normal University)

TransformerPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmarkChain-of-Thought

🎯 What it does: Designed and constructed the MathSight benchmark, which includes various visual variants such as original, hand-drawn, and photographed versions, as well as text-only conditions, to systematically evaluate the contribution of visual information in Vision-Language Models (VLM) for university-level mathematical reasoning.

MauBERT: Universal Phonetic Inductive Biases for Few-Shot Acoustic Units Discovery

Angelo Ortiz Tandazo (ENS), Emmanuel Dupoux (ENS)

RecognitionTransformerSupervised Fine-TuningAuto EncoderContrastive LearningAudio

🎯 What it does: A multilingual extension of HuBERT (MAUBERT) was constructed, further training language-agnostic and context-invariant speech representations through supervised learning using speech-to-articulatory features on 55 languages.

MavenCoder: Competitive Code Generation via Model Adaptive Planning Strategies and Multi-Perspective Verification Enhancement

ZhenChun Xu, Jiexin Wang (South China University of Technology)

GenerationAI Code AssistantTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Designed a model-adaptive and multi-perspective verification enhanced framework called MavenCoder for competition-level code generation.

Maximizing Local Entropy Where It Matters: Prefix-Aware Localized LLM Unlearning

Naixin Zhai, Xun Yang (University Of Science And Technology Of China)

Safty and PrivacyExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBenchmark

🎯 What it does: Proposed the PALU framework, achieving efficient forgetting of sensitive knowledge in LLMs by maximizing local entropy only on sensitive prefix words in the temporal dimension and topK words in the vocabulary dimension.

MCP-Flow: Facilitating LLM Agents to Master Real-World, Diverse and Scaling MCP Tools

WenHao Wang, Siheng Chen (Shanghai Jiao Tong University)

Autonomous DrivingOptimizationComputational EfficiencyData-Centric LearningAI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringDiffusion modelScore-based ModelFlow-based ModelTextTabularRetrieval-Augmented Generation

🎯 What it does: Propose MCP-Flow, which automatically collects MCP servers and tools across multiple platforms, generating over 60k instruction-function call pairs, and uses them to train LLMs to master MCP skills.

MDP-GRPO: Stabilized Group Relative Policy Optimization for Multi-Constraint Instruction Following

Mohammad Mahdi Salmani-Zarchi (University of Tehran), Mohammad Javad Dousti (University of Tehran)

Reinforcement LearningPrompt EngineeringText

🎯 What it does: Proposed a reinforcement learning method called MDP-GRPO for multi-constraint instruction following.

MeanAudio: Fast and Faithful Text-to-Audio Generation with Mean Flows

Xiquan Li (Shanghai Jiao Tong University), Xie Chen (Shanghai Jiao Tong University)

GenerationData SynthesisTransformerDiffusion modelFlow-based ModelAuto EncoderContrastive LearningTextMultimodalityAudio

🎯 What it does: Proposes MeanAudio, a text-to-audio generation model based on Mean Flow, capable of achieving high-quality audio synthesis in a single step (1 NFE);

Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding

Wenkai Wang (Zhejiang University), Shengyu Zhang (Zhejiang University)

Computational EfficiencyRepresentation LearningTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsVision Language ModelImageTextMultimodalityBenchmark

🎯 What it does: Proposes a co-evolutionary Propose-then-Critic framework based on reinforcement learning, which utilizes a MLLM to generate multiple candidate points in a single inference and selects the most accurate GUI coordinates through a visualization critic.

Measuring Distribution Shift in User Prompts and Its Effects on LLM Performance

Parker Seegmiller (Dartmouth College), Sarah Masud Preum (Dartmouth College)

Domain AdaptationExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This paper proposes the LENS framework, which quantifies the natural prompt distribution shift encountered by LLMs in real deployment environments and evaluates its impact on model performance.

Measuring Human Contribution in AI-Assisted Content Generation

Yueqi Xie (Princeton University), Fangzhao Wu (Microsoft Research Asia)

GenerationExplainability and InterpretabilityAdversarial AttackData-Centric LearningTransformerLarge Language ModelPrompt EngineeringText

🎯 What it does: This paper proposes an information-theoretic metric to quantitatively evaluate the proportion of human contribution in AI-assisted generated text, and estimates the minimum human contribution when the human input is unknown.

Measuring Social Bias in Vision-Language Models with Face-Only Counterfactuals from Real Photos

Haodong Chen (Harbin Institute of Technology), Jun Yu (Harbin Institute of Technology)

Explainability and InterpretabilityData-Centric LearningPrompt EngineeringVision Language ModelDiffusion modelContrastive LearningImageTextMultimodality

🎯 What it does: Studied the measurement of social bias in vision-language models using real photos with only minor facial modifications.

Measuring User’s Mental Models of Speech Translation in Human-AI Collaboration

HyoJung Han (University of Maryland), Marine Carpuat (University of Maryland)

Explainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented GenerationAudio

🎯 What it does: This paper proposes an experimental framework based on cross-lingual question answering to measure and improve users' mental models in speech translation;

Measuring What Matters!! Assessing Therapeutic Principles in Mental-Health Conversation

Abdullah Mazhar (IIIT Delhi), Md Shad Akhtar

ClassificationExplainability and InterpretabilityKnowledge DistillationRepresentation LearningTransformerLarge Language ModelTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Developed a framework called CARE to measure the consistency of therapists' responses in AI-generated mental health dialogues with six core therapeutic principles (non-judgmental acceptance, warm encouragement, respect for autonomy, active listening, reflection of feelings, contextual appropriateness), and released the FAITH-M fine-grained ordinal scoring benchmark based on expert annotations.

MECH: A Cost-Effective Multi-Task Cascade Framework for Classroom Opinion Evolution Recognition

Yancui Li (Henan Normal University), Fang Kong (Soochow University)

RecognitionTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsContrastive LearningText

🎯 What it does: Propose the classroom opinion evolution identification task and the corresponding COED dataset, and design the MECH framework to achieve efficient and low-cost automated analysis.

Mechanisms of Prompt-Induced Hallucination in Vision–Language Models

William Rudman (University of Texas at Austin), Kyle Mahowald (University of Texas at Austin)

Explainability and InterpretabilityRepresentation LearningTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodality

🎯 What it does: This paper investigates the hallucination phenomenon in vision-language models (VLMs) caused by conflicts between prompts and visual information in controlled object counting tasks, and significantly reduces such hallucinations by identifying and ablating a few key attention heads while not impairing normal counting ability.

Mechanistic Interpretability Should Prioritize Feature Consistency in Sparse Autoencoders

Xiangchen Song (Carnegie Mellon University), Kun Zhang (Carnegie Mellon University)

Explainability and InterpretabilityRepresentation LearningLarge Language ModelAuto EncoderContrastive LearningTextTabularTime SeriesSequential

🎯 What it does: This study investigates the feature consistency issue of sparse autoencoders (SAE) in terms of mechanistic interpretability, proposes and evaluates the Pairwise Dictionary Mean Correlation Coefficient (PW-MCC) metric based on dictionary matching, and theoretically proves and experimentally verifies its ability to achieve high consistency in TopK SAE.

MED-COREASONER: Reducing Language Disparities in Medical Reasoning via Language-Informed Co-Reasoning

Fan Gao (University of Tokyo), Irene Li (University of Tokyo)

Recommendation SystemFederated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningData-Centric LearningDrug DiscoveryTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBiomedical DataElectronic Health RecordsBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed and implemented the MED-COREASONER framework, leveraging parallel reasoning chains in English and local languages, concept extraction and fusion, and retrieval enhancement to improve the quality of medical multilingual reasoning.

MedEinst: Benchmarking the Einstellung Effect in Medical LLMs through Counterfactual Differential Diagnosis

Wenting Chen (Stanford University), Zhongrui Zhu (Xi'an Jiaotong University)

Explainability and InterpretabilityDrug DiscoveryTransformerLarge Language ModelAgentic AITextBiomedical DataBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This study proposes MedEinst, a counterfactual benchmark for detecting the Einstellung effect in medical LLMs during differential diagnosis, and based on this, designs ECR-Agent, aiming to reduce such biases through structured causal reasoning and experience accumulation.

MedForge: Interpretable Medical Deepfake Detection via Forgery-aware Reasoning

Zhihui Chen (National University of Singapore), Mengling Feng (Hunan University)

Anomaly DetectionExplainability and InterpretabilityTransformerLarge Language ModelReinforcement LearningPrompt EngineeringVision Language ModelImageMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasound

🎯 What it does: Propose a medical deep forgery detection framework called MedForge-Reasoner, which is based on pre-reasoning and interpretability, first locating suspicious regions and then generating medically verifiable reasoning.

MediEval: A Unified Medical Benchmark for Patient-Contextual and Knowledge-Grounded Reasoning in LLMs

Zhan Qu (TU Dresden), Michael Färber (TU Dresden)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningDrug DiscoveryTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringMixture of ExpertsTextBiomedical DataElectronic Health RecordsBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed the MediEval benchmark, combining MIMIC-IV medical records with the UMLS knowledge base, to evaluate LLMs in terms of factual accuracy and patient context consistency, and designed the CoRFu fine-tuning method based on this.

Mediocrity is the key for LLM as a Judge Anchor Selection

Shachar Don-Yehiya (Hebrew University of Jerusalem), Omri Abend (Hebrew University of Jerusalem)

TransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Investigated the impact of using anchors (benchmark models) for pairwise comparisons on ranking reliability when large language models (LLMs) act as judges (LLM-as-a-Judge, LMJ), systematically evaluating the performance of 22 anchors on the Arena-Hard-v2.0 dataset.

MedMCP-Calc: Benchmarking LLMs for Realistic Medical Calculator Scenarios via MCP Integration

Yakun Zhu (Shanghai Jiao Tong University), Xiaofan Zhang (Shanghai Jiao Tong University)

Drug DiscoveryTransformerLarge Language ModelReinforcement LearningAgentic AITextTabularBiomedical DataElectronic Health RecordsBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Constructed the MedMCP-Calc benchmark to evaluate the performance of LLMs in real medical calculator workflows using fuzzy task descriptions, EHR data interaction, and MCP tool integration, and on this basis proposed and trained the CalcMate model.

MedVerse: Efficient and Reliable Medical Reasoning via DAG-Structured Parallel Execution

Jianwen Chen (University Of North Carolina Chapel Hill), Huaxiu Yao (University Of North Carolina Chapel Hill)

Computational EfficiencyDrug DiscoveryGraph Neural NetworkTransformerLarge Language ModelTextGraphBiomedical DataElectronic Health RecordsBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose the MedVerse framework, which restructures medical reasoning as a directed acyclic graph (DAG) and enables parallel inference;

MeepleLM: A Virtual Playtester Simulating Diverse Subjective Experiences

Zizhen Li (Shanda AI Research), Kaipeng Zhang (Shanda AI Research)

Recommendation SystemReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Construct structured rule books and player review data, train MeepleLM as a virtual tester, capable of generating subjective experience evaluations based on different player personas;

MegaRAG: Multimodal Knowledge Graph-Based Retrieval Augmented Generation

Chi-Hsiang Hsiao (National Taiwan University), Chu-song Chen

RetrievalExplainability and InterpretabilityComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerLarge Language ModelVision Language ModelContrastive LearningImageTextMultimodalityGraphRetrieval-Augmented Generation

🎯 What it does: Propose MegaRAG, an end-to-end framework for automatically constructing a multi-modal knowledge graph (MMKG) and using it for retrieval-augmented generation;

Mem-Gallery: Benchmarking Multimodal Long-Term Conversational Memory for MLLM Agents

Yuanchen Bei (University of Illinois Urbana Champaign), Hanghang Tong (University of Illinois Urbana Champaign)

TransformerLarge Language ModelAgentic AIVision Language ModelImageTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose the Mem-Gallery benchmark to evaluate the multi-modal long-term conversational memory capability;

Mem^2Evolve: Towards Self-Evolving Agents via Co-Evolutionary Capability Expansion and Experience Distillation

Zihao Cheng (Beihang University), Yunhong Wang (Beijing Institute Of Technology)

Autonomous DrivingOptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningRobotic IntelligenceMeta LearningAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextSequentialBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose the Mem 2 Evolve framework, achieving dual mechanisms of asset memory (Asset Memory) and experience memory (Experience Memory), constructing a forward reasoning and backward evolution loop, realizing a self-evolving language model agent;

Mem2ActBench: A Benchmark for Evaluating Long-Term Memory Utilization in Task-Oriented Autonomous Agents

Yiting Shen (Chinese Academy of Sciences), Songlin Hu (Chinese Academy of Sciences)

TransformerLarge Language ModelAgentic AITextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Designed and constructed the MEM2ACTBENCH benchmark to evaluate the ability of large language model agents to proactively utilize memory to perform tool calls in long-term conversations.

MemBuilder: Reinforcing LLMs for Long-Term Memory Construction via Attributed Dense Rewards

Zhiyu Shen (Sun Yat-sen University), Yanghui Rao (Tencent)

Reinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningTextRetrieval-Augmented Generation

🎯 What it does: Trained and deployed a lightweight 4B LLM, utilizing the reinforcement learning framework MemBuilder to automatically construct and manage multi-dimensional (Core, Episodic, Semantic, Procedural) memories, thereby enhancing the consistency and traceability of long-term conversations.

MemCoRL: Alternating Co-Optimization of Memory Retrieval and Utilization via Collaborative Reinforcement Learning

Yuewen Liu (Beijing University of Posts and Telecommunications), Yutong Zhang (Beijing University of Posts and Telecommunications)

RetrievalOptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Designed a two-stage alternating cooperative reinforcement learning framework, MemCoRL, for simultaneously optimizing the external memory retrieval and utilization of large language models (LLMs).

Memorization, Emergence, and Explaining Reversal Failures: A Controlled Study of Relational Semantics in LLMs

Yihua Zhu, Hidetoshi Shimodaira

Data SynthesisExplainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningDiffusion modelTextGraph

🎯 What it does: Train and evaluate autoregressive LLMs from scratch to learn relational word logic semantics on synthetic knowledge graph data, and investigate the causes of reverse failure.

Memory efficiency and resource-rational encoding in sentence processing

Weijie Xu (University of California, Irvine), Richard Futrell (University of California, Irvine)

Computational EfficiencyRepresentation LearningTransformerLarge Language ModelContrastive LearningText

🎯 What it does: Introduce working memory constraints for resource rationalization into Transformer language models by injecting learnable Gaussian noise into self-attention and combining it with a hybrid loss function to regulate memory precision, and evaluate its impact on reading time prediction and the context representation space.

Memory-Augmented LLM-based Multi-Agent System for Automated Feature Generation on Tabular Data

FengXian Dong, Enhong Chen (University of Science and Technology of China)

OptimizationExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelAgentic AIPrompt EngineeringTabularRetrieval-Augmented Generation

🎯 What it does: A framework based on multi-agent and memory-enhanced large language models (MALMAS) was constructed for automated feature generation, supporting multi-round iterations and guiding feature generation and evaluation through multi-level memory (procedural memory, feedback memory, conceptual memory, and global conceptual memory).

Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning

Sikuan Yan (Ludwig Maximilian University of Munich), Yunpu Ma (Ludwig Maximilian University of Munich)

TransformerLarge Language ModelReinforcement LearningAgentic AITextRetrieval-Augmented Generation

🎯 What it does: Propose Memory‑R1, a framework for large language models (LLMs) that learns memory management and utilization through reinforcement learning.

MemRec: Collaborative Memory-Augmented Agentic Recommender System

Weixin Chen (Hong Kong Baptist University), Yongfeng Zhang (Rutgers University)

Recommendation SystemGraph Neural NetworkTransformerLarge Language ModelAgentic AIContrastive LearningTextGraphBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposed the MemRec framework, which improves the performance of LLM agent recommendation systems by utilizing collaborative memory and asynchronous propagation mechanisms.

MemSearch-o1: Empowering Large Language Models with Reasoning-Aligned Memory Growth in Agentic Search

Sheng Zhang (City University of Hong Kong), Xiangyu Zhao (City University of Hong Kong)

RetrievalExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelAgentic AIPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposes a agentic search framework based on large language models called MemSearch-o1, which utilizes part-of-speech seed tokens in queries to grow fine-grained memory fragments and reorganizes them into semantically smooth memory paths through backtracking retrieval and contribution functions, thereby significantly alleviating memory dilution and improving reasoning quality.

MentalSeek-Dx: Towards Progressive Hypothetico-Deductive Reasoning for Real-world Psychiatric Diagnosis

Xiao Sun (Chongqing University), Kaiwen Wei (Chongqing University)

Explainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningTextElectronic Health RecordsBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This paper proposes the ICD-11 level mental diagnosis benchmark MentalDx Bench based on real clinical electronic medical records, and develops the LLM framework MentalSeek-Dx with hierarchical hypothesis-deductive reasoning capabilities.

Merging Triggers, Breaking Backdoors: Defensive Poisoning for Instruction-Tuned Language Models

San Kim (POSTECH), Gary Lee

Federated LearningSafty and PrivacyExplainability and InterpretabilityComputational EfficiencyAdversarial AttackData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: Proposes the MB-Defense two-stage defense framework: first, merge the attacker's and defender's triggers into a unified backdoor representation through Defensive Poisoning, and then fine-tune via Backdoor Neutralization to disrupt the representation and restore the model's normal behavior.

MERIT Feedback Elicits Better Bargaining in LLM Negotiators

Jihwan Oh (KAIST AI), Taehyeon Kim (LG AI Research)

OptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextBenchmarkFinance RelatedChain-of-Thought

🎯 What it does: This paper constructs a new multi-scenario negotiation benchmark called AGORABENCH and proposes a multi-dimensional evaluation metric based on human preferences called MERIT, in order to improve the strategic depth and human consistency of large language models in negotiation tasks.

Merlin’s Whisper: Enabling Efficient Reasoning in Large Language Models via Black-box Persuasive Prompting

Heming Xia (Hong Kong Polytechnic University), Wenjie Li (Hong Kong Polytechnic University)

Computational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextBenchmark

🎯 What it does: Studied reducing overthinking in large language models through black-box persuasive prompting, and proposed the WHISPER framework to achieve efficient reasoning.

MessToClean: Evidence-Grounded Structure-Preserving Reconstruction for Real-World Degraded Exam Paper Images

Jiayi Tuo (University of Science and Technology of China), Ziwei Zhao (Technical University of Munich)

RecognitionImage TranslationRestorationTransformerLarge Language ModelSupervised Fine-TuningDiffusion modelImageMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose a structure-preserving reconstruction framework called MessToClean, which is evidence-driven, to convert real-world degraded exam images (RDEI) into traceable and auditable structured representations;

MetaBench: A Multi-task Benchmark for Assessing LLMs in Metabolomics

Yuxing Lu (Georgia Institute of Technology), May Dongmei Wang

TransformerLarge Language ModelTextBiomedical DataBenchmarkRetrieval-Augmented Generation

🎯 What it does: Constructed and released MetaBench, the first LLM capability evaluation benchmark for metabolomics, covering five major capabilities: knowledge, understanding, normalization, reasoning, and research.

Metaphor Reasoning is Meta-reasoning

Qianyu He (Fudan University), Yanghua Xiao (Fudan University)

GenerationMeta LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AIPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: This paper constructs an automated system called METAR to generate high-quality metaphor riddles and uses these riddles to train large language models to enhance their reasoning capabilities;