arXivSub Start free trial

ACL 2026 Papers with Code β€” Page 4

Annual Meeting of the Association for Computational Linguistics Β· 557 papers

MaDS: Long-Horizon GUI Automation via Synergizing Dual-Layer Memory and Multi-Round Debate

Pengchen Chen (Zhejiang University), Wei Xiang (Zhejiang University)

CodeAutonomous DrivingFederated LearningComputational EfficiencyRepresentation LearningRobotic IntelligenceReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIVision Language ModelVision-Language-Action ModelImageTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose a MaDS framework for long-sequence GUI automation, combining dual-layer memory with multi-round debates to achieve pre-task verification and experience cycles.

MAESTRO: Meta-learning Adaptive Estimation of Scalarization Trade-offs for Reward Optimization

Yang Zhao (Harbin Institute of Technology), Ting Liu (Harbin Institute of Technology)

CodeOptimizationMeta LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextChain-of-Thought

🎯 What it does: Developed and implemented the MAESTRO framework, achieving multi-objective optimization in the alignment task of open-domain large language models (LLMs) through dynamic adaptive reward weighting.

MAGMA: A Multi-Graph based Agentic Memory Architecture for AI Agents

Dongming Jiang (University of Texas at Dallas), Bingzhe Li (University of Texas at Dallas)

CodeRetrievalExplainability and InterpretabilityComputational EfficiencyGraph Neural NetworkTransformerLarge Language ModelAgentic AIPrompt EngineeringTextGraphBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposes the MAGMA multi-graph agent memory architecture, achieving structured retrieval and reasoning of external memory through four types of graphs: semantic, temporal, causal, and entity.

Make Mechanistic Interpretability Auditable: A Call to Develop Guidelines via Continuous Collaborative Reviewing

Michael Lan (Martian), Amir Abdullah

CodeSafty and PrivacyExplainability and InterpretabilityAgentic AIPrompt EngineeringTextReview/Survey PaperRetrieval-Augmented Generation

🎯 What it does: Propose a Mechanism Interpretability (MI) audit framework, which includes a continuous collaborative review platform, community-developed 'living document' guidelines, and source-code-based audit systems.

MalruleLib: Large-Scale Executable Misconception Reasoning with Step Traces for Modeling Student Thinking in Mathematics

Xinghe Chen (Rice University), Shashank Sonkar (University of Central Florida)

CodeLarge Language ModelPrompt EngineeringTextBenchmarkChain-of-Thought

🎯 What it does: Proposed the MALRULELIB framework, which converts 101 learning science-based mathematical misconceptions (malrule) into executable programs, and generates dual-path (correct and incorrect) step-by-step problem-solving trajectories through 498 parameterized templates, thereby constructing a large-scale cross-template student error reasoning benchmark;

Mango: Multi-Agent Web Navigation via Global-View Optimization

Weixi Tong (Purdue University), Tianyi Zhang (Purdue University)

CodeAutonomous DrivingOptimizationTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringMixture of ExpertsTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose the MANGO framework, which utilizes the global structure of websites and multi-armed bandits to achieve efficient web navigation.

Map of Encoders – Mapping Sentence Encoders using Quantum Relative Entropy

Gaifan Zhang (University of Liverpool), Danushka Bollegala (University of Liverpool)

CodeRepresentation LearningTransformerContrastive LearningTextBenchmark

🎯 What it does: Propose a sentence encoder mapping method based on quantum relative entropy (QRE), constructing a 2D visualization map of 1101 sentence encoders.

MARCH: Multi-Agent Reinforced Check for Hallucination

Zhuo Li (Qwen Large Model Application Team, Alibaba), Guanjun Jiang (Qwen Large Model Application Team, Alibaba)

CodeGenerationData SynthesisExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringMixture of ExpertsTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposed a retrieval-augmented generation framework called MARCH based on multi-agent reinforcement learning, aimed at eliminating hallucinations in large language models during retrieval-augmented generation (RAG) tasks.

Market-Bench: Benchmarking Large Language Models on Economic and Trade Competition

Yushuo Zheng (Shanghai Jiao Tong University), Guangtao Zhai (Shanghai Jiao Tong University)

CodeReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkFinance Related

🎯 What it does: Propose Market-Bench, a closed-loop multi-agent supply chain economic simulation environment, for evaluating the economic decision-making capabilities of large language models in procurement auctions, pricing, and marketing language generation.

MASPO: Unifying Gradient Utilization, Probability Mass, and Signal Reliability for Robust and Sample-Efficient LLM Reasoning

Xiaoliang Fu (Meituan), Xunliang Cai (Meituan)

CodeOptimizationComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsTextBenchmark

🎯 What it does: A unified RLVR algorithm called MASPO is studied, which addresses three major bottlenecks: gradient utilization, probability mass sensitivity, and signal reliability.

Massively Multilingual Joint Segmentation and Glossing

Michael Ginn (University of Colorado Boulder), Alexis Palmer (University of Colorado Boulder)

CodeSegmentationTransformerLarge Language ModelReinforcement LearningPrompt EngineeringContrastive LearningText

🎯 What it does: Proposed and trained a multilingual joint segmentation and interlinear glossing model called POLYGLOSS, which can output morphological segmentation of words and corresponding glosses in one go.

MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems

Shuhang Chen (Zhejiang University), Yi Yang (Zhejiang University)

CodeRecognitionExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmarkChain-of-Thought

🎯 What it does: Proposes the FlowVerse benchmark for fine-grained evaluation of perception and reasoning abilities in visual math problems, and designs a multi-module solution called MathFlow that separates perception and reasoning. A specialized model, MathFlow-P-7B, is trained to improve the extraction and description of graphics.

MauBERT: Universal Phonetic Inductive Biases for Few-Shot Acoustic Units Discovery

Angelo Ortiz Tandazo (ENS), Emmanuel Dupoux (ENS)

CodeRecognitionTransformerSupervised Fine-TuningAuto EncoderContrastive LearningAudio

🎯 What it does: A multilingual extension of HuBERT (MAUBERT) was constructed, further training language-agnostic and context-invariant speech representations through supervised learning using speech-to-articulatory features on 55 languages.

MCP-Flow: Facilitating LLM Agents to Master Real-World, Diverse and Scaling MCP Tools

WenHao Wang, Siheng Chen (Shanghai Jiao Tong University)

CodeAutonomous DrivingOptimizationComputational EfficiencyData-Centric LearningAI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringDiffusion modelScore-based ModelFlow-based ModelTextTabularRetrieval-Augmented Generation

🎯 What it does: Propose MCP-Flow, which automatically collects MCP servers and tools across multiple platforms, generating over 60k instruction-function call pairs, and uses them to train LLMs to master MCP skills.

MeanAudio: Fast and Faithful Text-to-Audio Generation with Mean Flows

Xiquan Li (Shanghai Jiao Tong University), Xie Chen (Shanghai Jiao Tong University)

CodeGenerationData SynthesisTransformerDiffusion modelFlow-based ModelAuto EncoderContrastive LearningTextMultimodalityAudio

🎯 What it does: Proposes MeanAudio, a text-to-audio generation model based on Mean Flow, capable of achieving high-quality audio synthesis in a single step (1 NFE);

Measuring Social Bias in Vision-Language Models with Face-Only Counterfactuals from Real Photos

Haodong Chen (Harbin Institute of Technology), Jun Yu (Harbin Institute of Technology)

CodeExplainability and InterpretabilityData-Centric LearningPrompt EngineeringVision Language ModelDiffusion modelContrastive LearningImageTextMultimodality

🎯 What it does: Studied the measurement of social bias in vision-language models using real photos with only minor facial modifications.

Mechanistic Interpretability Should Prioritize Feature Consistency in Sparse Autoencoders

Xiangchen Song (Carnegie Mellon University), Kun Zhang (Carnegie Mellon University)

CodeExplainability and InterpretabilityRepresentation LearningLarge Language ModelAuto EncoderContrastive LearningTextTabularTime SeriesSequential

🎯 What it does: This study investigates the feature consistency issue of sparse autoencoders (SAE) in terms of mechanistic interpretability, proposes and evaluates the Pairwise Dictionary Mean Correlation Coefficient (PW-MCC) metric based on dictionary matching, and theoretically proves and experimentally verifies its ability to achieve high consistency in TopK SAE.

MED-COREASONER: Reducing Language Disparities in Medical Reasoning via Language-Informed Co-Reasoning

Fan Gao (University of Tokyo), Irene Li (University of Tokyo)

CodeRecommendation SystemFederated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningData-Centric LearningDrug DiscoveryTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBiomedical DataElectronic Health RecordsBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed and implemented the MED-COREASONER framework, leveraging parallel reasoning chains in English and local languages, concept extraction and fusion, and retrieval enhancement to improve the quality of medical multilingual reasoning.

MediEval: A Unified Medical Benchmark for Patient-Contextual and Knowledge-Grounded Reasoning in LLMs

Zhan Qu (TU Dresden), Michael FΓ€rber (TU Dresden)

CodeExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningDrug DiscoveryTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringMixture of ExpertsTextBiomedical DataElectronic Health RecordsBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed the MediEval benchmark, combining MIMIC-IV medical records with the UMLS knowledge base, to evaluate LLMs in terms of factual accuracy and patient context consistency, and designed the CoRFu fine-tuning method based on this.

Mediocrity is the key for LLM as a Judge Anchor Selection

Shachar Don-Yehiya (Hebrew University of Jerusalem), Omri Abend (Hebrew University of Jerusalem)

CodeTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Investigated the impact of using anchors (benchmark models) for pairwise comparisons on ranking reliability when large language models (LLMs) act as judges (LLM-as-a-Judge, LMJ), systematically evaluating the performance of 22 anchors on the Arena-Hard-v2.0 dataset.

MedMCP-Calc: Benchmarking LLMs for Realistic Medical Calculator Scenarios via MCP Integration

Yakun Zhu (Shanghai Jiao Tong University), Xiaofan Zhang (Shanghai Jiao Tong University)

CodeDrug DiscoveryTransformerLarge Language ModelReinforcement LearningAgentic AITextTabularBiomedical DataElectronic Health RecordsBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Constructed the MedMCP-Calc benchmark to evaluate the performance of LLMs in real medical calculator workflows using fuzzy task descriptions, EHR data interaction, and MCP tool integration, and on this basis proposed and trained the CalcMate model.

MedVerse: Efficient and Reliable Medical Reasoning via DAG-Structured Parallel Execution

Jianwen Chen (University Of North Carolina Chapel Hill), Huaxiu Yao (University Of North Carolina Chapel Hill)

CodeComputational EfficiencyDrug DiscoveryGraph Neural NetworkTransformerLarge Language ModelTextGraphBiomedical DataElectronic Health RecordsBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose the MedVerse framework, which restructures medical reasoning as a directed acyclic graph (DAG) and enables parallel inference;

MegaRAG: Multimodal Knowledge Graph-Based Retrieval Augmented Generation

Chi-Hsiang Hsiao (National Taiwan University), Chu-song Chen

CodeRetrievalExplainability and InterpretabilityComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerLarge Language ModelVision Language ModelContrastive LearningImageTextMultimodalityGraphRetrieval-Augmented Generation

🎯 What it does: Propose MegaRAG, an end-to-end framework for automatically constructing a multi-modal knowledge graph (MMKG) and using it for retrieval-augmented generation;

Mem^2Evolve: Towards Self-Evolving Agents via Co-Evolutionary Capability Expansion and Experience Distillation

Zihao Cheng (Beihang University), Yunhong Wang (Beijing Institute Of Technology)

CodeAutonomous DrivingOptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningRobotic IntelligenceMeta LearningAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextSequentialBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose the Memβ€―2β€―Evolve framework, achieving dual mechanisms of asset memory (Asset Memory) and experience memory (Experience Memory), constructing a forward reasoning and backward evolution loop, realizing a self-evolving language model agent;

Memory-Augmented LLM-based Multi-Agent System for Automated Feature Generation on Tabular Data

FengXian Dong, Enhong Chen (University of Science and Technology of China)

CodeOptimizationExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelAgentic AIPrompt EngineeringTabularRetrieval-Augmented Generation

🎯 What it does: A framework based on multi-agent and memory-enhanced large language models (MALMAS) was constructed for automated feature generation, supporting multi-round iterations and guiding feature generation and evaluation through multi-level memory (procedural memory, feedback memory, conceptual memory, and global conceptual memory).

MemRec: Collaborative Memory-Augmented Agentic Recommender System

Weixin Chen (Hong Kong Baptist University), Yongfeng Zhang (Rutgers University)

CodeRecommendation SystemGraph Neural NetworkTransformerLarge Language ModelAgentic AIContrastive LearningTextGraphBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposed the MemRec framework, which improves the performance of LLM agent recommendation systems by utilizing collaborative memory and asynchronous propagation mechanisms.

Merlin’s Whisper: Enabling Efficient Reasoning in Large Language Models via Black-box Persuasive Prompting

Heming Xia (Hong Kong Polytechnic University), Wenjie Li (Hong Kong Polytechnic University)

CodeComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextBenchmark

🎯 What it does: Studied reducing overthinking in large language models through black-box persuasive prompting, and proposed the WHISPER framework to achieve efficient reasoning.

Metaphor Reasoning is Meta-reasoning

Qianyu He (Fudan University), Yanghua Xiao (Fudan University)

CodeGenerationMeta LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AIPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: This paper constructs an automated system called METAR to generate high-quality metaphor riddles and uses these riddles to train large language models to enhance their reasoning capabilities;

METER: Evaluating Multi-Level Contextual Causal Reasoning in Large Language Models

Pengfeng Li (Sichuan University), See-Kiong Ng (National University of Singapore)

CodeExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringTextMultimodalityBenchmarkChain-of-Thought

🎯 What it does: Proposed and implemented the METER benchmark, which evaluates the performance of large language models in three-layer causal reasoning (causal discovery, intervention, counterfactual) under a unified context, and constructed a multiple-choice dataset with 4,145 samples.

MetFuse: Figurative Fusion between Metonymy and Metaphor

Saptarshi Ghosh (University of Cincinnati), Tianyu Jiang (University of Cincinnati)

CodeGenerationData SynthesisTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBenchmark

🎯 What it does: A framework for converting textual sentences into metonymy, metaphor, and mixed sentences was constructed, and this framework was used to generate the first MetFuse dataset containing mixed expressions of metonymy and metaphor (totaling 4,000 sentences).

Min-k Sampling: Decoupling Truncation from Temperature Scaling via Relative Logit Dynamics

Yuanhao Ding (Henan University), Chongsheng Zhang (Henan University)

CodeGenerationExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose a new decoding sampling method called Mink Sampling, which can dynamically identify 'semantic cliffs' in the logit distribution and truncate the candidate set without relying on temperature parameters.

Miner: Mining Intrinsic Mastery for Data-Efficient RL in Large Reasoning Models

Shuyang Jiang (Fudan University), Yu Wang (Shanghai Jiao Tong University)

CodeComputational EfficiencyData-Centric LearningTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBiomedical DataBenchmark

🎯 What it does: Propose the MINER framework, which generates self-supervised rewards by leveraging the intrinsic uncertainty from positive homogeneous (PH) rollout, significantly improving the data efficiency of RLVR on large-scale reasoning models.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Long-tail Knowledge

Jie He (University of Edinburgh), Jeff Z. Pan (University of Edinburgh)

CodeTransformerLarge Language ModelTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This paper proposes MINTQA, a multi-hop question answering benchmark that evaluates the performance of large language models in multi-hop reasoning and retrieval fusion by combining two dimensions: new/old knowledge and popular/unpopular knowledge.

Mirroring Users: Towards Building Preference-aligned User Simulator with User Feedback in Recommendation

Tianjun Wei (Nanyang Technological University), Jie Zhang (Nanyang Technological University)

CodeRecommendation SystemExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringText

🎯 What it does: By constructing the USERMIRRORER framework, the decision-making process of user feedback in recommendation systems is utilized to achieve fine-grained alignment with LLMs, thereby realizing more accurate user simulation.

MIRTH: Mutual-Information Reasoning with Temporal Hubs for Vision-Language-Action Agents

Hao Sun (Ritsumeikan University), Yen-wei Chen

CodeExplainability and InterpretabilityComputational EfficiencyRepresentation LearningRobotic IntelligenceTransformerReinforcement LearningPrompt EngineeringMixture of ExpertsVision-Language-Action ModelContrastive LearningImageVideoTextMultimodality

🎯 What it does: Proposes a unified Vision-Language-Action framework called MIRTH, addressing the issues of short-term temporal vision, inference gap, and low inference efficiency in single-frame VLA models.

Mitigating Context Interference for Reliable and Efficient Search Agents

Boyang Xue (Chinese University of Hong Kong), Aldo Lipani (University College London)

CodeComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Systematically study context interference in multi-round search agents, finding that it mainly comes from the latest retrieved documents, and propose a context refiner based on distillation, which is then embedded into the reinforcement learning training process to improve the reliability and efficiency of the search agent.

Mitigating Lost in Multi-turn Conversation via Curriculum RL with Verifiable Accuracy and Abstention Rewards

Ming Li (University of Maryland), Bing Yin (Amazon)

CodeOptimizationAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringText

🎯 What it does: Designed and trained a reinforcement learning-based framework called RLAAR, encouraging LLMs to both answer correctly and determine when to give up in multi-turn conversations

MixKVQ: Query-Aware Mixed-Precision KV Cache Quantization for Long-Context Reasoning

Tao Zhang (South China University Of Technology), Cen Chen (Beihang University)

CodeComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelPrompt EngineeringText

🎯 What it does: Propose MixKVQ, a low-bit quantization method for the KV cache of large language models, aiming to significantly reduce memory usage while maintaining the accuracy of long-text inference.

MM-StanceDet: Retrieval-Augmented Multi-modal Multi-agent Stance Detection

Weihai Lu, Huan He (Brown University)

CodeClassificationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelAgentic AIPrompt EngineeringImageTextMultimodalityRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed the MM-StanceDet multi-agent framework, achieving more robust multi-modal stance detection through four stages: retrieval enhancement, specialized multi-modal analysis, debate, and self-reflection.

MMAC: A Multilingual, Multimodal Alignment Framework for Cultural Grounding Evaluation

Weihua Zheng (Agency for Science, Technology and Research), Nancy F. Chen (Singapore University of Technology and Design)

CodeExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-ThoughtAudio

🎯 What it does: Proposed the MMAC framework and the MMAC-bench dataset to systematically evaluate the cultural cognition and reasoning capabilities of large language models in multilingual, multimodal (text, image, voice) environments.

Mnemis: Dual-Route Retrieval on Hierarchical Graphs for Long-Term LLM Memory

Zihao Tang (Microsoft), Qi Zhang (Microsoft)

CodeRetrievalComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringTextGraphBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposed a long-term memory framework called Mnemis, which integrates traditional similarity retrieval (System-1) with global hierarchical retrieval (System-2) routing;

MoA: Heterogeneous Mixture of Adapters for Parameter-Efficient Fine-Tuning of Large Language Models

Jie Cao (Zhejiang University), Yueting Zhuang (Zhejiang University)

CodeComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsText

🎯 What it does: Designed a heterogeneous hybrid adapter (MoA) for parameter-efficient fine-tuning of large language models.

Modeling LLM Unlearning as an Asymmetric Two-Task Learning Problem

Zeguan Xiao (Shanghai University of Finance and Economics), Guanhua Chen (Southern University of Science and Technology)

CodeFederated LearningSafty and PrivacyComputational EfficiencyTransformerLarge Language ModelContrastive LearningTextBenchmark

🎯 What it does: View LLM unlearning as asymmetric multi-task learning, and propose a gradient synthesis framework prioritizing retention.

MONETA: Multimodal Industry Classification through Geographic Information with Multi Agent Systems

Arda YΓΌksel (Technical University of Darmstadt), Ivan Habernal (Ruhr University Bochum)

CodeClassificationTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed the MONETA multimodal industry classification benchmark, which uses text (websites, Wikipedia, Wikidata) and geospatial information (OpenStreetMap, satellite images) to perform NACE classification on 1,000 European companies, aiming to replace manual expert verification;

MQM Re-Annotation: A Technique for Collaborative Evaluation of Machine Translation

Parker Riley (Google), Markus Freitag (Google)

CodeData-Centric LearningPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposed and experimentally tested the MQM re-annotation method, allowing reviewers to delete, modify, or add errors based on existing error annotations, thereby improving fine-grained quality assessment in machine translation evaluation;

MSMO-ABSA: Multi-Scale and Multi-Objective Optimization for Cross-Lingual Aspect-Based Sentiment Analysis

Chengyan Wu (South China Normal University), Liu Xiaoyong

CodeDomain AdaptationOptimizationKnowledge DistillationRepresentation LearningTransformerSupervised Fine-TuningContrastive LearningText

🎯 What it does: Propose the MSMO framework, combining sentence-level adversarial training, aspect-level consistency training, and multi-objective optimization, to achieve cross-lingual ABSA feature alignment and fine-grained alignment, and then perform knowledge distillation based on this.

MT^{3}: A Synergistic Multi-Task RL Framework for Specializing MLLMs in Text Image Machine Translation

Zhaopeng Feng (Zhejiang University), Zuozhu Liu (Zhejiang University)

CodeImage TranslationTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision Language ModelImageTextMultimodalityBenchmark

🎯 What it does: Propose the MT3 framework, which utilizes multi-task reinforcement learning to specialize multimodal large language models (MLLMs) into end-to-end text-image machine translation (TIMT) expert models.

MTR-Suite: A Framework for Evaluating and Synthesizing Conversational Retrieval Benchmarks

Junhao Ruan (Northeastern University), JingBo Zhu

CodeData SynthesisRetrievalTransformerLarge Language ModelAgentic AIPrompt EngineeringContrastive LearningTextBenchmarkFinance RelatedRetrieval-Augmented Generation

🎯 What it does: Propose the MTR-Suite framework, integrating evaluation (MTR-EVAL), multi-agent synthesis (MTR-PIPELINE), and a new dialogue retrieval benchmark (MTR-BENCH).

MTSQL-R1: Towards Long-Horizon Multi-Turn Text-to-SQL via Agentic Training

Taicheng Guo (University of Notre Dame), Chandan K. Reddy (Amazon)

CodeExplainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringTextTabularBenchmarkChain-of-Thought

🎯 What it does: Proposed the MTSQL-R1 framework, modeling multi-turn Text-to-SQL as a Markov Decision Process, supporting agent-based execution, verification, and self-correction;

Multi-Granularity Semantic Revision for Large Language Model Distillation

Xiaoyu Liu (University of Science and Technology of China), Yunhe Wang (Huawei)

CodeComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: Proposes a multi-granularity semantic revision method to improve the knowledge distillation process of large language models.

Multi-View Attention Multiple-Instance Learning Enhanced by LLM Reasoning for Cognitive Distortion Detection

Jun Seo Kim (Gachon University), Hye Hyeon Kim (Yonsei University)

CodeClassificationExplainability and InterpretabilityTransformerLarge Language ModelText

🎯 What it does: Propose a method that splits each sentence into three parts: emotion, logic, and behavior (ELB), then uses a large language model (LLM) to generate multiple instances of cognitive distortions, and finally classifies them using a multi-instance learning (MIL) framework with multi-perspective gated attention.

Multilingual and Cross-Lingual Citation Needed Detection on Wikipedia for Lower-Resource Languages

Gerrit Quaremba (King's College London), Elena Simperl (King's College London)

CodeClassificationKnowledge DistillationTransformerLarge Language ModelSupervised Fine-TuningTextBenchmark

🎯 What it does: Constructed a cross-lingual 'Need to Cite' detection dataset (MCN), and trained and evaluated a small decoder model on 18 languages with different resource levels.

Multimodal Safety Evaluation in Generative Agent Social Simulations

Alhim Adonai Vera Gonzalez (University of Cincinnati), Bernard Ghanem

CodeSafty and PrivacyTransformerLarge Language ModelAgentic AIPrompt EngineeringVision Language ModelImageTextMultimodalityRetrieval-Augmented Generation

🎯 What it does: A reproducible multimodal safety evaluation framework was constructed, and generative agents were used in social simulation environments to detect and correct unsafe plans, analyzing safety improvements and social dynamics.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding

Fuwen Luo (Tsinghua University), Yang Liu (Tsinghua University)

CodeComputational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringVision Language ModelVideoTextMultimodality

🎯 What it does: Proposed a multi-segment temporal alignment method called MUSEG based on reinforcement learning, enhancing the temporal reasoning ability of multi-modal large language models.

MUTANT: A Recipe for Multilingual Tokenizer Design

Souvik Rana (Krutrim AI), Shubham Agarwal (Krutrim AI)

CodeComputational EfficiencyData-Centric LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextMultimodalityBenchmark

🎯 What it does: Designed a multilingual tokenizer training process called MUTANT, and constructed MUTANT-Indic tailored for Indian languages.

Native Hybrid Attention for Efficient Sequence Modeling

Jusen Du (Tsinghua University), Yu Cheng (Chinese University of Hong Kong)

CodeComputational EfficiencyRepresentation LearningRecurrent Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningTextSequential

🎯 What it does: Propose Native Hybrid Attention (NHA), a hybrid attention architecture that simultaneously integrates linear RNN memory and sliding window Softmax attention at the same level.

NaturalGAIA: A Verifiable Benchmark and Hierarchical Framework for Long-Horizon GUI Tasks

Zihan Zheng (South China Normal University), Qianglong Chen (Zhejiang University)

CodeRobotic IntelligenceReinforcement Learning from Human FeedbackTransformerLarge Language ModelVision-Language-Action ModelTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposes the NaturalGAIA benchmark and the LightManus-Jarvis hierarchical framework for evaluating and enhancing the performance of LLM agents in long-term GUI tasks.

Neo-Classic: A Benchmark for Evaluating Linguistic-Aesthetic Reasoning in Classical Chinese Poetry

Han Zhang (Shanghai Jiao Tong University), Cheng Hua (Shanghai Jiao Tong University)

CodeTransformerLarge Language ModelPrompt EngineeringTextBenchmarkChain-of-Thought

🎯 What it does: Proposes the NEO-CLASSIC benchmark, which uses strictly metered poems created by modern experts to evaluate the language aesthetic reasoning ability of LLMs.

NeoAMT: Neologism-Aware Agentic Machine Translation with Reinforcement Learning

Zhongtao Miao (The University of Tokyo), Yoshimasa Tsuruoka (The University of Tokyo)

CodeTransformerLarge Language ModelReinforcement LearningAgentic AITextRetrieval-Augmented Generation

🎯 What it does: Constructed a multilingual neologism machine translation dataset called Neko, and proposed the NeoAMT framework, which uses reinforcement learning and dictionary retrieval to translate sentences containing neologisms.

Neuron-Aware Active Few-Shot Learning for LLMs

Zhuowei Chen (University of Pittsburgh), Xiang Lorraine Li (University of Pittsburgh)

CodeClassificationExplainability and InterpretabilityMeta LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningText

🎯 What it does: Proposes NEUFS, an active few-shot learning framework based on dynamic neuronal activation, specifically designed to select the most valuable few-shot examples for annotation from unlabeled data in specialized domains for large language models (LLMs);

NOSE: Neural Olfactory-Semantic Embedding with Tri-Modal Orthogonal Contrastive Learning

Yanyi Su (Xiamen University), Jun Cheng (Xiamen University)

CodeDrug DiscoveryGraph Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningTextMultimodalityGraph

🎯 What it does: Propose a tri-modal (molecular structure, receptor sequence, semantic description) aligned olfactory embedding framework called NOSE, achieving modal information decoupling and fusion through orthogonal injection and weak positive sample contrastive learning.

Not All Tokens Matter: Towards Efficient LLM Reasoning via Token Significance in Reinforcement Learning

Hanbing Liu (Tsinghua University), Dongmei Zhang (Microsoft)

CodeComputational EfficiencyTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkChain-of-Thought

🎯 What it does: Propose the BINGO framework, which improves the efficiency of large model chain-of-thought reasoning by utilizing token importance awareness and dynamic length rewards.

OASIS: Online Sample Selection for Continual Instruction Tuning

Minjae Lee (Seoul National University), Jonghyun Choi (Seoul National University)

CodeOptimizationComputational EfficiencyData-Centric LearningTransformerSupervised Fine-TuningPrompt EngineeringContrastive LearningTextMultimodality

🎯 What it does: OASIS proposes an online adaptive sample selection framework in continuous instruction tuning (CIT), which can real-time select the most informative samples from the data stream, significantly shortening training time while maintaining the model's real-time adaptation capability.

OctoBench: Benchmarking Scaffold-Aware Instruction Following in Repository-Grounded Agentic Coding

Deming Ding (Fudan University), Tao Gui (Fudan University)

CodeAI Code AssistantTransformerLarge Language ModelAgentic AITextBenchmark

🎯 What it does: This paper proposes the OCTOBENCH benchmark to evaluate whether models can follow multi-source, persistent instruction constraints in warehouse-level agent-based coding.

OLA: Output Language Alignment in Code-Switched LLM Interactions

Juhyun Oh (KAIST), Alice Oh (KAIST)

CodeExplainability and InterpretabilityData-Centric LearningAI Code AssistantTransformerLarge Language ModelPrompt EngineeringTextMultimodalityBenchmarkChain-of-Thought

🎯 What it does: The study investigates the failure of large language models (LLMs) to implicitly output language alignment during code-switching interactions, constructs the OLA benchmark, and proposes the CS-DPO method based on preference alignment.

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Models

Qiguang Chen (Harbin Institute of Technology), Wanxiang Che (Harbin Institute of Technology)

CodeExplainability and InterpretabilityComputational EfficiencyRepresentation LearningDrug DiscoveryTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmarkPhysics RelatedRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose OMIBench, a multi-graph Olympiad-level reasoning benchmark, to evaluate the ability of large vision-language models in multi-graph reasoning.

Omni-I2C: A Holistic Benchmark for High-Fidelity Image-to-Code Generation

Jiawei Zhou (Wuhan University), Jing Zhang (Wuhan University)

CodeGenerationAI Code AssistantTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose Omni-I2C, a large-scale, cross-lingual, and multi-domain Image-to-Code benchmark, to evaluate the ability of large multimodal models to convert complex digital graphics into executable code.

On the Emotion Understanding of Synthesized Speech

Yuan Ge (Northeastern University), Tong Xiao (Kunming University of Science and Technology)

CodeRecognitionDomain AdaptationTransformerLarge Language ModelContrastive LearningTextAudio

🎯 What it does: Systematically evaluate the generalization ability of synthetic speech emotion recognition models, and analyze the gap between human speech and synthetic speech in emotional understanding;

On-policy Reinforcement Fine-tuning with Offline reward for Multi-step Embodied Planning

Di Wu (Tongji University), Bo Jin (Tongji University)

CodeRobotic IntelligenceTransformerSupervised Fine-TuningReinforcement LearningVision-Language-Action ModelImageTextMultimodality

🎯 What it does: Propose the ORBIT framework, which uses on-policy reinforcement fine-tuning (RFT) with offline expert trajectory rewards to address the high-cost interaction and sparse reward problems in multi-step embodied planning.

One Battle After Another: Probing LLMs’ Limits on Multi-Turn Instruction Following with a Benchmark Evolving Framework

Qi Jia (Shanghai Artificial Intelligence Laboratory), Guangtao Zhai (Shanghai Jiao Tong University)

CodeExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringFlow-based ModelTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposed an scalable multi-turn instruction-following evaluation framework and the EvolIF benchmark, utilizing a three-layer tracking mechanism and a query synthesis agent driven by large language models (LLMs) to dynamically generate dialogues, and introducing a patience threshold based on Flow theory and process-oriented evaluation metrics;

One Persona, Many Cues, Different Results: How Sociodemographic Cues Impact LLM Personalization

Franziska Weeber (University of Stuttgart), Sebastian PadΓ³ (University of Stuttgart)

CodeExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextBenchmark

🎯 What it does: Investigated the impact of six commonly used persona prompts (name, explicit description, conversation history) on the personalization bias of large language models, and constructed a scalable multi-task evaluation benchmark.

One-step Nonautoregressive Natural Language Generation with Shortcut Flow Matching Models

JΔ™drzej WarczyΕ„ski (Poznan University of Technology), Mateusz Lango (Poznan University of Technology)

CodeGenerationTransformerDiffusion modelScore-based ModelFlow-based ModelRectified FlowText

🎯 What it does: This paper proposes a first-order non-autoregressive natural language generation method that completes text generation in a single step using the shortcut flow matching model.

Open Your Model’s Eyes: Video and Context-Aware Multimodal Backchannel Prediction

Min-Jae Kim (Korea University), Gyeong-Moon Park (Korea University)

CodeClassificationTransformerSupervised Fine-TuningVision Language ModelContrastive LearningVideoTextMultimodalityAudio

🎯 What it does: A novel multimodal behind-channel prediction framework called CAMA-BC was studied, which integrates audio, text, and video information and achieves modality alignment through hierarchical cross-attention, addressing issues such as visual information bias, temporal deviation, and imbalance between context and response.

Outcome-Grounded Advantage Reshaping for Fine-Grained Credit Assignment in Mathematical Reasoning

Ziheng Li (Fudan University), Hongcheng Guo (Fudan University)

CodeOptimizationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsTextBenchmarkChain-of-Thought

🎯 What it does: Proposes a fine-grained credit assignment framework called OAR based on the impact of final answers, aimed at improving reward propagation in Group Relative Policy Optimization (GRPO) for long reasoning tasks.

Overcoming Copyright Barriers in Corpus Distribution Through Non-Reversible Hashing

Arthur Amalvy (Academia Sinica), Hen-Hsen Huang (Academia Sinica)

CodeSafty and PrivacyPrompt EngineeringText

🎯 What it does: Propose a method that utilizes non-reversible hashing to share copyrighted text annotations, allowing users with the original text to legally obtain annotations.

PAR: Training-Free Positional Perturbation and Attention Recycling for Faithful OCR

Yao Yao (Shanghai Jiao Tong University), Hai Zhao (Shanghai Jiao Tong University)

CodeRecognitionTransformerPrompt EngineeringVision Language ModelDiffusion modelContrastive LearningImageTextMultimodalityBenchmark

🎯 What it does: Propose a training-agnostic, inference-time intervention framework called PAR, designed to suppress hallucinations caused by language priors in visual language models during OCR tasks, thereby improving the visual consistency of text recognition.

PARIF: Pushing the Pareto Frontier of Instruction Following and Reasoning with Curriculum Reinforcement Learning

Rongchuan Mu (Harbin Institute of Technology), Bing Qin (Harbin Institute of Technology)

CodeOptimizationComputational EfficiencyRepresentation LearningData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringDiffusion modelAuto EncoderContrastive LearningTextChain-of-Thought

🎯 What it does: Propose a two-stage RLVR curriculum learning framework called PARIF to enhance the comprehensive performance of large reasoning models in instruction following and reasoning capabilities.

PEAP: Proactive Embodied Action Sequence Planning with Joint Understanding of Vision and Audio Perception

Tianwei Lan (Beijing Institute Of Technology), Yuhang Guo (Beihang University)

CodeAutonomous DrivingComputational EfficiencyRepresentation LearningRobotic IntelligenceTransformerLarge Language ModelPrompt EngineeringVision-Language-Action ModelContrastive LearningImageMultimodalityAudio

🎯 What it does: Studied the task of planning active avatar action sequences in a multimodal (visual + audio) setting, and constructed the PEAP dataset.

PEAR: Pairwise Evaluation for Automatic Relative Scoring in Machine Translation

Lorenzo Proietti (Sapienza University of Rome), Matt Post (Microsoft)

CodeClassificationRepresentation LearningData-Centric LearningTransformerSupervised Fine-TuningContrastive LearningText

🎯 What it does: Propose PEAR, a contrast-based quality estimation (QE) metric that can perform graded relative quality difference assessment between two candidate translations of the same source text.

PersonalAlign: Hierarchical Implicit Intent Alignment for Personalized GUI Agent with Long-Term User-Centric Records

Yibo Lyu (Harbin Institute of Technology), Liqiang Nie (Harbin Institute of Technology)

CodeRecommendation SystemReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringMixture of ExpertsContrastive LearningTextMultimodalitySequentialBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose the PersonalAlign task, construct the AndroidIntent benchmark, and design the HIM-Agent memory framework for hierarchical implicit intent alignment of long-term user records.

PersonalityDBench: A Dataset for Personality Disorders - from Modeling to Controlled Generation

Federico Ravenda (UniversitΓ  della Svizzera italiana), Andrea Raballo (UniversitΓ  della Svizzera italiana)

CodeClassificationRecognitionGenerationData SynthesisTransformerLarge Language ModelContrastive LearningTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Constructed the PersonalityDBench dataset, which includes clinically annotated Reddit language samples (PRISMA) and a benchmark (PersonaDSteering) for evaluating the controllability of LLMs in generating behaviors related to personality disorders. The feasibility of diagnosing personality disorders in natural language, HiTOP dimension features, and LLM directional control were validated from this dataset.

PIArena: A Platform for Prompt Injection Evaluation

Runpeng Geng (Pennsylvania State University), Jinyuan Jia (Pennsylvania State University)

CodeSafty and PrivacyAdversarial AttackTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposed the PIArena unified platform for evaluating prompt injection attacks and defenses, and designed an adaptive strategy attack;

PIAST: Rapid Prompting with In-context Augmentation for Scarce Training data

PaweΕ‚ Batorski (Heinrich Heine UniversitΓ€t DΓΌsseldorf), Paul Swoboda (Heinrich Heine UniversitΓ€t DΓΌsseldorf)

CodeClassificationComputational EfficiencyData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: Propose a fast automated prompt construction method called PIAST, which utilizes LLMs to generate and iteratively improve a few few-shot examples, thereby enhancing the performance of gradient-free updated LLMs.

PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media

Zoher Kachwala (Indiana University), Filippo Menczer (Indiana University)

CodeTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposed and constructed the PLURULE benchmark for detecting whether comments violate specific community rules in a multilingual, multimodal community environment, simulating real moderators' decision-making through multiple-choice questions.

PodBench: A Comprehensive Benchmark for Instruction-Aware Audio-Oriented Podcast Script Generation

Chenning Xu (Tencent), Mingyang Song (Tencent)

CodeGenerationData SynthesisTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-ThoughtAudio

🎯 What it does: Constructed PodBench benchmark, focusing on instruction-aware and context-driven long-form multi-speaker podcast script generation tasks, providing 800 long-context samples and multi-dimensional instructions;

PRBench: Large-Scale Expert Rubrics for Evaluating High-Stakes Professional Reasoning

Afra Feyza AkyΓΌrek (Scale AI), Yunzhong He (Scale AI)

CodeLarge Language ModelPrompt EngineeringTextBenchmarkFinance RelatedRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This work constructs and publicly releases a large professional reasoning benchmark called PRBench, which includes 1,100 real-world task scenarios written by financial and legal experts, along with 18,711 finely crafted expert rubrics, corresponding multi-turn dialogues, and economic impact annotations.

Preference Heads in Large Language Models: A Mechanistic Framework for Interpretable Personalization

Weixu Zhang (McGill University), Haolun Wu (McGill University)

CodeExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringContrastive LearningText

🎯 What it does: Propose a training-free differential preference-driven (DPS) method that identifies sparse 'preference heads' based on mechanism interpretation and controls them during decoding to achieve interpretable personalization.

PrefixNLI: Detecting Factual Inconsistencies as Soon as They Arise

Sapir Harary (Bar Ilan University), Ido Dagan (Bar Ilan University)

CodeGenerationExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: Propose the PrefixNLI task to detect factual inconsistencies during the text generation process.

PRIME: A Process-Outcome Alignment Benchmark for Verifiable Reasoning in Mathematics and Engineering

Xiangfeng Wang (University of Science and Technology of China), Daxin Jiang (Stepfun)

CodeExplainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed the PRIME benchmark for evaluating process-result consistency verification of large reasoning models in the fields of mathematics and engineering.

PRInTS: Reward Modeling for Long-Horizon Information Seeking

Jaewoo Lee (University of North Carolina at Chapel Hill), Mohit Bansal (University of North Carolina at Chapel Hill)

CodeAutonomous DrivingOptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningAdversarial AttackHyperparameter SearchData-Centric LearningRobotic IntelligenceMeta LearningDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackNeural Architecture SearchProtein Structure PredictionTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AIMixture of ExpertsDiffusion modelScore-based ModelFlow-based ModelRectified FlowNeural Radiance FieldAuto EncoderGenerative Adversarial NetworkContrastive LearningGaussian SplattingSimultaneous Localization and MappingWorld ModelOptical FlowTextBenchmarkRetrieval-Augmented GenerationChain-of-ThoughtStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: This paper proposes PRInTS, a generative process reward model that combines information gain scoring with recursive trajectory summarization, enabling fine-grained evaluation of each step (reasoning + tool call) in long-term information-seeking tasks, and providing guidance for selection during testing for LLM agents.

PRiSM: Benchmarking Phone Realization in Speech Models

Shikhar Bharadwaj (Carnegie Mellon University), David R. Mortensen (Carnegie Mellon University)

CodeRecognitionConvolutional Neural NetworkRecurrent Neural NetworkTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBenchmarkAudio

🎯 What it does: Proposes the PRiSM benchmark for evaluating systems that transcribe speech into phonemes, assessing them along two major dimensions: intrinsic (PFER) and extrinsic (downstream tasks).

PRISMA: Preference-Reinforced Self-Training Approach for Interpretable Emotionally Intelligent Negotiation Dialogues

Prajwal Vijay Kajare (Indian Institute of Technology Jodhpur), Asif Ekbal (Indian Institute of Technology Patna)

CodeExplainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This paper proposes PRISMA, an interpretable emotion-intelligent negotiation dialogue system, capable of generating emotionally appropriate and interpretable responses through emotion perception and strategy selection;

PRISP: Privacy-Safe Few-Shot Personalization via Lightweight Adaptation

Junho Park (Seoul National University), Taesup Moon (Seoul National University)

CodeFederated LearningSafty and PrivacyComputational EfficiencyMeta LearningTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose a privacy-safe, lightweight few-shot personalized framework called PRISP, which can achieve user-level personalization for large language models without requiring task data or sharing user parameters.

Privacy-R1: Privacy-Aware Multi-LLM Agent Collaboration via Reinforcement Learning

Zheng Hui (University of Cambridge), Nigel Collier (University College London)

CodeSafty and PrivacyTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringTextBiomedical Data

🎯 What it does: Proposes a privacy-aware multi-LLM agent collaboration framework called Privacy-R1, which dynamically routes text blocks to local or remote models while maintaining task performance and reducing the leakage of sensitive information.

Programming over Thinking: Efficient and Robust Multi-Constraint Planning

Derrick Goh Xin Deik (Nanyang Technological University), Wenya Wang (Nanyang Technological University)

CodeOptimizationAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringTextBenchmarkChain-of-Thought

🎯 What it does: Proposes the SCOPE framework, which decomposes multi-constraint planning into query-specific reasoning and general code execution, automatically generating reusable solver functions.

Projecting Out the Malice: A Global Subspace Approach to LLM Detoxification

Zenghao Duan (State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences), Xueqi Cheng (State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences)

CodeFederated LearningExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringText

🎯 What it does: A lightweight method named GLOSS is proposed, which achieves model detoxification by identifying and eliminating the global toxic subspace of parameters in the Feed-Forward network of LLMs.

ProMed: Shapley Information Gain Guided Reinforcement Learning for Proactive Medical LLMs

Hongxin Ding (Peking University), Yasha Wang (Peking University)

CodeExplainability and InterpretabilityDrug DiscoveryTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBiomedical DataBenchmark

🎯 What it does: Implementing an interactive diagnostic framework called ProMed in medical LLMs, transitioning from passive answering to active questioning.

Prosody as Supervision: Bridging the Non-Verbal–Verbal for Multilingual Speech Emotion Recognition

Girish (UPES), Muskaan Singh (Ulster University)

CodeRecognitionConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningOptical FlowAudio

🎯 What it does: Proposes an unsupervised cross-lingual emotion recognition transfer framework, called NOVA-ARC, that transfers from annotated non-linguistic sounds (such as laughter, crying, sighing) to linguistically vocalized speech.

Prune as You Generate: Online Rollout Pruning for Faster and Better RLVR

Haobo Xu (University Of Illinois At Urbana Champaign), Hanghang Tong (University Of Illinois At Urbana Champaign)

CodeComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsTextBenchmark

🎯 What it does: Propose an online episode trimming method called ARROL, which can predict the success probability of partial episodes during the generation process based on a lightweight quality head and trim them in advance, maintaining a near 0.5 ratio of positive and negative samples within the episode group;

Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness

Amin Banayeeanzade (University of Southern California), Sai Praneeth Karimireddy (University of Southern California)

CodeExplainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: Created and evaluated the PsySET benchmark, systematically comparing the effectiveness and reliability of multiple LLM emotion regulation methods (prompt engineering, vector injection, parameter-efficient fine-tuning, DPO) on emotional and personality dimensions.