arXivSub Start free trial

ACL 2026 Papers — Page 16

Annual Meeting of the Association for Computational Linguistics · 2296 papers

PEAP: Proactive Embodied Action Sequence Planning with Joint Understanding of Vision and Audio Perception

Tianwei Lan (Beijing Institute Of Technology), Yuhang Guo (Beihang University)

Autonomous DrivingComputational EfficiencyRepresentation LearningRobotic IntelligenceTransformerLarge Language ModelPrompt EngineeringVision-Language-Action ModelContrastive LearningImageMultimodalityAudio

🎯 What it does: Studied the task of planning active avatar action sequences in a multimodal (visual + audio) setting, and constructed the PEAP dataset.

PEAR: Pairwise Evaluation for Automatic Relative Scoring in Machine Translation

Lorenzo Proietti (Sapienza University of Rome), Matt Post (Microsoft)

ClassificationRepresentation LearningData-Centric LearningTransformerSupervised Fine-TuningContrastive LearningText

🎯 What it does: Propose PEAR, a contrast-based quality estimation (QE) metric that can perform graded relative quality difference assessment between two candidate translations of the same source text.

PEARL: Self-Evolving Assistant for Time Management with Reinforcement Learning

Bingxuan Li (University of Illinois at Urbana-Champaign), Heng Ji (University of Illinois at Urbana-Champaign)

Autonomous DrivingOptimizationFederated LearningAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringTextTime SeriesBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This paper proposes a calendar conflict resolution method based on large language models and reinforcement learning, which automates the handling of overlapping calendar invitations, gradually reasoning and executing user preferences;

Perception, Understanding and Reasoning: A Multimodal Benchmark for Video Fake News Detection

Cui Yakun (The Hong Kong University of Science and Technology), Sirui Han (The Hong Kong University of Science and Technology)

TransformerLarge Language ModelSupervised Fine-TuningVideoTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Designed and constructed a process-oriented evaluation benchmark for multi-modal video fake news detection, named POVFNDB, and conducted fine-grained assessment and baseline development based on it.

Perplexity-Aware Data Scaling Law: Perplexity Landscapes Predict Performance for Continual Pre-training

Lei Liu (State Key Laboratory of Blockchain and Data Security Zhejiang University), Kui Ren (State Key Laboratory of Blockchain and Data Security Zhejiang University)

Recommendation SystemData-Centric LearningDrug DiscoveryTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBiomedical DataBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose a continuous pre-training (CPT) data scaling method based on model perplexity, which predicts final performance by calculating the perplexity distribution of domain text and selects an efficient training subset accordingly.

Persona-E²: A Human-Grounded Dataset for Personality-Shaped Emotional Responses to Textual Events

Yuqin Yang (South China University of Technology), Zhanpeng Jin (South China University of Technology)

ClassificationData SynthesisExplainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextSequentialChain-of-Thought

🎯 What it does: Constructed the Persona-E2 dataset, collecting news, social media, and life experience texts, and having annotators with MBTI and BFI personality traits provide emotional labels. Subsequently, analyzed emotional differences and evaluated the effectiveness of LLMs in simulating personalized emotional expressions.

Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations

Prerna Juneja (Seattle University), Lika Lomidze (Seattle University)

Safty and PrivacyTransformerLarge Language ModelPrompt EngineeringTextReview/Survey PaperRetrieval-Augmented Generation

🎯 What it does: Proposes a multi-turn dialogue safety evaluation framework based on artificial roles, simulating interactions between high-risk individuals and AI companions to systematically capture emotions and harmful behaviors;

PersonalAlign: Hierarchical Implicit Intent Alignment for Personalized GUI Agent with Long-Term User-Centric Records

Yibo Lyu (Harbin Institute of Technology), Liqiang Nie (Harbin Institute of Technology)

Recommendation SystemReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringMixture of ExpertsContrastive LearningTextMultimodalitySequentialBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose the PersonalAlign task, construct the AndroidIntent benchmark, and design the HIM-Agent memory framework for hierarchical implicit intent alignment of long-term user records.

PersonalityDBench: A Dataset for Personality Disorders - from Modeling to Controlled Generation

Federico Ravenda (Università della Svizzera italiana), Andrea Raballo (Università della Svizzera italiana)

ClassificationRecognitionGenerationData SynthesisTransformerLarge Language ModelContrastive LearningTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Constructed the PersonalityDBench dataset, which includes clinically annotated Reddit language samples (PRISMA) and a benchmark (PersonaDSteering) for evaluating the controllability of LLMs in generating behaviors related to personality disorders. The feasibility of diagnosing personality disorders in natural language, HiTOP dimension features, and LLM directional control were validated from this dataset.

Personalizing LLMs with Binary Feedback: A Preference-Calibrated Optimization Framework

Xilai Ma (Harbin Institute of Technology), Jing Li (Harbin Institute of Technology)

Recommendation SystemOptimizationFederated LearningData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose a preference calibration optimization framework called C-BPO based on binary feedback, using the target user's history as positive samples and other users' history as implicit negative samples to personalize LLMs;

PExA: Parallel Exploration Agent for Complex Text-to-SQL

Tanmay Parekh (University of California Los Angeles), Yunmo Chen (Bloomberg)

OptimizationComputational EfficiencyAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextTabular

🎯 What it does: Propose a framework that treats the text-to-SQL task as a software testing coverage problem, utilizing parallel exploration for test case generation, execution, and synthesis of the final SQL;

Phonemes to the Rescue: Multilingual Tokenization Based on International Phonetic Alphabet

Milan Miletić (University of Amsterdam), Ekaterina Shutova (University of Amsterdam)

Computational EfficiencyRepresentation LearningData-Centric LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextAudio

🎯 What it does: This paper proposes using the International Phonetic Alphabet (IPA) as the input representation for multilingual tokenization, and constructs a corresponding subword tokenizer.

PHOTON: Hierarchical Autoregressive Modeling for Lightspeed and Memory-Efficient Language Generation

Yuma Ichikawa (Fujitsu Limited), Akira Sakai (Fujitsu Limited)

GenerationComputational EfficiencyTransformerLarge Language ModelFlow-based ModelAuto EncoderText

🎯 What it does: Propose a hierarchical autoregressive language model called PHOTON, which replaces the horizontal token-by-token scanning of Transformer with vertical multi-resolution context scanning, achieving more efficient language generation.

Phun-Bench: Evaluating LLMs on Phonological Understanding in Chinese

Xing Yue (Zhejiang University), Weiming Lu (Zhejiang University)

TransformerLarge Language ModelPrompt EngineeringTextMultimodalityBenchmarkChain-of-ThoughtAudio

🎯 What it does: Propose the concept of Chinese phonological understanding, and based on this, design a three-dimensional evaluation benchmark called Phun-Bench, covering three tasks: homonym recognition, prosody sentence generation, and speech similarity comparison, systematically evaluating the phonological reasoning capabilities of various large language models.

PIArena: A Platform for Prompt Injection Evaluation

Runpeng Geng (Pennsylvania State University), Jinyuan Jia (Pennsylvania State University)

Safty and PrivacyAdversarial AttackTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposed the PIArena unified platform for evaluating prompt injection attacks and defenses, and designed an adaptive strategy attack;

PIAST: Rapid Prompting with In-context Augmentation for Scarce Training data

Paweł Batorski (Heinrich Heine Universität Düsseldorf), Paul Swoboda (Heinrich Heine Universität Düsseldorf)

ClassificationComputational EfficiencyData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: Propose a fast automated prompt construction method called PIAST, which utilizes LLMs to generate and iteratively improve a few few-shot examples, thereby enhancing the performance of gradient-free updated LLMs.

PICTURE: Enhancing Theory-of-Mind in Large Language Models by Revealing, Not Hiding, Characters’ Lack of Knowledge

Eojin Jeon (Korea University), SangKeun Lee (Korea University)

Explainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextBenchmarkChain-of-Thought

🎯 What it does: Studies how to suppress responses to unknown events of characters in Theory of Mind reasoning in large language models, and proposes the PICTURE prompting method

Piece of Table: A Divide-and-Conquer Approach for Selecting Subtables in Table Question Answering

Wonjin Lee (POSTECH), Kwang In Kim

Explainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextTabularRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose a divide-and-conquer subtable selection framework called PieTa, which utilizes a language model to progressively extract and merge relevant cells within a window, ultimately generating a compact subtable tailored to answer the question;

PII-Bench: Evaluating Query-Aware Privacy Protection Systems

Hao Shen (Fudan University), Hongfeng Chai (Fudan University)

Safty and PrivacyRecurrent Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextBenchmarkChain-of-Thought

🎯 What it does: Proposed a query-oriented strategy for obscuring personally identifiable information (PII), and constructed the PII-Bench evaluation framework.

PILOT: Planning via Internalized Latent Optimization Trajectories for Large Language Models

Haoyu Zheng (Zhejiang University), Jun Xiao (Zhejiang University)

OptimizationComputational EfficiencyKnowledge DistillationRepresentation LearningAI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringMixture of ExpertsTextChain-of-Thought

🎯 What it does: Inject high-level planning capabilities into small LLMs through internalized latent guidance vectors, significantly improving their stability and accuracy in multi-step reasoning.

Placing Puzzle Pieces Where They Matter: A Question Augmentation Framework for Reinforcement Learning

Yangyi Fang (Tsinghua University), Haolin Shi (Tsinghua University)

TransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmark

🎯 What it does: This paper proposes the PieceHint framework, which utilizes value-driven prompt injection to address the challenges of sparse rewards in RL training.

Planning-Guided Tutoring with Assessment-Driven Memory for Pedagogical LLM Tutors

Zechen Li (Beijing Normal University), Hua Huang (Beijing Normal University)

AI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose ScaffoldLM, a planning-oriented, assessment-driven memory framework for multi-turn mathematical dialog-based teaching;

PLAWBENCH: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal Practice

Yuzhen Shi (Alibaba Group), HU Wei (Alibaba Group)

TransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Built a real-world legal workflow-based evaluation platform called PLAWBENCH, used to assess the performance of LLMs in three tasks: public legal consultation, case analysis, and legal document generation.

Please refuse to answer me! Mitigating Over-Refusal in Large Language Models via Adaptive Contrastive Decoding

Yupeng Qi (Sun Yat-sen University), Feng Xia (RMIT University)

Safty and PrivacyComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBenchmark

🎯 What it does: Proposed the AdaCD (Adaptive Contrastive Decoding) method, which extracts the distribution of rejection words through extreme safety prompts and dynamically switches decoding modes during inference to alleviate the over-rejection phenomenon in large language models (LLMs) while maintaining safety.

PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media

Zoher Kachwala (Indiana University), Filippo Menczer (Indiana University)

TransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposed and constructed the PLURULE benchmark for detecting whether comments violate specific community rules in a multilingual, multimodal community environment, simulating real moderators' decision-making through multiple-choice questions.

PodBench: A Comprehensive Benchmark for Instruction-Aware Audio-Oriented Podcast Script Generation

Chenning Xu (Tencent), Mingyang Song (Tencent)

GenerationData SynthesisTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-ThoughtAudio

🎯 What it does: Constructed PodBench benchmark, focusing on instruction-aware and context-driven long-form multi-speaker podcast script generation tasks, providing 800 long-context samples and multi-dimensional instructions;

PolitNuggets: Benchmarking Agentic Discovery of Long-Tail Political Facts

Yifei Zhu (University of Hong Kong)

TransformerLarge Language ModelAgentic AITextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Created the PolitNuggets benchmark to evaluate the ability of agent-based models to discover and merge long-tail political facts from the internet;

POLYCHARTQA: Benchmarking Large Vision-Language Models with Multilingual Chart Question Answering

Yichen Xu (Renmin University of China), Qin Jin (Renmin University of China)

Data SynthesisTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmark

🎯 What it does: Proposed POLYCHARTQA, the first large-scale chart question answering benchmark covering 10 languages, constructed with 22,606 charts and 26,151 QA pairs;

Polymorphic Universal Transformer

Yilong Chen (Chinese Academy of Sciences), Bryan Dai (IQuest Research)

Computational EfficiencyRepresentation LearningTransformerLarge Language ModelMixture of ExpertsContrastive LearningText

🎯 What it does: Propose a Polymorphic Universal Transformer (PUT), achieving functional polymorphism and depth sparsity within a shared parameter framework through conditional sparse subspaces, SiLU attention, and adaptive depth scheduling.

PosterForest: Hierarchical Multi-Agent Collaboration for Scientific Poster Generation

Jiho Choi (KAIST), Hyunjung Shim (KAIST)

GenerationData SynthesisOptimizationTransformerLarge Language ModelAgentic AIPrompt EngineeringDiffusion modelImageTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposes an untrained multi-agent framework called PosterForest, which utilizes a hierarchical Poster Tree to achieve collaborative optimization between content and layout, enabling the automatic generation of scientific posters.

Powerful Training-Free Membership Inference Against Fine-Tuned Autoregressive Language Models

David Ilić (JetBrains Research), Kostadin Cvejoski (JetBrains Research)

Safty and PrivacyAdversarial AttackTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringText

🎯 What it does: Proposed a membership inference attack called EZ-MIA, which determines whether text is part of the fine-tuning training set by utilizing probability shifts at error positions (tokens where the model makes prediction errors).

Powering Verifiable Learning via Automated Evolutionary Data Synthesis

He Du (Fudan University), Dacheng Tao (Nanyang Technological University)

OptimizationKnowledge DistillationData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringContrastive LearningTextSequentialBenchmark

🎯 What it does: Proposes EvoSyn, a general framework that automatically learns filtering strategies and generates verifiable data through evolutionary algorithms, aiming to enhance the learning effectiveness of LLMs in executable verifiable tasks.

POWSM: A Phonetic Open Whisper-Style Speech Foundation Model

Chin-Jou Li (Carnegie Mellon University), Shinji Watanabe (Carnegie Mellon University)

RecognitionTransformerMixture of ExpertsContrastive LearningAudio

🎯 What it does: Proposed POWSM, a unified speech foundation model capable of simultaneously performing telephone recognition (PR), automatic speech recognition (ASR), audio-guided grapheme-to-phoneme (G2P), and phoneme-to-grapheme (P2G) tasks.

PPA-Plan: Proactive Pitfall Avoidance for Reliable Planning in Long-Context LLM Reasoning

Byeongjin Kim (Hanyang University), Seo Yeon Park (Hanyang University)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposes a prospective error-avoidance planning framework called PPA-Plan, which first predicts logical pitfalls to generate negative constraints, then conditionally plans during the planning phase to avoid these constraints, ensures the executability of the plan through an error-correction module, and finally executes the generated reliable plan for long-context reasoning.

PR-XAI: PageRank-Based Feature Attribution for Transformers

Behrooz Azarkhalili (Simon Fraser University), Maxwell W. Libbrecht (Simon Fraser University)

ClassificationExplainability and InterpretabilityTransformerLarge Language ModelContrastive LearningText

🎯 What it does: Propose the PR-XAI method, which combines the PageRank algorithm with the transformer's attention mechanism to generate token-level feature attribution.

PRBench: Large-Scale Expert Rubrics for Evaluating High-Stakes Professional Reasoning

Afra Feyza Akyürek (Scale AI), Yunzhong He (Scale AI)

Large Language ModelPrompt EngineeringTextBenchmarkFinance RelatedRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This work constructs and publicly releases a large professional reasoning benchmark called PRBench, which includes 1,100 real-world task scenarios written by financial and legal experts, along with 18,711 finely crafted expert rubrics, corresponding multi-turn dialogues, and economic impact annotations.

Preconditioned Test-Time Adaptation for Out-of-Distribution Debiasing in Narrative Generation

Hanwen Shen (Stevens Institute of Technology), Shanshan Wang (University of Macau)

GenerationDomain AdaptationSafty and PrivacyTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: This paper proposes CAP-TTA, a context-aware preconditional LoRA fine-tuning framework triggered by thresholds during narrative generation, used to instantly correct bias and toxicity issues of large language models during testing.

Pref-CTRL: Preference Driven LLM Alignment using Representation Editing

Imranul Ashrafi (University of Technology Sydney), Massimo Piccardi (University of Technology Sydney)

OptimizationRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringText

🎯 What it does: Proposed a test-time alignment framework called Pref-CTRL based on representation editing, which trains a value function using a multi-objective loss, enabling large language models to achieve better alignment during inference by adjusting hidden states via gradient ascent.

Preference Heads in Large Language Models: A Mechanistic Framework for Interpretable Personalization

Weixu Zhang (McGill University), Haolun Wu (McGill University)

Explainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringContrastive LearningText

🎯 What it does: Propose a training-free differential preference-driven (DPS) method that identifies sparse 'preference heads' based on mechanism interpretation and controls them during decoding to achieve interpretable personalization.

Prefix Parsing is Just Parsing

Clemente Pasti (ETH Zürich), Tim Vieira (ETH Zürich)

Computational EfficiencyRepresentation LearningDiffusion modelScore-based ModelFlow-based ModelRectified FlowNeural Radiance FieldAuto EncoderGenerative Adversarial NetworkContrastive LearningText

🎯 What it does: This paper proposes a general method to convert the prefix parsing problem into standard parsing — prefix grammar transformation, and provides an efficient algorithm for computing prefix probabilities and the next-word weight vector;

PrefixNLI: Detecting Factual Inconsistencies as Soon as They Arise

Sapir Harary (Bar Ilan University), Ido Dagan (Bar Ilan University)

GenerationExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: Propose the PrefixNLI task to detect factual inconsistencies during the text generation process.

PRIME: A Process-Outcome Alignment Benchmark for Verifiable Reasoning in Mathematics and Engineering

Xiangfeng Wang (University of Science and Technology of China), Daxin Jiang (Stepfun)

Explainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed the PRIME benchmark for evaluating process-result consistency verification of large reasoning models in the fields of mathematics and engineering.

PRInTS: Reward Modeling for Long-Horizon Information Seeking

Jaewoo Lee (University of North Carolina at Chapel Hill), Mohit Bansal (University of North Carolina at Chapel Hill)

Autonomous DrivingOptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningAdversarial AttackHyperparameter SearchData-Centric LearningRobotic IntelligenceMeta LearningDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackNeural Architecture SearchProtein Structure PredictionTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AIMixture of ExpertsDiffusion modelScore-based ModelFlow-based ModelRectified FlowNeural Radiance FieldAuto EncoderGenerative Adversarial NetworkContrastive LearningGaussian SplattingSimultaneous Localization and MappingWorld ModelOptical FlowTextBenchmarkRetrieval-Augmented GenerationChain-of-ThoughtStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: This paper proposes PRInTS, a generative process reward model that combines information gain scoring with recursive trajectory summarization, enabling fine-grained evaluation of each step (reasoning + tool call) in long-term information-seeking tasks, and providing guidance for selection during testing for LLM agents.

PRiSM: Benchmarking Phone Realization in Speech Models

Shikhar Bharadwaj (Carnegie Mellon University), David R. Mortensen (Carnegie Mellon University)

RecognitionConvolutional Neural NetworkRecurrent Neural NetworkTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBenchmarkAudio

🎯 What it does: Proposes the PRiSM benchmark for evaluating systems that transcribe speech into phonemes, assessing them along two major dimensions: intrinsic (PFER) and extrinsic (downstream tasks).

PRISM: Probabilistic Reward Model with Inherent Structural Modeling

Yuhang Zhou (Fudan University), Guangnan Ye (Fudan University)

Reinforcement Learning from Human FeedbackTransformerReinforcement LearningMixture of ExpertsGaussian SplattingTextBenchmark

🎯 What it does: Construct a probabilistic reward model called PRISM, treating evaluation as a Gaussian mixture distribution, and learn expert heads and routers through two-stage training to capture subjective preferences and uncertainty.

PRISM: Probing Reasoning, Instruction, and Source Memory in LLM Hallucinations

Yuhe Wu (Hong Kong University of Science and Technology (Guangzhou)), Zhuang Liu (Dongbei University of Finance and Economics)

Explainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose the PRISM benchmark, constructing 9,448 samples, decomposing the hallucination issues of LLMs into four dimensions: knowledge missing, knowledge error, reasoning error, and instruction following error.

PRISMA: Preference-Reinforced Self-Training Approach for Interpretable Emotionally Intelligent Negotiation Dialogues

Prajwal Vijay Kajare (Indian Institute of Technology Jodhpur), Asif Ekbal (Indian Institute of Technology Patna)

Explainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This paper proposes PRISMA, an interpretable emotion-intelligent negotiation dialogue system, capable of generating emotionally appropriate and interpretable responses through emotion perception and strategy selection;

PRISP: Privacy-Safe Few-Shot Personalization via Lightweight Adaptation

Junho Park (Seoul National University), Taesup Moon (Seoul National University)

Federated LearningSafty and PrivacyComputational EfficiencyMeta LearningTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose a privacy-safe, lightweight few-shot personalized framework called PRISP, which can achieve user-level personalization for large language models without requiring task data or sharing user parameters.

Privacy Collapse: Benign Fine-Tuning Can Break Contextual Privacy in Language Models

Anmol Goel (Parameter Lab), Martin Gubri (Parameter Lab)

Safty and PrivacyExplainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringContrastive LearningTextBenchmark

🎯 What it does: Studied the phenomenon of 'privacy collapse' that occurs in language models after benevolent fine-tuning, leading to the loss of contextual privacy judgment capabilities while maintaining safety performance.

Privacy-preserving Prosody Representation Learning

Kevin Everson (University of Washington), Mari Ostendorf (University of Washington)

Safty and PrivacyRepresentation LearningTransformerAuto EncoderGenerative Adversarial NetworkContrastive LearningAudio

🎯 What it does: This study proposes a privacy-preserving prosody representation learning framework, designing a self-supervised encoder based on glottal source estimation, and decoupling speaker information through acoustic feature normalization and adversarial source identification;

Privacy-R1: Privacy-Aware Multi-LLM Agent Collaboration via Reinforcement Learning

Zheng Hui (University of Cambridge), Nigel Collier (University College London)

Safty and PrivacyTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringTextBiomedical Data

🎯 What it does: Proposes a privacy-aware multi-LLM agent collaboration framework called Privacy-R1, which dynamically routes text blocks to local or remote models while maintaining task performance and reducing the leakage of sensitive information.

ProactiveEval: A Unified Evaluation Framework for Proactive Dialogue Agents

Tianjian Liu (Sun Yat-sen University), Xiaojun Quan (Shenzhen Loop Area Institute)

Data SynthesisRecommendation SystemAutonomous DrivingOptimizationFederated LearningSafty and PrivacyExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningAdversarial AttackHyperparameter SearchData-Centric LearningRobotic IntelligenceMeta LearningDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackProtein Structure PredictionTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed the ProactiveEval unified evaluation framework, automatically generating 328 proactive dialogue evaluation environments across six domains, and using this framework to evaluate the performance of 22 LLMs in goal planning and dialogue guidance.

ProActor: Timing-Aware Reinforcement Learning for Proactive Task Scheduling Agents

Lei Ding (University of California Santa Cruz), Yang Liu (University of California Santa Cruz)

Autonomous DrivingOptimizationFederated LearningComputational EfficiencyKnowledge DistillationRepresentation LearningDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackNeural Architecture SearchProtein Structure PredictionTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AIPrompt EngineeringMixture of ExpertsVision Language ModelVision-Language-Action ModelDiffusion modelScore-based ModelFlow-based ModelRectified FlowNeural Radiance FieldAuto EncoderGenerative Adversarial NetworkContrastive LearningGaussian SplattingSimultaneous Localization and MappingWorld ModelOptical FlowTextSequentialFinance RelatedRetrieval-Augmented GenerationChain-of-ThoughtStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Propose the ProActor framework to achieve temporal-aware proactive behavior training for conversational task scheduling agents.

Probing Audio-Visual Reasoning in Multimodal Language Models through the Lens of Audio

Kaixiong Gong (Chinese University of Hong Kong), Xiangyu Yue (Chinese University of Hong Kong)

RecognitionExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringVision Language ModelVision-Language-Action ModelContrastive LearningImageVideoTextMultimodalityBenchmarkAudio

🎯 What it does: This paper systematically evaluates the audio processing and audio-visual integration capabilities of current multimodal large language models (MLLMs) by constructing two new evaluation tools: DeafTest (evaluating low-level audio perception) and AV-Odyssey (evaluating cross-modal audio-visual reasoning).

Probing for Reading Times

Eleftheria Tsipidi (ETH Zürich), Ryan Cotterell (ETH Zürich)

Explainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningText

🎯 What it does: By comparing the internal representations of multi-layer language models with eye movement fixation durations through regularized linear regression, the study explores their predictive power over the human reading process.

Probing Semantic Alignment, Lexical Invariance, and Syntactic Influence in LLM Metaphor Processing

Fengying Ye (University of Macau), Derek F. Wong (University of Macau)

Explainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringContrastive LearningText

🎯 What it does: This study designed three diagnostic experiments to systematically examine the semantic attribute alignment, lexical invariance, and syntactic sensitivity of LLMs in metaphor processing, analyzing the model's internal behavior from a geometric perspective and through controlled perturbations.

Probing the Safety Robustness of LLMs in Latent Space

Tianle Gu (Tsinghua University), Yingchun Wang (Shanghai Artificial Intelligence Laboratory)

Safty and PrivacyExplainability and InterpretabilityAdversarial AttackTransformerLarge Language ModelPrompt EngineeringContrastive LearningText

🎯 What it does: This paper proposes to evaluate the security robustness of large language models (LLMs) in the latent space by injecting precisely normalized perturbations into the hidden layers (Activation Steering Attack, ASA), and uses negative log-likelihood (NLL) as the diagnostic signal.

Process Reward Models Meet Planning: Generating Precise and Scalable Datasets for Step-Level Rewards

Raffaele Pisano (Babelscape), Roberto Navigli (Babelscape)

OptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextSequentialBenchmarkChain-of-Thought

🎯 What it does: Propose a process reward model (PRM) training dataset that is scalable and fine-grained, generated based on the planning language PDDL, to achieve step-level reward evaluation in the chain-of-thought of LLMs.

Profiling-Free Mixed-Precision Quantization for MoE LLMs via Fuzzy Rule Interpolation

Huachen Qi (Xiamen University), Qiang Shen (Aberystwyth University)

Computational EfficiencyKnowledge DistillationTransformerLarge Language ModelMixture of ExpertsContrastive LearningText

🎯 What it does: This paper proposes a tuning-free hybrid precision quantization framework based on fuzzy rule interpolation (FRI), specifically designed for large language models with sparse expert networks (MoE), which can predict quantization error and latency without conducting a full evaluation of each expert individually.

Programming over Thinking: Efficient and Robust Multi-Constraint Planning

Derrick Goh Xin Deik (Nanyang Technological University), Wenya Wang (Nanyang Technological University)

OptimizationAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringTextBenchmarkChain-of-Thought

🎯 What it does: Proposes the SCOPE framework, which decomposes multi-constraint planning into query-specific reasoning and general code execution, automatically generating reusable solver functions.

Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question Answering

Changin Choi (Seoul National University), Wonjong Rhee (Seoul National University)

RetrievalExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextMultimodalityRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed a progressive multi-modal retrieval and reasoning framework called PMSR for knowledge-intensive visual question answering; iteratively retrieves and constructs compressed reasoning records in heterogeneous knowledge bases through dual-range queries, forming structured reasoning trajectories, and finally provides answers by LLM.

ProgressLM: Towards Progress Reasoning in Vision-Language Models

Jianshu Zhang (Northwestern University), Han Liu (Northwestern University)

Robotic IntelligenceReinforcement Learning from Human FeedbackTransformerSupervised Fine-TuningReinforcement LearningPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose Progress-Bench, a progress reasoning benchmark, to explore the ability of Vision-Language Models (VLMs) to estimate progress under single observations, and achieve training-free and base model training through two-stage reasoning (first retrieving anchors and then performing mental simulation) with PROGRESSLM.

Projecting Out the Malice: A Global Subspace Approach to LLM Detoxification

Zenghao Duan (State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences), Xueqi Cheng (State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences)

Federated LearningExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringText

🎯 What it does: A lightweight method named GLOSS is proposed, which achieves model detoxification by identifying and eliminating the global toxic subspace of parameters in the Feed-Forward network of LLMs.

ProMed: Shapley Information Gain Guided Reinforcement Learning for Proactive Medical LLMs

Hongxin Ding (Peking University), Yasha Wang (Peking University)

Explainability and InterpretabilityDrug DiscoveryTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBiomedical DataBenchmark

🎯 What it does: Implementing an interactive diagnostic framework called ProMed in medical LLMs, transitioning from passive answering to active questioning.

ProMedical: Hierarchical Fine-Grained Criteria Modeling for Medical LLM Alignment via Explicit Injection

He Geng (Xunfei Healthcare Technology Co., Ltd.), Xiaodong Tao (Xunfei Healthcare Technology Co., Ltd.)

Safty and PrivacyExplainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextBiomedical DataElectronic Health RecordsBenchmarkRetrieval-Augmented Generation

🎯 What it does: Built a ProMedical framework based on fine-grained clinical Rubrics, which includes the Preference-50k dataset, a reward model, and an evaluation benchmark.

PROMPRINT: Prompt Fingerprinting via First-Token Response for LLM App Cloning Detection

Jungmin Lee, Yeonjoon Lee

RecognitionAnomaly DetectionExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringText

🎯 What it does: Propose a fingerprint identification method called PROMPRINT, used to detect copied LLM applications without exposing system prompts;

Prosody as Supervision: Bridging the Non-Verbal–Verbal for Multilingual Speech Emotion Recognition

Girish (UPES), Muskaan Singh (Ulster University)

RecognitionConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningOptical FlowAudio

🎯 What it does: Proposes an unsupervised cross-lingual emotion recognition transfer framework, called NOVA-ARC, that transfers from annotated non-linguistic sounds (such as laughter, crying, sighing) to linguistically vocalized speech.

Protecting Bystander Privacy via Selective Hearing in Audio LLMs

Xiao Zhan (VRAIN, Universitat Politècnica de València), Phil Woodland (INGENIO)

Safty and PrivacyTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringBenchmarkAudio

🎯 What it does: Proposed the SH-Bench benchmark to evaluate the selective listening ability of audio LLMs in multi-speaker environments, protecting bystander privacy; simultaneously designed a unified evaluation metric called Selective Efficacy (SE) and a Bystander Privacy Fine-tuning (BPFT) scheme.

Protecting Language Models Against Unauthorized Distillation through Trace Rewriting

Xinhang Ma (Washington University in St. Louis), Yevgeniy Vorobeychik (Washington University in St. Louis)

Safty and PrivacyKnowledge DistillationTransformerLarge Language ModelPrompt EngineeringText

🎯 What it does: The study rewrites the reasoning traces generated by large language models (LLMs) to prevent unauthorized knowledge distillation while embedding verifiable watermarks;

Protecting multimodal large language models against misleading visualizations

Jonathan Tonglet (TU Darmstadt), Iryna Gurevych (KU Leuven)

Safty and PrivacyExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextMultimodalityTabularRetrieval-Augmented Generation

🎯 What it does: The study evaluates and mitigates the vulnerability of multimodal large language models to misleading visualizations.

Protein-STORY: Semantic Text-Oriented Representation Yields biologically meaningful Protein embeddings

Nabil Ibtehaz (Purdue University), Daisuke Kihara (Purdue University)

RetrievalRepresentation LearningDrug DiscoveryTransformerLarge Language ModelContrastive LearningTextBiomedical Data

🎯 What it does: Propose Protein-STORY, an unsupervised representation learning framework that compresses multi-source textual descriptions of proteins into unified vectors;

Provably Safe Offline-to-Online RL: Decoupling Learning from Data-Driven Safety Enforcement

Kaitong Cai (Sun Yat Sen University), Keze Wang (Sun Yat Sen University)

Reinforcement LearningTabularTime SeriesSequentialBenchmark

🎯 What it does: Propose the RLPD-GX framework, decoupling reward-driven policy learning from safe execution, using projection guardians to ensure safe execution and achieve safe value backup;

Prune as You Generate: Online Rollout Pruning for Faster and Better RLVR

Haobo Xu (University Of Illinois At Urbana Champaign), Hanghang Tong (University Of Illinois At Urbana Champaign)

Computational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsTextBenchmark

🎯 What it does: Propose an online episode trimming method called ARROL, which can predict the success probability of partial episodes during the generation process based on a lightweight quality head and trim them in advance, maintaining a near 0.5 ratio of positive and negative samples within the episode group;

Pruning Unsafe Tickets: A Resource-Efficient Framework for Safer and More Robust LLMs

Wai Man Si (CISPA Helmholtz Center for Information Security), Yang Zhang (CISPA Helmholtz Center for Information Security)

Safty and PrivacyComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelContrastive LearningTextMultimodality

🎯 What it does: Propose a resource-efficient, gradient-agnostic pruning framework that directly identifies and removes parameter subnetworks (unsafe tickets) causing unsafe behaviors in large models, while maintaining model performance.

Psyche-R1: Towards Reliable Psychological LLMs through Unified Empathy, Expertise, and Reasoning

Chongyuan Dai (Hefei University of Technology), Meng Wang (Hefei University of Technology)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Constructed the Chinese psychology LLM Psyche-R1, which integrates empathy, professional knowledge, and reasoning capabilities in a unified manner.

Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness

Amin Banayeeanzade (University of Southern California), Sai Praneeth Karimireddy (University of Southern California)

Explainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: Created and evaluated the PsySET benchmark, systematically comparing the effectiveness and reliability of multiple LLM emotion regulation methods (prompt engineering, vector injection, parameter-efficient fine-tuning, DPO) on emotional and personality dimensions.

Pub-LawBench: Public-Oriented Benchmarking for LegalAI

Qiaoyu Zheng (Nankai University), Qian Liu (University of Auckland)

TransformerLarge Language ModelPrompt EngineeringTextBenchmarkChain-of-Thought

🎯 What it does: Proposed Pub-LawBench, a public-oriented legal AI benchmark that includes two task types: instant question answering and legal text generation, and designed a three-dimensional evaluation metric consisting of content relevance, legal normativity, and format usability;

PubMed Reasoner: Dynamic Reasoning-based Retrieval for Evidence-Grounded Biomedical Question Answering

Yiqing Zhang (PayPal), Fabricio Murai (Worcester Polytechnic Institute)

RetrievalExplainability and InterpretabilityComputational EfficiencyDrug DiscoveryTransformerLarge Language ModelPrompt EngineeringTextBiomedical DataRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Constructed PubMed Reasoner, a three-stage biomedical question-answering agent that includes self-critical query optimization, batch reflective retrieval with early stopping, and evidence-based answer generation.

Punctuation-Steered Representation Fine-Tuning

Zheng Gong (Hong Kong University of Science and Technology), Zhefeng Wang (Huawei Cloud)

Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningTextBenchmark

🎯 What it does: Proposed a punctuation-based representation fine-tuning method called PuReFT for efficiently adapting large language models.

PUPPET: Neural-Symbolic Standardized Patients for Mental Health

Chen Xu (Beijing Institute of Technology), Bin Hu (Beijing Institute of Technology)

Explainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelTextSequentialChain-of-Thought

🎯 What it does: Proposed and implemented PUPPET, a neuro-symbolic virtual standardized patient and its training framework PUPPET-TRAINER, which simulates changes in patients' psychological states using the OBSERVE-THINK-BEHAVE architecture, and provides immediate feedback to mental health trainers through interpretable causal chains.

Putting Captions to the Test: Evaluating Video Caption Quality through Multiple-Choice Question Answering

Zizhen Wang (Apple), Xiaoming Simon Wang (Apple)

Explainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningVideoTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposed the CapQuiz benchmark for reference-free evaluation, assessing video caption information fidelity through human-validated fine-grained multiple-choice question answering.

Putting HUMANS first: Efficient LAM Evaluation with Human Preference Alignment

Woody Haosheng Gan (University Of Southern California), Diyi Yang (Stanford University)

Recommendation SystemComputational EfficiencyData-Centric LearningContrastive LearningTextBenchmarkAudio

🎯 What it does: Systematically study subset selection for evaluating large audio models (LAM), constructing a minimal evaluation set aligned with human preferences called HUMANS, and publicly releasing the data and regression models.

QA-MoE: Towards a Continuous Reliability Spectrum with Quality-Aware Mixture of Experts for Robust Multimodal Sentiment Analysis

Yitong Zhu (Hong Kong University of Science and Technology (Guangzhou)), Yuyang Wang (Hong Kong University of Science and Technology (Guangzhou))

ClassificationRecognitionTransformerMixture of ExpertsDiffusion modelScore-based ModelAuto EncoderContrastive LearningGaussian SplattingImageTextMultimodalityStochastic Differential EquationAudio

🎯 What it does: Propose the QA-MoE framework and continuous reliability spectrum to uniformly model dynamic noise and missing data in multimodal sentiment analysis and achieve robust inference.

QBridge: Bridging Natural Language and SQL via Gold Query Rewriting with Agentic Refinement

Zhensheng Luo (Zhejiang University), Xiu Tang (Zhejiang University)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextTabularBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose the QBridge framework, which achieves efficient zero-shot NL2SQL inference by transcribing natural language questions into Gold Query (structured intermediate representation), followed by feedback-driven query and SQL correction.

QiMeng-PRepair: Precise Code Repair via Edit-Aware Reward Optimization

Changxin Ke (State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences), Yunji Chen (State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences)

AI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringDiffusion modelScore-based ModelContrastive LearningTextBenchmark

🎯 What it does: Proposes the PRepair framework, addressing the over-editing problem in LLM-based code repair, achieving more accurate and minimal code fixes.

Quantifying Aleatoric Uncertainty of In-Context Learning for Robust Measure of LLM Prediction Confidence

Jinseok Chung (Pohang University of Science and Technology), Namhoon Lee (Pohang University of Science and Technology)

Explainability and InterpretabilityMeta LearningTransformerLarge Language ModelPrompt EngineeringText

🎯 What it does: Studied how to decompose predictive uncertainty into data uncertainty (aleatoric) and model uncertainty (epistemic) in an In-Context Learning environment, and proposed using self-function vectors to directly estimate aleatoric uncertainty from internal representations of the model.

Quantifying and Improving the Robustness of Retrieval-Augmented Language Models Against Spurious Features in Grounding Data

Shiping Yang (Simon Fraser University), Dongmei Zhang (Microsoft)

RetrievalExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose the SURE evaluation framework, systematically constructing robustness tests for RAG models against spurious features, generating synthetic datasets such as SURE_Wiki and SIG_Wiki, conducting experiments on multiple LLMs, and further enhancing robustness through two training strategies: SFT and DPO.

Quantifying and Mitigating Socially Desirable Responding in LLMs: A Desirability-Matched Graded Forced-Choice Psychometric Study

Kensuke Okada (University of Tokyo), Kyosuke Bunji (Kobe University)

Explainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelPrompt EngineeringText

🎯 What it does: This paper proposes a psychometric framework to quantify and mitigate social desirability response (SDR) bias in large language models (LLMs) during questionnaire assessments.

Quantifying and Understanding Uncertainty in Large Reasoning Models

Yangyi Li (Iowa State University), Mengdi Huai (Iowa State University)

Explainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelContrastive LearningTextMultimodalityChain-of-Thought

🎯 What it does: Propose the CoRAP framework, which combines Conformal Prediction with the inference-answer structure to provide a unified uncertainty quantification for the generation process of large reasoning models (LRMs); meanwhile, design a hierarchical example-to-step explanation method based on Shapley values, which can efficiently locate training samples and inference steps that are critical to coverage guarantees.

Quantifying Metric and Model Agreement in Bias Evaluation of Large Language Models

Arash Asgari (York University), Laleh Seyyed-Kalantari (York University)

Explainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelContrastive LearningTextBenchmark

🎯 What it does: Propose Metric Agreement Score (MeAS) and Model Agreement Score (MoAS) to quantify the consistency of different bias evaluation metrics and different large language models in bias evaluation results;

Quantifying the Impact of Translation Errors on Multilingual LLM Evaluation

Klaudia Thellmann, Jens Lehmann (Amazon)

Explainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: This paper systematically evaluates the impact of translation errors in machine translation datasets on the evaluation of multilingual LLMs, and constructs professional human references and Span-ACESRef as gold standards. It compares the performance of LLM self-assessment with xCOMET-XXL in span-level error detection; further, it quantifies the relationship between translation errors and model accuracy through logistic regression, revealing that translation errors lead to a 6–11 percentage point drop in accuracy.

QuantileMark: A Message-Symmetric Multi-bit Watermark for LLMs

Junlin Zhu (Peking University), Xiaojun Wan (Peking University)

GenerationSafty and PrivacyData-Centric LearningTransformerLarge Language ModelPrompt EngineeringScore-based ModelText

🎯 What it does: Propose QuantileMark, a multi-bit watermarking scheme that uniformly partitions continuous probability intervals and samples within corresponding partitions during the LLM generation process;

QuantumQA: Enhancing Scientific Reasoning via Physics-Consistent Dataset and Verification-Aware Reinforcement Learning

Songxin Qu (University of Science and Technology of China), Zhao-Yun Chen (Hefei Comprehensive National Science Center)

Reinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextPhysics RelatedRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed a large-scale, physics-consistency-verified quantum mechanics reasoning dataset called QUANTUMQA, and developed the RLVR method based on verifiable rewards to enhance the reliability of models in scientific reasoning.

QuDAR: Query-Wise Dual-Perspective Adaptive Retrieval

Joeun Kim (Korea Advanced Institute of Science and Technology), Jae-Gil Lee (Korea Advanced Institute of Science and Technology)

RetrievalTransformerLarge Language ModelContrastive LearningTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose a Query-wise Dual-Perspective Adaptive Retrieval (QUDAR) framework that dynamically fuses four retrieval signals: sparse retrieval, dense retrieval, original query, and expanded query;

Query-Aware Knowledge Retrieval via Hyperbolic Structuring

Chuang Zhou (Hong Kong Polytechnic University), Xiao Huang (Hong Kong Polytechnic University)

RetrievalComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextGraphRetrieval-Augmented Generation

🎯 What it does: Proposes HyperRAG, a query-based knowledge retrieval framework that can dynamically construct a semantic and reasoning structured graph embedded in a Poincaré ball (hyper-sphere) for each query, and perform retrieval-augmented generation in this space.

Query-Efficient Agentic Graph Extraction Attacks on GraphRAG Systems

Shuhua Yang (Pennsylvania State University), Suhang Wang (Pennsylvania State University)

Safty and PrivacyAdversarial AttackGraph Neural NetworkTransformerLarge Language ModelAgentic AIPrompt EngineeringTextGraphAgriculture RelatedRetrieval-Augmented Generation

🎯 What it does: This paper investigates privacy leakage in the GraphRAG system under black-box query budget constraints, and proposes an efficient query-based attack framework called AGEA.

Question Difficulty Estimation for Large Language Models via Answer Plausibility Scoring

Jamshid Mozafari (University of Innsbruck), Adam Jatowt (University of Innsbruck)

Explainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: Propose a question-answering difficulty estimation method called Q-DAPS based on the entropy of the suspiciousness of large model answers. It utilizes the suspiciousness distribution of candidate answers and removes popularity bias to obtain interpretable difficulty scores;

Question Tells You Where the Answer Is: Intention-aware Long-Context KV Cache Compression

Liang Zhao (Harbin Institute of Technology), Bing Qin (Harbin Institute of Technology)

CompressionComputational EfficiencyTransformerLarge Language ModelText

🎯 What it does: Propose an intent-based KV cache compression method called IntentKV, which identifies and retains critical KV pairs for subsequent generation by leveraging the attention distribution differences of intent tokens.

R^3: End-to-End Reasoning-based Planning for Multi-step Retrosynthesis via Reinforcement Learning

YiFei Wang, Hao Zhou (Institute for AI Industry Research (AIR), Tsinghua University)

Knowledge DistillationDrug DiscoveryTransformerLarge Language ModelReinforcement LearningTextBenchmarkChain-of-Thought

🎯 What it does: Propose the R3 framework, which transforms multi-step retrosynthesis planning from traditional search to end-to-end generative reasoning, directly generating complete synthetic routes.

R^3AG: Retriever Routing for Retrieval-Augmented Generation

Tong Zhao (Renmin University of China), Zhicheng Dou (Renmin University of China)

GenerationRetrievalRepresentation LearningTransformerLarge Language ModelContrastive LearningTextRetrieval-Augmented Generation

🎯 What it does: Built a retrieval router framework called R3AG, which can dynamically select the most suitable retriever and decide whether to perform retrieval based on the query, thereby improving the overall effectiveness of retrieval augmented generation (RAG).