arXivSub Start free trial

ACL 2026 Papers — Page 11

Annual Meeting of the Association for Computational Linguistics · 2296 papers

InsAT: Instance-aware Semantic Alignment and Transfer from Human–Object Keypoints for Zero-to-Few-shot Action Understanding

Kazuki Tsutsukawa (Konicaminolta)

ClassificationRecognitionDomain AdaptationTransformerLarge Language ModelVision Language ModelContrastive LearningImageVideoText

🎯 What it does: This study proposes the InsAT framework, which aligns human and object key points with instance-level language descriptions to achieve zero/ few-shot action recognition and adaptation;

Inside Out: Evolving User-Centric Core Memory Trees for Long-Term Personalized Dialogue Systems

Jihao Zhao (Renmin University of China), Zhiyu li

Recommendation SystemData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningTextRetrieval-Augmented Generation

🎯 What it does: Proposes the Inside Out framework, which manages core memories in long-term personalized dialogues by constructing an evolvable PersonaTree (a hierarchical structure based on the bio-psycho-social model), and uses a lightweight MemListener (reinforcement learning based on process rewards) to convert dialogues into tree operations. Finally, it adopts an adaptive generation strategy (fast mode and agent recall) to improve response quality.

InsideOut: Measuring and Mitigating Insider–Outsider Bias in Interview Script Generation

Yixin Wan (University of California, Los Angeles), Kai-Wei Chang (University of California, Los Angeles)

GenerationExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose the INSIDEOUT benchmark and three dimensional metrics (CEP, CPD, CAG) to quantify the insider-outsider bias of LLMs across different cultural contexts, and effectively alleviate this bias through an agent-based MFA framework (SA, HA, Plan);

InsLogicBench: An Argumentation Logic Grounded Benchmark for Complex Insurance Claims Adjudication

Jin Liu (Fudan University), Yanghua Xiao (Fudan University)

Explainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringTextBenchmarkFinance RelatedRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose the InsLogicBench benchmark, which provides complete reasoning chains for insurance claim adjudication, and employs the Nested Toulmin model for controllable synthesis.

Instant Personalized Large Language Model Adaptation via Hypernetwork

Zhaoxuan Tan (University Of Notre Dame), Meng Jiang (Amazon Com Inc)

Domain AdaptationRecommendation SystemFederated LearningSafty and PrivacyComputational EfficiencyTransformerLarge Language ModelSupervised Fine-TuningTextRetrieval-Augmented Generation

🎯 What it does: Propose the Profile-to-PEFT framework, which generates LoRA parameters based on user profiles using a hypernetwork, enabling instant LLM personalization

InstructDiff: Domain-Adaptive Data Selection via Contrastive Entropy for Efficient LLM Fine-Tuning

Junyou Su (Peking University), Guanhua Chen (Southern University of Science and Technology)

Domain AdaptationComputational EfficiencyData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningText

🎯 What it does: Proposes a unified data selection framework called InstructDiff based on contrastive entropy, for efficient LLM supervised fine-tuning;

Instruction Data Selection via Answer Divergence

Bo Li (Peking University), Wei Ye (Hebei University of Technology)

OptimizationRepresentation LearningData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: Proposed a method called ADG for instruction data selection based on the geometric structure of multi-sample answers, aiming to improve the effectiveness of instruction tuning under a fixed data budget.

Instructions are all you need: Self-supervised Reinforcement Learning for Instruction Following

Qingyu Ren (Fudan University), Yanghua Xiao (Fudan University)

OptimizationRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextBenchmarkChain-of-Thought

🎯 What it does: Propose a self-supervised reinforcement learning framework that does not require external supervision, specifically aimed at improving the performance of large language models in multi-constraint instruction following tasks.

Integrating Data Validation with Large Language Models for Regulation-Guided Tabular Anomaly Detection

Haoliang Huang (Hunan University), Changjian Chen (Hunan University)

Anomaly DetectionData-Centric LearningTransformerLarge Language ModelAgentic AIPrompt EngineeringTabularFinance RelatedRetrieval-Augmented Generation

🎯 What it does: Propose the regulation-based table anomaly detection task RTAD, and design a training-agnostic method RegValidator, which transforms regulations into detection ideas and converts them into SQL queries, further improving detection accuracy through rule validation.

Interleaved Latent Visual Reasoning with Selective Perceptual Modeling

Shuai Dong (China University of Geosciences), Zhongyu Wei (Shanghai Innovation Institute)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelVision Language ModelContrastive LearningImageTextMultimodalityChain-of-Thought

🎯 What it does: Propose an Interleaved Latent Visual Reasoning (ILVR) framework, where the model alternately updates between text generation and latent space visual representations, enabling dynamic visual state tracking for multi-step reasoning.

Interleaved Tool-Call Reasoning for Protein Function Understanding

Chuanliu Fan (Soochow University), Guohong Fu (Soochow University)

Explainability and InterpretabilityDrug DiscoveryProtein Structure PredictionTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AITextBiomedical DataBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed a tool-enhanced protein function reasoning agent, PFUA, which generates verifiable intermediate evidence by alternately invoking domain-specific tools (such as MMseqs2, Pfam, TMbed, etc.) during the reasoning process, thereby enabling answers to multi-dimensional questions about protein function, catalytic activity, domains, etc.

Interpretable Coreference Resolution Evaluation Using Explicit Semantics

Bruno Gatti (Sapienza University of Rome), Roberto Navigli (Sapienza University of Rome)

RecognitionExplainability and InterpretabilityTransformerPrompt EngineeringContrastive LearningText

🎯 What it does: This paper proposes a semantic-enhanced coreference evaluation framework based on Concept and Named Entity Recognition (CNER), which employs a two-step annotation and propagation technique to assign fine-grained semantic labels to coreference clusters, and performs diagnosis using category-specific Mention F1 and Link F1 metrics;

Interpretable Safety Alignment via SAE-Constructed Low-Rank Subspace Adaptation

Dianyun Wang, Zhaofeng He (Beijing University of Posts and Telecommunications)

Safty and PrivacyExplainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelAuto EncoderContrastive LearningText

🎯 What it does: This paper proposes an interpretable safe alignment method called SAILS based on sparse autoencoders (SAE), which constructs a low-rank safe subspace using the decoding direction of SAE for initializing the LoRA adapter, thereby achieving efficient and interpretable safe alignment.

Interpretable Traces, Unexpected Outcomes: Investigating the Disconnect in Trace-Based Knowledge Distillation

Siddhant Bhambri (Arizona State University), Subbarao Kambhampati (Arizona State University)

Explainability and InterpretabilityKnowledge DistillationTransformerSupervised Fine-TuningTextChain-of-Thought

🎯 What it does: This paper generates verifiable intermediate reasoning trajectories by decomposing problems through regularization, uses these trajectories to perform knowledge distillation on small models, and evaluates the correlation between the semantic correctness and interpretability of the trajectories and the accuracy of the final answers in QA tasks.

Into the Gray Zone: Domain Contexts Can Blur LLM Safety Boundaries

Ki Sen Hung (Hong Kong University of Science and Technology), Yangqiu Song (Hong Kong University of Science and Technology)

Safty and PrivacyAdversarial AttackTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Studied how domain-specific contexts can lead large language models (LLMs) to exhibit gray areas on safety boundaries, proposing and evaluating the JARGON framework: leveraging security research contexts and multi-turn dialogues to generate 'academic-style' attacks, significantly improving the success rate of breaking through safety defenses.

IntrAgent: An LLM Agent for Content-Grounded Information Retrieval through Literature Review

Fengbo Ma (University of Georgia), Zhen Xiang (University of Georgia)

RetrievalDrug DiscoveryTransformerLarge Language ModelAgentic AIPrompt EngineeringTextReview/Survey PaperBenchmarkPhysics RelatedRetrieval-Augmented Generation

🎯 What it does: Proposes a fine-grained information retrieval task called IntraView for scientific literature, designs an LLM agent named IntrAgent that mimics human reading, and constructs a cross-domain benchmark called IntraBench.

Inverting the Shield: Systematically Generating Safety Tests from Policy Specifications

Xiaoyue Lu (Shenzhen Campus of Sun Yat-sen University), Jin Song Dong (National University of Singapore)

Safty and PrivacyExplainability and InterpretabilityAdversarial AttackTransformerLarge Language ModelPrompt EngineeringGenerative Adversarial NetworkTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose the POLARIS framework, which utilizes the formalization of security policies to achieve safety testing of LLMs, generating traceable and comprehensive attack queries;

Investigating Counterfactual Unfairness in LLMs towards Identities through Humor

Shubin Kim (Yonsei University), Youngjae Yu (Seoul National University)

GenerationExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: By swapping the speaker and listener identities, the study investigates unfair behaviors of large language models in three tasks: humor generation, intent inference, and social impact prediction.

Investigating Cross-Modal Skill Injection: Scenarios, Methods, and Hyperparameters

Zhiyu Xu (Peking University), Xu Sun (Peking University)

Representation LearningHyperparameter SearchData-Centric LearningTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsVision Language ModelImageTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: Systematically evaluate and implement cross-modal skill injection, analyzing three-dimensional factors of scenarios, methods, and hyperparameters;

Investigating More Explainable and Partition-Free Compositionality Estimation for LLMs: A Rule-Generation Perspective

Ziyao Xu (Peking University), Houfeng Wang (Peking University)

Explainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringText

🎯 What it does: This paper proposes a new perspective based on rule generation for evaluating the compositional ability of large language models (LLMs), and explicitly reveals the model's understanding of compositional properties through generated programs.

Investigating the Representation of Backchannels and Fillers in Fine-tuned Language Models

Yu Wang (Bielefeld University), Hendrik Buschmeier (Bielefeld University)

Representation LearningTransformerLarge Language ModelSupervised Fine-TuningText

🎯 What it does: Studied the impact of three fine-tuning strategies (MASK, NTP, TTP) on the representation of backchannels and fillers in Transformer language models during conversations, and evaluated model performance through clustering and NLG generation.

IRIS: Interleaved Reinforcement with Incremental Staged Curriculum for Cross-Lingual Mathematical Reasoning

Navya Gupta (Singapore Institute of Technology), Rajiv Ratn Shah (IIIT Delhi)

Representation LearningData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextMultimodalityReview/Survey PaperBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed the IRIS framework, combining vertical (increasing problem difficulty) and horizontal (gradually reducing given reasoning steps) dual-axis curriculum learning, along with composite rewards and GRPO reinforcement learning to achieve cross-lingual mathematical reasoning.

Is a Document Educational or Just Wikipedia-Style? — Pitfalls of Classifier-Based Quality Filtering

Mateusz Klimaszewski (Warsaw University of Technology), Piotr Andruszkiewicz (Warsaw University of Technology)

ClassificationFederated LearningSafty and PrivacyExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: By rewriting web content in a Wikipedia-style and using the Classifier-based Quality Filtering (CQF) model to score the text before and after rewriting, the study reveals the security vulnerabilities and biases of CQF when filtering pre-trained corpora.

Is Chain-of-Thought Really Not Explainability? Chain-of-Thought Can Be Faithful without Hint Verbalization

Kerem Zaman (UNC Chapel Hill), Shashank Srivastava (UNC Chapel Hill)

Explainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringTextChain-of-Thought

🎯 What it does: This paper conducts a multi-dimensional interpretability evaluation of chain-of-thought (CoT) in large language models, proposing a new 'faithful@k' metric. It combines Logit Lens with causal mediation analysis to explore the influence mechanism of CoT on predictions, even when the prompt words are not explicitly mentioned. Additionally, it verifies the generalizability of the results on larger-scale reasoning models.

Is EEG-to-Text Feasible in Real-World Scenarios? An In-Depth Analysis Using a Neuropsychology-Inspired Benchmark

Zihan Zhang (Harbin Institute of Technology), Kai Xiong (Nanjing University)

RecognitionRepresentation LearningTransformerSupervised Fine-TuningContrastive LearningTextTime SeriesBiomedical DataBenchmark

🎯 What it does: Constructed the COFETT dataset and proposed a teacher-forcing-free evaluation framework to verify the feasibility of EEG-to-text.

Is Human-Like Text Liked by Humans? Multilingual Human Detection and Preference Against AI

Yuxia Wang (MBZUAI), Preslav Nakov (MBZUAI)

ClassificationRecognitionExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: This paper conducts large-scale human detection experiments across languages and domains, evaluating experts' accuracy in identifying texts generated by the latest large language models versus human texts, and exploring how prompting strategies narrow the gap and influence human preferences.

Is the Attention Matrix Really the Key to Self-Attention in Multivariate Long-Term Time Series Forecasting?

Xinyu Li (Fudan University), Xin Wang (Fudan University)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerContrastive LearningTime SeriesBenchmark

🎯 What it does: Through extensive ablation experiments and the construction of a minimalist multi-branch MLP, the true role of the attention matrix in self-attention mechanisms for multivariate long-range time series forecasting is explored.

Is this chart lying to me? Automating the detection of misleading visualizations

Jonathan Tonglet (TU Darmstadt), Iryna Gurevych (KU Leuven)

Anomaly DetectionData-Centric LearningTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageMultimodalityTabularBenchmarkRetrieval-Augmented Generation

🎯 What it does: This paper constructs two large datasets, Misviz and Misviz-synth, to detect 12 types of misleading features in charts and systematically evaluates existing models.

IS-CoT: Breaking the Long-form Generation Collapse via Interleaved Structural Thinking

Zechen Sun (Soochow University), Min Zhang (Soochow University)

GenerationData SynthesisKnowledge DistillationTransformerLarge Language ModelSupervised Fine-TuningTextBenchmarkFinance RelatedRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This paper studies the 'length collapse' phenomenon in large language models during long-text generation, proposes a dynamic recursive Interleaved Structural Chain-of-Thought (IS-CoT) framework, and constructs a 5k high-quality interleaved reasoning dataset based on this. Subsequently, supervised fine-tuning is performed on Qwen3-8B to obtain IS-Writer-8B, which demonstrates excellent performance in long-text generation.

It’s High Time: A Survey of Temporal Question Answering

Bhawna Piryani (University of Innsbruck), Adam Jatowt (University of Innsbruck)

TransformerLarge Language ModelTextTime SeriesReview/Survey PaperRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: A systematic review of the Temporal Question Answering (TQA) field, constructing a unified framework based on three-dimensional time-aware aspects: corpus, question, and model, and classifying and providing a comprehensive evaluation of datasets, tasks, and model methods.

It’s Not What You Say, It’s How You Say It: Evaluating LLM Responses to Expressions of Belief

Kevin Du (ETH Zürich), Alex Warstadt (UC San Diego)

TransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This paper proposes the EoBench benchmark, which systematically evaluates the impact of different expressions of belief (EoB) on large language models (LLMs) in terms of their ability to accept contextual information and prior knowledge.

iTAG: Inverse Design for Natural Text Generation with Accurate Causal Graph Annotations

Wenshuo Wang (South China University of Technology), Wei Li (Chinese Academy of Sciences)

GenerationData SynthesisExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringTextGraphTabularBenchmarkFinance RelatedChain-of-Thought

🎯 What it does: Proposes the iTAG framework, which assigns real-world concepts to large language models (LLMs) using inverse design and chain-of-thought (CoT) reasoning, thereby generating text with natural language and high-accuracy causal graph annotations.

Iterative Dual-Model Alignment for Story Evaluation

Bruce Qin (Purdue University), Dan Goldwasser (Purdue University)

Recommendation SystemExplainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningTextRetrieval-Augmented Generation

🎯 What it does: Propose an iterative dual-model learning framework, training the story evaluator α and the explanation generator β in a closed loop to improve story preference prediction and explanation quality.

IterCOMP: Reasoning-aware Adaptive Prompt Compression for Multi-hop Question Answering

JungMin Yun, YoungBin Kim

CompressionComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose IterCOMP, an iterative prompting compression framework for multi-hop question answering, which can dynamically evaluate answerability during compression and generate follow-up questions to iteratively collect key information.

IUQ: Interrogative Uncertainty Quantification for Long-Form Large Language Model Generation

Haozhi Fan (University of Pennsylvania), Kaidi Xu (City University of Hong Kong)

GenerationExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed a framework called IUQ based on question-answering probing for uncertainty quantification in long-text generation, used to evaluate the reliability of LLMs in long-text generation.

J4R: Learning to Judge with Equivalent Initial State Group Relative Policy Optimization

Austin Xu (Salesforce AI Research), Shafiq Joty (Salesforce AI Research)

OptimizationTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningTextBenchmark

🎯 What it does: This paper proposes a method for training a judgment model using reinforcement learning, specifically addressing the judgment problem in reasoning tasks, and constructs a lightweight judgment model named J4R.

Jailbreak-Zero: A Path to Pareto Optimal Red Teaming for Large Language Models

Kai Hu (Meta Superintelligence Labs), Akash Bharadwaj (Meta Superintelligence Labs)

Safty and PrivacyAdversarial AttackReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringMixture of ExpertsDiffusion modelScore-based ModelText

🎯 What it does: Proposed the automatic red teaming (Jailbreak-Zero) framework, which supports evaluation from examples to strategies, and implements two red teaming modes: human-free strategy generation, zero-shot, and fine-tuning.

Jailbreaking Multimodal Large Language Models using Multi-Clip Video

Choongwon Kang (Sungkyunkwan University), Jang Hyun Kim (Sungkyunkwan University)

Safty and PrivacyAdversarial AttackTransformerLarge Language ModelPrompt EngineeringVision Language ModelDiffusion modelVideoTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: This paper proposes the Multi-Clip Video SafetyBench, a benchmark for systematically evaluating the jailbreak security of multi-modal large language models (MLLMs) under video inputs, and through experiments reveals that video diversity and dynamic content significantly enhance jailbreak success rates.

Jakiro: Boosting Speculative Decoding via Decoupled MoE

Haiduo Huang (Xi'an Jiaotong University), Pengju Ren (Xi'an Jiaotong University)

Computational EfficiencyAI Code AssistantTransformerLarge Language ModelMixture of ExpertsContrastive LearningText

🎯 What it does: By introducing a separated Mixture of Experts (MoE) architecture in the draft model and combining it with Contrastive Enhanced Parallel Decoding (CEPD), candidate word generation is decoupled within the same tree layer, thereby improving the speed and diversity of large language model inference.

JanusMM: A Benchmark for Self-Deprecation Understanding in Real-World Multimodal Conversations

Xinyi Xu (Sun Yat-sen University), Shihan Dou (Fudan University)

ClassificationRecognitionTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBenchmark

🎯 What it does: Proposed and constructed the JanusMM benchmark to evaluate the ability of multimodal large language models to recognize and reason about self-deprecating expressions in real-world scenarios;

JARVIS or Ultron? A Survey on the Safety and Security Threats of Computer-Using Agents

Ada Chen (Carnegie Mellon University), Shuai Wang (Hong Kong University of Science and Technology)

Safty and PrivacyImageVideoTextReview/Survey PaperBenchmark

🎯 What it does: A systematic review of the security and threats of computer usage agents (CUA), defining the CUA framework, constructing classifications of intrinsic and extrinsic threats, corresponding defense systems, and summarizing evaluation benchmarks and metrics.

Joint Knowledge Base Completion and Question Answering by Combining Large Language Models and Small Language Models

Yinan Liu (Northeastern University), Bin Wang (Northeastern University)

Recommendation SystemExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningAgentic AITextGraphBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed the JCQL framework, which jointly completes knowledge graph completion (KBC) and knowledge-based question answering (KBQA), improving the performance of both tasks through iterative interaction between LLM and SLM.

JoPR: Joint Emotion Perception and Reasoning for Conversational Emotion Recognition

Yumeng Fu (Harbin Institute of Technology), Bingquan Liu (Harbin Institute of Technology)

RecognitionTransformerSupervised Fine-TuningReinforcement LearningTextChain-of-Thought

🎯 What it does: Proposed the JoPR framework, which integrates multi-dimensional curriculum learning, long-chain thinking instruction fine-tuning, human preference alignment, and emotion-level reward grading to improve the accuracy of dialogue emotion recognition.

JPU: Bridging Jailbreak Defense and Unlearning via On-Policy Path Rectification

Xi Wang (National University of Defense Technology), Jie Yu (National University of Defense Technology)

Safty and PrivacyReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringText

🎯 What it does: Propose a new defense framework JPU based on machine unlearning, using dynamic path correction technology to enhance the jailbreak resistance of large language models.

JurisBench: A Deep Benchmark for Assessing Large Language Models in Professional Legal Practice

Ziang Chen (BIGAI), Bin Ling (Peking University)

TransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Constructed JurIBench, a legal evaluation benchmark focused on the entire process of civil litigation in China, with a vertical depth of analysis.

Just Pass Twice: Efficient Token Classification with LLMs for Zero-Shot NER

Ahmed Ewais (WitnessAI), Amr Ali (WitnessAI)

RecognitionTransformerLarge Language ModelPrompt EngineeringText

🎯 What it does: By duplicating the input sequence twice, causal LLMs can access the complete sentence during the second pass, enabling bidirectional token-level zero-shot named entity recognition.

JW-SVD: Bridging the Cross-Modal Mismatch in Post-Training MLLM Compression

Runchao Li (Case Western Reserve University), Kenneth A. Loparo (Case Western Reserve University)

CompressionComputational EfficiencyRepresentation LearningTransformerLarge Language ModelVision Language ModelContrastive LearningImageTextMultimodality

🎯 What it does: This work focuses on post-training compression of multi-modal large language models (MLLMs), proposing the Joint Whitening Singular Value Decomposition (JW-SVD) method to avoid the deletion of visual features caused by text-only SVD, thereby maintaining the collaborative capabilities between visual and linguistic modalities;

K-Merge: Online Continual Merging of Adapters for On-device Large Language Models

Donald Shenaj (Samsung Research and Development Institute), Umberto Michieli (Samsung Research and Development Institute)

Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextMultimodality

🎯 What it does: The paper proposes an online incremental merging method for low-rank adapters (LoRA) on mobile devices, aiming to gradually incorporate adapters from new tasks under limited storage budgets while maintaining the performance of existing tasks.

KARL: Reinforcement Learning for LLM Agents on Multi-Turn Knowledge-Intensive Agentic Tasks

Xueqiao Sun (Tsinghua University), Jie Tang

OptimizationKnowledge DistillationRepresentation LearningTransformerLarge Language ModelReinforcement LearningAgentic AITextTabularBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Designed and implemented the KARL framework, enabling LLM agents to actively explore structured knowledge and perform tool calls during multi-round interactions to accomplish knowledge-intensive tasks.

KASER: Knowledge-Aligned Student Error Simulator for Open-Ended Coding Tasks

Zhangqi Duan (University of Massachusetts Amherst), Andrew Lan (University of Massachusetts Amherst)

AI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextSequential

🎯 What it does: KASER proposes a student error simulation method based on reinforcement learning, generating error code that matches the student's knowledge level by aligning student code with knowledge components.

KCVR: Knowledge-Centric Video Reconstruction for Structured Pedagogical Summarization via Dynamic Graph Planning

Jingjiang Liu (Zhejiang Normal University), Jiajie Xu (Zhejiang Normal University)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningData-Centric LearningGraph Neural NetworkTransformerLarge Language ModelVision Language ModelDiffusion modelScore-based ModelContrastive LearningVideoTextGraphRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Design and implement a structured teaching reconstruction framework KCVR for educational videos, which generates a single structured summary that conforms to prior logic and is supported by blackboard visual evidence.

KG-MuLQA: A Framework for KG-based Multi-Level QA Extraction and Long-Context LLM Evaluation

Nikita Tatarinov (Georgia Institute of Technology), Sudheer Chava (Georgia Institute of Technology)

Graph Neural NetworkTransformerLarge Language ModelPrompt EngineeringTextGraphBenchmarkFinance RelatedRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Construct the KG-MULQA framework to systematically generate multi-level QA pairs from credit agreements using a knowledge graph;

KG-ViP: Bridging Knowledge Grounding and Visual Perception in Multi-modal LLMs for Visual Question Answering

Zhiyang Li (University of Science and Technology of China), Xike Xie (University of Science and Technology of China)

RecognitionImage TranslationRestorationRetrievalRecommendation SystemExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningAdversarial AttackHyperparameter SearchData-Centric LearningReinforcement Learning from Human FeedbackNeural Architecture SearchGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityGraphRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose the KG-ViP framework, which integrates scene graphs with common-sense knowledge graphs to jointly enhance the knowledge grounding and fine-grained visual perception of multimodal large models in visual question answering.

KinyaProp: Fine-Grained Propaganda Annotation in Kinyarwanda

Manzi Fabrice Niyigaba (Dartmouth College), Soroush Vosoughi (Dartmouth College)

ClassificationData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringText

🎯 What it does: Created the KinyaProp dataset and conducted experiments on fine-grained propaganda detection in the low-resource language Kinyarwanda.

Know the Known and the Unknown: Reasonable Answer Generation with Knowledge-Informed Citations

Yichi Zhang (Zhejiang University), Huajun Chen (Zhejiang University)

GenerationExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextBenchmarkChain-of-Thought

🎯 What it does: Designed the KFC benchmark and an automated data construction pipeline to generate QA data containing references to both known and unknown entities, and proposed the SELF-KFC self-correcting training strategy to enhance the traceability and reliability of LLMs in long-text QA tasks.

Know Your Place: Diagnosing Implicit Social Adaptation Failures in Chinese Large Language Models

Yu Tian (Inner Mongolia University), Xiangdong Su (Inner Mongolia University)

Explainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Diagnose and quantify the failure of large Chinese language models in adapting to implicit social cues during natural interactions, introduce the concept of Social Agnosia, and construct the C-ISA evaluation framework.

Knowing When Not to Answer: Lightweight KB-Aligned OOD Detection for Safe RAG

Ilias Triantafyllopoulos, João Sedoc (New York University)

Anomaly DetectionExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningTextTabularSequentialBenchmarkRetrieval-Augmented Generation

🎯 What it does: This paper proposes a lightweight, knowledge-base-based out-of-distribution (OOD) detection method. By performing PCA on KB document embeddings, queries are projected into a low-dimensional subspace, and discrimination is conducted using geometric rules or lightweight classifiers, thereby achieving a 'when not to answer' gating mechanism in retrieval-augmented generation (RAG) systems.

Knowledge Beyond Language: Bridging the Gap in Multilingual Machine Unlearning Evaluation

Kyomin Hwang (Seoul National University), Nojun Kwak (Seoul National University)

Federated LearningSafty and PrivacyExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposed two cross-lingual evaluation metrics, KSS (Knowledge Separability) and KPS (Knowledge Persistence), and evaluated various machine forgetting methods on a synthetic question-answer dataset spanning 10 languages.

Knowledge Injection Exists in MoE? Exploring Expert-Aware Contrast Decoding in MoE for Mitigating LLMs’ Hallucinations

Xinyue Fang (National University of Defense Technology), Dongsheng Li (National University of Defense Technology)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsContrastive LearningText

🎯 What it does: Studied the knowledge injection phenomenon in Mixture-of-Experts (MoE) models and found differences across architectures; proposed Expert-Aware Adaptive Contrastive Decoding (EAACD) based on expert activation differences to reduce hallucinations in large language models (LLMs).

Knowledge is Not Enough: Injecting RL Skills for Continual Adaptation

Pingzhi Tang (Peking University), Muhan Zhang (Peking University)

Domain AdaptationComputational EfficiencyKnowledge DistillationTransformerSupervised Fine-TuningReinforcement LearningPrompt EngineeringMixture of ExpertsTextTabular

🎯 What it does: Propose Parametric Skill Transfer (PaST), which extracts 'skill vectors' learned by RL from the source domain, and linearly injects them into the target domain after only light-weight SFT, thereby achieving decoupling and efficient transfer of knowledge updating and reasoning skills.

Knowledge Poisoning Attacks on Medical Multi-Modal Retrieval-Augmented Generation

Peiru Yang (Tsinghua University), Tao Qi (Beijing University of Posts and Telecommunications)

RetrievalAdversarial AttackData-Centric LearningTransformerLarge Language ModelPrompt EngineeringImageTextMultimodalityBiomedical DataElectronic Health RecordsRetrieval-Augmented Generation

🎯 What it does: Proposes a knowledge poisoning attack framework called M3Att for medical multimodal retrieval-augmented generation systems, assuming that the attacker only knows the database distribution information and does not need to query specific knowledge.

Knowledge Vector of Logical Reasoning in Large Language Models

Zixuan Wang (University of Florida), Yuanyuan Lei (University of Florida)

Explainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningAuto EncoderText

🎯 What it does: Investigate the linear knowledge vectors of three types of logical reasoning (deductive, inductive, abductive) in the activation space of large language models, and propose a fine-tuning framework with complementary subspace constraints to enhance the complementarity and specificity of these vectors, thereby improving reasoning performance.

Knowledge without Wisdom: Measuring Misalignment between LLMs and Intended Impact

Michael Hardy (Stanford University), Yunsung Kim (Stanford University)

Explainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelPrompt EngineeringVideoTextTabular

🎯 What it does: Measuring the alignment of large language models in elementary math classroom assessments, comparing their relationship with expert evaluations and student learning gains (VAM).

Knowledge-driven Augmentation and Retrieval for Integrative Temporal Adaptation

Weisi Liu (University of Memphis), Xiaolei Huang (University of Memphis)

ClassificationDomain AdaptationTransformerLarge Language ModelSupervised Fine-TuningTextBiomedical DataElectronic Health RecordsReview/Survey PaperRetrieval-Augmented Generation

🎯 What it does: To address text classification tasks with temporal drift, the KARITA framework is proposed, combining multi-dimensional drift detection, source data retrieval, and knowledge-driven augmentation to achieve adaptive temporal evolution.

Knowledge-to-Verification: Exploring RLVR for LLMs in Knowledge-Intensive Domains

Zhonghang Yuan (Shanghai Artificial Intelligence Laboratory), Nanqing Dong (Shanghai Artificial Intelligence Laboratory)

Reinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextGraphAgriculture RelatedRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: In knowledge-intensive domains (such as agriculture, law, and medicine), leveraging RLVR to enhance the reasoning capabilities of large language models, proposing the K2V framework to achieve automatic synthesis of verifiable data and verification of the reasoning process.

KnowMe-Bench: Benchmarking Person Understanding for Lifelong Digital Companions

Tingyu Wu (University of Chinese Academy of Sciences), Ronghao Chen (QuantaAlpha)

TransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose KnowMe-Bench, a benchmark for human understanding constructed based on long-form autobiographical texts;

KnowRL: Exploring Knowledgeable Reinforcement Learning for Factuality

Baochang Ren (Zhejiang University), Huajun Chen (Zhejiang University)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelReinforcement LearningTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose the KnowRL framework, which combines reinforcement learning (RL) with knowledge verification to provide process-level supervision during chain-of-thought reasoning, significantly reducing hallucinations.

KoCo-Bench: Can Large Language Models Leverage Domain Knowledge in Software Development?

Xue Jiang (Peking University), Yihong Dong (Peking University)

AI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AITextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposes the KOCO-Bench benchmark, combining domain-specific knowledge corpora with corresponding multi-granularity evaluation tasks to assess the ability of LLMs to acquire and apply new domain knowledge.

KoCo: Conditioning Language Model Pre-training on Knowledge Coordinates

Yudong Li (Shenzhen University), Linlin Shen (Shenzhen University)

Representation LearningData-Centric LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningText

🎯 What it does: This paper proposes Knowledge Coordinate Conditioning (KoCo), which maps documents to a three-dimensional (Source, Content, Stability) knowledge coordinate and incorporates this coordinate as a prefix into pretraining, endowing the model with explicit contextual awareness.

KV-Embedding: Training-free Text Embedding via Internal KV Re-routing in Decoder-only LLMs

Yixuan Tang (Hong Kong University of Science and Technology), Yi Yang (Hong Kong University of Science and Technology)

RetrievalComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBenchmark

🎯 What it does: Propose an untrained text embedding method called KV-Embedding, which reroutes the KV states of the final token inside a decoder-only LLM as a prefix, allowing all tokens to obtain global context in a single forward pass.

L2Dir: Integrating L\_2-Norm and Directional Alignment for Unsupervised Contrastive Representation Learning in Multimodal Retrieval

Tianyu Zong (University of Chinese Academy of Sciences), Jungang Xu (University of Chinese Academy of Sciences)

RetrievalRepresentation LearningTransformerVision Language ModelContrastive LearningMultimodality

🎯 What it does: Proposed the L2Dir framework, which jointly optimizes L2 norm alignment and directional consistency to enhance contrastive learning for multi-modal retrieval.

Label and Explanation Variation in LLM-Based Annotation: a Case Study in Natural Language Inference

Artur Kulmizev (Uclouvain), Marie-Catherine de Marneffe (Uclouvain)

Explainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringText

🎯 What it does: This paper investigates the diversity of labels and explanations generated by large language models (LLMs) in natural language inference (NLI) tasks, evaluating their similarity to the label distributions and explanation styles produced by human annotators.

Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge

Xin Sun (National Institute of Informatics), Saku Sugawara (National Institute of Informatics)

Federated LearningExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextBiomedical DataReview/Survey PaperRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Through a controlled experiment comparing source labels (Human vs AI) in health-related Q&A content, this study investigates the dependence of humans and large language models (LLM-as-a-Judge) on labels during trust evaluation, and explores the mechanisms behind label effects by combining eye-tracking with LLM internal attention/entropy analysis.

LaCo: Layer-wise Compensation for Pruned Large Language Models

Yingen Liu (Hunan University), Kenli Li (Hunan University)

CompressionComputational EfficiencyRepresentation LearningTransformerLarge Language ModelText

🎯 What it does: Propose a layer-wise compensation framework LaCo, which restores hidden representations layer by layer after pruning using mask-constrained regression;

LAD-RAG: Layout-aware Dynamic RAG for Visually-Rich Document Understanding

Zhivar Sourati (University of Southern California), Dan Roth (University of Southern California)

RetrievalGraph Neural NetworkTransformerAgentic AIPrompt EngineeringVision Language ModelImageTextMultimodalityGraphBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposes the LAD-RAG framework, which combines layout-aware symbolic document graphs with neural indexes to achieve dynamic, query-driven evidence retrieval in retrieval-augmented generation.

LADR: Locality-Aware Dynamic Rescue for Efficient Text-to-Image Generation with Diffusion Large Language Models

Chenglin Wang (East China Normal University), Kai Zhang (Nanjing University)

GenerationComputational EfficiencyTransformerPrompt EngineeringDiffusion modelContrastive LearningImageTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: This paper proposes a training-agnostic acceleration method called LADR, which dynamically rescues generated frontier image pixels to improve the inference speed of discrete diffusion language models.

LAFaCT: Attribution-based Localization and Focused Sequential Analysis of Fact-Critical Tokens for Hallucination Detection

Xin Wang (University of Science and Technology of China), Zhendong Mao (University of Science and Technology of China)

Anomaly DetectionExplainability and InterpretabilityRecurrent Neural NetworkTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBenchmark

🎯 What it does: Developed the LAFaCT framework, which detects hallucinations in LLM-generated text by first locating key factual words and then performing sequence analysis.

LAMCL: A Length-aware Momentum Contrastive Learning Framework for Multiscale Machine-Revised Text Detection

Bing Zhou (Chongqing University), Zhou Yongcheng

ClassificationAnomaly DetectionTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextRetrieval-Augmented Generation

🎯 What it does: This paper proposes the LAMCL framework for detecting machine-revised text. It first enhances the text using an LLM, then distinguishes between human-generated and machine-generated text through consistency measurement, and further improves the discriminative ability of multi-scale text by combining length-aware momentum contrastive learning.

LaMI: Augmenting Large Language Models via Late Multi-Image Fusion

Guy Yariv (Hebrew University of Jerusalem), Sagie Benaim (Hebrew University of Jerusalem)

ClassificationRecognitionGenerationData SynthesisExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelVision Language ModelDiffusion modelContrastive LearningImageTextMultimodalityRetrieval-Augmented Generation

🎯 What it does: Propose LAMI, which enhances visual common sense capabilities on text LLMs by utilizing multi-image generation and late fusion methods, while keeping text reasoning performance unaffected.

LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint Guidance

Yuchun Fan (Northeastern University), Tong Xiao (Northeastern University)

Representation LearningData-Centric LearningTransformerReinforcement LearningPrompt EngineeringTextMultimodality

🎯 What it does: Aiming at multilingual reasoning tasks, the LANG framework is proposed, which enhances multilingual reasoning accuracy and linguistic consistency through language-adaptive prompting guided reinforcement learning.

LangSAE Editing: Improving Multilingual Information Retrieval via Post-hoc Language Identity Removal

Dongjun Kim (Korea University), Heuiseok Lim (Korea University)

RetrievalRepresentation LearningTransformerAuto EncoderContrastive LearningTextMultimodality

🎯 What it does: Designed and validated the post-hoc sparse autoencoder LANGSAE editing, which removes language identity signals from multilingual retrieval embeddings to improve cross-lingual retrieval performance.

Language Acquisition Device in Large Language Models

Masato Mita (University of Tokyo), Yohei Oseki (University of Tokyo)

Representation LearningData-Centric LearningTransformerLarge Language ModelPrompt EngineeringText

🎯 What it does: Proposed the MP-STRUCT pre-pretraining framework based on the LAD (Language Acquisition Device) concept, using synthetic sequences to induce large language models to acquire structured language biases before natural language pretraining.

Language Model as Planner and Formalizer under Constraints

Cassie Huang (Drexel University), Li Zhang

Explainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Constructed the CoPE (Constrained Planning Environments) benchmark, adding fine-grained, natural language constraints to classical planning domains (BlocksWorld, CoinCollector), and evaluated the performance of LLMs in planning and formalization tasks;

Language Models Don’t Know What You Want: Evaluating Personalization in Deep Research Needs Real Users

Nishant Balepur (University of Maryland), Aakanksha Naik (Allen Institute for Artificial Intelligence)

Recommendation SystemExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelTextReview/Survey PaperBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposes MYSCHOLARQA, a deep research assistant that builds editable personal profiles based on user research papers, generates personalized action lists, and writes multi-chapter reports, evaluated through offline synthetic data with LLM judgment and online interviews.

Language Models Learn Universal Representations of Numbers and Here’s Why You Should Care

Michal Štefánik (R&D Centre for Large Language Models, National Institute of Informatics), Pontus Stenetorp (R&D Centre for Large Language Models, National Institute of Informatics)

Explainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextTabularBenchmark

🎯 What it does: This study investigates how large language models (LLMs) encode numbers, quantifies the prevalence of sinusoidal structures in different models, layers, and natural language contexts, and develops a new numerical probe based on sinusoidal parameterization (param-sin), further exploring the representation of multi-tag numbers, output tracking, and ordinal data.

Language Models Struggle to Use Representations Learned In-Context

Michael A. Lepori (Google DeepMind), Katja Filippova

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningMeta LearningTransformerLarge Language ModelDiffusion modelScore-based ModelRectified FlowWorld ModelTextSequentialChain-of-Thought

🎯 What it does: This paper systematically evaluates whether the representations learned through context can be flexibly deployed in downstream tasks, combining two experiments: next-token prediction and adaptive world modeling.

Language of Thought Shapes Output Diversity in Large Language Models

Shaoyang Xu (Singapore University of Technology and Design), Wenxuan Zhang (Singapore University of Technology and Design)

GenerationExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringTextBenchmarkChain-of-Thought

🎯 What it does: Improve the diversity of the generated output by controlling the thinking language (i.e., the language used in intermediate reasoning) during the inference process of large language models.

Language on Demand, Knowledge at Core: Composing LLMs with Encoder-Decoder Translation Models for Extensible Multilinguality

Mengyu Bu (Key Laboratory of Intelligent Information Processing, Institute of Computing Technology, Chinese Academy of Sciences), Yang Feng (Key Laboratory of Intelligent Information Processing, Institute of Computing Technology, Chinese Academy of Sciences)

Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringContrastive LearningTextMultimodality

🎯 What it does: Built a framework called XBridge, which combines a pre-trained multilingual encoder-decoder NMT model with a large language model (LLM), utilizing the LLM as the core knowledge processing unit for English, while the external NMT handles multilingual understanding and generation;

Language Reconstruction with Brain Predictive Coding from fMRI Data

Congchi Yin (Nanjing University of Aeronautics and Astronautics), Piji Li (Fudan University)

GenerationRepresentation LearningConvolutional Neural NetworkTransformerLarge Language ModelAuto EncoderContrastive LearningBiomedical DataMagnetic Resonance Imaging

🎯 What it does: This paper constructs the PREDFT model, which decodes continuous natural language from brain signals by encoding fMRI signals and combining them with a lateral network of brain predictive coding.

Language Statistics and False Belief Reasoning: Evidence from 41 Open-Weight LMs

Sean Trott (Rutgers University), Pamela D. Rivière (Rutgers University)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextChain-of-Thought

🎯 What it does: This study replicates and extends previous work on false belief reasoning, evaluating the sensitivity of 41 open-weight language models (from five model families) to knowledge states and knowledge cues in 192 scenarios based on Trott et al.'s 2023 design of false belief tasks. Model behaviors are compared with human experimental data; simultaneously, the study explores the impact of model parameters, training scale, and whether instruction fine-tuning is applied on task performance and the model's predictive power of human behavior. Furthermore, the study proposes and validates a new hypothesis about the generation of false belief bias by non-truth-conditional verbs (e.g., 'think') in false belief reasoning using model results.

Language-Aware Token Boosting: LLM Language Confusion Reduction Without Tuning

Trapoom Ukarapol (SCB DataX), Nut Chukamphaeng (SCB DataX)

GenerationData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextMultimodality

🎯 What it does: Propose a language alignment method without fine-tuning called LATB and Adaptive-LATB, which enhances the consistency of multilingual generation by perturbing the logits of target language tokens.

LaoBench: A Large-Scale Multidimensional Lao Benchmark for Large Language Models

Jian Gao (China-ASEAN Information Harbor Co., Ltd.), Yonghua Lin (Beijing Academy of Artificial Intelligence)

Large Language ModelPrompt EngineeringMixture of ExpertsTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Constructed LaoBench, a multidimensional benchmark for Lao with over 17,000 samples, covering cultural knowledge, K12 education, and trilingual translation.

Large Language Model-Enhanced Multi-Armed Bandits

Jiahang Sun (Chinese University of Hong Kong Shenzhen), Zhongxiang Dai (Chinese University of Hong Kong)

OptimizationTransformerLarge Language ModelReinforcement LearningPrompt EngineeringText

🎯 What it does: Embedding large language models (LLM) as reward predictors into classical multi-armed bandit (MAB) algorithms (Thompson Sampling and regression oracle) to improve the balance between exploration and exploitation.

Large Language Models Are Bad Dice Players: LLMs Struggle to Generate Random Numbers from Statistical Distributions

Minda Zhao (Harvard University), Mengyu Wang (Harvard University)

GenerationData SynthesisTransformerLarge Language ModelPrompt EngineeringText

🎯 What it does: This paper evaluates the performance of 11 state-of-the-art LLMs in generating random numbers that conform to 15 statistical distributions through large-scale experiments, and extends the results to practical tasks such as multiple-choice question generation and text-to-image prompt generation.

LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM Safety

Junxiao Yang (Tsinghua University), Minlie Huang (Tsinghua University)

Safty and PrivacyReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningTextMultimodality

🎯 What it does: In multilingual safe alignment, a method called LASA is proposed, which performs safe alignment at the semantic bottleneck layer of LLMs.

Late Code Chunking: A Code Chunking Strategy for Repository-Level Code Completion

Seungmin Oh (Sungkyunkwan University), Eunseok Lee (Sungkyunkwan University)

AI Code AssistantTransformerLarge Language ModelPrompt EngineeringTextSequentialBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose the Late Code Chunking (LC 2) dual context strategy to enhance semantic understanding and generation quality for repository-level code completion.

Latent Agents: A Post-Training Procedure for Internalized Multi-Agent Debate

John Seon Keun Yi (Boston University), Dokyun Lee (Boston University)

Explainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AIContrastive LearningText

🎯 What it does: Proposes Internalized Multi-Agent Debate (IMAD), which internalizes the multi-agent debate process into a single large language model through two-stage fine-tuning, achieving efficient reasoning;

Latent Attention Denoising: A Training-Free Energy-Based Framework for Mitigating Hallucinations in Vision-Language Models

Zhiwen Luo (Huazhong University of Science and Technology), Kun He (Huazhong University of Science and Technology)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerVision Language ModelScore-based ModelImageTextMultimodalityBenchmarkStochastic Differential Equation

🎯 What it does: Proposes Latent Attention Denoising (LAD), a training-free, energy-based single-step denoising framework that calibrates the attention of large vision-language models during inference, thereby reducing hallucinations.

Latent-Condensed Transformer for Efficient Long Context Modeling

Zeng You (South China University of Technology), Mingkui Tan (South China University of Technology)

Computational EfficiencyRepresentation LearningTransformerLarge Language ModelContrastive LearningTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposes Latent-Condensed Attention (LCA), performing semantic aggregation and position information anchor selection in the low-dimensional latent space of MLA to achieve efficient inference for long context modeling.

Layer-Wise High-Impact Parameter Ratio Optimization in Post-Training Quantization for Large Language Models

Cuong Pham (Monash University), Thanh-Toan Do (Monash University)

OptimizationFederated LearningComputational EfficiencyKnowledge DistillationHyperparameter SearchTransformerLarge Language ModelText

🎯 What it does: Proposes a post-training quantization method with hierarchical high-impact parameter ratio optimization for large-scale language models