arXivSub Start free trial

ACL 2026 Papers — Page 4

Annual Meeting of the Association for Computational Linguistics · 2296 papers

Can Persona-Prompted LLMs Emulate Subgroup Values? An Empirical Analysis of Generalisability and Fairness in Cultural Alignment

Bryan Chen Zhengyu Tan (Singapore University of Technology and Design), Roy Ka-Wei Lee (Singapore University of Technology and Design)

Federated LearningExplainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextTabularReview/Survey Paper

🎯 What it does: This paper constructs a structured dataset containing over 20,000 (question, subgroup) pairs to study the value orientation simulation of large language models among diverse subpopulations in Singapore, and tests the model's generalization ability on unknown subgroups and open-ended generation tasks.

Can Reasoning Path still be Effective as Input? Bridging Post-Reasoning to Chain-of-Thought Compression

Chengzhengxu Li (Xi'an Jiaotong University), Chao Shen (Xi'an Jiaotong University)

Computational EfficiencyKnowledge DistillationTransformerLarge Language ModelPrompt EngineeringTextChain-of-Thought

🎯 What it does: Proposed post-reasoning and the UCoT framework, which reduces the length of the reasoning output by adding compressed CoT to the input of the LLM.

Can Spectral-Clipping Enable Better Learning While Forgetting Less for Low-Rank Adaptation?

Hyowon Wi (Korea Advanced Institute of Science and Technology), Noseong Park (Korea Advanced Institute of Science and Technology)

Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningText

🎯 What it does: Perform singular value decomposition (SVD) on the weights of pre-trained models, revealing that principal singular components are easily transferable, while secondary singular components require extensive task adaptation; propose SCLoRA based on LoRA, introducing spectral clipping in the low-rank adapter that is based on the pre-trained singular value distribution, controlling singular value growth to improve downstream task performance while preserving pre-trained knowledge; also provide theoretical proof that singular value growth leads to catastrophic forgetting.

Can We Predict Before Executing Machine Learning Agents?

Jingsheng Zheng (Zhejiang University), Ningyu Zhang (Zhejiang University)

Autonomous DrivingOptimizationHyperparameter SearchData-Centric LearningRobotic IntelligenceAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringWorld ModelTextTabularBenchmarkChain-of-Thought

🎯 What it does: Construct the 'Data Center Solution Preference' task and create 18,438 comparative samples to verify the ability of large language models to predict the superiority and inferiority of two ML solutions without physical execution. Based on this, design the FOREAGENT Predict-then-Verify cycle, significantly accelerating the search process of ML agents.

Can You Make It Sound Like You? Post-Editing LLM-Generated Text for Personal Style

Connor Baumler (University of Maryland), Hal Daumé Iii

GenerationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: This study investigates whether users can reshape their personal writing style through post-editing drafts generated by LLMs in writing tasks that require personal style, and evaluates the effectiveness of this approach.

CAP: Controllable Alignment Prompting for Unlearning in LLMs

Zhaokun Wang (University of Electronic Science and Technology of China), Wenhong Tian (University of Electronic Science and Technology of China)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringText

🎯 What it does: Proposed a full-process controllable forgetting framework based on prompts, named CAP, which achieves selective knowledge forgetting without modifying model parameters;

Capability Decomposition for Unified Information Extraction via Hierarchical Mixture-of-Experts

Jing Zhou (Southeast University), Yao He (Southeast University)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningMixture of ExpertsText

🎯 What it does: Propose a unified information extraction framework based on large language models, UC-UIE, which describes all IE tasks through a unified framework-slot (schema) and explicitly decomposes the extraction reasoning into three general capabilities: judgment, localization, and association. A hierarchical MoE adapter is used to achieve parameter-efficient fine-tuning.

CAPC-CG: A Large-Scale, Expert-Directed LLM-Annotated Corpus of Adaptive Policy Communication in China

Bolun Sun (Johns Hopkins University), Pingxu Hao (Johns Hopkins University)

ClassificationTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextBenchmark

🎯 What it does: Constructed and made public the CAPC-CG corpus, which includes paragraph-level texts of policy directives from the Chinese central government from 1949 to 2023, and proposed an expert-oriented LLM annotation method.

CAPruner: Conceptual-Adjacent Scene Graph Pruner for Enhancing 3D Spatial Reasoning of Large Language Models

Shengli Zhou (Southern University of Science and Technology), Feng Zheng (Southern University of Science and Technology)

Computational EfficiencyRepresentation LearningGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningTextPoint CloudGraph

🎯 What it does: A lightweight concept adjacency scene graph pruning model called CAPruner is proposed to retain the most important spatial relationships in 3D vision-language tasks under limited budget, thereby improving the efficiency and accuracy of large language models in 3D spatial reasoning.

CAR-bench: Evaluating the Consistency and Limit-Awareness of LLM Agents under Real-World Uncertainty

Johannes Kirmayr (BMWGroup Research and Technology), Elisabeth Andre

Autonomous DrivingExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelAgentic AIPrompt EngineeringTextTabularTime SeriesBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Construct and evaluate CAR-bench, a multi-turn interaction benchmark for in-vehicle assistants, used to test the consistency, self-awareness of uncertainty and capabilities of large language models in dynamic environments.

CARES: Context-Aware Resolution Selector for VLMs

Moshe Kimhi (Technion), Eli Schwartz (IBM Research)

Computational EfficiencyRepresentation LearningTransformerSupervised Fine-TuningPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodality

🎯 What it does: Proposed a context-aware resolution selector called CARES, which can dynamically determine the minimal sufficient resolution at the frontend of VLM based on the image and query, thereby significantly reducing the number of visual tokens and computational load.

CaRL-EM: Cost-Aware Reinforcement Learning for Entity Matching with LLMs

Chaohui Guo (Vrije Universiteit Amsterdam), Zhisheng Huang (Vrije Universiteit Amsterdam)

Data-Centric LearningTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmark

🎯 What it does: Propose CaRL-EM, a reinforcement learning-based controller that can dynamically select LLM operations (MATCH/COMPARE/SELECT/DECIDE) in multi-candidate entity matching tasks, and supports seamless switching between different LLM backends.

CaseFacts: A Benchmark for Legal Fact-Checking and Precedent Retrieval

Akshith Reddy Putta (University of Texas at Arlington), Chengkai Li (University of Texas at Arlington)

RetrievalTransformerLarge Language ModelSupervised Fine-TuningTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Constructed a legal fact-checking benchmark called CaseFacts, targeting U.S. Supreme Court precedents, covering three types of spoken legal claims: supporting, refuting, and being overturned.

CASPER in the Machine: Insights into Character Variety in LLM-Generated Stories

Anneliese Brei (UNC Chapel Hill), Snigdha Chaturvedi (UNC Chapel Hill)

ClassificationGenerationTransformerLarge Language ModelPrompt EngineeringTextReview/Survey PaperRetrieval-Augmented Generation

🎯 What it does: Proposed and implemented the CASPER framework to automatically classify and compare narrative dimensions of characters in short stories generated by LLMs and written by humans;

CASS: Nvidia to AMD Transpilation with Data, Models, and Benchmark

Ahmed Heakl (Mohamed bin Zayed University of Artificial Intelligence), Abdulrahman Mahmoud (Mohamed bin Zayed University of Artificial Intelligence)

Data SynthesisOptimizationAI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningTextBenchmark

🎯 What it does: Propose the CASS dataset and model for source code and assembly-level translation between Nvidia (CUDA/SASS) and AMD (HIP/RDNA3)

Cat-MoD: Accelerating Multimodal Alignment via Caption Token Guided Asymmetric Mixture-of-Depths

YiJie Huang, Daling Wang (Northeastern University)

Computational EfficiencyRepresentation LearningTransformerMixture of ExpertsVision Language ModelContrastive LearningImageTextMultimodalityBenchmark

🎯 What it does: This paper proposes the Cat-MoD framework, which achieves efficient multi-modal alignment by integrating language-prior Guide Tokens with randomly generated Explorer Tokens in the Q-Former.

Causal-ESC: Reliable Policy Learning for Emotional Support Conversation via Causal Inference

Xv Wang (South China University Of Technology), Rui Zhang (Huya Inc)

Explainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: Learn emotional support strategies from offline dialogue logs using doubly robust learning, and combine them with LLM-style rewriting to generate empathetic responses.

Causal2Vec: Improving Decoder-only LLMs as Embedding Models through a Contextual Token

Ailiang Lin (Institute of Science Tokyo), Manabu Okumura (Institute of Science Tokyo)

RetrievalComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningText

🎯 What it does: Proposes Causal2Vec, a method that generates a single Contextual Token using a lightweight bidirectional encoder while preserving the original causal attention. This Contextual Token is then pre-inserted into the input sequence of the decoder-only LLM. Finally, the hidden states of Contextual and EOS are concatenated to form the final text embedding.

CE-GPPO: Coordinating Entropy via Gradient-Preserving Clipping Policy Optimization in Reinforcement Learning

Zhenpeng Su (Kuaishou Technology), Guorui Zhou (Kuaishou Technology)

Reinforcement LearningTextBenchmark

🎯 What it does: Proposes the CE-GPPO algorithm, which achieves fine-grained control over policy entropy by retaining and adjusting the gradients of clipped low-probability tokens in PPO.

CEBC: Conformal Evidence-Bounded Control for Low-Hallucination Vision–Language Generation

Ashish Mishra (Hewlett Packard Labs), Martin Foltin (Hewlett Packard Labs)

GenerationExplainability and InterpretabilityComputational EfficiencyTransformerPrompt EngineeringVision Language ModelImageTextMultimodalityRetrieval-Augmented Generation

🎯 What it does: Propose the CEBC framework, which uses evidence-constrained minimal editing to reduce hallucinations (false object mentions) in vision-language models during inference, without requiring retraining;

CEDAR: A Chinese Evaluation Dataset for Computational Argumentation

Tian Lan (Inner Mongolia University), Xiangdong Su (Inner Mongolia Key Laboratory of Multilingual Artificial Intelligence Technology)

ClassificationRecognitionTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextBenchmarkAudio

🎯 What it does: Proposed and released the CEDAR Chinese spoken debate dataset, containing 600 debates, 318 topics, and 250k sentences;

Cell-Based Representation of Relational Binding in Language Models

Qin Dai, Kentaro Inui (MBZUAI)

Explainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelText

🎯 What it does: By constructing narrative texts with multiple sentences and cross-entity, multi-relational structures, this study investigates how large language models bind entities and relations at the discourse level, and proposes a Cell-Based Binding Representation (CBR) subspace to represent this binding mechanism.

Chain-of-Memory: Lightweight Memory Construction with Dynamic Evolution for LLM Agents

Xiucheng Xu (State Key Laboratory of AI Safety), Huawei Shen (State Key Laboratory of AI Safety)

RetrievalExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Designed a lightweight out-of-memory in-memory construction and dynamic linked memory evolution framework called CoM, to enhance the long-term memory and reasoning capabilities of LLM agents.

Chain-of-Thought as a Lens: Evaluating Structured Reasoning Alignment between Human Preferences and Large Language Models

Boxuan Wang (University of Liverpool), Yi Dong (University of Liverpool)

Explainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose Alignment Score, which quantifies the alignment between the chain-of-thought reasoning generated by large language models and human preference reference chains using a semantic entropy matrix, and design alignment-based chain sampling and selection methods (ACSS, SC-Align) based on this.

Chain-of-Thought Compression Should Not Be Blind: V-Skip for Efficient Multimodal Reasoning via Dual-Path Anchoring

Dongxu Zhang (Xi'an Jiaotong University), Haijun Zhang (University of Science and Technology Beijing)

CompressionComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This paper proposes a compression method called V-Skip for multi-modal chain-of-thought reasoning, which can significantly shorten the reasoning length while preserving the visual foundation.

Chain-of-Thought Degrades Visual Spatial Reasoning Capabilities of Multimodal LLMs

Sai Srinivas Kancheti (Indian Institute of Technology Hyderabad), Tanuja Ganu (Microsoft Research India)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerSupervised Fine-TuningReinforcement LearningPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmarkChain-of-Thought

🎯 What it does: A systematic evaluation of the Chain-of-Thought (CoT) reasoning approach in Multimodal Large Language Models (MRM) on visual spatial reasoning tasks reveals that CoT significantly reduces the model's spatial reasoning performance.

Chaining the Evidence: Robust Reinforcement Learning for Deep Search Agents with Citation-Aware Rubric Rewards

Jiajie Zhang (Tsinghua University), Juanzi Li (Zhipu AI)

Explainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposes two techniques, Citation-aware Rubric Rewards (CaRR) and C-GRPO, to introduce fine-grained, interpretable reward signals into the reinforcement learning training of deep search agents, thereby improving the agents' comprehensiveness of reasoning, factual accuracy, and coherence of evidence.

CHAIRO: Contextual Hierarchical Analogical Induction and Reasoning Optimization for LLMs

Haotian Lu (Shenzhen University), Bingzhe Wu (Shenzhen University)

ClassificationOptimizationExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose a content moderation framework called CHAIRO based on large language models, which achieves more accurate moderation decisions through analogy retrieval, rule induction, and hierarchical reasoning.

Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models

Haoxiang Sun (Renmin University of China), Ji-Rong Wen (Renmin University of China)

Large Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This paper constructs a new bilingual (English-Chinese) Olympiad-level mathematics benchmark called OlymMATH, which includes 200 arithmetic problems (EASY/HARD) verifiable by Sympy and 150 formal proof problems using Lean4;

Challenging the Explanation Based on Preceding Tokens: Discovering Transferable Non-Literal Biasing

Yuchen Huang (Shanghai Jiao Tong University), Quanshi Zhang (Shanghai Jiao Tong University)

Explainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: It was found that generated preceding words, although semantically unrelated to the answer, can significantly increase the probability of LLMs generating the target answer and can transfer across prompts.

Characterizing and Evaluating Working Emotion Vocabularies in Multilingual Large Language Models

Nicholas Deas (Columbia University), Kathleen McKeown (Amherst College)

RecognitionTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringText

🎯 What it does: This paper systematically evaluates the understanding ability of multilingual large language models toward fine-grained emotional vocabulary through three tasks: emotional state recognition, emotional state expression, and emotional state verification.

Characterizing the Expressivity of Local Attention in Transformers

Jiaoda Li (ETH Zurich), Ryan Cotterell (ETH Zurich)

RecognitionExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelText

🎯 What it does: Formally analyze the local attention expression capability of Transformer, proving its correspondence with a fragment of linear temporal logic.

Chart-MRAG: Benchmarking Multimodal Retrieval Augmented Generation on Chart-based Documents

Ymyang, Kaiwen Wei (Chongqing University)

GenerationRetrievalTransformerLarge Language ModelVision Language ModelContrastive LearningImageTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposed the Chart-based MRAG task, focusing on retrieval-augmented generation for chart documents, and built the CHARGE framework for automatic generation, based on which the Chart-MRAG Bench benchmark was created.

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering

Rachneet Kaur (J P Morgan AI Research), Sumitra Ganesh (J P Morgan AI Research)

Explainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelAgentic AIVision Language ModelImageTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed the ChartAgent framework, which performs visual-rooted reasoning on charts through multi-round interactive tool calls to answer natural language questions.

CharTide: Data-Centric Chart-to-Code Generation via Tri-Perspective Tuning and Inquiry-Driven Evolution

Xiangxi Zheng (Nanjing University), Alex Jinpeng Wang (Central South University)

Data-Centric LearningAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision Language ModelImageTextMultimodalityRetrieval-Augmented Generation

🎯 What it does: Proposes a framework called CharTide for high-precision Chart-to-Code generation, which first decomposes visual perception, code logic, and modal fusion through a three-dimensional split SFT training, and then performs verifiable alignment through inquiry-driven RL.

ChartVerse: Scaling Chart Reasoning via Reliable Programmatic Synthesis from Scratch

Zheng Liu (Shanghai AI Laboratory), Lijun Wu

GenerationData SynthesisKnowledge DistillationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringGenerative Adversarial NetworkContrastive LearningImageTextMultimodalityBenchmarkChain-of-Thought

🎯 What it does: A scalable chart reasoning data synthesis framework called ChartVerse was constructed, capable of generating diverse and complex charts as well as high-quality question-answer data from scratch.

ChatAnime: Towards User-Centered Emotional Support in LLM-based Virtual Character Chat

Lanlan Qiu (Shanghai Qi Zhi Institute), Tianxing He (Shanghai Qi Zhi Institute)

Recommendation SystemExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: Constructed the first multi-turn role-playing dataset for emotional support, named ChatAnime, and proposed the Emotionally Supportive Role-Playing (ESRP) framework based on users' emotional needs;

ChatHLS: Towards Systematic Design Automation and Optimization for High-Level Synthesis

Runkai Li (Southeast University), Xi Wang (Southeast University)

OptimizationAI Code AssistantTransformerLarge Language ModelAgentic AIPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: ChatHLS proposes a multi-agent framework that leverages LLM to achieve HLS code generation, instruction optimization, and debugging, significantly improving generation success rate and hardware performance.

ChatR1: Reinforcement Learning for Conversational Reasoning and Retrieval Augmented Question Answering

Simon Lupart (University of Amsterdam), Evangelos Kanoulas (University of Amsterdam)

RetrievalOptimizationTransformerLarge Language ModelReinforcement LearningTextRetrieval-Augmented Generation

🎯 What it does: Developed a dialogue-based question answering framework called ChatR1 based on reinforcement learning, which can dynamically interact and retrieve information and reason in multi-turn conversations.

Check Your Work: Structured Checklist Feedback for Improving Large Language Models

Jonathan Cook (University of Oxford), Alex Wang (Cohere)

Reinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextBenchmark

🎯 What it does: This paper proposes decomposing AI feedback into fine-grained, task-specific checklists, and uses dynamic variance weighting (DIVA) to incorporate them as reward signals for RL fine-tuning, while also using the checklists as a structured self-correction mechanism during inference;

CheckMIABench: Firm Foundations For Membership Inference Attacks on Language Models

Jeffrey George Wang (Harvard University), Seth Neel (Harvard Business School)

Safty and PrivacyTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: A clean benchmark based on training checkpoints was constructed to evaluate membership inference attacks on language models, and the effectiveness of various published attacks was validated on Pythia and OLMo.

CheckRLM: Effective Knowledge–Thought Coherence Checking in Retrieval-Augmented Reasoning

Dingling Xu (Beijing Normal University), Maosong Sun (Chinese Academy Of Sciences)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: Proposes the CheckRLM framework, which reduces error accumulation during RLM inference by promptly checking and correcting factual errors through retrieval-augmented generation (RAG).

ChemReason-Bench: Benchmarking Large Language Models for Procedural Reasoning in Experimental Chemistry

Jinwei Zhang (Shanghai Jiao Tong University), Yanyan Xu (Shanghai Jiao Tong University)

Drug DiscoveryTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Constructed and released two verifiable chemical experiment procedure reasoning benchmarks and training sets, ChemReason-Bench and ChemReason-Tune, to evaluate the capabilities of multiple models across six tasks.

ChessArena: A Chess Testbed for Evaluating Strategic Reasoning Capabilities of Large Language Models

Jincheng Liu (Nanjing University), Yuan Yao (Nanjing University)

TransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningTextBenchmark

🎯 What it does: Created the ChessArena platform, which uses complete games to evaluate the strategic reasoning capabilities of large language models, and designed four gameplay modes and fine-grained evaluation tasks.

ChiKhaPo: A Large-Scale Multilingual Benchmark for Evaluating Lexical Comprehension and Generation in Large Language Models

Emily Chang (Toyota Technological Institute at Chicago), Niyati Bafna (Johns Hopkins University)

TransformerLarge Language ModelTextMultimodalityBenchmark

🎯 What it does: Proposed a multilingual lexical-level evaluation benchmark called ChiKhaPo, covering more than 2700 languages, which includes 8 subtasks (word translation, contextual word translation, translation-conditioned language modeling, bag-of-words machine translation) to evaluate the lexical understanding and generation capabilities of large language models (LLMs).

CHIMERA: A Knowledge Base of Scientific Idea Recombinations for Research Analysis and Ideation

Noy Sternlicht (Hebrew University of Jerusalem), Tom Hope (Hebrew University of Jerusalem)

Recommendation SystemExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningTextReview/Survey PaperBenchmarkRetrieval-Augmented Generation

🎯 What it does: Established and made public the CHIMERA knowledge base, which automatically extracts two types of innovative recombination instances, namely 'blend' and 'inspiration,' from the abstracts of scientific papers, and provides the extraction model and annotated dataset; meanwhile, it demonstrates the application of this KB in metascientific analysis and scientific creativity prediction.

ChipSeek: Optimizing Verilog Generation via EDA-Integrated Reinforcement Learning

Zhirong Chen (Chinese Academy of Sciences), Ying Wang (Chinese Academy of Sciences)

OptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextChain-of-Thought

🎯 What it does: Proposes the ChipSeek framework, which utilizes reinforcement learning combined with complete EDA toolchain feedback to directly optimize the functional correctness of RTL (Verilog) code and PPA (Power, Performance, Area) metrics.

CHOIR: Harmonizing Structured Persona Diversity for Robust Collaborative LLM Reasoning

Xiangjue Dong (Texas A&M University), James Caverlee (Texas A&M University)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsTextBenchmarkChain-of-Thought

🎯 What it does: Propose a test-time framework called CHOIR, which generates multiple personas by perturbing structured demographic attributes (such as gender, race, religion, disability, age, etc.) and enhances the reasoning robustness of large language models through token-level collaborative weight fusion during decoding.

Chunks as Arms: Multi-Armed Bandit-Guided Sampling for Long-Context LLM Preference Optimization

Shaohua Duan (Northeastern University), Maosong Sun (Microsoft Research Asia)

OptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningMixture of ExpertsText

🎯 What it does: A framework named LongMab is constructed, which utilizes a multi-armed bandit (MAB) strategy to perform chunk sampling on long texts, generating high-quality and diverse preference pairs for Direct Preference Optimization (DPO) fine-tuning, thereby enhancing the long-context reasoning ability of large language models (LLMs).

CIA: Inferring the Communication Topology from LLM-based Multi-Agent Systems

Yongxuan Wu (Chinese Academy of Sciences), Yanan Cao (Chinese Academy of Sciences)

Explainability and InterpretabilityAdversarial AttackReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringTextSequentialRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose an attack method called Communication Inference Attack (CIA) to infer the communication topology of multi-agent systems (MAS) in large language models under a black-box setting;

CIG: Measuring Conversational Information Gain in Deliberative Dialogues with Semantic Memory Dynamics

Ming-Bin Chen (University of Melbourne), Lea Frermann (University of Melbourne)

Explainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: Propose the Conversational Information Gain (CIG) framework, which measures information gain in conversations using semantic memory.

CiPO: Counterfactual Unlearning for Large Reasoning Models through Iterative Preference Optimization

Junyi Li (Hong Kong University of Science and Technology (Guangzhou)), Ningning Ding (Hong Kong University of Science and Technology (Guangzhou))

OptimizationSafty and PrivacyExplainability and InterpretabilityTransformerLarge Language ModelReinforcement LearningPrompt EngineeringSimultaneous Localization and MappingTextBenchmarkChain-of-Thought

🎯 What it does: Propose a counterfactual iterative preference optimization (CiPO) mechanism for large reasoning models (LRM), achieving machine unlearning of chain-of-thought (CoT) and answers;

CIRAG: Construction–Integration Retrieval and Adaptive Generation for Multi-hop Question Answering

Zili Wei (Northeastern University), Yifei Zhang (Northeastern University)

GenerationRetrievalKnowledge DistillationTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextMultimodalityRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose a multi-modal retrieval-augmented generation framework called CIRAG, which uses structured triplets as retrieval units and gradually focuses on core evidence through an iterative construction-integration process;

CIS-BWE: Chaos-Informed Speech Bandwidth Extension

Tarikul Islam Tamiti (Chittagong University of Engineering and Technology), Anomadarshi Barua (Chittagong University of Engineering and Technology)

RestorationDepth EstimationCompressionTransformerDiffusion modelGenerative Adversarial NetworkContrastive LearningAudio

🎯 What it does: Proposed a bandwidth extension model called CIS-BWE based on complex-valued GAN, which integrates Chaos-informed discriminator MRLD and MSDFA, as well as a dual-stream ConformerNeXt generator.

CITE: Benchmarking Heterogeneous Text-Attributed Graph Models

Chenghao Zhang (Chinese Academy of Sciences), Yi Du (Chinese Academy of Sciences)

ClassificationRecommendation SystemGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextGraphBenchmark

🎯 What it does: This paper constructs the first large-scale heterogeneous text-attribute graph dataset, CITE, containing 438K nodes, 1.22 million edges, four node types (papers, authors, journals, keywords), and four relationship types. Based on this dataset, various learning paradigms (homogeneous GNN, heterogeneous GNN, LLM, and LLM+Graph) are systematically evaluated for node classification and link prediction tasks.

CiteGuard: Faithful Citation Attribution for LLMs via Retrieval-Augmented Validation

Yee Man Choi (University of Waterloo), Qingyun Wang (University of Illinois Urbana-Champaign)

RetrievalExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelAgentic AITextRetrieval-Augmented Generation

🎯 What it does: Proposed a retrieval-enhanced verification-based LLM agent called CiteGuard for achieving citation attribution in scientific writing.

CityCube: Benchmarking Cross-view Spatial Reasoning on Vision-Language Models in Urban Environments

Haotian Xu (National University of Defense Technology), Yong Li (Tsinghua University)

Autonomous DrivingComputational EfficiencyRepresentation LearningTransformerSupervised Fine-TuningPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBenchmarkChain-of-Thought

🎯 What it does: Construct the CityCube benchmark, which includes 18.1k urban multi-view images and 5,022 cross-view spatial reasoning QA, evaluating and fine-tuning 33 VLMs.

CityVG: Contrastive Fine-Tuning and Reward-Based Chain-of-Thought Reasoning for Zero-Shot City-Scale 3D Visual Grounding

Jianjun Zhang (Tongji University), Hanli Wang (Tongji University)

RecognitionRetrievalComputational EfficiencyRepresentation LearningTransformerLarge Language ModelReinforcement LearningVision Language ModelContrastive LearningTextPoint CloudChain-of-Thought

🎯 What it does: Proposed the first zero-shot, city-scale 3D visual localization framework, CityVG, which can locate target objects in massive urban point clouds without annotations.

CL^2GEC: A Multi-Discipline Benchmark for Continual Learning in Chinese Literature Grammatical Error Correction

Shang Qin (Tsinghua University), Hong-Gee Kim (Seoul National University)

TransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextBenchmark

🎯 What it does: Proposed CL GEC 2, the first multidisciplinary Chinese Grammar Error Correction (CGEC) continual learning benchmark, simulating syntactic differences across disciplines and the need for continuous adaptation in academic writing.

ClaimDB: A Fact Verification Benchmark over Large Structured Data

Michael Theologitis (University of Washington), Dan Suciu (University of Washington)

Large Language ModelAgentic AIPrompt EngineeringTabularBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Designed a scalable structured data fact verification benchmark called CLAIMDB;

CLAOCS-TX: Cross-Lingual Triplet Extraction with Aspect-Opinion-Aware Code-Switched Prompting and LLM-Guided Contrastive Distillation

Lipika Dewangan (Indian Institute of Technology Indore), Chandresh Kumar Maurya (Indian Institute of Technology Indore)

Domain AdaptationKnowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningText

🎯 What it does: Achieve zero-shot aspect-term sentiment triplet extraction (ASTE) in unsupervised cross-lingual scenarios by generating pseudo labels via large language models as semantic teachers, complemented by multi-variant consistency filtering and triplet-level contrastive distillation to train a single student model for cross-lingual reasoning.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts

Jiuheng Lin (Peking University), Yansong Feng (Peking University)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextTabularReview/Survey PaperBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposes the CLARITY framework, which utilizes consistency rewards and a two-phase refine-then-monitor training process, and enhances the reasoning consistency and accuracy of multiple-choice question (MCQ) reinforcement learning (RL) training through dynamic data reconstruction, relying solely on a small general-purpose LLM, without the need for large PRMs or expert annotations.

CLEAR: Cross-Lingual Enhancement in Retrieval via Reverse-training

Seungyoon Lee (Korea University), Heuiseok Lim (Kyungpook National University)

RetrievalDomain AdaptationRepresentation LearningTransformerSupervised Fine-TuningContrastive LearningText

🎯 What it does: This paper proposes a cross-lingual retrieval loss function called CLEAR based on inverse training, which leverages English paragraphs as bridges to enhance multi-lingual retrieval performance.

CloneMem: Benchmarking Long-Term Memory for AI Clones

Sen Hu (Peking University), Huacan Wang (University of Chinese Academy of Sciences)

TransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose the CLONEMEM benchmark, which evaluates the long-term memory capability of AI clones using non-conversational digital traces (diaries, social media posts, emails, etc.) over a span of 1–3 years; coherent life trajectories are constructed through hierarchical generation (macro-level life arc → meso-level stages → micro-level daily traces), and multi-level reasoning tasks are designed.

Closing the Modality Reasoning Gap for Speech Large Language Models

Chaoren Wang (Chinese University of Hong Kong), Zhizheng Wu (Chinese University of Hong Kong)

Representation LearningData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningContrastive LearningTextMultimodalityAudio

🎯 What it does: Propose a trajectory alignment framework based on reinforcement learning called TARS, aimed at closing the performance gap between speech and text in reasoning.

Closing the Spatial Execution Gap in Digital Whiteboards via Verifiable Reinforcement Learning

Chang Liu (Colorado School of Mines), Bo Wu (Colorado School of Mines)

OptimizationComputational EfficiencyAI Code AssistantTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringVision-Language-Action ModelImageTextMultimodalityBenchmark

🎯 What it does: This paper proposes a 'System 2' reasoning protocol based on verifiable reinforcement learning, aiming to address the spatial execution gap in digital whiteboard interaction for multimodal large language models.

Clozing the Gap: Exploring Why Language Model Surprisal Outperforms Cloze Surprisal

Sathvik Nair (University of Maryland), Byung-Doh Oh (Nanyang Technological University)

Explainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringText

🎯 What it does: This paper explores why language model (LM) surprisal can better predict reading time than human cloze experiments, and experimentally verifies the reasons for its advantages.

ClusterRAG: Cluster-Based Collaborative Filtering for Personalized Retrieval-Augmented Generation

Gibson Nkhata (University of Arkansas), Susan Gauch (University of Arkansas)

RetrievalRecommendation SystemTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextRetrieval-Augmented Generation

🎯 What it does: Propose ClusterRAG, a clustering-based collaborative filtering method aimed at enhancing the effectiveness of personalized retrieval-augmented generation (RAG);

CNSL-bench: Benchmarking the Sign Language Understanding Capabilities of MLLMs on Chinese National Sign Language

Rui Zhao (Xiamen University), Yidong Chen (Xiamen University)

TransformerLarge Language ModelPrompt EngineeringVision Language ModelImageVideoTextMultimodalityBenchmarkChain-of-Thought

🎯 What it does: Propose CNSL-bench — a multi-modal (text, image, video) evaluation benchmark based on an official sign language dictionary, used to assess the understanding ability of multi-modal large language models (MLLMs) towards Chinese national sign language.

CO-EVO: Co-evolving Semantic Anchoring and Style Diversification for Federated DG-ReID

Fengchunzhang, Jianwei Hu (QiYuan Lab)

RecognitionDomain AdaptationFederated LearningTransformerVision Language ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningImage

🎯 What it does: This paper proposes a federated domain generalization framework for person re-identification called CO‑EVO, which enhances recognition performance on unseen target domains through collaborative evolution of semantic anchoring and style diversification.

CoAct: Co-Active LLM Preference Learning with Human-AI Synergy

Ruiyao Xu (Northwestern University), Kaize Ding (Northwestern University)

Data-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose the COACT framework, which combines self-reward and active learning through human-computer collaboration, generating self-consistent preference pairs for self-labeling, and improving LLM alignment through oracle screening and instruction augmentation.

CoCoNUTS: Concentrating on Content while Neglecting Uninformative Textual Styles for AI-Generated Peer Review Detection

Yihan Chen (University of Chinese Academy of Sciences), Le Sun (Chinese Information Processing Laboratory, Institute of Software, Chinese Academy of Sciences)

ClassificationAnomaly DetectionTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextBenchmark

🎯 What it does: Propose the CoCoNUTS benchmark and the CoCoDet detector, focusing on content rather than text style, achieving precise identification of AI-generated peer reviews.

CodeEvo: Interaction-Driven Synthesis of Code-centric Data through Hybrid and Iterative Feedback

Qiushi Sun (University of Hong Kong), Fei Yuan (University of Hong Kong)

Data SynthesisAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringText

🎯 What it does: Propose CodeEvo, an interactive framework based on dual agents (Coder and Reviewer), for automatically generating high-quality, executable, and logically complex instruction-code pairs, and build the CodeEvo-100K dataset.

CodeFlowBench: A Multi-turn, Iterative Benchmark for Complex Code Generation

Sizhe Wang (Peking University), Wentao Zhang (Peking University)

AI Code AssistantTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Constructed the CodeFlowBench benchmark to evaluate the performance of LLMs in multi-round, iterative code generation (codeflow).

CodeHacker: Automated Test Case Generation for Detecting Vulnerabilities in Competitive Programming Solutions

Jingwei Shi (Shanghai University of Finance and Economics), Jinman Zhao (Shanghai University of Finance and Economics)

Adversarial AttackAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed and implemented CodeHacker, an automated adversarial test case generation framework for uncovering vulnerabilities in competitive programming code.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks

Hongchao Jiang (ASUS Intelligent Cloud Services), Robby T. Tan (ASUS Intelligent Cloud Services)

AI Code AssistantTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBenchmark

🎯 What it does: Construct the CodeJudgeBench benchmark to evaluate the judgment capabilities of LLMs in code generation, code repair, and unit test generation tasks.

CodeRipple: Wavelet-Based Detection of LLM-Generated Code

Xingyu Yao, Quan Wang (Beijing University of Posts and Telecommunications)

Anomaly DetectionAI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningTextBenchmark

🎯 What it does: Proposed and implemented a training-free LLM code generation detection framework called CodeRipple, which distinguishes between LLM-generated code and human-written code by performing local waveform analysis on the Token Perplexity Sequence (TPS).

CODERL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

Xue Jiang (Peking University), Ge Li (Peking University)

AI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningTextSequentialBenchmark

🎯 What it does: In the post-training phase of code generation, the authors propose the CODERL+ method, which uses execution semantics alignment to train the model with reinforcement learning;

CODESTRUCT: Code Agents over Structured Action Spaces

Myeongsoo Kim (AWS AI Labs), Murali Krishna Ramanathan (AWS AI Labs)

AI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose the CODESTRUCT framework, allowing code agents to read and write structured named entities (such as files, classes, functions, methods) through AST, rather than traditional text strings;

CoDial: Interpretable Task-Oriented Dialogue Systems Through Dialogue Flow Alignment

Radin Shayanfar (Queen's University), Xiaodan Zhu (Queen's University)

Explainability and InterpretabilityAI Code AssistantGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringTextGraphRetrieval-Augmented Generation

🎯 What it does: Propose the CoDial framework, which automatically converts the task schema of task-oriented dialogue (TOD) systems into interpretable procedural Guardrail code, achieving zero-shot cross-task deployment.

CoEvolve: Training LLM Agents via Agent-Data Mutual Evolution

Shidong Yang (Alibaba Group), Xiangxiang Chu (Alibaba Group)

TransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringText

🎯 What it does: Propose the CoEvolve framework to achieve closed-loop co-evolution between LLM agents and training data;

CoG: Controllable Graph Reasoning via Relational Blueprints and Failure-Aware Refinement over Knowledge Graphs

Yuanxiang Liu (Zhejiang University), Wen Zhang (Zhejiang University)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringTextGraphBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed the CoG framework, which employs a dual-process (intuition + reasoning) mechanism to achieve controllable, multi-hop reasoning on knowledge graphs.

CogEvolve: A Multimodal Benchmark for Evaluating Relational Reasoning in Semantic Extension

Jingjie Zeng (Dalian University of Technology), Hongfei Lin (Dalian University of Technology)

Explainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelPrompt EngineeringVision Language ModelTextMultimodalityGraphBenchmarkChain-of-Thought

🎯 What it does: Propose the CogEvolve benchmark, specifically designed to evaluate models' generative reasoning capabilities in semantic evolution (analogy, metaphor, metonymy);

Cognitive Alpha Mining via LLM-Driven Code-Based Evolution

Fengyuan Liu (University of Hong Kong), Qi Liu (University of Hong Kong)

Explainability and InterpretabilityTransformerLarge Language ModelAgentic AIPrompt EngineeringTabularTime SeriesFinance Related

🎯 What it does: Constructed the CogAlpha framework, which utilizes large language models (LLMs) combined with a seven-layer agent hierarchy and evolutionary search to generate interpretable and robust alpha factors from raw OHLCV data in financial markets and iteratively evolve them.

Cognitive Policy-Driven LLM for Diagnosis and Intervention of Cognitive Distortions in Emotional Support Conversation

Lin Zhong (Harbin Institute of Technology), Qing Liao (Harbin Institute of Technology)

OptimizationFederated LearningSafty and PrivacyExplainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextBiomedical DataReview/Survey PaperBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose the CoPoLLM framework, combining CBT theory to construct the first dataset with cognitive dissonance labels, CogBiasESC, and achieve diagnosis and intervention of cognitive dissonance in emotional support dialogues through cognitive strategy reinforcement learning and dual-stream conditional optimization.

Cognitive Scaffold: From Fluid Context to Crystallized Memory for Long-Horizon DeepResearch Agents

Qiuyuan Ai (Peking University), Guannan He (Peking University)

CompressionExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextGraphBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose the Cognitive Scaffold framework, which separates the short-term reasoning context from the long-term memory knowledge graph;

CogToM: A Comprehensive Theory of Mind Benchmark inspired by Human Cognition for Large Language Models

Haibo Tong (Chinese Academy of Sciences), Yi Zeng (Chinese Academy of Sciences)

Large Language ModelPrompt EngineeringTextBenchmarkChain-of-Thought

🎯 What it does: Constructed a large-scale, multi-paradigm, bilingual theory of mind evaluation benchmark called CogToM, and systematically evaluated 22 large language models.

CoLA: A Choice Leakage Attack Framework to Expose Privacy Risks in Subset Training

Qi Li (King Abdullah University of Science and Technology), Di Wang (Johns Hopkins University)

Safty and PrivacyAdversarial AttackConvolutional Neural NetworkTransformerPrompt EngineeringAuto EncoderContrastive LearningImageText

🎯 What it does: This paper proposes a novel attack framework called CoLA to reveal privacy leakage risks during subset training. It investigates two types of privacy surfaces: training membership inference (TM-MIA) and selection participation inference (SP-MIA), and designs a multi-throw membership inference method tailored for subset-aware side-channel attack and black-box attack scenarios.

Collaboration of Fusion and Independence: Hypercomplex-driven Robust Multi-Modal Knowledge Graph Completion

Zhiqiang Liu (Zhejiang University), Wen Zhang (Ant Group)

Knowledge DistillationRepresentation LearningGraph Neural NetworkMixture of ExpertsContrastive LearningImageTextMultimodalityGraph

🎯 What it does: Proposes a multi-modal knowledge graph completion model called M-Hyper based on biquaternion space, combining fusion and independent modes to achieve collaborative representation of multi-modal information.

Collaborative Chain-of-Agents for Parametric-Retrieved Knowledge Synergy

Yi Jiang (Harbin Institute of Technology), Bing Qin (Harbin Institute of Technology)

RetrievalRecommendation SystemExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelSupervised Fine-TuningAgentic AITextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose the CoCoA framework, which adopts multi-agent collaborative reasoning and long-chain training, explicitly fusing model parameter knowledge with retrieval knowledge;

Collision to Cognition: Hash-Driven Graph Construction for Efficient RAG

Chuang Zhou (Hong Kong Polytechnic University), Xiao Huang (Hong Kong Polytechnic University)

RetrievalExplainability and InterpretabilityComputational EfficiencyGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextGraphRetrieval-Augmented Generation

🎯 What it does: Propose a lightweight graph construction framework called MeshRAG based on Locality-Sensitive Hashing (LSH), which is used to quickly build an interpretable two-layer knowledge graph in Retrieval-Augmented Generation (RAG), and directly supports LLM generation.

Combining Distantly Supervised Models with In Context Learning for Monolingual and Cross-Lingual Relation Extraction

Vipul Kumar Rathore (Indian Institute of Technology New Delhi), Mausam (Indian Institute of Technology New Delhi)

TransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextReview/Survey PaperBenchmarkRetrieval-Augmented Generation

🎯 What it does: This paper proposes the HYDRE framework, which combines candidate relation generation based on DSRE with context learning of LLMs, utilizing dynamically retrieved high-quality examples for relation discrimination.

CoMeT: Collaborative Memory Transformer for Efficient Long Context Modeling

Runsong Zhao (Northeastern University), Bo Zheng (Alibaba)

RetrievalComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: Proposed the CoMeT model as a pluggable module, enabling large language models to process text of arbitrary length with constant memory and linear time.

Common to Whom? Regional Cultural Commonsense and LLM Bias in India

Sangmitra Madhusudan (Emory University), Ali Emami (Emory University)

Large Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposed the INDICA benchmark to evaluate cultural common sense differences across regions in India, and constructed a question-answering dataset for five major regions through human annotations.

CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data

Pedro Ortiz Suarez (Common Crawl Foundation), Sarah K. K. Luger

ClassificationRecognitionTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextBenchmark

🎯 What it does: Constructed CommonLID, a community-driven, human-annotated language identification benchmark for the Web domain covering 109 languages.

Communicating in Emergent Language with an Induced Morphological Phrasebook

Brendon Boldt (Carnegie Mellon University), David R. Mortensen

Representation LearningData-Centric LearningRecurrent Neural NetworkReinforcement LearningAuto EncoderGenerative Adversarial NetworkContrastive LearningSequential

🎯 What it does: This paper constructs rule-based senders and receivers using a morphology dictionary induced by CSAR in a reconstruction game, verifying their communication effectiveness in Emergent Language (EL), and compares them with senders/receivers based on neural network genetics; meanwhile, through ablation experiments on the dictionary and sender/receiver algorithms, it reveals the morphological and syntactic features in EL and proposes a new compositional metric called morpheme bijectivity.

Communication-Efficient Desire Alignment for Proactive Embodied Human–Agent Interaction

Yuanfei Wang (Peking University), Hao Dong (Peking University)

Computational EfficiencyRobotic IntelligenceTransformerLarge Language ModelPrompt EngineeringVision-Language-Action ModelImageTextMultimodalityRetrieval-Augmented Generation

🎯 What it does: Proposed the HA-Desire environment and the FAMER framework for achieving fast and efficient user desire alignment in embedded scenarios such as home robots.

Community-Aware Assessment of Social Textual Engagement and Resonance: A Human-Centric Perspective on User-Generated Content Evaluation

Tianjiao Li (Bilibili Inc), Huyang Sun (Bilibili Inc)

Recommendation SystemExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision Language ModelVideoTextMultimodalityBenchmarkChain-of-Thought

🎯 What it does: This paper proposes a user-generated content (UGC) quality assessment framework called MEDEA, centered on community resonance, which uses social chain-of-thought (Social-CoT) to simulate diverse audience reactions and thereby determine whether UGC can generate positive resonance within a community.

Comparative Analysis of the Intrinsic Metrics for Tokenizers and their effect on Downstream Tasks for Hindi and Marathi

Shagun Dwivedi (FLAME University), Kaushik Gopalan (FLAME University)

Computational EfficiencyRepresentation LearningData-Centric LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: This paper investigates the impact of various UTF-8-based tokenizers on the performance of pre-trained language models for Hindi and Marathi in downstream tasks (question answering, transcription, phonetic transcription, noise robustness), and proposes a tokenizer based on visual units (grapheme clusters).

Comparing human and language models sentence processing difficulties on complex structures

Samuel Joseph Amouyal (Tel Aviv University), Jonathan Berant (Tel Aviv University)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextChain-of-Thought

🎯 What it does: Conducted a sentence comprehension experiment on seven syntactic structures (including four types of garden-path sentences, double-center embedding, deep negation, similarity interference, etc.), collecting human and answer data from 31 LLMs (five major clans).