ACL 2026 Papers — Page 2
Annual Meeting of the Association for Computational Linguistics · 2296 papers
Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts Inference
Baihui Liu (National University of Defense Technology), Dongsheng Li (National University of Defense Technology)
Computational EfficiencyTransformerMixture of ExpertsText
🎯 What it does: Propose the Alloc-MoE framework, which collaboratively allocates expert activations at both the hierarchical and token levels through a global activation budget in sparse Mixture-of-Experts inference, significantly reducing inference latency while maintaining model performance.
AlphaContext: An Evolutionary Tree-based Psychometric Context Generator for Creativity Assessment
Yixuan Wang (East China Normal University), Jiajun Guo (East China Normal University)
GenerationData SynthesisOptimizationTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation
🎯 What it does: Propose AlphaContext, an evolutionary psychometric context generator that combines hierarchical planning, MCTS generation, MAP-Elites optimization, and virtual participant evaluation for creativity assessment.
Amadeus: Autoregressive Model with Bidirectional Attribute Modelling for Symbolic Music
Hongju Su (Beijing University of Posts and Telecommunications), Yi-Zhe Song (University of Surrey)
GenerationData SynthesisTransformerLarge Language ModelDiffusion modelAuto EncoderTextSequentialAudio
🎯 What it does: Proposed the Amadeus framework, which employs a two-tier architecture: first autoregressively generating a sequence of notes, and then using a bidirectional discrete diffusion model to decode note attributes, achieving symbolic music generation.
AMATA: Adaptive Multi-Agent Trajectory Alignment for Knowledge-Intensive Question Answering
Taolin Zhang (Hefei University of Technology), Richang Hong (Hefei University of Technology)
Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringMixture of ExpertsTextMultimodalityBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposed an adaptive multi-agent trajectory alignment framework named AMATA for knowledge-intensive question answering, which improves the interpretability and factual accuracy of answers by dynamically integrating external knowledge.
Among Us: Language of Conspiracy Theorists on Mainstream Reddit
Francesco Corso (Politecnico di Milano), Gianmarco De Francisci Morales (CENTAI)
ClassificationAnomaly DetectionExplainability and InterpretabilitySupervised Fine-TuningContrastive LearningText
🎯 What it does: The study investigates whether users participating in conspiracy theory communities (such as r/conspiracy) on mainstream Reddit subcommunities exhibit detectable differences in their language use, constructing user vectors using psycholinguistic features extracted from the LIWC-22 dictionary, and employing a random forest classifier for discrimination;
Among Us: Measuring and Mitigating Malicious Contributions in Model Collaboration Systems
Ziyuan Yang (University of Washington), Yulia Tsvetkov (University of Washington)
Federated LearningSafty and PrivacyAdversarial AttackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringMixture of ExpertsContrastive LearningTextTabularSequentialBenchmark
🎯 What it does: Systematically evaluate the impact of malicious models in multi-model collaborative systems and propose two mitigation strategies (unsupervised and supervised)
An Existence Proof for Neural Language Models That Can Explain Garden-Path Effects via Surprisal
Ryo Yoshida (University of Tokyo), Tatsuki Kuribayashi (MBZUAI)
Explainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningText
🎯 What it does: This paper demonstrates that there exists a neural LM capable of explaining both the garden path effect and natural reading time phenomena by fine-tuning the GPT-2 language model on garden path sentences.
An Experimental Study on the Influence of Culture on Cross-Lingual Sentiment Transfer
Ahao Liu (Central China Normal University), Wang Chuanrong
ClassificationDomain AdaptationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringContrastive LearningText
🎯 What it does: This paper quantifies the impact of cultural distance on cross-lingual sentiment transfer through over 400 large-scale experiments.
An Exploration of Mamba for Speech Self-Supervised Models
Tzu-Quan Lin (National Taiwan University), Hung-yi Lee (National Taiwan University)
RecognitionRecurrent Neural NetworkTransformerSupervised Fine-TuningContrastive LearningAudio
🎯 What it does: This paper systematically explores the self-supervised speech model constructed by replacing the traditional Transformer blocks of HuBERT with the Mamba state space model, and comprehensively evaluates its performance on long sequences, streaming ASR, and SUPERB evaluation tasks.
An Information-Theoretic Foundation for the Subregular Hierarchy
Mai Phan Quoc Hung (National Economics University), Tuan Do (B0Labs N2TP Technology Solutions JSC)
Information TheoryClassificationExplainability and InterpretabilityRepresentation LearningContrastive LearningTextReview/Survey PaperAudio
🎯 What it does: This paper clarifies the sub-regular hierarchy from an information theory perspective, proving the equivalence between finite-order Markov sources and finite-type shifts (SFT), demonstrating that non-sub-regular patterns such as head-tail assimilation cannot be realized by Markov sources, and using mutual information (MI) metrics to verify the distinguishability between SL and TSL classes on synthetic and real corpora.
Analytical FFN-to-MoE Restructuring via Activation Pattern Analysis
Zehua Pei (The Chinese University of Hong Kong), Bei Yu (The Chinese University of Hong Kong)
Computational EfficiencyKnowledge DistillationTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringMixture of ExpertsText
🎯 What it does: During the inference phase of large-scale language models, dense feed-forward networks (FFN) are transformed into sparse Mixture-of-Experts (MoE) architectures. By utilizing a small amount of calibration data, activation pattern analysis is performed to quickly build shared experts and routing experts, and performance improvements can be achieved through minimal fine-tuning.
Analyzing and Internalizing Complex Policy Documents for LLM Agents
Jiateng Liu (University Of Illinois Urbana Champaign), Heng Ji (Amazon)
Data SynthesisRecommendation SystemExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This paper explores how large language model agents can accurately execute complex business rules without explicitly inputting complete policy documents, and proposes a controllable complexity policy document generator, CC-Gen, for systematic evaluation and data generation; meanwhile, it designs a Category-Aware Policy Continued Pretraining (CAP-CPT) method, which automatically classifies policy specifications and generates customized training data for different types, combining continued pretraining with supervised fine-tuning to achieve efficient internalization;
Anchor: Branch-Point Data Generation for GUI Agents
Jinbiao Wei (Yale University), Arman Cohan (Yale University)
Data SynthesisAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposed and implemented the ANCHOR framework, which automatically expands desktop GUI trajectories by leveraging branch points from seed demonstrations, generating diverse and high-quality training data.
Anchored Sliding Window: Toward Robust and Imperceptible Linguistic Steganography
Ruiyi Yan (Kyoto University), Yugo Murawaki (Kyoto University)
Safty and PrivacyTransformerLarge Language ModelPrompt EngineeringContrastive LearningText
🎯 What it does: Propose the Anchored Sliding Window (ASW) framework to improve the robustness and undetectability of hidden text in language models;
Anchoring Depends on Confidence and Post-Training in Language Models
Hillary N. Owusu (University of Maryland), Naomi H. Feldman (University of Maryland)
Explainability and InterpretabilityKnowledge DistillationTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringText
🎯 What it does: This paper investigates the extent to which large language models are influenced by the anchoring effect in numerical judgments, and explores how model confidence and accuracy predict anchoring susceptibility.
Anchoring the Affective Manifold: Learning Canonical and Disentangled Representations via Generative Cross-Modal Alignment
Weibin Li (South China Normal University), Chi Man Vong (University of Macau)
GenerationData SynthesisExplainability and InterpretabilityRepresentation LearningTransformerAuto EncoderGenerative Adversarial NetworkContrastive LearningVideoTextMultimodalityAudio
🎯 What it does: By constructing a twin-space framework based on VAE, we learn the shared emotional subspace (VAD) and private subspace, and achieve cross-modal alignment and dual anchoring to generate structured emotional manifolds.
Anchoring the Cache: Mitigating Contextual Hallucination in KV-Compressed Long-Context Summarization
Yu Fu (University of California, Riverside), Yue Dong (Amazon)
CompressionExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation
🎯 What it does: A systematic study on the hallucination problem caused by KV cache compression techniques in long-text summarization is conducted, and a strategy is proposed to clear the KV cache of key retrieval heads during the decoding phase, called HalluKV. This strategy effectively anchors the retrieval heads' attention to the source text, thereby reducing the hallucination rate.
AnchorSeg: Language Grounded Query Banks for Reasoning Segmentation
Rui Qian (Fudan University), Dejing Dou (Fudan University)
SegmentationTransformerLarge Language ModelSupervised Fine-TuningVision Language ModelContrastive LearningImageTextMultimodality
🎯 What it does: Propose AnchorSeg, which achieves reasoning segmentation through a language-guided query bank, separating semantic reasoning from spatial localization;
Android Coach: Improve Online Agentic Training Efficiency with Single State Multiple Actions
Guo Gan (Zhejiang University), Hong Zhou (Zhejiang University)
TransformerReinforcement LearningAgentic AIPrompt EngineeringText
🎯 What it does: Propose the ANDROiD COACH framework, which introduces a single-state multi-action (SSMA) paradigm in online reinforcement learning. By sampling multiple actions in bulk and using the Critic for evaluation, it improves sample efficiency and accelerates training without requiring additional simulation interactions.
Annotating Dimensions of Social Perception in Text: A Sentence-Level Dataset of Warmth and Competence
Mutaz Ayesh (Cardiff University), Nedjma Ousidhoum (Cardiff University)
ClassificationRecommendation SystemData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextBenchmark
🎯 What it does: Constructed the first sentence-level social perception dataset, W&C-Sent, containing 1,633 English sentence-target pairs, with seven-point ratings for trust, likeability, and competence
Answering the Wrong Question: Reasoning Trace Inversion for Abstention in LLMs
Abinitha Gourabathina (Mit), Prasanna Sattigeri
Explainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose a method for abstention based on reasoning trajectory inversion called TRACE INVERSION. First, generate the Chain-of-Thought trajectory of the LLM, then reconstruct the query that the model believes from the trajectory, and finally compare the original query with the reconstructed query in terms of similarity. If the similarity is low, trigger abstention.
Anti-Length Shift: Dynamic Outlier Truncation for Training Efficient Reasoning Models
Wei Wu (University of Science and Technology of China), Hui Xiong (Hong Kong University of Science and Technology)
Computational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringText
🎯 What it does: During the reinforcement learning training phase, dynamic outlier truncation (DOT) is applied to fully correct answers to suppress redundant reasoning, thereby reducing the reasoning length; meanwhile, KL normalization and predictive dynamic sampling are introduced to maintain training stability.
APB-V: Accelerating Long-Video Understanding via Sequence-Parallelism-aware Approximate Attention
Yuxiang Huang (Tsinghua University), Zhiyuan Liu (Tsinghua University)
Computational EfficiencyTransformerLarge Language ModelVideo
🎯 What it does: Propose the APB-V framework, which accelerates long video inference on multiple GPUs through sequence parallel approximate attention;
Apertus: Democratizing Open and Compliant LLMs for Global Language Environments
Alejandro Hernández-Cano (EPFL), Imanol Schlag (ETH Zurich)
Safty and PrivacyExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringContrastive LearningTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposed the Apertus series of fully open-source large language models with 8B and 70B parameters, achieving systematic improvements in data compliance, memory suppression, multilingual coverage, and model transparency during both pre-training and post-training stages.
APEX-MEM: Agentic Semi-Structured Memory with Temporal Reasoning for Long-Term Conversational AI
Pratyay Banerjee (Amazon), Ankit Chadha (Amazon)
Recommendation SystemAutonomous DrivingFederated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningAdversarial AttackHyperparameter SearchData-Centric LearningRobotic IntelligenceMeta LearningDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackNeural Architecture SearchProtein Structure PredictionGraph Neural NetworkTransformerLarge Language ModelAgentic AIPrompt EngineeringTextGraphSequentialBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Designed a long-term dialogue memory framework based on attribute graphs, named APEX-MEM, which can construct and query spatiotemporal events and entity information in real time.
APPSI-139: A Parallel Corpus of English Application Privacy Policy Summarization and Interpretation
Pengyun Zhu (Tianjin University), Kui Ren (Tianjin University)
Safty and PrivacyExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelSupervised Fine-TuningMixture of ExpertsTextRetrieval-Augmented Generation
🎯 What it does: Constructed the APPSI-139 English academic privacy policy parallel corpus and proposed the TCSI-pp-V2 multi-task framework to achieve summarization and interpretation of privacy policies.
AraVQA: Building a New Arabic Factoid Visual Question Answering Dataset from Wikipedia
Sultan Alrowili (IBM Research AI), Mathan Kumar Eswaran (IBM Research AI)
Data SynthesisTransformerLarge Language ModelSupervised Fine-TuningVision Language ModelImageTextMultimodalityBenchmark
🎯 What it does: Utilized Arabic Wikipedia template tags and large language models to automatically generate over 50,000 multiple-choice visual question answering (VQA) data, constructing the AraVQA dataset and providing corresponding benchmark evaluations.
ARCHITECT: Uncertainty-Aware Dynamic Tool Learning via Causal Intervention for Open-World Agents
Zhangyi Wang (Nanyang Technological University), Zongze Li (Nanyang Technological University)
Autonomous DrivingExplainability and InterpretabilityAI Code AssistantReinforcement Learning from Human FeedbackGraph Neural NetworkTransformerSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextGraphTabularBenchmark
🎯 What it does: This paper proposes the ARCHITECT framework, which predicts the reliability of dynamically generated tools and performs root cause attribution by constructing a causal structure model called CTD (Causal Tool Diagnosis). The attribution results are then used to guide targeted repairs.
ARCQuant: Boosting NVFP4 Quantization with Augmented Residual Channels for LLMs
Haoqian Meng (Tianjin University), Xindian Ma (Tianjin University)
Computational EfficiencyKnowledge DistillationTransformerLarge Language ModelText
🎯 What it does: Propose the ARCQuant framework, utilizing incremental residual channels to achieve W4A4 quantization under NVFP4, in order to improve LLM inference efficiency.
Are Emotion and Rhetoric Neurons in LLM? Neuron Recognition and Adaptive Masking for Emotion-Rhetoric Prediction Steering
Li Zheng (Wuhan University), Zhuang Li (RMIT University)
RecognitionExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringContrastive LearningText
🎯 What it does: Systematically identify neurons related to emotion and rhetoric in large language models, and achieve controllable regulation of emotional and rhetorical expressions through adaptive masking and activation injection, verifying that rhetorical neurons can enhance emotional recognition;
Are LLMs Reliable Rankers? Rank Manipulation via Two-Stage Token Optimization
Tiancheng Xing (National University Of Singapore), Xiyang Hu (Arizona State University)
Recommendation SystemOptimizationAdversarial AttackTransformerLarge Language ModelPrompt EngineeringText
🎯 What it does: Proposed a two-stage token optimization method called RAF (Rank Anything First), which uses natural language text to induce target items to improve rankings in LLM rerankers;
Are they lovers or friends? Evaluating LLMs’ Social Reasoning in English and Korean Dialogues
Eunsu Kim (KAIST), Najoung Kim (Boston University)
ClassificationExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Evaluate the ability of LLMs to reason about speakers' social relationships (e.g., friends, lovers, etc.) in dialogues on the SCRIPTS dataset;
Are We Using the Right Benchmark: An Evaluation Framework for Visual Token Compression Methods
Chenfei Liao (Hong Kong University of Science and Technology), Xuming Hu (Hong Kong University of Science and Technology)
CompressionVision Language ModelAuto EncoderMultimodalityBenchmark
🎯 What it does: This paper proposes an evaluation framework for visual token compression methods in multimodal large models and constructs a new benchmark called VTC-Bench;
ARES: Adaptive Red-Teaming and End-to-End Repair of Policy-Reward System
Jiacheng Liang (Stony Brook University), Charith Peris (Amazon Nova Responsible AI)
Adversarial AttackReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextRetrieval-Augmented Generation
🎯 What it does: Developed the ARES framework, achieving adaptive red teaming and closed-loop dual-phase repair for bidirectional failures of core models and reward models in the RLHF system.
ARF-RLHF: Adaptive Reward-Following for RLHF through Emotion-Driven Self-Supervision and Trace-Biased Dynamic Optimization
Yuxuan Zhang (South China Normal University)
OptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningContrastive LearningText
🎯 What it does: Developed a self-supervised RLHF framework called ARF-RLHF, which continuously models user free feedback rewards through sentiment-driven self-supervision and dynamic optimization with TraceBias.
Arg-LLaDA: Argument Summarization via Large Language Diffusion Models and Sufficiency-Aware Refinement
Hao Li (University of Manchester), Goran Nenadic (University of Manchester)
GenerationExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringDiffusion modelTextRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposed a method for argumentative summarization based on large language diffusion models called Arg-LLaDA, which gradually improves the quality of the summary through iterative remasking and generation strategies.
ArgGenBench: Benchmarking the Complex Controlled Argument Generation Capability of Large Language Models
Bojun Jin (Harbin Institute of Technology), Ruifeng Xu (Harbin Institute of Technology)
GenerationTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Constructed the ArgGenBench benchmark, which contains 803 multi-dimensional control instructions (topics, stance, style, strategy, audience, key points, length, etc.) and human-validated reference arguments, and systematically evaluated the control argument generation of 15 large language models under zero-shot, SFT, and DPO settings.
Arguments that Alter Minds: LLM Rationales Sway Human (and LLM) Notions of Plausibility
Shramay Palta (University of Maryland), Rachel Rudinger (University of Maryland)
Explainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringTextChain-of-Thought
🎯 What it does: This study explores the impact of reasoning justifications generated by LLMs (GPT-4o) on the feasibility evaluation of humans and LLMs in commonsense multiple-choice questions (SIQA and CommonsenseQA), by generating three types of justifications: PRO, CON, and PRO+CON.
ARK: Answer-Centric Retriever Tuning via KG-augmented Curriculum Learning
Jiawei Zhou (Shanghai Jiao Tong University), Haiyun Jiang (Shanghai Jiao Tong University)
RetrievalData-Centric LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextGraphBenchmarkFinance Related
🎯 What it does: Developed the ARK retriever fine-tuning framework, optimizing the retriever's ability to assess answer sufficiency through knowledge graph-driven curriculum learning and contrastive learning, thereby improving the quality of long-text retrieval.
ART: Attention Replacement Technique to Improve Factuality in LLMs
Ziqin Luo (Fudan University), Chen Shen (Shanghai Jiao Tong University)
Explainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextChain-of-Thought
🎯 What it does: By analyzing the distribution of shallow attention heads in LLMs, this paper finds that a unified attention pattern leads to hallucinations, and proposes a training-agnostic attention replacement technique (ART), which replaces the unified attention heads in the shallow layers with local attention heads, thereby improving the model's truthfulness and inference performance.
arXiv2Table: Toward Realistic Benchmarking and Evaluation for LLM-Based Literature-Review Table Generation
Weiqi Wang (Johns Hopkins University), Daniel Khashabi (Johns Hopkins University)
GenerationData SynthesisTransformerLarge Language ModelPrompt EngineeringTextTabularBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Studied a more realistic scientific literature review table generation task, constructed the ARXIV2TABLE benchmark, and proposed an iterative batch generation method.
Ascending the Infinite Ladder: Benchmarking Spatial Deformation Reasoning in Vision-Language Models
Jiahuan Zhang (Autolab, Westlake University), Kaicheng Yu (Autolab, Westlake University)
Explainability and InterpretabilityComputational EfficiencySupervised Fine-TuningVision Language ModelDiffusion modelAuto EncoderImageTextMultimodalityPoint CloudMeshBenchmarkChain-of-Thought
🎯 What it does: Propose the Inf-Bench evaluation framework, which designs 2D–3D spatial deformation reasoning tasks based on Shapez and Rubik’s Cube, including forward and backward reasoning, and evaluates the reasoning depth of VLMs using an infinite staircase competition format.
Assessing the Belief Consistency of Large Language Models on the Logical Conversation Process
Tomoki Tsujimura (National Institute of Advanced Industrial Science and Technology), Hiroya Takamura (National Institute of Advanced Industrial Science and Technology)
Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningAI Code AssistantTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This paper designs a multiple-choice question answering (MCQA) experiment inspired by the 20 questions approach, using free generation sampling to estimate the output tendencies of large language models (LLMs) under different contexts, and measures the consistency of model beliefs through Jensen-Shannon divergence.
ASTRA: Adaptive Semantic Tree Reasoning Architecture for Complex Table Question Answering
Xiaoke Guo (Zhejiang University), Wen Zhang (Zhejiang University)
Explainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextTabularBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposes the ASTRA framework, combining adaptive semantic tree serialization (AdaSTR) and dual-mode tree reasoning (DuTR), for complex table question answering;
ASTRA: An Automated Framework for Strategy Discovery, Retrieval, and Evolution for Jailbreaking LLMs
Xu Liu (East China University of Science and Technology), Huiqun Yu (East China University of Science and Technology)
Adversarial AttackTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation
🎯 What it does: Built an automated, closed-loop jailbreak framework called ASTRA for continuously discovering, retrieving, and evolving attack strategies.
AT²PO: Agentic Turn-based Policy Optimization via Tree Search
Zefang Zong (Tencent Inc), Jie Jiang (Shenzhen MSU-BIT University)
Reinforcement LearningAgentic AITextBenchmark
🎯 What it does: Proposes the AT 2 PO framework, which uniformly addresses the exploration diversity, sparse reward credit assignment, and policy update mismatch issues in agentic RL.
ATGL: An Adaptive-Threshold Global Loss for Document-level Relation Extraction
Huangming Xu (Northeastern University), Jingwei Cheng (Northeastern University)
ClassificationRepresentation LearningData-Centric LearningTransformerSupervised Fine-TuningContrastive LearningText
🎯 What it does: In the document-level relation extraction task, a new global threshold loss function called ATGL is proposed to address the issues of threshold instability and class imbalance.
ATIR: Towards Audio-Text Interleaved Contextual Retrieval
Tong Zhao (Renmin University of China), Zhicheng Dou (Renmin University of China)
RetrievalTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextMultimodalityBenchmarkRetrieval-Augmented GenerationAudio
🎯 What it does: Propose the Audio-Text Interleaved Context Retrieval (ATIR) task, construct the first corresponding benchmark dataset, and design the ATIR-Qwen-3B model to achieve efficient retrieval.
Attention as Selector: Unlocking VLM Attention for Long Document Page Retrieval
Minfeng Zhu (Zhejiang University), Linchao Zhu (Zhejiang University)
RetrievalTransformerMixture of ExpertsVision Language ModelContrastive LearningImageTextMultimodalityRetrieval-Augmented Generation
🎯 What it does: Propose the CAPS framework, which utilizes cross-modal attention from VLM for long document page retrieval, and enhances attention retrieval capability, expert head selection, and adaptive filtering through contrastive learning.
Attention Basin: Why Contextual Position Matters in Large Language Models
Zihao Yi, Ying Shen (Sun Yat Sen University)
RetrievalExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation
🎯 What it does: The study investigates and discovers the 'attention basin' phenomenon in large language models when processing structured contexts, where models tend to focus on the beginning and end positions of document sequences, and based on this, proposes AttnRank, a two-stage, no-training, lightweight re-ranking framework;
Attention Sinks Are Provably Necessary in Softmax Transformers: Evidence from Trigger-Conditional Tasks
Yuval Ran-Milo (Tel Aviv University)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningSequential
🎯 What it does: Investigate and demonstrate that in softmax Transformer, the trigger condition task (outputting the average of all non-BOS vectors before the trigger position, and zero elsewhere) must produce an attention sink to achieve the default no-op behavior, and verify this conclusion through experiments.
Attention Under Attack: Analog Noise Effects and Mechanistic Vulnerabilities in Transformer Models
Mafizur Rahman (Prairie View A&M University), Lijun Qian (Prairie View A&M University)
ClassificationExplainability and InterpretabilityComputational EfficiencyAdversarial AttackTransformerLarge Language ModelContrastive LearningText
🎯 What it does: A fine-grained robustness and mechanism analysis of pre-trained Transformer models on simulated analog in-memory computing (AIMC) hardware reveals the impact of hardware noise on attention projection, hierarchical distribution, and attention head behavior.
Attention Weights as an Indicator: Analyzing and Improving Document Utilization in Retrieval-Augmented Generation
Jing Jin (Peking University), Houfeng Wang (Peking University)
RetrievalExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: This paper studies how retrieval augmented generation (RAG) models utilize retrieved documents internally, proposing to use attention weights for document ranking, placement, and filtering to improve the model's utilization of documents.
Attn-GS: Attention-Guided Context Compression for Efficient Personalized LLMs
Shenglai Zeng (Michigan State University), Hui Liu (Amazon.com)
Recommendation SystemComputational EfficiencyTransformerLarge Language ModelSupervised Fine-TuningTextChain-of-Thought
🎯 What it does: Proposed a context compression framework called Attn-GS based on the LLM attention mechanism for efficient context compression in personalized LLMs.
AttnPO: Attention-Guided Process Supervision for Efficient Reasoning
Shuaiyi Nie (Institute of Information Engineering Chinese Academy of Sciences), Tingwen Liu (Baidu Inc)
Computational EfficiencyAI Code AssistantTransformerLarge Language ModelReinforcement LearningTextBenchmarkChain-of-Thought
🎯 What it does: Propose a low-cost process supervision reinforcement learning framework called ATTNP, which utilizes key attention heads in model attention to allocate credit to reasoning steps, reducing redundant thinking.
Attribution, Citation, and Quotation: A Survey of Evidence-based Text Generation with Large Language Models
Tobias Schreieder (TU Dresden), Michael Färber (TU Dresden)
Explainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringTextReview/Survey PaperBenchmarkRetrieval-Augmented Generation
🎯 What it does: This paper reviews and systematically analyzes 134 papers on evidence-driven text generation with large language models (LLMs), proposing a unified taxonomy, evaluation dimensions, and metrics, as well as organizing datasets and benchmarks.
Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models
Zirui Song (Mohamed bin Zayed University of Artificial Intelligence), Xiuying Chen (Mohamed bin Zayed University of Artificial Intelligence)
Safty and PrivacyAdversarial AttackTransformerLarge Language ModelPrompt EngineeringGenerative Adversarial NetworkContrastive LearningTextMultimodalityBenchmarkAudio
🎯 What it does: Proposed AJailBench — the first open-source jailbreak evaluation benchmark for large audio-language models (LAM);
Audio MultiChallenge: A Multi-Turn Evaluation of Spoken Dialogue Systems on Natural Human Interaction
Advait Gosai (Scale AI), Yunzhong He (Scale AI)
TransformerLarge Language ModelPrompt EngineeringContrastive LearningMultimodalityBenchmarkRetrieval-Augmented GenerationAudio
🎯 What it does: This study constructs the Audio MultiChallenge benchmark, which includes 452 unscripted multi-turn speech dialogues, to systematically evaluate the performance of end-to-end speech dialogue systems in four dimensions: reasoning memory, instruction retention, self-consistency, and voice editing.
Augur: Modeling Covariate Causal Associations in Time Series via Large Language Models
Zhiqing Cui (University of Science and Technology of China), Yang Wang (University of Science and Technology of China)
Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTime SeriesFinance Related
🎯 What it does: Propose the Augur framework, which leverages large language models (LLMs) to automatically learn causal association graphs in multivariate time series, and inputs causal summaries as text prompts to drive a lightweight student model to complete prediction tasks.
Author-in-the-Loop Response Generation and Evaluation: Integrating Author Expertise and Intent in Responses to Peer Review
Qian Ruan (Technical University of Darmstadt), Iryna Gurevych (Technical University of Darmstadt)
GenerationData SynthesisRecommendation SystemTransformerLarge Language ModelPrompt EngineeringTextReview/Survey PaperBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose an author-involving framework for paper response generation and evaluation (REspGen, REspEval) and the corresponding dataset (Re Align 3), achieving input, controllable generation, and iterative optimization of author expertise and intent.
Authorship Attribution in Multilingual Machine-Generated Texts
Lucio La Cava (University of Calabria), Andrea Tagarelli (University of Calabria)
ClassificationAnomaly DetectionData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningTextBenchmark
🎯 What it does: Studying the authorship attribution of machine-generated text in multiple languages, evaluating the cross-lingual applicability of various authorship attribution methods across 18 languages and 7 LLM models.
AutoGraph-R1: End-to-End Reinforcement Learning for Knowledge Graph Construction
Hong Ting Tsang (Hong Kong University of Science and Technology), Yangqiu Song (Hong Kong University of Science and Technology)
RetrievalOptimizationData-Centric LearningTransformerLarge Language ModelReinforcement LearningTextGraphRetrieval-Augmented Generation
🎯 What it does: AutoGraph-R1 directly optimizes the knowledge graph construction process through reinforcement learning, aligning the constructed knowledge graph with the performance of downstream retrieval-augmented generation (RAG) tasks.
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs
Xuanwen Ding (Fudan University), Zhongyu Wei (Fudan University)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelAgentic AIVision Language ModelImageTextMultimodalityBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose an adaptive evaluation framework called AutoJudger for efficiently evaluating multi-modal large language models (MLLM) under limited budget constraints, and implement its agent-based instance A^2-Judger, which can dynamically construct capability dimensions, estimate model capabilities, and adaptively select the most informative test samples during the evaluation process.
Automated Creativity Evaluation of Language Models Across Open-Ended Tasks
Tan Min Sen (Raffles Institution), Alvin Chan (Nanyang Technological University)
Explainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Designed a fully automated, domain-general framework for creative evaluation, using semantic entropy to measure the divergent creativity of LLMs, and assessing their convergent task completion through a retrieval-based multi-agent evaluation framework, conducting large-scale benchmarking across three types of creative tasks;
Automatic and Reliable Evaluation for Academic Caption-to-Figure Generation with LMMs
Guanghui Ye (Hunan University), Zhihua Jiang (Jinan University)
GenerationData SynthesisExplainability and InterpretabilityComputational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringVision Language ModelDiffusion modelContrastive LearningImageTextMultimodalityBenchmark
🎯 What it does: Propose the HE4AFG dataset and the automatic evaluation models HE4AFG-E/R for comprehensive and reliable assessment of academic title-image generation
Automatic Correction of Writing Anomalies in Hausa Texts
Ahmad Mustapha Wali (University of Bucharest), Sergiu Nisioi (University of Bucharest)
Data SynthesisAnomaly DetectionTransformerLarge Language ModelSupervised Fine-TuningText
🎯 What it does: Constructed approximately 400k pairs of noisy-clean Hausa parallel corpora, and fine-tuned multiple transformers on this data for writing anomaly correction.
Autonomous Knowledge Graph Exploration with Adaptive Breadth-Depth Retrieval
Joaquin Polonuer, Marinka Zitnik (Harvard Medical School)
RetrievalFederated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningReinforcement Learning from Human FeedbackGraph Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningAgentic AIPrompt EngineeringTextGraphBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Designed and implemented a tool-use knowledge graph retrieval framework called ARK, which utilizes large language models to achieve adaptive control over the breadth and depth of retrieval through two tools: global retrieval and neighborhood exploration;
AutoRAN: Automated Hijacking of Safety Reasoning in Large Reasoning Models
Jiacheng Liang (Stony Brook University), Ting Wang (Stony Brook University)
Safty and PrivacyExplainability and InterpretabilityAdversarial AttackReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringTextBenchmarkChain-of-Thought
🎯 What it does: Propose the AutoRAN framework, which uses a weak model to simulate inference execution and iteratively optimizes to automatically hijack the secure inference paths of large inference models, bypassing their security defenses.
AutoReproduce: Automatic AI Experiment Reproduction with Paper Lineage
Xuanle Zhao (Tsinghua University), Maosong Sun (Tsinghua University)
Computational EfficiencyData-Centric LearningAI Code AssistantReinforcement Learning from Human FeedbackLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposes AUTOREPRODUCE, a multi-agent framework and paper lineage algorithm for executable reproduction of experimental code in automated experiments;
AutoSchemaKG: Autonomous Knowledge Graph Construction through Dynamic Schema Induction from Web-Scale Corpora
Jiaxin Bai (Hong Kong University of Science and Technology), Yangqiu Song (Hong Kong University of Science and Technology)
Recommendation SystemKnowledge DistillationRepresentation LearningData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextGraphRetrieval-Augmented Generation
🎯 What it does: Propose the AutoSchemaKG framework, enabling the automatic construction of knowledge graphs without the need for predefined schemas;
AutoTaskEval: Towards Domain-Specific and Fine-Grained Evaluation for LLMs
Qingqing Lyu (Zhejiang University), Weiming Lu (Zhejiang University)
Domain AdaptationData-Centric LearningTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose the AUTOTASKEVAL automatic evaluation framework, which can automatically discover tasks, expand context, and generate high-quality evaluation instances with progressive difficulty from unstructured text, achieving fine-grained LLM evaluation with domain adaptation.
AV-Dialog: Spoken Dialogue Models with Audio-Visual Input
Tuochao Chen (University of Washington), Shyamnath Gollakota (University of Washington)
RecognitionObject TrackingGenerationTransformerVision Language ModelAuto EncoderContrastive LearningVideoTextMultimodalityRetrieval-Augmented GenerationAudio
🎯 What it does: Propose the first real-time audio-visual dialogue framework, AV-Dialog, which combines audio and video to achieve target speaker tracking, conversation rhythm prediction, and response generation;
Awakening Dormant Experts:Counterfactual Routing to Mitigate MoE Hallucinations
Wentao Hu (Xi'an Jiaotong University), Xuelong Li (China Telecom)
Explainability and InterpretabilityComputational EfficiencyLarge Language ModelPrompt EngineeringMixture of ExpertsTextRetrieval-Augmented Generation
🎯 What it does: Proposed a training-agnostic inference framework called Counterfactual Routing (CoR), which wakes up dormant experts through causal analysis, thereby improving the factual accuracy of MoE models on long-tail knowledge.
AwarenessBench: Assessing Cognitive Capabilities of Language Models
Xiaojian Li (Tsinghua University), Wei Xu (Tsinghua University)
Large Language ModelTextBenchmark
🎯 What it does: This paper proposes the AwarenessBench benchmark, which systematically evaluates the cognitive functions of language models across four dimensions—metacognition, self-awareness, social cognition, and situational cognition—covering 15 specific cognitive functions, and conducts comparative experiments with 18 mainstream models and three groups of human participants.
Backdoor Collapse: Eliminating Unknown Threats Via Known Backdoor Aggregation In Language Models
Liang Lin (Institute of Information Engineering, Chinese Academy of Sciences), Qingsong Wen (Squirrel Ai Learning)
ClassificationSafty and PrivacyAdversarial AttackTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringText
🎯 What it does: Propose a two-stage defense framework called Locphylax that does not rely on known triggers. It eliminates unknown backdoors by injecting known reverse triggers to aggregate unknown backdoors in the model and then correcting the aggregated backdoor outputs to normal results during the recovery fine-tuning stage.
Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward
Weiyang Guo (Harbin Institute of Technology), Jing Li (Harbin Institute of Technology)
Adversarial AttackReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextChain-of-Thought
🎯 What it does: Proposes a backdoor attack called Asymmetric Chained Backdoor (ACB) targeting the RLVR training framework, which injects a small number (<2%) of poisoned samples with trigger words into the training data and exploits the RLVR's verifiable reward mechanism to induce the model to generate harmful outputs.
BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers?
Fengqing Jiang (University of Washington), Radha Poovendran (University of Washington)
GenerationData SynthesisAnomaly DetectionExplainability and InterpretabilityTransformerLarge Language ModelAgentic AIPrompt EngineeringTextReview/Survey PaperBenchmarkRetrieval-Augmented Generation
🎯 What it does: Construct the BadScientist framework, exploring the interaction between adversarial paper generation and multi-model LLM reviewing.
BanHADEX: Towards Explainable HAte Speech Detection in Bangla Using Human Annotated EXplanation
Faisal Hossain Raquib (Rajshahi University of Engineering and Technology), Akmmahbubur Rahman
ClassificationExplainability and InterpretabilityTransformerSupervised Fine-TuningPrompt EngineeringTextChain-of-Thought
🎯 What it does: Constructed the BANHADEX dataset, collecting 19,203 YouTube Balinese comments, providing binary classification labels, seven subcategories of hate speech, seven victim groups, and human-annotated short explanations.
BaseCal: Unsupervised Confidence Calibration via Base Model Signals
Hexiang Tan (State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences), Xueqi Cheng (University of Chinese Academy of Sciences)
Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelPrompt EngineeringContrastive LearningText
🎯 What it does: Proposes an unsupervised framework called BaseCal based on the signal of a base model to calibrate the confidence of post-training large language models (LLMs), restoring the overconfidence problem of post-training LLMs without modifying model parameters.
BatonVoice: An Operationalist Framework for Enhancing Controllable Speech Synthesis with Linguistic Intelligence from LLMs
Yue Wang (Soochow University), Min Zhang (Soochow University)
GenerationData SynthesisTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextRetrieval-Augmented GenerationAudio
🎯 What it does: Propose the BATONVOICE framework, which decouples user instruction understanding from speech generation. It first uses an LLM to generate quantifiable text-based speech feature plans, and then a specialized BATONTTS model synthesizes audio based on these plans.
Bayesian Social Deduction with Graph-Informed Language Models
Shahab Rahimirad (Purdue University), Joseph Campbell (Purdue University)
Explainability and InterpretabilityReinforcement Learning from Human FeedbackGraph Neural NetworkTransformerLarge Language ModelTextGraph
🎯 What it does: Proposes GRAIL, a social reasoning framework that combines LLMs with graph-structured Bayesian inference, specifically designed for the Avalon game.
BEFT: Bias-Efficient Fine-Tuning of Language Models in Low-Data Regimes
Baichuan Huang (Lund University), Amir Aminifar (Lund University)
Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningText
🎯 What it does: This paper proposes a technique called BEFT, which fine-tunes only the bias term (especially the value bias b_v) of large language models (LLMs), aiming to achieve extremely high parameter efficiency in low-data scenarios.
Behavior Knowledge Merge in Reinforced Agentic Models
Xiangchi Yuan (Georgia Institute Of Technology), Wenke Lee (Georgia Institute Of Technology)
Knowledge DistillationTransformerSupervised Fine-TuningReinforcement LearningAgentic AIText
🎯 What it does: This paper proposes a model merging method for task-specific agent models trained through reinforcement learning (RL), which can merge multiple task-specific models into a single general-purpose agent without significantly losing task-specific capabilities.
Believing without Seeing: Quality Scores for Contextualizing Vision-Language Model Explanations
Keyu He (University of Southern California), Swabha Swayamdipta (University of Southern California)
Explainability and InterpretabilityLarge Language ModelVision Language ModelContrastive LearningImageTextMultimodality
🎯 What it does: Proposes two new metrics for evaluating the interpretability quality of vision-language models (VLMs) — Visual Fidelity and Contrastiveness — and demonstrates that they can help users better determine the reliability of model predictions;
BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks
Nishant Balepur (University of Maryland), Jordan Lee Boyd-Graber
Large Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose BenchMarker, a tool based on an LLM discriminator, for identifying three types of defects in multiple-choice questions (MCQA): leakage, shortcut, and writing errors;
Benchmarking and Enabling Efficient Chinese Medical Retrieval via Asymmetric Encoders
Angqing Jiang (University of Science and Technology of China), Defu Lian (University of Science and Technology of China)
RetrievalComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringContrastive LearningTextBiomedical DataBenchmark
🎯 What it does: This paper proposes the Chinese medical text embedding benchmark CMedTEB and implements efficient retrieval based on the heterogeneous retrieval architecture CARE.
Benchmarking and Learning Real-World Customer Service Dialogue
Tianhong Gao (ByteDance), Huiyu Yu (ByteDance)
Reinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AITextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Designed and implemented an industrial-level customer service dialogue evaluation benchmark, OLABENCH, as well as a multi-stage reinforcement learning training framework, OLAMIND.
Benchmarking Deflection and Hallucination in Large Vision-Language Models
Nicholas Moratelli (University of Modena and Reggio Emilia), Gonzalo Iglesias (Amazon AGI)
TransformerVision Language ModelImageTextMultimodalityBenchmarkRetrieval-Augmented Generation
🎯 What it does: Developed a dynamic and sustainable updating multi-modal retrieval-enhanced question answering benchmark to evaluate model bias and self-denial behavior when facing missing or conflicting knowledge.
Benchmarking Egocentric Clinical Intent Understanding Capability for Medical Multimodal Large Language Models
Shaonan Liu (Shenzhen University), Linlin Shen (Shenzhen University)
TransformerLarge Language ModelPrompt EngineeringVision Language ModelVideoTextMultimodalityBiomedical DataBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose and release the MedGaze-Bench benchmark to evaluate the perspective-subjective, gaze-based clinical intent understanding capabilities of medical multimodal LLMs, and construct a three-dimensional intent framework and a trap QA mechanism.
Benchmarking Fine-Grained Error Detection in Multimodal Reasoning
Chi-Min Chan (Hong Kong University of Science and Technology), Yike Guo (Hong Kong University of Science and Technology)
Explainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextMultimodalityBenchmarkChain-of-Thought
🎯 What it does: Propose the PRMBench-V multimodal error detection benchmark, constructing 907 multimodal reasoning questions and 8,163 samples with 9 fine-grained error labels, and systematically evaluate the error identification capabilities of 16 open-source and closed-source models;
Benchmarking Large Vision-Language Models on CFMME: A Comprehensive Chinese Financial Multimodal Evaluation Dataset
Qian Chen (Qwen DianJin Team, Alibaba Cloud Computing), Chi Zhang (Qwen DianJin Team, Alibaba Cloud Computing)
TransformerPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmarkFinance Related
🎯 What it does: Propose and construct CFMME—a comprehensive evaluation benchmark containing 6,052 Chinese financial multimodal instances, covering 8 image types and 4 core tasks, and systematically evaluate 14 large audio-visual language models (LVLMs) in zero-shot scenarios.
Benchmarking LLM’s Capability in Reasoning over Conflicting Web References
Yizhen Yuan (Tsinghua University), Yunxin Liu (Tsinghua University)
TransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Construct the CONFRAG benchmark dataset and evaluate the reasoning and structured answering capabilities of LLMs under conflicting network references.
Benchmarking Post-Training Quantization of Large Language Models under Microscaling Floating Point Formats
Manyi Zhang (Huawei Technologies), Xianzhi Yu (Huawei Technologies)
Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelVision Language ModelContrastive LearningTextMultimodalityBenchmark
🎯 What it does: This paper systematically evaluates post-training quantization (PTQ) of large-scale language models and multimodal language models in the micro-scale floating-point (MXFP) format, covering 7 PTQ algorithms, 15 evaluation benchmarks, and 3 model families.
Benchmarking Web Agent Safety under E-commerce Deceptive Interfaces
Zijing Shi (University of Technology Sydney), Ling Chen (University of Technology Sydney)
Safty and PrivacyTransformerLarge Language ModelAgentic AIPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmark
🎯 What it does: Propose the WebDecept framework, which injects various deceptive interface patterns into e-commerce websites, and systematically evaluates the security of multi-modal web agents in real deceptive environments.
Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs
Shei Pern Chua (Tsinghua University), Xiaolin Hu (Shanghai Jiao Tong University)
Safty and PrivacyExplainability and InterpretabilityAdversarial AttackTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextRetrieval-Augmented Generation
🎯 What it does: This paper proposes a new multi-round red team attack method called TRIAL, revealing that LLMs are susceptible to being induced to generate harmful outputs during ethical reasoning, and based on this, constructs the ERR defense framework, which can prevent such attacks while maintaining ethical reasoning capabilities.
Beyond "I Don’t Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty
Jingyi Ren (Tsinghua University), Yang Liu (Tsinghua University)
Explainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose the UA-Bench benchmark to evaluate the self-awareness of LLMs in identifying data uncertainty and model uncertainty, and enhance this ability through lightweight reinforcement learning.
Beyond Accuracy: Unveiling Inefficiency Patterns in Tool-Integrated Reasoning
Qisheng Su (University of Science and Technology of China), Feng Zhao (University of Science and Technology of China)
Computational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposes a hardware-aware tool-integrated inference efficiency metric called PTE (Prefill Token Equivalents), unifying internal inference and external tool call costs, and verifying its correlation with actual latency in industrial scenarios.
Beyond Atomic Characters: Glyph-Aware Sub-character Alignment for Low-Resource Multilingual OCR
Mengxiao Zhu (North China University of Technology), Ge Shi (Beijing Institute of Technology)
RecognitionData SynthesisConvolutional Neural NetworkTransformerSupervised Fine-TuningVision Language ModelContrastive LearningImageTextBenchmark
🎯 What it does: Proposed the BASA OCR framework, integrating high-resolution visual encoders, Glyph-Aware Fine-grained Adapter (GAFA) sub-character alignment module, two-stage curriculum learning, and Glyph-Aware Reverse Synthesis data generation technique with zero-cost sub-character labels, and constructed the BASA-Bench benchmark containing 11 low-resource languages and 23 real-world scenarios.
Beyond Chunking: Discourse-Aware Hierarchical Retrieval for Long Document Question Answering
Huiyao Chen (Harbin Institute of Technology), Min Zhang (Harbin Institute of Technology)
RetrievalTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation
🎯 What it does: Propose DISRetrieval, which constructs a hierarchical retrieval framework using discourse structure (RST) to improve the performance of long document question answering.
Beyond Code Pairs: Dialogue-Based Data Generation for LLM Code Translation
Le Chen (Argonne National Laboratory), Chunhua Liao (Lawrence Livermore National Laboratory)
Data SynthesisAI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextSequentialRetrieval-Augmented Generation
🎯 What it does: This paper proposes an automated data generation pipeline based on a dual LLM (Questioner–Solver) dialogue framework for code translation between low-resource programming languages (such as Fortran) and emerging frameworks (such as CUDA), generating multi-round dialogue data verified by unit tests.