ACL 2026 Papers — Page 14
Annual Meeting of the Association for Computational Linguistics · 2296 papers
METER: Evaluating Multi-Level Contextual Causal Reasoning in Large Language Models
Pengfeng Li (Sichuan University), See-Kiong Ng (National University of Singapore)
Explainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringTextMultimodalityBenchmarkChain-of-Thought
🎯 What it does: Proposed and implemented the METER benchmark, which evaluates the performance of large language models in three-layer causal reasoning (causal discovery, intervention, counterfactual) under a unified context, and constructed a multiple-choice dataset with 4,145 samples.
MetFuse: Figurative Fusion between Metonymy and Metaphor
Saptarshi Ghosh (University of Cincinnati), Tianyu Jiang (University of Cincinnati)
GenerationData SynthesisTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBenchmark
🎯 What it does: A framework for converting textual sentences into metonymy, metaphor, and mixed sentences was constructed, and this framework was used to generate the first MetFuse dataset containing mixed expressions of metonymy and metaphor (totaling 4,000 sentences).
METRO: Towards Strategy Induction from Expert Dialogue Transcripts for Non-collaborative Dialogues
Haofu Yang (Beihang University), See-Kiong Ng (National University of Singapore)
Autonomous DrivingOptimizationExplainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Automatically extract and structure strategy actions and planning logic from expert dialogue records using large language models, constructing a 'strategy forest' to guide non-cooperative dialogue agents;
MicroC-KT: Modeling Community Effect via Learning Micro-Environment for Evidence-Grounded Explainable Knowledge Tracing
Zhiyi Duan (Inner Mongolia University), Qi Wang (Jilin University)
Explainability and InterpretabilityKnowledge DistillationRepresentation LearningGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextTabularRetrieval-Augmented Generation
🎯 What it does: Proposes MicroC-KT, a training-free knowledge tracing framework that centers on learning micro-environments. It constructs a learning environment hypergraph, identifies learning communities, generates dual-grained summaries, and retrieves peer evidence from the same community. Finally, it inputs students, questions, and micro-environment evidence as prompts into LLMs to jointly generate predictions and interpretable reports.
Min-k Sampling: Decoupling Truncation from Temperature Scaling via Relative Logit Dynamics
Yuanhao Ding (Henan University), Chongsheng Zhang (Henan University)
GenerationExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose a new decoding sampling method called Mink Sampling, which can dynamically identify 'semantic cliffs' in the logit distribution and truncate the candidate set without relying on temperature parameters.
Mind Reader: Latent User Demand-Guided Content Optimization for Generative Search Engine
Tong Chen (SenseTime Research), Zhaoran Fan (SenseTime Research)
Recommendation SystemOptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextGraphBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose Mind Reader, which optimizes the content visibility of generative search engines through two modules: query decomposition and recombination, and reasoning coverage.
Mind the (DH) Gap! A Contrast in Risky Choices Between Reasoning and Conversational LLMs
Luise Ge (Washington University in St. Louis), Yevgeniy Vorobeychik (Washington University in St. Louis)
Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsTextBenchmarkChain-of-Thought
🎯 What it does: Evaluate the risk decision-making behavior of 20 LLMs under two information presentation modes (explicit description vs implicit history) and whether explanations are required, and compare them with human and economic rational benchmarks.
Mind the Gap in Cultural Alignment: Task-Aware Culture Management for Large Language Models
Binchi Zhang (University of Virginia), Zhengzhang Chen (NEC Laboratories America)
Data SynthesisDomain AdaptationData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextRetrieval-Augmented Generation
🎯 What it does: Propose CultureManager, a task-aware cultural alignment pipeline that generates culture-specific data through search and achieves alignment via multi-cultural adapters and a router.
Mind the Pause: Disfluency-Aware Objective Tuning for Multilingual Speech Correction with LLMs
Deepak Kumar (Indian Institute of Technology Patna), Asif Ekbal (Indian Institute of Technology Patna)
Computational EfficiencyRepresentation LearningData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningTextAudio
🎯 What it does: Propose a multilingual speech disfluency correction pipeline: first use MuRIL for word-level disfluency annotation, then feed the annotated results along with the input speech transcription to instruction-tuned large language models. By utilizing contrastive loss to suppress the generation of disfluent words, fluent text is generated.
MIND Your Reasoning: A Meta-Cognitive Intuitive-Reflective Network for Dual-Reasoning in Multimodal Stance Detection
Bingbing Wang (Harbin Institute of Technology), Min Zhang (Harbin Institute of Technology)
ClassificationExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelVision Language ModelImageTextMultimodalityRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposed the MIND framework, which achieves multimodal stance detection by imitating human dual-process cognition (intuition + reflection);
MIND: From Passive Mimicry to Active Reasoning through Capability-Aware Multi-Perspective CoT Distillation
Jin Cui (Xi'an Jiaotong University), Pengju Ren (Xi'an Jiaotong University)
Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose the MIND framework, which transfers multi-perspective knowledge with chain-of-thought reasoning to small models, enhancing their reasoning capabilities.
Mind’s Eye: A Benchmark of Visual Abstraction, Transformation and Composition for Multimodal LLMs
Rohit Sinha (IIT Hyderabad), Tanuja Ganu (Microsoft Research)
TransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmarkChain-of-Thought
🎯 What it does: Proposes the Mind's Eye benchmark, which evaluates the performance of multimodal large language models on three-dimensional cognitive tasks such as visual abstraction, relational mapping, and spatial transformation using a multiple-choice format;
Miner: Mining Intrinsic Mastery for Data-Efficient RL in Large Reasoning Models
Shuyang Jiang (Fudan University), Yu Wang (Shanghai Jiao Tong University)
Computational EfficiencyData-Centric LearningTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBiomedical DataBenchmark
🎯 What it does: Propose the MINER framework, which generates self-supervised rewards by leveraging the intrinsic uncertainty from positive homogeneous (PH) rollout, significantly improving the data efficiency of RLVR on large-scale reasoning models.
Minimal Free Resolution Guided Adaptive Tree Reasoning
Dezhao Tang (China University of Mining and Technology), Qiuyan Yan (China University of Mining and Technology)
Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose the SyRA framework, which achieves adaptive tree-like reasoning through a knowledge verifier, controller, and reflector based on the minimal free resolution (MFR) theory;
MiniRAG: A Lightweight RAG system with Small Language Models
Tianyu Fan (University of Hong Kong), Chao Huang (University of Hong Kong)
RetrievalComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerPrompt EngineeringContrastive LearningTextGraphBenchmarkRetrieval-Augmented Generation
🎯 What it does: Designed and implemented MiniRAG, a lightweight retrieval-augmented generation system tailored for small language models, employing semantic-aware heterogeneous graph indexing and topology-based lightweight retrieval.
MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Long-tail Knowledge
Jie He (University of Edinburgh), Jeff Z. Pan (University of Edinburgh)
TransformerLarge Language ModelTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This paper proposes MINTQA, a multi-hop question answering benchmark that evaluates the performance of large language models in multi-hop reasoning and retrieval fusion by combining two dimensions: new/old knowledge and popular/unpopular knowledge.
MirageBackdoor: A Stealthy Attack that Induces Think-Well-Answer-Wrong Reasoning
Yizhe Zeng (Chinese Academy of Sciences), Yuling Liu (Chinese Academy of Sciences)
Federated LearningSafty and PrivacyExplainability and InterpretabilityComputational EfficiencyAdversarial AttackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningTextChain-of-Thought
🎯 What it does: Propose MirageBackdoor (MirageBD), a backdoor attack that achieves 'correct thinking but incorrect answer' in chain-of-thought (CoT) reasoning.
MirrorCAPTCHA: Wild CAPTCHA, Wild Distribution, Wild Web-based Platform Meet Multimodal LLM Agents
Xiangyu Wu (Nanjing University of Science and Technology), Jianfeng Lu (Nanjing University of Science and Technology)
TransformerLarge Language ModelSupervised Fine-TuningAgentic AIVision Language ModelImageTextMultimodalityBenchmark
🎯 What it does: MirrorCAPTCHA is an evaluation benchmark based on real-world website CAPTCHA collection, clustering, and random walk estimation of visit frequency, providing two metrics: weighted pass rate and completion degree;
Mirroring Users: Towards Building Preference-aligned User Simulator with User Feedback in Recommendation
Tianjun Wei (Nanyang Technological University), Jie Zhang (Nanyang Technological University)
Recommendation SystemExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringText
🎯 What it does: By constructing the USERMIRRORER framework, the decision-making process of user feedback in recommendation systems is utilized to achieve fine-grained alignment with LLMs, thereby realizing more accurate user simulation.
MirrorQA: Benchmarking Multimodal LLMs on Mirror-Orientation Reasoning
Jingping Liu, Xiaofeng Jia (Bowling Green State University)
TransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmark
🎯 What it does: Proposed the MirrorQA benchmark to evaluate the ability of multimodal large language models to reason about left and right (mirror direction) in mirrored contexts;
MIRTH: Mutual-Information Reasoning with Temporal Hubs for Vision-Language-Action Agents
Hao Sun (Ritsumeikan University), Yen-wei Chen
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningRobotic IntelligenceTransformerReinforcement LearningPrompt EngineeringMixture of ExpertsVision-Language-Action ModelContrastive LearningImageVideoTextMultimodality
🎯 What it does: Proposes a unified Vision-Language-Action framework called MIRTH, addressing the issues of short-term temporal vision, inference gap, and low inference efficiency in single-frame VLA models.
Mitigating Action-Relation Hallucinations in LVLMs via Relation-aware Visual Enhancement
Zhenxin Qin (Tongji University), Wen Shen (Tongji University)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodality
🎯 What it does: Propose a training-free framework that alleviates action relationship hallucinations in LVLMs by locating action-related image regions and enhancing attention.
Mitigating Catastrophic Forgetting in Target Language Adaptation of LLMs via Source-Shielded Updates
Atsuki Yamaguchi (University of Sheffield), Nikolaos Aletras (University of Sheffield)
Domain AdaptationComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningText
🎯 What it does: Propose a method called Source Shielded Updates (SSU), which utilizes source language data to evaluate parameter importance and freezes important parameters column-wise in advance, thereby actively protecting source language knowledge during adaptation using only unlabeled target language text to fine-tune large instruction models, significantly reducing catastrophic forgetting.
Mitigating Context Interference for Reliable and Efficient Search Agents
Boyang Xue (Chinese University of Hong Kong), Aldo Lipani (University College London)
Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Systematically study context interference in multi-round search agents, finding that it mainly comes from the latest retrieved documents, and propose a context refiner based on distillation, which is then embedded into the reinforcement learning training process to improve the reliability and efficiency of the search agent.
Mitigating Hallucinations in Large Vision-Language Models without Performance Degradation
Xingyu Zhu (University of Science and Technology of China), Xiangnan He (University of Science and Technology of China)
Explainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerSupervised Fine-TuningVision Language ModelContrastive LearningImageTextMultimodalityBenchmark
🎯 What it does: Proposes the MPD two-stage framework to eliminate hallucinations in large-scale vision-language models while maintaining generation performance.
Mitigating Legal Hallucinations via Symbolic Constraints and Analogical Precedents
Zixuan Huang (University Of Sydney), Chang Xu (University Of Sydney)
ClassificationRetrievalExplainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: This paper proposes the AALawyer framework, which combines symbolic constraint retrieval (SCR) and analogy precedent retrieval (APR), using the legal syllogism theory to reduce hallucinations, semantic drift, and citation inadaptability issues in large language models during legal reasoning.
Mitigating Lost in Multi-turn Conversation via Curriculum RL with Verifiable Accuracy and Abstention Rewards
Ming Li (University of Maryland), Bing Yin (Amazon)
OptimizationAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringText
🎯 What it does: Designed and trained a reinforcement learning-based framework called RLAAR, encouraging LLMs to both answer correctly and determine when to give up in multi-turn conversations
Mitigating Over-Refusal in Aligned Large Language Models via Inference-Time Activation Energy
Eric Hanchen Jiang (UCLA), Xinfeng Li (NTU)
Safty and PrivacyExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelContrastive LearningTextBenchmark
🎯 What it does: Designed an inference-time activation energy field steering (Energy Landscape Steering, ELS) without fine-tuning, which assigns energy to LLM hidden states by training a lightweight energy model and dynamically adjusts hidden states in real-time during inference through gradient descent, significantly reducing the over-rejection rate while maintaining safety and general capabilities.
Mitigating Safety Context Amnesia in Multimodal Reasoning Models via Intent-Guided Safety Reasoning
Xiyao Dong (Huazhong University of Science and Technology), Kun He (Huazhong University of Science and Technology)
Safty and PrivacyExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmarkRetrieval-Augmented Generation
🎯 What it does: This paper proposes an intent-guided safe reasoning (IGSR) framework during the inference phase, aiming to alleviate the safety context forgetting (SCA) problem in multi-modal large reasoning models.
Mitigating Selection Bias in Large Language Models via Permutation-Aware GRPO
Jinquan Zheng (East China Normal University), Guoxiu He (East China Normal University)
OptimizationData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextBenchmark
🎯 What it does: Proposes Permutation-Aware Group Relative Policy Optimization (PA-GRPO), which debiases the order and label bias of options in multiple-choice questions and LLM-as-a-Judge tasks through reinforcement learning.
Mitigating Spurious Correlations in Text Classification Using Latent Space Geometry
Jiasen Gao, Yajun Du (Xihua University)
ClassificationTransformerPrompt EngineeringContrastive LearningText
🎯 What it does: Propose a prototype-guided bias elimination framework called PRIFT based on natural language prompts, which enhances the robustness of text classification on out-of-distribution and minority groups by constructing geometric anchors in the latent space and adopting central projection to remove instance-specific confounding bias.
Mitigating Structural Knowledge Collapse in Domain-Specific LLMs via Morpheme-Aware KV-Aggregation
Yuxuan Si (Zhejiang University), Fei Wu (Zhejiang University)
Drug DiscoveryTransformerLarge Language ModelSupervised Fine-TuningTextBiomedical DataReview/Survey PaperBenchmark
🎯 What it does: This paper investigates the structural knowledge collapse (SKC) caused by excessive subword tokenization in professional domains such as medicine and law in large language models, and proposes a lightweight Morpheme-aware KV-aggregation Attention (MorphKA) adapter. It dynamically merges fragmented tokens through input-layer morpheme aggregation and deep context-aware KV aggregation, improving performance on domain tasks and reducing catastrophic forgetting.
Mitigating Tokenization-Induced Distance Distortion in Long-Context Multilingual Machine Translation
Khotso Selialia (University of Massachusetts Amherst), Fatima M. Anwar (University of Massachusetts Amherst)
TransformerLarge Language ModelSupervised Fine-TuningTextMultimodalityBenchmark
🎯 What it does: Proposed a dynamic position encoding method called DCARPE based on input tokenization statistics, significantly improving the robustness of multilingual long-context translation.
MixKVQ: Query-Aware Mixed-Precision KV Cache Quantization for Long-Context Reasoning
Tao Zhang (South China University Of Technology), Cen Chen (Beihang University)
Computational EfficiencyKnowledge DistillationTransformerLarge Language ModelPrompt EngineeringText
🎯 What it does: Propose MixKVQ, a low-bit quantization method for the KV cache of large language models, aiming to significantly reduce memory usage while maintaining the accuracy of long-text inference.
Mixture-of-Experts with Intermediate CTC Supervision for Accented Speech Recognition
Wonjun Lee (POSTECH), Gary Lee
RecognitionTransformerSupervised Fine-TuningMixture of ExpertsAudio
🎯 What it does: In this paper, the authors propose a speech recognition model called MOE-CTC based on Mixture-of-Experts (MoE), which utilizes CTC supervision in the intermediate layer to guide the training of expert networks, and achieves transition from routing with dialect labels to unlabeled general routing through a two-stage training approach;
Mixture-of-Minds: Multi-Agent Reinforcement Learning for Table Understanding
Yuhang Zhou (Meta AI), Lizhu Zhang (Meta AI)
Autonomous DrivingOptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyRepresentation LearningDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackNeural Architecture SearchProtein Structure PredictionTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringMixture of ExpertsTabularBenchmarkFinance Related
🎯 What it does: Propose a multi-agent framework called MIXTURE-OF-MINDS, which decomposes table reasoning into three specialized roles: planning, encoding, and answering, and enhances itself through reinforcement learning;
MM-JudgeBias: A Benchmark for Evaluating Compositional Biases in MLLM-as-a-Judge
Sua Lee (Seoul National University), Jinbae Im (NAVER Cloud AI)
Explainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmark
🎯 What it does: Propose and quantify the compositional bias of MLLM-as-a-Judge, and construct the MM-JudgeBias benchmark for evaluation.
MM-PoisonRAG: Disrupting Multimodal RAG with Local and Global Knowledge Poisoning Attacks
Hyeonjeong Ha (University of Illinois Urbana Champaign), Heng Ji (University of Illinois Urbana Champaign)
RetrievalRepresentation LearningAdversarial AttackTransformerPrompt EngineeringDiffusion modelContrastive LearningImageTextMultimodalityRetrieval-Augmented Generation
🎯 What it does: This paper systematically studies the vulnerability of multimodal retrieval-augmented generation (Multimodal RAG) systems under knowledge poisoning attacks, proposing and evaluating two attack strategies: Local Poisoning Attack (LPA) and Global Poisoning Attack (GPA).
MM-StanceDet: Retrieval-Augmented Multi-modal Multi-agent Stance Detection
Weihai Lu, Huan He (Brown University)
ClassificationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelAgentic AIPrompt EngineeringImageTextMultimodalityRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposed the MM-StanceDet multi-agent framework, achieving more robust multi-modal stance detection through four stages: retrieval enhancement, specialized multi-modal analysis, debate, and self-reflection.
MMAC: A Multilingual, Multimodal Alignment Framework for Cultural Grounding Evaluation
Weihua Zheng (Agency for Science, Technology and Research), Nancy F. Chen (Singapore University of Technology and Design)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-ThoughtAudio
🎯 What it does: Proposed the MMAC framework and the MMAC-bench dataset to systematically evaluate the cultural cognition and reasoning capabilities of large language models in multilingual, multimodal (text, image, voice) environments.
MMCLIP: Cross-Modal Attention Masked Modelling for Medical Language-Image Pre-Training
Biao Wu (University of Technology Sydney), Qi Wu (Adelaide University)
ClassificationRepresentation LearningTransformerVision Language ModelAuto EncoderContrastive LearningImageTextBiomedical DataElectronic Health Records
🎯 What it does: Propose the MMCLIP framework, combining cross-modal attention masking techniques to achieve contrastive learning between medical images and text, image reconstruction, and entity-driven text masking.
MMErroR: A Benchmark for Erroneous Reasoning in Vision-Language Models
Yang Shi (Guangdong University of Technology), Zhiqi Huang (Peking University)
TransformerLarge Language ModelVision Language ModelMultimodalityBenchmarkChain-of-Thought
🎯 What it does: Proposes the MMErroR benchmark for evaluating error detection and error type classification in the reasoning chains of vision-language models.
MMSciCode: Real-world Evaluation of Multilingual Multi-Discipline Scientific Research Coding
Xue Xia (HKUST), Yilun Zhao (Yale University)
AI Code AssistantLarge Language ModelPrompt EngineeringTextBenchmarkChain-of-Thought
🎯 What it does: Proposed the MMSciCode multilingual, multidisciplinary scientific code generation benchmark, and evaluated the performance of 23 state-of-the-art models and 2 code agents within it.
MMSearch-R1: Incentivizing LMMs to Search
Jinming Wu (ByteDance), Ziwei Liu (S-Lab)
RetrievalReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningVision Language ModelImageTextMultimodalityRetrieval-Augmented Generation
🎯 What it does: This paper proposes MMSearch-R1, an end-to-end framework based on reinforcement learning, enabling large multi-modal models (LMMs) to perform multi-round image and text retrieval on demand in real network environments and answer visual question answering (VQA) questions accordingly.
MMTutorBench: The First Multimodal Benchmark for AI Math Tutoring
Tengchao Yang (Tongji University), Meng Jiang (University of Notre Dame)
TransformerLarge Language ModelPrompt EngineeringVideoTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposes MMTutorBench, the first benchmark for multi-modal math tutoring, which evaluates the ability of LLMs to identify key insights, formulate operations, and execute steps.
Mnemis: Dual-Route Retrieval on Hierarchical Graphs for Long-Term LLM Memory
Zihao Tang (Microsoft), Qi Zhang (Microsoft)
RetrievalComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringTextGraphBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposed a long-term memory framework called Mnemis, which integrates traditional similarity retrieval (System-1) with global hierarchical retrieval (System-2) routing;
MoA: Heterogeneous Mixture of Adapters for Parameter-Efficient Fine-Tuning of Large Language Models
Jie Cao (Zhejiang University), Yueting Zhuang (Zhejiang University)
Computational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsText
🎯 What it does: Designed a heterogeneous hybrid adapter (MoA) for parameter-efficient fine-tuning of large language models.
MOA: Multi-Objective Alignment for Role-Playing Agents
Chonghua Liao (Tsinghua University), Yongbin Li (Tongyi Lab)
OptimizationTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose the MOA framework, which uses multi-objective reinforcement learning to perform fine-grained, multi-dimensional reward alignment and optimization for role-playing agents;
Mobile-R1: Towards Interactive Capability for VLM-Based Mobile Agent via Systematic Training
Jihao Gu (Alibaba Group), Bo Zheng (Alibaba Group)
Computational EfficiencyRobotic IntelligenceTransformerSupervised Fine-TuningReinforcement LearningPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmarkRetrieval-Augmented Generation
🎯 What it does: Enhancing the interaction and self-correction capabilities of VLM mobile agents through a three-stage hierarchical training approach
MobileWorld: Benchmarking Autonomous Mobile Agents in Agent-User Interactive and MCP-Augmented Environments
Quyu Kong (Tongyi Lab, Alibaba Group), Yue Wang (Tongyi Lab, Alibaba Group)
Autonomous DrivingRobotic IntelligenceReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringVision-Language-Action ModelTextMultimodalityBenchmarkRetrieval-Augmented Generation
🎯 What it does: This work proposes the MobileWorld benchmark, aiming to provide a more challenging evaluation environment for mobile GUI agents. It covers 201 tasks across 20 applications, with task types including GUI-only operations, agent-user interactions, and MCP (Model Context Protocol) enhanced tasks. The average number of steps per task is approximately 27.8, with 62.2% of tasks spanning across applications.
MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings
Haonan Chen (Renmin University of China), Zhicheng Dou (Microsoft Corporation)
RetrievalRepresentation LearningTransformerVision Language ModelAuto EncoderContrastive LearningImageTextMultimodality
🎯 What it does: This paper proposes a two-stage framework called MoCa, which converts a pre-trained VLM into a bidirectional multi-modal embedding model. It first undergoes modality-aware continual pre-training, followed by heterogeneous contrastive fine-tuning.
Mock Worlds, Real Skills: Building Small Agentic Language Models with Synthetic Tasks, Simulated Environments, and Rubric-Based Rewards
Yuanjie Lyu (University Of Science And Technology Of China), Tong Xu (Xi'an Jiaotong University)
Autonomous DrivingOptimizationData-Centric LearningRobotic IntelligenceReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringWorld ModelTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Construct synthetic tools using task and simulation environments to train the agentic capabilities of small LLM agents.
Modal Dependency Parsing as Structured Prediction over Source-Cue Scope
Jayeol Chun (Brandeis University), Nianwen Xue (Brandeis University)
RecognitionRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextMultimodality
🎯 What it does: Propose a structured prediction framework based on large language models, transforming modal dependency parsing (MDP) into the identification of source-cue-scope triplets and edge prediction at the event level;
Model Internal Sleuthing: Finding Lexical Identity and Inflectional Features in Modern Language Models
Michael Li (Carnegie Mellon University), Nishant Subramani (Carnegie Mellon University)
RecognitionExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringContrastive LearningText
🎯 What it does: Systematically investigate how morphological and morphological information are encoded in 25 modern Transformer models.
Model-Based Imaginative Planning for Embodied Agents
Junru Song (Shanghai Jiao Tong University), Wen Yao (Intelligent Game and Decision Laboratory)
Robotic IntelligenceTransformerLarge Language ModelReinforcement LearningPrompt EngineeringVision Language ModelWorld ModelImageTextMultimodality
🎯 What it does: The paper proposes the IMPLEMENT framework, which uses a frozen LLM combined with a lightweight interpretable world model for imaginative planning, to achieve vision-driven embodied agents.
Modeling Annotator Disagreement with Demographic-Aware Experts and Synthetic Perspectives
Yinuo Xu (University of Michigan), David Jurgens (University of Michigan)
Explainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringMixture of ExpertsText
🎯 What it does: Propose the DEM-MOE model, which models annotator disagreement in subjective tasks through a mixture-of-experts routing based on annotator demographics, and enhances scarce perspectives by generating synthetic annotations via LLM zero-shot prompting, ultimately improving diversity perspective modeling by fusing real and synthetic data.
Modeling Human-Like Cognition for Stance Detection: Integrating Intuitive Judgment and Analytical Reasoning
Zhaodan Zhang (University of Chinese Academy of Sciences), Xueqi Cheng (University of Chinese Academy of Sciences)
ClassificationExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextBenchmarkChain-of-Thought
🎯 What it does: A stance detection framework CDSD based on Kahneman's dual-process theory was designed, combining fast intuitive judgment with deep reasoning to achieve more reliable identification of textual stance.
Modeling LLM Unlearning as an Asymmetric Two-Task Learning Problem
Zeguan Xiao (Shanghai University of Finance and Economics), Guanhua Chen (Southern University of Science and Technology)
Federated LearningSafty and PrivacyComputational EfficiencyTransformerLarge Language ModelContrastive LearningTextBenchmark
🎯 What it does: View LLM unlearning as asymmetric multi-task learning, and propose a gradient synthesis framework prioritizing retention.
ModeX: Evaluator-Free Best-of-N Selection for Open-Ended Generation
Hyeong Kyu Choi (University of Wisconsin-Madison), Sharon Li (University of Wisconsin-Madison)
GenerationAI Code AssistantTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsText
🎯 What it does: Propose ModeX, an estimator-free Best-of-N selection framework for open-ended text generation;
MoEC: A Memory-Routed Mixture-of-Experts Controller for Adaptive Minecraft Control
Hui Wu (Aerospace Information Research Institute Chinese Academy of Sciences), Emad Barsoum (Advanced Micro Devices Inc)
Robotic IntelligenceTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsTextMultimodalityBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose MoEC, a hybrid expert controller based on memory routing, which replaces a single control strategy to support executing multiple subgoals in the Minecraft environment.
MolMem: Memory-Augmented Agentic Reinforcement Learning for Sample-Efficient Molecular Optimization
Ziqing Wang (Northwestern University), Kaize Ding (Northwestern University)
OptimizationDrug DiscoveryReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningAgentic AIGraphTabularBiomedical DataRetrieval-Augmented Generation
🎯 What it does: This paper proposes a molecular optimization framework called MolMem based on multi-round agent-based reinforcement learning, which achieves high sample efficiency in optimizing molecular properties under a limited oracle call budget by utilizing a dual memory system.
MONETA: Multimodal Industry Classification through Geographic Information with Multi Agent Systems
Arda Yüksel (Technical University of Darmstadt), Ivan Habernal (Ruhr University Bochum)
ClassificationTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposed the MONETA multimodal industry classification benchmark, which uses text (websites, Wikipedia, Wikidata) and geospatial information (OpenStreetMap, satellite images) to perform NACE classification on 1,000 European companies, aiming to replace manual expert verification;
Monotonic Scaffolding as a Diagnostic Lens for Legal Reasoning in LLMs
Pedro Calais (UFMG), Wagner Meira Jr. (UFMG)
Explainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextChain-of-Thought
🎯 What it does: Proposed a process-aware evaluation framework based on monotonic scaffolding, which gradually injects expert-annotated legal information through the FIRAC structure, tracking the reasoning trajectory of LLMs at different levels of guidance;
More Aligned, Less Diverse? Analyzing the Grammar and Lexicon of Two Generations of LLMs
Adrián Gude (Universidade da Coruña), Olga Zamaraeva (Universidade da Coruña)
GenerationTransformerLarge Language ModelPrompt EngineeringTextReview/Survey Paper
🎯 What it does: This paper compares the differences in syntactic and lexical diversity between two generations of LLMs and human news writing.
More Than Meets the Eye: Measuring the Semiotic Gap in Vision-Language Models via Semantic Anchorage
Wei He (University of Exeter)
Explainability and InterpretabilityRepresentation LearningPrompt EngineeringVision Language ModelDiffusion modelContrastive LearningImageTextMultimodalityBenchmark
🎯 What it does: Propose the DIVA benchmark, which evaluates the understanding of metaphorical noun phrases by visual language models by abstracting high-fidelity images into iconized (low-detail) images; meanwhile, define the semantic alignment gap ∆ and the directed literal bias b to measure the model's preference for literal versus metaphorical meanings.
More Thinking, Less Talking: Internalizing Deliberative Safety into LLM Parameters
Guan Wang (Institute of Information Engineering Chinese Academy of Sciences), Songlin Hu (Institute of Information Engineering Chinese Academy of Sciences)
Safty and PrivacyComputational EfficiencyTransformerLarge Language ModelText
🎯 What it does: Proposes the HIAR framework, which internalizes the SCoT security reasoning process into the forward propagation of LLMs, eliminating publicly visible reasoning traces.
MoRI: Learning Motivation-Grounded Reasoning for Scientific Ideation in Large Language Models
Chenyang Gu (East China Normal University), Guoxiu He (East China Normal University)
Federated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningMeta LearningDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackNeural Architecture SearchProtein Structure PredictionTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AIPrompt EngineeringMixture of ExpertsTextReview/Survey PaperBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposes the MoRI framework, enabling large language models to automatically generate scientific ideas by learning the reasoning process from research motivations to methods.
MORPHOGEN: A Multilingual Benchmark for Evaluating Gender-Aware Morphological Generation
Mehul Agarwal (Indraprastha Institute of Information Technology Delhi), Anubha Gupta (Indraprastha Institute of Information Technology Delhi)
GenerationData SynthesisTransformerLarge Language ModelPrompt EngineeringTextMultimodalityBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposed and implemented the MORPHOGEN benchmark to evaluate the capability of multilingual LLMs in gender-aware morphological generation.
MPBoCo: Multimodal Prompt-based Boundary-enhanced Continual Framework for Joint Entity and Relation Extraction
Guanglu Sun (Harbin University of Science and Technology), Ming Liu (Harbin Institute of Technology)
Federated LearningRepresentation LearningData-Centric LearningTransformerPrompt EngineeringVision Language ModelContrastive LearningTextMultimodality
🎯 What it does: Propose the MPBoCo framework to achieve continuous multimodal entity relationship joint extraction
MPR-GUI: Benchmarking and Enhancing Multilingual Perception and Reasoning in GUI Agents
Ruihan Chen (Harbin Institute of Technology), Bing Qin (Huawei Technologies Co., Ltd)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningData-Centric LearningTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBenchmarkRetrieval-Augmented Generation
🎯 What it does: This paper proposes the MPR-GUI-Bench, a multilingual fine-grained perception and reasoning (P&R) GUI benchmark, and designs a training-free cross-lingual intervention method, GUI-XLI, to narrow the performance gap between English and non-English languages.
MQM Re-Annotation: A Technique for Collaborative Evaluation of Machine Translation
Parker Riley (Google), Markus Freitag (Google)
Data-Centric LearningPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposed and experimentally tested the MQM re-annotation method, allowing reviewers to delete, modify, or add errors based on existing error annotations, thereby improving fine-grained quality assessment in machine translation evaluation;
MSEarth: A Multimodal Benchmark for Earth Science Phenomenon Discovery with MLLMs
Xiangyu Zhao (Hong Kong Polytechnic University), Xiao-Ming Wu (Hong Kong Polytechnic University)
TransformerLarge Language ModelAgentic AIPrompt EngineeringImageTextMultimodalityBenchmarkRetrieval-Augmented Generation
🎯 What it does: Constructed a multi-modal benchmark called MSEarth for graduate-level Earth science, including scientific chart captions, open-ended questions, and multiple-choice questions;
MSMO-ABSA: Multi-Scale and Multi-Objective Optimization for Cross-Lingual Aspect-Based Sentiment Analysis
Chengyan Wu (South China Normal University), Liu Xiaoyong
Domain AdaptationOptimizationKnowledge DistillationRepresentation LearningTransformerSupervised Fine-TuningContrastive LearningText
🎯 What it does: Propose the MSMO framework, combining sentence-level adversarial training, aspect-level consistency training, and multi-objective optimization, to achieve cross-lingual ABSA feature alignment and fine-grained alignment, and then perform knowledge distillation based on this.
MT^{3}: A Synergistic Multi-Task RL Framework for Specializing MLLMs in Text Image Machine Translation
Zhaopeng Feng (Zhejiang University), Zuozhu Liu (Zhejiang University)
Image TranslationTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision Language ModelImageTextMultimodalityBenchmark
🎯 What it does: Propose the MT3 framework, which utilizes multi-task reinforcement learning to specialize multimodal large language models (MLLMs) into end-to-end text-image machine translation (TIMT) expert models.
MTA: Multi-Granular Trajectory Alignment for Large Language Model Distillation
Pham Khanh Chi (Hanoi University of Science and Technology), Thanh Hong Nguyen
Knowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningText
🎯 What it does: Propose a multi-granularity trajectory alignment (MTA) framework for aligning hierarchical representations of teacher and student models along depth-evolving trajectories during knowledge distillation.
MTA:A Merge-then-Adapt Framework for Personalized Large Language Models
Xiaopeng Li (City University of Hong Kong), Xiangyu Zhao (City University of Hong Kong)
Federated LearningComputational EfficiencyTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposed a Merge-then-Adapt framework (MTA), which achieves scalable and efficient fine-tuning of personalized large language models by constructing a Meta-LoRA bank, dynamically fusing multiple baseline LoRAs, and stacking an extremely low-rank LoRA on top of them.
MTAVG-Bench: A Diagnostic Benchmark for Multi-Talker Dialogue-Centric Audio-Video Generation
Yanghao Zhou, Yousheng Feng (Inkeverse Group Limited)
GenerationData SynthesisTransformerLarge Language ModelPrompt EngineeringVision Language ModelDiffusion modelVideoTextMultimodalityBenchmarkRetrieval-Augmented GenerationAudio
🎯 What it does: This paper proposes MTAVG-Bench, constructing a diagnostic benchmark for multi-speaker dialogue-centric audio-visual generation, and achieving fine-grained failure diagnosis through a four-level evaluation framework and human-annotated question-answer pairs.
MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation
Xiaoyuan Li (University of Science and Technology of China), Dayiheng Liu (Alibaba Group)
Large Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposed MTR-Bench, a multi-turn reasoning benchmark covering 4 categories, 40 tasks, and 3600 instances, and built a fully automated generation-monitoring-evaluation framework;
MTR-Suite: A Framework for Evaluating and Synthesizing Conversational Retrieval Benchmarks
Junhao Ruan (Northeastern University), JingBo Zhu
Data SynthesisRetrievalTransformerLarge Language ModelAgentic AIPrompt EngineeringContrastive LearningTextBenchmarkFinance RelatedRetrieval-Augmented Generation
🎯 What it does: Propose the MTR-Suite framework, integrating evaluation (MTR-EVAL), multi-agent synthesis (MTR-PIPELINE), and a new dialogue retrieval benchmark (MTR-BENCH).
MTRouter: Cost-Aware Multi-Turn LLM Routing with History–Model Joint Embeddings
Yiqun Zhang (Northeastern University), Shuyue Hu (Shanghai Artificial Intelligence Laboratory)
Computational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningMixture of ExpertsContrastive LearningTextSequential
🎯 What it does: Propose MTRouter, a routing framework that dynamically selects models in multi-turn interactions based on joint embeddings of history and model;
MTSQL-R1: Towards Long-Horizon Multi-Turn Text-to-SQL via Agentic Training
Taicheng Guo (University of Notre Dame), Chandan K. Reddy (Amazon)
Explainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringTextTabularBenchmarkChain-of-Thought
🎯 What it does: Proposed the MTSQL-R1 framework, modeling multi-turn Text-to-SQL as a Markov Decision Process, supporting agent-based execution, verification, and self-correction;
Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation
Jiaju Chen (Northeastern University), Dakuo Wang (Northeastern University)
Explainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextRetrieval-Augmented Generation
🎯 What it does: Propose the MAJ-EVAL framework, which automatically extracts stakeholder dimensions from literature to generate LLM roles, and allows multiple LLM agents within the same group to debate, ultimately aggregating multi-dimensional scores to simulate real human evaluation;
Multi-component Causal Tracing in Large Language Models
Zirui Yan (Rensselaer Polytechnic Institute), Ali Tajer (Rensselaer Polytechnic Institute)
Explainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringContrastive LearningText
🎯 What it does: Proposes a unified framework for causal tracing of multiple components in LLMs (attention heads, MLP neurons) to identify the subset most influential on the metrics.
Multi-Granularity Semantic Revision for Large Language Model Distillation
Xiaoyu Liu (University of Science and Technology of China), Yunhe Wang (Huawei)
Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextRetrieval-Augmented Generation
🎯 What it does: Proposes a multi-granularity semantic revision method to improve the knowledge distillation process of large language models.
Multi-Task Representation Alignment on Language Understanding: A Mutual Information Perspective
Dou Hu (Communication University of China), Yuan Zhang (Communication University of China)
ClassificationRepresentation LearningTransformerContrastive LearningTextBenchmark
🎯 What it does: Propose a Multi-Task Representation Alignment (MTRA) framework, which reduces task interference and enhances the task relevance of shared representations through mutual information maximization and self-alignment dual mechanisms.
Multi-View Attention Multiple-Instance Learning Enhanced by LLM Reasoning for Cognitive Distortion Detection
Jun Seo Kim (Gachon University), Hye Hyeon Kim (Yonsei University)
ClassificationExplainability and InterpretabilityTransformerLarge Language ModelText
🎯 What it does: Propose a method that splits each sentence into three parts: emotion, logic, and behavior (ELB), then uses a large language model (LLM) to generate multiple instances of cognitive distortions, and finally classifies them using a multi-instance learning (MIL) framework with multi-perspective gated attention.
MultiFinBen: Benchmarking Large Language Models for Multilingual and Multimodal Financial Application
Xueqing Peng (Fin AI), Qianqian Xie (Fin AI)
TransformerLarge Language ModelPrompt EngineeringImageTextMultimodalityBenchmarkFinance RelatedRetrieval-Augmented GenerationAudio
🎯 What it does: Proposed and implemented MULTIFINBEN, a multilingual and multimodal financial evaluation benchmark, covering three modalities (text, vision, audio) and five languages, and introduced two new tasks: multilingual financial question answering and OCR.
Multilingual and Cross-Lingual Citation Needed Detection on Wikipedia for Lower-Resource Languages
Gerrit Quaremba (King's College London), Elena Simperl (King's College London)
ClassificationKnowledge DistillationTransformerLarge Language ModelSupervised Fine-TuningTextBenchmark
🎯 What it does: Constructed a cross-lingual 'Need to Cite' detection dataset (MCN), and trained and evaluated a small decoder model on 18 languages with different resource levels.
Multilingual Idioms in Sentences and Conversations Across High-, Medium-, and Low-Resource Languages
Saeed Almheiri (Mohamed bin Zayed University of Artificial Intelligence), Fajri Koto (Mohamed bin Zayed University of Artificial Intelligence)
Large Language ModelTextBenchmark
🎯 What it does: Constructed a multilingual idiom understanding benchmark called MIDI, covering 18 high, medium, and low-resource languages, providing metaphorical and literal usages in both sentence and dialogue contexts, along with multiple-choice questions to evaluate the idiom comprehension ability of large language models.
Multilingual Language Models Encode Script Over Linguistic Structure
Aastha A K Verma (Indian Institute of Technology Delhi), Tanmoy Chakraborty (Indian Institute of Technology Delhi)
Explainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelPrompt EngineeringAuto EncoderContrastive LearningTextMultimodality
🎯 What it does: Systematically investigate the internal structure of language-related units in multilingual language models using LAPE, SAE-LAPE, linear probing, and causal intervention methods, exploring their sensitivity to orthography, syntax, word order, and linguistic features.
Multimodal Large Language Models for Multi-Subject In-Context Image Generation
Yucheng Zhou (University of Macau), Jianbing Shen (University of Macau)
GenerationData SynthesisTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelDiffusion modelContrastive LearningImageTextMultimodalityBenchmarkChain-of-Thought
🎯 What it does: Propose MUSIC, a multimodal large language model specifically designed for image generation in multi-agent contexts;
Multimodal Safety Evaluation in Generative Agent Social Simulations
Alhim Adonai Vera Gonzalez (University of Cincinnati), Bernard Ghanem
Safty and PrivacyTransformerLarge Language ModelAgentic AIPrompt EngineeringVision Language ModelImageTextMultimodalityRetrieval-Augmented Generation
🎯 What it does: A reproducible multimodal safety evaluation framework was constructed, and generative agents were used in social simulation environments to detect and correct unsafe plans, analyzing safety improvements and social dynamics.
MulVul: Retrieval-augmented Multi-Agent Code Vulnerability Detection via Cross-Model Prompt Evolution
Zihan Wu (City University of Hong Kong), Xiaohua Jia (City University of Hong Kong)
Anomaly DetectionExplainability and InterpretabilityComputational EfficiencyRepresentation LearningAI Code AssistantTransformerLarge Language ModelAgentic AIPrompt EngineeringContrastive LearningTextRetrieval-Augmented Generation
🎯 What it does: A multi-agent, retrieval-enhanced vulnerability detection framework called MulVul was constructed, which employs a coarse-to-fine router and specialized detectors to perform multi-class vulnerability detection on code.
MUR: Momentum Uncertainty guided Reasoning for Large Language Models
Hang Yan (Xi'an Jiaotong University), Jun Liu (Xi'an Jiaotong University)
Explainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextBenchmarkChain-of-Thought
🎯 What it does: Proposed an untrained momentum uncertainty guided reasoning (MUR) framework, which dynamically allocates additional computational resources only to key reasoning steps during inference, thereby reducing excessive reasoning.
MuSe: Multi-Stage Graph Reasoning via Vision-Language Models
Guanyu Wang (Peking University), Weiping Li (Peking University)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerSupervised Fine-TuningReinforcement LearningVision Language ModelImageTextGraph
🎯 What it does: The paper proposes a multi-stage graph reasoning framework called MuSe, based on a vision-language model, which performs reasoning by progressively sampling subgraphs and visualizing them.
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding
Fuwen Luo (Tsinghua University), Yang Liu (Tsinghua University)
Computational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringVision Language ModelVideoTextMultimodality
🎯 What it does: Proposed a multi-segment temporal alignment method called MUSEG based on reinforcement learning, enhancing the temporal reasoning ability of multi-modal large language models.
Musical Score Understanding Benchmark: Evaluating Large Language Models’ Comprehension of Complete Musical Scores
Congren Dai (Central Conservatory of Music), Maosong Sun (Tsinghua University)
TransformerLarge Language ModelSupervised Fine-TuningVision Language ModelImageTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposes a music score understanding benchmark called MSU-Bench, specifically designed to evaluate the reasoning capabilities of large language models and audio-visual models on complete musical scores.
MUTANT: A Recipe for Multilingual Tokenizer Design
Souvik Rana (Krutrim AI), Shubham Agarwal (Krutrim AI)
Computational EfficiencyData-Centric LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextMultimodalityBenchmark
🎯 What it does: Designed a multilingual tokenizer training process called MUTANT, and constructed MUTANT-Indic tailored for Indian languages.
MVP: Enhancing Video Large Language Models via Self-supervised Masked Video Prediction
Xiaokun Sun (University of Science and Technology of China), Linli Xu (University of Science and Technology of China)
Data SynthesisExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelReinforcement LearningPrompt EngineeringVision Language ModelAuto EncoderContrastive LearningVideoTextMultimodalityBenchmark
🎯 What it does: By introducing the Mask Video Prediction (MVP) task during the post-training phase of video LLMs, the model's understanding of video temporal and causal relationships is enhanced.
N-GLARE: An Non-Generative Latent Representation-Efficient LLM Safety Evaluator
Zheyu Lin (University of California Riverside), Yao Guan (Fudan University)
Safty and PrivacyExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelAuto EncoderContrastive LearningText
🎯 What it does: Designed and validated a non-generative, latent representation-based LLM security evaluation framework called N-GLARE, which evaluates model safety by leveraging the geometric separation degree of hidden layer trajectory.