CodeAutonomous DrivingFederated LearningComputational EfficiencyRepresentation LearningRobotic IntelligenceReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIVision Language ModelVision-Language-Action ModelImageTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought
π― What it does: Propose a MaDS framework for long-sequence GUI automation, combining dual-layer memory with multi-round debates to achieve pre-task verification and experience cycles.
MAESTRO: Meta-learning Adaptive Estimation of Scalarization Trade-offs for Reward Optimization
Yang Zhao (Harbin Institute of Technology), Ting Liu (Harbin Institute of Technology)
CodeOptimizationMeta LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextChain-of-Thought
π― What it does: Developed and implemented the MAESTRO framework, achieving multi-objective optimization in the alignment task of open-domain large language models (LLMs) through dynamic adaptive reward weighting.
MAGMA: A Multi-Graph based Agentic Memory Architecture for AI Agents
Dongming Jiang (University of Texas at Dallas), Bingzhe Li (University of Texas at Dallas)
CodeRetrievalExplainability and InterpretabilityComputational EfficiencyGraph Neural NetworkTransformerLarge Language ModelAgentic AIPrompt EngineeringTextGraphBenchmarkRetrieval-Augmented Generation
π― What it does: Proposes the MAGMA multi-graph agent memory architecture, achieving structured retrieval and reasoning of external memory through four types of graphs: semantic, temporal, causal, and entity.
Make Mechanistic Interpretability Auditable: A Call to Develop Guidelines via Continuous Collaborative Reviewing
Michael Lan (Martian), Amir Abdullah
CodeSafty and PrivacyExplainability and InterpretabilityAgentic AIPrompt EngineeringTextReview/Survey PaperRetrieval-Augmented Generation
π― What it does: Propose a Mechanism Interpretability (MI) audit framework, which includes a continuous collaborative review platform, community-developed 'living document' guidelines, and source-code-based audit systems.
MalruleLib: Large-Scale Executable Misconception Reasoning with Step Traces for Modeling Student Thinking in Mathematics
Xinghe Chen (Rice University), Shashank Sonkar (University of Central Florida)
CodeLarge Language ModelPrompt EngineeringTextBenchmarkChain-of-Thought
π― What it does: Proposed the MALRULELIB framework, which converts 101 learning science-based mathematical misconceptions (malrule) into executable programs, and generates dual-path (correct and incorrect) step-by-step problem-solving trajectories through 498 parameterized templates, thereby constructing a large-scale cross-template student error reasoning benchmark;
CodeAutonomous DrivingOptimizationTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringMixture of ExpertsTextMultimodalityBenchmarkRetrieval-Augmented Generation
π― What it does: Propose the MANGO framework, which utilizes the global structure of websites and multi-armed bandits to achieve efficient web navigation.
π― What it does: Propose a sentence encoder mapping method based on quantum relative entropy (QRE), constructing a 2D visualization map of 1101 sentence encoders.
MARCH: Multi-Agent Reinforced Check for Hallucination
Zhuo Li (Qwen Large Model Application Team, Alibaba), Guanjun Jiang (Qwen Large Model Application Team, Alibaba)
CodeGenerationData SynthesisExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringMixture of ExpertsTextBenchmarkRetrieval-Augmented Generation
π― What it does: Proposed a retrieval-augmented generation framework called MARCH based on multi-agent reinforcement learning, aimed at eliminating hallucinations in large language models during retrieval-augmented generation (RAG) tasks.
CodeReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkFinance Related
π― What it does: Propose Market-Bench, a closed-loop multi-agent supply chain economic simulation environment, for evaluating the economic decision-making capabilities of large language models in procurement auctions, pricing, and marketing language generation.
MASPO: Unifying Gradient Utilization, Probability Mass, and Signal Reliability for Robust and Sample-Efficient LLM Reasoning
Xiaoliang Fu (Meituan), Xunliang Cai (Meituan)
CodeOptimizationComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsTextBenchmark
π― What it does: A unified RLVR algorithm called MASPO is studied, which addresses three major bottlenecks: gradient utilization, probability mass sensitivity, and signal reliability.
Massively Multilingual Joint Segmentation and Glossing
Michael Ginn (University of Colorado Boulder), Alexis Palmer (University of Colorado Boulder)
CodeSegmentationTransformerLarge Language ModelReinforcement LearningPrompt EngineeringContrastive LearningText
π― What it does: Proposed and trained a multilingual joint segmentation and interlinear glossing model called POLYGLOSS, which can output morphological segmentation of words and corresponding glosses in one go.
MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems
Shuhang Chen (Zhejiang University), Yi Yang (Zhejiang University)
CodeRecognitionExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmarkChain-of-Thought
π― What it does: Proposes the FlowVerse benchmark for fine-grained evaluation of perception and reasoning abilities in visual math problems, and designs a multi-module solution called MathFlow that separates perception and reasoning. A specialized model, MathFlow-P-7B, is trained to improve the extraction and description of graphics.
π― What it does: A multilingual extension of HuBERT (MAUBERT) was constructed, further training language-agnostic and context-invariant speech representations through supervised learning using speech-to-articulatory features on 55 languages.
π― What it does: Propose MCP-Flow, which automatically collects MCP servers and tools across multiple platforms, generating over 60k instruction-function call pairs, and uses them to train LLMs to master MCP skills.
π― What it does: Proposes MeanAudio, a text-to-audio generation model based on Mean Flow, capable of achieving high-quality audio synthesis in a single step (1 NFE);
Measuring Social Bias in Vision-Language Models with Face-Only Counterfactuals from Real Photos
Haodong Chen (Harbin Institute of Technology), Jun Yu (Harbin Institute of Technology)
CodeExplainability and InterpretabilityData-Centric LearningPrompt EngineeringVision Language ModelDiffusion modelContrastive LearningImageTextMultimodality
π― What it does: Studied the measurement of social bias in vision-language models using real photos with only minor facial modifications.
Mechanistic Interpretability Should Prioritize Feature Consistency in Sparse Autoencoders
Xiangchen Song (Carnegie Mellon University), Kun Zhang (Carnegie Mellon University)
CodeExplainability and InterpretabilityRepresentation LearningLarge Language ModelAuto EncoderContrastive LearningTextTabularTime SeriesSequential
π― What it does: This study investigates the feature consistency issue of sparse autoencoders (SAE) in terms of mechanistic interpretability, proposes and evaluates the Pairwise Dictionary Mean Correlation Coefficient (PW-MCC) metric based on dictionary matching, and theoretically proves and experimentally verifies its ability to achieve high consistency in TopK SAE.
MED-COREASONER: Reducing Language Disparities in Medical Reasoning via Language-Informed Co-Reasoning
Fan Gao (University of Tokyo), Irene Li (University of Tokyo)
CodeRecommendation SystemFederated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningData-Centric LearningDrug DiscoveryTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBiomedical DataElectronic Health RecordsBenchmarkRetrieval-Augmented GenerationChain-of-Thought
π― What it does: Proposed and implemented the MED-COREASONER framework, leveraging parallel reasoning chains in English and local languages, concept extraction and fusion, and retrieval enhancement to improve the quality of medical multilingual reasoning.
MediEval: A Unified Medical Benchmark for Patient-Contextual and Knowledge-Grounded Reasoning in LLMs
Zhan Qu (TU Dresden), Michael FΓ€rber (TU Dresden)
CodeExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningDrug DiscoveryTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringMixture of ExpertsTextBiomedical DataElectronic Health RecordsBenchmarkRetrieval-Augmented GenerationChain-of-Thought
π― What it does: Proposed the MediEval benchmark, combining MIMIC-IV medical records with the UMLS knowledge base, to evaluate LLMs in terms of factual accuracy and patient context consistency, and designed the CoRFu fine-tuning method based on this.
Mediocrity is the key for LLM as a Judge Anchor Selection
Shachar Don-Yehiya (Hebrew University of Jerusalem), Omri Abend (Hebrew University of Jerusalem)
CodeTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation
π― What it does: Investigated the impact of using anchors (benchmark models) for pairwise comparisons on ranking reliability when large language models (LLMs) act as judges (LLM-as-a-Judge, LMJ), systematically evaluating the performance of 22 anchors on the Arena-Hard-v2.0 dataset.
CodeDrug DiscoveryTransformerLarge Language ModelReinforcement LearningAgentic AITextTabularBiomedical DataElectronic Health RecordsBenchmarkRetrieval-Augmented GenerationChain-of-Thought
π― What it does: Constructed the MedMCP-Calc benchmark to evaluate the performance of LLMs in real medical calculator workflows using fuzzy task descriptions, EHR data interaction, and MCP tool integration, and on this basis proposed and trained the CalcMate model.
MedVerse: Efficient and Reliable Medical Reasoning via DAG-Structured Parallel Execution
Jianwen Chen (University Of North Carolina Chapel Hill), Huaxiu Yao (University Of North Carolina Chapel Hill)
CodeComputational EfficiencyDrug DiscoveryGraph Neural NetworkTransformerLarge Language ModelTextGraphBiomedical DataElectronic Health RecordsBenchmarkRetrieval-Augmented GenerationChain-of-Thought
π― What it does: Propose the MedVerse framework, which restructures medical reasoning as a directed acyclic graph (DAG) and enables parallel inference;
CodeRetrievalExplainability and InterpretabilityComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerLarge Language ModelVision Language ModelContrastive LearningImageTextMultimodalityGraphRetrieval-Augmented Generation
π― What it does: Propose MegaRAG, an end-to-end framework for automatically constructing a multi-modal knowledge graph (MMKG) and using it for retrieval-augmented generation;
Mem^2Evolve: Towards Self-Evolving Agents via Co-Evolutionary Capability Expansion and Experience Distillation
Zihao Cheng (Beihang University), Yunhong Wang (Beijing Institute Of Technology)
CodeAutonomous DrivingOptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningRobotic IntelligenceMeta LearningAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextSequentialBenchmarkRetrieval-Augmented GenerationChain-of-Thought
π― What it does: Propose the Memβ―2β―Evolve framework, achieving dual mechanisms of asset memory (Asset Memory) and experience memory (Experience Memory), constructing a forward reasoning and backward evolution loop, realizing a self-evolving language model agent;
Memory-Augmented LLM-based Multi-Agent System for Automated Feature Generation on Tabular Data
FengXian Dong, Enhong Chen (University of Science and Technology of China)
CodeOptimizationExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelAgentic AIPrompt EngineeringTabularRetrieval-Augmented Generation
π― What it does: A framework based on multi-agent and memory-enhanced large language models (MALMAS) was constructed for automated feature generation, supporting multi-round iterations and guiding feature generation and evaluation through multi-level memory (procedural memory, feedback memory, conceptual memory, and global conceptual memory).
MemRec: Collaborative Memory-Augmented Agentic Recommender System
Weixin Chen (Hong Kong Baptist University), Yongfeng Zhang (Rutgers University)
CodeRecommendation SystemGraph Neural NetworkTransformerLarge Language ModelAgentic AIContrastive LearningTextGraphBenchmarkRetrieval-Augmented Generation
π― What it does: Proposed the MemRec framework, which improves the performance of LLM agent recommendation systems by utilizing collaborative memory and asynchronous propagation mechanisms.
Merlinβs Whisper: Enabling Efficient Reasoning in Large Language Models via Black-box Persuasive Prompting
Heming Xia (Hong Kong Polytechnic University), Wenjie Li (Hong Kong Polytechnic University)
CodeComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextBenchmark
π― What it does: Studied reducing overthinking in large language models through black-box persuasive prompting, and proposed the WHISPER framework to achieve efficient reasoning.
Qianyu He (Fudan University), Yanghua Xiao (Fudan University)
CodeGenerationMeta LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AIPrompt EngineeringTextRetrieval-Augmented Generation
π― What it does: This paper constructs an automated system called METAR to generate high-quality metaphor riddles and uses these riddles to train large language models to enhance their reasoning capabilities;
METER: Evaluating Multi-Level Contextual Causal Reasoning in Large Language Models
Pengfeng Li (Sichuan University), See-Kiong Ng (National University of Singapore)
CodeExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringTextMultimodalityBenchmarkChain-of-Thought
π― What it does: Proposed and implemented the METER benchmark, which evaluates the performance of large language models in three-layer causal reasoning (causal discovery, intervention, counterfactual) under a unified context, and constructed a multiple-choice dataset with 4,145 samples.
MetFuse: Figurative Fusion between Metonymy and Metaphor
Saptarshi Ghosh (University of Cincinnati), Tianyu Jiang (University of Cincinnati)
CodeGenerationData SynthesisTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBenchmark
π― What it does: A framework for converting textual sentences into metonymy, metaphor, and mixed sentences was constructed, and this framework was used to generate the first MetFuse dataset containing mixed expressions of metonymy and metaphor (totaling 4,000 sentences).
CodeGenerationExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought
π― What it does: Propose a new decoding sampling method called Mink Sampling, which can dynamically identify 'semantic cliffs' in the logit distribution and truncate the candidate set without relying on temperature parameters.
Miner: Mining Intrinsic Mastery for Data-Efficient RL in Large Reasoning Models
Shuyang Jiang (Fudan University), Yu Wang (Shanghai Jiao Tong University)
CodeComputational EfficiencyData-Centric LearningTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBiomedical DataBenchmark
π― What it does: Propose the MINER framework, which generates self-supervised rewards by leveraging the intrinsic uncertainty from positive homogeneous (PH) rollout, significantly improving the data efficiency of RLVR on large-scale reasoning models.
MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Long-tail Knowledge
Jie He (University of Edinburgh), Jeff Z. Pan (University of Edinburgh)
CodeTransformerLarge Language ModelTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
π― What it does: This paper proposes MINTQA, a multi-hop question answering benchmark that evaluates the performance of large language models in multi-hop reasoning and retrieval fusion by combining two dimensions: new/old knowledge and popular/unpopular knowledge.
Mirroring Users: Towards Building Preference-aligned User Simulator with User Feedback in Recommendation
Tianjun Wei (Nanyang Technological University), Jie Zhang (Nanyang Technological University)
CodeRecommendation SystemExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringText
π― What it does: By constructing the USERMIRRORER framework, the decision-making process of user feedback in recommendation systems is utilized to achieve fine-grained alignment with LLMs, thereby realizing more accurate user simulation.
MIRTH: Mutual-Information Reasoning with Temporal Hubs for Vision-Language-Action Agents
Hao Sun (Ritsumeikan University), Yen-wei Chen
CodeExplainability and InterpretabilityComputational EfficiencyRepresentation LearningRobotic IntelligenceTransformerReinforcement LearningPrompt EngineeringMixture of ExpertsVision-Language-Action ModelContrastive LearningImageVideoTextMultimodality
π― What it does: Proposes a unified Vision-Language-Action framework called MIRTH, addressing the issues of short-term temporal vision, inference gap, and low inference efficiency in single-frame VLA models.
Mitigating Context Interference for Reliable and Efficient Search Agents
Boyang Xue (Chinese University of Hong Kong), Aldo Lipani (University College London)
CodeComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation
π― What it does: Systematically study context interference in multi-round search agents, finding that it mainly comes from the latest retrieved documents, and propose a context refiner based on distillation, which is then embedded into the reinforcement learning training process to improve the reliability and efficiency of the search agent.
Mitigating Lost in Multi-turn Conversation via Curriculum RL with Verifiable Accuracy and Abstention Rewards
Ming Li (University of Maryland), Bing Yin (Amazon)
CodeOptimizationAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringText
π― What it does: Designed and trained a reinforcement learning-based framework called RLAAR, encouraging LLMs to both answer correctly and determine when to give up in multi-turn conversations
MixKVQ: Query-Aware Mixed-Precision KV Cache Quantization for Long-Context Reasoning
Tao Zhang (South China University Of Technology), Cen Chen (Beihang University)
CodeComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelPrompt EngineeringText
π― What it does: Propose MixKVQ, a low-bit quantization method for the KV cache of large language models, aiming to significantly reduce memory usage while maintaining the accuracy of long-text inference.
CodeClassificationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelAgentic AIPrompt EngineeringImageTextMultimodalityRetrieval-Augmented GenerationChain-of-Thought
π― What it does: Proposed the MM-StanceDet multi-agent framework, achieving more robust multi-modal stance detection through four stages: retrieval enhancement, specialized multi-modal analysis, debate, and self-reflection.
MMAC: A Multilingual, Multimodal Alignment Framework for Cultural Grounding Evaluation
Weihua Zheng (Agency for Science, Technology and Research), Nancy F. Chen (Singapore University of Technology and Design)
CodeExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-ThoughtAudio
π― What it does: Proposed the MMAC framework and the MMAC-bench dataset to systematically evaluate the cultural cognition and reasoning capabilities of large language models in multilingual, multimodal (text, image, voice) environments.
Mnemis: Dual-Route Retrieval on Hierarchical Graphs for Long-Term LLM Memory
Zihao Tang (Microsoft), Qi Zhang (Microsoft)
CodeRetrievalComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringTextGraphBenchmarkRetrieval-Augmented Generation
π― What it does: Proposed a long-term memory framework called Mnemis, which integrates traditional similarity retrieval (System-1) with global hierarchical retrieval (System-2) routing;
MONETA: Multimodal Industry Classification through Geographic Information with Multi Agent Systems
Arda YΓΌksel (Technical University of Darmstadt), Ivan Habernal (Ruhr University Bochum)
CodeClassificationTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought
π― What it does: Proposed the MONETA multimodal industry classification benchmark, which uses text (websites, Wikipedia, Wikidata) and geospatial information (OpenStreetMap, satellite images) to perform NACE classification on 1,000 European companies, aiming to replace manual expert verification;
π― What it does: Proposed and experimentally tested the MQM re-annotation method, allowing reviewers to delete, modify, or add errors based on existing error annotations, thereby improving fine-grained quality assessment in machine translation evaluation;
π― What it does: Propose the MSMO framework, combining sentence-level adversarial training, aspect-level consistency training, and multi-objective optimization, to achieve cross-lingual ABSA feature alignment and fine-grained alignment, and then perform knowledge distillation based on this.
MT^{3}: A Synergistic Multi-Task RL Framework for Specializing MLLMs in Text Image Machine Translation
Zhaopeng Feng (Zhejiang University), Zuozhu Liu (Zhejiang University)
CodeImage TranslationTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision Language ModelImageTextMultimodalityBenchmark
π― What it does: Propose the MT3 framework, which utilizes multi-task reinforcement learning to specialize multimodal large language models (MLLMs) into end-to-end text-image machine translation (TIMT) expert models.
MTR-Suite: A Framework for Evaluating and Synthesizing Conversational Retrieval Benchmarks
Junhao Ruan (Northeastern University), JingBo Zhu
CodeData SynthesisRetrievalTransformerLarge Language ModelAgentic AIPrompt EngineeringContrastive LearningTextBenchmarkFinance RelatedRetrieval-Augmented Generation
π― What it does: Propose the MTR-Suite framework, integrating evaluation (MTR-EVAL), multi-agent synthesis (MTR-PIPELINE), and a new dialogue retrieval benchmark (MTR-BENCH).
MTSQL-R1: Towards Long-Horizon Multi-Turn Text-to-SQL via Agentic Training
Taicheng Guo (University of Notre Dame), Chandan K. Reddy (Amazon)
CodeExplainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringTextTabularBenchmarkChain-of-Thought
π― What it does: Proposed the MTSQL-R1 framework, modeling multi-turn Text-to-SQL as a Markov Decision Process, supporting agent-based execution, verification, and self-correction;
Multi-View Attention Multiple-Instance Learning Enhanced by LLM Reasoning for Cognitive Distortion Detection
Jun Seo Kim (Gachon University), Hye Hyeon Kim (Yonsei University)
CodeClassificationExplainability and InterpretabilityTransformerLarge Language ModelText
π― What it does: Propose a method that splits each sentence into three parts: emotion, logic, and behavior (ELB), then uses a large language model (LLM) to generate multiple instances of cognitive distortions, and finally classifies them using a multi-instance learning (MIL) framework with multi-perspective gated attention.
Multilingual and Cross-Lingual Citation Needed Detection on Wikipedia for Lower-Resource Languages
Gerrit Quaremba (King's College London), Elena Simperl (King's College London)
CodeClassificationKnowledge DistillationTransformerLarge Language ModelSupervised Fine-TuningTextBenchmark
π― What it does: Constructed a cross-lingual 'Need to Cite' detection dataset (MCN), and trained and evaluated a small decoder model on 18 languages with different resource levels.
Multimodal Safety Evaluation in Generative Agent Social Simulations
Alhim Adonai Vera Gonzalez (University of Cincinnati), Bernard Ghanem
CodeSafty and PrivacyTransformerLarge Language ModelAgentic AIPrompt EngineeringVision Language ModelImageTextMultimodalityRetrieval-Augmented Generation
π― What it does: A reproducible multimodal safety evaluation framework was constructed, and generative agents were used in social simulation environments to detect and correct unsafe plans, analyzing safety improvements and social dynamics.
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding
Fuwen Luo (Tsinghua University), Yang Liu (Tsinghua University)
CodeComputational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringVision Language ModelVideoTextMultimodality
π― What it does: Proposed a multi-segment temporal alignment method called MUSEG based on reinforcement learning, enhancing the temporal reasoning ability of multi-modal large language models.
Native Hybrid Attention for Efficient Sequence Modeling
Jusen Du (Tsinghua University), Yu Cheng (Chinese University of Hong Kong)
CodeComputational EfficiencyRepresentation LearningRecurrent Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningTextSequential
π― What it does: Propose Native Hybrid Attention (NHA), a hybrid attention architecture that simultaneously integrates linear RNN memory and sliding window Softmax attention at the same level.
NaturalGAIA: A Verifiable Benchmark and Hierarchical Framework for Long-Horizon GUI Tasks
Zihan Zheng (South China Normal University), Qianglong Chen (Zhejiang University)
CodeRobotic IntelligenceReinforcement Learning from Human FeedbackTransformerLarge Language ModelVision-Language-Action ModelTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought
π― What it does: Proposes the NaturalGAIA benchmark and the LightManus-Jarvis hierarchical framework for evaluating and enhancing the performance of LLM agents in long-term GUI tasks.
CodeTransformerLarge Language ModelPrompt EngineeringTextBenchmarkChain-of-Thought
π― What it does: Proposes the NEO-CLASSIC benchmark, which uses strictly metered poems created by modern experts to evaluate the language aesthetic reasoning ability of LLMs.
NeoAMT: Neologism-Aware Agentic Machine Translation with Reinforcement Learning
Zhongtao Miao (The University of Tokyo), Yoshimasa Tsuruoka (The University of Tokyo)
CodeTransformerLarge Language ModelReinforcement LearningAgentic AITextRetrieval-Augmented Generation
π― What it does: Constructed a multilingual neologism machine translation dataset called Neko, and proposed the NeoAMT framework, which uses reinforcement learning and dictionary retrieval to translate sentences containing neologisms.
Zhuowei Chen (University of Pittsburgh), Xiang Lorraine Li (University of Pittsburgh)
CodeClassificationExplainability and InterpretabilityMeta LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningText
π― What it does: Proposes NEUFS, an active few-shot learning framework based on dynamic neuronal activation, specifically designed to select the most valuable few-shot examples for annotation from unlabeled data in specialized domains for large language models (LLMs);
NOSE: Neural Olfactory-Semantic Embedding with Tri-Modal Orthogonal Contrastive Learning
Yanyi Su (Xiamen University), Jun Cheng (Xiamen University)
CodeDrug DiscoveryGraph Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningTextMultimodalityGraph
π― What it does: Propose a tri-modal (molecular structure, receptor sequence, semantic description) aligned olfactory embedding framework called NOSE, achieving modal information decoupling and fusion through orthogonal injection and weak positive sample contrastive learning.
Not All Tokens Matter: Towards Efficient LLM Reasoning via Token Significance in Reinforcement Learning
Hanbing Liu (Tsinghua University), Dongmei Zhang (Microsoft)
CodeComputational EfficiencyTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkChain-of-Thought
π― What it does: Propose the BINGO framework, which improves the efficiency of large model chain-of-thought reasoning by utilizing token importance awareness and dynamic length rewards.
π― What it does: OASIS proposes an online adaptive sample selection framework in continuous instruction tuning (CIT), which can real-time select the most informative samples from the data stream, significantly shortening training time while maintaining the model's real-time adaptation capability.
OctoBench: Benchmarking Scaffold-Aware Instruction Following in Repository-Grounded Agentic Coding
Deming Ding (Fudan University), Tao Gui (Fudan University)
CodeAI Code AssistantTransformerLarge Language ModelAgentic AITextBenchmark
π― What it does: This paper proposes the OCTOBENCH benchmark to evaluate whether models can follow multi-source, persistent instruction constraints in warehouse-level agent-based coding.
OLA: Output Language Alignment in Code-Switched LLM Interactions
Juhyun Oh (KAIST), Alice Oh (KAIST)
CodeExplainability and InterpretabilityData-Centric LearningAI Code AssistantTransformerLarge Language ModelPrompt EngineeringTextMultimodalityBenchmarkChain-of-Thought
π― What it does: The study investigates the failure of large language models (LLMs) to implicitly output language alignment during code-switching interactions, constructs the OLA benchmark, and proposes the CS-DPO method based on preference alignment.
OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Models
Qiguang Chen (Harbin Institute of Technology), Wanxiang Che (Harbin Institute of Technology)
CodeExplainability and InterpretabilityComputational EfficiencyRepresentation LearningDrug DiscoveryTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmarkPhysics RelatedRetrieval-Augmented GenerationChain-of-Thought
π― What it does: Propose OMIBench, a multi-graph Olympiad-level reasoning benchmark, to evaluate the ability of large vision-language models in multi-graph reasoning.
CodeGenerationAI Code AssistantTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmarkRetrieval-Augmented Generation
π― What it does: Propose Omni-I2C, a large-scale, cross-lingual, and multi-domain Image-to-Code benchmark, to evaluate the ability of large multimodal models to convert complex digital graphics into executable code.
On the Emotion Understanding of Synthesized Speech
Yuan Ge (Northeastern University), Tong Xiao (Kunming University of Science and Technology)
CodeRecognitionDomain AdaptationTransformerLarge Language ModelContrastive LearningTextAudio
π― What it does: Systematically evaluate the generalization ability of synthetic speech emotion recognition models, and analyze the gap between human speech and synthetic speech in emotional understanding;
π― What it does: Propose the ORBIT framework, which uses on-policy reinforcement fine-tuning (RFT) with offline expert trajectory rewards to address the high-cost interaction and sparse reward problems in multi-step embodied planning.
CodeExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringFlow-based ModelTextBenchmarkRetrieval-Augmented Generation
π― What it does: Proposed an scalable multi-turn instruction-following evaluation framework and the EvolIF benchmark, utilizing a three-layer tracking mechanism and a query synthesis agent driven by large language models (LLMs) to dynamically generate dialogues, and introducing a patience threshold based on Flow theory and process-oriented evaluation metrics;
One Persona, Many Cues, Different Results: How Sociodemographic Cues Impact LLM Personalization
Franziska Weeber (University of Stuttgart), Sebastian PadΓ³ (University of Stuttgart)
CodeExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextBenchmark
π― What it does: Investigated the impact of six commonly used persona prompts (name, explicit description, conversation history) on the personalization bias of large language models, and constructed a scalable multi-task evaluation benchmark.
π― What it does: This paper proposes a first-order non-autoregressive natural language generation method that completes text generation in a single step using the shortcut flow matching model.
Open Your Modelβs Eyes: Video and Context-Aware Multimodal Backchannel Prediction
Min-Jae Kim (Korea University), Gyeong-Moon Park (Korea University)
CodeClassificationTransformerSupervised Fine-TuningVision Language ModelContrastive LearningVideoTextMultimodalityAudio
π― What it does: A novel multimodal behind-channel prediction framework called CAMA-BC was studied, which integrates audio, text, and video information and achieves modality alignment through hierarchical cross-attention, addressing issues such as visual information bias, temporal deviation, and imbalance between context and response.
Outcome-Grounded Advantage Reshaping for Fine-Grained Credit Assignment in Mathematical Reasoning
Ziheng Li (Fudan University), Hongcheng Guo (Fudan University)
CodeOptimizationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsTextBenchmarkChain-of-Thought
π― What it does: Proposes a fine-grained credit assignment framework called OAR based on the impact of final answers, aimed at improving reward propagation in Group Relative Policy Optimization (GRPO) for long reasoning tasks.
Overcoming Copyright Barriers in Corpus Distribution Through Non-Reversible Hashing
Arthur Amalvy (Academia Sinica), Hen-Hsen Huang (Academia Sinica)
CodeSafty and PrivacyPrompt EngineeringText
π― What it does: Propose a method that utilizes non-reversible hashing to share copyrighted text annotations, allowing users with the original text to legally obtain annotations.
PAR: Training-Free Positional Perturbation and Attention Recycling for Faithful OCR
Yao Yao (Shanghai Jiao Tong University), Hai Zhao (Shanghai Jiao Tong University)
CodeRecognitionTransformerPrompt EngineeringVision Language ModelDiffusion modelContrastive LearningImageTextMultimodalityBenchmark
π― What it does: Propose a training-agnostic, inference-time intervention framework called PAR, designed to suppress hallucinations caused by language priors in visual language models during OCR tasks, thereby improving the visual consistency of text recognition.
PARIF: Pushing the Pareto Frontier of Instruction Following and Reasoning with Curriculum Reinforcement Learning
Rongchuan Mu (Harbin Institute of Technology), Bing Qin (Harbin Institute of Technology)
CodeOptimizationComputational EfficiencyRepresentation LearningData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringDiffusion modelAuto EncoderContrastive LearningTextChain-of-Thought
π― What it does: Propose a two-stage RLVR curriculum learning framework called PARIF to enhance the comprehensive performance of large reasoning models in instruction following and reasoning capabilities.
PEAP: Proactive Embodied Action Sequence Planning with Joint Understanding of Vision and Audio Perception
Tianwei Lan (Beijing Institute Of Technology), Yuhang Guo (Beihang University)
CodeAutonomous DrivingComputational EfficiencyRepresentation LearningRobotic IntelligenceTransformerLarge Language ModelPrompt EngineeringVision-Language-Action ModelContrastive LearningImageMultimodalityAudio
π― What it does: Studied the task of planning active avatar action sequences in a multimodal (visual + audio) setting, and constructed the PEAP dataset.
π― What it does: Propose PEAR, a contrast-based quality estimation (QE) metric that can perform graded relative quality difference assessment between two candidate translations of the same source text.
PersonalAlign: Hierarchical Implicit Intent Alignment for Personalized GUI Agent with Long-Term User-Centric Records
Yibo Lyu (Harbin Institute of Technology), Liqiang Nie (Harbin Institute of Technology)
CodeRecommendation SystemReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringMixture of ExpertsContrastive LearningTextMultimodalitySequentialBenchmarkRetrieval-Augmented Generation
π― What it does: Propose the PersonalAlign task, construct the AndroidIntent benchmark, and design the HIM-Agent memory framework for hierarchical implicit intent alignment of long-term user records.
PersonalityDBench: A Dataset for Personality Disorders - from Modeling to Controlled Generation
Federico Ravenda (UniversitΓ della Svizzera italiana), Andrea Raballo (UniversitΓ della Svizzera italiana)
CodeClassificationRecognitionGenerationData SynthesisTransformerLarge Language ModelContrastive LearningTextBenchmarkRetrieval-Augmented Generation
π― What it does: Constructed the PersonalityDBench dataset, which includes clinically annotated Reddit language samples (PRISMA) and a benchmark (PersonaDSteering) for evaluating the controllability of LLMs in generating behaviors related to personality disorders. The feasibility of diagnosing personality disorders in natural language, HiTOP dimension features, and LLM directional control were validated from this dataset.
PIArena: A Platform for Prompt Injection Evaluation
Runpeng Geng (Pennsylvania State University), Jinyuan Jia (Pennsylvania State University)
CodeSafty and PrivacyAdversarial AttackTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation
π― What it does: Proposed the PIArena unified platform for evaluating prompt injection attacks and defenses, and designed an adaptive strategy attack;
CodeClassificationComputational EfficiencyData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation
π― What it does: Propose a fast automated prompt construction method called PIAST, which utilizes LLMs to generate and iteratively improve a few few-shot examples, thereby enhancing the performance of gradient-free updated LLMs.
CodeTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBenchmarkRetrieval-Augmented Generation
π― What it does: Proposed and constructed the PLURULE benchmark for detecting whether comments violate specific community rules in a multilingual, multimodal community environment, simulating real moderators' decision-making through multiple-choice questions.
PodBench: A Comprehensive Benchmark for Instruction-Aware Audio-Oriented Podcast Script Generation
Chenning Xu (Tencent), Mingyang Song (Tencent)
CodeGenerationData SynthesisTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-ThoughtAudio
π― What it does: Constructed PodBench benchmark, focusing on instruction-aware and context-driven long-form multi-speaker podcast script generation tasks, providing 800 long-context samples and multi-dimensional instructions;
PRBench: Large-Scale Expert Rubrics for Evaluating High-Stakes Professional Reasoning
Afra Feyza AkyΓΌrek (Scale AI), Yunzhong He (Scale AI)
CodeLarge Language ModelPrompt EngineeringTextBenchmarkFinance RelatedRetrieval-Augmented GenerationChain-of-Thought
π― What it does: This work constructs and publicly releases a large professional reasoning benchmark called PRBench, which includes 1,100 real-world task scenarios written by financial and legal experts, along with 18,711 finely crafted expert rubrics, corresponding multi-turn dialogues, and economic impact annotations.
CodeExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringContrastive LearningText
π― What it does: Propose a training-free differential preference-driven (DPS) method that identifies sparse 'preference heads' based on mechanism interpretation and controls them during decoding to achieve interpretable personalization.
PrefixNLI: Detecting Factual Inconsistencies as Soon as They Arise
Sapir Harary (Bar Ilan University), Ido Dagan (Bar Ilan University)
CodeGenerationExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextRetrieval-Augmented Generation
π― What it does: Propose the PrefixNLI task to detect factual inconsistencies during the text generation process.
PRIME: A Process-Outcome Alignment Benchmark for Verifiable Reasoning in Mathematics and Engineering
Xiangfeng Wang (University of Science and Technology of China), Daxin Jiang (Stepfun)
CodeExplainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
π― What it does: Proposed the PRIME benchmark for evaluating process-result consistency verification of large reasoning models in the fields of mathematics and engineering.
PRInTS: Reward Modeling for Long-Horizon Information Seeking
Jaewoo Lee (University of North Carolina at Chapel Hill), Mohit Bansal (University of North Carolina at Chapel Hill)
CodeAutonomous DrivingOptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningAdversarial AttackHyperparameter SearchData-Centric LearningRobotic IntelligenceMeta LearningDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackNeural Architecture SearchProtein Structure PredictionTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AIMixture of ExpertsDiffusion modelScore-based ModelFlow-based ModelRectified FlowNeural Radiance FieldAuto EncoderGenerative Adversarial NetworkContrastive LearningGaussian SplattingSimultaneous Localization and MappingWorld ModelOptical FlowTextBenchmarkRetrieval-Augmented GenerationChain-of-ThoughtStochastic Differential EquationOrdinary Differential Equation
π― What it does: This paper proposes PRInTS, a generative process reward model that combines information gain scoring with recursive trajectory summarization, enabling fine-grained evaluation of each step (reasoning + tool call) in long-term information-seeking tasks, and providing guidance for selection during testing for LLM agents.
PRiSM: Benchmarking Phone Realization in Speech Models
Shikhar Bharadwaj (Carnegie Mellon University), David R. Mortensen (Carnegie Mellon University)
CodeRecognitionConvolutional Neural NetworkRecurrent Neural NetworkTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBenchmarkAudio
π― What it does: Proposes the PRiSM benchmark for evaluating systems that transcribe speech into phonemes, assessing them along two major dimensions: intrinsic (PFER) and extrinsic (downstream tasks).
PRISMA: Preference-Reinforced Self-Training Approach for Interpretable Emotionally Intelligent Negotiation Dialogues
Prajwal Vijay Kajare (Indian Institute of Technology Jodhpur), Asif Ekbal (Indian Institute of Technology Patna)
CodeExplainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought
π― What it does: This paper proposes PRISMA, an interpretable emotion-intelligent negotiation dialogue system, capable of generating emotionally appropriate and interpretable responses through emotion perception and strategy selection;
PRISP: Privacy-Safe Few-Shot Personalization via Lightweight Adaptation
Junho Park (Seoul National University), Taesup Moon (Seoul National University)
CodeFederated LearningSafty and PrivacyComputational EfficiencyMeta LearningTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation
π― What it does: Propose a privacy-safe, lightweight few-shot personalized framework called PRISP, which can achieve user-level personalization for large language models without requiring task data or sharing user parameters.
Privacy-R1: Privacy-Aware Multi-LLM Agent Collaboration via Reinforcement Learning
Zheng Hui (University of Cambridge), Nigel Collier (University College London)
CodeSafty and PrivacyTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringTextBiomedical Data
π― What it does: Proposes a privacy-aware multi-LLM agent collaboration framework called Privacy-R1, which dynamically routes text blocks to local or remote models while maintaining task performance and reducing the leakage of sensitive information.
CodeOptimizationAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringTextBenchmarkChain-of-Thought
π― What it does: Proposes the SCOPE framework, which decomposes multi-constraint planning into query-specific reasoning and general code execution, automatically generating reusable solver functions.
Projecting Out the Malice: A Global Subspace Approach to LLM Detoxification
Zenghao Duan (State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences), Xueqi Cheng (State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences)
CodeFederated LearningExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringText
π― What it does: A lightweight method named GLOSS is proposed, which achieves model detoxification by identifying and eliminating the global toxic subspace of parameters in the Feed-Forward network of LLMs.
ProMed: Shapley Information Gain Guided Reinforcement Learning for Proactive Medical LLMs
Hongxin Ding (Peking University), Yasha Wang (Peking University)
CodeExplainability and InterpretabilityDrug DiscoveryTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBiomedical DataBenchmark
π― What it does: Implementing an interactive diagnostic framework called ProMed in medical LLMs, transitioning from passive answering to active questioning.
π― What it does: Proposes an unsupervised cross-lingual emotion recognition transfer framework, called NOVA-ARC, that transfers from annotated non-linguistic sounds (such as laughter, crying, sighing) to linguistically vocalized speech.
Prune as You Generate: Online Rollout Pruning for Faster and Better RLVR
Haobo Xu (University Of Illinois At Urbana Champaign), Hanghang Tong (University Of Illinois At Urbana Champaign)
CodeComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsTextBenchmark
π― What it does: Propose an online episode trimming method called ARROL, which can predict the success probability of partial episodes during the generation process based on a lightweight quality head and trim them in advance, maintaining a near 0.5 ratio of positive and negative samples within the episode group;
Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness
Amin Banayeeanzade (University of Southern California), Sai Praneeth Karimireddy (University of Southern California)
CodeExplainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextMultimodalityBenchmarkRetrieval-Augmented Generation
π― What it does: Created and evaluated the PsySET benchmark, systematically comparing the effectiveness and reliability of multiple LLM emotion regulation methods (prompt engineering, vector injection, parameter-efficient fine-tuning, DPO) on emotional and personality dimensions.