arXivSub Start free trial

ACL 2026 Papers — Page 8

Annual Meeting of the Association for Computational Linguistics · 2296 papers

EvoNarrator: Modeling Scientific Evolution for Feasible Hypothesis Generation

Xiaoying Le (National University of Defense Technology), Bo Ding (National University of Defense Technology)

Explainability and InterpretabilityKnowledge DistillationRepresentation LearningDrug DiscoveryGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringTextGraphReview/Survey PaperRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose the EvoNarrator framework, which generates feasible scientific hypotheses with temporal causality and structural logic through evolutionary narrative and SocketMatch mechanisms.

EvoRoute: Experience-Driven Self-Routing LLM Agent Systems

Guibin Zhang (National University of Singapore), Shuicheng Yan (National University of Singapore)

OptimizationFederated LearningComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringMixture of ExpertsTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposes EvoRoute, a self-evolving model routing framework that dynamically selects the optimal LLM in multi-step agent workflows, addressing the Agent System Trilemma (trade-off between performance, cost, and efficiency)

EvoSci: A Bio-Inspired Multi-Agent Framework for the Evolution of Scientific Discovery

Xiaoyu Xiong (Tianjin University), Deyi Xiong (Tianjin University)

Meta LearningDrug DiscoveryReinforcement Learning from Human FeedbackGraph Neural NetworkTransformerLarge Language ModelAgentic AIPrompt EngineeringTextGraphRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose EvoSci, a multi-agent collaborative framework that completes the construction of scientific problems, research execution, evaluation, and evolution cycle through roles such as mentors, researchers, and reviewers, to achieve continuous generation of scientific ideas.

EvoSpark: Endogenous Interactive Agent Societies for Unified Long-Horizon Narrative Evolution

Shiyu He (Xinjiang University), Tingxiang Gu (Xinjiang University)

GenerationReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringGenerative Adversarial NetworkWorld ModelTextRetrieval-Augmented Generation

🎯 What it does: Propose the EVOSPARK framework to achieve long-term narrative evolution in LLM-based multi-agent systems.

EVOTOOL: Self-Evolving Tool-Use Policy Optimization in LLM Agents via Blame-Aware Mutation and Diversity-Aware Selection

Shuo Yang (University of Melbourne), Eduard Hovy (University of Melbourne)

OptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringTextBenchmark

🎯 What it does: Proposes EVOTOOL, a self-evolving framework for optimizing tool usage strategies of LLM agents;

Example Quality Matters: Multi-Aspects Example Augmentation for Private Library Programming

Yuhao Li (Beijing University of Posts and Telecommunications), Jingyu Wang (Beijing University of Posts and Telecommunications)

AI Code AssistantTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Studied the impact of example complexity, readability, and correctness on the effectiveness in private library code generation, and proposed an automated example enhancement method called ComboPrompt;

EXCEEDS: Extracting Complex Events via Nugget-based Grid Modeling in Scientific Domain

Yi-Fan Lu (Beijing Institute Of Technology), Heyan Huang (Beijing Institute Of Technology)

ClassificationRecognitionConvolutional Neural NetworkRecurrent Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningText

🎯 What it does: Constructed the SciEvents scientific literature event extraction dataset and proposed the EXCEEDS end-to-end grid modeling framework;

Execution as Verification: Fine-Grained Self-Correcting Reasoning for Complex KBQA

Minghan Zhang (Anhui University), Shu Zhao (Anhui University)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelTextGraphRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposes the Execution as Verification (EVER) framework, which reconstructs semantic parsing as an iterative self-correcting reasoning process driven by execution feedback, aiming to improve the executability of queries and the accuracy of answers in complex knowledge base question answering (KBQA).

ExecVerify: White-Box RL with Verifiable Stepwise Rewards for Code Execution Reasoning

Lingxiao Tang (Zhejiang University), Lingfeng Bao (Zhejiang University)

AI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextSequentialChain-of-Thought

🎯 What it does: By constructing executable programs and using white-box reinforcement learning to enhance the execution reasoning ability of code LLMs

Exons-Detect: Identifying and Amplifying Exonic Tokens via Hidden-State Discrepancy for Robust AI-Generated Text Detection

Xiaowei Zhu (Chinese Academy of Sciences), Li Guo (Chinese Academy of Sciences)

Anomaly DetectionExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBenchmark

🎯 What it does: Designed an AI text detection method called Exons-Detect that does not require training, utilizing differences in hidden states to identify and enhance 'exon' tokens to improve detection accuracy.

Expect the Unexpected? Testing the Surprisal of Salient Entities

Jessica Lin (Georgetown University), Amir Zeldes (Georgetown University)

Explainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringText

🎯 What it does: Investigate how the importance of global discourse entities (based on abstracts) affects the surprisal computed by language models, and explore its role in text predictability.

Experience Retrieval-Augmentation with Electronic Health Records Enables Accurate Discharge QA

Justice Ou (Yale University), Rex Ying (University of Illinois Urbana-Champaign)

RetrievalRecommendation SystemExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringTextBiomedical DataElectronic Health RecordsRetrieval-Augmented Generation

🎯 What it does: Proposes a retrieval-augmented generation framework called EXPRAG based on EHR for medical question answering

Experience-driven Multi-turn Reinforcement Learning for GUI Agents

Zhengxi Lu (Zhejiang University), Yueting Zhuang (Zhejiang University)

Robotic IntelligenceTransformerSupervised Fine-TuningReinforcement LearningPrompt EngineeringWorld ModelTextSequential

🎯 What it does: Propose an experience-driven multi-round reinforcement learning framework called EMPO based on expert trajectories, used to train GUI agents to complete complex tasks through multi-round interactions

Experiments or Outcomes? Probing Scientific Feasibility in Large Language Models

Seyedali Mohammadi (University of Maryland, Baltimore County), Francis Ferraro (University of Maryland, Baltimore County)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelPrompt EngineeringTextBenchmark

🎯 What it does: Studied the performance of large language models in scientific feasibility assessment, and designed a controlled evidence framework to distinguish between hypotheses, experiments, and results.

ExPerT: Personalizing LLM Responses to Users’ Domain Expertise via Query-Wise Semantic and Keystroke Behavioral Cues

Yeji Park (UNIST), Taesik Gong (UNIST)

Recommendation SystemExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextSequentialRetrieval-Augmented Generation

🎯 What it does: Propose the ExPerT framework, which infers user domain expertise based on dual clues at the query level: semantic content and keystroke dynamics, and generates personalized LLM responses accordingly.

Explain the Synth: Interpretable Evaluation of LLM Data Synthesis

Yue Yang (Monash University), Hao Wang (Monash University)

Data SynthesisExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringGenerative Adversarial NetworkContrastive LearningTabularBenchmark

🎯 What it does: Develop a rule-based interpretable auditing framework for evaluating synthetic tabular data generated by LLMs.

Explainable and Fine-Grained Safeguarding of LLM Multi-Agent Systems via Bi-Level Graph Anomaly Detection

Junjun Pan (Griffith University), Shirui Pan (Griffith University)

Anomaly DetectionSafty and PrivacyExplainability and InterpretabilityGraph Neural NetworkTransformerLarge Language ModelAgentic AIContrastive LearningTextGraph

🎯 What it does: A framework named XG-Guard for unsupervised graph anomaly detection is proposed to identify and explain malicious agents in LLM multi-agent systems

Explainable Disentangled Representation Learning for Generalizable Authorship Attribution in the Era of Generative AI

Hieu Man (University of Oregon), Thien Huu Nguyen (University of Oregon)

ClassificationAnomaly DetectionExplainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelPrompt EngineeringAuto EncoderContrastive LearningText

🎯 What it does: Designed the Explainable Authorship Variational Autoencoder (EAVAE) framework, which achieves explicit and interpretable separation of author writing style and text content through two-stage pre-training and VAE fine-tuning, and applied it to authorship attribution and AI-generated text detection.

Explaining Sources of Uncertainty in Automated Fact-Checking

Jingyi Sun (University of Copenhagen), Isabelle Augenstein (University of Copenhagen)

Explainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextRetrieval-Augmented Generation

🎯 What it does: Developed the CLUE framework, which generates natural language explanations for model uncertainty in automatic fact-checking by identifying conflicts and consistencies between claims and multiple pieces of evidence.

Explicit Trait Inference for Multi-Agent Coordination

Suhaib Abdurahman (University of Southern California), Yi Zhang (AWS Agentic AI Labs)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringTextChain-of-Thought

🎯 What it does: Propose the Explicit Trait Inference (ETI) framework, enabling multi-agent systems to infer and track partners' warmth (trust) and competence traits from interaction history, thereby guiding decision-making;

Exploring and Distilling Multi-Dimensional Clues for Interpretable Social Bot Detection

Yi Han (Harbin Institute of Technology), Ting Liu (Harbin Institute of Technology)

ClassificationAnomaly DetectionExplainability and InterpretabilityKnowledge DistillationTransformerLarge Language ModelReinforcement LearningTextGraphBenchmark

🎯 What it does: Propose a four-dimensional clue framework, train a clue generator based on reinforcement learning, and then fuse multi-dimensional clues to achieve interpretable social robot detection;

Exploring Attention Attractors in Large Language Models

Ziheng Wang (Renmin University of China), Qin Jin (Renmin University of China)

Explainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelText

🎯 What it does: This paper systematically studies the 'attention attractors' in large language models — that is, the tokens that receive significantly high attention in attention weights — and deeply analyzes them from three perspectives: function, distribution, and mechanism. It finds that they play an important role in information aggregation, clustering of low-semantic tokens, and correlation with activation dimensions.

Exploring Concreteness Through a Figurative Lens

Saptarshi Ghosh (University of Cincinnati), Tianyu Jiang (University of Cincinnati)

Explainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelPrompt EngineeringText

🎯 What it does: Investigate how large language models internally represent and process lexical concreteness, and analyze its relationship with metaphorical language.

Exploring Cross-Client Memorization of Training Data in Large Language Models for Federated Learning

Tinnakit Udsa (VISTEC), Norrathep Rattanavipanon (Prince of Songkla University)

Federated LearningSafty and PrivacyTransformerLarge Language ModelContrastive LearningTextRetrieval-Augmented Generation

🎯 What it does: Propose a contrast-based cross-client memory detection framework, which can finely evaluate the degree to which models remember training data under federated learning, including two memory patterns: intra-client and cross-client.

Exploring Layer Activation Dynamic of CoT via Knowledge Probe

Chuanxin Zhang (Southeast University), Zhaoyu Yang (Northeastern University)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextPhysics RelatedChain-of-Thought

🎯 What it does: This study proposes a multi-stage detection framework that analyzes the information flow in LLMs during the reasoning process through structured CoT tasks such as keyword extraction, theorem generation, and parameter replacement.

Extending First-Order Logic for Factual Reasoning over Knowledge Graphs

Yuanzhen Hao (University of Chinese Academy of Sciences), Desheng Wu (University of Chinese Academy of Sciences)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextGraphBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This paper proposes an extended first-order logic FOLX-KG for fact reasoning on knowledge graphs (KG), and constructs the Fact-FOLX-KG dataset based on this.

EyeMulator: Improving Code Language Models by Mimicking Human Visual Attention

Yifan Zhang (Vanderbilt University), Yu Huang (Vanderbilt University)

Explainability and InterpretabilityComputational EfficiencyAI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringDiffusion modelContrastive LearningTextSequentialBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposes EYEMULATOR, a model-agnostic technique that injects human visual attention signals into CodeLLM by mapping human programmers' eye movement scan paths to AST tokens;

FactCorrector: A Graph-Inspired Approach to Long-Form Factuality Correction of Large Language Models

Javier Carnerero-Cano (IBM Research Europe - Ireland), Elizabeth M. Daly (IBM Research Europe - Ireland)

Explainability and InterpretabilityComputational EfficiencyData-Centric LearningGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringTextGraphRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This paper proposes a post-correction framework called FACTCORRECTOR, which rewrites long-form answers generated by large language models using structured factual feedback;

FACTrial: Factorized Clinical Contrastive Training for Scalable Patient-Trial Retrieval

Xuanren Chen (Beihang University), Shuai Ma (Beihang University)

RetrievalDrug DiscoveryTransformerLarge Language ModelContrastive LearningTextBiomedical DataElectronic Health RecordsBenchmark

🎯 What it does: We constructed FACTRIAL, a patient–clinical trial retrieval framework that utilizes large language models to generate diagnostic factors and perform factorized contrastive learning.

Factual Retrieval in LLMs Is a Redundant, Distributed and Non-Contiguous Process

Hail Hochman (Bar Ilan University), Yoav Goldberg (Bar Ilan University)

RetrievalExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: Studied how entity representations are transformed into retrievable attributes during fact retrieval in large language models, and proposed the concept of 'attribute computation path'.

FactVerse: A Benchmark for Factual Consistency in Interleaved Image–Text Generation

Yubo Shan (Zhengzhou University), Yuanzhuo Wang (Chinese Academy of Sciences)

GenerationData SynthesisReinforcement Learning from Human FeedbackTransformerSupervised Fine-TuningVision Language ModelContrastive LearningImageTextMultimodalityBenchmark

🎯 What it does: Propose the FactVerse benchmark for evaluating factual consistency in interactive image-text generation, and design a three-dimensional evaluation framework (FactJudge, semantic-anchored VQA, and rule constraints) to systematically evaluate 10 baseline models.

Failure Modes in Multi-Hop QA: The Weakest Link Effect and the Recognition Bottleneck

Meiru Zhang (University of Cambridge), Nigel Collier (University of Cambridge)

RecognitionExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringTextChain-of-Thought

🎯 What it does: This paper investigates the positional information bias in LLMs during multi-hop reasoning, revealing that the performance is determined by the lowest visible evidence position through the introduction of multi-focused attention instructions (MFAI) to detect, identify, and synthesize errors, and to uncover the weak-link effect.

Fair-CCD: Mitigating Bias in Large Language Models for Tabular Classification Through Context-Contrastive Decoding

Donghan Liu (East China Normal University), Min Zhang (East China Normal University)

ClassificationTransformerLarge Language ModelPrompt EngineeringContrastive LearningTabularBenchmark

🎯 What it does: Propose a method to regulate fairness in table prediction during large language model inference by using structured bias templates and contrastive decoding

FAIRGAMER: Evaluating Social Biases in LLM-Based Video Game NPCs

Bingkang Shi (Chinese Academy of Sciences), Zhongjiang Yao (Chinese Academy of Sciences)

Federated LearningExplainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningTextMultimodalityBenchmarkChain-of-Thought

🎯 What it does: Proposed the FAIRGAMER benchmark to evaluate social bias in LLM-driven game NPCs under different interaction modes;

FairQE: Multi-Agent Framework for Mitigating Gender Bias in Translation Quality Estimation

Jinhee Jang (Chung-Ang University), YoungBin Kim

Federated LearningExplainability and InterpretabilityComputational EfficiencyRepresentation LearningData-Centric LearningTransformerLarge Language ModelAgentic AIPrompt EngineeringMixture of ExpertsTextMultimodalityBenchmark

🎯 What it does: Proposed the FairQE framework, which uses a multi-agent approach to eliminate gender bias in quality estimation, while considering both gender-ambiguous and explicit scenarios;

Faithful Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution Guidance

Bar Alon (Tel Aviv University), Lior Wolf (Tel Aviv University)

Explainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringText

🎯 What it does: This paper evaluates the empirical faithfulness of text explanations generated by large language models (LLMs) and proposes a training-agnostic method called 'Faithfulness Serum,' which significantly improves the consistency between explanations and internal evidence by injecting PE-LRP heatmaps into the attention layer.

False Friends in the Shell: Unveiling the Emoticon Semantic Confusion in Large Language Models

Weipeng Jiang (Xi'an Jiaotong University), Yang Liu (Nanyang Technological University)

Safty and PrivacyExplainability and InterpretabilityAI Code AssistantTransformerLarge Language ModelPrompt EngineeringTextChain-of-Thought

🎯 What it does: Studied and systematically evaluated the semantic confusion risks generated by large language models when processing ASCII emoticons, finding that they are prone to misinterpreting emoticons as executable code, constructed a corpus containing 3,757 test cases, and conducted quantitative experiments on six mainstream LLMs, revealing a confusion rate as high as 38.6% and a 'silent failure' risk exceeding 90%.

False Friends or Cognates? A Cross-lingual Semantic Ambiguity Evaluation for Galician, Portuguese and Spanish

Marta Vázquez Abuín (Centro Singular de Investigación en Tecnoloxías Intelixentes (CiTIUS) Universidade de Santiago de Compostela), Marcos Garcia (Centro Singular de Investigación en Tecnoloxías Intelixentes (CiTIUS) Universidade de Santiago de Compostela)

ClassificationRecognitionExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBenchmark

🎯 What it does: This paper constructs a cross-lingual word sense disambiguation dataset based on true/false cognates (cognates and false friends) in Galician, Portuguese, and Spanish, and systematically evaluates the performance of various encoders and large language models on this task;

Fast and Accurate Fisher-Guided Quantization via Efficient Kronecker Factorization

Viktoriia A. Chekalina (FusionBrain Lab), Evgeny Frolov (AXXX)

CompressionOptimizationComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelDiffusion modelAuto EncoderContrastive LearningText

🎯 What it does: This paper proposes FastKron, an algorithm that utilizes efficient Kronecker decomposition to compute second-order Hessian information and embeds it into post-training quantization, significantly accelerating the factorization process;

FastV-RAG: Towards Fast and Fine-Grained Video QA with Retrieval-Augmented Generation

Gen Li (University of Electronic Science and Technology of China), Peiyu Liu (University of International Business and Economics)

RetrievalComputational EfficiencyTransformerPrompt EngineeringVision Language ModelVideoTextMultimodalityRetrieval-Augmented Generation

🎯 What it does: Proposes VideoSpeculateRAG, a retrieval-augmented generation framework that combines a lightweight draft model with a powerful verification model for efficient and precise video question answering.

FC-TTS: Style and Timbre Control in Zero-Shot Text-to-Speech with Disentangled Speech Representations

Yoonhyung Lee (Qualcomm Technologies Inc), Jinkyu Lee (Qualcomm Technologies Inc)

GenerationData SynthesisTransformerDiffusion modelFlow-based ModelAuto EncoderContrastive LearningTextAudio

🎯 What it does: Propose FC-TTS, a zero-shot text-to-speech model that can independently control the speaker's voice and speaking style through two different reference sentences.

FedMental: Evaluating Federated Learning for Mental Health Detection from Social Media Data

Nuredin Ali Abdelkadir (University of Minnesota), Stevie Chancellor (University of Minnesota)

ClassificationFederated LearningSafty and PrivacyTransformerSupervised Fine-TuningText

🎯 What it does: Evaluate the feasibility and performance of federated learning and differential privacy federated learning in detecting depression and suicide risk from social media text.

FedProxy: Federated Fine-Tuning of LLMs via Proxy SLMs and Heterogeneity-Aware Fusion

Tao Fan (Hong Kong University of Science and Technology), Qiang Yang (WeBank Co., Ltd)

CompressionFederated LearningComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningText

🎯 What it does: Propose the FedProxy framework for securely and efficiently fine-tuning large language models in a federated learning scenario;

Feeling Right vs. Being Right: How AI Sycophancy Affects Value-Laden Deliberation

Jeongwoo Ryu (Seoul National University), Bongwon Suh (Seoul National University)

Explainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This study explores the impact of sycophantic behavior exhibited by artificial intelligence in value-oriented moral dilemmas on users' decision-making confidence and open-mindedness;

Feeling Rules in Language Models: Mapping Norms of Emotional Appropriateness Across Roles, Institutions, and Intensity

Guangrui Fan (Taiyuan University of Science and Technology), Pan Lihu (Universiti Malaya)

TransformerLarge Language ModelPrompt EngineeringTextBenchmarkChain-of-Thought

🎯 What it does: This paper constructs a sentiment norm benchmark for large language models called FEELING RULES ATLAS, used to measure the model's judgment of emotional appropriateness under different roles, institutions, audiences, and emotional intensity levels.

Fiction Flows: A Replication and Reinterpretation of Narrative Sequentiality

Andrew Piper (McGill University), Federico Pianzola (University of Groningen)

Explainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextReview/Survey PaperRetrieval-Augmented Generation

🎯 What it does: Replicate and reinterpret the conclusion proposed by Sap et al. that 'imagined narratives are smoother than recalled narratives' on a large scale, systematically evaluate the impact of models, formulas, and datasets on the Sequentiality metric, and explore its effectiveness and limitations;

FIGMA: Towards FIne-Grained Music retrievAl

Nishit Anand (University of Maryland), Ramani Duraiswami (University of Maryland)

RetrievalTransformerContrastive LearningTextMultimodalityAudio

🎯 What it does: Proposed the FIGMA model, achieving fine-grained music retrieval capable of understanding long text descriptions containing musical attributes such as rhythm, key, and chord progression.

Figure It Out: Improve the Frontier of Reasoning with Executable Visual States

Meiqi Chen (WeChat AI, Tencent Inc), Jie Zhou (WeChat AI, Tencent Inc)

OptimizationExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelReinforcement LearningPrompt EngineeringVision Language ModelDiffusion modelScore-based ModelGenerative Adversarial NetworkTextMultimodalityBenchmarkChain-of-Thought

🎯 What it does: Proposed and implemented the FIGR system, which can dynamically construct and interactively update geometric figures through executable code during multi-round reasoning, achieving visual reasoning.

FinCall-Surprise: A Large Scale Multi-modal Benchmark for Earning Surprise Prediction

Dong Shu (Northwestern University), Mengnan Du (Chinese University of Hong Kong)

ClassificationTransformerLarge Language ModelAuto EncoderTextMultimodalityBenchmarkFinance RelatedRetrieval-Augmented GenerationAudio

🎯 What it does: Proposes FinCall-Surprise, a multimodal dataset containing text, audio, and presentation slides from corporate earnings call conferences, for predicting earnings surprises;

FinChain: A Symbolic Benchmark for Verifiable Chain-of-Thought Financial Reasoning

Zhuohan Xie (MBZUAI), Preslav Nakov (MBZUAI)

Explainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringTextTabularBenchmarkFinance RelatedRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Constructed the FINCHAIN symbolic financial reasoning benchmark, covering 58 financial topics across 12 domains, using executable Python templates to generate verifiable Chain-of-Thought reasoning chains; and proposed the CHAINEVAL evaluation metric based on semantic + numerical consistency and dynamic time alignment.

FinChart-Bench: Benchmarking Financial Chart Comprehension in Vision-Language Models

Dong Shu (Northwestern University), Mengnan Du (Chinese University of Hong Kong)

TransformerLarge Language ModelSupervised Fine-TuningVision Language ModelImageTextMultimodalityBenchmarkFinance Related

🎯 What it does: Established the FinChart-Bench benchmark for financial chart understanding, containing 1,200 real financial charts with 7,016 manually annotated questions.

Find Your Optimal Teacher: Personalized Data Synthesis via Router-Guided Multi-Teacher Distillation

Hengyuan Zhang (University of Hong Kong), Ngai Wong (University of Hong Kong)

Data SynthesisKnowledge DistillationTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringMixture of ExpertsText

🎯 What it does: Propose the PerSyn method, which dynamically assigns the optimal teacher from a teacher model ensemble for each prompt through a router, generating personalized synthetic data for distilling a small student model.

Fine-Grained Analysis of Shared Syntactic Mechanisms in Language Models

Ryoma Kumon (University of Tokyo), Hitomi Yanaka (University of Tokyo)

Explainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringContrastive LearningText

🎯 What it does: Investigate whether language models share internal mechanisms in syntactic phenomena such as fillers-gap dependencies and negative polarity item licensing, using an unsupervised activation patching method to analyze layers and attention heads.

Fine-Mem: Fine-Grained Feedback Alignment for Long-Horizon Memory Management

Weitao Ma (Harbin Institute of Technology), Bing Qin (Harbin Institute of Technology)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposed the Fine-Mem framework, which uses reinforcement learning to train a learnable memory manager to address memory management issues in large language models during long-sequence tasks.

Fine-tuning vs. In-context Learning in Large Language Models: A Formal Language Learning Perspective

Bishwamittra Ghosh (Max Planck Institute for Software Systems), Evimaria Terzi (Boston University)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningMeta LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringText

🎯 What it does: Compare two learning modes of large language models—fine-tuning and in-context learning—by evaluating their language proficiency and inductive bias through formal language learning tasks;

FineLAP: Taming Heterogeneous Supervision for Fine-grained Language-Audio Pretraining

Xiquan Li (Shanghai Jiao Tong University), Xie Chen (Shanghai Jiao Tong University)

ClassificationRecognitionData SynthesisRetrievalTransformerSupervised Fine-TuningPrompt EngineeringMixture of ExpertsAuto EncoderContrastive LearningTextMultimodalityAudio

🎯 What it does: Propose FineLAP, a heterogeneous supervised training framework that combines segment-level and frame-level audio-text alignment;

FineSteer: A Unified Framework for Fine-Grained Inference-Time Steering in Large Language Models

Zixuan Weng (University of California, Los Angeles), Yuan Tian (University of California, Los Angeles)

Safty and PrivacyExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsText

🎯 What it does: Propose the FineSteer framework to achieve fine-grained adjustment of the internal representations of large language models during inference, thereby enhancing safety and correctness.

Fingerprinting LLMs via Prompt Injection

Yuepeng Hu (Duke University), Neil Zhenqiang Gong (Duke University)

Anomaly DetectionSafty and PrivacyTransformerLarge Language ModelPrompt EngineeringText

🎯 What it does: This paper proposes a framework for LLM lineage detection based on prompt injection, named LLMPrint, which generates unique and robust fingerprints using optimized prompts and verifies models in both gray-box and black-box environments.

FinKario: Event-Enhanced Automated Construction of Financial Knowledge Graph

Xiang Li (Hong Kong University of Science and Technology Guangzhou), Xiaowen Chu (Hong Kong University of Science and Technology Guangzhou)

Graph Neural NetworkTransformerLarge Language ModelTextGraphFinance RelatedRetrieval-Augmented Generation

🎯 What it does: Constructed a dynamic event-driven financial knowledge graph called FinKario, and proposed a two-stage graph retrieval augmented generation framework named FinKario-RAG for stock trend prediction.

FinSight: Towards Real-World Financial Deep Research

Jiajie Jin (Renmin University of China), Zhicheng Dou (Renmin University of China)

TransformerLarge Language ModelAgentic AIVision Language ModelTextMultimodalityBenchmarkFinance RelatedRetrieval-Augmented Generation

🎯 What it does: A Multimodal Automated System for Generating Professional Financial Depth Research Reports, FinSight

Fisher-Driven Adaptive Locating for Knowledge Editing in Large Language Models

Chenghao Xu (Hohai University), Cheng Deng (Hohai University)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningText

🎯 What it does: Propose the FiDAL method, which dynamically locates and edits knowledge layers in large language models by utilizing Fisher information;

FL-MSCL: A Unified Figurative Language Detection Model Driven by Multi-Type Signals and Contrastive Learning

Lu Shijia, Yoshimi Suzuki (University of Yamanashi)

ClassificationFederated LearningExplainability and InterpretabilityRepresentation LearningTransformerPrompt EngineeringContrastive LearningTextRetrieval-Augmented Generation

🎯 What it does: Propose a unified four-class sentence-level framework for detecting metaphors, personification, similes, and literal translations;

FLAIR: Steering LLM Mathematical Problem Solving based on A Fuzzy-Logic-AssIsted Reasoner

Hao Wu (Central China Normal University), Jie Yang (University of Wollongong)

OptimizationExplainability and InterpretabilityTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkChain-of-Thought

🎯 What it does: This paper proposes the FLAIR framework, which utilizes fuzzy theory to diagnose and adaptively control the mathematical reasoning process of LLMs. It can real-time evaluate error types during reasoning and dynamically trigger corresponding error-correction actions, ultimately optimizing rule weights through reinforcement learning.

FLARE: Fine-Grained Length-Aware Routing for Resource-Efficient Heterogeneous LLM Serving

Yujia Fu (Sun Yat-sen University), Yutong Lu (Sun Yat-sen University)

OptimizationComputational EfficiencyTransformerLarge Language ModelText

🎯 What it does: Proposes a fine-grained resource-aware multi-model routing framework called FLARE based on output length, which predicts the latency and cost of each query using its output length and selects the most suitable LLM under the premise of meeting the accuracy threshold through multi-objective optimization.

Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs

Wenrui Zhou (Provable Responsible AI and Data Analytics Lab), Di Wang (Provable Responsible AI and Data Analytics Lab)

Explainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningVideoTextMultimodalityBenchmark

🎯 What it does: Proposed the VISe benchmark, systematically evaluating and countering the sycophancy behavior of Video-LLMs;

FlexGuard: Continuous Risk Scoring for Strictness-Adaptive LLM Content Moderation

Zhihao Ding (Hong Kong Polytechnic University), Jieming Shi (Hong Kong Polytechnic University)

Safty and PrivacyExplainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextBenchmark

🎯 What it does: Propose FlexGuard, an LLM content moderation model that outputs continuous risk scores, and design FlexBench benchmark to evaluate robustness under different strictness levels.

Flip-Flop Consistency: Unsupervised Training for Robustness to Prompt Perturbations in LLMs

Parsa Hejabi (University of Southern California), Morteza Dehghani (University of Southern California)

ClassificationKnowledge DistillationTransformerLarge Language ModelPrompt EngineeringContrastive LearningText

🎯 What it does: Without using labeled data, training on multiple variants of prompts for the same input improves the semantic consistency and performance of large language models when facing prompt perturbations.

Flow-Based Page Unique Semantic Mapping Architecture for Document Visual Question Answering

Haosen Wang (Tianjin University), Zhiyong Feng (Tianjin University)

RetrievalRepresentation LearningTransformerMixture of ExpertsVision Language ModelFlow-based ModelImageTextMultimodalityBenchmark

🎯 What it does: Propose an end-to-end document visual question answering framework called FUMA, which achieves precise evidence page localization and answer generation by learning a bijective mapping between document pages and the unique semantics conditioned on the question.

FlowSearch: Advancing Deep Research with Dynamic Structured Knowledge Flow

Yusong Hu (Shanghai Artificial Intelligence Laboratory), Bo Zhang (Shanghai Artificial Intelligence Laboratory)

Drug DiscoveryGraph Neural NetworkTransformerLarge Language ModelAgentic AITextGraphBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposes FlowSearch, a multi-agent deep research framework based on dynamic structured knowledge flow, which can automatically plan, execute, and dynamically adjust knowledge flow in scientific exploration to achieve parallel reasoning and recursive refinement.

FocalOrder: Focal Preference Optimization for Reading Order Detection

Fuyuan Liu (Unisound AI Technology Co., Ltd.), Junnan Zhu (Chinese Academy of Sciences)

OptimizationTransformerLarge Language ModelVision Language ModelContrastive LearningTextMultimodalityBenchmark

🎯 What it does: This paper proposes the FocalOrder framework to address the Positional Disparity (position discrepancy) problem in reading order detection, emphasizing dynamic attention to intermediate segments of structurally complex documents during training to improve the model's global logical consistency.

Focus-dLLM: Accelerating Long-Context Diffusion LLM Inference via Confidence-Guided Context Focusing

Lingkun Long (Beihang University), Jianlei Yang (Hong Kong University of Science and Technology)

Computational EfficiencyTransformerLarge Language ModelDiffusion modelText

🎯 What it does: This paper proposes a training-agnostic sparse attention framework called Focus‑dLLM, which utilizes confidence predictions from previous steps to predict the tokens to be decoded in the next step, and performs context trimming based on this, significantly accelerating the inference of long-context diffusion LLMs;

Focusing Condition: Inference-Time Self-Contrastive Steering Elicits Better Conditional Text Embeddings in LLMs

Zifeng Cheng (Nanjing University), Qing Gu (Nanjing University)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningText

🎯 What it does: This paper proposes a self-contrast-driven (SCS) method that performs self-contrast during inference to unsupervisedly extract conditional text embeddings from large language models (LLMs).

FoE: Forest of Errors Makes the First Solution the Best in Large Reasoning Models

Kehan Jiang (Peking University), Guojie Song (Peking University)

OptimizationExplainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringTextChain-of-Thought

🎯 What it does: Investigate the error propagation phenomenon in large reasoning models, propose the FoE model, and design the RED method to improve reasoning performance and efficiency.

Follow the Flow: On Information Flow Across Textual Tokens in Text-to-Image Models

Guy Kaplan (Hebrew University of Jerusalem), Roy Schwartz (Hebrew University of Jerusalem)

GenerationExplainability and InterpretabilityData-Centric LearningTransformerPrompt EngineeringVision Language ModelDiffusion modelContrastive LearningImageTextMultimodality

🎯 What it does: Investigate the impact of text encoding on alignment in text-to-image models, using causal masking and visual language models to evaluate the distribution of subword information and cross-item interactions, and propose de-redundant token and patching techniques to improve alignment;

For-Value: Efficient Forward-Only Data Valuation for finetuning LLMs and VLMs

Wenlong Deng (University of British Columbia), Xiaoxiao Li (University of British Columbia)

Explainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningVision Language ModelContrastive LearningImageTextMultimodality

🎯 What it does: Proposes For-Value, an efficient data value evaluation framework that uses only forward inference to quickly assess the impact of each training sample on validation set performance during the fine-tuning of large models.

Foresight Optimization for Strategic Reasoning in Large Language Models

Jessie Wang (Hong Kong Polytechnic University), Wenjie Li (Hong Kong Polytechnic University)

OptimizationFederated LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextChain-of-Thought

🎯 What it does: Propose the FoPO (Foresight Policy Optimization) algorithm, which combines opponent modeling with policy optimization, leveraging self-play and offline reinforcement learning to enhance the strategic reasoning capabilities of large language models in multi-agent environments. Based on this, two dialogue datasets, Cooperative RSA and Competitive Taboo, are designed for training and evaluation.

Forest Before Trees: Latent Superposition for Efficient Visual Reasoning

Yubo Wang (Fudan University), Yuhan Liu (MBZUAI)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBenchmark

🎯 What it does: Developed the Laser model, which transforms visual reasoning into continuous latent superpositions, achieving hierarchical reasoning with 'forest before trees' through dynamic window alignment learning.

FOREVER: Forgetting Curve-Inspired Memory Replay for Language Model Continual Learning

Yujie Feng (Hong Kong Polytechnic University), Xiao-Ming Wu (Hong Kong Polytechnic University)

Computational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposes FOREVER, a replay-based continual learning framework, which defines model time by tracking the magnitude of parameter updates and maps the intervals of the human forgetting curve to model time, thereby replaying at appropriate times.

Forget What Matters, Keep the Rest: Selective Unlearning of Informative Tokens

Seunghee Koh (Korea Advanced Institute of Science and Technology), Junmo Kim (Korea Advanced Institute of Science and Technology)

Federated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringContrastive LearningTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposed an entropy-based token weighting mechanism (ETW), which imposes greater penalties on information-rich tokens during the selective forgetting process of large language models, thereby better balancing forgetting effectiveness and model performance.

FormalScience: Scalable Human-in-the-Loop Autoformalisation of Science with Agentic Code Generation in Lean

Jordan Meadows (Independent Researcher), Andre Freitas

Reinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkPhysics Related

🎯 What it does: Proposes a human-in-the-loop semi-automated pipeline called FormalScience for converting informal reasoning in physics into formal proofs in Lean4;

FormulaSPIN: Self-Play Fine-Tuning for Natural Language to Spreadsheet Formula Generation

Cy Xie

GenerationData SynthesisAI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextTabularBenchmark

🎯 What it does: This paper proposes FormulaSPIN, a self-play fine-tuning framework that achieves iterative self-improvement without additional annotations by leveraging formula executability;

Frame-Semantic Knowledge Injection for Event-Level Inference in LLMs

Shahid Iqbal Rai (University of Rome Tor Vergata), Roberto Basili (University of Rome Tor Vergata)

Explainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: Injecting the complete semantic knowledge of FrameNet into large language models using LoRA, and evaluating its effectiveness through NLI and SRL tasks.

Framing Political Bias in Multilingual LLMs Across Pakistani Languages

Afrozah Nadeem (Macquarie University), Usman Naseem (Macquarie University)

Explainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringTextMultimodality

🎯 What it does: Evaluated political bias in 13 large language models across five Pakistani languages by combining the Political Compass Test with a multi-layer narrative framework analysis.

Frankentext: Stitching random text fragments into long-form narratives

Chau Minh Pham (University of Maryland College Park), Mohit Iyyer (University of Maryland College Park)

GenerationData SynthesisTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: This work proposes and implements a long-form narrative generation framework called Frankentexts, which generates high-quality stories by having LLMs select, arrange, and connect text segments from a large collection of human-written fragments, resulting in stories that are mostly composed of original text.

Free Energy-Driven Reinforcement Learning with Adaptive Advantage Shaping for Unsupervised Reasoning in LLMs

Yiming Huang (Harbin Institute of Technology), Chuanyi Liu (Harbin Institute of Technology)

TransformerLarge Language ModelReinforcement LearningScore-based ModelTextMultimodalityBenchmark

🎯 What it does: This paper proposes FREIA, an unsupervised reinforcement learning framework that leverages the free energy principle and adaptive advantage shaping to enhance the reasoning capabilities of large language models.

From "Aha Moments" to Controllable Thinking: Toward Meta-Cognitive Reasoning in LRMs via Decoupled Reasoning and Control

Rui Ha (Beijing University of Posts and Telecommunications), Sen Su (Beijing University of Posts and Telecommunications)

Explainability and InterpretabilityComputational EfficiencyMeta LearningTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextBenchmark

🎯 What it does: Proposes the MERAN meta-cognitive framework, decoupling reasoning and control, enabling the model to self-monitor and dynamically decide to continue, backtrack, or terminate during the reasoning process.

From \log \pi to \pi: Taming Divergence in Soft Clipping via Bilateral Decoupled Decay of Probability Gradient Weight

Xiaoliang Fu (Meituan), Ke Zeng (Meituan)

OptimizationRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextChain-of-Thought

🎯 What it does: Propose the DGPO algorithm, which addresses the gradient divergence and exploration limitations caused by soft clipping in RLVR through the use of probabilistic gradients and a bidirectional attenuation soft clipping method; significantly improves mathematical reasoning performance on large language models.

From 1,000,000 Users to Every User: Scaling Up Personalized Preference for User-level Alignment

Jia-Nan Li (Renmin University of China), Rui Yan (Wuhan University)

Recommendation SystemExplainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringText

🎯 What it does: Proposed an scalable personalized alignment framework and constructed the ALIGNX large-scale dataset covering a 90-dimensional psychological and behavioral preference space (approximately 1.3M samples), thus achieving precise alignment at the user level;

From Answers to Arguments: Toward Trustworthy Clinical Diagnostic Reasoning with Toulmin-Guided Curriculum Goal-Conditioned Learning

Chen Zhan (Shanghai University of Engineering Science), Xihe Qiu (Shanghai University of Engineering Science)

Explainability and InterpretabilityDrug DiscoveryTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextBiomedical DataElectronic Health Records

🎯 What it does: Proposes the CGCL course-based goal-oriented learning framework and the T-Eval evaluation tool, enhancing the transparency and reliability of LLMs in clinical diagnosis reasoning.

From AR to Diffusion: Efficiently Adapting Large Language Models with Strictly Causal and Elastic Horizons

Xiangyu Ma (Wuhan University), Lefei Zhang (Wuhan University)

GenerationComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelSupervised Fine-TuningDiffusion modelContrastive LearningTextStochastic Differential Equation

🎯 What it does: Propose the FLUID framework, which transfers pre-trained autoregressive (AR) language models into diffusion models, achieving efficient parallel text generation through strict causal alignment and entropy-driven elastic horizons.

From Charts to Code: A Hierarchical Benchmark for Multimodal Models

Jiahao Tang (Central South University), Alex Jinpeng Wang (Central South University)

GenerationAI Code AssistantTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextMultimodalityTabularBenchmarkRetrieval-Augmented Generation

🎯 What it does: Constructed a hierarchical chart generation benchmark called Chart2Code, evaluating the ability of multimodal models to generate executable drawing code from visual charts or long tables.

From Competition to Synergy: Unlocking Reinforcement Learning for Subject-Driven Image Generation

Ziwei Huang (Zhejiang University), Leilei Gan (Zhejiang University)

GenerationTransformerReinforcement LearningPrompt EngineeringDiffusion modelImageText

🎯 What it does: Proposed an online reinforcement learning framework called Customized-GRPO, specifically designed to address the competitive degradation problem between identity fidelity and prompt compliance in agent-driven image generation.

From Completion to Editing: Unlocking Context-Aware Code Infilling via Search-and-Replace Instruction Tuning

Jiajun Zhang (USTC), Junyang Lin (Alibaba Group)

AI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextSequentialBenchmark

🎯 What it does: Proposes the Search-Replace Insertion (SRI) framework, transforming code completion from static insertion to context-aware editing, and achieving safe and efficient real-time completion on chat LLMs.

From Curated Data to Scalable Models: Continual Pre-training of Dense and MoE Large Language Models for Tibetan

Lei Yang (Tianjin University), Deyi Xiong (Tianjin University)

Federated LearningComputational EfficiencyKnowledge DistillationRepresentation LearningData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringMixture of ExpertsTextMultimodalityBenchmark

🎯 What it does: A 72GB high-quality Tibetan corpus was constructed, and through methods such as continuous pre-training, vocabulary expansion, and instruction fine-tuning, a multi-lingual large language model with 7B dense and 50B Mixture-of-Experts was trained.

From Data-Centric to Sample-Centric: Enhancing LLM Reasoning via Progressive Optimization

Xinjie Chen (Zhejiang University), Xinggao Liu (Zhejiang University)

OptimizationData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmark

🎯 What it does: Propose a sample-based RLVR framework called LPPO, combining dynamic gradient weighting based on learning progress and prefix-guided sampling, to enhance the performance of large language models on reasoning tasks.

From Detection to Understanding: Multi-Turn Reasoning for Video Misinformation Analysis

Zhi Zeng (Xi'an Jiaotong University), Zihan Ma (Xi'an Jiaotong University)

Explainability and InterpretabilityTransformerLarge Language ModelAgentic AIVideoTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-ThoughtAudio

🎯 What it does: This paper proposes the MISVIDEOQA multi-round video misinformation analysis benchmark and designs the MISAGENT multi-agent framework to help models achieve end-to-end reasoning from perception to motivation, emotion, and persuasiveness in videos.

From Experts to Bases: Orthogonal Subspace Mixture for Continual Multimodal Instruction Tuning

Pei Chen (Zhejiang University), Lingyun Sun (Zhejiang University)

Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringMixture of ExpertsVision Language ModelImageTextMultimodalityBenchmark

🎯 What it does: Propose the MoBLoRA framework, which suppresses catastrophic forgetting and improves parameter efficiency in multi-modal continual instruction tuning by utilizing a shared orthogonal basis pool and task-specific hybrid matrices.

From Factuality to Meta-Factivity: A Cognitive Blueprint for Trustworthy LLMs

Daohuan Liu (Huazhong University of Science and Technology), Xuri Tang (Huazhong University of Science and Technology)

Explainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed the Meta-Factivity framework (MFF), shifting from static classification in event factuality judgment to dynamic metacognitive regulation;

From Form to Logic: Masked Reconstruction and Reasoning Distillation for Short Video Fake News Detection

Qingyan Wang (Northwestern Polytechnical University), Yaxiong Wang (Hefei University of Technology)

Anomaly DetectionKnowledge DistillationRepresentation LearningTransformerLarge Language ModelContrastive LearningVideoTextMultimodalityRetrieval-Augmented Generation

🎯 What it does: Proposed the PCDD framework, combining perception and cognition dual drivers to detect fake news in short videos.

From ID to LLM: Rethinking Representation Learning for Recommendation

Song-Li Wu (Tsinghua University), Xianquan Wang (Tsinghua University)

Recommendation SystemRepresentation LearningTransformerLarge Language ModelReinforcement LearningContrastive LearningTextSequential

🎯 What it does: Proposes the Profile-then-Embedding (PtE) framework, which first uses a large language model to perform bidirectional reasoning on user-item interactions to generate semantic user/item 'profiles', and then maps these profiles into low-dimensional, interpretable recommendation vectors through a task-aligned embedding phase.