arXivSub Start free trial

ACL 2026 Papers — Page 18

Annual Meeting of the Association for Computational Linguistics · 2296 papers

REVEALER: Reinforcement-Guided Visual Reasoning for Element-Level Text-Image Alignment Evaluation

FuLin Shi, Binchen

OptimizationExplainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBenchmark

🎯 What it does: Propose the REVEALER framework to achieve element-level alignment evaluation for text-to-image generation, using a structured grounding-reasoning-conclusion process

Revealing Procedural Reasoning Structures in Chain-of-Thought Training via Span-Level Gradient Organization

Jia Liu (South China University of Technology), Min Chen (Huazhong University of Science and Technology)

Explainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningTextChain-of-Thought

🎯 What it does: This paper proposes the Gradient-based Structural Developer (GSD) framework by analyzing the gradient dynamics of large language models during the chain-of-thought (CoT) fine-tuning process in an unsupervised manner. The framework can cluster and align gradients at the span level, revealing the implicit reasoning program structures that emerge during training.

Revealing the Seen, Imagining the Beyond: A Survey of Image-Grounded Chain-of-Thought Reasoning in Multimodal LLMs

Qihua Dong (Northeastern University), Yun Fu (Northeastern University)

Explainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerSupervised Fine-TuningReinforcement LearningPrompt EngineeringVision Language ModelImageTextMultimodalityReview/Survey PaperBenchmarkChain-of-Thought

🎯 What it does: Reviews methods, classification, and evaluation of image-driven chain-of-thought (IG-CoT) in multimodal large language models;

ReviewGrounder: Improving Review Substantiveness with Rubric-Guided, Tool-Integrated Agents

Zhuofeng Li (Texas Aandm University), Yu Zhang (Texas Aandm University)

TransformerLarge Language ModelAgentic AIPrompt EngineeringTextReview/Survey PaperBenchmarkRetrieval-Augmented Generation

🎯 What it does: Built a rubric-based evaluation benchmark called REVIEWBENCH, and proposed a multi-agent framework called REVIEWGROUNDER to generate review comments supported by solid evidence.

Revisit What You See: Revealing Visual Semantics in Vision Tokens to Guide LVLM Decoding

Beomsik Cho (Yonsei University), Jaehyung Kim (Yonsei University)

Explainability and InterpretabilityComputational EfficiencyTransformerPrompt EngineeringMixture of ExpertsVision Language ModelImageTextMultimodalityBenchmark

🎯 What it does: This paper proposes a decoding strategy called ReVisiT that does not require additional training, leveraging semantic information from visual tokens to guide large vision-language models in generating more accurate and less hallucinatory text.

Revisiting a Pain in the Neck: A Semantic Reasoning Benchmark for Language Models

Yang Liu (University of Science and Technology Beijing), Qiankun Liu (University of Science and Technology Beijing)

TransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Construct and release the SEMANTICQA benchmark to evaluate the ability of language models in processing the semantics of multi-word expressions (MWE);

Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation

Wei-Cheng Tseng (University of Texas at Austin), Dong Yu (Tencent AI Lab Seattle)

ClassificationRecognitionRetrievalRepresentation LearningTransformerVision Language ModelContrastive LearningTextMultimodalityAudio

🎯 What it does: This paper constructs CaptionStew by aggregating 10.7M multi-source audio-text pairs, and systematically compares the performance of contrastive learning and captioning as two pre-training objectives on speech, music, and environmental sound tasks.

Revisiting Evaluation of Question Answering Systems in Low-Resource Indic Languages: Bridging Human and Metric Alignment

Anuj Kumar (Indian Institute of Technology Jammu), Virendra Singh (Indian Institute of Technology Bombay)

Recommendation SystemData-Centric LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: This paper proposes and verifies a multi-dimensional evaluation metric called LRM²QAS for evaluating question-answering systems in low-resource Indian languages.

Revisiting Metric Reliability for Fine-grained Evaluation of Machine Translation and Summarization in Indian Languages

Amir Hossein Yari (Sharif University of Technology), Fajri Koto (Mohamed bin Zayed University of Artificial Intelligence)

TransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: This paper proposes the ITEM (Indian Text Evaluation Metrics Testbed) benchmark, systematically evaluating the alignment between 29 automatic metrics for machine translation (MT) and text summarization (TS) and human assessments in six Indian languages (Hindi, Bengali, Marathi, Gujarati, Tamil, Telugu).

Revisiting Model Interpolation for Efficient Reasoning

Taiqiang Wu (University of Hong Kong), Ngai Wong (University of Hong Kong)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelPrompt EngineeringTextBenchmarkChain-of-Thought

🎯 What it does: This paper achieves efficient inference by directly interpolating the weights of the 'Instruct' model and the 'Thinking' model (MI), systematically analyzing the impact of the interpolation coefficient λ on model behavior, proposing a three-stage evolutionary paradigm, and proving that MI can outperform more complex model fusion methods on multiple benchmarks.

Revisiting Non-Verbatim Memorization in Large Language Models: The Role of Entity Surface Forms

Yuto Nishida (Nara Institute of Science and Technology), Taro Watanabe (Nara Institute of Science and Technology)

Explainability and InterpretabilityKnowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Constructed a entity question-answering dataset called RedirectQA based on Wikipedia redirect information, and used it to evaluate the factual memory consistency of LLMs under different entity surface forms.

Revisiting the Reliability of Language Models in Instruction-Following

Jianshuo Dong (Tsinghua University), Han Qiu (Tsinghua University)

Explainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposes the concept of 'nuance-oriented reliability', defines the reliability metric reliable@k, and evaluates the robustness of large language models (LLMs) by generating 'a set of similar but subtly different cousin prompts' through automated data augmentation.

Revitalizing Black-Box Interpretability: Actionable Interpretability for LLMs via Proxy Models

Junhao Liu (Peking University), Xin Zhang (Peking University)

OptimizationExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringContrastive LearningText

🎯 What it does: Propose a Screen-and-Apply framework based on a proxy model, which generates high-fidelity post-hoc explanations on large models using cost-effective proxy LLMs, and verifies its operability in practical optimization tasks such as prompt compression and detoxification examples.

Reviving Iterative Refinement in Diffusion-based NER with an Initializer-Restorer Approach

Long Hai Trieu, Makoto Miwa (Cancer Research UK Cambridge Institute, University of Cambridge)

RecognitionTransformerDiffusion modelContrastive LearningTextBiomedical Data

🎯 What it does: Studied the iterative refinement mechanism of diffusion models in Named Entity Recognition (NER), finding that models often perform only one-time generation when trained with EMA; to address this, we proposed a two-stage framework called Initializer-Restorer, which first generates candidate entities using a conventional NER model and then refines them through a multi-step conditional diffusion model;

Reward Alignment Optimization: A Direct Point-wise Alignment Approach

Zelin Li (Beijing Institute of Technology), Yangen Hu (Meituan Inc)

OptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringText

🎯 What it does: Propose Reward Alignment Optimization (RAO), achieving point-wise direct alignment through an explicit reward model and prefix consistency, avoiding the likelihood shift problem in DPO.

Reward Modeling for Scientific Writing Evaluation

Furkan Şahinuç (Technical University of Darmstadt), Iryna Gurevych (Technical University of Darmstadt)

Reinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextReview/Survey PaperBenchmark

🎯 What it does: Proposed a two-stage reward model, SCIRM and SCIRM-REF, for scientific writing evaluation, capable of multi-dimensional assessment according to task specifications and self-reflective error correction.

RExBench: Can coding agents autonomously implement AI research extensions?

Nicholas Edwards (University of Vienna), Najoung Kim (Boston University)

AI Code AssistantTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposed and implemented the REXBENCH benchmark to evaluate the ability of LLM agents to autonomously complete AI research expansion without human intervention.

REZE: Representation Regularization for Domain-adaptive Text Embedding Pre-finetuning

Seungmin Lee (Yonsei University), Mingi Sung (PwC Korea)

Domain AdaptationRepresentation LearningTransformerSupervised Fine-TuningContrastive LearningText

🎯 What it does: Proposed a representation regularization framework called REZE, which explicitly controls the shift in the representation space during the pre-fine-tuning stage of text embedding models by performing feature decomposition on the anchor positive sample relationship vectors and applying adaptive soft shrinkage to the bias introduced by the task.

RFS-Guard: Detecting Reasoning Hallucinations via Cross-Phase Routing Focus in Large Reasoning Models

Zihang Liu (Zhejiang University), Haishuai Wang (Zhejiang University)

Anomaly DetectionExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextChain-of-Thought

🎯 What it does: This paper proposes a hallucination detection framework called RFS-Guard for large reasoning models, which can locate errors in the reasoning process without the need for external verification or multiple sampling.

Rhetorical Questions in LLM Representations: A Linear Probing Study

Louie Hong Yao (Independent Researcher), Tianyu Jiang (University Of Cincinnati)

Explainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelText

🎯 What it does: This work utilizes linear probes to analyze how large language models internally encode rhetorical questions, exploring the separability and cross-dataset transferability of rhetorical questions versus informational questions in the representation space. At the same time, by comparing trained (Logistic, SVM) and untrained (diffMean) linear probes, it reveals that rhetorical intent in the model is not a single linear direction, but rather a multidimensional and semantically diverse expression. It also conducts quantitative and qualitative analyses of differences in hierarchical, ranking, and directional aspects among different probe directions.

Right at My Level: A Unified Multilingual Framework for Proficiency-Aware Text Simplification

Jinhong Jeong (Yonsei University), Youngjae Yu (Seoul National University)

Reinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextMultimodality

🎯 What it does: Proposes a unified multilingual text simplification framework called Re-RIGHT, which does not require parallel corpora and supports four languages. It can control vocabulary difficulty according to learners' standards such as CEFR, JLPT, TOPIK, and HSK.

RIMRULE: Improving Tool-Using Language Agents via MDL-Guided Rule Learning

Xiang Gao (Intuit AI Research), Kamalika Das (Intuit AI Research)

Explainability and InterpretabilityRobotic IntelligenceReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposes a rule learning framework called RIMRULE based on failure tracing, which enhances the performance of LLMs in tool usage tasks by dynamically injecting interpretable rules distilled from errors during inference.

RISK: A Framework for GUI Agents in E-commerce Risk Management

Renqi Chen (Ant International Ant Group), Shuai Chen (Ant International Ant Group)

Recommendation SystemAnomaly DetectionAutonomous DrivingOptimizationData-Centric LearningRobotic IntelligenceTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringMixture of ExpertsVision Language ModelTextMultimodalitySequentialBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Built the RISK framework, which includes a GUI agent specifically designed for e-commerce risk management, covering three modules: RISK-Data (large-scale single-step and multi-step interactive trajectories), RISK-Bench (benchmark task set), and RISK-R1 (reinforcement learning fine-tuning method based on GRPO).

River-LLM: Large Language Model Seamless Exit Based on KV Share

Yingtao Shen (Shanghai Jiao Tong University), An Zou (Shanghai Jiao Tong University)

Computational EfficiencyKnowledge DistillationTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: Propose River-LLM, a training-free seamless token-level early exit framework, which addresses the KV Cache Absence issue in decoder-only LLMs through KV-Shared Exit River;

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization

Yihong Dong (Peking University), Ge Li (Alibaba Group)

OptimizationAI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringText

🎯 What it does: Propose RL-PLUS, a hybrid strategy optimization framework that integrates internal exploration with external data, aimed at enhancing the reasoning capabilities of large language models (LLMs) in reinforcement learning with human feedback (RLVR), and addressing the capability boundary collapse problem.

RLSeek: Evidence-Grounded Reasoning for RAG Hallucination Detection

Zhaoheng Huang (Renmin University of China), Fangzhao Wu (Microsoft)

Explainability and InterpretabilityData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Improve the accuracy of hallucination span detection by introducing a reinforcement learning-based evidence citation framework into the retrieval-augmented generation (RAG) system.

Robertha: Eigenspectrum Regularized Attention for Robust Natural Language Understanding

Andreia Podasca (Drexel University), Anup Das (Drexel University)

Adversarial AttackTransformerText

🎯 What it does: Analyzes the 'asymmetric fragility' where low-norm vocabulary in the embedding layer is more susceptible to damage under noise attacks, and proposes an attention mechanism called Robertha. It utilizes modern Hopfield networks and spectral regularization to achieve adaptive iterative reconstruction of damaged embeddings, thereby significantly enhancing the robustness of natural language understanding models in noisy environments.

RoboFailRing: Retrieval-Augmented and Language Grounding Failure Detection for VLM-enabled Robotic Manipulation

Chenduo Ying (Zhejiang University), Peng Cheng (Zhejiang University)

Anomaly DetectionRobotic IntelligenceTransformerVision Language ModelContrastive LearningImageVideoTextMultimodalityRetrieval-Augmented Generation

🎯 What it does: Proposes the RoboFailRing framework for rapidly detecting failures in robotic manipulation tasks and performing language-based failure cause reasoning through generating grounded failure reports.

RoBSA: RoPE-based Blockwise Sparse Multi-head Latent Attention

Xinyu Shi (Tsinghua University), Wenguang Chen (Tsinghua University)

Computational EfficiencyTransformerLarge Language ModelText

🎯 What it does: Proposed a RoPE-based Blockwise Sparse Attention (RoBSA) algorithm tailored for the MLA architecture, achieving sparse attention during the decoding phase to reduce computational costs

Robust Membership Inference for Large Language Models under Adversarial Generative Corruption

Yuanhong Huang (Beijing University of Posts and Telecommunications), Tao Qi (Beijing University of Posts and Telecommunications)

Safty and PrivacyExplainability and InterpretabilityAdversarial AttackTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsTextRetrieval-Augmented Generation

🎯 What it does: Studied the interference caused by AI-generated text on membership inference attacks (MIA) in large language models, and proposed a Mixture-of-Experts framework called MoMIA to enhance membership inference robustness under adversarial generated text.

Robust Tool Use via Fission-GRPO: Learning to Recover from Execution Errors

Zhiwei Zhang (Chinese University of Hong Kong), Kam-Fai Wong (Chinese University of Hong Kong)

Robotic IntelligenceAI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringMixture of ExpertsTextSequentialBenchmark

🎯 What it does: Through the FISSION-GRPO framework, execution errors are converted into on-policy corrective supervision, and diagnostic feedback is generated using an error simulator. Furthermore, error recovery capability in multi-round tool calls is enhanced through a fission-based resampling approach.

Rose-SQL: Role-State Evolution Guided Structured Reasoning for Multi-Turn Text-to-SQL

Le Zhou (National University of Defense Technology), Boyan Xu (Guangdong University of Technology)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningAI Code AssistantTransformerLarge Language ModelPrompt EngineeringTextTabularBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposes the Rose-SQL framework, which utilizes prompt learning with a small large-scale inference model to achieve multi-turn Text-to-SQL generation, avoiding the costs of inference and fine-tuning of large models;

ROSE: An Intent-Centered Evaluation Metric for NL2SQL

Wenqi Pei (Hong Kong University of Science and Technology), Yuyu Luo (Hong Kong University of Science and Technology)

AI Code AssistantTransformerLarge Language ModelPrompt EngineeringGenerative Adversarial NetworkTextTabularBenchmark

🎯 What it does: This paper proposes ROSE, an intent-oriented NL2SQL evaluation metric, which determines the semantic correctness of predicted SQL through an adversarial Prover-Refuter cascade.

Route to Rome Attack: Directing LLM Routers to Expensive Models via Adversarial Suffix Optimization

Haochun Tang (Jilin University), Enyan Dai (Hong Kong University of Science and Technology (Guangzhou))

OptimizationComputational EfficiencyAdversarial AttackTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Study black-box LLM router attacks, design a suffix optimization method to induce the router to route simple queries to expensive high-performance models.

RouteMoA: Dynamic Routing without Pre-Inference Boosts Efficient Mixture-of-Agents

Jize Wang (Shanghai Jiao Tong University), Dacheng Tao

Computational EfficiencyAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringMixture of ExpertsText

🎯 What it does: The paper proposes RouteMoA, an efficient hybrid agent framework that selects high-potential LLMs and performs multi-round collaborative reasoning without executing full inference through dynamic routing.

Routing with Generated Data: Annotation-Free LLM Skill Estimation and Expert Selection

Tianyi Niu (University Of North Carolina Chapel Hill), Mohit Bansal (University Of North Carolina Chapel Hill)

Recommendation SystemExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsContrastive LearningTextRetrieval-Augmented Generation

🎯 What it does: In scenarios lacking human-annotated data, we propose using query and answer data generated by generative LLMs for LLM routing, constructing a routing framework based on generated data called RGD. Based on this, we propose an annotation-free query-only routing method called CASCAL, which can identify fine-grained skills and experts of models through model consistency evaluation and hierarchical clustering.

RPC-Bench: A Fine-grained Benchmark for Research Paper Comprehension

Yelin Chen (Xinjiang University), Juanzi Li (Tsinghua University)

TransformerLarge Language ModelPrompt EngineeringTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: Constructed RPC-Bench, a large-scale benchmark for research paper understanding, consisting of 61,300 QA pairs from 4,150 papers based on real OpenReview review-response pairs, and designed a fine-grained task classification framework along with an LLM-human collaborative annotation and LLM-as-Judge evaluation framework.

RRAtention: Dynamic Block Sparse Attention via Per-Head Round-Robin Shifts for Long-Context Inference

Siran Liu (Baidu Inc.), Chao Yang (Peking University)

Computational EfficiencyTransformerLarge Language ModelMixture of ExpertsVideoTextMultimodality

🎯 What it does: Proposed a dynamic sparse attention mechanism called RRAttention, which achieves efficient sparse computation for long context reasoning through head-round-robin sampling, reducing complexity;

RSDA: Restoring Stale Data Affinity via Dynamic Renovation Strategy for Mitigating Data Scarcity

Yidan Liang (Zhejiang Normal University), Jiajie Xu (Southeast University)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextBenchmark

🎯 What it does: Propose the RSDA framework, which quantifies the reconstruction value of samples through potential entropy and dynamically selects component-level renovation strategies to enhance the adaptability of scarce data;

rSIM: Incentivizing Reasoning Capabilities of LLMs via Reinforced Strategy Injection

Sijia Chen (University of Toronto), Di Niu (University of Alberta)

OptimizationComputational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkChain-of-Thought

🎯 What it does: Designed and evaluated a reinforcement strategy injection mechanism called rSIM, enabling large language models to dynamically inject human-designed strategies during reasoning through a small planner, significantly enhancing reasoning capabilities.

RSMeM: Knowledge-Enhanced Memory Evolution for Remote Sensing Agents with Systematic Evaluation

Bingxian Wu (Institute of Geographic Sciences and Natural Resources Research, Chinese Academy of Sciences), Maosong Sun (Tsinghua University)

Autonomous DrivingOptimizationComputational EfficiencyKnowledge DistillationRepresentation LearningRobotic IntelligenceReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringVision-Language-Action ModelImageTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Designed the RSMeM mechanism, leveraging knowledge-enhanced memory evolution to improve the performance of remote sensing agents in tool usage and answer quality.

RST-Guarder: Enhancing Long-Context Robustness for Safeguards via RST Parsing and Probabilistic Inference

Xu Zhang (Peking University), Xiaojun Wan (Peking University)

Safty and PrivacyExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose a method called RST-Guarder for enhancing the safety detection of long texts during inference. It utilizes Rhetorical Structure Theory (RST) to parse and construct discourse hierarchy structures, and performs hierarchical probabilistic reasoning on this basis, thereby improving the detection accuracy of existing Guardrail models on long texts.

RubricBench: Aligning Model-Generated Rubrics with Human Standards

Junyi Zhou (City University of Hong Kong), Chen Ma (City University of Hong Kong)

Data-Centric LearningAI Code AssistantLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Construct and use RubricBench benchmark to evaluate the reliability of models in rubric-based assessment.

RubricHub: A Comprehensive and Highly Discriminative Rubric Dataset via Automated Coarse-to-Fine Generation

Sunzhu Li (Li Auto Inc.), Chen Wei

GenerationData SynthesisReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringMixture of ExpertsTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: This paper proposes an automatic coarse-to-fine rubric generation framework and constructs a large-scale multi-domain RubricHub dataset for fine-grained evaluation of open-ended generation tasks; subsequently, it implements a two-stage post-training process (RuFT+RuRL) using this dataset, significantly improving the performance of LLMs across multiple domains.

RV-HATE: Reinforced Multi-Module Voting for Implicit Hate Speech Detection

Yejin Lee (Yonsei University), Yo-Sub Han (Yonsei University)

ClassificationExplainability and InterpretabilityComputational EfficiencyTransformerReinforcement LearningPrompt EngineeringMixture of ExpertsContrastive LearningText

🎯 What it does: Designed the RV-HATE framework, which detects implicit hate speech by using reinforcement learning to dynamically weight votes from four specialized modules.

S^4: Operationalizing Speech Act Theory for Strategic Semi-Structured Psychiatric Interview

Guanqun Bi (Tsinghua University), Minlie Huang (Tsinghua University)

OptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningTextBiomedical DataReview/Survey PaperRetrieval-Augmented Generation

🎯 What it does: Designed and implemented the S4 framework based on speech act theory to optimize strategies and language generation in semi-structured interviews in psychiatry, and achieved long-term therapeutic effectiveness optimization through reinforcement learning

S2G-RAG: Structured Sufficiency and Gap Judging for Iterative Retrieval-Augmented QA

Minghan Li (Soochow University), Guodong Zhou (Soochow University)

RetrievalExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: Proposes S2G-RAG, an iterative question-answering framework that controls multi-round retrieval through explicit sufficiency judgment and structured gap prediction.

S2O: Early Stopping for Sparse Attention via Online Permutation

Yu Zhang (ByteDance), Xing Wang (ByteDance)

Computational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: Propose an online swapping and early stopping mechanism called S2O to achieve efficient inference of sparse attention in long contexts;

S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models

Feng Jiang (Shenzhen University of Advanced Technology), Haizhou Li (Chinese University of Hong Kong, Shenzhen)

TransformerLarge Language ModelPrompt EngineeringContrastive LearningTextMultimodalityBenchmarkRetrieval-Augmented GenerationAudio

🎯 What it does: Propose the S2S-Arena benchmark to evaluate the ability of speech-to-speech models in instruction following and phonetic expression.

Saber: Efficient Sampling with Adaptive Acceleration and Backtracking Enhanced Remasking for Diffusion Language Model in Code Generation

Yihong Dong (Peking University), Ge Li (Peking University)

GenerationAI Code AssistantTransformerPrompt EngineeringDiffusion modelTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose Saber, a training-agnostic sampling algorithm for efficient sampling in diffusion language models on code generation tasks.

SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization

Wenxi Chen (Shanghai Jiao Tong University), Xie Chen (Shanghai Jiao Tong University)

CompressionTransformerSupervised Fine-TuningAuto EncoderGenerative Adversarial NetworkContrastive LearningAudio

🎯 What it does: Proposed a semantic-acoustic dual-stream neural speech codec (SAC), which decouples and separately optimizes semantic content and acoustic details by using a pre-trained semantic tokenizer and an acoustic quantization module respectively, ultimately achieving high-quality reconstruction and semantic expression;

SACTOR: LLM-Driven Correct and Idiomatic C to Rust Translation with Static Analysis and FFI-Based Verification

Tianyang Zhou (University of Illinois Urbana-Champaign), Varun Chandrasekaran (University of Illinois Urbana-Champaign)

AI Code AssistantTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose a two-stage LLM-driven C-to-Rust translation tool called SACTOR, which first generates non-idiomatic Rust that preserves the interface, and then rewrites it into safe and idiomatic Rust code through static analysis and FFI verification.

SAD: A Large-Scale Strategic Argumentative Dialogue Dataset

YongKang Liu (Northeastern University), Hinrich Schuetze (LMU Munich)

GenerationData SynthesisTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextSequentialRetrieval-Augmented Generation

🎯 What it does: Constructed and publicly released a large multi-turn argumentative dialogue dataset called SAD, and defined a strategy-controlled argument generation task on this dataset.

SADA: Bridging In-Context Learning and Fine-Tuning via State-Aligned Distillation Adapters

Wenhao Gao (Tencent Inc.), Zhou Xiao

Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringContrastive LearningTextRetrieval-Augmented Generation

🎯 What it does: Propose the SADA (State-Aligned Distillation Adapter) framework, which dynamically compresses long contextual information into low-rank parameter updates through attention output, replacing the full KV cache during inference, thus balancing the flexibility of ICL and the efficiency of FT.

Safe-FedLLM: Delving into the Safety of Federated Large Language Models

Mingxiang Tao (Hainan University), Xiangyan Tang (Hainan University)

Federated LearningSafty and PrivacyTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningText

🎯 What it does: Explores the security of large language models in a federated learning environment, proposing a lightweight attack detection and defense framework called Safe‑FedLLM based on LoRA updates, and verifying its robustness under various attack scenarios.

SAFE-QAQ: End-to-End Slow-Thinking Audio-Text Fraud Detection via Reinforcement Learning

Peidong Wang (Northeastern University), Daling Wang (Northeastern University)

Anomaly DetectionExplainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextMultimodalityChain-of-ThoughtAudio

🎯 What it does: Built an end-to-end slow-thinking audio-text fraud detection framework SAFE-QAQ based on reinforcement learning, which can directly process raw audio and provide interpretable multi-step reasoning and final judgment.

SafeAgent: Safeguarding LLM Agents via an Automated Risk Simulator

Xueyang Zhou (Huazhong University of Science and Technology), Lichao Sun (Lehigh University)

Safty and PrivacyReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningAgentic AIPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose the SafeAgent framework, which uses an automated risk simulator to align large language model (LLM) agents with safety standards;

SafeConstellations: Mitigating Over-Refusals in LLMs Through Task-Aware Representation Steering

Utsav Maskey (Macquarie University), Usman Naseem (Macquarie University)

Safty and PrivacyExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: This paper proposes a task-aware representation-guided method called SafeConstellations, which significantly reduces the over-rejection of LLMs on harmless tasks while maintaining safety.

Safeguarding LLM Fine-tuning via Push-Pull Distributional Alignment

Haozhong Wang (Jilin University), Dandan Guo (Jilin University)

OptimizationSafty and PrivacyExplainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningText

🎯 What it does: This paper proposes a method to protect the safety of large language models during the fine-tuning process through a distribution-level optimal transport (Optimal Transport) framework, called Safety Optimal Transport (SOT), which achieves a balance between safety and usefulness by learning sample weights.

SafeMCP: Proactive Power Regulation for LLM Agent Defense via Environment-Grounded Look-Ahead Reasoning

Lichao Wang (Beijing Institute of Technology), Juntao Dai (Beijing Academy of Artificial Intelligence)

Safty and PrivacyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsWorld ModelTextBenchmarkChain-of-Thought

🎯 What it does: Proposed a server-side defense plugin called SafeMCP, which utilizes an internal world model to perform prospective reasoning, actively filtering and instantly intercepting tool acquisition by LLM agents under the Model Context Protocol (MCP), in order to prevent catastrophic risks caused by power-seeking behavior.

SafeMT: Multi-turn Safety for Multimodal Language Models

Han Zhu (Hong Kong University of Science and Technology), Yike Guo (Hong Kong University of Science and Technology)

Safty and PrivacyTransformerPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: A multi-turn dialogue safety evaluation benchmark called SafeMT was constructed, along with a safety index (SI) and a pluggable dialogue safety regulator.

Safety-Utility Conflicts Are Not Global: Surgical Alignment via Head-Level Diagnosis

Wang Cai (Peking University), Yunfang Wu (Peking University)

OptimizationSafty and PrivacyExplainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringText

🎯 What it does: Propose the CAST framework, which addresses the conflict between the safety and general capabilities of large language models by performing conflict diagnosis and sparse fine-tuning at the attention head level;

SafetyMem: Adaptive Jailbreak Defense via Dual-Component Safety Memory

Hao Wang (Beihang University), Lei Sha (University of Chinese Academy of Sciences)

Safty and PrivacyExplainability and InterpretabilityComputational EfficiencyAdversarial AttackTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextRetrieval-Augmented Generation

🎯 What it does: Proposed the SafetyMem dual safety memory framework, which uses semantic safety memory (SSM) to record attack patterns and episodic safety memory (ESM) to record defense rules, achieving adaptive defense of LLM against jailbreak attacks;

SAFO: Stable Adaptive Fairness Optimization for LLM-Based Social Survey Simulation

Chenxi Lin (Zhejiang University), Yiquan Wu (Alibaba Group)

OptimizationFederated LearningExplainability and InterpretabilityAdversarial AttackData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningTabularBenchmark

🎯 What it does: Propose a dynamic fairness and stability optimization framework named SAFO for training large language models in social survey simulations;

SAGE: A Search-AuGmented Evaluation of Large Language Models on Free-Form QA

Sher Badshah (Dalhousie University), Hassan Sajjad (Emory University)

Explainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: Proposes the SAGE framework, which uses an LLM judge to actively retrieve, summarize, and reflect on web evidence to evaluate the factual accuracy of free-form question-answering answers.

SAGE: Sparse Adaptive Guidance for Dependency-Aware Tabular Data Generation

Shuo Yang (Technical University of Munich), Gjergji Kasneci (Technical University of Munich)

GenerationData SynthesisTransformerLarge Language ModelPrompt EngineeringScore-based ModelTabularBenchmark

🎯 What it does: Proposed the SAGE framework, which utilizes sparse adaptive guidance to achieve dynamic dependency modeling of LLMs in table data generation

SAGE: Synergistic Adaptive Gating of Experts for Hateful Video Detection

Jie Huang (State Key Laboratory of Complex System Modeling and Simulation Technology), Qing Wang (State Key Laboratory of Complex System Modeling and Simulation Technology)

ClassificationAnomaly DetectionTransformerLarge Language ModelMixture of ExpertsContrastive LearningVideoTextMultimodalityAudio

🎯 What it does: Propose the SAGE framework for detecting hate videos, adopting a decoupled expert and instance-level decision arbitration approach, rather than traditional feature fusion;

SAHM: A Benchmark for Arabic Financial and Shari’ah-Compliant Reasoning

Rania Elbadry (Mohamed bin Zayed University of Artificial Intelligence), Zhuohan Xie (Mohamed bin Zayed University of Artificial Intelligence)

Domain AdaptationTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextBenchmarkFinance RelatedRetrieval-Augmented Generation

🎯 What it does: Constructed and publicly released the first multi-task benchmark for Arabic financial text, SAHM, covering seven tasks (AAOIFI standard QA, fatwa QA, accounting and business multiple-choice questions, financial sentiment analysis, extractive summarization, and event-cause reasoning), and evaluated 20 LLMs, proposing that domain adaptation can significantly enhance Arabic financial reasoning capabilities.

SAIR-Comb : A Structure-Aware Iterative Refinement Framework for Combinatorics Autoformalization

Weijie Jiang (East China Normal University), Zhengfeng Yang (East China Normal University)

AI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposes the SAIR-Comb framework for structure-aware iterative automated formalization of combinatorial problems in the Lean 4 environment.

SAM3-I: Segment Anything with Instructions

Jingjing Li (University of Alberta), Li Cheng (Dalian University of Technology)

SegmentationTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodality

🎯 What it does: Designed and implemented SAM3-I, expanding the capabilities of SAM3 to enable direct instance segmentation based on natural language instructions, and constructed a hierarchical, multi-granular HMPL-Instruct dataset.

SAME: Safety-Aware Model Editing Guided by Safety Transformation

Jiayi Wang (Xi'an Jiaotong University), Jian Sun (Xi'an Jiaotong University)

Safty and PrivacyExplainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringText

🎯 What it does: Propose the SAME method for model editing of large language models with safety awareness, achieving safe editing under the locate-then-edit framework.

SAME: Signer-Aware Mixture-of-Experts for Test-Time Adaptation in Sign Language Translation

Lujia Yang (Zhejiang University), Jianwei Yin (Zhejiang University)

Domain AdaptationTransformerMixture of ExpertsVideoMultimodality

🎯 What it does: Proposes SAME — a signer-aware sparse mixture-of-experts network tailored for sign language translation, capable of adaptively addressing domain drift during unsupervised testing;

SARA: Selective and Adaptive Retrieval-augmented Generation with Context Compression

Yiqiao Jin (Georgia Institute of Technology), Srijan Kumar (Georgia Institute of Technology)

GenerationData SynthesisRetrievalCompressionTransformerPrompt EngineeringAuto EncoderTextRetrieval-Augmented Generation

🎯 What it does: Built the SARA framework, combining natural language fragments with semantic compressed vectors to achieve retrieval-augmented generation.

SAT: Balancing Reasoning Accuracy and Efficiency with Stepwise Adaptive Thinking

Weiyang Huang (Harbin Institute of Technology), Min Zhang (Harbin Institute of Technology)

Computational EfficiencyAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextSequentialBenchmarkChain-of-Thought

🎯 What it does: Proposed the Stepwise Adaptive Thinking (SAT) framework, which dynamically controls the depth of reasoning steps through a finite state machine, achieving a balance between reasoning accuracy and efficiency.

SATQuest: A Verifier for Logical Reasoning Evaluation and Reinforcement Fine-Tuning of LLMs

Yanxiao Zhao (Chengdu Institute of Computer Applications Chinese Academy of Sciences), Kaiwen Long (Li Auto)

Data SynthesisOptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextBenchmark

🎯 What it does: Built a CNF-based SAT problem generator and validator called SATQuest, used to automatically generate multi-dimensional logical reasoning tasks and conduct verifiable evaluations, and on this basis, implemented reward-based reinforcement fine-tuning.

Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical Reasoning

Zelin Tan (University of Science and Technology of China), Lei Bai (Shanghai AI Laboratory)

Computational EfficiencyData-Centric LearningTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkChain-of-Thought

🎯 What it does: Studying the scaling behavior of large language models after reinforcement learning training on mathematical reasoning tasks

Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding

Sensen Gao (MBZUAI), Mingming Gong (MBZUAI)

RetrievalExplainability and InterpretabilityComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerAgentic AIVision Language ModelContrastive LearningImageTextMultimodalityGraphTabularReview/Survey PaperBenchmarkFinance RelatedRetrieval-Augmented Generation

🎯 What it does: Reviews the development path of Multimodal Retrieval-Augmented Generation (Multimodal RAG) in document understanding, constructs a classification framework based on domain, retrieval modality, and granularity, and systematically summarizes the integration and application of graph structures and agent frameworks in this task.

Scaling External Knowledge Input Beyond Context Windows of LLMs via Multi-Agent Collaboration

Zijun Liu (Tsinghua University), Yang Liu (Tsinghua University)

Recommendation SystemExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed a multi-agent framework called EXTAGENTS, which processes massive external knowledge beyond the context window of large language models (LLMs) in a sharded parallel manner without expanding the LLM context window, and dynamically integrates this knowledge into answer generation during inference;

Scaling Law for Multimodal Large Language Model Supervised Fine-Tuning

YiFan Zhang (CASIA), Rong Jin (Meta AI)

OptimizationRepresentation LearningData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningVision Language ModelImageTextMultimodalityBenchmark

🎯 What it does: This paper proposes and verifies a scaling law framework for supervised fine-tuning (SFT) of multi-modal large language models (MLLMs), covering both scenarios from scratch training and pre-training. It provides an interaction formula among model scale, pre-training data, and SFT data, and establishes a high correlation between training loss and downstream performance.

Scaling Laws for Code: A More Data-Hungry Regime

Xianzhen Luo (Harbin Institute of Technology), Wanxiang Che (Harbin Institute of Technology)

AI Code AssistantTransformerLarge Language ModelText

🎯 What it does: Conducted large-scale experiments on 117 code LLMs of varying sizes, constructed and fitted scaling laws for code languages, and systematically evaluated the impact of mixed code and natural language data on scaling laws.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models

Tingchen Fu (Renmin University of China), Yu Cheng (Chinese University of Hong Kong)

Explainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextBenchmarkChain-of-Thought

🎯 What it does: This paper proposes MathIF, an evaluation framework specifically designed for large reasoning models, which systematically assesses the model's ability to follow user instructions in mathematical reasoning tasks and explores the trade-off between reasoning depth and instruction following;

Scaling Test-Time Compute to Achieve IOI Gold Medal with Open-Weight Models

Mehrzad Samadi (NVIDIA), Boris Ginsburg (NVIDIA)

Computational EfficiencyAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelMixture of ExpertsTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed the GENCLUSTER framework, utilizing large-scale candidate generation, behavior clustering, tournament ranking, and polling submission, helping open-source weight models win a gold medal in IOI 2025.

SCAN: Structured Capability Assessment and Navigation for LLMs

Zongqi Wang (Tsinghua University), Yujiu Yang (Tsinghua University)

ClassificationExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Established the SCAN framework for structured, fine-grained capability evaluation and navigation of LLMs.

SceneAlign: Aligning Multimodal Reasoning to Scene Graphs in Complex Visual Scenes

Chuhan Wang (University of California San Diego), Jingbo Shang (University of California San Diego)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityGraphBenchmarkChain-of-Thought

🎯 What it does: Propose the SceneAlign framework, which utilizes scene graphs to structurally align multi-modal reasoning, thereby enhancing the authenticity and reliability of visual reasoning.

Schoenfeld’s Anatomy of Mathematical Reasoning by Language Models

Ming Li (University of Maryland College Park), Tianyi Zhou (University of Maryland College Park)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextReview/Survey PaperChain-of-Thought

🎯 What it does: Propose the ThinkARM framework, which divides the reasoning process of LLMs into eight categories (Read, Analyze, Plan, Implement, Explore, Verify, Monitor, Answer) based on Schoenfeld's Episode Theory, achieving sentence-level automatic annotation;

ScholaWrite: A Dataset of End-to-End Scholarly Writing

Khanh Chi Le (University of Minnesota), Dongyeop Kang (University of Minnesota)

GenerationData SynthesisFederated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningAI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextSequentialBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Constructed and released the SCHOLAWRITE dataset, which records the complete keyboard operations and cognitive writing intentions of five computer science preprints from the draft to the final version; evaluated and fine-tuned large language models on the alignment between intent prediction and text generation during the writing process based on this dataset.

SciCoQA: Quality Assurance for Scientific Paper–Code Alignment

Tim Baumgärtner (TU Darmstadt), Iryna Gurevych (TU Darmstadt)

Large Language ModelTextMultimodalityBenchmarkPhysics Related

🎯 What it does: Designed and released the SCICOQA dataset to evaluate the performance of large language models in detecting inconsistencies between scientific papers and code.

SciCustom: A Framework for Custom Evaluation of Scientific Capabilities in Large Language Models

Yiyang Gu (Peking University), Ming Zhang (Peking University)

TransformerLarge Language ModelPrompt EngineeringMixture of ExpertsContrastive LearningTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Built a customizable evaluation framework called SCICUSTOM based on a scientific ontology, which can automatically map large-scale scientific question-answer data to fine-grained knowledge units. It retrieves and filters data through techniques such as voting consensus and binary search according to user needs, ultimately generating a high-quality multiple-choice benchmark without the need for manual annotation.

SciFlow-Bench: Evaluating Structure-Aware Scientific Diagram Generation via Inverse Parsing

Tong Zhang (Peking University), Wentao Zhang (Peking University)

GenerationData SynthesisGraph Neural NetworkTransformerAgentic AIPrompt EngineeringVision Language ModelDiffusion modelContrastive LearningImageTextGraphBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose the SciFlow-Bench benchmark, constructing a structural prior evaluation framework for generating scientific diagrams from text to image, with the core being the use of inverse parsing to recover the generated image into a graph structure and compare it with the ground truth graph.

SciMDR: Advancing Scientific Multimodal Document Reasoning

Ziyu Chen (University of Chicago), Arman Cohan (Yale University)

Federated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningData-Centric LearningDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackNeural Architecture SearchProtein Structure PredictionTransformerLarge Language ModelPrompt EngineeringVision Language ModelDiffusion modelScore-based ModelFlow-based ModelRectified FlowNeural Radiance FieldAuto EncoderGenerative Adversarial NetworkContrastive LearningGaussian SplattingSimultaneous Localization and MappingWorld ModelOptical FlowImageTextMultimodalityReview/Survey PaperBenchmarkRetrieval-Augmented GenerationChain-of-ThoughtStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: This paper proposes a two-stage framework called 'synthesize-and-reground,' which first generates high-quality QA pairs and reasoning chains in simplified atomic contexts, and then re-embeds them into complete scientific papers to construct a large-scale multi-modal scientific reasoning dataset called SCIMDR;

SciPedia: Unlocking the Value of Scientific Data for Pre-training

Yiwei Qin (Shanghai Innovation Institute), Pengfei Liu (Shanghai Innovation Institute)

Knowledge DistillationData-Centric LearningTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsContrastive LearningTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Constructed a high-quality scientific corpus named SciPedia with a scale of 900B words. A two-stage processing pipeline (content cleaning + pedagogical incremental learning) was designed to address the learning challenges of raw scientific texts. A controlled validation framework was also built (SciPedia-Eval benchmark, 600B continuous pre-training with 3B/7B base models from scratch) to evaluate the value of scientific data.

SCOPE: Boosting LLM Efficiency with Scoped Position Encoding

Qingguo Qi (Zhejiang University), Zhao Li (Zhejiang Lab)

Computational EfficiencyTransformerLarge Language ModelText

🎯 What it does: Proposes SCOPE, a framework that achieves implicit position encoding through exponentially expanded attention视野.

SCOUT: Selective Coupling via Optimal Unbalanced Transport for Interpretable Text Classification

Junhao Jia (Zhejiang University), Lei Wu (Zhejiang University)

ClassificationExplainability and InterpretabilityTransformerContrastive LearningText

🎯 What it does: Proposed the SCOUT model, achieving interpretable prototype reasoning in text classification through shard alignment based on Unbalanced Optimal Transport.

SCVQ: Sparse-Compensated Vector Quantization for Large Language Models

Zixuan Zhou (Beijing University of Posts and Telecommunications), Zhaofeng He (Beijing University of Posts and Telecommunications)

Computational EfficiencyKnowledge DistillationTransformerLarge Language ModelAuto EncoderText

🎯 What it does: Propose a post-training quantization method called SCVQ, which combines sparse compensation and vector quantization to significantly reduce model size and inference latency while maintaining the performance of LLMs.

SDAR-VL: Stable and Efficient Block-wise Diffusion for Vision-Language Understanding

Shuang Cheng (Zhejiang University), Bowen Zhou (Shanghai AI Laboratory)

ClassificationRecognitionImage TranslationRestorationSegmentationGenerationData SynthesisComputational EfficiencyRepresentation LearningTransformerPrompt EngineeringVision Language ModelDiffusion modelContrastive LearningImageVideoTextMultimodalityBenchmarkChain-of-Thought

🎯 What it does: A block-level discrete diffusion framework called SDAR-VL for visual-language understanding was constructed, and a unified training framework was proposed to enhance its stability and efficiency.

SDE-SQL: Enhancing Text-to-SQL Generation in Large Language Models via Self-Driven Exploration with SQL Probes

Wenxuan Xie (Fudan University), Wenhao Jiang (Guangdong Laboratory of AI and Digital Economy)

GenerationAI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextTabularStochastic Differential Equation

🎯 What it does: Enable LLM to proactively query databases during the inference phase through self-driven exploration (SQL probe), improving text-to-SQL generation.

SDiaReward: Modeling and Benchmarking Spoken Dialogue Rewards with Modality and Colloquialness

Jingyu Lu (Zhejiang University), Zhou Zhao (Zhejiang University)

Reinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningContrastive LearningTextMultimodalityBenchmarkAudio

🎯 What it does: Proposed SDiaReward, a reward model for multi-turn speech dialogue, addressing two major evaluation challenges: differences in speech patterns and differences in colloquial expressions.

SEA-BED: How Do Embedding Models Represent Southeast Asian Languages?

Wuttikorn Ponwitayarat (Vidyasirimedhi Institute of Science and Technology), Sarana Nutanong (Vidyasirimedhi Institute of Science and Technology)

Representation LearningTransformerLarge Language ModelContrastive LearningTextBenchmark

🎯 What it does: Constructed SEA-BED, a large-scale benchmark for systematically evaluating multilingual embedding models across 10 Southeast Asian languages and 9 task types; evaluated 17 multilingual embedding models on this benchmark.

SearchGym: Bootstrapping Real-World Search Agents via Cost-Effective and High-Fidelity Environment Simulation

Xichen Zhang (Hong Kong University of Science and Technology), Jiaya Jia (University of Hong Kong)

OptimizationData-Centric LearningTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextGraphBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Designed a low-cost, high-fidelity simulation environment called SearchGym for training search agents, and proposed the SearchGym-RL curriculum-based reinforcement learning framework;

SEARL: Joint Optimization of Policy and Tool Graph Memory for Self-Evolving Agents

Xinshun Feng (Shanghai Artificial Intelligence Laboratory), Jing Shao (Shanghai Artificial Intelligence Laboratory)

OptimizationTransformerLarge Language ModelReinforcement LearningTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposes the SEARL framework, which jointly optimizes LLM strategies and tool graph memory, enabling self-evolving agents to continuously create, retrieve, and reuse tools while learning through reinforcement learning.