arXivSub Start free trial

ACL 2026 Papers — Page 19

Annual Meeting of the Association for Computational Linguistics · 2296 papers

SecureGate: Learning When to Reveal PII Safely via Token-Gated Dual-Adapters for Federated LLMs

Mohamed Shaaban (Washington State University), Mohamed Elmahallawy (Washington State University)

Federated LearningSafty and PrivacyTransformerLarge Language ModelSupervised Fine-TuningTextReview/Survey Paper

🎯 What it does: Proposes the SECUREGATE framework for secure fine-tuning of large language models in federated learning, combining dual adapters and token gating to achieve on-demand PII leakage.

SeCuRepair: Semantics-Aligned, Curriculum-Driven, and Reasoning-Enhanced Vulnerability Repair Framework

Chengran Yang (Singapore Management University), David Lo (Singapore Management University)

Knowledge DistillationAI Code AssistantTransformerLarge Language ModelReinforcement LearningPrompt EngineeringContrastive LearningText

🎯 What it does: Propose a reinforcement learning-based automatic vulnerability repair framework called SeCuRepair, combining expert-style reasoning, semantic-aware rewards, and difficulty-progressive training.

SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing Scenarios

Junkai Chen (Singapore Management University), David Lo (Singapore Management University)

Safty and PrivacyAI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose SECUREVIBEBENCH, a C/C++ benchmark for evaluating the security coding capabilities of code agents, focusing on real-world scenarios where vulnerabilities are introduced;

SED-SFT: Selectively Encouraging Diversity in Supervised Fine-Tuning

Yijie Chen (Tencent Inc), Fandong Meng (Tencent Inc)

OptimizationData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningTextBenchmark

🎯 What it does: This paper proposes a method called SED‑SFT, which encourages generation diversity during the supervised fine-tuning (SFT) stage through selective entropy regularization.

SeDev: Structured Semantic Exploration for LLM-Driven Code Generation

Ronghui Yang (South China University of Technology), Mengchen Zhao (South China University of Technology)

GenerationAI Code AssistantTransformerLarge Language ModelAgentic AIPrompt EngineeringText

🎯 What it does: Proposed the SeDev framework, using a multi-agent structure and structured prompts for LLM-driven code generation.

See the Forest for the Trees: Loosely Speculative Decoding via Visual-Semantic Guidance for Efficient Inference of Video LLMs

Yicheng Ji (State Key Laboratory of Blockchain and Data Security, Zhejiang University), Huan Li (State Key Laboratory of Blockchain and Data Security, Zhejiang University)

Computational EfficiencyRepresentation LearningTransformerLarge Language ModelVision Language ModelVideoText

🎯 What it does: Proposes a loose inference framework called LVSPEC based on visual semantic guidance, which significantly accelerates the autoregressive inference of video LLMs without training a draft model;

SEE: Signal Embedding Energy for Quantifying Noise Interference in Large Audio Language Models

Yuanhe Zhang (Beijing University of Posts and Telecommunications), Sen Su (Chongqing University of Posts and Telecommunications)

RestorationTransformerLarge Language ModelAuto EncoderContrastive LearningAudio

🎯 What it does: Propose the Signal Embedding Energy (SEE) metric to quantify the internal interference of LALM to noise, and develop a training-agnostic Signal Embedding Energy Neutralization (SEEN) method based on this to suppress noise.

See2Refine: Vision-Language Feedback Improves LLM-Based eHMI Action Designers

Ding Xia, Takeo Igarashi (University of Tokyo)

Autonomous DrivingOptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringVision Language ModelImageVideoTextMultimodality

🎯 What it does: Built a closed-loop system called SEE2REFINE, which uses the perceptual evaluation of a vision-language model (VLM) as automated, human-free feedback to iteratively improve the action design of the external human-machine interface (eHMI) generated by the large language model (LLM).

Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-Experts

Haolei Xu (Zhejiang University), Yueting Zhuang (Zhejiang University)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerPrompt EngineeringMixture of ExpertsVision Language ModelContrastive LearningImageTextMultimodalityBenchmark

🎯 What it does: This paper investigates the 'see but do not think' phenomenon in visual understanding and reasoning using multimodal sparse mixture-of-experts models, and proposes the routing interference hypothesis;

Seeing No Evil: Blinding Large Vision-Language Models to Safety Instructions via Adversarial Attention Hijacking

Jingru Li, Tianqing Zhu

Safty and PrivacyAdversarial AttackTransformerVision Language ModelImageTextMultimodalityBenchmark

🎯 What it does: Achieve visual destruction attacks on large vision-language models by incorporating auxiliary objectives that suppress system prefix attention and enhance image attention in visual perturbation optimization.

Seeing the Whole Elephant: A Benchmark for Failure Attribution in LLM-based Multi-Agent Systems

Mengzhuo Chen (State Key Laboratory of Complex System Modeling and Simulation Technology), Qing Wang (State Key Laboratory of Complex System Modeling and Simulation Technology)

Explainability and InterpretabilityTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmark

🎯 What it does: Designed and released TraceElephant, a failure attribution benchmark for large language model (LLM)-based multi-agent systems (MAS), which collects complete executable execution trajectories and reproducible environments; and systematically evaluated multiple attribution methods on this benchmark;

Segment, Embed, and Align: A Universal Recipe for Aligning Subtitles to Signing

Zifan Jiang (University of Oxford), Andrew Zisserman (University of Zurich)

RecognitionSegmentationRetrievalRecurrent Neural NetworkTransformerVision Language ModelContrastive LearningOptical FlowVideoTextMultimodality

🎯 What it does: Propose the SEA method, which automatically aligns captions with continuous sign language videos through three steps: segmentation, embedding, and alignment.

SegTune: Structured and Fine-Grained Control for Song Generation

Yuejiao Wang (Kuaishou Technology), Pengfei Wan (Kuaishou Technology)

GenerationData SynthesisReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringDiffusion modelAuto EncoderTextAudio

🎯 What it does: Propose SegTune, a non-autoregressive singing and accompaniment generation framework based on Diffusion Transformer, supporting global and segment-level text control, and achieving precise lyric alignment by using LLM to predict sentence-level lyric durations.

SeLaR: Selective Latent Reasoning in Large Language Models

Renyu Fu, Guibo Luo (Peking University)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBenchmarkChain-of-Thought

🎯 What it does: This paper studies a training-free implicit reasoning framework called SeLaR, which activates soft embeddings in uncertain reasoning steps through an entropy gating mechanism, and combines entropy-aware contrastive regularization to achieve more stable and efficient chain-of-thought reasoning.

Select Before Use: On the Importance of Reference Model Selection in Preference Alignment

Muyang Li (Sydney AI Centre University of Sydney), Tongliang Liu (Sydney AI Centre University of Sydney)

Reinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningTextBenchmark

🎯 What it does: Studied how to select the most suitable SFT checkpoint as a reference model for preference alignment during the post-training phase of large language models, and proposed an evaluation metric called RewardRank based on initial rewards for model selection without training.

SELECting over Tokens: Curating Pre-training Data at Scale via Token Classification

Xin Tong (Alibaba), Bo Zheng (Alibaba)

OptimizationData-Centric LearningTransformerLarge Language ModelContrastive LearningText

🎯 What it does: Transform the pre-training data cleaning task into token-level classification, and achieve efficient data optimization on large-scale corpora through the SELECT framework.

Selective Contrastive Learning For Gloss Free Sign Language Translation

Chang Hao Lai, Yidong Chen (Xiamen University)

Image TranslationRepresentation LearningTransformerPrompt EngineeringContrastive LearningVideoTextMultimodality

🎯 What it does: Proposed SCL-SLT, which uses selective contrastive learning in sign language translation without gloss;

Selective Knowledge Distillation: Fusing LLM Semantic Strengths with DNN Efficiency for Binary Code Similarity Detection

Shize Zhou (Zhejiang University), Wenhai Wang (Hangzhou Dianzi University)

RetrievalComputational EfficiencyKnowledge DistillationConvolutional Neural NetworkTransformerLarge Language ModelContrastive LearningText

🎯 What it does: Proposed the BinSKD framework, which transfers the high-level semantic knowledge of large language models to lightweight deep neural networks through selective distillation, thereby improving the accuracy and robustness of binary code similarity detection (BCSD).

Selective Span-Level Unlearning for Large Language Models

Chaewon Yoon (Jeonbuk National University), Hyun-Je Song (Jeonbuk National University)

Federated LearningSafty and PrivacyExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningAdversarial AttackTransformerLarge Language ModelSupervised Fine-TuningAuto EncoderContrastive LearningTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Designed a selective span-level forgetting method based solely on internal model information, calculating token importance by comparing gradients of forgotten and retained data, and using self-consistency to generate stable spans as forgetting targets.

Selective Test-Time Debiasing for CLIP via Reward Gating

Jaeho Han (Chung-Ang University), Junyeong Kim (Chung-Ang University)

Domain AdaptationFederated LearningExplainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerReinforcement LearningVision Language ModelContrastive LearningImageTextMultimodality

🎯 What it does: Propose a test-time adaptation method called RG-TTA based on reward gating, which selectively debiases input bias sensitivity in the CLIP vision-language model.

Self-Guided Alignment: Adaptive Preference Sensing for Multi-Objective Generation

Ning Wang (Jiangnan University), Haojie Zhou (Jiangnan University)

GenerationReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextBenchmark

🎯 What it does: This paper proposes the Self-Guided Alignment (SGA) framework, which unifies preference space learning and conditional generation through a dual-head structure, enabling adaptive preference perception and self-guided generation during inference without requiring manual input.

Self-Reflective Generation at Test Time

Jian Mu (Hong Kong University of Science and Technology (Guangzhou)), Yao Shu (Hong Kong University of Science and Technology (Guangzhou))

GenerationExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextSequentialRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose a generation framework called SRGen that self-reflects during testing, dynamically detecting high-uncertainty tokens and instantly optimizing correction vectors to improve the reasoning reliability of large language models.

Self-SoftCoT: A Self-Consistent Framework via Position-Aware Latent Space Reinforcement Learning

Liangliang Dong (Jilin University), Shuaimin Li (Shenzhen Institutes of Advanced Technology)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextChain-of-Thought

🎯 What it does: Propose Self-SoftCoT, a continuous latent chain-of-thought method implemented through a single-stream closed-loop within a frozen large language model, enabling the model to perform reasoning without relying on external auxiliary models.

SelFusion: Self-distillation for Diffusion Language Models

Hyeongsoo Lim, Ji Won Yoon (Chung-Ang University)

GenerationKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningDiffusion modelContrastive LearningText

🎯 What it does: Proposed SelFusion, a self-distillation framework that enhances the generation quality of diffusion language models through bidirectional distillation between the 'easy mode' and 'hard mode' of the same model under different masking ratios.

Semantic Hardness Is Not Visual Hardness: Sign-Aware Hard Negative Mining for Sign Language Retrieval

Junmyeong Lee (Korea Advanced Institute of Science and Technology), KyungTae Lim (Korea Advanced Institute of Science and Technology)

RetrievalTransformerVision Language ModelContrastive LearningVideoTextMultimodality

🎯 What it does: Proposed and implemented a vision similarity-based hard negative sampling framework (SAN) for sign language retrieval, which improves fine-grained retrieval performance by selecting visually confusing negative samples in the sign language embedding space.

Semantic-Aware Logical Reasoning via a Semiotic Framework

Yunyao Zhang (Huazhong University of Science and Technology), Zikai Song (Huazhong University of Science and Technology)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Built a multi-perspective logical reasoning system called LogicAgent based on a semantic symbol framework, and proposed a new semantic-logical complexity evaluation benchmark called RepublicQA

Semantically Comprehensive Token Pruning in LVLMs via Maximizing Concept Coverage

Xueting Li (Xidian University), Cheng Deng (Xidian University)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerVision Language ModelAuto EncoderContrastive LearningImageTextMultimodality

🎯 What it does: Propose a visual token pruning framework called SCTS based on semantic integrity, which uses a sparse autoencoder to decompose visual features into atomic semantic concepts, and iteratively selects tokens through Maximum Concept Coverage (MCC) and Marginal Semantic Gain (MSG) strategies to achieve semantic integrity and efficient inference under high compression rates.

SemCSE-Multi: Multifaceted and Decodable Embeddings for Aspect-Specific and Interpretable Scientific Domain Mapping

Marc Brinner (Bielefeld University), Sina Zarrieß (Bielefeld University)

Explainability and InterpretabilityKnowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextReview/Survey Paper

🎯 What it does: This paper proposes an unsupervised framework called SemCSE-Multi, which is used to generate multi-faceted embeddings from scientific abstracts and can decode the embeddings back into natural language descriptions.

Sense and Sensitivity: Examining the Influence of Semantic Recall on Long Context Code Understanding

Adam Štorek (Columbia University), Suman Jana (Columbia University)

AI Code AssistantTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: This paper distinguishes two types of memory capabilities of large language models (LLMs) in long-context code understanding: lexical recall and semantic recall. By embedding target code at different positions and adding a large number of irrelevant placeholders, the performance of 10 advanced LLMs in lexical recall, semantic recall, and high-sensitivity semantic recall tasks is evaluated. The concept of 'semantic recall sensitivity' is proposed, and a line-deletion control experiment is designed to measure sensitivity; further, a high-sensitivity output prediction benchmark called SemTrace is created to eliminate pattern-matching shortcuts. Experimental results show that lexical recall is almost unaffected by position, while semantic recall significantly decreases when the code is located in the middle of the context; in contrast, SemTrace exhibits stronger position dependence, indicating that existing evaluations underestimate the severity of semantic recall failure.

SenseRel: A Sense-Level Benchmark for Denotational and Connotational Meaning Relations

Pierluigi Cassotti (University of Gothenburg), Nina Tahmasebi (University of Gothenburg)

TransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextBenchmark

🎯 What it does: Proposed and released the first benchmark dataset, SenseRel, at the semantic level, for evaluating semantic relationships between word senses (including referential and affective dimensions), and for assessing large language models (LLMs);

ServImage: An Image Generation and Editing Benchmark from Real-world Commercial Imaging Services

Fengxian Ji (MBZUAI), Xiuying Chen (MBZUAI)

GenerationData SynthesisRecommendation SystemTransformerSupervised Fine-TuningPrompt EngineeringVision Language ModelDiffusion modelImageMultimodalityTabularBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose the ServImage benchmark to evaluate the performance of image generation and editing models in real paid commercial design tasks.

SFTMix: Elevating Language Model Instruction Tuning with Mixup Recipe

Yuxin Xiao (Massachusetts Institute of Technology), Wenxuan Zhou (Zoom Video Communications)

Computational EfficiencyRepresentation LearningData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningTextBiomedical DataBenchmark

🎯 What it does: This paper proposes an SFTMix optimization scheme based on Mixup, aimed at improving the performance of large language models during the instruction tuning phase;

SGPVT: Self-Generated Proximal Visual Tokens for Mitigating Proximal Collateral Damage in MLLM Unlearning

Jiaqi Li (Southeast University), Guilin Qi (Southeast University)

Safty and PrivacyExplainability and InterpretabilityKnowledge DistillationPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBenchmark

🎯 What it does: This paper proposes a novel MLLM forgetting mechanism that utilizes self-generated approximate visual tokens (SGPVT) to forget target concepts while minimizing damage to related concepts.

SGT: Securing Open-Source LLMs Against Malicious Fine-tuning via Safety Guidance Trigger

Sunguk Shin (Korea University), Sungwon Park (Max Planck Institute for Security and Privacy)

Safty and PrivacyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringContrastive LearningText

🎯 What it does: Before releasing open-weight LLMs, first train a soft safety trigger to align the model's representation space to a safe manifold, then achieve safe outputs without the trigger, thereby enhancing robustness against malicious fine-tuning.

SGVEF-LOOP: Coverage-Guided Progressive Topological Exploration and Fact-Grounded Metamorphic Evaluation for MCP Agents

Zenghao Liu (Renmin University of China), Yansong Zhang (Renmin University of China)

Autonomous DrivingOptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyRobotic IntelligenceAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringTextTabularBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposes SGVEF-LOOP, a coverage-driven closed-loop framework for adaptive exploration of the MCP tool space and generating mutation test pairs without oracles;

Shanks: Simultaneous Hearing and Thinking for Spoken Language Models

Cheng-Han Chiang (National Taiwan University), Lijuan Wang (Microsoft)

Human-Computer InteractionTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-ThoughtAudio

🎯 What it does: Proposed the SHANKS framework, enabling speech-language models (SLM) to synchronously generate internal chain reasoning while the user is speaking, achieving real-time interaction and tool invocation.

SHAPE: Stage-aware Hierarchical Advantage via Potential Estimation for LLM Reasoning

Zhengyang Ai (Huawei Taylor Lab), Pinyan Lu (Huawei Taylor Lab)

Computational EfficiencyKnowledge DistillationTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkChain-of-Thought

🎯 What it does: Designed the SHAPE process supervision framework, leveraging potential estimation to achieve hierarchical advantage rewards, thereby improving the accuracy and token efficiency of LLMs in mathematical reasoning.

SHAPE: Unifying Safety, Helpfulness and Pedagogy for Educational LLMs

Sihang Zhao (New York University), Hongyi Wen (Chinese University of Hong Kong)

Safty and PrivacyExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningAdversarial AttackGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringTextGraphTabularBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This study unifies the three behaviors of safety, helpfulness, and pedagogy in educational large language models (LLMs) and constructs an evaluation framework based on a knowledge mastery graph. Subsequently, the SHAPE benchmark was proposed, containing 9,087 student-question pairs covering the concept map of linear algebra and students' mastery status. To enhance the model's robustness under adversarial attacks, the authors designed a graph-enhanced pedagogical pipeline: first, parse the prerequisite concepts corresponding to the question and compare them with the student's mastery status, deciding whether to directly answer or generate guided pedagogical Q&A. Experiments show that this pipeline significantly improves the safety rate of most LLMs under answer-inducing jailbreaks while maintaining high levels of helpfulness and pedagogy.

SharedRequest: Privacy-Preserving Model-Agnostic Inference for Large Language Models

Peihua Mai (National University of Singapore (Chongqing) Research Institute), Yan Pang (National University of Singapore (Chongqing) Research Institute)

Federated LearningSafty and PrivacyComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringAuto EncoderGenerative Adversarial NetworkContrastive LearningTextBenchmark

🎯 What it does: Propose a model-agnostic privacy-preserving LLM inference framework called SharedRequest, which achieves privacy and efficient inference by obfuscating original prompts at the batch level and sharing noisy queries;

SHARP: Self-adaptive Harmful Category-aware Prompt Generation for Black-box Jailbreaking

Yingjie Xue (Wuhan University), Fei Li (Wuhan University)

Safty and PrivacyAdversarial AttackTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: Developed SHARP, the first category-aware black-box jailbreak framework, capable of adaptively generating prompts that bypass security mechanisms based on the semantic categories of harmful questions.

SharVeT: Similarity-aware Parameter Sharing with Vector-based Tuning for Efficient LLM Compression

Jeongin Yun (Seoul National University), U Kang (Seoul National University)

CompressionComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelText

🎯 What it does: Propose the SharVeT parameter-sharing framework, which achieves efficient compression of LLMs through module-level similarity clustering, adaptive parameter allocation, and lightweight correction.

Sherry: Hardware-Efficient 1.25-Bit Ternary Quantization via Fine-grained Sparsification

Hong Huang (City University of Hong Kong), Dapeng Wu (City University of Hong Kong)

Computational EfficiencyKnowledge DistillationTransformerLarge Language ModelText

🎯 What it does: Proposed the Sherry framework, achieving 1.25-bit hardware-friendly ternary quantization, utilizing a 3:4 sparse structure to compress four weights into five bits, and addressing the weight trap in sparse ternary training through the Annealing Residual Synapse (Arenas) mechanism;

ShopSimulator: Evaluating and Exploring RL-Driven LLM Agent for Shopping Assistants

Pei Wang (Alibaba Group), Bo Zheng (Alibaba Group)

Recommendation SystemReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AIText

🎯 What it does: Propose ShopSimulator, a unified simulation environment for multi-round interaction, personalization, and fine-grained product retrieval in Chinese e-commerce, used to evaluate and train LLM shopping assistants.

Shuttle Between Symbolic Instructions and Neural Parameters of Large Language Models

Wangtao Sun (Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences), Kang Liu (Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences)

Representation LearningData-Centric LearningAI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringAuto EncoderText

🎯 What it does: Proposed the SHIP framework, achieving bidirectional mapping between symbolic instructions and LLM parameters, and verified its effectiveness in tasks such as reasoning and induction.

Sigmoid Head for Quality Estimation under Language Ambiguity

Tu Anh Dinh (Karlsruhe Institute of Technology), Jan Niehues (Karlsruhe Institute of Technology)

Data-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningTextMultimodalityBenchmark

🎯 What it does: Train and insert an additional Sigmoid Head on top of a pre-trained language model for unsupervised quality estimation (QE).

Sign-Language Datasets at Scale: A Comprehensive Survey on Resources, Benchmarks, and Annotation Standards

Yiming Ni (University of Washington), Wei Cheng (University of Washington)

Data-Centric LearningSupervised Fine-TuningContrastive LearningVideoTextMultimodalityReview/Survey PaperBenchmark

🎯 What it does: This paper provides a comprehensive review and systematic analysis of 120 publicly available sign language datasets, covering 35 sign languages, and proposes a unified 24-field Datasheet template along with a public GitHub repository. It also deeply discusses issues such as modality, annotation granularity, and signer bias in the datasets, and provides cross-dataset benchmark results.

SILO-BENCH: A Scalable Environment for Evaluating Distributed Coordination in Multi-Agent LLM Systems

Yuzhe Zhang (Beijing University of Technology), Wenyuan Jiang (ETH Zürich)

TransformerLarge Language ModelAgentic AITextBenchmark

🎯 What it does: Proposed a role-agnostic, scalable multi-agent LLM evaluation benchmark called SILO-BENCH, designed to measure agents' distributed coordination capabilities under information silos.

SiLP: Enhancing Non-Dominant Language Capabilities with a Selective Bidirectional Language Projection Framework

Junpeng Liu (Hong Kong University of Science and Technology), Hui Xiong (Hong Kong University of Science and Technology)

Representation LearningData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextMultimodality

🎯 What it does: Designed and implemented a selective bidirectional language projection framework called SiLP, which aligns non-primary languages with the primary language at specific layers using intrinsic model parameters, thereby enhancing the understanding and generation capabilities of non-primary languages.

SimPBL: A Multi-Agent Framework for Project-Based Learning

Daniel Zhang-Li (Tsinghua University), Juanzi Li (Tsinghua University)

AI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextSequentialRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Construct and experiment a multi-agent framework (SIMPBL), where learners collaborate with agent companions to complete project-based learning (PBL), organizing the learning process through project goals, driving questions, and role matching.

Simple Agents, Biased Judges: Efficient Multi-Party Dialogue Generation & The Evaluation Gap

Kunal Samanta (Université Laval), Richard Khoury (Université de Sherbrooke)

GenerationExplainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringTextReview/Survey PaperRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: We propose a lightweight multi-party dialogue generation framework called MPOD, and conduct comparative experiments with human evaluation and LLM judgments.

Simple-VGC: Enhancing Visual Grounding in Multimodal Reasoning via Adaptive Tool Composition

Ye Wang (Fudan University), Zhongyu Wei (International Digital Economy Academy)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmarkChain-of-Thought

🎯 What it does: Proposes a multi-step visual reasoning framework called SIMPLE‑VGC based on tools, aiming to address the visual grounding failures of multi-modal large language models in long-sequence reasoning, fine-grained identification, and region localization.

Simulated Students in Tutoring Dialogues: Substance or Illusion?

Alexander Scarlatos (University of Massachusetts Amherst), Andrew Lan (University of Massachusetts Amherst)

TransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextBenchmark

🎯 What it does: This paper proposes a set of metrics to evaluate the performance of LLMs simulating students in tutoring dialogues, and conducts benchmark evaluations of multiple simulation methods on real mathematical dialogue data.

Single-Pass, Depth-Selective Reading for Multi-Aspect Sentiment Analysis

Yan Xia (Universiti Malaya), Chee Seng Chan (Universiti Malaya)

Computational EfficiencyRepresentation LearningTransformerSupervised Fine-TuningContrastive LearningTextBenchmark

🎯 What it does: Propose the DABS framework, which achieves querying and adjustable-depth reading of multi-target sentiments through a shared deep substructure after a single encoding, significantly reducing the cost of multi-aspect reasoning.

Situated Embedding Models for Context-Aware Dense Retrieval

Junjie Wu (Hong Kong University Of Science And Technology), Mo Yu (Tencent)

RetrievalTransformerContrastive LearningTextRetrieval-Augmented Generation

🎯 What it does: Proposes a scenario-aware embedding model, studying how to utilize contextual information from paragraphs to improve retrieval quality in retrieval-augmented generation tasks.

Skill-Aware Data Selection and Fine-Tuning for Data-Efficient Reasoning Distillation

Lechen Zhang (University of Michigan), Lu Wang (University of Michigan)

Knowledge DistillationData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextChain-of-Thought

🎯 What it does: Proposed a skill-centric data selection and training framework for efficiently distilling large reasoning models using only 1,000 samples.

SkMTEB: Slovak Massive Text Embedding Benchmark and Model Adaptation

Marek Suppa, Viktória Ondrejová (Cisco Systems)

RetrievalRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningTextBenchmark

🎯 What it does: Designed and released SkMTEB, the first benchmark providing 31 datasets covering 7 task categories for the Slovak language, and built locally deployable embedding models with 45M~365M parameters based on the multilingual E5 model through vocabulary trimming and fine-tuning.

SLICEFORMER: Static Program Slicing Using Language Models With Dataflow-Aware Pretraining and Constrained Decoding

Pengfei He (University of Manitoba), Muhammad Asaduzzaman (University of Windsor)

AI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextSequential

🎯 What it does: Propose SLICEFORMER, a static program slicing method that utilizes a small language model combined with data-flow-aware pre-training and constrained decoding.

SlideAgent: Hierarchical Agentic Framework for Multi-Page Visual Document Understanding

Yiqiao Jin (Georgia Institute of Technology), Srijan Kumar (J P Morgan Ai Research)

RetrievalRecommendation SystemExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelAgentic AIVision Language ModelImageTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose SlideAgent, a multi-level agent-based framework for understanding multi-page visual documents (e.g., slides). The framework is divided into three levels: global, page-level, and element-level. During the knowledge construction phase, specialized agents generate query-free structured knowledge bases, and during the reasoning phase, answers are dynamically generated by calling the corresponding agents based on queries.

SLoRA: Balancing Plasticity and Forgetting in Large Language Models for Continual Learning

Lina Yang (Shanghai Jiao Tong University), Yu Wang (Shanghai Jiao Tong University)

Computational EfficiencyRepresentation LearningTransformerLarge Language ModelTextBenchmark

🎯 What it does: This paper studies the problem of catastrophic forgetting in large language models during continual learning, and proposes an SLoRA framework based on subspace denoising, which automatically removes noisy components by leveraging the subspace similarity in LoRA low-rank updates, thereby mitigating forgetting.

SLR: Automated Synthesis for Scalable Logical Reasoning

Lukas Helff (TU Darmstadt), Kristian Kersting (TU Darmstadt)

Data SynthesisExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextBenchmarkChain-of-Thought

🎯 What it does: Proposed the SLR (Scalable Logical Reasoning) framework, which automatically generates verifiable inductive logical reasoning tasks, constructs a 20-layer progressive SLR-BENCH benchmark, and uses this framework to train and evaluate LLMs.

SMART: Evaluating LLMs’ Mathematical Reasoning via a Human Cognitive Process-Inspired Benchmark

Yujie Hou (Beijing Normal University), Hua Huang (Beijing Normal University)

Large Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed the SMART benchmark to fine-grainedly evaluate the mathematical reasoning ability of LLMs from four cognitive dimensions: semantic understanding, mathematical reasoning, arithmetic calculation, and reflection improvement.

SMARTER: A Data-efficient Framework to Improve Toxicity Detection with Explanation via Self-augmenting Large Language Models

Huy Nghiem (University of Maryland), Hal Daumé Iii

ClassificationExplainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextChain-of-Thought

🎯 What it does: Propose the SMARTER framework, which leverages large language models to self-enhance and generate explanations for achieving low-data explainable toxicity detection;

SMILE-Next: Teaching Large Language Models to Detect, Classify, and Reason about Laughter

Lee Jung-Mok (KAIST), Tae-Hyun Oh (KAIST)

ClassificationRecognitionTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsVideoTextMultimodalityAudio

🎯 What it does: This paper constructs a multimodal laughter understanding dataset named SMILE-Next and designs three tasks: detection, type classification, and reasoning, aiming to achieve a comprehensive understanding of laughter in real social contexts.

SOAPTriage: SOAP-Guided Multi-View Clinical Text Modeling Framework for Automated ESI Prediction

Enming Wang (Shandong Computer Science Center (National Supercomputer Center in Jinan)), Wenpeng Lu (Harbin Institute of Technology)

ClassificationExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsTextTabularBiomedical DataElectronic Health RecordsRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposes SOAPTriage, a multi-perspective clinical text modeling method integrated with the SOAP framework, for automated prediction of the Emergency Severity Index (ESI), and generates natural language triage text from structured emergency department records to alleviate data scarcity.

SOAR: Supervision from Observation for Agentic Reinforcement Learning

Meng Li (Renmin University of China), Zang Li (Tencent)

OptimizationTransformerReinforcement LearningAgentic AIPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose the SOAR method, which in Agentic Reinforcement Learning treats environmental observations as learning signals, assigns positive advantages to observation tokens, and uses the negative entropy of the previous action as weights, encouraging the agent to consider the results of actions during learning;

SOCIA-EVO: Automated Simulator Construction via Dual-Anchored Bi-Level Optimization

Yuncheng Hua (University of New South Wales), Flora D. Salim (University of New South Wales)

Autonomous DrivingOptimizationAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringTextTabularTime SeriesSequentialRetrieval-Augmented Generation

🎯 What it does: Developed the SOCIA-EVO framework for automatically constructing simulators that satisfy distributed fidelity through large language models.

Social Dynamics as Critical Vulnerabilities that Undermine Objective Decision-Making in LLM Collectives

Changgeon Ko (Korea Advanced Institute of Science and Technology), Jong C. Park (Korea Advanced Institute of Science and Technology)

Explainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextMultimodalityBenchmark

🎯 What it does: Studied the social psychological weaknesses of large language model (LLM) agents in multi-agent decision-making, constructed a centralized framework to systematically evaluate the impact of social consensus, professional perception, speaking length, and rhetorical persuasion on the decision-making of a single representative agent.

Social Story Frames: Contextual Reasoning about Narrative Intent and Reception

Joel Mire (Carnegie Mellon University), Maarten Sap (Carnegie Mellon University)

ClassificationExplainability and InterpretabilityKnowledge DistillationTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: Constructed the Social Story Frames (SSF) framework, defined a 10-dimensional reader response classification, and implemented story reasoning and classification;

Soft Orthogonal Low-Rank Adaptation for Knowledge Sharing in Large Language Model Continual Learning

Yitong Wang (JIUTIAN Research, China Mobile), Junlan Feng (JIUTIAN Research, China Mobile)

ClassificationFederated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringMixture of ExpertsTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose the ASO-LoRA framework to achieve parameter-efficient continual learning on large language models, utilizing attribution scores to dynamically adjust the soft orthogonal constraints of the LoRA subspace, balancing catastrophic forgetting suppression and cross-task knowledge transfer.

Solve-Detect-Verify: Inference-Time Scaling with Flexible Generative Verifier

Jianyuan Zhong (Chinese University of Hong Kong), Qiang Xu (Chinese University of Hong Kong)

Computational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringGenerative Adversarial NetworkTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed a generative verifier called FlexiVe with a Solve-Detect-Verify (SDV) pipeline that dynamically allocates computational resources during inference, achieving efficient and reliable LLM inference.

SOS-LoRA: Static Orthogonal-Subspace Low-Rank Adaptation with Fixed Multi-Scale Scaling

Yupeng Chang (Jilin University), Yi Chang (Jilin University)

Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningMixture of ExpertsTextBenchmark

🎯 What it does: Propose a novel parameter-efficient fine-tuning method called SOS-LoRA based on LoRA, which splits the total rank into multiple static experts and introduces multi-scale scaling and cross-expert orthogonalization.

SoundBreak: A Systematic Study of Audio-Only Adversarial Attacks on Trimodal Models

Aafiya Shamshad Hussain (Virginia Tech), Chris Thomas (Virginia Tech)

Adversarial AttackTransformerMultimodalityAudio

🎯 What it does: This paper studies untargeted adversarial attacks that only modify the audio input for audio-visual-language trimodal models.

SPAGBias: Uncovering and Tracing Structured Spatial Gender Bias in Large Language Models

Binxian Su (Beijing Language and Culture University), Pengyuan Liu (Beijing Language and Culture University)

Recommendation SystemExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposed the SPAGBias framework to systematically evaluate gender bias in large language models within urban microspaces.

SPARD: Self-Paced Curriculum for RL Alignment via Integrating Reward Dynamics and Data Utility

Xuyang Zhi (University of Science and Technology of China), Enhong Chen (University of Science and Technology of China)

OptimizationData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningContrastive LearningTextBenchmark

🎯 What it does: Propose an adaptive RL alignment framework called SPARD, which dynamically adjusts multi-objective reward weights and data importance through real-time learning progress, achieving self-paced curriculum learning;

SPARK: Strategic Policy-Aware Exploration via Dynamic Branching for Long-Horizon Agentic Learning

Jinyang Wu (Tsinghua University), Jianhua Tao (Tsinghua University)

OptimizationRobotic IntelligenceTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AIPrompt EngineeringWorld ModelText

🎯 What it does: Propose the SPARK framework, which realizes dynamic branch exploration based on key decision points, utilizing the LLM's own <explore> signal to adaptively allocate exploration budget in long-horizon tasks;

SPARKLE: A Structured and Plug-and-play Agentic Retrieval Policy for Adaptive RAG Models

Jinyuan Fang (University of Glasgow), Craig Macdonald (University of Glasgow)

RetrievalKnowledge DistillationTransformerReinforcement LearningAgentic AIPrompt EngineeringTextGraphRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose SPARKLE, a lightweight agent model that achieves pluggable adaptive retrieval strategies through structured knowledge graph reasoning;

Sparrow: Text-Anchored Window Attention with Visual-Semantic Glimpsing for Speculative Decoding in Video LLMs

Libo Zhang (National University of Defense Technology), Dongsheng Li (National University of Defense Technology)

Computational EfficiencyRepresentation LearningTransformerLarge Language ModelMixture of ExpertsVision Language ModelDiffusion modelVideoTextMultimodality

🎯 What it does: Propose the Sparrow framework, which offloads visual computation to the target model through visual semantic internalization and hidden state reuse, achieving lossless acceleration for long video inference.

Sparse Feature Coactivation Reveals Causal Semantic Modules in Large Language Models

Ruixuan Deng (Georgia Institute of Technology), Joyce Chai (University of Michigan)

Explainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelAuto EncoderContrastive LearningText

🎯 What it does: Automatically identify modules in large language models related to concepts and relationships through sparse autoencoder feature co-activation patterns, and control model outputs by ablating/amplifying these modules.

Sparse-RL: Breaking the Memory Wall in LLM Reinforcement Learning via Stable Sparse Rollouts

Sijia Luo (Renmin University of China), Jing Zhang (Ant Group)

TransformerLarge Language ModelReinforcement LearningText

🎯 What it does: Studying how to use KV cache compression for sparse replay in large language model reinforcement learning to eliminate memory bottlenecks

Spatial-Agent: Agentic Geo-spatial Reasoning with Scientific Core Concepts

Riyang Bao (Emory University), Liang Zhao (Rutgers University)

Recommendation SystemAutonomous DrivingOptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningAgentic AIPrompt EngineeringTextGraphTabularBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Developed Spatial-Agent, a conceptual transformation framework based on core GIS concepts and functional roles, which can parse geographic problems into executable GeoFlow graphs and complete spatial reasoning through tool calls.

SPEAK: Spiking Neurons as an Entropy-Aware Tokenizer for Large Language Models

Ming Chen (Zhejiang University), Gang Pan (Zhejiang University)

CompressionRepresentation LearningSpiking Neural NetworkTransformerLarge Language ModelContrastive LearningTextMultimodality

🎯 What it does: Propose a gradient-trainable tokenizer called SPEAK based on spiking neurons, which can explicitly utilize and selectively retain historical tokenization information when making decisions about token boundaries.

SpeakerSleuth: Can Large Audio-Language Models Judge Speaker Consistency across Multi-turn Dialogues?

Jonggeun Lee (Seoul National University), Yohan Jo (Seoul National University)

ClassificationRecognitionAnomaly DetectionTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextMultimodalityBenchmarkAudio

🎯 What it does: Constructed the SpeakerSleuth benchmark to evaluate the ability of large audio-language models to judge speaker consistency in multi-turn dialogues.

Spec-o3: A Tool-Augmented Vision-Language Agent for Rare Celestial Object Candidate Vetting via Automated Spectral Inspection

Minghui Jia (Chinese Academy of Sciences), Dongbin Zhao (Chinese Academy of Sciences)

ClassificationAnomaly DetectionReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AIVision Language ModelImageTextMultimodalityBenchmarkPhysics RelatedRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose Spec-o3, an audio-visual language agent enhanced by an interactive tool, achieving multimodal chain-of-thought reasoning similar to astronomers' 'image thinking' in the spectral inspection of rare celestial candidate objects.

SpecAgent: A Speculative Retrieval and Forecasting Agent for Code Completion

George Ma (University of California Berkeley), Murali Krishna Ramanathan (Amazon Web Services)

RetrievalAI Code AssistantTransformerLarge Language ModelAgentic AITextSequentialRetrieval-Augmented Generation

🎯 What it does: Design SpecAgent to pre-build repository-level cross-file context during indexing, reducing retrieval latency during inference and improving code completion quality.

SpecCache: Speculative KV Cache Reuse for Efficient RAG Serving

Zijian Wen (University of Science and Technology of China), Jianyang

RetrievalComputational EfficiencyTransformerLarge Language ModelTextRetrieval-Augmented Generation

🎯 What it does: Propose SpecCache, a method that utilizes a lightweight inference model to guide the selective recomputation of KV cache in retrieval-augmented generation (RAG), significantly reducing prefill latency while maintaining generation quality.

Specializing Large Models for Oracle Bone Script Interpretation via Component-Grounded Multimodal Knowledge Augmentation

Jianing Zhang (Jilin University), Xi Yang (Jilin University)

RecognitionRetrievalExplainability and InterpretabilityTransformerLarge Language ModelAgentic AIVision Language ModelImageTextGraphRetrieval-Augmented Generation

🎯 What it does: This paper proposes an agent-driven retrieval-augmented generation framework (Agentic RAG), which achieves visual recognition of oracle bone characters, knowledge graph retrieval, component relationship reasoning, and semantic explanation generation, and constructs the OB-Radix component-level dataset containing 1,022 character images and 1,853 component images.

SpecMind: Cognitively Inspired, Interactive Multi-Turn Framework for Postcondition Inference

Cuong Chi Le (University of Texas at Dallas), Tien N. Nguyen (University of Texas at Dallas)

AI Code AssistantTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose an interactive multi-round LLM reasoning framework called SPECMIND, which generates more accurate and complete postconditions for programs using feedback-driven exploration and self-adaptive stopping.

Spectral Characterization and Mitigation of Sequential Knowledge Editing Collapse

Chi Zhang (Shandong University), Zhumin Chen (Shandong University)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelText

🎯 What it does: Spectral analysis of sequential knowledge editing in large language models reveals that the principal singular subspace is closely related to the model's general capabilities. Subsequently, the REVIVE framework is proposed, which filters updates based on the singular vectors of the original weights, preserving the principal singular subspace and significantly improving editing effectiveness while maintaining general capabilities.

Spectral Disentanglement: Rank-Aware Task Adaptation for Rehearsal-free Continual Learning in LLMs

Huanxuan Liao (Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences), Kang Liu (Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences)

Federated LearningComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningMixture of ExpertsText

🎯 What it does: Propose a replay-free continual learning framework called SPARTA, which decomposes shared and task-specific knowledge into low-rank and high-rank branches through spectral decomposition, achieving efficient adaptation for large language models.

Speculative End-Turn Detector for Efficient Speech Chatbot Assistant

Hyunjong Ok (POSTECH), Jaeho Lee (POSTECH)

RecognitionComputational EfficiencyRecurrent Neural NetworkTransformerPrompt EngineeringContrastive LearningTextAudio

🎯 What it does: Proposed an open-source dataset called OpenETD for end-turn detection, and introduced the SpeculativeETD two-stage inference framework to achieve efficient real-time detection;

Speech-Hands: A Self-Reflection Voice Agentic Approach to Speech Recognition and Audio Reasoning with Omni Perception

Zhen Wan (NVIDIA), Yu-Chiang Frank Wang (NVIDIA)

RecognitionTransformerSupervised Fine-TuningAgentic AIPrompt EngineeringTextMultimodalityRetrieval-Augmented GenerationAudio

🎯 What it does: Propose the Speech-Hands framework, endowing multimodal models with self-reflective decision-making capabilities to determine whether to use internal perception, external suggestions, or rewrite outputs, implemented in speech recognition and audio question-answering tasks.

SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation

Hui Wang (Nankai University), Yong Qin (Nankai University)

ClassificationExplainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextChain-of-ThoughtAudio

🎯 What it does: Propose the SpeechLLM-as-Judges framework, which uses large language models for multi-task, interpretable speech quality assessment; construct the SpeechEval dataset; train the SQ-LLM to perform four evaluation tasks (quality assessment, comparison, improvement suggestions, deepfake detection).

SpeechMedAssist: Efficiently and Effectively Adapting Speech Language Models for Medical Consultation

Sirry Chen (Fudan University), Zhongyu Wei (Peking University)

Domain AdaptationComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextBiomedical DataElectronic Health RecordsBenchmarkRetrieval-Augmented GenerationAudio

🎯 What it does: This paper proposes SpeechMedAssist, an end-to-end medical speech-language model that supports real-time multi-round voice-based medical consultations.

SPENCE: A Syntactic Probe for Detecting Contamination in NL2SQL Benchmarks

Mohammadtaher Safarzadeh (Oracle AI), Dan Roth

Explainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextTabularBenchmarkRetrieval-Augmented Generation

🎯 What it does: Through the SPENCE framework, systematically generate and rank syntactic variants, evaluate the execution accuracy of NL2SQL models at different syntactic distances, thereby detecting and quantifying dataset contamination phenomena.

SpiderFlow: Efficient Topology-Aware Scheduling for LLM Training Across Decentralized GPU Clusters

Zihan Chang (Zhejiang University), Zhe Pan (Zhejiang University)

OptimizationFederated LearningComputational EfficiencyGraph Neural NetworkTransformerLarge Language ModelReinforcement LearningContrastive LearningGraphTabular

🎯 What it does: SpiderFlow is a scheduling system for LLM training in decentralized GPU clusters, which achieves topology-aware scheduling at the task level and provides an automated model parallelization solution.

SpidR-Adapt: A Universal Speech Representation Model for Few-Shot Adaptation

Mahi Luthra (Meta AI), Emmanuel Dupoux (Meta AI)

Domain AdaptationRepresentation LearningMeta LearningTransformerSupervised Fine-TuningMixture of ExpertsContrastive LearningAudio

🎯 What it does: Propose SpidR-Adapt, an end-to-end speech representation model that leverages meta-learning, MAdaPT (multi-task adaptive pre-training), FOBLO (first-order bi-level optimization), and cross-supervision, enabling the model to quickly adapt to new languages with only a small amount of unlabeled audio.

SPIO: Ensemble and Selective Strategies via LLM-Based Multi-Agent Planning in Automated Data Science

Wonduk Seo (Enhans), Yi Bu (Peking University)

OptimizationHyperparameter SearchData-Centric LearningTransformerLarge Language ModelAgentic AIPrompt EngineeringMixture of ExpertsTabularChain-of-Thought

🎯 What it does: Proposed a multi-path planning and integration framework called SPIO based on LLM, which automates the construction of four modules in the machine learning pipeline: data preprocessing, feature engineering, model selection, and hyperparameter tuning. It supports two operational modes: single-path (SPIO-S) and ensemble (SPIO-E).

Splits! Flexible Sociocultural Linguistic Investigation at Scale

Eylon Caplan (Purdue University), Dan Goldwasser (Purdue University)

Explainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningTextRetrieval-Augmented Generation

🎯 What it does: Built an expandable Reddit social language research 'sandbox' — the SPLITS! dataset, and proposed a two-stage filtering process based on a dictionary to quickly screen out meaningful sociolinguistic hypotheses.

SPPO: Sequence-Level PPO for Long-Horizon Reasoning Tasks

Tianyi Wang (Southern University of Science and Technology), Guanhua Chen (Southern University of Science and Technology)

TransformerReinforcement LearningTextBenchmarkChain-of-Thought

🎯 What it does: Propose Sequence-Level PPO (SPPO), transforming long-chain reasoning tasks from token-level MDP into Sequence-Level Contextual Bandit, using a single scalar value function and unifying the advantage signal, solving the temporal credit assignment problem under sparse rewards in traditional PPO.

SQL-Trail: Multi-Turn Reinforcement Learning with Interleaved Feedback for Text-to-SQL

Harper Hua (Stanford University), Huzefa Rangwala (Amazon Web Services)

OptimizationAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextTabularBenchmark

🎯 What it does: Propose SQL-TRAIL, a multi-round reinforcement learning framework that iteratively improves the generation of text-to-SQL by interacting with databases and leveraging execution feedback;