ICML 2026 Papers — Page 52
International Conference on Machine Learning · 6554 papers
ScDiVa: Masked Discrete Diffusion for Joint Modeling of Single-Cell Identity and Expression
Mingxuan Wang (Renmin University of China), Yanbiao Ma (Renmin University of China)
GenerationRepresentation LearningTransformerDiffusion modelContrastive LearningBiomedical Data
🎯 What it does: Proposed scDiVa, a single-cell RNA-seq generation and representation learning framework based on masked discrete diffusion, capable of jointly modeling gene identity and expression levels;
Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models
Zhiwei Yang (Tencent Youtu Lab), Shouhong Ding (Tencent Youtu Lab)
Explainability and InterpretabilityRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision Language ModelImageTextMultimodalityGraphChain-of-Thought
🎯 What it does: Propose the Scene Graph Thinking (SaGe) framework, enabling multi-modal large language models (MLLMs) to perform fine-grained, structured visual reasoning through explicit scene graphs;
SceneDirector: Bridging Explicit Geometry and Generative Priors for Unified Driving Scene Editing
Yiyuan Liang (Huazhong University of Science and Technology), Changxin Gao (Huazhong University of Science and Technology)
Autonomous DrivingTransformerDiffusion modelRectified FlowAuto EncoderGenerative Adversarial NetworkVideoPoint Cloud
🎯 What it does: Propose SceneDirector, a unified method for simultaneously editing objects in driving videos (insertion, deletion, replacement, repositioning) and the owner's trajectory, achieving multi-view scene modification with a single inference.
ScenePilot: Controllable Boundary-Driven Critical Scenario Generation for Autonomous Driving
Qiyu Ruan (University of Macau), Cheng-zhong Xu
Autonomous DrivingOptimizationReinforcement LearningDiffusion modelScore-based ModelContrastive LearningPoint CloudTabularTime SeriesSequentialBenchmark
🎯 What it does: Proposes ScenePilot, a controllable boundary-driven scene generation framework guided by two signals: physical feasibility and AV risk, which generates high-risk scenarios that can cause deployed autonomous driving stacks to fail while ensuring physical feasibility.
SceneSmith: Agentic Generation of Simulation-Ready Indoor Scenes
Nicholas Pfaff (Massachusetts Institute of Technology), Russ Tedrake (Massachusetts Institute of Technology)
GenerationData SynthesisTransformerLarge Language ModelAgentic AIVision Language ModelDiffusion modelNeural Radiance FieldGenerative Adversarial NetworkImageTextMultimodalityPoint CloudMeshRetrieval-Augmented Generation
🎯 What it does: This paper proposes SceneSmith, a natural language generation system for indoor scenes based on a hierarchical agent framework;
Scheduling LLM Inference with Uncertainty-Aware Output Length Predictions
Haoyu Zheng (Wuhan University), Jiawei Jiang (Wuhan University)
OptimizationComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringText
🎯 What it does: Proposed the TIE scheduling algorithm, which models the output length of LLMs using a log-t distribution and combines CVaR with Tail-Inflated Expectation to achieve non-preemptive LLM inference scheduling.
Scheduling Thoughts: Learning the Order of Thought in Diffusion Language Models
Jiawei Xu (University of Maryland), Furong Huang (University of Maryland)
OptimizationAI Code AssistantReinforcement Learning from Human FeedbackTransformerSupervised Fine-TuningReinforcement LearningDiffusion modelTextChain-of-Thought
🎯 What it does: Propose a Self-Aware Scheduling (SAS) framework that learns and optimizes the 'thinking order' (unmasking schedule) during the decoding process of masked diffusion language models.
SCHUR-A*: Layer-wise Optimal Expert Pruning for MoEs via Schur-Complement Guided A* Search
Zheng Chen (University of Electronic Science and Technology of China), Buhui Yao (University of Electronic Science and Technology of China)
OptimizationComputational EfficiencyKnowledge DistillationTransformerMixture of ExpertsText
🎯 What it does: Propose a Schur-A* search algorithm based on Schur complement for hierarchical optimal expert pruning in MoE models.
SciAgentGym: Benchmarking Multi-Step Scientific Tool-Use in LLM Agents
Yujiong Shen (Fudan University), Yu-Gang Jiang (Fudan University)
Drug DiscoveryTransformerLarge Language ModelSupervised Fine-TuningAgentic AIPrompt EngineeringTextMultimodalityBenchmarkPhysics RelatedRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This paper constructs the SciAgentGym environment, integrating 1,780 professional scientific tools, and proposes the SciAgentBench evaluation suite to quantify the scientific reasoning ability of multi-step tool calls.
Scientific logicality enriched methodology for LLM reasoning: A practice in physics
Zhaoxin Yu (State Key Laboratory of Multimodal Artificial Intelligence Systems Institute of Automation Chinese Academy of Sciences), Wenji Mao (State Key Laboratory of Multimodal Artificial Intelligence Systems Institute of Automation Chinese Academy of Sciences)
TransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringContrastive LearningTextBenchmarkPhysics Related
🎯 What it does: Designed and implemented a logic evaluation and training method for large language models, constructing a physics problem set and the first multidimensional logic benchmark, PHYSLOGIC;
SciNet: Evaluating AI Agents in Relation-Aware Scientific Literature Retrieval
Chenyang Shao (Tsinghua University), Yong Li (Tsinghua University)
RetrievalRepresentation LearningGraph Neural NetworkTransformerLarge Language ModelAgentic AIPrompt EngineeringTextGraphTabularBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose the SciNet dataset to evaluate relation-aware capabilities in scientific literature retrieval, covering ego-centric, dual relationship, and path-level retrieval tasks;
SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences?
Udari Madhushani Sehwag (Scale AI), Bing Liu (Scale AI)
TransformerLarge Language ModelPrompt EngineeringTextBenchmarkPhysics RelatedRetrieval-Augmented Generation
🎯 What it does: This study proposes and constructs the SciPredict benchmark to evaluate the ability of large language models (LLMs) in predicting the results of natural science experiments;
SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models
Andong Deng (University of Central Florida), Xiaohan Wang (Stanford University)
Autonomous DrivingDrug DiscoveryReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringVision Language ModelVideoTextMultimodalityBenchmarkPhysics RelatedRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposed SCIVIDEOBENCH, a benchmark dataset for scientific video reasoning, and systematically evaluated the performance of various large multimodal models (LMMs) on this benchmark.
SCNS: Continual Personalization of Diffusion Models via Submodular Concept Neuron Selection
Zijie Peng (Sun Yat-sen University), Li Shen (Sun Yat-sen University)
GenerationOptimizationComputational EfficiencyTransformerDiffusion modelImageText
🎯 What it does: Propose a sparse neuron selection framework based on submodular optimization (SCNS), achieving continuous personalization for diffusion models, avoiding model fusion and catastrophic forgetting.
SCoA: Revisiting Domain Generalized Object Detection with Style-Conditioned Adaptation
Han Jiang (University of Science and Technology of China), Yongdong Zhang (University of Science and Technology of China)
Object DetectionDomain AdaptationConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImage
🎯 What it does: Propose the SCoA framework, achieving domain-generalizable object detection under single-source domain through dynamic and style-aware adaptation.
SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text
Miaobo Hu (Chinese Academy of Sciences), Daren Zha (Chinese Academy of Sciences)
Data-Centric LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextGraphBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposes the SCOPE benchmark and the SCION pipeline for inducing and fusing structured schemas from raw corpora based solely on training text;
SCOPE: Evolving Symbolic World for Planning in Open-Ended Environments
Yundaichuan Zhan (Zhejiang University), Yueting Zhuang (Zhejiang University)
Robotic IntelligenceTransformerReinforcement LearningWorld ModelImageTextMultimodalityBenchmark
🎯 What it does: Proposed the SCOPE framework, which dynamically evolves the symbolic world and adaptively enhances long-term planning capabilities through a closed-loop of perception, symbolic verification, and real-time execution.
SCOPE: Selective Conformal Optimized Pairwise LLM Judging
Sher Badshah (Dalhousie University), Hassan Sajjad (Dalhousie University)
ClassificationOptimizationExplainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringContrastive LearningTextBenchmark
🎯 What it does: Propose a reliability control framework for using LLMs in binary classification judgment, allowing acceptance of judgments under user-defined error rate thresholds and automatically abandoning uncertain cases.
Score Based Error Correcting Code Decoder
Alon Helvits (Ben Gurion University), Eliya Nachmani (Ben Gurion University)
OptimizationComputational EfficiencyData-Centric LearningTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningTabularTime SeriesPhysics RelatedStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Proposes a soft decoder based on score, named SB-ECC, which treats decoding as a continuous-time denoising process, iteratively recovering codewords in the symbol space through a probability flow ODE.
Score Correction for Generative Models with Probabilistic Constraints
Shishang Wu (Purdue University), Vinayak Rao (Purdue University)
GenerationData SynthesisOptimizationTransformerDiffusion modelScore-based ModelContrastive LearningImageTabularStochastic Differential Equation
🎯 What it does: This paper proposes a framework called DUALSCORE, which can generate distributions that satisfy given probabilistic constraints on the marginal distribution under random transformations, without retraining the original model, by applying a learnable additive correction to the pre-trained score function.
Score-Repellent Monte Carlo: Toward Efficient Non-Markovian Sampler with Constant Memory in General State Spaces
Jie Hu (Oakland University), Do Young Eun (North Carolina State University)
OptimizationComputational EfficiencyReinforcement Learning from Human FeedbackDiffusion modelScore-based ModelImagePoint CloudGraphTabularStochastic Differential Equation
🎯 What it does: Propose the Score-Repellent Monte Carlo (SRMC) framework, which utilizes score averaging to achieve constant memory history recording, and implements non-Markovian sampling through an exponentially tilted surrogate objective;
SCORE: A Unified Framework for Overshoot Refund in Online FDR Control
Qi Kuang (Fudan University), Yin Xia (Fudan University)
OptimizationFederated LearningComputational EfficiencyData-Centric LearningScore-based ModelTabularBiomedical Data
🎯 What it does: Proposed a unified online FDR control framework called SCORE, which enhances the power of hypothesis testing by compensating for the excess evidence in e-values.
ScoreMatchingRiesz: Score Matching for Debiased Machine Learning and Policy Path Estimation
Masahiro Kato (Mizuho-DL Financial Technology Co Ltd)
OptimizationFederated LearningExplainability and InterpretabilityData-Centric LearningReinforcement Learning from Human FeedbackDiffusion modelScore-based ModelTabularTime SeriesBenchmark
🎯 What it does: Proposed a Riesz representative estimator framework based on score matching, named ScoreMatchingRiesz, to assist in estimating causal and structural parameters in debiased learning.
ScoreMix: Synthetic Data Generation by Score Composition in Diffusion Models Improves Recognition
Parsa Rahimi (EPFL), Sébastien Marcel (EPFL)
RecognitionData SynthesisConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelContrastive LearningImage
🎯 What it does: Leverage the linear combination of scores in diffusion models to self-train and generate synthetic data to enhance the training of face recognition models
Scout Before You Attend: Sketch-and-Walk Sparse Attention for Efficient LLM Inference
Hoang Anh Duy Le (Rice University), Anshumali Shrivastava (Rice University)
Computational EfficiencyTransformerLarge Language ModelText
🎯 What it does: Propose a training-agnostic, unified sparse attention method called Sketch&Walk, which can significantly reduce computational and memory costs during the prefill and decoding phases in long-context reasoning of large language models (LLMs).
SCOUT: Active Information Foraging for Long-Text Understanding with Decoupled Epistemic States
Zhenliang Zhang (Peking University), Xiaojun Wan (Peking University)
Autonomous DrivingExplainability and InterpretabilityComputational EfficiencyData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposed the SCOUT framework, achieving efficient understanding and reasoning of million-scale long texts through active information acquisition and state decoupling.
SCOUT: Cyclic Causal Discovery Under Soft Interventions with Unknown Targets
Alpar Turkoglu (Georgia Institute of Technology), Faramarz Fekri (Georgia Institute of Technology)
OptimizationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningGraph Neural NetworkFlow-based ModelContrastive LearningGraphTabularBiomedical Data
🎯 What it does: Propose the SCOUT framework, which can learn nonlinear directed cyclic causal graphs from soft intervention data and simultaneously infer intervention targets without knowing the intervention objectives.
ScreenParse: Moving Beyond Sparse Grounding with Complete Screen Parsing Supervision
A. Said Gurbuz (IBM Research), Peter W. J. Staar
Object DetectionSegmentationComputational EfficiencyRepresentation LearningData-Centric LearningTransformerSupervised Fine-TuningVision Language ModelContrastive LearningImageMultimodality
🎯 What it does: Constructed a large-scale ScreenParse dataset and trained a lightweight ScreenVLM model for dense parsing of full-screen content.
SCRWKV: Ultra-Compact Structure-Calibrated Vision-RWKV for Topological Crack Segmentation
Hanxu Zhang (Tianjin University of Technology), Shengyong Chen (Tianjin University of Technology)
SegmentationConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningImage
🎯 What it does: Proposed an extremely lightweight SCRWKV network for pixel-level segmentation of structural cracks.
SD-MoE: Spectral Decomposition for Effective Expert Specialization
Ruijun Huang (Fudan University), Li Shang (Fudan University)
Computational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningMixture of ExpertsText
🎯 What it does: Analyze and address the issue of expert specialization failure in Mixture-of-Experts (MoE) models, proposing the Spectral-Decoupled MoE (SD-MoE) architecture.
SDiD:Shared diffusion prior for efficient distributed stereo image compression
Yichong Xia (Tsinghua University), Haoqian Wang (Tsinghua University)
Depth EstimationCompressionTransformerDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningImage
🎯 What it does: Propose the SDiD distributed 3D image compression framework, which achieves efficient compression by leveraging a shared pre-trained diffusion model prior
SE-GA: Memory-Augmented Self-Evolution for GUI Agents
Shilong Jin (Tianjin University), Zhuosheng Zhang (Shanghai Jiao Tong University)
Autonomous DrivingOptimizationRobotic IntelligenceTransformerReinforcement LearningAgentic AIVision Language ModelImageTextMultimodalityRetrieval-Augmented Generation
🎯 What it does: Propose the SE-GA framework, integrating Test-Time Memory Expansion (TTME) with Self-Evolving Training (MASE), enabling GUI agents to dynamically retrieve historical episodes, semantic, and experiential memories during multi-step tasks, and continuously learn and optimize through online interaction;
SE(3)-Equivariant Flow Matching with Gaussian Process Priors for Geometric Trajectory Prediction
Xuyang Wang (Shanghai Jiao Tong University), Jianping He (Shanghai Jiao Tong University)
Graph Neural NetworkFlow-based ModelGaussian SplattingGraphTime SeriesPhysics RelatedStochastic Differential Equation
🎯 What it does: Constructed a SE(3)-equivariant flow matching model GP-EquiFlow based on vector-valued Gaussian processes, for predicting the future trajectories of N-body systems from historical observations.
SE(n)-Invariant Flow Matching: A General Framework with Application to Object Reassembly
Gaël Heck (Universite Paris Saclay), Nicolas Lermé (Universite Paris Saclay)
RestorationGenerationFlow-based ModelImagePoint CloudMesh
🎯 What it does: A generative framework is constructed that performs flow matching on the SE(n) shape space, used for the task of assembling multi-fragment objects.
SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience
Zeyi Sun (Shanghai Jiao Tong University), Jiaqi Wang (Shanghai Innovation Institute)
Autonomous DrivingRobotic IntelligenceTransformerReinforcement LearningAgentic AIMixture of ExpertsVision Language ModelWorld ModelImageTextMultimodality
🎯 What it does: Designed and implemented SEAgent, a computer usage agent capable of self-evolution through autonomous exploration and experience learning within unknown software.
Search for Truth from Reasoning: A Dynamic Representation Editing Framework for Steering LLM Trajectories
Tianlong Wang (Peking University), Liantao Ma (Peking University)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningAI Code AssistantTransformerLarge Language ModelReinforcement LearningPrompt EngineeringContrastive LearningTextBenchmark
🎯 What it does: Proposed the dynamic representation editing framework DynaSteer, which utilizes entropy monitoring to identify high-uncertainty inference branch points, projects purified truth vectors at these nodes, and dynamically injects them to correct the LLM's reasoning trajectory, thereby reducing hallucinations.
Search or Accelerate: Confidence-Switched Position Beam Search for Diffusion Language Models
Mingyu Cao (University of Surrey), Lu Yin (University of Surrey)
Computational EfficiencyAI Code AssistantTransformerLarge Language ModelDiffusion modelScore-based ModelText
🎯 What it does: Proposed a confidence-adaptive position information beam search algorithm called SOAR, which dynamically switches between position search and parallel decoding during DLM decoding to achieve a better balance between quality and speed.
Search Space Synthesis for Parametric Functions
Felix Laarmann (TU Dortmund University), Jakob Rehof (TU Dortmund University)
OptimizationHyperparameter SearchData-Centric LearningNeural Architecture SearchReinforcement LearningContrastive LearningGaussian SplattingTabularTime Series
🎯 What it does: This paper proposes a framework based on Finite Combinatory Logic with Parameters (FCLP) and Para-construction, used to automatically synthesize the search space of parameterized functions and perform search within this space.
Search-R2: Enhancing Search-Integrated Reasoning via Actor-Refiner Collaboration
Bowei He (Mohamed bin Zayed University of Artificial Intelligence), Irwin King (Chinese University of Hong Kong)
Explainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningTextRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposes Search‑R2, a framework that enhances search-integrated reasoning through the collaboration of Actor and Refiner.
SecCodePRM: A Process Reward Model for Code Security
Weichen Yu (Carnegie Mellon University), Matt Fredrikson (Carnegie Mellon University)
Anomaly DetectionAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringContrastive LearningTextSequential
🎯 What it does: Proposes SecCodePRM, a process reward model that provides security scoring for each step in the code generation process, supporting vulnerability detection for both complete and partial code, as well as the generation of secure code.
Second-Order Bilevel Optimization with Accelerated Convergence Rates
Sheng Yang (University of California Riverside), John C.S. Lui (Chinese University of Hong Kong)
OptimizationHyperparameter SearchMeta LearningImageTabular
🎯 What it does: Proposes a full second-order Bayesian approximation method (FSBA) and its lazy variant (LFSBA), and generalizes them to non-convex-strongly convex Bayesian optimization and non-convex-strongly concave min-max problems, significantly improving the iteration and computational complexity.
Second-Order Smooth Planning with Optimal-Transport Bellman Smoothing
Tuan Quang Dam (Hanoi University of Science and Technology)
OptimizationReinforcement Learning
🎯 What it does: This paper proposes a Bellman backup based on optimal transport (OT) smoothing, and builds upon this to design the second-order SmoothCruiser algorithm, which can estimate the root state value under a generator model, theoretically achieving a better sample complexity.
Secure Multi-agent Reinforcement Learning for Service Systems with Affinity and Byzantine Nodes: Stability Analysis and Protection Design
Yifan Jiang (Shanghai Jiao Tong University), Li Jin (Shanghai Jiao Tong University)
Federated LearningSafty and PrivacyTransformerReinforcement LearningAgentic AIMixture of ExpertsTabularTime SeriesSequential
🎯 What it does: In service networks, a decentralized multi-agent reinforcement learning (MARL) framework is proposed that can handle work-server affinity issues in the presence of Byzantine nodes, while ensuring queue stability and learning convergence;
Securing Multi-Agent Systems Against Corruptions via Node Contribution Backpropagation
Chengcan Wu (Peking University), Meng Sun (Peking University)
Anomaly DetectionSafty and PrivacyReinforcement Learning from Human FeedbackGraph Neural NetworkLarge Language ModelPrompt EngineeringContrastive LearningTextGraph
🎯 What it does: A dynamic defense framework based on directed acyclic graphs (DAG) and backward propagation of node contributions (BPD) is proposed to address misinformation caused by malicious agents in multi-agent systems (MAS).
Securing Multimodal AI through Internal Information Decomposition
Jehyeok Yeon (University of Illinois Urbana Champaign), Heng Ji (University of Illinois Urbana Champaign)
Anomaly DetectionSafty and PrivacyTransformerVision Language ModelFlow-based ModelContrastive LearningImageTextMultimodality
🎯 What it does: Detect cross-modal attacks in multi-modal large language models by monitoring the internal information consistency between text and images during inference;
Security--Fidelity Tradeoffs: The Hidden Cost of Prompt Injection Defense
Mitchell Hermon (University of Illinois Urbana Champaign), Haohan Wang (University of Illinois Urbana Champaign)
Safty and PrivacyExplainability and InterpretabilityComputational EfficiencyData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextBenchmark
🎯 What it does: This paper proposes the SECFID benchmark, which distinguishes three behaviors of LLMs when facing indirect prompt injection: execution (EXECUTED), processing as data (PROCESSED), and ignoring (IGNORED), thereby enabling the simultaneous measurement of safety and fidelity.
SEDRAS: Symbolically Evaluated Deep Research And Science
Fredrik Carlsson (Research Institutes of Sweden), Joakim Nivre (Uppsala University)
Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningData-Centric LearningDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringAuto EncoderTextMultimodalityTabularBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposed a modular and procedurally generated scientific discovery evaluation framework called SEDRAS, which combines controllable latent data distributions, different representation methods, and interactive dynamics, requiring LLMs to generate interpretable theories in context and translate them into code for verification.
See First, Reason Later: Mutual Information-Guided Reinforcement Learning for Vision-Language Models
Yin Zhang (East China Normal University), Yilei Shao (East China Normal University)
Computational EfficiencyRepresentation LearningReinforcement LearningVision Language ModelContrastive LearningImageTextMultimodalityBenchmarkChain-of-Thought
🎯 What it does: This work proposes a reinforcement learning framework called MIRL based on mutual information, aimed at improving the visual perception quality and reasoning accuracy of vision-language models.
See More, Forecast Better and Faster: Enhancing Time Series Foundation Models via Inference-Time Plug-and-Play Downsampling
Longlong Xu (Tsinghua University), Dan Pei (Tsinghua University)
Computational EfficiencyTransformerPrompt EngineeringAuto EncoderContrastive LearningTime SeriesBenchmark
🎯 What it does: Propose a training-agnostic, plug-and-play framework called SPRINT, which achieves long-term time series forecasting by utilizing downsampling, trend-seasonal decomposition, resolution interpolation, and pattern replication.
See the Emotion: A Facial Emoji Proxy Modeling for EEG Emotion Recognition
Jingjing Hu (Hefei University of Technology), Meng Wang (Hefei University of Technology)
RecognitionExplainability and InterpretabilityConvolutional Neural NetworkDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningMultimodalityTime SeriesBiomedical DataElectrocardiogram
🎯 What it does: A cross-modal model that generates facial expression emojis from EEG signals was constructed, and the generation process was used as an interpreter to visualize emotional states.
See What Matters: Differentiable Grid Sample Pruning for Generalizable Vision-Language-Action Model
Yixu Feng (University of Sydney), Chang Xu (University of Sydney)
Computational EfficiencyRobotic IntelligenceTransformerSupervised Fine-TuningVision-Language-Action ModelDiffusion modelFlow-based ModelRectified FlowContrastive LearningImageVideoText
🎯 What it does: Proposes a differentiable grid sampler (GridS) module, which performs geometry-aware continuous resampling of visual tokens in Vision-Language-Action (VLA) models, achieving high compression rates without performance loss.
See, Act, Adapt: Active Perception for Unsupervised Cross-Domain Visual Adaptation via Personalized VLM-Guided Agent
Tianci Tang (Zhejiang University), Gaoang Wang (Zhejiang University)
Domain AdaptationRobotic IntelligenceTransformerReinforcement LearningAgentic AIVision Language ModelContrastive LearningImageMultimodality
🎯 What it does: By combining pre-trained vision and language models (VLM), an agent capable of actively controlling the camera perspective in indoor environments was constructed. The agent performs unsupervised reinforcement learning using scalar feedback (confidence and geometric consistency) from a frozen perception module, enabling adaptive performance across visual tasks.
Seeing is Solving: Unlocking Efficient Multimodal RL via View Alignment
Qinsi Wang (Duke University), Wentian Zhao (Adobe Inc)
Computational EfficiencyRepresentation LearningData-Centric LearningTransformerSupervised Fine-TuningReinforcement LearningVision Language ModelContrastive LearningGaussian SplattingImageTextMultimodalityChain-of-Thought
🎯 What it does: This paper proposes a reinforcement learning fine-tuning framework called FOCUS-RL based on view alignment, which utilizes dynamic text-visual alignment information from VLM to accelerate the training of multi-modal reasoning models.
Seeing is Understanding: Unlocking Causal Attention into Modality-Mutual Attention for Multimodal LLMs
Wei-Yao Wang (Sony Group Corporation), Yoshiyuki Kobayashi (Sony Group Corporation)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningVision Language ModelImageTextMultimodality
🎯 What it does: Proposed a Modal Mutual Attention (MMA), which unlocks the causal attention in the decoder, allowing image tokens to attend to text tokens, thereby alleviating the vision-language mismatch and target misreporting issues in Multimodal Large Language Models (MLLMs).
Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation
Chenyu Hui (Shanghai Jiao Tong University), Chang Xu (University Of Sydney)
Data SynthesisDomain AdaptationRobotic IntelligenceTransformerPrompt EngineeringVision-Language-Action ModelDiffusion modelScore-based ModelGenerative Adversarial NetworkContrastive LearningImageVideoTextMultimodality
🎯 What it does: This paper proposes an efficient video transfer framework that converts VLA videos from simulated environments into realistic videos to enhance training data while preserving task semantics and action trajectories.
Seeing Symbols, Missing Structure: A Real-World Handwritten Mathematical Expression Recognition Benchmark for Large Models
Sheng Jiang (Beijing Normal University), Hua Huang (Beijing Normal University)
RecognitionTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageMultimodalityBenchmarkChain-of-Thought
🎯 What it does: Proposed a structured diagnostic dimension-based benchmark for real handwritten mathematical expressions, HMER-Bench, and introduced an untrained Schema-Anchored Structure-Aware reasoning framework, SASR, on this benchmark, significantly improving the recognition stability of large models on complex two-dimensional structures.
Seeing the Unseen: Physics-as-Representation for Generalizable Gaze Perception
Yunfeng Xiao (Tianjin University), Erwei Yin (Tianjin Artificial Intelligence Innovation Center)
Pose EstimationDomain AdaptationExplainability and InterpretabilityRepresentation LearningConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageVideoPhysics Related
🎯 What it does: Proposed the SG-Gaze framework, adopting the physics-as-representation approach, and learning a unified structural and geometric consistent representation (SGR) through dual branches (analytical and structural), achieving interpretable and well-generalizable pupil direction estimation.
Seeing to Generalize: How Visual Data Corrects Binding Shortcuts
Nicolás Buzeta, Rodrigo Toro Icarte (Pontificia Universidad Católica)
RetrievalExplainability and InterpretabilityRepresentation LearningTransformerVision Language ModelContrastive LearningImageTextMultimodality
🎯 What it does: Investigate the performance of vision-language models (VLM) in pure text retrieval tasks, and reveal how image training enhances the generalization ability of text reasoning through comparative experiments and interpretability methods.
Seeing Without Understanding: Disentangling Perception, Reasoning, and Simulation in VLM Gameplay
Dingyang Jin (Northeastern University), Ryan Rad (Northeastern University)
Explainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerPrompt EngineeringVision Language ModelImageTextMultimodalityChain-of-Thought
🎯 What it does: Proposed a two-stage diagnostic framework to separate the perception, rule reasoning, and simulation processes of VLMs in games, and evaluated model performance through systematic perception tests and rule complexity staircases.
Seeking Commonality, Preserving Specificity: A Spectral-Aware Hierarchical Framework for Cross-City Road Representation Learning
Jingtian Ma (Beihang University), Leong Hou U (University of Macau)
Autonomous DrivingRepresentation LearningGraph Neural NetworkTransformerAuto EncoderContrastive LearningGraphSequential
🎯 What it does: Propose the CoSpec framework to achieve unified road representation learning across cities, utilizing a hierarchical structure to decompose road networks into low-frequency common features and high-frequency city-specific details, and achieving spectral decoupling through a dual-path reconstruction.
SEER: Transformer-based Robust Time Series Forecasting via Automated Patch Enhancement and Replacement
Xiangfei Qiu (ECNU), Jilin Hu (ECNU)
Anomaly DetectionOptimizationTransformerMixture of ExpertsTabularTime SeriesBenchmarkFinance Related
🎯 What it does: Built SEER, a robust time series forecasting framework based on Transformer, which can automatically enhance and replace low-quality patches to improve prediction accuracy.
Seg-ReSearch: Segmentation with Interleaved Reasoning and External Search
Tianming Liang (Sun Yatsen University), Wei-Shi Zheng (Sun Yatsen University)
SegmentationOptimizationData-Centric LearningTransformerLarge Language ModelReinforcement LearningPrompt EngineeringVision Language ModelImageVideoTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose Seg‑ReSearch, an agent-based segmentation framework that can locate and segment targets in videos/images through multi-round reasoning and interaction with external search.
Segment Anything with Robust Uncertainty-Accuracy Correlation
Hongyou Zhou (Technical University Of Berlin), Zihan Ye (University Of Chinese Academy Of Sciences)
SegmentationDomain AdaptationAnomaly DetectionTransformerDiffusion modelGenerative Adversarial NetworkContrastive LearningImageVideo
🎯 What it does: Building a robust uncertainty-accuracy correlation (RUAC) framework by adding a Bayesian mask decoder to SAM2 and combining it with bio-inspired adversarial style and deformation perturbations;
Segment-Aligned Policy Optimization for Multi-Modal Reasoning
Lei Gao (Fudan University), Xuelong Li (China Telecom)
OptimizationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelReinforcement LearningPrompt EngineeringScore-based ModelContrastive LearningTextMultimodalityChain-of-Thought
🎯 What it does: This paper proposes a Segment-Aligned Policy Optimization (SAPO) framework, which uses thinking steps as the basic unit of RL for reinforcement learning of large language models in multimodal reasoning tasks.
Segment-driven Structural Induction and Semantic Alignment for Heterogeneous Tabular Representation
Woojun Jung (Korea University), Susik Yoon (Korea University)
Representation LearningData-Centric LearningTransformerContrastive LearningTabular
🎯 What it does: Proposes a table pre-training framework called NAVE based on head-value segments, aiming to enhance cross-table domain-specific semantics through structured segment masking and entropy-driven segment alignment learning.
SegPVSG: Panoptic Video Scene Graph Generation via Temporal Focusing and Generative Augmentation
YiKai Li, Shuangping Huang (South China University of Technology)
Object DetectionObject TrackingSegmentationGenerationData SynthesisConvolutional Neural NetworkTransformerLarge Language ModelPrompt EngineeringDiffusion modelGenerative Adversarial NetworkContrastive LearningGaussian SplattingImageVideoTextMultimodalityRetrieval-Augmented Generation
🎯 What it does: Propose a new framework called SegPVSG for Panoptic Video Scene Graph Generation (PVSG), which includes two modules: Temporal Focus Network (TFN) and Relation-Centric Generative Video Augmentation (RGVA);
SeisMark: A Large-Scale Open Benchmark for Robust 3D Seismic Fault Detection
Min Jun Park (X Moonshot Factory), Kevin F. Smith (X Moonshot Factory)
SegmentationData SynthesisConvolutional Neural NetworkDiffusion modelScore-based ModelGenerative Adversarial NetworkContrastive LearningPoint CloudMeshBenchmark
🎯 What it does: Proposes SeisMark, a large-scale open benchmark for 3D seismic fault detection;
Seizure-Semiology-Suite($S^3$): A Clinically Multimodal Dataset, Benchmark, and Models for Seizure Semiology Understanding
Lina Zhang (University of California, Los Angeles), Vwani Roychowdhury (University of California, Los Angeles)
ClassificationRecognitionAnomaly DetectionConvolutional Neural NetworkRecurrent Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringVideoTextMultimodalityBiomedical DataBenchmark
🎯 What it does: Proposed Seizure‑Semiology‑Suite (S³), which includes 438 seizure video clips and over 35,000 dense ILAE semiology labels, and designed a seven-level clinical evaluation benchmark along with the Seizure RQI metric;
Select to Think: Unlocking SLM Potential with Local Sufficiency
Wenxuan Ye (Technical University of Munich), Yunpu Ma (Ludwig Maximilian University of Munich)
Computational EfficiencyKnowledge DistillationTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringText
🎯 What it does: This paper proposes the SELECT TO THINK (S2T) framework, which converts the assistance of large language models (LLMs) during the reasoning process into a discrete selection among top-K candidate words generated by small language models (SLMs), thereby eliminating the need for high-latency calls to LLMs. Subsequently, the paper further introduces S2T-LOCAL, which internalizes the selection logic into SLMs via ZIP technology, achieving fully teacher-free reasoning. It also verifies the 'local sufficiency' hypothesis, which states that in most critical positions, the optimal vocabulary of LLMs is often already included in the top-K of SLMs.
Selecting Samples on Graphs: A Unified Dataset Pruning Framework for Lossless Training Acceleration
Dongyue Wu (Huazhong University of Science and Technology), Changxin Gao (Huazhong University of Science and Technology)
ClassificationComputational EfficiencyData-Centric LearningGraph Neural NetworkContrastive LearningImageGraph
🎯 What it does: This paper proposes UGIES, a unified graph-based sample screening framework that enables lossless pruning of large training datasets to accelerate model training;
Selective Coupling of Decoupled Informative Regions: Masked Attention Alignment for Data-Free Quantization of Vision Transformers
Biao Qian (Tsinghua University), Jungong Han (Tsinghua University)
ClassificationData SynthesisCompressionComputational EfficiencyTransformerAuto EncoderContrastive LearningImage
🎯 What it does: Propose a data-agnostic quantization method called MaskAQ, which utilizes information regions in self-attention to perform masked attention alignment, thereby generating high-quality synthetic samples for low-bit quantization of Vision Transformers.
Selective Deferred Routing: Enabling Cost-Efficient Collaboration between Local SLMs and Remote LLMs
Qijun Miao (Tsinghua University), Zhixuan Fang (Tsinghua University)
Recommendation SystemFederated LearningComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelMixture of ExpertsTextRetrieval-Augmented Generation
🎯 What it does: Propose the Selective Deferred Routing (SDR) scheme: first use a local small language model (SLM) to generate an initial answer and provide semantic features such as hidden states, and then use a lightweight decision module to determine whether to directly return the local result or route the request to the most suitable remote large language model (LLM) to obtain a higher quality answer.
Selective Disclosure Watermarking for Large Language Models
Xuyang Chen (University of Pennsylvania), Qi Long (University of Pennsylvania)
Safty and PrivacyTransformerLarge Language ModelPrompt EngineeringGenerative Adversarial NetworkText
🎯 What it does: Propose a multi-bit watermarking method based on hierarchical vocabulary routing (HeRo), which supports embedding hidden metadata that can be progressively decoded according to permission levels in text generated by LLMs.
Self-Augmenting Retrieval for Diffusion Language Models
Paul Jünger (Cornell University), Kilian Q Weinberger (Cornell University)
GenerationRetrievalComputational EfficiencyTransformerPrompt EngineeringDiffusion modelTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: A self-enhanced retrieval framework (SARDI) for discrete diffusion language models was designed, which dynamically refreshes retrieval results at each denoising step by using unconfirmed low-confidence predictions as retrieval queries, ultimately submitting only high-confidence words during generation.
Self-Calibrated Consistency can Fight Back for Adversarial Robustness in Vision-Language Models
Jiaxiang Liu (Guangdong Institute of Intelligence Science and Technology), Mingkun Xu (Guangdong Institute of Intelligence Science and Technology)
Representation LearningAdversarial AttackPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBenchmark
🎯 What it does: Proposed an effective test-time defense method called Self-Calibrated Consistency (SCC) to enhance the robustness of vision-language models (VLMs) against adversarial attacks.
Self-Captioning Multimodal Interaction Tuning: Amplifying Exploitable Redundancies for Robust Vision Language Models
Yuriel Ryan (Singapore University Of Technology And Design), Roy Ka-Wei Lee (Singapore University Of Technology And Design)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerSupervised Fine-TuningVision Language ModelContrastive LearningImageTextMultimodality
🎯 What it does: This paper proposes a self-descriptive multimodal interaction optimization framework, which utilizes the Multimodal Interaction Gate to convert visually unique information into shared information, thereby systematically increasing redundancy in the training data;
Self-correcting for Debiasing Large Language Models
Xuan Feng (Jinan University), Bo An (Nanyang Technological University)
Federated LearningExplainability and InterpretabilityData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningContrastive LearningTextBenchmarkChain-of-Thought
🎯 What it does: This paper proposes the SELFDEBIAS framework, which introduces a self-correction mechanism into the chain-of-thought reasoning of large language models to achieve bias mitigation against social prejudices;
Self-CriTeach: LLM Self-Teaching and Self-Critiquing for Improving Robotic Planning via Automated Domain Generation
Jinbang Huang (Huawei Noah's Ark Lab), Yingxue Zhang (Huawei Noah's Ark Lab)
Robotic IntelligenceTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningDiffusion modelTextChain-of-Thought
🎯 What it does: Propose the Self-CriTeach framework, which utilizes LLMs to automatically generate PDDL planning domains. It generates symbolic task-plan data for supervised fine-tuning and uses the same domain as a structured reward for RL self-critique, thereby enhancing the LLM's robotic planning capabilities without human annotation.
Self-Distillation Enables Continual Learning
Idan Shenfeld (Massachusetts Institute Of Technology), Pulkit Agrawal (Massachusetts Institute Of Technology)
Computational EfficiencyKnowledge DistillationRepresentation LearningMeta LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringContrastive LearningTextTabularSequentialBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose a self-distillation fine-tuning (SDFT) method that leverages the model's own in-context learning capability to convert expert demonstrations into on-policy training signals, thereby achieving continuous learning.
Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
Siyan Zhao (Ucla), Aditya Grover (Ucla)
Computational EfficiencyKnowledge DistillationRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextChain-of-Thought
🎯 What it does: By allowing a single large language model to simultaneously act as both teacher and student in different contexts, the paper achieves self-supervised on-policy inductive learning by aligning the student's own generation trajectory word-by-word using the 'privileged information' of question-answer pairs.
Self-evolving LLM agents with in-distribution Optimization
Yudi Zhang (Eindhoven University of Technology), Mykola Pechenizkiy (Eindhoven University of Technology)
OptimizationTransformerLarge Language ModelReinforcement LearningAgentic AITextSequentialRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposes Q-Evolve, a self-evolving LLM agent framework that can automatically generate process-level rewards and perform policy updates within the same distribution based solely on terminal rewards in long-sequence tasks.
Self-Guidance: Enhancing Neural Codecs via Decoder Manifold Alignment
Xiang Li (Tsinghua University), Hui Wang (Pengcheng Laboratory)
CompressionAdversarial AttackConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningAudio
🎯 What it does: Introduce a self-guidance mechanism into the neural speech codec based on VQ-VAE. During training, the internal features of the decoder are kept consistent when processing quantized and unquantized latent vectors, thereby reducing the impact of quantization error on reconstruction quality.
Self-Prompting Diffusion Transformer for Open-Vocabulary Scene Text Edit via In-Context Learning
Hongxi Li (Meitu Inc.), Ting Liu (Meitu Inc.)
Image TranslationRestorationGenerationTransformerPrompt EngineeringDiffusion modelContrastive LearningImageTextMultimodality
🎯 What it does: This paper proposes a Self-Prompting Diffusion Transformer, which achieves open-vocabulary editing of scene text by directly constructing style and glyph prompts from the original image, while preserving the original text's color, font, and texture.
Self-Prophetic Decoding to Unlock Visual Search in LVLMs
Zhendong He (Sun Yat-sen University), Sibei Yang (Sun Yat-sen University)
RetrievalTransformerPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmarkChain-of-Thought
🎯 What it does: Proposes the SeProD framework, which achieves multi-step reasoning in visual search tasks through self-fulfilling decoding of pre-trained and post-trained large vision-language models.
Self-Refining Video Sampling
Sangwon Jang (KAIST), Sung Ju Hwang (KAIST)
GenerationData SynthesisTransformerDiffusion modelScore-based ModelFlow-based ModelRectified FlowAuto EncoderVideoStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Propose a sampling method called Self-Refining Video Sampling, which uses the pre-trained video generator itself during the inference phase to iteratively refine at each time step, leveraging the denoising autoencoder properties of the flow matching model.
Self-Soupervision: Cooking Model Soups without Labels
Anthony Fuller (Carleton University), Evan Shelhamer (Carleton University)
ClassificationDomain AdaptationRepresentation LearningData-Centric LearningConvolutional Neural NetworkTransformerSupervised Fine-TuningMixture of ExpertsAuto EncoderContrastive LearningImage
🎯 What it does: Propose the SelfSouping idea, generating diverse model 'ingredients' using self-supervised learning on unlabeled data, and then fine-tuning on labeled tasks and linearly mixing parameters to form a 'model soup' that improves robustness and transfer performance.
Self-Supervised Dynamical System Representations for Physiological Time-Series
Yenho Chen (Georgia Institute of Technology), Christopher John Rozell (Georgia Institute of Technology)
Anomaly DetectionComputational EfficiencyRepresentation LearningData-Centric LearningRecurrent Neural NetworkTransformerAuto EncoderContrastive LearningTime SeriesBiomedical DataElectrocardiogramPhysics RelatedStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: A new self-supervised pre-training framework called PULSE is studied, which utilizes dynamic system models and cross-reconstruction tasks to extract transferable system information and suppress sample-specific noise, specifically for physiological time series.
Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis
Hila Chefer (Black Forest Labs), Robin Rombach
GenerationData SynthesisRepresentation LearningTransformerDiffusion modelScore-based ModelFlow-based ModelRectified FlowAuto EncoderContrastive LearningImageVideoTextMultimodalityAudio
🎯 What it does: Propose Self-Flow, a self-supervised flow matching framework that can simultaneously learn strong representations and high-quality generation across three modalities—image, video, and audio—without relying on external pre-trained encoders;
Self-supervised Hierarchical Visual Reasoning with World Model
Yuanfei Xu (University of Science and Technology of China), Houqiang Li (University of Science and Technology of China)
Robotic IntelligenceRecurrent Neural NetworkTransformerReinforcement LearningAuto EncoderContrastive LearningWorld ModelImageVideo
🎯 What it does: Built a self-supervised hierarchical world model called ResDreamer, which achieves layer-by-layer visual prediction through residual learning.
Self-Supervised Learning as Discrete Communication
Kawtar Zaher (INRIA, LIRMM, Université de Montpellier), Alexis Joly (INRIA, LIRMM, Université de Montpellier)
ClassificationRetrievalDomain AdaptationRepresentation LearningTransformerAuto EncoderContrastive LearningImageVideo
🎯 What it does: View self-supervised learning as a discrete communication process where semantic information is transmitted between a teacher and student through a fixed-capacity binary channel, with the student learning representations via a multi-label binary objective
Self-Supervised Weight Templates for Scalable Vision Model Initialization
Yucheng Xie (Southeast University), Xin Geng (Southeast University)
ClassificationObject DetectionSegmentationGenerationConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningImage
🎯 What it does: Proposes the SWEET framework, which learns cross-size and cross-task weight templates via self-supervised Tucker constraints, and achieves scalable initialization through width random scaling.
SelfJudge: Faster Speculative Decoding via Self-Supervised Judge Verification
Kanghoon Yoon (KAIST), Dongsoo Lee (NAVER Cloud)
GenerationComputational EfficiencyTransformerLarge Language ModelScore-based ModelTextRetrieval-Augmented Generation
🎯 What it does: Propose SelfJudge, a Speculative Decoding method that accelerates LLM inference by training a discriminator in a self-supervised manner
Selling Data as a Digital Good with Scaling Valuations
Ningyuan Li (Peking University), Jie Zhang (University of Bath)
OptimizationFederated LearningData-Centric Learning
🎯 What it does: Studies how to sell and price digital goods under the scale effect of data volume;
Sem-Detect: Semantic Level Detection of AI Generated Peer-Reviews
Andre Vicente Duarte (Carnegie Mellon University), Lei Li (Carnegie Mellon University)
ClassificationAnomaly DetectionExplainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningTextReview/Survey PaperBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose the Sem-Detect method to identify the authorship of peer review texts (human-written, LLM-polished, fully AI-generated).
SEMA: a Scalable and Efficient Mamba like Attention via Token Localization and Averaging
Nhat Thanh Tran (University of California, Irvine), Jack Xin (University of California, Irvine)
ClassificationObject DetectionSegmentationComputational EfficiencyTransformerMixture of ExpertsContrastive LearningImageVideo
🎯 What it does: Propose a scalable and efficient Mamba-style attention mechanism called SEMA, which utilizes window localization to avoid attention dispersion and achieves global information fusion through arithmetic averaging (mean mixing), addressing the problem of attention focus loss in long sequences, and verifying its effectiveness in various visual tasks.
Semantic Cache Distillation: Efficient State Transfer via Reuse and Selective Patching
Qianli Ma (Beijing Normal University), Weijia Jia (Beijing Normal University)
Computational EfficiencyKnowledge DistillationTransformerLarge Language ModelAuto EncoderContrastive LearningTextRetrieval-Augmented Generation
🎯 What it does: In separated LLM inference, a Semantic Cache Distillation (SCD) framework is constructed between producers and consumers with shared architecture but different weights, achieving efficient KV cache transmission and semantic alignment;
Semantic Editing with Coupled Stochastic Differential Equations
Jianxin Zhang (University of Michigan), Clayton Scott (University of Michigan)
GenerationTransformerDiffusion modelFlow-based ModelImageStochastic Differential Equation
🎯 What it does: A method for image semantic editing based on coupled stochastic differential equations (SDEs) is proposed, aiming to guide the sampling process of pre-trained generative models through shared reverse Brownian motion paths, thereby generating new images that conform to target prompts while maintaining visual similarity to the source image.
Semantic Granularity Navigation in Image Editing
Liangsi Lu (Guangdong University of Technology), Yang Shi (Guangdong University of Technology)
Image TranslationImage HarmonizationRestorationGenerationTransformerPrompt EngineeringDiffusion modelScore-based ModelFlow-based ModelImageBenchmark
🎯 What it does: Propose a training-free inference-time controller called NaviEdit, which achieves semantically stronger image editing while maintaining structural integrity by decoupling the editing progress from model scale and reallocating computational resources to effective scale windows.
Semantic Impact–Driven Visual Scheduling in Vision-Language Models
Xuan Wang (Beijing University of Posts and Telecommunications), Xiaojie Wang (Beijing University of Posts and Telecommunications)
Explainability and InterpretabilityComputational EfficiencyTransformerVision Language ModelContrastive LearningImageTextMultimodality
🎯 What it does: Propose a training-agnostic inference-time visual KV scheduling method called SIVS, which quantifies the semantic impact of visual information on model decisions, dynamically retaining the most important visual tokens for the current inference step.
Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression
Xiang Liu (Hong Kong University of Science and Technology Guangzhou), Xiaowen Chu (Hong Kong University of Science and Technology Guangzhou)
CompressionComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This work proposes the KVFundaBench benchmark to systematically evaluate the impact of KV cache compression on high-density inference and fundamental capabilities. Based on this, it is found that semantic integrity is crucial for inference performance, leading to the design of the ShotKV compression scheme. In the prefill stage, complete semantic units are maintained at the shot level, while dynamic token-level compression is used during the decoding stage. This approach significantly improves accuracy and reduces latency in various long-context inference and generation tasks.