arXivSub Start free trial

ICML 2026 Papers — Page 54

International Conference on Machine Learning · 6554 papers

Skewness-Robust Causal Discovery in Location-Scale Noise Models

Daniel Klippert (TU Dortmund University), Alexander Marx (TU Dortmund University)

OptimizationExplainability and InterpretabilityComputational EfficiencyData-Centric LearningDiffusion modelScore-based ModelFlow-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningOptical FlowGraphTabularTime SeriesSequentialBenchmarkStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Proposed a location-scale noise model (LSNM) based on the skew-normal distribution and the corresponding causal direction inference method SKEWD.

Ski Rental with Distributional Predictions of Unknown Quality

Qiming Cui (Johns Hopkins University), Michael Dinitz (Johns Hopkins University)

OptimizationTabularTime SeriesBenchmark

🎯 What it does: Designed an online algorithm that utilizes distribution prediction to address the ski rental problem, and provides the optimal leasing timing based on the predicted distribution.

Skill Neologisms: Towards Skill-based Continual Learning

Antonin Berthon (University of Cambridge), Mihaela van der Schaar (University of Cambridge)

Computational EfficiencyKnowledge DistillationRepresentation LearningData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringContrastive LearningText

🎯 What it does: By adding soft words (skill neologisms) to the vocabulary of pre-trained large language models and training using a skill-centralized dataset without updating model parameters, new compositional skills are learned and their zero-shot composition ability is verified on synthetic tasks and natural language tasks.

Skill-Based Mixture-of-Experts: Adaptive Routing for Heterogeneous Reasoning via Inferred Skills

Justin Chen, Mohit Bansal (University Of North Carolina Chapel Hill)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningMeta LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringMixture of ExpertsTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose SKILL-MOE, a Mixture-of-Experts framework that utilizes existing pre-trained large language models during the inference phase. It achieves instance-level expert selection by inferring the discrete skills required for each question, and synthesizes the final answer by aggregating the Chain-of-Thought results generated by each expert.

Skill-Pro: Learning Reusable Skills from Experience via Non-Parametric PPO for LLM Agents

Qirui Mi (Shanghai University of Finance and Economics), Jun Wang (University College London)

Autonomous DrivingComputational EfficiencyRepresentation LearningRobotic IntelligenceTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Designed the Skill-Pro framework, which automatically learns and reuses executable procedural skills (Skill) in LLM agents using non-parametric PPO.

SkillNet: Hierarchical Skill Modeling for Compositional Generalization in Vision-Language Action Models

Senwei Xie (Chinese Academy of Sciences), Xilin CHEN

Robotic IntelligenceReinforcement Learning from Human FeedbackTransformerMixture of ExpertsVision-Language-Action ModelFlow-based ModelContrastive LearningImageVideoTextMultimodality

🎯 What it does: Propose SkillNet, a visual-language-action model that achieves compositional generalization through hierarchical skill modeling (mechanical attributes and semantic attributes) combined with a skill-contextualized Mixture-of-Experts (SCMoE).

SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems

Yunhao Feng (National University of Defense Technology), Wenke Huang (Wuhan University)

Safty and PrivacyAdversarial AttackLarge Language ModelAgentic AIPrompt EngineeringTextTabularElectronic Health Records

🎯 What it does: This paper proposes a backdoor attack method called SkillTrojan for skill-based agent systems, which embeds encrypted fragments into reusable skill implementations and reassembles and executes malicious payloads under trigger conditions.

Skip a Layer or Loop It? Learning Program-of-Layers in LLMs

Ziyue Li (University of Maryland), Tianyi Zhou (MBZUAI)

Computational EfficiencyAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningMixture of ExpertsText

🎯 What it does: The study dynamically skips or loops layers during LLM inference to construct input-specific program execution.

Skip-It? Theoretical Conditions for Layer Skipping in Vision–Language Models

Max Hartman (University of Illinois), Lav R. Varshney (Stony Brook University)

Explainability and InterpretabilityComputational EfficiencyTransformerVision Language ModelContrastive LearningMultimodality

🎯 What it does: Propose a theoretical framework that measures redundancy in intermediate layers and tokens of visual language models from four levels: geometry, proximity, functionality, and information redundancy, and derive conditions for safe layer skipping (early exit/late insertion) based on this; subsequently, verify which layers can be skipped across different tasks by calculating metrics such as inter-layer cosine distance and VAR; on models such as LLaVA 1.5/13B, LLaVA NeXT, DeepSeek-VL, and Qwen 2.5-VL, compare the drop in accuracy/F1 metrics before and after layer skipping using various datasets including VQA, Captioning, and MM-Reasoning, proving that the loss is negligible when redundancy conditions are met; and provide an algorithm that can locate layers that can be skipped using only a small amount of unlabeled data.

Skipping the Zeros in Diffusion Models for Sparse Data Generation

Phil Ostheimer, Sophie Fellenz (RPTU University Kaiserslautern-Landau)

GenerationData SynthesisComputational EfficiencyTransformerDiffusion modelScore-based ModelAuto EncoderImageTabularBiomedical DataPhysics Related

🎯 What it does: Proposed a sparse data generation diffusion model called SED, which can skip zero values during both training and inference, performing diffusion only on non-zero dimensions, thereby achieving accurate preservation of sparsity and improved computational efficiency.

SL-VC: A Benchmark and Automated Framework for Separation Logic Verification Condition Proving

Hanyang Wang (Shanghai Jiao Tong University), Qinxiang Cao (Shanghai Jiao Tong University)

AI Code AssistantTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: This paper constructs the SL-VC benchmark, proposes the SPLIT framework, and conducts experiments on this benchmark.

SlaClip: Gradient Norm Slacks can be Indicator for Adaptive Clipping in DP-SGD

Shuyan Zou (University of Southampton), Han Wu (University of Southampton)

OptimizationSafty and PrivacyImageText

🎯 What it does: Propose SlaClip, an adaptive clipping method that utilizes gradient 'slack' information in a single DP-SGD release, compatible with traditional DP-SGD;

SLAE: Strictly Local All-atom Environment for Protein Representation

Yilin Chen (Stanford University), Po-Ssu Huang (Stanford University)

Protein Structure PredictionGraph Neural NetworkTransformerSupervised Fine-TuningContrastive LearningGraphBiomedical Data

🎯 What it does: Proposes the SLAE framework, which encodes the local atomic environment of each residue using SE(3)-equivariant graph networks and reconstructs full-atom structures from these compact residue embeddings using a Transformer decoder, while jointly optimizing coordinate reconstruction, sequence recovery, and Rosetta energy regression during pre-training.

SLAP: The Semantic Least Action Principle for Variational Video-Language Modeling

Xiang Fang (Huazhong University of Science and Technology), Wanlong Fang (Nanyang Technological University)

OptimizationComputational EfficiencyRepresentation LearningTransformerVision Language ModelDiffusion modelScore-based ModelContrastive LearningVideoTextMultimodalityRetrieval-Augmented Generation

🎯 What it does: Proposed the Semantic Least Action Principle (SLAP), addressing the time blind spots caused by sparse frame sampling in video-language models by constructing physical dynamics on the semantic manifold;

SLASH the Sink: Sharpening Structural Attention Inside LLMs

Yiming Liu (Shanghai Jiao Tong University), Meng Jin (Shanghai Jiao Tong University)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationDrug DiscoveryTransformerLarge Language ModelTextGraphBenchmark

🎯 What it does: Proposed a training-free method called SLASH, which enhances the understanding of graph structures within large language models by redistributing attention internally.

SLAT: Segment-Level Adaptive Trimming for Efficient CoT Reasoning

Jian Yao (Hong Kong Polytechnic University), KC Tan

Computational EfficiencyTransformerReinforcement LearningTextBenchmarkChain-of-Thought

🎯 What it does: This paper proposes a reinforcement learning framework based on Segment-Level Adaptive Trimming (SLAT), aiming to significantly compress the reasoning length while maintaining the accuracy of chain-of-thought (CoT) reasoning.

SleepLM: Natural-Language Intelligence for Human Sleep

Zongzhe Xu (University of California Los Angeles), Yuzhe Yang (University of California Los Angeles)

GenerationRetrievalRepresentation LearningTransformerLarge Language ModelAuto EncoderContrastive LearningTextMultimodalityTime SeriesBiomedical DataRetrieval-Augmented Generation

🎯 What it does: Propose SleepLM, a sleep-language foundation model capable of describing, retrieving, and interacting with multimodal sleep data using natural language.

SleepMaMi: A Universal Sleep Foundation Model for Integrating Macro- and Micro-structures

Keondo Park (Seoul National University), Hyung-Sin Kim (Seoul National University)

ClassificationSegmentationAnomaly DetectionExplainability and InterpretabilityRepresentation LearningTransformerSupervised Fine-TuningMixture of ExpertsAuto EncoderContrastive LearningMultimodalityTime SeriesBiomedical DataElectrocardiogram

🎯 What it does: Built SleepMaMi, a dual-encoder sleep foundation model that pretrains on complete overnight multimodal PSG using both macro and micro structures, and evaluated on multiple downstream tasks including sleep staging, respiratory disorder segmentation, and disease prediction.

SlerpFlow: Spherical Trajectory Correction for Rectified Flow Inversion

Wenbin Duan (University of International Relations), Binyang Li (University of International Relations)

Image TranslationRestorationGenerationTransformerDiffusion modelScore-based ModelRectified FlowOptical FlowImageBenchmarkOrdinary Differential Equation

🎯 What it does: Propose a zero-training SlerpFlow algorithm to correct discretization errors that occur during the reverse process of rectified flow, thereby achieving higher fidelity image reconstruction and more precise semantic editing.

SliceFine: The Universal Winning-Slice Hypothesis for Pretrained Networks

Md Kowsher (Meta), Chen Chen (University of Central Florida)

Computational EfficiencyRepresentation LearningTransformerSupervised Fine-TuningContrastive LearningImageVideoTextMultimodality

🎯 What it does: Propose the Universal Winning Slice Hypothesis and design SliceFine, a parameter-efficient fine-tuning method that only updates structured row/column slices in the pre-trained network without adding new parameters.

SlideSparse: Fast and Flexible (2N-2):2N Structured Sparsity

Yingbo Hao (South China University of Technology), Furu Wei (Microsoft Research)

Computational EfficiencyTransformerLarge Language ModelText

🎯 What it does: Propose the SlideSparse system, which decomposes and maps the (2N-2):2N structured sparse model through a sliding window into the 2:4 sparse format, enabling acceleration of LLM inference on existing NVIDIA GPUs using Sparse Tensor Cores.

SLIM: Secure and Efficient Inference for Large Language Models on Untrusted Devices via TEEs

Wei Wang (National University of Defense Technology), Bao Li (National University of Defense Technology)

Safty and PrivacyComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: Propose SLIM, a framework for secure large language model inference using TEE

SLIP-RS: Structured-Attribute Language-Image Pre-Training for Remote Sensing Object Detection

Chenxu Wang (Nankai University), Qibin Hou (Nankai University)

Object DetectionConvolutional Neural NetworkTransformerSupervised Fine-TuningVision Language ModelContrastive LearningImageText

🎯 What it does: Propose SLIP-RS, which pretrains remote sensing images through structured attribute disentanglement, attribute contrastive learning, and confidence reliability engine to achieve fine-grained object detection.

SLQ: Bridging Modalities via Shared Latent Queries for Retrieval with Frozen MLLMs

Haoran Lou (Beijing University of Posts and Telecommunications), Xu Tang (Xiaohongshu Inc)

RetrievalDomain AdaptationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBenchmark

🎯 What it does: Propose the SLQ framework, which efficiently transfers the model into a retriever by adding a small number of shared latent queries (Shared Latent Queries) behind a frozen multimodal large language model, preserving pre-trained knowledge and inference capabilities.

SMAC: Score-Matched Actor-Critics for Robust Offline-to-Online Transfer

Nathan Samuel de Lara (University of Toronto), Florian Shkurti (University of Toronto)

Reinforcement LearningDiffusion modelScore-based Model

🎯 What it does: Propose the SMAC method, which regularizes the gradient of the Q-function in offline RL through score-matching, making the actor-critic trained offline easier to fine-tune online.

Small Agent Group is the Future of Digital Health

Yuqiao Meng (State University of New York at Binghamton), Zhaohan Xi (State University of New York at Binghamton)

OptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIMixture of ExpertsTextBiomedical DataElectronic Health RecordsBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose a Small Agent Group (SAG) framework, which replaces a single large model with collaborative small models to enhance the efficacy and reliability of digital health decision support.

Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models

Yun Qu (Tsinghua University), Xiangyang Ji (Tsinghua University)

OptimizationComputational EfficiencyTransformerLarge Language ModelReinforcement LearningPrompt EngineeringContrastive LearningText

🎯 What it does: Proposes an online prompt selection method called Generalizable Predictive Prompt Selection (GPS), which significantly improves training and inference efficiency when training large reasoning models after reinforcement learning.

Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO

Yiming Ren (Tsinghua University), Ruihang Chu (Tsinghua University)

Computational EfficiencyKnowledge DistillationTransformerLarge Language ModelReinforcement LearningMixture of ExpertsContrastive LearningText

🎯 What it does: This paper proposes using a small model as an explorer to train a large model, thereby improving the performance of GRPO on mathematical reasoning tasks and reducing computational costs.

SMART: Scalable Mesh‑free Aerodynamic Simulations from Raw Geometries using a Transformer‑based Surrogate Model

Jan Hagnberger (University of Stuttgart), Mathias Niepert (University of Stuttgart)

OptimizationComputational EfficiencyTransformerAuto EncoderContrastive LearningPoint CloudMeshPhysics Related

🎯 What it does: Proposes SMART, a grid-free aerodynamic simulation surrogate model based on Transformer, capable of predicting physical quantities using only point cloud geometry and query points, completely independent of simulation grids.

SmartThinker: Progressive Chain-of-Thought Length Calibration for Efficient Large Language Model Reasoning

Chenzhi Hu (Shanghai Jiao Tong University), Guihai Chen (Shanghai Jiao Tong University)

Computational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextChain-of-Thought

🎯 What it does: Propose SmartThinker, an efficient inference method based on GRPO, which compresses the reasoning trajectory of large language models by dynamically estimating the optimal length of chain-of-thought and introducing dynamic length rewards.

SMD: Multi-view Safety-Critical Driving Video Generation in the Real-world Domain

Jiawei Zhou (Harbin Institute of Technology), Yu Li (Zhejiang University)

GenerationData SynthesisAutonomous DrivingTransformerSupervised Fine-TuningVision Language ModelDiffusion modelVideoMultimodality

🎯 What it does: Proposed the SMD framework, which can generate multi-perspective safety-critical driving videos in the real-world domain, and evaluate and train end-to-end driving systems using the generated safety-critical scenarios.

SMILE: Extended Deep Submodular Function-Based Instruction and In-context Learning Demonstration Selection

Zihan Chen (University of Virginia), Cong Shen (University of Virginia)

Recommendation SystemOptimizationComputational EfficiencyData-Centric LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextRetrieval-Augmented Generation

🎯 What it does: Propose the SMILE method, which jointly selects task instructions and example sets to optimize prompts for large language models, directly generating the best instruction-example pairs at the query level.

SMM Transformer: Leveraging Spiking Neural Networks for Multimodal Tasks

Xiubo Liang (Zhejiang University), Hongzhi Wang (Zhejiang University)

Computational EfficiencyRepresentation LearningSpiking Neural NetworkTransformerMixture of ExpertsContrastive LearningImageTextMultimodality

🎯 What it does: A multi-modal Transformer framework based on Spiking Neural Networks (SNN), named SMMTransformer, is proposed to address the issues of unstable training in deep spiking networks and the mismatch between dense softmax attention and spiking-based communication.

Smooth Dynamic Cutoffs for Machine Learning Interatomic Potentials

Kevin Han (Carnegie Mellon University), Amir Barati Farimani

Computational EfficiencyGraph Neural NetworkAuto EncoderContrastive LearningGraphTabularBenchmarkPhysics Related

🎯 What it does: Proposed a dynamic cutoff method for adaptively selecting the number of neighbors in machine learning interatomic potentials (MLIP), achieving graph sparsification, significantly reducing memory usage and inference time.

Smooth Multi-Policy Causal Effect Estimation in Longitudinal Settings

Wenxin Chen (Cornell University), Fei Wang (Cornell University)

Drug DiscoveryRecurrent Neural NetworkTransformerReinforcement LearningContrastive LearningTabularTime SeriesBiomedical DataElectronic Health Records

🎯 What it does: Jointly estimate the causal effects of multiple dynamic treatment strategies in longitudinal clinical data, addressing variance inflation and information waste caused by traditional separate estimation.

Smoothie: Smoothing Diffusion on Token Embeddings for Text Generation

Alexander Shabalin (Constructor University), Dmitry Vetrov (Constructor University)

GenerationTransformerDiffusion modelScore-based ModelText

🎯 What it does: Proposed a new text diffusion model called SMOOTHIE, which generates text by performing smooth diffusion in the word embedding space.

Smoothing Slot Attention Iterations and Recurrences

Rongzhen Zhao (Aalto University), Joni Pajarinen (Aalto University)

Object DetectionObject TrackingSegmentationTransformerContrastive LearningImageVideo

🎯 What it does: Propose SmoothSA, which addresses the cold start and homogeneity issues of frame-to-frame transformation in the first frame of images/videos by introducing preheated queries with different iteration counts across frames, thereby improving the quality of object-centric learning.

Smoothness Errors in Dynamics Models and How to Avoid Them

Edward Berman (Northeastern University), Robin Walters (Northeastern University)

OptimizationGraph Neural NetworkTransformerDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningMeshGraphTime SeriesPhysics Related

🎯 What it does: The study investigates and proposes relaxed unit convolutions (Taylor truncation and composite smooth destruction) to address over-smoothing and under-smoothing issues in graph neural networks for dynamic models, and extends the approach to grid structures.

SmoothSpike: Spiking Transformer with Learnable Hadamard Transformation

Zijian Zhou (University of Electronic Science and Technology of China), Haizhou Li (Shenzhen Loop Area Institute)

ClassificationComputational EfficiencyRepresentation LearningSpiking Neural NetworkTransformerContrastive LearningText

🎯 What it does: To address the issue of information homogenization caused by neuron saturation in Spiking Neural Networks (SNN) for language modeling, the SmoothSpike framework is proposed, which applies Hadamard transform to the activation before spiking neurons to smooth the input, thereby reducing the saturation rate and enhancing the expressive power.

Sobolev Regularized Score Difference Estimation in Diffusion Models

Chenghan Xie (Stanford University), Renyuan Xu (Stanford University)

GenerationData SynthesisDomain AdaptationSupervised Fine-TuningDiffusion modelScore-based ModelContrastive LearningImageBiomedical DataElectrocardiogramAudio

🎯 What it does: Proposed a Sobolev-regularized fractional difference estimation method for transfer learning and post-training in diffusion models.

Social Hippocampus Memory Learning

Liping Yi (Tianjin University), Qinghua Hu (Tianjin University)

ClassificationFederated LearningKnowledge DistillationContrastive LearningImage

🎯 What it does: Proposed a memory-based social machine learning framework called SoHip, which enables collaborative learning of heterogeneous models by exchanging only long-term memory in a federated learning environment.

SoftBinary Coding: A New Information-Theoretic Paradigm for Neural Compression via Fast Channel Simulation

Ezgi Ozyilkan (New York University), Jona Ballé (New York University)

CompressionScore-based ModelFlow-based ModelAuto EncoderContrastive LearningTabularTime SeriesSequentialStochastic Differential Equation

🎯 What it does: This paper proposes SoftBinary Coding (SBC), an end-to-end neural compression method based on random binary latent space and polarization channel simulation.

SoftJAX & SoftTorch: Empowering Automatic Differentiation Libraries with Informative Gradients

Anselm Paulus (University of Tübingen), Georg Martius (University of Tübingen)

OptimizationComputational EfficiencyReinforcement Learning from Human FeedbackDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningBenchmark

🎯 What it does: Proposed two open-source libraries, SoftJAX and SoftTorch, which uniformly implement differentiable alternatives to various hard discrete operations, covering element-wise and axis-wise operators.

SoftMatcha 2: A Fast and Soft Pattern Matcher for Trillion-Scale Corpora

Masataka Yoneda (University of Tokyo), Sho Yokoi (National Institute for Japanese Language and Linguistics)

Computational EfficiencyData-Centric LearningText

🎯 What it does: Research and implement SoftMatcha 2, a high-speed soft pattern matching algorithm for tera-scale corpora, supporting semantic variations such as word substitution, insertion, and deletion;

Softmax as Linear Attention in the Large-Prompt Regime: a Measure-based Perspective

Etienne Boursier (Université Paris-Saclay), Claire Boyer (Université Paris-Saclay)

OptimizationExplainability and InterpretabilityRepresentation LearningTransformerPrompt EngineeringTabularReview/Survey Paper

🎯 What it does: Study the behavior of softmax attention in the case of long prompts (large-prompt), and construct a measure-based framework, proving that it degenerates into linear attention in the infinite-prompt limit, thereby providing non-asymptotic convergence bounds and comparing training dynamics with linear attention;

SoftMoE: Soft Differentiable Routing for Mixture-of-Experts in LLMs

Mikołaj Zasada (AGH University of Krakow), Marcin Kurdziel (AGH University of Krakow)

Computational EfficiencyRepresentation LearningTransformerLarge Language ModelMixture of ExpertsContrastive LearningText

🎯 What it does: Propose SoftMoE, a sparse Mixture-of-Experts routing mechanism that uses LapSum soft topk approximation; achieve adaptive allocation of expert capacity through differentiable soft routing; and reduce the number of activated experts while maintaining autoregressive compatibility.

Softplus Attention with Re-weighting Boosts Length Extrapolation in Large Language Models

Bo Gao (Nanyang Normal University), Letizia Gionfrida (King's College London)

RetrievalComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: Proposes a two-stage attention mechanism called LSSAR, which first normalizes using Softplus and a length scaling factor (LSSA), and then reweights through Shift-ReLU and a power transformation, addressing the numerical instability of Softmax and its insufficient generalization for long sequences.

Softsign: Smooth Sign in Your Optimizer For Better Parameter Heterogeneity Handling

Dmitrii Feoktistov (HSE University), Aleksandr Beznosikov (MIRAI)

OptimizationFederated LearningRepresentation LearningReinforcement Learning from Human FeedbackRecurrent Neural NetworkGraph Neural NetworkTransformerLarge Language ModelMixture of ExpertsDiffusion modelScore-based ModelContrastive LearningTextGraphTabularSequential

🎯 What it does: Propose two smoothed sign-based and spectral LMO optimizers, SoftSignum and SoftMuon, and provide an adaptive decay scheme with a tunable temperature, achieving parameter-level transition from fixed amplitude to gradient-size-based updating.

SOLAR for Offline MARL: Plateau-Triggered Potential Shaping under World-Model Uncertainty

Jusheng Zhang (Sun Yat-sen University), Keze Wang (Sun Yat-sen University)

TransformerReinforcement LearningWorld ModelTabularSequential

🎯 What it does: Propose the SOLAR framework, which utilizes a world model for low-cost roll-out in offline multi-agent reinforcement learning. It detects whether learning has entered a steady state before activating reward shaping, employing potential-based reward shaping combined with an uncertainty threshold for suppression.

SOLAR: Self-supervised Joint Learning for Symmetric Multimodal Retrieval

Wenjie Yang (Ant Group), Peng Di (Ant Group)

RetrievalKnowledge DistillationRepresentation LearningTransformerLarge Language ModelVision Language ModelAuto EncoderContrastive LearningImageTextMultimodalityBenchmark

🎯 What it does: Designed and implemented a self-supervised two-stage framework called SOLAR for symmetric multi-modal to multi-modal retrieval tasks.

Solving Imperfect-Recall Games via Sum-of-Squares Optimization

Rui Zheng (Singapore University of Technology and Design), Antonios Varvitsiotis (Singapore University of Technology and Design)

OptimizationReinforcement Learning

🎯 What it does: Propose a framework based on the Sum-of-Squares (SOS) hierarchy to solve for optimal behavioral strategies and Nash equilibria in incomplete memory generalized form games (IREFG).

Solving Inverse Problems with Flow-based Models via Model Predictive Control

George Webber (King's College London), Andrew J. Reader

RestorationGenerationOptimizationDiffusion modelFlow-based ModelImageBiomedical DataComputed TomographyReview/Survey Paper

🎯 What it does: Proposes the MPC-Flow framework, which utilizes model predictive control (MPC) during inference by solving subproblems of a pre-trained flow model, enabling untrained conditional generation and solving inverse problems.

Solving Positive Linear Programs with Differential Privacy

Alina Ene (Boston University), Adrian Vladu (Université Paris Cité)

OptimizationSafty and Privacy

🎯 What it does: The study investigates approximate algorithms for solving positive linear programming (Packing, Covering, and mixed Packing-Covering) under a differential privacy framework, and proposes a private solver that can maintain privacy while violating only a controllable number of constraints.

Solving Spatial-Spectral Fusion with Latent Spectral Operators

Wei Li (Zhejiang University of Technology), Jianwei Zheng (Zhejiang University of Technology)

Super ResolutionTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningImage

🎯 What it does: A spatial-spectral fusion framework based on a latent space is designed, which compresses high-dimensional hyperspectral images into latent representations using cross-attention projection, and performs multi-scale fusion through a hierarchical patch structure.

Solving Stochastic Variational Inequalities without the Bounded Variance Assumption

Ahmet Alacaoglu (University of British Columbia), JunHyun Kim

OptimizationDiffusion modelScore-based ModelContrastive LearningStochastic Differential Equation

🎯 What it does: This paper proposes three stochastic variational inequality (SVI) solving algorithms without the assumption of bounded variance, achieving optimal complexity under unconstrained domains, non-monotonic (weak Minty VI) conditions, and constraints.

Solving the Offline and Online Min-Max Problem of Non-smooth Submodular-Concave Functions: A Zeroth-Order Approach

Amir Ali Farzin (Australian National University), Iman Shames (University of Melbourne)

OptimizationImage

🎯 What it does: This paper studies the max-min and min-max problems involving submodular functions that may be non-smooth with respect to the minimizer and concave functions with respect to the maximizer, and proposes a solution based on zeroth-order methods.

Solving Time-Dependent Differential Equations with Physical Dynamical Systems

Chuan Liu (Rice University), Tony Geng (Rice University)

Computational EfficiencySpiking Neural NetworkTime SeriesBenchmarkPhysics RelatedStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Proposed a time-dependent differential equation solver based on physical dynamic systems (DSMs), named DS-TS, which can efficiently and accurately provide dynamic process trajectories in continuous time.

SoMA: A Real-to-Sim Neural Simulator for Robotic Soft-Body Manipulation

Mu Huang (Fudan University), Jiangmiao Pang (Shanghai Artificial Intelligence Laboratory)

Robotic IntelligenceGraph Neural NetworkDiffusion modelGaussian SplattingOptical FlowImageVideo

🎯 What it does: Developed SoMA, a real-time neural simulator based on Gaussian splatting, which can learn and achieve realistic-to-simulation mapping of soft objects in robotic manipulation under multi-view RGB video.

Sonar-TS: Search-Then-Verify Natural Language Querying for Time Series Databases

Zhao Tan (Jiangxi University of Finance and Economics), Ming Jin (Griffith University)

Recommendation SystemAnomaly DetectionOptimizationTransformerLarge Language ModelPrompt EngineeringTextTime SeriesBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose the Sonar-TS framework, which combines LLM planning with SQL+Python code to implement a 'search-then-verify' workflow for natural language queries on time series databases, and construct the NLQTSBench long-historical query benchmark.

SONAR: Spectral‑Contrastive Audio Residuals for Generalizable Deepfake Detection

Ido Nitzan Hidekel (Tel Aviv University), Dan Raviv (Tel Aviv University)

Anomaly DetectionTransformerAuto EncoderContrastive LearningAudio

🎯 What it does: Propose the SONAR framework, which processes low-frequency content and high-frequency residual branches in parallel, and explicitly aligns low-frequency semantics with high-frequency residuals in the latent space using Jensen-Shannon alignment loss to detect deepfake audio.

SonicMaster: Towards Controllable All-in-One Music Restoration and Mastering

Jan Melechovsky (Singapore University of Technology and Design), Dorien Herremans (Singapore University of Technology and Design)

RestorationGenerationTransformerPrompt EngineeringDiffusion modelRectified FlowAuto EncoderTextMultimodalityAudio

🎯 What it does: Proposes SonicMaster, a unified text-driven generative model for simultaneously repairing 19 common distortions in music (such as reverb, distortion, attenuation, dynamic compression, stereo imbalance, etc.), and supports users to achieve controllable repair through natural language instructions or automatic global repair;

SOPE: Situation-Aware and Statistically Indistinguishable Privacy Exfiltration for MCP-enabled Agents

Ruixiao Lin (Zhejiang University), Shouling Ji (Zhejiang University)

Safty and PrivacyAdversarial AttackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: This paper proposes the SOPE framework, which automatically transforms ordinary MCP servers into privacy leakage tools, achieving zero-click, no-user-interaction LLM agent-based personal information theft.

SORA: Free Second-Order Attacks in Fast Adversarial Training

Mazdak Teymourian (Sharif University of Technology), Mohammad Hossein Rohban (Sharif University of Technology)

OptimizationComputational EfficiencyAdversarial AttackSupervised Fine-TuningContrastive LearningImageBiomedical Data

🎯 What it does: Proposed an adaptive step-size single-step adversarial training method called SORA, and introduced two theoretical tools, Epsilon Overfitting and PertAlign, to predict and avoid catastrophic overfitting;

SorryDB: Can AI Provers Complete Real-World Lean Theorems?

Austin Letson (Axiomatic AI), Lenny Taelman (University of Amsterdam)

AI Code AssistantTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmark

🎯 What it does: Proposes a dynamically updated Lean task benchmark SorryDB, aiming to evaluate the practicality of AI provers in real formalization projects.

SOTAlign: Semi-Supervised Alignment of Unimodal Vision and Language Models via Optimal Transport

Simon Roschmann (Helmholtz Munich), Zeynep Akata (Helmholtz Munich)

RetrievalDomain AdaptationRepresentation LearningTransformerVision Language ModelContrastive LearningImageTextMultimodality

🎯 What it does: We construct a semi-supervised alignment framework called SOTAlign on pre-trained unimodal vision and language encoders, using only a small number of image-text paired samples and a large amount of unpaired image/text data. It first performs rough alignment via a linear teacher, and then refines cross-modal embeddings using KLOT divergence based on Optimal Transport.

Source-Free Open-World RF Fingerprint Identification

Kunling Li (Shanghai Jiao Tong University), Pengwenlong Gu (Conservatoire national des arts et metiers)

ClassificationRecognitionAnomaly DetectionDiffusion modelScore-based ModelNeural Radiance FieldAuto EncoderGenerative Adversarial NetworkContrastive LearningOptical FlowAudio

🎯 What it does: Proposed a source-unavailable open-world wireless spectrum fingerprint recognition framework that addresses the stability-plasticity conflict under mixed streams of new and old classes.

SP-Mind: An Autonomous Reasoning Agent for Spatial Proteomics Analysis

YuCheng Yuan, Ruijiang Li (Stanford University)

Explainability and InterpretabilityComputational EfficiencyDrug DiscoveryTransformerLarge Language ModelAgentic AIPrompt EngineeringImageBiomedical DataBenchmarkChain-of-Thought

🎯 What it does: Propose SP-Mind, an AI agent capable of autonomous reasoning, which can automatically complete the entire spatial proteomics analysis workflow, including image preprocessing, registration, cell segmentation, quantification, clustering, etc., starting from raw multi-tissue imaging data, and provide interpretable results.

SPA: A Simple but Tough-to-Beat Baseline for Knowledge Injection

Kexian Tang (Tsinghua University), Kaifeng Lyu (Tsinghua University)

Data SynthesisExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: By utilizing seven manually designed prompt templates based on principles of cognitive learning, LLMs are repeatedly used to rewrite a small-scale professional corpus, generating a large-scale synthetic corpus. Continued pre-training on this synthetic corpus enables knowledge injection.

SpaCeFormer: Fast Proposal-Free Open-Vocabulary 3D Instance Segmentation

Chris Choy (NVIDIA), Jan Kautz (NVIDIA)

SegmentationTransformerVision Language ModelContrastive LearningTextPoint Cloud

🎯 What it does: Propose SpaCeFormer — a fast, box-free open-vocabulary 3D instance segmentation network, and construct the SpaCeFormer-3M dataset containing 604K instances and 3M multi-view consistent captions.

SpaceVista: All-Scale Visual Spatial Reasoning from mm to km

Peiwen Sun (Chinese University of Hong Kong), Xiangyu Yue (Chinese University of Hong Kong)

Data SynthesisRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsVision Language ModelContrastive LearningImageVideoTextMultimodalityBenchmark

🎯 What it does: Propose SpaceVista: a visual spatial reasoning dataset (SpaceVista-1M) and benchmark (SpaceVista-Bench) covering scales from millimeters to kilometers, and train a multi-scale expert spatial reasoning model, SpaceVista-7B, based on this.

SPADA: A Verifiable Test-Driven Agent for Controllable Parametric CAD Assembly Generation

Keyou Zheng (Guangdong University of Technology), Jiewu Leng (Guangdong University of Technology)

GenerationData SynthesisAI Code AssistantTransformerLarge Language ModelAgentic AIPrompt EngineeringTextMultimodalityMeshBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposes SPADA, a test-driven self-inspecting parametric assembly generation agent that generates editable and constraint-satisfying CAD code in a sandbox through an iterative compile-test-repair loop.

SpaEF: Spatially Resolved Transcriptomics Data Element-Wise Denoising Framework Powered by Large Models

Zekuan Shang (Jilin University), You Zhou (Jilin University)

RestorationGraph Neural NetworkTransformerLarge Language ModelAuto EncoderContrastive LearningGraphBiomedical Data

🎯 What it does: Propose and implement SpaEF, a framework that constructs spot and gene graphs using large models and achieves denoising of spatial transcriptomics data through element-wise graph autoencoders.

SpanNorm: Reconciling Training Stability and Performance in Deep Transformers

Wang Chao (Meituan Inc), Xunliang Cai (Meituan Inc)

OptimizationComputational EfficiencyRepresentation LearningTransformerLarge Language ModelMixture of ExpertsText

🎯 What it does: Propose the SpanNorm mechanism, which introduces residual connections across the entire block within the Transformer block, combining the stability of PreNorm with the performance advantages of PostNorm.

SPAR: Support-Preserving Action Rectification

Jiaxin Zhao (Zhejiang University), Binbin Lin (Zhejiang University)

OptimizationExplainability and InterpretabilityComputational EfficiencyRobotic IntelligenceRecurrent Neural NetworkTransformerReinforcement LearningDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningMultimodalityTabularTime SeriesSequentialBenchmark

🎯 What it does: Designed and implemented a residual reinforcement learning framework SPAR based on a behavioral cloning benchmark, used for local correction of actions with support preservation in offline reinforcement learning.

SPARC: Separating Perception And Reasoning Circuits for Test-time Scaling of VLMs

Niccolò Avogaro (ETH Zürich), Mattia Rigotti (IBM Research)

Explainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBenchmark

🎯 What it does: Proposed the SPARC framework, which separates the two major functions of visual perception and reasoning into two independent processing stages, forming a two-stage reasoning process: first, locating problem-related regions through implicit relevance detection (IRD), and then performing reasoning on these high-resolution cropped regions.

SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance–Diversity Data Selection

Shuhao Chen (Southern University of Science and Technology), Yu Zhang (Southern University of Science and Technology)

Safty and PrivacyTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringContrastive LearningText

🎯 What it does: The study addresses the risk of large language models losing safety alignment during fine-tuning, and proposes a novel defense framework called SPARD. It enhances model safety when facing harmful fine-tuning attacks by combining safety projection alternating optimization (SPAG) with relevance-diversity data selection.

SPARe: Stacked Parallelism with Adaptive Reordering for Fault-Tolerant LLM Pretraining Systems with 100k+ GPUs

Jin Lee (University of California, Santa Barbara), Zheng Zhang (University of California, Santa Barbara)

OptimizationFederated LearningComputational EfficiencyTransformerLarge Language ModelText

🎯 What it does: Propose the SPARe framework, which achieves fault tolerance in large LLM pretraining by utilizing stacked parallelism and adaptive reordering, avoiding frequent global restarts.

SPARKLING: Balancing Signal Preservation and Symmetry Breaking for Width-Progressive Learning

Qifan Yu (State Key Laboratory of General Artificial Intelligence, Peking University), Di He (State Key Laboratory of General Artificial Intelligence, Peking University)

OptimizationComputational EfficiencyRepresentation LearningTransformerLarge Language ModelMixture of ExpertsContrastive LearningText

🎯 What it does: This paper proposes the SPARKLING framework, which reduces computational cost during the mid-phase of pre-training while maintaining model performance through width expansion;

Sparks of Cooperative Reasoning: LLMs as Strategic Hanabi Agents

Mahesh Ramesh (University of Wisconsin-Madison), Emmanouil-Vasileios Vlatakis-Gkaragkounis

Recommendation SystemOptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningAdversarial AttackHyperparameter SearchData-Centric LearningRobotic IntelligenceMeta LearningDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackNeural Architecture SearchProtein Structure PredictionTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringMixture of ExpertsWorld ModelTextSequentialBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This paper conducts a large-scale, transparent, and reproducible evaluation of the performance of 17 top large language models (LLMs) in the Hanabi cooperative card game, investigates the impact of context engineering (Watson, Sherlock, Mycroft) on cooperative reasoning, and releases two new datasets, HanabiLogs and HanabiRewards; subsequently, these datasets are used to perform instruction fine-tuning and reinforcement learning fine-tuning on the 4B open-source model Qwen3-Instruct, significantly narrowing the performance gap with the strongest closed-source reasoning models, and demonstrating its transferability on other cooperative and reasoning benchmark tasks.

Sparse ActionGen: Accelerating Diffusion Policy with Real-time Pruning

Kangye Ji (Tsinghua University), Zhi Wang (Tsinghua University)

Computational EfficiencyRobotic IntelligenceTransformerPrompt EngineeringDiffusion modelImageVideoSequentialRetrieval-Augmented Generation

🎯 What it does: Propose Sparse ActionGen (SAG), which accelerates Diffusion Policy through real-time sparse pruning and global cache reuse, achieving extremely high inference speed.

Sparse and Faithful Local Explanations with Piecewise Linear Surrogates

Yixin Wang (Sichuan University), Yucheng Dong (Sichuan University)

Explainability and InterpretabilityTabular

🎯 What it does: Proposes the PL-LIME framework, which uses instance-anchored piecewise linear proxies to provide local interpretability for black-box models;

Sparse Autoencoders are Topic Models

Leander Girrbach (Technical University of Munich), Zeynep Akata (Technical University of Munich)

Explainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelAuto EncoderContrastive LearningImageTextMultimodality

🎯 What it does: Theoretical formulation of sparse autoencoders (SAE) as a continuous topic model (CTM), and the proposal of the SAE-TM framework: first pretrain SAE on large-scale text or image embeddings, then interpret SAE features as word distributions using a posterior word emission matrix, and finally merge them into interpretable topics through clustering, thus achieving cross-modal topic discovery and analysis.

Sparse Autoencoders for Interpretable Emotion Control in Text-to-Speech

Hongfei Du (William & Mary), Ashley Gao

GenerationExplainability and InterpretabilityTransformerLarge Language ModelAuto EncoderContrastive LearningTextAudio

🎯 What it does: In text-to-speech systems based on LLMs, sparse autoencoders (SAEs) are introduced to analyze the semantic hidden layers, revealing that emotional information is distributed across a few sparse features. By increasing or decreasing these features, bidirectional induction and suppression of emotions are achieved.

Sparse Bayesian Deep Functional Learning with Structured Region Selection

Xiaoxian Zhu (Shanghai University of Finance and Economics), Mengyun Wu (Shanghai University of Finance and Economics)

Explainability and InterpretabilityComputational EfficiencyData-Centric LearningDiffusion modelScore-based ModelContrastive LearningTabularTime SeriesElectrocardiogramAudio

🎯 What it does: Proposed a sparse Bayesian deep functional neural network (sBayFDNN) for nonlinear scalar-to-function regression and automatic identification of functional subregions.

Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders

David Chanin (University College London), Adrià Garriga-Alonso (MATSResearch)

Explainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelAuto EncoderContrastive LearningTextSequential

🎯 What it does: Investigates how improper L0 settings (the average number of activated latent dimensions per input) in sparse autoencoders (SAE) can lead to feature mixing, thereby compromising uniqueness, and proposes a proxy metric based on the decoder's diagonal cosine similarity (c_dec) to determine the correct L0.

Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs

Yukun Jiang (CISPA Helmholtz Center for Information Security), Yang Zhang (CISPA Helmholtz Center for Information Security)

Safty and PrivacyAdversarial AttackTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsTextRetrieval-Augmented Generation

🎯 What it does: Study the impact of routing on safety in Mixture-of-Experts LLMs, introduce the concept of unsafe routes, and demonstrate how unsafe answers can be induced by modifying routing.

Sparse Regression with $\ell_0$ Constraints for $\alpha$-Mixing Time Series: Algorithms and Guarantees

Ruoxin Yuan (Fudan University), Lijun Ding (University of California San Diego)

OptimizationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTime Series

🎯 What it does: In time series data, the ℓ0-constrained least squares problem for sparse linear regression on α-mixing Gaussian processes is solved, and high probability RSC/RSS properties are provided along with sample and iteration complexity analysis for exact sparse algorithms such as IHT, CoSaMP, and Subspace Pursuit; subsequently, the theory is applied to sparse VAR models, and the method is validated on synthetic VAR data and real New York City ride-sharing time series.

Sparse Relaxed-Lasso Steering: Automatic Sparse Autoencoder Feature Selection for Precise Image Editing

Zongxin Liu (Key Laboratory of System Software Chinese Academy of Sciences), Lijun Zhang (Key Laboratory of System Software Chinese Academy of Sciences)

Image TranslationGenerationOptimizationRepresentation LearningTransformerDiffusion modelAuto EncoderContrastive LearningImageText

🎯 What it does: Proposes Sparse Relaxed-Lasso Steering (SRLS), which automatically selects features in the sparse autoencoder space through sparse optimization and uses Bayesian optimization to determine the editing intensity, achieving precise image editing that is independent of training.

Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization

Zheng Fang (Wuhan University), Zhijin Ge (Xidian University)

OptimizationAdversarial AttackTransformerLarge Language ModelPrompt EngineeringTextAudio

🎯 What it does: Proposed and verified a sparse token optimization method (TAGO) for audio-language models, enabling jailbreak attacks on the model by updating only a small number of audio tokens with high gradient energy;

Sparse Topology-Aware Pairwise Scoring for Large-Scale Multi-Agent Reinforcement Learning

Zhibo Deng (Shenzhen MSU-BIT University), Xiping Hu (Shenzhen MSU-BIT University)

Graph Neural NetworkTransformerReinforcement LearningContrastive LearningGraphBenchmark

🎯 What it does: Proposed an scalable sparse communication mechanism called SOPS, which achieves dynamic sparse connections in large-scale multi-agent reinforcement learning through an exponential graph backbone and learnable subgraphs.

SparseInfer: Accelerating Large Language Model Inference with Semantics-Inspired Adaptive Sparse Activation

Qinsi Wang (Duke University), Yiran Chen (Duke University)

Computational EfficiencyTransformerLarge Language ModelText

🎯 What it does: Propose SparseInfer, a sparse inference framework based on sentence-level core neurons, which eliminates the MLP prediction of activation layers and maintains the core neurons unchanged during inference.

SparseOpt: Addressing Normalization-induced Gradient Skew in Sparse Training

Mohammed Adnan (University of Calgary), Yani Ioannou (University of Calgary)

ClassificationOptimizationComputational EfficiencyConvolutional Neural NetworkImage

🎯 What it does: Studied the negative impact of batch normalization (BN) on gradient direction in dynamic sparse training (DST), and proposed a sparse-aware preconditioned optimizer, SparseOpt, to correct gradient imbalance, improving the convergence speed and generalization performance of sparse training.

Sparser Block-Sparse Attention via Token Permutation

Xinghao Wang (Fudan University), Xipeng Qiu (Fudan University)

Computational EfficiencyTransformerLarge Language ModelText

🎯 What it does: Proposed a pluggable Permuted Block-Sparse Attention (PBS-Attn) mechanism, which improves block-level sparsity by segmenting and rearranging key tokens, thereby significantly reducing computational and memory costs during the forward pass for long texts.

Sparser, Faster, Lighter Transformer Language Models

Edoardo Cetin (Sakana AI), Llion Jones (Sakana AI)

Computational EfficiencyRepresentation LearningTransformerLarge Language ModelMixture of ExpertsText

🎯 What it does: Introduce sparse representations in the feed-forward layers of large language models, and propose the TwELL sparse format and its CUDA kernel that can be efficiently executed on NVIDIA GPUs;

SparseSSM: Efficient Selective Structured State Space Models Can Be Pruned in One-Shot

Kaiwen Tuo (Westlake University), Huan Wang (Westlake University)

Computational EfficiencyKnowledge DistillationTransformerLarge Language ModelText

🎯 What it does: Proposed a no-training one-shot sparsification framework called SparseSSM for the Mamba model, which prunes time-shared SSM parameters by extending the OBS method.

SPATIA: Multimodal Generation and Prediction of Spatial Cell Phenotypes

Zhenglun Kong (Harvard Medical School), Marinka Zitnik (Harvard Medical School)

GenerationData SynthesisTransformerDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageMultimodalityBiomedical Data

🎯 What it does: This work proposes the SPATIA model, which integrates cell morphology images, transcriptomic expression, and spatial position information to achieve multi-modal generation and prediction of cell phenotypes.

Spatial Conformal Inference through Localized Quantile Regression

Hanyang Jiang (Georgia Institute of Technology), Yao Xie (Georgia Institute of Technology)

TabularAgriculture RelatedFinance Related

🎯 What it does: Proposed an LSCP framework based on local quantile regression for achieving distribution-free prediction intervals on non-exchangeable spatial data.

Spatial Deconfounder: Interference-Aware Deconfounding for Spatial Causal Inference

Ayush Khot (University of Illinois at Urbana-Champaign), Xihaier Luo (Brookhaven National Laboratory)

Convolutional Neural NetworkGraph Neural NetworkAuto EncoderTabularBenchmark

🎯 What it does: Studied the problem of causal inference in spatial data while simultaneously addressing interference and unobserved spatial confounding.

Spatial Memory for Out-of-Vision Manipulation in Vision-Language-Action

Pengteng Li (Hong Kong University of Science and Technology), Hui Xiong (Hong Kong University of Science and Technology)

Robotic IntelligenceTransformerVision-Language-Action ModelDiffusion modelAuto EncoderContrastive LearningImageTextMultimodalityRetrieval-Augmented Generation

🎯 What it does: Proposes the SOMA framework, providing persistent spatial memory for Vision-Language-Action models, enabling reasoning and manipulation when the target is not in the current field of view;

Spatial Priors via Space Filling Curves for Small and Limited Data Vision Transformers

Leyla Naz Candogan (Ecole Polytechnique Federale De Lausanne), Volkan Cevher (Ecole Polytechnique Federale De Lausanne)

ClassificationObject DetectionSegmentationComputational EfficiencyTransformerAuto EncoderContrastive LearningImage

🎯 What it does: Introduce a lightweight masked attention mechanism called VIOLIN in Vision Transformer, which enhances the model's spatial prior by attenuating attention through spatial filling curves (SFC).