arXivSub Start free trial

ICML 2026 Papers — Page 55

International Conference on Machine Learning · 6554 papers

Spatial-Aware Reduction Framework: Towards Efficient and Faithful Visual State Space Models

Jindi Lv (Sichuan University), Jiancheng Lv (Sichuan University)

ClassificationObject DetectionSegmentationComputational EfficiencyTransformerAuto EncoderContrastive LearningImage

🎯 What it does: Propose STORM, a spatially aware token reduction framework that can maintain the 2D grid structure without training;

SpatialJB: How Text Distribution Art Becomes The "Jailbreak Key" for LLM Guardrails

Zhiyi Mou (Zhejiang University), Kui Ren (Zhejiang University)

Safty and PrivacyAdversarial AttackTransformerLarge Language ModelPrompt EngineeringText

🎯 What it does: This paper proposes SpatialJB, a jailbreak method that disrupts the semantic continuity of Transformers by rearranging text into a two-dimensional spatial layout;

Spatially-Adaptive Gradient Re-parameterization for 3D Large Kernel Optimization

Ho Hin Lee (Vanderbilt University), Bennett Allan Landman

SegmentationOptimizationComputational EfficiencyConvolutional Neural NetworkBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose a 3D large convolutional network called Rep3D based on spatial adaptive gradient reparameterization, which addresses the instability in training large convolutions and improves the performance of 3D medical image segmentation.

Spatially-Regularized Entropy for Discriminative Token Merging in Fine-Grained Re-Identification

Shangze Li (Beihang University), Yingbo Qu (Beihang University)

RecognitionRetrievalComputational EfficiencyTransformerAuto EncoderContrastive LearningImage

🎯 What it does: Propose a training-free, spatial entropy-based token merging framework (SRE-Merge), specifically designed for fine-grained person re-identification tasks, significantly reducing the computational cost of ViT.

SpatialReward: Bridging the Perception Gap in Online RL for Image Editing via Explicit Spatial Reasoning

Yancheng Long (Harbin Institute of Technology), Shuo Yang (Harbin Institute of Technology)

Image TranslationGenerationTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringVision Language ModelContrastive LearningImageMultimodalityBenchmark

🎯 What it does: Propose SpatialReward, a method that enhances online RL image editing reward models through explicit spatial reasoning, addressing the attention collapse problem in existing evaluators.

Spatio-Temporal LLM: Reasoning about Environments and Actions

Haozhen Zheng (University of Illinois Urbana-Champaign), Alex Schwing

Autonomous DrivingRepresentation LearningRobotic IntelligenceTransformerLarge Language ModelVision Language ModelVision-Language-Action ModelContrastive LearningVideoTextMultimodalityPoint Cloud

🎯 What it does: Constructed a QA dataset named REA that simultaneously includes global point clouds and local perspective videos, and designed two LLM models, STLLM-3D and STLLM-Aligner, that integrate spatial and temporal information.

SpatioLM: Towards General Physical Spatial Intelligence in Vision-Language Models

jing wu, Long Chen (Xiaomi)

Autonomous DrivingKnowledge DistillationRobotic IntelligenceTransformerVision Language ModelDiffusion modelContrastive LearningImageTextPoint CloudBenchmark

🎯 What it does: Propose the SpatioLM framework, which integrates the Spatio-Vision module in a plug-and-play manner on top of a frozen VLM, achieving spatial reasoning without requiring additional 3D priors.

Spatiotemporal Imputation with Graph-Informed Flow Matching

Zepeng Zhang (EPFL), Olga Fink (EPFL)

RestorationComputational EfficiencyGraph Neural NetworkTransformerDiffusion modelFlow-based ModelAuto EncoderGraphTime Series

🎯 What it does: Propose a spatiotemporal missing value imputation framework called GiFlow based on graph information flow matching

SPEAR: A Unified SSL Framework for Learning Speech and Audio Representations

Xiaoyu Yang (University of Cambridge), Phil Woodland (University of Cambridge)

Knowledge DistillationRepresentation LearningTransformerAuto EncoderContrastive LearningMultimodalityAudio

🎯 What it does: Integrate speech and general audio representations through self-supervised learning, proposing the SPEAR framework, which utilizes multi-codebook vector quantization to generate fine-grained discrete tokens and performs masked prediction, constructing a unified acoustic feature encoder.

SpecExit: Accelerating Large Reasoning Model via Speculative Exit

Rubing Yang (Tencent), Peng Chen (Tencent)

Computational EfficiencyAI Code AssistantTransformerLarge Language ModelPrompt EngineeringTextChain-of-Thought

🎯 What it does: Propose the SpecExit framework, which uses the hidden states of the LRM to predict future tokens and dynamically early stops, thereby reducing unnecessary inference processes.

SpecForge: A Flexible and Efficient Open-Source Training Framework for Speculative Decoding

Shenggui Li (Nanyang Technological University), Tianwei Zhang (Nanyang Technological University)

Computational EfficiencyKnowledge DistillationTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsContrastive LearningTextRetrieval-Augmented Generation

🎯 What it does: This paper proposes the SpecForge framework and the SpecBundle toolkit, aiming to efficiently train and deploy speculative decoding draft models.

SpecMD: A Comprehensive Study On Speculative Expert Prefetching

Duc N.M Hoang (Apple), Minsik Cho (Apple)

OptimizationComputational EfficiencyMixture of ExpertsTextBenchmark

🎯 What it does: Developed the SpecMD benchmark and systematically evaluated MoE caching strategies, proposing the Least-Stale eviction strategy and verifying its significant performance improvement.

SpecPL: Disentangling Spectral Granularity for Prompt Learning

Jingtao Zhou (City University of Hong Kong), Lai Man Po

ClassificationRecognitionDomain AdaptationRepresentation LearningAdversarial AttackTransformerPrompt EngineeringVision Language ModelDiffusion modelAuto EncoderContrastive LearningImageTextMultimodalityBenchmark

🎯 What it does: By decomposing low-frequency semantics and high-frequency details in the frozen VAE latent space, and introducing adversarial detail supervision during training, improving VLM prompt learning to eliminate modality asymmetry;

SpecPrune-VLA: Accelerating Vision-Language-Action Models via Action-Aware Self-Speculative Pruning

Hanzhen Wang (Shanghai Jiao Tong University), Guohao Dai (Shanghai Jiao Tong University)

Computational EfficiencyRobotic IntelligenceTransformerLarge Language ModelVision-Language-Action ModelContrastive LearningImageVideoTextMultimodality

🎯 What it does: A training-agnostic acceleration framework called SpecPrune-VLA is proposed for Vision-Language-Action models, significantly reducing the number of visual tokens through self-speculative pruning.

Spectra: Rethinking Optimizers for LLMs Under Spectral Anisotropy

Zhendong Huang (Fudan University), Li Shang (Fudan University)

OptimizationTransformerLarge Language ModelText

🎯 What it does: This paper investigates the spectral heterogeneity of gradient signals in large language model (LLM) training, finding that gradient energy is concentrated in a few low-rank peak directions, while the long-tail directions are sparse in information and easily suppressed. Based on this, the paper proposes the Spectra optimizer, which significantly improves long-tail learning and convergence speed by suppressing the dominant peak subspace without amplifying the noise-sensitive tail.

Spectral Bridge Variational Inference: Dynamic LoRA via Bures-Wasserstein Gradient Flows

Yuhang Xi (Guangzhou University), Zhao-Rong Lai (Guangzhou University)

OptimizationComputational EfficiencyRepresentation LearningAI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningTextStochastic Differential Equation

🎯 What it does: Propose a dynamic LoRA framework SBVI based on Bures-Wasserstein gradient flow, allowing low-rank adapters to adaptively evolve with task gradients during training.

Spectral Collapse Drives Loss of Plasticity in Deep Continual Learning

Arjun Prakash (Brown University), George Konidaris (Brown University)

OptimizationComputational EfficiencyRepresentation LearningMeta LearningReinforcement LearningAuto EncoderGenerative Adversarial NetworkContrastive LearningImageTextMultimodalityBenchmark

🎯 What it does: Investigated the fundamental cause of plasticity loss in deep continual learning and proposed a method to maintain network plasticity by preventing Hessian spectral collapse.

Spectral Evolution Search: Efficient Inference-Time Scaling for Reward-Aligned Image Generation

Jinyan Ye (East China Normal University), Yingda Chen (Alibaba Group)

GenerationComputational EfficiencyReinforcement Learning from Human FeedbackTransformerDiffusion modelFlow-based ModelImageTextOrdinary Differential Equation

🎯 What it does: Propose Spectral Evolution Search (SES), a framework that aligns image generation models with rewards during inference by optimizing low-frequency noise.

Spectral Flow Matching: Stabilizing Stochastic GFlowNets via Frequency-Domain Regularization

Nadhir Hassen (Adelaide University), Johan W. Verjans (Adelaide University)

OptimizationMeta LearningReinforcement Learning from Human FeedbackReinforcement LearningMixture of ExpertsDiffusion modelScore-based ModelFlow-based ModelContrastive LearningTabularTime SeriesSequentialBiomedical DataStochastic Differential Equation

🎯 What it does: Propose Spectral Time-Dependent GFlowNets (ST-GFN), which shifts the GFlowNet training objective to the frequency domain. It suppresses high variance and improves stability using spectral consistency loss and low-pass filter regularization, and achieves structured exploration by constructing an autocorrelation intrinsic reward based on the Wiener-Khinchin theory.

Spectral Gradient Descent Mitigates Anisotropy-Driven Misalignment: A Case Study in Phase Retrieval

Guillaume Braun (RIKEN AIP), Masaaki Imaizumi (University of Tokyo)

OptimizationContrastive LearningImageTabularReview/Survey PaperPhysics RelatedStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Studied the training dynamics of Spectral Gradient Descent (SpecGD) in the phase retrieval model with high-dimensional Gaussian inputs and covariance with sharp peaks, and compared it with traditional gradient descent (GD), revealing how SpecGD suppresses bias driven by variance.

Spectral Guidance for Flexible and Efficient Control of Diffusion Models

Gabriel Moreira (Institute for Systems and Robotics, Instituto Superior Técnico), Chenyan Xiong (Carnegie Mellon University)

GenerationData SynthesisComputational EfficiencyTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningImage

🎯 What it does: Propose the Spectral Guidance framework, achieving flexible control without training by learning low-dimensional spectral coordinates of the diffusion process;

Spectral Heat Flow for Conservative Token Condensation in Vision-Language Models

Zhaoyang Li (University of Science and Technology of China), Tianzhu Zhang (University of Science and Technology of China)

CompressionComputational EfficiencyRepresentation LearningTransformerVision Language ModelFlow-based ModelContrastive LearningOptical FlowImageVideoTextMultimodality

🎯 What it does: Proposed a training-agnostic, conservative visual token compression framework called SpecFlow, which can significantly reduce the number of visual tokens during VLM inference.

Spectral Imbalance Causes Forgetting in Low-Rank Continual Adaptation

Hao Gu (Southeast University), Tong Wei (Southeast University)

OptimizationComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerVision Language ModelMultimodalityBenchmark

🎯 What it does: Proposes the EBLoRA method, which reduces forgetting in continual learning by balancing energy distribution in low-rank adaptation.

Spectral Reach: Understanding Neural Scaling as Progress into the Spectral Tail

Konstantin Nikolaou (University of Stuttgart), Christian Holm (University of Stuttgart)

ClassificationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelAuto EncoderContrastive LearningImageText

🎯 What it does: Proposes an scalable LNP decomposition that utilizes spectral position indicators to monitor spectral learning dynamics during the training of large models.

Spectral-Informed Neural Networks Outperform Spectral methods in High-dimensional PDEs

Tianchi Yu, Ivan Oseledets

OptimizationConvolutional Neural NetworkRecurrent Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningTabularTime SeriesSequentialPhysics RelatedStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Propose and implement an improved Spectral-Informed Neural Networks (Modified SINNs) for solving high-dimensional partial differential equations (PDEs), and can predict missing spectral coefficients in medium- and high-dimensional problems, significantly improving the accuracy of solutions.

Spectral-Progressive Thought Flow for Lightweight Multimodal Reasoning

Yixian Shen (University of Amsterdam), Anuj Pathania (University of Amsterdam)

Computational EfficiencyRepresentation LearningTransformerVision Language ModelDiffusion modelFlow-based ModelImageTextMultimodalityChain-of-ThoughtOrdinary Differential Equation

🎯 What it does: Propose Spectral-Progressive Thought Flow (SpecFlow), achieving lightweight multi-modal spatial reasoning by low-frequency compression of visual thoughts in the discrete cosine domain and recursive inference via flow matching.

Spectrally-Guided Diffusion Noise Schedules

Carlos Esteves (Google Research), Ameesh Makadia (Google Research)

GenerationDiffusion modelScore-based ModelAuto EncoderContrastive LearningImageText

🎯 What it does: Designed and implemented an adaptive noise scheduling based on the spectral features of each image, and validated its effectiveness in a single-stage pixel-level diffusion model.

Speculative Coupled Decoding for Training-Free Lossless Acceleration of Autoregressive Visual Generation

Junhyuk So (POSTECH), Eunhyeok Park (POSTECH)

GenerationComputational EfficiencyTransformerDiffusion modelImageVideo

🎯 What it does: Propose Speculative Coupled Decoding (SCD), a training-free, lossless acceleration method for autoregressive visual generation models.

Speculative Safety Honeypot: Toward Proactive Defense Against Multi-turn Agent Attacks

Zezhong WANG, Heqing Huang (Huawei Technologies Co., Ltd)

Safty and PrivacyReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmark

🎯 What it does: Propose the Speculative Safety Honeypot framework, which utilizes multi-agent asynchronous simulation to predict the future behavior of LLM agents in advance, thereby achieving proactive defense in multi-round attacks.

Speculative Sampling For Faster Molecular Dynamics

Arthur Kosmala (Meta), Brandon M. Wood (Meta)

Drug DiscoveryScore-based ModelContrastive LearningGaussian SplattingTabularTime SeriesBiomedical DataPhysics RelatedStochastic Differential Equation

🎯 What it does: This paper proposes Langevin Speculative Dynamics (LSD), a distributed explicit sampling method that utilizes a fast draft model to generate steps and parallel verification at the backend, specifically designed to accelerate molecular dynamics simulations based on machine learning potentials.

Speech-Audio Compositional Attacks on Multimodal LLMs and Their Defense with SALMONN-Guard

Yudong Yang (Tsinghua University), Chao Zhang (Tsinghua University)

Safty and PrivacyAdversarial AttackTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringDiffusion modelAuto EncoderTextMultimodalityBenchmarkRetrieval-Augmented GenerationAudio

🎯 What it does: This paper constructs the SACRED-Bench benchmark, proposes three speech-audio combination attacks (speech overlap, speech-non-speech mixture, multi-speaker conversation), and designs corresponding queries; subsequently, it develops the SALMONN-Guard security model, which can simultaneously analyze speech, audio, and text, achieving multi-modal security protection.

SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding

Talor Abramovich (NVIDIA), Yonatan Geifman (NVIDIA)

Computational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose SPEED-Bench, a unified and diverse benchmark for evaluating the real-world performance of Speculative Decoding (SD) acceleration techniques across different semantic domains, input lengths, and concurrency levels.

SPEED: Sharpened-Teacher Distillation for Parallel Decoding of Diffusion Language Models

Qiuhong Shen (National University of Singapore), Xinchao Wang (National University of Singapore)

Computational EfficiencyKnowledge DistillationAI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringDiffusion modelContrastive LearningText

🎯 What it does: Propose a framework called SPEED that combines teacher reinforcement distillation and Jensen-Shannon divergence for accelerating parallel decoding in diffusion language models.

SpeedCP: Fast Kernel-based Conditional Conformal Prediction

Yating Liu (University of Chicago), Claire Donnat (University of Chicago)

Anomaly DetectionFederated LearningComputational EfficiencyRepresentation LearningDrug DiscoveryDiffusion modelScore-based ModelAuto EncoderContrastive LearningImageTextTabularBiomedical DataMagnetic Resonance ImagingBenchmark

🎯 What it does: Propose SpeedCP, an efficient and adjustable method for constructing prediction intervals through kernel RKHS conditional conformal prediction.

Speedup Patch: Learning a Plug-and-Play Policy to Accelerate Embodied Manipulation

Zhichao Wu (Nanjing University), Yang Yu (Nanjing University)

Computational EfficiencyRobotic IntelligenceRecurrent Neural NetworkTransformerReinforcement LearningWorld ModelTabularTime SeriesSequential

🎯 What it does: Proposes the Speedup Patch (SuP) framework, which accelerates multiple body posture learning strategies by adaptively downsampling action blocks in offline data through an external scheduler, without retraining the base policy.

SpeedVFI: One-step Diffusion for Efficient Video Frame Interpolation

Ganggui Ding (Zhejiang University), Chunhua Shen (Zhejiang University)

GenerationComputational EfficiencyKnowledge DistillationTransformerDiffusion modelScore-based ModelAuto EncoderVideo

🎯 What it does: Propose SpeedVFI, a one-step diffusion model framework that can perform multi-frame video interpolation in a single inference step.

SPHERE: Mitigating the Loss of Spectral Plasticity in Mixture-of-Experts for Deep Reinforcement Learning

Lirui Luo (Peking University), Qing Li (State Key Laboratory of General Artificial Intelligence)

Reinforcement LearningMixture of ExpertsTabularTime Series

🎯 What it does: This paper proposes the SPHERE regularization method to alleviate spectral plasticity loss in Mixture-of-Experts during continuous reinforcement learning.

Spherical Procrustes Alignment for Reliable Medical Audio Diagnosis

Ying Wang (Macao Polytechnic University), Xiaochen Yuan (Macao Polytechnic University)

ClassificationAnomaly DetectionKnowledge DistillationData-Centric LearningContrastive LearningBiomedical DataAudio

🎯 What it does: Propose a Spherical Procrustes Alignment (SPA) method, which improves the accuracy and reliability of medical audio diagnosis through spherical constraints and dynamic ETF alignment.

Spherical SO(3) Equivariant Local Attention

Yusuke Sekikawa (DENSO IT Lab), Ruka Eto (DENSO IT Lab)

ClassificationSegmentationPose EstimationDepth EstimationComputational EfficiencyTransformerContrastive LearningImageMeshPhysics Related

🎯 What it does: Proposed a SO(3) equivariant local attention mechanism named S oLA without position encoding, which can achieve full three-dimensional rotational equivariance on spherical signals.

Spherical Steering: Geometry-Aware Activation Rotation for Language Models

Zejia You (Tufts University), Hanjie Chen (Rice University)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelContrastive LearningText

🎯 What it does: Propose a geometry-based method for activation rotation during inference called Spherical Steering, which utilizes contrastive samples to construct directional prototypes and achieves control over the behavior of language models by performing norm-preserving rotations on the spherical surface in the hidden layer;

SphericalDreamer: Generating Navigable Immersive 3D Worlds with Panorama Fusion

Antoine Schnepf (Criteo AI Lab), Andrew I. Comport (Université Côte d'Azur)

Image HarmonizationGenerationData SynthesisDepth EstimationTransformerPrompt EngineeringDiffusion modelNeural Radiance FieldGaussian SplattingImageTextPoint CloudMesh

🎯 What it does: Generate a 3D world that allows long-distance navigation and full immersion through text prompts.

Spik4lite: Refactoring Neuromorphic Sparsity for Efficient Spiking Neural Networks on Commodity Edge Devices

Yongzhi She (Shenzhen University), Jingcai Guo (Hong Kong Polytechnic University)

ClassificationImage TranslationRestorationComputational EfficiencySpiking Neural NetworkTransformerContrastive LearningImage

🎯 What it does: Proposed a lightweight plug-in called Spik4lite, which dynamically sparsifies neural channels during training through an energy-aware gating mechanism, ultimately pruning inefficient channels at the channel level;

Spike Camera Autofocus via Frequency-Domain Spectral-Centroid Migration

Xijie Xiang (Peking University), Yonghong Tian (Peking University)

OptimizationComputational EfficiencyOptical FlowImageVideo

🎯 What it does: By analyzing the energy migration phenomenon in the time-frequency domain when the spiking camera focuses, an automatic focusing method called CEN based on the frequency domain spectral centroid is proposed.

Spike-HTR: Spiking Neural Transformer for Handwritten Text Recognition

Xiubo Liang (Zhejiang University), Hongzhi Wang (Zhejiang University)

RecognitionSpiking Neural NetworkTransformerDiffusion modelContrastive LearningImage

🎯 What it does: A pulse Transformer named Spike-HTR is designed, which achieves efficient hand-written text line recognition by utilizing two budgets: short time slot numbers and width sequence length.

SpikeCLR: Self-Supervised Contrastive Learning for Visual Representations with Spiking Neural Networks

Chengwei Zhou (Case Western Reserve University), Gourav Datta (Case Western Reserve University)

Computational EfficiencyRepresentation LearningSpiking Neural NetworkTransformerContrastive LearningImage

🎯 What it does: This paper proposes a fully self-supervised contrastive learning framework called SpikeCLR, specifically designed to train Spiking Neural Networks (SNNs) for sparse event-driven representations in visual tasks.

Spiked-CFR: Causal Representation Learning from LLMs via Wasserstein Projection Pursuit

Fan Wang (Zhejiang University), Shuiguang Deng (Zhejiang University)

OptimizationExplainability and InterpretabilityRepresentation LearningData-Centric LearningTransformerLarge Language ModelGenerative Adversarial NetworkContrastive LearningTextTabularElectronic Health RecordsBenchmark

🎯 What it does: Propose the SPIKED-CFR framework, which achieves causal representation learning by Wasserstein projection tracking in the frozen LLM embedding space, for estimating treatment effects from text.

SpikeNet: Sparse Spike-Driven Mask Vector Transformer for Energy-Efficient and Stable Spiking Point Cloud Processing

Zhiming Zhou (Anhui University), Ajmal Saeed Mian (University Of Western Australia)

ClassificationSegmentationSpiking Neural NetworkTransformerSupervised Fine-TuningContrastive LearningPoint Cloud

🎯 What it does: Propose SpikeNet, a point cloud classification and segmentation framework based on spiking neural networks;

SpikeVLA: Vision-Language-Action Models with Spiking Neural Networks

Ruiqi Song (Tongji University), Long Chen (Chinese Academy of Sciences)

Computational EfficiencyRobotic IntelligenceSpiking Neural NetworkTransformerReinforcement LearningVision-Language-Action ModelDiffusion modelContrastive LearningImageTextMultimodality

🎯 What it does: Developed an end-to-end Spiking Vision-Language-Action model called SpikeVLA, which utilizes event-driven SNNs to achieve visual encoding, multi-modal language modeling, and action strategies, significantly reducing energy consumption and computational costs.

SpikingLM: Towards Fully Spiking Language Model

Yu Liang (University of Electronic Science and Technology of China), Haizhou Li (Shenzhen Loop Area Institute)

Computational EfficiencyRepresentation LearningSpiking Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningText

🎯 What it does: Built a fully spiking pulse language model called SpikingLM, which addresses the issues of 'dead neurons' and insufficient attention competition in deep spiking neural networks.

Spiral RoPE: Rotate Your Rotary Positional Embeddings in the 2D Plane

Haoyu Liu (University of California Berkeley), Feng Wang (John Hopkins University)

ClassificationSegmentationGenerationTransformerDiffusion modelContrastive LearningImage

🎯 What it does: Proposes Spiral RoPE, a technique that extends rotational position encoding to two-dimensional planes, aiming to eliminate the directional bias of traditional axial 2D RoPE.

SplAttN: Bridging 2D and 3D with Gaussian Soft Splatting and Attention for Point Cloud Completion

Zhaoyang Li (Southwest Jiaotong University), Tianrui Li (Southwest Jiaotong University)

Data SynthesisAutonomous DrivingRepresentation LearningGraph Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningGaussian SplattingImageMultimodalityPoint Cloud

🎯 What it does: Proposed a multi-modal point cloud completion framework called SplAttN based on differentiable Gaussian soft splatting and attention mechanisms, which can effectively bridge cross-modal information between 2D images and 3D point clouds;

Split Group Knockoffs: Controlling False Discovery Rate in Transformational Group Sparsity

Siqi Chen (University of Pennsylvania), Xinwei Sun (Fudan University)

Anomaly DetectionOptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyRepresentation LearningDrug DiscoveryTabularBiomedical DataAlzheimer's DiseaseBenchmark

🎯 What it does: Proposed a Split Group Knockoffs (SGK) method for group variable selection under a group sparse structure after linear transformation.

Split Personality Training: Revealing Latent Knowledge Through Alternate Personalities

Florian Dietz (Saarland University), Dietrich Klakow (Saarland University)

Explainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningText

🎯 What it does: Propose Split Personality Training (SPT), achieving model internal 'honest personality' auditing through LoRA fine-tuning;

SPLIT-VLM: Salience-Guided Partitioning towards Local Coverage for Importance-Aware Token Dropping in Vision-Language Models

Seungil Lee (Seoul National University of Science and Technology), Hyun Kim (Seoul National University of Science and Technology)

Computational EfficiencyRepresentation LearningTransformerPrompt EngineeringVision Language ModelContrastive LearningImageVideoTextMultimodality

🎯 What it does: Propose a visual–language model efficient token reduction framework called SPLIT, which utilizes time offset estimation for salience, local budget allocation, and diversity scores for token selection.

Sponge Tool Attack: Stealthy Denial-of-Efficiency against Tool-Augmented Agentic Reasoning

Qi Li (National University of Singapore), Xinchao Wang (National University of Singapore)

Federated LearningSafty and PrivacyExplainability and InterpretabilityComputational EfficiencyAdversarial AttackTransformerLarge Language ModelPrompt EngineeringVision Language ModelTextMultimodalityRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Designed and implemented a malicious prompt rewriting framework called Sponge Tool Attack (STA), which, without modifying the model or tools, induces tool-enhanced large language model (LLM) agents to generate redundant and inefficient reasoning trajectories by subtly rewriting input prompts, thereby significantly increasing computational costs while maintaining task semantics.

SPR-RAFT: Parameter-Efficient Regression-Aware Fine-Tuning for Biomedical LLM Regression

Yuanlin Yang (City University of Hong Kong), Haodong Liu (Nanjing University)

Drug DiscoveryTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringBiomedical DataBenchmark

🎯 What it does: For medical regression tasks, the SPR-RAFT framework is proposed, which efficiently transforms frozen large language models into numerical regressors through soft prompts and [REG] regression heads.

SPR: A Structured Prompt Refinement Network for Modality Missing

Hao Chen (Beijing Institute of Technology), Xia Wu (Beijing Institute of Technology)

ClassificationConvolutional Neural NetworkRecurrent Neural NetworkTransformerPrompt EngineeringContrastive LearningImageTextMultimodality

🎯 What it does: A structured prompt refinement network (SPR) is proposed, which significantly improves classification performance in multimodal missing scenarios by structurally refining the prompt vectors on the frozen CLIP model in global, local, and channel dimensions.

SpreadsheetArena: Decomposing Preference in LLM Generation of Spreadsheet Workbooks

Srivatsa R Kundurthy, John Ling (Longitude Labs Inc.)

GenerationData SynthesisRecommendation SystemExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringTabularBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposed the SPREADSHEETARENA platform to evaluate the performance of LLMs in generating complete spreadsheet workbooks, and made the benchmark data and preference voting publicly available.

SPUR: Scale-Partitioned Uncertainty Rectification for Robust UAV-on-UAV Interception

Chenqi Yan (Shanghai Jiao Tong University), Nanyang Ye (Shanghai Jiao Tong University)

Object DetectionAutonomous DrivingComputational EfficiencyRobotic IntelligenceConvolutional Neural NetworkContrastive LearningOptical FlowImageVideoStochastic Differential Equation

🎯 What it does: A scalable and real-time detection framework named SPUR is proposed for UAV-to-UAV interception tasks, addressing issues such as scale drift, scale imbalance, and robustness degradation caused by flight noise.

Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training

Christian Moya (Purdue University), Elliott Thornley (Massachusetts Institute of Technology)

OptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningContrastive LearningTextBenchmark

🎯 What it does: Studied the false correlation learning mechanism caused by surface-level correlations in training data within preference optimization (e.g., DPO), its irreversible impact on model error during deployment, and proposed a theoretical and experimental solution to alleviate this issue by adding equivalent-value preference pairs (Tie Training).

Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs

Lecheng Yan (Southern University of Science and Technology), Chenyang Lyu (Alibaba Group)

Explainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelText

🎯 What it does: By providing mechanistic explanations of LLMs under RLVR (Reinforcement Learning Verification Reward), we identified and located the 'Anchor-Adapter' circuit, elucidating why random or incorrect rewards can activate the model's memory shortcut and lead to improved reasoning performance;

Spurious Rewards: Rethinking Training Signals in RLVR

Rulin Shao (University of Washington), Luke Zettlemoyer (University of Washington)

OptimizationExplainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmark

🎯 What it does: Studied the phenomenon where reinforcement learning with verifiable rewards (RLVR) can significantly improve the mathematical reasoning performance of large language models even when using no information or incorrect reward signals, and explained the underlying mechanisms.

Squeezing More from the Stream : Learning Representation Online for Streaming Reinforcement Learning

Nilaksh (Chandar Research Lab), Sarath Chandar (Chandar Research Lab)

OptimizationRepresentation LearningReinforcement LearningContrastive LearningImageVideo

🎯 What it does: Introduce self-supervised representation learning (SPR) into the stream reinforcement learning framework, maximizing information extraction through single updates per frame transition.

SRPO: Self-Reflective Policy Optimization for Long-Horizon Reasoning

Jialong Liu (Wuhan University), Zuchao Li (Wuhan University)

OptimizationKnowledge DistillationReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringAuto EncoderTextChain-of-Thought

🎯 What it does: Propose Self-Reflective Policy Optimization (SRPO), which enables large language models (LLMs) to generate concise 'reflection patches' after completing a full inference, converting sparse terminal rewards into dense token-level supervision, and achieving post-training long-sequence reasoning capabilities through self-teacher self-distillation.

SS-TPT: Stability and Suitability-Guided Test-Time Prompt Tuning for Adversarially Robust Vision-Language Models

Sunoh Kim (Dankook University), Daeho Um (University of Seoul)

ClassificationDomain AdaptationComputational EfficiencyRepresentation LearningAdversarial AttackTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodality

🎯 What it does: Propose a test-time prompt tuning method (SS-TPT) based on view stability and applicability scores, enhancing the robustness of CLIP under adversarial attacks and distribution shifts.

SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space

Zhenyi Shen (King's College London), Xing Sun (Tencent)

Computational EfficiencyRepresentation LearningTransformerLarge Language ModelContrastive LearningTextBenchmark

🎯 What it does: Studied a training framework named SSA, which integrates sparse attention with full attention, and enhances performance under both sparse and full inference modes through bidirectional attention output alignment.

SSDCN: Spatial-Spectral Dual-Clustering-based Network for Hyperspectral Image Super-resolution

Yong Yang (Tiangong University), Hangyuan Lu (Jinhua University of Vocational Technology)

Super ResolutionConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningImage

🎯 What it does: Proposed a spatial-spectral dual clustering network (SSDCN) for single image super-resolution of hyperspectral images.

SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning

Xiaojun Guo (Peking University), Yisen Wang (Peking University)

Representation LearningTransformerReinforcement LearningVision Language ModelAuto EncoderContrastive LearningImageTextMultimodality

🎯 What it does: Propose a framework called SSL4RL, which transforms self-supervised learning tasks into verifiable reinforcement learning rewards to enhance the reasoning and visual understanding of vision-language models.

SSR-Merge: Subspace Signal Routing for Training-Free LoRA Merging in Diffusion Models

Zhengxuan Wei (Nanjing University), Qi Fan (Nanjing University)

GenerationData SynthesisTransformerDiffusion modelImage

🎯 What it does: Designed an untrained LoRA merging framework called SSR-Merge, which can merge multiple task LoRAs into a single model without retraining, and eliminate parameter conflicts through subspace signal routing.

ST-TGExplainer: Disentangling Stability and Transition Patterns for Temporal GNN Interpretability

Hongjiang Chen (Hangzhou Dianzi University), Shirui Pan (Griffith University)

Explainability and InterpretabilityGraph Neural NetworkTransformerDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningGraphTime SeriesSequential

🎯 What it does: Propose ST-TGEXPLAINER, a self-explaining temporal graph neural network, which generates more trustworthy explanations and improves prediction performance by decoupling stable patterns and transition patterns.

ST-Veto: Spatio-Temporal Token Veto for Diffusion MLLMs via Taylor Prediction and Visual Grounding

Keuntae Kim (Hanyang University), Yong Suk Choi (Hanyang University)

GenerationExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelVision Language ModelDiffusion modelImageTextMultimodalityChain-of-Thought

🎯 What it does: This paper proposes a training-agnostic decoding strategy called ST-Veto, which improves reasoning by leveraging temporal stability and visual alignment signals within diffusion-based multimodal large language models.

Stability Analysis of Sharpness-Aware Minimization

Hoki Kim (Chung-Ang University), Jaewook Lee (Seoul National University)

OptimizationContrastive LearningImageStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: This paper investigates the convergence instability of Sharpness-Aware Minimization (SAM) near saddle points, proving that under certain conditions, saddle points can become attractors of SAM, and that the diffusion speed of SAM escaping saddle points in a random noise environment is slower than that of standard gradient descent.

Stability and Generalization of Nonconvex Optimization with Heavy-Tailed Noise

Hongxu Chen (Fudan University), Luo Luo (Fudan University)

Optimization

🎯 What it does: This paper proposes deriving upper bounds on generalization error for non-convex optimization under heavy-tailed noise by leveraging algorithm stability;

Stability beyond Bounded Differences: Sharp Generalization Bounds under Finite $L_p$ Moments

Qianqian Lei (University of Chicago), Wei Biao Wu (University of Chicago)

OptimizationFederated LearningExplainability and InterpretabilityRepresentation LearningMeta LearningReinforcement LearningContrastive LearningTabularTime SeriesBenchmark

🎯 What it does: Proposes an algorithm stability framework under the limited Lp moment condition, and provides corresponding concentration inequalities and high-probability generalization bounds; meanwhile, applies this framework to three major learning paradigms: empirical risk minimization, transductive regression, and meta-learning;

Stability-Aware Feature Design for Robust Watermark Detection in Machine-Generated Text

Sina Mansouri (George Mason University), Abolfazl Safikhani (George Mason University)

ClassificationAnomaly DetectionTransformerLarge Language ModelSupervised Fine-TuningText

🎯 What it does: This paper proposes a watermark detection method called PSS based on stability awareness, which can robustly distinguish between human-generated and model-generated texts even under multiple rounds of rewriting and short text conditions.

Stabilizing Equation Learning via Zero-Point Constraints

Sannyuya Liu (Central China Normal University), Jianwen Sun (Central China Normal University)

OptimizationExplainability and InterpretabilityComputational EfficiencyNeural Architecture SearchTransformerTabularTime SeriesSequentialPhysics Related

🎯 What it does: Propose an improved neuro-symbolic regression framework EQL-Z, which utilizes zero-point constraints and adaptive structural search to enhance the stability and interpretability of equation learning.

Stabilizing In-Context Multi-Source Domain Adaptation for Biomedical Images Through Controls

Ana Sanchez-Fernandez (Johannes Kepler University), Günter Klambauer (Johannes Kepler University)

ClassificationDomain AdaptationMeta LearningDrug DiscoveryImageBiomedical Data

🎯 What it does: This study proposes the CS-ARM-BN method, which stabilizes batch effects in multi-source domain adaptation by incorporating negative control samples into BN statistics during meta-learning and testing;

Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers

Wenhan Ma (Peking University), Zhifang Sui (Peking University)

TransformerReinforcement LearningPrompt EngineeringMixture of ExpertsTextChain-of-Thought

🎯 What it does: This paper proposes a mechanism called Rollout Routing Replay (R3), which records the MoE routing distribution during the inference phase and replays it during training to achieve consistency between training and inference routing, significantly improving the stability of RL training and preventing model collapse.

Stabilizing Native Low-Rank LLM Pretraining

Paul Janson (Concordia University), Eugene Belilovsky (Concordia University)

OptimizationComputational EfficiencyRepresentation LearningTransformerLarge Language ModelContrastive LearningText

🎯 What it does: Train large language models from scratch using low-rank decomposition, completely eliminating full-rank auxiliary weights, achieving end-to-end low-rank pre-training;

Stabilizing PPO via Latent-Space Regularization and KDE-Driven Exploration

Meiyu Du (Tongji University), Wei Wang (Tongji University)

Reinforcement LearningContrastive LearningImageTabular

🎯 What it does: Propose SPPO, which introduces three latent-space regularization methods (CKA alignment, no-flip constraint, KDE exploration shaping) on top of PPO to improve training stability and performance in continuous control tasks.

Stabilizing Recurrent Dynamics for Test-Time Scalable Latent Reasoning in Looped Language Models

Xiao-Wen Yang (Nanjing University), Yu-Feng Li (Nanjing University)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningRecurrent Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringText

🎯 What it does: Investigate and address the issue of unstable scalability of LoopLM during testing, achieving reliable deep reasoning through the STARS training framework.

Stabilizing Reinforcement Learning for Diffusion Language Models

Jianyuan Zhong, Qiang Xu (Chinese University of Hong Kong)

TransformerReinforcement LearningDiffusion modelScore-based ModelText

🎯 What it does: Implement full-parameter reinforcement learning training on discrete diffusion large language models (dLLM) by proposing the StableDRL framework;

Stabilizing the Q-Gradient Field for Policy Smoothness in Actor-Critic Methods

Jeong Woon Lee (Kyung Hee University), Hyoseok Hwang (Kyung Hee University)

Reinforcement Learning

🎯 What it does: The PAVE framework is proposed to stabilize the Q gradient field and achieve smooth policy in Actor-Critic methods by regularizing the geometric properties of the Critic.

Stable Asynchrony: Variance-Controlled Off-Policy RL for LLMs

Luke J. Huang (Massachusetts Institute Of Technology), Song Han (Massachusetts Institute Of Technology)

TransformerLarge Language ModelReinforcement LearningText

🎯 What it does: This paper proposes the VCPO method, aiming to address the issues of variance explosion and training collapse caused by policy lag in asynchronous reinforcement learning for large-scale LLMs.

Stable Deep Reinforcement Learning via Isotropic Gaussian Representations

Ali Saheb Pasand (Mila Quebec AI Institute), Pablo Samuel Castro (Mila Quebec AI Institute)

Reinforcement LearningContrastive LearningImageVideoTabularTime Series

🎯 What it does: In non-stationary deep reinforcement learning environments, we propose actively approximating the feature distribution to an isotropic Gaussian distribution through 'Sketched Isotropic Gaussian Regularization (SIGReg)' to improve learning stability and performance.

Stable Localized Conformal Prediction via Transduction

Yinjie Min (Nankai University), Changliang Zou (University of Melbourne)

ClassificationDomain AdaptationAnomaly DetectionTabularBiomedical Data

🎯 What it does: Proposed a stabilization method called Stable Conformal Prediction (StCP), which utilizes source task labels and unlabelled target data to stabilize the size of prediction sets through transfer learning;

Stable Spectral Copula Alignment for Robust Multimodal Learning

Hongkang Zhang (Tsinghua University), Ercan Engin KURUOGLU (Tsinghua University)

Domain AdaptationRepresentation LearningDiffusion modelScore-based ModelContrastive LearningGaussian SplattingImageTextMultimodalityAudio

🎯 What it does: Designed a 'Stable Spectral Copula Alignment (SSCA)' protocol that maintains stable multi-modal alignment under deployment drift.

Stable Velocity: A Variance Perspective on Flow Matching

Donglin Yang (University of Hong Kong), Renjie Liao (University of British Columbia)

GenerationData SynthesisTransformerDiffusion modelScore-based ModelFlow-based ModelRectified FlowImageVideoTextStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Proposes the Stable Velocity framework based on analysis of variance, unifying the training and sampling of flow matching

Stable-GFlowNet: Toward Diverse and Robust LLM Red-Teaming via Contrastive Trajectory Balance

Minchan Kwon (KAIST), Junmo Kim (KAIST)

Adversarial AttackReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningFlow-based ModelGenerative Adversarial NetworkContrastive LearningText

🎯 What it does: This paper proposes the Stable-GFN framework, improving Generative Flow Networks (GFN) for red team attacks on large language models (LLM), addressing the instability of the partition function Z estimation and the mode collapse caused by noisy rewards in traditional GFN.

STABLE: Simulation-Ready Tabletop Layout Generation via a Semantics–Physics Dual System

Zhen Luo (SUSTech), Yanwei Fu (SII)

GenerationData SynthesisOptimizationTransformerLarge Language ModelSupervised Fine-TuningDiffusion modelScore-based ModelFlow-based ModelTextPoint CloudMeshRetrieval-Augmented Generation

🎯 What it does: Propose a dual-system framework called STABLE, where an LLM first generates a rough semantic layout, and then a physics-aware flow model finely corrects the pose, generating desktop scenes that comply with task instructions and are physically simulatable.

StableI2I: Spotting Unintended Changes in Image-to-Image Transition

Jiayang Li (Peking University), Yihao Liu (Shanghai Artificial Intelligence Laboratory)

Image TranslationRestorationAnomaly DetectionTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision Language ModelImageTextMultimodalityBenchmarkChain-of-Thought

🎯 What it does: Proposes StableI2I, a unified and dynamic evaluation framework that explicitly measures content fidelity and coherence in image-to-image (I2I) tasks without relying on reference images, and constructs the StableI2I-Bench benchmark to systematically evaluate the multi-dimensional consistency judgment ability of large models.

STABLEVAL: Disagreement-Aware and Stable Evaluation of AI Systems

Akash Bonagiri (University of California, Davis), Houman Homayoun (University of California, Davis)

Explainability and InterpretabilityData-Centric LearningTextBenchmark

🎯 What it does: Propose the STABLEVAL framework, which performs inconsistency-aware and stable evaluation of human annotations with multiple reviewers.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data

YIYANG FU, Daquan Zhou (Peking University)

Computational EfficiencyRepresentation LearningAdversarial AttackRobotic IntelligenceTransformerVision-Language-Action ModelContrastive LearningImageVideoTextMultimodality

🎯 What it does: This paper addresses the robustness issue of vision-language-action (VLA) models under real-world visual perturbations, proposing a lightweight information bottleneck adapter (IB-Adapter) without additional data, and building the StableVLA model based on it.

Stage-wise Distortion–Perception Traversal in Zero-shot Inverse Problems with Diffusion Models

Jiawei Zhang (Tsinghua University), Yuantao Gu (Tsinghua University)

RestorationSuper ResolutionDiffusion modelImage

🎯 What it does: Propose a two-stage strategy based on a single diffusion model (MAP estimation + re-noised posterior sampling), achieving traversal of the distortion-perception trade-off in zero-shot inverse problems.

STAND: Self-Aware Precondition Induction for Interactive Task Learning

Daniel Weitekamp (Georgia Institute of Technology), Christopher J. MacLellan (Georgia Institute of Technology)

ClassificationOptimizationComputational EfficiencyRepresentation LearningMeta LearningReinforcement Learning from Human FeedbackTransformerReinforcement LearningPrompt EngineeringContrastive LearningTabular

🎯 What it does: Proposed the STAND method for small-sample conditional pre-induction in interactive task learning, with the ability to self-assess learning progress.

Star Elastic: Many-in-One Reasoning LLMs with Efficient Budget Control

Ali Taghibakhshi (Nvidia Corporation), Pavlo Molchanov (Nvidia Corporation)

Computational EfficiencyKnowledge DistillationTransformerLarge Language ModelMixture of ExpertsText

🎯 What it does: Propose a post-training method called Star Elastic, which generates multiple-sized nested sub-models through a single training process and achieves elastic budget control during the inference phase.

STAR-KV: Low-Rank KV Cache Compression via Soft Thresholding for Adaptive Rank Control

Priyansh Bhatnagar (University of California San Diego), Mingu Kang (University of California San Diego)

Computational EfficiencyKnowledge DistillationTransformerLarge Language ModelAuto EncoderContrastive LearningText

🎯 What it does: STAR-KV compresses the KV cache through thresholded low-rank decomposition with Microsoft thresholding, and combines low-rank-aware mixed-precision quantization to achieve dynamic low-rank control for each attention head and decoding block.

STAR-VAE: Structured Topology-Aware Regularization for Audio Reconstruction and Generation

Huadai Liu (Hong Kong University of Science and Technology), Wei Xue (Hong Kong University of Science and Technology)

RestorationGenerationCompressionConvolutional Neural NetworkTransformerDiffusion modelFlow-based ModelAuto EncoderContrastive LearningAudio

🎯 What it does: Proposed Structured Topology-Aware Regularization (STAR) and CNN-Mamba-based STAR-VAE to address the compression, reconstruction, and topology trilemma in audio VAEs.

STAR: Rethinking MoE Routing as Structure-Aware Subspace Learning

Sumin Park (Korea Advanced Institute of Science and Technology), Noseong Park (Korea Advanced Institute of Science and Technology)

OptimizationFederated LearningComputational EfficiencyRepresentation LearningData-Centric LearningLarge Language ModelSupervised Fine-TuningMixture of ExpertsContrastive LearningImageTextMultimodalityTabularBenchmark

🎯 What it does: A STAR (Structure-Aware Routing) framework is proposed within Mixture-of-Experts (MoE) models, which utilizes online Principal Subspace learning (Generalized Hebbian Algorithm, GHA) to model the input structure, and combines it with traditional linear gating to achieve more stable and specialized expert routing.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits

Foivos Paraperas Papantoniou (Imperial College London), Stefanos Zafeiriou (Imperial College London)

GenerationData SynthesisConvolutional Neural NetworkTransformerSupervised Fine-TuningDiffusion modelAuto EncoderContrastive LearningImageVideoAudio

🎯 What it does: Achieved speech-driven avatar generation with facial animation and continuous viewpoint control in a single video diffusion framework, without requiring explicit 3D models.