arXivSub Start free trial

ICML 2026 Papers — Page 53

International Conference on Machine Learning · 6554 papers

Semantic Robustness Certification for Vision-Language Models

Peiyu Yang (University of Melbourne), Sarah Monazam Erfani

Explainability and InterpretabilityAdversarial AttackPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodality

🎯 What it does: Proposes a provably robustness certification framework for vision-language models (VLMs) under semantic-level transformations.

Semantic Router: On the Feasibility of Hijacking MLLMs via a Single Adversarial Perturbation

Changyue Li (Chinese University of Hong Kong), Pinjia He (Chinese University of Hong Kong)

Autonomous DrivingAdversarial AttackRobotic IntelligenceTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodality

🎯 What it does: The study investigates and demonstrates that a single semantic-aware universal perturbation (SAUP) in multimodal large language models (MLLMs) can route different attack targets based on the semantics of the input image, thereby hijacking multi-step stateless decisions in one go.

Semantic Tube Prediction: Beating LLM Data Efficiency with JEPA

Hai Huang (Palona AI), Randall Balestriero (Brown)

Computational EfficiencyKnowledge DistillationRepresentation LearningData-Centric LearningTransformerLarge Language ModelContrastive LearningTextStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: This paper proposes the 'Geodesic Hypothesis,' which suggests that the hidden state trajectory of word sequences in language models is locally linear on the semantic manifold. Based on this hypothesis, a new regularization objective—Semantic Tube Prediction (STP)—is designed to improve the model's signal-to-noise ratio and generation diversity by constraining the hidden states to evolve within a 'semantic tube.'

Semantic-Aware Motion Encoding for Topology-Agnostic Character Animation

Zongye Zhang (Beihang University), Yunhong Wang (Beihang University)

GenerationData SynthesisPose EstimationGraph Neural NetworkTransformerLarge Language ModelVision-Language-Action ModelAuto EncoderContrastive LearningVideoTextMultimodality

🎯 What it does: Proposes a semantic-aware, topology-agnostic motion encoding framework called SATA, which can map actions from any skeletal structure into a unified latent space and enable reproduction and cross-species retargeting.

Semantic-Enriched Latent Visual Reasoning

Tianrun Xu (Tsinghua University), Jing Liu (Zhongguancun Academy)

RecognitionExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerReinforcement LearningVision Language ModelContrastive LearningImageTextMultimodalityBenchmark

🎯 What it does: Proposes a two-stage semantic-enhanced latent visual reasoning framework called SLVR, which learns region latent vectors containing fine-grained attribute semantic information.

Semantic-level Backdoor Attack against Text-to-Image Diffusion Models

Tianxin Chen (Fudan University), Cheng Huang (Fudan University)

GenerationSafty and PrivacyKnowledge DistillationAdversarial AttackTransformerPrompt EngineeringDiffusion modelImageTextMultimodality

🎯 What it does: Design and inject semantic-level backdoors into text-to-image diffusion models, using continuous semantic representations to trigger the backdoors while maintaining consistency across different surface texts.

SemanticNVS: Improving Semantic Scene Understanding in Generative Novel View Synthesis

Xinya Chen (Max Planck Institute for Informatics), Jan Eric Lenssen (Max Planck Institute for Informatics)

GenerationData SynthesisTransformerDiffusion modelContrastive LearningImage

🎯 What it does: Propose the SemanticNVS method, which integrates pre-trained semantic features (DINO) into a multi-view diffusion model to enhance generation quality and consistency.

SemBind: Binding Diffusion Watermarks to Semantics Against Black-Box Forgery Attacks

Xin Zhang (University of Science and Technology of China), Nenghai Yu (University of Science and Technology of China)

GenerationSafty and PrivacyTransformerDiffusion modelContrastive LearningImageText

🎯 What it does: Propose SemBind, which binds watermark signals of latent diffusion models to image semantics via semantic masks, thereby defending against black-box forgery attacks.

Semi-knockoffs: a model-agnostic conditional independence testing method with finite-sample guarantees

Angel David REYERO LOBO (University of Toulouse), Pierre Neuvial (University of Paris-Saclay)

OptimizationFederated LearningExplainability and InterpretabilityData-Centric LearningContrastive LearningTabularBiomedical Data

🎯 What it does: This paper proposes a model-agnostic conditional independence testing method called Semi-knockoffs, which can directly use any pre-trained machine learning model without requiring training-test splits; it provides finite sample type-I error and FDR control in high-dimensional settings; and provides theoretical support, including the optimization stability and double robustness of regularized learners with irrelevant features.

Semi-LAR: Semi-supervised Contrastive Learning with Linear Attention for Removal of Nighttime Flares

Xiyu Zhu (Wuhan University of Science and Technology), Zhengguo Li (Institute for Infocomm Research)

RestorationTransformerContrastive LearningImage

🎯 What it does: Propose a semi-supervised framework named Semi‑LAR, which removes glare from night-time camera images by utilizing a reliable pseudo-label repository and flare-aware contrastive learning, with the model adopting the RaLiFormer linear attention architecture.

Semi-Supervised Gaze Estimation via Disentangled Subspace Contrastive Learning

Qida Tan (Sichuan University), Wenchao Du (Sichuan University)

Pose EstimationDomain AdaptationConvolutional Neural NetworkSupervised Fine-TuningContrastive LearningImage

🎯 What it does: Propose a semi-supervised learning framework called DSCL, which achieves gaze estimation by utilizing a small amount of labeled data and a large amount of unlabeled data. It mainly achieves feature separation through Jacobian regularization, and then improves generalization performance by performing contrastive learning and sequence ranking in each subspace.

Semi-Supervised Hypothesis Testing by Betting on Predictions

Yaniv Tenzer (Technion - Israel Institute of Technology), Yaniv Romano (Technion - Israel Institute of Technology)

Domain AdaptationAnomaly DetectionFederated LearningData-Centric LearningLarge Language ModelReinforcement LearningContrastive LearningTextTabularBiomedical DataBenchmark

🎯 What it does: Propose a semi-supervised sequential hypothesis testing framework based on predictive betting, which enhances the power of testing by utilizing predictions from unlabelled data.

Semi-Supervised Learning for Molecular Graphs via Ensemble Consensus

Rasmus Hannibal Tirsgaard (Technical University of Denmark), Mikkel N. Schmidt (Technical University of Denmark)

Drug DiscoveryGraph Neural NetworkMixture of ExpertsContrastive LearningGraph

🎯 What it does: Propose a semi-supervised learning framework based on integrated consensus, leveraging unlabeled molecular graph data to improve prediction performance.

Semi-Supervised Learning with Noisy Proxy Covariates: Generalization Bounds and Distribution Regression

Kwangho Kim (Korea University), Jisu Kim (Seoul National University)

OptimizationFederated LearningRepresentation LearningData-Centric LearningContrastive LearningTabularTime SeriesSequentialBiomedical DataFinance RelatedPhysics Related

🎯 What it does: Propose a two-stage semi-supervised regression framework, learning kernel features on all proxy covariates, and then performing ridge regression using labeled samples;

Semi-Supervised Neural Super-Resolution for Mesh-Based Simulations

Jiyeon Kim (Yonsei University), Won-Yong Shin (Yonsei University)

Super ResolutionOptimizationGraph Neural NetworkSupervised Fine-TuningContrastive LearningMeshGraphPhysics Related

🎯 What it does: Designed and implemented SuperMeshNet, a semi-supervised neural super-resolution framework based on MPNN, which can predict high-quality high-resolution (HR) solutions from low-resolution (LR) grids using only a small number of HR training samples.

Semi-Supervised Noise Adaptation: Transferring Knowledge from Noise Domain

Yuan Yao (Teleinfo, CAICT), Yu Zhang (Southern University of Science and Technology)

ClassificationDomain AdaptationConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningImageText

🎯 What it does: In a semi-supervised learning setting, a new task called semi-supervised noise adaptation (SSNA) is proposed, which utilizes a synthetic noise domain to enhance the generalization ability of the target domain.

SEMIR: Semantic Minor-Induced Representation Learning on Graphs for Visual Segmentation

Luke James Miller (University of Missouri-Kansas City), Yugyung Lee (University of Missouri-Kansas City)

SegmentationOptimizationComputational EfficiencyRepresentation LearningGraph Neural NetworkAuto EncoderContrastive LearningImageGraphBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Propose the SEMIR framework, which learns task-adapted graph differential representations, compressing high-resolution medical images into boundary-aligned sparse graphs, and then performs segmentation and precisely recovers pixel-level results on these graphs.

SemRep : Generative Code Representation Learning with Code Transformations

Weichen Li (University of Chicago), Kexin Pei (University of Chicago)

Representation LearningAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextSequential

🎯 What it does: Propose the SEMREP framework, which first lets a large model generate semantically preserved code variants (intermediate representations), and then performs specified code transformations based on this representation, thus explicitly learning code semantics.

SENDAI: A Hierarchical Sparse-measurement, EfficieNt Data AssImilation Framework

Xingyue Zhang (University of Washington), J. Nathan Kutz (University of Washington)

Data SynthesisComputational EfficiencyRepresentation LearningRecurrent Neural NetworkDiffusion modelNeural Radiance FieldAuto EncoderGenerative Adversarial NetworkContrastive LearningTabularTime SeriesSequentialAgriculture RelatedPhysics Related

🎯 What it does: A hierarchical sparse measurement and data assimilation framework named SENDAI is constructed, which reconstructs the complete spatial state using observations from extremely low-density sensors;

Sentinel-VLA: A Metacognitive VLA Model with Active Status Monitoring for Dynamic Reasoning and Error Recovery

Wenhao Li (University of Sydney), Chang Xu (University of Sydney)

Robotic IntelligenceReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsVision-Language-Action ModelImageTextMultimodality

🎯 What it does: Propose Sentinel-VLA, a metacognitive Vision-Language-Action model with active state monitoring and dynamic reasoning capabilities.

Separating Representation from Reconstruction Enables Scalable Text Encoders

Megi Dervishi (FAIR Meta), Yann LeCun (New York University)

Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringAuto EncoderContrastive LearningText

🎯 What it does: Propose the CrossBERT architecture, which separates representation learning from reconstruction tasks, and achieves higher training efficiency and better frozen performance through high masking rates and complementary masking strategies.

SEPS: Semantic-Enhanced Patch Slimming Framework for Fine-Grained Cross-Modal Alignment

Xinyu Mao (University of Electronic Science and Technology of China), Ming Sun (University of Electronic Science and Technology of China)

RetrievalRepresentation LearningTransformerLarge Language ModelVision Language ModelContrastive LearningImageTextMultimodality

🎯 What it does: Proposed the SEPS framework, addressing the semantic sparsity bias in cross-modal alignment of vision-language tasks through dual-granularity semantic calibration and significance-guided metric aggregation.

Sequential Group Composition: A Window into the Mechanics of Deep Learning

Giovanni Luca Marchetti (KTH Royal Institute of Technology), Nina Miolane (UC Santa Barbara)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningConvolutional Neural NetworkRecurrent Neural NetworkTransformerContrastive LearningSequential

🎯 What it does: Propose the 'sequential group composition' task, where the network learns to compute the cumulative product through group elements in a sequence; and explain via Fourier analysis how the network gradually learns irreducible representations from small initialization to gradient descent, revealing how deep models leverage the associative law to significantly improve efficiency.

Sequential Kernel-based Conditional Independence Testing via Adaptive Betting

Zheng He (University of British Columbia), Danica J. Sutherland (University of British Columbia)

Anomaly DetectionOptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencySupervised Fine-TuningReinforcement LearningContrastive LearningImageTabularTime SeriesSequentialBiomedical DataBenchmark

🎯 What it does: Propose a sequential conditional independence test method based on the betting framework and self-normalized kernel conditional independence statistic.

SERA: Soft-Verified Efficient Repository Agents

Ethan Shen (Allen Institute of Artificial Intelligence), Tim Dettmers (Allen Institute of Artificial Intelligence)

Data SynthesisComputational EfficiencyRepresentation LearningAI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningAgentic AITextSequentialRetrieval-Augmented Generation

🎯 What it does: Propose a method for generating synthetic trajectories through Soft-Verified Generation (SVG), utilizing line-level recall for soft verification to eliminate the dependency on unit tests, thereby enabling fast and low-cost training of code agents for any codebase.

Server-Proximal Aggregation for Federated Domain-Incremental Learning under Partial Participation: Task-Uniform Convergence and Backward Transfer

Longtao Xu (Stony Brook University), Jian Li (Stony Brook University)

Domain AdaptationOptimizationFederated LearningRepresentation LearningConvolutional Neural NetworkTransformerSupervised Fine-TuningAgentic AIContrastive LearningImageText

🎯 What it does: Proposes the SPECIAL (Server-Proximal Efficient Continual Aggregation for Learning) algorithm, specifically designed for Federated Domain-Incremental Learning (FDIL) in partial participation environments. The algorithm introduces a lightweight 'anchor' regularization on the server side, which performs a quadratic weighted average between the current aggregated model and the global model of the previous task after each aggregation round, thereby suppressing cross-task drift without requiring experience replay or task-specific network heads.

Set Diffusion: Interpolating Token Orderings between Autoregression and Diffusion for Fast and Flexible Decoding

Marianne Arriola (Cornell University), Volodymyr Kuleshov (Cornell University)

GenerationComputational EfficiencyRepresentation LearningTransformerLarge Language ModelDiffusion modelScore-based ModelText

🎯 What it does: Propose the Set Diffusion model, which achieves a flexible decoding that can approximate autoregressive (AR) while also being diffusion-based, by performing factorization on token sets with variable length and variable positions within a discrete diffusion framework.

Set-Coupled Guidance: Set-Level Coordination in Diffusion-Based Dataset Distillation

Ziang Gan (Beijing Normal University), Libao Zhang (Beijing Normal University)

Knowledge DistillationData-Centric LearningTransformerDiffusion modelScore-based ModelContrastive LearningImage

🎯 What it does: Propose a plug-and-play controller called Set-Coupled Guidance (SCG), which enables diffusion-based dataset distillation to use set-level feedback at every step, thus achieving collaborative evolution of samples within the set.

Set-Preserving Calibration from Conformal P-Values to E-Values

Nabil Alami (Mohamed bin Zayed University of Artificial Intelligence), Souhaib Ben Taieb (Mohamed bin Zayed University of Artificial Intelligence)

Federated LearningExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTabularBenchmark

🎯 What it does: Proposed a Set-Preserving calibration method from conformal p-value to e-value (P2E), and applied it to cross-conformal prediction (CCP) and conformal aggregation (CA).

SetPO: Set-Level Policy Optimization for Diversity-Preserving LLM Reasoning

Chenyi Li (Peking University), Nan Duan (JD Explore Academy)

OptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkChain-of-Thought

🎯 What it does: Propose SetPO, a set-level diversity preservation strategy optimization method, aimed at improving the diversity and accuracy of large language models (LLMs) in reasoning tasks.

SF-Mamba: Rethinking State Space Model for Vision

Masakazu Yoshimura (Sony Group Corporation), Takeshi Ohashi (Sony Group Corporation)

ClassificationObject DetectionSegmentationTransformerImage

🎯 What it does: Proposed the SF-Mamba model, which combines a hybrid architecture of visual Mamba and Transformer, optimized for speed bottlenecks caused by short sequence lengths.

SFCLTA: Spectral Fusion Contrastive Learning with Topology-Adaptive Graph Augmentation

Zhuo Xu (Beijing Normal University), Yue Wang (Central University of Finance and Economics)

Representation LearningGraph Neural NetworkContrastive LearningGraph

🎯 What it does: This paper proposes an unsupervised graph contrastive learning framework called SFCLTA, which combines spectral fusion and adaptive topological augmentation to improve the representation learning effect on heterogeneous graphs (highly heterogeneous graphs).

SFedPO: Streaming Federated Learning with a Prediction Oracle under Temporal Shifts

Jinrui Zhou (University of Science and Technology of China), Mingjun Xiao (University of Science and Technology of China)

OptimizationFederated LearningComputational EfficiencyContrastive LearningImageTabular

🎯 What it does: Propose a streaming federated learning framework called SFedPO, which utilizes a prediction oracle to capture the temporal evolution of client data distributions and achieve dynamic sampling and aggregation;

SG2Loc: Sequential Visual Localization on 3D Scene Graphs

Nicole Damblon (ETH Zurich), Daniel Barath (ETH Zurich)

Pose EstimationOptimizationComputational EfficiencyGraph Neural NetworkSupervised Fine-TuningNeural Radiance FieldContrastive LearningSimultaneous Localization and MappingOptical FlowPoint CloudMeshGraphSequential

🎯 What it does: This paper proposes a sequential visual localization method based on 3D scene graphs and particle filters, utilizing semantic features and coarse grids to achieve efficient localization.

SGERA: Stein-Guided ECG-Report Alignment for ECG Representation Learning

Jian Chen (University of Hong Kong), Edith Cheuk-Han Ngai (University of Hong Kong)

Domain AdaptationRepresentation LearningConvolutional Neural NetworkTransformerScore-based ModelContrastive LearningTextMultimodalityElectronic Health RecordsElectrocardiogram

🎯 What it does: Proposed the SGERA framework, which uses the Stein kernel to achieve dual-level alignment between ECG and reports, addressing the limitations of traditional CLIP-style alignment in cross-modal distribution differences;

SGMD: Score Gradient Matching Distillation for Few-Step Video Diffusion Distillation

Zhuguanyu Wu (Beihang University), Xianglong Liu (Beihang University)

GenerationData SynthesisFederated LearningKnowledge DistillationTransformerDiffusion modelScore-based ModelContrastive LearningVideoText

🎯 What it does: The study proposes a new few-step video diffusion model distillation method called SGMD, which directly optimizes pseudo scores and achieves collaborative tracking between the generator and the score network through dual potentials.

SHAP-Guided Kernel Actor-Critic for Explainable Reinforcement Learning

Na Li (Zhejiang University), Xinyu Li (Huazhong University of Science and Technology)

Explainability and InterpretabilityReinforcement LearningTabularSequential

🎯 What it does: Propose a dual-time-scale Actor-Critic algorithm called RSA2C based on RKHS-SHAP, which adaptively weights the policy gradient and advantage objective using state feature importance, thereby achieving interpretable reinforcement learning.

ShapCCS: Shapley-Driven Client Coreset Selection in Federated Learning

Shuo Ji (Southwest Jiaotong University), Jie Xu (University of Leeds)

Federated LearningExplainability and InterpretabilityComputational EfficiencyContrastive LearningImage

🎯 What it does: Proposes ShapCCS, a client core sample selection method based on Shapley values, aimed at significantly reducing computational and communication costs in federated learning.

Shape of Thought: Progressive Object Assembly via Visual Chain-of-Thought

Yu Huo (Chinese University of Hong Kong), Xiaoying Tang (Chinese University of Hong Kong)

GenerationData SynthesisExplainability and InterpretabilityTransformerVision Language ModelRectified FlowAuto EncoderImageTextMultimodalityBenchmarkChain-of-Thought

🎯 What it does: Propose Shape-of-Thought (SoT), a generative framework that progressively assembles target objects in the 2D rendering domain through visual Chain-of-Thought.

Shapley Neuron Values for Continual Learning: Which Neurons Matter Most?

Mohammad Ali Vahedifar (Aarhus University), Qi Zhang (Aarhus University)

ClassificationExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationConvolutional Neural NetworkContrastive LearningImage

🎯 What it does: Propose the Shapley Neuron Values (SNV) framework, which evaluates the importance of each neuron using Shapley values, freezes the most important neurons during continuous learning, and avoids using buffers and expanding the network.

Shapley Regularized Neural Granger Causality

Maolin Yang (Xiamen University), MUYI LI

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningRecurrent Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningTabularTime SeriesBiomedical DataBenchmarkPhysics Related

🎯 What it does: Proposed an information theory-based global feature importance measure (Info-Shap) and corresponding differentiable regularizations (Shap and F-Shap), embedding them into neural network training to improve neural Granger causal discovery.

Shared Lexical Task Representations Explain Behavioral Variability In LLMs

Zhuonan Yang (Brown University), Ellie Pavlick (Brown University)

Explainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: Studied the sensitivity of prompts on the performance of large language models (LLMs), finding that the shared 'lexical task heads' across different prompt styles can encode the task itself and explain differences in model behavior.

Shared Semantics, Divergent Mechanisms: Unsupervised Feature Discovery by Aligning Semantics and Mechanisms

Hyunjin Cho (Yonsei University), Jaehyung Kim (Yonsei University)

Explainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelContrastive LearningText

🎯 What it does: Propose a distribution-level unsupervised feature discovery method that combines semantic embeddings and sequence-level mechanism attribution to cluster the diverse continuations generated by LLMs;

Sharp Concentration Bounds for Bundle-Valued Statistics on Manifolds

Swagatam Das (Indian Statistical Institute), Vaclav Snasel (VSB Technical University of Ostrava)

OptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyRepresentation LearningDrug DiscoveryReinforcement Learning from Human FeedbackNeural Architecture SearchProtein Structure PredictionDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageVideoTextMultimodalityPoint CloudMeshGraphTabularTime SeriesSequentialBiomedical DataFibre Orientation DistributionDiffusion Tensor ImagingReview/Survey PaperBenchmarkPhysics RelatedStochastic Differential EquationOrdinary Differential EquationAudio

🎯 What it does: This paper proposes a non-asymptotic concentration theory for vector bundle-valued statistics on manifolds. By translating sample observations to a common reference fiber and taking their average, explicit high-probability upper bounds of Hoeffding and Bernstein are provided, and the global distortion (holonomy) error caused by the non-uniqueness of paths is quantified. Additionally, the paper presents the minimax error lower bound in the limit, the resampling robust estimator (median-of-means), and the central limit theorem.

Sharp description of local minima in the loss landscape of high-dimensional two-layer ReLU neural networks

Jie Huang (Chalmers University of Technology and University of Gothenburg), Stefano Sarao Mannelli (Chalmers University of Technology and University of Gothenburg)

OptimizationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningAuto EncoderContrastive LearningPhysics RelatedStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Under the teacher-student high-dimensional Gaussian input framework, the overall loss landscape of a two-layer ReLU network is studied, providing an exact representation of its low-dimensional summary statistics and revealing the discrete hierarchical structure of local minima.

Sharp Empirical Bernstein Inequalities for the Variance of Bounded Random Variables

Diego Martinez-Taboada (Carnegie Mellon University), Aaditya Ramdas (Carnegie Mellon University)

OptimizationExplainability and InterpretabilityComputational EfficiencyTabularTime SeriesReview/Survey Paper

🎯 What it does: This paper proposes a fully empirical Bessel inequality for the variance of bounded random variables, providing confidence intervals and confidence sequences for both batch and sequential settings.

Sharp Inequalities between Total Variation and Hellinger Distances for Gaussian Mixtures

Joonhyuk Jung (University of Chicago), Chao Gao (University of Chicago)

OptimizationExplainability and InterpretabilityRepresentation LearningContrastive Learning

🎯 What it does: Proves a sharp inequality between total variation and Hellinger distance for Gaussian mixture models, and constructs matching extreme examples.

SHARP-Q: Spectral Hessian Alignment and Rectification for Post-training Quantization

Menghao Lv (Zhejiang University), Mingli Song (Zhejiang University)

ClassificationSuper ResolutionOptimizationComputational EfficiencyConvolutional Neural NetworkTransformerContrastive LearningImage

🎯 What it does: Propose a post-training quantization framework called SHARP-Q based on information geometry, which uses Hessian-Aware Rectification (HAR) and Dynamic Fisher-Subspace Compensation (DFSC) to perform geometric correction and precise approximation of the Fisher information matrix during the quantization process.

Sharper Generalization Guarantees for Asynchronous SGD: Beyond Lipschitzness, Smoothness and Data Homogeneity

Xie Yufeng, Yunwen Lei (University of Hong Kong)

OptimizationFederated LearningContrastive LearningTabularBenchmarkStochastic Differential Equation

🎯 What it does: This paper provides an upper bound on the generalization error (i.e., expected risk) of asynchronous stochastic gradient descent (ASGD) after a fine-grained analysis of its algorithmic stability, under relaxed assumptions that remove common conditions such as Lipschitz, smoothness, and data uniformity.

Sharpness-Aware Minimization Can Hallucinate Minimizers

Chanwoong Park (Seoul National University), Insoon Yang (Seoul National University)

OptimizationConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningImage

🎯 What it does: This paper investigates the phenomenon where the standard Sharpness-Aware Minimization (SAM) update rule may get stuck at non-critical points of the original loss function when using a large radius ρ, and names this phenomenon illusory minimizers.

Sharpness-Aware Pretraining Mitigates Catastrophic Forgetting

Ishaan Watts (Carnegie Mellon University), Aditi Raghunathan (Carnegie Mellon University)

OptimizationComputational EfficiencyKnowledge DistillationRepresentation LearningData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningText

🎯 What it does: This paper studies how pre-training optimization strategies affect the stability of models during subsequent fine-tuning, quantization, and other modification processes, focusing on reducing loss curvature (sharpness) to alleviate catastrophic forgetting.

Sheaf Neural Networks on SPD Manifolds: Second-Order Geometric Representation Learning

Yuhan Peng (Nanyang Technological University), Kelin Xia (Nanyang Technological University)

Representation LearningDrug DiscoveryGraph Neural NetworkContrastive LearningGraphBiomedical Data

🎯 What it does: Developed the first Sheaf neural network capable of performing direct computations on the SPD (symmetric positive definite) manifold, used to learn second-order geometric representations of molecular graphs.

SHERPA: Fine-tuning Segment Anything Models with Task-relevant Guidance

Jingcheng Xie (University of Science and Technology of China), Zhiwei Xiong (University of Science and Technology of China)

SegmentationKnowledge DistillationRepresentation LearningTransformerSupervised Fine-TuningContrastive LearningImageVideoBiomedical Data

🎯 What it does: Propose the SHERPA framework, which utilizes task-related features from a small SAM to guide the fine-tuning of a large SAM, thereby enhancing performance on specialized tasks while maintaining its generality.

Shift-Dependent Asymmetry: Orthogonal Inverse Low-Rank Adaptation for Federated Medical Segmentation

Xingyue Zhao (Chinese Academy of Medical Sciences), Bo XU

SegmentationFederated LearningTransformerBiomedical Data

🎯 What it does: For the federated learning scenario in medical image segmentation, a reverse asymmetric fine-tuning framework based on LoRA (IAT) and subspace orthogonal regularization (SOR) are proposed, achieving structured separation and gradient isolation of shared and personalized parameters.

Shifting the Breaking Point of Flow Matching for Multi-Instance Editing

Carmine Zaccagnino (Unimore), Silvia Cascianelli (Unimore)

Image TranslationImage HarmonizationGenerationTransformerPrompt EngineeringVision Language ModelDiffusion modelFlow-based ModelImageTextMultimodalityBenchmark

🎯 What it does: This paper proposes an instance-decoupled attention mechanism to achieve multi-instance text-guided image editing within a flow matching framework, enabling multi-region editing in a single inference without causing semantic interference.

SHINE: A Scalable In-Context Hypernetwork for Mapping Context to LoRA in a Single Pass

Yewei Liu (Peking University), Muhan Zhang (Peking University)

Computational EfficiencyKnowledge DistillationRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringDiffusion modelTextRetrieval-Augmented Generation

🎯 What it does: Proposed a scalable context-to-LoRA hypernetwork called SHINE, which can map any semantic context to high-quality LoRA adapters in a single forward pass, and then directly use them for LLM inference without needing to access the context again.

Short Chains, Deep Thoughts: Balancing Reasoning Efficiency and Intra-Segment Capability via Split-Merge Optimization

Runquan Gui (University of Science and Technology of China), Feng Wu (University of Science and Technology of China)

OptimizationComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextChain-of-Thought

🎯 What it does: Designed a structured optimization framework called CoSMo based on a split-merge approach to eliminate redundant reasoning paragraphs while maintaining reasoning depth.

Shortcut-Resistant CAM Distillation for Long-Tailed Recognition

Wenhai Wan (National Engineering Research Center for Big Data Technology and System), Songcan Chen (MIIT Key Laboratory of Pattern Analysis and Machine Intelligence)

ClassificationRecognitionKnowledge DistillationSupervised Fine-TuningContrastive LearningImage

🎯 What it does: This paper proposes a framework called Shortcut-Resistant CAM Distillation (SRCD), which utilizes Class Activation Maps (CAM) to transfer the object-centric attention learned from the head classes with large sample sizes to the tail classes with scarce samples, thereby suppressing the model's reliance on spurious features in long-tailed scenarios.

Should I Have Expressed a Different Intent? Counterfactual Generation for LLM-Based Autonomous Control

Amirmohammad Farzaneh (Northeastern University), Osvaldo Simeone (Northeastern University)

Autonomous DrivingOptimizationRobotic IntelligenceReinforcement Learning from Human FeedbackTransformerLarge Language ModelTextTime SeriesRetrieval-Augmented Generation

🎯 What it does: This paper proposes a reliable counterfactual generation framework (CCG) for answering the question 'What would happen if I expressed different intentions' in large language model-driven autonomous control systems.

Show, Don't Tell: Morphing Latent Reasoning into Image Generation

Harold Haodong Chen (Hong Kong University Of Science And Technology (Gz)), Ying-Cong Chen (Hong Kong University Of Science And Technology (Gz))

GenerationData SynthesisTransformerReinforcement LearningPrompt EngineeringVision Language ModelDiffusion modelImageTextMultimodality

🎯 What it does: Achieve dynamic self-correction by introducing implicit latent reasoning in the text-to-image generation process.

Shrinking the Variance: Shrinkage Baselines for Reinforcement Learning with Verifiable Rewards

Guanning Zeng (Carnegie Mellon University), Andrea Zanette (Carnegie Mellon University)

OptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringContrastive LearningTextMultimodality

🎯 What it does: Proposes an unbiased reward baseline estimation method based on the James-Stein shrinkage idea to reduce the gradient variance of the large-scale inference model RLVR.

Shuffle the Context: RoPE-Perturbed Self-Distillation for Long-Context Adaptation

Zichong Li (Georgia Institute of Technology), Weizhu Chen (Microsoft)

Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: In long context adaptation, robustness of LLMs in long sequence reasoning tasks is enhanced by perturbing RoPE indices and leveraging self-distillation consistency regularization.

Shuffling-Aware Optimization for Private Vector Mean Estimation

Shun Takagi (LY Corporation), Seng Pei Liew (LY Corporation)

OptimizationFederated LearningSafty and PrivacyContrastive LearningGaussian Splatting

🎯 What it does: This paper studies unbiased high-dimensional mean estimation under the single-message shuffling model and proposes an optimization framework centered around the shuffling index.

SI-IGCL: Subject Invariance-aware Inverse Graph Contrastive Learning for Psychiatric Disorder Identification

Jiayu Lu (Taiyuan University of Technology), Bin Wang (Taiyuan University of Technology)

ClassificationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerAuto EncoderContrastive LearningGraphBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease

🎯 What it does: Proposed a two-stage SI-IGCL framework, first using self-supervised inverse graph contrastive learning and structural preservation reconstruction constraints to remove individual differences, and then fine-tuning the classifier to complete mental illness identification.

SiameseNorm: Breaking the Barrier to Reconciling Pre/Post-Norm

Tianyu Li (Tsinghua University), Gao Huang (Tsinghua University)

OptimizationComputational EfficiencyRepresentation LearningTransformerLarge Language ModelMixture of ExpertsDiffusion modelContrastive LearningImageText

🎯 What it does: Propose a two-branch (Siamese) residual structure called SiameseNorm, which balances the gradient stability of Pre-Norm and the representational power of Post-Norm.

SIGMA-PPG: Statistical-prior Informed Generative Masking Architecture for PPG Foundation Model

Zongheng Guo (Politecnico di Milano), Manuela Ferrario (Politecnico di Milano)

Representation LearningData-Centric LearningTransformerReinforcement LearningDiffusion modelAuto EncoderContrastive LearningTime SeriesBiomedical DataElectrocardiogram

🎯 What it does: Developed a generative occlusion framework called SIGMA-PPG driven by statistical priors, which discretizes PPG signals using VQ-VAE and enhances robustness to noise through semantic consistency constraints, while utilizing reinforcement learning teacher-student adversarial occlusion to learn physiologically relevant structures.

Sign Lock-In: Randomly Initialized Weight Signs Persist and Bottleneck Sub-Bit Model Compression

Akira Sakai (Fujitsu Limited), Yuma Ichikawa (Fujitsu Limited)

CompressionOptimizationComputational EfficiencyRepresentation LearningConvolutional Neural NetworkTransformerContrastive LearningImageText

🎯 What it does: Studied the phenomenon of weight sign locking in sub-bit compression and proposed theoretical explanations and practical methods.

Signal Strength Estimation in Logistic Regression Using Data Splitting

Weihao Li (Tsinghua University), Jun S. Liu (Tsinghua University)

ClassificationExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTabularBiomedical Data

🎯 What it does: A consistent estimation method for signal strength κ in high-dimensional logistic regression is established by utilizing data splitting and non-decomposable regularization.

Signature-Informed Transformer for Asset Allocation

Yoontae Hwang (Pusan National University), Stefan Zohren (University of Oxford)

OptimizationTransformerTime SeriesFinance Related

🎯 What it does: Propose a unified end-to-end asset allocation model (Signature-Informed Transformer, SIT) that combines path signatures with Transformer, which performs feature extraction and directly outputs portfolio weights, avoiding error amplification in the traditional prediction-then-optimization process;

SIKA-GP: Accelerating Gaussian Process Inference with Sparse Inducing Kernel Approximations for Bayesian Deep Learning

Wenyuan Zhao (Texas A&M University), Chao Tian (Texas A&M University)

Computational EfficiencyRepresentation LearningImageTextTabular

🎯 What it does: Propose SIKA-GP, which achieves efficient GP inference and training by constructing sparse Laplace kernel basis functions on Dyadic grids, supporting deep feature learning.

Sim2Reason: Solving Physics Olympiad via Reinforcement Learning on Physics Simulators

Mihir Prabhudesai (Carnegie Mellon University), Deepak Pathak (Carnegie Mellon University)

OptimizationTransformerLarge Language ModelReinforcement LearningPrompt EngineeringDiffusion modelScore-based ModelGenerative Adversarial NetworkTextBenchmarkPhysics Related

🎯 What it does: Generate a large number of physics question-answer pairs using a physics simulator, and perform post-training of large language models using reinforcement learning;

SimGFM: Simplifying Discrete Flow Matching for Graph Generation

Chunyu Luo (Beihang University), Lei Shi (Beihang University)

GenerationDrug DiscoveryGraph Neural NetworkDiffusion modelScore-based ModelFlow-based ModelGraph

🎯 What it does: Propose a graph generation method called SimGFM based on Discrete Flow Matching (DFM), achieving efficient sampling through a minimalist design.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models

Sultan Alshehri (Carnegie Mellon University), Marios Savvides (Carnegie Mellon University)

RetrievalExplainability and InterpretabilityComputational EfficiencyPrompt EngineeringVision Language ModelScore-based ModelContrastive LearningImageTextMultimodalityBenchmark

🎯 What it does: Researchers address the limitations of dual-encoder vision-language models in logical reasoning by proposing a method called decomposed reasoning and logical constraint score editing (LCSE), enabling the model to correctly perform Boolean constraints such as negation, conjunction, and disjunction while maintaining retrieval performance.

SIMPC: Learning Self-Induced Mirror-Point Consistency for Unsupervised Point Cloud Denoising

Chengwei Zhang (Chinese Academy of Sciences), Longyong Chen (Chinese Academy of Sciences)

RestorationGraph Neural NetworkTransformerScore-based ModelAuto EncoderContrastive LearningPoint Cloud

🎯 What it does: Proposed an unsupervised point cloud denoising framework named SIMPC, which achieves one-to-one correspondence between points and underlying surfaces through self-induced mirror point consistency, thereby accurately locating and removing noise.

Simple Algorithms for Bad Triangle Transversals with Applications to Correlation Clustering

Florian Adriaens (University of Helsinki), Nikolaj Tatti (University of Helsinki)

ClassificationRecommendation SystemOptimizationContrastive LearningGraphReview/Survey Paper

🎯 What it does: The paper proposes multiple 2-approximation algorithms for the Bad Triangle Transversal (BTT) problem and reveals the close relationship between BTT and Correlation Clustering (CC) on complete graphs.

Simple Approximation and Derivative Free Inference-Time Scaling for Diffusion Models via Sequential Monte Carlo on Path Measures

Chenyang Wang (Peking University), Yiping Lu (Northwestern University)

GenerationData SynthesisDiffusion modelScore-based ModelImageTextStochastic Differential Equation

🎯 What it does: Propose a gradient-free, near-error-free diffusion model sampling method named URGE during inference, which utilizes the Girsanov transformation to perform importance weighting on the path space and resamples at each step, thereby achieving an unbiased approximation of task-specific reward functions.

Simple Denoising Diffusion Language Models

Huaisheng Zhu (Penn State University), Teng Xiao (University of Washington)

GenerationData SynthesisTransformerLarge Language ModelDiffusion modelText

🎯 What it does: Proposed a simplified denoising diffusion language model (SDDLM), which improves training efficiency by optimizing only the tokens replaced by noise.

Simple Policy Gradients for Reasoning with Diffusion Language Models

Anthony Zhan (Stanford University)

TransformerLarge Language ModelReinforcement LearningPrompt EngineeringDiffusion modelText

🎯 What it does: The AGRPO algorithm is proposed for discrete diffusion language models, utilizing a step-level Markov decision process to achieve post-training, significantly improving inference consistency and quality.

Simple yet Effective: Low-Rank Spatial Attention for Neural Operators

Zherui Yang (University of Science and Technology of China), Ligang Liu (University of Science and Technology of China)

OptimizationComputational EfficiencyTransformerDiffusion modelScore-based ModelFlow-based ModelRectified FlowNeural Radiance FieldAuto EncoderGenerative Adversarial NetworkContrastive LearningPoint CloudMeshGraphTabularTime SeriesBenchmarkPhysics Related

🎯 What it does: Propose a Low-Rank Spatial Attention (LRSA) module, implementing low-rank global mixing using standard Transformer primitives to construct neural operators.

SimpleGPT: Improving GPT via A Simple Normalization Strategy

Marco Chen (Tsinghua University), Rong Xiao (Intellifusion)

Computational EfficiencyRepresentation LearningTransformerLarge Language ModelContrastive LearningText

🎯 What it does: Proposed a simple strategy called SimpleNorm, which inserts a normalization layer immediately after each linear mapping in the Transformer, and designed a new GPT model (SimpleGPT) based on this strategy.

SimpleMem: Efficient Lifelong Memory for LLM Agents

Jiaqi Liu (UNC Chapel Hill), Huaxiu Yao (UNC Chapel Hill)

RetrievalCompressionComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose SimpleMem, an efficient lifelong memory framework for LLM agents

SimulCost: A Cost-Aware Benchmark and Toolkit for Automating Physics Simulations with LLMs

Yadi Cao (University of California San Diego), Rose Yu (University of California San Diego)

OptimizationComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextBenchmarkPhysics Related

🎯 What it does: Proposes the SIMULCOST benchmark to evaluate the success rate and computational cost of LLMs in parameter tuning during physical simulations.

Simultaneous Confidence Bounds for Aggregated Effects via Exact Subset Optimization

Weihang Xu (University of Washington), Xinshang Wang (Alibaba Group)

OptimizationTabularBenchmark

🎯 What it does: Proposes a joint confidence interval estimation method for the aggregation effect of downward closed subset families, utilizing bootstrap-calibrated subset maximization standardized statistics to achieve effective confidence lower bounds for data selection subsets.

Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards

Yiran Jenny Shen (University of California San Diego), Prithviraj Ammanabrolu (University of California San Diego)

Reinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsText

🎯 What it does: Propose the MAHALO framework, combining multi-action-head DPO training with process reward model (PRM)-guided decoding, to achieve multi-objective alignment for both verifiable and non-verifiable rewards.

Simultaneous Speech-to-Speech Translation Without Aligned Data

Tom Labiausse (Kyutai), Neil Zeghidour (Gradium)

Data SynthesisKnowledge DistillationRecurrent Neural NetworkTransformerSupervised Fine-TuningReinforcement LearningPrompt EngineeringContrastive LearningTextMultimodalityAudio

🎯 What it does: Propose a parallel speech translation system called HibikiZero that does not require word-level alignment data. First, the system trains a base model using sentence-level alignment, and then optimizes real-time performance and translation quality through process reward based on BLEU and GRPO reinforcement learning.

Single-Head Attention in High Dimensions: A Theory of Generalization, Weights Spectra, and Scaling Laws

Fabrizio Boncoraglio (Ecole polytechnique federale de Lausanne), Lenka Zdeborová (Ecole polytechnique federale de Lausanne)

OptimizationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelContrastive LearningTabularSequentialPhysics Related

🎯 What it does: Analyze the empirical risk minimization of single-head attention layers in high dimensions, derive high-dimensional analytical formulas for training and test errors, spectral structure, and generalization, and explain the mechanisms of feature recovery and scaling laws.

Single-Rollout Hidden-State Dynamics for Training-Free RLVR Data Selection

Jianghao Wu (Monash University), Yasmeen George (Monash University)

Data-Centric LearningReinforcement Learning from Human FeedbackTransformerReinforcement LearningPrompt EngineeringContrastive LearningTextBiomedical DataRetrieval-Augmented Generation

🎯 What it does: This paper proposes SHIFT, a one-time, no-training, no-label RLVR data selection method.

Singular Bayesian Neural Networks

Mame Diarra Toure (McGill University), David A. Stephens (McGill University)

ClassificationExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTextTabularBiomedical DataBenchmark

🎯 What it does: A low-rank variational inference framework is proposed for Bayesian neural networks, parameterizing the weights as W = AB⊤ to learn the posterior distribution of factors, thereby enabling structured correlation modeling of weights.

Singular Proxies for Adaptive Caching in Diffusion Language Models

Wenhao Sun (Nanyang Technological University), Dacheng Tao (Nanyang Technological University)

Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelDiffusion modelContrastive LearningTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose SPA-Cache, an efficient caching framework for diffusion language models, achieving sparse updates and adaptive budget allocation.

Singular Vectors of Attention Heads Align with Features

Gabriel Franco (Boston University), Mark Crovella (Boston University)

Explainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelAuto EncoderContrastive LearningText

🎯 What it does: This paper studies the alignment between singular vectors of Transformer attention heads and internal feature vectors of the model (SVF) phenomenon, and provides theoretical explanations and experimental verification.

Singularity-aware Optimization via Randomized Geometric Probing: Towards Stable Non-smooth Optimization

Ruoran Xu (Xi'an Jiaotong-Liverpool University), Qiufeng Wang (Xi'an Jiaotong-Liverpool University)

OptimizationImage

🎯 What it does: Proposed a novel adaptive optimizer called S-Adam, which utilizes random geometric probing to dynamically adjust the learning rate, thus achieving more stable training on non-smooth deep learning loss surfaces.

Sinkhorn Normalization of Diffusion Kernels

Nathan Kessler (ENS Paris-Saclay), Jean Feydy (Inria)

OptimizationComputational EfficiencyRepresentation LearningDiffusion modelContrastive LearningGaussian SplattingMultimodalityPoint CloudMeshGraph

🎯 What it does: Designed a symmetric Sinkhorn normalization method that transforms any similarity matrix into a heat diffusion-type smoothing operator, enabling Laplace-like smoothing on irregular geometric data such as point clouds and voxels without grid structures.

Sinkhorn Treatment Effects: A Causal Optimal Transport Measure

Medha Agarwal (University of Washington), Alex Luedtke (Harvard Medical School)

OptimizationExplainability and InterpretabilityComputational EfficiencyData-Centric LearningContrastive LearningImageTabularBenchmark

🎯 What it does: Proposes a distribution effect measure based on entropy-regularized optimal transport called Sinkhorn Treatment Effect (STE), which captures the complete differences between the potential outcome distributions of treatment and control groups, and provides an estimable unbiased estimator and testing method.

SINQ: Sinkhorn-Normalized Quantization for Calibration-Free Low-Precision LLM Weights

Lorenz K Muller, Lukas Cavigelli (Huawei Technologies)

Computational EfficiencyKnowledge DistillationTransformerLarge Language ModelContrastive LearningText

🎯 What it does: Propose the SINQ method, introducing the second-axis scale and Sinkhorn-Knopp normalization when quantizing LLM weights, to improve the accuracy of low-precision quantization;

SIPO: Stabilized and Improved Preference Optimization for Aligning Diffusion Models

Xiaomeng Yang (Shanghai Academy of AI for Science), Hao Li (Fudan University)

GenerationOptimizationReinforcement Learning from Human FeedbackTransformerReinforcement LearningDiffusion modelImageVideoText

🎯 What it does: This paper proposes an algorithm called SIPO, which stably and efficiently optimizes human preferences in diffusion models, thereby improving the quality of image and video generation and the alignment with human preferences.

Size Transferability of Graph Convolutional Networks across Sparsity: A Generalized Graphon Perspective

Qinji Shu (Fudan University), Bo Hu (Fudan University)

Computational EfficiencyRepresentation LearningConvolutional Neural NetworkGraph Neural NetworkAuto EncoderContrastive LearningGraph

🎯 What it does: Studied the scale transferability of graph convolutional networks under different sparsity levels, and proposed GWCN as the non-zero limit for sparse graphs, providing a unified upper bound on error.

SJD-SV: Speculative Jacobi Decoding with Semantics Verification for Autoregressive Image Generation

Baoquan Zhang (Harbin Institute of Technology), Yunming Ye (Harbin Institute of Technology)

GenerationTransformerPrompt EngineeringDiffusion modelImageText

🎯 What it does: Propose a Speculative Jacobi Decoding with Semantic Verification (SJD-SV) method for semantic-level verification in autoregressive image generation, aiming to improve token verification rate and accelerate generation.

SkelHCC: A Hyperbolic CLIP-Driven Cache Adaptation Framework for Skeleton-based One-Shot Action Recognition

Yanan Liu (Yunnan University), Qiuhong Ke (Monash University)

RecognitionTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningTextGraph

🎯 What it does: Propose the SkelHCC framework, which utilizes hyperbolic CLIP for hierarchical alignment between skeleton and language, and achieves context-aware one-shot skeleton action recognition without training by employing LLM-guided multi-granularity voting cache during testing.

Sketch-Based Low-Rank Model Merging with Shared Circulant Transforms

Zhiming Zhang (Beihang University), Dan Meng (Chinese Academy of Sciences)

Computational EfficiencyRepresentation LearningTransformerImageText

🎯 What it does: This paper proposes the CircuMerge framework, which efficiently merges multi-task LoRA adapters by utilizing shared circulant matrix transformations and compressed sampling (sketching);

SKETCH: Semantic Key-Point Conditioning for Long-Horizon Vessel Trajectory Prediction

Linyong Gan (Chinese University of Hong Kong), Shuhang Chen (COSCO SHIPPING Advanced Technology Institute)

Autonomous DrivingOptimizationComputational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackTransformerDiffusion modelAuto EncoderContrastive LearningTime SeriesSequentialRetrieval-Augmented Generation

🎯 What it does: Propose a hierarchical ship trajectory prediction framework based on semantic key points (NKP), decomposing long-term trajectory prediction into two steps: global intent and local motion.