ICML 2026 Papers — Page 66
International Conference on Machine Learning · 6554 papers
WinDeskGround: A Benchmark for Robust GUI Grounding in Complex Multi-Window Desktop Environments
Haoren Zhao (Hangzhou Dianzi University), Zhen Wang (Hangzhou Dianzi University)
Data SynthesisTransformerLarge Language ModelPrompt EngineeringVision Language ModelDiffusion modelImageTextMultimodalityBenchmark
🎯 What it does: This paper proposes the WinDeskGround benchmark, which utilizes a parameterized multi-window synthesis framework to generate complex desktop scenes across dimensions such as multi-window, occlusion, and semantic similarity, in order to evaluate the GUI localization robustness of multimodal large language models (MLLMs).
Winformer: Transcending Pairwise Similarity for Time-series Generation
Haoyi Zhou (Beihang University), Jianxin Li (Beihang University)
GenerationData SynthesisAutonomous DrivingOptimizationTransformerDiffusion modelContrastive LearningTime SeriesFinance RelatedPhysics Related
🎯 What it does: Propose Winformer, a diffusion model based on window attention, for cross-domain time series generation;
WinQ: Accelerating Quantization-Aware Training of Language Models Around Saddle Points
Dongyue Li (Northeastern University), Steven Li (Meta AI)
OptimizationComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelContrastive LearningTextStochastic Differential Equation
🎯 What it does: Propose the WINQ method to accelerate low-precision language model quantization-aware training through weight linear interpolation reset and noise injection;
WISE: World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation
Yuwei Niu (Shenzhen Graduate School, Peking University), Li Yuan (Shenzhen Graduate School, Peking University)
GenerationTransformerLarge Language ModelPrompt EngineeringDiffusion modelImageTextMultimodalityBenchmarkChain-of-Thought
🎯 What it does: This paper proposes a benchmark called WISE to evaluate the ability of text-to-image models in world knowledge reasoning and complex semantic understanding.
With Argus Eyes: Assessing Retrieval Gaps via Uncertainty Scoring to Detect and Remedy Retrieval Blind Spots
Zeinab Sadat Taghavi (Lucerne University of Applied Sciences and Arts), Andreas Marfurt (Lucerne University of Applied Sciences and Arts)
RetrievalTransformerLarge Language ModelContrastive LearningTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Design and evaluate the RPS metric for measuring retrieval blind spot risks, and propose the ARGUS solution to improve RAG retrieval performance by pre-indexing supplementary knowledge.
WMVLM: Evaluating Diffusion Model Image Watermarking via Vision-Language Models
Zijin Yang (University of Science and Technology of China), Nenghai Yu (University of Science and Technology of China)
Safty and PrivacyExplainability and InterpretabilityKnowledge DistillationTransformerSupervised Fine-TuningReinforcement LearningVision Language ModelDiffusion modelContrastive LearningImageTextMultimodality
🎯 What it does: Proposed the WMVLM framework, which uses a vision-language model to conduct unified and interpretable evaluation of residual watermarks and semantic watermarks generated by diffusion models;
Words & Weights: Streamlining Multi-Turn Interactions via Co-Adaptation
Chenxing Wei (Shenzhen University), Yao Shu (Hong Kong University Of Science And Technology)
OptimizationComputational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextSequentialBenchmark
🎯 What it does: Propose a joint optimization framework for test-time strategy adaptation called ROSA2, which adjusts both prompts and model parameters simultaneously during multi-round interactions.
Words Towards Explainability: Caption Label-Free Learning via Dual Loop Agentic Time Series Captioning
Difei Hou (Zhejiang University), Chunhui Zhao (Zhejiang University)
Anomaly DetectionExplainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringTabularTime SeriesBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposed a label-free caption generation method for time series (Caption Label‑Free Learning, CLFL), achieving autonomous exploration and feedback optimization of time series captions through the Dual Loop Agentic Captioning (DLAC) framework.
World Guidance: World Modeling in Condition Space for Action Generation
Yue Su (University of Hong Kong), Xihui Liu (University of Hong Kong)
Autonomous DrivingOptimizationRobotic IntelligenceTransformerVision Language ModelVision-Language-Action ModelRectified FlowAuto EncoderContrastive LearningWorld ModelImageVideoTextMultimodality
🎯 What it does: Propose the WoG framework, which compresses future observations into a compact conditional space, helping Vision-Language-Action models generate fine-grained actions based solely on current observations.
World Models in Pieces: Structural Certification for General Agents
Yikai Lu (Chinese University of Hong Kong, Shenzhen), Tongxin Li (Chinese University of Hong Kong, Shenzhen)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningReinforcement LearningWorld ModelTabularTime SeriesSequential
🎯 What it does: This paper proposes a structural certification method for certifying general agents in specific tasks, addressing the problem that general agents cannot maintain consistent performance in complex environments.
World-Model Inspired Emotion-aware Token Refinement for Training-Free Multimodal Emotion Recognition
Kejun Liu (China University of Geosciences (Wuhan)), Hongyan Zhang (China University of Geosciences (Wuhan))
RecognitionTransformerLarge Language ModelVision Language ModelContrastive LearningWorld ModelMultimodality
🎯 What it does: Propose a training-free, plug-and-play inference-time token refinement framework called WETR, which enhances the accuracy and stability of emotion recognition by selecting and weighting tokens within a frozen multimodal large language model.
World-R1: Reinforcing 3D Constraints for Text-to-Video Generation
Weijie Wang (Zhejiang University), Bohan Zhuang (Zhejiang University)
GenerationReinforcement Learning from Human FeedbackTransformerReinforcement LearningVision Language ModelDiffusion modelFlow-based ModelRectified FlowGaussian SplattingVideoText
🎯 What it does: Post-training existing text-to-video generation models to generate geometrically consistent and 3D reconstructable videos under text prompts.
World-Shaper: A Unified Framework for 360° Panoramic Editing
Dong Liang (City University of Hong Kong), Rynson W. H. Lau (City University of Hong Kong)
GenerationData SynthesisTransformerSupervised Fine-TuningDiffusion modelImageBenchmark
🎯 What it does: Studied 360° panoramic image editing in the equirectangular projection (ERP) domain, and proposed a unified World-Shaper framework that supports five types of editing operations, including object addition, deletion, modification, and movement.
WorldCache: Accelerating World Models for Free via Heterogeneous Token Caching
Weilun Feng (Chinese Academy of Sciences), Chuanguang Yang (Chinese Academy of Sciences)
GenerationComputational EfficiencyTransformerDiffusion modelScore-based ModelWorld ModelImageVideo
🎯 What it does: This paper proposes WorldCache, a training-agnostic acceleration framework that significantly improves the inference speed of diffusion world models (such as HunyuanVoyager and Aether) through heterogeneous token caching.
WorldComp2D: Spatio-semantic Representations of Object Identity and Location from Local Views
Seongmin Jin (Hanyang University), Doo Seok Jeong (Hanyang University)
Pose EstimationComputational EfficiencyRepresentation LearningConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningImage
🎯 What it does: Proposed the WorldComp2D framework, which maps local observations to a structured spatial semantic latent space and achieves object localization through a local locator.
WorldCompass: Reinforcement Learning for Long-Horizon World Models
Zehan Wang (Zhejiang University), Zhou Zhao (Zhejiang University)
GenerationTransformerReinforcement LearningVision-Language-Action ModelDiffusion modelFlow-based ModelWorld ModelImageVideoText
🎯 What it does: This paper proposes WorldCompass, a post-training reinforcement learning framework for video world models, enabling the model to more accurately follow actions and maintain visual quality during long-sequence interactive generation.
WorldMirror: Universal 3D World Reconstruction with Any-Prior Prompting
Yifan Liu (Chinese University of Hong Kong), Chunchao Guo (Tencent)
GenerationData SynthesisPose EstimationDepth EstimationTransformerPrompt EngineeringNeural Radiance FieldGaussian SplattingImagePoint CloudMesh
🎯 What it does: Propose a unified feed-forward 3D reconstruction model that can receive any geometric prior (camera pose, intrinsic parameters, depth) and predict point clouds, depth, normals, camera parameters, and 3D Gaussian distributions in one go to achieve novel view synthesis.
WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling
Wenqiang Sun (Hong Kong University of Science and Technology), Chunchao Guo (Tencent Hunyuan)
GenerationData SynthesisKnowledge DistillationTransformerDiffusion modelAuto EncoderContrastive LearningVideo
🎯 What it does: Designed and implemented WorldPlay, a real-time interactive world modeling framework capable of generating 720p 24FPS video in real-time under user control, while maintaining long-term geometric consistency.
WorldTravel: A Realistic Multimodal Travel-Planning Benchmark with Tightly Coupled Constraints
Zexuan Wang (ByteDance Seed), Wenhao Huang (ByteDance Seed)
TransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmarkChain-of-Thought
🎯 What it does: Proposes the WorldTravel benchmark and its multimodal environment WorldTravel-Webscape for evaluating travel itinerary planning under real-world web interfaces.
WS-GRPO: Weakly-Supervised Group-Relative Policy Optimization for Rollout-Efficient Reasoning
Gagan Mundada (University of California San Diego), Junda Wu (University of California San Diego)
OptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkChain-of-Thought
🎯 What it does: Propose the WS-GRPO method, which utilizes a weakly supervised preference model to convert terminal correctness into prefix-level rewards, significantly reducing the length of chain-of-thought reasoning while maintaining high accuracy.
WUSH: Near-Optimal Adaptive Transforms for LLM Quantization
Jiale Chen (Institute of Science and Technology Austria), Dan Alistarh (Institute of Science and Technology Austria)
OptimizationComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelTextBenchmark
🎯 What it does: This paper proposes a block-level linear transformation-based LLM quantization method called WUSH, which can approximately optimally reduce the impact of extremes on low-precision representation in weight-activation quantization;
X-EviProbe: Post-hoc Parameter-Free Evidential Uncertainty Quantification for Frozen Graph Neural Networks
Chenghua Guo (Beijing University of Posts and Telecommunications), Xi Zhang (Beijing University of Posts and Telecommunications)
Anomaly DetectionExplainability and InterpretabilityComputational EfficiencyRepresentation LearningGraph Neural NetworkAuto EncoderContrastive LearningGraph
🎯 What it does: Propose X-EviProbe, a post-hoc parameter-free framework for evidence uncertainty quantification, which converts a frozen graph neural network into Dirichlet distribution predictions, taking into account both prior and geometric information.
X-MoGe: A Cross-Modal Adaptation Framework with Mixture-of-Experts and Geometry Guidance for Heterogeneous Collaborative Perception
Wenkai Lin (Xiamen University), Chenglu Wen (Xiamen University)
Domain AdaptationAutonomous DrivingRepresentation LearningTransformerMixture of ExpertsContrastive LearningOptical FlowImageMultimodalityPoint Cloud
🎯 What it does: Propose the X-MoGe framework, which utilizes pixel-level Mixture-of-Experts (P-MoE) and geometry-guided fusion to achieve cross-modal semantic alignment and geometric consistency for heterogeneous multi-agent collaborative perception;
XDomainBench: Diagnosing Reasoning Collapse in High-Dimensional Scientific Knowledge Composition
Gong Zhiren (Nanyang Technological University), Wei Yang Bryan Lim (Nanyang Technological University)
Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningDrug DiscoveryTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringMixture of ExpertsTextTabularSequentialBenchmarkPhysics RelatedRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This paper proposes XDomainBench, an interactive interdisciplinary benchmark for diagnosing reasoning breakdowns in high-dimensional combinations of scientific knowledge;
xKV: Cross-Layer KV-Cache Compression via Aligned Singular Vector Extraction
Chi-Chih Chang (Cornell University), Mohamed S. Abdelfattah (Cornell University)
CompressionComputational EfficiencyTransformerLarge Language ModelScore-based ModelAuto EncoderText
🎯 What it does: This paper proposes an untrained, cross-layer KV-Cache compression method called xKV, which significantly reduces the storage requirements of KV-Cache during long context inference.
XPERT: Expert Knowledge Transfer for Effective Training of Language Models
Chang Liu (Southeast University), Xin Geng (Southeast University)
Knowledge DistillationTransformerLarge Language ModelMixture of ExpertsText
🎯 What it does: Propose the XPERT framework, which leverages the general expert knowledge of pre-trained MoE LLMs for extraction, compression, and parameter-level adaptation, thereby providing pre-training initialization for language models of different scales.
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations
Shichao Fan (Beijing Innovation Center of Humanoid Robotics), Jian Tang (Beijing Innovation Center of Humanoid Robotics)
Representation LearningRobotic IntelligenceTransformerSupervised Fine-TuningVision Language ModelVision-Language-Action ModelAuto EncoderContrastive LearningImageVideoTextMultimodality
🎯 What it does: Propose the XR-1 framework, which realizes a vision-language-action model across robots and tasks by leveraging a unified visual-motion code (UVMC).
XRPO: Pushing the Limits of GRPO with Targeted Exploration and Exploitation
Udbhav Bamba (Unaffiliated), Fan Lai (University Of Illinois Urbana Champaign)
AI Code AssistantTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextSequentialChain-of-Thought
🎯 What it does: Propose the XRPO framework, improving rollout allocation based on GRPO, introducing ICL seeds and advantage sharpening techniques to enhance the reasoning and programming performance of large language models.
XSkill: Continual Learning from Experience and Skills in Multimodal Agents
Guanyu Jiang (Hong Kong University of Science and Technology), Yi R. Fung (Hong Kong University of Science and Technology)
Federated LearningComputational EfficiencyKnowledge DistillationRepresentation LearningRobotic IntelligenceReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsVision Language ModelVision-Language-Action ModelImageTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose the XSKILL framework, which utilizes experience and skills under visual contexts to achieve continuous learning and adaptive tool usage for multi-modal agents.
XSpecMesh: Quality-Preserving Auto-Regressive Mesh Generation Acceleration via Multi-Head Speculative Decoding
Dian Chen (Xiamen University), Shengchuan Zhang (Xiamen University)
GenerationComputational EfficiencyKnowledge DistillationTransformerPrompt EngineeringDiffusion modelMesh
🎯 What it does: Designed and implemented the XSpecMesh framework, achieving acceleration in autoregressive mesh generation through multi-head speculative decoding, cross-attention, validation resampling, knowledge distillation, and LoRA fine-tuning.
XTransfer: Modality-Agnostic Few-Shot Model Transfer for Human Sensing at the Edge
Yu Zhang (Macquarie University), Tao Gu (Macquarie University)
ClassificationDomain AdaptationComputational EfficiencyRepresentation LearningMeta LearningNeural Architecture SearchConvolutional Neural NetworkRecurrent Neural NetworkGraph Neural NetworkSpiking Neural NetworkTransformerPrompt EngineeringAuto EncoderGenerative Adversarial NetworkContrastive LearningImageTextMultimodalityTime SeriesAudio
🎯 What it does: Proposes a cross-modal, few-shot, resource-efficient model transfer framework called XTransfer, which can transfer publicly pre-trained models to human perception tasks with only a small amount of target sensor data.
XYZFlow: Scaling Multidimensional Shortcut Flows for Efficient Generative Modeling
Jinxiu Liu (Chinese University of Hong Kong), Weiyang Liu (Chinese University of Hong Kong)
GenerationKnowledge DistillationTransformerScore-based ModelFlow-based ModelRectified FlowImage
🎯 What it does: Propose the XYZFlow framework, achieving scalable generation through multi-dimensional conditioning for flow matching.
You Can Learn Tokenization End-to-End with Reinforcement Learning
Sam Dauncey (ETH Zürich), Roger Wattenhofer (ETH Zürich)
Computational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningScore-based ModelContrastive LearningText
🎯 What it does: Studies how to use reinforcement learning methods end-to-end to learn token boundaries within large language models, proposing a tokenization strategy based on score-function estimation.
You Don’t Need All That Attention: Surgical Memorization Mitigation in Text-to-Image Diffusion Models
Kairan Zhao (University of Warwick), Peter Triantafillou (University of Warwick)
GenerationData SynthesisSafty and PrivacyTransformerPrompt EngineeringDiffusion modelContrastive LearningImageTextMultimodality
🎯 What it does: This paper proposes a memory elimination framework called GUARD for text-to-image diffusion models during inference, and implements a dynamic weakening mechanism on cross-attention called CA-in-GUARD, which can significantly reduce the phenomenon of models memorizing training data without compromising generation quality.
You Don't Protect if You Don't Expect: Breaking the Key Assumption behind CLIP's Test-Time Defenses
Ruize Zhang (Institute of Computing Technology, Chinese Academy of Sciences), Sheng Tang (Institute of Computing Technology, Chinese Academy of Sciences)
Adversarial AttackTransformerVision Language ModelContrastive LearningImageTextMultimodality
🎯 What it does: This paper investigates the key assumptions of the CLIP model in test-time defense and proposes the CLIP-MAD attack strategy, revealing the vulnerability of existing defense methods when facing adaptive attacks.
You Need Better Attention Priors
Elon Litman (Stanford University), Gabe Guo (Stanford University)
Computational EfficiencyRepresentation LearningTransformerLarge Language ModelContrastive LearningImageTextBiomedical Data
🎯 What it does: This paper proposes a new attention mechanism called GOAT (Generalized Optimal Transport Attention with Trainable priors), which treats attention as an entropy-regularized Optimal Transport problem and introduces learnable prior distributions to replace the traditional implicit uniform prior.
Z-Erase: Enabling Concept Erasure in Single Stream Diffusion Transformers
Nanxiang Jiang (Beihang University), wenjun wu
RestorationGenerationData SynthesisTransformerSupervised Fine-TuningDiffusion modelFlow-based ModelAuto EncoderContrastive LearningImageTextMultimodality
🎯 What it does: Proposes Z-Erase, a concept elimination method specifically designed for single-stream text-image diffusion transformers, addressing the problem of generation collapse caused by directly transferring traditional elimination methods.
Zero Sum SVD: Balancing Loss Sensitivity for Low Rank LLM Compression
Ali Abbasi (Vanderbilt University), Soheil Kolouri (Vanderbilt University)
CompressionOptimizationComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelText
🎯 What it does: Propose Zero Sum SVD (ZS-SVD), a post-training low-rank compression method, which achieves heterogeneous rank allocation for large language models through a global zero-sum selection rule combined with activation whitening and gradient direction sensitivity, with an optional single projection gradient correction followed by re-truncation.
Zero-Flow Encoders
Yakun Wang (University of Bristol), Taiji Suzuki (RIKEN)
Representation LearningData-Centric LearningConvolutional Neural NetworkRecurrent Neural NetworkSpiking Neural NetworkTransformerScore-based ModelFlow-based ModelRectified FlowAuto EncoderContrastive LearningImageMultimodalityGraphTabularTime SeriesFinance RelatedStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Propose a Zero-Flow Encoder, which utilizes the zero-flow condition to detect conditional independence, extracting sufficient representations from data to achieve non-parametric representation learning.
Zero-Shot 3D Question Answering via Hierarchical View-to-Token Transportation
Dongsheng Wang (Shenzhen University), Hui Huang (Shenzhen University)
Computational EfficiencyRepresentation LearningData-Centric LearningTransformerPrompt EngineeringVision Language ModelContrastive LearningVideoTextMultimodalityPoint CloudRetrieval-Augmented Generation
🎯 What it does: Propose a zero-shot 3D question answering method called KeyVT, which utilizes hierarchical key views (Key-View) and key tokens (Key-Token) selection techniques to enable 2D vision-language models (VLMs) to obtain as rich 3D context as possible under limited input budgets.
Zero-shot Active Mapping via Fused 360-BEV Representations and Vision–Language Models
Yuanze Wang (Shanghai Jiao Tong University), Mengzhu Wang (Hebei University of Technology)
Autonomous DrivingOptimizationRepresentation LearningTransformerPrompt EngineeringVision Language ModelSimultaneous Localization and MappingImageTextMultimodalityBenchmark
🎯 What it does: This paper proposes a zero-shot active mapping method based on a frozen vision-language model (VLM).
Zero-Shot Off-Policy Learning
Arip Asadulaev (Mohamed bin Zayed University of Artificial Intelligence), Martin Takáč (Mohamed bin Zayed University of Artificial Intelligence)
Reinforcement LearningDiffusion modelScore-based ModelFlow-based ModelRectified FlowAuto EncoderGenerative Adversarial NetworkContrastive LearningTabularTime SeriesSequential
🎯 What it does: This paper proposes a zero-shot offline reinforcement learning method called ZOL, which rapidly adapts a pre-trained behavioral base model without training or interaction during testing by using the stationary density ratio derived from forward-backward successor representations.
Zero-Shot Rankability: Revealing Latent Ordinal Structure in Multimodal Large Language Models via Language
Nam Hyeon-Woo (POSTECH), Tae-Hyun Oh (KAIST)
Explainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityAudio
🎯 What it does: Investigate whether zero-shot rankability exists in the embedding space of multimodal large language models (MLLMs) and whether it can be directly extracted through text prompts
Zero-Shot Text-to-Motion Evaluation using Video Language Models
Yuwen Ji (Zhejiang University), Yue Zhang (Westlake University)
GenerationData SynthesisRepresentation LearningTransformerPrompt EngineeringVision Language ModelVideoTextMultimodality
🎯 What it does: Propose VeMo, a zero-shot text-motion evaluation framework;
Zero-source LLM Hallucination Detection with Human-like Criteria Probing
Jiahao Yang (South China University of Technology), Mingkui Tan (South China University of Technology)
ClassificationExplainability and InterpretabilityTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringTextRetrieval-Augmented Generation
🎯 What it does: Proposed a zero-shot hallucination detection framework called HCPD, which utilizes a pre-trained LLM to perform multi-dimensional interpretable scoring of question-answer pairs through 'human-like criterion probing,' and obtains the final correctness judgment through multi-sampling aggregation.
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
Jonathan Roberts (University of Cambridge), Samuel Albanie (University of Hong Kong)
TransformerPrompt EngineeringVision Language ModelImageMultimodalityBenchmarkChain-of-Thought
🎯 What it does: Propose ZeroBench — a benchmark consisting of 100 manually crafted, multi-step reasoning visual reasoning tasks, which are unsolvable by all state-of-the-art LMMs at the time of release;
ZeroDiff: Zero-Shot Time Series Reconstruction via Informed-Prior Diffusion
Yingda Fan (University of Pittsburgh), Xiaowei Jia (University of Pittsburgh)
RestorationGenerationData SynthesisRecurrent Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningTabularTime Series
🎯 What it does: Propose ZeroDiff, a zero-shot time series reconstruction method, which first estimates the distribution moments (mean, variance) of the target sequence using exogenous variables and learns a dynamics model in a standardized space, then calibrates the prior using information-theoretic prior diffusion, and finally reconstructs the target time series at unobserved locations.
Zeroth-Order Forward-Only SNN Training Inspiring Neuromorphic On-Chip Learning
Mingyue Qin (Shanghai Jiao Tong University), Fei Wen (Shanghai Jiao Tong University)
ClassificationOptimizationComputational EfficiencySpiking Neural NetworkImageVideo
🎯 What it does: Based on pre-trained models, we propose a zeroth-order forward-only trained spiking neural network method (SZO) to achieve online learning.
Zeroth-Order Non-Log-Concave Sampling with Variance Reduction and Applications to Inverse Problems
M. Berk Sahin (Purdue University), Abolfazl Hashemi (Purdue University)
OptimizationScore-based ModelImageBiomedical DataMagnetic Resonance ImagingPhysics RelatedStochastic Differential Equation
🎯 What it does: Propose a variance-reduced zeroth-order Langevin Monte Carlo (ZO-LMC) method for sampling from non-log-convex distributions, and extend it to posterior sampling in black-box inverse problems;
Zeroth-Order Optimization at the Edge of Stability
Minhak Song (KAIST), Sewoong Oh (University of Washington)
OptimizationConvolutional Neural NetworkRecurrent Neural NetworkTransformerContrastive LearningGaussian SplattingImageSequential
🎯 What it does: Proposes the mean-square linear stability theory for zeroth-order (ZO) optimization methods, and verifies that ZO methods often reside at the mean-square edge of stability (EoS) during deep learning training.
ZeroUnlearn: Few-Shot Knowledge Unlearning in Large Language Models
Yujie Lin (Xiamen University), Jinsong Su (Xiamen University)
Safty and PrivacyComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelPrompt EngineeringContrastive LearningText
🎯 What it does: Proposes the ZeroUnlearn framework for knowledge forgetting with few samples, which can quickly remove sensitive information from LLMs without retraining or significant fine-tuning.
Zeus: Towards Tuning-Free Foundation Model for Time Series Analysis
Yisong Fu (Institute of Computing Technology Chinese Academy of Sciences), Fei Wang (Institute of Computing Technology Chinese Academy of Sciences)
ClassificationAnomaly DetectionOptimizationComputational EfficiencyTransformerLarge Language ModelContrastive LearningTime Series
🎯 What it does: Propose ZEUS, a general-purpose time series foundation model based on multi-scale Transformer, capable of performing five tasks—forecasting, imputation, anomaly detection, and classification—without any fine-tuning.
ZipMoE: Efficient On-Device MoE Serving via Lossless Compression and Cache-Affinity Scheduling
Yuchen Yang (Nanjing University), Zhi-Hua Zhou (Nanjing University)
CompressionComputational EfficiencyTransformerMixture of ExpertsText
🎯 What it does: ZipMoE, a system for implementing Mixture-of-Experts (MoE) inference on mobile devices, achieves efficient inference through lossless compression and cache-affinity scheduling.
Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception
Lai Wei (Shanghai Jiao Tong University), Weiran Huang (Shanghai Jiao Tong University)
Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerReinforcement LearningVision Language ModelContrastive LearningImageTextMultimodalityBenchmark
🎯 What it does: Proposes the Region-to-Image Distillation method, transforming the inference-time 'Zooming' into preprocessing during training, achieving fine-grained perception during single forward inference.