ICML 2026 Papers — Page 49
International Conference on Machine Learning · 6554 papers
Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting
Howard Chen (Princeton University), Danqi Chen (Princeton University)
Federated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningAdversarial AttackHyperparameter SearchData-Centric LearningMeta LearningTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningText
🎯 What it does: Investigate and systematically compare the catastrophic forgetting behavior of supervised fine-tuning (SFT) and reinforcement learning (RL) in the post-training of language models, and explore the fundamental reasons for their differences and mitigation strategies.
Rethink the Role of Neural Decoders in Quantum Error Correction
Ge Yan (Nanyang Technological University), Yuxuan Du (Nanyang Technological University)
Computational EfficiencyConvolutional Neural NetworkRecurrent Neural NetworkGraph Neural NetworkTransformerTabularTime SeriesPhysics Related
🎯 What it does: This paper systematically evaluates and implements five neural network decoders for surface codes, seeking a trade-off between accuracy and microsecond-level latency, and provides an FPGA-deployable compressed pipeline.
Rethinking 1-bit Optimization Leveraging Pre-trained Large Language Models
Zhijun Tu (Huawei Technologies Co Ltd), Yunhe Wang (Huawei Technologies Co Ltd)
OptimizationComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelText
🎯 What it does: This paper proposes a 1-bit quantization framework called BinaryLLM based on pre-trained large language models, combining consistent progressive training, binary-aware initialization, and dual-scale compensation to achieve high-performance 1-bit LLMs.
Rethinking 3D Shape Generation: Diffusion over Superquadrics
Zhiyang Liu (National University of Singapore), Marcelo H Ang Jr (National University of Singapore)
GenerationData SynthesisTransformerDiffusion modelScore-based ModelPoint CloudMesh
🎯 What it does: Migrate the diffusion process for 3D shape generation from high-dimensional dense geometry (voxels, point clouds, meshes) to sparse, explicit superquadric (Superquadric) parameter sets, performing denoising directly in these parameter spaces.
Rethinking Attention in Spiking Transformers: Overcoming Density Bias with Set Similarity
JinGyo Lim (Seoul National University of Science and Technology), Seong-Eun Kim (Seoul National University of Science and Technology)
ClassificationSpiking Neural NetworkTransformerContrastive LearningAudio
🎯 What it does: Proposed DiceFormer, a spiking neural network Transformer for audio classification, addressing the bias in traditional spiking attention towards firing density, and capturing frequency-time features through the SADA module;
Rethinking Calibration for Early-Exit Neural Networks
Piotr Kubaty (Jagiellonian University), Bartosz Wójcik (Jagiellonian University)
ClassificationImage TranslationRestorationComputational EfficiencyRepresentation LearningConvolutional Neural NetworkTransformerPrompt EngineeringContrastive LearningImage
🎯 What it does: This paper proposes a failure prediction method for early exit neural networks (EENN), called EEFP, and uses this method to achieve dynamic inference without relying on traditional confidence thresholds.
Rethinking Code Complexity Through the Lens of Large Language Models
Chen Xie (Shanghai Jiao Tong University), Beijun Shen (Shanghai Jiao Tong University)
Explainability and InterpretabilityComputational EfficiencyAI Code AssistantTransformerLarge Language ModelPrompt EngineeringContrastive LearningText
🎯 What it does: The study investigates the correlation between traditional code complexity metrics and the performance of large language models (LLMs), and proposes a new LLM-focused code complexity metric called LM-CC based on token entropy and semantic hierarchy.
Rethinking Contrastive Learning for Graph Collaborative Filtering: Limitations and a Simple Remedy
Geon Lee (KAIST), Kijung Shin (KAIST)
Recommendation SystemHyperparameter SearchGraph Neural NetworkContrastive LearningGraph
🎯 What it does: By analyzing the prediction mechanism of graph collaborative filtering (GCF), this study investigates the role of neighbor pairs in prediction scores and proposes a new contrastive learning objective, NT-SSM, for type-aware weighted updates of neighbor pairs during training.
Rethinking Convergence in MoE Training: The Role of Routing Sparsity
Weihao Zhu (Nanjing University of Science and Technology), Haixia Zhang (Shandong University)
OptimizationComputational EfficiencyRepresentation LearningTransformerSupervised Fine-TuningMixture of ExpertsImageText
🎯 What it does: Study the impact of routing sparsity on convergence speed in Mixture-of-Experts training, and provide the theoretically optimal number of activated experts K*.
Rethinking Depth Pruning for Vision Transformers: A Heterogeneity-Aware Perspective
Zhenfeng Su (Chinese University of Hong Kong), Wenxuan Wang (Renmin University of China)
ClassificationComputational EfficiencyKnowledge DistillationTransformerContrastive LearningImage
🎯 What it does: Proposes a heterogeneity-aware deep pruning method for visual Transformers called HetDPT, and extends it to HetDPT+ for simultaneous pruning of width and depth
Rethinking Efficient Graph Coarsening via a Non-Selfishness Principle
Xu Bai (Shanghai Jiao Tong University), Meng Jin (Shanghai Jiao Tong University)
ClassificationComputational EfficiencyRepresentation LearningGraph Neural NetworkLarge Language ModelContrastive LearningTextGraph
🎯 What it does: Proposed a non-selfish graph coarsening method called NOPE and its fast version NOPE*, based on the neighborhood interference index, achieving linear memory and near-linear time text-attribute graph coarsening.
Rethinking Evaluation Paradigms in IBP-based Certified Training
Konstantin Kaulen (RWTH Aachen University), Holger H. Hoos (RWTH Aachen University)
ClassificationAdversarial AttackHyperparameter SearchConvolutional Neural NetworkImage
🎯 What it does: Construct a complete natural accuracy–validation accuracy Pareto frontier through multi-objective hyperparameter search, systematically evaluate IBP-based certified training methods, and rediscover and quantify their performance boundaries.
Rethinking Feature Alignment in Generalist Graph Anomaly Detection: A Relational Fingerprint-based Approach
Yujing Liu (Griffith University), Shirui Pan (Griffith University)
Domain AdaptationAnomaly DetectionGraph Neural NetworkTransformerContrastive LearningGraph
🎯 What it does: Propose REFI-GAD, a general graph anomaly detection framework that aligns heterogeneous graph features through the construction of relational fingerprints.
Rethinking Federated Prompt Learning for Medical Images: From Textual Tuning to Visual Manifold Anchoring
Yipan Wei (Wuhan University), Bo Du (Wuhan University)
ClassificationFederated LearningRepresentation LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasoundRetrieval-Augmented Generation
🎯 What it does: This paper reconsiders the application of federated prompt learning in medical imaging, shifting from traditional text tuning to visual manifold anchoring.
Rethinking Forgery Attacks on Semantic Watermarks in Black-Box Settings: A Geometric Distortion Perspective
Cheng-Yi Lee (Academia Sinica), Jun-Cheng Chen (Academia Sinica)
Anomaly DetectionAdversarial AttackTransformerDiffusion modelScore-based ModelContrastive LearningImageText
🎯 What it does: Studies the theoretical and detection methods of semantic watermark forgery attacks under black-box conditions.
Rethinking Gating Mechanism in Sparse MoE: Handling Arbitrary Modality Inputs with Confidence-Guided Gate
Liangwei Nathan Zheng (Adelaide University), Weitong Chen (Adelaide University)
ClassificationRepresentation LearningData-Centric LearningTransformerSupervised Fine-TuningMixture of ExpertsContrastive LearningTextMultimodalityBiomedical DataElectronic Health RecordsAudio
🎯 What it does: Propose ConfSMoE, a confidence-guided sparse mixture-of-experts model for robust handling of missing modalities in multi-modal learning.
Rethinking Generative Image Pretraining: How Far Are We From Scaling Up Next-Pixel Prediction?
Xinchen Yan (Google Deepmind), Quoc V Le
ClassificationRestorationGenerationTransformerDiffusion modelAuto EncoderContrastive LearningImage
🎯 What it does: Explored the scalability of autoregressive pixel-level prediction across different tasks (pixel loss, ImageNet classification, image completion) and resolutions (16×16, 32×32, 64×64), providing the optimal ratio between computational resources, data volume, and model scale.
Rethinking Genomic Modeling Through Optical Character Recognition
Hongxin Xiang (Hunan University), xiangxiang Zeng
RecognitionComputational EfficiencyRepresentation LearningConvolutional Neural NetworkTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataBenchmarkAgriculture Related
🎯 What it does: Modeling genomic sequences as OCR-style document understanding by rendering DNA sequences into structured 2D visual layouts and training visual-language models.
Rethinking GNNs and Missing Features: Challenges, Evaluation and a Robust Solution
Francesco Ferrini (University of Trento), Manfred Jaeger (Aalborg University)
ClassificationExplainability and InterpretabilityRepresentation LearningData-Centric LearningGraph Neural NetworkSupervised Fine-TuningPrompt EngineeringContrastive LearningGraphTabularReview/Survey PaperBenchmark
🎯 What it does: This paper re-examines the robustness of graph neural networks under the scenario of missing node features, proposing a more practically meaningful evaluation process, four dense feature datasets, and multiple missing mechanisms, and conducting experiments using the lightweight method GNNmim based on missing indicators.
Rethinking Graph Transformers as Graph Signal Denoisers: The Role of Block-Diagonal Priors
Jiaming Zhuo (Hebei University of Technology), Liang Yang
ClassificationGraph Neural NetworkTransformerScore-based ModelContrastive LearningGraph
🎯 What it does: Proposed a new graph Transformer model called BDFormer, which achieves efficient node classification by using potential anchors in the global channel and introducing a block diagonal prior
Rethinking Human Intent-to-CAD: Parametric CAD Model Generation via Cooperative Multi-Task Alignment and Spatial-Aware Reinforcement Learning
Qingwang Zhang (Fudan University), Xiangdong Zhou (Fudan University)
GenerationData SynthesisAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringVision Language ModelDiffusion modelScore-based ModelContrastive LearningImageTextMultimodalityMesh
🎯 What it does: This paper proposes a unified framework called HiCAD, which directly maps hand-drawn sketches and text descriptions into executable CadQuery code, achieving the generation of parameterized CAD models from human intent.
Rethinking Instruction Drift as a Sampling Error: SNR-Aware Power Distributions for Long-Horizon Robotic Planning
Kewei Chen (Chongqing Institute of Green and Intelligent Technology, Chinese Academy of Sciences), Mingsheng Shang (Chongqing Institute of Green and Intelligent Technology, Chinese Academy of Sciences)
OptimizationRobotic IntelligenceTransformerPrompt EngineeringMixture of ExpertsScore-based ModelMultimodalityTime SeriesSequential
🎯 What it does: This paper proposes treating instruction drift as sampling error, enhancing the robustness of long-term tasks by using SNR-aware power distribution for adaptive MCMC search during inference.
Rethinking KV Cache Eviction via a Unified Information-Theoretic Objective
Jiaming Yang (Sichuan University), Jiancheng Lv (Sichuan University)
Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelText
🎯 What it does: This paper redefines the eviction problem of KV caches from an information theory perspective, constructs a mutual information maximization objective based on the information bottleneck theory, and proposes the CAPKV algorithm to achieve capacity-aware KV cache eviction.
Rethinking LLM Ensembling from the Perspective of Mixture Models
Jiale Fu (Southeast University), Xu Yang (Southeast University)
Computational EfficiencyKnowledge DistillationTransformerLarge Language ModelMixture of ExpertsText
🎯 What it does: Integrating large language models (LLMs) as a hybrid model, and proposing the Mixture-model-like Ensemble (ME) scheme, which can achieve the same generation quality as traditional ensembles by calling only a single model.
Rethinking Loss Reweighting for Imbalance Learning as an Inverse Problem: A Neural Collapse Point of View
Jinping Wang (Wenzhou-Kean University), Zhiqiang Gao (Wenzhou-Kean University)
ClassificationOptimizationContrastive LearningImage
🎯 What it does: Propose a loss reweighting method based on the inverse problem perspective of the neural collapse (NC) theory, aiming to make the average loss of each class equal, thereby alleviating the imbalance caused by long-tailed distribution;
Rethinking Low-Confidence Pseudo Labels: Influence-Aware Semi-Supervised Fine-Tuning for Hyperspectral Change Detection
Keyun Zhao (Northwestern Polytechnical University), Ying Li (Northwestern Polytechnical University)
SegmentationTransformerSupervised Fine-TuningContrastive LearningImageBenchmark
🎯 What it does: Propose a semi-supervised fine-tuning framework IA-SFT based on influence evaluation, and design an Adaptive Fusion Change Decoder (AFCD) that integrates global semantics and local details, aiming to improve the quality of pseudo-labels and model adaptability in hyperspectral change detection (HSICD).
Rethinking Multimodal Time-Series Forecasting Evaluation
Haoxin Liu (Georgia Institute of Technology), Abhimanyu Das (Google Research)
Data-Centric LearningTransformerLarge Language ModelAgentic AIMixture of ExpertsTextMultimodalityTime SeriesBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposed a new multi-modal time series forecasting benchmark called TimesX, focusing on real-world large-scale cross-domain data and high-quality textual context;
Rethinking Neural Network Learning Rates: A Stackelberg Perspective
Sihan Zeng (JPMorgan AI Research), Sumitra Ganesh (JPMorgan AI Research)
OptimizationReinforcement LearningImageTextTabularBenchmark
🎯 What it does: This paper studies non-uniform learning rates in neural networks, proposing to understand their effectiveness from the perspective of Stackelberg optimization, particularly using smaller learning rates for the lower layers and larger learning rates for the final layer during training.
Rethinking Parameter Sharing as Graph Coloring for Structured Compression
Boyang Zhang (Institute of Computing Technology, Chinese Academy of Sciences), Fangming Liu (Pengcheng Laboratory)
CompressionOptimizationComputational EfficiencyGraph Neural NetworkTransformerLarge Language ModelAuto EncoderContrastive LearningImageText
🎯 What it does: Propose a curvature-aware graph coloring method (CGC), which treats model layers as graph nodes and shared low-rank bases as colors. It utilizes Hessian feature vectors to guide parameter sharing between layers, thereby achieving high compression ratios across layers.
Rethinking Personalization in Large Language Models at the Token Level
Chenheng Zhang, Zhouchen Lin
GenerationExplainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelContrastive LearningTextBenchmark
🎯 What it does: Propose a token-level personalized evaluation and training method based on causal intervention
Rethinking Pretraining Data Detection for LLMs: From Local to Global
Chenye Ke (University of Science and Technology of China), Qi Liu (University of Science and Technology of China)
Anomaly DetectionData-Centric LearningTransformerLarge Language ModelContrastive LearningText
🎯 What it does: Propose a pre-trained data detection framework called AECA from a global sequence perspective, which utilizes self-information calibration and convolutional filtering to capture fluctuation signals in probability sequences;
Rethinking Serialization in Linear 3D Vision: Decoupling Anisotropic Geometry from Isotropic Semantics
YinYun Yan (AnnLab, Institute of Semiconductors, Chinese Academy of Sciences), Xin Ning (AnnLab, Institute of Semiconductors, Chinese Academy of Sciences)
SegmentationTransformerScore-based ModelAuto EncoderContrastive LearningGaussian SplattingPoint Cloud
🎯 What it does: Designed and implemented the AnIsoNet framework, which decouples local directional geometric modeling from global isotropic semantic aggregation, addressing the serialization bias in 3D State-Space models.
Rethinking Sparse Mixture of Experts from a Unified Perspective
Giang Do (Deakin University), Truyen Tran (Deakin University)
OptimizationComputational EfficiencyRepresentation LearningTransformerLarge Language ModelMixture of ExpertsContrastive LearningImageTextMultimodality
🎯 What it does: Propose a unified sparse expert mixture of experts model (USMoE) to address the limitations of traditional Token Choice and Expert Choice.
Rethinking Temporal Consistency in Video Object-Centric Learning: From Prediction to Correspondence
Zhiyuan Li (Aalto University), Joni Pajarinen (Aalto University)
Object DetectionObject TrackingSegmentationRepresentation LearningTransformerContrastive LearningOptical FlowVideo
🎯 What it does: Propose the Grounded Correspondence framework, which utilizes frozen self-supervised visual Transformer features for slot initialization based on significant peak saliency, and employs the Hungarian algorithm for discrete matching between frames, completely without requiring a learned temporal dynamic prediction module.
Rethinking the Design Space of Reinforcement Learning for Diffusion Models: On the Importance of Likelihood Estimation Beyond Loss Design
Jaemoo Choi (Georgia Institute of Technology), Yongxin Chen (Georgia Institute of Technology)
GenerationTransformerReinforcement LearningDiffusion modelScore-based ModelImageStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: This paper systematically analyzes the design space of RL in diffusion models and finds that the likelihood estimation of ELBO based on the final generated samples is key to improving training efficiency and performance.
Rethinking the Flow-based Gradual Domain Adaptation: A Semi-Dual Optimal Transport Perspective
Zhichao Chen (Peking University), Zhouchen Lin (Peking University)
Domain AdaptationFlow-based ModelImageTabular
🎯 What it does: Proposed a flow-based progressive domain adaptation framework called E-SUOT based on semi-dual unbalanced optimal transport, which can generate intermediate domains and achieve step-by-step transfer from source to target without explicitly estimating the target domain probability density.
Rethinking the Hardness of PbRL: A Provable General Regret Bound
Chenjie Mao (Washington University in Saint Louis), Chongjie Zhang (Washington University in Saint Louis)
Reinforcement Learning from Human FeedbackReinforcement Learning
🎯 What it does: Proposes a preference reinforcement learning algorithm called RTPQ based on the value function, which can directly utilize trajectory-level preference feedback and achieve a regret upper bound of square root T under the general function approximation framework.
Rethinking the Reranker: Boundary-Aware Evidence Selection for Robust Retrieval-Augmented Generation
Jiashuo Sun (University of Illinois Urbana Champaign), Jiawei Han (University of Illinois Urbana Champaign)
GenerationRetrievalData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextRetrieval-Augmented Generation
🎯 What it does: Proposes the boundary-aware evidence selection framework BAR-RAG, making retrieval-augmented generation (RAG) more robust in noisy retrieval environments, capable of selecting evidence sets that are neither too easy nor too difficult at the generator's capability boundary.
Rethinking the Trust Region in LLM Reinforcement Learning
Penghui Qi (Sea AI Lab), Wee Sun Lee (National University of Singapore)
Reinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningMixture of ExpertsText
🎯 What it does: This paper rethinks the trust region mechanism of PPO in reinforcement learning for large-scale language models, proposing Divergence Proximal Policy Optimization (DPPO). It replaces the traditional ratio clipping with directly estimating the TV or KL divergence of the policy distribution (using binary and Top-K approximations), thereby more reasonably constraining policy updates.
Rethinking Thinking Tokens: LLMs as Improvement Operators
Lovish Madaan (Meta), Anirudh Goyal (Meta)
Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose a framework that treats LLM reasoning as an improvement operation, designing two reasoning strategies: parallel-distillation-refinement (PDR) and sequential refinement (SR), and training the model using reinforcement learning to better utilize this operation;
Rethinking Time-Series Imputation as Conditional Inference along Temporal Evolution
Yu Fan (Tsinghua University), Pengjun Wang (Tsinghua University)
RestorationData SynthesisAnomaly DetectionTransformerPrompt EngineeringDiffusion modelScore-based ModelAuto EncoderContrastive LearningTabularTime SeriesBenchmark
🎯 What it does: Proposed the Conditional Temporal Inference Paradigm (CTIP) and implemented the Cross-Block Imputer Transformer (CBiT) for missing value imputation in time series with long historical contexts.
Rethinking Video Generation Model for the Embodied World
Yufan Deng (Peking University Shenzhen Graduate School), Daquan Zhou (ByteDance Seed)
Data SynthesisRobotic IntelligenceTransformerLarge Language ModelPrompt EngineeringVision Language ModelVision-Language-Action ModelDiffusion modelVideoTextMultimodalityBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposed the robot video generation evaluation benchmark RBench and a large-scale dataset RoVid-X, and conducted a systematic evaluation of 25 video generation models.
Rethinking Visual Autoregressive Sampling with Information-Grounding Guidance
Ky Dan Nguyen (University of Sydney), Chang Xu (University of Sydney)
GenerationTransformerDiffusion modelScore-based ModelImage
🎯 What it does: This paper proposes a guiding method called IGG for the scale-recursive autoregressive (SwAR) image generation model, aiming to address the issues of scattered guidance signals and semantic mismatch in traditional classifier-free guidance (CFG) within SwAR.
Rethinking Visual Intelligence: Insights from Video Pretraining
Pablo Acuaviva (University of Bern), Paolo Favaro (University of Bern)
RestorationSegmentationGenerationData SynthesisCompressionOptimizationComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringDiffusion modelScore-based ModelContrastive LearningImageVideoTextTabularTime SeriesBenchmark
🎯 What it does: Propose a unified framework that compares pre-trained visual diffusion models (VDM) and large language models (LLM) on the same task through lightweight adapters (LoRA); rewrite image-to-image visual tasks as time conversion in short videos, using video generation priors to solve structured visual problems.
Retrieval-Aware Distillation for Transformer-SSM Hybrids
Aviv Bick (Carnegie Mellon University), Albert Gu (Carnegie Mellon University)
Computational EfficiencyKnowledge DistillationTransformerLarge Language ModelTextRetrieval-Augmented Generation
🎯 What it does: This paper proposes a Retrieval-Aware Distillation method, distilling a pre-trained Transformer model into a Transformer-SSM hybrid model;
Retriever Portfolios: A Principled Approach to Adaptive RAG
Miltiadis Stouras (EPFL), Ola Svensson (EPFL)
RetrievalOptimizationComputational EfficiencyTransformerMixture of ExpertsTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposed a method that automatically selects a small and diverse combination of retrievers (portfolio) from a large pool of candidate retrievers to cover different query distribution regions.
Retro-Expert: Collaborative Reasoning for Interpretable Retrosynthesis
Xinyi Li (Wuhan University), Yu Wu (Wuhan University)
Explainability and InterpretabilityDrug DiscoveryReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextGraphTabularRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Designed an interpretable retrosynthetic prediction framework called Retro-Expert, which leverages LLM and specialized models for collaborative reasoning, and optimizes the generation of natural language explanations for reaction pathways through pure reinforcement learning.
RetrOrchestrator: A Multi-Step Retrosynthesis Agent Dynamically Orchestrating Single-Step Transition Models
Liao Chang (Nanyang Technological University), Ying Wei (Zhejiang University)
Drug DiscoveryReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringMixture of ExpertsTextGraph
🎯 What it does: Proposed a LLM-driven multi-step retrosynthesis planning agent named RetrOrchestrator, which achieves adaptive navigation of the search space by dynamically selecting different single-step retrosynthesis models (SSR), and models the entire process as a partially observable Markov decision process (POMDP)
Return of Frustratingly Easy Unsupervised Video Domain Adaptation
Pengfei Wei (Magellan Technology Research Institute), Lawrence B. Hsieh (Magellan Technology Research Institute)
RecognitionDomain AdaptationTransformerAuto EncoderGenerative Adversarial NetworkContrastive LearningVideo
🎯 What it does: Propose MetaTrans, an unsupervised video domain adaptation framework that uses only source supervision loss and domain adversarial loss, leveraging a time-static subtraction module to achieve separation of spatial and temporal domain differences.
Return-Critic: Bridging Goal Discrepancy for Efficient Visual Reinforcement Learning
Ruyi Lu (China University of Mining and Technology), Yuhu Cheng (China University of Mining and Technology)
Representation LearningConvolutional Neural NetworkTransformerReinforcement LearningContrastive LearningImage
🎯 What it does: Proposed the Return-Critic (RC) auxiliary framework, which uses the return prediction across the entire episode to guide the visual encoder in learning more relevant representations;
Return-to-Go Is More Than a Number: Q-Guided Alignment for Return-Conditioned Supervised Learning
Yuxiao Yang (University of North Carolina at Chapel Hill), Weitong Zhang (University of North Carolina at Chapel Hill)
TransformerReinforcement LearningTabularSequentialBenchmark
🎯 What it does: Proposed a new Q-Guided Alignment Decision Transformer (Q-ALIGN DT), which forces the model to align the monotonic relationship between return-to-go (RTG) and Q-values, enabling the model to generate high-reward trajectories under different RTG conditions;
Reuse your FLOPs: Scaling RL on Hard Problems by Conditioning on Very Off-Policy Prefixes
Amrith Setlur (Carnegie Mellon University), Sang Michael Xie (Meta)
TransformerLarge Language ModelReinforcement LearningPrompt EngineeringText
🎯 What it does: This paper proposes the PrefixRL method, which uses previously generated offline correct prefixes as conditions to guide LLMs in completing reasoning tasks within on-policy reinforcement learning.
Reusing Trajectories in Policy Gradients Enables Fast Convergence
Alessandro Montenegro (Politecnico di Milano), Alberto Maria Metelli (Politecnico di Milano)
Reinforcement Learning
🎯 What it does: Proposed the RT-PG algorithm, which achieves fast convergence based on policy gradient by reusing trajectories collected from the recent ω iterations, combined with importance sampling estimation corrected by multiple power mean averaging;
Revealing Behavioral Plasticity in Large Language Models: A Token-Conditional Perspective
Liyuan Mao (Qwen Team, Alibaba Group), Junyang Lin (Qwen Team, Alibaba Group)
Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This paper explores the behavioral plasticity of large language models and proposes the ToCoRL framework, which combines token-conditional generation with reinforcement learning to achieve behavioral control of the model during inference.
Revealing Long-context Potential of Attention Heads via Frequency Kernels
Senyu Han (Shanghai Jiao Tong University), Lu Chen (Shanghai Jiao Tong University)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation
🎯 What it does: Studied the frequency kernel characteristics of Transformer self-attention heads, proposed Long-context Potential Score (LPS) based on frequency kernel to statically evaluate the potential of heads in handling long contexts, and further improved the performance of long-context tasks by enhancing low-frequency kernels.
Revealing Scaling Paradox in Large-scale Time Series Models: Implications for More Efficient and Accurate Forecasting
Xin Qiu (Ningbo Institute of Digital Twin, Eastern Institute of Technology), Xiaoyu Shen (Ningbo Institute of Digital Twin, Eastern Institute of Technology)
Anomaly DetectionComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelTabularTime SeriesBenchmarkFinance RelatedPhysics Related
🎯 What it does: Investigate the scale paradox of large time series models, reveal and quantify that larger models are not necessarily better, and propose a pruning method that retains only key layers;
RevealLayer: Disentangling Hidden and Visible Layers via Occlusion-Aware Image Decomposition
Binhao Wang (Wenzhou University), Yuhui Yin (360 AI Research)
Image TranslationRestorationSegmentationGenerationTransformerPrompt EngineeringDiffusion modelAuto EncoderImageMultimodalityBenchmark
🎯 What it does: Propose the RevealLayer framework, which decomposes natural images into background and multiple RGBA foreground layers under user-provided bounding boxes, achieving controllable layer decomposition and occlusion completion.
Revenue Guarantees of No-Swap-Regret Dynamics in First Price Auctions
Anders Bo Ipsen (Aarhus University), Stratis Skoulakis (Aarhus University)
OptimizationReinforcement LearningTabularFinance Related
🎯 What it does: Study the revenue guarantee of approximate correlated equilibrium in discrete first-price auctions, and provide a polynomial convergence rate for non-switching regret dynamics.
Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies
Zeyang Li (Massachusetts Institute of Technology), Navid Azizan (Massachusetts Institute of Technology)
Reinforcement LearningDiffusion modelScore-based ModelFlow-based ModelTabularTime SeriesStochastic Differential Equation
🎯 What it does: Proposes the Reverse Flow Matching (RFM) framework for training diffusion and flow-based policies in online reinforcement learning to approximate the Boltzmann distribution.
Reverse-Engineering Model Editing on Language Models
Zhiyu Sun (Shanghai Qi Zhi Institute), Tianxing He (Tsinghua University)
Explainability and InterpretabilityAdversarial AttackTransformerLarge Language ModelPrompt EngineeringTextBenchmark
🎯 What it does: Proposed a reverse engineering attack (KSTER) targeting the locate-then-edit editing paradigm for LLMs, along with corresponding defense strategies (subspace camouflage).
REVIS: Sparse Latent Steering to Mitigate Object Hallucination in Large Vision-Language Models
Jialin Wu (Ant Group), Zhou Yang (Ant Group)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodality
🎯 What it does: Proposes REVIS, a training-free sparse latent modulation framework that uses orthogonal projection to separate pure visual information vectors from deep activations of vision-language models, and dynamically injects this vector into the automatically located optimal layer to alleviate object hallucination;
Revisiting Anisotropy in Language Transformers: The Geometry of Learning Dynamics
Raphael Bernas (CentraleSupelec, Universite ParisSaclay), CELINE HUDELOT
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelAuto EncoderContrastive LearningText
🎯 What it does: This paper explores the mechanism of anisotropy formation in text representations during Transformer training through differential geometry and frequency bias analysis, and evaluates gradient energy and anisotropy by using activation concept subspaces during training.
Revisiting Asymmetries in Black-box Link Stealing against Graph Neural Networks
Paul Agbaje (University of Texas at Arlington), Habeeb Olufowobi (University of Texas at Arlington)
Safty and PrivacyAdversarial AttackGraph Neural NetworkScore-based ModelAuto EncoderContrastive LearningGraph
🎯 What it does: This paper re-evaluates attack methods that only utilize posterior probabilities for link stealing attacks on black-box GNN models, and deeply investigates their tail reliability and internal geometric bottlenecks.
Revisiting Coding-Based Approaches to Overcome the Curse of Dimensionality in Learning-Based Watermarking
Yupeng Qiu (National University of Singapore), Ee-Chien Chang (National University of Singapore)
Safty and PrivacyConvolutional Neural NetworkFlow-based ModelImage
🎯 What it does: Propose the OrthoMark framework, which separates robust feature extraction from encoding-based watermark embedding. It learns a noise-invariant feature space through a reversible network and uses orthogonal projection with QIM for high-capacity watermark encoding and decoding in this space.
Revisiting Distribution Correction Estimation for Offline Imitation Learning with Suboptimal Dataset
Quang Anh Pham (Singapore Management University), Akshat Kumar (Singapore Management University)
Adversarial AttackReinforcement LearningContrastive LearningTabularTime SeriesSequentialBenchmark
🎯 What it does: This paper proposes an offline imitation learning method called ReDICE, which uses a distribution correction based on a mixture of suboptimal data and expert data to eliminate the strict assumption of expert coverage in traditional DICE methods, while achieving stable dual optimization through Gumbel regression and introducing a τ-scaled policy extraction mechanism.
Revisiting Efficiency–Accuracy Scaling in Mixture-of-Experts Architectures
Venmugil Elango (NVIDIA Corporation), Bita Darvish Rouhani (NVIDIA Corporation)
Computational EfficiencyAI Code AssistantTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsTextRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This paper proposes the LatentMoE architecture, which reduces expert computation and communication overhead in the MoE framework by projecting inputs into a low-dimensional latent space, and enhances model expressiveness by increasing the number of experts and the topK value.
Revisiting ML Training under Fully Homomorphic Encryption: Convergence Guarantees, Differential Privacy, and Efficient Algorithms
Yvonne Zhou (University of Maryland), Dana Dachman-Soled (University of Maryland)
Federated LearningSafty and PrivacyComputational EfficiencyImageTabular
🎯 What it does: Theoretical convergence analysis of machine learning training is conducted under fully homomorphic encryption (FHE), and a differential privacy gradient descent algorithm that does not require gradient clipping is designed.
Revisiting Neural Processes via Fourier Transform and Volterra Series
Peiman Mohseni (Texas A&M University), Raymond K. W. Wong (Texas A&M University)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningConvolutional Neural NetworkTransformerImagePoint CloudTabularTime Series
🎯 What it does: Proposed a new framework for constructing translation equivariant conditional neural processes (CNP) through convolution and Fourier transform;
Revisiting OOD Generalization in Programmatic RL
Amirhossein Rajabpour (University of Alberta), Levi Lelis
Convolutional Neural NetworkRecurrent Neural NetworkReinforcement LearningPrompt EngineeringContrastive LearningSequentialBenchmark
🎯 What it does: This paper re-evaluates the gap between procedural strategies and neural networks in generalization to out-of-distribution (OOD) discrete distributions, finding that the gap mainly stems from experimental settings rather than representational differences, and demonstrates that by adjusting the reward function, observation sparsification, and other modifications, neural networks can achieve comparable generalization performance to procedural strategies;
Revisiting Padded Transformer Expressivity: Which Architectural Choices Matter and Which Don't
Anej Svete (ETH Zürich), Ashish Sabharwal (Allen Institute for AI)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerSupervised Fine-TuningAuto EncoderGenerative Adversarial NetworkContrastive Learning
🎯 What it does: Conduct a comprehensive theoretical analysis of the expressive power of padded Transformers, clarifying the impact of design choices such as attention type, model width, numerical precision, looping, and uniformity (L-uniform versus fully uniform) on expressiveness, and providing precise equivalence results.
Revisiting Parameter-Based Knowledge Editing in Large Language Models: Theoretical Limits and Empirical Evidence
Wanying Ren (East China Normal University), Aixin Sun (Nanyang Technological University)
Federated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringContrastive LearningTextRetrieval-Augmented Generation
🎯 What it does: This paper studies methods for knowledge editing in large language models through local parameter modification, analyzing their effectiveness and safety in real-world applications.
Revisiting Photometric Ambiguity for Accurate Gaussian-Splatting Surface Reconstruction
Jiahe Li (Beihang University), Gim Hee Lee (National University of Singapore)
RestorationDepth EstimationSuper ResolutionNeural Radiance FieldContrastive LearningGaussian SplattingImagePoint Cloud
🎯 What it does: Propose the AmbiSuR framework to address the photometric ambiguity problem that occurs during the Gaussian Splatting process, thereby improving the accuracy of surface reconstruction.
Revisiting Positive Samples in Graph Contrastive Learning: From the Perspective of Message Passing
Lianze Shan (Tianjin University), Dongxiao He (Tianjin University)
Representation LearningGraph Neural NetworkContrastive LearningGraph
🎯 What it does: Studied the role of positive samples in graph contrastive learning, and proposed the SPGCL method, which separates features using Dirichlet energy and improves propagation and positive sample sampling.
Revisiting Pre-Propagation GNNs: Robust Diffusion Operators and Hidden-State Re-Propagation
Zichao Yue (Cornell University), Zhiru Zhang (Cornell University)
Computational EfficiencyRepresentation LearningRecurrent Neural NetworkGraph Neural NetworkDiffusion modelNeural Radiance FieldAuto EncoderContrastive LearningOptical FlowGraphStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: This study addresses the accuracy bottleneck of pre-propagation graph neural networks (PP-GNN), proposing a more robust diffusion operator and a hidden state re-propagation mechanism to enhance performance on both heterogeneous and homogeneous graphs.
Revisiting Regularized Policy Optimization for Stable and Efficient Reinforcement Learning in Two-Player Games
Kazuki Ota (University of Tokyo), Tatsuya Harada (University of Tokyo)
Convolutional Neural NetworkRecurrent Neural NetworkTransformerReinforcement LearningGraphTabular
🎯 What it does: Proposed and analyzed a strategy optimization method using reverse KL regularization and entropy regularization (KLENT) in two-player zero-sum games, and implemented a reinforcement learning algorithm without search and pure model
Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control
Yonghui Yang (National University of Singapore), Tat-Seng Chua (National University of Singapore)
OptimizationSafty and PrivacyReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningTextBenchmark
🎯 What it does: Propose a robust safe alignment framework called ShaPO based on selective geometric control, addressing the vulnerability caused by optimization geometry during the preference optimization process.
Revisiting Spectral Representations in Generative Diffusion Models
Yuehao Wang (University of Texas at Austin), Zhangyang Wang (University of Texas at Austin)
GenerationData SynthesisTransformerDiffusion modelScore-based ModelRectified FlowContrastive LearningImagePoint Cloud
🎯 What it does: Propose a regularized method based on self-supervised spectral representation alignment, enabling diffusion models to adaptively learn time-varying spectral embeddings during training, thereby improving generation quality.
Revisiting the Bertrand Paradox via Equilibrium Analysis of No-regret Learners
Arnab Maiti (University of Washington), Lillian J. Ratliff (University of Washington)
OptimizationReinforcement LearningFinance Related
🎯 What it does: Studied the discrete Bertrand pricing game with non-decreasing demand functions, analyzing the equilibrium outcomes that may arise under different no-regret learning guarantees.
Revisiting the Platonic Representation Hypothesis: An Aristotelian View
Fabian Gröger (EPFL), Maria Brbic (EPFL)
Representation LearningTransformerAgentic AIPrompt EngineeringContrastive LearningImageVideoTextMultimodalityReview/Survey PaperRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This paper re-examines the Plato representation hypothesis, proposing a permutation-based null calibration framework that corrects the influence of network width and depth on representation similarity, and based on the calibration results, proposes the Aristotle representation hypothesis.
Revisiting the Role of Pretrained Weights in Model Merging: On Near-Optimality within the Core Subspace
Wenju Sun (Beijing Jiaotong University), Boyang Li
ClassificationImage TranslationRecommendation SystemOptimizationKnowledge DistillationRepresentation LearningTransformerContrastive LearningImageTextMultimodality
🎯 What it does: Explore the role of pre-trained weights in model merging and propose a strategy to maintain the core subspace during merging.
Revisiting the Volume Hypothesis
Ari Pakman (BenGurion University of Negev), Yakir Berchenko (BenGurion University of Negev)
ClassificationConvolutional Neural NetworkImage
🎯 What it does: This paper uses the Replica Exchange Wang-Landau algorithm to estimate the joint density of training/test accuracy under different numbers of training samples for binary networks, re-examining the 'volume hypothesis'.
Revisiting Uncertainty: On Evidential Learning for Partially Relevant Video Retrieval
Jun Li (Tsinghua University), Shu-Tao Xia (Tsinghua University)
RetrievalTransformerContrastive LearningVideoTextMultimodality
🎯 What it does: Designed and implemented a hierarchical evidence learning framework called Holmes to address query uncertainty in partial relevance video retrieval (PRVR) and the sparse supervision problem in multi-instance learning.
Revisiting Zeroth-Order Hessian Approximation: A Single-Step Policy Optimization Lens
Junbin Qiu (Hong Kong University Of Science And Technology (Guangzhou)), Yao Shu (Hong Kong University Of Science And Technology (Guangzhou))
OptimizationAdversarial AttackSupervised Fine-TuningReinforcement LearningGaussian SplattingImageTextTabular
🎯 What it does: The ZoVH framework is proposed by reformulating the zeroth-order Hessian estimation as the Hessian of single-step policy optimization, designing a low-variance Hessian estimator with average baseline optimality and reusable queries, and further constructing curvature-aware zeroth-order optimizer for the invertible Hessian and its gradient product.
ReViT: Rotational-equivariant Vision Transformers for Neural PDE Solvers
Hao Wei (Technical University of Munich), Nils Thuerey (Technical University of Munich)
OptimizationTransformerVision-Language-Action ModelContrastive LearningImagePoint CloudPhysics RelatedStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Proposed ReViT, a visual Transformer capable of achieving rotational equivariance on lattice physical fields;
REViT: Roto-reflection Equivariant Convolutional Vision Transformer
Sheir A. Zaheer (KC Machine Learning Lab), Chan Y. Park (KC Machine Learning Lab)
ClassificationRecognitionConvolutional Neural NetworkTransformerVision-Language-Action ModelContrastive LearningImage
🎯 What it does: Proposed a visual Transformer with discrete rotation-reflection group equivariance (REViT), achieving equivariance without position encoding through convolutional self-attention;
Reviving Error Correction in Modern Deep Time-Series Forecasting
Minh Hoang Nguyen (Deakin University), Hung Le (Deakin University)
Computational EfficiencyTransformerTabularTime SeriesBenchmark
🎯 What it does: Propose a post-error correction framework (UEC-STD), which can recursively correct the errors in deep time series prediction without retraining any frontend predictor, significantly improving the accuracy of long-term predictions.
ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning
Yiming Zhang (Simon Fraser University), Angel X Chang
Prompt EngineeringVision Language ModelContrastive LearningImageVideoPoint CloudBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Constructed a more reliable visual spatial intelligence evaluation benchmark, ReVSI, by re-annotating and calibrating 3D objects and geometry in videos, generating question-answer pairs under an observable framework, and designing evaluation and visualization diagnosis based on a frame budget.
Reward and Guidance through Rubrics: Promoting Exploration to Improve Multi-Domain Reasoning
Baolong Bi (State Key Laboratory of AI Safety, Institute of Computing Technology, CAS), Xueqi Cheng (State Key Laboratory of AI Safety, Institute of Computing Technology, CAS)
OptimizationExplainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkPhysics RelatedRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This paper proposes a reinforcement learning framework called RGR-GRPO based on Rubrics, aiming to enhance the exploration and performance of large language models in multi-domain reasoning tasks.
Reward Auditor: Inference on Reward Modeling Suitability in Real-World Perturbed Scenarios
Jianxiang Zang (Fudan University), Hui Liu (Shanghai University of International Business and Economics)
Reinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningContrastive LearningTextBenchmark
🎯 What it does: Proposes the Reward Auditor framework, which infers the applicability of reward models and quantifies the severity of vulnerabilities by conducting paired hypothesis testing on the confidence levels of preferences under various real-world perturbations.
Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use
Kunvar Thaman (Independent Researcher)
TransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposed and implemented the Reward Hacking Benchmark (RHB), a benchmark for evaluating language model agents that use multi-step tools, aimed at quantifying reward hacking behaviors;
Reward Learning through Ranking Mean Squared Error
Chaitanya Kharyal (University of Alberta), Matthew E. Taylor (University of Alberta)
OptimizationReinforcement Learning from Human FeedbackReinforcement LearningContrastive LearningTextTabular
🎯 What it does: Propose a reinforcement learning method called R4 based on multi-class human ratings, which learns the reward function using ranking mean squared error (rMSE) loss.
Reward Modeling from Natural Language Human Feedback
Zongqi Wang (Tsinghua University), Yongbin Li (Qwen-Character, Tongyi Lab, Alibaba)
Reinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringText
🎯 What it does: Proposes using natural language human feedback process rewards to improve generative reward models (RM-NLHF), and designs an scalable online MetaRM to predict missing process rewards.
Reward Redistribution for CVaR MDPs using a Bellman Operator on L-infinity
Aneri Muni (Universite De Montreal), Erick Delage (Mila - Quebec Ai Institute)
Reinforcement LearningTabular
🎯 What it does: Study the Bellman operator for static CVaR MDP, propose a reward redistribution method, and provide value iteration and Q-learning algorithms.
Reward Shaping Control Variates for Off-Policy Evaluation Under Sparse Rewards
Ritam Majumdar (Imperial College London), Sonali Parbhoo (Imperial College London)
Recurrent Neural NetworkTransformerReinforcement LearningAuto EncoderContrastive LearningTabularTime SeriesSequentialBiomedical DataBenchmarkFinance Related
🎯 What it does: In sparse reward environments, a class of control variables constructed using the potential underlying reward shape is proposed for offline policy evaluation (OPE), significantly reducing estimation variance while maintaining unbiasedness.
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective
Haichuan Wang (Harvard University), Milind Tambe (Harvard University)
OptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmark
🎯 What it does: The study addresses the problem of designing a reward model for alignment during inference of large language models under KL regularization constraints, proposing to view the optimization of the reward model as a Stackelberg leader problem and providing a threshold reward shaping scheme.
Reward Under Attack: Analyzing the Robustness and Hackability of Process Reward Models
Rishabh Tiwari (University Of California Berkeley), Amir Gholami (University Of California Berkeley)
Adversarial AttackReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringScore-based ModelTextBenchmark
🎯 What it does: This paper proposes a three-tier diagnostic framework to conduct robustness evaluation of process reward models (PRM), and reveals vulnerabilities that PRMs are susceptible to through three attack methods: static perturbation, gradient attacks, and reinforcement learning reward hacking. It also publicly releases PRM-BiasBench and evaluation tools.
Reward-free Alignment for Conflicting Objectives
Peter Chen (Columbia University), Tianyi Lin (Columbia University)
OptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningContrastive LearningTextBenchmark
🎯 What it does: Propose a reward-free alignment framework called RACO, which directly utilizes pairwise preference data and addresses gradient conflicts in LLM fine-tuning through a clipped version of conflict-avoiding gradient descent (CAGrad-Clip).
Reward-Preserving Counterfactual State Editing for Offline Reinforcement Learning
Siyu Wang (Macquarie University), Lina Yao (University of New South Wales)
Data SynthesisTransformerReinforcement LearningDiffusion modelAuto EncoderContrastive LearningTabularTime SeriesSequential
🎯 What it does: In strict offline reinforcement learning, the authors propose the CSET (Counterfactual State Editing Transformer) method, which uses reward-preserving counterfactual state editing to augment training data, and combines a causality-guided hybrid Transformer architecture to enhance the robustness of offline policies.
Rewiring Experts on the Fly: Continuous Rerouting for Better Online Adaptation in Mixture-of-Expert Models
Guinan Su (Max Planck Institute for Intelligent Systems & ELLIS Institute Tubingen & Tubingen AI Center), Jonas Geiping (Max Planck Institute for Intelligent Systems & ELLIS Institute Tubingen & Tubingen AI Center)
OptimizationComputational EfficiencyAI Code AssistantTransformerSupervised Fine-TuningPrompt EngineeringMixture of ExpertsText
🎯 What it does: Propose a data-agnostic, online testing rerouting method that improves inference performance by self-supervised optimization of expert selection in MoE models.
Rex: A Family of Reversible Exponential (Stochastic) Runge-Kutta Solvers
Zander W. Blasingame (AITHYRA), Chen Liu (Clarkson University)
GenerationData SynthesisDiffusion modelImageTextBiomedical DataStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: A set of reversible exponential (stochastic) Runge–Kutta solvers, named Rex, is proposed to achieve high-order accurate inversion in the ODE/SDE of diffusion models.