arXivSub Start free trial

ICML 2026 Papers — Page 47

International Conference on Machine Learning · 6554 papers

Quaternion Self-Attention with Shared Scores

Shogo Yamauchi (Asahi Shimbun Company), Hideaki Tamori (Asahi Shimbun Company)

Computational EfficiencyRepresentation LearningTransformerImageTextAudio

🎯 What it does: Proposed a quaternion self-attention mechanism with shared scores

QUATRO: Query-Adaptive Trust Region Policy Optimization for LLM Fine-tuning

Doyeon Lee (Seoul National University), Jaemoo Choi (Seoul National University)

OptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningTextBenchmark

🎯 What it does: Proposes a query-adaptive trust region policy optimization (QUATRO) for fine-tuning large language models, explicitly constraining KL divergence for each query to eliminate heuristic importance sampling trimming, ensuring stable updates while maintaining entropy diversity.

Query Circuits: Explaining How Language Models Answer User Prompts

Tung-Yu Wu (University of Oxford), Fazl Barez (University of Oxford)

Explainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringAuto EncoderText

🎯 What it does: Proposes the Query Circuit discovery task, which directly tracks the computational flow of individual prompts within large language models and provides corresponding explanation methods.

Query Lens: Interpreting Sparse Key-Value Features with Indirect Effects

Hwiyeong Lee (Hanyang University), Taeuk Kim (Hanyang University)

Explainability and InterpretabilityTransformerLarge Language ModelAuto EncoderText

🎯 What it does: This work proposes the Query Lens method to explain the causal relationships between input and output for sparse key-value features;

Query-Based Asymmetric Modeling with Decoupled Input–Output Rates for Speech Restoration

Ui-Hyeop Shin (Sogang University), Hyung-Min Park (Sogang University)

RestorationTransformerAuto EncoderContrastive LearningAudio

🎯 What it does: Propose a query-driven asymmetric time-frequency Transformer framework, TF-Restormer, for speech restoration (including noise reduction, echo removal, bandwidth expansion, etc.) under mismatched input and output sampling rates.

Query-efficient model evaluation using cached responses

Hayden Helm (Helivan), Carey Priebe (Johns Hopkins University)

Computational EfficiencyData-Centric LearningMixture of ExpertsTextBenchmark

🎯 What it does: This study proposes a query-efficient model evaluation method based on cached responses, which can predict the scores of new models on full benchmarks by constructing a data kernel perspective space (DKPS) on the cached responses of existing models.

Questioning the Coverage-Length Metric in Conformal Prediction: When Shorter Intervals Are Not Better

Yizhou Min (Shanghai University of Finance and Economics), Jiaye Teng (Shanghai University of Finance and Economics)

Anomaly DetectionExplainability and InterpretabilityComputational EfficiencyTextTabular

🎯 What it does: This study investigates the limitations of traditional coverage-length evaluation metrics in Conformal Prediction (CP) and proposes a Prejudicial Trick (PT) method that 'tricks' the system into producing shorter prediction intervals through randomization. Subsequently, the paper introduces an Interval Stability metric to detect such issues.

QuITE: Query-Based Irregular Time Series Embedding

Junghoon Lim (SK Shieldus)

ClassificationAnomaly DetectionRepresentation LearningTransformerAuto EncoderContrastive LearningTabularTime SeriesBiomedical DataElectronic Health RecordsBenchmark

🎯 What it does: A query-based input embedding module called QuITE is proposed to handle irregular multivariate time series (IMTS), enabling existing regular multivariate time series models to directly process IMTS.

R-Diverse: Mitigating Diversity Illusion in Self-Play LLM Training

Gengsheng Li (Foundation Model Research Center Institute of Automation Chinese Academy of Sciences), Jinqiao Wang (Foundation Model Research Center Institute of Automation Chinese Academy of Sciences)

OptimizationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose the R-Diverse framework, which enhances the reasoning ability of LLMs through self-play loops.

R$^3$L: Reasoning 3D Layouts from Relative Spatial Relations

Zhifeng Gu (Hong Kong Polytechnic University), Bing WANG

OptimizationTransformerLarge Language ModelPrompt EngineeringDiffusion modelScore-based ModelTextMultimodalityPoint CloudMeshRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Leverage multi-modal large language models (MLLM) to reason about relative spatial relationships in 3D layouts and convert them into executable layouts; improve the reliability and consistency of reasoning through three major techniques; further optimize the layout to ensure physical feasibility.

R1-SyntheticVL: Is Synthetic Data from Generative Models Ready for Multimodal Large Language Model?

Jingyi Zhang (Hong Kong Polytechnic University), Jiaxing Huang (Hong Kong Polytechnic University)

Data SynthesisOptimizationRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsVision Language ModelGenerative Adversarial NetworkImageTextMultimodalityRetrieval-Augmented Generation

🎯 What it does: Proposes the Collective Adversarial Data Synthesis (CADS) method for automatically generating high-quality, diverse, and challenging multimodal training data, and trains the R1-SyntheticVL model based on this.

R2-Router: A New Paradigm for LLM Routing with Reasoning

Jiaqi Xue (University of Central Florida), Heng Huang (University of Maryland)

OptimizationComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose R2-ROUTER, which improves LLM routing by predicting the quality-cost curve for each model under different output length budgets through reasoning, and selecting the optimal model and budget based on this; meanwhile, construct the R2-BENCH dataset to record model performance across multiple lengths;

R2R2: Robust Representation for Intensive Experience Reuse via Redundancy Reduction in Self-Predictive Learning

Sanghyeob Song (Seoul National University), Sungroh Yoon (Seoul National University)

Reinforcement LearningContrastive LearningTabularSequential

🎯 What it does: Propose the R2R2 regularization method, which stabilizes representations by non-centralized redundancy reduction in self-predictive learning (SPL), improving sample efficiency in high UTD environments.

RA-Det: Towards Universal Detection of AI-Generated Images via Robustness Asymmetry

Xinchang Wang (Jiangnan University), Hui Li (Jiangnan University)

Image TranslationRestorationGenerationAnomaly DetectionTransformerPrompt EngineeringDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningImage

🎯 What it does: Propose a behavior-driven detection framework RA-Det based on robustness asymmetry, which amplifies the drift differences between real and synthetic images in the feature space by applying small semantic-preserving perturbations to images, and discriminates through a multi-branch network that aggregates semantic, difference, and low-level residual signals.

RA-VLA: Retrieval-Augmented VLA for Test-Time Adaptation

Sanghwan Jang (POSTECH), Hwanjo Yu (POSTECH)

Domain AdaptationRobotic IntelligenceTransformerVision-Language-Action ModelDiffusion modelContrastive LearningImageTextMultimodalityRetrieval-Augmented Generation

🎯 What it does: RA-VLA proposes a retrieval-enhanced Vision-Language-Action framework, achieving test-time adaptation without weight updates, enabling action generation on new tasks through a few expert demonstrations.

RaBiT: Residual Aware Binarization Training for Accurate and Efficient LLMs

Youngcheon You (Samsung Research), Dongkyu Kim (Samsung Research)

CompressionComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelText

🎯 What it does: Propose a method for high-precision, low-cost LLM compression under 2-bit extreme quantization by utilizing the residual binary training framework RaBiT, addressing the co-adaptation problem in parallel paths.

RaBitQCache: Rotated Binary Quantization for KVCache in Long Context LLM Inference

Wenhao Li (Renmin University of China), Xiaoyong Du (Renmin University of China)

Computational EfficiencyTransformerLarge Language ModelText

🎯 What it does: Propose RaBitQCache, a sparse attention framework for KVCache achieved through random rotation binary quantization;

RACER: Risk-Aware Calibrated Efficient Routing for Large Language Models

Sai Hao (Southern University of Science and Technology), Bingyi Jing (Chinese University of Hong Kong Shenzhen)

Federated LearningSafty and PrivacyExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose a post-training, risk-aware, calibration-efficient routing method called RACER, which transforms traditional single-model routing into confidence-based ensemble model routing, and improves the final answer quality through aggregation.

RAD: Retrieval High-quality Demonstrations to Enhance Decision-making

Lu Guo (Jilin University), Yi Chang (Jilin University)

Data SynthesisOptimizationTransformerReinforcement LearningDiffusion modelTabularTime SeriesSequentialRetrieval-Augmented Generation

🎯 What it does: This paper proposes an offline reinforcement learning method called RAD, which uses a retrieval mechanism to dynamically obtain high-reward and reachable target states, and combines diffusion models to generate sub-trajectories to improve decision-making performance;

RADAR: Defending RAG Dynamically against Retrieval Corruption

Ziyuan Chen (Nanjing University), Tieniu Tan (Nanjing University)

RetrievalAdversarial AttackData-Centric LearningGraph Neural NetworkTransformerPrompt EngineeringTextTime SeriesRetrieval-Augmented Generation

🎯 What it does: Proposed the RADAR framework for dynamically defending against retrieval corruption in RAG systems.

RADAR: Redundancy-Aware Diffusion for Multi-Agent Communication Structure Generation

Zhen Zhang (Nanjing University), Wei Ji (Nanjing University)

GenerationOptimizationComputational EfficiencyAI Code AssistantGraph Neural NetworkTransformerReinforcement LearningPrompt EngineeringDiffusion modelScore-based ModelGenerative Adversarial NetworkTextGraphBenchmark

🎯 What it does: Propose a multi-agent communication topology generation framework called RADAR based on iterative redundant-aware diffusion, which can automatically construct adaptive collaboration graphs given specific tasks, significantly improving collaborative performance and reducing token consumption.

RADE: Random Add-Drop Edge as a Regularizer

Danial Saber (Ontario Tech University), Amirali Salehi-Abari (Ontario Tech University)

ClassificationGraph Neural NetworkContrastive LearningGraph

🎯 What it does: A random edge addition and deletion graph data augmentation method called RADE is designed to uniformly address the overfitting and overcompression issues in GNNs.

Radial Scaling Voxelization for Accurate Small Object 3D Detection

Hao Liu (East China Normal University), Yanni Ma (Sun Yat-sen University)

Object DetectionAutonomous DrivingDiffusion modelScore-based ModelFlow-based ModelRectified FlowNeural Radiance FieldAuto EncoderGenerative Adversarial NetworkContrastive LearningPoint Cloud

🎯 What it does: Propose the Radial Scaling Voxelization (RSV) scheme, which applies a continuous radial scaling function to transform point cloud coordinates, replacing traditional uniform voxelization to achieve non-uniform discretization that maintains fine-grained resolution at close distances and preserves original resolution at far distances;

RADIO1D: Elastic Representations for Condensed Vision Modeling

Greg Heinrich (NVIDIA), Pavlo Molchanov (NVIDIA)

RetrievalCompressionKnowledge DistillationRepresentation LearningTransformerAuto EncoderContrastive LearningImageMultimodality

🎯 What it does: Propose RADIO1D, a visual encoder that compresses images into variable-length 1D token sequences through multi-teacher distillation and autoencoder design, and integrate it into VLM to support efficient scene understanding and retrieval.

RAG without Forgetting: Continual Query-Infused Key Memory

Yuntong Hu (Emory University), Liang Zhao (Emory University)

RetrievalTransformerTextRetrieval-Augmented Generation

🎯 What it does: Propose an untrained, continuously updated retrieval-augmented generation framework called Evolving Retrieval Memory (ERM), which converts extended information during queries into persistent document key updates;

RAGEN-2: Reasoning Collapse in Agentic RL

Zihan Wang, Manling Li (Northwestern)

Explainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextSequentialRetrieval-Augmented Generation

🎯 What it does: Studied the template collapse problem that arises in multi-turn LLM agents during reinforcement learning training, and proposed MI agent diagnostics and SNR-aware filtering to alleviate it.

RaGEP: Rank-aware Geometric Expert Pruning for Mixture-of-Experts Language Models

Wentao Hu (Xi'an Jiaotong University), Jiayin Wang (Xi'an Jiaotong University)

Computational EfficiencyKnowledge DistillationTransformerLarge Language ModelMixture of ExpertsTextRetrieval-Augmented Generation

🎯 What it does: Post-training pruning of large sparse Mixture-of-Experts (MoE) language models, achieving structural compression by leveraging the geometric characteristics of expert activations.

RAIGen: Rare Attribute Identification in Text-to-Image Generative Models

Silpa Vadakkeeveetil Sreelatha (University Of Surrey), Anjan Dutta (University Of Surrey)

GenerationExplainability and InterpretabilityTransformerPrompt EngineeringDiffusion modelAuto EncoderContrastive LearningImageTextMultimodality

🎯 What it does: Propose the RAIGen framework, which utilizes the Matryoshka sparse autoencoder to discover rare attributes in the internal representations of text-to-image diffusion models.

RAMAC: Multimodal Risk-Aware Offline Reinforcement Learning and the Role of Behavior Regularization

Kai Fukazawa (University of California), Iman Soltani (University of California)

Reinforcement LearningDiffusion modelScore-based ModelFlow-based ModelMultimodality

🎯 What it does: Proposes the RAMAC framework, which combines an expression generation actor with a distributed critic, utilizing a single objective to optimize behavioral cloning and CVaR, achieving multi-modal risk awareness in offline RL.

Ramba: Selective State-Space Models for Relational Deep Learning

Yiming Liu (Renmin University of China), Yueguo Chen (Renmin University of China)

Recommendation SystemComputational EfficiencyData-Centric LearningGraph Neural NetworkTransformerPrompt EngineeringMixture of ExpertsContrastive LearningTabularBenchmark

🎯 What it does: Propose a new selective state space model called Ramba, specifically designed for relational databases;

Random Process Flow Matching: Generative Implicit Representations of Multivariate Random Fields

Julien Lalanne (Navier, CNRS, Univ Gustave Eiffel, ENPC, IP Paris), Jean-Michel Pereira (Navier, CNRS, Univ Gustave Eiffel, ENPC, IP Paris)

RestorationGenerationSuper ResolutionConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelFlow-based ModelImagePoint CloudTabularPhysics Related

🎯 What it does: Propose RP Flow, a generative framework based on flow matching, which utilizes implicit neural representations to map sparse observations of a random field from a Gaussian source process to a target random field, achieving interpolation and uncertainty quantification.

Random Scaling of Emergent Capabilities

Rosie Zhao (Harvard University), Naomi Saphra (Harvard University)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningData-Centric LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningText

🎯 What it does: Analyze the scalability of language models across different scales, pointing out that breakthrough performance originates from distribution bimodality caused by random seeds rather than a single threshold

Random Selection Reveals Implicit Knowledge Consensus in Code Generation

Ren-Biao Liu (Nanjing University), Ming Li (Nanjing University)

Explainability and InterpretabilityComputational EfficiencyData-Centric LearningAI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextSequentialRetrieval-Augmented Generation

🎯 What it does: Systematically evaluate the effectiveness of different selection strategies in multi-solution code generation training data, finding that random sampling performs stably across various models and representation spaces without additional computational cost; by comparing multiple complex selection methods (K-Center, Facility Location, K-Means, AST Coverage, Kernel Herding, IFD Ranking) with random sampling, propose an 'implicit knowledge consensus' perspective to explain the effectiveness of random sampling; and analyze when random sampling can be surpassed under different difficulty and constraint scenarios.

Randomized Advantage Transformation (RAT): Computing Natural Policy Gradients via Direct Backpropagation

Mingfei Sun (University of Manchester)

OptimizationReinforcement LearningContrastive LearningTabularTime Series

🎯 What it does: Propose Randomized Advantage Transformation (RAT), which utilizes the Woodbury formula to transform the Tikhonov-regularized natural policy gradient into an advantage function transformation, and directly estimates this gradient via backpropagation on mini-batch samples using the stochastic block Kaczmarz iteration.

Randomized Feasibility Methods for Constrained Optimization with Adaptive Step Sizes

Abhishek Chakraborty (Arizona State University), Angelia Nedich (Arizona State University)

OptimizationImageTextTabular

🎯 What it does: Developed a method combining random feasibility algorithms with adaptive step sizes to solve multi-constrained convex optimization problems; achieves linear convergence for strongly convex smooth objectives and parameter-free O(1/√T) convergence for convex non-smooth objectives.

Rank-Aware Spectral Bounds on Attention Logits for Stable Low-Precision Training

Seyed Morteza Emadi (UNC-Chapel Hill)

OptimizationComputational EfficiencyTransformerLarge Language ModelContrastive LearningText

🎯 What it does: This paper studies the risk of attention logit overflow in low-precision FP8 training, and proposes a scaling prediction method based on weight geometry without activation observation.

Rank-guided Diffusion for Noise Few-Shot Learning

Zelei Wu (Ningbo University), Jieyu Zhao (Ningbo University)

ClassificationData-Centric LearningMeta LearningTransformerDiffusion modelAuto EncoderContrastive LearningImageText

🎯 What it does: Propose a few-shot learning framework based on low-rank guided diffusion (CRDProto), which detects noisy samples through differentiable low-rank decomposition (SVE) and generates high-quality alternative features using a diffusion model with low-rank constraints, ultimately constructing a clean and consistent support set.

Rank-Learner: Orthogonal Ranking of Treatment Effects

Henri Arno (Ghent University - imec), Stefan Feuerriegel (LMU Munich)

Recommendation SystemFederated LearningExplainability and InterpretabilityComputational EfficiencyData-Centric LearningContrastive LearningTabularBiomedical DataBenchmark

🎯 What it does: Propose Rank-Learner, a two-stage orthogonal learning framework that directly learns the ranking of treatment effects from observed data.

Ranking Free RAG: Replacing Re-ranking with Selection in RAG for Sensitive Domains

Yash Saxena (University of Maryland, Baltimore County), Manas Gaur (University of Maryland, Baltimore County)

RetrievalSafty and PrivacyExplainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextFinance RelatedRetrieval-Augmented Generation

🎯 What it does: Propose the METEORA framework, which replaces traditional RAG re-ranking with explainable reasoning, achieving interpretable evidence selection and verification in sensitive domains.

Ranking Time Series using a Time Warping Ideal Point Model

Lucas Zoroddu (Universite Paris Saclay), Laurent Oudre (Universite Paris Saclay)

Anomaly DetectionOptimizationComputational EfficiencySupervised Fine-TuningContrastive LearningTabularTime Series

🎯 What it does: Propose an ideal point model (IPM) based on time elastic distances (DTW, TWED), utilizing pairwise comparisons to learn the global ranking of time series, and provide theoretical convergence guarantees and a differentiable soft-TWED optimization scheme;

Rapid Poison: Practical Poisoning Attacks Against the Rapid Response Framework

David Huang (Princeton University), Chawin Sitawarin (Google DeepMind)

Safty and PrivacyAdversarial AttackTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: This paper investigates how to utilize Prompt Injection to inject a small number of poisoned jailbreak samples during the synthetic proliferation process of the Rapid Response (RR) framework, thereby achieving two types of attacks on the safety classifier: one is false-positive attacks targeting specific formats, domains, entities, or distributions; the other is false-negative attacks constructed using Omission Attack, which makes jailbreak samples with trigger words be misjudged as safe.

RAPNet: Accelerating Algebraic Multigrid with Learned Sparse Corrections

Yali Fink (Ben Gurion University Of Negev), Eran Treister (Ben Gurion University Of Negev)

OptimizationComputational EfficiencyRepresentation LearningGraph Neural NetworkContrastive LearningMeshGraphPhysics Related

🎯 What it does: Generate sparse incremental corrections for algebraic multigrid (AMG) during the preconditioning phase for solving sparse linear systems using graph neural networks, thereby improving the solver's convergence speed.

Rare Event Analysis of Large Language Models

Jake McAllister Dorman (University of Nottingham), Juan P. Garrahan (University of Nottingham)

Explainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelTextPhysics Related

🎯 What it does: Built and implemented a complete rare event analysis framework for large language models (LLMs), systematically evaluating from defining rare events, estimating probabilities, to exploring their structures.

RAST-MoE-RL: A Regime-Aware Spatio-Temporal MoE Framework for Deep Reinforcement Learning in Ride-Hailing

Yuhan Tang (Massachusetts Institute of Technology), Jinhua Zhao (Massachusetts Institute of Technology)

Autonomous DrivingOptimizationTransformerReinforcement LearningMixture of ExpertsAuto EncoderTabularTime Series

🎯 What it does: Designed and implemented a spatiotemporal reinforcement learning framework based on Mixture-of-Experts, named RAST-MoE, to address the adaptive delay matching problem in shared mobility platforms.

RAT+: Train Dense, Infer Sparse - Recurrence Augmented Attention for Dilated Inference

Xiuying Wei (EPFL), Caglar Gulcehre (EPFL)

Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelContrastive LearningText

🎯 What it does: Propose a dense pre-trained RAT+ model that utilizes full-sequence recursion and active recursion learning, enabling flexible switching to different sparse dilated attention modes during inference while maintaining performance close to full attention;

Rate or Fate? RLV$^{\varepsilon}$R: Reinforcement Learning with Verifiable Noisy Rewards

Ali Rad (Cognichip AI), Ehsan Kamalinejad (Cognichip AI)

Explainability and InterpretabilityTransformerLarge Language ModelReinforcement LearningMixture of ExpertsContrastive LearningText

🎯 What it does: The study investigates how learning direction and convergence speed of large language models are affected under noisy verifiable reward (RLVR), and provides an interpretable theoretical explanation.

Ratio-Variance Regularized Policy Optimization

Yu Luo (Huawei), Dong Li (Huawei)

OptimizationReinforcement LearningTextTabular

🎯 What it does: Proposes an R VPO strategy optimization method based on proportional variance regularization, replacing traditional hard clipping, which simultaneously controls upper and lower variance and achieves dual optimization;

Rational Neural Networks have Expressivity Advantages

Maosen Tang (Cornell University), Alex Townsend (Cornell University)

ClassificationTransformerReinforcement LearningVision-Language-Action ModelContrastive LearningImageVideoTabular

🎯 What it does: Studied neural networks with trainable low-order rational activation functions, and theoretically proved their superiority in expressiveness and parameter efficiency over modern fixed activation functions; subsequently verified practical performance improvements on visual classification and offline reinforcement learning tasks.

Rational Transductors

Mehryar Mohri (Google)

OptimizationComputational EfficiencyRepresentation LearningRecurrent Neural NetworkTransformerLarge Language ModelAuto EncoderContrastive LearningSequentialBenchmark

🎯 What it does: Proposed a dual-stream architecture called Rational Transductors, which combines the self-attention mechanism of Transformer with matrix recursion based on weighted finite automata (WFA), constructing a linear recursive head that can be computed in parallel, and achieving full-layer state information injection through deep Rational injection.

Rationality Measurement and Theory for Reinforcement Learning Agents

Kejiang Qian (University of Edinburgh), Fengxiang He (University of Edinburgh)

Reinforcement Learning

🎯 What it does: Proposes a rationality measure and theoretical framework for reinforcement learning agents, defining perfectly rational actions and quantifying rationality risk.

Rays as Pixels: Learning A Joint Distribution of Video and Camera Trajectories

Wonbong Jang (Meta AI), Tao Xiang (Meta AI)

GenerationPose EstimationTransformerDiffusion modelFlow-based ModelAuto EncoderVideo

🎯 What it does: Train a unified Video Diffusion Model that learns the joint distribution of video frames and camera trajectories, supporting camera pose estimation, camera-controlled video generation, and joint generation of both;

RBCBF: Decoding Time Safety Alignment via Risk Guided Rollback and Barrier Control

Tianxiang Chen (National University of Singapore), Kaidi Yang (National University of Singapore)

OptimizationSafty and PrivacyTransformerLarge Language ModelReinforcement LearningText

🎯 What it does: Proposes a rollback-based inference-time safe alignment framework called RBCBF, which utilizes risk aggregation and control barrier functions to locate and perform distribution-level correction on the generation process;

RC-FCL: Combating Asynchronous Concept Drift in Federated Continual Learning via Retrospective Calibration

Hang Su (Huazhong University of Science and Technology), Imran Razzak (Mohamed bin Zayed University of Artificial Intelligence)

ClassificationDomain AdaptationFederated LearningConvolutional Neural NetworkGenerative Adversarial NetworkContrastive LearningImage

🎯 What it does: Propose the RC-FCL framework to address the problem of asynchronous concept drift in federated continual learning. The framework generates a reference distribution using a generative adversarial network, detects drift with Sinkhorn distance, and achieves local adaptation and global updates through sample weighting and drift-aware aggregation.

RDT2: Exploring the Scaling Limit of UMI Data Towards Zero-Shot Cross-Embodiment Generalization

Songming Liu (Tsinghua University), Jun Zhu (Tsinghua University)

Knowledge DistillationRobotic IntelligenceTransformerVision Language ModelDiffusion modelFlow-based ModelRectified FlowVideoTextMultimodality

🎯 What it does: This paper proposes and trains RDT2, a robot foundation model based on a 7B parameter VLM, capable of achieving zero-shot generality on unseen objects, scenes, instructions, and robot platforms.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning

Renos Zabounidis (AWS Agentic AI), Stefano Soatto (AWS Agentic AI)

OptimizationComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsTextChain-of-Thought

🎯 What it does: Proposes Re‑FORC, a method that dynamically stops, selects models and reasoning lengths, and allocates computational resources during chain-of-thought reasoning by predicting the relationship between future rewards and computational costs through an adapter.

RE-TRAC: REcursive TRAjectory Compression for Deep Search Agents

jialiang zhu, Baining Guo (Southeast University)

Computational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningAgentic AITextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose the Re-TRAC framework, which generates structured state representations through recursive trajectory compression, enabling deep search agents to reflect and plan across trajectories.

Reading Between the Tokens: Improving Preference Predictions through Mechanistic Forecasting

Sarah Ball (LMU Munich), Niklas Kühl (University of Bayreuth)

Explainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextTabular

🎯 What it does: This study proposes a mechanistic forecasting method that aggregates and estimates voters' party preferences by utilizing hidden representations within large language models, thereby improving the prediction of voting outcomes.

Reading the Cell, Designing the Cure: Perturbation-Conditioned Molecular Diffusion for Function-Oriented Drug Design

ZIYU XU, Liang Wang (Chinese Academy of Sciences)

Drug DiscoveryGraph Neural NetworkTransformerPrompt EngineeringDiffusion modelContrastive LearningTabularBiomedical Data

🎯 What it does: Proposed and implemented a drug design framework called CURE based on transcriptome perturbation, which can generate new molecules that meet functional requirements by utilizing diffusion models under the condition of given prognostic transcriptome changes.

ReaForest: Fostering Generative Video Reasoning for Spatial Planning

Kun Ouyang (Peking University), Xu Sun (Peking University)

GenerationOptimizationReinforcement Learning from Human FeedbackTransformerSupervised Fine-TuningPrompt EngineeringVision Language ModelDiffusion modelImageVideoTextMultimodalityBenchmarkChain-of-Thought

🎯 What it does: Propose the ReaForest framework, which enhances the chain-frame reasoning capability of video generation models in spatial planning tasks through text-image alignment activation, test-time multi-branch search, and self-correction.

Real Data Lies: Unveiling and Closing the Quality Shortcut in Generalizable AI-Generated Video Detection

Ziyuan Fang (University of Science and Technology of China), Wenbo Zhou (University of Science and Technology of China)

Data SynthesisAnomaly DetectionTransformerDiffusion modelGenerative Adversarial NetworkContrastive LearningVideo

🎯 What it does: Explore and eliminate quality bias in AI-generated video detection, proposing a data augmentation training strategy based on quality matching and full-spectrum coverage.

Real-Time Aligned Reward Model beyond Semantics

Zixuan Huang (Beihang University), deqing wang

OptimizationComputational EfficiencyReinforcement Learning from Human FeedbackTransformerReinforcement LearningPrompt EngineeringContrastive LearningText

🎯 What it does: Propose a lightweight RLHF framework called R2M, which dynamically aligns the reward model by utilizing real-time feedback from the hidden layers of the policy model, reducing the phenomenon of reward over-optimization.

Real-Time Monitoring and Calibration of Chain-of-Thought Sycophancy in Large Reasoning Models

Jingyu Hu (University of Bristol), Di Wang (King Abdullah University of Science and Technology)

Explainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerSupervised Fine-TuningPrompt EngineeringContrastive LearningTextChain-of-Thought

🎯 What it does: Constructed the MONICA framework, which dynamically corrects sycophancy behavior in the chain-of-thought process of large-scale inference models in real-time.

Real-Time Visual Attribution Streaming in Thinking Model

Seil Kang (Yonsei University), Seong Jae Hwang (Yonsei University)

Explainability and InterpretabilityComputational EfficiencyTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityChain-of-Thought

🎯 What it does: Designed a real-time streamable visualization attribution framework called VSTREAM for multimodal reasoning models, which learns a lightweight linear estimator to predict the causal impact of semantic regions from attention features, enabling immediate tracking of the visual evidence relied upon during text generation in the reasoning process.

Real-World Unsupervised Models Generalize to Predict Brain Responses to Out-of-Distribution Stimuli

Chenggang Chen (Johns Hopkins University), Xiaoqin Wang (Johns Hopkins University)

Representation LearningTransformerContrastive LearningVideoBiomedical DataMagnetic Resonance ImagingAudio

🎯 What it does: This paper trains multiple self-supervised models based on a large amount of unlabeled natural data, and systematically compares their predictive performance for human auditory and visual cortices with traditional supervised models.

REAL: Regression-Aware Reinforcement Learning for LLM-as-a-Judge

Yasi Zhang (University of California Los Angeles), Michal Lukasik (Google Research)

OptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkChain-of-Thought

🎯 What it does: Propose a regression-aware reinforcement learning framework called REAL, which is oriented towards numerical scoring and integrates chain-of-thought (CoT) exploration with numerical prediction refinement.

REAL: Resolving Knowledge Conflicts in Knowledge-Intensive Visual Question Answering via Reasoning-Pivot Alignment

Kai Ye (Zhejiang University), Jiajun Bu (Zhejiang University)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningData-Centric LearningTransformerSupervised Fine-TuningPrompt EngineeringContrastive LearningImageTextMultimodalityRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This paper proposes a framework based on reasoning pivot alignment (REAL) to address knowledge conflicts in knowledge-intensive visual question answering (KI-VQA).

RealisMotion: Decomposed Human Motion Control and Video Generation in the World Space

Jingyun Liang (Damo Academy Alibaba Group), Fan Wang (Damo Academy Alibaba Group)

GenerationPose EstimationDepth EstimationTransformerSupervised Fine-TuningPrompt EngineeringDiffusion modelImageVideoPoint Cloud

🎯 What it does: This paper proposes a framework that decouples human motion and video generation in a 3D world space, enabling separate control over the foreground subject, background video, motion trajectory, and action patterns;

REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations

Buyun Liang (University of Pennsylvania), Rene Vidal

Adversarial AttackTransformerLarge Language ModelPrompt EngineeringGenerative Adversarial NetworkTextStochastic Differential Equation

🎯 What it does: This paper proposes a REALISTA framework based on latent space for adversarial attacks, which can generate input prompts that induce hallucinations in LLMs while maintaining semantic equivalence and coherence.

Realizable Bayes-Consistency for General Metric Losses

Dan Tsir Cohen (Ben Gurion University of Negev), Aryeh Kontorovich (Ben Gurion University of Negev)

OptimizationExplainability and InterpretabilityRepresentation LearningContrastive Learning

🎯 What it does: This paper provides necessary and sufficient conditions for strong Bayes-consistency (RSUBC) in the realizable scenario for general metric loss learning,

RealtimeTool: Parallel Decoding for Real-Time LLM Function Calling

Xiaoxin Shi (Shanghai Jiao Tong University), Zengfeng Huang (Shanghai Innovation Institute)

Computational EfficiencyKnowledge DistillationAI Code AssistantTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsAuto EncoderContrastive LearningTextRetrieval-Augmented Generation

🎯 What it does: Proposed RealtimeTool, a real-time LLM function calling framework that compresses low-entropy information through special tokens and achieves parallel decoding of function names and parameters.

REAR: Test-time Preference Realignment through Reward Decomposition

Fuxiang Zhang (Nanyang Technological University), Bo An (Nanyang Technological University)

Reinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringText

🎯 What it does: Propose a framework called REAR that achieves preference alignment during inference by decomposing rewards, enabling the adjustment of LLM generation according to user preferences without additional training.

Reason with Thumbnails, Answer with Focus: An Efficient and Effective Paradigm for Multimodal Grounded Visual Reasoning

An-Lan Wang (Sun Yat-sen University), Kun-Yu Lin (University of Hong Kong)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerReinforcement LearningPrompt EngineeringVision Language ModelImageTextMultimodality

🎯 What it does: Proposes a two-stage multimodal reasoning framework called 'Reason with Thumbnails, Answer with Focus (RTAF)', which first uses low-resolution thumbnails to locate key regions and then uses high-resolution cropping to obtain the final answer.

Reason, Then Re-reason: Cross-view Revisiting Improves Spatial Reasoning

Chaofan Ma (Shanghai Jiao Tong University), Jiangchao Yao (Shanghai Jiao Tong University)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerPrompt EngineeringVision Language ModelNeural Radiance FieldAuto EncoderGenerative Adversarial NetworkImageVideoPoint CloudBenchmarkChain-of-Thought

🎯 What it does: To address the perspective limitations in spatial reasoning from single-view videos, this paper proposes ReRe, a two-stage framework that does not require training. It first forms spatial hypotheses under the original perspective and then verifies or refines the answers by synthesizing new perspective videos.

ReasonEdit: Editing Vision--Language Models using Human Reasoning

Jiaxing Qiu (University of Virginia), Thomas Hartvigsen (University of Virginia)

Explainability and InterpretabilityRepresentation LearningData-Centric LearningTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This paper proposes ReasonEdit, a framework that utilizes human reasoning to edit vision-language models (VLMs), enabling the correction of errors in VLMs during reasoning tasks without altering irrelevant behaviors;

Reasoning about Reasoning: BAPO Bounds on Chain-of-Thought Token Complexity in LLMs

Kiran Tomlinson (Microsoft Research), Jennifer Neville (Microsoft Research)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextGraphChain-of-Thought

🎯 What it does: This paper analyzes the required reasoning token length for chain-of-thought (CoT) using the BAPO model, proving that for three classes of BAPO-hard problems (binary majority, ternary matching, graph connectivity), at least linear (Ω(n)) CoT tokens are needed; meanwhile, it provides matching upper bounds and experimental verification.

Reasoning as an Attack Surface: Adaptive Evolutionary CoT Jailbreaks for LLMs

Jianan Li (Northeastern University), Xiaochun Cao (Sun Yat-sen University)

Explainability and InterpretabilityComputational EfficiencyAdversarial AttackTransformerLarge Language ModelPrompt EngineeringTextChain-of-Thought

🎯 What it does: Propose an Adaptive Evolutionary Chain-of-Thought (AE-CoT) framework that generates and optimizes jailbreak triggers for large reasoning models using teacher-style rewriting and fragmentation strategies.

Reasoning Cache: Continual Improvement Over Long Horizons via Short-Horizon RL

Ian Wu (Carnegie Mellon University), Aviral Kumar (Carnegie Mellon University)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed an iterative decoding algorithm called Reasoning Cache (RC), which replaces traditional autoregressive decoding, allowing large language models to utilize longer reasoning time during testing and continuously improve answers.

Reasoning Can Be Restored by Correcting a Few Decision Tokens

Changshuo Shen (University of Science and Technology of China), An Zhang (University of Science and Technology of China)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelPrompt EngineeringTextBenchmarkChain-of-Thought

🎯 What it does: Studied the token-level disagreements between base models and large reasoning models during the autoregressive generation process, quantified the disagreements, and performed sparse interventions based on the disagreement rate to restore reasoning capabilities.

Reasoning Compartmentalization: Bridging the Concretization Gap via Abstraction-based Routing

Ling-I Wu (Shanghai Jiao Tong University), Guoqiang Li (Shanghai Jiao Tong University)

Federated LearningExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsContrastive LearningTextChain-of-Thought

🎯 What it does: The study investigates the differences in reasoning between abstract forms (FL) and natural language forms (NL) of the same logical task in large language models, and proposes an abstract alignment intervention achieved through translation training;

Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge

Wenbo Zhang (University of California), Hengrui Cai (University of California)

Recommendation SystemOptimizationComputational EfficiencyTransformerLarge Language ModelText

🎯 What it does: Compare reasoning and non-reasoning LLM discriminators, and propose a robust routing method called RACER that adaptively activates reasoning modes under a fixed budget.

Reasoning LLM Improves Speaker Recognition in Long-form TV Dramas

Yuxuan Li (Tsinghua University), Qi Tian (Guangdong Laboratory of Artificial Intelligence and Digital Enonomy)

RecognitionTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AIVideoTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-ThoughtAudio

🎯 What it does: This paper proposes the DramaSR-532K benchmark and the DramaSR-LRM model, focusing on the speaker recognition task in long-form television dramas.

Reasoning Models Are Test Exploiters: Rethinking Multiple Choice

Narun Krishnamurthi Raman (University of British Columbia), Kevin Leyton-Brown (University of British Columbia)

Explainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelTextBenchmarkChain-of-Thought

🎯 What it does: Examines the bias in multiple-choice question (MCQA) evaluation for large language model (LLM) reasoning capabilities, systematically evaluates 15 benchmarks with 27 models, and analyzes how option exposure is utilized by the models.

Reasoning Models Struggle to Control their Chains of Thought

Chen Yueh-Han (New York University), Tomek Korbak (OpenAI)

Explainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkChain-of-Thought

🎯 What it does: This paper proposes and constructs the CoT-Control evaluation suite, which is used to systematically measure the ability of reasoning models to control their Chain-of-Thought (CoT) while meeting task constraints. Large-scale experiments are conducted on 12 state-of-the-art reasoning models to explore the impact of factors such as model scale, RL training, and situational awareness on CoT controllability.

Reasoning on the Manifold: Bidirectional Consistency for Self-Verification in Diffusion Language Models

Jiaoyang Ruan (Fudan University), Jian Pu (Fudan University)

Explainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelDiffusion modelTextBenchmarkChain-of-Thought

🎯 What it does: Propose an untrained, unsupervised geometric consistency metric called BMC to evaluate the effectiveness of sequences generated by diffusion large language models;

Reasoning over Boundaries: Enhancing Specification Alignment via Test-time Deliberation

Haoran Zhang (Shanghai Jiao Tong University), Yu Cheng (Chinese University of Hong Kong)

Safty and PrivacyExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed the concept of 'norm alignment' and constructed a unified SPECBENCH benchmark to evaluate the performance of LLMs under dual constraints of safety norms and behavioral norms; simultaneously proposed a lightweight single-round inference method called ALIGN3, which improves norm compliance through phased behavioral optimization, safety guidance, and overall review.

Reasoning Quality Emerges Early: Data Curation for Reasoning Models

Hongyi Henry Jin (University of California Los Angeles), Baharan Mirzasoleiman (University of California Los Angeles)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringContrastive LearningTextBiomedical DataBenchmarkChain-of-Thought

🎯 What it does: Designed a method that utilizes the loss of the top 100 chain-of-thought tokens on randomly perturbed points from a pre-trained model to filter high-difficulty and diverse training samples. It further clusters and selects samples with low gradient similarity by evaluating the loss of the first 1k tokens through small perturbations along the fine-tuning direction, thus constructing a high-quality supervised fine-tuning (SFT) dataset.

Reasoning Structure of Large Language Models

Frédéric Berdoz (ETH Zurich), Roger Wattenhofer (ETH Zurich)

Explainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextGraphBenchmarkChain-of-Thought

🎯 What it does: Propose an scalable 2D grid puzzle benchmark and build a pipeline that converts large language model reasoning trajectories into verifiable reasoning graphs, while defining a reasoning efficiency metric η.

Reasoning Theater: Disentangling Model Beliefs from Chain-of-Thought

Siddharth Boppana (Goodfire AI), Jack Merullo (Goodfire AI)

Explainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextBenchmarkChain-of-Thought

🎯 What it does: The study investigates whether the internal beliefs and external expressions of large language models are consistent during chain-of-thought reasoning, finding that easy questions often exhibit 'performative reasoning,' while difficult questions are closer to actual reasoning;

Reasoning to Edit: Hypothetical Instruction-Based Image Editing with Visual Reasoning

Qingdong He (Tencent Youtu Lab), Yabiao Wang (Zhejiang University)

Image TranslationGenerationData SynthesisTransformerLarge Language ModelPrompt EngineeringVision Language ModelDiffusion modelImageTextMultimodalityRetrieval-Augmented Generation

🎯 What it does: Proposed the Hypothetical Instruction Image Editing task (HI-IE), and constructed the Reason50K dataset containing 51,039 triplets, as well as the ReasonBrain framework based on MLLM and diffusion models, which can perform reasoning and editing on implicit hypothetical instructions;

Reasoning-preserved Efficient Distillation of Large Language Models via Activation-aware Initialization

Junlin He (Hong Kong Polytechnic University), Wei Ma (Hong Kong Polytechnic University)

Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningTextChain-of-Thought

🎯 What it does: The RED method is proposed to address the reasoning collapse caused by the drop in representation rank (eRank collapse) due to width compression in large language models, by using activation-aware initialization during the width compression process.

Reasoning-VLA: An Efficient and Spatial-Guided General Vision-Language-Action Reasoning Model for Autonomous Driving

Dapeng Zhang (National University of Singapore), Tat-Seng Chua (National University of Singapore)

Autonomous DrivingOptimizationTransformerSupervised Fine-TuningReinforcement LearningVision Language ModelVision-Language-Action ModelImageVideoTextMultimodalityChain-of-Thought

🎯 What it does: Propose the Reasoning-VLA framework, which utilizes learnable action queries and implicit spatial guidance combined with reasoning-enhanced vision-language models to generate continuous driving trajectories in parallel during a single forward pass.

ReAugment: Targeted Few-Shot Time Series Augmentation via Model Zoo-Guided Reinforcement Learning

Haochen Yuan (Shanghai Jiao Tong University), Xiaokang Yang (Shanghai Jiao Tong University)

Data SynthesisTransformerReinforcement LearningAuto EncoderGenerative Adversarial NetworkContrastive LearningTime Series

🎯 What it does: For few-shot time series forecasting, a closed-loop reinforcement learning framework called ReAugment is proposed to adaptively generate augmented samples;

RECAST: Model Reconstruction via Counterfactual-Aware Wasserstein Geometry under Limited Data

Xuan Zhao (Forschungszentrum Jülich), Ira Assent (Forschungszentrum Jülich)

ClassificationExplainability and InterpretabilityData-Centric LearningContrastive LearningTabularBenchmark

🎯 What it does: Propose a black-box model reconstruction method called RECAST under limited data and one-time counterfactual explanations, which constructs Wasserstein centroid prototypes to approximate class distributions and generate approximate surrogates.

ReCoG: Relational and Compact Context Graph Learning for Few-shot Molecular Property Prediction

Zeyu Wang (Zhejiang University of Technology), Shirui Pan (Griffith University)

OptimizationRepresentation LearningMeta LearningDrug DiscoveryGraph Neural NetworkMixture of ExpertsContrastive LearningGraphTabularBiomedical DataBenchmark

🎯 What it does: Construct and learn relational and compact context graphs to address the few-shot molecular property prediction problem.

Recognize Your Orchestrator: An Entropy Dynamics Perspective for LLM Multi-Agent Systems

Junze Zhu (Nanjing University), Xinyu Dai (Nanjing University)

Anomaly DetectionAutonomous DrivingOptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyMeta LearningDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringWorld ModelTextSequentialReview/Survey PaperBenchmarkChain-of-ThoughtStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Studied the behavior of large language models (LLMs) as centralized orchestrators in multi-agent systems (MAS), modeling the orchestration process as Mean-Field entropy dynamics, and generating observable step-level logs through inverse workflow generation (IWG) to validate the model and reveal the 'Reasoning Trap' problem.

Reconstructing Template-Memorized Images from Natural Prompts

Sol Yarkoni (Tel Aviv University), Roi Livni (Tel Aviv University)

Image TranslationRestorationSegmentationGenerationData SynthesisAdversarial AttackTransformerPrompt EngineeringDiffusion modelContrastive LearningImageTextMultimodalityRetrieval-Augmented Generation

🎯 What it does: Proposed a low-resource template memory image reconstruction attack, which can extract template images memorized during training from text-to-image models using common natural prompts;

Reconstruction Outcomes Look Similar but Processes Differ: Improving Context Consistency and Coverage in Graph Masked Auto-Encoder

Geng Tang (Jiangsu University of Science and Technology), Yuhua Qian (Shanxi University)

ClassificationRepresentation LearningGraph Neural NetworkTransformerAuto EncoderContrastive LearningGraph

🎯 What it does: Propose a graph mask autoencoder, C2-GMAE, which considers both neighborhood context consistency and coverage, improving the reconstruction process of traditional GMAE.

Recontextualization Mitigates Specification Gaming Without Modifying the Specification

Ariana Azarbal (MATS), Alexander Matt Turner

Data-Centric LearningAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringText

🎯 What it does: Propose a training method called 'Recontextualization,' which reduces the normative gaming phenomenon when the reward signal is incomplete, by using inhibitory prompts during generation and allowing or encouraging erroneous behavior during training.

RECOVER: Reliable Detection of Unauthorized Data Usage in Text-to-Image Diffusion Models via Inversion Robustness

Yanhao Wei (Wuhan University), Run Wang (Wuhan University)

Anomaly DetectionSafty and PrivacyTransformerDiffusion modelScore-based ModelAuto EncoderImageText

🎯 What it does: A non-intrusive copyright detection framework called RECOVER is proposed for text-to-image diffusion models, which can identify whether unauthorized data was used for fine-tuning without requiring pre-finetuning of the model or the use of watermarks.

Recovering Hidden Reward in Diffusion-Based Policies

Yanbiao Ji (Shanghai Jiao Tong University), Hongtao Lu (Shanghai Jiao Tong University)

Robotic IntelligenceReinforcement LearningDiffusion modelScore-based ModelFlow-based ModelTabularTime SeriesSequentialOrdinary Differential Equation

🎯 What it does: This paper parameterizes diffusion policies by using an energy function as a scalar of state-action, enabling denoising score matching to both generate actions and implicitly extract rewards in maximum entropy inverse reinforcement learning.