arXivSub Start free trial

ICML 2026 Papers — Page 13

International Conference on Machine Learning · 6554 papers

Cross-Modal Knowledge Distillation without Paired Data: Theoretical Foundation and Algorithm

T. Khiem Tran, Trong Nghia Hoang (Washington State University)

Knowledge DistillationRepresentation LearningMeta LearningConvolutional Neural NetworkTransformerContrastive LearningImageVideoMultimodalityAudio

🎯 What it does: Propose a cross-modal knowledge distillation framework without sample-level pairing, and provide theoretical upper bounds on generalization error. Design a distribution alignment loss based on feature and label alignment, and achieve unpaired distillation through bi-level optimization.

Cross-Modal Semantic Decoupling and Transfer for Text-to-Visible-Infrared Person Re-Identification

Ziang Zhang (Wuhan University), Mang Ye (Wuhan University)

RecognitionRetrievalDomain AdaptationTransformerVision Language ModelDiffusion modelScore-based ModelFlow-based ModelRectified FlowAuto EncoderGenerative Adversarial NetworkContrastive LearningImageTextMultimodality

🎯 What it does: This paper proposes a cross-modal semantic disentanglement and transfer framework, CSDT, for cross-modal person re-identification (TVI-ReID) tasks involving text retrieval of visible and infrared images, addressing the issue of insufficient visible light information in nighttime surveillance scenarios.

Cross-Subject Modeling for Widefield Calcium Imaging via Atlas-Aligned Spatiotemporal Tokenization

Mohammad Hosseini (University of Southern California), Maryam M. Shanechi (University of Southern California)

Representation LearningData-Centric LearningConvolutional Neural NetworkTransformerSupervised Fine-TuningAuto EncoderContrastive LearningBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose WiCAT, a multi-agent wide-field calcium imaging model that achieves zero-shot behavior decoding and brain region reconstruction through shared representations across subjects.

Cross-Tactile Sensor Representation Learning

Yan Zhang (Tongji University), Heng Tao Shen (Tongji University)

ClassificationRecognitionPose EstimationDomain AdaptationRepresentation LearningTransformerVision Language ModelAuto EncoderContrastive LearningImageTextMultimodality

🎯 What it does: This paper proposes a cross-tactile sensor representation learning framework called CTSRL, which achieves sensor-agnostic tactile representations by utilizing a cross-sensor modulator and a two-stage training approach.

Cross-task Calibration for Asynchronous Federated Continual Learning

Yichen Li (Huazhong University of Science and Technology), Ruixuan Li (Huazhong University of Science and Technology)

Federated LearningContrastive LearningImage

🎯 What it does: Proposes a cross-task calibration framework called C-AFCL2 to address client drift and task drift problems in asynchronous federated continual learning.

Cross-View Lewis Weight Fusion Empowering Exemplar Replay for Federated Class-Incremental Learning

Zhuang Qi (Shandong University), Xiangxu Meng (Shandong University)

ClassificationFederated LearningRepresentation LearningConvolutional Neural NetworkTransformerContrastive LearningImage

🎯 What it does: Propose the CLIF framework, combining cross-perspective Lewis weight fusion with frequency-weighted training to improve sample selection and replay in federated class-incremental learning;

CrossQ: Task-Aligned Cross-Token Conditional Quantization for Late Interaction Retrieval

Rohit Kumar Salla (Virginia Tech), Ramya Manasa Amancherla (Columbia University)

RetrievalCompressionKnowledge DistillationTransformerAuto EncoderContrastive LearningText

🎯 What it does: Proposed a context-aware multi-codebook quantization method called CrossQ, used to compress the document-side token embeddings in Late-Interaction retrieval models while maintaining the original query-side full precision;

Crowd4D: Scene-Aware Monocular 4D Crowd Reconstruction

Hongbo Kang (Tianjin University), Kun Li (Tianjin University)

Pose EstimationOptimizationConvolutional Neural NetworkTransformerDiffusion modelNeural Radiance FieldAuto EncoderContrastive LearningOptical FlowImageVideoPoint Cloud

🎯 What it does: Proposed a monocular camera 4D crowd reconstruction framework called Crowd4D for large-scale complex scenarios, achieving scene-consistent 4D reconstruction by jointly optimizing the crowd and the scene.

CRPO: Character-centric Group Relative Policy Optimization for Role-aware Reasoning in Role-playing Agents

Yihong Tang (Harbin Institute of Technology), Min Zhang (Harbin Institute of Technology)

OptimizationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringContrastive LearningText

🎯 What it does: Proposes the CRPO framework for role-playing agents, improving the optimization objective of traditional GRPO to perform reinforcement learning centered on roles, thereby enhancing role consistency and reasoning depth in dialogues.

CryoACE: An Atom-centric Framework for Accurate and Automated Model Building in Cryo-EM

Minzhang Li (ShanghaiTech University), Jingyi Yu (ShanghaiTech University)

Protein Structure PredictionConvolutional Neural NetworkTransformerMixture of ExpertsDiffusion modelScore-based ModelContrastive LearningBiomedical Data

🎯 What it does: Proposed CryoACE, an end-to-end atomic center framework for automatically building high-precision atomic models from cryo-EM density maps.

CSD: Content-aware Speculative Decoding for Efficient Image Generation

Mingcheng Wang (East China Normal University), Shaohui Lin (East China Normal University)

GenerationTransformerDiffusion modelImage

🎯 What it does: Proposes a content-aware speculative decoding (CSD) to accelerate autoregressive image generation.

CSG: Cognitive Structure Generation for Intelligent Education

Hengnian Gu (Northeast Normal University), Dongdai Zhou (Northeast Normal University)

Knowledge DistillationRepresentation LearningReinforcement Learning from Human FeedbackGraph Neural NetworkTransformerReinforcement LearningDiffusion modelScore-based ModelGraphSequential

🎯 What it does: A cognitive structure generation framework (CSG) based on a graph diffusion probability model was constructed, which can generate a cognitive network of concepts and their relationships from student interaction logs.

CSOR: Coreset Selection for Object Re-identification via Class Pruning

Minyoung Oh (Ulsan National Institute of Science and Technology), Jae-Young Sim (Ulsan National Institute of Science and Technology)

RetrievalOptimizationData-Centric LearningContrastive LearningImage

🎯 What it does: Proposes a core set selection (CSOR) method for object re-identification (ReID), jointly performing class pruning and sample selection;

CSPLoRA: Confidence-Guided Structure Planning for Low-Rank Adaptation

Huiming Ding (University of Science and Technology of China), Zhenyu Tan (University of Science and Technology of China)

Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBenchmark

🎯 What it does: This paper proposes the CSPLoRA framework, which achieves structural planning for LoRA's low-rank adapters by predicting uncertainty-weighted samples and utilizing Taylor expansion to estimate module importance.

CSPO: Constraint-Sensitive Policy Optimization for Safe Reinforcement Learning

Ayoub Belouadah (University of Luxembourg), YVES LE TRAON

OptimizationReinforcement Learning

🎯 What it does: Propose a first-order primal-dual algorithm called CSPO, which uses a constraint gradient norm adaptive correction strategy to update, enabling rapid recovery of safety while maintaining KKT optimality.

CUARewardBench: A Benchmark for Evaluating Reward Models for Computer-Using Agents

Haojia Lin (Tencent Youtu Lab), Xing Sun (Tencent Youtu Lab)

Reinforcement Learning from Human FeedbackTransformerPrompt EngineeringMixture of ExpertsVision Language ModelMultimodalityBenchmark

🎯 What it does: Constructed the CUARewardBench benchmark to evaluate reward models for computer usage agents, including trajectory success and step correctness.

CUPID in the Model Zoo: Online Matchmaking for Selecting Your Dream LLM

Son Nguyen (Arizona State University), Ransalu Senanayake (Arizona State University)

Recommendation SystemOptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringText

🎯 What it does: Proposes the CUPID framework, which uses dual-arm bandit to iteratively select pairs of LLMs, collects user preference comparison feedback, and matches the most suitable LLM for users under cost and time budget constraints through Bayesian learning.

Curated Synthetic Data Doesn’t Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences

Ali Falahati (University of Waterloo), Lukasz Golab (University of Waterloo)

GenerationData SynthesisTransformerReinforcement LearningFlow-based ModelGenerative Adversarial NetworkImageText

🎯 What it does: This paper studies the selection of synthetic data using multiple reward functions within a recursive self-training loop to prevent mode collapse in generative models.

Curating the Future: A Scalable Recipe for Training Open-Ended Forecasters

Nikhil Chandak (Max Planck Institute For Intelligent Systems), Jonas Geiping (Max Planck Institute For Intelligent Systems)

Data-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningTextRetrieval-Augmented Generation

🎯 What it does: This study constructs the OpenForesight dataset by automatically generating open-ended forecasting question-answer pairs from daily news, and fine-tunes large language models using reinforcement learning on this dataset, achieving an 8B-scale predictor that can compete with large proprietary models.

Cure-SFT: Diagnostic-Guided Data Curation for Instruction Tuning

Yuankang Fu (South China University of Technology), Kaixiang Yang (South China University of Technology)

Data-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposed a diagnostic-driven instruction data refinement method called Cure-SFT, aimed at enhancing the instruction-following ability of large language models.

CURE: Consistency-under-Unified Semantic Regularization for Generalized Category Discovery

Yuwei Bian (Nanjing University of Science and Technology), Haofeng Zhang (Nanjing University of Science and Technology)

ClassificationDomain AdaptationRepresentation LearningTransformerAuto EncoderContrastive LearningOptical FlowImage

🎯 What it does: Proposes the CURE two-stage framework, which first learns class prototypes using labeled data while maintaining semantic coherence between prototypes, and then extends to unlabeled data using consistency regularization and semantic exploration energy, achieving general category discovery under a unified semantic structure.

CURE: Context-driven Diffusion with Progressive Expansion for Single Domain Generalization in Time Series Classification

Yuhang Pei (Northeastern University), Xiao Luo (University of Wisconsin-Madison)

ClassificationDomain AdaptationTransformerDiffusion modelAuto EncoderContrastive LearningTime Series

🎯 What it does: Propose a single-domain generalization framework called CURE based on conditional diffusion models, which utilizes semantic-aware and semantic-agnostic contexts for data augmentation, and achieves progressive diversity expansion through a memory bank and boundary filtering.

cuRegOT: A GPU-Accelerated Solver for Entropic-Regularized Optimal Transport

Yixuan Qiu (Shanghai University of Finance and Economics)

OptimizationImagePoint CloudTabular

🎯 What it does: Proposes cuRegOT, a GPU-accelerated entropy-regularized optimal transport solver based on the SPLR approximation of second-order methods.

Curriculum Reinforcement Learning for Black-Box Prompt Tuning via Large Language Models

Shuai Gong (Shandong University of Finance and Economics), Linwei Fan (Shandong University of Finance and Economics)

OptimizationExplainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringImageText

🎯 What it does: Propose the CRL-BPT framework, which combines curriculum reinforcement learning in black-box prompt tuning to guide LLMs from imitating reference prompts to generating innovative prompts, thus generating interpretable and high-performance prompts in scenarios with limited API calls.

Curriculum-Guided Layer Scaling for Language Model Pretraining

Karanpartap Singh (Stanford University), Ehsan Adeli (Stanford University)

Computational EfficiencyKnowledge DistillationRepresentation LearningData-Centric LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBenchmark

🎯 What it does: Proposed and implemented the Curriculum-Guided Layer Scaling (CGLS) framework, which integrates progressive layer stacking with curriculum learning by simultaneously increasing model depth and training data difficulty during pre-training.

CURVE: Learning Causality-Inspired Invariant Representations for Robust Scene Understanding via Uncertainty-Guided Regularization

Yue Liang (Tongji University), Hong Chen (Tongji University)

Domain AdaptationAutonomous DrivingExplainability and InterpretabilityComputational EfficiencyRepresentation LearningGraph Neural NetworkContrastive LearningImageGraph

🎯 What it does: Propose the CURVE framework, which utilizes variational uncertainty modeling and prototype-based soft backdoor adjustment to achieve interpretable causal sparse structures in scene graphs, thereby enhancing robustness.

CurvZO: Adaptive Curvature-Guided Sparse Zeroth-Order Optimization for Efficient LLM Fine-Tuning

Shuo Wang (Southern University of Science and Technology), Ming Tang (Southern University of Science and Technology)

OptimizationComputational EfficiencyTransformerLarge Language ModelText

🎯 What it does: Propose a method to efficiently fine-tune large language models by using online curvature signals to guide sparse zeroth-order optimization (CurvZO).

Cutting LLM Evaluation Costs with SySRs: A Bandit Algorithm That Provably Exploits Model Similarity

Zifan Lyu (ETH Zurich), Florian E. Dorner (Max-Planck Institute for Intelligent Systems)

OptimizationComputational EfficiencyHyperparameter SearchTransformerLarge Language ModelReinforcement LearningTextBenchmark

🎯 What it does: Propose a new parameter-free best model identification algorithm called SySRs, which reduces the evaluation cost of LLMs by leveraging the similarity between models.

CVE-Factory: Scaling Expert-Level Agentic Tasks for Code Security Vulnerability

Xianzhen Luo (Harbin Institute of Technology), Wanxiang Che (Harbin Institute of Technology)

Anomaly DetectionOptimizationFederated LearningSafty and PrivacyAI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Automatically generate executable security vulnerability repair tasks, construct the continuously updated LiveCVEBench benchmark, and produce over 1,000 executable training environments.

CVSearch: Empowering Multimodal LLMs with Cognitive Visual Search for High-Resolution Image Perception

Liupeng Li (Harbin Institute of Technology), Yaowei Wang (Harbin Institute of Technology)

Image TranslationRestorationObject DetectionSegmentationSuper ResolutionRetrievalTransformerLarge Language ModelPrompt EngineeringVision Language ModelDiffusion modelContrastive LearningImageMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: Designed and implemented a training-free CVSearch framework that provides cognitive visual search with high-resolution images for multimodal large language models through an Assess-then-Search workflow.

CyberGym-E2E: Scalable Real-World Benchmark for AI Agents' End-to-End Cybersecurity Capabilities

Tianneng Shi (University Of California Berkeley), Dawn Song (University Of California Berkeley)

AI Code AssistantLarge Language ModelAgentic AIPrompt EngineeringTextTabularBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose the CyberGym-E2E benchmark to evaluate the end-to-end capabilities of AI agents throughout the vulnerability lifecycle (discovery, PoC generation, patch generation).

CyberJurors: A Multi-Agent Simulation Task for E-Commerce Disputes Verdict

Yanhui Sun (University of Science and Technology of China), Yongdong Zhang (University of Science and Technology of China)

Recommendation SystemAnomaly DetectionExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelAgentic AIPrompt EngineeringImageVideoTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed a task called 'E-commerce Dispute Verdicts (EDV)' aimed at e-commerce transaction disputes, and implemented intelligent adjudication through multi-agent simulation.

Cycle-of-Science: Reliable Reasoning through Counterfactual Verification for Agent Decision Making

Ruojie Zhang (University of Electronic Science and Technology of China), dayong zhu

OptimizationExplainability and InterpretabilityRobotic IntelligenceReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AIPrompt EngineeringTextSequentialBenchmarkChain-of-Thought

🎯 What it does: Propose the Cycle-of-Science framework, which utilizes the hypothesis–experiment–validation cycle and adversarial causal preference optimization, enabling LLM-driven agents to actively verify causal relationships during the decision-making process.

D-ARL: A Distribution-Matched Asynchronous Reinforcement Learning Framework for Language Reasoning

Yinqi Bai (University of Science and Technology of China), Feng Wu (University of Science and Technology of China)

Large Language ModelReinforcement LearningTextBenchmark

🎯 What it does: Proposes a distribution-matching asynchronous reinforcement learning framework, D-ARL, for post-training in language reasoning.

D-CORE: Incentivizing Task Decomposition in Large Reasoning Models for Complex Tool Use

Bowen Xu (Alibaba Cloud Computing, Alibaba Group), Bin Yang (Alibaba Cloud Computing, Alibaba Group)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkChain-of-Thought

🎯 What it does: Propose the D-CORE framework, aiming to enhance the task decomposition and reasoning capabilities of large reasoning models (LRM) in complex tool usage scenarios, addressing the 'Lazy Reasoning' problem.

D-FUSEr: Diverse Failure, Unified Success via Error-Distribution Shaping in LLM Reasoning

David D. Baek (Massachusetts Institute of Technology), Tao Wang

OptimizationExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsDiffusion modelScore-based ModelTextBenchmarkChain-of-Thought

🎯 What it does: Propose the D-FUSEr framework, which enhances the performance of majority voting and iterative refinement during inference by shaping the error distribution of LLMs.

D-Judge: Disrupting Multi-Turn Jailbreaks using Semantics-Preserving Output Rewriting

Huanli Gong (University of California, Berkeley), N. Benjamin Erichson (International Computer Science Institute)

Safty and PrivacyReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringText

🎯 What it does: In multi-round jailbreak attacks, D-Judge first rewrites the output of the victim LLM while preserving its semantics before it is evaluated by an external judgment model, thereby disrupting the attacker's feedback-driven rewriting loop;

D$^2$O: A Dual Debiasing Operator for Training-Free Test-Time Adaptation of Vision–Language Models

Yihong Luo (Fujian University of Technology), Zhuo-Xu Cui (Shenzhen Institutes of Advanced Technology)

Domain AdaptationComputational EfficiencyRepresentation LearningPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityRetrieval-Augmented Generation

🎯 What it does: Proposes a training-free dual debiasing operation D²O for adaptive testing of vision-language models under style/environment distribution drift.

d$^2$p: Structured Soft Attention Is All You Need

Casey Sumagaysay Mogilevsky (Orikata Bio PBC), Kimberly Liang (Independent Researcher)

OptimizationProtein Structure PredictionDiffusion modelScore-based ModelTextBiomedical DataBenchmark

🎯 What it does: Proposed the d2p library, viewing dynamic programming (DP) as a temperature-regulated structured attention, and implemented second-order gradient computation for learnable DP hyperparameters (gap, edit cost, temperature);

D$^3$: Dynamic Directional Graph-Constrained Data Scheduling for LLM Training

Yuanjian Xu (HKUSTGZ), Zhong Li (Microsoft Research)

OptimizationComputational EfficiencyRepresentation LearningData-Centric LearningTransformerLarge Language ModelContrastive LearningTextGraph

🎯 What it does: Propose the D3 framework, which improves the training order of LLMs by constraining the training data with a dynamic directional influence graph.

d2: Improving Reasoning in Diffusion Language Models via Trajectory Likelihood Estimation

Guanghan Wang (Cornell University, Cornell Tech), Volodymyr Kuleshov (Cornell University, Cornell Tech)

Computational EfficiencyRepresentation LearningTransformerReinforcement LearningPrompt EngineeringDiffusion modelScore-based ModelTextBenchmarkChain-of-Thought

🎯 What it does: Proposes d2, a reinforcement learning inference framework tailored for masked diffusion language models (DLMs), specifically addressing the trajectory likelihood estimation problem to enhance inference capabilities.

D²Evo: Dual Difficulty-Aware Self-Evolution for Data-Efficient Reinforcement Learning

Ru Zhang (Zhejiang University), Xiangxiang Chu (AMAP, Alibaba Group)

OptimizationData-Centric LearningTransformerLarge Language ModelReinforcement LearningPrompt EngineeringGenerative Adversarial NetworkContrastive LearningTextBenchmark

🎯 What it does: Designed and implemented the D2 Evo framework, which uses self-evolving questioners and solvers to mine medium-difficulty real samples as anchors in each iteration. The questioner is trained to generate diverse questions that match the current capability of the solver, and the solver is further trained using a mixture of generated and anchor data, forming a closed-loop co-evolution process.

d3LLM: Ultra-Fast Diffusion LLM using Pseudo-Trajectory Distillation

Yu-Yang Qian (University of California, San Diego), Hao Zhang (University of California, San Diego)

Computational EfficiencyKnowledge DistillationTransformerLarge Language ModelPrompt EngineeringDiffusion modelScore-based ModelTextChain-of-Thought

🎯 What it does: The paper proposes a d3LLM framework that combines pseudo-trajectory distillation and entropy-driven multi-block decoding to achieve a balance between high parallelism and high accuracy.

DADP: Domain Adaptive Diffusion Policy

Pengcheng Wang (University of California, Berkeley), Yixiao Wang (University of California, Berkeley)

Domain AdaptationTransformerReinforcement LearningDiffusion modelAuto EncoderContrastive LearningTabularTime Series

🎯 What it does: Propose DADP (Domain Adaptive Diffusion Policy), a strategy that achieves zero-shot cross-domain adaptation by using unsupervised decoupled domain representations and injecting domain information into the diffusion generation process.

DAG-MoE: From Simple Mixture to Structural Aggregation in Mixture-of-Experts

Jiarui Feng (Meta MRS), Yixin Chen (Washington University in St. Louis)

Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelMixture of ExpertsTextChain-of-Thought

🎯 What it does: This study investigates replacing traditional weighted sum aggregation with structured aggregation (DAG) in Mixture-of-Experts (MoE) models, proposing a learnable DAG-MoE framework that enhances expressive power and enables multi-step reasoning without altering the experts or router.

DAG: A Dual Correlation Network for Time Series Forecasting with Exogenous Variables

Xiangfei Qiu (East China Normal University), Jilin Hu (East China Normal University)

Autonomous DrivingOptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyData-Centric LearningReinforcement Learning from Human FeedbackNeural Architecture SearchTransformerSupervised Fine-TuningAuto EncoderContrastive LearningTabularTime SeriesSequentialBenchmarkFinance Related

🎯 What it does: Propose a bidirectional correlation network named DAG, specifically designed for time series forecasting by combining historical and future exogenous variables.

DAISI: Data Assimilation with Inverse Sampling using Stochastic Interpolants

Martin Andrae (Linköping University), Fredrik Lindsten (Linköping University)

OptimizationComputational EfficiencyData-Centric LearningDiffusion modelScore-based ModelFlow-based ModelImageVideoTime SeriesPhysics RelatedStochastic Differential Equation

🎯 What it does: Proposes DAISI, an expandable filtering framework based on flow generative models, which integrates predictions into the latent variables of the generative model through an inverse sampling step, and achieves observation fusion via guided conditional sampling.

DAL: A Practical Prior-Free Black-Box Framework for Piecewise Stationary Bandits

Argyrios Gerogiannis (University of Illinois at UrbanaChampaign), Venugopal Veeravalli

Reinforcement LearningTabularTime SeriesBenchmark

🎯 What it does: Proposes a black-box framework called Detection Augmented Learning (DAL) for handling piecewise-stationary bandit problems under unknown non-stationarity; DAL combines any optimal stationary bandit algorithm with change detectors and forced exploration to achieve adaptive learning in non-stationary environments.

DANCE: Dynamic, Available, Neighbor-gated Condensation for Federated Text-Attributed Graphs

Zekai Chen (Beijing Institute of Technology), Guoren Wang (Beijing Institute of Technology)

Federated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningGraph Neural NetworkTransformerLarge Language ModelAuto EncoderContrastive LearningTextGraph

🎯 What it does: Proposes the DANCE framework, introducing graph distillation, neighbor gating, and self-expressive topology reconstruction in federated text attribute graph learning, achieving dynamic, usable, and interpretable graph compression and model refresh.

DAPD: Dependency-Aware Parallel Decoding via Attention for Diffusion LLMs

BumJun Kim, Albert No (Yonsei University)

Computational EfficiencyAI Code AssistantGraph Neural NetworkTransformerLarge Language ModelDiffusion modelText

🎯 What it does: Propose a training-agnostic parallel decoding method called DAPD, which builds a dependency graph using model self-attention and selects an approximately independent set of tokens at each step for simultaneous decoding.

DARC: Disagreement-Aware Alignment via Risk-Constrained Decoding

mingxi Zou, Zenglin Xu (Shanghai Academy of AI for Science)

Reinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringContrastive LearningText

🎯 What it does: Proposes a method called DARC (Disagreement-Aware Alignment via Risk-Constrained Decoding), which performs response selection during the inference phase, aiming to address the heterogeneity and inconsistency of human preferences.

DART: Distribution-Aware Adaptive Relational Transfer for Adversarial Attacks against Closed-Source MLLMs

Kaidi Hu (Shanghai Jiao Tong University), Ruigang Yang (Shanghai Jiao Tong University)

Adversarial AttackGraph Neural NetworkTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodality

🎯 What it does: The study targets goal-oriented adversarial attacks against closed-source multimodal large language models (MLLMs), and proposes an attack framework that can be trained on open-source proxy models and transferred to closed-source target models.

DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning

Yujie Wang (Peking University), Bin CUI

Computational EfficiencyNeural Architecture SearchTransformerLarge Language ModelReinforcement LearningPrompt EngineeringDiffusion modelTextMultimodality

🎯 What it does: Propose the DARTS framework, which significantly accelerates the reinforcement learning training of large language models (LLMs) by actively reshaping distributions, distribution-aware trajectory sampling, and adaptive redundancy allocation, addressing the bottleneck in episode generation caused by long-tailed distributions.

Darwinian Memory: A Training-Free Self-Regulating Memory System for GUI Agent Evolution

Hongze Mi (Didichuxing Co Ltd), Naiqiang Tan (Didichuxing Co Ltd)

Autonomous DrivingOptimizationComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringMixture of ExpertsVision Language ModelImageTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: Designed and implemented a training-free, self-adjusting memory system called DMS to enhance the long-term task performance of multi-modal large language models in GUI automation.

DASH: Faster Shampoo via Batched Block Preconditioning and Efficient Inverse-Root Solvers

Ionut-Vlad Modoranu (Institute of Science and Technology Austria), Dan Alistarh (Institute of Science and Technology Austria)

OptimizationTransformerLarge Language ModelText

🎯 What it does: Propose DASH, an improved distributed Shampoo optimizer, achieving 3D block stacking and more efficient inverse matrix root computation.

Data Agent: Learning to Select Data via End-to-End Dynamic Optimization

Suorong Yang (Nanjing University), Soujanya Poria (Nanyang Technological University)

OptimizationData-Centric LearningRecurrent Neural NetworkTransformerReinforcement LearningAgentic AIPrompt EngineeringImageTextMultimodality

🎯 What it does: Designed the Data Agent framework, which utilizes reinforcement learning to achieve training-aware online dynamic data selection, co-evolving with model weights.

Data Augmentation of Contrastive Learning is Estimating Positive-incentive Noise

Hongyuan Zhang (University of Hong Kong), Xuelong Li (China Telecom)

Information TheoryData SynthesisRepresentation LearningConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningImageTextMultimodalityTabularTime Series

🎯 What it does: Designed and implemented PiNDA, a contrastive learning framework that automatically generates augmented noise by learning Positive-incentive Noise (π-Noise).

Data Difficulty and the Generalization–Extrapolation Tradeoff in LLM Fine-Tuning

Siyuan Liu (Tsinghua University), Jingzhao Zhang (Tsinghua University)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningTextChain-of-Thought

🎯 What it does: Studied the impact of training data difficulty and data scale on the performance of large language model supervised fine-tuning (SFT), conducted systematic experiments and provided theoretical explanations.

Data Provenance Auditing of Fine-Tuned Large Language Models with a Text-Preserving Technique

Yanming Li (Inria), Seifeddine Ghozzi (Institut Polytechnique de Paris)

Safty and PrivacyExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringText

🎯 What it does: A technique is proposed to construct 'prompt/response' pairs using invisible Unicode characters for watermarking, which can detect whether training data has been used through black-box interaction after fine-tuning large language models, and control the false positive rate through ranking tests.

Data Reconstruction: Identifiability and Optimization with Sample Splitting

Yujie Shen (Tsinghua University), Qi Lei (New York University)

Data SynthesisOptimizationExplainability and InterpretabilityConvolutional Neural NetworkImage

🎯 What it does: The study investigates the identifiability and optimization methods for recovering original training data from trained neural network parameters, proposing a sample splitting technique to enhance reconstruction quality.

Data Selection for Fine-tuning Vision Language Models via Cross Modal Alignment Trajectories

Nilay Naharas (University of California Los Angeles), Baharan Mirzasoleiman (Google)

Computational EfficiencyData-Centric LearningTransformerSupervised Fine-TuningVision Language ModelContrastive LearningImageTextMultimodality

🎯 What it does: Proposes the XMAS data selection method for LVLM based on cross-modal attention trajectories, improving the efficiency of instruction fine-tuning by clustering attention trajectories and sampling balanced subsets.

Data- and Variance-dependent Regret Bounds for Online Tabular MDPs

Mingyi Li (University of Tokyo), Kenji Yamanishi (University of Tokyo)

OptimizationReinforcement LearningTabular

🎯 What it does: This paper designs two types of best-of-both-worlds algorithms, global optimization and policy optimization, for finite-layer table Markov decision processes with known transition probabilities. These algorithms provide data-oriented improvements to the regret upper bounds in both adversarial and random environments.

Data-driven Mixed Integer Optimization through Probabilistic Multi-variable Branching

Yanguang Chen (Shanghai University of Finance and Economics), Yinyu Ye (Shanghai Jiao Tong University)

OptimizationGraph Neural NetworkContrastive LearningTabularBenchmark

🎯 What it does: This paper proposes a hybrid integer programming acceleration method based on probabilistic multivariate branching (PMVB).

Data-Source Adaptive Online Learning under Heteroscedastic Noise

Amith Bhat Hosadurga Anand (University of Illinois Chicago), Aadirupa Saha (University of Illinois Chicago)

Recommendation SystemOptimizationReinforcement LearningTextTabular

🎯 What it does: A multi-armed bandit model with multi-source heteroscedastic noise is proposed, and the SOAR algorithm is designed to dynamically select low-variance sources and optimal arms without knowing the source variances, achieving online learning.

DataGuard: A Non-intrusive Dataset Auditing Framework via Differential Information Forensics

Jiadong Lou (Rowan University), Xu Yuan (University of Delaware)

Anomaly DetectionSafty and PrivacyData-Centric LearningConvolutional Neural NetworkTransformerContrastive LearningImage

🎯 What it does: Proposes DataGuard, a non-intrusive dataset auditing framework that utilizes differential information forensics and statistical testing to identify whether a model has used a target dataset.

Dataset Distillation Efficiently Encodes Low-Dimensional Representations from Gradient-Based Learning of Non-Linear Tasks

Yuri Kinoshita (University of Tokyo), Taro Toyoizumi (University of Tokyo)

Computational EfficiencyKnowledge DistillationRepresentation LearningConvolutional Neural NetworkRecurrent Neural NetworkContrastive LearningImageTabular

🎯 What it does: This paper provides a theoretical analysis of the gradient learning process of dataset distillation on two-layer ReLU networks and multi-index models, proving that distilled data can efficiently encode the low-dimensional structure of tasks, and gives an upper bound on memory complexity of ˜(Θ(r²d + L));

DAVE: Distribution-Aware Attribution via ViT Gradient Decomposition

Adam Wróbel (Jagiellonian University), Dawid Damian Rymarczyk (Jagiellonian University)

ClassificationExplainability and InterpretabilityTransformerContrastive LearningImage

🎯 What it does: Proposes DAVE, a distribution-aware attribution method for Vision Transformers, which eliminates structural artifacts through gradient decomposition, achieving high-resolution and stable pixel-level explanations.

daVinci-Dev: Agent-native Mid-training for Software Engineering

Ji Zeng (Shanghai Jiao Tong University), Pengfei Liu (Generative Ai Research Lab)

AI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningAgentic AIPrompt EngineeringTextSequentialRetrieval-Augmented Generation

🎯 What it does: This paper proposes the 'Agent-native Mid-Training' method in the field of software engineering, which involves introducing specially constructed agent-native data during the mid-stage of large model training;

DB-KSVD: Scalable Alternating Optimization for Disentangling High-Dimensional Embedding Spaces

Romeo Valentin (Stanford University), Mykel Kochenderfer

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelAuto EncoderContrastive LearningImageText

🎯 What it does: Proposed an scalable dictionary learning algorithm called DB‑KSVD for disentangling monosemantic features from large transformer embeddings.

DC-LA: Difference-of-Convex Langevin Algorithm

Hoang Phuc Hau Luu (Nanyang Technological University), Zhongjian Wang (Nanyang Technological University)

OptimizationDiffusion modelScore-based ModelImageBiomedical DataComputed TomographyStochastic Differential Equation

🎯 What it does: Proposed a forward-backward Langevin sampling algorithm called DC-LA, which can sample when the potential energy of the target distribution is composed of a Lipschitz smooth term plus a difference-of-convex (DC) non-smooth regularization, and provided its Wasserstein distance convergence upper bound under remote dissipativity conditions;

DC-Leap: Training-Free Acceleration of dLLMs via Draft-Guided Contiguous Leaping Decoding

Yanhua Jiao (Harbin Institute of Technology), Min Zhang (Harbin Institute of Technology)

GenerationComputational EfficiencyTransformerLarge Language ModelDiffusion modelText

🎯 What it does: Propose a training-agnostic parallel decoding framework called DC-Leap, which achieves high-speed inference in diffusion large language models through dynamic continuous verification and draft-guided leap decoding.

DC-W2S: Dual-Consensus Weak-to-Strong Training for Reliable Process Reward Modeling in Biological Reasoning

Chi-Min Chan (Hong Kong University of Science and Technology), Gabriele Scalia (Genentech)

Drug DiscoveryTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningTextBiomedical DataChain-of-Thought

🎯 What it does: Train a bio-reasoning system based on a process reward model (PRM), learning the correctness of reasoning steps using a large amount of noisy weak labels rather than manually annotated expert labels.

DDGA: Dirichlet Distributional Gradient Aggregation for Transferable Vision-Language Adversarial Attacks

Yiwei You (University of International Business and Economics), Bo Wang (University of International Business and Economics)

GenerationRetrievalAdversarial AttackTransformerVision Language ModelDiffusion modelContrastive LearningImageTextMultimodality

🎯 What it does: Proposes the DDGA framework, which explicitly learns the Dirichlet distribution on the simple shape of the Adversarial Evolution Triangle (AET), directly optimizing the expected adversarial objective instead of random sampling;

DDIM Inversion as a Perturbation Amplifier: Breaking Mimicry Protection via Reconstruction Error Minimization

Huming Qiu (Fudan University), Min Yang (Fudan University)

GenerationSafty and PrivacyTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningImage

🎯 What it does: This paper studies the amplification effect of minor protective perturbations during the DDIM inversion process, and proposes a perturbation removal method called DIRP based on minimizing the DDIM reconstruction error, which is used to eliminate imitative protective perturbations in image generation models;

DDP-WM: Disentangled Dynamics Prediction for Efficient World Models

Shicheng Yin (Sun Yat-sen University), Liang Lin (Sun Yat-sen University)

OptimizationComputational EfficiencyRobotic IntelligenceReinforcement Learning from Human FeedbackConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningWorld ModelOptical FlowImageVideoPoint Cloud

🎯 What it does: Propose a world model called DDP-WM, which achieves efficient visual dynamic prediction by separating sparse main dynamics from background context updates.

DDSVM: A Differentiable Framework for Deep Support Vector Machines with Iterative Geometry-Aware Optimization

Yirun Ding (Shenzhen University), Zhihui Lai (Shenzhen University)

ClassificationOptimizationComputational EfficiencyConvolutional Neural NetworkContrastive LearningImage

🎯 What it does: Proposed a differentiable deep support vector machine framework called DDSVM, which dynamically guides feature learning by alternatingly training SVM and neural networks.

De-attribute to Forget for LLM Unlearning

Xinyang Lu (National University of Singapore), Bryan Kian Hsiang Low (National University of Singapore)

OptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringContrastive LearningText

🎯 What it does: Proposed a learning-free framework called DareU based on data de-attribution for LLMs, which uses reinforcement learning to de-attribute model-generated responses and eliminate the influence of forgotten data.

De-Linearizing Agent Traces: Bayesian Inference of Latent Partial Orders for Efficient Execution

Dongqing li, Quyu Kong

OptimizationComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringTextGraph

🎯 What it does: De-linearize the linear execution logs generated by LLM agents, infer their underlying partial order graph (i.e., concurrency and precedence dependencies), and compile this graph into an efficient frontier-based executor, reducing redundant reasoning and token consumption.

De4D-SLAM: Gradient-Isolated Static-Dynamic Decoupling for Monocular SLAM in Dynamic Environments

Zhicheng Fan (Xidian University), Bo REN

Autonomous DrivingOptimizationSafty and PrivacyComputational EfficiencyTransformerContrastive LearningGaussian SplattingSimultaneous Localization and MappingOptical FlowImageVideoPoint Cloud

🎯 What it does: Proposed a novel monocular dynamic SLAM framework called De4D-SLAM, which can simultaneously perform localization and complete 4D (spatiotemporal) reconstruction in dynamic environments;

Debate with Images: Detecting Deceptive Behaviors in Multimodal Large Language Models

Sitong Fang (Peking University), Jiaming Ji (Peking University)

Anomaly DetectionSafty and PrivacyExplainability and InterpretabilityTransformerLarge Language ModelAgentic AIPrompt EngineeringVision Language ModelVision-Language-Action ModelDiffusion modelContrastive LearningImageTextMultimodalityBenchmarkChain-of-Thought

🎯 What it does: Propose the MM-DeceptionBench benchmark and the Debate-with-Images multi-agent visual adversarial evaluation framework for detecting deceptive behaviors in multi-modal large language models.

Debate2Create: Robot Co-design via Multi-Agent LLM Debate

Kevin Qiu (University of Warsaw), Marek Cygan (University of Warsaw)

OptimizationRobotic IntelligenceReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringTextTabular

🎯 What it does: Propose a robot co-design framework DEBATE2CREATE (D2C) based on multi-agent LLM debate, which simultaneously optimizes robot morphology and reward functions through structured argumentation and physical simulation evaluation.

Debiased Model-based Representations for Sample-efficient Continuous Control

Jiafei Lyu (Tencent Hunyuan), Deheng Ye (Tencent Hunyuan)

Convolutional Neural NetworkRecurrent Neural NetworkTransformerReinforcement LearningContrastive LearningTabularTime Series

🎯 What it does: In continuous control reinforcement learning, sample efficiency is improved by learning model-based representations and combining them with experience replay.

DecAEvolve: Decompose, Adapt, and Evolve for Effective LLM-based Scientific Equation Discovery

Pouya Behzadifar (Sharif University of Technology), Chandan K. Reddy

OptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningAuto EncoderTabularTime SeriesSequentialPhysics Related

🎯 What it does: Combining LLM generative capabilities, symbolic decomposition, reinforcement learning adaptation, and evolutionary search to achieve scientific equation discovery.

Decentralized and Disentangled Task–Role Representation Learning for Generalizable Offline Multi-Agent Meta Reinforcement Learning

Lei Yuan (Nanjing University), Yang Yu (Nanjing University)

Federated LearningKnowledge DistillationRepresentation LearningMeta LearningTransformerLarge Language ModelReinforcement LearningContrastive LearningTabularSequential

🎯 What it does: This study proposes the DTR framework, which addresses offline multi-task multi-agent meta-reinforcement learning by designing a decentralized and decoupled task and role representation learning method to enhance generalization to unseen tasks.

Decentralized Bandits without Global Clock for Dynamic Matching Market

Mengtong Gao (Tsinghua University), Jing Chen (Tsinghua University)

OptimizationFederated LearningReinforcement LearningTabularTime SeriesBenchmark

🎯 What it does: In a fully decentralized, dynamic bilateral matching market without a global clock, two algorithms are proposed: one-sided learning Way-SE and two-sided learning Way-SE-2S, which are proven to achieve sublinear regret (i.e., converge to players' optimal stable matching) under any sequence of player arrivals and departures.

Decentralized Instruction Tuning: Conflict-Aware Splitting and Weight Merging

MINSIK CHOI, Geewook Kim (NAVER Cloud AI)

Federated LearningComputational EfficiencyKnowledge DistillationTransformerSupervised Fine-TuningPrompt EngineeringMixture of ExpertsTextMultimodality

🎯 What it does: Propose a decentralized instruction fine-tuning pipeline called MERIT, which first estimates gradient conflicts for each task at a shared 'mergeable' initialization point, uses PCA to reduce the dimensionality of the conflict structure, recursively partitions the task set along the main conflict axis, independently fine-tunes on non-interactive branches, and finally merges the model once using token-weighted averaging.

Decentralized Online Convex Optimization with Efficient Communication: Improved Algorithm and Lower Bounds

Sifan Yang (Nanjing University), Lijun Zhang (Nanjing University)

OptimizationFederated Learning

🎯 What it does: Study algorithms for compressed communication in distributed online convex optimization, and provide better regret bounds.

DecepChain: Inducing Deceptive Reasoning in Large Language Models

Wei Shen (University of Illinois Urbana-Champaign), Huan Zhang (University of Illinois Urbana-Champaign)

Explainability and InterpretabilityAdversarial AttackReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextBenchmarkChain-of-Thought

🎯 What it does: This paper studies how to induce large language models to generate surface-credible but ultimately erroneous chains of reasoning, and proposes DecepChain as an induction framework.

DecFus: Decentralized Layer-wise Fusion with Dynamic Exploration and Exploitation

Li Yang (Durham University), Bo Liu (Shenzhen University of Advanced Technology)

OptimizationFederated LearningComputational EfficiencyConvolutional Neural NetworkContrastive LearningImage

🎯 What it does: Propose DecFus, a framework that unifies layer-level exchange and averaging in decentralized federated learning, to dynamically balance exploration and exploitation, thereby improving model convergence and generalization.

Decision Transformers As Zero-Shot Learners via Text-Behavior Alignment

Xin Zhang (San Diego State University), Yingxue Zhang (Binghamton University)

Meta LearningTransformerLarge Language ModelReinforcement LearningPrompt EngineeringContrastive LearningTextSequential

🎯 What it does: Proposes a Text-Guided Decision Transformer (TG-DT), achieving zero-shot task generalization in offline meta reinforcement learning through natural language task descriptions.

Decision Tree Learning on Product Spaces

Arshia Soltani Moakhar (University of Maryland), MohammadTaghi Hajiaghayi (University of Maryland)

ClassificationOptimizationExplainability and InterpretabilityComputational EfficiencyTabular

🎯 What it does: This paper extends the theoretical analysis of decision tree learning, generalizing the classic top-down greedy heuristic method from uniform distributions to arbitrary product distributions, providing broader theoretical guarantees.

Decision-Focused Learning via Tangent-Space Projection of Prediction Error

Junhyeong Lee (Ulsan National Institute of Science and Technology), Yongjae Lee (Ulsan National Institute of Science and Technology)

OptimizationFederated LearningComputational EfficiencyReinforcement Learning from Human FeedbackDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningTabularTime SeriesSequentialBenchmarkFinance Related

🎯 What it does: Propose a new decision-focused learning method called PEAR, which directly obtains the regret gradient by projecting the prediction error onto the tangent space of active constraints, thereby training the predictor.

Decision-focused Sparse Tangent Portfolio Optimization

Haeun Jeon (Korea Advanced Institute of Science and Technology), Woo Chang Kim (Korea Advanced Institute of Science and Technology)

OptimizationTabularTime SeriesFinance Related

🎯 What it does: Propose an end-to-end decision-focused learning framework that achieves Sharpe ratio maximization under cardinality constraints through differentiable sparse tangential portfolio optimization.

DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile Adapter

Xukun Li (XYZ Embodied AI), Zhenguo Sun (XYZ Embodied AI)

Robotic IntelligenceTransformerReinforcement LearningDiffusion modelImageVideoMultimodalityTime Series

🎯 What it does: Propose the DECO framework, which utilizes a decoupled multi-modal diffusion Transformer to achieve dexterous bimanual manipulation, along with a plug-in tactile adapter.

DeCoDe: Decoupling Binding Position and Molecular Conformation in 3D Ligand Diffusion for Structure-Based Drug Design

Julong Yang (Sichuan University), Jian Peng (Sichuan University)

Drug DiscoveryGraph Neural NetworkTransformerDiffusion modelScore-based ModelGraphBiomedical Data

🎯 What it does: Proposed the DeCoDe framework, which decouples the diffusion process of the binding position of the molecular ligand within the protein binding pocket from the internal conformation of the molecule;

DecoderTCR: Compositional Pretraining and Entropy-Guided Decoding for TCR-pMHC Interactions

Boqiao Lai (Biohub), Aly A Khan

Drug DiscoveryProtein Structure PredictionTransformerLarge Language ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningSequentialBiomedical Data

🎯 What it does: Developed a sequence model called DecoderTCR for TCR-pMHC interactions, capable of achieving zero-shot binding prediction and TCR design.

DecodeShare: Tracing the Shared Subspace of LLM Decode-Time Decisions

Zishan Shao (Duke University), Hai Helen Li (Duke University)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextBenchmark

🎯 What it does: Propose the DECODESHARE protocol, which identifies and intervenes in the low-dimensional subspace shared across tasks during the KV-cache decoding phase of LLMs.

Decompose and Recompose: Reasoning New Skills from Existing Abilities for Cross-Task Robotic Manipulation

Xitie Zhang (Tianjin University), Yahong Han (Tianjin University)

Robotic IntelligenceReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringVision Language ModelVision-Language-Action ModelImageTextMultimodalityRetrieval-Augmented Generation

🎯 What it does: Propose a framework based on skill splitting and recombination, utilizing atomic skills and action pairs from original demonstrations for zero-shot cross-task robot manipulation.

Decompose, Structure, and Repair: A Neuro-Symbolic Framework for Autoformalization via Operator Trees

Xiaoyang Liu (Shanghai Jiao Tong University), Tao Luo (Shanghai Jiao Tong University)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringMixture of ExpertsTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Decompose natural language propositions into logical components, map them to Lean code and corresponding operator trees, and utilize the tree structure for precise error localization and repair, thereby achieving automated formalization.

DecomPose: Disentangling Cross-Category Optimization Contention for Category-Level 6D Object Pose Estimation

Yifan Gao (Wuhan Institute of Technology), Guoping Wang (Peking University)

Pose EstimationOptimizationConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelRectified FlowNeural Radiance FieldAuto EncoderContrastive LearningGaussian SplattingImageMultimodalityPoint Cloud

🎯 What it does: The paper proposes the DecomPose framework, which addresses the cross-category optimization competition problem in category-level 6D object pose estimation. It first quantifies the competition using gradient diagnosis, and then achieves module-level decoupling by implementing difficulty-aware static grouping and asymmetric branches corresponding to the modules, thereby improving the stability of multi-class joint learning.