ICML 2026 Papers — Page 16
International Conference on Machine Learning · 6554 papers
Distribution Alignment for One-Shot Federated Learning via Optimal Transport
Daniele Berardini (Italian Institute of Technology), Vittorio Murino (University of Verona)
Domain AdaptationFederated LearningTransformerContrastive LearningImageMultimodality
🎯 What it does: Propose a single-round feature alignment framework called SLOT-ALIGN based on optimal transport, aimed at addressing feature distribution mismatch caused by domain shift and label shift in single-round federated learning.
Distribution Matching Variational AutoEncoder
Sen Ye (Peking University), Han Hu (Tencent)
GenerationRepresentation LearningTransformerDiffusion modelScore-based ModelFlow-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningImage
🎯 What it does: Propose Distribution Matching VAE (DMVAE), which explicitly aligns the aggregated posterior of the autoencoder to any reference distribution through distribution matching constraints, thereby enabling controllable shaping of the latent space structure.
Distribution Transformers: Fast Approximate Bayesian Inference With On-The-Fly Prior Adaptation
George Whittle (University of Oxford), Michael A Osborne
OptimizationComputational EfficiencyRepresentation LearningTransformerMixture of ExpertsScore-based ModelContrastive LearningGaussian SplattingTabularTime SeriesSequentialPhysics Related
🎯 What it does: Propose a Distribution Transformer based on Transformer, which can map any prior (represented as a Gaussian Mixture Model, GMM) to a posterior (also a GMM) in a single forward pass, thereby enabling approximate Bayesian inference and supporting the modification of the prior at any time.
Distribution-Calibrated Inference Time Compute for Thinking LLM-as-a-Judge
Hamid Dadkhahi (Google), Mehdi Mirzazadeh (Google Research)
Recommendation SystemExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: The paper proposes an inference-time computation (ITC) aggregation scheme for thinking large language models (Thinking LLMs) serving as judges. After generating multiple independent thinking-evaluation samples for each item to be assessed, the method employs a distribution calibration aggregation approach based on the Bradley-Terry-Davidson (BTD) model to obtain more reliable preference judgments.
Distributional Active Inference
Abdullah Akgül (University Of Southern Denmark), Melih Kandemir (University Of Southern Denmark)
TransformerReinforcement LearningAuto EncoderImageTabularTime Series
🎯 What it does: Propose a framework that integrates Active Inference (AI) with Distributional Reinforcement Learning (Distributional RL), named Distributional Active Inference (DAIF), which achieves model-free transfer learning through the push-forward process, reducing reliance on forward dynamic models;
Distributional Alignment Games for Answer-Level Fine-Tuning
Mehryar Mohri (Google Research), Yifan Wu (Microsoft Research)
OptimizationData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringContrastive LearningText
🎯 What it does: Proposes a Distributional Alignment Games framework, transforming the answer layer fine-tuning (ALFT) problem into a two-player game between strategy and target distribution, thereby avoiding the need for massive summation over implicit reasoning paths;
Distributional Inverse Reinforcement Learning
Feiyang Wu (Georgia Institute of Technology), Anqi Wu (Georgia Institute of Technology)
Reinforcement LearningTabularTime SeriesBiomedical Data
🎯 What it does: This paper proposes a distributed inverse reinforcement learning (DistIRL) framework that learns the full distribution of the reward function and trains risk-sensitive policies based on this distribution.
Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook
Jaehyeok Lee (Sungkyunkwan University), JinYeong Bak (Sungkyunkwan University)
OptimizationExplainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelContrastive LearningText
🎯 What it does: This paper proposes the DOVE framework for distributed evaluation of cultural value orientations in large language models.
Distributionally Robust Causal Abstractions
Yorgos Felekis (University of Warwick), Paris Giampouras (University of Warwick)
Domain AdaptationOptimizationFederated LearningExplainability and InterpretabilityImageTabular
🎯 What it does: Proposed a distributed robust causal abstraction framework (pρ,ιq abstraction) and its learning algorithm DIROCA, which maintains intervention consistency across different levels of causal models and is robust to environmental uncertainty.
Distributionally Robust Markov Games with Average Reward
Zachary Andrew Roch (University of Central Florida), Yue Wang (University of Central Florida)
OptimizationReinforcement Learning
🎯 What it does: Proposed and analyzed distributed robust Markov games (DR-MGs) under average reward settings, proving the existence of steady-state Nash equilibria under two structural conditions: irreducibility and weak communication, and provided two convergence algorithms (Robust Nash Iteration and Robust TD Descent).
Distributionally Robust Reinforcement Learning from Human Feedback
Debmalya Mandal (University of Warwick), Goran Radanovic
Domain AdaptationOptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningContrastive LearningText
🎯 What it does: Designed and implemented a distributed robust reinforcement learning (DRO) framework for reward model learning and policy optimization in human feedback (RLHF), enhancing model robustness under distribution shift.
Distributionally Robust Set Representation Learning Under Inference-Time Element Corruption
Yankai Chen (MBZUAI), Xue Liu (MBZUAI)
Data SynthesisOptimizationRepresentation LearningDiffusion modelScore-based ModelContrastive LearningImageTextMultimodalityPoint Cloud
🎯 What it does: Proposes a distributed robust set representation learning framework called SW-DRSO based on the Sliced-Wasserstein distance, which can perform robust optimization against sparse element corruptions during inference.
DITING: A Weak Degradation Listener for Battery Lifetime Early Prediction
Hao Miao (Taiyuan University of Technology), Li Wang (Taiyuan University of Technology)
Anomaly DetectionComputational EfficiencyRepresentation LearningRecurrent Neural NetworkTransformerDiffusion modelScore-based ModelContrastive LearningTabularTime Series
🎯 What it does: Developed a weak degradation listener called DITING based on the binaural effect, used to extract weak degradation signals from early cycle battery data and predict lifespan in noisy environments.
DITRON: Distributed Multi-level Tiling Compiler for Parallel Tensor Programs
Size Zheng (ByteDance Seed), Xin Liu (ByteDance Seed)
OptimizationComputational EfficiencyLarge Language ModelText
🎯 What it does: Propose DITRON, a distributed multi-level block compiler that adopts a three-tier block abstraction (core/device/task), achieving high-performance distributed execution of tensor programs.
DIVA: Harnessing the Representation Divergence in Unified Multimodal Models for Mutual Reinforcement
Renjie Lu (Ping An Technology (Shenzhen) Co., Ltd.), Shangfei Wang (University of Science and Technology of China)
GenerationData SynthesisRepresentation LearningTransformerSupervised Fine-TuningVision Language ModelDiffusion modelContrastive LearningImageTextMultimodality
🎯 What it does: This paper proposes DIVA, a post-training framework designed to eliminate mutual interference caused by the biases of the understanding branch and the generation branch in unified multi-modal models (UMMs), and to achieve complementary improvement by decomposing shared and unique information.
Dive into the Scene: Breaking the Perceptual Bottleneck in Vision-Language Decision Making via Focus Plan Generation
Boyuan Xiao (Zhejiang University), Kun Zhou (Zhejiang University)
Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningRobotic IntelligenceTransformerPrompt EngineeringVision Language ModelVision-Language-Action ModelContrastive LearningImageTextGraph
🎯 What it does: This paper proposes the SceneDiver method, which utilizes scene graphs to generate coarse-to-fine hierarchical attention plans for VLM, significantly reducing visual hallucinations and attention drift in embodied vision-language decision-making; and designs a lightweight adapter to transfer this attention capability to VLA, achieving real-time responsive control.
DIVE: Scaling Diversity in Agentic Task Synthesis for Generalizable Tool Use
Aili Chen (Fudan University), Yanghua Xiao (Fudan University)
Recommendation SystemAutonomous DrivingOptimizationFederated LearningRobotic IntelligenceMeta LearningDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AIPrompt EngineeringTextBenchmarkFinance RelatedRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Developed the DIVE framework, which generates executable and verifiable agentic training data by first executing diverse real-world tools, collecting evidence, and then reverse-engineering tasks, thereby enhancing the generalization ability of LLMs in tool usage.
DIVER: Diving Deeper into Distilled Data via Expressive Semantic Recovery
Qianxin Xia (University of Electronic Science and Technology of China), Guoming Lu (University of Electronic Science and Technology of China)
Data SynthesisKnowledge DistillationConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningImage
🎯 What it does: A two-stage framework called DIVER is proposed based on data set distillation (DD). In the second stage, a pre-trained diffusion model is used to recover the semantics of the distilled data, thereby enhancing cross-architecture generalization ability.
Divergence Decoding: Inference-Time Unlearning via Auxiliary Models
Humzah Merchant (University of Chicago), Bradford Levy (University of Chicago)
Federated LearningSafty and PrivacyComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringMixture of ExpertsImageText
🎯 What it does: Propose the Divergence Decoding (DD) mechanism, which corrects the logits of the target model during the inference phase through two small auxiliary models, thereby achieving 'forgetting' of specific training data.
Diversity Matters: Revisiting Test-Time Compute in Vision-Language Models
Yijie Tong (ETH Zuerich), Mrinmaya Sachan (ETH Zuerich)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerMixture of ExpertsVision Language ModelContrastive LearningImageTextMultimodalityBenchmarkChain-of-Thought
🎯 What it does: Systematically studied the effect of test-time computation (TTC) on vision-language models (VLMs), proposed an entropy-based TTC (ETTC) method, and conducted experiments on seven VLMs and six multiple-choice visual reasoning datasets.
Diversity Over Frequency: Rethinking Tool Use in Visual Chain-of-Thought Agents
Dong-Hee Kim (Korea University), Donghyun Kim (Korea University)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerReinforcement LearningPrompt EngineeringVision-Language-Action ModelImageTextMultimodalityBenchmarkChain-of-Thought
🎯 What it does: Investigate the tool usage behavior of visual chain-of-thought agents in more complex visual reasoning tasks, and reveal the tool usage collapse phenomenon through experiments.
Diversity-Aware Recursive Feature Multiple Kernel Learning
Nan Cao (Huazhong University of Science and Technology), Teng Zhang (Huazhong University of Science and Technology)
ClassificationAuto EncoderContrastive LearningGaussian SplattingOptical FlowTabularBenchmarkStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Proposed a multi-kernel learning method called DARFMMKL, which combines the Recursive Feature Machine (RFM) kernel family with a kernel selector that simultaneously optimizes diversity and quality, enabling the learning of feature importance from data and effectively selecting diverse base kernels.
Diversity-aware Weight Perturbation Promotes Robust Adaptation
Zibo Chen (Harbin Institute of Technology (Shenzhen)), Zilu Wang (Harbin Institute of Technology (Shenzhen))
Domain AdaptationComputational EfficiencyConvolutional Neural NetworkSupervised Fine-TuningContrastive LearningImageTextMultimodalityBenchmarkPhysics RelatedAudio
🎯 What it does: Proposes Diversity-aware Weight Perturbation (DWP), a training method inspired by the immune system's affinity selection, to enhance the robustness of deep neural networks against weight noise on Compute-In-Memory (CIM) edge accelerators.
Diversity-Driven Offline Multi-Objective Optimization via Nested Pareto Set Learning
Yiyi Zhu (East China Normal University), Hong Qian (East China Normal University)
OptimizationTabularBenchmark
🎯 What it does: Proposes the DOMOO method, which combines diversity-driven offline multi-objective optimization with nested Pareto set learning (NPSL), to address the problem of reduced diversity and convergence caused by OOD errors.
Diversity-Preserved Distribution Matching Distillation for Fast Visual Synthesis
Tianhe Wu (City University of Hong Kong), Kede Ma (City University of Hong Kong)
GenerationData SynthesisKnowledge DistillationTransformerDiffusion modelScore-based ModelFlow-based ModelGenerative Adversarial NetworkImageText
🎯 What it does: Proposed a distribution matching distillation method called DP-DMD based on role separation, specifically designed for simultaneous preservation of sample diversity and visual quality in few-step image generation.
Divide and Conquer: Reliable Multi-View Evidential Learning for Deepfake Detection
Xiaolu Kang (Wuhan University), Qian Wang (Wuhan University)
Anomaly DetectionTransformerSupervised Fine-TuningVision Language ModelAuto EncoderContrastive LearningImageVideo
🎯 What it does: Propose a depth-forgery detection framework called DiCoME based on the 'divide and conquer' strategy, which achieves reliable detection by separating the semantic perspective from the structural perspective and fusing multi-perspective evidence.
Divide and Contrast: Learning Robust Temporal Features without Augmentation
Abdul-Kazeem Shamba (Norwegian University of Science and Technology), Gavin Taylor (United States Naval Academy)
ClassificationRetrievalAnomaly DetectionRepresentation LearningRecurrent Neural NetworkTransformerAuto EncoderContrastive LearningTabularTime Series
🎯 What it does: Proposes Di-COT, a temporal representation learning framework that does not rely on data augmentation or multiple encoder passes, learning robust temporal features by randomly partitioning overlapping sub-blocks within each window and contrasting neighboring sub-blocks.
Divide and Learn: Multi-Objective Combinatorial Optimization at Scale
Esha Singh (University of California San Diego), Yian Ma
OptimizationMixture of ExpertsGraphTabularBenchmark
🎯 What it does: Proposes Divide and Learn (D&L), a multi-objective combinatorial optimization framework based on decomposition and multi-expert online learning;
Divide-and-Denoise: A Game-Theoretic Method for Fairly Composing Diffusion Models
Abhi Gupta (MIT), Tommi Jaakkola (MIT)
GenerationData SynthesisOptimizationTransformerPrompt EngineeringMixture of ExpertsDiffusion modelImageTextMultimodality
🎯 What it does: Propose a method called Divide‑and‑Denoise, which fairly divides the game dynamics during the sampling process of diffusion models to allocate workloads, guiding pre-trained models to denoise in corresponding regions, thus achieving collaborative generation among multiple models.
Diving into Kronecker Adapters: Component Design Matters
Jiayu Bai (Huazhong University of Science and Technology), Zenan Ling (Huazhong University of Science and Technology)
ClassificationOptimizationComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningImageText
🎯 What it does: This paper conducts a fine-grained analysis of the component design of Kronecker adapters and proposes Component Designed Kronecker Adapters (CDKA) based on theoretical derivations, providing design guidelines and training stabilization strategies under a fixed parameter budget.
Divisiveness-Consistent Label Distribution Learning
Yunan Lu (Hong Kong Polytechnic University), Xiuyi Jia (Nanjing University of Science and Technology)
ClassificationRepresentation LearningContrastive LearningImageTextMultimodalityAudio
🎯 What it does: Proposes a splitting-consistent label distribution learning framework that quantifies and preserves the polarity splitting information in label distributions.
DIYHealth Suite: Dataset, Model, and Benchmark for Health Management at Home
Changshuo Liu (National University of Singapore), Beng Chin Ooi (Zhejiang University)
Data SynthesisTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsImageTextMultimodalityBiomedical DataBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose DIYHealth Suite, which includes a large-scale home-use multimodal dataset DIYHealth-900K, an adaptive foundation model for home health management called DIYHealthGPT, and a unified evaluation benchmark DIYHealthBench.
DLEBench: Evaluating Small-scale Object Editing Ability for Instruction-based Image Editing Model
Shibo Hong (Fudan University), Yixin Cao (Fudan University)
Image TranslationImage HarmonizationTransformerLarge Language ModelPrompt EngineeringDiffusion modelGenerative Adversarial NetworkContrastive LearningImageTextMultimodalityBenchmark
🎯 What it does: Constructed a benchmark called DLEBench focused on small-scale object editing, and proposed an evaluation method based on objective scoring criteria and a dual-mode evaluation framework.
dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive Caching
Zhiyuan Liu (Shanghai Jiao Tong University), Linfeng Zhang (Shanghai Jiao Tong University)
Computational EfficiencyTransformerLarge Language ModelDiffusion modelText
🎯 What it does: Proposes dLLM-Cache, a no-training, model-agnostic adaptive caching framework for accelerating inference in diffusion-based large language models;
DLLMQuant: A Post-Training Quantization Framework Tailored for Diffusion-Based Large Language Models
Chen Xu (Houmo AI), Dawei Yang (Houmo AI)
GenerationComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelDiffusion modelContrastive LearningText
🎯 What it does: A post-training quantization (PTQ) framework called DLLMQuant is proposed for diffusion-based large language models (DLLMs), addressing the issue of a sharp performance drop of traditional PTQ on DLLMs.
DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders
Xu Wang (University of Hong Kong), Difan Zou (University of Hong Kong)
Explainability and InterpretabilityTransformerLarge Language ModelDiffusion modelAuto EncoderText
🎯 What it does: Propose DLM-Scope, a mechanism interpretation framework based on sparse autoencoders (SAE), used to parse the internal representations of diffusion language models (DLM) and achieve controllable interventions.
DLM: Unified Decision Language Model for Offline Multi-Agent Sequential Decision Making
Zhuohui Zhang (Tongji University), Bin He (Tongji University)
TransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningTextSequential
🎯 What it does: Proposed a unified decision language model (DLM) that achieves sequential decision-making for offline multi-agent systems through conversational sequence modeling.
DLO-Lab: Benchmarking Deformable Linear Object Manipulations with Differentiable Physics
Junyi Cao (University of Massachusetts Amherst), Chuang Gan (University of Massachusetts Amherst)
OptimizationTransformerReinforcement LearningVision Language ModelDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageVideoPoint CloudMeshBenchmarkPhysics Related
🎯 What it does: Developed a differentiable DLO simulator and built the DLO-Lab benchmark (10 tasks) based on it, designed a specialized DLO agent for grasp point proposal and task decomposition, evaluated multiple RL/optimization methods, and verified zero-shot sim-to-real.
DMCO: Budget-Aware Co-Optimization of Data Cleaning and AutoML
Xiaoou Ding (Harbin Institute of Technology), Jianmin Wang (Tsinghua University)
OptimizationData-Centric LearningTabular
🎯 What it does: Under a fixed budget, a unified DMCO framework is proposed, jointly optimizing data cleaning and AutoML, and splitting the two into time slices executed in an interleaved manner.
DNA: Uncovering Universal Latent Forgery Knowledge
Jingtong Dou (University of Sydney), Tat-Seng Chua (National University of Singapore)
Anomaly DetectionConvolutional Neural NetworkTransformerSupervised Fine-TuningDiffusion modelContrastive LearningImageBenchmark
🎯 What it does: The DNA framework is proposed to detect forged images by mining sparse 'forgery discriminative units' (FDUs) from pre-trained visual models.
DNACHUNKER: Learnable Tokenization for DNA Language Models
Taewon Kim (KAIST), Insu Han (KAIST)
Computational EfficiencyRepresentation LearningDrug DiscoveryTransformerLarge Language ModelContrastive LearningBiomedical DataBenchmark
🎯 What it does: Proposed a learnable DNA language model called DNACHUNKER, which achieves dynamic tokenization and masked language model training on DNA sequences through a variable-length, context-aware chunker.
dnaHNet: A Scalable and Hierarchical Foundation Model for Genomic Sequence Learning
Arnav Shah (University of Toronto), Albert Gu (Cartesia AI)
Computational EfficiencyRepresentation LearningDrug DiscoveryTransformerLarge Language ModelAuto EncoderContrastive LearningBiomedical Data
🎯 What it does: Propose dnaHNet, an end-to-end subword-free autoregressive genomic learning model that achieves efficient compression while maintaining nucleotide-level granularity through a differentiable dynamic blocking mechanism.
Do Activation Verbalization Methods Convey Privileged Information?
Millicent Li (Northeastern University), Byron C Wallace (Northeastern University)
Safty and PrivacyExplainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextRetrieval-Augmented Generation
🎯 What it does: This paper systematically evaluates methods that use a second LLM to convert the activation of a target model into natural language descriptions (activation decoding), and examines whether these methods truly reveal 'privileged information' inside the target model, rather than merely repeating the input text.
Do Audio LLMs Listen or Read? Analyzing and Mitigating Paralinguistic Failures with VoxParadox
Jiacheng Pang (University of Southern California), Mohammad Soleymani (University of Southern California)
RecognitionTransformerLarge Language ModelPrompt EngineeringTextBenchmarkAudio
🎯 What it does: Propose the VoxParadox adversarial benchmark to evaluate the understanding of audio non-linguistic information by Audio LLMs, and improve the model's paralinguistic recognition ability through Prompt-Conditioned Layer Mixer (PCLM) and Direct Preference Optimization (DPO).
Do Language Models Track Entities Across State Changes?
Zilu Tang (Boston University), Najoung Kim (Boston University)
Explainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This study explores how Transformer language models handle multiple state changes in the Entity Tracking (ET) task, revealing that the models do not incrementally build world states, but instead aggregate relevant information in parallel at query time.
Do LLMs “Feel”? Emotion Circuits Discovery and Control
Chenxi Wang (Mohamed bin Zayed University of Artificial Intelligence), Xiuying Chen (Mohamed bin Zayed University of Artificial Intelligence)
Explainability and InterpretabilityTransformerLarge Language ModelText
🎯 What it does: By analyzing the internal mechanisms of large language models, the system identified and verified hierarchical neurons and attention heads that form emotional circuits, and utilized these circuits to achieve precise control over emotional outputs;
Do LLMs Signal When They’re Right? Evidence from Neuron Agreement
Kang Chen (Fudan University), Yixin Cao (Fudan University)
Explainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringContrastive LearningText
🎯 What it does: Analyze the internal neuron activation patterns of LLMs, using activation sparsity and cross-sample consistency as unsupervised evaluation criteria, and propose Neuron-Agreement Decoding (NAD) to achieve best-of-N sampling and early stopping.
Do Neural Operators Forget Geometry? The Forgetting Hypothesis in Deep Operator Learning
Yanming Xia (Tsinghua University), Angelica I Aviles-Rivero
OptimizationExplainability and InterpretabilityComputational EfficiencyTransformerMeshGraphBenchmarkPhysics Related
🎯 What it does: Propose and verify the 'geometric forgetting hypothesis' that neural operators tend to forget geometric information in deep layers, and alleviate this issue through geometric memory injection.
Do Sparse Autoencoders Identify Reasoning Features in Language Models?
George Ma (University of California Berkeley), Somayeh Sojoudi (University of California Berkeley)
Explainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelAuto EncoderContrastive LearningTextChain-of-Thought
🎯 What it does: Investigate the reliability of sparse autoencoders (SAE) in identifying inference-related internal features within large language models.
Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs
Xin Gao (University of California San Diego), Taylor Berg-Kirkpatrick (University of California San Diego)
GenerationData SynthesisExplainability and InterpretabilityTransformerPrompt EngineeringVision Language ModelDiffusion modelImageTextMultimodalityBenchmarkChain-of-Thought
🎯 What it does: Propose the UNIKE benchmark to evaluate the cross-modal knowledge editing effectiveness of unified multi-modal models, and systematically test whether text editing can be transferred to image generation.
Do Transformers Need Three Projections? Systematic Study of QKV Variants
Ali Kayyam (BrainChip Inc), M Anthony Lewis (BrainChip Inc)
ClassificationSegmentationAnomaly DetectionComputational EfficiencyTransformerLarge Language ModelContrastive LearningImageTextBiomedical Data
🎯 What it does: Systematically evaluated shared variants of Q, K, V projections in Transformers (Q=K=V, Q=K=V, Q=K=V), and experimentally compared their performance and memory overhead across synthetic, visual, and language modeling tasks.
Do Vision and Text Cues Exhibit Evidential Coupling? UFO: A Benchmark for Compositional Multimodal Reasoning in Unified Models
Zhongyu Yang (Heriotwatt University), Yingfang Yuan (Heriotwatt University)
Explainability and InterpretabilityTransformerLarge Language ModelVision Language ModelDiffusion modelImageTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposed the UFO (Unified Foundation) benchmark to evaluate whether unified foundation models can generate and use image and text clues as grounded evidence for compositional multimodal reasoning during generation and understanding processes.
Do We Need Adam? Surprisingly Strong and Sparse Reinforcement Learning with SGD in LLMs
Sagnik Mukherjee (University of Illinois), Hao Peng (University of Illinois)
OptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningTextSequential
🎯 What it does: Studied the performance comparison between SGD and AdamW in large language model reinforcement learning (RLVR), demonstrating that SGD can match or even surpass AdamW, while significantly reducing memory consumption and parameter update sparsity.
Do-Prompt: Causal Interventions Meet Variational Prompt Bottlenecks
Xueting Chen (National University of Defense Technology), Wenjing Yang (National University of Defense Technology)
Domain AdaptationRepresentation LearningTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodality
🎯 What it does: Introduce variational information bottleneck and causal intervention into multi-modal prompt learning to limit the capacity of prompts and suppress environment-related shortcuts, thereby enhancing cross-domain and base-new generalization performance.
Doc-to-LoRA: Learning to Instantly Internalize Contexts
Rujikorn Charakorn (Sakana AI), Robert Tjarko Lange (Sakana AI)
Computational EfficiencyKnowledge DistillationRepresentation LearningMeta LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextMultimodalityBenchmarkRetrieval-Augmented Generation
🎯 What it does: Designed a lightweight hypernetwork called Doc-to-LoRA (D2L), which can rapidly internalize arbitrary contextual information into LoRA parameters during a single forward pass, and subsequently answer queries directly without needing the original context.
DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents
Zhuoran Yu (University of Wisconsin Madison), Yong Jae Lee (University of Wisconsin Madison)
Data SynthesisTransformerLarge Language ModelPrompt EngineeringImageTextMultimodalityTabularBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposed the DOCHOP benchmark, specifically designed to evaluate multi-hop reasoning capabilities that integrate multiple charts with textual narratives.
DOCKSMITH: Scaling Reliable Coding Environments via an Agentic Docker Builder
Jiaran Zhang (StepFun), Xiangyu Zhang (StepFun)
AI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningAgentic AIPrompt EngineeringTextSequentialRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Designed and trained a specialized Docker environment building agent called DOCKSMITH to address the bottleneck of environment setup in software engineering agents.
DocOS: Towards Proactive Document-Guided Actions in GUI Agents
Jingjing Liu (Beihang University), Haifeng Wang (Baidu Inc)
Autonomous DrivingAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelVision Language ModelVision-Language-Action ModelTextMultimodalityBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposed the Proactive Document-Guided Action (Proactive Document-Guided Action) paradigm and designed the DocOS benchmark based on it, which is used to evaluate GUI agents' ability to proactively retrieve, understand, and execute instructions from online documents in a fully interactive environment.
DocVAL: Validated Chain-of-Thought Distillation for Grounded Document VQA
Pinaki Prasad Guha Neogi (Ohio State University), Rajiv Ramnath (Ohio State University)
Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerSupervised Fine-TuningPrompt EngineeringVision Language ModelImageTextMultimodalityChain-of-Thought
🎯 What it does: Propose the DocVAL framework, which leverages verified chain-of-thought (CoT) distillation to transfer spatial reasoning from large teacher models to small, deployable vision-language models, achieving precise localization of answers and their positions in document visual question answering.
Does a Hybrid Space-Aware Randomized Defense Improve Empirical and Certified Adversarial Robustness?
Joy Dhar (Indian Institute of Technology Ropar), Pietro Lio (University of Cambridge)
ClassificationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningAdversarial AttackConvolutional Neural NetworkTransformerContrastive LearningImageBiomedical DataBenchmark
🎯 What it does: Under the existing randomized smoothing framework, we propose HySCAN, which achieves hybrid randomized defense by injecting attention-controlled random noise into the weight space and feature space.
Does AI Reviewer See the Full Picture? Attacking and Defending Multimodal Peer Review
Xinyu Zhao (University of North Carolina at Chapel Hill), Tianlong Chen (University of North Carolina at Chapel Hill)
Anomaly DetectionExplainability and InterpretabilityComputational EfficiencyAdversarial AttackData-Centric LearningTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose the PaperGuard benchmark, systematically evaluate and defend against adversarial attacks in multi-modal AI peer review, covering text injection and image perturbation.
Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning
Maggie Ziyu Huan (Carnegie Mellon University), Xiang Yue (Carnegie Mellon University)
Explainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringContrastive LearningTextBenchmarkChain-of-Thought
🎯 What it does: This paper investigates whether improving mathematical reasoning ability can be transferred to other tasks of LLMs, and systematically evaluates the transferability of 20+ publicly available weight reasoning models;
Does Reasoning Improve Seeing? Understanding When Vision-Language Models Benefit from Thinking
Jing Bi (University of Rochester), Chenliang Xu (University of Rochester)
Explainability and InterpretabilityComputational EfficiencyTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityChain-of-Thought
🎯 What it does: Investigates when visual language models need to think and how much thinking (token length) is required, and achieves adaptive control of computational resources and reasoning strategies during testing through linear probing and causal regulation of attention heads.
Does Reinforcement Fine-Tuning Improve Generalization of LLM Agents? An Empirical Study
Zhiheng Xi (Fudan University), Xuanjing Huang (Fudan University)
Autonomous DrivingOptimizationFederated LearningSafty and PrivacyExplainability and InterpretabilityComputational EfficiencyRobotic IntelligenceMeta LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AIPrompt EngineeringWorld ModelTextSequentialChain-of-Thought
🎯 What it does: This paper studies the generalization ability of Reinforcement Fine-tuning (RFT) in large language model (LLM) agents, systematically evaluating the performance of RFT across task difficulty levels within the same environment, cross-environment migration, and sequential training across multiple environments.
Does Your Reasoning Model Implicitly Know When to Stop Thinking?
Zixuan Huang (Beihang University), Yikun Ban (Beihang University)
Computational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkChain-of-Thought
🎯 What it does: This paper reveals the inherent 'when to stop thinking' capability in large reasoning models (LRMs) and designs new sampling methods, SAGE (Self-aware Guided Efficient Reasoning) and SAGE-RL (integrating SAGE into the RLVR framework), enabling the model to terminate early during the generation process by exploring the space more thoroughly, thereby improving reasoning accuracy and significantly compressing output length.
Domain Adaptation with Adaptive $f$-Divergence: Tighter Variational Representation and Generalization Bounds
Zhe Cheng (Southwestern University of Finance and Economics), Jiaolong Wang (Southwestern University of Finance and Economics)
ClassificationDomain AdaptationAdversarial AttackContrastive LearningImage
🎯 What it does: Propose an unsupervised domain adaptation framework that further tightens the variational lower bound of f-divergence by utilizing a learnable monotone L-Lipschitz transformation τ, and adaptively selects the divergence family through a likelihood method.
Domain Adaptive Object Detection via Dynamic Causal Refinement
Zeyu Ma (University of Electronic Science and Technology of China), Heng Tao Shen (Tongji University)
Object DetectionDomain AdaptationTransformerAuto EncoderContrastive LearningImage
🎯 What it does: A dynamic causal refinement framework (DCR) is proposed for domain adaptive object detection, suppressing domain-agnostic pseudo factors through semantic prediction consistency (SPC) and difference-guided causal refinement (DGCR).
Domain Restriction via SAE Multi-Layer Transitions
Elias Shaheen (Technion Israel Institute of Technology), Avi Mendelson (Technion Israel Institute of Technology)
Domain AdaptationAnomaly DetectionExplainability and InterpretabilityTransformerLarge Language ModelAuto EncoderContrastive LearningText
🎯 What it does: A lightweight domain restriction (OOD detection) method that uses unlabeled, in-domain data only is achieved by performing sequential anomaly scoring on the feature sequences extracted from sparse autoencoders (SAE) learned between internal layers of large language models.
Domain Transfer Becomes Identifiable via a Single Alignment
Sagar Shrestha (Oregon State University), Xiao Fu (Oregon State University)
Domain AdaptationGenerative Adversarial NetworkContrastive LearningImageBiomedical Data
🎯 What it does: A method is proposed in unsupervised domain adaptation that identifies true transfer mappings by leveraging the structured Jacobian sparsity and using a single paired anchor sample, further achieving high-dimensional sparse regularization through randomized masked finite differences.
Domain-Shift-Aware Conformal Prediction for Large Language Models
Zhexiao Lin (University of California), Michael Berger
Domain AdaptationExplainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelContrastive LearningTextBenchmark
🎯 What it does: Proposes a domain-aware adaptive conformal prediction framework (DSCP) for large language models, aiming to achieve reliable uncertainty quantification under domain shift.
DomED: Redesigning Ensemble Distillation for Domain Generalization
Ziang Song (Wuhan University), Zijun Zhang (Wuhan University)
Domain AdaptationKnowledge DistillationContrastive LearningImageBenchmark
🎯 What it does: This paper proposes DomED, an ensemble distillation framework for domain generalization, which transfers the generalization capability and uncertainty information of integrated models to a single student model by training teacher models on different source domains and performing complementary distillation on unseen domains.
Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference
Mostafa Elhoushi (Cerebras Systems), Joel Hestness (Cerebras Systems)
OptimizationComputational EfficiencyTransformerLarge Language ModelTextStochastic Differential Equation
🎯 What it does: In large-scale language model pretraining, systematically evaluate and optimize the hierarchical Dropout (Stochastic Depth) strategy, proving that it can achieve significant computational savings and performance improvements during both training and inference phases.
Don't Force the Fit: Bounded Log-Likelihood Loss for Enhanced Reasoning in Large Language Models
Feng Zhao (Huazhong University of Science and Technology), Guandong Xu (Education University of Hong Kong)
OptimizationExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningTextChain-of-Thought
🎯 What it does: This paper proposes a new training loss called Bounded Log-Likelihood Loss (BLL-Loss), aiming to alleviate token-level overfitting in large language models during reasoning tasks and improve reasoning generalization performance.
Don't Forget Why You Started: Tackling Dual Forgetting in Vision-Language Continual Learning
Borui Kang (Nanjing University), Yang Gao (Nanjing University)
ClassificationComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerSupervised Fine-TuningMixture of ExpertsVision Language ModelContrastive LearningImageTextMultimodality
🎯 What it does: Propose the DFA-CIL framework and the SCR metric to strictly evaluate two types of forgetting in VLM during continual learning (IKF and PKF), and based on this, design the DFA-MoE parameter-efficient fine-tuning method, which balances knowledge retention and new task adaptation.
Don't Ignore the Tail: Decoupled Distillation Produces Top Maths Students on an Academic Budget
Sayantan Dasgupta (University of Melbourne), Timothy Baldwin (University of Melbourne)
Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningText
🎯 What it does: Proposes the Tail-aware Distillation (TAD) method, which separates the TopK probabilities from the tail probabilities of the teacher model, and enhances tail learning through normalized KL divergence, achieving pre-training and supervised knowledge distillation, particularly significantly improving the performance of small models on mathematical reasoning tasks.
Don't Overthink with Pixels: Efficient Reasoning for Segmentation
Song Wang (Zhejiang University), Xinchao Wang (National University of Singapore)
SegmentationComputational EfficiencyTransformerLarge Language ModelReinforcement LearningPrompt EngineeringVision Language ModelImageMultimodalityBenchmark
🎯 What it does: Propose PIXELTHINK, an efficient segmentation method that regulates the inference length based on task difficulty and model uncertainty, significantly reducing the number of inference tokens used and improving segmentation accuracy.
Don't Reinvent the Wheel, Just Realign the Spokes: Resource-Efficient Federated Fine-Tuning via Rank-Wise Expert Assembly
Yebo Wu (University of Macau), Li Li (University of Macau)
Federated LearningComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelSupervised Fine-TuningMixture of ExpertsText
🎯 What it does: This paper proposes SMARTFED, a resource-efficient federated fine-tuning framework that leverages existing LoRA modules for knowledge reuse, avoiding the need to train LLMs from scratch, significantly reducing computational and communication costs on edge devices.
Don't Walk the Line: Boundary Guidance for Filtered Generation
Sarah Ball (LMU Munich), Andreas Haupt (Stanford University)
GenerationSafty and PrivacyReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringText
🎯 What it does: This paper proposes Boundary Guidance, a fine-tuning method based on reinforcement learning, which directly introduces boundary information from a safety classifier during the generation process to avoid generating content that falls on the classifier's decision boundary, thereby improving the collaborative performance between the generative model and the safety filter.
Doppler Prompting for Stable mmWave-based Human Pose Estimation
Shuntian Zheng (University of Warwick), Yu Guan (University of Warwick)
Pose EstimationTransformerPrompt EngineeringContrastive LearningPoint CloudTime Series
🎯 What it does: Achieving spatiotemporal stability in mmWave human pose estimation through Doppler hints control.
DOT-MoE: Differentiable Optimal Transport for MoEfication
Udbhav Bamba (Amazon), Deepak Gupta (Amazon)
Computational EfficiencyKnowledge DistillationTransformerLarge Language ModelMixture of ExpertsText
🎯 What it does: Proposes a framework called DOT-MoE, which converts the feed-forward network (FFN) of dense Transformers into a sparse Mixture of Experts (MoE), enabling efficient inference without retraining the model weights.
Doubly Outlier-Robust Online Infinite Hidden Markov Model
Horace Yiu (University of Oxford), Gerardo Duran-Martin (Oxford-Man Institute of Quantitative Finance)
Anomaly DetectionFederated LearningComputational EfficiencyRepresentation LearningDiffusion modelScore-based ModelAuto EncoderContrastive LearningTabularTime SeriesSequentialFinance Related
🎯 What it does: Proposed a doubly robust online infinite hidden Markov model (BR-iHMM) that can simultaneously resist outliers in both the observation space and the hidden state space.
Doubly Regularized Markov Decision Processes for Robust Reinforcement Learning
Yiting He (Georgia Institute of Technology), Pan Xu (Duke University)
Reinforcement LearningAuto EncoderContrastive Learning
🎯 What it does: Propose a dual regularized MDP framework and the RSPVI online algorithm to achieve robust learning against reward and transition uncertainties.
Doubly Robust Distributionally Robust Offline Contextual Pricing
Min Xu (Nanjing University), Caihua Chen (Nanjing University)
OptimizationReinforcement LearningTabularFinance Related
🎯 What it does: This paper proposes a dual-robust distributionally robust offline policy evaluation and learning framework for continuous pricing.
DOUBT: Decoupled Object-level Understanding and Bridging via vMF-based Trustworthiness for Hallucination Detection in MLLMs
Kaiqi Chen (Sichuan University), Peng Hu (Sichuan University)
Anomaly DetectionExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmark
🎯 What it does: Propose a black-box hallucination detection framework called DOUBT, which first splits visual recognition and relation reasoning by object-level understanding and bridging (OUB), and then uses vMF to evaluate the consistency of multiple answers based on the directional credibility metric of vectors, thus determining hallucinations.
DP-KFC: Data-Free Preconditioning for Privacy-Preserving Deep Learning
Marc Molina Van den bosch, Luigi Serio (CERN)
OptimizationFederated LearningSafty and PrivacyConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageTextMultimodalityBiomedical Data
🎯 What it does: Proposes a data-free second-order preprocessing method called DP-KFC, which reconstructs the KFAC preprocessor using structured synthetic noise, and geometrically matches DP-SGD without consuming privacy budget or using public data.
DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and Its Loss' Convexity is Dispensable)
Wenxuan Zhou (Google DeepMind), Andrew Hard (Google DeepMind)
OptimizationReinforcement Learning from Human FeedbackLarge Language ModelSupervised Fine-TuningContrastive LearningText
🎯 What it does: This paper deepens the theoretical foundation of Direct Preference Optimization (DPO), proposing a general normative framework (KLST*), and proves that as long as this framework is satisfied, the core derivation of DPO no longer depends on specific BTL choice models, allowing any differentiable non-convex loss function and human choice model to be freely combined;
DPsurv: Dual-Prototype Evidential Fusion for Uncertainty-Aware and Interpretable Whole Slide Image Survival Prediction
Yucheng Xing (National University of Singapore), Mengling Feng (National University of Singapore)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningData-Centric LearningConvolutional Neural NetworkTransformerMixture of ExpertsContrastive LearningGaussian SplattingImageBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasound
🎯 What it does: Proposed the DPsurv network, which uses dual-prototype evidence fusion to perform survival prediction on whole-slide images, outputs prediction intervals with uncertainty, and achieves end-to-end multi-level interpretability.
DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research
Rulin Shao (University Of Washington), Pang Wei Koh
OptimizationComputational EfficiencyRepresentation LearningData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningTextMultimodalityBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposed DR Tulu‑8B, an open-source LLM designed for long-form deep research tasks, and developed a Reinforcement Learning with Evolving Rubrics (RLER) training framework that optimizes policies using dynamically generated and updated evaluation criteria (rubrics).
DR-MMSearchAgent: Deepening Reasoning in Multimodal Search Agents
Shengqin Wang (East China Normal University), Yuan Xie (East China Normal University)
Recommendation SystemAutonomous DrivingOptimizationTransformerLarge Language ModelReinforcement LearningVision Language ModelGaussian SplattingImageTextMultimodality
🎯 What it does: This paper proposes DR-MMSearchAgent, which utilizes reinforcement learning to achieve deep reasoning and multi-round tool interaction for multi-modal search agents.
Dr. Kernel: Reinforcement Learning Done Right for Triton Kernel Generations
Wei Liu (Hong Kong University of Science and Technology), Junxian He (Hong Kong University of Science and Technology)
OptimizationComputational EfficiencyAI Code AssistantTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsContrastive LearningTextSequentialBenchmark
🎯 What it does: Built an extensible RL training framework called KERNELGYM for the task of generating GPU core code, and trained the DR.KERNEL-14B model using multi-round RL combined with unbiased advantage estimation TRLOO, aligned reward PR/PRS, and MRS training-inference mismatch correction, significantly improving the Fast@1/1.2 metrics of the benchmark KernelBench.
DR$^2$Seg: Decomposed Two-Stage Rollouts for Efficient Reasoning Segmentation in Multimodal Large Language Models
Yulin He (National University of Defense Technology), Wenjing Yang (National University of Defense Technology)
SegmentationComputational EfficiencyReinforcement Learning from Human FeedbackTransformerSupervised Fine-TuningReinforcement LearningVision Language ModelImageTextMultimodality
🎯 What it does: Designed the DR Seg framework, which adopts a two-stage rollout to divide reasoning into multi-modal reasoning and reference segmentation, and improves reasoning efficiency and segmentation accuracy through self-rewarding.
Draft-and-Audit Reinforcement Learning for Optimization Modeling
Zeping Min (Alibaba Damo Academy), Xinshang Wang (Alibaba Damo Academy)
OptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmark
🎯 What it does: Train a two-round draft-review reinforcement learning framework that converts natural language to optimization modeling tasks into an interactive workflow, leveraging terminal validation to improve decoding quality.
Dream-MPC: Gradient-Based Model Predictive Control with Latent Imagination
Jonathan Spieler (University of Bonn), Sven Behnke (University of Bonn)
OptimizationConvolutional Neural NetworkRecurrent Neural NetworkTransformerReinforcement LearningAuto EncoderWorld ModelImageVideoTabularBenchmark
🎯 What it does: Propose a method called Dream-MPC, which generates candidate trajectories in the latent space using a learned world model and policy network, and optimizes each trajectory via gradient ascent, finally applying the first action of the optimal trajectory to the environment.
DREAM-R: Multimodal Speculative Reasoning with RL-Based Refined Drafting, Precise Verification, and Fully Parallel Execution
Yunhai Hu (New York University), Sai Qian Zhang (New York University)
Computational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringContrastive LearningTextMultimodality
🎯 What it does: Proposes the DREAM-R framework for explicit acceleration of multi-modal reasoning, enhancing inference efficiency through speculative alignment and fully parallel execution.
DREAM: A Unified Framework for Drift-Corrected Federated Multi-Objective Learning
Yuan Zhou (Southeast University), Xinli Shi (Southeast University)
OptimizationFederated LearningContrastive LearningImage
🎯 What it does: Propose the DREAM framework to address client drift and aggregation bias in federated multi-objective learning.
DREAM: Dual-Standard Semantic Homogeneity with Dynamic Optimization for Graph Learning with Label Noise
Yusheng Zhao (Peking University), Ming Zhang (Peking University)
ClassificationAnomaly DetectionOptimizationRepresentation LearningGraph Neural NetworkContrastive LearningGraph
🎯 What it does: Propose a dual-standard semantic homogenization and dynamic optimization framework (DREAM) for handling label noise in graph learning, which dynamically evaluates label reliability by combining the semantic proximity of nodes with graph topology relationships and performs weighted optimization on the loss.
DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos
Shenyuan Gao (NVIDIA), Linxi Fan (NVIDIA)
Autonomous DrivingKnowledge DistillationRobotic IntelligenceReinforcement Learning from Human FeedbackTransformerVision-Language-Action ModelDiffusion modelScore-based ModelAuto EncoderContrastive LearningWorld ModelVideoMultimodality
🎯 What it does: Developed a general-purpose robot world model called DreamDojo, which is pre-trained on a large scale using 44k hours of self-shot human videos, learning diverse interactions and continuous action control, and achieving real-time inference through fine-tuning and distillation.
DreamID-Omni: Unified Framework for Controllable Human-Centric Audio-Video Generation
Xu Guo (Tsinghua University), Xiangwang Hou (Tsinghua University)
GenerationData SynthesisTransformerPrompt EngineeringVision Language ModelDiffusion modelVideoTextMultimodalityAudio
🎯 What it does: Propose a unified framework called DreamID-Omni, achieving controllable generation of human reference audio-video (R2AV), video editing (RV2AV), and audio-driven video animation (RA2V), supporting identity and voice synchronization for single/multiple persons.
Dreaming in Code for Curriculum Learning in Open-Ended Worlds
Konstantinos Mitsides (Imperial College London), Antoine Cully (Imperial College London)
Data SynthesisTransformerLarge Language ModelReinforcement LearningPrompt EngineeringDiffusion modelScore-based ModelWorld ModelTextBenchmark
🎯 What it does: Propose a framework for environment code generation based on foundational models called Dreaming in Code (DiCode), which generates executable program-level code in an open world (Craftax) to build adaptive curricula that enhance long-term skill learning in RL agents.
DRFusion: Drift-Resilient Temporally Consistent Infrared–Visible Video Fusion
Xingyuan Li (Zhejiang University), Jinyuan Liu (Dalian University of Technology)
Image TranslationObject TrackingTransformerDiffusion modelAuto EncoderOptical FlowVideoMultimodality
🎯 What it does: Propose a drift-robust and temporally consistent infrared-visible video fusion method based on 3D diffusion models (DRFusion).