arXivSub Start free trial

ICML 2026 Papers — Page 15

International Conference on Machine Learning · 6554 papers

Diamond Maps: Efficient Reward Alignment via Stochastic Flow Maps

Peter Holderrieth (MIT), Max Simchowitz (Carnegie Mellon University)

GenerationComputational EfficiencyKnowledge DistillationReinforcement Learning from Human FeedbackTransformerSupervised Fine-TuningDiffusion modelScore-based ModelFlow-based ModelRectified FlowGaussian SplattingImageTextMultimodality

🎯 What it does: Propose Diamond Maps, a stochastic flow graph model that can instantly align with arbitrary rewards during inference, providing one-time sampling and efficiently estimating the value function, applicable to search, SMC, and guidance.

DiasR: Dual-Modal Identity-Anchored Sparse Routing for Efficient Multi-Subject Video Generation

Yang-yang Li, Guoqing Jin (People's Daily Online)

GenerationData SynthesisComputational EfficiencyTransformerVision Language ModelDiffusion modelAuto EncoderContrastive LearningVideoTextMultimodality

🎯 What it does: Proposed an efficient multi-agent personalized video generation framework called DiasR, which can generate multi-agent interaction videos while maintaining cross-frame identity consistency.

Dichotomy of Feature Learning and Unlearning: Fast-Slow Analysis on Neural Networks with Stochastic Gradient Descent

Shota Imai (University of Tokyo), Masaaki Imaizumi (University of Tokyo)

OptimizationRepresentation LearningContrastive LearningTabularStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: This paper studies the macroscopic dynamics of two-layer neural networks under large-batch SGD training, revealing the fast and slow time-scale mechanisms of feature learning and forgetting, and deriving the corresponding ordinary differential equations in the infinite-width limit;

DiFA: Inference-Time Forward-Process Alignment for Diffusion Models

Shigui Li (South China University of Technology), Delu zeng

GenerationComputational EfficiencyKnowledge DistillationDiffusion modelScore-based ModelImage

🎯 What it does: Proposes DiFA, an untrained inference-time framework that improves the clean signal prediction of diffusion models through temporal consistency filtering and bias guidance.

Diff-MN: Diffusion Parameterized MoE-NCDE for Continuous Time Series Generation with Irregular Observations

Xu Zhang (Microsoft Research), Jiang Bian (Microsoft Research)

GenerationData SynthesisRecurrent Neural NetworkTransformerMixture of ExpertsDiffusion modelAuto EncoderTabularTime SeriesElectrocardiogramFinance Related

🎯 What it does: Propose the Diff-MN framework, which utilizes a Mixture-of-Experts (MoE)-enhanced Neural Controlled Differential Equation (NCDE) to generate continuous high-resolution time series from irregular observations;

DiffCrossGait: Trajectory-Level Alignment for 2D-3D Cross-Modal Gait Recognition via Latent Diffusion

Zhiyang Lu (Xiamen University), Ming Cheng (Xiamen University)

RecognitionTransformerDiffusion modelContrastive LearningImageMultimodalityPoint Cloud

🎯 What it does: Propose the DiffCrossGait framework, which utilizes a shared Gaussian noise-driven latent diffusion process to achieve trajectory-level alignment between 2D imagery and 3D LiDAR views;

Difference-Aware Decision Learning for Multimodal Image Fusion

Hao Pan (Southwest University of Science and Technology), Xingfeng Li (Southwest University of Science and Technology)

Image HarmonizationRestorationConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageMultimodality

🎯 What it does: Proposes the IDEAL framework, which redefines the multi-modal image fusion problem as decision learning under observational conditions, explicitly modeling the contribution of local modalities to decision variables.

Differentiable Conformal Training for LLM Reasoning Factuality

Nathan Hittesdorf (University of Illinois at Chicago), Lu Cheng (University of Illinois at Chicago)

OptimizationExplainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelTextGraphBenchmarkChain-of-Thought

🎯 What it does: Propose a differentiable Coherent Factuality (DCF) framework that allows the factuality of multi-step reasoning generated by LLMs to be both learnable and optimized, while retaining the original statistical coverage guarantees during testing.

Differentiable Optimization Layers for Guaranteed Fairness in Deep Learning

David Troxell (University of California Los Angeles), Guido Montufar (University of California Los Angeles)

OptimizationFederated LearningSafty and PrivacyExplainability and InterpretabilityDiffusion modelScore-based ModelFlow-based ModelRectified FlowNeural Radiance FieldAuto EncoderGenerative Adversarial NetworkContrastive LearningImageTabularTime SeriesSequentialFinance Related

🎯 What it does: Proposes a differentiable fairness layer, embedding fairness constraints directly into the network output to ensure group fairness during both training and inference processes;

Differentiable Weightless Controllers: Learning Logic Circuits for Continuous Control

Fabian Kresse (Institute of Science and Technology Austria), Christoph H. Lampert (Institute of Science and Technology Austria)

Computational EfficiencyRobotic IntelligenceReinforcement LearningAuto EncoderContrastive LearningTabularTime Series

🎯 What it does: A differentiable weightless controller (DWC) was developed, replacing traditional matrix multiplication with a logic gate network to achieve low-latency, low-energy consumption continuous control strategies.

Differential Smoothing Mitigates Sharpening and Improves LLM Reasoning

Jingchu Gai (Carnegie Mellon University), Aditi Raghunathan (Carnegie Mellon University)

OptimizationTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringText

🎯 What it does: Studied the diversity collapse problem caused by RL fine-tuning and proposed the Differential Smoothing (DS-GRPO) method to simultaneously improve correctness and diversity.

Differential syntactic and semantic encoding in LLMs

Santiago Acevedo (Scuola Internazionale Superiore di Studi Avanzati), Marco Baroni (Catalan Institute of Research and Advanced Studies)

Explainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelContrastive LearningTextMultimodality

🎯 What it does: This paper investigates the encoding and separation of syntactic and semantic information in sentence representations of internal layers of large language models by constructing and ablating the 'center points' of sentence syntax and semantics in linear projections.

Differentially Private Continual Release with Relative Error

Bo Li (Hong Kong University of Science and Technology), Peng Ye (Hong Kong University of Science and Technology)

Federated LearningSafty and Privacy

🎯 What it does: Under the continuous release model, four basic tasks—MaxSum, MinSum, MaxSelect, and MinSelect—are studied, and relative error (relative error) is introduced beyond the traditional purely additive error framework to significantly reduce the error.

Differentially Private Cross-Silo Recommendation from Implicit Feedback

Xun Ran (Hong Kong Polytechnic University), Haibo Hu (Hong Kong Polytechnic University)

Recommendation SystemFederated LearningSafty and PrivacyContrastive LearningTabular

🎯 What it does: Proposes a cross-machine group privacy differential privacy implicit feedback matrix factorization framework, DPIMF, which can complete collaborative recommendation without sharing the original interaction data.

Differentially Private Geodesic Regression

Aditya Kulkarni (University of Massachusetts Amherst), Carlos J Soto (University of Massachusetts Amherst)

Safty and PrivacyPoint CloudGraphTabularTime SeriesSequentialBiomedical Data

🎯 What it does: This paper proposes a differential privacy geodesic regression method aimed at protecting sensitive data during regression analysis on manifolds. The method applies differential privacy to the parameters of geodesic regression using the K-Norm Gradient mechanism.

Differentially Private Preference Data Synthesis for Large Language Model Alignment

Fengyu Gao (University of Virginia), Jing Yang (University of Virginia)

Data SynthesisSafty and PrivacyReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringDiffusion modelText

🎯 What it does: Propose the DPPrefSyn algorithm, which synthesizes preference data using differential privacy to support preference alignment in LLMs;

Differentially Private Range Subgraph Counting

Xian Chen (University of Science and Technology of China), Pan Peng (University of Science and Technology of China)

Safty and PrivacyGraph

🎯 What it does: This paper studies differential privacy counting of subgraph occurrences within a multi-dimensional attribute space.

Differentially Private Submodular Maximization with a Knapsack Constraint

Ron Zadicario (Tel Aviv University), Tova Milo (Tel Aviv University)

OptimizationSafty and PrivacyReinforcement LearningPrompt EngineeringMixture of ExpertsTabularChain-of-Thought

🎯 What it does: Proposes an algorithm to achieve differential privacy in submodular function maximization with a knapsack constraint, separately for monotonic and non-monotonic objectives.

Differentially Private Synthetic Data via APIs 4: Tabular Data

Toan Tran (Emory University), Sergey Yekhanin (Microsoft Research)

Data SynthesisSafty and PrivacyComputational EfficiencyDiffusion modelScore-based ModelGaussian SplattingTabular

🎯 What it does: This paper proposes a Tab-PE algorithm based on a Private Evolution framework for generating high-quality tabular synthetic data under differential privacy constraints.

Diffract: Spectral View of LLM Domain Adaptation

Nikita Borodin (Risk Ai Research Lab), Dmitry Vinichenko (Risk Ai Research Lab)

Domain AdaptationComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsTextTabularReview/Survey PaperBenchmarkRetrieval-Augmented Generation

🎯 What it does: This paper studies the adaptation mechanism of continuous pre-training (CPT) for large language models (LLMs), and reveals the weight update characteristics during the CPT phase through spectral analysis of the weight matrix using singular value decomposition (SVD).

DiffStyle3D: Consistent 3D Gaussian Stylization via Attention Optimization

Yitong Yang (Shanghai University of Finance and Economics), Shuting He (Shanghai University of Finance and Economics)

GenerationOptimizationTransformerPrompt EngineeringDiffusion modelGaussian SplattingImagePoint CloudMesh

🎯 What it does: This paper proposes a 3D Gaussian Splatting (3DGS) style transfer method based on direct optimization in the latent space of diffusion models, which can generate high-quality artistic renderings with multi-view consistency for 3D scenes or objects while preserving the original geometry.

DiffThinker: Towards Generative Multimodal Reasoning with Diffusion Models

Zefeng He (Shanghai AI Laboratory), Yu Cheng (Chinese University of Hong Kong)

GenerationOptimizationExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringDiffusion modelScore-based ModelFlow-based ModelRectified FlowAuto EncoderContrastive LearningImageTextMultimodalityChain-of-Thought

🎯 What it does: Propose DiffThinker, a framework that transforms multimodal reasoning into an image-to-image diffusion model generation process.

DiffuMamba: High-Throughput Diffusion LMs with Mamba Backbone

Vaibhav Singh (ServiceNow Research), Torsten Scholak (ServiceNow Research)

GenerationComputational EfficiencyRepresentation LearningTransformerLarge Language ModelMixture of ExpertsDiffusion modelText

🎯 What it does: This paper proposes two diffusion language models based on Mamba, named DiffuMamba and DiffuMamba-H, to replace traditional Transformer backbones for sequence denoising.

DiffuReason: Enhancing Reasoning Ability for Diffusion Language Models via Monte Carlo Tree Search

YIPING SONG (National University of Defense Technology), Chenping Hou (National University of Defense Technology)

GenerationOptimizationExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelReinforcement LearningMixture of ExpertsDiffusion modelScore-based ModelTextBenchmarkChain-of-Thought

🎯 What it does: Propose DiffuReason, a reasoning framework that embeds Monte Carlo Tree Search (MCTS) into diffusion large language models, achieving parallel generation and system 2 reasoning.

Diffuse to Detect: Bi-Level Sample Rebalancing with Pseudo-Label Diffusion for Point-Supervised Infrared Small-Target Detection

Zhu Liu (Dalian University of Technology), Risheng Liu (Dalian University of Technology)

Object DetectionMeta LearningConvolutional Neural NetworkTransformerDiffusion modelContrastive LearningImage

🎯 What it does: This paper proposes a dual-layer dual-update framework that utilizes point-level supervision for infrared small target detection, and improves detection performance through physics-guided diffusion annotation and sample re-balancing.

Diffusing to Coordinate: Efficient Online Multi-Agent Diffusion Policies

Zhuoran Li (Tsinghua University), Longbo Huang (Tsinghua University)

Reinforcement LearningDiffusion modelTabularTime SeriesSequential

🎯 What it does: Propose an online multi-agent diffusion strategy, OMAD, achieving high expressiveness in online reinforcement learning.

Diffusion Bridge or Flow Matching? A Unifying Framework and Comparative Analysis

Kaizhen Zhu (ShanghaiTech University), Ye Shi (ShanghaiTech University)

Image TranslationRestorationGenerationSuper ResolutionTransformerDiffusion modelFlow-based ModelImageStochastic Differential Equation

🎯 What it does: This paper studies two mainstream methods for distribution-to-distribution transformation—Diffusion Bridge and Flow Matching, and proposes a unified theoretical framework while comparing their advantages and disadvantages.

Diffusion Controller: Framework, Algorithms and Parameterization

Tong Yang (Carnegie Mellon University), Bo Dai (Google DeepMind)

GenerationReinforcement Learning from Human FeedbackTransformerSupervised Fine-TuningReinforcement LearningDiffusion modelScore-based ModelImageTextMultimodalityStochastic Differential Equation

🎯 What it does: View the reverse diffusion sampling as state control under a linear solvable Markov decision process (LS-MDP), and propose the Diffusion Controller (DiffCon) framework. Based on this, a reinforcement learning (RL) fine-tuning algorithm is designed using policy gradient, PPO, and reward-weighted methods, along with a gray-box parameterization scheme that freezes the pre-trained model and adds a lightweight side network.

Diffusion Differentiable Resampling

Jennifer R. Andersson, Zheng Zhao (Linköping University)

GenerationData SynthesisOptimizationComputational EfficiencyReinforcement Learning from Human FeedbackDiffusion modelScore-based ModelAuto EncoderImageTime SeriesSequentialStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: This paper proposes a differentiable resampling method based on untrained diffusion models, which can achieve particle resampling within the sequential Monte Carlo (SMC) framework while maintaining differentiability, thereby supporting gradient optimization.

Diffusion Flow Matching: Dimension-Improved KL Bounds and Wasserstein Guarantees

Marta Gentiloni Silveri (Ecole Polytechnique), Alain Oliviero Durmus (Ecole Polytechnique)

GenerationOptimizationComputational EfficiencyRepresentation LearningDiffusion modelScore-based ModelFlow-based ModelRectified FlowContrastive LearningTabularTime SeriesSequentialReview/Survey PaperStochastic Differential Equation

🎯 What it does: The paper conducts an in-depth theoretical analysis of the Diffusion Flow Matching (DFM) model based on the Brownian bridge, providing non-asymptotic convergence error upper bounds under the Kullback–Leibler (KL) and Wasserstein-2 (W₂) metrics, and separately provides error analysis for different training strategies (no early stopping, early stopping, and a novel step size scheduling).

Diffusion Language Model Parallel Decoding via Product-of-Experts Bridge

Juntong Shi (Stanford University), Minkai Xu (Stanford University)

GenerationComputational EfficiencyTransformerReinforcement LearningPrompt EngineeringMixture of ExpertsDiffusion modelText

🎯 What it does: Propose PoE-Bridge, a method that constructs a Product-of-Experts (PoE) bridging distribution between diffusion language models (DLM) and autoregressive language models (AR), enabling parallel decoding while maintaining AR-level generation quality.

Diffusion Models Are Statistically Optimal for Learning Low-Dimensional Multi-Modal Distributions

Jingda Wu (University of Michigan), Changxiao Cai (University of Michigan)

GenerationData SynthesisDiffusion modelScore-based ModelTabularStochastic Differential Equation

🎯 What it does: Proposes a theoretical analysis of the sample complexity for diffusion models when learning low-dimensional multi-modal distributions (i.e., distributions supported on the union of several linear subspaces), and provides a corresponding kernel-based score estimator along with sampling error analysis.

Diffusion Models Preferentially Memorize Prototypical Examples or: Why Does My Diffusion Model Love Slop?

Marta Aparicio Rodriguez (Imperial College London), Daniel James Korchinski (Ecole Polytechnique Federale De Lausanne)

GenerationTransformerDiffusion modelImage

🎯 What it does: Investigate how diffusion models memorize samples during training and reveal their preference for typical samples and the partial memory phase (slop)

Diffusion-based learning framework for Constrained Nonconvex Optimization with Weighted Bootstrapped Refinement

Shutong Ding (ShanghaiTech University), Ye Shi (ShanghaiTech University)

OptimizationDiffusion modelScore-based ModelContrastive LearningTabularTime SeriesSequentialBenchmark

🎯 What it does: Propose a self-supervised learning framework called DiOpt based on diffusion models for solving constrained non-convex optimization problems.

DiL: Discrete-anchored Representation Alignment for Semi-Supervised Continual Learning

Nanyi Wang (Guizhou University), Qi Wang (Guizhou University)

ClassificationKnowledge DistillationRepresentation LearningContrastive LearningImage

🎯 What it does: This paper proposes a semi-supervised continual learning framework called DiL based on discrete anchors, which aligns features using interpretable discrete prototypes to reduce feature drift caused by pseudo-label noise and improve inter-class separation.

DiLA: Disentangled Latent Action World Models

Tianqiu Zhang (Peking University), Si Wu (Peking University)

GenerationRepresentation LearningRobotic IntelligenceTransformerVision-Language-Action ModelDiffusion modelAuto EncoderContrastive LearningWorld ModelVideo

🎯 What it does: Designed and implemented DiLA, a latent action world model that separates content from structure, capable of learning abstract actions from unannotated videos and achieving high-fidelity video generation.

Dimension-free convergence of diffusion models for approximate Gaussian mixtures

Gen Li (Chinese University of Hong Kong), Yuting Wei (University of Pennsylvania)

GenerationData SynthesisDiffusion modelScore-based ModelTabularStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Conducts a theoretical convergence analysis of DDPM (Denoising Diffusion Probabilistic Model) under the scenario where the target distribution is approximated as a Gaussian Mixture Model (GMM), proving that its iterative complexity is independent of the dimension d and the number of mixture components K, and only grows as ˜O(1/ε) with the error ε.

Dimension-Free Multimodal Sampling via Preconditioned Annealed Langevin Dynamics

Lorenzo Baldassari (University of Basel), Maarten V. de Hoop (Rice University)

OptimizationComputational EfficiencyRepresentation LearningDiffusion modelScore-based ModelMultimodalityTabularReview/Survey PaperBenchmarkPhysics RelatedStochastic Differential Equation

🎯 What it does: Studied the dimension-agnostic sampling performance of continuous-time Annealed Langevin Dynamics (ALD) under multi-modal objectives, and provided spectral design conditions that ensure stability in high-dimensional and error-perturbed scenarios.

Dimension-Independent Convergence of Underdamped Langevin Monte Carlo in KL Divergence

Shiyuan Zhang (University of California, Los Angeles), Quanquan Gu (University of California, Los Angeles)

OptimizationStochastic Differential Equation

🎯 What it does: Proposes a dimension-independent convergence proof for the Underdamped Langevin Monte Carlo (ULMC) and the Randomized Midpoint Method (RMD) based on low-order discretization under the KL divergence.

Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning

Junxuan Wang (Shanghai Innovation Institute), Xipeng Qiu (Shanghai Innovation Institute)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelAuto EncoderContrastive LearningText

🎯 What it does: This paper investigates the phenomenon where the attention outputs of Transformers exhibit low-dimensional folding (effective rank around 60%) in the hidden space, and reveals the fundamental cause of the large number of 'dead features' in sparse dictionary learning (e.g., sparse autoencoders). Based on this, the authors propose the Active Subspace Initialization (ASI) method, which aligns the initialization direction of sparse features with the activated principal subspace, thereby significantly reducing dead features and improving reconstruction quality; and extend this idea to sparse replacement models (LoRA).

Dimensionality Reduction with Point-distributions Similarity Invariant

Hang Zhang (Nanjing University), Kai Ming Ting (Nanjing University)

Computational EfficiencyRepresentation LearningData-Centric LearningAuto EncoderContrastive LearningImageTextTabularBiomedical DataBenchmark

🎯 What it does: Proposed the point-distribution similarity invariant and designed a linear-time dimensionality reduction algorithm Ψ-DR (PSIDR), achieving the preservation of point and distribution similarity through a two-stage process involving kernel mean embedding, classical MDS, and multilateration.

DiP-G: Discrete Prompting for Graph Neural Networks

Yumeng Zhao (University of Science and Technology of China), Bei Hua (University of Science and Technology of China)

ClassificationComputational EfficiencyRepresentation LearningMeta LearningGraph Neural NetworkPrompt EngineeringAuto EncoderContrastive LearningGraph

🎯 What it does: To address few-shot adaptation for pre-trained graph neural networks, the framework DiP-G is proposed, which directly learns hard k-sparse topological prompts in the discrete space.

Direct 3D-Aware Object Insertion via Decomposed Visual Proxies

Jingbo Gong (Nankai University), Chen Change Loy (Nanyang Technological University)

Image TranslationGenerationData SynthesisPose EstimationTransformerPrompt EngineeringVision Language ModelDiffusion modelAuto EncoderGenerative Adversarial NetworkImageMultimodalityPoint CloudMeshRetrieval-Augmented Generation

🎯 What it does: Proposes a framework called DIRECT that can insert a reference object into a target background according to a user-specified 6-DoF pose.

Direct Flow Q-Learning

Shicheng Cao (Nankai University), Shengbo Eben Li

TransformerReinforcement LearningDiffusion modelFlow-based ModelTabularTime SeriesSequentialBenchmark

🎯 What it does: Propose an offline reinforcement learning framework called Direct Flow Q-Learning (DFQL), which directly injects the terminal Q gradient into the flow matching velocity field, achieving reinforcement learning optimization for multi-step generative models without requiring BPTT or policy distillation.

DirectEdit: Step-Level Accurate Inversion for Flow-Based Image Editing

Desong Yang (Wuhan University), Mang Ye (Wuhan University)

Image TranslationImage HarmonizationRestorationGenerationTransformerPrompt EngineeringVision Language ModelDiffusion modelScore-based ModelFlow-based ModelRectified FlowAuto EncoderImageTextMultimodalityBenchmark

🎯 What it does: Proposes a no-training, step-level accurate inversion flow model image editing method called DirectEdit, which can achieve precise editing while preserving the integrity of the background and structure.

Directional Neural Collapse Explains Few-Shot Transfer in Self-Supervised Learning

Achleshwar Luthra (Texas A&M University), Tomer Galanti (Texas A&M University)

ClassificationRepresentation LearningConvolutional Neural NetworkTransformerContrastive LearningImage

🎯 What it does: The paper investigates the transfer performance of self-supervised pre-trained representations under few-shot settings and demonstrates that directional class-wise variance (directional CDNV) is the key geometric quantity explaining this performance.

Directly Optimizing Natural Language Explanations for Behavioral Faithfulness: Simulatability and Recoverability

Advaith Malladi (IIIT Hyderabad), Shashank Srivastava (UNC Chapel Hill)

ClassificationOptimizationExplainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextTabular

🎯 What it does: Propose the OPEX model, which directly optimizes natural language explanations through reinforcement learning to achieve reproducibility and predictability of classification model behavior.

Dirichlet-Prior Shaping: Guiding Expert Specialization in Upcycled MoEs

Leyla Mirvakhabova (Qualcomm AI Research), Paul N. Whatmough

ClassificationRecognitionImage TranslationRestorationObject DetectionSegmentationGenerationData SynthesisPose EstimationDepth EstimationSuper ResolutionCompressionDomain AdaptationOptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningReinforcement Learning from Human FeedbackNeural Architecture SearchTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringMixture of ExpertsVision Language ModelContrastive LearningImageTextMultimodalityBenchmark

🎯 What it does: In sparse Mixture-of-Experts (MoE) models, researchers propose and apply the Dirichlet-Prior Shaping Loss (DPSL), significantly enhancing the specialization of experts in MoE models upcycled from pre-trained dense models by regularizing the probability distribution of the router's output.

DISCO: Mitigating Bias in Deep Learning with Conditional Distance Correlation

Emre Kavak (Technical University of Munich), Christian Wachinger (Technical University of Munich)

Domain AdaptationExplainability and InterpretabilityData-Centric LearningContrastive LearningImageText

🎯 What it does: Propose the DISCO method, which mitigates bias by utilizing conditional distance correlation within a counterfactual framework.

DiscoForcing: A Unified Framework for Real-Time Audio-Driven Character Control with Diffusion Forcing

Kaiyang Ji (ShanghaiTech University), Jingya Wang (ShanghaiTech University)

GenerationPose EstimationComputational EfficiencyTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningVideoSequentialAudio

🎯 What it does: Studied strictly causal, finite-latency real-time audio-driven character control, and proposed an end-to-end deployable framework called DiscoForcing.

DiScoFormer: Plug-In Density and Score Estimation with Transformers

Vasily Ilin (University of Washington), Ranjay Krishna (Allen Institute for Artificial Intelligence)

Anomaly DetectionOptimizationRepresentation LearningTransformerScore-based ModelAuto EncoderContrastive LearningTabularTime SeriesSequential

🎯 What it does: Propose a novel DiScoFormer model that directly predicts the density and gradient (score) of the potential distribution from i.i.d. samples using Transformer, and can instantly infer at any query point;

Discontinuous Galerkin Neural Operator for Pathology Defocus Deblurring

Shaoqing Duan (East China Normal University), Yan Wang (East China Normal University)

RestorationTransformerDiffusion modelContrastive LearningImageBiomedical DataBenchmark

🎯 What it does: Proposed a discrete Galerkin neural operator (DGNO) for spatially varying, locally discontinuous defocus deblurring tasks in pathological microscopy images;

Discounted Beta–Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards

Haechan Kim (KAIST), Eunho Yang (KAIST)

OptimizationComputational EfficiencyTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmark

🎯 What it does: This paper addresses the reward estimation problem in the Verifiable Reward (RLVR) framework of reinforcement learning, proposing a Discounted Beta-Bernoulli (DBB) reward estimation method that uses historical reward information to reduce variance and prevent variance collapse, thereby improving sample efficiency.

Discovering Differences in Strategic Behavior between Humans and LLMs

Caroline Wang (University of Texas at Austin), Pablo Samuel Castro (Google DeepMind)

Explainability and InterpretabilityReinforcement Learning from Human FeedbackRecurrent Neural NetworkTransformerLarge Language ModelReinforcement LearningPrompt EngineeringAuto EncoderTextSequential

🎯 What it does: Compare strategic behaviors of humans and large language models (LLMs) in iterative rock-paper-scissors (IRPS), and use AlphaEvolve to automatically discover interpretable behavioral models, revealing structural differences between the two in value learning and opponent modeling.

Discovering Implicit Large Language Model Alignment Objectives

Edward Chen (Stanford University), Carlos Guestrin (Stanford University)

Explainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: This paper proposes the Obj-Disco framework, which automatically decomposes LLM alignment reward signals into interpretable natural language objectives.

Discovering Interpretable Algorithms by Decompiling Transformers to RASP

Xinting Huang (Saarland Informatics Campus, Saarland University), Michael Hahn (Saarland Informatics Campus, Saarland University)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelPrompt EngineeringContrastive LearningText

🎯 What it does: This paper proposes a general method that can extract interpretable RASP language programs from trained Transformer models, demonstrating evidence that Transformers internally implement simple and interpretable algorithms.

Discovering Ordinary Differential Equations with LLM-Based Qualitative and Quantitative Evaluation

Sum Kyun Song (Chung Ang University), Jae Yong Lee (Chung Ang University)

OptimizationExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelAgentic AIPrompt EngineeringTabularTime SeriesBenchmarkStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Propose a multi-agent framework DoLQ, which uses LLM to qualitatively and quantitatively evaluate candidate ODEs, ultimately automatically discovering interpretable ordinary differential equations from observational data.

Discovering Scaling Exponents with Physics-Informed Müntz-Szász Networks

Gnankan Landry Regis N'guessan (Axiom Research Group), Bum Jun Kim (University of Tokyo)

TabularTime SeriesBenchmarkPhysics RelatedOrdinary Differential Equation

🎯 What it does: Proposed the physics-informed Müntz-Szász network (MSN-PINN), which explicitly learns the power-law exponents in the network to directly recover the scaling exponents from PDE constraints.

Discovering Symmetry Groups with Flow Matching

Yuxuan Chen (Northeastern University), Robin Walters (Northeastern University)

GenerationRepresentation LearningScore-based ModelFlow-based ModelRectified FlowPoint Cloud

🎯 What it does: Learn symmetric distributions on Lie groups using flow matching, automatically discovering the continuous and discrete symmetric subgroups of the data.

DiscoverLLM: From Executing Intents to Discovering Them

Tae Soo Kim (Korea Advanced Institute Of Science And Technology), Juho Kim (Korea Advanced Institute Of Science And Technology)

Recommendation SystemAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: This study proposes the DISCOVERLLM framework, which trains LLMs to explore and help users discover their unformed intentions through interaction, rather than merely executing explicitly stated instructions.

Discrete Adjoint Schrödinger Bridge Sampler

Wei Guo (Georgia Institute of Technology), Jaemoo Choi (Georgia Institute of Technology)

OptimizationDiffusion modelScore-based ModelGraphTabularPhysics RelatedStochastic Differential Equation

🎯 What it does: Proposed a Schrödinger bridge sampler for discrete state spaces, DASBS, and generalized the adjoint matching (AM) from the continuous domain to the discrete domain;

Discrete Diffusion Samplers and Bridges: Off-Policy Algorithms and Applications in Latent Spaces

Arran Carter (University of Edinburgh), Esmeralda S. Whitammer (University of Edinburgh)

GenerationData SynthesisReinforcement LearningDiffusion modelScore-based ModelImageGraphTabularSequential

🎯 What it does: Studied diffusion samplers in discrete spaces and Schrödinger bridges, proposed a training method for discrete offline reinforcement learning, and applied it to posterior sampling in discrete latent spaces.

Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies

Zhixuan Liang (University of Hong Kong), Ping Luo (University of Hong Kong)

Robotic IntelligenceReinforcement Learning from Human FeedbackTransformerVision-Language-Action ModelDiffusion modelContrastive LearningImageTextMultimodality

🎯 What it does: In robot vision-language-action models, a method is proposed to integrate discrete diffusion processes into a unified Transformer backbone for action generation.

Discrete Diffusion with Physical Mass Constraints for \emph{De Novo} Peptide Sequencing

An Zeyu (Hong Kong Polytechnic University), Wanyu Lin (Hong Kong Polytechnic University)

Drug DiscoveryProtein Structure PredictionTransformerDiffusion modelBiomedical DataBenchmark

🎯 What it does: Redefine de novo peptide sequencing as a discrete diffusion process with physical mass constraints, achieving global iterative reasoning and self-correction.

Discrete Survival Knowledge Distillation for Competing Risks Analysis

Feiyang Deng (University of Michigan), Kevin He (University of Michigan)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningConvolutional Neural NetworkRecurrent Neural NetworkTransformerSupervised Fine-TuningContrastive LearningTabularTime SeriesBiomedical DataElectronic Health Records

🎯 What it does: Propose a knowledge distillation framework called DiSKD for competing risks discrete-time survival analysis.

Discrete Tilt Matching

Yuyuan Chen (Harvard University), Michael Samuel Albergo (Harvard University)

OptimizationComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningDiffusion modelScore-based ModelTextSequential

🎯 What it does: The study proposes a discrete tilted matching (DTM) method for reward fine-tuning of masked diffusion large language models under a no-likelihood condition.

Discretely-Refined Multi-view Clustering via Aligned Anchor Learning

Yuemeng Huang (Dalian Maritime University), Jiqing Zhang (Dalian Maritime University)

OptimizationRepresentation LearningContrastive LearningMultimodalityGraph

🎯 What it does: This paper proposes the Discretely-Refined Multi-view Clustering (DRMC) framework, which constructs a sample-level similarity graph using anchor point learning and introduces a discrete feedback mechanism to achieve high-quality clustering of multi-view data.

Discretized Density-Guided Source-Free Domain Adaptation for Regression

Gezheng Xu (Western University), Boyu Wang (Western University)

Domain AdaptationConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningImageTabularTime Series

🎯 What it does: In the absence of source data, a source-agnostic domain adaptation framework (MERCI) is proposed for regression tasks. It estimates the discrete density of target domain labels through a self-supervised histogram head, enabling uncertainty-aware pseudo-label generation and adaptive training of the regression model.

Discriminative Attribute Graph Clustering Through Topology-Guided Contrastive Learning

Ling Ding (Tianjin University), Cuiying Huo (Tianjin University)

Representation LearningGraph Neural NetworkAuto EncoderContrastive LearningGraph

🎯 What it does: Proposes a self-supervised attribute graph clustering method based on topological reconstruction and contrastive learning without augmentation, utilizing dual perspectives from graph autoencoder and standard autoencoder to suppress redundant information and enhance the discriminativeness of node representations.

Discriminative Mixture-of-Experts on Graphs with Reliable Expert Fusion

Haoyue Deng (Beihang University), Xiao Wang (Beihang University)

ClassificationGraph Neural NetworkMixture of ExpertsContrastive LearningGraph

🎯 What it does: Propose a new graph mixture of experts framework, C2GMoE, which utilizes contrastive routing learning and confidence-aware fusion to improve the performance of graph neural networks in node classification tasks.

Discriminative Visual Process Rewards for Scaling Thinking at Test-Time with Images

Bo-Wen Yin (Nankai University), Qibin Hou (Nankai University)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackTransformerSupervised Fine-TuningReinforcement LearningPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmarkChain-of-Thought

🎯 What it does: Proposed DiscPRM, which evaluates the visual and textual steps in the thinking-image process to improve the reasoning quality of multimodal models.

Disease-Centric Vision-Language Pretraining with Hybrid Visual Encoding for 3D Computed Tomography

Bowen Shi (DAMO Academy, Alibaba Group), Jianpeng Zhang (DAMO Academy, Alibaba Group)

ClassificationImage TranslationAnomaly DetectionData-Centric LearningConvolutional Neural NetworkTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataComputed TomographyRetrieval-Augmented Generation

🎯 What it does: Proposed a disease-centric vision-language pre-training framework (CT-DiagVLM) for 3D CT, achieving zero-shot diagnosis and report generation through cross-modal alignment.

Disentangling Consensus and Value-Specific Representations for Controllable Pluralistic Value Alignment of LLMs

JianKui Zhou, Tun Lu (Fudan University)

Federated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsContrastive LearningText

🎯 What it does: Propose the DisAlign framework, which decomposes the value expert model into a consensus part and a value-specific part, thereby achieving controllable alignment across multiple values.

Disentangling Geometry, Performance, and Training in Language Models

Atharva Kulkarni (University of Southern California), Swabha Swayamdipta (University of Southern California)

Representation LearningHyperparameter SearchData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningText

🎯 What it does: Trained and systematically evaluated the relationship between the effective rank of the unembedding matrix in 108 OLMo-style language models and model performance, generalization, fine-tuning, quantization, and saturation phenomena.

Disentangling Intent from Role: Adversarial Self-Play for Persona-Invariant Safety Alignment

Jiajia Li (Northwestern Polytechnical University), Zhen Wang (Northwestern Polytechnical University)

Safty and PrivacyReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningGenerative Adversarial NetworkContrastive LearningText

🎯 What it does: Propose the Persona-Invariant Alignment (PIA) framework, combining the attack-side Persona Lineage Evolution (PLE) and the defense-side Persona-Invariant Consistency Learning (PICL), to achieve automatic discovery and robust defense against jailbreak attacks based on personas in large language models (LLMs);

Disentangling Latent Risk Pathways via Bayesian Hypergraph Inference

Shengxian Ding (Yale University), Yize Zhao (Yale University)

Explainability and InterpretabilityDrug DiscoveryGraph Neural NetworkTabularBiomedical DataElectronic Health Records

🎯 What it does: Propose a multi-disease modeling framework based on Bayesian hypergraph inference (BHPI), which jointly predicts disease risk through high-order disease pathways regulated by risk factors.

DisjunctiveNet: Neural Symbolic Learning via Differentiable Convexified Optimization Layers

Shraman Pal (Purdue University), Can Li (Purdue University)

OptimizationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningReinforcement LearningDiffusion modelScore-based ModelFlow-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningTabularBiomedical Data

🎯 What it does: Propose a differentiable convexification optimization layer that can enforce input-related mixed logical and linear constraints within neural networks, achieving hard constraint satisfaction during both training and inference phases.

Dismantling Pathological Shortcuts: A Causal Framework for Faithful LVLM Decoding

Liu Yu (University of Electronic Science and Technology of China), Gillian Dobbie (University of Auckland)

Explainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringVision Language ModelDiffusion modelScore-based ModelAuto EncoderContrastive LearningImageTextMultimodality

🎯 What it does: Propose a training-agnostic causal framework named Fox to eliminate object hallucinations in large vision-language models during inference;

Dismantling the Illusion of Vision-Language-Action Models Competence via Explicit Distributional Shifts

Xueyang Zhou (Huazhong University of Science and Technology), Yongchao Chen (Tsinghua University)

Domain AdaptationComputational EfficiencyRepresentation LearningRobotic IntelligenceReinforcement Learning from Human FeedbackTransformerPrompt EngineeringVision-Language-Action ModelDiffusion modelFlow-based ModelContrastive LearningImageTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: This paper proposes the LIBERO-Gen benchmark, systematically constructing multi-dimensional, hierarchical distribution shift evaluations on Vision-Language-Action (VLA) models, distinguishing three scenarios: within-training distribution, compositional transfer, and domain transfer.

Dispersion Loss Counteracts Embedding Condensation and Improves Generalization in Small Language Models

Chen Liu (Yale University), Smita Krishnaswamy (Yale University)

Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelContrastive LearningText

🎯 What it does: Investigate the geometric phenomena of word embeddings in Transformer models, discovering that small models exhibit 'embedding consolidation,' leading to highly similar representation directions;

DisPOSE: Projected Polystochastic Diffusion for Self-Supervised Multi-View 3D Human Pose Estimation

Tony Danjun Wang (Technical University of Munich), Lennart Bastian (Technical University of Munich)

Pose EstimationGraph Neural NetworkDiffusion modelContrastive LearningImagePoint CloudBenchmark

🎯 What it does: Propose a self-supervised multi-view 3D human pose estimation framework called DisPOSE, which transforms the multi-view person identity association problem into a diffusion process of multivariate Polystochastic tensors, enabling training without 3D annotations.

DisPPO: Quantile-Based Distributional Reinforcement Learning for Large Language Models

Zhijian Zhou (Fudan University), Yuan Qi (Fudan University)

OptimizationComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningTextSequential

🎯 What it does: Propose a distributed reinforcement learning framework called DisPPO based on quantiles, integrating distributed value functions into the PPO training of LLMs.

Dissect and Prune: Enhancing Robustness in AI-Generated Image Detection

Dahye Kim (Korea AI Safety Institute), Jang-Ho Choi (ETRI)

Anomaly DetectionExplainability and InterpretabilityComputational EfficiencyRepresentation LearningConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningImage

🎯 What it does: Propose the DEAR method, which enhances the robustness of AI-generated image detection by aligning features on inpainted images and performing bidirectional pruning.

Dissecting Causal Mechanism Shifts via FANS: Function And Noise Separation

Gyeongdeok Seo (Yonsei University), Kyungwoo Song (Yonsei University)

Anomaly DetectionExplainability and InterpretabilityGraph Neural NetworkFlow-based ModelContrastive LearningImageTabularBiomedical Data

🎯 What it does: Propose the FANS framework, which detects and separates functional changes from noise changes in causal mechanism transitions by utilizing causal regularized flow.

Dissecting Multimodal In-Context Learning: Modality Asymmetries and Circuit Dynamics in modern Transformers

Yiran Huang (Technical University Of Munich), Zeynep Akata (Technical University Of Munich)

ClassificationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningMeta LearningTransformerPrompt EngineeringMixture of ExpertsVision Language ModelContrastive LearningImageTextMultimodality

🎯 What it does: This paper systematically investigates the emergence mechanisms of multimodal in-context learning (ICL) and its relationship with data and architectural factors by training small Transformers on controlled synthetic classification tasks.

Dissecting Post-Training: Uncovering the Complementary Roles of SFT and RL for Document Parsing

Jun-Peng Jiang (Nanjing University), Han-Jia Ye (Nanjing University)

Reinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision Language ModelImageTextMultimodality

🎯 What it does: Studied and compared the performance of supervised fine-tuning (SFT) and reinforcement learning (RL) during the post-training phase of multi-modal large language models (MLLMs) for document parsing, finding that the two methods complement each other in structural learning and content refinement, and proposed a joint post-training strategy that combines their advantages.

Dissecting Quantization Error: A Concentration-Alignment Perspective

Marco Federici (Qualcomm AI Research), Markus Nagel (Qualcomm AI Research)

Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelContrastive LearningText

🎯 What it does: Proposes a quantitative framework based on signal-to-quantization-noise ratio (SQNR), decomposing the quantization error of linear layers into the concentration of weights/activations and the alignment of directions, and based on this, designs a trainable block-level alignment and concentration transformation (CAT) to simultaneously improve these two factors.

Dissecting the Safety Circuit: Neuronal Intervention for Transferable Adversarial Attacks on VLMs

Chunlong Xie (Chongqing University), Tao Xiang (Chongqing University)

Safty and PrivacyExplainability and InterpretabilityAdversarial AttackSupervised Fine-TuningVision Language ModelDiffusion modelContrastive LearningImageTextMultimodality

🎯 What it does: Identify safety-related neural circuits in Vision-Language Models (VLMs) through linear probing, and propose the Safety Circuit Intervention Attack (SCIA) framework, which utilizes dual-objective adversarial optimization to suppress defensive circuits and amplify transferable circuits, achieving transferable attacks in black-box and cross-model scenarios;

DISSOLVR: An Interpretable and Fast Framework for Aqueous and Organic Solubility Prediction

Vansh Ramani (Indian Institute of Technology Delhi), Tarak Karmakar (Indian Institute of Technology Delhi)

Explainability and InterpretabilityComputational EfficiencyDrug DiscoveryGraph Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringGraphTabularRetrieval-Augmented Generation

🎯 What it does: Developed an interpretable solubility prediction framework named DISSOLVR, which achieves rapid prediction through manually crafted physically grounded features and gradient boosting trees, and combines post-hoc explanation generation with LLM to produce chemically readable narratives.

DistFlow: A Fully Distributed RL Framework for Scalable and Efficient LLM Post-Training

Zhixin Wang (Zhejiang University), Yuan Cheng (Shanghai Innovation Institute)

Federated LearningComputational EfficiencyTransformerLarge Language ModelReinforcement LearningVision Language ModelTextMultimodality

🎯 What it does: Proposes DISTFLOW, a fully distributed RL framework that decouples data and control flow, achieving high throughput and linear scalability for large-scale LLM and VLM training.

Distillation Models are Good Samplers for Diffusion Reinforcement Learning

Zunxu Liu (University of Science and Technology of China), Tao Mei (HiDream.ai Inc)

GenerationData SynthesisComputational EfficiencyKnowledge DistillationTransformerReinforcement LearningDiffusion modelScore-based ModelContrastive LearningImageVideoTextMultimodality

🎯 What it does: Propose the DMSampler framework, which replaces the time-consuming multi-step sampling with a co-advancing distillation model, significantly accelerating the reinforcement learning process of diffusion models.

Distilling Linearized Behavior into Non-linear Fine-Tuning for Effective Task Arithmetic

Thomas Sommariva (University of Modena and Reggio Emilia), Angelo Porrello (University of Modena and Reggio Emilia)

Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerSupervised Fine-TuningDiffusion modelAuto EncoderContrastive LearningImageText

🎯 What it does: A model that can maintain nonlinear expressiveness during inference while possessing the task vector separability and editability of linearized models is constructed by incorporating distillation of linearized models into nonlinear fine-tuning.

Distilling Neuro-Symbolic Programs into 3D Multi-modal LLMs

Wentao Mo (Peking University), Yang Liu (Peking University)

Explainability and InterpretabilityKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision Language ModelTextMultimodalityPoint CloudChain-of-Thought

🎯 What it does: Propose the APEIRIA framework, distilling the reasoning mode of neuro-symbolic programs into a 3D multimodal LLM in the form of chain-of-thought (CoT), achieving interpretable spatial reasoning.

Distilling Task-Level Coordination Policies for Generalizable Multi-Agent Cooperation

Zimo Zhai (Beijing Institute of Technology), Wei Liang (Beijing Institute of Technology)

Knowledge DistillationTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringText

🎯 What it does: Proposes a self-supervised SynCoord pipeline that distills task-level coordination knowledge generated by large language models into lightweight agents, achieving efficient and transferable multi-agent collaboration.

Distinguishable Deletion: Unifying Knowledge Erasure and Refusal for Large Language Model Unlearning

Puning Yang (MBZUAI), Xiuying Chen (MBZUAI)

Safty and PrivacyExplainability and InterpretabilityTransformerLarge Language ModelScore-based ModelContrastive LearningTextBenchmark

🎯 What it does: Propose the Distinguishable Deletion (D²) framework to achieve unified knowledge forgetting and safe rejection in LLMs;

Distinguishing Imitation Error from Intrinsic Motion Learning Difficulty

Zhaorui Meng (Xiamen University), Yipeng Qin (Cardiff University)

Pose EstimationAnomaly DetectionRobotic IntelligenceReinforcement LearningContrastive LearningOptical FlowPoint CloudSequentialBenchmark

🎯 What it does: Proposed and implemented Torque Variation Score (TVS) — a physics metric based on rigid body dynamics that is strategy-independent, used to quantify the intrinsic learning difficulty of action sequences and for error attribution.

DistMatch: Adaptive Binning via Distribution Matching for Robust Sequential Conformal Prediction

Enver Menadjiev (Korea Advanced Institute of Science and Technology), Jaesik Choi (Korea Advanced Institute of Science and Technology)

Anomaly DetectionFederated LearningExplainability and InterpretabilityComputational EfficiencyScore-based ModelContrastive LearningTabularTime SeriesSequentialFinance Related

🎯 What it does: This paper proposes DistMatch, an adaptive binning method based on the recursive partitioning of residuals using the Kolmogorov-Smirnov statistic, for achieving verifiable sequential conformal prediction in time series with distribution drift.

Distortion of AI Alignment Revisited: RLHF is a Decent Utilitarian Aligner

Kazusato Oko (University of California, Berkeley), Han Bao (Institute of Statistical Mathematics and Graduate University for Advanced Studies)

Reinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningContrastive LearningTextBenchmark

🎯 What it does: This paper provides a detailed theoretical analysis of the average distortion of RLHF in a multi-user preference environment, and presents upper and lower bounds.

Distributed Direct Preference Optimization

Zhanhong Jiang (Iowa State University)

OptimizationFederated LearningReinforcement Learning from Human FeedbackSupervised Fine-TuningReinforcement LearningText

🎯 What it does: Studied the convergence properties and time complexity of distributed Direct Preference Optimization (DPO) in federated learning and decentralized learning environments.

Distributed Stochastic $K$-Level Optimization Over Networks

Xinwen Zhang (Temple University), Heng Huang (University of Maryland)

OptimizationFederated LearningTabular

🎯 What it does: Propose a decentralized multi-layer stochastic optimization algorithm that combines recursive Hessian inverse vector product, variance reduction, and error feedback compression to achieve efficient computation and communication;