arXivSub Start free trial

ICML 2026 Papers — Page 17

International Conference on Machine Learning · 6554 papers

DRIFT-BENCH: Diagnosing CoopeRative Breakdowns in LLM Agents under Input Faults via Multi-Turn Interaction

Han Bao (University of Notre Dame), Yanfang Ye (University of Notre Dame)

Safty and PrivacyExplainability and InterpretabilityTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Designed and implemented the DRIFT-BENCH benchmark to diagnose multi-round clarification and execution safety issues of LLM agents under input failures.

DRIFT: Decoupled Rollouts and Importance-Weighted Fine-Tuning for Efficient Multi-Turn Optimization

Jian Mu (Hong Kong University of Science and Technology), Yao Shu (Hong Kong University of Science and Technology)

OptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningText

🎯 What it does: Proposes the DRIFT method, which achieves multi-round self-correction optimization for LLMs by offline sampling multiple rounds of interaction trajectories under a fixed reference policy, calculating importance weights, and then using weighted supervised fine-tuning.

DRIVE: Best Data Scheduling Practices for Reinforcement Learning with Verifiable Reward in Competitive Code Generation

Speed Zhu (Tencent), Wiggin Zhou (Tencent)

OptimizationData-Centric LearningAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextSequentialChain-of-Thought

🎯 What it does: Propose a complete RLVR training pipeline based on hard sample prioritization supervised fine-tuning and two-stage reinforcement learning, focusing on data scheduling and curriculum design to enhance competitive programming code generation performance.

DRIVE: Distributional and Retrieval-Augmented Bidding with Value Evaluation

Miduo Cui (Shandong University), Zhiwei Xu (Shandong University)

Recommendation SystemOptimizationTransformerReinforcement LearningTabularSequentialRetrieval-Augmented Generation

🎯 What it does: Propose a unified framework called DRIVE for offline ad bidding, which integrates distributed action modeling, retrieval-enhanced candidate generation, and value-based evaluation to generate and select multi-modal bidding strategies.

DriveWorld-VLA: Unified Latent-Space World Modeling with Vision–Language–Action for Autonomous Driving

Feiyang Jia, Long Chen

Autonomous DrivingOptimizationRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringVision-Language-Action ModelDiffusion modelFlow-based ModelContrastive LearningWorld ModelImageTextMultimodality

🎯 What it does: Proposes DriveWorld-VLA, an end-to-end framework that jointly models and plans visual-language-action (VLA) and world models (WM) within a shared latent space;

DRL-STAF: A Deep Reinforcement Learning Framework for State-Aware Forecasting of Complex Multivariate Hidden Markov Processes

Manrui Jiang (Tsinghua University), Chen Zhang (Tsinghua University)

OptimizationExplainability and InterpretabilityRepresentation LearningRecurrent Neural NetworkGraph Neural NetworkTransformerReinforcement LearningAuto EncoderContrastive LearningTabularTime SeriesSequentialBenchmarkFinance Related

🎯 What it does: Proposes the DRL-STAF framework, which uses deep reinforcement learning to directly estimate the discrete hidden states of a multivariate hidden Markov process and jointly predict the next observation.

DroneDINO: Towards Heterogeneous Routed Mixture of Experts for Drone-based Unified Object Detection

Dongdong Li (National University Of Defense Technology), Pengfei Zhu (National University Of Defense Technology)

Object DetectionTransformerMixture of ExpertsImageMultimodality

🎯 What it does: This paper proposes a unified drone object detection framework called DroneDINO, which can handle three input formats—RGB, IR, and RGB-IR—in the same model and achieve multi-task learning.

Drop-in Circulant Structural Priors for Transformer Decoding of Cyclic Codes

Shuai Xiao (Shandong University), Qiaosheng Zhang (Shanghai Innovation Institute)

Computational EfficiencyRepresentation LearningTransformerContrastive LearningGraphTabularBenchmarkPhysics Related

🎯 What it does: For cyclic codes, this paper proposes a pluggable cyclic structure prior method, tightly coupling the algebraic structure of cyclic codes with the Transformer decoder to achieve explicit encoding of the decoding process.

Dropout Universality: Scaling Laws and Optimal Scheduling at the Edge-of-Chaos

Lucas Fernandez-Sarmiento

ClassificationOptimizationComputational EfficiencyConvolutional Neural NetworkTransformerContrastive LearningImage

🎯 What it does: This paper studies the role of Dropout at the 'edge of chaos' in neural networks based on mean-field theory, derives its perturbation, criticality, and cross-scaling laws for signal propagation, and proposes a forward-weighted Dropout scheduling scheme, proving that it can significantly reduce test loss under a fixed budget.

DropoutTS: Sample-Adaptive Dropout for Robust Time Series Forecasting

Siru Zhong (Hong Kong University of Science and Technology), Yuxuan Liang (Hong Kong University of Science and Technology)

Anomaly DetectionComputational EfficiencyRepresentation LearningTransformerScore-based ModelAuto EncoderTime Series

🎯 What it does: This paper proposes the DropoutTS plugin, which enhances robustness in time series forecasting through a sample-adaptive Dropout mechanism.

DRPBench: Evaluating LLMs in Concurrent Code Comprehension via Fine-Grained Data Race Prediction

Yuqi Guo (Key Laboratory of System Software, Chinese Academy of Sciences), Yan Cai (Key Laboratory of System Software, Chinese Academy of Sciences)

AI Code AssistantTransformerLarge Language ModelAgentic AIPrompt EngineeringTextSequentialBenchmarkRetrieval-Augmented Generation

🎯 What it does: Constructed the DRPBench benchmark to evaluate the data race prediction capability of large language models in concurrent code understanding, and conducted a fine-grained evaluation of 15 LLMs.

DSB: Dynamic Sliding Block Scheduling for Diffusion LLMs

Lizhuo Luo (Nanyang Technological University), Tianwei Zhang (Nanyang Technological University)

Computational EfficiencyAI Code AssistantTransformerLarge Language ModelDiffusion modelText

🎯 What it does: Proposed an untrained dynamic sliding block scheduling (DSB) and the corresponding KV cache scheme (DSB Cache) to improve the semi-autoregressive inference of diffusion-based large language models (dLLMs).

DSENet: A Novel Dual-Stream Enhancement Network for Multi-Scale Non-Stationary Time Series Forecasting

Yuhan Wang (Shanghai Jiao Tong University), Jinhong Guo (Shanghai Jiao Tong University)

Explainability and InterpretabilityComputational EfficiencyTransformerMixture of ExpertsAuto EncoderContrastive LearningTime SeriesBiomedical DataElectronic Health Records

🎯 What it does: Proposes the DSENet dual-stream enhanced network, which models the problem of local mutations in long sequences for blood glucose prediction.

DSGCR: Decomposed Spectral Geometry-Aware Cross-Modal Semantic Representation for 3D Visual Grounding

Jing He (Xidian University), Long Sun (Xidian University)

RecognitionDomain AdaptationRepresentation LearningTransformerSupervised Fine-TuningVision Language ModelContrastive LearningTextMultimodalityPoint Cloud

🎯 What it does: In the 3D visual localization task, two modules, Text-Aware Feature Tuning (TFT) and Decomposed Spectral Geometry (DSG), are proposed, significantly enhancing cross-modal semantic alignment and spatial reasoning capabilities.

DSGym: A Standardized and Holistic Framework for Evaluating and Training Data Science Agents

Fan Nie (Stanford University), James Zou (Stanford University)

Data-Centric LearningAI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningAgentic AIPrompt EngineeringTextTabularBiomedical DataBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose the DSGym unified framework, using containerized Jupyter environments to evaluate and train data science agents, and construct task sets such as DSGym-TASKS, DSBIO, and DSPREDICT;

DTKG: Dual-Track Knowledge Graph-Verified Reasoning Framework for Multi-Hop QA

Changhao Wang (Beihang University), Yunfeng Lu (Beihang University)

RetrievalExplainability and InterpretabilityComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerPrompt EngineeringMixture of ExpertsTextGraphBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This paper proposes the DTKG framework, which dynamically classifies multi-hop questions into two categories: parallel verification and chain reasoning, and adopts targeted strategies for knowledge graph retrieval and fact verification respectively.

DTop-p MoE: Sparsity-Controlled Dynamic Top-p MoE for Foundation Model Pre-training

Can Jin (Rutgers University), Dimitris N. Metaxas (Rutgers University)

Computational EfficiencyRepresentation LearningTransformerLarge Language ModelMixture of ExpertsImageText

🎯 What it does: Proposed a controllable sparse dynamic top-p routing mechanism DTOP-p MoE, used to achieve adaptive expert selection at the token level and layer level while maintaining the global computational budget;

dTRPO : Trajectory Reduction in Policy Optimization of Diffusion Large Language Models

Wenxuan Zhang (Meta AI), Wei Wen (Meta AI)

OptimizationComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringDiffusion modelTextBenchmark

🎯 What it does: For strategy optimization of diffusion-based large language models (dLLMs), the trajectory reduction technique (dTRPO) is proposed, achieving efficient offline alignment training by reducing the number of forward inference steps.

DTS: Enhancing Large Reasoning Models via Decoding Tree Sketching

Zicheng Xu (Johns Hopkins University), Vladimir Braverman (Johns Hopkins University)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose a no-training, plug-and-play decoding framework called Decoding Tree Sketching (DTS), which systematically explores the reasoning tree of large inference models and selects the shortest and most reliable reasoning trajectory by expanding branches according to probability distributions at decision points and terminating early during decoding.

Dual Latent Memory for Visual Multi-agent System

Xinlei Yu (National University of Singapore), Shuicheng YAN (National University of Singapore)

Computational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackTransformerVision Language ModelAuto EncoderContrastive LearningImageVideoTextMultimodalityRetrieval-Augmented Generation

🎯 What it does: Proposes the L‑VMAS2 framework, which utilizes dual latent memories (perception memory and thinking memory) to achieve efficient collaboration in visual multi-agent systems, and implements on-demand memory invocation through an entropy-driven active triggering mechanism.

Dual Mechanisms of Value Expression: Intrinsic vs. Prompted Values in Large Language Models

Jongwook Han (Seoul National University), Yohan Jo (Seoul National University)

Explainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextMultimodalityBenchmark

🎯 What it does: This paper investigates two mechanisms by which large language models express values: intrinsic expression and prompted expression, and decomposes these mechanisms at the activation vector and neuron levels.

Dual Optimal Transport for Multi-Concept Composition: Structural Alignment and Texture Injection in Diffusion Models

Hao Fu (Shandong University), Tian Gan (Shandong University)

GenerationData SynthesisTransformerPrompt EngineeringDiffusion modelContrastive LearningImageTextBenchmark

🎯 What it does: Train-free multi-concept personalized generation of diffusion models through dual optimal transport (Barycentric Soft-Transport and Geometry-Guided Transport) in three stages (layout planning, structural sketching, and texture rendering).

Dual Quaternion SE(3) Synchronization with Recovery Guarantees

Jianing Zhao (Chinese University of Hong Kong), Anthony Man-Cho So (Chinese University of Hong Kong)

Pose EstimationOptimizationContrastive LearningGaussian SplattingSimultaneous Localization and MappingOptical FlowPoint CloudMesh

🎯 What it does: A two-stage algorithm is proposed by mapping the SE(3) synchronization problem to the unit dual quaternion space: first, spectral initial values are obtained using power iteration, and then refined using the dual quaternion generalized power method (DQGPM) with projection at each step, along with an upper bound on the finite iteration error.

Dual-branch Robust Unlearnable Examples

Xianlong Wang (City University of Hong Kong), Xiaohua Jia (City University of Hong Kong)

ClassificationAdversarial AttackConvolutional Neural NetworkContrastive LearningImage

🎯 What it does: Proposed a dual-branch robust non-learnable example (DUNE) generation framework, which reduces the model's learnability by optimizing perturbations in the spatial and color domains separately to form a mapping between features and labels.

Dual-Calibration Multi-View Clustering via Compact Anchor Learning

Huibing Wang (Dalian Maritime University), Ximing Li (Jilin University)

Representation LearningContrastive LearningImageText

🎯 What it does: A bidirectional calibration multi-view clustering method called DCMC is proposed, which significantly improves the quality of anchors by jointly learning anchors and clustering assignments.

Dual-channel Dynamic Graph Neural Networks with Adaptive Adjacency Learning and Multi-scale Representation Fusion

Youqing Wang (Beijing University of Chemical Technology), Jipeng Guo (Beijing University of Chemical Technology)

ClassificationRepresentation LearningGraph Neural NetworkContrastive LearningGraph

🎯 What it does: Propose a dual-channel dynamic graph neural network, DCD-GNN, which integrates static structure and dynamic self-attention adjacency learning, and fuses multi-scale features to enhance node classification performance.

Dual-Latent Memory Routing for Vision-Language Reasoning

Hao-Xuan Ma (Nanjing University), Han-Jia Ye (Nanjing University)

Explainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelSupervised Fine-TuningVision Language ModelMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed a parameter-efficient dual latent memory routing mechanism (DLMR) that provides separable visual memory and reasoning memory for multi-modal large language models, and dynamically selects when and how much memory to inject during the reasoning process;

Dual-stage Contrastive Learning-enhanced Multi-view Variational Clustering

Yanxi Liu (Harbin Institute of Technology), Guoqing Chao (Harbin Institute of Technology)

Representation LearningTransformerAuto EncoderContrastive LearningImageMultimodalityTabular

🎯 What it does: This paper proposes a dual-stage contrastive learning enhanced multi-view variational clustering model, DCL-MVC, which jointly integrates view fusion and potential space contrastive regularization to achieve cross-view consistency and denoising of boundary samples.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

John Won (Korea Advanced Institute of Technology), Jinwoo Shin (Korea Advanced Institute of Technology)

Robotic IntelligenceTransformerVision-Language-Action ModelDiffusion modelFlow-based ModelContrastive LearningWorld ModelImageVideoTextMultimodality

🎯 What it does: Propose the Dual-Stream Diffusion (DUST) framework, combining world models with vision-language-action (VLA) models to achieve joint prediction of actions and vision and causal learning.

Dual-View Predictive Diffusion: Lightweight Speech Enhancement via Spectrogram-Image Synergy

Ke Xue (Beijing Institute of Technology), Jianping An (Beijing Institute of Technology)

RestorationConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningImageStochastic Differential EquationAudio

🎯 What it does: Proposed a highly lightweight dual-perspective prediction diffusion model (DVPD) for speech enhancement, which takes into account both the visual texture and physical frequency domain characteristics of the spectrum.

DualCOIL: Offline Imitation Learning from Contrasting Demonstrations

Huy Hoang (Singapore Management University), Tanvi Verma (Institute of High Performance Computing, Agency for Science, Technology and Research)

Reinforcement LearningContrastive LearningTabularTime SeriesSequential

🎯 What it does: DualCOIL simultaneously utilizes expert demonstrations and explicit poor demonstrations in offline imitation learning, learning a strategy that can both imitate excellent behaviors and avoid bad behaviors.

DualOptim+: Bridging Shared and Decoupled Optimizer States for Better Machine Unlearning in Large Language Models

Xuyang Zhong (City University of Hong Kong), Chen Liu (City University of Hong Kong)

OptimizationSafty and PrivacyTransformerLarge Language ModelSupervised Fine-TuningText

🎯 What it does: Proposes the DualOptim+ optimization framework to improve the unlearning performance of large models.

DualTimesField: Rethinking Time Series as Continuous-Time Trends and Events

Wencheng Zhang (Northwest Normal University), Wanghu Chen (Northwest Normal University)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerNeural Radiance FieldAuto EncoderContrastive LearningTabularTime SeriesBiomedical DataElectrocardiogramBenchmarkStochastic Differential Equation

🎯 What it does: Propose DualTimesField, which models time series by decomposing them into continuous trends and discrete events through dual implicit neural fields.

DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing

Lei Gao (University of Southern California), Murali Annavaram (University of Southern California)

OptimizationComputational EfficiencyTransformerLarge Language ModelText

🎯 What it does: DuetServe is an adaptive LLM service framework that can dynamically perform fine-grained partitioning of SMs on a single GPU, decoupling the prefill and decode stages when needed, thus maintaining high throughput in aggregated mode while providing discretized isolation when facing TBT (Time-Between-Tokens) SLO risks;

DuRP: Dual-Stage Physics-Embedded Learning for Joint Radiance and Polarization Restoration

Zhenshuo Yang (State Key Laboratory of Robotics and Intelligent Systems, Shenyang Institute of Automation, Chinese Academy of Sciences), Jiandong Tian (State Key Laboratory of Robotics and Intelligent Systems, Shenyang Institute of Automation, Chinese Academy of Sciences)

RestorationConvolutional Neural NetworkTransformerSupervised Fine-TuningDiffusion modelContrastive LearningImagePhysics Related

🎯 What it does: Proposes the DuRP framework, which utilizes a two-stage physical embedded learning approach to jointly restore scene irradiance and polarization information under foggy conditions.

Dustin: Draft-Augmented Sparse Verification for Efficient Long-Context Generation with Speculative Decoding

WenHung Lee (National Yang Ming Chiao Tung University), Kai-Chiang Wu (National Yang Ming Chiao Tung University)

GenerationComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: Propose a sparse verification framework called DUSTIN to accelerate speculative decoding for multi-batch long context.

DV-World: Benchmarking Data Visualization Agents in Real-World Scenarios

Jinxiang Meng (Chinese Academy of Sciences), Kang Liu (Chinese Academy of Sciences)

Data-Centric LearningAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringWorld ModelTextMultimodalityTabularBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Constructed the DV-World benchmark, which covers the complete lifecycle of data visualization (DV) in real workflows: native spreadsheet visualization (creation, diagnostic repair, dashboard composition), cross-framework visualization evolution (image-to-code migration), and proactive intent alignment (multi-round interaction with a user simulator).

DyCon: Dynamic Reasoning Control via Evolving Difficulty Modeling

Tengyao Tu (Harbin Institute of Technology), Min Zhang (Harbin Institute of Technology)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextBenchmarkChain-of-Thought

🎯 What it does: Through the untrained DyCon framework, it leverages the hidden layer representations of large-scale inference models to estimate the task difficulty in real-time during the inference process, thereby dynamically adjusting the logical scores of reflection keywords, reducing redundant reasoning steps, and improving inference efficiency.

DyGRO-VLA: Cross-Task Scaling of Vision–Language–Action Models via Dynamic Grouped Residual Optimization

Sixu Lin (School of Data Science, Chinese University of Hong Kong (Shenzhen)), Guiliang Liu (School of Data Science, Chinese University of Hong Kong (Shenzhen))

OptimizationRobotic IntelligenceReinforcement Learning from Human FeedbackTransformerReinforcement LearningMixture of ExpertsVision-Language-Action ModelContrastive LearningImageVideoTextMultimodality

🎯 What it does: Proposes the DyGRO-VLA framework, combining information bottleneck and dynamic hybrid residual RL experts to achieve cross-task scalability and fine-grained optimization of VLA models;

DyLLM: Efficient Diffusion LLM Inference via Saliency-based Token Selection and Partial Attention

Younjoo Lee (Seoul National University), Jung Ho Ahn (Seoul National University)

GenerationComputational EfficiencyTransformerLarge Language ModelDiffusion modelText

🎯 What it does: Propose DyLLM, an untrained inference framework that achieves efficient inference for Diffusion LLM by identifying and only computing the 'significant words' that change significantly during diffusion steps.

Dyn-VPP: Video Prediction Policy Optimization for Improved Visual Dynamics

Zirui Ge (Zhejiang University), Donglin Wang (Westlake University)

OptimizationRobotic IntelligenceTransformerReinforcement LearningDiffusion modelAuto EncoderVideoSequential

🎯 What it does: This work proposes a post-training framework called Dyn-VPP, which further improves the visual dynamics accuracy of video prediction models in robotic manipulation by treating multi-step diffusion denoising as a policy optimization process.

DynaMem: Consistent Long Video Generation via Hierarchical Memory and Motion Priors

Jingyu Lin (Monash University), Zhiqiang Shen (Mohamed bin Zayed University of Artificial Intelligence)

GenerationData SynthesisTransformerDiffusion modelContrastive LearningOptical FlowVideoText

🎯 What it does: This paper proposes DYNAMEM, a unified framework that maintains semantic consistency, motion coherence, and appearance stability in autoregressive long video generation.

Dynamic Compression Flows for Neuroscience Data

Ganchao Wei (Duke University), John Pearson (Duke University)

CompressionRepresentation LearningConvolutional Neural NetworkRecurrent Neural NetworkDiffusion modelScore-based ModelFlow-based ModelAuto EncoderImageVideoBiomedical DataStochastic Differential EquationOrdinary Differential EquationAudio

🎯 What it does: Propose the Dynamic Compression Flows (DCF) framework, which utilizes dual flow-matching to learn compression flows and dynamic flows, directly extracting low-dimensional, time-dynamic latent representations from high-dimensional data.

Dynamic Decision Learning: Test-Time Evolution for Abnormality Grounding in Rare Diseases

Jun Li (Technical University of Munich), Julia Schnabel (Technical University of Munich)

Anomaly DetectionExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringVision Language ModelDiffusion modelContrastive LearningImageBiomedical DataMagnetic Resonance ImagingBenchmark

🎯 What it does: Propose the Dynamic Decision Learning (DDL) framework, which enables test-time dynamic decision-making for rare disease anomaly localization on frozen vision-language models through prompt optimization and visual consistency verification.

Dynamic Fractal Mamba: A Neural Renormalization Group Flow for Scale-Invariant Sequence Modeling

Shenglei Fang (Cardiff University), You Zhou (Cardiff University)

ClassificationRecognitionComputational EfficiencyRepresentation LearningTransformerAuto EncoderContrastive LearningTabularTime SeriesBiomedical DataBenchmark

🎯 What it does: Proposed a recursive parameter-sharing state space model called DF-Mamba, which can maintain dynamic consistency across multiple scales, achieve an exponentially expanded receptive field, while maintaining linear computational complexity.

Dynamic High-Dimensional Facility Location with Low Recourse

Sayan Bhattacharya (University of Warwick), Nikos Parotsidis (Google)

OptimizationContrastive Learning

🎯 What it does: This paper proposes a dynamic facility location algorithm tailored for non-uniform facility opening costs in high-dimensional Euclidean spaces, which can maintain an approximately optimal solution under real-time insertions/deletions of customers, achieving a constant approximation ratio, logarithmic amortized scheduling cost, and geometric amortized update time.

Dynamic Linear Attention

Xin Wang (Ohio State University), Mi Zhang (Ohio State University)

Computational EfficiencyRepresentation LearningTransformerLarge Language ModelTextSequential

🎯 What it does: Propose a dynamic linear attention framework called DLA, which uses information-aware dynamic state merging and capacity-limited memory modeling to achieve efficient processing of long sequences.

Dynamic Multimodal Evaluation via Knowledge-Enhanced Benchmark Evolution

Junzhe Zhang (Peking University), Xiaojun Wan (Peking University)

Data-Centric LearningGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextMultimodalityGraphBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposes the KBE-DME framework, which dynamically evolves VQA benchmarks using graph structures and knowledge-enhancement techniques to achieve controllable difficulty in multimodal evaluation.

Dynamic Optimizations of LLM Ensembles with Two-Stage Reinforcement Learning Agents

Selim Furkan Tekin (Georgia Institute of Technology), Ling Liu (Georgia Institute of Technology)

OptimizationFederated LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningAgentic AIMixture of ExpertsTextBenchmark

🎯 What it does: This paper proposes a two-stage reinforcement learning framework, RL-Focal, for dynamically selecting a subset of small-scale LLMs and fusing their outputs, thereby improving reasoning performance in multi-task environments.

Dynamic Programming for Epistemic Uncertainty in Markov Decision Processes

Axel Benyamine (Institut Polytechnique de Paris), Alain Oliviero Durmus

Reinforcement Learning

🎯 What it does: Propose a novel fuzzy Markov decision process framework that treats transition probabilities as random variables and evaluates rewards using risk measures, unifying various uncertainty models.

Dynamic Regret via Discounted-to-Dynamic Reduction with Applications to Curved Losses and Adam Optimizer

Yan-Feng Xie (Nanjing University), Zhi-Hua Zhou (Nanjing University)

OptimizationFederated LearningMeta LearningReinforcement Learning from Human FeedbackReinforcement LearningContrastive LearningImage

🎯 What it does: This paper studies dynamic loss minimization in non-stationary online learning by proposing a modular discounted-dynamic reduction method, and applies this method to curvature loss (online linear regression and online logistic regression) as well as the convergence analysis of the Adam optimizer.

Dynamic Relational Priming Improves Transformer in Multivariate Time Series

Hunjae Lee (Southern Methodist University), Corey Clark (Southern Methodist University)

Computational EfficiencyTransformerPrompt EngineeringContrastive LearningTime SeriesBenchmark

🎯 What it does: Proposed a dynamic relationship modulation mechanism called Prime Attention for multivariate time series (MTS) prediction;

Dynamic Stratified Contrastive Learning with Upstream Augmentation for MILP Branching

Tongkai Lu (Beihang University), Chongyang Tao (Beihang University)

OptimizationGraph Neural NetworkReinforcement LearningContrastive LearningGraphTabularBenchmark

🎯 What it does: This paper proposes SC-MILP, a dynamic hierarchical contrastive learning framework for branching in mixed integer linear programming (MILP).

Dynamic Symmetric Point Tracking: Tackling Non-ideal Reference in Analog In-memory Training

Quan Xiao (Cornell University), Tianyi Chen (Cornell University)

OptimizationComputational EfficiencyRepresentation LearningConvolutional Neural NetworkAuto EncoderContrastive LearningImage

🎯 What it does: A dynamic symmetric point tracking algorithm (RIDER) and its enhanced version (E-RIDER) are proposed for real-time calibration and tracking of the device's symmetric points during training on the analog in-memory computing (AIMC) platform, compensating for system drift caused by asymmetric updates.

Dynamic Thinking-Token Selection for Efficient Reasoning in Large Reasoning Models

Zhenyuan Guo (Zhejiang University), Wenzhi CHEN

Computational EfficiencyKnowledge DistillationTransformerLarge Language ModelTextChain-of-Thought

🎯 What it does: This paper proposes a dynamic thinking-token selection method (DYNTS), which compresses the memory and computation of large inference models by predicting the importance of each token during the thinking phase to the final answer and retaining only key KV caches.

Dynamic TMoE: A Drift-Aware Dynamic Mixture of Experts Framework for Non-Stationary Time Series Forecasting

Jiawen Zhu (Zhejiang University), Yingcai Wu (Zhejiang University)

Anomaly DetectionRecurrent Neural NetworkTransformerMixture of ExpertsDiffusion modelScore-based ModelContrastive LearningGaussian SplattingTabularTime SeriesFinance Related

🎯 What it does: Propose a Drift-Aware Dynamic Mixture of Experts (Dynamic TMoE) framework based on a dynamic expert pool, aimed at addressing distribution drift issues in non-stationary time series.

Dynamics and Representation Structure of Local Approximations to Gradient-Based Learning in Linear Recurrent Neural Networks

Ezekiel Williams (Universite de Montreal), Guillaume Lajoie (Universite de Montreal)

OptimizationComputational EfficiencyRepresentation LearningRecurrent Neural NetworkContrastive LearningTime SeriesSequentialStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Studies the learning dynamics, stability, and convergence speed of linear recurrent neural networks (RNNs) under local gradient descent approximations (such as RFLO, tBPTT) and full BPTT, and analyzes the impact of these local rules on the rank of the learned weight matrices.

Dynamics Are Learned, Not Told: Semi-Supervised Discovery of Latent Dynamics Geometries For Zero-Shot Policy Adaptation

Zhiming Xu (Tongji University), Chenpeng Yao (Tongji University)

Domain AdaptationRepresentation LearningTransformerReinforcement LearningAuto EncoderContrastive LearningTabularTime Series

🎯 What it does: A framework is studied that directly learns the latent dynamic geometry from interaction trajectories, achieving zero-shot policy adaptation to unseen dynamic environments.

Dynamics of neural scaling laws in random feature regression with powerlaw-distributed kernel eigenvalues

Jakob Kramp (Jülich Research Centre), Moritz Helias (Jülich Research Centre)

OptimizationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTabularPhysics RelatedStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: This paper establishes a dynamic mean-field theory to describe the learning dynamics in random feature regression with power-law distributed feature kernels, and uses it to explain the scaling laws of neural networks, early stopping strategies, and the evolution of error over time;

Dynamics Reveals Structure: Challenging the Linear Propagation Assumption

Hoyeon Chang (KAIST), Seong Joon Oh (KAIST)

OptimizationExplainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelContrastive LearningText

🎯 What it does: This study investigates whether large language models can maintain logical consistency during one-time, linear parameter updates, and proposes the Systematic Linear Propagation (SLP) theory by combining relational algebra with gradient geometry.

Dynamics Within Latent Chain-of-Thought: An Empirical Study of Causal Structure

Zirui Li (Harbin Institute of Technology), Min Zhang (Harbin Institute of Technology)

Explainability and InterpretabilityTransformerLarge Language ModelTextChain-of-Thought

🎯 What it does: This paper treats latent chain-of-thought as a controllable causal system, and systematically quantifies the causal importance of each step, information propagation paths, and internal pattern preservation through methods such as single-step 'do' interventions on hidden steps, early-stopping decoding, influence matrix inference, and multi-modal placeholder evaluation.

DynaSchedBench: Calibrated Dynamic Scheduling Benchmarks and Observability Paradox in LLM-based Scheduling Agents

Shijie Cao (Beihang University), Jing Liu (Xidian University)

OptimizationTransformerLarge Language ModelAgentic AIPrompt EngineeringTabularSequentialBenchmarkChain-of-Thought

🎯 What it does: Built DynaSchedBench, a complete experimental framework for generating, calibrating, and evaluating dynamic flexible job shop scheduling (DFJSP), and evaluated the scheduling performance of LLM agents under this framework.

DynaTok: Token-Based 4D Reconstruction from Partial Point Clouds

Weirong Chen (Technische Universitaet Muenchen), Federico Tombari (Technische Universitaet Muenchen)

GenerationRepresentation LearningTransformerFlow-based ModelAuto EncoderContrastive LearningPoint CloudOrdinary Differential Equation

🎯 What it does: DynaTok encodes partially unordered and uncorrelated point cloud sequences into compact spatiotemporal tokens, utilizes Transformer for cross-frame aggregation, and achieves geometric and motion decoupling through residual tokens, thereby enabling globally consistent 4D point cloud reconstruction without images or explicit correspondences.

DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics

Silin Gao (EPFL), Antoine Bosselut (EPFL)

GenerationRepresentation LearningTransformerLarge Language ModelMixture of ExpertsVision Language ModelDiffusion modelContrastive LearningWorld ModelImageVideoText

🎯 What it does: Propose DynaVieW, a world model based on hierarchical JSON schema, which jointly learns visual state simulation and state transition prediction.

DynVLA: Learning World Dynamics for Action Reasoning in Autonomous Driving

Shuyao Shang (Institute of Automation, Chinese Academy of Sciences), Tieniu Tan (Yinwang Intelligent Technology Co. Ltd.)

Autonomous DrivingExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerSupervised Fine-TuningReinforcement LearningPrompt EngineeringVision-Language-Action ModelContrastive LearningVideoTextMultimodalityChain-of-Thought

🎯 What it does: Design and implement the DynVLA model, proposing the Dynamics Chain-of-Thought (Dynamics CoT) mechanism within the Vision-Language-Action (VLA) framework. This mechanism first compresses and infers the future dynamics of the world, and then generates driving actions based on the predicted dynamics.

DyPE: Dynamic Position Extrapolation for Ultra High Resolution Diffusion

Noam Issachar (Hebrew University of Jerusalem), Raanan Fattal (Hebrew University of Jerusalem)

GenerationTransformerDiffusion modelImage

🎯 What it does: Propose a dynamic position extrapolation method (DYPE) that requires no training and no additional sampling cost, enabling existing diffusion Transformers to generate images at ultra-high resolutions (up to 16 million pixels).

Dywave: Event-Aligned Dynamic Tokenization for Heterogeneous IoT Sensing Signals

Tomoyoshi Kimura (University of Illinois Urbana Champaign), Tarek F. Abdelzaher (University of Illinois Urbana Champaign)

ClassificationAnomaly DetectionComputational EfficiencyRepresentation LearningConvolutional Neural NetworkTransformerSupervised Fine-TuningAuto EncoderContrastive LearningMultimodalityTabularTime SeriesBiomedical Data

🎯 What it does: Propose a dynamic event-aligned tokenization framework called Dywave based on wavelet decomposition, which converts heterogeneous IoT sensing signals into compact and semantically aligned input tokens.

E-mem: Multi-Agent Based Episodic Context Reconstruction for LLM Agent Memory

Kaixiang Wang (Shanghai Jiao Tong University), Jie LI (Shanghai Jiao Tong University)

Explainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: Propose the E-mem framework, which achieves context reconstruction based on episodic memory through a multi-agent hierarchical structure, replacing traditional preprocessing-based retrieval;

E-VAds: An E-commerce Short Videos Understanding Benchmark for MLLMs

Xianjie Liu (Alimama Tech, Taobao & Tmail Group of Alibaba), Bo Zheng (Alimama Tech, Taobao & Tmail Group of Alibaba)

Recommendation SystemReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningVision Language ModelVideoTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-ThoughtAudio

🎯 What it does: Designed and released the E-VAds benchmark specifically for e-commerce short videos, and trained a reinforcement learning-based reasoning model called E-VAds-R1 on this benchmark.

E2Former-V2: On-the-Fly Equivariant Attention with Linear Activation Memory

Lin Huang (IQuest Research Ubio Team), Jia Zhang (IQuest Research Ubio Team)

Computational EfficiencyDrug DiscoveryProtein Structure PredictionGraph Neural NetworkTransformerDiffusion modelContrastive LearningGraphBiomedical Data

🎯 What it does: Proposed E2Former-V2, an efficient 3D atomic model that achieves equivariant attention and linear activation memory at the node level.

E²I-VRWKV: Explicit EPI-Representation and Interaction-Aware Vision-RWKV for Light Field Semantic Segmentation

Wei Zhang (Tianjin University of Technology), Shengyong Chen (Tianjin University of Technology)

SegmentationTransformerAuto EncoderContrastive LearningImageBenchmark

🎯 What it does: Propose a light field semantic segmentation framework called E²I‑VRWKV based on a linear complexity Vision‑RWKV backbone network, explicitly injecting EPI geometric priors into the network and interactively fusing them through GC‑Gate.

EAGer: Entropy-Aware GEneRation for Adaptive Inference-Time Scaling

Daniel Scalena (University of Groningen), Ahmet Üstün (Cohere Labs)

GenerationComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed a training-free, token-entropy-based generation method called EAGER, which can dynamically determine where to branch and reallocate unused computational resources during inference, thereby reducing redundant computations.

EAKV: An Entropy-Driven Adaptive KV Compression Framework for Long Video Understanding

Hengrui Hu (University of Science and Technology of China), Lan Zhang (University of Science and Technology of China)

CompressionComputational EfficiencyTransformerLarge Language ModelVision Language ModelContrastive LearningVideo

🎯 What it does: Proposes EAKV, a training-free KV cache compression framework based on attention entropy, to address the memory bottleneck in long video inference.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance

Siyao Song (State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), Institute of Automation, Chinese Academy of Sciences), Kai Jia (ByteDance BandAI)

OptimizationAI Code AssistantTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsTextTabularSequentialBenchmarkChain-of-Thought

🎯 What it does: Optimize the reasoning strategy by dynamically invoking external experts during the reinforcement learning training phase, and achieve fully autonomous reasoning at test time.

EARL: Towards a Unified Analysis-Guided Reinforcement Learning Framework for Egocentric Interaction Reasoning and Pixel Grounding

Yuejiao Su (Hong Kong Polytechnic University), Yi Wang (Hong Kong Polytechnic University)

Image TranslationObject DetectionSegmentationPose EstimationComputational EfficiencyRepresentation LearningData-Centric LearningRobotic IntelligenceTransformerReinforcement LearningPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose a unified analysis-guided reinforcement learning framework called EARL for first-person perspective interactive reasoning and pixel-level localization.

Early Decisions Matter: Proximity Bias and Initial Trajectory Shaping in Non-Autoregressive Diffusion Language Models

Jiyeon Kim (KAIST AI), Minjoon Seo (KAIST AI)

GenerationOptimizationComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringDiffusion modelText

🎯 What it does: This paper systematically analyzes the temporal dynamics of non-autoregressive diffusion language models (dLLMs) during the inference process, revealing a failure mode where early unmasked positions, determined by 'proximity bias,' dominate the entire generation trajectory. It proposes two minimal intervention strategies based on a lightweight planner and EOS temperature annealing, significantly improving inference quality.

Easier to Judge than to Find: Predicting In-Context Learning Success for Demonstration Selection

Haochun Wang (Harbin Institute of Technology), Ting Liu (Harbin Institute of Technology)

ClassificationTransformerLarge Language ModelPrompt EngineeringTextBenchmark

🎯 What it does: Propose a example selection framework DiSP based on judgment rather than search, using a router and hierarchical judge to test the feasibility of prompt context within a limited budget.

EasyBalance: Cross-Layer Load Balancing in Distributed MoE Inference

Yize Wu (Intelligent Software Research Center, Institute of Software, Chinese Academy of Sciences), Yanjun Wu (Intelligent Software Research Center, Institute of Software, Chinese Academy of Sciences)

OptimizationComputational EfficiencyTransformerLarge Language ModelMixture of ExpertsTextBenchmark

🎯 What it does: Achieve cross-layer load balancing in distributed Mixture-of-Experts inference by scheduling micro-batches across different layers to improve GPU utilization;

Eating for a Sustainable Planet: Personalized Sustainable Diet Recommendation via Constraint-Aware Decision-Making Modeling

Ying Jin (University of Chinese Academy of Sciences), Shuqiang Jiang (Institute of Computing Technology)

Recommendation SystemOptimizationTransformerMixture of ExpertsContrastive LearningTextTabularAgriculture Related

🎯 What it does: The study proposes a personalized sustainable diet recommendation framework based on constraint-aware decision-making, which achieves individualized recommendations by converting sustainability into learnable user constraints, and combining multi-expert Transformers, interest attention, and multi-task sustainability prediction.

ECA: Efficient Continual Alignment for Open-Ended Image-to-Text Generation

Jiangtao Kong (William And Mary), Huajie Shao (William And Mary)

GenerationComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerMixture of ExpertsVision Language ModelContrastive LearningImageTextMultimodality

🎯 What it does: Propose the ECA method to achieve an image-to-text generation model for sample-free incremental learning, and define the concept of continuous alignment.

ECCO: Evidence-Driven Causal Reasoning for Compiler Optimization

Haolin Pan (Institute of Software at Chinese Academy of Sciences), Yanjun Wu (Institute of Software at Chinese Academy of Sciences)

OptimizationExplainability and InterpretabilityComputational EfficiencyAI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextTabularChain-of-Thought

🎯 What it does: Designed and implemented the ECCO framework, which combines LLMs with interpretable causal reasoning and genetic algorithms for compiler optimization.

ECG-R1: Protocol-Guided and Modality-Agnostic MLLM for Reliable ECG Interpretation

Jiarui Jin (Peking University), Shenda Hong (Peking University)

Anomaly DetectionExplainability and InterpretabilityTransformerLarge Language ModelReinforcement LearningPrompt EngineeringVision Language ModelMultimodalityTime SeriesElectronic Health RecordsElectrocardiogram

🎯 What it does: A protocol-guided multi-modal large language model, ECG-R1, was constructed to reliably interpret electrocardiograms (ECGs) and reduce hallucinations and errors.

ECHO: Elastic Speculative Decoding with Sparse Gating for High-Concurrency Scenarios

Xinyi. Hu, Mingcheng Wan (Qwen Applications Business Group Of Alibaba)

Computational EfficiencyAI Code AssistantTransformerLarge Language ModelMixture of ExpertsTextBenchmark

🎯 What it does: Propose the ECHO framework, which achieves inference acceleration in high-concurrency scenarios through sparse confidence gating and elastic budget scheduling.

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning

Chu Zhao (Northeastern University), Guibing Guo (Northeastern University)

TransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsTextMultimodalityChain-of-Thought

🎯 What it does: During testing, unsupervised reinforcement learning is achieved by combining entropy-confidence tree search with online pruning, as well as confidence-adaptive trimming and advantage shaping, to enhance the performance of large language models in mathematical and multimodal reasoning.

EchoAttention: Exploiting Token-Pair Redundancy and Frame-Block Similarity for Efficient Video Generation

Yifei Xia (Peking University), Bin CUI

GenerationComputational EfficiencyKnowledge DistillationTransformerPrompt EngineeringDiffusion modelScore-based ModelContrastive LearningVideo

🎯 What it does: Propose the EchoAttention framework, which integrates sparse attention with the Echo operator based on frame block similarity, to accelerate inference in video Diffusion Transformers;

Echoes within the Reasoning: Stealthy and Effective Watermarking via Chain of Thought

Jiacheng Lu (Shanghai Jiao Tong University), Jiaheng Zhang (National University Of Singapore)

Safty and PrivacyExplainability and InterpretabilityTransformerPrompt EngineeringTextChain-of-Thought

🎯 What it does: Propose BiCoT, a framework that embeds ownership watermarks internally within the Chain-of-Thought process

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs

Chao Gong (Fudan University), Jingjing Chen (Fudan University)

Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerPrompt EngineeringVision Language ModelContrastive LearningVideoTextMultimodalityAudio

🎯 What it does: This paper proposes the EchoingPixels framework to address the localization aliasing problem that occurs during the sparsification process of AVLLMs, achieving high performance with only 5-20% of the original tokens.

EchoRL: Reinforcement Learning via Rollout Echoing

Jinhe Bi (Huawei Heisenberg Research Center), Yunpu Ma (LMU Munich)

Computational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextTabularBenchmarkChain-of-Thought

🎯 What it does: Propose EchoRL, a lightweight module that identifies key segments (EchoClip) from verified successful rollouts by utilizing step-level entropy, and injects it as an auxiliary supervisory signal into RLVR training to address the gradient vanishing problem caused by advantage degradation.

ECo-MoE: Embodiment-Conditioned Mixture of Experts Increases the Evolvability of Robots

Yibin Wang (Northwestern University), Sam Kriegman (Northwestern University)

Robotic IntelligenceReinforcement LearningMixture of ExpertsAuto EncoderGenerative Adversarial NetworkContrastive LearningPoint CloudTabular

🎯 What it does: A framework for simultaneously evolving robot morphology and control strategies is studied, using a latent space to generate robot structures and controlling behavior through a mixture of experts network determined by morphology.

ECO: Quantized Training without Full-Precision Master Weights

Mahdi Nikdan (Google Research), Vahab Mirrokni (Institute of Science and Technology Austria)

OptimizationComputational EfficiencyRepresentation LearningTransformerLarge Language ModelMixture of ExpertsText

🎯 What it does: Proposed a new optimizer called ECO (Error-Compensating Optimizer), which eliminates the dependency on high-precision master weights during training by injecting quantization errors into the momentum, achieving zero additional memory overhead for low-precision (FP8/INTX) training.

EcoVLA: Environment-Aware Adaptive Pruning with Interleaved Inference Orchestration for Vision-Language-Action Models

Yuting Huang (University of Science and Technology of China), Yanyong Zhang (University of Science and Technology of China)

Computational EfficiencyRobotic IntelligenceReinforcement LearningVision Language ModelVision-Language-Action ModelContrastive LearningImageTextMultimodality

🎯 What it does: Proposes a training-free, plug-and-play adaptive sparsification framework called EcoVLA for Vision-Language-Action models under real-time control.

ECSEL: Explainable Classification via Signomial Equation Learning

Adia C. Lumadjeng (University of Amsterdam), Erman Acar (University of Amsterdam)

ClassificationExplainability and InterpretabilityComputational EfficiencyTabularBenchmarkPhysics Related

🎯 What it does: Propose ECSEL, an interpretable classification method that learns signomial function expressions, capable of performing classification and directly providing readable mathematical formulas.

EDCO: Dynamic Curriculum Orchestration for Domain-specific Large Language Model Fine-tuning

Jing-Cheng Pang (Huawei Technologies Co Ltd), Xin Chen

Computational EfficiencyRepresentation LearningData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringText

🎯 What it does: Proposed the EDCO framework, which dynamically schedules training samples based on inference entropy to achieve adaptive adjustment of the model's learning state.

Edge-colored Clustering in Hypergraphs: A MaxECC Approximation

Aravind Srinivasan (University of Maryland), Jiayi Wu (University of Maryland)

OptimizationGraph

🎯 What it does: Studied hypergraphs with given edge colors, proposed the maximum edge clustering problem (MAXECC) to maximize the satisfaction of edges, and provided an approximation algorithm.

Edit-Based Refinement for Parallel Masked Diffusion Language Models

Houxing Ren (CUHK), Hongsheng Li (CUHK)

GenerationAI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningDiffusion modelText

🎯 What it does: Proposes the ME-DLM refinement framework based on editing, which utilizes minimal edit operations to perform full-sequence correction after mask diffusion generation, addressing the issue of consistency degradation in multi-word parallel generation.

Editable Proof Sketch for Automated Theorem Proving

Zikai Xiao, Shing-Tung Yau (Tsinghua University)

AI Code AssistantTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Construct an editable proof-sketch structure called EditableSketch, and implement the SketchRefine iterative ATP framework based on it.

EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation

Jingzhe Lin (Beijing Normal University), Fangwei Zhong (Beijing Normal University)

Explainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: A multi-agent-based simulation platform called EduMirror was established to simulate and analyze social dynamics in educational environments, and to provide actionable intervention and measurement tools.

EEG-Based Multimodal Learning via Hyperbolic Mixture-of-Curvature Experts

Runhe Zhou (Nanyang Technological University), Cuntai Guan (Nanyang Technological University)

ClassificationRecognitionMixture of ExpertsContrastive LearningMultimodalityBiomedical Data

🎯 What it does: This paper proposes an EEG multi-modal learning framework called EEG-MoCE based on hyperbolic mixture of experts, which is used for emotion recognition, sleep staging, and cognitive assessment by fusing EEG, speech, video, and other multi-modal data;

EEG-FM-Bench: A Comprehensive Benchmark for the Systematic Evaluation and Diagnostic Analyses of EEG Foundation Models

Wei Xiong (Tongji University), Changjun Jiang (Tongji University)

ClassificationExplainability and InterpretabilityRepresentation LearningTransformerSupervised Fine-TuningContrastive LearningTime SeriesBiomedical DataBenchmark

🎯 What it does: Proposed EEG-FM-Bench, a unified benchmark framework for systematically evaluating and diagnosing the performance of EEG foundation models.

EEmo-Logic: A Unified Dataset and Multi-Stage Framework for Comprehensive Image-Evoked Emotion Assessment

Lancheng Gao (Shanghai Jiao Tong University), Xiongkuo Min (Shanghai Jiao Tong University)

ClassificationRecognitionRecommendation SystemExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityChain-of-Thought

🎯 What it does: This paper proposes the EEmo-Logic framework, which first constructs the largest image-evoked emotion understanding dataset, EEmoDB (containing QA and fine-grained evaluation parts), and then achieves unified and fine-grained understanding of emotion QA, emotion ranking, emotion description, and fine-grained emotion assessment through a two-stage training approach (LoRA SFT + GRPO) on multimodal large language models.