arXivSub Start free trial

ICML 2026 Papers with Code β€” Page 3

International Conference on Machine Learning Β· 1032 papers

Convex Optimization for Alignment and Preference Learning on a Single GPU

Miria Feng (Stanford University), Mert Pilanci (Stanford University)

CodeOptimizationComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelAuto EncoderContrastive LearningText

🎯 What it does: This paper proposes COALA, a single GPU language model preference alignment method based on convex optimization.

Coordinated Disentanglement with Iterative Mode Discovery Under Hidden Correlations

Rong Hu (Zhejiang University), Ling Chen (Zhejiang University)

CodeDomain AdaptationRepresentation LearningGenerative Adversarial NetworkContrastive LearningImageMultimodalityTabularTime Series

🎯 What it does: Propose the CoDID framework, which decomposes hidden-related attributes by utilizing iterative pattern discovery and conditional mutual information minimization.

CoPE: A Framework for Optimizing Coordination between Planning and Execution in LLM-based Agents

Huanxi Liu (National University of Defense Technology), Huaimin Wang (National University of Defense Technology)

CodeAutonomous DrivingOptimizationRobotic IntelligenceReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AIMixture of ExpertsTextSequential

🎯 What it does: Propose the CoPE framework, which utilizes self-improving MCTS to collect high-quality plan-execution data, evaluates the executability of planning and the adherence of execution, and uses it as sample weights for LLM fine-tuning.

CoPE: Continual Probe-guided Expansion for Large Vision-Language Models

Ziqin Wang (Beihang University), Si Liu (Beihang University)

CodeFederated LearningComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerMixture of ExpertsVision Language ModelImageTextMultimodalityBenchmark

🎯 What it does: Proposed a continuous learning framework called CoPE, aimed at addressing the issues of parameter growth and knowledge forgetting in large-scale vision-language models when handling sequential tasks.

CORAL: Uncertainty-Aware Regulation of Exposure Concentration in Recommender Systems

Nitin Bisht (University of Technology), Guandong Xu (Education University of Hong Kong)

CodeRecommendation SystemOptimizationExplainability and InterpretabilityTransformerReinforcement LearningMixture of ExpertsContrastive LearningTabularSequential

🎯 What it does: This paper proposes the CORAL framework, which provides interpretable, model-free risk-aware regulation of exposure concentration in recommendation systems by constructing a self-exciting intensity model and an upper confidence bound.

CORE: Conflict-Oriented Reasoning for General Multimodal Manipulation Detection

Jinjie Shen (Hefei University of Technology), Zhun Zhong (Hefei University of Technology)

CodeAnomaly DetectionExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringContrastive LearningImageTextMultimodality

🎯 What it does: Construct a Conflict Attribution Corpus (CAC) and train with conflict awareness to enable multimodal large language models to capture and reason about conflicts, thereby achieving rapid detection of multimodal fake news.

Corrected Samplers for Discrete Flow Models

Zhengyan Wan (East China Normal University), Guang Cheng (University of California Los Angeles)

CodeGenerationData SynthesisDiffusion modelScore-based ModelFlow-based ModelImageTextStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: This paper proposes a time-corrected sampler and a position-corrected sampler for discrete flow models to reduce discretization errors;

Correcting in Hindsight: Editing Past Key-Value States for Robust LLM Reasoning

Mengfei Zhang (Zhejiang University), Leijing Zhou (Zhejiang University)

CodeExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose HEdit, which corrects early errors in autoregressive inference by real-time detection of trigger points and backtracking to edit key value caches.

CorrectionPlanner: Self-Correction Planner with Reinforcement Learning in Autonomous Driving

Yihong Guo (Johns Hopkins University), Xianming Liu (XPENG Motors)

CodeAutonomous DrivingTransformerReinforcement LearningWorld ModelPoint Cloud

🎯 What it does: Proposed a planner called CorrectionPlanner that performs self-correction in the motion token space, capable of proactively identifying and correcting unsafe actions before execution;

CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features

Seonglae Cho (Holistic Ai), Adriano Koshiyama (Holistic Ai)

CodeOptimizationExplainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelPrompt EngineeringAuto EncoderTextBenchmark

🎯 What it does: Propose a model optimization method called CorrSteer, which does not require a contrastive dataset or backpropagation, and instead utilizes features from sparse autoencoders (SAE) activated during generation, and automatically selects features through a two-stage process of correlation plus intervention.

CountsDiff: A diffusion model on the natural numbers for generation and imputation of count-based data

Renzo Soatto, Maria Skoularidou (Massachusetts Institute of Technology)

CodeGenerationData SynthesisTransformerDiffusion modelScore-based ModelImageTabularBiomedical Data

🎯 What it does: Proposes CountsDiff, a diffusion model that operates on the set of natural numbers, for generating count data and imputing biological count data such as single-cell RNA-seq.

Coupled Variational Reinforcement Learning for Language Model General Reasoning

Xueru Wen (University of Chinese Academy of Sciences), Debing Zhang (Xiaohongshu Inc)

CodeComputational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsText

🎯 What it does: This study proposes the Coupled Variational Reinforcement Learning (CoVRL) method, which combines prior (question-only) and posterior (answer-guided) generation modes, training the reasoning ability of large language models through a mixture distribution and mixed sampling.

Courtroom Analogy: New Perspective on Uncertainty-Aware Classification

Taeseong Yoon (Korea Advanced Institute of Science and Technology), Heeyoung Kim (Korea Advanced Institute of Science and Technology)

CodeClassificationExplainability and InterpretabilityConvolutional Neural NetworkMixture of ExpertsImage

🎯 What it does: Proposed a single-channel uncertainty estimation framework based on analogies from courtroom debates, and implemented the Mixture of Dirichlet Experts (MoDEX) model.

Creat3r: Confidence Reaggregation for Exploration-aware Active 3D Reconstruction

Chih-Jung Tsai (National Tsing Hua University), Tyng-Luh Liu (Academia Sinica)

CodeOptimizationComputational EfficiencyNeural Radiance FieldGaussian SplattingSimultaneous Localization and MappingOptical FlowImagePoint Cloud

🎯 What it does: Propose Creat3r, an iterative next best view selection framework based on 3D Gaussian Splatting, which constructs a lightweight geometric proxy and guides view selection through confidence and exploration maps.

Credit-assigned Policy Gradient for Early Stage Retrieval in Two-stage Ranking

Haruka Kiyohara (Cornell University), Udi Weinsberg (Meta)

CodeRecommendation SystemReinforcement LearningMixture of ExpertsVideoTabular

🎯 What it does: Studied the policy gradient training of early retrievers in two-stage retrieval, proposing a credit assignment policy gradient (CA-PG) method to reduce variance and address the credit assignment problem.

CriticalKV: Optimizing KV Cache Eviction from an Output Perturbation Perspective

Yuan Feng (University of Science and Technology of China), Xike Xie (University of Science and Technology of China)

CodeCompressionOptimizationComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: To address the compression problem of KV cache in large language models, this paper proposes a key cache entry selection algorithm based on worst-case perturbation constraints from the perspective of output perturbation.

Cross-Modal Semantic Decoupling and Transfer for Text-to-Visible-Infrared Person Re-Identification

Ziang Zhang (Wuhan University), Mang Ye (Wuhan University)

CodeRecognitionRetrievalDomain AdaptationTransformerVision Language ModelDiffusion modelScore-based ModelFlow-based ModelRectified FlowAuto EncoderGenerative Adversarial NetworkContrastive LearningImageTextMultimodality

🎯 What it does: This paper proposes a cross-modal semantic disentanglement and transfer framework, CSDT, for cross-modal person re-identification (TVI-ReID) tasks involving text retrieval of visible and infrared images, addressing the issue of insufficient visible light information in nighttime surveillance scenarios.

CSD: Content-aware Speculative Decoding for Efficient Image Generation

Mingcheng Wang (East China Normal University), Shaohui Lin (East China Normal University)

CodeGenerationTransformerDiffusion modelImage

🎯 What it does: Proposes a content-aware speculative decoding (CSD) to accelerate autoregressive image generation.

CSPO: Constraint-Sensitive Policy Optimization for Safe Reinforcement Learning

Ayoub Belouadah (University of Luxembourg), YVES LE TRAON

CodeOptimizationReinforcement Learning

🎯 What it does: Propose a first-order primal-dual algorithm called CSPO, which uses a constraint gradient norm adaptive correction strategy to update, enabling rapid recovery of safety while maintaining KKT optimality.

cuRegOT: A GPU-Accelerated Solver for Entropic-Regularized Optimal Transport

Yixuan Qiu (Shanghai University of Finance and Economics)

CodeOptimizationImagePoint CloudTabular

🎯 What it does: Proposes cuRegOT, a GPU-accelerated entropy-regularized optimal transport solver based on the SPLR approximation of second-order methods.

CURVE: Learning Causality-Inspired Invariant Representations for Robust Scene Understanding via Uncertainty-Guided Regularization

Yue Liang (Tongji University), Hong Chen (Tongji University)

CodeDomain AdaptationAutonomous DrivingExplainability and InterpretabilityComputational EfficiencyRepresentation LearningGraph Neural NetworkContrastive LearningImageGraph

🎯 What it does: Propose the CURVE framework, which utilizes variational uncertainty modeling and prototype-based soft backdoor adjustment to achieve interpretable causal sparse structures in scene graphs, thereby enhancing robustness.

CurvZO: Adaptive Curvature-Guided Sparse Zeroth-Order Optimization for Efficient LLM Fine-Tuning

Shuo Wang (Southern University of Science and Technology), Ming Tang (Southern University of Science and Technology)

CodeOptimizationComputational EfficiencyTransformerLarge Language ModelText

🎯 What it does: Propose a method to efficiently fine-tune large language models by using online curvature signals to guide sparse zeroth-order optimization (CurvZO).

CyberJurors: A Multi-Agent Simulation Task for E-Commerce Disputes Verdict

Yanhui Sun (University of Science and Technology of China), Yongdong Zhang (University of Science and Technology of China)

CodeRecommendation SystemAnomaly DetectionExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelAgentic AIPrompt EngineeringImageVideoTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed a task called 'E-commerce Dispute Verdicts (EDV)' aimed at e-commerce transaction disputes, and implemented intelligent adjudication through multi-agent simulation.

DAG-MoE: From Simple Mixture to Structural Aggregation in Mixture-of-Experts

Jiarui Feng (Meta MRS), Yixin Chen (Washington University in St. Louis)

CodeComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelMixture of ExpertsTextChain-of-Thought

🎯 What it does: This study investigates replacing traditional weighted sum aggregation with structured aggregation (DAG) in Mixture-of-Experts (MoE) models, proposing a learnable DAG-MoE framework that enhances expressive power and enables multi-step reasoning without altering the experts or router.

DAISI: Data Assimilation with Inverse Sampling using Stochastic Interpolants

Martin Andrae (LinkΓΆping University), Fredrik Lindsten (LinkΓΆping University)

CodeOptimizationComputational EfficiencyData-Centric LearningDiffusion modelScore-based ModelFlow-based ModelImageVideoTime SeriesPhysics RelatedStochastic Differential Equation

🎯 What it does: Proposes DAISI, an expandable filtering framework based on flow generative models, which integrates predictions into the latent variables of the generative model through an inverse sampling step, and achieves observation fusion via guided conditional sampling.

DASH: Faster Shampoo via Batched Block Preconditioning and Efficient Inverse-Root Solvers

Ionut-Vlad Modoranu (Institute of Science and Technology Austria), Dan Alistarh (Institute of Science and Technology Austria)

CodeOptimizationTransformerLarge Language ModelText

🎯 What it does: Propose DASH, an improved distributed Shampoo optimizer, achieving 3D block stacking and more efficient inverse matrix root computation.

Data Augmentation of Contrastive Learning is Estimating Positive-incentive Noise

Hongyuan Zhang (University of Hong Kong), Xuelong Li (China Telecom)

CodeInformation TheoryData SynthesisRepresentation LearningConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningImageTextMultimodalityTabularTime Series

🎯 What it does: Designed and implemented PiNDA, a contrastive learning framework that automatically generates augmented noise by learning Positive-incentive Noise (Ο€-Noise).

Data Provenance Auditing of Fine-Tuned Large Language Models with a Text-Preserving Technique

Yanming Li (Inria), Seifeddine Ghozzi (Institut Polytechnique de Paris)

CodeSafty and PrivacyExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringText

🎯 What it does: A technique is proposed to construct 'prompt/response' pairs using invisible Unicode characters for watermarking, which can detect whether training data has been used through black-box interaction after fine-tuning large language models, and control the false positive rate through ranking tests.

Data-driven Mixed Integer Optimization through Probabilistic Multi-variable Branching

Yanguang Chen (Shanghai University of Finance and Economics), Yinyu Ye (Shanghai Jiao Tong University)

CodeOptimizationGraph Neural NetworkContrastive LearningTabularBenchmark

🎯 What it does: This paper proposes a hybrid integer programming acceleration method based on probabilistic multivariate branching (PMVB).

De-Linearizing Agent Traces: Bayesian Inference of Latent Partial Orders for Efficient Execution

Dongqing li, Quyu Kong

CodeOptimizationComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringTextGraph

🎯 What it does: De-linearize the linear execution logs generated by LLM agents, infer their underlying partial order graph (i.e., concurrency and precedence dependencies), and compile this graph into an efficient frontier-based executor, reducing redundant reasoning and token consumption.

Debiased Model-based Representations for Sample-efficient Continuous Control

Jiafei Lyu (Tencent Hunyuan), Deheng Ye (Tencent Hunyuan)

CodeConvolutional Neural NetworkRecurrent Neural NetworkTransformerReinforcement LearningContrastive LearningTabularTime Series

🎯 What it does: In continuous control reinforcement learning, sample efficiency is improved by learning model-based representations and combining them with experience replay.

DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile Adapter

Xukun Li (XYZ Embodied AI), Zhenguo Sun (XYZ Embodied AI)

CodeRobotic IntelligenceTransformerReinforcement LearningDiffusion modelImageVideoMultimodalityTime Series

🎯 What it does: Propose the DECO framework, which utilizes a decoupled multi-modal diffusion Transformer to achieve dexterous bimanual manipulation, along with a plug-in tactile adapter.

Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding

Hee Suk Yoon (Korea Advanced Institute of Science and Technology), Chang D. Yoo (Korea Advanced Institute of Science and Technology)

CodeKnowledge DistillationRepresentation LearningTransformerSupervised Fine-TuningReinforcement LearningVision Language ModelContrastive LearningImageTextMultimodalityBenchmark

🎯 What it does: This paper proposes a decomposed on-policy distillation method for visual language reasoning and designs a Visual Gradient Guidance (VGS) to enhance the model's visual perception ability.

Decomposing Query-Key Feature Interactions Using Contrastive Covariances

Andrew Lee (Harvard University), Martin Wattenberg (Harvard University)

CodeExplainability and InterpretabilityRepresentation LearningTransformerContrastive LearningText

🎯 What it does: This paper proposes a contrastive covariance decomposition of the query-key (QK) space in attention heads, decomposing it into interpretable low-rank subspaces, and verifies the effectiveness of this method in toy models and large language models.

Decomposition-Based Modular Conformal Prediction for Two-Stage Modeling

William Zhang (Massachusetts Institute of Technology), Georgia Perakis (Massachusetts Institute of Technology)

CodeAnomaly DetectionFederated LearningExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTabularTime SeriesSequentialFinance Related

🎯 What it does: Proposes a synthetic prediction framework for two-stage modular models, achieving uncertainty attribution at each pipeline stage through stage-wise decomposition of residuals, and determining scaling parameters via FWER-controlled risk calibration methods;

Decouple and Cache: KV Cache Construction for Streaming Video Understanding

Zhanzhong Pang (National University of Singapore), Angela Yao (National University of Singapore)

CodeComputational EfficiencyRepresentation LearningTransformerLarge Language ModelVision Language ModelVideoTextMultimodalityRetrieval-Augmented Generation

🎯 What it does: Designed and implemented a training-free KV cache construction framework called DSCache, which is used in resource-constrained streaming video understanding tasks to real-time construct and maintain accumulated cache and instant cache, thereby achieving efficient inference for infinitely long video streams.

Decoupling The "What" and "Where" With Polar Coordinate Positional Embedding

Anand Gopalakrishnan (Swiss AI Lab IDSIA USI SUPSI), Michael Curtis Mozer (University of Colorado)

CodeComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBiomedical DataAudio

🎯 What it does: Proposed a new relative position encoding method called PoPE, aimed at separating the mixed information of 'content-what' and 'position-where' in Transformers, addressing the difficulties encountered by RoPE during independent matching;

Deep Ensemble Clustering for Visual Representation Learning

Yuwei Wang (Nanjing University of Science and Technology), Wenguan Wang (Zhejiang University)

CodeClassificationObject DetectionSegmentationRepresentation LearningTransformerDiffusion modelAuto EncoderContrastive LearningImage

🎯 What it does: Propose ENFORMER, a visual backbone that integrates clustering into the visual feature extraction process, leveraging multiple differentiable clustering methods to learn richer and more interpretable visual representations.

Deep sequence models tend to memorize geometrically; it is unclear why

Shahriar Noroozizadeh (Carnegie Mellon University), Sanjiv Kumar (Google Research)

CodeExplainability and InterpretabilityComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerSupervised Fine-TuningContrastive LearningGraph

🎯 What it does: Investigate the implicit reasoning capabilities of deep sequence models (such as Transformer and Mamba) in graph structure memory tasks, discovering that models form global geometric memory rather than traditional associative memory.

DeepAnalyze: Agentic Large Language Models for Autonomous Data Science

Shaolei Zhang (Renmin University of China), Xiaoyong Du (Renmin University of China)

CodeOptimizationData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringMixture of ExpertsTextTabularBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposes DeepAnalyze, an agent-based large language model capable of autonomously completing the full workflow of data science from raw data to research reports.

Degradation-Aware Metric Prompting for Hyperspectral Image Restoration

Binfeng Wang (Beijing Institute of Technology), Jing Zhang (Wuhan University)

CodeRestorationExplainability and InterpretabilityComputational EfficiencyConvolutional Neural NetworkTransformerPrompt EngineeringMixture of ExpertsContrastive LearningImage

🎯 What it does: A unified hyperspectral image restoration framework named DAMP is proposed, which utilizes interpretable spatial-spectral metrics as degradation hints and dynamically routes experts through degradation-adaptive Mixture-of-Experts to achieve unified restoration for various degradations.

Delving into Muon and Beyond: Deep Analysis and Extensions

Xianbiao Qi (Intellifusion Inc), Rong Xiao (Intellifusion Inc)

CodeOptimizationTransformerLarge Language ModelTextPhysics Related

🎯 What it does: This paper studies the Muon optimizer and its variants through a spectral transformation framework, systematically evaluating their stability and performance on matrix parameters;

Demystifying the Optimal Fair Classifier in Multi-Class Classification

Li Zhang (Zhejiang University), Chaochao Chen (Zhejiang University)

CodeClassificationOptimizationFederated LearningSafty and PrivacyExplainability and InterpretabilityImageTabular

🎯 What it does: Propose the OptFair framework, theoretically derive the Bayes optimal discriminator for multi-class fair classification, and provide an implementable attribute-blind algorithm to achieve the Pareto frontier between accuracy and fairness.

Denoising without Diffusion: Fixed-Noise Denoiser Anomaly Detection in Tabular Data

Manuel Hirth (Daimler Truck AG), Enkelejda Kasneci (Technical University of Munich)

CodeAnomaly DetectionTransformerDiffusion modelScore-based ModelContrastive LearningTabularBenchmarkFinance Related

🎯 What it does: Proposes a single-step fixed-noise denoising anomaly detection method for tabular data called DenoiserAD, which utilizes a self-supervised denoiser and calculates stability scores through multiple noise perturbations during inference.

Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration

Zhicheng Yang (Hong Kong University of Science and Technology (Guangzhou)), Jing Tang (Hong Kong University of Science and Technology (Guangzhou))

CodeOptimizationComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsTextBenchmark

🎯 What it does: This paper investigates the interaction between depth (problem difficulty) and breadth (number of training instances) in the Reinforcement Learning with Verifiable Rewards (RLVR) framework, and proposes Difficulty-Adaptive Episode Sampling (DARS) as well as the DARS-Breadth method combining DARS with large-batch training, to enhance the reasoning performance of large language models (LLMs).

Di-BiLPS: Denoising induced Bidirectional Latent-PDE-Solver under Sparse Observations

Zhonghao Li (Harbin Institute of Technology), Zhang Qian

CodeOptimizationTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningGraphTabularBenchmarkPhysics Related

🎯 What it does: Proposed a unified neural network framework called Di-BiLPS for simultaneously solving forward and inverse problems of PDEs under extremely sparse observations.

Differential Smoothing Mitigates Sharpening and Improves LLM Reasoning

Jingchu Gai (Carnegie Mellon University), Aditi Raghunathan (Carnegie Mellon University)

CodeOptimizationTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringText

🎯 What it does: Studied the diversity collapse problem caused by RL fine-tuning and proposed the Differential Smoothing (DS-GRPO) method to simultaneously improve correctness and diversity.

Differential syntactic and semantic encoding in LLMs

Santiago Acevedo (Scuola Internazionale Superiore di Studi Avanzati), Marco Baroni (Catalan Institute of Research and Advanced Studies)

CodeExplainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelContrastive LearningTextMultimodality

🎯 What it does: This paper investigates the encoding and separation of syntactic and semantic information in sentence representations of internal layers of large language models by constructing and ablating the 'center points' of sentence syntax and semantics in linear projections.

Differentially Private Preference Data Synthesis for Large Language Model Alignment

Fengyu Gao (University of Virginia), Jing Yang (University of Virginia)

CodeData SynthesisSafty and PrivacyReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringDiffusion modelText

🎯 What it does: Propose the DPPrefSyn algorithm, which synthesizes preference data using differential privacy to support preference alignment in LLMs;

Differentially Private Submodular Maximization with a Knapsack Constraint

Ron Zadicario (Tel Aviv University), Tova Milo (Tel Aviv University)

CodeOptimizationSafty and PrivacyReinforcement LearningPrompt EngineeringMixture of ExpertsTabularChain-of-Thought

🎯 What it does: Proposes an algorithm to achieve differential privacy in submodular function maximization with a knapsack constraint, separately for monotonic and non-monotonic objectives.

Diffract: Spectral View of LLM Domain Adaptation

Nikita Borodin (Risk Ai Research Lab), Dmitry Vinichenko (Risk Ai Research Lab)

CodeDomain AdaptationComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsTextTabularReview/Survey PaperBenchmarkRetrieval-Augmented Generation

🎯 What it does: This paper studies the adaptation mechanism of continuous pre-training (CPT) for large language models (LLMs), and reveals the weight update characteristics during the CPT phase through spectral analysis of the weight matrix using singular value decomposition (SVD).

Diffusing to Coordinate: Efficient Online Multi-Agent Diffusion Policies

Zhuoran Li (Tsinghua University), Longbo Huang (Tsinghua University)

CodeReinforcement LearningDiffusion modelTabularTime SeriesSequential

🎯 What it does: Propose an online multi-agent diffusion strategy, OMAD, achieving high expressiveness in online reinforcement learning.

Diffusion Differentiable Resampling

Jennifer R. Andersson, Zheng Zhao (LinkΓΆping University)

CodeGenerationData SynthesisOptimizationComputational EfficiencyReinforcement Learning from Human FeedbackDiffusion modelScore-based ModelAuto EncoderImageTime SeriesSequentialStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: This paper proposes a differentiable resampling method based on untrained diffusion models, which can achieve particle resampling within the sequential Monte Carlo (SMC) framework while maintaining differentiability, thereby supporting gradient optimization.

Direct Flow Q-Learning

Shicheng Cao (Nankai University), Shengbo Eben Li

CodeTransformerReinforcement LearningDiffusion modelFlow-based ModelTabularTime SeriesSequentialBenchmark

🎯 What it does: Propose an offline reinforcement learning framework called Direct Flow Q-Learning (DFQL), which directly injects the terminal Q gradient into the flow matching velocity field, achieving reinforcement learning optimization for multi-step generative models without requiring BPTT or policy distillation.

Discontinuous Galerkin Neural Operator for Pathology Defocus Deblurring

Shaoqing Duan (East China Normal University), Yan Wang (East China Normal University)

CodeRestorationTransformerDiffusion modelContrastive LearningImageBiomedical DataBenchmark

🎯 What it does: Proposed a discrete Galerkin neural operator (DGNO) for spatially varying, locally discontinuous defocus deblurring tasks in pathological microscopy images;

Discounted Beta–Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards

Haechan Kim (KAIST), Eunho Yang (KAIST)

CodeOptimizationComputational EfficiencyTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmark

🎯 What it does: This paper addresses the reward estimation problem in the Verifiable Reward (RLVR) framework of reinforcement learning, proposing a Discounted Beta-Bernoulli (DBB) reward estimation method that uses historical reward information to reduce variance and prevent variance collapse, thereby improving sample efficiency.

Discovering Ordinary Differential Equations with LLM-Based Qualitative and Quantitative Evaluation

Sum Kyun Song (Chung Ang University), Jae Yong Lee (Chung Ang University)

CodeOptimizationExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelAgentic AIPrompt EngineeringTabularTime SeriesBenchmarkStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Propose a multi-agent framework DoLQ, which uses LLM to qualitatively and quantitatively evaluate candidate ODEs, ultimately automatically discovering interpretable ordinary differential equations from observational data.

DisPOSE: Projected Polystochastic Diffusion for Self-Supervised Multi-View 3D Human Pose Estimation

Tony Danjun Wang (Technical University of Munich), Lennart Bastian (Technical University of Munich)

CodePose EstimationGraph Neural NetworkDiffusion modelContrastive LearningImagePoint CloudBenchmark

🎯 What it does: Propose a self-supervised multi-view 3D human pose estimation framework called DisPOSE, which transforms the multi-view person identity association problem into a diffusion process of multivariate Polystochastic tensors, enabling training without 3D annotations.

Dissecting Causal Mechanism Shifts via FANS: Function And Noise Separation

Gyeongdeok Seo (Yonsei University), Kyungwoo Song (Yonsei University)

CodeAnomaly DetectionExplainability and InterpretabilityGraph Neural NetworkFlow-based ModelContrastive LearningImageTabularBiomedical Data

🎯 What it does: Propose the FANS framework, which detects and separates functional changes from noise changes in causal mechanism transitions by utilizing causal regularized flow.

Distillation Models are Good Samplers for Diffusion Reinforcement Learning

Zunxu Liu (University of Science and Technology of China), Tao Mei (HiDream.ai Inc)

CodeGenerationData SynthesisComputational EfficiencyKnowledge DistillationTransformerReinforcement LearningDiffusion modelScore-based ModelContrastive LearningImageVideoTextMultimodality

🎯 What it does: Propose the DMSampler framework, which replaces the time-consuming multi-step sampling with a co-advancing distillation model, significantly accelerating the reinforcement learning process of diffusion models.

Distilling Linearized Behavior into Non-linear Fine-Tuning for Effective Task Arithmetic

Thomas Sommariva (University of Modena and Reggio Emilia), Angelo Porrello (University of Modena and Reggio Emilia)

CodeComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerSupervised Fine-TuningDiffusion modelAuto EncoderContrastive LearningImageText

🎯 What it does: A model that can maintain nonlinear expressiveness during inference while possessing the task vector separability and editability of linearized models is constructed by incorporating distillation of linearized models into nonlinear fine-tuning.

DistMatch: Adaptive Binning via Distribution Matching for Robust Sequential Conformal Prediction

Enver Menadjiev (Korea Advanced Institute of Science and Technology), Jaesik Choi (Korea Advanced Institute of Science and Technology)

CodeAnomaly DetectionFederated LearningExplainability and InterpretabilityComputational EfficiencyScore-based ModelContrastive LearningTabularTime SeriesSequentialFinance Related

🎯 What it does: This paper proposes DistMatch, an adaptive binning method based on the recursive partitioning of residuals using the Kolmogorov-Smirnov statistic, for achieving verifiable sequential conformal prediction in time series with distribution drift.

Distribution Transformers: Fast Approximate Bayesian Inference With On-The-Fly Prior Adaptation

George Whittle (University of Oxford), Michael A Osborne

CodeOptimizationComputational EfficiencyRepresentation LearningTransformerMixture of ExpertsScore-based ModelContrastive LearningGaussian SplattingTabularTime SeriesSequentialPhysics Related

🎯 What it does: Propose a Distribution Transformer based on Transformer, which can map any prior (represented as a Gaussian Mixture Model, GMM) to a posterior (also a GMM) in a single forward pass, thereby enabling approximate Bayesian inference and supporting the modification of the prior at any time.

DITING: A Weak Degradation Listener for Battery Lifetime Early Prediction

Hao Miao (Taiyuan University of Technology), Li Wang (Taiyuan University of Technology)

CodeAnomaly DetectionComputational EfficiencyRepresentation LearningRecurrent Neural NetworkTransformerDiffusion modelScore-based ModelContrastive LearningTabularTime Series

🎯 What it does: Developed a weak degradation listener called DITING based on the binaural effect, used to extract weak degradation signals from early cycle battery data and predict lifespan in noisy environments.

DIVA: Harnessing the Representation Divergence in Unified Multimodal Models for Mutual Reinforcement

Renjie Lu (Ping An Technology (Shenzhen) Co., Ltd.), Shangfei Wang (University of Science and Technology of China)

CodeGenerationData SynthesisRepresentation LearningTransformerSupervised Fine-TuningVision Language ModelDiffusion modelContrastive LearningImageTextMultimodality

🎯 What it does: This paper proposes DIVA, a post-training framework designed to eliminate mutual interference caused by the biases of the understanding branch and the generation branch in unified multi-modal models (UMMs), and to achieve complementary improvement by decomposing shared and unique information.

Diversity-Preserved Distribution Matching Distillation for Fast Visual Synthesis

Tianhe Wu (City University of Hong Kong), Kede Ma (City University of Hong Kong)

CodeGenerationData SynthesisKnowledge DistillationTransformerDiffusion modelScore-based ModelFlow-based ModelGenerative Adversarial NetworkImageText

🎯 What it does: Proposed a distribution matching distillation method called DP-DMD based on role separation, specifically designed for simultaneous preservation of sample diversity and visual quality in few-step image generation.

Divide and Conquer: Reliable Multi-View Evidential Learning for Deepfake Detection

Xiaolu Kang (Wuhan University), Qian Wang (Wuhan University)

CodeAnomaly DetectionTransformerSupervised Fine-TuningVision Language ModelAuto EncoderContrastive LearningImageVideo

🎯 What it does: Propose a depth-forgery detection framework called DiCoME based on the 'divide and conquer' strategy, which achieves reliable detection by separating the semantic perspective from the structural perspective and fusing multi-perspective evidence.

Divide and Contrast: Learning Robust Temporal Features without Augmentation

Abdul-Kazeem Shamba (Norwegian University of Science and Technology), Gavin Taylor (United States Naval Academy)

CodeClassificationRetrievalAnomaly DetectionRepresentation LearningRecurrent Neural NetworkTransformerAuto EncoderContrastive LearningTabularTime Series

🎯 What it does: Proposes Di-COT, a temporal representation learning framework that does not rely on data augmentation or multiple encoder passes, learning robust temporal features by randomly partitioning overlapping sub-blocks within each window and contrasting neighboring sub-blocks.

DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders

Xu Wang (University of Hong Kong), Difan Zou (University of Hong Kong)

CodeExplainability and InterpretabilityTransformerLarge Language ModelDiffusion modelAuto EncoderText

🎯 What it does: Propose DLM-Scope, a mechanism interpretation framework based on sparse autoencoders (SAE), used to parse the internal representations of diffusion language models (DLM) and achieve controllable interventions.

Do Neural Operators Forget Geometry? The Forgetting Hypothesis in Deep Operator Learning

Yanming Xia (Tsinghua University), Angelica I Aviles-Rivero

CodeOptimizationExplainability and InterpretabilityComputational EfficiencyTransformerMeshGraphBenchmarkPhysics Related

🎯 What it does: Propose and verify the 'geometric forgetting hypothesis' that neural operators tend to forget geometric information in deep layers, and alleviate this issue through geometric memory injection.

Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs

Xin Gao (University of California San Diego), Taylor Berg-Kirkpatrick (University of California San Diego)

CodeGenerationData SynthesisExplainability and InterpretabilityTransformerPrompt EngineeringVision Language ModelDiffusion modelImageTextMultimodalityBenchmarkChain-of-Thought

🎯 What it does: Propose the UNIKE benchmark to evaluate the cross-modal knowledge editing effectiveness of unified multi-modal models, and systematically test whether text editing can be transferred to image generation.

Do We Need Adam? Surprisingly Strong and Sparse Reinforcement Learning with SGD in LLMs

Sagnik Mukherjee (University of Illinois), Hao Peng (University of Illinois)

CodeOptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningTextSequential

🎯 What it does: Studied the performance comparison between SGD and AdamW in large language model reinforcement learning (RLVR), demonstrating that SGD can match or even surpass AdamW, while significantly reducing memory consumption and parameter update sparsity.

Domain-Shift-Aware Conformal Prediction for Large Language Models

Zhexiao Lin (University of California), Michael Berger

CodeDomain AdaptationExplainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelContrastive LearningTextBenchmark

🎯 What it does: Proposes a domain-aware adaptive conformal prediction framework (DSCP) for large language models, aiming to achieve reliable uncertainty quantification under domain shift.

Don't Forget Why You Started: Tackling Dual Forgetting in Vision-Language Continual Learning

Borui Kang (Nanjing University), Yang Gao (Nanjing University)

CodeClassificationComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerSupervised Fine-TuningMixture of ExpertsVision Language ModelContrastive LearningImageTextMultimodality

🎯 What it does: Propose the DFA-CIL framework and the SCR metric to strictly evaluate two types of forgetting in VLM during continual learning (IKF and PKF), and based on this, design the DFA-MoE parameter-efficient fine-tuning method, which balances knowledge retention and new task adaptation.

DOUBT: Decoupled Object-level Understanding and Bridging via vMF-based Trustworthiness for Hallucination Detection in MLLMs

Kaiqi Chen (Sichuan University), Peng Hu (Sichuan University)

CodeAnomaly DetectionExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmark

🎯 What it does: Propose a black-box hallucination detection framework called DOUBT, which first splits visual recognition and relation reasoning by object-level understanding and bridging (OUB), and then uses vMF to evaluate the consistency of multiple answers based on the directional credibility metric of vectors, thus determining hallucinations.

DSB: Dynamic Sliding Block Scheduling for Diffusion LLMs

Lizhuo Luo (Nanyang Technological University), Tianwei Zhang (Nanyang Technological University)

CodeComputational EfficiencyAI Code AssistantTransformerLarge Language ModelDiffusion modelText

🎯 What it does: Proposed an untrained dynamic sliding block scheduling (DSB) and the corresponding KV cache scheme (DSB Cache) to improve the semi-autoregressive inference of diffusion-based large language models (dLLMs).

DSENet: A Novel Dual-Stream Enhancement Network for Multi-Scale Non-Stationary Time Series Forecasting

Yuhan Wang (Shanghai Jiao Tong University), Jinhong Guo (Shanghai Jiao Tong University)

CodeExplainability and InterpretabilityComputational EfficiencyTransformerMixture of ExpertsAuto EncoderContrastive LearningTime SeriesBiomedical DataElectronic Health Records

🎯 What it does: Proposes the DSENet dual-stream enhanced network, which models the problem of local mutations in long sequences for blood glucose prediction.

Dual Mechanisms of Value Expression: Intrinsic vs. Prompted Values in Large Language Models

Jongwook Han (Seoul National University), Yohan Jo (Seoul National University)

CodeExplainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextMultimodalityBenchmark

🎯 What it does: This paper investigates two mechanisms by which large language models express values: intrinsic expression and prompted expression, and decomposes these mechanisms at the activation vector and neuron levels.

Dual Quaternion SE(3) Synchronization with Recovery Guarantees

Jianing Zhao (Chinese University of Hong Kong), Anthony Man-Cho So (Chinese University of Hong Kong)

CodePose EstimationOptimizationContrastive LearningGaussian SplattingSimultaneous Localization and MappingOptical FlowPoint CloudMesh

🎯 What it does: A two-stage algorithm is proposed by mapping the SE(3) synchronization problem to the unit dual quaternion space: first, spectral initial values are obtained using power iteration, and then refined using the dual quaternion generalized power method (DQGPM) with projection at each step, along with an upper bound on the finite iteration error.

DualOptim+: Bridging Shared and Decoupled Optimizer States for Better Machine Unlearning in Large Language Models

Xuyang Zhong (City University of Hong Kong), Chen Liu (City University of Hong Kong)

CodeOptimizationSafty and PrivacyTransformerLarge Language ModelSupervised Fine-TuningText

🎯 What it does: Proposes the DualOptim+ optimization framework to improve the unlearning performance of large models.

DyCon: Dynamic Reasoning Control via Evolving Difficulty Modeling

Tengyao Tu (Harbin Institute of Technology), Min Zhang (Harbin Institute of Technology)

CodeExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextBenchmarkChain-of-Thought

🎯 What it does: Through the untrained DyCon framework, it leverages the hidden layer representations of large-scale inference models to estimate the task difficulty in real-time during the inference process, thereby dynamically adjusting the logical scores of reflection keywords, reducing redundant reasoning steps, and improving inference efficiency.

DyLLM: Efficient Diffusion LLM Inference via Saliency-based Token Selection and Partial Attention

Younjoo Lee (Seoul National University), Jung Ho Ahn (Seoul National University)

CodeGenerationComputational EfficiencyTransformerLarge Language ModelDiffusion modelText

🎯 What it does: Propose DyLLM, an untrained inference framework that achieves efficient inference for Diffusion LLM by identifying and only computing the 'significant words' that change significantly during diffusion steps.

Dynamic Stratified Contrastive Learning with Upstream Augmentation for MILP Branching

Tongkai Lu (Beihang University), Chongyang Tao (Beihang University)

CodeOptimizationGraph Neural NetworkReinforcement LearningContrastive LearningGraphTabularBenchmark

🎯 What it does: This paper proposes SC-MILP, a dynamic hierarchical contrastive learning framework for branching in mixed integer linear programming (MILP).

Dynamic Thinking-Token Selection for Efficient Reasoning in Large Reasoning Models

Zhenyuan Guo (Zhejiang University), Wenzhi CHEN

CodeComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelTextChain-of-Thought

🎯 What it does: This paper proposes a dynamic thinking-token selection method (DYNTS), which compresses the memory and computation of large inference models by predicting the importance of each token during the thinking phase to the final answer and retaining only key KV caches.

Dynamic TMoE: A Drift-Aware Dynamic Mixture of Experts Framework for Non-Stationary Time Series Forecasting

Jiawen Zhu (Zhejiang University), Yingcai Wu (Zhejiang University)

CodeAnomaly DetectionRecurrent Neural NetworkTransformerMixture of ExpertsDiffusion modelScore-based ModelContrastive LearningGaussian SplattingTabularTime SeriesFinance Related

🎯 What it does: Propose a Drift-Aware Dynamic Mixture of Experts (Dynamic TMoE) framework based on a dynamic expert pool, aimed at addressing distribution drift issues in non-stationary time series.

Dynamics Are Learned, Not Told: Semi-Supervised Discovery of Latent Dynamics Geometries For Zero-Shot Policy Adaptation

Zhiming Xu (Tongji University), Chenpeng Yao (Tongji University)

CodeDomain AdaptationRepresentation LearningTransformerReinforcement LearningAuto EncoderContrastive LearningTabularTime Series

🎯 What it does: A framework is studied that directly learns the latent dynamic geometry from interaction trajectories, achieving zero-shot policy adaptation to unseen dynamic environments.

DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics

Silin Gao (EPFL), Antoine Bosselut (EPFL)

CodeGenerationRepresentation LearningTransformerLarge Language ModelMixture of ExpertsVision Language ModelDiffusion modelContrastive LearningWorld ModelImageVideoText

🎯 What it does: Propose DynaVieW, a world model based on hierarchical JSON schema, which jointly learns visual state simulation and state transition prediction.

E-VAds: An E-commerce Short Videos Understanding Benchmark for MLLMs

Xianjie Liu (Alimama Tech, Taobao & Tmail Group of Alibaba), Bo Zheng (Alimama Tech, Taobao & Tmail Group of Alibaba)

CodeRecommendation SystemReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningVision Language ModelVideoTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-ThoughtAudio

🎯 What it does: Designed and released the E-VAds benchmark specifically for e-commerce short videos, and trained a reinforcement learning-based reasoning model called E-VAds-R1 on this benchmark.

EARL: Towards a Unified Analysis-Guided Reinforcement Learning Framework for Egocentric Interaction Reasoning and Pixel Grounding

Yuejiao Su (Hong Kong Polytechnic University), Yi Wang (Hong Kong Polytechnic University)

CodeImage TranslationObject DetectionSegmentationPose EstimationComputational EfficiencyRepresentation LearningData-Centric LearningRobotic IntelligenceTransformerReinforcement LearningPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose a unified analysis-guided reinforcement learning framework called EARL for first-person perspective interactive reasoning and pixel-level localization.

EchoRL: Reinforcement Learning via Rollout Echoing

Jinhe Bi (Huawei Heisenberg Research Center), Yunpu Ma (LMU Munich)

CodeComputational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextTabularBenchmarkChain-of-Thought

🎯 What it does: Propose EchoRL, a lightweight module that identifies key segments (EchoClip) from verified successful rollouts by utilizing step-level entropy, and injects it as an auxiliary supervisory signal into RLVR training to address the gradient vanishing problem caused by advantage degradation.

EEmo-Logic: A Unified Dataset and Multi-Stage Framework for Comprehensive Image-Evoked Emotion Assessment

Lancheng Gao (Shanghai Jiao Tong University), Xiongkuo Min (Shanghai Jiao Tong University)

CodeClassificationRecognitionRecommendation SystemExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityChain-of-Thought

🎯 What it does: This paper proposes the EEmo-Logic framework, which first constructs the largest image-evoked emotion understanding dataset, EEmoDB (containing QA and fine-grained evaluation parts), and then achieves unified and fine-grained understanding of emotion QA, emotion ranking, emotion description, and fine-grained emotion assessment through a two-stage training approach (LoRA SFT + GRPO) on multimodal large language models.

EffGen: Enabling Small Language Models as Capable Autonomous Agents

Gaurav Srivastava (Virginia Tech), Xuan Wang (Virginia Tech)

CodeAutonomous DrivingOptimizationExplainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose the EFFGEN framework, enabling small language models (SLM) to efficiently and securely perform multi-tool calling, task decomposition, and memory management, constructing a locally deployable agent system.

Efficient Adaptive Testing via Gradient Path Matching Subset Selection for AI Education

Yan Zhuang (Nanjing University of Aeronautics and Astronautics), Daoqiang Zhang (Nanjing University of Aeronautics and Astronautics)

CodeRecommendation SystemOptimizationComputational EfficiencySpiking Neural NetworkTransformerReinforcement LearningContrastive LearningTextTabularBenchmark

🎯 What it does: Proposes the Gradient Path Matching (GPM) framework, which utilizes gradient path matching for subset selection of questions, thereby achieving more accurate ability estimation with fewer questions in AI education adaptive assessment.

Efficient Code Analysis via Graph Representation Learning-Guided Large Language Models

Hang Gao (Key Laboratory of System Software, Institute of Software, Chinese Academy of Sciences), Jian Zhang (Key Laboratory of System Software, Institute of Software, Chinese Academy of Sciences)

CodeAnomaly DetectionExplainability and InterpretabilityComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextGraph

🎯 What it does: Propose a malicious code detection framework called GMLLM that combines graph representation learning with large language models (LLMs). It uses graph neural networks (GNNs) to identify key code snippets and then guides the LLM to perform in-depth analysis on them.

Efficient Distributed MLLM Training with Cornstarch

Insu Jang (University of Michigan), Mosharaf Chowdhury (University of Michigan)

CodeOptimizationComputational EfficiencyTransformerLarge Language ModelMixture of ExpertsVision Language ModelContrastive LearningImageTextMultimodalityAudio

🎯 What it does: This paper proposes the Cornstarch framework, specifically designed to achieve efficient distributed training for multi-modal large language models (MLLM), addressing performance bottlenecks caused by frozen parameters, non-causal attention patterns, and model/data heterogeneity.

Efficient Equivariant High-Order Crystal Tensor Prediction via Cartesian Local-Environment Many-Body Coupling

Dian Jin (Hong Kong Polytechnic University), Xiaoming Tao (Hong Kong Polytechnic University)

CodeComputational EfficiencyDrug DiscoveryGraph Neural NetworkTransformerAuto EncoderContrastive LearningGraphTabularPhysics Related

🎯 What it does: Proposed a new equivariant higher-order crystal tensor prediction framework, CEITNet, which can directly and end-to-end predict second- to fourth-order tensors (dielectric, piezoelectric, elastic tensors) from crystal structures.

Efficient Hallucination Detection for LLMs Using Uncertainty-Aware Attention Heads

Artem Vazhentsev (Mohamed bin Zayed University of Artificial Intelligence), Artem Shelmanov (Mohamed bin Zayed University of Artificial Intelligence)

CodeAnomaly DetectionExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelText

🎯 What it does: Propose an unsupervised, single-forward-pass hallucination detection framework named RAUQ, which identifies factual errors in LLM-generated text by leveraging the change in attention head focus on the previous token from the Transformer.

Efficient Learning of Deep State Space Models via Importance Smoothing

John-Joseph Brady (King's College London), Yunpeng Li (King's College London)

CodeGenerationData SynthesisOptimizationComputational EfficiencyReinforcement Learning from Human FeedbackDiffusion modelScore-based ModelAuto EncoderContrastive LearningTabularTime SeriesSequentialFinance RelatedPhysics RelatedStochastic Differential Equation

🎯 What it does: Proposed and implemented a parallel differentiable particle smoothing method called PVMC for efficiently training deep state space models, avoiding the serial limitations of traditional particle filters.

Efficient numeracy in language models through single-token number embeddings

Linus Kreitner (Technical University of Munich), Martin J. Menten (Technical University of Munich)

CodeComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringText

🎯 What it does: Propose a numerical encoding for individual tokens called BitTokens, enabling language models to perform arithmetic operations in a single forward pass, thereby improving the model's numerical reasoning efficiency and accuracy.

Efficient Prediction of SO(3)-Equivariant Hamiltonian Matrices via SO(2) Local Frames

Haiyang Yu (Texas A&M University), Shuiwang Ji (Texas A&M University)

CodeComputational EfficiencyDrug DiscoveryGraph Neural NetworkContrastive LearningGraphTabularBenchmarkPhysics Related

🎯 What it does: Propose the QHNetV2 network, which achieves global equivariance under SO(3) symmetry through the SO(2) local frame, and predicts the quantum Hamiltonian matrix, thereby significantly accelerating the computation.