arXivSub Start free trial

ICML 2026 Papers — Page 36

International Conference on Machine Learning · 6554 papers

Mixture of Distributions Matters: Dynamic Sparse Attention for Efficient Video Diffusion Transformers

Yuxi Liu (Peking University), Kun Yuan (Peking University)

GenerationComputational EfficiencyTransformerDiffusion modelVideo

🎯 What it does: Propose MOD-DiT, a dynamic sparse attention framework that requires no training and no sampling, to accelerate the inference of video diffusion Transformers while maintaining or improving video generation quality.

Mixture of Horizons in Action Chunking

Dong Jing (Renmin University of China), Mingyu Ding (University of North Carolina at Chapel Hill)

Robotic IntelligenceTransformerReinforcement LearningMixture of ExpertsVision-Language-Action ModelContrastive LearningVideoSequential

🎯 What it does: Propose the Mixture of Horizons (MoH) strategy, which processes action segments of different lengths in parallel, alleviating the trade-off between action block length and model performance, and achieving dynamic inference.

Mixture Prototype Flow Matching for Open-Set Supervised Anomaly Detection

Fuyun Wang (Nanjing University of Science and Technology), Zhen Cui (Beijing Normal University)

Anomaly DetectionFlow-based ModelImageBiomedical Data

🎯 What it does: Propose a flow matching framework based on Gaussian Mixture Prototypes, named MPFM, for open-supervised anomaly detection;

Mixtures Closest To A Given Measure: A Semidefinite Programming Approach

Srecko Durasinovic, Victor Magron (Université de Toulouse)

OptimizationImageTabular

🎯 What it does: Propose a hierarchical relaxation method based on semidefinite programming (SDP), which approximates a mixture distribution composed of parameterized distribution families (such as Gaussian, exponential, Poisson) by utilizing the finite-order moments of the target measure. The objective is to minimize the 2-Wasserstein or total variation distance.

Mixtures of geodesic factor analyzers on Riemannian homogeneous spaces

Hengchao Chen (University of Toronto), Qiang Sun (University of Toronto)

Anomaly DetectionExplainability and InterpretabilityComputational EfficiencyRepresentation LearningDiffusion modelScore-based ModelContrastive LearningPoint CloudMeshBiomedical DataAlzheimer's DiseaseReview/Survey Paper

🎯 What it does: Propose a Mixed Geographical Factor Analysis (MGFA) model on Riemannian homogeneous spaces for clustering and modeling manifold-valued data.

ML-Agent: Reinforcing LLM Agents for Autonomous Machine Learning Engineering

Zexi Liu (Shanghai Jiao Tong University), Siheng Chen (Shanghai Jiao Tong University)

Autonomous DrivingOptimizationFederated LearningHyperparameter SearchData-Centric LearningRobotic IntelligenceAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AIPrompt EngineeringMixture of ExpertsImageVideoTextMultimodalityTabularTime SeriesAudio

🎯 What it does: Proposed and implemented a reinforcement learning-based LLM agent training framework for autonomous machine learning engineering, completing the full cycle from code generation to experimental feedback.

ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World

Ziyin Zhang (Shanghai Jiao Tong University), Rui Wang (Shanghai Jiao Tong University)

RetrievalCompressionComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelContrastive LearningTextMultimodalityBenchmark

🎯 What it does: Proposed the ML-Embed series of models, utilizing the 3-D Matryoshka Learning (3D-ML) framework to achieve multi-dimensional compression of the embedding layer, network depth, and representation size, enabling significant savings in training, inference, and storage;

MLLM-4D: Towards Visual-based Spatial-Temporal Intelligence

Xingyilang Yin (University of Macau), Xiaodong Cun (Great Bay University)

Autonomous DrivingRobotic IntelligenceReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision Language ModelVideoTextMultimodalityBenchmarkChain-of-Thought

🎯 What it does: Developed a multi-modal large language model framework (MLLM-4D) for visual spatiotemporal intelligence, enhancing the model's understanding and reasoning capabilities of 3D space evolution over time through automated stereo video to 4D data conversion and post-training strategies.

MLUBench: A Benchmark for Lifelong Unlearning Evaluation in MLLMs

He Li (National University of Defense Technology), Bo Han (Hong Kong Baptist University)

Federated LearningSafty and PrivacyExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningAdversarial AttackData-Centric LearningTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsVision Language ModelContrastive LearningImageTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposes a new benchmark for lifelong unlearning in multimodal large language models (MLLM), named MLUBench, and achieves sustainable forgetting through the LUMoE method based on a Mixture-of-Experts (MoE) + LoRA framework.

MM-DeepResearch: A Simple and Effective Multimodal Agentic Search Baseline

Huanjin Yao (ByteDance), Jiaxing Huang (Hong Kong Polytechnic University)

RetrievalRecommendation SystemOptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AIMixture of ExpertsVision Language ModelImageTextMultimodalityRetrieval-Augmented GenerationChain-of-ThoughtAudio

🎯 What it does: Developed a multi-modal deep research agent called MM-DeepResearch, capable of explicit reasoning, planning, multi-tool calling, and fusing cross-modal information, supporting multi-round searches.

MM-Snowball: Evaluating and Mitigating Hallucination Snowballing in Multimodal Multi-Turn Dialogue

Yue Jiang (Fudan University), Dingkang Yang (Fudan University)

GenerationData SynthesisExplainability and InterpretabilityComputational EfficiencyRepresentation LearningAdversarial AttackTransformerLarge Language ModelPrompt EngineeringVision Language ModelGenerative Adversarial NetworkContrastive LearningImageTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: Construct the MM-Snowball benchmark and propose the CAVR method to evaluate and mitigate the hallucination avalanche phenomenon in multi-modal multi-turn dialogues.

MM-Spectrum: Multimodal Multi-spectral Molecular Structural Elucidation with a Stable MoE Framework

Hai-tao Yu (Hong Kong University of Science and Technology), Jun Xia (Hong Kong University of Science and Technology)

Drug DiscoveryTransformerMixture of ExpertsMultimodalityBiomedical DataMagnetic Resonance ImagingReview/Survey Paper

🎯 What it does: Propose MM-Spectrum, a sparse Mixture-of-Experts (MoE) framework, to integrate multi-modal (NMR, IR, MS) spectra and predict molecular structures.

MMBench-Live: A Continuously Evolving Benchmark for Multimodal Models

Yuanzhi Liu (Beijing University of Posts and Telecommunications), Zhanyu Ma (Beijing University of Posts and Telecommunications)

TransformerLarge Language ModelAgentic AIPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: Constructed a sustainable and continuously updated multimodal evaluation benchmark called MMBench-Live, utilizing a multi-agent automated pipeline to persistently generate new instances.

mmBERT: A Modern Multilingual Encoder with Annealed Language Learning

Marc Marone (Johns Hopkins University), Benjamin Van Durme (Johns Hopkins University)

Representation LearningData-Centric LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextMultimodality

🎯 What it does: Proposed MMBERT, a multilingual encoder model designed for over 1800 languages;

MMClima: A Framework for Multimodal Climate Science Data and Evaluation

Muhammad Umer Sheikh (Mohamed bin Zayed University of Artificial Intelligence), Muhammad Haris Khan (Mohamed bin Zayed University of Artificial Intelligence)

TransformerLarge Language ModelPrompt EngineeringVision Language ModelImageVideoTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: Constructed MMCLIMA, a multimodal climate QA benchmark covering text, video transcripts, and scientific charts, containing over 104k expert-verified question-answer pairs;

MMD Guidance: Training-Free Distribution Adaptation for Diffusion Models via Maximum Mean Discrepancy Guidance

Matina Mahdizadeh Sani (University Of Waterloo), Farzan Farnia (Chinese University Of Hong Kong)

GenerationDomain AdaptationDiffusion modelScore-based ModelContrastive LearningImageTextMultimodality

🎯 What it does: Proposed a training-agnostic inference-time guidance method called MMDGuidance, which uses the gradient of the Maximum Mean Discrepancy (MMD) to guide the generated samples in the reverse sampling process of diffusion models to align with the target distribution defined by a small number of reference samples.

MME-Reasoning: A Broad-Spectrum Benchmark for Evaluating Logical Reasoning in MLLMs

Jiakang Yuan (Fudan University), Bo Zhang (Shanghai Artificial Intelligence Laboratory)

TransformerLarge Language ModelReinforcement LearningTextMultimodalityBenchmarkChain-of-Thought

🎯 What it does: Designed and constructed the MME-Reasoning benchmark to systematically evaluate the logical reasoning capabilities of multi-modal large language models (MLLMs), covering three types of reasoning: inductive, deductive, and abductive, and providing various question types (multiple-choice, free-response, rule-based) and difficulty levels.

MMKU-Bench: A Multimodal Update Benchmark for Diverse Visual Knowledge

Baochen Fu (Shandong University), Yi Wan (Shandong University)

TransformerSupervised Fine-TuningReinforcement LearningVision Language ModelImageTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposed MMKU-Bench, a systematic evaluation framework for assessing the performance of multimodal models in updating existing knowledge and learning new knowledge, and constructed a VQA dataset containing 25,792 knowledge instances and 49,000 images.

MMPD-Bench: Bridging Multimodal Fission with Multi-Polarimetric Modalities Decomposition

Yi He (University College London), Yukun Hu (University College London)

TransformerSupervised Fine-TuningDiffusion modelContrastive LearningImageBiomedical DataBenchmarkPhysics Related

🎯 What it does: Proposed MMPD-Bench, redefining the Mueller matrix multi-polarization decomposition task as a multi-modal separation (Modality Fission) problem, and constructing a unified benchmark framework.

MobileFusion: Mobile-Friendly Infrared and Visible Image Fusion via Structural Re-parameterization

Yufa Duan (Xiamen University), Xiaotong Tu (Xiamen University)

Image TranslationRestorationComputational EfficiencyConvolutional Neural NetworkAuto EncoderGenerative Adversarial NetworkContrastive LearningImageMultimodality

🎯 What it does: Proposed a lightweight, real-time mobile infrared-visible image fusion framework called MobileFusion

Mobility-Embedded POIs: Learning What A Place Is and How It Is Used from Human Movement

Maria Despoina Siampou (University of Southern California), Cyrus Shahabi (University of Southern California)

Recommendation SystemFederated LearningSafty and PrivacyRepresentation LearningData-Centric LearningTransformerSupervised Fine-TuningAuto EncoderContrastive LearningGaussian SplattingTextTabularTime SeriesSequential

🎯 What it does: Proposes the ME-POIS framework, which combines static text embeddings with large-scale human mobility data to generate global embeddings that contain both POI identity and functionality through contrastive learning and multi-scale distribution migration.

MOC: Multi-Order Communication in LLM-based Multi-Agent Systems

Yao Guan (Fudan University), Qiang Duan (Pennsylvania State University)

OptimizationFederated LearningComputational EfficiencyKnowledge DistillationReinforcement Learning from Human FeedbackGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringTextGraphRetrieval-Augmented Generation

🎯 What it does: Propose a Multi-Order Communication (MOC) scheme to transmit raw information with multi-hop dependencies in a structured manner to the target agent within large language model (LLM)-driven multi-agent systems, and design a Semantic-Topological Merging mechanism to compress redundant information and improve communication efficiency.

MoCL: Metabolic Optimization for Curvature-Aware Continual Learning

Jiajun Lai (South China University of Technology), Huaiguang Jiang (South China University of Technology)

OptimizationComputational EfficiencyRepresentation LearningMeta LearningContrastive LearningImageBenchmark

🎯 What it does: Proposes the MoCL framework, using metabolic gating and decomposition subspace approximation to mitigate catastrophic forgetting in continual learning.

MoCo-EA: Exploiting Adversarial Mode Connectivity for Efficient Evolutionary Attacks

Hyo Seo Kim (Illinois Institute of Technology), Ren Wang (Illinois Institute of Technology)

Adversarial AttackConvolutional Neural NetworkTransformerImage

🎯 What it does: Propose MoCo-EA evolutionary attack, which efficiently generates adversarial examples in white-box scenarios by achieving cross-parent crossover through Bézier curves.

MOD-SR: Unifying Multimodal Learning and Direct Optimization with Gradient-Guided Diffusion Model for Symbolic Regression

Chuyang Xiang (Shanghai Jiao Tong University), Junchi Yan (Shanghai Jiao Tong University)

GenerationOptimizationRepresentation LearningTransformerDiffusion modelScore-based ModelContrastive LearningMultimodalityTabularBenchmarkPhysics RelatedStochastic Differential Equation

🎯 What it does: Propose the MOD-SR framework, which unifies multi-modal distribution learning with direct optimization, generating symbolic expressions through gradient-guided diffusion models.

MoDA: Modulation Adapter for Fine-Grained Visual Understanding in Instructional MLLMs

Wayner Barrios (Dartmouth College), Bernard Ghanem (King Abdullah University of Science and Technology)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerSupervised Fine-TuningPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBenchmark

🎯 What it does: Proposed a lightweight channel-level modulation adapter called MoDA, which dynamically modulates visual features with language instructions to enhance fine-grained visual understanding in multi-modal large language models and reduce hallucinations.

Modality-Decoupled Online Recursive Editing

Siyuan Li (Harbin Institute Of Technology), Jing Li (Harbin Institute Of Technology)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelVision Language ModelContrastive LearningImageTextMultimodalityRetrieval-Augmented Generation

🎯 What it does: Propose an online recursive editing framework called M-ORE for multi-modal large language models (MLLM), which can continuously correct model knowledge under limited computational and memory budgets.

Mode Seeking meets Mean Seeking for Fast Long Video Generation

Shengqu Cai (Stanford University), Arash Vahdat (NVIDIA Research)

GenerationTransformerSupervised Fine-TuningDiffusion modelVideo

🎯 What it does: Propose a long video generation training paradigm that utilizes Decoupled Diffusion Transformer (DDT) to separate local details from long-term temporal consistency. It supervises the global structure through flow matching (Mean Seeking) on scarce long videos, while aligning with frozen short video teachers on sliding windows through distribution matching (Mode Seeking), ultimately achieving high-quality minute-level video generation with few steps.

Model Fusion via Retrofitting

Phoomraphee Luenam (ETH Zürich), Sidak Pal Singh (ETH Zürich)

ClassificationFederated LearningExplainability and InterpretabilityKnowledge DistillationConvolutional Neural NetworkTransformerMixture of ExpertsAuto EncoderContrastive LearningImageBiomedical Data

🎯 What it does: This paper proposes a fusion framework based on neuron clustering and alignment (Retrofitting), which first performs importance-weighted clustering on the intermediate neurons of the parent model, and then trains the subnetwork of the fusion model to approximate the cluster centers, achieving fusion for any hierarchically structured DAG model.

Model Merging Scaling Laws in Large Language Models

Yuanyi Wang (Hong Kong Polytechnic University), Hongxia Yang (Hong Kong Polytechnic University)

Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelMixture of ExpertsText

🎯 What it does: Empirically investigate scaling laws in expert merging of large language models, propose a unified floor+tail formula, and conduct large-scale validation across multiple models, methods, and domains.

Model Monotonicity in Autobidding Auctions: When Do Better Predictions Lead to Better Outcomes?

Ashwinkumar Badanidiyuru (Uber Technologies, Inc)

OptimizationFederated LearningReinforcement Learning from Human FeedbackReview/Survey PaperBenchmarkFinance Related

🎯 What it does: Explores the impact of model improvements in advertising auctions and provides a formal definition of model improvements along with a systematic study on the monotonicity of platform-level evaluation metrics (ECM)

MODEL SOUPS NEED ONLY ONE INGREDIENT

Alireza Abdollahpoorrostam (École Polytechnique Fédérale de Lausanne), Pascal Frossard (École Polytechnique Fédérale de Lausanne)

ClassificationDomain AdaptationComputational EfficiencyKnowledge DistillationTransformerSupervised Fine-TuningContrastive LearningImageTextMultimodality

🎯 What it does: Propose MonoSoup, a single-model post-processing method that enhances the model's performance balance within-distribution (ID) and out-of-distribution (OOD) by performing singular value decomposition on fine-tuned weight updates and adaptively reweighting them.

Model-Based Diffusion Sampling for Predictive Control in Offline Decision Making

Haldun Balim (Harvard University), Yilun Du (Harvard University)

OptimizationRobotic IntelligenceTransformerReinforcement LearningDiffusion modelAuto EncoderImageVideoTabularTime Series

🎯 What it does: Propose Model Predictive Diffuser (MPDiffuser), which generates task-consistent and dynamically feasible trajectories by combining a diffusion model with a planner and a dynamics model through alternating sampling, and uses a sorter to select trajectories that satisfy constraints.

Model-Dowser: Data-Free Importance Probing to Mitigate Catastrophic Forgetting in Multimodal Large Language Models

Hyeontaek Hwang (KAIST), Daeyoung Kim (KAIST)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelImageTextMultimodality

🎯 What it does: Propose a sparse fine-tuning method called Model-Dowser, which measures and freezes the parameters that have the greatest impact on model outputs through data-agnostic sensitivity detection, thereby reducing catastrophic forgetting when adapting multi-modal large language models (MLLMs).

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis

Zachary Andrew Roch (University of Central Florida), Yue Wang (University of Central Florida)

Reinforcement Learning

🎯 What it does: Propose a model-free robust average reward reinforcement learning algorithm called Robust Halpern Iteration (RHI), and provide its sample complexity analysis under the generative model setting.

Model-Preserving Adaptive Rounding

Albert Tseng (Cornell University), Christopher De Sa (Cornell University)

OptimizationComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelContrastive LearningText

🎯 What it does: Propose a new adaptive quantization algorithm called YAQA, which can directly optimize the end-to-end KL divergence during the quantization of LLM weights, keeping the minimal change in the model output distribution.

Modeling Attributional Style at Scale: A Dataset and Analysis for Psychological Attribution Assessment and Reframing

Qiang Zhou (University of Pittsburgh), Jingtong Hu (University of Pittsburgh)

ClassificationRecognitionData SynthesisExplainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: A large-scale Attributed Style Transfer Dataset (ASTD) was constructed and released, and the tasks of attributed style identification and rewriting were evaluated based on this dataset.

Modeling Covariate Transition for Efficient Estimation of Longitudinal Treatment Effects in Randomized Experiments

Naoki Chihara (SANKEN University of Osaka), Shota Yasui (CyberAgent)

OptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyContrastive LearningTabularTime SeriesBenchmarkPhysics Related

🎯 What it does: This study proposes a regression adjustment framework that utilizes covariance shift kernels and recursive forward integration for efficient estimation of long-term treatment effects in randomized experiments.

Modeling Hierarchical Thinking in Large Reasoning Models

G M Shahariar (University of California Riverside), Nael Abu-Ghazaleh (University of California Riverside)

Explainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningContrastive LearningTextBenchmarkChain-of-Thought

🎯 What it does: This study investigates hierarchical thinking in large reasoning models (LRM), abstracting Chain-of-Thought into six cognitive states using a finite state machine (FSM), and designing a sparse activation guided control method without weighted updates during training through transition advantage matrices and Q-Value iteration;

Modeling Long-Tail Relations in the Operating Room via In-Context Multimodal Learning

Boqiang Xu (Centre for Artificial Intelligence and Robotics, Hong Kong Institute of Science and Innovation, Chinese Academy of Sciences), Zhen Lei (University of Chinese Academy of Sciences)

RecognitionGenerationRetrievalComputational EfficiencyTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityPoint CloudRetrieval-Augmented GenerationAudio

🎯 What it does: Propose a model called SGG-ICL for generating scene graphs in operating room scenarios, specifically addressing long-tail relationships, using selective reasoning and multi-modal retrieval to improve prediction performance for long-tail categories.

Modeling Spectral Energy Shifts in Spatio-Temporal Graph Anomaly Detection

Yilin Liu (Vanderbilt University), Meiyi Ma (Vanderbilt University)

Anomaly DetectionGraph Neural NetworkAuto EncoderContrastive LearningGraphTime SeriesBenchmark

🎯 What it does: Proposes a graph anomaly detection method called EGNN based on node-level spectral energy, which can simultaneously capture high-variance anomalies and masked low-variance anomalies, and achieve energy-driven message passing in spatiotemporal graphs without requiring a dedicated sequence module;

Modeling Temporal scRNA-seq Data with Latent Gaussian Process and Optimal Transport

Mehmet Yigit Balik (Aalto University), Harri Lähdesmäki (Aalto University)

GenerationData SynthesisRepresentation LearningDiffusion modelScore-based ModelFlow-based ModelAuto EncoderContrastive LearningGaussian SplattingTime SeriesSequentialBiomedical DataStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Built a generative model based on implicit Gaussian processes and optimal transport for reconstructing continuous temporal dynamics from static snapshots of single-cell RNA sequencing data;

Modelling Attention with Aitchison Geometry: Token Distinguishability and Temperature Scaling

Sam Hilton-Jones (University of Southampton), Zhanxing Zhu (University of Southampton)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningText

🎯 What it does: The paper models the softmax dilution problem in Transformer, proposing to use Aitchison distance to measure the distinguishability of attention distribution tokens, and theoretically analyzes how it changes with context length and embedding dimension.

Models Under SCOPE: Scalable and Controllable Routing via Pre-hoc Reasoning

Qi Cao (University of California San Diego), Pengtao Xie (University of California San Diego)

Recommendation SystemOptimizationComputational EfficiencyTransformerSupervised Fine-TuningReinforcement LearningMixture of ExpertsTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposes SCOPE, a scalable and controllable model routing framework based on pre-reasoning, which uses model behavior fingerprints to predict the accuracy and cost of each candidate model, and dynamically selects the optimal model according to user budget.

ModernVBERT: Towards Smaller Visual Document Retrievers

Paul Teiletche (Illiun Technology), Manuel Faysse

RetrievalComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodality

🎯 What it does: This paper systematically studies the core design and training process of visual document retrieval (VDR) models, proposing and training a lightweight, from-scratch multi-modal encoder called ModernVBERT (250M parameters), which achieves retrieval performance comparable to large-scale models while maintaining extremely low latency.

Modular Pretraining Enables Access Control

Ethan Roland (AE Studio), Alex Cloud (Anthropic)

Federated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningData-Centric LearningMeta LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningMixture of ExpertsTextTabularTime SeriesSequential

🎯 What it does: By introducing a gradient routing auxiliary module (GRAM) into the MLP layer of the Transformer, multiple detachable capability configurations are achieved in a single pre-training process, allowing the model to control its dual-purpose capabilities during inference by activating or masking different modules.

MODUS: Decoder-only Any-to-Any Modeling of Diverse Modalities

Mingqiao Ye (EPFL), Amir Zamir (EPFL)

GenerationData SynthesisRepresentation LearningTransformerLarge Language ModelMixture of ExpertsVision Language ModelDiffusion modelFlow-based ModelAuto EncoderImageTextMultimodality

🎯 What it does: Propose MODUS, a unified decoder-only model that achieves arbitrary mapping and generation between any modalities.

MOES-Pred: Molecular Structural Representation Learning by Adaptive Energy-Sentinel Vibration for Generalized Property Prediction

Zhiran Hou (Jiangsu Ocean University), Ming Li (Jiangsu Ocean University)

Representation LearningDrug DiscoveryGraph Neural NetworkSupervised Fine-TuningAuto EncoderContrastive LearningGraphBiomedical Data

🎯 What it does: Self-supervised denoising pre-training on 3D molecular structures is conducted through an energy sentinel mechanism and a molecule-specific noise scheme to learn molecular force fields.

MolAlign3D: Enhancing Fixed-Dimensional E(3)-Equivariant Latent Space for High-Fidelity 3D Molecular Reconstruction and Editing

Zitao Chen (Tsinghua University), Yanyan Lan (Tsinghua University)

Drug DiscoveryProtein Structure PredictionGraph Neural NetworkTransformerDiffusion modelFlow-based ModelAuto EncoderGraphBiomedical Data

🎯 What it does: Propose MolAlign3D, a fixed-dimension, E(3)-equivariant latent space, combined with a pre-trained molecular encoder to achieve high-fidelity 3D molecular reconstruction and editable features.

MoLF: Mixture-of-Latent-Flow for Pan-Cancer Spatial Gene Expression Prediction from Histology

Susu Hu (National Center for Tumor Diseases), Stefanie Speidel (National Center for Tumor Diseases)

Drug DiscoveryTransformerMixture of ExpertsDiffusion modelFlow-based ModelAuto EncoderBiomedical DataStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Proposed the MoLF model, which utilizes Mixture-of-Latent-Flow to predict spatial gene expression from tissue sections at a pan-cancer level, taking into account multi-sample heterogeneity and generative modeling;

MoLoRA: Composable Specialization via Per-Token Adapter Routing

Shrey Shah (Microsoft), Justin Wagle (Microsoft)

Computational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringMixture of ExpertsTextMultimodality

🎯 What it does: Proposes the MoLoRA multi-adapter service framework based on per-token routing, which can dynamically select the most suitable LoRA adapter for each token within a single model deployment, achieving mixed inference of multi-modal and multi-capability;

Moment Matching Q-Learning

Yiyan Liang Edgar, Weitong Zhang (University of North Carolina at Chapel Hill)

Reinforcement LearningDiffusion modelFlow-based ModelTabularTime SeriesSequentialBenchmark

🎯 What it does: This paper proposes the Moment Matching Q-Learning (MoMa QL) framework, which utilizes the Maximum Mean Discrepancy (MMD) regularization to match the distribution of generative policies, thereby significantly accelerating sampling and improving performance in offline reinforcement learning and offline-to-online fine-tuning.

Momentum Further Constrains Sharpness at the Edge of Stochastic Stability

Arseniy Andreyev (Princeton University), Pierfrancesco Beneventano (Massachusetts Institute of Technology)

OptimizationConvolutional Neural NetworkTransformerContrastive LearningImageTextStochastic Differential Equation

🎯 What it does: This study investigates the 'edge of stability' (EOSS) behavior of stochastic gradient descent with momentum (SGDM/SGDN) during mini-batch training, revealing that momentum leads to different curvature saturation levels and self-organizes at unstable edges under varying batch sizes.

Monitorability as a Free Gift: How RLVR Spontaneously Aligns Reasoning

Zidi Xiong (Harvard University), Himabindu Lakkaraju (Harvard University)

Explainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkChain-of-Thought

🎯 What it does: This paper systematically evaluates the evolution of monitorability during the RLVR training process, revealing its strong dependence on data distribution and its incomplete correlation with reasoning ability.

Monitoring Monitorability

Melody Y. Guan (OpenAI), Bowen Baker (OpenAI)

Explainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringTextBenchmarkChain-of-Thought

🎯 What it does: This paper proposes an evaluation framework for monitoring monitorability, defines a new g-mean2 metric, and constructs three types of evaluation scenarios (intervention, process, and outcome attributes), conducting systematic experiments on the chain-of-thought monitoring effectiveness of state-of-the-art reasoning models using multiple datasets.

MonoScale: Scaling Multi-Agent System with Monotonic Improvement

Shuai Shao (Shanghai Jiao Tong University), Weinan Zhang (Shanghai Jiao Tong University)

Autonomous DrivingOptimizationFederated LearningRobotic IntelligenceTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsTextRetrieval-Augmented Generation

🎯 What it does: The paper proposes the MonoScale framework, which addresses the incremental expansion of multi-agent systems by generating targeted familiarization tasks for newly added agents and updating an editable text memory using collected evidence of success/failure, ensuring that the router's performance monotonically improves during incremental expansion.

Monotonic Variational Gaussian Process for Efficient Data Collection

Donghyun Lee (Pohang University of Science and Technology), Young Myoung Ko (Pohang University of Science and Technology)

OptimizationData-Centric LearningContrastive LearningGaussian SplattingImageVideoTabularStochastic Differential Equation

🎯 What it does: Propose a learning curve model and expected shortage objective based on monotonic variational Gaussian processes for multi-source data acquisition optimization.

MOOSE-Star: Unlocking Tractable Training for Scientific Discovery by Breaking the Complexity Barrier

Zonglin Yang (MiroMind AI), Lidong Bing (MiroMind AI)

OptimizationComputational EfficiencyData-Centric LearningDrug DiscoveryTransformerLarge Language ModelPrompt EngineeringTextReview/Survey PaperRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose the MOOSE-Star framework, which decomposes the exponential search problem of directly training P(h|b) into subtasks that are linear or even logarithmic, enabling a trainable model for scientific discovery;

MORALISE: A Structured Benchmark for Moral Alignment in Visual Language Models

Xiao Lin (University of Illinois Urbana-Champaign), Hanghang Tong (University of Illinois Urbana-Champaign)

Explainability and InterpretabilityComputational EfficiencyPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBenchmark

🎯 What it does: Proposes the MORALISE benchmark, using real image-text pairs to evaluate the alignment of vision-language models on morally sensitive tasks.

More Capable, Less Cooperative? When LLMs Fail at Zero-Cost Collaboration

Advait Yadav (MATS), Oliver Sourbut (Future of Life Foundation)

Autonomous DrivingOptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyRobotic IntelligenceMeta LearningDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: The study constructs a cost-free, non-competitive multi-agent LLM collaboration environment to test the cooperative performance of eight mainstream LLMs in helping behaviors.

More Edits, More Stable: Understanding the Lifelong Normalization in Sequential Model Editing

Xin Ma (University Of Science And Technology Of China), Enhong Chen (University Of Science And Technology Of China)

OptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencySupervised Fine-TuningReinforcement LearningContrastive LearningTextBiomedical DataBenchmark

🎯 What it does: This paper investigates the stability issue in Lifelong Model Editing (LME), and for the first time reveals the theoretical mechanism of Lifelong Normalization (LN), and based on this, designs the STABLEEDIT editor to improve the stability of long-term editing.

More Sail than Ballast: Addressing Harmful Knowledge Leakage in the Expansive Reasoning Space of LRMs

Qibing Ren (Shanghai Jiao Tong University), Jing Shao (Shanghai Artificial Intelligence Laboratory)

Data SynthesisSafty and PrivacyTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmark

🎯 What it does: Studies the problem of unintentionally leaking harmful knowledge during long-chain reasoning in large inference models and proposes a solution.

MORE: A Multilingual Document Parsing Benchmark and Evaluation

Long Xu (Tencent), feng zhang

TransformerPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose the MORE benchmark for multilingual document parsing evaluation.

MoRGen: Mixture-of-Resolutions Generative Forecasting for Irregularly Sampled Medical Time-Series Data

Nassim Oufattole (Massachusetts Institute of Technology), Collin Stultz (Harvard Medical School)

GenerationData SynthesisAnomaly DetectionTransformerMixture of ExpertsTabularTime SeriesBiomedical DataElectronic Health Records

🎯 What it does: Proposed the MoRGen method, which performs zero-shot medical time series risk prediction by fusing generative predictors with different temporal resolutions.

MoSA: Motion-constrained Stress Adaptation for Mitigating Real-to-Sim Gap in Continuum Dynamics via Learning Residual Anisotropy

Jiaxu Wang (Hong Kong University of Science and Technology), Renjing Xu (Hong Kong University of Science and Technology)

Robotic IntelligenceNeural Radiance FieldGaussian SplattingOptical FlowImageVideoPhysics Related

🎯 What it does: Proposes a framework called MoSA based on residual stress adaptation, for learning continuous dynamics of real-world objects from multi-view videos.

Mosaic: Runtime-Efficient Multi-Agent Embodied Planning

Kunjal Panchal (University of Massachusetts), Hui Guan (University of Massachusetts)

OptimizationComputational EfficiencyRobotic IntelligenceReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextMultimodalityChain-of-Thought

🎯 What it does: Proposes the MOSAIC framework to enhance the runtime efficiency and success rate of LLM-driven multi-agent systems in embodied environments

Mosaic: Unlocking Over 30$\times$ Context Length for Diffusion LLMs Inference via Global Memory Planning and Dynamic Peak Taming

Liang Zheng (Tianjin University), Keqiu Li (Tianjin University)

Computational EfficiencyTransformerLarge Language ModelDiffusion modelText

🎯 What it does: MOSAIC enhances long-context inference for diffusion LLMs through global dynamic memory management and lazy chunking

MoSE: Mixture of Slimmable Experts for Efficient and Adaptive Language Models

Nurbek Tastan (Mohamed bin Zayed University of Artificial Intelligence), Samuel Horváth (Mohamed bin Zayed University of Artificial Intelligence)

Computational EfficiencyRepresentation LearningTransformerLarge Language ModelMixture of ExpertsText

🎯 What it does: Propose Mixture of Slimmable Experts (MoSE), combining Mixture-of-Experts with Slimmable networks, enabling each expert to dynamically adjust its width during inference, thus achieving bi-axial variable computation.

MoshiRAG: Asynchronous Knowledge Retrieval for Full-Duplex Speech Language Models

Chung-Ming Chien (Toyota Technological Institute at Chicago), Alexandre Défossez (Kyutai)

RetrievalExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelTextRetrieval-Augmented GenerationAudio

🎯 What it does: Built Moshi A, a full-duplex speech-language model, incorporating asynchronous retrieval-augmented generation (RAG) functionality to enhance the factual accuracy of answers while maintaining real-time interaction.

MoSSP: A Momentum-Based Single-Loop Stochastic Penalty Method for Nonconvex Constrained DC-regularized Optimization

Luxuan Li (Beihang University), Xiao Wang (Sun Yat-sen University)

OptimizationTabular

🎯 What it does: Study the differential convex (DC) regularization optimization problem under non-convex constraints, propose a single-loop stochastic penalty framework called MoSSP, and provide two variants: Polyak momentum and recursive momentum;

MoST: Mixing Speech and Text with Modality-Aware Mixture of Experts

Yuxuan Lou (National University of Singapore), Yang You (National University of Singapore)

RecognitionGenerationData SynthesisRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningMixture of ExpertsTextMultimodalityAudio

🎯 What it does: Built the MoST model, which integrates speech and text into a unified large-scale language model.

MotiMotion: Motion-Controlled Video Generation with Visual Reasoning

Lee Hsin-Ying (University of California Merced), Zhixin Shu (Adobe Research)

GenerationData SynthesisTransformerPrompt EngineeringVision Language ModelDiffusion modelFlow-based ModelImageVideoTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposes a motion control video generation framework called MotiMotion based on visual reasoning, which can transform sparse trajectories and text prompts into realistic video dynamics.

Motion Attribution for Video Generation

Xindi Wu (NVIDIA), Jonathan Lorraine (NVIDIA)

GenerationExplainability and InterpretabilityComputational EfficiencyTransformerSupervised Fine-TuningDiffusion modelAuto EncoderOptical FlowVideo

🎯 What it does: Propose the MOTIVE framework, which identifies motion-related training segments in video generation models through gradient attribution analysis and uses these segments for targeted fine-tuning.

Motion Dynamics Learning for Few-Shot Embodied Adaptation

Sibo He (Xidian University), Yunsong Li (Xidian University)

Robotic IntelligenceMeta LearningReinforcement Learning from Human FeedbackTransformerVision Language ModelVision-Language-Action ModelDiffusion modelFlow-based ModelContrastive LearningVideoMultimodalitySequentialRetrieval-Augmented Generation

🎯 What it does: Designed the DynVLA system, achieving few-shot robot adaptation through trajectory-level motion dynamics modeling.

Motion Planning in Compressed Representation Spaces

Lukas Lao Beyer (Massachusetts Institute of Technology), Sertac Karaman (Massachusetts Institute of Technology)

Autonomous DrivingOptimizationRepresentation LearningTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningPoint CloudSequentialRetrieval-Augmented Generation

🎯 What it does: This paper proposes a general framework for motion planning in a compressed representation space. It first learns highly compressed discrete hierarchical tokens through an environment-conditioned autoencoder, and then performs decode-score search based on arbitrary objective functions in this latent space, achieving adaptive trajectory generation and controllable motion planning in multi-agent scenarios.

Motion-Aware Caching for Efficient Autoregressive Video Generation

Jing Xu (Xiamen University), Songwei Liu (ByteDance)

GenerationComputational EfficiencyTransformerDiffusion modelFlow-based ModelOptical FlowVideo

🎯 What it does: Proposes MotionCache, a motion-aware caching framework for autoregressive video generation models, which uses frame differences as motion features to dynamically decide whether to recalculate each token or directly use cached residuals, thereby significantly reducing the number of iterative denoising steps.

Motion-Residual Conflict-Aware Time Reversal for Generative Inbetweening

Zhenbang Zhang (Mohamed bin Zayed University of Artificial Intelligence), zhiqiang xu

GenerationData SynthesisTransformerDiffusion modelScore-based ModelOptical FlowVideo

🎯 What it does: Propose a framework called MR-CATR that performs motion residual conflict-aware alignment with time-reversed sampling during inference to improve the temporal consistency of generated intermediate frames.

MotionGRPO: Overcoming Low Intra-Group Diversity in GRPO-Based Egocentric Motion Recovery

Nanjie Yao (Hong Kong University of Science and Technology (Guangzhou)), Hao Wang (Hong Kong University of Science and Technology (Guangzhou))

Pose EstimationReinforcement Learning from Human FeedbackTransformerReinforcement LearningDiffusion modelScore-based ModelContrastive LearningVideoPoint CloudStochastic Differential Equation

🎯 What it does: Propose MOTIONGRPO, a framework for post-training RL on Diffusion models, used to recover high-fidelity full-body 3D motions from head trajectories.

MotionMAR: Multi-scale Auto-Regressive Human Motion Reconstruction from Sparse Observations

Yuhua Luo (Xiamen University), Cheng Wang (Xiamen University)

GenerationData SynthesisPose EstimationTransformerVision-Language-Action ModelDiffusion modelAuto EncoderContrastive LearningTime SeriesSequential

🎯 What it does: Proposed and implemented the MotionMAR framework, which reconstructs high-quality full-body motion from extremely sparse VR/AR tracking signals using a multi-scale autoregressive method.

Move-Then-Operate: Behavioral Phasing for Human-Like Robotic Manipulation

Haoming Xu (University of Chinese Academy of Sciences), Rui-Qi Wang (University of Science and Technology Beijing)

Robotic IntelligenceTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsVision Language ModelVision-Language-Action ModelFlow-based ModelVideoTextMultimodalityBenchmark

🎯 What it does: This paper proposes the Move-Then-Operate framework, which decomposes robotic grasping and detailed manipulation into two stages: movement and operation, and employs a stage-routed dual-expert strategy for control.

MoVie: Multimodal Video Compression with Text Guidance

Jiaqi Hu (Zhejiang University), Lianrui Mu (Zhejiang University)

CompressionConvolutional Neural NetworkTransformerVision Language ModelAuto EncoderContrastive LearningVideoTextMultimodality

🎯 What it does: Proposes a text-guided multi-modal video compression framework called MoVie, which can achieve high perceptual quality video reconstruction at low bitrates.

Moving Out: Physically-grounded Human-AI Collaboration

Xuhui Kang (University of Virginia), Yen-Ling Kuo (University of Virginia)

Robotic IntelligenceRecurrent Neural NetworkTransformerReinforcement LearningDiffusion modelAuto EncoderOptical FlowTextSequentialBenchmark

🎯 What it does: Proposes the Moving Out benchmark for studying human-robot collaboration under physical constraints.

MPFM: Cross Multi-Domain Prototype Flow Matching for Log Anomaly Detection

Jing Zhang (Shandong Normal University), Chao Luo (Shandong Normal University)

Domain AdaptationAnomaly DetectionTransformerFlow-based ModelAuto EncoderContrastive LearningTextOrdinary Differential Equation

🎯 What it does: Propose a cross-domain log anomaly detection framework MPFM, which identifies anomalies by utilizing shared-private prototypes and dual evidence from stream matching;

MRPO: Magnitude-Regularized Policy Optimization via L1 Constraints

Wei Han (Harbin Institute of Technology), Ting Liu (Harbin Institute of Technology)

OptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringContrastive LearningText

🎯 What it does: This paper proposes an MRPO method based on L1 norm constraints to address two major limitations of traditional KL constraints in RL training for large language models.

MSP: Probabilistically Consistent Multi-Scale Action Generation

Zhixuan Lin (Northeastern University), Fei Wang (Northeastern University)

GenerationRobotic IntelligenceTransformerDiffusion modelFlow-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningSequential

🎯 What it does: Propose a framework called MSP for multi-scale coarse-to-fine hierarchical action generation in a continuous latent space, to address the problems of long-term action continuity and cross-scale consistency.

MTNL: A Unified Modeling Perspective for Enhancing Tensor Network Learning

Junhua Zeng (Inner Mongolia University), Guoxu Zhou (Guangdong University of Technology)

RestorationComputational EfficiencyRepresentation LearningImageText

🎯 What it does: Propose a hybrid tensor network learning (MTNL) framework that unifies unsupervised and supervised learning in tensor networks;

MuCO: Generative Peptide Cyclization Empowered by Multi-stage Conformation Optimization

Yitian Wang (Renmin University of China), Hongteng Xu (Renmin University of China)

Drug DiscoveryProtein Structure PredictionGraph Neural NetworkTransformerDiffusion modelScore-based ModelFlow-based ModelBiomedical Data

🎯 What it does: Propose a multi-stage generative cyclic peptide conformation optimization method called MuCO, which can efficiently and accurately generate diverse low-energy cyclic peptide conformations.

MulFCoder: Framework-conditioned Multi-agent for MLLM-based Multi-framework Front-end Code Generation

Jie Wu (Communication University of China), Jiechao Gao (Stanford University)

AI Code AssistantTransformerLarge Language ModelAgentic AIPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose a multi-agent, framework-conditioned frontend code generation framework called MulFCoder, which addresses the cross-framework differences in generating executable React/Vue/Angular code from UI screenshots.

MuLoCo: Muon is a Practical Inner Optimizer for DiLoCo

Benjamin Thérien (FAIR at Meta), Eugene Belilovsky (Mila)

CompressionOptimizationFederated LearningComputational EfficiencyTransformerLarge Language ModelContrastive LearningText

🎯 What it does: Proposes MuLoCo, which uses Muon as the inner optimizer of DiLoCo to enhance the effectiveness of distributed training

Multi-Adapter Representation Interventions via Energy Calibration

Manjiang Yu (University of Queensland), Lijie Hu (Mohamed bin Zayed University of Artificial Intelligence)

Explainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelPrompt EngineeringText

🎯 What it does: Propose an input-adaptive representation intervention framework called MARI, which does not modify model weights, combining competitive multi-adapter and energy gating to precisely regulate the alignment behavior of LLMs.

Multi-agent imitation learning with function approximation: linear Markov games and beyond

Luca Viano (EPFL), Giorgia Ramponi (University of Zurich)

Reinforcement LearningTabularSequential

🎯 What it does: This paper provides the first theoretical analysis of multi-agent imitation learning (MAIL) in linear Markov games, and presents the corresponding sample complexity.

Multi-Agent Reinforcement Learning with Submodular Reward

Wenjing Chen (Texas A&M University), Victoria G. Crawford (Texas A&M University)

Reinforcement Learning

🎯 What it does: This paper studies cooperative multi-agent reinforcement learning (MARL), where the joint reward exhibits submodularity, a property that naturally captures the diminishing marginal returns when adding agents to a team. Unlike standard additive reward MARL, the submodular reward model better fits real-world scenarios, such as multi-drone surveillance and collaborative exploration.

Multi-Agent Teams Hold Experts Back

Aneesh Pappu (Stanford University), James Zou (Stanford University)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningAdversarial AttackHyperparameter SearchData-Centric LearningRobotic IntelligenceMeta LearningDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackNeural Architecture SearchProtein Structure PredictionTransformerLarge Language ModelAgentic AIPrompt EngineeringMixture of ExpertsTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: The study investigates whether self-organizing multi-agent LLM teams can leverage expert knowledge to achieve strong collaboration, finding that they often fail to match expert performance, even when explicitly pointing out the expert, as they are not good at utilizing the expert's knowledge.

Multi-Distribution Robust Conformal Prediction

Yuqi YANG, Ying Jin (University of Pennsylvania)

OptimizationFederated LearningData-Centric LearningImageTabularTime Series

🎯 What it does: This paper proposes a multi-source robust quantile prediction framework called MDCP, which constructs prediction sets that satisfy the pre-set information coverage on all sources by utilizing max-p aggregation and self-learning consistency scores.

Multi-Head Attention as a Source of Catastrophic Forgetting in MoE Transformers

Anrui Chen (Fudan University), Li Shang (Fudan University)

Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsContrastive LearningTextBenchmark

🎯 What it does: This paper explains the problem of catastrophic forgetting in MoE Transformer models during continual learning by analyzing the feature combination conflicts caused by the mixing of routing inputs in multi-head attention. It proposes an improved solution, MH-MoE, which performs head-level routing on each sub-representation to reduce combination collisions and improve memory retention.

Multi-Head LatentMoE and Head Parallel: Communication-Efficient and Deterministic MoE Parallelism

Chenwei Cui (Arizona State University), Hannah Kerner (Arizona State University)

Computational EfficiencyRepresentation LearningTransformerLarge Language ModelMixture of ExpertsText

🎯 What it does: This paper proposes two technologies, Multi-Head LatentMoE and Head Parallel (HP), to construct a new sparse Mixture of Experts training parallel architecture;

Multi-Integration of Labels Across Categories for Component Identification in Multi-trial Time Series

Noga Mudrik (Johns Hopkins University), Adam Shabti Charles

Explainability and InterpretabilityRepresentation LearningTime Series

🎯 What it does: Propose the MILCCI method to discover sparse and interpretable components in multi-experiment, multi-label time series, integrating label information and capturing variations between experiments and label effects;

Multi-Label Learning with Contrastive Cluster Self-Supervision for 3D Hierarchical Semantic Segmentation

Shuyu Cao (Southwest Jiaotong University), Na Zhao (Singapore University of Technology and Design)

SegmentationConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningPoint Cloud

🎯 What it does: Propose a multi-label learning framework called ML3DHS for 3D hierarchical semantic segmentation, addressing issues of multi-level conflicts and class imbalance.

Multi-Label Test-Time Adaptation with Bayesian Conditional Priors

QiRu Li, Qing Gu (Nanjing University)

ClassificationRecognitionDomain AdaptationVision Language ModelContrastive LearningImageMultimodality

🎯 What it does: Proposed a gradient-free test-time adaptation method called BCL (Bayesian Conditional Priors), which improves the performance of multi-label recognition under distribution drift by adding anchor conditional priors to the zero-shot logits of CLIP for correction.

Multi-Level Strategic Classification: Incentivizing Improvement through Promotion and Relegation Dynamics

Ziyuan Huang (University of Michigan), Mingyan Liu (University of Michigan)

ClassificationOptimizationReinforcement LearningTabularFinance Related

🎯 What it does: Propose a multi-level strategy classification framework, design a threshold sequence to encourage individuals to achieve long-term honesty rather than cheating; achieve arbitrary high attributes through theoretical proof and experimental validation;