ICML 2026 Papers — Page 34
International Conference on Machine Learning · 6554 papers
MADE: Benchmark Environments for Closed-Loop Materials Discovery
Shreshth A Malik, Yarin Gal
OptimizationDrug DiscoveryReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringDiffusion modelGraphTabularBenchmark
🎯 What it does: Proposed the MADE (Materials Discovery Environments) framework for evaluating the efficiency and effectiveness of closed-loop materials discovery pipelines under limited query budgets;
MAFE: Enabling Equitable Algorithm Design in Multi-Agent Multi-Stage Decision-Making Systems
Zachary McBride Lazri (University of Maryland), Min Wu (University of Maryland)
OptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyData-Centric LearningDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackNeural Architecture SearchProtein Structure PredictionGraph Neural NetworkTransformerReinforcement LearningDiffusion modelScore-based ModelFlow-based ModelRectified FlowNeural Radiance FieldAuto EncoderGenerative Adversarial NetworkContrastive LearningTabularTime SeriesBiomedical DataReview/Survey PaperBenchmarkFinance Related
🎯 What it does: This paper proposes the Multi-Agent Fair Environment (MAFE) framework and constructs scalable, real-data-based simulation environments in three major social domains—health, loans, and education—for long-term fairness research.
MAGIC: A Co-Evolving Attacker–Defender Adversarial Game for Robust LLM Safety
Xiaoyu Wen (Shanghai Jiao Tong University), Qiaosheng Zhang (Shanghai Artificial Intelligence Laboratory)
Safty and PrivacyTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkChain-of-Thought
🎯 What it does: Proposes MAGIC, a multi-round multi-agent reinforcement learning framework involving both attackers and defenders, treating the safety alignment of large language models as an asymmetric sequential game, supporting attackers to continuously rewrite prompts and approach defenders.
MAGIC: Multi-Granularity Language-Informed Image Clustering
Xiaohan Zhang (Nanjing University), Huaxiong Li (Nanjing University)
Image TranslationRetrievalRepresentation LearningTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodality
🎯 What it does: This paper proposes the MAGIC framework, which enhances unsupervised image clustering by utilizing multi-grained language descriptions;
Magnitude Distance: A Geometric Measure of Dataset Similarity
Sahel Torkamani (University of Edinburgh), Rik Sarkar (University of Edinburgh)
GenerationData SynthesisDiffusion modelScore-based ModelGenerative Adversarial NetworkContrastive LearningImage
🎯 What it does: Proposed a scale-adjustable magnitude distance metric to measure the geometric similarity between finite datasets
Making Expert Reasoning Learnable with Self-Distillation
Ethan Mendes (Georgia Institute of Technology), Alan Ritter (Georgia Institute of Technology)
Computational EfficiencyKnowledge DistillationRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringContrastive LearningTextChain-of-Thought
🎯 What it does: Proposes a two-step self-distillation method called DAIL, which enables large language models to learn reasoning abilities from a small number of high-quality expert answers.
Making Learner Weakness Actionable for Learning from Demonstration with Novice Teachers
Yuqing Zhu (King's College London), Matthew Howard (King's College London)
Robotic IntelligenceReinforcement Learning from Human FeedbackTransformerReinforcement LearningPrompt EngineeringVision-Language-Action ModelDiffusion modelContrastive LearningTabularTime SeriesSequential
🎯 What it does: Proposes the CLASP framework, which provides actionable teaching guidance for novice teachers in Learning from Demonstration (LfD) by constructing region maps based on teacher demonstrations and learner difficulties, helping them rapidly improve robot task performance within a limited demonstration budget.
Making Models Unmergeable via Scaling-Sensitive Loss Landscape
Minwoo Jang (POSTECH), Jungseul Ok (POSTECH)
Federated LearningSafty and PrivacyComputational EfficiencyKnowledge DistillationRepresentation LearningAdversarial AttackConvolutional Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageTextMultimodalityBenchmark
🎯 What it does: To address the governance gap caused by model merging, the authors design a training-time protection framework called TRAP 2, which enables the model to function normally at the authorized scale (s=1), but its performance drops sharply when it is illegally rescaled (s≠1) or merged, thus achieving model inmergeability.
MALICE: Memory-aware Loop Invariants Generation on Symbolic Execution Traces
Tong Chen (Shanghai Jiao Tong University), Qinxiang Cao (Shanghai Jiao Tong University)
OptimizationComputational EfficiencyAI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextBenchmarkChain-of-Thought
🎯 What it does: This study proposes the MALICE framework, which leverages symbolic execution traces to guide LLMs in automatically generating memory shape loop invariants, and improves accuracy through proxy iterative refinement.
MalTree: Tracing Malware Evolution using Embeddings at Scale
Akash Amalan (Delft University of Technology), Tom Julian Viering
Anomaly DetectionRepresentation LearningMixture of ExpertsContrastive LearningMultimodalityTabularTime Series
🎯 What it does: Propose the MalTree framework, which constructs and verifies the evolution tree of malware using multi-modal embeddings and large-scale phylogenetic methods.
MaMa: A Game-Theoretic Approach for Designing Safe Agentic Systems
Jonathan Nöther (Max Planck Institute for Software Systems), Goran Radanovic (Max Planck Institute for Software Systems)
OptimizationSafty and PrivacyAdversarial AttackReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringTextTabularTime SeriesSequentialReview/Survey PaperBenchmarkFinance Related
🎯 What it does: Design an automated method called MaMa based on Stackelberg security games to build secure multi-agent systems in the presence of agents under attack
MAMBO-G: Magnitude-Aware Mitigation for Boosted Guidance
Shangwen Zhu (Shanghai Jiao Tong University), Fan Cheng (Shanghai Jiao Tong University)
GenerationComputational EfficiencyTransformerDiffusion modelScore-based ModelImageVideoTextOrdinary Differential Equation
🎯 What it does: Proposes MAMBO-G, an untrained adaptive acceleration framework that dynamically adjusts the strength of classifier-free guidance (CFG).
MaMi-HOI: Harmonizing Global Kinematics and Local Geometry for Human-Object Interaction Generation
Hao Wang (South China University of Technology), Qi Liu (South China University of Technology)
GenerationData SynthesisPose EstimationTransformerVision-Language-Action ModelDiffusion modelScore-based ModelContrastive LearningImageVideoTextMultimodalityPoint CloudMesh
🎯 What it does: Propose the MaMi-HOI framework for generating human-robot interaction animations in 3D scenes that are both semantically intent-aligned and physically accurate in contact.
MAnchors: Memorization-Based Acceleration of Anchors via Rule Reuse and Transformation
Haonan Yu (Peking University), Xin Zhang (Peking University)
Explainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelImageTextTabular
🎯 What it does: Accelerate the Anchors explanation method by using memoization storage and rule transformation, significantly reducing computational costs.
Manifold-Aligned Guided Integrated Gradients for Reliable Feature Attribution
Soyeon Kim (South Korea Advanced Institute of Science and Technology), Jaesik Choi (South Korea Advanced Institute of Science and Technology)
ClassificationExplainability and InterpretabilityAuto EncoderImage
🎯 What it does: Propose an interpretation method called MA-GIG that constructs integral paths in the latent space of a variational autoencoder, addressing the problem of IG paths deviating from the data manifold.
Manifold-Aware Perturbations for Constrained Generative Modeling
Katherine Keegan (Emory University), Lars Ruthotto (Emory University)
GenerationData SynthesisOptimizationDiffusion modelScore-based ModelFlow-based ModelImagePoint CloudMeshGraphTabular
🎯 What it does: Propose a method to solve the numerical instability and distribution distortion caused by dimension mismatch in equality-constrained generative models, by perturbing the original distribution with noise in the regular normal direction and projecting it back onto the constraint manifold.
Manifold-Optimal Guidance: A Unified Riemannian Control View of Diffusion Guidance
Zexi Jia (WeChat AI, Tencent Inc.), Jie Zhou (WeChat AI, Tencent Inc.)
GenerationOptimizationComputational EfficiencyTransformerDiffusion modelScore-based ModelRectified FlowImageText
🎯 What it does: This paper proposes a Diffusion-guided framework based on manifold optimization (MOG), which improves the out-of-manifold drift problem caused by the Euclidean extrapolation of traditional CFG by performing natural gradient updates on the data manifold, and provides a closed-form solution that can be directly integrated into existing samplers.
ManifoldKV: Training-Free KV Cache Compression via Euclidean Outlier Detection
Debajyoti Datta (Hippocratic AI), Subhabrata Mukherjee (Hippocratic AI)
CompressionAnomaly DetectionComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation
🎯 What it does: Propose an untrained KV cache compression method called ManifoldKV, which scores tokens based on the Euclidean distance from key vectors to their mean, and selects important tokens via TopK; meanwhile, introduce windowed local mean (WindowedManifoldKV) to address the center dilution problem in long contexts.
ManiSoft: Towards Vision-Language Manipulation for Soft Continuum Robotics
Ziyu Wei (Beihang University), Si Liu (Beihang University)
Robotic IntelligenceTransformerReinforcement LearningPrompt EngineeringVision Language ModelVision-Language-Action ModelDiffusion modelImageVideoTextMultimodalityBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose the ManiSoft benchmark, aiming to evaluate the performance of soft continuum robot arms under visual-language manipulation, including a simulator, tasks, data generation pipeline, and expert trajectories;
Mantis: Lightweight Foundation Model for Time Series Classification
Vasilii Feofanov (Huawei Noah's Ark Lab), Ievgen Redko (Huawei Noah's Ark Lab)
ClassificationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerContrastive LearningTime Series
🎯 What it does: Proposed a lightweight time series classification foundation model called Mantis, and achieved zero-shot feature extraction;
Many Experiments, Few Repetitions, Unpaired Data, and Sparse Effects: Is Causal Inference Possible?
Felix Schur (ETH Zurich), Jonas Peters (ETH Zurich)
TabularReview/Survey Paper
🎯 What it does: Estimate causal effects under hidden confounding in an experimental setting with many conditions, few repetitions per condition, and unpaired data using an instrumental variable framework.
Many Needles in a Haystack: Active Hit Discovery for Perturbation Experiments
Andrea Rubbi (Wellcome Sanger Institute), Mohammad Lotfollahi
Drug DiscoveryBiomedical Data
🎯 What it does: This study proposes a new experimental design method called Probability-of-Hit, aiming to efficiently identify high-effect gene perturbations, especially under limited budget conditions.
Many-Shot CoT-ICL: Making In-Context Learning Truly Learn
Tsz Ting Chung (Hong Kong University Of Science And Technology), Dit-Yan Yeung (Hong Kong University Of Science And Technology)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningMeta LearningTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This paper studies the extension behavior of Many-Shot CoT-ICL in reasoning tasks and proposes an optimization method for example ordering.
MapDream: Task-Driven Map Learning for Vision-Language Navigation
Guoxin Lian (Renmin University of China), Zhaoxin Fan (Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing)
Autonomous DrivingRobotic IntelligenceTransformerSupervised Fine-TuningReinforcement LearningVision Language ModelVision-Language-Action ModelDiffusion modelSimultaneous Localization and MappingWorld ModelImageTextMultimodality
🎯 What it does: Propose the MapDream framework, which jointly learns map generation and navigation strategies using task-driven bird's-eye view (BEV) image generation;
MAPS: Memory-Aware Predictive Scheduling Framework for Large Language Model Serving
Tiancheng Zhang (Tianjin University), Xiaofei Wang (Tianjin University)
OptimizationComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsTextRetrieval-Augmented Generation
🎯 What it does: Proposes MAPS, a memory-aware prediction scheduling framework based on device-side inference prediction and uncertainty calibration, for split large language model services.
MapUQ: Map with Uncertainty Quantification for Robust BEV Vectorized Construction
Shaoyuan Mo (Chongqing University), Ke Wang (Chongqing University)
Autonomous DrivingExplainability and InterpretabilityComputational EfficiencyTransformerContrastive LearningSimultaneous Localization and MappingImageVideoPoint Cloud
🎯 What it does: Proposes the MapUQ framework, integrating uncertainty quantification into BEV vectorized map generation to enhance model robustness in complex traffic scenarios.
Margin-Adaptive Confidence Ranking for Reliable LLM Judgement
Gaojie Jin (University of Macau), Tianjin Huang (University of Exeter)
Explainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringContrastive LearningText
🎯 What it does: Provide reliable confidence estimates for LLMs as judges, learn a margin-based ranking network to replace traditional heuristic confidence signals, and propose an adaptive margin training scheme based on this.
MarketSim: Simulating Stock Markets with Large-Scale Generative Agents
Jinghua Piao (Tsinghua University), Yong Li (Tsinghua University)
TransformerLarge Language ModelAgentic AITextTabularTime SeriesFinance RelatedRetrieval-Augmented Generation
🎯 What it does: Built MarketSim — a stock market simulation framework based on large-scale generative agents, simulating over 15k institutional and background participants, executing at nanosecond levels in a Nasdaq-style continuous double auction (CDA) market, and reproducing real price dynamics and market structures.
Markov Chain Monte Carlo without Evaluating the Target: an Auxiliary Variable Approach
Wei Yuan (Rutgers University), Guanyang Wang (Rutgers University)
OptimizationComputational EfficiencyData-Centric LearningImageTabularStochastic Differential Equation
🎯 What it does: Proposed a MCMC framework based on auxiliary variables, unifying algorithms such as exchange, PoissonMH, TunaMH, and designed new mini-batch sampling methods incorporating gradient information under this framework (Poisson-Barker, Poisson-MALA, Tuna-SGLD).
Marrying Generative Model of Healthcare Events with Digital Twin of Social Determinants of Health for Disease Reasoning
Ziquan Wei (UNC Chapel Hill), Guorong Wu (UNC Chapel Hill)
Drug DiscoveryTransformerSupervised Fine-TuningDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkTabularTime SeriesSequentialBiomedical DataElectronic Health Records
🎯 What it does: A novel disease reasoning framework, DiffDT, was constructed, combining autoregressive (AR) models with conditional diffusion generative models. It utilizes social determinants of health (SDoH) proxies based on ICD codes to generate digital twins (DT) of multiple organs, incorporating physiological mechanisms when predicting future diseases.
MARS-SQL: A Multi-Agent Reinforcement Learning Framework For Text-To-SQL
Haolin Yang (Hong Kong University Of Science And Technology), Yi R. Fung (Hong Kong University Of Science And Technology)
AI Code AssistantTransformerLarge Language ModelReinforcement LearningAgentic AITextTabularChain-of-Thought
🎯 What it does: Proposed MARS-SQL, a multi-agent reinforcement learning framework that decomposes the Text-to-SQL task into three roles: schema grounding, query generation, and solution validation, and trains the generator through interactive RL.
MARS: Modular Agent with Reflective Search for Automated AI Research
Jiefeng Chen (Google Cloud AI Research), Jinsung Yoon (Google Cloud AI Research)
AI Code AssistantReinforcement Learning from Human FeedbackNeural Architecture SearchTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringMixture of ExpertsTextTabularBenchmark
🎯 What it does: Propose the MARS framework to address the bottleneck of machine learning engineering tasks in automated AI research, capable of generating maintainable multi-module code within a limited budget.
MAS-Architect: Declarative Multi-Agent System Design via Separation of Concerns
Jing Huang (Zhejiang University), Qiang Zhu (Zhejiang University)
Computational EfficiencyKnowledge DistillationAI Code AssistantTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose the MAS-Architect framework, which automatically generates multi-agent systems from task queries through a declarative MAS paradigm.
MAS-Orchestra: Understanding and Improving Multi-Agent Reasoning Through Holistic Orchestration and Controlled Benchmarks
Zixuan Ke (Salesforce Research), Shafiq Joty (Salesforce Research)
OptimizationTransformerReinforcement LearningAgentic AIPrompt EngineeringTextSequentialBenchmarkChain-of-Thought
🎯 What it does: Proposed the MAS-Orchestra framework, achieving global one-time orchestration of multi-agent systems through function-call-based reinforcement learning during training, and constructed MASBench for systematic comparison between MAS and single-agent systems.
MAS-ProVe: Understanding the Process Verification of Multi-Agent Systems
Vishal Venkataramani (Rutgers University), Shafiq Joty (Salesforce AI Research)
Autonomous DrivingOptimizationFederated LearningExplainability and InterpretabilityMeta LearningDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackNeural Architecture SearchTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextSequentialBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Conduct systematic experiments on process verification in multi-agent systems, proposing a pluggable MAS-ProVe framework to evaluate the effectiveness of verifiers under different granularities and context management strategies.
MASH: Modeling Abstention via Selective Help-Seeking
Mustafa Omer Gul (Cornell University), Tanya Goyal (Cornell University)
Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextRetrieval-Augmented Generation
🎯 What it does: Propose a training framework called MASH, which uses reinforcement learning to enable large language models to invoke retrieval tools only when necessary when answering questions, thus indirectly learning to identify their own knowledge boundaries and achieving abstention.
Masked Multi-path Contrast with Confidence-Gated Semantic Imputation for Incomplete Multi-view Clustering
Fan Yang (Nanjing University of Finance and Economics), Haikun Xu (Nanjing University of Finance and Economics)
Representation LearningTransformerAuto EncoderContrastive LearningImage
🎯 What it does: Propose the MAGIC framework, which addresses multi-view clustering with high missing rates. It first robustly learns semantic representations through multi-path contrastive consistency learning, and then infers missing views via confidence-gated semantic transmission.
Masks Can Be Distracting: On Context Comprehension in Diffusion Language Models
Julianna Piskorz (University of Cambridge), Christos Louizos (Qualcomm AI Research)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningData-Centric LearningTransformerSupervised Fine-TuningPrompt EngineeringDiffusion modelContrastive LearningText
🎯 What it does: This paper systematically evaluates the performance of Masked Diffusion Language Models (MDLMs) in contextual understanding, revealing their locality bias and the interference of the number of masks on performance, and proposes a mask-agnostic fine-tuning method to alleviate these issues.
MASPO: Joint Prompt Optimization for LLM-based Multi-Agent Systems
Zhexuan Wang (Harbin Institute of Technology), Min Zhang (Harbin Institute of Technology)
OptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextChain-of-Thought
🎯 What it does: Proposed a framework called MASPO for automatically jointly optimizing role prompts in large language model (LLM)-driven multi-agent systems (MAS), addressing the alignment of local goals with global system objectives and credit allocation issues.
MASPOB: Bandit-Based Prompt Optimization for Multi-Agent Systems with Graph Neural Networks
Zhi Hong (Chinese University of Hong Kong), Zhongxiang Dai (Chinese University of Hong Kong)
OptimizationAI Code AssistantGraph Neural NetworkTransformerReinforcement LearningPrompt EngineeringText
🎯 What it does: Optimize prompts for a confirmed multi-agent system (MAS), proposing a Bandit-based framework called MASPOB.
MAST: Motif-Augmented Diffusion with Search Tree for Spectroscopic Molecular Structure Elucidation
Chenghao Jia (Chinese Academy of Sciences), Xilin CHEN
GenerationData SynthesisDrug DiscoveryGraph Neural NetworkTransformerDiffusion modelGraphTabular
🎯 What it does: Proposed a 2D-3D molecular structure inference framework called MAST based on spectroscopic data, combining diffusion generation with tree search to achieve spectral consistency generation.
MatchFixAgent: Language-Agnostic Autonomous Repository-Level Code Translation Validation and Repair
Ali Reza Ibrahimzada (University of Illinois UrbanaChampaign), Daniel Kroening (Amazon)
AI Code AssistantTransformerLarge Language ModelAgentic AIPrompt EngineeringTextRetrieval-Augmented Generation
🎯 What it does: Developed a code translation verification and repair framework called MATCHFIXAGENT based on large language models and a multi-agent architecture, which can automatically verify and repair code translations between different programming languages at the repository level.
MathlibLemma: Folklore Lemma Generation and Benchmark for Formal Mathematics
Xinyu Liu (University of Virginia), Shangtong Zhang (University of Virginia)
AI Code AssistantTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposes MATHLIBLEMMA, a modular pipeline based on LLMs for automatically discovering, formalizing, and proving mathematical folklore lemmas, and constructs a corresponding benchmark dataset;
Matrix-Free GPU Semidefinite Programming for Quantum Ordered Search at the k=6 Frontier
Yancheng Wu (Shanghai Jiao Tong University), Yinyu Ye (Shanghai Institute for Mathematics and Interdisciplinary Sciences)
OptimizationPhysics Related
🎯 What it does: By constructing a matrix-free GPU solver, the memory bottleneck in the quantum ordered search problem with k=6 was overcome, successfully solving semi-definite programming (SDP) problems of scale N≈10⁵, and pushing the maximum feasible list size to between 90,000 and 94,000;
Matroid Algorithms Under Size-Sensitive Independence Oracles
Kiarash Banihashem (University of Maryland), Danny Mittal (University of Maryland)
OptimizationGraphReview/Survey Paper
🎯 What it does: Proposed a size-sensitive independence query model, and under this model, studied three major problems: basis, rank estimation, and partition size, providing almost tight upper and lower bounds. A sub-quadratic level maximum weight basis algorithm was presented for Matroid with bounded girth.
MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference
Raphaël Baur (ETH Zurich), Thomas Kleine Buening (ETH Zurich)
Reinforcement Learning from Human FeedbackTransformerReinforcement LearningAuto EncoderContrastive LearningTabularSequential
🎯 What it does: A shared reward function is constructed based on multiple human feedback types (demonstrations, comparisons, ratings, stopping), learned uniformly within a Bayesian inference framework.
Maximin Relative Improvement: Fair Learning as a Bargaining Problem
Jiwoo Han (University of Michigan), Yuekai Sun (University of Michigan)
ClassificationOptimizationFederated LearningExplainability and InterpretabilitySupervised Fine-TuningContrastive LearningTabularReview/Survey PaperBenchmark
🎯 What it does: This paper proposes a fair learning framework based on maximizing relative improvement (maximin relative improvement), treating the multi-group fairness problem as a cooperative bargaining problem;
Maximizing Mutual Information Between Prompt and Response Improves LLM Performance With No Additional Data
HyunJi Nam, Natasha Jaques (University of Washington)
OptimizationRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringContrastive LearningText
🎯 What it does: Propose a self-training method called Mutual Information Preference Optimization (MIPO), which generates positive and negative samples through adversarial data augmentation and directly maximizes the point-wise mutual information of the model under a reference LLM using Direct Preference Optimization (DPO), thereby enhancing the personalization and reasoning capabilities of LLMs without requiring additional human labels or external verifiers.
Maximum Likelihood Reinforcement Learning
Fahim Tajwar (Carnegie Mellon University), Andrea Zanette (Carnegie Mellon University)
OptimizationComputational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackLarge Language ModelReinforcement LearningPrompt EngineeringContrastive LearningImageTextTabularTime SeriesChain-of-Thought
🎯 What it does: Proposes the MaxRL framework, which treats reinforcement learning as an approximation of maximum likelihood in non-differentiable tasks with only binary success feedback. It achieves interpolation optimization of the likelihood objective through computable sampling order.
Maximum-Likelihood Learning of Latent Dynamics Without Reconstruction
Samo Hromadka (University College London), Maneesh Sahani (University College London)
RecognitionOptimizationRepresentation LearningScore-based ModelAuto EncoderContrastive LearningGaussian SplattingImageVideoTime Series
🎯 What it does: Propose the Recognition-Parametrized Gaussian State-Space Model (RP-GSSM), which eliminates the explicit decoder and uses maximum likelihood learning to infer latent dynamics, directly reasoning over the observation sequence;
MaxSAT-Based Compression for Tsetlin Machines
Stefan Szeider (TU Wien)
ClassificationCompressionKnowledge DistillationContrastive LearningTabular
🎯 What it does: This paper proposes a compression method based on maximum satisfiability (MaxSAT), which compresses a trained Tsetlin Machine (TM) from hundreds of clauses to just a few dozen, while maintaining prediction accuracy;
MC-HNN: Learning Latent Structural Semantics and High-Rank Representations for Hypergraph Neural Networks
Shuyang Fang (Xiamen University), Xiaoping Min (Xiamen University)
Representation LearningGraph Neural NetworkAgentic AIPrompt EngineeringMixture of ExpertsContrastive LearningGraphTabularBiomedical DataBenchmark
🎯 What it does: Proposed a multi-channel hypergraph neural network (MC-HNN), which enhances the rank and semantic independence of hypergraph representations through multi-channel information passing and implicit hyperedge type encoding.
MCCE: A Framework for Multi-LLM Collaborative Search in Discrete Spaces with Similarity-Filtered Preference Learning
Nian Ran (University of Manchester), Xiaoguang Zhao (Chinese Academy of Sciences)
OptimizationDrug DiscoveryTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningContrastive LearningTabularBiomedical DataBenchmark
🎯 What it does: Propose the Multi-LLM collaborative evolution framework MCCE, which jointly searches the discrete multi-objective optimization space using frozen large models and trainable small models.
MCP-Persona: Benchmarking LLM Agents on Real-World Personal Applications via Environment Simulation
WenHao Wang, Siheng Chen (Shanghai Jiao Tong University)
Autonomous DrivingFederated LearningSafty and PrivacyExplainability and InterpretabilityComputational EfficiencyDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackNeural Architecture SearchProtein Structure PredictionTransformerLarge Language ModelPrompt EngineeringWorld ModelTextSequentialBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose the MCP-Persona benchmark, simulating a realistic personalized MCP tool environment to evaluate the performance of LLM agents in social and collaborative applications
MDGMIX: Boundary-Aware Subgraph Mixing for Multi-Domain Graph Pre-Training
Ziyu Zheng (Xidian University), Xinyan Huang (Xidian University)
Domain AdaptationRepresentation LearningGraph Neural NetworkPrompt EngineeringContrastive LearningGraph
🎯 What it does: Propose the MDGMIX framework, which achieves multi-domain graph pre-training through boundary-aware subgraph mixing and hierarchical domain discrimination, significantly reducing data redundancy and improving cross-domain generalization.
MDN: Parallelizing Stepwise Momentum for Delta Linear Attention
Yulong Huang (Hong Kong University of Science and Technology (Guangzhou)), Bojun Cheng (Hong Kong University of Science and Technology (Guangzhou))
OptimizationComputational EfficiencyTransformerLarge Language ModelTextBenchmark
🎯 What it does: Propose Momentum DeltaNet (MDN), a model that parallelizes stepwise momentum within linear attention
MEAL: A Benchmark for Continual Multi-Agent Reinforcement Learning
Tristan Tomilin (Eindhoven University Of Technology), Meng Fang (University Of Liverpool)
TransformerReinforcement LearningContrastive LearningTabularSequentialBenchmark
🎯 What it does: This paper proposes the MEAL benchmark, providing a multi-agent continual reinforcement learning framework that can train 100 consecutive tasks on a single GPU, and implements efficient vectorized simulation and training using JAX.
Mean Flow Distillation: Robust and Stable Distillation for Flow Matching Models
An Zhao (Zhejiang University), Lingyun Sun (Zhejiang University)
GenerationOptimizationComputational EfficiencyKnowledge DistillationTransformerDiffusion modelScore-based ModelFlow-based ModelImageVideoTextOrdinary Differential Equation
🎯 What it does: Propose a Mean Flow Distillation (MFD) for Flow Matching models to achieve high-quality generation in a single step.
Mean Flow Policy Optimization
Xiaoyi Dong (Institute of Automation, Chinese Academy of Sciences), Jian Cheng (Institute of Automation, Chinese Academy of Sciences)
OptimizationReinforcement LearningScore-based ModelFlow-based ModelOptical FlowTabularTime SeriesSequential
🎯 What it does: This paper proposes a policy optimization method called MFPO based on the MeanFlow model, which uses the average flow to generate policies to achieve efficient exploration and learning.
Mean-Shift PCA by Knockoff Mean
Mengda Li (Chinese University of Hong Kong), Jianfeng Yao (Chinese University of Hong Kong)
Anomaly DetectionComputational EfficiencyRepresentation LearningData-Centric LearningTabularTime SeriesSequentialBenchmark
🎯 What it does: Propose a two-stage Mean-Shift PCA (MS-PCA) algorithm, which utilizes artificially added 'knockoff' mean perturbations and leverages random matrix theory (RMT) to detect and eliminate mean shift noise in high-dimensional data, ultimately recovering the principal components of the original data;
Measurement-Consistent Langevin Corrector for Stabilizing Latent Diffusion Inverse Problem Solvers
Lee Hyoseok (KAIST), Tae-Hyun Oh (KAIST)
RestorationSuper ResolutionDiffusion modelScore-based ModelImageStochastic Differential Equation
🎯 What it does: Proposed a Measurement Consistency Langevin Corrector (MCLC) to stabilize the dynamics of latent diffusion inverse problem solvers and improve reconstruction quality.
Measuring Agents in Production
Melissa Pan, Marquita Ellis (Ibm Research)
Autonomous DrivingOptimizationFederated LearningSafty and PrivacyExplainability and InterpretabilityComputational EfficiencyRobotic IntelligenceDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextReview/Survey Paper
🎯 What it does: This study systematically investigates LLM agents deployed in production environments, combining 20 in-depth interviews and questionnaires from 86 deployed systems, clarifying application motivations, model and architecture choices, evaluation methods, and deployment challenges.
Measuring and Mitigating Post-Hoc Rationalization in Reverse Chain-of-Thought Generation
Guangyue Peng (Peking University), Houfeng Wang (BOSS Zhipin)
Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This paper proposes a new evaluation framework to quantify the impact of pre-given answers on the generation of reasoning chains during reverse chain-of-thought reasoning (RCG), referred to as 'post-rationalization.' Based on this, we introduce Structure-Skeleton Guided Reasoning (SSR) and its distillation version SSR-D to alleviate answer anchoring, thereby improving reasoning quality and generalization ability.
Measuring Intent Comprehension in LLMs
Nadav Kunievsky (University of Chicago), James Evans (University of Chicago)
GenerationExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: A framework is proposed to evaluate the intent understanding capability of large language models (LLMs). The model output is decomposed into intent sensitivity (IS), expression sensitivity (AS), and model uncertainty (MU) through variance decomposition, thereby quantifying the model's robustness under different prompt languages/expressions.
Measuring Meta-Cultural Competency: A Spectral Framework for LLM Knowledge Structures
Sougata Saha (Mohamed bin Zayed University of Artificial Intelligence), Monojit Choudhury (Mohamed bin Zayed University of Artificial Intelligence)
Recommendation SystemTransformerLarge Language ModelPrompt EngineeringContrastive LearningText
🎯 What it does: Propose a macro-structural evaluation framework based on spectral analysis to assess the macro-structural understanding of large language models (LLMs) regarding cross-cultural knowledge;
MEC: Machine-Learning-Assisted Generalized Entropy Calibration for Semi-Supervised Mean Estimation
Se Yoon Lee (Texas AandM University), Jae Kwang Kim (Iowa State University)
OptimizationRepresentation LearningData-Centric LearningContrastive LearningTabular
🎯 What it does: Proposed a machine learning-assisted generalized entropy calibration (MEC) method to improve prediction-driven inference (PPI) in semi-supervised mean estimation.
MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks
Yadong Niu (Xiaomi Inc), Jian Luan (Xiaomi Inc)
TransformerLarge Language ModelMixture of ExpertsTextBenchmarkChain-of-ThoughtAudio
🎯 What it does: Introduces the MECAT benchmark, providing multi-expert-constructed multi-perspective audio descriptions and QA data, and proposes the DATE evaluation metric.
Mechanisms of Introspective Awareness
Uzay Macar (Anthropic Fellows), Jack Lindsey (Anthropic)
Anomaly DetectionExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringText
🎯 What it does: Studied the introspection capability of large language models when injecting steering vectors, and located their internal implementation through various analytical mechanisms.
Mechanistic Anomaly Detection via Functional Attribution
Hugo Lyons Keenan (University of Melbourne), Sarah Monazam Erfani (University of Melbourne)
Anomaly DetectionExplainability and InterpretabilityTransformerContrastive LearningImageText
🎯 What it does: Propose a mechanism anomaly detection (MAD) method based on functional attribution to identify whether the model uses abnormal internal mechanisms during testing.
Mechanistic Data Attribution: Tracing the Training Origins of Interpretable LLM Units
Jianhui Chen (Peking University), Liangming Pan (Peking University)
Explainability and InterpretabilityTransformerLarge Language ModelText
🎯 What it does: Proposed and implemented the 'Mechanism Data Attributions' (MDA) framework, which uses influence functions to trace and quantify the training sources of interpretable units within large language models (LLMs), and verified their causal effects by deleting or enhancing high-impact samples.
Mechanistic Interpretability as Statistical Estimation: A Variance Analysis
Maxime Méloux (Université Grenoble Alpes), Maxime Peyrard (Université Grenoble Alpes)
Explainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringText
🎯 What it does: This paper systematically analyzes the circuit discovery process in mechanism interpretability (MI) from a statistical estimation perspective, revealing the high variance and instability of CMA and its approximation methods, and evaluates the impact of data resampling, hyperparameters, and counterfactual definitions on circuit structures.
MechVQA: Benchmarking and Enhancing Multimodal LLMs on Comprehensive Mechanical Drawing Understanding
Qian Kou (Beijing Academy of Artificial Intelligence), Cao Dongxing (Beijing University of Technology)
RecognitionImage TranslationRestorationAnomaly DetectionReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision Language ModelImageTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose the MechVQA mechanical drawing understanding benchmark and the MechVL model, enhancing the recognition, reasoning, and judgment capabilities of multi-modal LLMs on mechanical drawings.
Med-Scout: Curing MLLMs' Geometric Blindness in Medical Perception via Geometry-Aware RL Post-Training
Anglin Liu (Hong Kong University of Science and Technology), Jintai Chen (Hong Kong University of Science and Technology)
Data-Centric LearningTransformerLarge Language ModelReinforcement LearningVision Language ModelImageMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyBenchmarkChain-of-Thought
🎯 What it does: Construct three categories of geometric proxy tasks and enhance the geometric perception of medical multimodal LLM through reinforcement learning post-training
Med-SegLens: Latent-Level Model Diffing for Interpretable Medical Image Segmentation
Salma J. Ahmed (Wilfrid Laurier University), Azam Asilian Bidgoli (Wilfrid Laurier University)
SegmentationDomain AdaptationExplainability and InterpretabilityTransformerDiffusion modelAuto EncoderContrastive LearningImageBiomedical DataMagnetic Resonance Imaging
🎯 What it does: This paper proposes the Med‑SegLens framework, which interprets the intermediate activations of medical image segmentation models using sparse autoencoders and performs model diffing across different datasets to visualize, diagnose, and correct segmentation failures caused by dataset shifts.
MEDA: Medical-Oriented Activation Editing for Hallucination Mitigation in Medical Large Vision-Language Model
Tianbo Wang (Beihang University), Xianglong Liu (Beihang University)
Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataElectronic Health RecordsRetrieval-Augmented Generation
🎯 What it does: This paper proposes a medical large audio-visual model hallucination mitigation method called MEDA, which dynamically guides the model to acquire medical knowledge during inference using activation editing technology, thereby significantly reducing the hallucination rate.
MedCoG: Maximizing LLM Inference Density in Medical Reasoning via Meta-Cognitive Regulation
Yu Zhao (Nankai University), Dacheng Tao (Nanyang Technological University)
Explainability and InterpretabilityComputational EfficiencyDrug DiscoveryTransformerLarge Language ModelPrompt EngineeringTextBiomedical DataBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposes the MedCoG framework, which dynamically schedules procedural, situational, and factual knowledge by leveraging the metacognitive evaluation of LLMs (complexity, familiarity, knowledge density) to achieve more efficient medical reasoning.
MedCRP-CL: Continual Medical Image Segmentation via Bayesian Nonparametric Semantic Modality Discovery
Ziyuan Gao (University College London)
SegmentationTransformerPrompt EngineeringMixture of ExpertsVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasound
🎯 What it does: Propose MedCRP-CL, a framework for online task structure discovery in medical image segmentation and achieve continual learning without replay.
MedMamba: Multi-View State Space Models with Adaptive Graph Learning for Medical Time Series Classification
Da Zhang (Northwest Polytechnical University), Xuelong Li (TeleAI, China Telecom)
ClassificationGraph Neural NetworkTransformerAuto EncoderContrastive LearningTime SeriesBiomedical DataElectrocardiogramReview/Survey PaperBenchmark
🎯 What it does: Propose the MedMamba structure, integrating multi-scale convolutional embeddings, a three-branch differential state space encoder, and adaptive spatial graph Mamba, to achieve efficient classification of medical time series.
MedMosaic: A Challenging Large Scale Benchmark of Diverse Medical Audio
Harshit Rajgarhia (Centific Global Solutions Inc), Prasanna Desikan (Centific Global Solutions Inc)
TransformerLarge Language ModelPrompt EngineeringTextMultimodalityElectronic Health RecordsBenchmarkRetrieval-Augmented GenerationAudio
🎯 What it does: Constructed MedMosaic, a multimodal medical audio question-answer dataset containing 46,701 pairs, covering various audio types such as heart and lung sounds, dialogues, and synthetic audio.
MedREK: Retrieval-Based Editing for Medical LLMs with Key-Aware Prompts
Shujun Xia (Columbia University), Quanzheng Li (Harvard Medical School)
RetrievalKnowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBiomedical DataBenchmarkRetrieval-Augmented Generation
🎯 What it does: Constructed the MedVersa medical knowledge editing benchmark and proposed the MedREK retrieval-based editing framework for batch updating of factual knowledge in medical LLMs.
MedScope: Incentivizing "Think with Videos" for Clinical Reasoning via Coarse-to-Fine Tool Calling
Wenjie Li (Shanghai Jiao Tong University School of Medicine), Yankai Jiang (Chinese University of Hong Kong)
Explainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision Language ModelVision-Language-Action ModelVideoTextMultimodalityBiomedical DataBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose a clinical video reasoning model called MedScope based on tool calling, which achieves hierarchical retrieval and verification of long videos through a coarse-to-fine visual chain of thought (VCoT).
MedSIGHT: Towards Grounded Visual Comprehension in Medical Large Vision-Language Models
Aofei Chang (Pennsylvania State University), Cao Xiao (GE Healthcare)
RecognitionSegmentationTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelContrastive LearningImageTextBiomedical DataBenchmark
🎯 What it does: Propose MEDSIGHT, a unified medical large vision-language model that can perform visual understanding and pixel-level segmentation of medical images within the same framework.
MEDUSA: Motion Elimination in Diffusion Using Spectral Attack
Hongwei Yu (University of Science and Technology Beijing), Jiansheng Chen (University of Science and Technology Beijing)
GenerationAdversarial AttackTransformerDiffusion modelOptical FlowImageVideo
🎯 What it does: A spectral attack that applies nuclear norm optimization on the time attention matrix of video diffusion models generates adversarial perturbations to freeze the video generation process and eliminate motion semantics.
Meerkat-VL: Implicit Risk Safety Alignment in Multimodal LLMs via Perceptual Reasoning and Self-Verification
Peicheng Zhou (University of Science and Technology of China), Hongtao Xie (University of Science and Technology of China)
Safty and PrivacyReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision Language ModelVision-Language-Action ModelImageTextMultimodality
🎯 What it does: Propose an implicit risk safety alignment framework called Meerkat-VL for multimodal large language models, enhancing the model's ability to perceive and respond safely to implicit risks through perceptual reasoning, model self-verification, and dual-objective consistency alignment.
MEG-XL: Data-Efficient Brain-to-Text via Long-Context Pre-Training
Dulhan Jayalath (University of Oxford), Oiwi Parker Jones (University of Oxford)
Representation LearningData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Propose MEG-XL, a long-context (2.5 minutes) self-supervised pretraining framework for brain-text decoding, followed by fine-tuning on a small amount of labeled data.
Mem-T: Densifying Rewards for Long-Horizon Memory Agents
Yanwei Yue (Peking University), Yan Zhang (Peking University)
OptimizationComputational EfficiencyTransformerLarge Language ModelReinforcement LearningTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Developed the Mem‑T memory agent and the MOT‑GRPO training framework to achieve end-to-end optimization for long-term memory management.
Membership Inference Attacks for Unseen Classes
Pratiksha Thaker (Carnegie Mellon University), Virginia Smith (Carnegie Mellon University)
Safty and PrivacyAdversarial AttackConvolutional Neural NetworkTransformerContrastive LearningImageTextTabular
🎯 What it does: Propose a membership inference attack in the scenario of 'unseen classes' where target class samples are unavailable, and systematically evaluate attack methods under this setting.
MemCast: Memory-Driven Time Series Forecasting with Experience-Conditioned Reasoning
Xiaoyu Tao (University of Science and Technology of China), Shijin Wang (iFLYTEK Research)
TransformerLarge Language ModelPrompt EngineeringTime SeriesFinance RelatedRetrieval-Augmented Generation
🎯 What it does: Propose the MemCast framework, which redefines time series forecasting as an experience-conditioned reasoning task, guiding LLM reasoning through text-based multi-level memory (historical patterns, reasoning wisdom, and general rules).
MemEvolve: Meta-Evolution of Agent Memory Systems
Guibin Zhang (National University of Singapore), Shuicheng YAN
OptimizationMeta LearningNeural Architecture SearchTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringMixture of ExpertsTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose MemEvolve: a dual-layer evolutionary framework that can simultaneously evolve the experiential knowledge of LLM agents and the architecture of their memory systems, and provide EvolveLab, a unified code library that modularly implements 12 representative memory systems.
MemIncept: Steering LLM Agents via Cooperative Stealthy Memory Injections
Nan Yan (Northwestern University), Jiarong Xing (Rice University)
Safty and PrivacyAdversarial AttackReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextTabularBiomedical DataBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose MemIncept, a method that manipulates LLM agent behavior by generating collaborative and stealthy memory injection records.
MEMO: Memory-Augmented Model Context Optimization for Robust Multi-Turn Multi-Agent LLM Games
Yunfei Xie (Rice University), Zhangyang Wang (University of Texas at Austin)
Autonomous DrivingOptimizationFederated LearningRobotic IntelligenceMeta LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringWorld ModelTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: By self-play and memory enhancement, the reasoning context of LLMs in multi-round multi-agent text games is optimized, significantly improving the win rate and reducing runtime variance.
MemOCR: Layout-Aware Visual Memory for Efficient Long-Horizon Reasoning
Yaorui Shi (University of Science and Technology of China), An Zhang (University of Science and Technology of China)
Computational EfficiencyRepresentation LearningTransformerReinforcement LearningPrompt EngineeringVision Language ModelTextMultimodalityRetrieval-Augmented Generation
🎯 What it does: Propose MemOCR, a 2D memory management method that leverages visual layout to achieve efficient context compression for long-sequence reasoning.
Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity
Menglin Xia (Microsoft), Saravan Rajmohan (Microsoft)
RetrievalComputational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextSequentialBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose the MEMORA harmonic memory architecture, which balances the abstraction and concreteness of memory by utilizing main abstraction, prompt anchor points, and policy-based retrieval.
Memory as a Markov Matrix: Sample Efficient Knowledge Expansion via Token-to-Dictionary Mapping
Kaustubh Pethkar (New Jersey Institute of Technology), Yingcong Li (New Jersey Institute of Technology)
Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelText
🎯 What it does: Model the memory of large language models as a Markov matrix, and propose to achieve knowledge expansion while avoiding catastrophic forgetting through a token-to-dictionary mapping approach.
Memory as Dynamics: Learning Reliability-Guided Predictive Models for Online Video Perception
Minwoo Kim (Kookmin University), Sang Min Yoon (Kookmin University)
Object TrackingSegmentationTransformerPrompt EngineeringVision-Language-Action ModelDiffusion modelVideo
🎯 What it does: Built a reliability-guided predictive memory framework called RPM for online video perception.
Memory Caching: RNNs with Growing Memory
Ali Behrouz (Google Research), Vahab Mirrokni (Google Research)
RetrievalComputational EfficiencyRepresentation LearningRecurrent Neural NetworkTransformerLarge Language ModelMixture of ExpertsTextRetrieval-Augmented Generation
🎯 What it does: Propose the Memory Caching (MC) technique, allowing the memory capacity of recurrent neural networks to grow with the sequence length, and implement four caching strategies (Residual, Gated Residual, Memory Soup, Sparse Selective).
Memory is Reconstructed, Not Retrieved: Graph Memory for LLM Agents
Shuo Ji (National University of Singapore), Bryan Hooi (National University of Singapore)
RetrievalComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerLarge Language ModelAgentic AIPrompt EngineeringTextGraphBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposed MRAgent, which utilizes a Cue-Tag-Content graph structure based on associated labels and actively reconstructs external memory in multiple steps during inference.
Memory Savings at What Cost? A Study of Alternatives to Backpropagation
Kunjal Panchal (University of Massachusetts), Hui Guan (University of Massachusetts)
OptimizationComputational EfficiencyTransformerLarge Language ModelVision Language ModelAuto EncoderTextMultimodality
🎯 What it does: Compare BP, FMAD, ZO, and activation checkpointing in terms of memory, accuracy, convergence speed, and computational cost during LLM and vision-language model training, providing a unified theoretical and experimental analysis.
Memory-Distilled Selection for Noise-Robust Anomaly Detection
Sirojbek Safarov (AIVEX Inc), Octavia Camps (Northeastern University)
Anomaly DetectionKnowledge DistillationTransformerAuto EncoderContrastive LearningImageBenchmark
🎯 What it does: Propose a training framework named Memory‑Distilled Selection (MeDS), which achieves robust image-level and pixel-level anomaly detection in unsupervised defect detection scenarios with noisy (contaminated) samples.
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design
Athanasios Glentis (University of Minnesota), Mingyi Hong (University of Minnesota)
OptimizationComputational EfficiencyTransformerLarge Language ModelText
🎯 What it does: This paper proposes SCALE, an optimizer with extremely low memory usage, specifically designed for the pretraining of large language models;