arXivSub Start free trial

ICML 2026 Papers with Code β€” Page 6

International Conference on Machine Learning Β· 1032 papers

Learn-to-learn on Arbitrary Textual Conditioning: A Hypernetwork-Driven Meta-gated LLM

Luo Ji (Geely AI Lab, Geely Auto Group), Hongyan Li (Geely AI Lab, Geely Auto Group)

CodeMeta LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringText

🎯 What it does: Propose MeGan, which utilizes a hypernetwork to generate adaptive β-SwiGLU gating, injecting meta-learning control signals into the FFN of LLMs, enabling fast adaptation under any text conditions.

Learnability-Informed Fine-Tuning of Diffusion Language Models

Shubham Parashar (Texas A&M University), Shuiwang Ji (Texas A&M University)

CodeComputational EfficiencyRepresentation LearningData-Centric LearningTransformerSupervised Fine-TuningDiffusion modelAuto EncoderText

🎯 What it does: Propose a learnable mask fine-tuning method called LIFT, which adaptively masks training for token learning difficulty and timing in diffusion language models.

Learning Compressed Shape-Aware Molecular Representations for Virtual Screening

Robin Winter (Pfizer), Djork-ArnΓ© Clevert (Pfizer)

CodeRetrievalCompressionRepresentation LearningDrug DiscoveryGraph Neural NetworkAuto EncoderContrastive LearningGraphBiomedical Data

🎯 What it does: Utilize the SAND framework to learn compressed shape-aware representations based on 2D molecular graphs, enabling shape similarity retrieval from a 1-billion-level molecular library;

Learning Flexible Generalization in Video Quality Assessment by Bringing Device and Viewing Condition Distributions

Nikolay Safonov (MSU Institute for Artificial Intelligence), Dmitriy S. Vatolin

CodeDomain AdaptationComputational EfficiencyData-Centric LearningConvolutional Neural NetworkTransformerSupervised Fine-TuningAuto EncoderContrastive LearningImageVideo

🎯 What it does: Constructed a subjective comparative evaluation dataset covering more than 300 Android devices, containing display parameters and environmental information, and used the Blade-Chest model to aggregate preference votes. Subsequently, a lightweight conditional adaptation network was trained to improve the prediction of existing VQA metrics across different devices and environments.

Learning High-Frequency Continuous Action Chunks in Latent Space

Kunyun Wang (Shanghai Jiao Tong University), Wenchao Ding (Fudan University)

CodeRobotic IntelligenceTransformerReinforcement LearningFlow-based ModelRectified FlowAuto EncoderTime SeriesSequential

🎯 What it does: In this study, the authors propose using a variational autoencoder (VAE) in the latent space to learn high-frequency continuous action blocks, and design a 'Reuse-then-Refine' (RTR) strategy to maintain continuity between action blocks under asynchronous inference, thereby enabling smooth and continuous execution by robots under high-frequency control.

Learning Long Range Spatio-Temporal Representations over Continuous Time Dynamic Graphs with State Space Models

Ayushman Raghuvanshi (Indian Institute of Science), Mahesh Chandran (Fujitsu Research India)

CodeRepresentation LearningGraph Neural NetworkTransformerGraphTime SeriesSequential

🎯 What it does: Propose a state-space model based on continuous-time dynamic graphs (CTDG) (CTDG-SSM), which utilizes topology-aware HiPPO (CTT-HiPPO) together with graph Laplacian polynomial filters to achieve efficient memory and update of long-term and multi-hop spatial information.

Learning Protein Structure-Function Relationships through Knowledge-guided Representation Decomposition

Mingqing Wang (Tsinghua University), Zhixiang Ren (Pengcheng Laboratory)

CodeExplainability and InterpretabilityRepresentation LearningProtein Structure PredictionTransformerAuto EncoderContrastive LearningBiomedical Data

🎯 What it does: Propose the ProtDiS framework, which utilizes knowledge-guided representation decomposition to split pre-trained protein microenvironment embeddings into independent channels aligned with biophysical attributes, thereby improving the interpretability and predictive performance of structure-function relationships.

Learning Query-Aware Budget-Tier Routing for Runtime Agent Memory

Haozhen Zhang (Nanyang Technological University), Wenya Wang (Nanyang Technological University)

CodeComputational EfficiencyTransformerLarge Language ModelReinforcement LearningAgentic AITextRetrieval-Augmented Generation

🎯 What it does: Propose BudgetMem, a modular framework that decomposes runtime memory extraction into adjustable budget levels, and learns a lightweight router to make decisions across different budget levels, achieving controllable balance between memory extraction cost and quality in LLM agents.

Learning Randomized Reductions

Ferhat Erata (Yale University), Ruzica Piskac (Yale University)

CodeOptimizationData-Centric LearningAI Code AssistantLarge Language ModelSupervised Fine-TuningAgentic AIPrompt EngineeringTextTabularBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposes a framework for automatically learning random self-reductions (RSR) from known programs and implements a system called Bitween.

Learning Rewrite-Invariant Reasoning with Targeted Alternation Training

Mousa Arraf (Technion Israel Institute of Technology), Kira Radinsky

CodeExplainability and InterpretabilityComputational EfficiencyData-Centric LearningGraph Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextChain-of-Thought

🎯 What it does: By sampling and aggregating multiple reasoning trajectories of large language models on semantically preserving rewriting tasks, we construct a reasoning graph for each problem; using this graph, we identify the 'Solution Boundary Cut' (SBC), which marks the transition from recoverable to unrecoverable states, and generate a small number of targeted positive and negative example pairs for model fine-tuning or in-context learning, thereby improving the model's reasoning robustness under semantically preserving rewriting.

Learning syntax without semantics: Disentangled tiny language models

Ezra Winston (Carnegie Melon University), J Zico Kolter

CodeComputational EfficiencyKnowledge DistillationRepresentation LearningData-Centric LearningTransformerLarge Language ModelPrompt EngineeringText

🎯 What it does: By training extremely small language models with constrained paraphrasing (SAMBAL) on text, they learn syntactic structures while suppressing the influence of semantics and world knowledge.

Learning Taxonomic Trees with Hierarchical Representation Regularization for Large Multimodal Models

Hulingxiao He (Peking University), Yuxin Peng (Peking University)

CodeRecognitionRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningVision Language ModelContrastive LearningImageTextMultimodality

🎯 What it does: Propose HiR 2, a parameter-free hierarchical representation regularization method, which utilizes non-parametric cross-attention to construct semantic visual trees from intermediate layers of LMMs, and enhances hierarchical visual recognition (HVR) consistency through hyperbolic implication loss and spherical angular dispersion loss.

Learning Tight Rejection Boundaries without Negatives for Strict One-Class Audio Deepfake Detection

Yuze Zhao (Harbin Institute of Technology), Wei Jiang (Harbin Institute of Technology)

CodeAnomaly DetectionTransformerAuto EncoderContrastive LearningAudio

🎯 What it does: Propose a strictly one-class learning audio deepfake detection method called CA-SOADD, which utilizes the distribution shift view without negative samples as a boundary detector, and achieves a clear delineation of the compact distribution of real audio and the rejection boundary through triple-objective central anchor point learning.

Learning to Rank by Directly Optimizing Full-Order Probabilities

Yongxiang Tang (Kuaishou Technology), Peng Jiang (Kuaishou Technology)

CodeRecommendation SystemOptimizationComputational EfficiencyScore-based ModelContrastive LearningTextTabularBenchmark

🎯 What it does: This paper proposes Full-Order Bound (FOB), which approximates and directly optimizes the probability of complete ranking events by introducing segmented thresholds to the latent Gaussian scores.

Learning to Refine: Spectral-Decoupled Iterative Refinement Framework for Precipitation Nowcasting

Yunlong Zhou (Nanjing University), Xiaotong Yuan

CodeOptimizationRecurrent Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningVideoTime SeriesPhysics Related

🎯 What it does: A deterministic, spectrum-decoupled iterative refinement framework called SDIR is proposed for high-resolution precipitation nowcasting.

Learning to Route Languages for Multilingual Policy Optimization

Geyang Guo (Georgia Institute of Technology), Wei Xu (Georgia Institute of Technology)

CodeOptimizationTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsTextMultimodalityBenchmark

🎯 What it does: This study proposes a language routing-based multilingual policy optimization framework called LRPO, which allows the model to actively select the response language during training and generate diverse training signals through multilingual rollouts.

Learning to Share: Selective Memory for Efficient Parallel Agentic Systems

Joseph Fioresi (University of Central Florida), Mubarak Shah (University of Central Florida)

CodeComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningAgentic AITextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposed and implemented Learning to Share (LTS), a method that uses a global shared memory and a learnable controller in parallel agent systems to selectively share intermediate results, thereby reducing redundant computations and improving execution efficiency.

Learning with Admissibility: Robust Fuzzy Hashing for Cross-Modal Retrieval with Noisy Labels

Xincheng Sun (Sichuan University), Yuan Sun (Sichuan University)

CodeRetrievalContrastive LearningMultimodality

🎯 What it does: Proposes a Robust Fuzzy Cross-Modal Hashing (RFCMH) framework based on fuzzy set theory to address the label noise problem in cross-modal retrieval;

Learning-Augmented Scalable Linear Assignment Problem Optimization via Neural Dual Warm-Starts

Ilay Yavlovich (Technion Israel Institute of Technology), Jose Yallouz (Technion Israel Institute of Technology)

CodeOptimizationGraph Neural NetworkTransformerReinforcement LearningAuto EncoderContrastive LearningGraphTabular

🎯 What it does: Use deep learning models to predict dual variables for linear assignment problems to warm start traditional exact solvers.

LEGO: An LLM-Enabled Hierarchical Optimizer for Tensor Computation Graphs with Structure-Aware Search and Compositional Synthesis

Ruiyuan Xu (Chinese Academy of Sciences), Huimin Cui (Chinese Academy of Sciences)

CodeOptimizationComputational EfficiencyAI Code AssistantTransformerLarge Language ModelAgentic AIPrompt EngineeringMixture of ExpertsImageTextGraph

🎯 What it does: Construct a hierarchical optimization framework called LEGO based on large language models, used to automatically synthesize and optimize end-to-end tensor computation graphs on GPUs.

Less Is More: Fast and Accurate Reasoning with Cross-Head Unified Sparse Attention

Lijie Yang (Princeton University), Ravi Netravali (Princeton University)

CodeComputational EfficiencyAI Code AssistantTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsText

🎯 What it does: Proposes LessIsMore, a training-agnostic sparse attention mechanism that improves decoding efficiency in long reasoning models by leveraging cross-head unified token selection and stable nearest window.

Less is More: Geometric Unlearning for LLMs with Minimal Data Disclosure

Chenchen Tan (Monash University), Longxiang Gao (Qilu University of Technology)

CodeSafty and PrivacyExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBenchmark

🎯 What it does: Propose a geometry-based LLM unlearning method (Geometric Unlearning), which eliminates specific knowledge by projecting and aligning target entities in the prompt-conditioned hidden state space, without needing access to the original training corpus.

Less Precise Can Be More Reliable: A Systematic Evaluation of Quantization’s Impact on VLMs Beyond Accuracy

Aymen Bouguerra (UniversitΒ΄ e Paris-Saclay), Fabio Arnez (UniversitΒ΄ e Paris-Saclay)

CodeClassificationCompressionExplainability and InterpretabilityComputational EfficiencyRepresentation LearningContrastive LearningImageTextMultimodalityBenchmark

🎯 What it does: This paper conducts a systematic quantitative evaluation of Vision-Language Models (VLMs), designing over 700k evaluation runs covering five reliability dimensions (robustness, calibration, OOD detection, distribution drift, and pulse correlation), and analyzes the impact of different quantization strategies on these metrics.

Less Token, More Signal: MoE Expert Pruning via Critical Token Selection

Zeliang Zong (Hikvision Research Institute), Jilin Hu (East China Normal University)

CodeComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelMixture of ExpertsText

🎯 What it does: Propose a Mixture-of-Experts (MoE) expert pruning method called STEP based on key token selection, which uses attention-guided token importance to filter noise, and combines dual-factor expert importance scoring and expert-to-bias knowledge preservation techniques to achieve efficient compression;

Leveraging Gauge Freedom for Learning Non-Gradient Population Dynamics of Stochastic Systems

Jules Berman (New York University), Benjamin Peherstorfer (New York University)

CodeDiffusion modelScore-based ModelAuto EncoderContrastive LearningPoint CloudTabularTime SeriesPhysics RelatedStochastic Differential Equation

🎯 What it does: Propose an algorithm called NGIF that learns non-gradient population dynamics by utilizing the weak continuity equation and gauge freedom

Leveraging Machine Unlearning for Cost-Efficient Preference Alignment

XiaoHua Feng, Chaochao Chen (Zhejiang University)

CodeComputational EfficiencyData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningContrastive LearningText

🎯 What it does: This paper proposes a U2A framework that combines Machine Unlearning with Preference Alignment, utilizing a two-layer optimization to precisely select and weight negative samples for unlearning, thereby significantly improving the preference alignment of LLMs without requiring a large number of positive samples.

LIFT: A Novel Framework for Enhancing Long-Context Understanding of LLMs via Long Input Fine-Tuning

Yansheng Mao (Peking University), Muhan Zhang (Peking University)

CodeComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose the LIFT framework, which enables short-context LLMs to answer questions without full input in long-context tasks by generating synthetic QA and parameter fine-tuning during testing on long texts.

LightningRL: Breaking the Accuracy–Parallelism Trade-off of Block-wise dLLMs via Reinforcement Learning

Yanzhe Hu (Shanghai Jiao Tong University), Zhijie Deng (Shanghai Jiao Tong University)

CodeOptimizationComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringContrastive LearningText

🎯 What it does: Post-training a pre-trained block-wise diffusion language model (dLLM) using reinforcement learning, aiming to simultaneously improve parallel generation speed and generation quality.

LILO: Bayesian Optimization with Natural Language Feedback

Kasia Kobalczyk, Eytan Bakshy (Meta)

CodeOptimizationReinforcement Learning from Human FeedbackNeural Architecture SearchTransformerLarge Language ModelPrompt EngineeringImageTextTabular

🎯 What it does: Proposes a framework called LILO (Language-in-the-Loop Optimization) based on Bayesian optimization, which leverages large language models (LLMs) to convert free-text feedback from decision-makers into structured pairwise preferences, and performs Bayesian optimization using a Gaussian process (GP) surrogate.

LIMSSR: LLM-Driven Sequence-to-Score Reasoning under Training-Time Incomplete Multimodal Observations

Huangbiao Xu (Fuzhou University), Yuxin Peng (Peking University)

CodeRestorationRepresentation LearningData-Centric LearningTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsAuto EncoderContrastive LearningVideoTextMultimodality

🎯 What it does: Propose a framework called LIMSSR based on large language models to address learning tasks where multi-modal missing data exists during the training phase, converting missing multi-modal reasoning into a conditional sequence-to-score reasoning process.

LLM4Branch: Large Language Model for Discovering Efficient Branching Policies of Integer Programs

Zhinan Hou (Tsinghua University), Keyou You (Tsinghua University)

CodeOptimizationTransformerLarge Language ModelPrompt EngineeringText

🎯 What it does: Leverage large language models to automatically generate and optimize branching strategies in order to improve the efficiency of solving mixed integer linear programming problems

Local Linearity of LLMs Enables Activation Steering via Model-Based Linear Optimal Control

Julian Skifstad (Georgia Institute of Technology), Glen Chou (Georgia Institute of Technology)

CodeOptimizationExplainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningTextOrdinary Differential Equation

🎯 What it does: This paper proposes a model-based linear optimal control framework (A-LQR) for real-time adjustment of the activation layer in LLMs, and demonstrates that Transformer layers can be approximated as locally linear near reachable activations, enabling closed-loop feedback control;

Local MAP Sampling for Diffusion Models

Shaorong Zhang (University of California, Riverside), Greg Ver Steeg (University of California, Riverside)

CodeRestorationOptimizationDiffusion modelScore-based ModelImagePoint CloudMagnetic Resonance ImagingComputed TomographyBenchmark

🎯 What it does: Propose and implement Local MAP Sampling (LMAPS), a method that solves inverse problems during the inference process of diffusion models by iteratively solving local MAP subproblems.

Localize and Neutralize: Gradient-Guided Token Suppression Against Visual Prompt Injection Attack

Dongpeng Zhang (University of Chinese Academy of Sciences), Qingming Huang (University of Chinese Academy of Sciences)

CodeSafty and PrivacyExplainability and InterpretabilityComputational EfficiencyAdversarial AttackTransformerPrompt EngineeringVision Language ModelImageTextMultimodality

🎯 What it does: Proposes a lightweight inference-time defense mechanism called Gradient Token Masking (GTM) to counteract perturbation-based visual prompt injection attacks.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining

Qiuwu Chen (AIGCode), Mingkui Tan (South China University Of Technology)

CodeComputational EfficiencyRepresentation LearningTransformerLarge Language ModelMixture of ExpertsText

🎯 What it does: Designed a novel large language model architecture called LoKiFormer, which combines Local Fusion Attention (LFA) and Knowledge Memory Module (KMM) to improve pre-training efficiency.

LORD-GoF: A Robust Online Detection Approach for LLM Watermarks in Sparse and Mixed Streams

Jiade Xu (Lanzhou University), Zhouping Li (Lanzhou University)

CodeAnomaly DetectionData-Centric LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningText

🎯 What it does: Proposes the LORD‑GoF framework, combining the Goodness‑of‑Fit (GoF) statistic with LORD online false discovery rate (FDR) control, to achieve watermark detection in sparse and mixed human-machine text streams.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation

Jiarui Wang (Shanghai Jiao Tong University), Xiongkuo Min (Shanghai Jiao Tong University)

CodeGenerationData SynthesisExplainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision Language ModelVideoTextMultimodalityBenchmark

🎯 What it does: Constructed the largest AI-generated video evaluation dataset, AIGVE-60K, and proposed the LMM-driven LOVE evaluation framework and LOVE-Reward reward strategy to multidimensionally evaluate and improve text-to-video generation models;

Lower Bounds for Frank-Wolfe on Strongly Convex Sets

Jannis Halbey (Zuse Institute Berlin), Sebastian Pokutta (Zuse Institute Berlin)

CodeOptimization

🎯 What it does: Minimize a strongly convex quadratic function on the Euclidean unit ball, construct the worst-case Frank-Wolfe trajectory, and prove that the lower bound of FW on strongly convex sets is Ω(1/√Ρ).

MA$^3$S: Model-Agnostic Active Annotation Strategy for Crowdsourcing

Wenjun Zhang (Zhongnan University of Economics and Law), Shanshan Si (China University of Geosciences)

CodeFederated LearningData-Centric LearningContrastive LearningImageTextTabular

🎯 What it does: Propose a model-agnostic active repeated annotation strategy, MA S, aiming to reduce label redundancy and instance redundancy in crowdsourcing annotation, and improve annotation efficiency through online updates.

Machine Learning Hamiltonians are Accurate Energy-Force Predictors

Seongsu Kim (Korea Advanced Institute of Science and Technology), Sungsoo Ahn (Korea Advanced Institute of Science and Technology)

CodeDrug DiscoveryGraph Neural NetworkTransformerFlow-based ModelGraphTabularBenchmarkPhysics Related

🎯 What it does: Proposed a machine learning Hamiltonian model QHFlow2 that can be directly used for computing energy and forces, and constructed a unified benchmark for direct evaluation.

MADE: Benchmark Environments for Closed-Loop Materials Discovery

Shreshth A Malik, Yarin Gal

CodeOptimizationDrug DiscoveryReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringDiffusion modelGraphTabularBenchmark

🎯 What it does: Proposed the MADE (Materials Discovery Environments) framework for evaluating the efficiency and effectiveness of closed-loop materials discovery pipelines under limited query budgets;

MaMa: A Game-Theoretic Approach for Designing Safe Agentic Systems

Jonathan NΓΆther (Max Planck Institute for Software Systems), Goran Radanovic (Max Planck Institute for Software Systems)

CodeOptimizationSafty and PrivacyAdversarial AttackReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringTextTabularTime SeriesSequentialReview/Survey PaperBenchmarkFinance Related

🎯 What it does: Design an automated method called MaMa based on Stackelberg security games to build secure multi-agent systems in the presence of agents under attack

MAMBO-G: Magnitude-Aware Mitigation for Boosted Guidance

Shangwen Zhu (Shanghai Jiao Tong University), Fan Cheng (Shanghai Jiao Tong University)

CodeGenerationComputational EfficiencyTransformerDiffusion modelScore-based ModelImageVideoTextOrdinary Differential Equation

🎯 What it does: Proposes MAMBO-G, an untrained adaptive acceleration framework that dynamically adjusts the strength of classifier-free guidance (CFG).

MaMi-HOI: Harmonizing Global Kinematics and Local Geometry for Human-Object Interaction Generation

Hao Wang (South China University of Technology), Qi Liu (South China University of Technology)

CodeGenerationData SynthesisPose EstimationTransformerVision-Language-Action ModelDiffusion modelScore-based ModelContrastive LearningImageVideoTextMultimodalityPoint CloudMesh

🎯 What it does: Propose the MaMi-HOI framework for generating human-robot interaction animations in 3D scenes that are both semantically intent-aligned and physically accurate in contact.

Mantis: Lightweight Foundation Model for Time Series Classification

Vasilii Feofanov (Huawei Noah's Ark Lab), Ievgen Redko (Huawei Noah's Ark Lab)

CodeClassificationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerContrastive LearningTime Series

🎯 What it does: Proposed a lightweight time series classification foundation model called Mantis, and achieved zero-shot feature extraction;

MapUQ: Map with Uncertainty Quantification for Robust BEV Vectorized Construction

Shaoyuan Mo (Chongqing University), Ke Wang (Chongqing University)

CodeAutonomous DrivingExplainability and InterpretabilityComputational EfficiencyTransformerContrastive LearningSimultaneous Localization and MappingImageVideoPoint Cloud

🎯 What it does: Proposes the MapUQ framework, integrating uncertainty quantification into BEV vectorized map generation to enhance model robustness in complex traffic scenarios.

Margin-Adaptive Confidence Ranking for Reliable LLM Judgement

Gaojie Jin (University of Macau), Tianjin Huang (University of Exeter)

CodeExplainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringContrastive LearningText

🎯 What it does: Provide reliable confidence estimates for LLMs as judges, learn a margin-based ranking network to replace traditional heuristic confidence signals, and propose an adaptive margin training scheme based on this.

MAS-Orchestra: Understanding and Improving Multi-Agent Reasoning Through Holistic Orchestration and Controlled Benchmarks

Zixuan Ke (Salesforce Research), Shafiq Joty (Salesforce Research)

CodeOptimizationTransformerReinforcement LearningAgentic AIPrompt EngineeringTextSequentialBenchmarkChain-of-Thought

🎯 What it does: Proposed the MAS-Orchestra framework, achieving global one-time orchestration of multi-agent systems through function-call-based reinforcement learning during training, and constructed MASBench for systematic comparison between MAS and single-agent systems.

MASPOB: Bandit-Based Prompt Optimization for Multi-Agent Systems with Graph Neural Networks

Zhi Hong (Chinese University of Hong Kong), Zhongxiang Dai (Chinese University of Hong Kong)

CodeOptimizationAI Code AssistantGraph Neural NetworkTransformerReinforcement LearningPrompt EngineeringText

🎯 What it does: Optimize prompts for a confirmed multi-agent system (MAS), proposing a Bandit-based framework called MASPOB.

MCP-Persona: Benchmarking LLM Agents on Real-World Personal Applications via Environment Simulation

WenHao Wang, Siheng Chen (Shanghai Jiao Tong University)

CodeAutonomous DrivingFederated LearningSafty and PrivacyExplainability and InterpretabilityComputational EfficiencyDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackNeural Architecture SearchProtein Structure PredictionTransformerLarge Language ModelPrompt EngineeringWorld ModelTextSequentialBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose the MCP-Persona benchmark, simulating a realistic personalized MCP tool environment to evaluate the performance of LLM agents in social and collaborative applications

MDGMIX: Boundary-Aware Subgraph Mixing for Multi-Domain Graph Pre-Training

Ziyu Zheng (Xidian University), Xinyan Huang (Xidian University)

CodeDomain AdaptationRepresentation LearningGraph Neural NetworkPrompt EngineeringContrastive LearningGraph

🎯 What it does: Propose the MDGMIX framework, which achieves multi-domain graph pre-training through boundary-aware subgraph mixing and hierarchical domain discrimination, significantly reducing data redundancy and improving cross-domain generalization.

Mean Flow Distillation: Robust and Stable Distillation for Flow Matching Models

An Zhao (Zhejiang University), Lingyun Sun (Zhejiang University)

CodeGenerationOptimizationComputational EfficiencyKnowledge DistillationTransformerDiffusion modelScore-based ModelFlow-based ModelImageVideoTextOrdinary Differential Equation

🎯 What it does: Propose a Mean Flow Distillation (MFD) for Flow Matching models to achieve high-quality generation in a single step.

MedCRP-CL: Continual Medical Image Segmentation via Bayesian Nonparametric Semantic Modality Discovery

Ziyuan Gao (University College London)

CodeSegmentationTransformerPrompt EngineeringMixture of ExpertsVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasound

🎯 What it does: Propose MedCRP-CL, a framework for online task structure discovery in medical image segmentation and achieve continual learning without replay.

MedMamba: Multi-View State Space Models with Adaptive Graph Learning for Medical Time Series Classification

Da Zhang (Northwest Polytechnical University), Xuelong Li (TeleAI, China Telecom)

CodeClassificationGraph Neural NetworkTransformerAuto EncoderContrastive LearningTime SeriesBiomedical DataElectrocardiogramReview/Survey PaperBenchmark

🎯 What it does: Propose the MedMamba structure, integrating multi-scale convolutional embeddings, a three-branch differential state space encoder, and adaptive spatial graph Mamba, to achieve efficient classification of medical time series.

Meerkat-VL: Implicit Risk Safety Alignment in Multimodal LLMs via Perceptual Reasoning and Self-Verification

Peicheng Zhou (University of Science and Technology of China), Hongtao Xie (University of Science and Technology of China)

CodeSafty and PrivacyReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision Language ModelVision-Language-Action ModelImageTextMultimodality

🎯 What it does: Propose an implicit risk safety alignment framework called Meerkat-VL for multimodal large language models, enhancing the model's ability to perceive and respond safely to implicit risks through perceptual reasoning, model self-verification, and dual-objective consistency alignment.

MemCast: Memory-Driven Time Series Forecasting with Experience-Conditioned Reasoning

Xiaoyu Tao (University of Science and Technology of China), Shijin Wang (iFLYTEK Research)

CodeTransformerLarge Language ModelPrompt EngineeringTime SeriesFinance RelatedRetrieval-Augmented Generation

🎯 What it does: Propose the MemCast framework, which redefines time series forecasting as an experience-conditioned reasoning task, guiding LLM reasoning through text-based multi-level memory (historical patterns, reasoning wisdom, and general rules).

Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity

Menglin Xia (Microsoft), Saravan Rajmohan (Microsoft)

CodeRetrievalComputational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextSequentialBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose the MEMORA harmonic memory architecture, which balances the abstraction and concreteness of memory by utilizing main abstraction, prompt anchor points, and policy-based retrieval.

Memory as Dynamics: Learning Reliability-Guided Predictive Models for Online Video Perception

Minwoo Kim (Kookmin University), Sang Min Yoon (Kookmin University)

CodeObject TrackingSegmentationTransformerPrompt EngineeringVision-Language-Action ModelDiffusion modelVideo

🎯 What it does: Built a reliability-guided predictive memory framework called RPM for online video perception.

MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems

Qingyao Ai (Tsinghua University), Yiqun LIU

CodeFederated LearningExplainability and InterpretabilityComputational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Constructed the MemoryBench benchmark to evaluate the memory and continuous learning capabilities of LLM systems;

Message Tuning Outshines Graph Prompt Tuning: A Prismatic Space Perspective

Yancheng Chen (Academy of Mathematics and Systems Science, Chinese Academy of Sciences), Chuan Zhou (Academy of Mathematics and Systems Science, Chinese Academy of Sciences)

CodeClassificationGraph Neural NetworkSupervised Fine-TuningPrompt EngineeringGraph

🎯 What it does: This paper proposes a framework for measuring the adaptability of graph models based on Prismatic Space Theory, and designs the Message Tuning method on this basis to enhance the adaptability of graph foundational models in downstream tasks.

MetaphorVU: Towards Metaphorical Video Understanding

Zhuoqun Li (Chinese Academy of Sciences), Le Sun (Chinese Academy of Sciences)

CodeExplainability and InterpretabilityKnowledge DistillationRepresentation LearningGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringVision Language ModelVideoTextGraphBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposed the MetaphorVU-Bench benchmark specifically for metaphor video understanding, systematically evaluating the capabilities of existing MLLMs in metaphor video reasoning, and introduced the MetaphorBoost approach based on a metaphor knowledge graph, enhancing cross-domain mapping during reasoning.

Mind the Gap: Mixtures of Gaussians in Approximate Differential Privacy

Huikang Liu (Shanghai Jiao Tong University), Wolfram Wiesemann (Imperial Business School)

CodeSafty and PrivacyGaussian SplattingTabularBenchmarkStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Propose a noise injection mechanism under (Ρ,δ) approximate differential privacy by mixing multiple Gaussian distributions with the same variance but different means, significantly reducing the expected noise magnitude and variance;

MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models

Chuang Yu (Shenyang Institute of Automation, Chinese Academy of Sciences), Xiangyu Yue (Chinese University of Hong Kong)

CodeExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningImageTextMultimodalityRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose a multi-reason integration discriminative reasoning framework based on a multi-modal large language model (MIND), achieving the model's 'understand β†’ re-examine β†’ correct' reasoning ability through three major technologies: automatically constructing multi-reason data, two-stage evolutionary learning, and contrastive alignment.

miniF2F-Dafny: LLM-Guided Mathematical Theorem Proving via Auto-Active Verification

Mantas Baksys, Sean B. Holden (University of Cambridge)

CodeOptimizationComputational EfficiencyAI Code AssistantTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This study proposes MINIF2F-DAFNY, which migrates the miniF2F mathematical proof benchmark to the automated verifier Dafny, and utilizes LLMs to assist in generating proof prompts;

MINIM: Privacy-Aware Minimal View for Agents via Trusted Local Sanitization

Hexuan Yu (Virginia Tech), Wenjing Lou (Virginia Tech)

CodeFederated LearningSafty and PrivacyGraph Neural NetworkTransformerPrompt EngineeringContrastive LearningTextGraph

🎯 What it does: Proposes MINIM, a trustworthy local agent that performs structured UI observation on the client side. It generates a ternary publication strategy by predicting sensitivity and task necessity for each UI element, thereby minimizing the observations sent to the remote inference server.

Mining Useful General Data for Low-Resource Domain Adaptation

Pingjie Wang (Shanghai Jiao Tong University), Yu Wang (Shanghai Jiao Tong University)

CodeDomain AdaptationData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningTextTabularFinance RelatedChain-of-Thought

🎯 What it does: Propose NTK-Selector, which improves the adaptation effect in low-resource domains by leveraging general domain data;

Mitigating Bias in Locally Constrained Decoding via Tractable Proposals

Meihua Dang (Stanford University), Stefano Ermon (Stanford University)

CodeGenerationData SynthesisComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelReinforcement LearningDiffusion modelScore-based ModelAuto EncoderContrastive LearningTextSequentialStochastic Differential Equation

🎯 What it does: In constraint generation tasks for large language models, we propose Global Constraint Decoding (GCD) and Probabilistic Global Constraint Decoding (P-GCD), and use them as the proposal and potential distributions in sequential Monte Carlo (SMC) sampling, thereby reducing the bias of local constraint decoding (LCD) and ensuring that constraints are satisfied within the given token limit.

Mitigating Plasticity Loss through Architectural Design in Continual Learning

Niklas Koeppe (KAIST), Sang Wan Lee (KAIST)

CodeOptimizationReinforcement LearningContrastive LearningImageVideo

🎯 What it does: Designed and evaluated a network layer called InterpLayer to alleviate plasticity loss in continual reinforcement learning.

Mixing Configurations for Downstream Prediction

Juntang Wang (Duke Kunshan University), Shixin Xu (Duke Kunshan University)

CodeClassificationRepresentation LearningData-Centric LearningTransformerMixture of ExpertsAuto EncoderContrastive LearningImageTextTabularBiomedical Data

🎯 What it does: Propose the MixConfig module, which can extract limited stable clustering configurations in any embedding space, and use an energy-aware selector to learn sample-level weighted mixing for downstream prediction.

ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World

Ziyin Zhang (Shanghai Jiao Tong University), Rui Wang (Shanghai Jiao Tong University)

CodeRetrievalCompressionComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelContrastive LearningTextMultimodalityBenchmark

🎯 What it does: Proposed the ML-Embed series of models, utilizing the 3-D Matryoshka Learning (3D-ML) framework to achieve multi-dimensional compression of the embedding layer, network depth, and representation size, enabling significant savings in training, inference, and storage;

MOC: Multi-Order Communication in LLM-based Multi-Agent Systems

Yao Guan (Fudan University), Qiang Duan (Pennsylvania State University)

CodeOptimizationFederated LearningComputational EfficiencyKnowledge DistillationReinforcement Learning from Human FeedbackGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringTextGraphRetrieval-Augmented Generation

🎯 What it does: Propose a Multi-Order Communication (MOC) scheme to transmit raw information with multi-hop dependencies in a structured manner to the target agent within large language model (LLM)-driven multi-agent systems, and design a Semantic-Topological Merging mechanism to compress redundant information and improve communication efficiency.

MoDA: Modulation Adapter for Fine-Grained Visual Understanding in Instructional MLLMs

Wayner Barrios (Dartmouth College), Bernard Ghanem (King Abdullah University of Science and Technology)

CodeExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerSupervised Fine-TuningPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBenchmark

🎯 What it does: Proposed a lightweight channel-level modulation adapter called MoDA, which dynamically modulates visual features with language instructions to enhance fine-grained visual understanding in multi-modal large language models and reduce hallucinations.

Model Fusion via Retrofitting

Phoomraphee Luenam (ETH ZΓΌrich), Sidak Pal Singh (ETH ZΓΌrich)

CodeClassificationFederated LearningExplainability and InterpretabilityKnowledge DistillationConvolutional Neural NetworkTransformerMixture of ExpertsAuto EncoderContrastive LearningImageBiomedical Data

🎯 What it does: This paper proposes a fusion framework based on neuron clustering and alignment (Retrofitting), which first performs importance-weighted clustering on the intermediate neurons of the parent model, and then trains the subnetwork of the fusion model to approximate the cluster centers, achieving fusion for any hierarchically structured DAG model.

Modeling Hierarchical Thinking in Large Reasoning Models

G M Shahariar (University of California Riverside), Nael Abu-Ghazaleh (University of California Riverside)

CodeExplainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningContrastive LearningTextBenchmarkChain-of-Thought

🎯 What it does: This study investigates hierarchical thinking in large reasoning models (LRM), abstracting Chain-of-Thought into six cognitive states using a finite state machine (FSM), and designing a sparse activation guided control method without weighted updates during training through transition advantage matrices and Q-Value iteration;

Modeling Temporal scRNA-seq Data with Latent Gaussian Process and Optimal Transport

Mehmet Yigit Balik (Aalto University), Harri LΓ€hdesmΓ€ki (Aalto University)

CodeGenerationData SynthesisRepresentation LearningDiffusion modelScore-based ModelFlow-based ModelAuto EncoderContrastive LearningGaussian SplattingTime SeriesSequentialBiomedical DataStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Built a generative model based on implicit Gaussian processes and optimal transport for reconstructing continuous temporal dynamics from static snapshots of single-cell RNA sequencing data;

MOOSE-Star: Unlocking Tractable Training for Scientific Discovery by Breaking the Complexity Barrier

Zonglin Yang (MiroMind AI), Lidong Bing (MiroMind AI)

CodeOptimizationComputational EfficiencyData-Centric LearningDrug DiscoveryTransformerLarge Language ModelPrompt EngineeringTextReview/Survey PaperRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose the MOOSE-Star framework, which decomposes the exponential search problem of directly training P(h|b) into subtasks that are linear or even logarithmic, enabling a trainable model for scientific discovery;

MORE: A Multilingual Document Parsing Benchmark and Evaluation

Long Xu (Tencent), feng zhang

CodeTransformerPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose the MORE benchmark for multilingual document parsing evaluation.

MoRGen: Mixture-of-Resolutions Generative Forecasting for Irregularly Sampled Medical Time-Series Data

Nassim Oufattole (Massachusetts Institute of Technology), Collin Stultz (Harvard Medical School)

CodeGenerationData SynthesisAnomaly DetectionTransformerMixture of ExpertsTabularTime SeriesBiomedical DataElectronic Health Records

🎯 What it does: Proposed the MoRGen method, which performs zero-shot medical time series risk prediction by fusing generative predictors with different temporal resolutions.

MoSE: Mixture of Slimmable Experts for Efficient and Adaptive Language Models

Nurbek Tastan (Mohamed bin Zayed University of Artificial Intelligence), Samuel HorvΓ‘th (Mohamed bin Zayed University of Artificial Intelligence)

CodeComputational EfficiencyRepresentation LearningTransformerLarge Language ModelMixture of ExpertsText

🎯 What it does: Propose Mixture of Slimmable Experts (MoSE), combining Mixture-of-Experts with Slimmable networks, enabling each expert to dynamically adjust its width during inference, thus achieving bi-axial variable computation.

MoshiRAG: Asynchronous Knowledge Retrieval for Full-Duplex Speech Language Models

Chung-Ming Chien (Toyota Technological Institute at Chicago), Alexandre DΓ©fossez (Kyutai)

CodeRetrievalExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelTextRetrieval-Augmented GenerationAudio

🎯 What it does: Built Moshi A, a full-duplex speech-language model, incorporating asynchronous retrieval-augmented generation (RAG) functionality to enhance the factual accuracy of answers while maintaining real-time interaction.

Motion-Aware Caching for Efficient Autoregressive Video Generation

Jing Xu (Xiamen University), Songwei Liu (ByteDance)

CodeGenerationComputational EfficiencyTransformerDiffusion modelFlow-based ModelOptical FlowVideo

🎯 What it does: Proposes MotionCache, a motion-aware caching framework for autoregressive video generation models, which uses frame differences as motion features to dynamically decide whether to recalculate each token or directly use cached residuals, thereby significantly reducing the number of iterative denoising steps.

Multi-Integration of Labels Across Categories for Component Identification in Multi-trial Time Series

Noga Mudrik (Johns Hopkins University), Adam Shabti Charles

CodeExplainability and InterpretabilityRepresentation LearningTime Series

🎯 What it does: Propose the MILCCI method to discover sparse and interpretable components in multi-experiment, multi-label time series, integrating label information and capturing variations between experiments and label effects;

Multi-Label Learning with Contrastive Cluster Self-Supervision for 3D Hierarchical Semantic Segmentation

Shuyu Cao (Southwest Jiaotong University), Na Zhao (Singapore University of Technology and Design)

CodeSegmentationConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningPoint Cloud

🎯 What it does: Propose a multi-label learning framework called ML3DHS for 3D hierarchical semantic segmentation, addressing issues of multi-level conflicts and class imbalance.

Multimodal Scaling Laws for Task & Data-Optimized Models of Visual Cortex

Abdulkadir Gokce (Γ‰cole Polytechnique FΓ©dΓ©rale de Lausanne), Martin Schrimpf (Γ‰cole Polytechnique FΓ©dΓ©rale de Lausanne)

CodeExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningConvolutional Neural NetworkTransformerSupervised Fine-TuningAuto EncoderContrastive LearningMultimodalityBiomedical DataMagnetic Resonance ImagingReview/Survey Paper

🎯 What it does: Systematically studied the scaling laws of model-brain alignment of visual models across multiple brain recording modalities (electrophysiology, fMRI, EEG, MEG), evaluating the impact of pre-training scale, neural fine-tuning, and the number of mapping training samples on alignment.

Multiple Choice Learning of Low-Rank Adapters for Language Modeling

Victor Letzelter (TΓ©lΓ©com Paris), Patrick Perez (Kyutai)

CodeGenerationData-Centric LearningTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsContrastive LearningImageTextMultimodalityAudio

🎯 What it does: Based on pre-trained language models, multiple low-rank adapters are added and a winner-takes-all loss is adopted to generate diverse and reasonable sentences for the same input.

Multiview Self-Representation Learning across Heterogeneous Views

Jie Chen (Sichuan University), Xi Peng (Sichuan University)

CodeDomain AdaptationRepresentation LearningConvolutional Neural NetworkTransformerContrastive LearningImage

🎯 What it does: Propose the MSRL method, which performs self-representation learning through heterogeneous views generated by multiple pre-trained models to learn invariant representations.

MutAtlas: A PDB-Wide Energy-Guided Atlas of Protein Mutation Effects

Ruihan Guo (University of Illinois Urbana-Champaign), Ge Liu (University of Illinois Urbana-Champaign)

CodeKnowledge DistillationDrug DiscoveryProtein Structure PredictionTransformerLarge Language ModelContrastive LearningGraphBiomedical DataBenchmark

🎯 What it does: Constructed a comprehensive PDB-wide benchmark dataset of single-point mutations, aligning the mutation signals from FoldX physical energy, ESM2 language model, and ESM-IF inverse folding model into a unified site-level 20-dimensional logit representation, and proposed a cross-source preference distillation framework without experimental labels, achieving multi-source information fusion through soft consistency weighted with inconsistency regularization;

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs

Huiyi Chen (University of Illinois Chicago), Lu Cheng (University of Illinois Chicago)

CodeExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringVision Language ModelDiffusion modelContrastive LearningImageTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This paper proposes MVI-Bench, a comprehensive benchmark for evaluating the robustness of large vision-language models (LVLMs) under misleading visual inputs, and designs the MVI-Sensitivity metric to quantify the model's sensitivity to visual deception.

NaRA: Noise-Aware LoRA for Parameter-Efficient Fine-Tuning of Diffusion LLMs

Shuaidi Wang (Southern University of Science and Technology), Yu Zhang (Southern University of Science and Technology)

CodeComputational EfficiencyRepresentation LearningData-Centric LearningTransformerLarge Language ModelPrompt EngineeringDiffusion modelImageText

🎯 What it does: Propose a noise-aware low-rank adaptation method (NaRA), which dynamically generates the core matrix according to the noise level during the diffusion process through a globally shared lightweight hypernetwork, achieving parameter-efficient fine-tuning of diffusion large language models.

Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning

Tong Wu (State Key Laboratory of General Artificial Intelligence Bigai), Zilong Zheng (State Key Laboratory of General Artificial Intelligence Bigai)

CodeComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelReinforcement LearningPrompt EngineeringText

🎯 What it does: Propose Native Parallel Reasoner (NPR), enabling LLM to self-evolve real parallel reasoning capabilities without teacher supervision.

NaviCache: Test-Time Self-Calibration Caching for Video Generation

Zheqi Lv (Zhejiang University), Fei Wu (Zhejiang University)

CodeGenerationComputational EfficiencyTransformerDiffusion modelVideo

🎯 What it does: Propose a framework called NaviCache for achieving test-time self-calibrating caching in video diffusion models, accelerating inference by dynamically tracking feature evolution.

NAVIGATE: Evaluating Visual-Guided Search Decision-Making on the Open Web

YaoQi Fan (Nanjing University), Tong Lu (Nanjing University)

CodeSupervised Fine-TuningVision Language ModelImageTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed the NAVIGATE benchmark for evaluating vision-guided open-network search decisions.

Needles in the Haystack: Addressing Signal Dilution Improves scRNA-seq Perturbation Response Modeling and Evaluation

Gabriel Mateo Mejia, Lucas Paulo de Lima Camillo (Shift Bioscience)

CodeData-Centric LearningDrug DiscoverySupervised Fine-TuningContrastive LearningBiomedical Data

🎯 What it does: Investigated the reasons why the average prediction baseline performs well in single-cell RNA sequencing perturbation response models, and proposed a weighted evaluation metric and training objective tailored for sparse signals.

Negatives-Dominant Contrastive Learning for Generalization in Imbalanced Domains

Meng Cao (Nanjing University of Aeronautics and Astronautics), Songcan Chen (Nanjing University of Aeronautics and Astronautics)

CodeClassificationDomain AdaptationContrastive LearningImage

🎯 What it does: This paper proposes a negative sample dominated contrastive learning framework called NDCL for imbalanced domain generalization (IDG), which can achieve more robust decision boundaries under different domains and class imbalance conditions;

Neural Low-Discrepancy Sequences

Michael Etienne Van Huffel (Max Plank Institute for Intelligent Systems), T. Konstantin Rusch (Max Plank Institute for Intelligent Systems)

CodeOptimizationHyperparameter SearchData-Centric LearningSupervised Fine-TuningTabularSequentialBenchmark

🎯 What it does: Propose a neural network-based low-discrepancy sequence (NEUROLDS) generation framework that can output low-discrepancy sequences at any length.

Neural QAOA$^2$: Differentiable Joint Graph Partitioning and Parameter Initialization for Quantum Combinatorial Optimization

Zubin Zheng (Southern University of Science and Technology), Shengcai Liu (Southern University of Science and Technology)

CodeOptimizationGraph Neural NetworkSupervised Fine-TuningGraphTabularBenchmarkPhysics Related

🎯 What it does: Propose Neural QAOA 2, a differentiable divide-and-conquer framework that jointly generates graph partitioning and QAOA parameters, addressing the partitioning metric mismatch and topology-unaware initialization issues in QAOA2.

Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights

Yulu Gan (Massachusetts Institute of Technology), Phillip Isola (Massachusetts Institute of Technology)

CodeOptimizationKnowledge DistillationRepresentation LearningHyperparameter SearchData-Centric LearningTransformerLarge Language ModelMixture of ExpertsContrastive LearningGaussian SplattingTextMultimodality

🎯 What it does: Investigated the structure of the parameter space around pre-trained models, discovering a large number of task experts near large models, and proposed the RandOpt algorithm, which combines random guessing and ensemble learning, using random perturbations to quickly locate and aggregate multi-task experts.

NeurIPS: Neuro-anatomical Inductive Priors for Sphere-based Brain Decoding

Sijin Yu (South China University Of Technology), Xin Zhang (South China University Of Technology)

CodeRestorationTransformerMixture of ExpertsDiffusion modelAuto EncoderImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: Propose the NeurIPS framework, achieving cross-subject image reconstruction through anatomy-prior-driven fMRI decoding on surface meshes.

NeurOCNN: A Neural-Operator-Based Model for Physiological Time Series

Daya Kumar (University of Western Ontario), Apurva Narayan (University of Western Ontario)

CodeClassificationConvolutional Neural NetworkTransformerContrastive LearningTime SeriesBiomedical DataElectrocardiogram

🎯 What it does: Proposes NeurOCNN, a physiological time series model based on neural operators, continuous-time spline convolution, Fourier projection pooling, and attention heads, for achieving function-to-label mapping of multi-channel physiological signals.

NITP: Next Implicit Token Prediction for LLM Pre-training

Xiangdong Zhang (Shanghai Jiao Tong University), Junchi Yan (Shanghai Jiao Tong University)

CodeRepresentation LearningTransformerLarge Language ModelMixture of ExpertsContrastive LearningText

🎯 What it does: Proposed the Next Implicit Token Prediction (NITP) objective, complementing the standard Next Token Prediction (NTP), enabling the model to learn the representation of the next token in a shallow semantic space, thereby reinforcing the geometric structure of hidden representations.