International Conference on Machine Learning Β· 1032 papers
Learn-to-learn on Arbitrary Textual Conditioning: A Hypernetwork-Driven Meta-gated LLM
Luo Ji (Geely AI Lab, Geely Auto Group), Hongyan Li (Geely AI Lab, Geely Auto Group)
CodeMeta LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringText
π― What it does: Propose MeGan, which utilizes a hypernetwork to generate adaptive Ξ²-SwiGLU gating, injecting meta-learning control signals into the FFN of LLMs, enabling fast adaptation under any text conditions.
π― What it does: Propose a learnable mask fine-tuning method called LIFT, which adaptively masks training for token learning difficulty and timing in diffusion language models.
CodeRetrievalCompressionRepresentation LearningDrug DiscoveryGraph Neural NetworkAuto EncoderContrastive LearningGraphBiomedical Data
π― What it does: Utilize the SAND framework to learn compressed shape-aware representations based on 2D molecular graphs, enabling shape similarity retrieval from a 1-billion-level molecular library;
π― What it does: Constructed a subjective comparative evaluation dataset covering more than 300 Android devices, containing display parameters and environmental information, and used the Blade-Chest model to aggregate preference votes. Subsequently, a lightweight conditional adaptation network was trained to improve the prediction of existing VQA metrics across different devices and environments.
π― What it does: In this study, the authors propose using a variational autoencoder (VAE) in the latent space to learn high-frequency continuous action blocks, and design a 'Reuse-then-Refine' (RTR) strategy to maintain continuity between action blocks under asynchronous inference, thereby enabling smooth and continuous execution by robots under high-frequency control.
π― What it does: Propose a state-space model based on continuous-time dynamic graphs (CTDG) (CTDG-SSM), which utilizes topology-aware HiPPO (CTT-HiPPO) together with graph Laplacian polynomial filters to achieve efficient memory and update of long-term and multi-hop spatial information.
Learning Protein Structure-Function Relationships through Knowledge-guided Representation Decomposition
Mingqing Wang (Tsinghua University), Zhixiang Ren (Pengcheng Laboratory)
CodeExplainability and InterpretabilityRepresentation LearningProtein Structure PredictionTransformerAuto EncoderContrastive LearningBiomedical Data
π― What it does: Propose the ProtDiS framework, which utilizes knowledge-guided representation decomposition to split pre-trained protein microenvironment embeddings into independent channels aligned with biophysical attributes, thereby improving the interpretability and predictive performance of structure-function relationships.
Learning Query-Aware Budget-Tier Routing for Runtime Agent Memory
Haozhen Zhang (Nanyang Technological University), Wenya Wang (Nanyang Technological University)
CodeComputational EfficiencyTransformerLarge Language ModelReinforcement LearningAgentic AITextRetrieval-Augmented Generation
π― What it does: Propose BudgetMem, a modular framework that decomposes runtime memory extraction into adjustable budget levels, and learns a lightweight router to make decisions across different budget levels, achieving controllable balance between memory extraction cost and quality in LLM agents.
CodeOptimizationData-Centric LearningAI Code AssistantLarge Language ModelSupervised Fine-TuningAgentic AIPrompt EngineeringTextTabularBenchmarkRetrieval-Augmented GenerationChain-of-Thought
π― What it does: Proposes a framework for automatically learning random self-reductions (RSR) from known programs and implements a system called Bitween.
Learning Rewrite-Invariant Reasoning with Targeted Alternation Training
Mousa Arraf (Technion Israel Institute of Technology), Kira Radinsky
CodeExplainability and InterpretabilityComputational EfficiencyData-Centric LearningGraph Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextChain-of-Thought
π― What it does: By sampling and aggregating multiple reasoning trajectories of large language models on semantically preserving rewriting tasks, we construct a reasoning graph for each problem; using this graph, we identify the 'Solution Boundary Cut' (SBC), which marks the transition from recoverable to unrecoverable states, and generate a small number of targeted positive and negative example pairs for model fine-tuning or in-context learning, thereby improving the model's reasoning robustness under semantically preserving rewriting.
CodeComputational EfficiencyKnowledge DistillationRepresentation LearningData-Centric LearningTransformerLarge Language ModelPrompt EngineeringText
π― What it does: By training extremely small language models with constrained paraphrasing (SAMBAL) on text, they learn syntactic structures while suppressing the influence of semantics and world knowledge.
Learning Taxonomic Trees with Hierarchical Representation Regularization for Large Multimodal Models
Hulingxiao He (Peking University), Yuxin Peng (Peking University)
CodeRecognitionRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningVision Language ModelContrastive LearningImageTextMultimodality
π― What it does: Propose HiR 2, a parameter-free hierarchical representation regularization method, which utilizes non-parametric cross-attention to construct semantic visual trees from intermediate layers of LMMs, and enhances hierarchical visual recognition (HVR) consistency through hyperbolic implication loss and spherical angular dispersion loss.
π― What it does: Propose a strictly one-class learning audio deepfake detection method called CA-SOADD, which utilizes the distribution shift view without negative samples as a boundary detector, and achieves a clear delineation of the compact distribution of real audio and the rejection boundary through triple-objective central anchor point learning.
π― What it does: This paper proposes Full-Order Bound (FOB), which approximates and directly optimizes the probability of complete ranking events by introducing segmented thresholds to the latent Gaussian scores.
Learning to Refine: Spectral-Decoupled Iterative Refinement Framework for Precipitation Nowcasting
Yunlong Zhou (Nanjing University), Xiaotong Yuan
CodeOptimizationRecurrent Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningVideoTime SeriesPhysics Related
π― What it does: A deterministic, spectrum-decoupled iterative refinement framework called SDIR is proposed for high-resolution precipitation nowcasting.
Learning to Route Languages for Multilingual Policy Optimization
Geyang Guo (Georgia Institute of Technology), Wei Xu (Georgia Institute of Technology)
CodeOptimizationTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsTextMultimodalityBenchmark
π― What it does: This study proposes a language routing-based multilingual policy optimization framework called LRPO, which allows the model to actively select the response language during training and generate diverse training signals through multilingual rollouts.
Learning to Share: Selective Memory for Efficient Parallel Agentic Systems
Joseph Fioresi (University of Central Florida), Mubarak Shah (University of Central Florida)
CodeComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningAgentic AITextBenchmarkRetrieval-Augmented Generation
π― What it does: Proposed and implemented Learning to Share (LTS), a method that uses a global shared memory and a learnable controller in parallel agent systems to selectively share intermediate results, thereby reducing redundant computations and improving execution efficiency.
Learning with Admissibility: Robust Fuzzy Hashing for Cross-Modal Retrieval with Noisy Labels
Xincheng Sun (Sichuan University), Yuan Sun (Sichuan University)
CodeRetrievalContrastive LearningMultimodality
π― What it does: Proposes a Robust Fuzzy Cross-Modal Hashing (RFCMH) framework based on fuzzy set theory to address the label noise problem in cross-modal retrieval;
LEGO: An LLM-Enabled Hierarchical Optimizer for Tensor Computation Graphs with Structure-Aware Search and Compositional Synthesis
Ruiyuan Xu (Chinese Academy of Sciences), Huimin Cui (Chinese Academy of Sciences)
CodeOptimizationComputational EfficiencyAI Code AssistantTransformerLarge Language ModelAgentic AIPrompt EngineeringMixture of ExpertsImageTextGraph
π― What it does: Construct a hierarchical optimization framework called LEGO based on large language models, used to automatically synthesize and optimize end-to-end tensor computation graphs on GPUs.
Less Is More: Fast and Accurate Reasoning with Cross-Head Unified Sparse Attention
Lijie Yang (Princeton University), Ravi Netravali (Princeton University)
CodeComputational EfficiencyAI Code AssistantTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsText
π― What it does: Proposes LessIsMore, a training-agnostic sparse attention mechanism that improves decoding efficiency in long reasoning models by leveraging cross-head unified token selection and stable nearest window.
Less is More: Geometric Unlearning for LLMs with Minimal Data Disclosure
Chenchen Tan (Monash University), Longxiang Gao (Qilu University of Technology)
CodeSafty and PrivacyExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBenchmark
π― What it does: Propose a geometry-based LLM unlearning method (Geometric Unlearning), which eliminates specific knowledge by projecting and aligning target entities in the prompt-conditioned hidden state space, without needing access to the original training corpus.
Less Precise Can Be More Reliable: A Systematic Evaluation of Quantizationβs Impact on VLMs Beyond Accuracy
Aymen Bouguerra (UniversitΒ΄ e Paris-Saclay), Fabio Arnez (UniversitΒ΄ e Paris-Saclay)
CodeClassificationCompressionExplainability and InterpretabilityComputational EfficiencyRepresentation LearningContrastive LearningImageTextMultimodalityBenchmark
π― What it does: This paper conducts a systematic quantitative evaluation of Vision-Language Models (VLMs), designing over 700k evaluation runs covering five reliability dimensions (robustness, calibration, OOD detection, distribution drift, and pulse correlation), and analyzes the impact of different quantization strategies on these metrics.
Less Token, More Signal: MoE Expert Pruning via Critical Token Selection
Zeliang Zong (Hikvision Research Institute), Jilin Hu (East China Normal University)
CodeComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelMixture of ExpertsText
π― What it does: Propose a Mixture-of-Experts (MoE) expert pruning method called STEP based on key token selection, which uses attention-guided token importance to filter noise, and combines dual-factor expert importance scoring and expert-to-bias knowledge preservation techniques to achieve efficient compression;
π― What it does: Propose an algorithm called NGIF that learns non-gradient population dynamics by utilizing the weak continuity equation and gauge freedom
Leveraging Machine Unlearning for Cost-Efficient Preference Alignment
XiaoHua Feng, Chaochao Chen (Zhejiang University)
CodeComputational EfficiencyData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningContrastive LearningText
π― What it does: This paper proposes a U2A framework that combines Machine Unlearning with Preference Alignment, utilizing a two-layer optimization to precisely select and weight negative samples for unlearning, thereby significantly improving the preference alignment of LLMs without requiring a large number of positive samples.
LIFT: A Novel Framework for Enhancing Long-Context Understanding of LLMs via Long Input Fine-Tuning
Yansheng Mao (Peking University), Muhan Zhang (Peking University)
CodeComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation
π― What it does: Propose the LIFT framework, which enables short-context LLMs to answer questions without full input in long-context tasks by generating synthetic QA and parameter fine-tuning during testing on long texts.
LightningRL: Breaking the AccuracyβParallelism Trade-off of Block-wise dLLMs via Reinforcement Learning
Yanzhe Hu (Shanghai Jiao Tong University), Zhijie Deng (Shanghai Jiao Tong University)
CodeOptimizationComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringContrastive LearningText
π― What it does: Post-training a pre-trained block-wise diffusion language model (dLLM) using reinforcement learning, aiming to simultaneously improve parallel generation speed and generation quality.
LILO: Bayesian Optimization with Natural Language Feedback
Kasia Kobalczyk, Eytan Bakshy (Meta)
CodeOptimizationReinforcement Learning from Human FeedbackNeural Architecture SearchTransformerLarge Language ModelPrompt EngineeringImageTextTabular
π― What it does: Proposes a framework called LILO (Language-in-the-Loop Optimization) based on Bayesian optimization, which leverages large language models (LLMs) to convert free-text feedback from decision-makers into structured pairwise preferences, and performs Bayesian optimization using a Gaussian process (GP) surrogate.
CodeRestorationRepresentation LearningData-Centric LearningTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsAuto EncoderContrastive LearningVideoTextMultimodality
π― What it does: Propose a framework called LIMSSR based on large language models to address learning tasks where multi-modal missing data exists during the training phase, converting missing multi-modal reasoning into a conditional sequence-to-score reasoning process.
LLM4Branch: Large Language Model for Discovering Efficient Branching Policies of Integer Programs
Zhinan Hou (Tsinghua University), Keyou You (Tsinghua University)
CodeOptimizationTransformerLarge Language ModelPrompt EngineeringText
π― What it does: Leverage large language models to automatically generate and optimize branching strategies in order to improve the efficiency of solving mixed integer linear programming problems
Local Linearity of LLMs Enables Activation Steering via Model-Based Linear Optimal Control
Julian Skifstad (Georgia Institute of Technology), Glen Chou (Georgia Institute of Technology)
CodeOptimizationExplainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningTextOrdinary Differential Equation
π― What it does: This paper proposes a model-based linear optimal control framework (A-LQR) for real-time adjustment of the activation layer in LLMs, and demonstrates that Transformer layers can be approximated as locally linear near reachable activations, enabling closed-loop feedback control;
π― What it does: Propose and implement Local MAP Sampling (LMAPS), a method that solves inverse problems during the inference process of diffusion models by iteratively solving local MAP subproblems.
Localize and Neutralize: Gradient-Guided Token Suppression Against Visual Prompt Injection Attack
Dongpeng Zhang (University of Chinese Academy of Sciences), Qingming Huang (University of Chinese Academy of Sciences)
CodeSafty and PrivacyExplainability and InterpretabilityComputational EfficiencyAdversarial AttackTransformerPrompt EngineeringVision Language ModelImageTextMultimodality
π― What it does: Proposes a lightweight inference-time defense mechanism called Gradient Token Masking (GTM) to counteract perturbation-based visual prompt injection attacks.
LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining
Qiuwu Chen (AIGCode), Mingkui Tan (South China University Of Technology)
CodeComputational EfficiencyRepresentation LearningTransformerLarge Language ModelMixture of ExpertsText
π― What it does: Designed a novel large language model architecture called LoKiFormer, which combines Local Fusion Attention (LFA) and Knowledge Memory Module (KMM) to improve pre-training efficiency.
LORD-GoF: A Robust Online Detection Approach for LLM Watermarks in Sparse and Mixed Streams
Jiade Xu (Lanzhou University), Zhouping Li (Lanzhou University)
CodeAnomaly DetectionData-Centric LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningText
π― What it does: Proposes the LORDβGoF framework, combining the GoodnessβofβFit (GoF) statistic with LORD online false discovery rate (FDR) control, to achieve watermark detection in sparse and mixed human-machine text streams.
LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation
Jiarui Wang (Shanghai Jiao Tong University), Xiongkuo Min (Shanghai Jiao Tong University)
CodeGenerationData SynthesisExplainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision Language ModelVideoTextMultimodalityBenchmark
π― What it does: Constructed the largest AI-generated video evaluation dataset, AIGVE-60K, and proposed the LMM-driven LOVE evaluation framework and LOVE-Reward reward strategy to multidimensionally evaluate and improve text-to-video generation models;
π― What it does: Propose a model-agnostic active repeated annotation strategy, MA S, aiming to reduce label redundancy and instance redundancy in crowdsourcing annotation, and improve annotation efficiency through online updates.
Machine Learning Hamiltonians are Accurate Energy-Force Predictors
Seongsu Kim (Korea Advanced Institute of Science and Technology), Sungsoo Ahn (Korea Advanced Institute of Science and Technology)
CodeDrug DiscoveryGraph Neural NetworkTransformerFlow-based ModelGraphTabularBenchmarkPhysics Related
π― What it does: Proposed a machine learning Hamiltonian model QHFlow2 that can be directly used for computing energy and forces, and constructed a unified benchmark for direct evaluation.
MADE: Benchmark Environments for Closed-Loop Materials Discovery
Shreshth A Malik, Yarin Gal
CodeOptimizationDrug DiscoveryReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringDiffusion modelGraphTabularBenchmark
π― What it does: Proposed the MADE (Materials Discovery Environments) framework for evaluating the efficiency and effectiveness of closed-loop materials discovery pipelines under limited query budgets;
MaMa: A Game-Theoretic Approach for Designing Safe Agentic Systems
Jonathan NΓΆther (Max Planck Institute for Software Systems), Goran Radanovic (Max Planck Institute for Software Systems)
CodeOptimizationSafty and PrivacyAdversarial AttackReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringTextTabularTime SeriesSequentialReview/Survey PaperBenchmarkFinance Related
π― What it does: Design an automated method called MaMa based on Stackelberg security games to build secure multi-agent systems in the presence of agents under attack
π― What it does: Proposes MAMBO-G, an untrained adaptive acceleration framework that dynamically adjusts the strength of classifier-free guidance (CFG).
π― What it does: Propose the MaMi-HOI framework for generating human-robot interaction animations in 3D scenes that are both semantically intent-aligned and physically accurate in contact.
MapUQ: Map with Uncertainty Quantification for Robust BEV Vectorized Construction
Shaoyuan Mo (Chongqing University), Ke Wang (Chongqing University)
CodeAutonomous DrivingExplainability and InterpretabilityComputational EfficiencyTransformerContrastive LearningSimultaneous Localization and MappingImageVideoPoint Cloud
π― What it does: Proposes the MapUQ framework, integrating uncertainty quantification into BEV vectorized map generation to enhance model robustness in complex traffic scenarios.
Margin-Adaptive Confidence Ranking for Reliable LLM Judgement
Gaojie Jin (University of Macau), Tianjin Huang (University of Exeter)
CodeExplainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringContrastive LearningText
π― What it does: Provide reliable confidence estimates for LLMs as judges, learn a margin-based ranking network to replace traditional heuristic confidence signals, and propose an adaptive margin training scheme based on this.
π― What it does: Proposed the MAS-Orchestra framework, achieving global one-time orchestration of multi-agent systems through function-call-based reinforcement learning during training, and constructed MASBench for systematic comparison between MAS and single-agent systems.
CodeAutonomous DrivingFederated LearningSafty and PrivacyExplainability and InterpretabilityComputational EfficiencyDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackNeural Architecture SearchProtein Structure PredictionTransformerLarge Language ModelPrompt EngineeringWorld ModelTextSequentialBenchmarkRetrieval-Augmented GenerationChain-of-Thought
π― What it does: Propose the MCP-Persona benchmark, simulating a realistic personalized MCP tool environment to evaluate the performance of LLM agents in social and collaborative applications
π― What it does: Propose the MDGMIX framework, which achieves multi-domain graph pre-training through boundary-aware subgraph mixing and hierarchical domain discrimination, significantly reducing data redundancy and improving cross-domain generalization.
MedCRP-CL: Continual Medical Image Segmentation via Bayesian Nonparametric Semantic Modality Discovery
Ziyuan Gao (University College London)
CodeSegmentationTransformerPrompt EngineeringMixture of ExpertsVision Language ModelContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasound
π― What it does: Propose MedCRP-CL, a framework for online task structure discovery in medical image segmentation and achieve continual learning without replay.
π― What it does: Propose the MedMamba structure, integrating multi-scale convolutional embeddings, a three-branch differential state space encoder, and adaptive spatial graph Mamba, to achieve efficient classification of medical time series.
Meerkat-VL: Implicit Risk Safety Alignment in Multimodal LLMs via Perceptual Reasoning and Self-Verification
Peicheng Zhou (University of Science and Technology of China), Hongtao Xie (University of Science and Technology of China)
CodeSafty and PrivacyReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision Language ModelVision-Language-Action ModelImageTextMultimodality
π― What it does: Propose an implicit risk safety alignment framework called Meerkat-VL for multimodal large language models, enhancing the model's ability to perceive and respond safely to implicit risks through perceptual reasoning, model self-verification, and dual-objective consistency alignment.
MemCast: Memory-Driven Time Series Forecasting with Experience-Conditioned Reasoning
Xiaoyu Tao (University of Science and Technology of China), Shijin Wang (iFLYTEK Research)
CodeTransformerLarge Language ModelPrompt EngineeringTime SeriesFinance RelatedRetrieval-Augmented Generation
π― What it does: Propose the MemCast framework, which redefines time series forecasting as an experience-conditioned reasoning task, guiding LLM reasoning through text-based multi-level memory (historical patterns, reasoning wisdom, and general rules).
CodeRetrievalComputational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextSequentialBenchmarkRetrieval-Augmented Generation
π― What it does: Propose the MEMORA harmonic memory architecture, which balances the abstraction and concreteness of memory by utilizing main abstraction, prompt anchor points, and policy-based retrieval.
MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems
Qingyao Ai (Tsinghua University), Yiqun LIU
CodeFederated LearningExplainability and InterpretabilityComputational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation
π― What it does: Constructed the MemoryBench benchmark to evaluate the memory and continuous learning capabilities of LLM systems;
Message Tuning Outshines Graph Prompt Tuning: A Prismatic Space Perspective
Yancheng Chen (Academy of Mathematics and Systems Science, Chinese Academy of Sciences), Chuan Zhou (Academy of Mathematics and Systems Science, Chinese Academy of Sciences)
π― What it does: This paper proposes a framework for measuring the adaptability of graph models based on Prismatic Space Theory, and designs the Message Tuning method on this basis to enhance the adaptability of graph foundational models in downstream tasks.
MetaphorVU: Towards Metaphorical Video Understanding
Zhuoqun Li (Chinese Academy of Sciences), Le Sun (Chinese Academy of Sciences)
CodeExplainability and InterpretabilityKnowledge DistillationRepresentation LearningGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringVision Language ModelVideoTextGraphBenchmarkRetrieval-Augmented Generation
π― What it does: Proposed the MetaphorVU-Bench benchmark specifically for metaphor video understanding, systematically evaluating the capabilities of existing MLLMs in metaphor video reasoning, and introduced the MetaphorBoost approach based on a metaphor knowledge graph, enhancing cross-domain mapping during reasoning.
Mind the Gap: Mixtures of Gaussians in Approximate Differential Privacy
Huikang Liu (Shanghai Jiao Tong University), Wolfram Wiesemann (Imperial Business School)
CodeSafty and PrivacyGaussian SplattingTabularBenchmarkStochastic Differential EquationOrdinary Differential Equation
π― What it does: Propose a noise injection mechanism under (Ξ΅,Ξ΄) approximate differential privacy by mixing multiple Gaussian distributions with the same variance but different means, significantly reducing the expected noise magnitude and variance;
MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models
Chuang Yu (Shenyang Institute of Automation, Chinese Academy of Sciences), Xiangyu Yue (Chinese University of Hong Kong)
CodeExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningImageTextMultimodalityRetrieval-Augmented GenerationChain-of-Thought
π― What it does: Propose a multi-reason integration discriminative reasoning framework based on a multi-modal large language model (MIND), achieving the model's 'understand β re-examine β correct' reasoning ability through three major technologies: automatically constructing multi-reason data, two-stage evolutionary learning, and contrastive alignment.
miniF2F-Dafny: LLM-Guided Mathematical Theorem Proving via Auto-Active Verification
Mantas Baksys, Sean B. Holden (University of Cambridge)
CodeOptimizationComputational EfficiencyAI Code AssistantTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
π― What it does: This study proposes MINIF2F-DAFNY, which migrates the miniF2F mathematical proof benchmark to the automated verifier Dafny, and utilizes LLMs to assist in generating proof prompts;
CodeFederated LearningSafty and PrivacyGraph Neural NetworkTransformerPrompt EngineeringContrastive LearningTextGraph
π― What it does: Proposes MINIM, a trustworthy local agent that performs structured UI observation on the client side. It generates a ternary publication strategy by predicting sensitivity and task necessity for each UI element, thereby minimizing the observations sent to the remote inference server.
π― What it does: In constraint generation tasks for large language models, we propose Global Constraint Decoding (GCD) and Probabilistic Global Constraint Decoding (P-GCD), and use them as the proposal and potential distributions in sequential Monte Carlo (SMC) sampling, thereby reducing the bias of local constraint decoding (LCD) and ensuring that constraints are satisfied within the given token limit.
Juntang Wang (Duke Kunshan University), Shixin Xu (Duke Kunshan University)
CodeClassificationRepresentation LearningData-Centric LearningTransformerMixture of ExpertsAuto EncoderContrastive LearningImageTextTabularBiomedical Data
π― What it does: Propose the MixConfig module, which can extract limited stable clustering configurations in any embedding space, and use an energy-aware selector to learn sample-level weighted mixing for downstream prediction.
CodeRetrievalCompressionComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelContrastive LearningTextMultimodalityBenchmark
π― What it does: Proposed the ML-Embed series of models, utilizing the 3-D Matryoshka Learning (3D-ML) framework to achieve multi-dimensional compression of the embedding layer, network depth, and representation size, enabling significant savings in training, inference, and storage;
MOC: Multi-Order Communication in LLM-based Multi-Agent Systems
Yao Guan (Fudan University), Qiang Duan (Pennsylvania State University)
CodeOptimizationFederated LearningComputational EfficiencyKnowledge DistillationReinforcement Learning from Human FeedbackGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringTextGraphRetrieval-Augmented Generation
π― What it does: Propose a Multi-Order Communication (MOC) scheme to transmit raw information with multi-hop dependencies in a structured manner to the target agent within large language model (LLM)-driven multi-agent systems, and design a Semantic-Topological Merging mechanism to compress redundant information and improve communication efficiency.
MoDA: Modulation Adapter for Fine-Grained Visual Understanding in Instructional MLLMs
Wayner Barrios (Dartmouth College), Bernard Ghanem (King Abdullah University of Science and Technology)
CodeExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerSupervised Fine-TuningPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBenchmark
π― What it does: Proposed a lightweight channel-level modulation adapter called MoDA, which dynamically modulates visual features with language instructions to enhance fine-grained visual understanding in multi-modal large language models and reduce hallucinations.
Phoomraphee Luenam (ETH ZΓΌrich), Sidak Pal Singh (ETH ZΓΌrich)
CodeClassificationFederated LearningExplainability and InterpretabilityKnowledge DistillationConvolutional Neural NetworkTransformerMixture of ExpertsAuto EncoderContrastive LearningImageBiomedical Data
π― What it does: This paper proposes a fusion framework based on neuron clustering and alignment (Retrofitting), which first performs importance-weighted clustering on the intermediate neurons of the parent model, and then trains the subnetwork of the fusion model to approximate the cluster centers, achieving fusion for any hierarchically structured DAG model.
Modeling Hierarchical Thinking in Large Reasoning Models
G M Shahariar (University of California Riverside), Nael Abu-Ghazaleh (University of California Riverside)
CodeExplainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningContrastive LearningTextBenchmarkChain-of-Thought
π― What it does: This study investigates hierarchical thinking in large reasoning models (LRM), abstracting Chain-of-Thought into six cognitive states using a finite state machine (FSM), and designing a sparse activation guided control method without weighted updates during training through transition advantage matrices and Q-Value iteration;
π― What it does: Built a generative model based on implicit Gaussian processes and optimal transport for reconstructing continuous temporal dynamics from static snapshots of single-cell RNA sequencing data;
MOOSE-Star: Unlocking Tractable Training for Scientific Discovery by Breaking the Complexity Barrier
Zonglin Yang (MiroMind AI), Lidong Bing (MiroMind AI)
CodeOptimizationComputational EfficiencyData-Centric LearningDrug DiscoveryTransformerLarge Language ModelPrompt EngineeringTextReview/Survey PaperRetrieval-Augmented GenerationChain-of-Thought
π― What it does: Propose the MOOSE-Star framework, which decomposes the exponential search problem of directly training P(h|b) into subtasks that are linear or even logarithmic, enabling a trainable model for scientific discovery;
MoRGen: Mixture-of-Resolutions Generative Forecasting for Irregularly Sampled Medical Time-Series Data
Nassim Oufattole (Massachusetts Institute of Technology), Collin Stultz (Harvard Medical School)
CodeGenerationData SynthesisAnomaly DetectionTransformerMixture of ExpertsTabularTime SeriesBiomedical DataElectronic Health Records
π― What it does: Proposed the MoRGen method, which performs zero-shot medical time series risk prediction by fusing generative predictors with different temporal resolutions.
MoSE: Mixture of Slimmable Experts for Efficient and Adaptive Language Models
Nurbek Tastan (Mohamed bin Zayed University of Artificial Intelligence), Samuel HorvΓ‘th (Mohamed bin Zayed University of Artificial Intelligence)
CodeComputational EfficiencyRepresentation LearningTransformerLarge Language ModelMixture of ExpertsText
π― What it does: Propose Mixture of Slimmable Experts (MoSE), combining Mixture-of-Experts with Slimmable networks, enabling each expert to dynamically adjust its width during inference, thus achieving bi-axial variable computation.
CodeRetrievalExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelTextRetrieval-Augmented GenerationAudio
π― What it does: Built Moshi A, a full-duplex speech-language model, incorporating asynchronous retrieval-augmented generation (RAG) functionality to enhance the factual accuracy of answers while maintaining real-time interaction.
π― What it does: Proposes MotionCache, a motion-aware caching framework for autoregressive video generation models, which uses frame differences as motion features to dynamically decide whether to recalculate each token or directly use cached residuals, thereby significantly reducing the number of iterative denoising steps.
Multi-Integration of Labels Across Categories for Component Identification in Multi-trial Time Series
Noga Mudrik (Johns Hopkins University), Adam Shabti Charles
CodeExplainability and InterpretabilityRepresentation LearningTime Series
π― What it does: Propose the MILCCI method to discover sparse and interpretable components in multi-experiment, multi-label time series, integrating label information and capturing variations between experiments and label effects;
π― What it does: Propose a multi-label learning framework called ML3DHS for 3D hierarchical semantic segmentation, addressing issues of multi-level conflicts and class imbalance.
CodeExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningConvolutional Neural NetworkTransformerSupervised Fine-TuningAuto EncoderContrastive LearningMultimodalityBiomedical DataMagnetic Resonance ImagingReview/Survey Paper
π― What it does: Systematically studied the scaling laws of model-brain alignment of visual models across multiple brain recording modalities (electrophysiology, fMRI, EEG, MEG), evaluating the impact of pre-training scale, neural fine-tuning, and the number of mapping training samples on alignment.
CodeGenerationData-Centric LearningTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsContrastive LearningImageTextMultimodalityAudio
π― What it does: Based on pre-trained language models, multiple low-rank adapters are added and a winner-takes-all loss is adopted to generate diverse and reasonable sentences for the same input.
π― What it does: Propose the MSRL method, which performs self-representation learning through heterogeneous views generated by multiple pre-trained models to learn invariant representations.
MutAtlas: A PDB-Wide Energy-Guided Atlas of Protein Mutation Effects
Ruihan Guo (University of Illinois Urbana-Champaign), Ge Liu (University of Illinois Urbana-Champaign)
CodeKnowledge DistillationDrug DiscoveryProtein Structure PredictionTransformerLarge Language ModelContrastive LearningGraphBiomedical DataBenchmark
π― What it does: Constructed a comprehensive PDB-wide benchmark dataset of single-point mutations, aligning the mutation signals from FoldX physical energy, ESM2 language model, and ESM-IF inverse folding model into a unified site-level 20-dimensional logit representation, and proposed a cross-source preference distillation framework without experimental labels, achieving multi-source information fusion through soft consistency weighted with inconsistency regularization;
MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs
Huiyi Chen (University of Illinois Chicago), Lu Cheng (University of Illinois Chicago)
CodeExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringVision Language ModelDiffusion modelContrastive LearningImageTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought
π― What it does: This paper proposes MVI-Bench, a comprehensive benchmark for evaluating the robustness of large vision-language models (LVLMs) under misleading visual inputs, and designs the MVI-Sensitivity metric to quantify the model's sensitivity to visual deception.
NaRA: Noise-Aware LoRA for Parameter-Efficient Fine-Tuning of Diffusion LLMs
Shuaidi Wang (Southern University of Science and Technology), Yu Zhang (Southern University of Science and Technology)
CodeComputational EfficiencyRepresentation LearningData-Centric LearningTransformerLarge Language ModelPrompt EngineeringDiffusion modelImageText
π― What it does: Propose a noise-aware low-rank adaptation method (NaRA), which dynamically generates the core matrix according to the noise level during the diffusion process through a globally shared lightweight hypernetwork, achieving parameter-efficient fine-tuning of diffusion large language models.
Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning
Tong Wu (State Key Laboratory of General Artificial Intelligence Bigai), Zilong Zheng (State Key Laboratory of General Artificial Intelligence Bigai)
CodeComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelReinforcement LearningPrompt EngineeringText
π― What it does: Propose Native Parallel Reasoner (NPR), enabling LLM to self-evolve real parallel reasoning capabilities without teacher supervision.
π― What it does: Propose a framework called NaviCache for achieving test-time self-calibrating caching in video diffusion models, accelerating inference by dynamically tracking feature evolution.
Needles in the Haystack: Addressing Signal Dilution Improves scRNA-seq Perturbation Response Modeling and Evaluation
Gabriel Mateo Mejia, Lucas Paulo de Lima Camillo (Shift Bioscience)
CodeData-Centric LearningDrug DiscoverySupervised Fine-TuningContrastive LearningBiomedical Data
π― What it does: Investigated the reasons why the average prediction baseline performs well in single-cell RNA sequencing perturbation response models, and proposed a weighted evaluation metric and training objective tailored for sparse signals.
π― What it does: This paper proposes a negative sample dominated contrastive learning framework called NDCL for imbalanced domain generalization (IDG), which can achieve more robust decision boundaries under different domains and class imbalance conditions;
π― What it does: Propose a neural network-based low-discrepancy sequence (NEUROLDS) generation framework that can output low-discrepancy sequences at any length.
Neural QAOA$^2$: Differentiable Joint Graph Partitioning and Parameter Initialization for Quantum Combinatorial Optimization
Zubin Zheng (Southern University of Science and Technology), Shengcai Liu (Southern University of Science and Technology)
CodeOptimizationGraph Neural NetworkSupervised Fine-TuningGraphTabularBenchmarkPhysics Related
π― What it does: Propose Neural QAOA 2, a differentiable divide-and-conquer framework that jointly generates graph partitioning and QAOA parameters, addressing the partitioning metric mismatch and topology-unaware initialization issues in QAOA2.
Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights
Yulu Gan (Massachusetts Institute of Technology), Phillip Isola (Massachusetts Institute of Technology)
CodeOptimizationKnowledge DistillationRepresentation LearningHyperparameter SearchData-Centric LearningTransformerLarge Language ModelMixture of ExpertsContrastive LearningGaussian SplattingTextMultimodality
π― What it does: Investigated the structure of the parameter space around pre-trained models, discovering a large number of task experts near large models, and proposed the RandOpt algorithm, which combines random guessing and ensemble learning, using random perturbations to quickly locate and aggregate multi-task experts.
NeurIPS: Neuro-anatomical Inductive Priors for Sphere-based Brain Decoding
Sijin Yu (South China University Of Technology), Xin Zhang (South China University Of Technology)
CodeRestorationTransformerMixture of ExpertsDiffusion modelAuto EncoderImageBiomedical DataMagnetic Resonance Imaging
π― What it does: Propose the NeurIPS framework, achieving cross-subject image reconstruction through anatomy-prior-driven fMRI decoding on surface meshes.
π― What it does: Proposes NeurOCNN, a physiological time series model based on neural operators, continuous-time spline convolution, Fourier projection pooling, and attention heads, for achieving function-to-label mapping of multi-channel physiological signals.
CodeRepresentation LearningTransformerLarge Language ModelMixture of ExpertsContrastive LearningText
π― What it does: Proposed the Next Implicit Token Prediction (NITP) objective, complementing the standard Next Token Prediction (NTP), enabling the model to learn the representation of the next token in a shallow semantic space, thereby reinforcing the geometric structure of hidden representations.