arXivSub Start free trial

ICML 2026 Papers — Page 37

International Conference on Machine Learning · 6554 papers

Multi-marginal temporal Schrödinger Bridge Matching from unpaired data

Thomas Gravier (IBENS, ENS Ulm, PSL), Auguste Genovesio (IBENS, ENS Ulm, PSL)

GenerationData SynthesisTransformerDiffusion modelScore-based ModelImageVideoTabularBiomedical DataStochastic Differential Equation

🎯 What it does: Propose a multi-boundary time Schrödinger bridge matching method (MMtSBM) to address dynamic inference and reconstruction from unpaired data

Multi-Objective Bayesian Optimization via Adaptive $\varepsilon$-Constraint Decomposition

Yaohong Yang (Aalto University), Samuel Kaski (Aalto University)

OptimizationHyperparameter SearchTabularTime SeriesSequentialBenchmark

🎯 What it does: Propose the STAGE-BO method, which achieves multi-objective Bayesian optimization through adaptive ε-constrained decomposition, aiming to uniformly cover the Pareto front while simultaneously handling constraints and preferences.

Multi-Objective Learning for Diffusion Models: A Statistical Theory under Semi-Supervised Learning

Ziheng Cheng (University of California, Berkeley), Xin Guo (University of California, Berkeley)

RestorationKnowledge DistillationRobotic IntelligenceConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelImageVideo

🎯 What it does: Propose a multi-objective learning framework for conditional diffusion models, achieving Pareto optimal multi-objective performance under semi-supervised settings by first training a lightweight specialist model with limited labeled data, and then distilling pseudo samples into a large-capacity generalist model.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models

Akhil Agnihotri (University of Southern California), Zheng Wen (Google DeepMind)

OptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringText

🎯 What it does: Proposed a multi-objective preference optimization method called MOPO, enabling LLMs to achieve Pareto optimality while satisfying multiple human preferences such as usefulness and safety simultaneously.

Multi-Objective Protein Design via Memory-Aware Test-Time Scaling in Diffusion Models

Ming Yang (Dalian University of Technology), Shirui Pan (Griffith University)

Drug DiscoveryProtein Structure PredictionTransformerDiffusion modelContrastive LearningBiomedical Data

🎯 What it does: Propose the MOMST framework, which utilizes memory-aware test-time diffusion models to achieve multi-objective protein design without retraining.

Multi-Round Human–AI Collaboration with User-Specified Requirements

Sima Noorani (University of Pennsylvania), George J. Pappas

Safty and PrivacyExplainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBiomedical DataBenchmark

🎯 What it does: This paper proposes a multi-round human-AI collaborative framework, which enables AI to provide uncertainty information in the form of a prediction set during the dialogue process by setting user-specified counterfactual harm and complementarity constraints.

Multi-scale Explainer for Graph Neural Networks

Lutong Wu (Shanxi University), Jiye Liang (Shanxi University)

ClassificationExplainability and InterpretabilityGraph Neural NetworkContrastive LearningGraph

🎯 What it does: Propose a multi-scale explainer, MSExplainer, to simultaneously extract key subgraphs at different granularities to explain the predictions of graph neural networks.

Multi-Scale Wavelet Transformers for Operator Learning of Dynamical Systems

Xuesong Wang (Commonwealth Scientific and Industrial Research Organisation), Edwin V. Bonilla (Commonwealth Scientific and Industrial Research Organisation)

OptimizationTransformerDiffusion modelAuto EncoderTime SeriesPhysics Related

🎯 What it does: Proposed a multi-scale wavelet transform Transformer (MSWT), which learns dynamics in the tokenized wavelet domain, using wavelet attention and downsampling/upsampling of retained frequency bands to alleviate spectral bias in neural operators;

Multi-Task Bayesian In-Context Learning

Qingyang Zhu (New York University), Kyunghyun Cho (New York University)

Computational EfficiencyRepresentation LearningMeta LearningTransformerPrompt EngineeringMixture of ExpertsFlow-based ModelTextTabularTime Series

🎯 What it does: This paper proposes a Multi-Task Bayesian In-Context Learning framework, which utilizes Transformer to explicitly encode prior information as a prefix dataset, and achieves efficient hierarchical Bayesian prediction that can be adjusted at test time through a single forward pass.

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks

Shyam Sundhar Ramesh (University College London Electrical And Electronic Engineering), Ilija Bogunovic (University College London Centre For Ai)

OptimizationComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsText

🎯 What it does: Propose the Multi-Task GRPO (MT-GRPO) algorithm to achieve reliable and balanced performance improvements for large language models across various reasoning tasks;

Multi-task Linear Regression without Eigenvalue Lower Bounds: Adaptivity, Robustness, and Safety

Seok-Jin Kim (Columbia University)

OptimizationTabularTime Series

🎯 What it does: A multi-task linear regression estimator suitable for tasks with outliers is proposed, using matrix-weighted regularization and introducing a balance constant to overcome the dependence on traditional lower bounds.

Multi-timescale Reinforcement Learning by Value Reconstruction

Zhan Su (Peking University), Fanqi Shen (Zhejiang University)

TransformerReinforcement LearningAuto EncoderTime SeriesSequential

🎯 What it does: Propose a multi-discount-factor multi-timescale reinforcement learning framework, achieving value learning consistency through multi-scale value estimation and reward reconstruction, and driving policy optimization by cross-attention adaptive fusion of Q-values across different time scales.

Multi-View Causal Discovery without Non-Gaussianity: Identifiability and Algorithms

Ambroise Heurtebise (Inria), Aapo Hyvarinen (University Of Helsinki)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningScore-based ModelFlow-based ModelAuto EncoderContrastive LearningGaussian SplattingTabularTime SeriesBiomedical DataMagnetic Resonance ImagingElectrocardiogramReview/Survey Paper

🎯 What it does: A multi-view linear structural equation model, LiMVAM, is proposed, achieving identifiability of causal ordering and parameters based solely on second-order statistics. Subsequently, two novel recursive residual algorithms, PairwiseLiMVAM and DirectLiMVAM, are introduced based on this model, and the existing ICA-LiMVAM method is improved.

Multi-view Consistent Latent Action Learning for World Modeling and Control

Shenghua Wan (Nanjing University), De-Chuan Zhan (Nanjing University)

Knowledge DistillationRepresentation LearningRobotic IntelligenceReinforcement Learning from Human FeedbackTransformerAuto EncoderGenerative Adversarial NetworkContrastive LearningWorld ModelImageVideo

🎯 What it does: Propose the MuCoLA framework, which achieves perspective-invariant, semantically consistent world model control through multi-perspective consistency learning for self-supervised latent action representation.

Multi-Way Representation Alignment

Akshit Achara (King's College London), Donato Crisostomi (Sapienza University of Rome)

RetrievalRepresentation LearningAuto EncoderGenerative Adversarial NetworkContrastive LearningImageTextMultimodality

🎯 What it does: This study investigates the unified alignment of multi-model representation spaces, proposing a framework that maps all models into a shared reference space, and on this basis, designs Geometry-Corrected Procrustes Alignment (GCPA) to balance geometric fidelity and retrieval consistency.

MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety

Jialin Song (Simon Fraser University), Jianfeng Gao (Microsoft Research)

Safty and PrivacyAdversarial AttackTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Constructed a large-scale, topic-diverse multi-round jailbreak benchmark called MultiBreak, which continuously generates high-quality and diverse attack dialogues by utilizing active learning and uncertainty-guided rewriting.

Multicalibration Yields Better Matchings

Riccardo Colini Baldeschi (Meta Central Applied Science), Niek Tax

OptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyTabular

🎯 What it does: This paper proposes a post-processing method using multicalibration to improve the decision-making of max-weight matching based on predictions, and provides theoretical guarantees and sample complexity; by iteratively enhancing the predictor, it ensures that the optimal matching based on calibrated predictions is within a pre-set error ε of the expected utility achieved by selecting the best algorithm from any given matching algorithm family C using the original predictor.

MultiHal: Multilingual Dataset for Knowledge-Graph Grounded Evaluation of LLM Hallucinations

Ernests Lavrinovics (Aalborg University), Johannes Bjerva (Aalborg University)

Explainability and InterpretabilityKnowledge DistillationData-Centric LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextGraphBenchmarkRetrieval-Augmented Generation

🎯 What it does: Constructed a multilingual, multi-hop knowledge graph supported LLM hallucination evaluation dataset called MultiHal, filling the gap left by existing English-centric evaluation datasets that lack structured knowledge and multilingual support.

Multilingual Safety Alignment via Representation-Space Separability

Dan Shi (Tianjin University), Deyi Xiong (Tianjin University)

OptimizationSafty and PrivacyRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelContrastive LearningTextMultimodalityBenchmark

🎯 What it does: Propose a security alignment method called SMO based on the separability of multilingual representation spaces, leveraging well-separated features with good security geometry in the main language (e.g., English) to enhance the rejection behavior of low-resource languages;

Multilingual Safety Alignment Via Sparse Weight Editing

Jiaming Liang (Xidian University), Handing Wang (Xidian University)

Safty and PrivacyTransformerLarge Language ModelPrompt EngineeringText

🎯 What it does: This paper proposes a training-free multilingual safe alignment method, which utilizes sparse weight editing to transfer safe representations from high-resource languages to low-resource languages.

Multilingual Unlearning in LLMs: Transfer, Dynamics, and Reversibility

Chaoyi Xiang (University of Melbourne), Lea Frermann (University of Melbourne)

Federated LearningSafty and PrivacyExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningAdversarial AttackTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringContrastive LearningTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This paper studies the 'unlearning' technique for multilingual large language models (LLMs), which involves removing sensitive facts learned by the model without retraining, and explores the transferability, dynamic changes, and reversibility of unlearning across different languages.

MultiLoReFT: Decoupling Shared and Modality-Specific Subspaces in Multimodal Learning via Low-Rank Representation Fine-Tuning

Sana Tonekaboni (Massachusetts Institute of Technology), Caroline Uhler (Massachusetts Institute of Technology)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerSupervised Fine-TuningAuto EncoderContrastive LearningImageTextMultimodalityAudio

🎯 What it does: Based on pre-trained unimodal encoders, low-rank projection is used to fine-tune multi-modal representations, automatically separating the shared subspace and modality-specific subspaces, generating interpretable and embeddable multi-modal features for downstream tasks.

Multimarginal flow matching with optimal transport potentials

Raghav Kansal (Bexorg Inc), Bradley R Parry

OptimizationComputational EfficiencyFlow-based ModelPoint CloudTime SeriesBiomedical DataStochastic Differential Equation

🎯 What it does: Introduces a soft potential-based multi-margin flow matching framework called OTP-FM, which learns continuous-time flow mappings through conditional OT's conditional solution learning;

Multimodal Crystal Flow: Any-to-Any Modality Generation for Unified Crystal Modeling

Kiyoung Seong (KAIST), Changyoung Park (LG AI Research)

GenerationDrug DiscoveryTransformerDiffusion modelScore-based ModelFlow-based ModelContrastive LearningMultimodalityGraphTabularBenchmarkPhysics Related

🎯 What it does: Proposes Multimodal Crystal Flow (MCFlow), a unified multimodal flow model for crystal generation tasks (structure prediction, de novo generation, atom-type generation under structural conditions), achieving conditional generation for any modality by separating the time axis.

Multimodal Fact-Level Attribution for Verifiable Reasoning

David Wan (UNC Chapel Hill), Mohit Bansal (UNC Chapel Hill)

Explainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringVision Language ModelVideoTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-ThoughtAudio

🎯 What it does: Propose the MURGAT benchmark, requiring multimodal large language models to generate answers with explicit reasoning processes and precise timestamps/modal references based on heterogeneous inputs such as videos, audio, and charts; meanwhile, design the MURGAT-SCORE automatic evaluation process.

Multimodal Function Vectors for Visual Relations

Shuhao Fu (University of California Los Angeles), Hongjing Lu (University of California Los Angeles)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerSupervised Fine-TuningPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodality

🎯 What it does: Proposed a visual relationship extraction and manipulation method based on function vectors (FV), and implemented controllable intervention for relationship reasoning in large multimodal models.

Multimodal Fusion via Self-Consistent Task-Gradient Fields

Jiayu Xiong (Huaqiao University), Zhouqiang Jiang (University of Osaka)

Representation LearningTransformerSupervised Fine-TuningAuto EncoderContrastive LearningImageVideoMultimodalityAudio

🎯 What it does: Propose Self-Consistent Field Autoencoder (SCFAE), which jointly optimizes task loss and reconstruction loss on the same representation, achieving gradient conflict-free and information-complete multi-modal fusion.

Multimodal Latent Language Modeling with Next-Token Diffusion

Yutao Sun (Tsinghua University), Furu Wei (Microsoft Research)

GenerationData SynthesisRepresentation LearningTransformerLarge Language ModelDiffusion modelScore-based ModelAuto EncoderImageTextMultimodalityAudio

🎯 What it does: This paper proposes LatentLM, which uses a causal Transformer combined with VAE and next-token diffusion to achieve unified generation and understanding of continuous and discrete modalities.

Multimodal Meta-Verifier with Explicit Structured Recalibration

Xinchen Zhang (Tsinghua University), Ling Yang (Princeton University)

Explainability and InterpretabilityMeta LearningTransformerReinforcement LearningAgentic AIPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodality

🎯 What it does: Propose a multi-modal meta-verification framework, and train the OmniVerifier-M1 visual verifier based on this framework, then develop the M1-TTS agent generation system to achieve region-level self-correction.

Multimodal Nested Learning for Decoupled and Coordinated Optimization

Yanglin Feng (Sichuan University), Peng Hu (Sichuan University)

OptimizationRepresentation LearningData-Centric LearningTransformerContrastive LearningImageVideoTextMultimodalityPoint CloudAudio

🎯 What it does: Proposed a multi-modal framework called MoNet based on multi-layer nested learning, which addresses the problem of modal imbalance by decomposing multi-modal learning into outer-layer decoupled encoding and inner-layer coordinated fusion.

Multimodal Scaling Laws for Task & Data-Optimized Models of Visual Cortex

Abdulkadir Gokce (École Polytechnique Fédérale de Lausanne), Martin Schrimpf (École Polytechnique Fédérale de Lausanne)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningConvolutional Neural NetworkTransformerSupervised Fine-TuningAuto EncoderContrastive LearningMultimodalityBiomedical DataMagnetic Resonance ImagingReview/Survey Paper

🎯 What it does: Systematically studied the scaling laws of model-brain alignment of visual models across multiple brain recording modalities (electrophysiology, fMRI, EEG, MEG), evaluating the impact of pre-training scale, neural fine-tuning, and the number of mapping training samples on alignment.

Multiple Choice Learning of Low-Rank Adapters for Language Modeling

Victor Letzelter (Télécom Paris), Patrick Perez (Kyutai)

GenerationData-Centric LearningTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsContrastive LearningImageTextMultimodalityAudio

🎯 What it does: Based on pre-trained language models, multiple low-rank adapters are added and a winner-takes-all loss is adopted to generate diverse and reasonable sentences for the same input.

Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining

Rupert Mitchell (TU Darmstadt), Kristian Kersting (TU Darmstadt)

Computational EfficiencyRepresentation LearningTransformerLarge Language ModelMixture of ExpertsContrastive LearningTextRetrieval-Augmented Generation

🎯 What it does: Proposes Multipole Semantic Attention (MuSe), a method that achieves efficient softmax attention in long-sequence pretraining through semantic clustering and multipole approximation.

MultiPriv: Benchmarking Individual-Level Privacy Reasoning in Vision-Language Models

Xiongtao Sun (Xidian University), Wei Yang Bryan Lim (Nanyang Technological University)

Safty and PrivacyExplainability and InterpretabilityTransformerVision Language ModelImageTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Construct the MultiPriv benchmark to systematically evaluate the ability of vision-language models in individual privacy reasoning.

Multivariate Distributional Reinforcement Learning Using Sliced Divergences

Baptiste Debes (KU Leuven), Tinne Tuytelaars (KU Leuven)

Reinforcement LearningScore-based ModelTabularTime SeriesSequentialBenchmark

🎯 What it does: Extend distributed reinforcement learning from single-dimensional rewards to multi-dimensional rewards, proposing a measurable and trainable multi-dimensional return distribution through projection slicing.

Multiview Self-Representation Learning across Heterogeneous Views

Jie Chen (Sichuan University), Xi Peng (Sichuan University)

Domain AdaptationRepresentation LearningConvolutional Neural NetworkTransformerContrastive LearningImage

🎯 What it does: Propose the MSRL method, which performs self-representation learning through heterogeneous views generated by multiple pre-trained models to learn invariant representations.

Muon in Associative Memory Learning: Training Dynamics and Scaling Laws

Kaifei Wang (Peking University), Liwei Wang (Peking University)

OptimizationComputational EfficiencyData-Centric LearningImageTextPhysics Related

🎯 What it does: This paper reveals the acceleration mechanism and scaling laws of the Muon optimizer under frequency-imbalanced data by analyzing its training dynamics in linear associative memory models.

MuonSSM: Orthogonalizing State Space Models for Sequence Modeling

Thai Khanh Nguyen, Cuong Pham (Posts and Telecommunications Institute of Technology)

ClassificationRecognitionSegmentationRetrievalComputational EfficiencyRepresentation LearningTransformerLarge Language ModelMixture of ExpertsContrastive LearningOptical FlowImageVideoTextMultimodalityTime SeriesPhysics RelatedStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Propose MuonSSM, a general framework for stabilizing long sequence learning by incorporating momentum paths and single-step Newton-Schulz normalization into state space models.

MUSA-PINN: Multi-scale Weak-form Physics-Informed Neural Networks for Fluid Flow in Complex Geometries

Weizheng Zhang (Shandong University), Lin Lu (Shandong University)

Diffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningMeshPhysics Related

🎯 What it does: Propose a multi-scale weak form physics-informed neural network, MUSA-PINN, for solving Navier-Stokes PDEs in internal flows within complex geometries.

MUSE: Resolving Manifold Misalignment in Visual Tokenization via Topological Orthogonality

Panqi Yang (Xi'an Jiao Tong University), Yongqiang Ma (Xi'an Jiao Tong University)

Image TranslationRestorationGenerationRepresentation LearningTransformerVision Language ModelDiffusion modelAuto EncoderContrastive LearningImageTextMultimodality

🎯 What it does: Propose the MUSE framework, which unifies visual segmentation by achieving topological orthogonality between structural and semantic gradients, resolving the zero-sum dilemma between reconstruction and understanding.

MusicDET: Zero-Shot AI-Generated Music Detection

Chaolei Han (Southeast University), Jie Gui (Southeast University)

Anomaly DetectionTransformerDiffusion modelScore-based ModelFlow-based ModelAudio

🎯 What it does: This paper proposes a zero-shot AI-generated music detection method called MusicDET, which is trained using only real music and learns the probability distribution of real music through frequency-guided normalizing flow models.

Must All Negatives Be Pushed Away Equally? Uncertainty-Aware Cross-View Geo-Localization via Normal Inverse Gamma Distribution

Songsong Ouyang (Shenzhen University), Yingying Zhu (Shenzhen University)

RetrievalDomain AdaptationRepresentation LearningTransformerContrastive LearningImage

🎯 What it does: Proposes a NIG distribution based on deep evidence regression to estimate the environmental complexity around positive samples, and dynamically softens the InfoNCE objective using this uncertainty to reduce the phenomenon of semi-positive samples being overly pulled apart, thus significantly improving the generalization performance of cross-view geolocation.

MutAtlas: A PDB-Wide Energy-Guided Atlas of Protein Mutation Effects

Ruihan Guo (University of Illinois Urbana-Champaign), Ge Liu (University of Illinois Urbana-Champaign)

Knowledge DistillationDrug DiscoveryProtein Structure PredictionTransformerLarge Language ModelContrastive LearningGraphBiomedical DataBenchmark

🎯 What it does: Constructed a comprehensive PDB-wide benchmark dataset of single-point mutations, aligning the mutation signals from FoldX physical energy, ESM2 language model, and ESM-IF inverse folding model into a unified site-level 20-dimensional logit representation, and proposed a cross-source preference distillation framework without experimental labels, achieving multi-source information fusion through soft consistency weighted with inconsistency regularization;

MV-FGAD: Towards Efficient and Effective Federated Graph Anomaly Detection via Multi-view Learning

Junyi Yan (National University of Defense Technology), Xinwang Liu (National University of Defense Technology)

Anomaly DetectionFederated LearningGraph Neural NetworkContrastive LearningGraph

🎯 What it does: Propose the MV-FGAD framework, which utilizes multi-view learning and Mahalanobis distance to achieve federated graph anomaly detection;

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs

Huiyi Chen (University of Illinois Chicago), Lu Cheng (University of Illinois Chicago)

Explainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringVision Language ModelDiffusion modelContrastive LearningImageTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This paper proposes MVI-Bench, a comprehensive benchmark for evaluating the robustness of large vision-language models (LVLMs) under misleading visual inputs, and designs the MVI-Sensitivity metric to quantify the model's sensitivity to visual deception.

MVISTA-4D: View-Consistent 4D World Model with Test-Time Action Inference for Robotic Manipulation

Jiaxu Wang (Chinese University of Hong Kong), Xiangyu Yue (Chinese University of Hong Kong)

GenerationOptimizationRobotic IntelligenceTransformerVision-Language-Action ModelDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningWorld ModelImageVideoMultimodalityPoint Cloud

🎯 What it does: Proposes a multi-view 4D world model based on cross-modal and cross-view fusion, which can generate view-consistent RGB-D sequences from single-view RGB-D inputs, and achieve robot manipulation through trajectory latent optimization and residual inverse dynamics.

MVP-LAM: Learning Action-Centric Latent Action via Cross-Viewpoint Reconstruction

Jung Min Lee (Seoul National University), Jungwoo Lee (Seoul National University)

Representation LearningRobotic IntelligenceTransformerVision Language ModelAuto EncoderContrastive LearningVideoMultimodality

🎯 What it does: Researchers propose a potential action learning framework for multi-view cross-view reconstruction called MVP-LAM, which learns discrete potential actions with high mutual information about real action information by training a vector-quantized VAE on multi-view videos.

MVR-cache: Optimizing Semantic Caching via Multi-Vector Retrieval and Learned Prompt Segmentation

Ali Noshad (Peking University), Yinjun Wu (Peking University)

RetrievalOptimizationRepresentation LearningRecurrent Neural NetworkTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: Designed and implemented a semantic cache system called MVR-cache, which improves cache hit rate while maintaining correctness guarantees by combining learning-based prompt segmentation with multi-vector retrieval (MVR).

N2M: Bridging Navigation and Manipulation by Learning Pose Preference from Rollout

Kaixin Chai (KAIST), Joseph J Lim

Robotic IntelligenceTransformerReinforcement LearningVision-Language-Action ModelContrastive LearningImagePoint Cloud

🎯 What it does: Propose the N2M module, which predicts the preferred robot base pose distribution for executable learning-based operation strategies using RGB point clouds, and directly performs supervision based on the strategy rollout results;

NAACA: Training-Free NeuroAuditory Attentive Cognitive Architecture with Oscillatory Working Memory for Salience-Driven Attention Gating

Zhongju Yuan (Ghent University), Dick B.M. Botteldooren (Ghent University)

Anomaly DetectionTransformerAuto EncoderContrastive LearningMultimodalityAudio

🎯 What it does: Propose an untrained neural auditory attention architecture, NAACA, which achieves adaptive attention gating for long audio streams through oscillatory working memory (OWM).

Names Don’t Matter: Symbol-Invariant Transformer for Open-Vocabulary Learning

İlker Işık (Boston University), Wenchao Li (Boston University)

Computational EfficiencyRepresentation LearningAI Code AssistantTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextSequential

🎯 What it does: Designed a symbol-invariant Transformer structure that achieves literal equivalence (alpha-equivalence) invariance for interchangeable symbols (such as bound variables, atomic propositions) and supports vocabulary expansion during inference.

NanoFLUX: Distillation-Driven Compression of Large Text-to-Image Generation Models for Mobile Devices

Ruchika Chavhan (Samsung AI Center), Abhinav Mehrotra (Samsung AI Center)

GenerationCompressionKnowledge DistillationTransformerDiffusion modelAuto EncoderImageText

🎯 What it does: The 17B parameter FLUX.1‑Schnell model is compressed to a 2.4B parameter NanoFLUX model through multi-stage distillation and structural pruning, achieving real-time generation of 512×512 images in 2.5 seconds on mobile devices.

NanoQuant: Efficient Sub-1-Bit Quantization of Large Language Models

Hyochan Chong (Samsung Research), Minseop Choi (Samsung Research)

CompressionComputational EfficiencyKnowledge DistillationReinforcement Learning from Human FeedbackTransformerLarge Language ModelAuto EncoderContrastive LearningText

🎯 What it does: Proposes NANOQUANT, a post-training quantization (PTQ) method that can compress LLMs to 1-bit and sub-1-bit;

NanoSpec: Accelerating Speculative Decoding using Minimalist In-Context Vocabularies

Zhiyang Chen (Peking University), Yun Ma (Peking University)

Computational EfficiencyAI Code AssistantTransformerLarge Language ModelPrompt EngineeringText

🎯 What it does: Accelerate the LM head computation during speculative decoding by dynamically constructing a minimal, context-aware vocabulary at each step, achieving significant improvements in inference speed without compromising the final output quality.

NaRA: Noise-Aware LoRA for Parameter-Efficient Fine-Tuning of Diffusion LLMs

Shuaidi Wang (Southern University of Science and Technology), Yu Zhang (Southern University of Science and Technology)

Computational EfficiencyRepresentation LearningData-Centric LearningTransformerLarge Language ModelPrompt EngineeringDiffusion modelImageText

🎯 What it does: Propose a noise-aware low-rank adaptation method (NaRA), which dynamically generates the core matrix according to the noise level during the diffusion process through a globally shared lightweight hypernetwork, achieving parameter-efficient fine-tuning of diffusion large language models.

Narrowing the ANN–SNN Gap for Continuous 1D Temporal Signal Classification with Multi-Scale Temporal Encoding and Sparsity-Regularized Transform Encoding

Qi Sun (Xidian University), Biao Hou (Xidian University)

ClassificationComputational EfficiencyConvolutional Neural NetworkRecurrent Neural NetworkSpiking Neural NetworkTransformerSupervised Fine-TuningContrastive LearningTime SeriesBiomedical DataBenchmarkAudio

🎯 What it does: The study implements an SNN for efficient classification of continuous one-dimensional temporal signals using multi-scale time encoding and sparse regularization transform encoding under short time step lengths.

Nash Equilibria in Games with Playerwise Concave Coupling Constraints: Existence and Computation

Philip Jordan (EPFL), Maryam Kamgarpour (EPFL)

OptimizationReinforcement Learning

🎯 What it does: Studied the existence and computational methods of Nash equilibria in games with concave coupling constraints among players.

Native Active Perception as Reasoning for Omni-Modal Understanding

Zhenghao Xing (Chinese University of Hong Kong), Pheng-Ann Heng (Chinese University of Hong Kong)

Autonomous DrivingComputational EfficiencyRepresentation LearningTransformerSupervised Fine-TuningReinforcement LearningAgentic AIVision-Language-Action ModelVideoTextMultimodalityAudio

🎯 What it does: Propose OmniAgent, which views multimodal video understanding as an active perception reasoning process, adopting an observe-think-act (OOTA) cycle to achieve adaptive information extraction.

Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning

Tong Wu (State Key Laboratory of General Artificial Intelligence Bigai), Zilong Zheng (State Key Laboratory of General Artificial Intelligence Bigai)

Computational EfficiencyKnowledge DistillationTransformerLarge Language ModelReinforcement LearningPrompt EngineeringText

🎯 What it does: Propose Native Parallel Reasoner (NPR), enabling LLM to self-evolve real parallel reasoning capabilities without teacher supervision.

Native Spatio-Temporal 4D Variational Autoencoder

Lihe Ding (Chinese University of Hong Kong), Tianfan Xue (Chinese University of Hong Kong)

GenerationData SynthesisRepresentation LearningTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningGaussian SplattingOptical FlowImageVideoPoint CloudMesh

🎯 What it does: Proposed a variational autoencoder (Native 4D VAE) that operates directly in the native 4D dynamic voxel space, capable of simultaneously encoding geometry and color, achieving high-quality reconstruction of complete or partial 4D content;

Natural Hypergradient Descent: Algorithm Design, Convergence Analysis, and Parallel Implementation

Deyi Kong (University of Minnesota), Shancong Mou (University of Minnesota)

OptimizationFederated LearningComputational EfficiencyHyperparameter SearchReinforcement LearningDiffusion modelScore-based ModelFlow-based ModelRectified FlowNeural Radiance FieldAuto EncoderGenerative Adversarial NetworkContrastive LearningImageTabularTime SeriesSequentialBiomedical DataBenchmarkPhysics Related

🎯 What it does: Proposed a two-level optimization algorithm called Natural Hypergradient Descent (NHGD), which uses the inverse of the empirical Fisher information matrix (EFIM) to approximate the inverse of the inner Hessian, thereby achieving a parallel optimize-and-approximate structure.

NaviAgent: Graph‑Driven Bilevel Planning for Scalable Tool Orchestration

Yan Jiang (JD.COM), Ai Han (JD.COM)

Autonomous DrivingOptimizationAI Code AssistantReinforcement Learning from Human FeedbackGraph Neural NetworkTransformerLarge Language ModelWorld ModelTextGraphBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposes the NaviAgent dual-layer planning framework, separating task planning from tool execution, and achieving scalable toolchain orchestration through a graph-driven Tool World Navigation Model (TWNM).

NaviCache: Test-Time Self-Calibration Caching for Video Generation

Zheqi Lv (Zhejiang University), Fei Wu (Zhejiang University)

GenerationComputational EfficiencyTransformerDiffusion modelVideo

🎯 What it does: Propose a framework called NaviCache for achieving test-time self-calibrating caching in video diffusion models, accelerating inference by dynamically tracking feature evolution.

NAVIGATE: Evaluating Visual-Guided Search Decision-Making on the Open Web

YaoQi Fan (Nanjing University), Tong Lu (Nanjing University)

Supervised Fine-TuningVision Language ModelImageTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed the NAVIGATE benchmark for evaluating vision-guided open-network search decisions.

Navigating Massive Visual Context in Retrieval-Augmented Generation via Multimodal Memory Graph

Qiuchen Wang (Nanyang Technological University), Ruixue Ding (Alibaba Group)

RetrievalComputational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackGraph Neural NetworkTransformerLarge Language ModelReinforcement LearningPrompt EngineeringVision Language ModelDiffusion modelScore-based ModelContrastive LearningImageVideoTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose the VimRAG framework, which unifies retrieval, perception, and reasoning in multi-modal retrieval-augmented generation tasks into a dynamic multi-modal memory graph, effectively managing massive visual contexts.

Navigating the Energy Landscape of Collaboration: Multi-Agent Communication Graph Generation via Score-Based Diffusion

GuanHao Zhao (University of Science and Technology of China), Enhong Chen (University of Science and Technology of China)

OptimizationAI Code AssistantReinforcement Learning from Human FeedbackGraph Neural NetworkTransformerDiffusion modelScore-based ModelTextBenchmark

🎯 What it does: This paper proposes a multi-agent communication graph generation framework called MAGE based on energy minimization.

Navigating the Flatlands: Dual Adaptive Sharpness-Aware Minimization for Domain Generalization

Junwen He (Huazhong Agricultural University), Yulong Wang (Huazhong Agricultural University)

ClassificationDomain AdaptationOptimizationRepresentation LearningConvolutional Neural NetworkTransformerContrastive LearningImageBenchmark

🎯 What it does: This paper proposes a dual-adaptive sharpness-aware minimization algorithm, DA-SAM, for domain generalization. The method dynamically adapts the difficulty of the source domain to generate domain-specific scaling factors (DAS) and performs flattening exploration in multiple directions (AMDF), ultimately obtaining an equi-orientation flattened minimum point.

Navigating the Pareto Frontier of Alignment: Spectrum-Adaptive Fine-Tuning for LLMs

Yaoyou Fan (Peking University), Xunliang Cai (Meituan)

OptimizationRepresentation LearningData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningTextBenchmark

🎯 What it does: Propose Spectrum‑Adaptive Fine‑Tuning (SAFT), which performs post-training of LLMs by interpolating between standard SFT and Dynamic Fine‑Tuning (DFT) to achieve adaptive gradient weighting.

NavOL: Navigation Policy with Online Imitation Learning

Xiaofei Wei (Fudan University), Li Zhang (Fudan University)

Autonomous DrivingRobotic IntelligenceTransformerReinforcement LearningAgentic AIDiffusion modelSimultaneous Localization and MappingImagePoint CloudMeshBenchmark

🎯 What it does: Designed a visual navigation framework called NavOL based on online imitation learning, which utilizes a global planner to provide expert trajectories in real time, collects experience online, and rapidly iterates training in IsaacLab.

NBCG: Nash-Bargained Causal Game for Long-Tailed Multi-Label NLP

Jing Yang (Sun Yat-sen University), Keze Wang (Sun Yat-sen University)

ClassificationGraph Neural NetworkTransformerContrastive LearningText

🎯 What it does: View long-tailed multi-label text classification as a cooperative bargaining process among label coalitions. By using a Neural Structural Equation Model (Neural SEM), the causal dependency structure of labels is first learned, and then the Nash bargaining solution is adopted as a global objective to dynamically allocate learning credits to each coalition, solving the problem of dominant coalition capture caused by traditional weighted sum methods.

nD-RoPE: A Generalized RoPE for n-Dimensional Position Embedding

Boyang Li (New York University), Takahiro Yabe (New York University)

Representation LearningTransformerDiffusion modelScore-based ModelFlow-based ModelRectified FlowNeural Radiance FieldAuto EncoderGenerative Adversarial NetworkContrastive LearningImageVideoPoint Cloud

🎯 What it does: Proposes nD-RoPE, a unified rotation position encoding method applicable in arbitrary dimensions, avoiding axial decomposition, and achieving isotropy using frequency vectors from regular simplices;

Near-Minimax Multi-Objective RL under Predictable Adversarial Preferences and Preference-Free Exploration in Linear MDPs

Mingxi Hu (Fudan University), Meiling Yu (Nankai University)

Reinforcement LearningTabularBenchmark

🎯 What it does: This paper proposes a protocol-safe multi-objective reinforcement learning method that takes into account predictable opponent preferences and free exploration under reward-free preferences, achieving near-optimal statistical guarantees within the linear MDP framework.

Near-Optimal and Efficient First-Order Algorithm for Multi-Task Learning with Shared Linear Representation

Shihong Ding (Peking University), Cong Fang (Peking University)

OptimizationFederated LearningRepresentation LearningTabular

🎯 What it does: Proposed a two-stage gradient descent algorithm that jointly learns multi-task shared linear representations and task-specific parameters, achieving near-optimal statistical error and extremely fast iterative convergence speed

Near-Optimal Convergence of Accelerated Gradient Methods under Generalized and $(L_0,L_1)$-Smoothness

Alexander Tyurin (AXXX)

Optimization

🎯 What it does: Proposes two accelerated gradient methods (Algorithm 1 and Algorithm 2), and proves that they can achieve optimal convergence with O(√ℓ(0)·R/√ε) gradient calls when ε is small, under the generalized smoothness ℓ-smoothness (including the classical L-smooth and (L,L₀,1)-smooth as special cases).

Near-Optimal Dynamic Matching via Coarsening with Application to Heart Transplantation

Itai Zilberstein (Carnegie Mellon University), Tuomas Sandholm (Carnegie Mellon University)

OptimizationTabularBiomedical Data

🎯 What it does: Proposed a clustering-based online matching algorithm and applied it to heart transplant allocation.

Near-Optimal Private Linear Regression via Iterative Hessian Mixing

Omri Lev (Massachusetts Institute of Technology), Ashia C. Wilson (Massachusetts Institute of Technology)

OptimizationSafty and PrivacyContrastive LearningGaussian SplattingTabularBenchmark

🎯 What it does: A new differentially private linear regression algorithm called Iterative Hessian Mixing (IHM) is proposed, aiming to improve traditional differentially private ordinary least squares (DP-OLS) methods.

Near-Optimal Regret for KL-Regularized Multi-Armed Bandits

Kaixuan Ji (University of California, Los Angeles), Quanquan Gu (University of California, Los Angeles)

OptimizationReinforcement LearningTabular

🎯 What it does: Studied multi-armed bandits (MAB) with KL regularization objectives, proposed and analyzed an improved version of the KL-UCB algorithm, and provided upper and lower bounds under different regularization strengths, proving its near-optimality;

Near-Optimal Regret for Policy Optimization in Contextual MDPs with General Offline Function Approximation

Orin Levy (Tel Aviv University), Yishay Mansour (Tel Aviv University)

OptimizationReinforcement Learning

🎯 What it does: This paper proposes an algorithm called OPO-CMDP, which is based on policy optimization and can achieve an approximate optimal regret upper bound in stochastic context Markov decision processes (CMDP) with general offline function approximation.

Near-Universal Multiplicative Updates for Nonnegative Einsum Factorization

John Hood (University of Chicago), Aaron Schein (University of Chicago)

OptimizationComputational EfficiencyRepresentation LearningGraphTabularTime Series

🎯 What it does: Propose a general non-negative tensor decomposition algorithm called NNEinFact based on einsum, which supports various loss functions and custom models. Users only need to write a string to define the model and quickly fit it;

Needles in the Haystack: Addressing Signal Dilution Improves scRNA-seq Perturbation Response Modeling and Evaluation

Gabriel Mateo Mejia, Lucas Paulo de Lima Camillo (Shift Bioscience)

Data-Centric LearningDrug DiscoverySupervised Fine-TuningContrastive LearningBiomedical Data

🎯 What it does: Investigated the reasons why the average prediction baseline performs well in single-cell RNA sequencing perturbation response models, and proposed a weighted evaluation metric and training objective tailored for sparse signals.

Negative Sampling From the Ground Up: A Redesign for Recommendation

Yanbang Wang (Cornell University), Yanhong Wu (Meta)

Recommendation SystemGraph Neural NetworkContrastive LearningGraph

🎯 What it does: This paper proposes a negative sampling method based on the distribution of real negative samples. Through theoretical derivation, the optimal sampling distribution is obtained, and corresponding implementations are provided for two scenarios: unbiased and biased observations.

Negatives-Dominant Contrastive Learning for Generalization in Imbalanced Domains

Meng Cao (Nanjing University of Aeronautics and Astronautics), Songcan Chen (Nanjing University of Aeronautics and Astronautics)

ClassificationDomain AdaptationContrastive LearningImage

🎯 What it does: This paper proposes a negative sample dominated contrastive learning framework called NDCL for imbalanced domain generalization (IDG), which can achieve more robust decision boundaries under different domains and class imbalance conditions;

NEMO: Execution-Aware Optimization Modeling via Autonomous Coding Agents

Yang Song (C3 AI), Graham Neubig (Carnegie Mellon University)

OptimizationAI Code AssistantTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose the NEMO system, which utilizes autonomous coding agents (ACAs) to transform natural language decision-making problems into executable optimization models, constructing a simulation-optimization asynchronous verification loop;

Nested birth-death processes are competitive with neural networks as time-dependent models of protein evolution

Annabel Large (University of California, Berkeley), Ian Holmes (University of California, Berkeley)

Protein Structure PredictionRecurrent Neural NetworkTransformerContrastive LearningSequentialBiomedical Data

🎯 What it does: Propose a nested birth-death process model that integrates TKF92 into a multi-layer hybrid structure (fragments, domains, sites) and combines it with neural sequence embeddings for time-dependent protein evolution modeling, and compare it with pure neural sequence-to-sequence (seq2seq) models.

Nested Spatio-Temporal Time Series Forecasting

YingHao Ai, Yuan Qi (Shanghai Academy of AI for Science)

Autonomous DrivingOptimizationComputational EfficiencyGraph Neural NetworkTransformerReinforcement LearningAuto EncoderContrastive LearningGraphTime SeriesBenchmarkStochastic Differential Equation

🎯 What it does: This paper proposes the NEST framework, which utilizes macro-level regional prediction to guide fine-grained node prediction, constructing a hierarchical spatiotemporal sequence prediction model.

NetDiff: Graph Diffusion with Improved Global Capabilities to Generate and Update Mobile Network Topologies

Félix Marcoccia (Inria Paris), Paul Mühlethaler

GenerationData SynthesisOptimizationGraph Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningGraph

🎯 What it does: The paper proposes NetDiff, a node-conditioned denoising diffusion model for generating and updating mobile ad hoc network topologies with dual-slot transmission/reception allocations.

Networked Information Aggregation for Binary Classification

Mohammadhossein Bateni, Shayan Taherijam (University of California)

ClassificationFederated LearningSupervised Fine-TuningContrastive LearningTabular

🎯 What it does: A sequential log passing protocol was designed in a directed acyclic graph (DAG): each agent observes only partial features and receives logs from parent nodes. Using local features and parent logs, a logistic regression model is trained, and then the agent passes its logs to child nodes. Finally, the predictions of downstream agents are compared with the optimal logistic regression model that uses full information and full features.

NeUQI: Near-Optimal Uniform Quantization Parameter Initialization for Low-Bit LLMs

Li Lin (Peking University), Xiaojun Wan (Peking University)

OptimizationComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelPrompt EngineeringTextBenchmark

🎯 What it does: Propose a new low-bitwidth LLM uniform quantization parameter initialization method called NeUQI, significantly improving the performance of the quantized model.

Neural Attention Search Linear: Towards Adaptive Token-Level Hybrid Attention Models

Difan Deng (Leibniz University Hannover), Marius Lindauer (Leibniz University Hannover)

RetrievalComputational EfficiencyRepresentation LearningTransformerLarge Language ModelMixture of ExpertsContrastive LearningText

🎯 What it does: Propose a token-level adaptive hybrid attention model called NAtS-L, which can dynamically decide whether to apply linear attention or softmax attention to each token within the same layer;

Neural Collapse by Design: Learning Class Prototypes on the Hypersphere

Panagiotis Koromilas (University of Athens), Yannis Panagakis (University of Athens)

ClassificationRepresentation LearningContrastive LearningImage

🎯 What it does: In traditional supervised learning, category prototypes are learned on the unit hypersphere, and new contrastive learning losses NTCE and NONL are proposed. It is proven that SCL has already internally optimized a prototype classifier weighted by class means.

Neural Concept Verifier: Scaling Prover-Verifier Games via Concept Encodings

Berkant Turan (Zuse Institute Berlin), Sebastian Pokutta (Zuse Institute Berlin)

ClassificationExplainability and InterpretabilityTransformerVision Language ModelGenerative Adversarial NetworkContrastive LearningImage

🎯 What it does: This paper proposes the Neural Concept Verifier (NCV), a framework that combines concept encoding with the Prover-Verifier Game (PVG), to achieve verifiable interpretable classification on high-dimensional image data.

Neural Control: Adjoint Learning Through Equilibrium Constraints

Dezhong Tong (University of Michigan), M. Khalid Jawed (University of California)

OptimizationRobotic IntelligenceReinforcement LearningContrastive LearningTime SeriesSequentialPhysics RelatedStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: This paper proposes a control framework based on the adjoint-adjacent gradient and receding horizon control (RHC) strategy to address the long-term learning and control problems in multi-stable potential systems (such as deformable linear objects).

Neural Dispersion on Graphs

Ryien Hosseini (University of Chicago), Henry Hoffmann (University of Chicago)

GenerationData SynthesisOptimizationGraph Neural NetworkDiffusion modelScore-based ModelGenerative Adversarial NetworkContrastive LearningGraph

🎯 What it does: Propose a neural network framework called Neural Graph Dispersion that directly optimizes graph structural diversity, generating diverse graph collections through trajectory optimization.

Neural Feature Geometry Evolves as Discrete Ricci Flow

Moritz Hehl (Max Planck Institute for Mathematics in the Sciences), Melanie Weber (Harvard University)

ClassificationRepresentation LearningGraph Neural NetworkDiffusion modelScore-based ModelFlow-based ModelRectified FlowAuto EncoderGenerative Adversarial NetworkContrastive LearningImageGraph

🎯 What it does: This paper constructs a geometric graph in the feature space after data transformation (k-NN or r-neighborhood graph), and investigates how deep feedforward networks evolve feature representations during training through geometric transformations, and compares them with the dynamics of discrete Ricci flow.

Neural Honeytrace: Plug&Play Watermarking Framework against Model Extraction Attacks

Yixiao Xu (Beijing University of Posts and Telecommunications), Zhihong Tian (Guangzhou University)

Information TheoryClassificationSafty and PrivacyContrastive LearningImage

🎯 What it does: Proposes a no-retraining, plug-and-play watermarking framework called Neural Honeytrace to defend against model extraction attacks.

Neural Implicit Action Fields: From Discrete Waypoints to Continuous Functions for Vision-Language-Action Models

Haoyun Liu (Nanjing University), Sheng Zhong (Nanjing University)

Robotic IntelligenceTransformerLarge Language ModelSupervised Fine-TuningVision-Language-Action ModelNeural Radiance FieldMultimodality

🎯 What it does: Propose Neural Implicit Action Fields (NIAF), transforming the prediction of robot actions from discrete waypoints into continuous-time analytic functions, enabling high-resolution, differentiable action trajectories;

Neural Logistic Bandits

Seoungbin Bae (KAIST), Dabeen Lee (Seoul National University)

Reinforcement LearningImageTextTabular

🎯 What it does: Studied the neural logical bandit problem, with the main task being to use neural networks to learn an unknown reward function, which is embedded within a logical link function.

Neural Low-Discrepancy Sequences

Michael Etienne Van Huffel (Max Plank Institute for Intelligent Systems), T. Konstantin Rusch (Max Plank Institute for Intelligent Systems)

OptimizationHyperparameter SearchData-Centric LearningSupervised Fine-TuningTabularSequentialBenchmark

🎯 What it does: Propose a neural network-based low-discrepancy sequence (NEUROLDS) generation framework that can output low-discrepancy sequences at any length.

Neural Minimum Weight Perfect Matching for Quantum Error Codes

Yotam Peled (Ben-Gurion University), Eliya Nachmani (Ben-Gurion University)

OptimizationGraph Neural NetworkTransformerContrastive LearningGraphTabularPhysics Related

🎯 What it does: Propose a hybrid neural network decoder NMWPM, which utilizes graph neural networks to extract local symbolic information and combines Transformer to capture global dependencies, dynamically predicting the edge weights of MWPM;

Neural Modular Physics for Elastic Simulation

Yifei Li (Massachusetts Institute of Technology), Wojciech Matusik (Massachusetts Institute of Technology)

Graph Neural NetworkTransformerSupervised Fine-TuningMeshPhysics Related

🎯 What it does: Propose a neural modular physics framework (NMP), which replaces the corresponding steps in traditional elastic simulations with learnable constitutive modules and integration modules, achieving physically credible simulation of elastic objects.