ICML 2026 Papers — Page 40
International Conference on Machine Learning · 6554 papers
Online Continual Learning with Dynamic Label Hierarchies
Xinrui Wang (Nanjing University of Aeronautics and Astronautics), Songcan Chen (Nanjing University of Aeronautics and Astronautics)
ClassificationRecognitionRepresentation LearningMeta LearningConvolutional Neural NetworkTransformerSupervised Fine-TuningReinforcement LearningContrastive LearningImage
🎯 What it does: Propose the Dynamic Hierarchical Online Continuous Learning (DHOCL) framework, which allows labels to appear at any level and enables the hierarchy to evolve over time, and introduce the HALO model that achieves rapid adaptation and long-term stability through hierarchical prototype regularization and prediction aggregation;
Online Contract Design With Unknown Technology
Matteo Bollini (Politecnico di Milano), Alberto Marchesi (Politecnico di Milano)
OptimizationReinforcement Learning
🎯 What it does: Designed an online learning algorithm to learn the optimal contract sequentially from a set of candidate actions with unknown agent technology, aiming to minimize cumulative penalties.
Online Fair Division with Additional Information
Tzeh Yuan Neoh (Harvard University), Nicholas Teh (University of Oxford)
OptimizationFederated LearningReinforcement Learning from Human Feedback
🎯 What it does: Study the problem of online fair division, analyze the fairness constraints that can be achieved under scenarios with no additional information, only given standardized information, or frequency prediction, and provide corresponding algorithms and lower bounds.
Online Learning and Inference for Cox Proportional Hazards Model Using Renewable Sieve Estimation
Mengtong Hu (University of Michigan), Peter Song (University of Michigan)
OptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTabularBiomedical DataElectronic Health Records
🎯 What it does: Propose an online learning framework COLSA, which uses Renewable Sieve Estimation to approximate the baseline risk through basis functions based on Bernstein polynomials, replacing the traditional Cox partial likelihood, achieving online updates without the need to store historical data or a global risk set;
Online Learning with Recency: Algorithms for Sliding-window Streaming Multi-armed Bandits
Vladimir Braverman (Johns Hopkins University), Samson Zhou (Texas A&M University)
Reinforcement LearningTabularTime SeriesBenchmark
🎯 What it does: Studied the sliding window streaming multi-armed bandit (MAB) model, provided theoretical analysis and algorithms for pure exploration and reward minimization, and quantified the trade-off between memory and reward under limited memory constraints.
Online Linear Programming for Multi-Objective Routing in LLM Serving
Zixi Chen (Peking University), Zijie Zhou (HKUST)
OptimizationExplainability and InterpretabilityComputational EfficiencyLarge Language ModelReinforcement LearningText
🎯 What it does: Proposed a multi-objective online linear programming framework for request routing in large model inference services;
Online Packet Scheduling with Deadlines and Learning
Gianmarco Genalti (Politecnico di Milano), Vianney Perchet (Ensae)
OptimizationReinforcement Learning
🎯 What it does: Proposed the learning-based online knapsack scheduling problem K-OPSD, and designed multiple deterministic and randomized algorithms for it, studying the upper and lower bounds of α-regret and competitive ratio.
Online Robust Reinforcement Learning with General Function Approximation
Debamita Ghosh (University of Central Florida), Yue Wang (University of Central Florida)
OptimizationReinforcement LearningContrastive LearningTabular
🎯 What it does: Propose an online distributed robust reinforcement learning framework RFLϕ, which directly learns the optimal robust policy from interactive data using general function approximation, avoiding the need for prior generative models or large-scale offline data.
Online Rubrics Elicitation from Pairwise Comparisons
MohammadHossein Rezaei (Scale AI), Afra Feyza Akyürek (Scale AI)
Reinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose a framework called OnlineRubrics that dynamically mines and updates evaluation criteria (rubrics) during the reinforcement learning process. It generates paired comparisons between the current policy's response and the control policy's response, and uses an LLM to extract differences and generate new evaluation standards, which are then merged with existing rubrics for reward calculation.
Online Social Welfare Function-based Resource Allocation
Kanad Shrikar Pardeshi (Carnegie Mellon University), Aarti Singh (Carnegie Mellon University)
OptimizationReinforcement Learning
🎯 What it does: This paper proposes an online resource allocation framework based on the social welfare function (SWF), aiming to address the problem of allocating limited resources to a fixed population across multiple time steps.
OOVDet: Low-Density Prior Learning for Zero-Shot Out-of-Vocabulary Object Detection
Binyi Su (Hebei University of Technology), Haiyong Chen (Hebei University of Technology)
Object DetectionAnomaly DetectionTransformerPrompt EngineeringVision Language ModelDiffusion modelContrastive LearningGaussian SplattingImageText
🎯 What it does: Propose the OOVDet framework to achieve detection of known categories and reliable rejection of unknown categories under zero-shot conditions.
Op-CAD: Benchmarking and Investigating Operation-oriented CAD Generation
Yixue Bai (Hong Kong University of Science and Technology (Guangzhou)), Zeke Xie (Hong Kong University of Science and Technology (Guangzhou))
GenerationData SynthesisOptimizationTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextMultimodalityMeshBenchmarkChain-of-Thought
🎯 What it does: Constructed the Op-CAD dataset and defined the operation-oriented CAD generation task and evaluation framework;
Open Materials Generation with Inference-Time Reinforcement Learning
Philipp Höllmer, Stefano Martiniani (New York University)
GenerationDrug DiscoveryReinforcement Learning from Human FeedbackProtein Structure PredictionGraph Neural NetworkReinforcement LearningDiffusion modelScore-based ModelFlow-based ModelGraphTabularPhysics RelatedStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Proposes OMatG-IRL, a method that applies policy gradient reinforcement learning to continuous time crystal generation models during inference, marking the first application of reinforcement learning to the crystal structure prediction (CSP) task.
Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence
Jiahao Meng (Peking University), Zhuochen Wang (ByteDance)
Explainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision Language ModelVideoTextMultimodalityChain-of-Thought
🎯 What it does: Propose Open-o3-Video, a multimodal model that directly outputs timestamps, bounding boxes, and chain-of-thought reasoning in a single forward inference, achieving traceability in video reasoning.
Open-Text Aerial Detection: A Unified Framework For Aerial Visual Grounding And Detection
Guoting Wei (Nanjing University of Science and Technology), Rong Xiao (Intellifusion)
Image TranslationRestorationObject DetectionSegmentationGenerationData SynthesisAutonomous DrivingComputational EfficiencyRepresentation LearningConvolutional Neural NetworkTransformerPrompt EngineeringVision Language ModelDiffusion modelContrastive LearningImageTextMultimodalityRetrieval-Augmented Generation
🎯 What it does: This paper proposes a unified framework called OTA-Det, which integrates open-vocabulary aerial detection (OVAD) and remote sensing visual grounding (RSVG), achieving multi-granularity semantic understanding and multi-object detection.
Open-World LLM Logical Reasoning
Ye Mo (Zhejiang University), Philip Torr (University of Oxford)
Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningMeta LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposed the OpenIKLR framework, which enables logical reasoning under conditions of incomplete knowledge in an open world.
OpenDeception: Learning Deception and Trust in Human–AI Interaction via Multi-Agent Simulation
Yichen Wu (Fudan University), Min Yang (Fudan University)
Anomaly DetectionFederated LearningSafty and PrivacyExplainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningAgentic AIContrastive LearningTextBenchmark
🎯 What it does: Propose the OpenDeception framework, which achieves real-time monitoring and early warning of deception risks in open-ended human-computer dialogues by jointly evaluating the deception intent of AI (IntentNet) and user trust (TrustNet).
OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing
zhihong Chen, YiFan Zhang
GenerationData SynthesisTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringDiffusion modelImageTextMultimodalityBenchmarkRetrieval-Augmented Generation
🎯 What it does: Constructed a large-scale multi-modal training dataset called OpenGPT-4o-Image, containing 80k instruction-image pairs with hierarchical task decomposition for image generation and editing, covering 11 major domains and 51 subtasks, particularly adding challenging scenarios such as scientific images, complex instruction following, and multi-round editing.
OpenHA: A Series of Open-Source Hierarchical Agentic Models in Minecraft
Zihao Wang (Peking University), Yitao Liang (Peking University)
Robotic IntelligenceTransformerSupervised Fine-TuningReinforcement LearningAgentic AIMixture of ExpertsVision Language ModelVision-Language-Action ModelDiffusion modelAuto EncoderImageVideoTextSequentialBenchmarkChain-of-Thought
🎯 What it does: Proposed and open-sourced multiple hierarchical agent models (OpenHA) in Minecraft, and conducted large-scale benchmark evaluations on over 800 manually designed tasks.
OpenMAG: A Comprehensive Benchmark for Multimodal-Attributed Graph
Chenxi Wan (Beijing Institute of Technology), Guoren Wang (Beijing Institute of Technology)
Recommendation SystemExplainability and InterpretabilityComputational EfficiencyData-Centric LearningGraph Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningVision Language ModelContrastive LearningImageTextMultimodalityGraphBenchmark
🎯 What it does: A unified multi-modal attribute graph (MAG) evaluation benchmark named OpenMAG was constructed, integrating 19 datasets across six domains, 16 modal encoders, 24 MAG task models, and covering eight downstream tasks. It provides a five-dimensional evaluation framework ranging from necessity, data quality, effectiveness, robustness to efficiency.
OpenSage: Self-programming Agent Generation Engine
Hongwei Li (University Of California Santa Barbara), Dawn Song (University Of California Berkeley)
AI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose OpenSage ADK, enabling LLM to automatically generate agent topology, tool sets, and hierarchical graphical memory, supporting dynamic creation and parallel execution of sub-agents.
OpenTSLM: Time-Series Language Models for Reasoning over Multivariate Medical Text- and Time-Series Data
Patrick Langer (Stanford University), Paul Schmiedmayer (Stanford University)
Anomaly DetectionExplainability and InterpretabilityComputational EfficiencyRepresentation LearningData-Centric LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextTime SeriesBiomedical DataElectronic Health RecordsElectrocardiogramReview/Survey PaperChain-of-Thought
🎯 What it does: Proposed an open-source time series language model (OpenTSLM) framework that can locally process multivariate medical time series data within large language models and perform reasoning through natural language.
Operationalising the Superficial Alignment Hypothesis via Task Complexity
Tomás Vergara Browne, Marius Mosbach (Mila Quebec AI Institute)
CompressionComputational EfficiencyRepresentation LearningData-Centric LearningAI Code AssistantTransformerLarge Language ModelPrompt EngineeringText
🎯 What it does: This paper quantifies the surface adaptability of pre-trained language models by defining task complexity (the shortest program length required to achieve a given performance), proving that pre-training can significantly reduce the program length needed to achieve high performance.
Operator Splitting with Hamilton-Jacobi-based Proximals
Nicholas Di (Rice University), Samy Wu Fung (Colorado School of Mines)
OptimizationTabularBiomedical Data
🎯 What it does: Studies the application of a gradient-free proximal operator based on the Hamilton–Jacobi approximation (HJ-Prox) in splitting algorithms, and provides a unified convergence theory.
Ophiuchus: Incentivizing Tool-augmented ''Think with Images'' for Joint Medical Segmentation, Understanding and Reasoning
Yankai Jiang (Zhejiang University), Shihui Zhen
Image TranslationSegmentationExplainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AIVision Language ModelVision-Language-Action ModelImageMultimodalityBiomedical DataBenchmarkChain-of-Thought
🎯 What it does: Propose a multi-modal LLM tool enhancement framework called Ophiuchus, which can adaptively invoke visual tools (such as SAM2, BiomedParse, and scaling) during the reasoning process for fine-grained image analysis, thereby improving the performance of medical image question answering and segmentation.
OPIC: Enhancing Language Model Merging via Optimizing In-Context Capability
Jie He (National University of Defense Technology), Ji Wang (National University of Defense Technology)
OptimizationKnowledge DistillationRepresentation LearningHyperparameter SearchTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposed the OPIC framework, achieving model merging without dependency on a validation set, and enhancing the performance after fusion by maintaining the model's in-context learning (ICL) capability.
Opportunistic Expert Activation: Batch-Aware Expert Routing for Faster Decode Without Retraining
Costin-Andrei Oncescu (Harvard University), Ben Athiwaratkun (Together AI)
Computational EfficiencyKnowledge DistillationTransformerLarge Language ModelMixture of ExpertsText
🎯 What it does: To address the memory-constrained latency during the medium-batch decoding phase of large sparse expert models, this paper proposes a batch-aware expert activation (OEA) method that requires no retraining.
OPT-Engine: Benchmarking the Limits of LLMs in Optimization Modeling via Complexity Scaling
Yitian Chen (Cardinal Operations), Dongdong Ge (Shanghai Jiao Tong University)
OptimizationTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposed an expandable optimization modeling benchmark framework, OPT-Engine, for evaluating the performance of large language models in optimization modeling tasks ranging from linear programming to mixed integer programming.
Opt-Miner: Empowering Information-Seeking Agent with Tree-Guided Data Synthesis for Optimization Modeling
Haoyang Liu (University of Science and Technology of China), Feng Wu (University of Science and Technology of China)
Data SynthesisOptimizationTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: A tree-structured data synthesis and reinforcement learning framework, Opt-Miner, was constructed to enable LLMs to proactively retrieve and integrate external knowledge to complete complex optimization modeling tasks.
Opt-Verifier: Unleashing the Power of LLMs for Optimization Modeling via Dual-Side Verification
Haoyang Liu (University of Science and Technology of China), Jianye HAO
OptimizationTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Designed Opt-Verifier, an optimization modeling framework that utilizes LLMs for dual-side (structure and solution) verification, to automatically generate and correct mathematical optimization models.
OptiFluence: Principled Design of Privacy Canaries
Mohammad Yaghini (University of Toronto), Florian Tramèr (ETH Zurich)
OptimizationSafty and PrivacyConvolutional Neural NetworkRecurrent Neural NetworkContrastive LearningImage
🎯 What it does: Proposes the OptiFluence framework, which automatically generates high-detectability privacy canaries through bi-level optimization, achieving near-perfect membership inference detection on multiple datasets;
Optimal and Scalable MAPF via Multi-Marginal Optimal Transport and Schrödinger Bridges
Usman Khan, Joseph W Durham
OptimizationRobotic IntelligenceGraph
🎯 What it does: This paper proposes to view anonymous multi-agent pathfinding (MAPF) as a multi-marginal optimal transport (MMOT) problem, and derives integer optimal solutions by proving total solvability; subsequently, a scalable fractional transport is obtained through entropy regularization of the Schrödinger bridge, and integer solutions are recovered via projection.
Optimal Anytime Algorithms for Online Convex Optimization with Adversarial Constraints
Dhruv Sarkar (Indian Institute of Technology), Abhishek Sinha (Tata Institute of Fundamental Research)
Optimization
🎯 What it does: Proposed an online algorithm that learns a sequence of convex cost functions while approximately satisfying a sequence of convex constraints, without prior knowledge of the time horizon.
Optimal Attention Temperature Improves the Robustness of In-Context Learning under Distribution Shift in High Dimensions
Samet Demir (Koc University), Zafer Dogan (Koc University)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningMeta LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningText
🎯 What it does: This paper studies the impact of the Transformer attention temperature on the robustness of in-context learning (ICL) under distribution shift, and provides a closed-form expression for the optimal temperature, proving that appropriately adjusting the temperature can significantly improve ICL performance.
Optimal Bayesian Stopping for Efficient Inference of Consistent LLM Answers
Jingkai Huang (New York University), Zhengyuan Zhou (New York University)
OptimizationComputational EfficiencyTransformerLarge Language ModelTextBenchmark
🎯 What it does: Studied an adaptive stopping strategy based on Bayesian priors to efficiently identify the most frequently occurring answer (mode) in large language model (LLM) inference, thereby reducing sampling costs and improving accuracy.
Optimal conversion from Rényi Differential Privacy to $f$-Differential Privacy
Anneliese Riess (Helmholtz Munich), Georgios Kaissis (University of Potsdam)
Safty and Privacy
🎯 What it does: Proves the conjecture proposed by Zhu et al. (2022) in Appendix F.3, which states that among all transformation rules mapping R\'enyi differential privacy (RDP) configurations to effective hypothesis testing trade-offs f, the rule based on the intersection of single-order RDP privacy regions is optimal.
Optimal Decision-Making Based on Prediction Sets
Tao Wang (University of Pennsylvania), Edgar Dobriban (University of Pennsylvania)
Anomaly DetectionAutonomous DrivingOptimizationImageBiomedical Data
🎯 What it does: A decision-theoretic framework based on prediction sets is proposed, minimizing the expected loss under the constraint of prediction set coverage, and corresponding minimax optimal decisions and prediction set constructions are provided;
Optimal Design for Multinomial Logit Model with Applications to Best Assortment Identification
Joongkyu Lee (Seoul National University), Min-hwan Oh (Seoul National University)
Recommendation SystemOptimizationReinforcement Learning from Human FeedbackTabularBenchmark
🎯 What it does: Study the optimal experimental design under the multinomial logit model (MNL), propose two efficient solution methods, and apply them to an assortment identification algorithm.
Optimal Domain-Aware Privacy Mechanisms for Synthetic Data Generation
Sajani Vithana (Harvard University), Haewon Jeong (University of California, Santa Barbara)
Data SynthesisSafty and PrivacyDiffusion modelScore-based ModelContrastive LearningTextTabular
🎯 What it does: Designed and analyzed a domain-aware differential privacy mechanism (PUBMIX) for synthetic data generation using public data.
Optimal Fair Aggregation of Crowdsourced Noisy Labels using Demographic Parity Constraints
Gabriel Singer (University Paris Saclay), Argyris Kalogeratos (University Paris Saclay)
Recommendation SystemAnomaly DetectionOptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyData-Centric LearningContrastive LearningTextTabularReview/Survey Paper
🎯 What it does: Studied the fairness of crowd-labeled noisy labels during the aggregation process, provided fairness error upper bounds for Majority Vote and Bayes aggregators, and proposed a post-processing algorithm called FairCrowd, which can enforce ε-fair racial equality constraints on any aggregation rule.
Optimal Learning from Label Proportions with General Loss Functions
Lorne Applebaum (Google Research), Tomer Koren (Google Research and Tel Aviv University)
ClassificationOptimizationContrastive LearningImageTabular
🎯 What it does: Proposed a low-variance, unbiased proportion label learning (LLP) method that can learn with any (even unbounded) loss function under the condition of obtaining only batch label proportions;
Optimal Pricing for Data-Augmented AutoML Marketplaces
Minbiao Han (University of Chicago), sainyam galhotra
OptimizationFederated LearningHyperparameter SearchData-Centric LearningReinforcement LearningTabularTime SeriesBenchmark
🎯 What it does: This study proposes a data augmentation AutoML market compatible with existing cloud AutoML platforms, automatically providing external data augmentation to buyers and pricing it based on improvements in model quality.
Optimal Quantum Speedups for Repeatedly Nested Expectation Estimation
Yihang Sun (Stanford University), Jose Blanchet (Stanford University)
OptimizationFinance RelatedPhysics Related
🎯 What it does: Proposed a quantum algorithm for achieving an estimate with error ε in constant-depth repeated nested expectation (RNE) problems.
Optimal Rates for Feasible Payoff Set Estimation in Games
Annalisa Barbara (Bocconi University), Andrea Celli (Bocconi University)
OptimizationReinforcement Learning from Human FeedbackContrastive LearningReview/Survey Paper
🎯 What it does: The study investigates the sample complexity of recovering all sets of payoff matrices compatible with observed equilibrium behaviors in two-player games (feasible payoff sets).
Optimal Regret for Policy Optimization in Contextual Bandits
Orin Levy (Tel Aviv University), Yishay Mansour (Tel Aviv University)
OptimizationReinforcement LearningContrastive LearningTabularBenchmark
🎯 What it does: Proposed an online contextual multi-armed bandit (CMAB) algorithm based on policy optimization (PO) called OPO-CMAB, and proved that it achieves an optimal scheduling loss upper bound of ˜O(√K|A|log|F|) with high probability under a general offline function approximation framework.
Optimal Regularization for Performative Learning
Edwige Cyffers (CNRS), Marco Mondelli (Institute of Science and Technology Austria)
OptimizationFederated LearningData-Centric LearningTabularBenchmark
🎯 What it does: This paper studies how regularization can mitigate the negative impact of data distribution drift caused by model deployment under the performative learning framework.
Optimal Self-Consistency for Efficient Reasoning with Large Language Models
Austin Feng (Yale University), Ievgen Redko (Noah's Ark Lab)
OptimizationExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextChain-of-Thought
🎯 What it does: Studied the sample efficiency of self-consistency (Self-Consistency) in the reasoning of large language models and proposed an adaptive self-consistency blend called Blend-ASC.
Optimal Splitting of Language Models from Mixtures to Specialized Domains
Skyler Seto (Apple), David Grangier (Apple)
OptimizationComputational EfficiencyRepresentation LearningTransformerLarge Language ModelMixture of ExpertsContrastive LearningText
🎯 What it does: Proposes a split model training method for multi-domain pre-training data, separating general pre-training and domain-specific continuous pre-training (CPT), and optimizing them according to computational budget;
Optimal Stopping in Latent Diffusion Models
Yu-Han Wu (Sorbonne Université), Pierre Marion (Inria)
GenerationOptimizationHyperparameter SearchDiffusion modelScore-based ModelAuto EncoderImage
🎯 What it does: This paper investigates the relationship between the optimal stopping time and the latent space dimension in latent diffusion models (LDM), proving that early stopping of diffusion can improve sample quality in certain cases, and provides theoretical support;
Optimal structure learning and conditional independence testing
Ming Gao (University of Chicago), Bryon Aragam (University of Chicago)
OptimizationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningGraph Neural NetworkReinforcement LearningContrastive LearningGraphTabularBenchmark
🎯 What it does: Studied the minimal sample complexity relationship between graph structure learning and conditional independence testing, and provided the optimal sample complexity expression on polynomial forest models;
Optimal Top-$k$ Identification from Pairwise Comparisons
Motti Goldberger (Yale University), Nils Rudi (Yale University)
Recommendation SystemOptimizationReinforcement LearningContrastive LearningTabularBenchmark
🎯 What it does: This paper proposes a top-k identification algorithm that achieves fixed confidence under noisy pairwise comparisons, capable of returning the correct set of top k items while minimizing the expected number of comparisons.
Optimal Transport for LLM Reward Modeling from Noisy Feedback
Licheng Pan (Zhejiang University), Hao Wang (Zhejiang University)
OptimizationData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelContrastive LearningText
🎯 What it does: This paper proposes SelectiveRM, a reward model training framework based on optimal transport, designed for denoising and robust training on preference data with instance-dependent noise.
Optimal Transport Group Counterfactual Explanations
Enrique Valero-Leal (Universidad Politecnica de Madrid), Giuseppe Casalicchio (Ludwig Maximilian University of Munich)
OptimizationExplainability and InterpretabilityTabular
🎯 What it does: Propose a group-level counterfactual explanation framework based on optimal transport (OT) mapping, which generates counterfactual points for the entire group using a single generalizable function.
Optimal Transport under Group Fairness Constraints
Linus Bleistein (EPFL), Aurélien Bellet (Universite de Montpellier)
OptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTabular
🎯 What it does: Proposes a group fairness definition based on the target matching probability, and under the optimal transport framework, presents three methods to achieve fairness: the exact fairness FairSinkhorn algorithm, a convex optimization method with a penalty term, and a bi-level optimization method through cost learning.
Optimal Transport with Symmetry Groups
Jiechao Zhang (Xi'an Jiaotong University), Wei Zeng (Xi'an Jiaotong University)
OptimizationComputational EfficiencyRepresentation LearningDiffusion modelScore-based ModelFlow-based ModelAuto EncoderContrastive LearningOptical FlowImagePoint CloudGraphTabular
🎯 What it does: Proposes utilizing finite group symmetry to automatically identify orbits and solve optimal transport in the orbit space, significantly reducing computational complexity.
Optimal Transport–Guided Stochastic Control for Graph Combinatorial Optimization
Yang Huang (Institute of Automation, Chinese Academy of Sciences), Jian Cheng (Institute of Automation, Chinese Academy of Sciences)
OptimizationGraph Neural NetworkReinforcement LearningScore-based ModelGraphStochastic Differential Equation
🎯 What it does: Propose an OT (Optimal Transport)-guided stochastic control sampling framework, using continuous multilinear relaxation to solve graph combinatorial optimization problems (such as maximum independent set, maximum clique, maximum cut)
Optimal Unconstrained Self-Distillation in Ridge Regression: Strict Improvements, Precise Asymptotics, and One-Shot Tuning
Hien Dang (University of Texas at Austin), Alessandro Rinaldo (University of Texas at Austin)
OptimizationKnowledge DistillationHyperparameter SearchImageTabular
🎯 What it does: Theoretical aspects of self-distillation (SD) in ridge regression are studied, proving that under any non-stationary regularization parameter, the predictive risk can be strictly improved through optimal mixing weights, and providing an analytical expression for these weights;
Optimality of FSQ Tokens for Continuous Diffusion for Categorical Data with Application to Text-to-Speech
Vadim Popov (Huawei Noah's Ark Lab), Assel Yermekova (Huawei Noah's Ark Lab)
GenerationData SynthesisComputational EfficiencyTransformerDiffusion modelScore-based ModelContrastive LearningTextAudio
🎯 What it does: This paper investigates the potential latent space structure of continuous diffusion models on classification data, demonstrating that the FSQ quantization scheme enables more optimal training of the CDCD model, and based on this, proposes the first efficient zero-shot text-to-speech (TTS) model.
Optimization, Generalization and Differential Privacy Bounds for Gradient Descent on Kolmogorov–Arnold Networks
Puyu Wang (RPTU KaiserslauternLandau), Marius Kloft (RPTU KaiserslauternLandau)
OptimizationSafty and PrivacyImageTabularBiomedical Data
🎯 What it does: This paper studies the optimization, generalization, and upper bounds of privacy utility of gradient descent (GD) and differential privacy gradient descent (DP-GD) on two-layer Kolmogorov-Arnold networks (KAN), and provides theoretical and experimental guidance on width and iteration counts.
Optimized Deferral for Imbalanced Settings
Corinna Cortes (Google Research), Yutao Zhong (Google Research)
ClassificationOptimizationFederated LearningMixture of ExpertsImageText
🎯 What it does: This paper studies the problem of learning to defer in two-stage learning under an imbalanced expert environment, and proposes the MILD algorithm based on margin loss.
Optimizing Agentic Reasoning with Retrieval via Synthetic Semantic Information Gain Reward
Senkang Hu (Hong Kong JC STEM Lab of Smart City), Yuguang Fang (Hong Kong JC STEM Lab of Smart City)
RetrievalOptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningAgentic AITextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Design and implement the InfoReasoner framework, which utilizes synthetic semantic information gain rewards to optimize retrieval-based reasoning models.
Optimizing Diversity and Quality through Base-Aligned Model Collaboration
Yichen Wang (University of Chicago), Mina Lee (University of Chicago)
GenerationOptimizationComputational EfficiencyData-Centric LearningTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsTextRetrieval-Augmented Generation
🎯 What it does: Proposed a token-level collaboration framework (BACO) during inference, which dynamically routes between the base LLM and its alignment model in real-time to balance output diversity and quality.
Optimizing Few-Step Generation with Adaptive Matching Distillation
Lichen Bai (Hong Kong University of Science and Technology), Zeke Xie (Hong Kong University of Science and Technology)
GenerationKnowledge DistillationReinforcement Learning from Human FeedbackTransformerDiffusion modelScore-based ModelContrastive LearningImageVideo
🎯 What it does: Propose Adaptive Matching Distillation (AMD), a few-step diffusion model distillation framework based on reward model self-diagnosis, dynamic gradient regulation, and rebound potential sharpening, aimed at solving Forbidden Zones in traditional DMD and improving generation quality and robustness.
Optimizing Inference-Time Compute for Medical Reasoning via Uncertainty Quantification
Shaohao Rui (Shanghai Jiao Tong University), Xiaosong Wang (Shanghai Innovation Institute)
OptimizationExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBiomedical DataElectronic Health RecordsBenchmarkChain-of-Thought
🎯 What it does: Propose AdaThink-Med, a medical large language model framework that achieves adaptive reasoning through uncertainty-guided length calibration, enabling the model to dynamically determine the thinking length during inference based on problem difficulty.
Optimizing Network Simulation: Enhancing Performance Prediction Accuracy via Neural Architecture Search
ShaoChen He (Beijing University of Posts and Telecommunications), Jingyu Wang (Beijing University of Posts and Telecommunications)
OptimizationHyperparameter SearchNeural Architecture SearchRecurrent Neural NetworkTransformerReinforcement LearningTabularTime SeriesSequential
🎯 What it does: To address the tail latency distortion problem in network simulation, the ANAS framework is proposed to automatically search for high-accuracy distribution-aware network models.
Optimizing Rank for High-Fidelity Implicit Neural Representations
Julian McGinnis (Technical University of Munich), Benedikt Wiestler (Technical University of Munich)
RestorationData SynthesisSuper ResolutionOptimizationRepresentation LearningDiffusion modelScore-based ModelNeural Radiance FieldAuto EncoderImagePoint CloudMeshBiomedical DataComputed TomographyReview/Survey PaperAudio
🎯 What it does: Explores the impact of optimizers on the expressiveness of implicit neural representations (INR), proposes using stable rank as a core metric to measure network expressiveness, and demonstrates that using the optimizer Muon, which maintains high rank and near-orthogonal updates, significantly improves the reconstruction quality of high-frequency details, even allowing standard ReLU MLPs to reach or exceed common architectures such as Fourier Features or SIREN;
Optimizing Visual Generative Models via Distribution-wise Rewards
Ruihang Li (University of Science and Technology of China), Wenjie Wang (University of Science and Technology of China)
GenerationOptimizationReinforcement Learning from Human FeedbackTransformerReinforcement LearningDiffusion modelScore-based ModelImageStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Fine-tune visual generation models with reinforcement learning using distributed rewards to reduce reward hacking, improve image diversity and quality; achieve efficient distribution-level FID rewards through subset-replace strategy, and apply them to model fine-tuning and post-model fusion.
OPTION: Optimal Transport–Guided Flow Matching for Incomplete and Unaligned Multi-View Clustering
Siyuan Zhou (Hebei Normal University), Zhibin Gu (Hebei Normal University)
Data SynthesisOptimizationRepresentation LearningTransformerDiffusion modelFlow-based ModelAuto EncoderContrastive LearningImageMultimodalityTabularOrdinary Differential Equation
🎯 What it does: Propose a unified framework called OPTION for handling multi-view clustering problems where both missing views and view misalignment exist simultaneously.
OptMaster: A DAG-Based Framework for Formulation and Heuristic Discovery in Optimization
Hang Lin (Tongji University), Weinan E (AI for Science Institute)
OptimizationTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposed the OptMaster framework, which uniformly converts natural language requirements into executable optimization models and heuristic algorithms, employing techniques such as DAG search, LLM code generation and verification, and cross-branch knowledge transfer.
OptProver: Bridging Olympiad and Optimization through Continual Training in Formal Theorem Proving
Chenyi Li (Peking University), Zaiwen Wen (Peking University)
OptimizationComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringMixture of ExpertsTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: OptProver is a formal theorem proving model specialized in undergraduate-level optimization problems, achieving cross-domain transfer through continuous training from Olympiad-level theorems to the field of optimization.
OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every Iteration
Shaobo Wang (Shanghai Jiao Tong University), Linfeng Zhang (Shanghai Jiao Tong University)
OptimizationComputational EfficiencyData-Centric LearningTransformerLarge Language ModelContrastive LearningTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose a dynamic data selection framework named OPUS, which determines which training samples are most valuable at each update step based on the utility induced by the optimizer's projection, and achieves large-scale efficient computation through Ghost, CountSketch, and Boltzmann sampling;
ORBIT: A Prognostic World Model for Ocular Reasoning Based on Imagined Trajectories
Jiangtao Yan (Wuhan University), Diping Song (Shanghai Artificial Intelligence Laboratory)
ClassificationRecognitionImage TranslationRestorationSegmentationGenerationData SynthesisOptimizationDrug DiscoveryTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision Language ModelVision-Language-Action ModelDiffusion modelScore-based ModelWorld ModelImageMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyUltrasoundElectronic Health Records
🎯 What it does: Proposed an ophthalmic disease progression world model called ORBIT based on visual counterfactual reasoning, for long-term fundus disease management
Orchestrating Spatial Semantics via a Zone-Graph Paradigm for Intricate Indoor Scene Generation
Meisheng Zhang (Peking University), Jiang Bian (Microsoft Research Asia)
GenerationData SynthesisOptimizationReinforcement Learning from Human FeedbackGraph Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringVision Language ModelDiffusion modelGenerative Adversarial NetworkImageTextMultimodalityGraphBenchmarkChain-of-Thought
🎯 What it does: Propose the ZoneMaestro framework, shifting indoor scene generation from object-based simple arrangement to space planning based on Zone-Graph (functional zones); and achieve high-density and geometrically valid scenes through internalized reasoning and geometric denoising cycle (Alternating Spatial Alignment).
OrchJail: Jailbreaking Tool-Calling Text-to-Image Agents by Orchestration-Guided Fuzzing
Jianming Chen (Institute of Software, Chinese Academy of Sciences), Fanjiang Xu (Institute of Software, Chinese Academy of Sciences)
Safty and PrivacyAdversarial AttackTransformerLarge Language ModelPrompt EngineeringDiffusion modelImageTextMultimodalityRetrieval-Augmented Generation
🎯 What it does: A black-box fuzzing framework named OrchJail based on toolchain abstraction and causal reasoning was constructed to jailbreak tool-call-based text-to-image (T2I) agents, leveraging the tool orchestration patterns from successful cases to guide the search and enhance attack efficiency.
Order Matters in Retrosynthesis: Structure-aware Generation via Reaction-Center-Guided Discrete Flow Matching
Chenguang Wang (Chinese University of Hong Kong), Tianshu Yu (Chinese University of Hong Kong)
Drug DiscoveryGraph Neural NetworkTransformerDiffusion modelFlow-based ModelContrastive LearningGraphSequential
🎯 What it does: Propose a structure-aware, template-free, one-step retrosynthesis model called RetroDiT, which achieves positional inductive bias by placing the reaction center atom at the beginning of the sequence, and combines discrete flow matching to enable efficient generation.
Order Matters: Unveiling the Hidden Impact of Macro Placement Sequences via Proxy-Guided LLM Evolution
Shibing Mo (Xidian University), Ruilin Wu (Xidian University)
OptimizationTransformerLarge Language ModelReinforcement LearningAgentic AITabularBenchmark
🎯 What it does: Proposed an agent-guided LLM evolution framework called OrderPlace, which automatically discovers macro placement order strategies, thereby significantly improving macro placement quality.
Order within Chaos: Capturing Intrinsic Energy Anomalies for AI-Manipulated Image Forgery Localization
Yiming Wang (Zhejiang University), Shouling Ji (Zhejiang University)
GenerationData SynthesisAnomaly DetectionConvolutional Neural NetworkTransformerPrompt EngineeringDiffusion modelScore-based ModelContrastive LearningImage
🎯 What it does: This paper proposes the FLAME framework, which utilizes the energy anomalies of diffusion models (implemented through the LAD graph and SAM adapter) to achieve pixel-level tampering localization in AI-generated images, and introduces EditStream as an automated data generation pipeline;
Origo: Interpretable Multi-physics PDE Foundation Model through Neural Operator Splitting
Li Sun (Beijing University of Posts and Telecommunications), Philip S. Yu (University of Illinois Chicago)
Explainability and InterpretabilityComputational EfficiencyConvolutional Neural NetworkTransformerMixture of ExpertsContrastive LearningTabularTime SeriesSequentialPhysics RelatedStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Propose Origo, a general multi-physics PDE foundation model based on neural operator splitting, which can learn physical mechanisms and perform automatic reasoning through context learning without relying on equation identifiers;
ORLoopBench: Solver-in-the-Loop Benchmarks for Self-Correction and Behavioral Rationality in Operations Research
Ruicheng Ao (Massachusetts Institute of Technology), Xinshang Wang (Alibaba Group)
OptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextTabularBenchmark
🎯 What it does: Propose a verifiable repair method for infeasible linear/integer programming models based on solver feedback (IIS) as a Markov decision process; construct two evaluation benchmarks, ORLOOPBENCH (OR-DEBUGBENCH and OR-BIASBENCH), for iterative debugging and decision rationality assessment, respectively.
Orthogonal Concept Erasure for Diffusion Models
Yuhao Sun (University of Science and Technology of China), Hongtao Xie (University of Science and Technology of China)
GenerationDiffusion modelImageText
🎯 What it does: Propose an orthogonal transformation-based concept elimination method called OCE, which achieves concept erasure by hierarchically rotating parameters through orthogonal transformations in diffusion models, preserving generation capability while effectively removing target concepts.
Orthogonal Hierarchical Decomposition for Structure-Aware Table Understanding with Large Language Models
Bin Cao (Zhejiang University Of Technology), JING FAN
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextMultimodalityTabularRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This paper proposes an Orthogonal Hierarchical Decomposition (OHD) framework, which utilizes Orthogonal Tree Induction (OTI) to decompose complex tables into row trees and column trees. It achieves structure-aware textual table representations through dual-path association and LLM semantic arbitration, significantly enhancing the performance of LLMs on table reasoning tasks.
Orthogonal Model Merging
Sihan Yang (Chinese University of Hong Kong), Weiyang Liu (Chinese University of Hong Kong)
OptimizationComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningMixture of ExpertsTextBenchmark
🎯 What it does: Propose Orthogonal Model Merging (OrthoMerge), which merges models fine-tuned for multi-task learning on the orthogonal group Riemannian manifold. It can handle orthogonal matrices obtained from OTF training and extract and merge updates from models not trained with OTF through Orthogonal-Residual Decoupling.
OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization
Zhikai Li (Institute of Automation Chinese Academy of Sciences), Qingyi Gu (Institute of Automation Chinese Academy of Sciences)
OptimizationComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelText
🎯 What it does: Propose an additive weight transformation based on the low-rank property of the Hessian (OSAQ), used for low-bit (e.g., 2-bit, 3-bit) LLM weight quantization, which can suppress weight-level outliers without affecting the original model's task loss.
OSCS: Online Selection with Provable FAR Control for LLM Safety
Zirui Hu (Nanyang Technological University), Dacheng Tao (Nanyang Technological University)
Safty and PrivacyTransformerLarge Language ModelText
🎯 What it does: Proposed an online control framework for FAR (False Acceptance Rate) of LLM malicious inputs called OSCS, which can make real-time accept/reject decisions without using malicious calibration samples.
OServe: Accelerating LLM Serving via Spatial-Temporal Workload Orchestration
Youhe Jiang (University of Cambridge), Eiko Yoneki (University of Cambridge)
OptimizationComputational EfficiencyTransformerLarge Language ModelFlow-based ModelTextTime Series
🎯 What it does: Designed and implemented the OSERVE system, providing heterogeneous model deployment, workload-aware scheduling, and spatiotemporal adaptive switching to improve the throughput and latency of large language model inference services.
OSF: On Pre-training and Scaling of Sleep Foundation Models
Zitao Shuai (University of California, Los Angeles), Yuzhe Yang (University of California, Los Angeles)
ClassificationRecognitionAnomaly DetectionRepresentation LearningTransformerMixture of ExpertsAuto EncoderContrastive LearningTime SeriesBiomedical DataBenchmark
🎯 What it does: Construct and systematically evaluate the self-supervised pre-training and scaling methods of the sleep foundation model (Sleep FM), proposing a pre-training strategy based on channel-invariant feature learning, and creating a large open-source sleep benchmark called SleepBench.
OSM+: Billion-Level Open Street Map Dataset for City-wide Experiments
Guanjie Zheng (Shanghai Jiao Tong University), Wen Ling (Shanghai Jiao Tong University)
Autonomous DrivingOptimizationFederated LearningComputational EfficiencyMeta LearningReinforcement Learning from Human FeedbackNeural Architecture SearchGraph Neural NetworkDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningGraphTabularTime SeriesSequentialBenchmark
🎯 What it does: Constructed the OSM+ global billion-scale road network graph dataset and provided cloud-based querying and benchmark tasks;
OSNIP: Balancing the Privacy-Utility-Efficiency Trilemma in LLM Inference via Obfuscated Semantic Null Space
Zhiyuan Cao (Shanghai Key Laboratory Of Computer Software Testing And Evaluating), Mingang Chen (Shanghai Key Laboratory Of Computer Software Testing And Evaluating)
Safty and PrivacyComputational EfficiencyTransformerLarge Language ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningText
🎯 What it does: Propose a lightweight client-side encryption framework called OSNIP, which utilizes a 'fuzzy semantic null space' to project original embeddings into a high-dimensional space that is approximately orthogonal and semantically preserved, thereby achieving privacy protection without affecting LLM inference;
OTora: A Unified Red Teaming Framework for Reasoning-Level Denial-of-Service in LLM Agents
Xinyu Li (University of Exeter), Gaojie Jin (University of Macau)
Federated LearningSafty and PrivacyExplainability and InterpretabilityComputational EfficiencyAdversarial AttackTransformerLarge Language ModelPrompt EngineeringGenerative Adversarial NetworkTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposes OTora, a unified red team framework for inducing reasoning layer denial-of-service (R-DoS) attacks in LLM agents.
Out-of-Distribution Evaluation of Rule-Based and Strategic Reasoning in Chess Transformers
Anna Mészáros (University of Cambridge), Ferenc Huszár (University of Cambridge)
OptimizationExplainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerSupervised Fine-TuningTabularSequentialBenchmarkChain-of-Thought
🎯 What it does: Studied the performance of Transformer-based chess decision models in rule reasoning and strategy adaptation, and evaluated them by constructing various OOD (out-of-distribution) chess positions and variants.
Outcome Rewards Do Not Guarantee Verifiable or Causally Important Reasoning
Qinan Yu (Stanford University), Christopher Potts (Stanford University)
Explainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningTextChain-of-Thought
🎯 What it does: This paper proposes two novel evaluation metrics—Causal Importance of Reasoning (CIR) and Sufficiency of Reasoning (SR)—to measure the causal importance and verifiability of reasoning chains in chain-of-thought reasoning. Based on these metrics, the paper systematically evaluates the impact of result-based RLVR on the quality of reasoning chains. Subsequently, two improvement strategies are proposed: (1) performing a small amount of supervised fine-tuning (SFT) with expert reasoning chains before RLVR, and (2) incorporating auxiliary signals of CIR/SR into the RLVR reward. These methods are shown to significantly improve CIR and SR while maintaining or enhancing task accuracy.
Outcome-Aware Spectral Feature Learning for Instrumental Variable Regression
Dimitri Meunier (University College London), Arthur Gretton (University College London)
Representation LearningTransformerReinforcement LearningAuto EncoderContrastive LearningImageTabularTime Series
🎯 What it does: Proposed an outcome-aware spectral feature learning framework (Augmented Spectral Feature Learning) for nonparametric instrumental variable regression, which improves the performance of traditional spectral feature learning in cases of spectral misalignment by introducing an augmented operator that incorporates outcome information.
Outrunning LLM Cutoffs: A Live Kernel Crash Resolution Benchmark for All
Chenxi Huang (Columbia University), Baishakhi Ray (Columbia University)
Autonomous DrivingOptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyAI Code AssistantReinforcement Learning from Human FeedbackNeural Architecture SearchTransformerLarge Language ModelAgentic AIPrompt EngineeringDiffusion modelScore-based ModelFlow-based ModelRectified FlowNeural Radiance FieldAuto EncoderContrastive LearningGaussian SplattingTextTabularTime SeriesSequentialBenchmarkRetrieval-Augmented GenerationChain-of-ThoughtStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: This paper proposes LIVE-KBENCH, a continuously updated Linux kernel crash repair benchmark, and implements KENV environment, enabling any Bash-based LLM agent to generate and verify patches on a unified compilation/test platform.
Over-Alignment vs Over-Fitting: The Role of Feature Learning Strength in Generalization
Taesun Yeom (Pohang University of Science and Technology), Jaeho Lee (Pohang University of Science and Technology)
ClassificationRepresentation LearningConvolutional Neural NetworkContrastive LearningImage
🎯 What it does: Investigate the impact of feature learning strength (FLS) on the generalization performance of deep network classification, and discover that there exists an optimal moderate FLS.
Overclocking Electrostatic Generative Models
Daniil Shlenskii (AXXX), Alexander Korotin (Applied AI Institute)
GenerationComputational EfficiencyKnowledge DistillationTransformerDiffusion modelScore-based ModelFlow-based ModelRectified FlowImageStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: This paper proposes an inverse Poisson flow matching (IPFM) method to compress the high-order ODE sampling of electrostatic generative models such as PFGM++ into a one-step or few-step generator requiring only a small number of network evaluations.
Overcoming PINNs Failure Modes In High Dimension With Low-Rank Fourier Sum
Natan Kaminsky (Technion - Israel Institute of Technology), Kira Radinsky (Technion - Israel Institute of Technology)
OptimizationAuto EncoderPhysics RelatedStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Propose the Low-Rank Fourier Sums (LoRFS) method, which directly represents PDE solutions using low-rank separable Fourier series, and computes physical loss and gradients via closed-form integration, addressing the training failure of PINNs in high-dimensional oscillatory, multi-scale, stiff, or long-time dynamical systems.
Overcoming the Incentive Collapse Paradox
Qichuan Yin (University of Chicago), Shuangning Li (University of Chicago)
OptimizationFederated LearningData-Centric LearningReinforcement Learning from Human FeedbackTabularBiomedical Data
🎯 What it does: This paper investigates the incentive collapse paradox in AI-assisted task delegation, proposes a sentinel-auditing payment mechanism, and constructs an incentive-friendly active statistical inference framework based on this mechanism.
Overcoming the Modality Gap in Context-Aided Forecasting
Vincent Zhihao Zheng (ServiceNow), Valentina Zantedeschi (ServiceNow)
Data-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextTime SeriesBenchmarkRetrieval-Augmented Generation
🎯 What it does: A verifiable time series context dataset CAF-7M with 7 million entries was constructed by generating and verifying contexts using LLMs, and a dual-modal model called DoubleCast was trained and evaluated on this dataset, demonstrating the positive impact of context on prediction.
Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets
Jack Hopkins (Anthropic Fellows Program), Fabien Roger (Anthropic)
Safty and PrivacyExplainability and InterpretabilityAdversarial AttackTransformerLarge Language ModelPrompt EngineeringTextMultimodalityRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose an 'Overthinking' method that induces language models to leak hidden information by amplifying the reasoning direction in the weight space, based on the weighted amplification of task vectors;
OvisOCR: End-to-End Document Parsing via Aligning Specialized Perception with General Reasoning
Jun-Peng Jiang (Nanjing University), Han-Jia Ye (Alibaba Group)
RecognitionComputational EfficiencyKnowledge DistillationData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision Language ModelContrastive LearningImageTextMultimodality
🎯 What it does: Proposes OvisOCR, a fully end-to-end multimodal large language model capable of directly mapping full-page document images to structured Markdown/JSON, eliminating the traditional cropping-identification-merging process.