π― What it does: Propose a single-stage sparse retrieval (SSR) method, which maps token embeddings into a high-dimensional sparse space via a sparse autoencoder, directly constructing a neuron-level inverted index, eliminating K-means clustering and multi-stage pruning;
Noise-corrected GRPO: From Noisy Rewards to Unbiased Gradients
Omar El mansouri, Salem Lahlou (Mohamed bin Zayed University of Artificial Intelligence)
CodeReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringContrastive LearningText
π― What it does: Proposes the noise-corrected GRPO/Dr.GRPO methods, which utilize a Bernoulli noise model to denoise rewards and recover unbiased gradients, thereby enhancing the robustness of policy optimization in RLHF/RLVR environments.
π― What it does: This paper proposes a new non-adversarial imitation learning method called Dual Q-DM, aiming to solve the composite error problem in existing methods and achieve better generalization ability through a value flow mechanism.
Tianyi Ma (University of Notre Dame), Yanfang Ye (University of Notre Dame)
CodeGenerationOptimizationComputational EfficiencyAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextSequential
π― What it does: Proposes N-MARS, a non-monotonic autoregressive sequence model that utilizes a special <UNDO> token to enable immediate corrections during the generation process.
Nonconvex Low-Rank Tensor Representation with Deep Priors for Multiview Subspace Clustering
Yao Fu (Southwest University), Zhi Wang (Southwest University)
CodeOptimizationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningData-Centric LearningAuto EncoderContrastive LearningImageMultimodalityTabularBenchmark
π― What it does: Proposed a multi-view subspace clustering model called NRDN-MvSC, which combines a non-convex low-rank tensor representation with a deep prior.
Nonparametric Distribution Regression Re-calibration
ΓdΓ‘m Jung (HUN-REN SZTAKI), Andras A Benczur
CodeExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTabularBenchmark
π― What it does: Proposed a nonparametric recalibration algorithm based on conditional kernel mean embedding, aiming to address the calibration issue between the predicted distribution and the true empirical uncertainty in probabilistic regression.
π― What it does: Propose the GraphNC framework, which uses a pre-trained teacher model to perform normality calibration for semi-supervised graph anomaly detection in the score space and representation space.
OC-space: a Unifying Perspective on Verification of Tree Ensembles
Timo Martens (KU Leuven), Jesse Davis (KU Leuven)
CodeExplainability and InterpretabilityComputational EfficiencyTabularBenchmark
π― What it does: This paper proposes a unified verification framework based on the output configuration space (OC-space) of tree ensemble models, which determines whether a model satisfies a given property by searching over all possible leaf combinations.
π― What it does: Propose an offline multi-agent reinforcement learning framework called OMSD, which utilizes sequential decomposition of behavioral policies and employs diffusion models to estimate conditional scores for behavioral regularization, encouraging agents to maintain coordination within offline data.
π― What it does: This paper proposes Generative Trajectory Policies (GTP), a general generative policy based on continuous-time ODEs for offline reinforcement learning.
OLion: Approaching the Hadamard Ideal by Intersecting Spectral and L inf Implicit Biases
Zixiao Wang (Peking University), Huishuai Zhang (Peking University)
CodeOptimizationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningMixture of ExpertsImageText
π― What it does: Proposed a new optimizer called OLion (Orthogonal Lion), which first performs Newton-Schulz orthogonalization on the gradient, then takes the sign of the orthogonalized direction, combining Muon's spectral structure control and Lion's ββ coordinate control, achieving Hadamard idealized updates for matrix parameters;
On Information Self-Locking in Reinforcement Learning for Active Reasoning of LLM agents
Deyu Zou (Chinese University of Hong Kong), James Cheng (Chinese University of Hong Kong)
CodeTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringTextMultimodalityBenchmark
π― What it does: Investigate and address the information self-locking problem caused by reward learning in large language model agents during active reasoning, and propose a directional criticism-based advantage reweighting method (AREW) to break this bottleneck.
π― What it does: Perform a non-asymptotic analysis of the convergence rate of original LoRA gradient descent without requiring parameter boundedness or Lipschitz smoothness assumptions, and propose an adaptive learning rate scheme based on theoretical insights;
On the Entropy Dynamics in Reinforcement Fine-Tuning of Large Language Models
Shumin Wang (University of Science and Technology of China), Yanyong Zhang (University of Science and Technology of China)
CodeOptimizationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningText
π― What it does: Proposes a theoretical framework regarding entropy changes during the process of reinforcement learning fine-tuning (RFT), and provides a first-order expression for entropy changes caused by single logit updates;
On the Generalization in Topology Optimization via Sensitivity-Conditioned Bernoulli Flow Matching
Mohammad Rashed (Technical University of Munich), Nils Thuerey (Technical University of Munich)
CodeOptimizationTransformerDiffusion modelFlow-based ModelImageTabularPhysics Related
π― What it does: This paper proposes a binary manifold generative model based on sensitivity conditions to accelerate topology optimization and improve generalization to out-of-distribution samples.
π― What it does: This paper studies the behavior of predictive coding networks (PCN) in the limits of infinite width and infinite depth, and proves that under parameterizations with stable width and depth and feature learning, PCN converges to backpropagation (BP). Subsequently, experiments verify the validity of this theory in nonlinear networks (such as CNN and Transformer).
On the Limits of LLM Adaptability: Impact of Model-Internalized Priors on Annotation Task Performance
Etienne Casanova (California Institute of Technology), R. Michael Alvarez (California Institute of Technology)
CodeClassificationExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsContrastive LearningText
π― What it does: This study investigates the interaction between internalized priors (solidified cognition about task definitions) within large language models and prompt instructions, evaluating the model's adaptability and error-correction ability across different definitions of toxicity.
π― What it does: The study investigates the role of class-level feature statistics in class-incremental learning under large-scale pre-trained models, and proposes two minimalist reference methods (SCLIP and SViT) to verify this.
π― What it does: Studied how to utilize steady-state observations and intervention data to recover the drift, bias, and diffusion parameters of multivariate Ornstein-Uhlenbeck processes, and proved that under specific SCC conditions, only one intervention per strongly connected component is needed to achieve parameter generic identifiability (subject only to global scale uncertainty constraints)
One LR Doesnβt Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs
Di He (Shenzhen Institutes of Advanced Technology Chinese Academy of Sciences), Shiwei Liu (Max Planck Institute for Intelligent Systems)
CodeOptimizationComputational EfficiencyRepresentation LearningTransformerLarge Language ModelAuto EncoderContrastive LearningText
π― What it does: Proposes a hierarchical learning rate allocation strategy LLR based on the Heavy-Tail Self-Regulation (HT-SR) theory, which automatically calculates the heavy-tail index of the weight spectrum for each layer of the Transformer language model and dynamically adjusts the learning rate.
CodeAI Code AssistantTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringText
π― What it does: Proposes RepoNavigator, an LLM agent that achieves repository-level code localization using only a single 'jump' tool, through reinforcement learning;
π― What it does: Proposed a single-step conditional sampling framework called CGMMD based on Maximum Mean Discrepancy (MMD), training the conditional generator by minimizing the empirical ECMMD loss.
OneSearch: A Preliminary Exploration of the Unified End-to-End Generative Framework for E-commerce Search
Ben Chen (Kuaishou Technology), Kun Gai (Kuaishou Technology)
CodeRetrievalRecommendation SystemTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringMixture of ExpertsTextTabularSequentialRetrieval-Augmented Generation
π― What it does: Designed and deployed OneSearch, an end-to-end generative retrieval and ranking framework to replace traditional multi-stage e-commerce search systems.
Online Bayesian Experimental Design for Partially Observed Dynamical Systems
Sara Perez-Vieites (University of Helsinki), Dominik Baumann (Aalto University)
CodeAutonomous DrivingOptimizationRobotic IntelligenceDrug DiscoveryReinforcement Learning from Human FeedbackTabularTime SeriesSequentialBenchmarkStochastic Differential Equation
π― What it does: Propose an online Bayesian experimental design method for partially observable dynamic systems, which can select the most informative experiment design after each update step.
Online Learning and Inference for Cox Proportional Hazards Model Using Renewable Sieve Estimation
Mengtong Hu (University of Michigan), Peter Song (University of Michigan)
CodeOptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTabularBiomedical DataElectronic Health Records
π― What it does: Propose an online learning framework COLSA, which uses Renewable Sieve Estimation to approximate the baseline risk through basis functions based on Bernstein polynomials, replacing the traditional Cox partial likelihood, achieving online updates without the need to store historical data or a global risk set;
Ye Mo (Zhejiang University), Philip Torr (University of Oxford)
CodeExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningMeta LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
π― What it does: Proposed the OpenIKLR framework, which enables logical reasoning under conditions of incomplete knowledge in an open world.
OpenTSLM: Time-Series Language Models for Reasoning over Multivariate Medical Text- and Time-Series Data
Patrick Langer (Stanford University), Paul Schmiedmayer (Stanford University)
CodeAnomaly DetectionExplainability and InterpretabilityComputational EfficiencyRepresentation LearningData-Centric LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextTime SeriesBiomedical DataElectronic Health RecordsElectrocardiogramReview/Survey PaperChain-of-Thought
π― What it does: Proposed an open-source time series language model (OpenTSLM) framework that can locally process multivariate medical time series data within large language models and perform reasoning through natural language.
OPIC: Enhancing Language Model Merging via Optimizing In-Context Capability
Jie He (National University of Defense Technology), Ji Wang (National University of Defense Technology)
CodeOptimizationKnowledge DistillationRepresentation LearningHyperparameter SearchTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsTextBenchmarkRetrieval-Augmented Generation
π― What it does: Proposed the OPIC framework, achieving model merging without dependency on a validation set, and enhancing the performance after fusion by maintaining the model's in-context learning (ICL) capability.
OPT-Engine: Benchmarking the Limits of LLMs in Optimization Modeling via Complexity Scaling
Yitian Chen (Cardinal Operations), Dongdong Ge (Shanghai Jiao Tong University)
CodeOptimizationTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
π― What it does: Proposed an expandable optimization modeling benchmark framework, OPT-Engine, for evaluating the performance of large language models in optimization modeling tasks ranging from linear programming to mixed integer programming.
OptiFluence: Principled Design of Privacy Canaries
Mohammad Yaghini (University of Toronto), Florian Tramèr (ETH Zurich)
CodeOptimizationSafty and PrivacyConvolutional Neural NetworkRecurrent Neural NetworkContrastive LearningImage
π― What it does: Proposes the OptiFluence framework, which automatically generates high-detectability privacy canaries through bi-level optimization, achieving near-perfect membership inference detection on multiple datasets;
Anneliese Riess (Helmholtz Munich), Georgios Kaissis (University of Potsdam)
CodeSafty and Privacy
π― What it does: Proves the conjecture proposed by Zhu et al. (2022) in Appendix F.3, which states that among all transformation rules mapping R\'enyi differential privacy (RDP) configurations to effective hypothesis testing trade-offs f, the rule based on the intersection of single-order RDP privacy regions is optimal.
CodeOptimizationComputational EfficiencyData-Centric LearningTransformerLarge Language ModelContrastive LearningTextBenchmarkRetrieval-Augmented Generation
π― What it does: Propose a dynamic data selection framework named OPUS, which determines which training samples are most valuable at each update step based on the utility induced by the optimizer's projection, and achieves large-scale efficient computation through Ghost, CountSketch, and Boltzmann sampling;
Yuhao Sun (University of Science and Technology of China), Hongtao Xie (University of Science and Technology of China)
CodeGenerationDiffusion modelImageText
π― What it does: Propose an orthogonal transformation-based concept elimination method called OCE, which achieves concept erasure by hierarchically rotating parameters through orthogonal transformations in diffusion models, preserving generation capability while effectively removing target concepts.
OTora: A Unified Red Teaming Framework for Reasoning-Level Denial-of-Service in LLM Agents
Xinyu Li (University of Exeter), Gaojie Jin (University of Macau)
CodeFederated LearningSafty and PrivacyExplainability and InterpretabilityComputational EfficiencyAdversarial AttackTransformerLarge Language ModelPrompt EngineeringGenerative Adversarial NetworkTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
π― What it does: Proposes OTora, a unified red team framework for inducing reasoning layer denial-of-service (R-DoS) attacks in LLM agents.
π― What it does: Propose the Low-Rank Fourier Sums (LoRFS) method, which directly represents PDE solutions using low-rank separable Fourier series, and computes physical loss and gradients via closed-form integration, addressing the training failure of PINNs in high-dimensional oscillatory, multi-scale, stiff, or long-time dynamical systems.
OVLR: Efficient, Scalable, and Robust Training via Output-Level Variance-Reduced Likelihood Ratio
Minhao Zou (Peking University), Yijie Peng (Peking University)
CodeOptimizationHyperparameter SearchData-Centric LearningRobotic IntelligenceReinforcement Learning from Human FeedbackImageTextTabularTime SeriesSequentialAudio
π― What it does: Propose the OVLR (Output-Level Variance-Reduced Likelihood Ratio) framework, which achieves direct and efficient optimization of gradient-agnostic (vanishing or black-box) objectives by adding noise to the model's output space and using the likelihood ratio method to estimate gradients.
π― What it does: Propose PACT, a self-evolving framework that performs post-training safety alignment on diffusion policies, leveraging self-replay and continuous constraint supervision to project policies into feasible physical safety regions.
PaperBanana: Automating Academic Illustration for AI Scientists
Dawei Zhu (Peking University), Jinsung Yoon (Google Cloud AI Research)
CodeGenerationData SynthesisTransformerLarge Language ModelAgentic AIPrompt EngineeringVision Language ModelDiffusion modelImageTextTabularBenchmarkRetrieval-Augmented Generation
π― What it does: Proposed the PAPERBANANA framework, which leverages multi-agent systems to automatically generate figures and statistical plots that meet academic publishing standards.
π― What it does: Proposed and implemented a parallel echo state network (ParalESN) based on diagonal complex linear recursion, which can process sequential data in parallel and construct a high-dimensional reservoir with low memory usage.
Parameters as Experts: Adapting Vision Models with Dynamic Parameter Routing
Meng Lou (University of Hong Kong), Yizhou Yu (University of Hong Kong)
CodeClassificationObject DetectionSegmentationConvolutional Neural NetworkTransformerPrompt EngineeringMixture of ExpertsContrastive LearningImage
π― What it does: Proposes ParaX, a parameter-efficient fine-tuning method for visual models that generates input-related low-rank adapters through dynamic parameter routing in a shared expert center.
π― What it does: Propose the Parametric Prior Mapping (PPM) framework, which combines learnable push mapping with parameterized priors to achieve probabilistic forecasting for non-stationary multivariate time series;
ParaTool: Shifting Tool Representations from Context to Parameters
Zekai Yu (Beijing University of Posts and Telecommunications), Cheng Yang (Beijing University of Posts and Telecommunications)
CodeComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation
π― What it does: Propose the ParaTool framework, which transfers the tool calling knowledge of LLMs from the context to loadable parameters, enabling tool calling without including tool documentation or examples in the prompt.
ParEVO: Synthesizing Code for Irregular Data: High-Performance Parallelism through Agentic Evolution
Liu Yang (Yale University), Quanquan C. Liu (Yale University)
CodeOptimizationComputational EfficiencyAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningAgentic AIPrompt EngineeringGraphTabularTime SeriesBenchmark
π― What it does: Designed and implemented an end-to-end system named ParEVO, which leverages large language models and evolutionary search techniques to automatically generate high-performance parallel code for irregular data structures (such as sparse graphs, non-uniform grids, etc);
Partial Identification under High-Dimensional Potential Outcomes and Confounders via Optimal Transport
Yunfeng Wang (Fudan University), Zijun Gao (University of Southern California)
CodeOptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTabularBiomedical DataElectronic Health Records
π― What it does: Propose a conditional subspace-slicing (CSS) estimator to achieve partial identification via optimal transport in high-dimensional potential outcomes and confounding variables settings;
Alain Rakotomamonjy (Criteo AI Lab), Liva Ralaivola (Criteo AI Lab)
CodeClassificationDomain AdaptationImageTabular
π― What it does: When learners only have the label proportion of each sample's bag, the paper proposes an algorithm based on particle flow and distribution alignment to learn instance-level classifiers.
PATRA: Pattern-Aware Alignment and Balanced Reasoning for Time Series Question Answering
Junkai Lu (East China Normal University), Bin Yang (East China Normal University)
CodeTransformerLarge Language ModelReinforcement LearningPrompt EngineeringContrastive LearningTextTime SeriesBenchmark
π― What it does: Propose the PATRA framework, which enhances deep reasoning in time series question answering by pattern-aware alignment and reinforcement learning with balanced rewards.
Ji Zhang (Beijing Institute of Technology), Kan Li (Beijing Institute of Technology)
CodeComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought
π― What it does: To address the low-bit quantization bottleneck of KV cache in LLM inference, the authors propose PatternKV, a lightweight scheme that online mines pattern vectors and quantizes KV residuals.
PepCompass: Navigating Peptide Embedding Spaces Using Riemannian Geometry
Marcin MoΕΌejko (University of Warsaw), Ewa Szczurek (Helmholtz Center Munich)
CodeDrug DiscoveryDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkBiomedical Data
π― What it does: Proposed a Riemannian geometry-based antimicrobial peptide design framework called PepCompass, which explores the global and local potential space of a generative model.
Perceptrons and Localization of Attentionβs Mean-Field Landscape
Antonio Γlvarez-LΓ³pez (Universidad AutΓ³noma de Madrid), DomΓ¨nec Ruiz-Balet (Universitat de Barcelona)
CodeExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerContrastive LearningStochastic Differential EquationOrdinary Differential Equation
π― What it does: This paper studies the mean field limit of Transformers under large context lengths, revealing that perceptron blocks localize the self-attention energy landscape, leading to stable points being only finite atomic distributions;
Persona-Pruner: Sculpting Lightweight Models for Role-Playing
Jinsu Kim (Korea University), Jongheon Jeong (Korea University)
CodeData SynthesisComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextRetrieval-Augmented Generation
π― What it does: Propose Persona-Pruner, which utilizes textual persona descriptions to perform structured pruning on large language models, generating lightweight role-playing agents.
Personalized Additive Modeling for Multi-level Federated Learning
Shutong Chen (University of Technology Sydney), Chengqi Zhang (Hong Kong Polytechnic University)
CodeClassificationRecommendation SystemFederated LearningTransformerSupervised Fine-TuningMixture of ExpertsContrastive LearningImageText
π― What it does: Proposes a multi-level incremental modeling framework called FeMAM, which constructs personalized predictions through residual additive combinations of multi-layer shared models (global, subgroup, client);
π― What it does: This paper proposes the Peak-Guided Calibration (PGC) framework, which leverages local peak features to calibrate global discriminability, thereby enhancing the generalization ability of AI-generated image detection.
PGD-NO: A Neural Operator with Precomputed Geometry Decomposition for 3D Million-Scale Physics Simulations
Weiheng Zhong (University of Illinois Urbana-Champaign), Hadi Meidani (University of Illinois Urbana-Champaign)
CodeGraph Neural NetworkTransformerContrastive LearningPoint CloudMeshGraphBenchmarkPhysics Related
π― What it does: Propose a neural operator called PGD-NO, which precomputes geometric encoding offline, enabling the learning and prediction of large-scale three-dimensional industrial PDE scenarios using precomputed geometric tokens without requiring complex geometric mapping on the GPU.
π― What it does: Developed a physics-informed diffusion model called PISD based on a spectral domain latent space, which can generate solutions and parameters of PDEs during inference by guiding Adam to satisfy PDE constraints and observational conditions.
Physiology-Aware Masked Cross-Modal Reconstruction for Biosignal Representation Learning
Hao Zhou (Pennsylvania State University), Sharanya Arcot Desai (Samsung Research America)
CodeAnomaly DetectionExplainability and InterpretabilityRepresentation LearningTransformerAuto EncoderContrastive LearningTime SeriesBiomedical DataElectrocardiogram
π― What it does: By pretraining PPG representations through masked cross-reconstruction of synchronized ECG and PPG, the generalization performance of single PPG signals in health tasks is enhanced.
PixCLIP: Towards Fine-grained Vision-Language Understanding via Any-granularity Pixel-Text Alignment
YiCheng Xiao, Jinqiao Wang (Institute of Automation Chinese Academy of Sciences)
CodeClassificationSegmentationRetrievalRepresentation LearningConvolutional Neural NetworkTransformerLarge Language ModelVision Language ModelContrastive LearningImageTextMultimodality
π― What it does: Develop the PixCLIP framework to achieve alignment between arbitrary-shaped pixel-level regions and arbitrary-length text, and construct a large-scale long-text mask-aligned dataset called LongGRIT.
Plan for Speed: Dilated Scheduling for Masked Diffusion Language Models
Omer Luxembourg (Ben-Gurion University of the Negev), Eliya Nachmani (Ben-Gurion University of the Negev)
CodeComputational EfficiencyAI Code AssistantTransformerLarge Language ModelPrompt EngineeringDiffusion modelText
π― What it does: Propose a sparse decoding scheduling method called Dilated Unmasking Scheduler (DUS), which is only used during inference. By dividing the positions within a block into sparse hierarchical groups and revealing them in parallel, the number of denoising calls per block is reduced from O(B) to O(log B), achieving significant acceleration and quality improvement.
π― What it does: Propose a framework that transforms any two-dimensional continuous representation into a strictly planar group symmetric and continuous representation.
π― What it does: Propose a plugin called PAPO based on the Polar Operator to activate suppressed gradient directions, improving the stability-plasticity balance in continual learning.
PlugMem: A Task-Agnostic Plugin Memory Module for LLM Agents
Ke Yang (University of Illinois Urbana Champaign), ChengXiang Zhai (University of Illinois Urbana Champaign)
CodeExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelTextGraphBenchmarkRetrieval-Augmented GenerationChain-of-Thought
π― What it does: This paper proposes PLUGMEM, a pluggable and task-agnostic memory module that can abstract the long-term experience of large language model agents into propositional and procedural knowledge, organizing it into a knowledge graph.
POLCA: Stochastic Generative Optimization with LLM
Xuanfei Ren (University of Wisconsin Madison), Ching-An Cheng (Google Research)
CodeOptimizationHyperparameter SearchAI Code AssistantTransformerLarge Language ModelTextBenchmarkRetrieval-Augmented GenerationStochastic Differential Equation
π― What it does: Propose an scalable LLM-based random generation optimization framework called POLCA, which automatically searches for optimal parameters under uncertain feedback using a priority queue and an Ξ΅-grid filter.
POLIA: Policy Optimization with Visual-Object-Level Intrinsic Advantage for Multimodal Reasoning
Yiran Zeng (South China University of Technology), Mengchen Zhao (South China University of Technology)
CodeOptimizationTransformerReinforcement LearningVision Language ModelMultimodalityBenchmarkChain-of-Thought
π― What it does: This paper proposes POLIA, a group-based reinforcement learning method that introduces visual object-level internal advantages in multi-modal reasoning, achieving finer-grained credit assignment.
Position: Improved Documentation is Necessary for Benchmarking AI Systems in Geometry
Anna Genevaux (Independent Researcher), Simon Frieder (University of Oxford)
CodeLarge Language ModelTextBenchmarkPhysics Related
π― What it does: Proposed a benchmark release standard for an executable DSL (JGEX) in the geometric domain, constructed and released the JgexDiv corpus containing 137 Euclidean geometry problems, and provided executable interface contracts, predicate support tables, version-fixed verification scripts, and document rewrite records.
Position: Predictive Uncertainty Is Not Enough β Joint Distribution for Full Uncertainty Representation
Adria Aldoma (Barcelona Supercomputing Center), Axel Brando (Barcelona Supercomputing Center)
CodeInformation TheoryClassificationAnomaly DetectionTransformerMixture of ExpertsFlow-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBenchmark
π― What it does: This paper argues that relying solely on predictive uncertainty (epistemic + aleatoric) is insufficient to fully assess model risk, and proposes three sources of uncertainty (domain, model, data noise), achieving comprehensive representation through the joint distribution p(x,y|D)=p(x|D)Β·p(y|x,D).
Position: Reasoning is a Learnable Rule-Based Process
Rachel Lawrence (Microsoft Research), Jacqueline R. M. A. Maasch (Cornell Tech)
CodeExplainability and InterpretabilityReinforcement LearningChain-of-Thought
π― What it does: This paper proposes a definable, learnable rule-driven reasoning process and systematically elaborates on its effectiveness and reliability.
Position: Robust AI Personalization Will Require a Human Context Protocol
Anand V. Shah (Massachusetts Institute of Technology), Michiel A. Bakker (Massachusetts Institute of Technology)
CodeRecommendation SystemSafty and PrivacyExplainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextRetrieval-Augmented Generation
π― What it does: Propose and elaborate on the 'Human Context Protocol (HCP)', providing a portable and user-governed preference layer for AI personalization; construct an open-source prototype in the paper to demonstrate the protocol's feasibility and security; and discuss the differences and advantages of HCP compared to existing personalization technologies.
π― What it does: This paper quantifies the energy consumption and carbon emissions during training and storage of ImageNet-1K, and experimentally verifies the feasibility of subset selection (coreset) techniques in maintaining accuracy, reducing energy consumption, and mitigating data bias.
Position: Sycophancy is an Educational Safety Risk: Why LLM Tutors Need Sycophancy Benchmarks
Enkelejda Kasneci (Technical University of Munich), Gjergji Kasneci (Technical University of Munich)
CodeSafty and PrivacyTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation
π― What it does: Propose a new risk in educational safety β in LLM tutoring, students' pressure (frame-switching, authority, face) can lead to model concession and reinforcement of misconceptions.
π― What it does: The paper studies the impact of measurement time intervals in single-cell perturbation prediction on computational complexity and model complexity, proposing a critical time threshold that transforms the problem from polynomially solvable to NP-hard;
CodeExplainability and InterpretabilityData-Centric LearningAI Code AssistantTransformerLarge Language ModelPrompt EngineeringTextReview/Survey PaperBenchmarkRetrieval-Augmented GenerationChain-of-Thought
π― What it does: The paper explores the validation issue of social simulations based on large language models (LLMs), emphasizing the lack of consensus and standardized evaluation methods in this field, and proposes a shift from expansion to integration, prioritizing the standardization of methodologies.
Position: Your VLM May Not Be Thinking with Interleaved Images
Wenjie Yang (Fudan University), Zengfeng Huang (Fudan University)
CodeExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerSupervised Fine-TuningReinforcement LearningVision Language ModelImageTextMultimodalityBenchmark
π― What it does: This paper systematically evaluates whether the interleaved interactive images in the 'Thinking with Images' paradigm are truly utilized by the model, mainly verifying the impact of images on performance through ablation experiments, attention visualization, and occlusion experiments.
π― What it does: Proposed a relative position encoding called SPE for Spiking Transformer, which utilizes PE-LIF neurons with position-related thresholds to encode relative position information while maintaining linear attention.
Possibilistic Predictive Uncertainty for Deep Learning
Yao Ni (Nanyang Technological University), Piotr Koniusz (University Of New South Wales)
CodeClassificationExplainability and InterpretabilityComputational EfficiencyConvolutional Neural NetworkTransformerContrastive LearningImageTextStochastic Differential Equation
π― What it does: Propose a deep learning framework for predicting uncertainty based on possibility theory, DAPPr, which uses the Dirichlet possibility function to approximate the possible predictive posterior;
Post-Training Language Models for Crosslingual Consistency
Tianyu Liu (ETH ZΓΌrich), Arianna Bisazza (University of Groningen)
CodeOptimizationFederated LearningComputational EfficiencyRepresentation LearningData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningContrastive LearningTextMultimodalityBenchmark
π― What it does: Proposes a post-training method called DCO (Direct Consistency Optimization) to improve the response consistency of multilingual models across different languages, and proves its equivalence to the more challenging PCO (Penalized Consistency Optimization).
Power-Boosted Granger-Causal Discovery for Large Heterogeneous Panel Data
Yiheng Gu (University of Notre Dame), Xiufan Yu (University of Notre Dame)
CodeTabularTime SeriesFinance Related
π― What it does: This paper proposes a Power-Enhanced Panel Granger Causality Test (PE-PGCT) for large-scale heterogeneous panel data, which enhances the power of the test under sparse signals by adding a power-enhancing component based on the maximum Wald statistic to existing panel Granger causality tests (such as DH or HPJ).
CodeFederated LearningSafty and PrivacyComputational EfficiencyContrastive LearningTabularBiomedical Data
π― What it does: In the federated learning scenario, a new independence test method called FedIT-CS is designed, which can perform effective independence tests even when the data distributions across clients are heterogeneous.
PRISM: Distribution-free Adaptive Computation of Matrix Functions for Accelerating Neural Network Training
Shenghao Yang (International Computer Science Institute), Michael W. Mahoney (International Computer Science Institute)
CodeOptimizationComputational EfficiencyImageText
π― What it does: Propose the PRISM framework, which accelerates the iterative computation of matrix functions (such as square roots, inverse roots, and polarization decomposition) through adaptive polynomial approximation and randomized projection, thereby improving the training efficiency of neural networks.
PRISM: Synergizing Vision Foundation Models via Self-organized Expert Specialization
Ying Tang (Huazhong University of Science and Technology), Wei Yang (Huazhong University of Science and Technology)
CodeKnowledge DistillationRepresentation LearningTransformerMixture of ExpertsContrastive LearningImage
π― What it does: Propose a dual-stream conditional Mixture-of-Experts framework PRISM for knowledge distillation in multi-teacher visual foundation models, achieving model fusion through self-organized expert specialization.
PRISM: Training-Free Video Anomaly Detection via Intrinsic Statistical Modeling
YUANTONG CHEN, YanFeng Shang
CodeAnomaly DetectionTransformerVision Language ModelAuto EncoderContrastive LearningVideoTextMultimodality
π― What it does: Proposes a training-free, real-time video anomaly detection framework called PRISM, which is based on a pre-trained multi-modal embedder. It statistically whitens textual anomaly descriptions to suppress common-mode noise, then aligns video features with the denoised anomaly semantic axis to generate anomaly scores.
Privileged Information Distillation for Language Models
Emiliano Penaloza (ServiceNow AI Research), Massimo Caccia (ServiceNow AI Research)
CodeFederated LearningSafty and PrivacyComputational EfficiencyKnowledge DistillationRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringMixture of ExpertsTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
π― What it does: This paper proposes two methods for utilizing privileged information (PI) during training for language model distillation: ΟβDistill (joint teacher-student objective) and OnβPolicy SelfβDistillation (OPSD), achieving the retention of advantages brought by PI even when PI is not available during testing;
Cristiana Diaconu (Polymathic AI Collboration), Payel Mukhopadhyay (Polymathic AI Collboration)
CodeComputational EfficiencyRepresentation LearningScore-based ModelContrastive LearningTabularTime SeriesPhysics Related
π― What it does: Propose a training-efficient probabilistic post-processing method to convert pre-trained deterministic PDE models into probabilistic models.
ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluation
Yizheng Huang (Google DeepMind), Zi Wang (Google DeepMind)
CodeAnomaly DetectionExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelGaussian SplattingTextMultimodalityBenchmark
π― What it does: Propose the PROEVAL framework, which utilizes transfer learning to construct an efficient Gaussian Process prior, achieving both active failure detection and sample-efficient performance estimation.
π― What it does: Proposed the ProMiSE benchmark for systematically evaluating the performance of protein multi-state structure prediction models under three biological contexts: intrinsic, multi-body induced, and ligand induced.
π― What it does: This paper investigates the prompt forgetting phenomenon in Multimodal Diffusion Transformer (MMDiT) and proposes a no-training, inference-time feasible Prompt Reinjection technique to help recover the semantic information of deep text features.
π― What it does: Proposed a framework called PromptDyG for test-time adaptation in dynamic graphs, which learns lightweight graph prompts on frozen dynamic graph models and corrects graph structure drift by leveraging unsupervised entropy minimization.
PromptRL: Prompt Matters in RL for Flow-Based Image Generation
Fu-Yun Wang (Chinese University of Hong Kong), Taesung Park (Reve)
CodeGenerationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringFlow-based ModelImageText
π― What it does: This paper proposes PromptRL, a reinforcement learning framework that jointly trains language models (LMs) with flow-matching image generation models (FMs), leveraging LMs to generate diverse prompts to improve exploration efficiency and suppress prompt overfitting;
Propose, Solve, Verify: Self-Play Through Formal Verification
Alex Wilf (Carnegie Mellon University), Sean Welleck (Carnegie Mellon University)
CodeAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringText
π― What it does: This paper proposes a self-play framework based on formal verification called PSV (Propose, Solve, Verify), which generates code problems in formal specifications and uses a formal verifier to determine whether the generated code meets the specification, thereby providing reliable reward signals for code generation training of large language models; during the iterative process, the model can generate more challenging specifications and improve itself through verified solutions.
Proteus: Lookup-Free Trellis-Coded Quantization by Lattice-Breaking Compute Codes for 2-Bit LLMs
Zhengwu Yang (Baidu Inc), Dianhai Yu (Baidu Inc)
CodeComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelDiffusion modelAuto EncoderContrastive LearningText
π― What it does: Designed and implemented Proteus, a fully lookup-table-free, bitshift trellis-based low-bit quantization framework for 2-bit weight post-training quantization (PTQ) of large language models.
π― What it does: Propose ProtoKV, a constant-sized KV cache mechanism that combines near-window precise storage with remote prototype memory to achieve delayed queries in streaming video understanding;
Prototype-Based Test-Time Adaptation of Vision-Language Models
Zhaohong Huang (Xiamen University), Rongrong Ji (Xiamen University)
CodeClassificationDomain AdaptationTransformerVision Language ModelContrastive LearningImageTextPoint Cloud
π― What it does: Proposes a prototype-based test-time adaptation method called PTA, which does not use backpropagation or caching. It leverages class-specific knowledge prototypes to accumulate information in real-time during the test stream, thereby improving the zero-shot performance of vision-language models such as CLIP.
Prototype-Grounded Concept Models for Verifiable Concept Alignment
Stefano Colamonaco (KU Leuven), Giuseppe Marra (KU Leuven)
CodeExplainability and InterpretabilityRepresentation LearningContrastive LearningImage
π― What it does: Prototype-Grounded Concept Models (PGCMs) are proposed by associating concepts with visual prototypes, achieving verifiable concept alignment.
π― What it does: Proposed a joint policy-reward co-pretraining adversarial imitation learning framework called CoPT-AIL, and provided theoretical proof.
Provably Protecting Fine-Tuned LLMs from Training Data Extraction while Preserving Utility
Tom Segal (Ben-Gurion University of Negev), Asaf Shabtai (Ben-Gurion University of Negev)
CodeSafty and PrivacyTransformerLarge Language ModelSupervised Fine-TuningText
π― What it does: To address training data extraction attacks in fine-tuning large language models, the paper proposes the SCP-Ξr algorithm based on the SpaRPS sparsity attribute and base model smoothing, achieving theoretical and empirical protection while maintaining practicality.