ICML 2026 Papers — Page 29
International Conference on Machine Learning · 6554 papers
InteractComp: Evaluating Search Agents With Ambiguous Queries
Mingyi Deng (DeepWisdom), Yuyu Luo (Hong Kong University of Science and Technology Guangzhou)
Explainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AITextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This paper proposes the INTERACTCOMP benchmark to evaluate whether search agents can identify and proactively engage in clarifying interactions when faced with ambiguous queries.
Interaction-Breaking Adversarial Learning Framework for Robust Multi-Agent Reinforcement Learning
Sunwoo Lee (UNIST), Seungyul Han (UNIST)
Adversarial AttackGraph Neural NetworkTransformerReinforcement LearningAuto EncoderGenerative Adversarial NetworkContrastive LearningTabularTime SeriesSequential
🎯 What it does: Propose and implement an interaction-breaking adversarial learning framework (IBAL) based on mutual information, which trains multi-agent reinforcement learning strategies that can maintain efficient collaboration even when interactions are disrupted through obscuring and perturbing observations and actions.
Interactive Person Retrieval via Multi-Turn Multimodal Conversation
Yang Bai (Wuhan University), Mang Ye (Wuhan University)
RetrievalTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodality
🎯 What it does: Propose a multi-round, multi-modal interactive person retrieval (MInterPR) framework that allows users to iteratively refine retrieval results through visual difference feedback.
Interactive Segmentation with Elaborate Focus Prior
Kangpeng Hu (Nanjing University of Science and Technology), Quansen Sun (Nanjing University of Science and Technology)
SegmentationConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningImage
🎯 What it does: Proposes EFPNet, an end-to-end interactive segmentation framework that can automatically learn focus views, perform triple affinity correction, and achieve local refinement through click-focus integration;
InteractScience: Programmatic and Visually-Grounded Evaluation of Interactive Scientific Demonstration Code Generation
Qiaosheng Chen (Nanjing University), Fei Yuan (Shanghai Artificial Intelligence Laboratory)
AI Code AssistantTransformerLarge Language ModelPrompt EngineeringVision Language ModelTextMultimodalityBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposed the InteractScience benchmark to evaluate the ability of large language models in generating interactive science demonstration code;
Interleaved Selective State Space Models for Efficient WiFi-Based 3D Multi-Person Pose Estimation
Quang-Anh N.D. (VinUniversity), Kok-Seng Wong (VinUniversity)
Pose EstimationTransformerContrastive LearningOptical FlowTime Series
🎯 What it does: This paper proposes a 3D multi-person human pose estimation method based on WiFi CSI called WiFi-Mamba.
Internalizing Safety Understanding in Large Reasoning Models via Verification
Yi Zhang (University of Science and Technology of China), An Zhang (University of Science and Technology of China)
Safty and PrivacyReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringText
🎯 What it does: Proposed and implemented the SInternal framework, which enables the model to internalize safety norms by training the LLM to verify the safety of its own generated answers;
Interpretability and Generalization Bounds for Learning Spatial Physics
Alejandro Francisco Queiruga (OpenAI), Shuai Jiang (Sandia National Laboratories)
Numerical AnalysisExplainability and InterpretabilityTabularPhysics Related
🎯 What it does: This paper investigates the impact of training data in different function spaces on the generalization ability of machine learning models when learning linear partial differential equations (such as the Poisson equation), and theoretically proves and experimentally validates the subspace constraints.
Interpretability Driven Evolutionary Approach for the Design of Biological Sequences
Akash Pandey (Northwestern University), Sinan Keten (Northwestern University)
OptimizationExplainability and InterpretabilityDrug DiscoveryReinforcement LearningContrastive LearningBiomedical Data
🎯 What it does: Integrate interpretable models to guide evolutionary mutations, accelerating the design of biological sequences (such as proteins, DNA) for target properties
Interpretability Transfer from Language to Vision via Sparse Autoencoders
Alexey Kravets (University of Bath), Vinay P. Namboodiri (University of Bath)
Explainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelAuto EncoderContrastive LearningImageTextMultimodality
🎯 What it does: Proposed the VISTA framework, which transfers knowledge from sparse autoencoders (SAE) in language models to visual inputs, enabling visual tokens to directly fall into the text SAE space of LLMs, achieving visual interpretability and controllability without training dedicated visual SAEs.
Interpretable Discovery of One-parameter Subgroups: A Modular Framework for Elliptical, Hyperbolic, and Parabolic Symmetries
Pavan Karjol (Indian Institute of Science), Prathosh AP (Indian Institute of Science)
OptimizationExplainability and InterpretabilityContrastive LearningTabularTime SeriesSequentialBiomedical DataPhysics Related
🎯 What it does: This paper proposes a modular data-driven framework that can end-to-end learn the objective function and automatically discover its underlying one-dimensional continuous symmetry subgroups from data, without requiring prior knowledge of symmetry.
Interpretable Embeddings with Sparse Autoencoders: A Data Analysis Toolkit
Nicholas Jiang, Neel Nanda
RetrievalAnomaly DetectionExplainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelPrompt EngineeringAuto EncoderContrastive LearningTextMultimodality
🎯 What it does: This study uses sparse autoencoders (SAE) to generate interpretable embeddings and applies them to four tasks: data discrepancy detection, concept association mining, controllable clustering, and attribute retrieval, exploring the relationship between model outputs and dataset characteristics.
Interpretable Functional Koopman Learning with Non-Markovian Closure for Spatiotemporal Systems
Wanfeng Lu (Fudan University), Qunxi Zhu (Fudan University)
Explainability and InterpretabilityComputational EfficiencyRecurrent Neural NetworkTransformerAuto EncoderImageTabularTime SeriesPhysics RelatedStochastic Differential Equation
🎯 What it does: Proposed MERLIN, an interpretable framework based on function Koopman learning, to efficiently predict the spatiotemporal evolution of PDEs and achieve arbitrary resolution reconstruction from random, partial, and irregular observations.
Interpretable Neural ODEs for Gene Regulatory Network Discovery under Perturbations
Zaikang Lin (New York Genome Center), David A. Knowles (Columbia University)
Explainability and InterpretabilityDrug DiscoveryRecurrent Neural NetworkGraph Neural NetworkDiffusion modelScore-based ModelContrastive LearningTabularTime SeriesBiomedical DataOrdinary Differential Equation
🎯 What it does: Modeling the dynamic evolution of gene expression under different genetic perturbations through an interpretable Neural ODE combined with a modular single-layer perceptron, directly inferring gene regulatory networks (GRN) from this, which are then used to predict the cellular states of unseen perturbations.
Interpretable Self-Supervised Learning via Representer Landmarks and Nyström Approximation
Maedeh Zarvandi (Technical University of Munich), Debarghya Ghoshdastidar (Technical University of Munich)
Explainability and InterpretabilityRepresentation LearningAuto EncoderContrastive LearningImageTabular
🎯 What it does: Proposes the KREPES framework, achieving interpretability of self-supervised learning model representations through Representer Landmarks and Nyström approximation.
Interpreting and Enhancing Emotional Circuits in Large Vision-Language Models via Cross-Modal Information Flow
Chengsheng Zhang (University of Science and Technology of China), Xinmei Tian (University of Science and Technology of China)
Explainability and InterpretabilityRepresentation LearningPrompt EngineeringVision Language ModelAuto EncoderContrastive LearningImageTextMultimodality
🎯 What it does: This paper studies the internal emotional processing mechanisms of large vision-language models (LVLMs), proposes a vector-based causal attribution framework, constructs an emotional contrast dataset, and reveals a three-stage emotional circuit named Adapt-Aggregate-Execute;
Interpreting and Steering State-Space Models via Activation Subspace Bottlenecks
Vamshi Sunku Mohan (AppViewX), Chandan Singh (Microsoft Research)
Explainability and InterpretabilityComputational EfficiencyTransformerAuto EncoderContrastive LearningTextBenchmark
🎯 What it does: This paper identifies and locates activation subspace bottlenecks in state space models such as Mamba through mechanism interpretability techniques, and proposes a posterior intervention that scales these bottlenecks during inference.
Interpreting Genomic Language Models using Sparse Autoencoders
Akira A Nair, Dokyoon Kim (University of Pennsylvania)
Explainability and InterpretabilityRepresentation LearningTransformerAuto EncoderContrastive LearningBiomedical Data
🎯 What it does: This paper trains sparse autoencoders (SAE) on the intermediate layer embeddings of genomic language models (such as Evo2) to extract interpretable features and align them with human genomic annotations;
Interpreting Physics in Video World Models
Sonia Joseph (FAIR, Meta Superintelligence Labs), Michael Rabbat (FAIR, Meta Superintelligence Labs)
Explainability and InterpretabilityTransformerContrastive LearningWorld ModelVideoPhysics Related
🎯 What it does: Conduct a mechanistic interpretability analysis of the internal representations in large video world models, locating and describing the distribution and organization of physical-related information.
Intervene When It Doubts: Conjunction-Guided Interactive Reasoning
Qianyue Wang (South China University of Technology), Mingkui Tan (South China University of Technology)
Explainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: To address the inefficiency caused by excessive thinking and overthinking in large reasoning models (LRM) during the reasoning process, the Conjunction-Guided Intervention (CGI) framework is proposed. It pauses reasoning at conjunctions, receives external feedback based on the reasoning state, and dynamically controls the reasoning length.
Interventional Processes For Causal Uncertainty Quantification
Hugh Dance (University College London), Arthur Gretton (University College London)
OptimizationFederated LearningExplainability and InterpretabilityDrug DiscoveryContrastive LearningGaussian SplattingTabularTime SeriesSequentialBenchmarkFinance Related
🎯 What it does: This paper proposes a framework based on Gaussian processes (IMPSPEC) for uncertainty quantification of causal functions represented via inner products in RKHS (such as conditional average treatment effects).
Intra-Modal Neighbors Never Lie: Rectifying Inter-Modal Noisy Correspondence via Graph-Based Intra-Modal Reasoning
Yang Liu (Sichuan University), Jiancheng Lv (Sichuan University)
Data SynthesisRetrievalDomain AdaptationRepresentation LearningGraph Neural NetworkTransformerVision Language ModelDiffusion modelContrastive LearningGaussian SplattingImageTextMultimodality
🎯 What it does: Proposes a continuous error correction framework called IN2R based on cross-modal neighbor information, which uses graph neural networks to aggregate intra-modal neighbors and synthesize continuous soft prototypes to correct noisy correspondences in web-collected data.
Intrinsic Credit Assignment for Long Horizon Interaction
Ilze Amanda Auzina (Tubingen Ai Center), Matthias Bethge
TransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringText
🎯 What it does: Proposes Delta Belief-RL, a reinforcement learning framework that utilizes the language model's internal belief changes about the target concept as dense reward signals, designed for long-term information-seeking tasks.
Intrinsic Gradient Suppression for Label-Noise Prompt Tuning in Vision–Language Models
Jia-yu Li, Xian-Sheng Hua
ClassificationRecognitionRepresentation LearningTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodality
🎯 What it does: Propose a dual Softmax prompt tuning method without hyperparameters to address the prompt tuning problem of CLIP under label noise
Intrinsic Task Symmetry Drives Generalization in Algorithmic Tasks
Hyeonbin Hwang (KAIST), Yeachan Park (Sejong University)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerPrompt EngineeringContrastive LearningGraphTabular
🎯 What it does: Studied the driving role of intrinsic symmetry in algorithmic tasks on model generalization, discovering a three-stage training dynamic from memorization to symmetry acquisition and then to geometric organization.
Introspection Adapters: Training LLMs to Report Their Learned Behaviors
Keshav Shenoy (Anthropic Fellows), Rowan Wang (Anthropic)
Explainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringText
🎯 What it does: Train the LoRA adapter (Introspection Adapter, IA) to enable large language models to accurately report the behaviors they have learned in natural language after being fine-tuned, thereby improving the interpretability of model auditing.
Invariant Representation Learning for Source-Free Time Series Forecasting with LLM-Centric Proxy Denoising
Kangjia Yan (East China Normal University), Bin Yang (East China Normal University)
Domain AdaptationKnowledge DistillationRepresentation LearningTransformerLarge Language ModelContrastive LearningTime Series
🎯 What it does: Proposes a source-free temporal forecasting framework called TimeID, combining LLM proxy denoising with dual-branch seasonal-trend invariant feature learning, achieving transfer and adaptation to sparse target domains.
Inverse Depth Scaling From Most Layers Being Similar
Yizhou Liu (Massachusetts Institute Of Technology), Jeff Gore (Massachusetts Institute Of Technology)
OptimizationComputational EfficiencyRepresentation LearningTransformerLarge Language ModelText
🎯 What it does: Analyze the impact of LLM layer depth on loss, finding that loss decreases inversely with depth, mainly due to integration averaging between layers rather than hierarchical abstraction;
Inverse Entropic Optimal Transport Solves Semi-supervised Learning via Data Likelihood Maximization
Mikhail Persiianov (Applied AI Institute), Alexander Korotin (Applied AI Institute)
ClassificationImage TranslationDomain AdaptationOptimizationTransformerDiffusion modelScore-based ModelFlow-based ModelAuto EncoderContrastive LearningGaussian SplattingOptical FlowImageTabularTime SeriesStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: This paper proposes a new semi-supervised learning framework called EBiEOT, which can simultaneously utilize limited paired samples and a large number of unpaired samples. It learns the conditional distribution π∗(·|x) through data likelihood maximization and connects this objective with the theory of inverse entropic optimal transport (Inverse Entropic OT).
Inverting Data Transformations via Diffusion Sampling
Jinwoo Kim (KAIST), Siamak Ravanbakhsh (Mila - Quebec Artificial Intelligence Institute)
OptimizationRepresentation LearningDiffusion modelScore-based ModelImagePoint CloudTabularTime SeriesStochastic Differential Equation
🎯 What it does: Studied how to reverse unknown data transformations by performing diffusion sampling on Lie groups, to enhance the equivariance and robustness of pre-trained networks during testing.
Investigating Advanced Reasoning of Large Language Models via Black-Box Environment Interaction
Congchi Yin (Nanjing University of Aeronautics and Astronautics), Piji Li (Nanjing University of Aeronautics and Astronautics)
Autonomous DrivingAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringWorld ModelTextBenchmarkChain-of-Thought
🎯 What it does: Designed and implemented a black-box environment interaction evaluation framework, constructing the ORACLE benchmark covering six types of tasks to assess the high-level reasoning capabilities of LLMs in unknown environments.
Investigating Component Contributions in Multi-Agent ML Systems
Junsung Kim (Celestra), Dylan Yihan Dai (Celestra)
Autonomous DrivingOptimizationFederated LearningData-Centric LearningRobotic IntelligenceReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextTabularTime SeriesReview/Survey PaperBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This paper systematically analyzes the actual contributions of five key components—iterative feedback, multi-agent collaboration, memory, planning, and retrieval—in enhancing agent performance through over four thousand experiments on an automated machine learning engineering agent system.
Investigating Memory in Model-Free RL with POPGym Arcade
Zekang Wang (University of Macau), Steven Morad (University of Macau)
Recurrent Neural NetworkTransformerReinforcement LearningImage
🎯 What it does: Propose the POPGym Arcade environment and four memory evaluation tools, systematically studying the role and limitations of memory in model-agnostic reinforcement learning.
InvGNN: Learning Invertible Node Representations on Graphs
Giannis Nikolentzos (University of Peloponnese), Nikolaos Nakis (Yale University)
Computational EfficiencyRepresentation LearningGraph Neural NetworkFlow-based ModelContrastive LearningGraph
🎯 What it does: Proposed an invertible graph neural network layer (INVGNN), which can achieve invertible transformations of node representations through matrix exponentiation operations;
IO-Adam: Rethinking Memory-Efficient Adaptive Optimizers from Gradient Computation
Yiting Chen (Shanghai Jiao Tong University), Junchi Yan (Shanghai Jiao Tong University)
OptimizationComputational EfficiencySupervised Fine-TuningContrastive LearningImageTextTabular
🎯 What it does: Proposed an adaptive optimizer called IO-Adam, which estimates first and second moments by separately tracking input and output gradients of weights, thereby achieving higher memory efficiency without storing the complete gradient matrix.
IPMark: A Sentence-Level Watermark for LLMs with Hierarchical Personalization and Efficient Detection
Wenbo An (Northwestern Polytechnical University), Zehao Wang (Northwestern Polytechnical University)
Safty and PrivacyExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation
🎯 What it does: This paper proposes a sentence-level watermarking framework called IPMark based on hierarchical IP address encoding, which can achieve personalized traceability at both the model and user levels when generating text with large language models.
IQA-Spider: Unifying Multi-Granularity Image Quality Assessment with Reasoning, Grounding and Referring
Xinge Peng (University of Science and Technology of China), Zhibo Chen (University of Science and Technology of China)
TransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelVision-Language-Action ModelContrastive LearningImageTextMultimodalityBenchmark
🎯 What it does: Proposes a unified multi-granularity image quality assessment framework called IQA-Spider, integrating reasoning, localization, and reference into a single large-scale multimodal model, and designing four tasks;
IRIS: Implicit Reward-Guided Internal Sifting for Mitigating Multimodal Hallucination
Yuanshuai li, Yaochu Jin (Westlake University)
Explainability and InterpretabilityComputational EfficiencyData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision Language ModelImageTextMultimodalityBenchmark
🎯 What it does: Propose a self-alignment framework called IRIS based on internal implicit rewards to reduce visual hallucinations in multi-modal large language models.
IRPM: Intergroup Relative Preference Modeling for Pointwise Generative Reward Models
Haonan Song (HUJING Digital Media & Entertainment Group), Fan Yang (Tsinghua University)
Reinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextBenchmarkChain-of-Thought
🎯 What it does: Propose the IRPM (Intergroup Relative Preference Modeling) method, which trains a pointwise generative reward model (GRM) using pairwise preference data and optimizes it via RL, significantly reducing computational complexity during the RLHF process.
Is Code Better Than Language for Algorithmic Reasoning?
Terry Tong (University of Pennsylvania), Dan Roth (University of Pennsylvania)
OptimizationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningAI Code AssistantTransformerLarge Language ModelPrompt EngineeringTextTabularBenchmark
🎯 What it does: Studied the differences in performance between natural language reasoning and code execution in algorithmic reasoning, and proposed a three-route framework (Route 1: natural language; Route 2: code + LLM simulation; Route 3: code + Python execution) to separate the representation and execution phases.
Is Data Shapley Not Better than Random in Data Selection? Ask NASH
Xiao Tian (National University of Singapore), Bryan Kian Hsiang Low (National University of Singapore)
ClassificationExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningTextTabularBenchmark
🎯 What it does: This paper investigates the issue where Data Shapley and its semi-value methods perform worse than random selection in certain scenarios in the data selection task, and proposes the NASH framework: first decomposing the target utility (e.g., validation accuracy) into sub-utilities that can be effectively measured by Shapley values, such as the correctness of individual validation samples, and then summing these sub-utilities using nonlinear aggregation (threshold distribution function) to enhance the effectiveness of subset selection.
Is Fixing Schema Graphs Necessary? Full-Resolution Graph Structure Learning for Relational Deep Learning
Yi Huang (Beihang University), Jianxin Li (Beihang University)
Recommendation SystemOptimizationData-Centric LearningGraph Neural NetworkSupervised Fine-TuningContrastive LearningGraphTabularBenchmark
🎯 What it does: Propose the FROG framework, which achieves full-resolution and optimizable graph structure learning in relational deep learning (RDL). It can end-to-end learn the role of tables in the graph (node or edge), and capture cross-table dependencies through relation-driven message passing.
Is Generation Required for Data-Efficient Perception?
Jack Brady (Max Planck Institute for Intelligent Systems), Wieland Brendel (Max Planck Institute for Intelligent Systems)
Representation LearningData-Centric LearningTransformerDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImage
🎯 What it does: This paper explores whether generative methods are necessary to build internal representations when achieving human-level data-efficient visual perception;
Is Graph Mixup Beneficial? Investigating Interpolation And Empirical Performance of Graph Mixup Methods
Simon Forbat (University of Mannheim), Rainer Gemulla (University of Mannheim)
ClassificationHyperparameter SearchData-Centric LearningGraph Neural NetworkContrastive LearningGraphBenchmark
🎯 What it does: Conduct an independent and unified experimental evaluation of graph Mixup methods in the task of graph classification, examining their impact on model generalization.
Is One Layer Enough? Understanding Inference Dynamics in Tabular Foundation Models
Amir Rezaei Balef (TU Dortmund University), Katharina Eggensperger (TU Dortmund University)
ClassificationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerReinforcement LearningPrompt EngineeringMixture of ExpertsContrastive LearningTabularBenchmark
🎯 What it does: This paper systematically analyzes the reasoning dynamics of six mainstream tabular foundation models (TabPFN, TabICL, etc.) along the depth dimension through large-scale mechanism studies, revealing characteristics such as iterative refinement and hierarchical redundancy;
Is Spurious Correlation Removal Always Learnable?
Yibo Zhou (Beihang University), Ruifan Zhang (Beihang University)
Domain AdaptationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningGaussian SplattingImageTabularBenchmark
🎯 What it does: This paper investigates the computational-statistical disentanglement of spurious correlations under a multi-environment linear Gaussian setting; by constructing samplable instances, it demonstrates that even if the invariant subspace is statistically identifiable, computational hardness may prevent constant-accuracy recovery from being achievable in polynomial time.
Is Task-Specific Training Necessary for Anomaly Detection?
Xingwu Zhang (Hunan University), ZIJUN LONG (Hunan University)
RetrievalAnomaly DetectionTransformerAuto EncoderContrastive LearningImageRetrieval-Augmented Generation
🎯 What it does: Proposes a retrieval-based, no-training, multi-class unsupervised anomaly detection framework called RAD, which utilizes a frozen encoder and a multi-layer anomaly-free feature memory to perform global context retrieval and multi-layer patch matching under spatial conditions, directly calculating the nearest neighbor distance as the anomaly score.
Is the Last Layer Sufficient for Uncertainty Quantification?
Joseph Wilson (University of Queensland), Fred Roosta (University of Queensland)
ClassificationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerImageTextTabular
🎯 What it does: This paper studies uncertainty quantification in deep neural networks, comparing Bayesian GLM methods that linearize only the final layer (LL-GLM) with those that linearize the entire network (DNN-GLM).
Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks
Songwen Zhao (Carnegie Mellon University), Lei Li (Carnegie Mellon University)
Safty and PrivacyAI Code AssistantTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmark
🎯 What it does: Studied the security of using large language models (LLM) to generate code in real-world software engineering, and proposed a new evaluation benchmark.
Is Your Diffusion Sampler Actually Correct? A Sampler-Centric Evaluation of Discrete Diffusion Language Models
Luhan Tang (University of California, Riverside), Greg Ver Steeg (University of California, Riverside)
GenerationData-Centric LearningTransformerDiffusion modelScore-based ModelContrastive LearningText
🎯 What it does: Propose an evaluation framework based on oracle, where the denoiser learned by the discrete diffusion model is replaced with the exact HMM posterior, allowing for the independent evaluation of the sampler's error.
Is Your LLM Overcharging You? Tokenization, Transparency, and Incentives
Ander Artola Velasco (Max Planck Institute for Software Systems), Manuel Gomez Rodriguez (Max Planck Institute for Software Systems)
OptimizationFederated LearningSafty and PrivacyExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This paper studies the billing mechanism of LLM-as-a-service through a principal-agent model, proving that token-based billing creates implicit profit motives for service providers, and proposes a incentive-compatible scheme based on character-based billing, while designing a heuristic algorithm to achieve excessive charging without being detected.
IsoCompute Playbook: Optimally Scaling Sampling Compute for LLM RL
Zhoujun Cheng (University of California San Diego), Aviral Kumar (Carnegie Mellon University)
OptimizationComputational EfficiencyTransformerLarge Language ModelReinforcement LearningPrompt EngineeringText
🎯 What it does: The study investigates how to optimally allocate the number of parallel episodes, the number of problems per batch, and the number of gradient update steps under a fixed sampling computational budget during RL after training of large language models (LLMs), in order to improve performance.
It's a TRAP! Task-Redirecting Agent Persuasion Benchmark for Web Agents
Karolina Korgul (University of Oxford), Adel Bibi (University of Oxford)
Safty and PrivacyAdversarial AttackTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Designed and released TRAP (Task-Redirecting Agent Persuasion Benchmark), an evaluation framework that systematically decomposes prompt injection attacks into four modular dimensions: interface, persuasion principles, LLM operation methods, injection location, and customization, applied to six real website clones to test the security of six large language model (LLM) agents.
It's TIME: Towards the Next Generation of Time Series Forecasting Benchmarks
Zhongzheng Qiao (Nanyang Technological University), Chenghao Liu (Datadog AI Research)
Autonomous DrivingOptimizationData-Centric LearningTransformerLarge Language ModelTabularTime SeriesBenchmarkFinance Related
🎯 What it does: Built a task-centric benchmark for time series foundational models called TIME, and conducted rigorous zero-shot evaluations.
Item Response Scaling Laws: A Measurement Theory Approach for Efficient and Generalizable Neural Scaling Estimation
Sang T. Truong (Stanford University), Sanmi Koyejo (Stanford University)
Computational EfficiencyRepresentation LearningData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningTextBenchmark
🎯 What it does: Proposed the Item Response Scaling Laws (IRSL) framework, combining Item Response Theory (IRT) with neural network scaling laws. It models the probability responses of the model (such as token probability, pass@1) using Beta-IRT and verifies it through scaling experiments during large-scale pre-training and testing.
Iterated Population Based Training with Task-Agnostic Restarts
Alexander Chebykin (Centrum Wiskunde & Informatica), Peter Bosman
Hyperparameter SearchReinforcement LearningImageVideo
🎯 What it does: Proposed an Iterative Population-Based Training (IPBT) framework that can automatically adjust the step size during training and reuse partially trained weights and hyperparameters through a restart strategy.
Iterative Refinement Neural Operators are Learned Fixed-Point Solvers: A Principled Approach to Spectral Bias Mitigation
Xiaotian Liu (Dartmouth College), Yaoqing Yang (Dartmouth College)
OptimizationComputational EfficiencyConvolutional Neural NetworkTransformerSupervised Fine-TuningDiffusion modelContrastive LearningImagePoint CloudTabularTime SeriesBenchmarkPhysics Related
🎯 What it does: Propose the Iterative Refinement Neural Operator (IRNO), which introduces a learnable iterative refinement module on top of a pre-trained neural operator, achieving multi-step residual correction to address high-frequency spectral bias issues.
Iterative Robust Satisficing: Minimizing Performance Degradation Under Distribution Shift
Enes Ağırman, Cem Tekin (Bilkent University)
Domain AdaptationOptimizationComputational EfficiencyContrastive LearningImageTabularTime Series
🎯 What it does: Propose a gradient-driven training method called IRS, which directly minimizes fragility in the robust satisfaction objective, thereby enhancing the model's robustness under distribution shifts.
iTryOn: Mastering Interactive Video Virtual Try-On with Spatial-Semantic Guidance
Jun Zheng (Shenzhen Campus of Sun Yat-sen University), Xiaodan Liang (Shenzhen Campus of Sun Yat-sen University)
Image TranslationGenerationPose EstimationTransformerPrompt EngineeringVision-Language-Action ModelDiffusion modelContrastive LearningImageVideoTextMultimodality
🎯 What it does: Propose the Interactive Virtual Try-On (Interactive VVT) task, and develop the iTryOn framework based on a large-scale video diffusion Transformer;
ITSPACE: Monotone Gaussian Optimal Transport Updates
Woojoo Na (Northeastern University), Jennifer Dy (Northeastern University)
OptimizationContrastive LearningGaussian SplattingTabularBenchmark
🎯 What it does: Proposed a method called ITSPACE for aligning covariance matrices under the Gaussian optimal transport objective.
iVGR: Internalizing Visually Grounded Reasoning for MLLMs with Reinforcement Learning
Chang-Bin Zhang (University of Hong Kong), Kai Han (University of Hong Kong)
Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationReinforcement Learning from Human FeedbackTransformerSupervised Fine-TuningReinforcement LearningPrompt EngineeringVision Language ModelImageTextMultimodalityChain-of-Thought
🎯 What it does: A dual-stream training framework based on reinforcement learning is proposed, internalizing visual localization capabilities, allowing multi-modal large language models to reason using only textual chain-of-thought during inference without explicitly outputting bounding boxes.
IVQ: Structured and Lightweight Vector Quantization via Binary Hierarchical Composition Inspired by $\textit{IChing}$
Heda Zuo (Zhejiang University), Weitao You (Zhejiang University)
CompressionAuto EncoderContrastive LearningImageVideoMultimodalityAudio
🎯 What it does: Propose a lightweight, structured vector quantization framework called IVQ, inspired by the I Ching;
IVQA-LD: Inclusive Multimodal Understanding for Population with Limb Deficiency
Yan Ke (University of Queensland), Xin Yu (Adelaide University)
ClassificationRecognitionPose EstimationTransformerSupervised Fine-TuningPrompt EngineeringVision Language ModelImageTextMultimodality
🎯 What it does: Constructed the first expert-annotated visual question answering dataset, IVQA-LD, which includes amputated limbs, prosthetics, and functional classification, covering daily life, rehabilitation, and sports scenarios for people with disabilities.
iWorld-Bench: A Benchmark for Interactive World Models with a Unified Action Generation Framework
Jianjie Fang (Tsinghua University), Yong Li (Tsinghua University)
GenerationAutonomous DrivingReinforcement Learning from Human FeedbackTransformerSupervised Fine-TuningVision Language ModelVision-Language-Action ModelDiffusion modelContrastive LearningWorld ModelVideoTextMultimodalityBenchmark
🎯 What it does: Proposed the iWorld-Bench benchmark, a unified action generation framework, and constructed a high-quality video dataset of 330k, designed 4.9k interactive tasks and 9-dimensional evaluation metrics to systematically assess the generation quality, trajectory following, and memory capabilities of interactive world models.
JADAI: Jointly Amortizing Adaptive Design and Bayesian Inference
Niels Leif Bracher, Stefan T. Radev (Rensselaer Polytechnic Institute)
OptimizationComputational EfficiencyData-Centric LearningRecurrent Neural NetworkTransformerReinforcement LearningDiffusion modelScore-based ModelFlow-based ModelImageMultimodalityTabular
🎯 What it does: Proposed the JADAI framework, which can simultaneously learn experimental design strategies, historical summary networks, and posterior inference networks in a single training process, achieving joint amortization of experimental design and Bayesian inference;
JADE: Bridging the Strategic-Operational Gap in Dynamic Agentic RAG
Yiqun Chen (Renmin University of China), Jiaxin Mao (Renmin University of China)
OptimizationTransformerLarge Language ModelReinforcement LearningAgentic AITextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposes the JADE framework, which unifies planning and execution, achieving end-to-end dynamic Agentic RAG joint optimization through a multi-agent game with shared parameters.
JADE: Expert-Grounded Dynamic Evaluation for Open-Ended Professional Tasks
Lanbo Lin (Alibaba International Digital Commerce Group), Guannan Zhang (Alibaba International Digital Commerce Group)
Explainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringTextTabularReview/Survey PaperBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposes JADE—a two-layer evaluation framework that generates query-level checklists using expert knowledge and performs real-time fact verification and reasoning validation on report-level statements, addressing the stability-adaptability dilemma in open-ended professional task evaluation.
JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments
Zhan Liu (Tsinghua University), Chao Zhang (Tsinghua University)
RecognitionComputational EfficiencyRepresentation LearningData-Centric LearningConvolutional Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningVision Language ModelDiffusion modelContrastive LearningImageMultimodalityPoint CloudBenchmarkRetrieval-Augmented GenerationAudio
🎯 What it does: Propose the JAEGER framework to achieve end-to-end 3D spatial localization and reasoning using RGB-D visual and FOA multi-channel audio.
Jailbreak Foundry: From Papers to Runnable Attacks for Reproducible Benchmarking
Zhicheng Fang (Shanghai Qi Zhi Institute), Wei Xu (Tsinghua University)
Adversarial AttackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextReview/Survey PaperBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Built an end-to-end system called JAILBREAK FOUNDRY (JBF), which automatically converts jailbreak papers into executable modules, and conducts repeatable benchmark evaluations under a unified harness for 30 types of attacks and 10 victim models.
Jailbreak to Protect: Buffering and Reinforcing via Temporary Jailbreaking for Safe Fine-Tuning in Large Language Models
Seokil Ham (Korea Advanced Institute of Science and Technology), Changick Kim (Korea Advanced Institute of Science and Technology)
Safty and PrivacyComputational EfficiencyTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringText
🎯 What it does: Propose the Buffer‑and‑Reinforce framework, which temporarily jailbreaks during user fine-tuning using an attachable BufferLoRA to saturate the safety degradation gradient, followed by reinforcing safety with ReinforceLoRA after fine-tuning, while maintaining task performance through QR orthogonal merging.
Jailbreaking Vision-Language Models Through the Visual Modality
Aharon Azulay (Independent), Yossi Gandelsman (Toyota Technological Institute at Chicago)
Safty and PrivacyAdversarial AttackTransformerPrompt EngineeringVision Language ModelGenerative Adversarial NetworkContrastive LearningImageTextMultimodalityBenchmark
🎯 What it does: Proposes four categories of attacks that exploit visual inputs to jailbreak VLMs, demonstrating the potential threat of the visual modality to safety alignment.
JANUS-LORA: A Balanced Low-Rank Adaptation for Continual Learning
Cheng Chen (University of Electronic Science and Technology of China), Jingkuan Song (Tongji University)
ClassificationRepresentation LearningTransformerSupervised Fine-TuningContrastive LearningImage
🎯 What it does: Propose a new continual learning framework called Janus-LoRA, addressing the issues of non-orthogonal parameter updates and feature space intrusion that exist in LoRA during continual learning
JanusPipe: Efficient Pipeline Parallel Training for Machine Learning Interatomic Potentials
Hongyu Wang (Institute of Computing Technology, Chinese Academy of Sciences), Guangming Tan (Institute of Computing Technology, Chinese Academy of Sciences)
Computational EfficiencyTransformerLarge Language ModelGraphTabularPhysics Related
🎯 What it does: A three-dimensional distributed training system named JanusPipe is proposed to address the bidirectional backward execution mode of conservative machine learning interatomic potentials (MLIP).
Joint Enhancement and Classification using Coupled Diffusion Models of Signals and Logits
Gilad Nurko (Technion Israel Institute of Technology), Joseph Keshet (Technion Israel Institute of Technology)
ClassificationRestorationTransformerDiffusion modelScore-based ModelImageMultimodalityStochastic Differential EquationAudio
🎯 What it does: Propose a joint framework that simultaneously performs signal enhancement and classification, achieving improved robustness without fine-tuning the classifier by coupling two diffusion models (one acting on the input signal, and the other acting on the logits output by the classifier).
Joint Geometric and Trajectory Consistency Learning for One-Step Real-World Super-Resolution
Chengyan Deng (University of Electronic Science and Technology of China), Wang Zhang (University of Electronic Science and Technology of China)
Super ResolutionConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelImage
🎯 What it does: Propose a GTASR method based on a consistency model to achieve first-order real-time super-resolution, utilizing trajectory alignment (TA) and dual reference structure correction (DRSR) to enhance image detail and structural consistency.
Joint Learning in the Gaussian Single Index Model
Loucas Pillaud-Vivien (CERMICS, CNRS, ENPC, Institut Polytechnique de Paris), Adrien Schertzer (UMPA, ENS Lyon)
OptimizationRepresentation LearningTabular
🎯 What it does: Studied the joint gradient flow dynamics of simultaneously learning the projection direction and the univariate function in a high-dimensional Gaussian single-index model, and proved that even when the initial direction is opposite to the target, convergence to the global optimum can be achieved when the information index is sufficiently large;
Joint Model and Data Sparsification via the Marginal Likelihood
Alexander Timans (University of Amsterdam), Eric Nalisnick (Johns Hopkins University)
OptimizationExplainability and InterpretabilityData-Centric LearningTabularTime SeriesSequential
🎯 What it does: Propose a joint ARD method that simultaneously sparsifies model weights and samples by maximizing the marginal likelihood to learn the relevance of each feature and each sample.
Joint Navigation and Manipulation Planning with 3D Interaction Chains
Keming Zhang (State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences), Shuqiang Jiang (University of Chinese Academy of Sciences)
Autonomous DrivingOptimizationRobotic IntelligenceConvolutional Neural NetworkTransformerSupervised Fine-TuningReinforcement LearningVision Language ModelVision-Language-Action ModelSimultaneous Localization and MappingImageTextMultimodalityPoint CloudBenchmarkChain-of-Thought
🎯 What it does: Proposed the 3D Interaction Chains (3D-IC) framework, achieving joint planning for long-term navigation and manipulation of target objects and storage locations by mobile robots in unseen environments.
Joint-Embedding Predictive Learning of Latent Market States in U.S. Equities
Simon Mahns (Johns Hopkins University), Mido Assran
Representation LearningTransformerAuto EncoderContrastive LearningTabularTime SeriesFinance Related
🎯 What it does: Use self-supervised Joint-Embedding Predictive Architecture (JEPA) to learn low-dimensional embedded representations of cross-asset features in the daily U.S. stock market.
Joint-Space Empowerment as a Theory of Dexterous Motor Coordination
James Heald (University College London), Maneesh Sahani (University College London)
Robotic IntelligenceReinforcement LearningTabularTime Series
🎯 What it does: Propose and implement the Joint-Space Empowerment (JoSE) objective to discover low-dimensional action manifolds in musculoskeletal over-actuated motion systems, and build the JoSEPi compositional policy based on this; meanwhile, demonstrate the method's manipulation performance on high-dimensional MyoHand and MyoArm.
Judging What We Cannot Solve: A Consequence-Based Approach for Oracle-Free Evaluation of Research-Level Math
Guijin Son (Seoul National University), Youngjae Yu (Seoul National University)
Explainability and InterpretabilityComputational EfficiencyAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposed a oracle-free evaluation method called Consequence-Based Utility (CBU), which assesses the correctness of candidate solutions by evaluating their problem-solving effectiveness on neighborhood-related verifiable questions, using the candidate solution as a context example.
Judgment Operators: A Composition-Invariant Substrate for Multi-Agent Action Spaces
Jun Li (Veyra)
Autonomous DrivingOptimizationExplainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This paper proposes Judgment Operators (JO), a decision-time projection framework designed to externalize governance and correction capabilities in multi-agent LLM systems, addressing the issue of fragmented knowledge.
Just Ask: Curious Code Agents Reveal System Prompts in Frontier LLMs
Xiang Zheng (City University of Hong Kong), Cong Wang (City University of Hong Kong)
Safty and PrivacyExplainability and InterpretabilityAI Code AssistantTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringText
🎯 What it does: Proposes a self-evolving system prompt extraction framework called JUSTASK, which can automatically discover and optimize extraction strategies through interaction on black-box large language models.
Just Noticeable Difference Modeling for Deep Visual Features
Rui Zhao (Nanyang Technological University), Weisi Lin (Nanyang Technological University)
ClassificationObject DetectionSegmentationCompressionConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningImage
🎯 What it does: Proposes a FeatJND model for perceptible differences in deep visual features, used to predict the maximum tolerable feature perturbation while maintaining performance on downstream tasks.
Just Y-Prediction: Enabling Historical Cumulative Inconsistency in Label Diffusion for Learning with Noisy Label
Senyu Hou (Shanxi University), Wenjian Wang (Key Laboratory of Data Intelligence and Cognitive Computing of Shanxi Province)
ClassificationConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelContrastive LearningImageTabular
🎯 What it does: Designed a novel label diffusion training paradigm (JYP), which achieves robust learning against noisy labels by directly predicting clean labels and leveraging historical cumulative inconsistency (HCI);
Just-In-Time Reinforcement Learning: Continual Learning in LLM Agents Without Gradient Updates
Yibo Li (National University Of Singapore), Bryan Hooi (National University Of Singapore)
Autonomous DrivingOptimizationFederated LearningRobotic IntelligenceTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose Just-In-Time Reinforcement Learning (JitRL), which instantly optimizes the policy of a frozen LLM by retrieving memories without performing gradient updates.
KAGE-Bench: Fast Known-Axis Visual Generalization Evaluation for Reinforcement Learning
Egor Cherepanov (AXXX), Alexey Kovalev (MIRAI)
Convolutional Neural NetworkReinforcement LearningContrastive LearningImageBenchmark
🎯 What it does: Proposed KAGE-Env (a JAX-native 2D platform game) and KAGE-Bench (containing 34 training/evaluation configurations with known visual axes) to systematically evaluate the visual generalization performance of pixel-level reinforcement learning.
Kalman Linear Attention: Parallel Bayesian Filtering For Efficient Language Modeling and State Tracking
Vaisakh Shaj (University of Edinburgh), Amos Storkey (University of Edinburgh)
Computational EfficiencyRepresentation LearningTransformerLarge Language ModelTextStochastic Differential Equation
🎯 What it does: Propose Kalman Linear Attention (KLA) — a probabilistic sequence mixture layer that views language modeling as Bayesian filtering, utilizing an information-form Kalman filter to achieve time-parallel inference, and explicitly maintaining uncertainty in the hidden states.
KANFIS: A Neuro-Symbolic Framework for Interpretable and Uncertainty-Aware Learning
Binbin Yong (Lanzhou University), Zhao Su (Lanzhou University)
ClassificationExplainability and InterpretabilityComputational EfficiencyDiffusion modelAuto EncoderContrastive LearningTabularTime SeriesBiomedical DataElectronic Health Records
🎯 What it does: Proposes the KANFIS neuro-symbolic framework, combining ANFIS with Kolmogorov-Arnold networks to achieve interpretable fuzzy reasoning and uncertainty awareness.
KAST-BAR: Knowledge-Anchored Semantically-Dynamic Topology Brain Autoregressive Modeling for Universal Neural Interpretation
Haoning Wang (Beihang University), Yang Li (Beihang University)
ClassificationAnomaly DetectionExplainability and InterpretabilityRepresentation LearningGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsAuto EncoderContrastive LearningTextMultimodalityTime SeriesBiomedical DataElectrocardiogramReview/Survey Paper
🎯 What it does: Proposes KAST-BAR, a knowledge-anchored and dynamic topology-aware EEG foundation model for cross-task neural decoding and interpretation.
KBQA-R1: Reinforcing Large Language Models for Knowledge Base Question Answering
Xin Sun (University of Science and Technology of China), Liang Wang (Institute of Automation, Chinese Academy of Sciences)
OptimizationKnowledge DistillationRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextGraphRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: By transforming KB QA into a multi-round interactive decision-making process, using structured action space and execution feedback to train LLM strategies, the KBQA-R1 framework is proposed;
Keep Everyone Happy: Online Fair Division of Numerous Items with Few Copies
Arun Verma (Singapore-Mit Alliance For Research And Technology Centre), Bryan Kian Hsiang Low (National University Of Singapore)
OptimizationReinforcement LearningTabular
🎯 What it does: Proposed an online fair allocation model for scenarios with a large number of items and only a small number of copies per item, transforming it into a contextual multi-armed bandit problem, and provided UCB and TS algorithms to achieve a balance between fairness and efficiency.
Keep It in Mind: User Centric Continual Spatial Intelligence Reasoning in Egocentric Video Streams
Yun wang, Angela Yao (National University of Singapore)
Object DetectionObject TrackingSegmentationDepth EstimationAutonomous DrivingRepresentation LearningData-Centric LearningRobotic IntelligenceTransformerLarge Language ModelVision Language ModelVideoMultimodalityBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose the UCS-Bench dataset and the DirectMe framework to evaluate and enhance user-centered continuous spatial reasoning in front-facing camera streaming video.
Keeping a Secret Requires a Good Memory: Unconditional Streaming Lower-Bounds for Differentially Private Algorithms
Alessandro Epasto (Google Research), Pasin Manurangsi (Google Research)
Safty and Privacy
🎯 What it does: This paper proposes an unconditional spatial lower bound in user-level differential privacy streaming algorithms, proving that under the privacy constraint, the algorithm must use at least Ω(T^{γ_w+γ_k-2γ_h}) memory to maintain accuracy;
Kernel-based Maximum-of-difference Test for Two-sample Comparison
Dan Pu (Southwestern University of Finance and Economics), Wei Lan (Southwestern University of Finance and Economics)
Anomaly DetectionComputational EfficiencyRepresentation LearningContrastive LearningImageTabular
🎯 What it does: Proposed a kernel-based two-sample test method called MOD (Maximum-of-Difference) test, and extended it to multi-sample scenarios; calculated the squared maximum value of the kernel distance difference between within-sample and between-sample for each observation point as the statistic; also designed a fused version that aggregates multi-kernel information.
KernelBand: Steering LLM-based Kernel Optimization via Hardware-Aware Multi-Armed Bandits
Dezhi Ran (Peking University), Tao Xie (Peking University)
OptimizationLarge Language ModelReinforcement LearningMixture of ExpertsTabularTime SeriesBenchmark
🎯 What it does: The optimization problem of GPU kernels is modeled as a context-aware multi-armed bandit (MAB) problem with hardware constraints, combining code generated by LLMs with MAB strategies.
KernelCraft: Benchmarking for Agentic Close-to-Metal Kernel Generation on Emerging Hardware
Jiayi Nie (University of Cambridge), Yiren Zhao (Imperial College London)
OptimizationTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This paper proposes KernelCraft, a low-level assembly kernel generation and evaluation framework tailored for emerging hardware, which leverages LLM agents to automatically generate, debug, and optimize custom ISA bare-metal kernels through tool calls and feedback loops;
KernelFoundry: Hardware-Aware Evolutionary GPU Kernel Optimization
Nina Wiedemann (Intel Corporation), Benjamin Ummenhofer (Intel Corporation)
OptimizationAI Code AssistantNeural Architecture SearchLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmark
🎯 What it does: KernelFoundry automatically generates high-performance SYCL/CUDA GPU kernels using large language models by combining MAP-Elites quality-diversity search, meta-prompt evolution, and templated parameter optimization.
KineFlow: Kinematic Second-Order Flow Matching for Time-Series Forecasting
Haiqi Jiang (Hong Kong University of Science and Technology), Hui Xiong (Hong Kong University of Science and Technology)
GenerationOptimizationRecurrent Neural NetworkTransformerDiffusion modelScore-based ModelFlow-based ModelTabularTime SeriesBenchmarkPhysics RelatedStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Propose KineFlow, a generative time series forecasting framework that utilizes second-order flow matching and phase-space neural acceleration fields.
Kinematics-Driven Gaussian Shape Deformation for Blurry Monocular Dynamic Scenes
Yeon-Ji Song (Seoul National University), Byoung-Tak Zhang (Seoul National University)
RestorationNeural Radiance FieldGaussian SplattingVideo
🎯 What it does: Proposed a kinematics-based 3D Gaussian splitting framework, Kinematics-GS, for reconstructing dynamic 3D scenes from monocular motion-blurred videos.