International Conference on Machine Learning Β· 1032 papers
Generalized Correctness Models: Learning Calibrated and Cross-Model Correctness Predictors from Historical Patterns
Hanqi Xiao (University of North Carolina Chapel Hill), Mohit Bansal (University of North Carolina Chapel Hill)
CodeExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought
π― What it does: Studied a cross-model confidence calibration method based on historical prediction data, called Generalized Correctness Models (GCM), used to predict the correctness of answers from multiple large language models (LLMs).
Generative Adaptation of Dynamics to Environmental Shifts via Weight-space Diffusion
Ruikun Li (Tsinghua University), Yong Li (Tsinghua University)
CodeGenerationData SynthesisDomain AdaptationComputational EfficiencyMeta LearningGraph Neural NetworkTransformerPrompt EngineeringMixture of ExpertsDiffusion modelAuto EncoderTime SeriesSequentialPhysics Related
π― What it does: Propose DynaDiff, which directly generates environment-specific dynamic prediction models on observed short sequences using weight space diffusion models, achieving rapid adaptation with zero gradient fine-tuning.
π― What it does: Proposes the SDE-VI framework, which uses SDE-induced continuous-discrete variational inference for generative modeling of irregular time series.
GeoAlign: Geometric Rollout Curation for Robust LLM Reinforcement Learning
Ting Zhou (Sun Yat Sen University), Daoyuan Chen (Alibaba Group)
CodeRepresentation LearningData-Centric LearningTransformerLarge Language ModelReinforcement LearningPrompt EngineeringContrastive LearningText
π― What it does: Proposes GEOALIGN, a lightweight episode generation curve screening plugin used to detect and correct directionally inconsistent high-reward episodes in LLM reinforcement learning.
CodeGenerationData SynthesisGraph Neural NetworkTransformerDiffusion modelScore-based ModelFlow-based ModelGraphTime Series
π― What it does: This paper proposes the GeoFlow framework for predicting and generating origin-destination (OD) flows, combining geographic attributes and graph neural networks to achieve more accurate modeling.
GeoPT: Scaling Physics Simulation via Lifted Geometric Pre-Training
Haixu Wu (MIT), Wojciech Matusik (MIT)
CodeTransformerGenerative Adversarial NetworkContrastive LearningOptical FlowPoint CloudMeshTabularPhysics Related
π― What it does: Propose GeoPT, which utilizes geometric data combined with randomly synthesized velocity fields for lifted geometric self-supervised pre-training, thereby providing dynamic-aware priors for neural physics simulators;
GePBench: Evaluating Fundamental Geometric Perception for Multimodal Large Language Models
Shangyu Xing (Nanjing University), Xinyu Dai (Nanjing University)
CodeExplainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmarkRetrieval-Augmented Generation
π― What it does: Constructed GePBench, a large-scale geometry-aware benchmark dataset, systematically evaluating the perception capabilities of multi-modal large language models (MLLMs) regarding geometric shapes and their spatial relationships, and enhancing model performance on downstream tasks through retraining on this dataset.
GFFMERGE: Efficient Merging of Graph Neural Force Fields and Beyond
Parth Verma (Indian Institute of Technology Delhi), Sayan Ranu (Indian Institute of Technology Delhi)
CodeOptimizationComputational EfficiencyKnowledge DistillationDrug DiscoveryGraph Neural NetworkSupervised Fine-TuningContrastive LearningGraphTabularTime SeriesPhysics Related
π― What it does: Propose the GFFMERGE framework, achieving closed-form linear merging of graph neural network force field models, and quickly restoring joint training performance through minimal subsequent fine-tuning; simultaneously extended to general GNNs as GNNMERGE.
π― What it does: Propose the GFMate framework, which achieves no pre-training coupling and fine-tuning with prompts at test time for graph foundation models, leveraging central point prompts, layer prompts, and test-time complementary learning to enhance cross-domain node and graph classification performance.
π― What it does: Proposes the GP2F dual-branch graph prompting learning framework, which integrates frozen pre-trained GNNs with a lightweight adapter branch for cross-domain few-shot graph tasks.
CodeClassificationExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelContrastive LearningGaussian SplattingText
π― What it does: This study proposes an scalable Gaussian process (GP) Bayesian LoRA framework called GPan-LoRA, which utilizes sparse GP approximation and autoregressive variational inference to achieve Bayesian low-rank fine-tuning of large language models and reliable uncertainty quantification.
Graph-GRPO: Training Graph Flow Models with Reinforcement Learning
Baoheng Zhu (Beijing University of Posts and Telecommunications), Xiao Wang (Beihang University)
CodeGenerationOptimizationDrug DiscoveryGraph Neural NetworkReinforcement LearningScore-based ModelFlow-based ModelGraphBiomedical Data
π― What it does: Propose an online reinforcement learning framework called Graph-GRPO, which aligns the discrete flow matching graph generative model (Graph Flow Model) with task-specific rewards, overcoming the issues of non-differentiability and low exploration efficiency in traditional GFMs during sampling.
CodeData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextSequentialElectronic Health RecordsRetrieval-Augmented Generation
π― What it does: Learning an active questioning strategy without a simulator, Learn-to-Ask, which generates dense and realistic rewards by 'looking back' using offline expert dialogue logs, training LLMs to decide what questions to ask and when to stop.
Grouter: Decoupling Routing from Representation for Accelerated MoE Training
Yuqi Xu (Peking University), Kun Yuan (Peking University)
CodeComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerMixture of ExpertsText
π― What it does: Propose a hybrid expert (MoE) training framework with pre-learned and fixed routing structure (GROUTER), decoupling routing from expert weight updates to accelerate and improve model convergence quality.
Michael Sullivan (Saarland University), Alexander Koller (Saarland University)
CodeOptimizationComputational EfficiencyReinforcement Learning from Human FeedbackLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkChain-of-Thought
π― What it does: Demonstrate that GRPOζ¬θ΄¨δΈ is inherently an implicit process reward model, and propose the Ξ»-GRPO improved algorithm based on this insight
π― What it does: Proposes the GRO K framework, which uses a cluster decision agent based on GRPO to autonomously estimate the unknown number of clusters K in multi-view clustering, forming a perception-decision-feedback closed loop.
GUDA: Counterfactual Group-wise Training Data Attribution for Diffusion Models via Unlearning
Naoki Murata (Sony AI), Yuki Mitsufuji (Sony AI)
CodeGenerationExplainability and InterpretabilityComputational EfficiencyData-Centric LearningReinforcement LearningDiffusion modelScore-based ModelContrastive LearningImageText
π― What it does: Proposed the GUDA framework, which uses machine unlearning to approximately perform group deletion (LOGO) on generative models and quantifies the impact of each training data group on the generation results;
π― What it does: Propose GUI-Spotlight, a graphical user interface visual localization model that utilizes multiple tools and iterative reinforcement learning for focused attention;
Guidance: Sentence-Level Citation Enforcement via Prefix-Tail Guidance during LLM Decoding
Yirui Zhan (Peking University), Jun Gao (Peking University)
CodeRetrievalComputational EfficiencyData-Centric LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBenchmarkRetrieval-Augmented Generation
π― What it does: Proposes a framework called Guidance, which offers training freedom and enforces sentence-level citations during the LLM decoding phase.
π― What it does: This paper proposes a contrastive learning framework called HCL for heterogeneous hypergraphs, which utilizes node-hyperedge heterogeneity to guide view generation and encoding, achieving robust representation learning.
π― What it does: Designed a role-asymmetric Hamiltonian Asymmetric Fusion (HAF) module to achieve a secure multi-step iterative fusion under modal imbalance.
π― What it does: Reintroduce hard labels as correction signals in dataset distillation to alleviate the problem of local semantic drift caused by limited soft labels.
π― What it does: Proposes an NSPSG framework that combines unconstrained diffusion models with discrete projection to enforce hard constraints in graph generation.
π― What it does: This paper transforms electroencephalogram (EEG) signals into structured spectrum videos (Spectrum Video) through time-frequency transformation and spatial mapping, and utilizes video MAE for self-supervised pre-training to achieve generalization across subjects and electrode layouts.
HASTE: Hardware-Aware Dynamic Sparse Training for Large Output Spaces
Nasib Ullah (Aalto University), Rohit Babbar (University of Bath)
CodeClassificationTextTabular
π― What it does: The HASTE framework is proposed for extreme multi-label classification, achieving efficient sparse training by grouping labels and sharing fixed fan-in connections, eliminating the need for auxiliary objectives.
Shijie Mei (Institute of Automation Chinese Academy of Sciences), Guoqi Li (Institute of Automation Chinese Academy of Sciences)
CodeComputational EfficiencyRepresentation LearningTransformerLarge Language ModelContrastive LearningText
π― What it does: This paper proposes the Head-in-Head structure, which partitions memory states within a single linear attention head, thereby enhancing the expressive power of linear attention models while maintaining a fixed memory size.
HELIX: Hybrid Encoding with Learnable Identity and Cross-dimensional Synthesis for Time Series Imputation
Fengming Zhang (Beijing Institute of Technology), Shen Qu (Beijing Institute of Technology)
CodeRestorationAnomaly DetectionRepresentation LearningTransformerMixture of ExpertsContrastive LearningTabularTime SeriesElectronic Health Records
π― What it does: A hybrid encoding framework called HELIX is proposed for missing value imputation in multivariate time series, with the core idea of introducing learnable feature identity embeddings to provide persistent semantic anchors for each feature.
π― What it does: Propose a framework called FedSSA that simultaneously addresses node feature heterogeneity and structural heterogeneity in graph federated learning.
Heterogeneous Customizable Personalized Federated Fine-Tuning Approach for Large Language Models
xin tong, Baojiang cui
CodeFederated LearningComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningText
π― What it does: This paper proposes a heterogeneous customizable personalized federated LoRA fine-tuning framework called Het-CPFLoRA, which allows each client to simultaneously learn shared and personalized knowledge through a single adapter, and dynamically adjust weights during inference.
π― What it does: Propose a training-agnostic hypergraph active learning framework, HIAL, which constructs a node selection strategy that maximizes information by utilizing high-order interaction weighted projection and linear diffusion.
Hidden in Plain Sight -- Class Competition Focuses Attribution Maps
Nils Philipp Walter (CISPA Helmholtz Center for Information Security), Jonas Fischer (Max Planck Institute for Informatics)
CodeClassificationObject DetectionSegmentationExplainability and InterpretabilityConvolutional Neural NetworkTransformerImageBiomedical Data
π― What it does: Proposed a lightweight post-processing method called Attribution Lens (AL), which transforms attribution results from a single logit into an attribution distribution across multiple classes, significantly improving the target specificity and interpretability of attribution while keeping the original method unchanged.
π― What it does: This paper proposes a hierarchical reinforcement learning (HRL) framework for searching non-Hirsch ideals in conformal algebra environments with sparse rewards, thereby constructing counterexamples to the Hirsch conjecture in Kalai algebra.
Hierarchical Representations for Cross-task Automated Heuristic Design using LLMs
Fei Liu (City University of Hong Kong), Qingfu Zhang (City University of Hong Kong)
CodeOptimizationMeta LearningReinforcement Learning from Human FeedbackNeural Architecture SearchTransformerLarge Language ModelPrompt EngineeringGraphTabularBenchmark
π― What it does: Designed a multi-task hierarchical search framework, MTHS, which leverages large language models to evolve general metaheuristics and task-specific procedures, enabling automatic heuristic algorithm design across tasks.
CodeRetrievalExplainability and InterpretabilityComputational EfficiencyTransformerContrastive LearningImageTextMultimodalityRetrieval-Augmented GenerationAudio
π― What it does: Proposes a binary tree-based retrieval framework called RETREEVER, which can achieve high accuracy, low latency, and provide interpretable hierarchical organization in large-scale retrieval tasks.
CodeComputational EfficiencyData-Centric LearningDrug DiscoveryTransformerVision Language ModelAuto EncoderContrastive LearningImageMultimodalityBiomedical DataBenchmark
π― What it does: Designed and implemented a HiST (Hierarchical Sparse Transformer) model for predicting gene expression in spatial transcriptomics (ST) from conventional H&E tissue sections.
HΓΆlder++: Improving Quality-Coherence Trade-off in Multimodal VAEs
Huyen Thuc Khanh Vo (Saarland University), Isabel Valera (Saarland University)
CodeGenerationRepresentation LearningMixture of ExpertsDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageMultimodality
π― What it does: This paper proposes Holder++, a multi-modal variational autoencoder, which improves the balance between generation quality and consistency by implementing symmetric Holder pooling, introducing shared and private latent subspaces (Holder+), and hierarchical inference (Holder++).
How does information access affect LLM monitors' ability to detect sabotage?
Rauno Arike (Aether Research), Francis Rhys Ward (Independent)
CodeAnomaly DetectionFederated LearningSafty and PrivacyExplainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsTextSequentialBenchmarkRetrieval-Augmented GenerationChain-of-Thought
π― What it does: This paper studies the impact of information access levels on the performance of large language model (LLM) monitors in detecting attackers' 'sabotage' behaviors, and proposes a hierarchical information filtering 'Extract-and-Evaluate (EaE)' monitoring scheme.
CodeExplainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringContrastive LearningTextChain-of-Thought
π― What it does: This paper designs a low-rank adapter method called TeleβLens to probe the internal hidden states of large language models (LLMs) layer by layer across 12 different types of tasks, revealing their planning horizon during the chain-of-thought (CoT) process. Based on this finding, we propose an adaptive uncertainty estimation that focuses only on key 'pivot' positions and a CoT skipping mechanism that leverages early answer clues.
How to Correctly Report LLM-as-a-Judge Evaluations
Chungpa Lee (Yonsei University), Kangwook Lee (University of WisconsinMadison)
CodeExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelTextBenchmark
π― What it does: Propose an unbiased estimation framework based on misclassification adjustment, using LLM as a judge to correct results and provide statistical confidence intervals, while designing an adaptive calibration sample allocation strategy to shorten the interval length.
How to Fine-Tune a Reasoning Model? A TeacherβStudent Cooperation Framework to Synthesize Student-Consistent SFT Data
Zixian Huang (Shanghai AI Laboratory), Qipeng Guo (Shanghai AI Laboratory)
CodeData SynthesisComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
π― What it does: Proposes the TESSY framework based on teacher-student collaboration, alternately generating capability and style texts for supervised fine-tuning of reasoning models.
π― What it does: This paper proposes a graph anomaly detection framework called HSMAD that simultaneously models heterogeneity in the spectral domain and the manifold domain, enhancing the ability to identify abnormal nodes by utilizing heterogeneity-weighted spectral filtering and manifold routing message updates.
Hyper-LLaVA: Hyperbolic Uncertainty-aware Modality-Balanced Routing for Multimodal Continual Instruction Tuning
Kunlun Xu (Wangxuan Institute of Computer Technology, Peking University), Jiahuan Zhou (Wangxuan Institute of Computer Technology, Peking University)
CodeOptimizationFederated LearningComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsVision Language ModelContrastive LearningImageTextMultimodalityBenchmark
π― What it does: Propose the Hyper-LLaVA framework for multi-modal continual instruction tuning, improving parameter routing
π― What it does: Propose Hyperbolic Associative Memory Networks (HAMNs), which migrate modern Hopfield networks to negative curvature hyperbolic spaces, achieving hierarchical memory retrieval based on arc-length energy.
CodeRecommendation SystemTransformerLarge Language ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningSequential
π― What it does: This paper proposes HG-Rec, a generative recommendation framework based on residual quantization and differential length codebook in hyperbolic space, aiming to improve codebook utilization and recommendation performance.
π― What it does: Propose a unified MS/HS fusion framework named SSA, which can simultaneously accommodate the number of spectral bands from different sensors and arbitrary spatial magnification scales, achieving multi-sensor joint training with a single model.
Identifiable Markov Switching Models with Instantaneous Effects and Exponential Families
Roel Hulsman (University of Amsterdam), Sara Magliacane (University of Amsterdam)
CodeFlow-based ModelTabularTime SeriesFinance Related
π― What it does: Constructing identifiable Markov switching models in non-stationary time series, allowing for instantaneous effects, nonlinear lagged effects, and restricting noise to the exponential family, while proposing a switching detection and causal structure discovery framework based on conditional regularized flows (FlowMSM).
Identifying and Correcting Label Noise for Robust GNNs via Influence Contradiction
Wei Ju (Sichuan University), Ming Zhang (Peking University)
CodeClassificationAnomaly DetectionExplainability and InterpretabilityComputational EfficiencyRepresentation LearningGraph Neural NetworkContrastive LearningGraphBenchmark
π― What it does: Propose an influence contradiction score-based graph neural network (ICGNN) for identifying and correcting label noise in graph data, and enhancing robustness in semi-supervised scenarios.
π― What it does: Propose a module called CPF that rewrites visual features through low-rank learnable primitives to enhance the performance of generalized class discovery (GCD).
iLoRA: Bayesian Low-Rank Adaptation with Latent Interaction Graphs for Microbiome Diagnosis
Yang Song (University of Copenhagen), Hengguan Huang (University of Copenhagen)
CodeClassificationExplainability and InterpretabilityDrug DiscoveryGraph Neural NetworkTransformerContrastive LearningGraphTabularBiomedical Data
π― What it does: Propose a Bayesian Graph Conditional LoRA framework (iLoRA), which simultaneously learns a prediction model and a sample-level microbial interaction network in the microbiome diagnosis task, and generates LoRA updates conditioned on this network.
π― What it does: This paper analyzes the approximation error of state representation based on the graph Laplacian spectrum in reinforcement learning, and proves that the error varies with the algebraic connectivity (Ξ»β) of the state graph.
IMPACT: Influence Modeling for Open-Set Time Series Anomaly Detection
Xiaohui Zhou (National Key Laboratory of Parallel and Distributed Computing), Guansong Pang (Singapore Management University)
CodeAnomaly DetectionRecurrent Neural NetworkTransformerAuto EncoderGenerative Adversarial NetworkContrastive LearningTime Series
π― What it does: For open-ended time series anomaly detection, the IMPACT framework is proposed, which evaluates the impact of training samples using influence functions, performs anomaly decontamination and pseudo-anomaly generation, and trains a dual-head model to achieve joint detection of known and unknown anomalies.
π― What it does: In cross-domain few-shot learning tasks without source domain data, we propose adaptive alignment of different image patches in the CLIP vision transformer: bringing the head tokens rich in semantic information closer, and pushing away the tail tokens with insufficient semantic information, to improve the classification performance in the target domain.
π― What it does: Proposed the Cumulative Memory Recurrent Unit (CMRU) and its relaxed version Ξ± CMRU, which solves the gradient blocking problem in BMRU during state updates, achieving a parallelizable trained persistent memory RNN.
π― What it does: Proposed a source-target video joint denoising framework called ReCo based on width concatenation, which utilizes natural language instructions for instructive video editing and ensures precise localization of the editing region and background preservation through regional constraints.
π― What it does: Proposes an efficient incremental update model called Transformer Neural Process (incTNP) that can operate on real-time streaming data;
Inference-Time Conformal Reasoning with Valid Factuality Control for Large Language Models
Ting Wang (University of Illinois Urbana Champaign), Huan Zhang (University of Illinois Urbana Champaign)
CodeExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelTextGraphBenchmarkChain-of-Thought
π― What it does: Proposes an Inference-Time Conformal Reasoning (ITCR) framework that applies conformal prediction in real-time during multi-step reasoning, achieving factual control over the reasoning graph and determining when to stop expanding.
π― What it does: Propose a self-attention model called IPAM, capable of jointly processing discrete and continuous values in sequences, achieving variable-length, infinite-precision vector graphics and layout generation.
π― What it does: This paper proposes InfoPO, a reinforcement learning framework tailored for user-centric multi-turn interactions, achieving finer-grained credit assignment by calculating contrastive information gain rewards at each step.
π― What it does: This paper proposes the 'Informational Asymmetric Actor-Critic' framework, which allows the critic to use arbitrary state-related privileged signals during training while maintaining unbiased policy gradients.
InteractComp: Evaluating Search Agents With Ambiguous Queries
Mingyi Deng (DeepWisdom), Yuyu Luo (Hong Kong University of Science and Technology Guangzhou)
CodeExplainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AITextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought
π― What it does: This paper proposes the INTERACTCOMP benchmark to evaluate whether search agents can identify and proactively engage in clarifying interactions when faced with ambiguous queries.
Interventional Processes For Causal Uncertainty Quantification
Hugh Dance (University College London), Arthur Gretton (University College London)
CodeOptimizationFederated LearningExplainability and InterpretabilityDrug DiscoveryContrastive LearningGaussian SplattingTabularTime SeriesSequentialBenchmarkFinance Related
π― What it does: This paper proposes a framework based on Gaussian processes (IMPSPEC) for uncertainty quantification of causal functions represented via inner products in RKHS (such as conditional average treatment effects).
π― What it does: This paper proposes a new semi-supervised learning framework called EBiEOT, which can simultaneously utilize limited paired samples and a large number of unpaired samples. It learns the conditional distribution Οβ(Β·|x) through data likelihood maximization and connects this objective with the theory of inverse entropic optimal transport (Inverse Entropic OT).
π― What it does: Studied how to reverse unknown data transformations by performing diffusion sampling on Lie groups, to enhance the equivariance and robustness of pre-trained networks during testing.
Investigating Component Contributions in Multi-Agent ML Systems
Junsung Kim (Celestra), Dylan Yihan Dai (Celestra)
CodeAutonomous DrivingOptimizationFederated LearningData-Centric LearningRobotic IntelligenceReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextTabularTime SeriesReview/Survey PaperBenchmarkRetrieval-Augmented GenerationChain-of-Thought
π― What it does: This paper systematically analyzes the actual contributions of five key componentsβiterative feedback, multi-agent collaboration, memory, planning, and retrievalβin enhancing agent performance through over four thousand experiments on an automated machine learning engineering agent system.
π― What it does: Proposed an invertible graph neural network layer (INVGNN), which can achieve invertible transformations of node representations through matrix exponentiation operations;
IPMark: A Sentence-Level Watermark for LLMs with Hierarchical Personalization and Efficient Detection
Wenbo An (Northwestern Polytechnical University), Zehao Wang (Northwestern Polytechnical University)
CodeSafty and PrivacyExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation
π― What it does: This paper proposes a sentence-level watermarking framework called IPMark based on hierarchical IP address encoding, which can achieve personalized traceability at both the model and user levels when generating text with large language models.
IQA-Spider: Unifying Multi-Granularity Image Quality Assessment with Reasoning, Grounding and Referring
Xinge Peng (University of Science and Technology of China), Zhibo Chen (University of Science and Technology of China)
CodeTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelVision-Language-Action ModelContrastive LearningImageTextMultimodalityBenchmark
π― What it does: Proposes a unified multi-granularity image quality assessment framework called IQA-Spider, integrating reasoning, localization, and reference into a single large-scale multimodal model, and designing four tasks;
π― What it does: Conduct an independent and unified experimental evaluation of graph Mixup methods in the task of graph classification, examining their impact on model generalization.
Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks
Songwen Zhao (Carnegie Mellon University), Lei Li (Carnegie Mellon University)
CodeSafty and PrivacyAI Code AssistantTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmark
π― What it does: Studied the security of using large language models (LLM) to generate code in real-world software engineering, and proposed a new evaluation benchmark.
Is Your LLM Overcharging You? Tokenization, Transparency, and Incentives
Ander Artola Velasco (Max Planck Institute for Software Systems), Manuel Gomez Rodriguez (Max Planck Institute for Software Systems)
CodeOptimizationFederated LearningSafty and PrivacyExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought
π― What it does: This paper studies the billing mechanism of LLM-as-a-service through a principal-agent model, proving that token-based billing creates implicit profit motives for service providers, and proposes a incentive-compatible scheme based on character-based billing, while designing a heuristic algorithm to achieve excessive charging without being detected.
Iterative Robust Satisficing: Minimizing Performance Degradation Under Distribution Shift
Enes AΔΔ±rman, Cem Tekin (Bilkent University)
CodeDomain AdaptationOptimizationComputational EfficiencyContrastive LearningImageTabularTime Series
π― What it does: Propose a gradient-driven training method called IRS, which directly minimizes fragility in the robust satisfaction objective, thereby enhancing the model's robustness under distribution shifts.
JADE: Bridging the Strategic-Operational Gap in Dynamic Agentic RAG
Yiqun Chen (Renmin University of China), Jiaxin Mao (Renmin University of China)
CodeOptimizationTransformerLarge Language ModelReinforcement LearningAgentic AITextBenchmarkRetrieval-Augmented Generation
π― What it does: Proposes the JADE framework, which unifies planning and execution, achieving end-to-end dynamic Agentic RAG joint optimization through a multi-agent game with shared parameters.
Jailbreaking Vision-Language Models Through the Visual Modality
Aharon Azulay (Independent), Yossi Gandelsman (Toyota Technological Institute at Chicago)
CodeSafty and PrivacyAdversarial AttackTransformerPrompt EngineeringVision Language ModelGenerative Adversarial NetworkContrastive LearningImageTextMultimodalityBenchmark
π― What it does: Proposes four categories of attacks that exploit visual inputs to jailbreak VLMs, demonstrating the potential threat of the visual modality to safety alignment.
Joint Navigation and Manipulation Planning with 3D Interaction Chains
Keming Zhang (State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences), Shuqiang Jiang (University of Chinese Academy of Sciences)
CodeAutonomous DrivingOptimizationRobotic IntelligenceConvolutional Neural NetworkTransformerSupervised Fine-TuningReinforcement LearningVision Language ModelVision-Language-Action ModelSimultaneous Localization and MappingImageTextMultimodalityPoint CloudBenchmarkChain-of-Thought
π― What it does: Proposed the 3D Interaction Chains (3D-IC) framework, achieving joint planning for long-term navigation and manipulation of target objects and storage locations by mobile robots in unseen environments.
Joint-Space Empowerment as a Theory of Dexterous Motor Coordination
James Heald (University College London), Maneesh Sahani (University College London)
CodeRobotic IntelligenceReinforcement LearningTabularTime Series
π― What it does: Propose and implement the Joint-Space Empowerment (JoSE) objective to discover low-dimensional action manifolds in musculoskeletal over-actuated motion systems, and build the JoSEPi compositional policy based on this; meanwhile, demonstrate the method's manipulation performance on high-dimensional MyoHand and MyoArm.
Just Ask: Curious Code Agents Reveal System Prompts in Frontier LLMs
Xiang Zheng (City University of Hong Kong), Cong Wang (City University of Hong Kong)
CodeSafty and PrivacyExplainability and InterpretabilityAI Code AssistantTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringText
π― What it does: Proposes a self-evolving system prompt extraction framework called JUSTASK, which can automatically discover and optimize extraction strategies through interaction on black-box large language models.
Just-In-Time Reinforcement Learning: Continual Learning in LLM Agents Without Gradient Updates
Yibo Li (National University Of Singapore), Bryan Hooi (National University Of Singapore)
CodeAutonomous DrivingOptimizationFederated LearningRobotic IntelligenceTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation
π― What it does: Propose Just-In-Time Reinforcement Learning (JitRL), which instantly optimizes the policy of a frozen LLM by retrieving memories without performing gradient updates.
Keep It in Mind: User Centric Continual Spatial Intelligence Reasoning in Egocentric Video Streams
Yun wang, Angela Yao (National University of Singapore)
CodeObject DetectionObject TrackingSegmentationDepth EstimationAutonomous DrivingRepresentation LearningData-Centric LearningRobotic IntelligenceTransformerLarge Language ModelVision Language ModelVideoMultimodalityBenchmarkRetrieval-Augmented Generation
π― What it does: Propose the UCS-Bench dataset and the DirectMe framework to evaluate and enhance user-centered continuous spatial reasoning in front-facing camera streaming video.
KITE: Knowledge-Guided Probabilistic Modeling for Time Series Forecasting with Exogenous Variables
Hanyin Cheng (East China Normal University), Chenjuan Guo (East China Normal University)
CodeGenerationData SynthesisTransformerDiffusion modelScore-based ModelFlow-based ModelContrastive LearningTabularTime SeriesFinance Related
π― What it does: Propose the KITE framework to achieve probabilistic time series forecasting with external variables, introducing historical conditions, knowledge guidance, and classifier-free guidance during the generation process.
Knapsack RL: Compute-Efficient Reinforcement Learning via Heterogeneous Rollout Allocation
Ziniu Li (Chinese University of Hong Kong), Zhi-Quan Luo (Chinese University of Hong Kong)
CodeOptimizationComputational EfficiencyTransformerLarge Language ModelReinforcement LearningPrompt EngineeringText
π― What it does: This paper studies how to heterogeneously allocate rollout budgets when fine-tuning large language models with reinforcement learning, and proposes the Knapsack RL framework;
Knowing Bias, Doing Better: Mitigating Social Bias in LLMs via Know-Bias Neuron Enhancement
Jinhao Pan (George Mason University), Ziwei Zhu (George Mason University)
CodeFederated LearningExplainability and InterpretabilityComputational EfficiencyRepresentation LearningData-Centric LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningText
π― What it does: To address social bias in large language models (LLMs), this paper proposes a lightweight method called KnowBias, which enhances the internal 'bias-aware' neurons (know-bias neurons) during inference, thereby suppressing biased outputs without compromising the model's overall capabilities.
KromHC: Manifold-Constrained Hyper-Connections with Kronecker-Product Residual Matrices
Wuyang Zhou (Imperial College London), Danilo Mandic (Imperial College London)
CodeOptimizationComputational EfficiencyRepresentation LearningTransformerLarge Language ModelText
π― What it does: A new framework called KromHC is proposed, which utilizes the Kronecker product of dual random matrices to address the training instability and parameter complexity issues of hyper-connection (HC) in neural networks.
L-CUBE: Isolating Long-Context Capacity from Knowledge with Controllable Mutual Information Scaling
Zhuo Chen (NSF AI Institute for Artificial Intelligence and Fundamental Interactions), Marin Soljacic
CodeData SynthesisComputational EfficiencyRepresentation LearningTransformerLarge Language ModelScore-based ModelContrastive LearningTextSequentialBenchmark
π― What it does: This paper proposes L-CUBE, a controllable information synthetic long sequence benchmark, used to separate the long context capture capability of language models from the confounding effects of semantic knowledge.
L-Drive: Beyond a Single MappingβLatent Context Drives Time Series Forecasting
Fan Zhang (Shandong Technology and Business University), Hua Wang (Ludong University)
CodeRecurrent Neural NetworkTransformerAuto EncoderContrastive LearningTime Series
π― What it does: Proposed the L-Drive framework, which achieves adaptive modeling of temporal changes by introducing latent context (L-Context) and patch-based relative position basis functions, thereby reducing prediction lag;
LAGEA: Language Guided Embodied Agents for Robotic Manipulation
Abdul Monaf Chowdhury (University of Dhaka), Rabeya Akter (University of Dhaka)
CodeRobotic IntelligenceReinforcement Learning from Human FeedbackTransformerReinforcement LearningPrompt EngineeringVision Language ModelVision-Language-Action ModelContrastive LearningImageTextMultimodality
π― What it does: Propose the LAGEA framework, which utilizes a vision-language model (VLM) to generate structured error reflection and converts it into a temporalized reward signal to guide reinforcement learning in robotic manipulation tasks.
CodeExplainability and InterpretabilityComputational EfficiencyGraph Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringGraphTime SeriesBenchmark
π― What it does: Propose the LagLLM framework, which utilizes LLMs to generate lead-lag graphs through prompting and perform structured token sorting, enabling the model to explicitly capture spatial-temporal dependencies to improve the accuracy of time series forecasting.
π― What it does: Construct a controllable and safe 3D mesh generation method by performing linear affine mixing in the aligned SDF decoder weight space from a small number of parameterized samples.
Laplacian Representations for Decision-Time Planning
Dikshant Shehmar (University of Alberta), Marlos C. Machado (University of Alberta)
CodeReinforcement LearningContrastive LearningWorld ModelOptical FlowGraphTabularTime Series
π― What it does: This paper proposes a decision-time planning algorithm called ALPS that utilizes Laplacian representations, achieving efficient planning and control in offline goal-conditional reinforcement learning tasks.
π― What it does: This paper proposes the LARA framework, which jointly trains the Latent Action Model and a diffusion-based Vision-Language-Action model, achieving complementarity through representation alignment.
LARFT: Closing the Cognition-Action Gap for Length Instruction Following in Large Language Models
Wei Zhang (Beijing University of Posts and Telecommunications), Sen Su (Beijing University of Posts and Telecommunications)
CodeOptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextBenchmark
π― What it does: Propose a length-aware reinforcement fine-tuning framework called LARFT, which enables large language models to internally recognize length while meeting length instructions and precisely output.
Large Language Models Explore by Latent Distilling
Yuanhao Zeng (ShanghaiTech University), Kan Ren (ShanghaiTech University)
CodeComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelTextBenchmark
π― What it does: Propose an online lightweight potential distiller that combines exploratory sampling (ESamp) to encourage semantic diversity during LLM decoding
Latent Thoughts Tuning: Bridging Context and Reasoning with Fused Information in Latent Tokens
Weihao Liu (University of Illinois Chicago), Lu Cheng (University of Illinois Chicago)
CodeComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextChain-of-Thought
π― What it does: Proposes the Latent Thoughts Tuning (LT-Tuning) framework, enabling large language models to perform stable and dynamic reasoning in a continuous latent space.
LATMiX: Learnable Affine Transformations for Microscaling Quantization of LLMs
Ofir Gordon (Arm), Hai Victor Habi (Arm)
CodeComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelText
π― What it does: Propose a learnable affine transformation (LATMiX) for micro-scale quantization in LLMs, reducing activation outliers and improving inference accuracy at low bit precision.
Learn from A Rationalist: Distilling Intermediate Interpretable Rationales
Jiayi Dai (University of Alberta), Randy Goebel (University of Alberta)
CodeExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningImageText
π― What it does: This paper proposes the REKD (Rationale Extraction with Knowledge Distillation) framework, which utilizes the interpretability and verifiable intermediate 'rationale' from the teacher model to guide the student model's feature selection and prediction, thus addressing the 'chicken and egg' dilemma faced by lightweight models during rationale extraction.