ICML 2026 Papers — Page 25
International Conference on Machine Learning · 6554 papers
GHOST: Unmasking Phantom States in Mamba2 via Grouped Hidden-state Output-aware Selection & Truncation
Michael Menezes (Rice University), Anastasios Kyrillidis (Rice University)
Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelScore-based ModelContrastive LearningText
🎯 What it does: A structured pruning method called GHOST is proposed to systematically identify and truncate the state dynamics by targeting the internal state space dimension of the Mamba2 semantic model.
GI-GCN: Global Interacted Graph Convolutional Networks via Dominant Sets for Graph Classification
Lu Bai (Beijing Normal University), Xin Jin (Central University of Finance and Economics)
ClassificationGraph Neural NetworkGraph
🎯 What it does: Proposed a Global Interaction Graph Convolutional Network (GI-GCN) based on Dominant Set for graph classification.
GICDM: Mitigating Hubness for Reliable Distance-Based Generative Model Evaluation
Nicolas Salvy (Inria), Bertrand Thirion (Inria)
GenerationData SynthesisRepresentation LearningDiffusion modelScore-based ModelContrastive LearningImageMultimodality
🎯 What it does: Proposed a method called GICDM to alleviate hubness in high-dimensional embedding spaces and improve the reliability of generative model evaluation.
GIFT: Bootstrapping Image-to-CAD Program Synthesis via Geometric Feedback
Giorgio Giannone (AI Innovation, Red Hat), Faez Ahmed (DeCoDE Lab, MIT)
GenerationData SynthesisOptimizationComputational EfficiencyTransformerSupervised Fine-TuningPrompt EngineeringVision Language ModelDiffusion modelImageTextMultimodalityMesh
🎯 What it does: Propose a data augmentation framework for image-to-CAD program synthesis called GIFT, which utilizes geometric feedback.
GIPO: Gaussian Importance Sampling Policy Optimization
Chengxuan Lu (IROOTECH TECHNOLOGY), Yang Liu (IROOTECH TECHNOLOGY)
Reinforcement LearningTabularTime SeriesSequential
🎯 What it does: Propose GIPO (Gaussian Importance Sampling Policy Optimization), a smooth proportional attenuation strategy that can replace the hard clipping of PPO when using outdated experience replay.
GIST: Targeted Data Selection for Instruction Tuning via Coupled Optimization Geometry
Guanghui Min (University of Virginia), Chen Chen (University of Virginia)
OptimizationComputational EfficiencyData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningText
🎯 What it does: Propose the GIST method, which performs SVD on the target gradient to extract a low-dimensional subspace, and selects data for targeted instruction tuning through projection alignment scoring, thereby more accurately capturing parameter coupling in LoRA and other parameter-efficient fine-tuning (PEFT) methods while significantly reducing computational and storage costs.
Giving Sensors a Voice: Multimodal JEPA for Semantic Time-Series Embeddings
Utsav Dutta (C3 AI), Henrik Ohlsson (C3 AI)
Anomaly DetectionRepresentation LearningTransformerPrompt EngineeringAuto EncoderContrastive LearningTextMultimodalityTime Series
🎯 What it does: Developed the CHARM model, integrating channel-level textual descriptions into time series embeddings, trained using the JEPA self-supervised framework;
GKD-Recruiter: Jointly Modeling Social and Task Heterogeneity for Spatial Crowdsourcing via Graph Knowledge Distillation
Yucen Gao (Northeastern University), Xiaofeng Gao (Shanghai Jiao Tong University)
Recommendation SystemKnowledge DistillationGraph Neural NetworkReinforcement LearningGraphTabular
🎯 What it does: Proposes the GKD-Recruiter framework, which jointly models social networks and task heterogeneity, and selects worker-task seed pairs in spatial crowdsourcing scenarios through graph knowledge distillation and Rainbow DQN to maximize effective task satisfaction (ETS).
GLAD: Bidirectional Structure-Attribute Alignment via Latent Graph Diffusion Models
Jiankai Zuo (Suzhou University of Science and Technology), Yaying Zhang (Tongji University)
RestorationRepresentation LearningGraph Neural NetworkDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningGraph
🎯 What it does: Propose GLAD, a bidirectional structure-attribute alignment framework based on latent diffusion models, for completing missing node attributes in graphs.
GLARE: Scalable Neuro-Symbolic Reward Shaping for LLM Agents via Group-Level Automata
Jingyuan Yan (University of Science and Technology of China), Jiahu Qin (University of Science and Technology of China)
TransformerLarge Language ModelReinforcement LearningPrompt EngineeringWorld ModelText
🎯 What it does: Propose the GLARE framework, which extracts trajectory events through neural networks, converts them into LTL formulas, compiles them into deterministic automata, and provides dense, reliable, and consistent reward signals for LLM agents in long-horizon tasks.
GLEAN: Guideline-Grounded Evidence Accumulation for High-Stakes Agent Verification
Yichi Zhang (Tsinghua University), Mihaela van der Schaar (University Of Cambridge)
Safty and PrivacyExplainability and InterpretabilityTransformerLarge Language ModelTextBiomedical DataElectronic Health RecordsRetrieval-Augmented Generation
🎯 What it does: Propose a high-risk agent verification framework called GLEAN based on professional guidelines, which progressively aligns guidelines to evaluate agent execution trajectories and accumulates the probability of correctness calibration;
Glimpse: Geometry Learning of Multi-scale Structural Priors for 3D Pose Estimation
Zhenhua Tang (University of Macau), Chi-Man Pun (University of Macau)
Pose EstimationGraph Neural NetworkTransformerDiffusion modelContrastive LearningGaussian SplattingImageVideo
🎯 What it does: Proposes the Glimpse framework, which learns multi-scale structural priors from single-frame images through structured sampling and geometric correction, achieving monocular 3D human pose estimation.
Global Convergence of Adaptive Sensing for Principal Eigenvector Estimation
Alex Saad-Falcon (Georgia Institute of Technology), Justin Romberg (Georgia Institute of Technology)
OptimizationTabularPhysics Related
🎯 What it does: This paper proposes a compressed Oja algorithm that uses only two adaptive linear measurements per sample to estimate the principal eigenvector of the covariance matrix in high-dimensional space, and provides its global convergence analysis;
Global Credit Assignment via Dynamical Criticality
Wentao Wang (Peking University), Guozhang Chen (Peking University)
ClassificationOptimizationComputational EfficiencyReinforcement Learning from Human FeedbackRecurrent Neural NetworkSpiking Neural NetworkAuto EncoderGenerative Adversarial NetworkContrastive LearningImageTextTabularTime SeriesSequentialBenchmark
🎯 What it does: Proposes COLA (Criticality-driven Online Local Alignment) — an online RNN training method that utilizes long-range spatiotemporal correlations in critical states to approximate BPTT gradients using local learning rules.
Global Directional Priors with Local Statistical Validation for Scalable Causal Discovery
Wei Yuan (Chinese Academy of Sciences), Shuhui Wang (Chinese Academy of Sciences)
OptimizationExplainability and InterpretabilityComputational EfficiencyData-Centric LearningContrastive LearningGraphTabularBiomedical DataBenchmark
🎯 What it does: Propose the OCMB framework, which first estimates directionality using global ranking, and then performs local Markov blanket tests on a restricted candidate set to control the dimensionality of the conditioning set in CI tests.
Global Geometry Is Not Enough for Vision Representations
Jiwan Chung (Yonsei University), Seon Joo Kim (Yonsei University)
Explainability and InterpretabilityRepresentation LearningAuto EncoderContrastive LearningImage
🎯 What it does: Investigated the relationship between global embedding geometry and the compositional binding ability in visual representations, finding that global geometry metrics cannot predict compositional binding performance, and proposed and validated Jacobian Effective Rank (JER) as a metric for functional sensitivity.
Global Merger-Arbitrage Forecasting with Language Models
Hinal Jajal (Balyasny Asset Management), Peter Anderson (Balyasny Asset Management)
TransformerLarge Language ModelSupervised Fine-TuningTextTabularFinance RelatedRetrieval-Augmented Generation
🎯 What it does: Establish an LLM prediction system for merger arbitrage, combining expert-driven retrieval with post-hoc guided fine-tuning to achieve probability prediction and report generation for the outcomes of publicly announced merger transactions.
Global Plane Waves from Local Gaussians: Periodic Charge Densities in a Blink
Jonas Elsborg (Technical University of Denmark), Arghya Bhowmik (Technical University of Denmark)
Computational EfficiencyGraph Neural NetworkTransformerAuto EncoderContrastive LearningGaussian SplattingGraphTabularBenchmarkPhysics Related
🎯 What it does: Developed the ELECTRAFI model for rapidly and differentiably predicting charge density in periodic crystals, serving as an initial guess for Kohn-Sham DFT.
Global Policy-Space Response Oracles for Two-Player Zero-Sum Games
Junyu Zhang (Tsinghua University), Xudong Zhang (Tsinghua University)
Graph Neural NetworkSpiking Neural NetworkTransformerReinforcement LearningTabularSequential
🎯 What it does: This paper proposes an extension of PSRO based on global population exploitability (Population Exploitability, PE), called Global PSRO. It directly evaluates the impact of candidate strategies on global approximate quality through an exploration-selection two-phase framework, thereby constructing a smaller and more accurate strategy family.
GO-PRE: Goal-Oriented Next-Best-View Selection via Predictive Rendering Entropy for Active 3D Reconstruction
Yan Song (Fudan University), Wenqiang Zhang (Fudan University)
Autonomous DrivingOptimizationComputational EfficiencyRepresentation LearningConvolutional Neural NetworkTransformerNeural Radiance FieldGaussian SplattingImagePoint CloudMeshBenchmark
🎯 What it does: Propose GO-PRE, a goal-oriented next-best-view selection framework based on predictive rendering entropy, for active 3D reconstruction.
Goal-Conditioned Agents that Learn Everything All at Once
Michael Matthews (University of Oxford), Jakob Nicolaus Foerster
Recurrent Neural NetworkTransformerReinforcement LearningAuto EncoderContrastive LearningTabularTime SeriesSequential
🎯 What it does: Propose an efficient full-target learning method called LEO, which can update all discrete targets at once; further propose Dual LEO, which enhances the performance of the UVFA network by leveraging the teacher network of LEO.
Goal-Oriented Lower-Tail Calibration of Gaussian Processes for Bayesian Optimization
Aurélien Pion (Transvalor S.A.), Emmanuel Vazquez (University Paris Saclay)
OptimizationScore-based ModelGaussian SplattingTabularTime Series
🎯 What it does: This paper studies the calibration of Gaussian process predictive distributions under low thresholds in Bayesian optimization, proposing a post-calibration method called TCGP and applying it within the EI sampling criterion.
GOCM: Single-Step Graph Outlier Synthesis via Origin Consistency Model
Yifan Li (Shandong University of Science and Technology), Peng Zhang (Shandong University of Science and Technology)
Data SynthesisAnomaly DetectionGraph Neural NetworkDiffusion modelScore-based ModelAuto EncoderContrastive LearningGraph
🎯 What it does: Propose a single-step graph anomaly sample generation framework called GOCM, aimed at alleviating the sample imbalance problem in graph anomaly detection.
Golden Goose: A Simple Trick to Synthesize Unlimited RLVR Tasks from Unverifiable Internet Text
Ximing Lu (NVIDIA), Yejin Choi (NVIDIA)
Data SynthesisReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose a simple pipeline (olden oose), which transforms unverifiable internet text into verifiable RLVR tasks by masking key reasoning steps and generating multiple-choice answers, and constructs a 0.7M-scale multiple-choice question-answering dataset.
Good SFT Optimizes for SFT, Better SFT Prepares for Reinforcement Learning
Dylan Zhang (University of Illinois), Hao Peng (University of Illinois)
OptimizationTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningContrastive LearningText
🎯 What it does: Introduce a reweighting scheme called PEAR based on importance sampling during the later stages of offline supervised fine-tuning (SFT) to correct the distribution mismatch between SFT and subsequent reinforcement learning (RL)
GoodDiffusion: Proactive Copyright Protection for Diffusion Bridge Models via Learnable Sample-specific Signatures
Shixi Qin (University of Chinese Academy of Sciences), Qingming Huang (University of Chinese Academy of Sciences)
GenerationSafty and PrivacyAdversarial AttackDiffusion modelScore-based ModelContrastive LearningImage
🎯 What it does: Designed an active copyright protection method called GoodDiffusion based on learnable sample-specific signatures to prevent unauthorized use of diffusion models.
GOTabPFN: From Feature Ordering to Compact Tokenization for Tabular Foundation Models on High-Dimensional Data
Al Zadid Sultan Bin Habib (West Virginia University), Donald Adjeroh
ClassificationExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationData-Centric LearningTransformerContrastive LearningTabularBiomedical DataBenchmark
🎯 What it does: In high-dimensional low-sample table prediction, the TabPFN model is compressed to be suitable for HDLSS data through graph-guided feature ranking and neuro-inspired sub-unit compression.
GP2F: Cross-Domain Graph Prompting with Adaptive Fusion of Pre-trained Graph Neural Networks
Dongxiao He (Tianjin University), Di Jin (Tianjin University)
ClassificationDomain AdaptationRepresentation LearningMeta LearningGraph Neural NetworkTransformerSupervised Fine-TuningPrompt EngineeringContrastive LearningGraph
🎯 What it does: Proposes the GP2F dual-branch graph prompting learning framework, which integrates frozen pre-trained GNNs with a lightweight adapter branch for cross-domain few-shot graph tasks.
gp2Scale: A Class of Compactly Supported Non-Stationary Kernels and Distributed Computing for Exact Gaussian Processes on 10 Million Data Points
Marcus M. Noack (Lawrence Berkeley National Laboratory), Ronald J. Pandolfi (Lawrence Berkeley National Laboratory)
OptimizationComputational EfficiencyData-Centric LearningMixture of ExpertsContrastive LearningGaussian SplattingImagePoint CloudTabularTime SeriesBenchmark
🎯 What it does: Propose the gp2Scale method, constructing a scalable non-stationary compactly supported kernel, combined with distributed sparse covariance computation and block MCMC, to achieve precise Gaussian process regression on data with over ten million points.
GPan-LoRA: Gaussian Process Amortized Networks for Bayesian Low-Rank Adaptation in Large Language Models
Weifeng Zhang (Texas A&M University), Xiaoning Qian (Texas A&M University)
ClassificationExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelContrastive LearningGaussian SplattingText
🎯 What it does: This study proposes an scalable Gaussian process (GP) Bayesian LoRA framework called GPan-LoRA, which utilizes sparse GP approximation and autoregressive variational inference to achieve Bayesian low-rank fine-tuning of large language models and reliable uncertainty quantification.
GR-LoRA: Gradient-Recycling Low-Rank Adaptation for Class-Incremental Learning
Yipeng Lin (Nanjing University of Science and Technology), Yang Yang (Nanjing University of Science and Technology)
ClassificationComputational EfficiencyRepresentation LearningMeta LearningTransformerSupervised Fine-TuningPrompt EngineeringContrastive LearningImageBenchmark
🎯 What it does: Designed a gradient recovery low-rank adaptation method called GR-LoRA to address the stability-plasticity trade-off in class-incremental learning.
Gradient Descent as a Perceptron Algorithm: Understanding Dynamics and Implicit Acceleration
Alexander Tyurin (Applied AI Institute)
ClassificationRecognitionImage TranslationOptimizationConvolutional Neural NetworkImage
🎯 What it does: Study the optimization dynamics of gradient descent in logistic regression and two-layer small convolutional networks, proving its equivalence to the perceptron algorithm and explaining the phenomenon of implicit acceleration.
Gradient Descent with Large Step Size Restores Symmetry in Deep Linear Networks with Multi-Pathway
Hee-Sung Kim (Hanyang University), Sungyoon Lee (Hanyang University)
OptimizationStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: This paper studies the training dynamics of multi-channel deep linear networks under large-step gradient descent, finding that contrary to the 'winner-takes-all' symmetric breaking predicted by continuous-time gradient flow, the network recovers symmetry in the margin-stable region and distributes features across multiple channels.
Gradient Flow Dynamics and Implicit Bias of Diagonal Linear Networks under Infinitesimal Initialization
Jiajie Zhao (Shanghai Jiao Tong University), Yaoyu Zhang (Shanghai Jiao Tong University)
OptimizationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningContrastive LearningTabularStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Studied the gradient flow dynamics of diagonal linear networks under infinitesimal initialization, extended the theorem from Pesme & Flammarion (2023), analyzed the training trajectories of deep diagonal linear networks and a broader class of two-layer diagonal linear networks, and proved that the training trajectories of these models can be equivalently represented by the proposed algorithm.
Gradient Flow Sampler-based Distributionally Robust Optimization
Zusen Xu (KTH Royal Institute of Technology), Jia-Jie Zhu (KTH Royal Institute of Technology)
Domain AdaptationOptimizationAdversarial AttackDiffusion modelScore-based ModelContrastive LearningImagePoint CloudTabularStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Propose and implement a distributionally robust optimization framework (GF-DRO) based on gradient flow samplers, by equivalently transforming the inner maximization problem of DRO into an entropy-regularized JKO operation, directly sampling from the worst distribution to solve DRO under non-convex losses.
Gradient Flow Through Diagram Expansions: Learning Regimes and Explicit Solutions
Dmitry Yarotsky (Applied AI Institute), Yaroslav Gusev (Applied AI Institute)
OptimizationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningContrastive LearningGraphTabularPhysics RelatedStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: A general mathematical framework is proposed for analyzing gradient flows (GF) in large-scale learning problems and deriving explicit analytical solutions.
Gradient Inversion Attacks Beyond SGD
Guangnian Wan (National University of Singapore), Xinchao Wang (National University of Singapore)
OptimizationFederated LearningSafty and PrivacyAdversarial AttackConvolutional Neural NetworkDiffusion modelContrastive LearningImage
🎯 What it does: The study investigates how to perform gradient inversion attacks to recover private client images and labels in federated learning scenarios using adaptive optimizers (such as Adam, RMSProp, AdaGrad), relying only on model updates rather than original gradients.
Gradient Preconditioning for Efficient and Reliable Reward-Guided Generation
Jisung Hwang (KAIST), Minhyuk Sung (KAIST)
GenerationOptimizationComputational EfficiencyTransformerReinforcement LearningDiffusion modelAuto EncoderContrastive LearningImageText
🎯 What it does: This paper proposes a gradient preprocessing method that achieves efficient and reliable reward-guided generation for first-order generative models by projecting the reward gradient into a feasible white noise domain.
Gradient Regularization Mitigates Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
Johannes Ackermann (University of Tokyo), Masashi Sugiyama (RIKEN AIP)
Reinforcement Learning from Human FeedbackSupervised Fine-TuningReinforcement LearningText
🎯 What it does: This paper proposes using gradient regularization (GR) to alleviate the reward model attack problem in RLHF and RLVR, and improves the model's accuracy in rewards through explicit and implicit implementation methods.
Gradient Smoothing: Coupling Layer-wise Updates for Improved Optimization
Haoming Meng (University of Toronto), Vardan Papyan (University of Toronto)
OptimizationConvolutional Neural NetworkRecurrent Neural NetworkTransformerImageText
🎯 What it does: A new optimization method called Gradient Smoothing is proposed, which improves the optimization process of deep neural networks by coupling hierarchical updates across depth.
Gradient Testing and Estimation by Comparisons
Xiwen Tao (Peking University), Tongyang Li (Peking University)
OptimizationReinforcement LearningContrastive Learning
🎯 What it does: The study proposes testing and estimation algorithms for gradient direction using a comparison oracle that can only provide pairwise comparison results. It achieves gradient testing with O(1) comparison queries and gradient estimation with O(n log 1/ε) comparison queries, and provides proofs of their optimality; in the quantum model, it achieves quantum gradient estimation with O(log(n/ε)) comparison queries.
Gradient Transformer: Learning to Generate Updates for LLMs
Binh-Nguyen Nguyen (New Jersey Institute of Technology), Issa Khalil (Qatar Computing Research Institute)
Federated LearningSafty and PrivacyComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelSupervised Fine-TuningText
🎯 What it does: Propose a data-agnostic weak-to-strong knowledge distillation framework called Gradient Transformer (GRAD-TRANSFORMER), which achieves LLM fine-tuning without exposing private data by converting the update vectors of the fine-tuned TinyLM into update vectors of the LLM.
Gradient-Aware Scheduling: Coupling Curriculum and Staleness for Async Reinforcement Learning
Xinyu Zhang (Anyscale)
OptimizationComputational EfficiencyAI Code AssistantTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsContrastive LearningTextSequential
🎯 What it does: Propose the Gradient‑Aware Scheduling (GAS) framework, combining adaptive curriculum learning with asynchronous scheduling, to address the contradiction between gradient staleness and task difficulty in reinforcement learning.
Gradient-Based Causal Tree Ensembles: A Backbone Architecture for Heterogeneous Treatment Effects
Yusuke Kano (Canon Inc.), Mihaela van der Schaar (University of Cambridge)
Federated LearningExplainability and InterpretabilityComputational EfficiencyMixture of ExpertsContrastive LearningTabularTime SeriesSequentialBenchmark
🎯 What it does: A framework called GRACE based on gradient-differentiable tree ensembles for heterogeneous treatment effect estimation was studied, which can be trained end-to-end and replace fully connected layers.
Gradient-Free Approaches is a Key to an Efficient Interaction with Markovian Stochasticity
Boris Prokhorov (EPFL), Aleksandr Beznosikov (Basic Research of Artificial Intelligence Laboratory)
OptimizationStochastic Differential Equation
🎯 What it does: This paper proposes a zeroth-order (gradient-free) optimization method tailored for Markov noise environments, which can achieve accelerated convergence in both strongly convex smooth and non-smooth problems;
Gradients with Respect to Semantics Preserving Embeddings Tell the Uncertainty of Large Language Model
Mingda Li (Harbin Institute of Technology), Ting Liu (Harbin Institute of Technology)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelContrastive LearningText
🎯 What it does: Proposes two gradient-based, sampling-free uncertainty quantification methods, SemGrad and HybridGrad, for free-text generation in large language models.
GradientStabilizer: Fix the Norm, Not the Gradient
Tianjin Huang (University of Exeter), Shiwei Liu (ELLIS Institute Tubingen)
OptimizationRepresentation LearningTransformerLarge Language ModelReinforcement LearningDiffusion modelContrastive LearningImageTextTime SeriesSequential
🎯 What it does: Propose a lightweight gradient transformation method called GradientStabilizer, which keeps the gradient direction unchanged and uses the moving average of gradient norms to stabilize the update magnitude.
GradMem: Learning to Write Context into Memory with Test-Time Gradient Descent
Yuri Kuratov (AXXX), Mikhail Burtsev (London Institute for Mathematical Sciences)
CompressionRepresentation LearningMeta LearningTransformerLarge Language ModelTextRetrieval-Augmented Generation
🎯 What it does: Propose GradMem, a mechanism that writes context into a limited memory vector through a few gradient descent steps during inference, achieving compressed memory for large language models.
GradPower: Powering Gradients for Faster Language Model Pre-Training
Jinbo Wang (Peking University), Lei Wu (Peking University)
OptimizationTransformerLarge Language ModelText
🎯 What it does: Proposes the GradPower gradient transformation technique for pre-training of LLMs
Gram2Token: Enabling Run-time GPU-Native Grammar-Constrained Decoding for LLMs
Hantao Hua (National University of Defense Technology), Feng Zhu (National University of Defense Technology)
Computational EfficiencyAI Code AssistantTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation
🎯 What it does: Proposed a GPU-native Grammar-Constrained Decoding framework called GRAM2TOKEN, which preprocesses byte-level grammar into token-level state transition tables, thereby avoiding byte-level traversal during inference.
Granularity-Aware Adaptive Classifier Expansion via Zero-Shot Learning
Xiangyu Wang (University of Science and Technology of China), Huanhuan Chen (University of Science and Technology of China)
ClassificationRepresentation LearningTransformerLarge Language ModelMixture of ExpertsVision Language ModelContrastive LearningImageTextBenchmark
🎯 What it does: Propose a granularity-aware adaptive framework that extends zero-shot classifiers in an unsupervised manner by leveraging multi-source semantic generation and structural discovery.
GRAPE: Let GRPO Supervise Query Rewriting by Ranking for Retrieval
Zhaohua Zhang (Dalian University of Technology), Zhixun Su (Dalian University of Technology)
RetrievalDomain AdaptationTransformerLarge Language ModelReinforcement LearningPrompt EngineeringContrastive LearningImageTextMultimodality
🎯 What it does: Rewrite user queries using large language models (LLMs), and supervise the rewriting strategy with ranking information from a frozen CLIP retriever to improve the accuracy of cross-modal retrieval.
Graph Alignment for Benchmarking Graph Neural Networks and Learning Positional Encodings
Adrien Lagesse (INRIA, École Normale Supérieure - PSL), Marc Lelarge (INRIA, École Normale Supérieure - PSL)
Data SynthesisRepresentation LearningGraph Neural NetworkTransformerAuto EncoderContrastive LearningGraphBenchmark
🎯 What it does: Based on the graph alignment task, a new GNN evaluation benchmark is proposed, which can generate multi-class alignment datasets ranging from easy to difficult through adjustable noise levels, and uses this task as a self-supervised pre-training objective to generate high-quality node position encodings (GAPE).
Graph Alignment via Dual-Pass Spectral Encoding and Latent Space Communication
Maysam Behmanesh (Ecole Polytechnique), Maks Ovsjanikov (Ecole Polytechnique)
Representation LearningGraph Neural NetworkAuto EncoderContrastive LearningMultimodalityGraph
🎯 What it does: Under unsupervised conditions, we propose a graph alignment framework called GADL, which achieves node correspondence through dual-channel spectral encoding and functional mapping.
Graph is a Natural Regularization: Revisiting Vector Quantization for Graph Representation Learning
Zian Zhai (University of New South Wales), Wenjie Zhang (University of New South Wales)
OptimizationComputational EfficiencyRepresentation LearningGraph Neural NetworkAuto EncoderContrastive LearningGraph
🎯 What it does: Systematic experiments on vector quantization (VQ) methods for graph-structured data reveal the codebook collapse problem and propose a Regularized Graph Vector Quantization (RGVQ) solution to improve codebook utilization and downstream task performance.
Graph is a Substrate Across Data Modalities
Ziming Li (University of Connecticut), Chuxu Zhang (University of Connecticut)
Computational EfficiencyRepresentation LearningData-Centric LearningMeta LearningDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsVision Language ModelContrastive LearningImageTextMultimodalityGraphChain-of-Thought
🎯 What it does: Propose the G-Substrate framework, treating graph structures as reusable intermediate structures across modalities and tasks, and enhancing performance through a unified structural pattern and alternating role training.
Graph Neural Dynamics via Learned Energy and Tangential Flows
Moshe Eliasof (University of Cambridge), Carola-Bibiane Schönlieb (University of Cambridge)
OptimizationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerAuto EncoderContrastive LearningGraphTabularTime SeriesSequentialBiomedical DataBenchmarkStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Proposes the TANGO framework, which views the feature evolution of graph neural networks as a combination of gradient descent and tangent flows (maintaining energy) under a learned energy landscape, achieving more flexible and stable graph representation learning.
Graph Neural Networks Are Not Continuous Across Graph Resolutions
Christian Koke (AITHYRA), Daniel Cremers (Munich Center for Machine Learning)
Representation LearningDrug DiscoveryGraph Neural NetworkContrastive LearningGraphTabular
🎯 What it does: This paper demonstrates that graph neural networks (GNNs) do not exhibit continuity across different graph resolutions, leading to potentially very different representations for similar graphs. In particular, for graphs representing the same underlying object but at different resolutions, GNNs generate distinctly different latent embeddings.
Graph of States: Solving Abductive Tasks with Large Language Models
Yu Luo (Nankai University), Dan Pei (Tsinghua University)
Autonomous DrivingOptimizationFederated LearningExplainability and InterpretabilityDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackNeural Architecture SearchProtein Structure PredictionGraph Neural NetworkTransformerLarge Language ModelAgentic AIPrompt EngineeringMixture of ExpertsTextGraphTabularBiomedical DataBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose Graph of States (GoS), a dual-layer neuro-symbolic framework for addressing inductive reasoning tasks in LLMs.
Graph Rewiring based on Flow Alignment for Improving Fluid Simulation
Zenong Li (Nanyang Technological University), Adams Wai-Kin Kong (Nanyang Technological University)
Graph Neural NetworkOptical FlowMeshGraphPhysics Related
🎯 What it does: Proposed a local graph reconnection method called FLARE based on flow direction alignment, aimed at improving fluid simulations using graph neural networks.
Graph-GRPO: Training Graph Flow Models with Reinforcement Learning
Baoheng Zhu (Beijing University of Posts and Telecommunications), Xiao Wang (Beihang University)
GenerationOptimizationDrug DiscoveryGraph Neural NetworkReinforcement LearningScore-based ModelFlow-based ModelGraphBiomedical Data
🎯 What it does: Propose an online reinforcement learning framework called Graph-GRPO, which aligns the discrete flow matching graph generative model (Graph Flow Model) with task-specific rewards, overcoming the issues of non-differentiability and low exploration efficiency in traditional GFMs during sampling.
Graph-Link: Bridging the Semantic-Structural Gap in Text-to-SQL via Constrained Subgraph Induction
Jianwei Zhong (Huazhong University of Science and Technology), Yijun Mo (Huazhong University of Science and Technology)
OptimizationComputational EfficiencyData-Centric LearningGraph Neural NetworkTransformerLarge Language ModelTextGraphTabularBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose the Graph-Link framework, treating Schema Linking as a subgraph induction task under topological constraints, addressing structural blind spots in multi-hop JOIN operations;
Graph-Preference Learning: Debiasing Network-Sampled Human Feedback for Target Welfare Estimation
Guangrui Fan (Taiyuan University of Science and Technology), Pan Lihu
OptimizationReinforcement Learning from Human FeedbackGraph Neural NetworkTransformerReinforcement LearningPrompt EngineeringDiffusion modelScore-based ModelContrastive LearningTextGraphRetrieval-Augmented Generation
🎯 What it does: This paper proposes the Graph-Preference Learning framework, addressing the bias caused by network sampling. It designs a graph personalized reward model and a graph balanced aggregation in two steps, aiming to let the reward model reflect the target welfare rather than the observed inclusion distribution.
Graph-R1: Towards Agentic GraphRAG Framework via End-to-end Reinforcement Learning
Haoran Luo (Beijing University of Posts and Telecommunications), Anh Tuan Luu (Nanyang Technological University)
RetrievalKnowledge DistillationRepresentation LearningGraph Neural NetworkTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringTextGraphBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposes Graph-R1, an agent-based GraphRAG framework based on end-to-end reinforcement learning, capable of achieving closed-loop reasoning from lightweight knowledge hypergraph construction to multi-round retrieval interaction and generation.
GraphFLEx: Unsupervised Structure Learning $\underline{\text{F}}$ramework for $\underline{\text{L}}$arge $\underline{\text{Ex}}$panding $\underline{\text{Graph}}$s
Mohit Kataria (Indian Institute of Technology Delhi), Sandeep Kumar (Indian Institute of Technology Delhi)
Federated LearningComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerMixture of ExpertsDiffusion modelScore-based ModelAuto EncoderContrastive LearningGaussian SplattingGraphTabularTime SeriesBiomedical DataBenchmark
🎯 What it does: Proposes the GraphFLEx framework, integrating graph clustering, graph coarsening, and unsupervised structural learning, supporting incremental structural learning for large-scale and dynamically scalable graphs.
GraphFlow: A Graph-Based Workflow Management for Efficient LLM-Agent Serving
Ao Li (Xi'an Jiaotong University), su zhou
Computational EfficiencyAI Code AssistantGraph Neural NetworkTransformerLarge Language ModelAgentic AIPrompt EngineeringTextGraphBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposed the GraphFlow framework, which achieves dynamic workflow generation and efficient execution of LLM agents by constructing a unified global operation graph (wGraph);
GraphP-FL: Personalized Federated Graph Learning via Dynamic Structure Awareness and Fisher Information Elastic Alignment
Haoyu Chen (Tianjin University Of Technology), Jianhao Li (Tianjin University Of Technology)
OptimizationFederated LearningRepresentation LearningGraph Neural NetworkAuto EncoderContrastive LearningGraph
🎯 What it does: This paper proposes the GraphP-FL framework for achieving personalized federated learning on distributed graph data.
GraphPFN: A Prior-Data Fitted Graph Foundation Model
Dmitry Eremeev (HSE University), Liudmila Prokhorenkova (Yandex Research)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningData-Centric LearningMeta LearningGraph Neural NetworkTransformerSupervised Fine-TuningPrompt EngineeringMixture of ExpertsContrastive LearningGraph
🎯 What it does: Propose GraphPFN, a graph foundation model based on Prior-Data Fitted Networks, which can achieve in-context learning and fine-tuning in node-level tasks;
GRASP: Awakening Latent Spatial Reasoning in LVLMs via Training-free Geometric Rectification
Jiadong Yan (Soochow University), Xizhao Luo (Soochow University)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerPrompt EngineeringVision Language ModelRectified FlowContrastive LearningImageVideoTextMultimodalityBenchmark
🎯 What it does: Propose a training-free inference framework GRASP for geometric correction during inference, leveraging geometric manifold search and dual-layer correction to activate the spatial reasoning capabilities of LVLMs.
GRASP: Graph Reasoning via Agentic Solving and Probing of LLMs
Xiaojun Guo (Peking University), Yisen Wang (Peking University)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AITextGraphRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposed the GRASP framework, which utilizes LLMs to explore graph structures through active agents and implements graph reasoning via two tools: neighbor retrieval and code interpreter.
Great Minds Think Alike: Contextual Tacit Communication for Decentralized LLM-Agent Cooperation
Yue Pei (Beihang University), Ziliang Chen (Peng Cheng Laboratory)
Autonomous DrivingOptimizationFederated LearningComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextMultimodalityBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose the Contextual Tacit Communication (Tact) protocol, enabling large language model (LLM) agents to implicitly collaborate without online message exchange by mining mismatches through residual banding during the offline phase and storing corrected behavioral biases in a retrievable Tacit Rule Memory (TRM).
Greedy Coordinate Diffusion: Effective and Semantically Coherent Adversarial Attacks via Diffusion Guidance
Bohdan Turbal (Princeton University), Aleksandra Korolova (Princeton University)
Adversarial AttackTransformerPrompt EngineeringDiffusion modelText
🎯 What it does: Propose Greedy Coordinate Diffusion (GCD), an adversarial prompt generation framework that utilizes discrete diffusion models as the proposal distribution and performs greedy coordinate optimization in a gray-box environment.
Grokking Finite-Dimensional Algebra
Pascal Junior Tikeng Notsawo, Guillaume Rabusseau (Université de Montréal)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerTabular
🎯 What it does: Study the grokking phenomenon that occurs when neural networks learn finite-dimensional algebraic multiplication, and build a tensor framework to analyze its relationship with algebraic structures.
Gromov-Wasserstein at Scale, Beyond Squared Norms
Guillaume Houry (Inria, Université Paris Cités, Inserm), François-Xavier Vialard (LIGM, Université Gustave Eiffel)
OptimizationComputational EfficiencyRepresentation LearningContrastive LearningPoint CloudMesh
🎯 What it does: This paper proposes an scalable Gromov-Wasserstein (GW) matching solver, specifically implementing linear memory and quadratic time complexity, and being differentiable, for conditional negative type (CNT) cost.
Grounded in Reality: Learning and Deploying Proactive LLM from Offline Logs
Fei Wei (Alibaba Group), Bolin Ding (Alibaba Group)
Data-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextSequentialElectronic Health RecordsRetrieval-Augmented Generation
🎯 What it does: Learning an active questioning strategy without a simulator, Learn-to-Ask, which generates dense and realistic rewards by 'looking back' using offline expert dialogue logs, training LLMs to decide what questions to ask and when to stop.
Grounding Functional Similarity by Invariance-Aware Model Stitching
Ioannis Athanasiadis (Linkoping University), Michael Felsberg (Linkoping University)
ClassificationExplainability and InterpretabilityRepresentation LearningConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningImage
🎯 What it does: Evaluate the functional similarity of deep networks through model stitching, and propose an unbiased model stitching method based on forward-backward compatibility.
Grounding LLMs in Scientific Discovery via Embodied Actions
Bo Zhang (Tsinghua University), Hongning Wang (Tsinghua University)
Autonomous DrivingOptimizationDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringVision-Language-Action ModelSimultaneous Localization and MappingWorld ModelOptical FlowTextMultimodalityTabularTime SeriesBenchmarkPhysics RelatedRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: By building the EmbodiedAct framework in the MATLAB environment, integrating LLM with simulation software, achieving real-time perception and intervention of LLM on physical simulation, promoting verifiable automation in scientific discovery and engineering design.
Grounding Multi-Hop Reasoning in Structural Causal Models via Group Relative Policy Optimization
Yunhan Bu (Beijing Institute of Technology), Shuai Lei (Academy of Military Science)
Explainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This paper addresses the multi-hop fact verification problem by proposing a structured reasoning framework based on Structural Causal Models (SCM). It dynamically regulates the length and depth of the reasoning chain through Group Relative Policy Optimization (GRPO) reinforcement learning, thereby achieving traceable and more accurate reasoning processes.
Group Cognition Learning: Making Everything Better Through Controlled Two-Stage Agents Collaboration
Chunlei Meng (Fudan University), Zhongxue Gan (Fudan University)
ClassificationExplainability and InterpretabilityRepresentation LearningTransformerAgentic AIMixture of ExpertsContrastive LearningVideoTextMultimodalityAudio
🎯 What it does: This paper proposes the Group Cognition Learning (GCL) framework, which explicitly governs multi-modal interactions through a two-stage collaborative agent (routing, auditing, public factor, aggregation), addressing the issues of modality dominance and pseudo coupling.
Group Distributionally Robust Optimization-Driven RL for LLM Reasoning
Kishan Panaganti (Tencent Frontier Lab), Dong Yu (Capital One)
OptimizationData-Centric LearningTransformerLarge Language ModelReinforcement LearningPrompt EngineeringText
🎯 What it does: This paper proposes two reinforcement learning post-training methods based on group distributionally robust optimization (GDRO), named Prompt-GDRO and Rollout-GDRO, which improve the performance of large language models on reasoning tasks by dynamically grouping difficulty levels to guide data sampling and computational resource allocation.
Group-wise Data Ordering: Enhancing Instruction Tuning of Large Language Models via Embedding Proximity
Yiwen Ye (Bytedance), Yong Xia (Ningbo No. 2 Hospital)
Computational EfficiencyRepresentation LearningData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningTextMultimodalityBenchmark
🎯 What it does: Designed and implemented an EP-Order method based on embedding similarity for group-level data sorting, aimed at improving the instruction fine-tuning effectiveness of large language models (LLMs).
Grouter: Decoupling Routing from Representation for Accelerated MoE Training
Yuqi Xu (Peking University), Kun Yuan (Peking University)
Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerMixture of ExpertsText
🎯 What it does: Propose a hybrid expert (MoE) training framework with pre-learned and fixed routing structure (GROUTER), decoupling routing from expert weight updates to accelerate and improve model convergence quality.
GRPO is Secretly a Process Reward Model
Michael Sullivan (Saarland University), Alexander Koller (Saarland University)
OptimizationComputational EfficiencyReinforcement Learning from Human FeedbackLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkChain-of-Thought
🎯 What it does: Demonstrate that GRPO本质上 is inherently an implicit process reward model, and propose the λ-GRPO improved algorithm based on this insight
GRPO-based Cluster Decision Agent for Unknown-$\boldsymbol{K}$ Multi-view Clustering
Xuqian Xue (Fudan University), Junping Zhang (Fudan University)
OptimizationRepresentation LearningTransformerReinforcement LearningAgentic AIAuto EncoderContrastive LearningImageMultimodality
🎯 What it does: Proposes the GRO K framework, which uses a cluster decision agent based on GRPO to autonomously estimate the unknown number of clusters K in multi-view clustering, forming a perception-decision-feedback closed loop.
GSFixer: Improving 3D Gaussian Splatting with Reference-Guided Video Diffusion Priors
Xingyilang Yin (University of Macau), Xiaodong Cun (GVC Lab, Great Bay University)
RestorationGenerationOptimizationTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningGaussian SplattingImageVideoPoint CloudMeshBenchmark
🎯 What it does: To address the artifact problem in 3D Gaussian Splatting (3DGS) under sparse viewpoints, the GSFixer framework is proposed. It repairs artifacts on new viewpoint images using a video diffusion model and backpropagates the results to 3DGS, achieving high-quality reconstruction and view synthesis under sparse viewpoints.
GSRQ: Gain-Shape Residual Quantization for Sub-1-bit KV Cache
Soosung Kim (Yonsei University), Jaeyong Chung (Yonsei University)
CompressionComputational EfficiencyRepresentation LearningTransformerLarge Language ModelAuto EncoderContrastive LearningText
🎯 What it does: A new KV cache compression method called GSRQ is proposed, which improves the directional fidelity and reconstruction accuracy of vector quantization by using Gain-Shape K-means (GSKM) at each stage and incorporating gradient weighting.
GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy
Hongze Tan (ByteDance China), Haihua Yang (ByteDance China)
OptimizationComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkChain-of-Thought
🎯 What it does: Propose a dynamic entropy-weighted reward shaping framework, and design two algorithms at the token level (GTPO) and sequence level (GRPO-S) based on GRPO, to achieve fine-grained reward allocation during the LLM generation process, thereby improving mathematical reasoning performance.
Guaranteed Optimal Compositional Explanations for Neurons
Biagio La Rosa (University of California Santa Cruz), Leilani H. Gilpin (University of California Santa Cruz)
OptimizationExplainability and InterpretabilityConvolutional Neural NetworkContrastive LearningImage
🎯 What it does: A complete framework is proposed for computing compositional explanations of neurons across the entire search space, along with an optimal algorithm that can complete the task within an acceptable time;
GUDA: Counterfactual Group-wise Training Data Attribution for Diffusion Models via Unlearning
Naoki Murata (Sony AI), Yuki Mitsufuji (Sony AI)
GenerationExplainability and InterpretabilityComputational EfficiencyData-Centric LearningReinforcement LearningDiffusion modelScore-based ModelContrastive LearningImageText
🎯 What it does: Proposed the GUDA framework, which uses machine unlearning to approximately perform group deletion (LOGO) on generative models and quantifies the impact of each training data group on the generation results;
GUI-Spotlight: Adaptive Iterative Focus Refinement for Enhanced GUI Visual Grounding
Bin Lei (University of Minnesota), Caiwen Ding (University of Minnesota)
RecognitionTransformerSupervised Fine-TuningReinforcement LearningVision-Language-Action ModelGaussian SplattingImageMultimodality
🎯 What it does: Propose GUI-Spotlight, a graphical user interface visual localization model that utilizes multiple tools and iterative reinforcement learning for focused attention;
Guidance: Sentence-Level Citation Enforcement via Prefix-Tail Guidance during LLM Decoding
Yirui Zhan (Peking University), Jun Gao (Peking University)
RetrievalComputational EfficiencyData-Centric LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposes a framework called Guidance, which offers training freedom and enforces sentence-level citations during the LLM decoding phase.
Guided Star-Shaped Masked Diffusion
Viacheslav Meshchaninov (Constructor University), Dmitry Vetrov (Constructor University)
GenerationData SynthesisTransformerDiffusion modelScore-based ModelText
🎯 What it does: Proposed a novel sampling algorithm called G-Star for pre-trained discrete diffusion models, which achieves efficient error correction by introducing a star sampling framework and a learnable error detector, significantly improving sample quality and sampling speed.
GuidedBridge: Training-freely Improving Bridge Models with Prior Guidance
Zehua Chen (Tsinghua University), Jun Zhu (Tsinghua University)
Image TranslationRestorationTransformerDiffusion modelScore-based ModelContrastive LearningImageStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: This paper proposes a training-agnostic bridge model guidance method called Prior Guidance (PG) and its frequency modulation version FMPG, and designs a CFG-FMPG framework that cascades with CFG, aiming to improve the generation quality and inference efficiency of bridge models in image translation and restoration tasks.
GXPO: Group Cross-Lingual Relative Policy Optimization for Code Generation
Linzheng Chai (Beihang University), Xianglong Liu (Beihang University)
AI Code AssistantTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextSequentialBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposed and implemented a multilingual reinforcement learning framework called GXPO for code generation. GXPO generates solutions in multiple programming languages for the same problem, forming a 'multilingual group,' and calculates three types of advantages (monolingual advantage, cross-lingual advantage, group advantage) for each solution, thereby achieving cross-lingual alignment, balanced optimization, and knowledge transfer during training.
H$^2$CL: Heterogeneity-Aware Hypergraph Contrastive Learning for Robust Representation Learning
Kaixuan Yao (Shanxi University), Ming Li (Zhejiang Normal University)
Representation LearningGraph Neural NetworkContrastive LearningGraph
🎯 What it does: This paper proposes a contrastive learning framework called HCL for heterogeneous hypergraphs, which utilizes node-hyperedge heterogeneity to guide view generation and encoding, achieving robust representation learning.
h1: Bootstrapping LLMs to Reason over Longer Horizons via Reinforcement Learning
Alesia Ivanova (University of Oxford), Charles London (University of Oxford)
Data-Centric LearningTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextChain-of-Thought
🎯 What it does: Proposed a method to train LLM's long-term reasoning ability by utilizing existing short-term reasoning data, generating multi-step chain questions through serial concatenation, and adaptively improving model performance through stage-wise reinforcement learning (RL).
Hair-Trigger Alignment: Black-Box Evaluation Cannot Guarantee Post-Update Alignment
Yavuz Faruk Bakman, Sai Praneeth Karimireddy (University of Southern California)
Safty and PrivacyExplainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringText
🎯 What it does: Studied whether black-box evaluation can guarantee model alignment after LLM updates, proving that even if the model passes black-box evaluation for alignment, it cannot be guaranteed to remain aligned after updates, and further demonstrated that a single gradient update can activate hidden misalignment behaviors.
Hallucination Detection from Structural Reasoning Model
Jianbo Sun (Tsinghua University), Pengkun Yang (Tsinghua University)
Anomaly DetectionExplainability and InterpretabilityGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringTextGraphRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose a structured reasoning model (SRM) that detects hallucinations in large language models by generating condition-step dependency directed acyclic graphs and performing local verification.
Hallucination is a Consequence of Space-Optimality: A Rate-Distortion Theorem for Membership Testing
Anxin Guo (Northwestern University), Jingwei Li (Columbia University)
OptimizationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningTextRetrieval-Augmented Generation
🎯 What it does: This paper formalizes the random fact memory problem of large language models as membership testing, and proposes an information-theoretic rate-distortion theorem under sparse fact scenarios, explaining why models inevitably produce high-confidence hallucinations under limited capacity.
HALO: A Unified Vision-Language-Action Model for Embodied Multimodal Chain-of-Thought Reasoning
Quanxin Shou (Hong Kong University of Science and Technology), Song Guo (Hong Kong University of Science and Technology)
Representation LearningRobotic IntelligenceReinforcement Learning from Human FeedbackTransformerLarge Language ModelMixture of ExpertsVision Language ModelVision-Language-Action ModelAuto EncoderImageVideoTextMultimodalityChain-of-Thought
🎯 What it does: Proposes HALO — a unified vision-language-action model that realizes the 'embodied mental chain-of-thought' (EM-CoT) for text reasoning, visual foresight, and action prediction through a Mixture-of-Experts Transformer, and designs an automated EM-CoT data synthesis pipeline and staged training strategy.