ICML 2026 Papers — Page 23
International Conference on Machine Learning · 6554 papers
From Flat Facts to Sharp Hallucinations: Detecting Stubborn Errors via Gradient Sensitivity
Liew Yee Zhing (Xi'an Jiaotong-Liverpool University), Anwar P.P. Abdul Majeed (Sunway University)
Anomaly DetectionExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringContrastive LearningText
🎯 What it does: Propose a method that detects high-confidence errors (Stubborn Hallucinations) in large language models by utilizing gradient sensitivity (EPGS) after input embedding perturbation, distinguishing robust facts from fragile memories by probing the local curvature of the loss surface.
From Generalist to Specialist Representation
Yujia Zheng (Carnegie Mellon University), Kun Zhang (Carnegie Mellon University)
Representation LearningAuto EncoderGenerative Adversarial NetworkContrastive LearningImageVideoText
🎯 What it does: This paper proves, under a fully nonparametric setup, how to learn task-specific representations from a general model, and provides theoretical guarantees for identifiability. It first proves that cross-time task structures can be identified in an unsupervised manner, and then separates task-related latent variables at each step through sparse regularization.
From Generative to Episodic: Sample-Efficient Replicable Reinforcement Learning
Max Hopkins (Institute of Advanced Study), Yuichi Yoshida (National Institute of Informatics)
Reinforcement Learning
🎯 What it does: A replicable reinforcement learning algorithm is designed for discrete MDPs with finite time steps, achieving near-optimal sample efficiency without using a generative model.
From geometry to dynamics: Learning overdamped Langevin dynamics from sparse observations with geometric constraints
Dimitra Maoutsa (Technical University of Berlin)
OptimizationFederated LearningComputational EfficiencyRepresentation LearningDrug DiscoveryReinforcement Learning from Human FeedbackScore-based ModelTabularTime SeriesSequentialPhysics RelatedStochastic Differential Equation
🎯 What it does: This paper proposes a geometry-guided path augmentation framework for learning the drift term of overdamped Langevin dynamics from sparse observations.
From Growing to Looping: A Unified View of Iterative Computation in LLMs
Ferdinand Kapl (Technical University of Munich), Stefan Bauer (Technical University of Munich)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Investigate the impact of two techniques, Looping and Depth-Grown, on the reasoning performance of large language models (LLMs), propose a unified mechanism perspective, and demonstrate that the two techniques can be complementary and combined.
From Guessing to Placeholding: A Cost-Theoretic Framework for Uncertainty-Aware Code Completion
Liang Zhu (Tencent), Xian Wu (Tencent)
AI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringText
🎯 What it does: Propose the Adaptive Placeholder Completion (APC) framework, allowing large language models to insert placeholders at positions of high uncertainty, thereby reducing the editing cost for developers later on.
From Historical Patches to Repair Plans: Outcome-Conditioned Reasoning for Repository-Level Program Repair
Chenglin Li (Concordia University), Tse-Hsun Chen
AI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose a framework named ConRAD, which guides new warehouse-level program repair tasks through conditionalized reasoning plans that reverse reconstruct historical repair instances.
From Holo Pockets to Electron Density: GPT-style Drug Design with Density
Jiahao Chen (Renmin University of China), Bo Huang (NeoPrimeTech Biology)
Drug DiscoveryProtein Structure PredictionTransformerLarge Language ModelSupervised Fine-TuningDiffusion modelScore-based ModelAuto EncoderPoint CloudBiomedical Data
🎯 What it does: Proposed EDMolGPT, a decoder autoregressive model based on low-resolution electron density point clouds, for generating 3D drug molecules that conform to the protein binding environment.
From Human Labels to Literature: Semi-Supervised Learning of NMR Chemical Shifts at Scale
Yongqi Jin (Peking University), Weinan E (Peking University)
Data-Centric LearningDrug DiscoveryGraph Neural NetworkTransformerSupervised Fine-TuningContrastive LearningTextGraphMagnetic Resonance Imaging
🎯 What it does: By using semi-supervised learning, combine a large number of unannotated NMR spectra from literature with a small amount of annotated data to build a high-precision chemical shift prediction model.
From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation
Yajie Li (Fudan University), Li Zhang (Fudan University)
GenerationData SynthesisRobotic IntelligenceTransformerMixture of ExpertsVision-Language-Action ModelDiffusion modelFlow-based ModelImageVideoMultimodality
🎯 What it does: Propose the MoLA framework, which converts future images generated from videos into executable latent actions through a pre-trained multi-modal inverse dynamics model, thereby achieving action-based reasoning about imagined futures;
From Individual Calibration to Reliable Classifiers: ALD Parameterization with mPAIC Guarantees
Deming Sheng (Duke University), Ricardo Henao (Duke University)
ClassificationExplainability and InterpretabilityComputational EfficiencyData-Centric LearningScore-based ModelContrastive LearningImageTextTabular
🎯 What it does: Propose a multi-classifier based on the asymmetric Laplace distribution (HALD), and achieve reliable probabilistic prediction through individualized calibration (HICALD);
From Inpainting to Editing: Unlocking Robust Mask-Free Visual Dubbing via Generative Bootstrapping
Xu He (Tsinghua University), Zhiyong Wu (Tsinghua University)
GenerationData SynthesisTransformerDiffusion modelScore-based ModelFlow-based ModelVideoMultimodalityBenchmarkAudio
🎯 What it does: Propose a two-stage generative guidance framework called X-Dub, which uses mask-based inpainting to generate pseudo-paired data, and then trains an unmasked editing model to complete visual dubbing.
From Interaction Trajectories to Prompt Rules: Credit Assignment for Multi-Agent Prompt Optimization
Bin Wu (University College London), Qiang Zhang (Zhejiang University)
OptimizationExplainability and InterpretabilityAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmark
🎯 What it does: Proposed a rule-based credit estimation framework called TRUCE, which is based on interaction trajectories, for optimizing natural language prompts in multi-agent LLM systems.
From Interactions to Principles: Experience-Driven Self-Distillation for Evolving LLM Agents
Rong Wu (Zhejiang University), Botian Shi (Shanghai Artificial Intelligence Laboratory)
Knowledge DistillationTransformerLarge Language ModelReinforcement LearningAgentic AITextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose the EvolveR framework, constructing a complete closed-loop experience lifecycle, including offline self-distillation to generate abstract principles, online interaction to retrieve experiences, and iterative optimization of strategies through reinforcement learning.
From Internal Diagnosis to External Auditing: A VLM-Driven Paradigm for Data-Free Online Backdoor Defense
Binyan Xu (Chinese University of Hong Kong), Kehuan Zhang (Chinese University of Hong Kong)
Anomaly DetectionTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityRetrieval-Augmented Generation
🎯 What it does: Propose PRISM, an online backdoor defense framework that does not require access to training data or modification of model weights, utilizing external Vision-Language Models to perform semantic auditing on model predictions.
From Kepler to Newton: Inductive Biases Guide Learned World Models in Transformers
Ziming Liu (Stanford University), Andreas S. Tolias (Stanford University)
TransformerSupervised Fine-TuningContrastive LearningWorld ModelTabularTime SeriesPhysics Related
🎯 What it does: The study investigates why generic Transformers fail to learn Newtonian world models from planetary motion data, and addresses this limitation by introducing three minimal priors (spatial smoothness, spatial stability, and temporal locality).
From Knowledge to Inference: Formalizing Specialized Public Health Reasoning on GlobalHealthAtlas
Zhaokun Yan (China Academy of Information and Communications Technology), Tongning Wu (China Academy of Information and Communications Technology)
Data SynthesisTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextTabularBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Constructed and released the GlobalHealthAtlas, a global public health reasoning dataset, and designed a multi-stage data generation and quality control pipeline based on LLMs; simultaneously developed a six-dimensional domain-specific evaluator and performed specialized fine-tuning on qwen3-8b to form Public-Model.
From LLM-Generated Conjectures to Lean Formalizations: Automated Polynomial Inequality Proving via Sum-of-Squares Certificates
Ruobing Zuo (East China Normal University), Jianlin Wang (Henan University)
OptimizationExplainability and InterpretabilityComputational EfficiencyAI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringAuto EncoderTextTabularBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Automatically proving multivariate polynomial inequalities using a neuro-symbolic framework: the LLM first generates an approximate SOS (Sum-of-Squares) structure, and then symbolic computation is used to obtain the exact Gram matrix and generate machine-checkable proofs in Lean.
From Lyapunov Analysis to Algorithm Design in two-sided PL Minimax Optimization
Mansi Rankawat (Université de Montréal), Damien Scieur (Samsung AI Lab)
Optimization
🎯 What it does: Proposed a single-loop algorithm called TALDA based on the Lyapunov function for solving smooth non-convex non-concave min-max problems that satisfy the bidirectional Polyak-Lojasiewicz (S2PL) condition, and provided a linear convergence analysis.
From Memorization to Parameter Interference: How Overtraining Experts Harms Model Merging
Stefan Horoi (Université de Montréal), Gintare Karolina Dziugaite (Google DeepMind)
ClassificationOptimizationKnowledge DistillationConvolutional Neural NetworkTransformerSupervised Fine-TuningPrompt EngineeringMixture of ExpertsContrastive LearningImageText
🎯 What it does: Investigate the negative impact of over-training expert models on model merging and propose an early stopping strategy to improve merging performance
From Moments to Models: Graphon-Mixture Learning for Mixup and Contrastive Learning
Ali Azizpour (Rice University), Santiago Segarra (Rice University)
ClassificationRepresentation LearningGraph Neural NetworkMixture of ExpertsDiffusion modelContrastive LearningGraphBenchmark
🎯 What it does: Proposes a framework for estimating graphon mixtures based on graph matrix (motif density), and applies the inferred multiple generative models to graph mixture (GMAM) and model-aware graph contrastive learning (MGCL) to improve the performance of downstream tasks.
From Muon to Gluon: Bridging Theory and Practice of LMO-based Optimizers for LLMs
Artem Riabinin (King Abdullah University of Science and Technology), Peter Richtárik (King Abdullah University of Science and Technology)
OptimizationLarge Language ModelText
🎯 What it does: Proposed a hierarchical optimization framework called Gluon based on the linear minimization orbit (LMO), and introduced a hierarchical (L, L₀¹)-smoothness model to characterize the geometric structure of deep networks, providing a more realistic convergence analysis;
From Noise to Control: Parameterized Diffusion Policies
Renhao Zhang (University of Massachusetts), Bruno Castro da Silva (University of Massachusetts)
OptimizationRobotic IntelligenceTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningMultimodalityPoint Cloud
🎯 What it does: Propose Parameterized Diffusion Policy (PDP), which parameterizes traditional diffusion policies by learning a geometry-aligned behavioral latent space, enabling precise control of behavior on low-dimensional latent variables;
From Noise to Intent: Anchoring Generative VLA Policies with Residual Bridges
Yiming Zhong (ShanghaiTech University), Yuexin Ma (ShanghaiTech University)
Robotic IntelligenceTransformerReinforcement LearningVision Language ModelVision-Language-Action ModelDiffusion modelFlow-based ModelRectified FlowMultimodality
🎯 What it does: Propose the ResVLA framework, which adopts a generative VLA strategy combining low-frequency intent anchoring with high-frequency residual diffusion bridge, achieving a fundamental shift from 'Generation-from-Noise' to 'Refinement-from-Intent'.
From Observations to States: Latent Time Series Forecasting
Jie Yang (University of Illinois Chicago), Philip S. Yu (University of Illinois Chicago)
Information TheoryAnomaly DetectionComputational EfficiencyRepresentation LearningTransformerAuto EncoderContrastive LearningTabularTime Series
🎯 What it does: Propose LatentTSF, which transforms time series forecasting from direct observation regression to first mapping via AutoEncoder to a latent space, then performing state prediction in that space, and finally decoding the predicted latent states back to observations.
From Optimization to Generalization under Heavy-Tailed Data: The Role of Gradient Clipping
Aleksandr Shestakov (Basic Research of Artificial Intelligence Laboratory), Eduard Gorbunov (Mohamed bin Zayed University of Artificial Intelligence)
Optimization
🎯 What it does: This paper studies the impact of gradient clipping on the optimization and generalization of empirical risk minimization (ERM) algorithms when the data follows a heavy-tailed distribution, and provides corresponding convergence and generalization theories;
From Out-of-Distribution Detection to Hallucination Detection: A Geometric View
Litian Liu (Qualcomm AI Research), Roland Memisevic (Qualcomm AI Research)
Anomaly DetectionExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringScore-based ModelText
🎯 What it does: Reformulate hallucination detection in large language models as an outlier detection problem, and propose a single-sample, zero-training-cost method based on the geometric OOS detector.
From Outcomes to Actions: Leveraging Hindsight for Long-Horizon Language Agent Training
Zishang Jiang (Fudan University), Yanghua Xiao (Fudan University)
OptimizationTransformerLarge Language ModelReinforcement LearningText
🎯 What it does: This paper proposes a hindsight intention space policy gradient method called HPO for training long-horizon language agents.
From Parameter Dynamics to Risk Scoring: Quantifying Sample-Level Safety Degradation in LLM Fine-tuning
Xiao Wang (Northeastern University), Daling Wang (Northeastern University)
Safty and PrivacyExplainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringText
🎯 What it does: This paper reveals the fundamental mechanism behind the decline in safety by tracking the dynamic evolution of model parameters during the fine-tuning process, and proposes a sample-level risk quantification method called SQSD based on parameter projection;
From Parameters to Data: A Task-Parameter-Guided Fine-Tuning Pipeline for Efficient LLM Alignment
Hao Chen (Zhejiang University), Junbo Zhao (Zhejiang University)
Computational EfficiencyKnowledge DistillationData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringText
🎯 What it does: Constructed the P2D framework, which uses task-related attention heads as bidirectional guides to achieve a collaborative pipeline for data mining and sparse parameter fine-tuning;
From Parameters to Feature Space: Task Arithmetic for Backdoor Mitigation in Model Merging
Zhenqian Zhu (Harbin Institute of Technology), Wenjian Luo (Shenzhen University)
OptimizationFederated LearningSafty and PrivacyExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningAdversarial AttackTransformerSupervised Fine-TuningDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageMultimodality
🎯 What it does: Proposes a backdoor defense framework called LFPM based on feature space task arithmetic, aimed at eliminating backdoors while maintaining clean task performance during model merging.
From Per-Image Low-Rank to Encoding Mismatch: Rethinking Feature Distillation in Vision Transformers
Huiyuan Tian (Zhejiang University), Shijian Li (Zhejiang University)
Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerAuto EncoderContrastive LearningImage
🎯 What it does: This paper explores the fundamental reasons behind the failure of feature distillation in Vision Transformers during the compression process, and provides feasible improvement solutions for this issue;
From Perception to Planning: Evolving Ego-Centric Task-Oriented Spatiotemporal Reasoning via Curriculum Learning
Xiaoda Yang (Zhejiang University), Zhou Zhao (Zhejiang University)
Autonomous DrivingComputational EfficiencyRepresentation LearningRobotic IntelligenceTransformerSupervised Fine-TuningVision Language ModelContrastive LearningImageVideoTextMultimodalityChain-of-Thought
🎯 What it does: Design and implement the EgoTSR framework, which adopts a three-stage curriculum learning (CoT → Tag → LongTag) to progressively achieve egocentric task-based spatiotemporal reasoning, moving from explicit spatial awareness, internalized judgment, to long-term planning.
From Pixels to Tokens: A Systematic Study of Latent Action Supervision for Vision-Language-Action Models
Yihan Lin (Renmin University of China), Jing Zhang (Renmin University of China)
Robotic IntelligenceReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningVision Language ModelVision-Language-Action ModelContrastive LearningImageVideoTextMultimodality
🎯 What it does: Studied the effect of using latent action supervision in vision-language-action (VLA) models, and proposed and implemented four different supervision strategies.
From Player to Master: Enhancing Test-Time Learning of LLM Agents via Reinforcement Learning over Memory
Yishuo Cai (Peking University), Xu Sun (Peking University)
TransformerLarge Language ModelReinforcement LearningAgentic AITextRetrieval-Augmented Generation
🎯 What it does: Propose a trainable memory update module called MEMOPILOT, enabling frozen LLM players to achieve test-time learning during multi-round interactions.
From Poisoned to Aware: Fostering Backdoor Self-Awareness in LLMs
Guangyu Shen (Purdue University), Xiangyu Zhang (Purdue University)
Safty and PrivacyExplainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringText
🎯 What it does: This paper proposes a post-training framework based on reinforcement learning, aiming to enable poisoned LLMs to self-identify and articulate their hidden backdoor triggers, and achieve more effective backdoor elimination and detection based on this.
From Prior to Pro: Efficient Skill Mastery via Distribution Contractive RL Finetuning
Zhanyi Sun (Stanford University), Shuran Song (Stanford University)
Robotic IntelligenceSupervised Fine-TuningReinforcement LearningDiffusion modelFlow-based ModelVideoTabularSequentialBenchmark
🎯 What it does: Propose Distribution Contractive Reinforcement Learning (DICE-RL), which improves the success rate of long-horizon tasks with sparse rewards by performing distribution contraction fine-tuning on a pre-trained generative behavioral cloning (BC) policy.
From Prompts to Responses: Dual-Sided Data Leakage and Defense in Split Large Language Models
Zixuan GU (Tsinghua University), Yang Liu (Hong Kong Polytechnic University)
Safty and PrivacyTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextFinance Related
🎯 What it does: Proposed the bilateral leakage attack PIDI and the defense framework ADMI targeting both ends of leakage, revealing that Split-LLM is vulnerable to reverse engineering in both input prompts and generated responses, posing privacy risks.
From Prompts to Tokens: Internalizing Causal Supervision in Vision-Language Model for Multi-Image Causal Reasoning
Haoping Yu (Case Western University), Jing Ma (Case Western University)
Explainability and InterpretabilityRepresentation LearningTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityGraphBenchmark
🎯 What it does: Propose the BridgeVLM framework, which learns and internalizes causal graphs (DAGs) from multi-graph inputs, generates Causal Tokens, and utilizes RAMP for directed information propagation, enabling vision-language models to perform internal supervision for multi-graph causal reasoning (intervention and counterfactuals).
From Rashomon Theory to PRAXIS: Efficient Decision Tree Rashomon Sets
Zakk Heile (Duke University), Cynthia Rudin (Duke University)
ClassificationOptimizationExplainability and InterpretabilityComputational EfficiencyTabularTime SeriesSequential
🎯 What it does: Propose the PRAXIS algorithm, which efficiently approximates the Rashomon set of sparse decision trees using proxy subroutines and budget recursion.
From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning
Lingjing Kong (Carnegie Mellon University), Zhengzhong Liu (Institute of Foundation Models)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextChain-of-Thought
🎯 What it does: In large language model inference tasks, the authors explain through theoretical analysis and controlled experiments the mechanism of the joint pipeline of supervised fine-tuning (SFT) and reinforcement learning (RL), demonstrating that they play complementary roles in identifying and reusing composable atomic modules (skills and routing mechanisms).
From Representation to Action: A Unified Laplacian Framework for Spatial Representation and Path Planning
Junfeng Zuo (Peking University), Si Wu (Peking University)
OptimizationRobotic IntelligenceConvolutional Neural NetworkReinforcement LearningContrastive LearningImagePoint CloudGraph
🎯 What it does: Integrate the Laplace operator to construct a unified framework, derive spatial encoding as Laplace features, and design a navigation strategy based on Green's functions.
From Retrieval to Translation: Translating Query into Graph-level Clues for Retrieval-Augmented Generation
Qichuan Liu (Xiamen University), Zhihong Zhang (Xiamen University)
GenerationRetrievalRepresentation LearningGraph Neural NetworkTransformerSupervised Fine-TuningPrompt EngineeringTextGraphRetrieval-Augmented Generation
🎯 What it does: Propose KG-Translator, which transforms the retrieval task into a graph-level clue translation. First, a lightweight NER + syntactic parsing is used to build ParseKG. Then, a constrained decoding LLM translates the query into graph clues. Finally, the clues are tracked back to the original paragraph for retrieval and answering.
From Reward-Free Representations to Preferences: Rethinking Offline Preference-Based Reinforcement Learning
Jun-Jie Yang (National Yang Ming Chiao Tung University), Ping-Chun Hsieh (National Yang Ming Chiao Tung University)
Representation LearningRobotic IntelligenceReinforcement Learning from Human FeedbackSupervised Fine-TuningReinforcement LearningContrastive LearningPoint CloudTabularTime SeriesSequentialBenchmark
🎯 What it does: This paper proposes a new offline preference reinforcement learning framework, FB-PbRL, which combines reward-agnostic representation learning with preference learning. It first pretrains the FB representation on reward-free data, and then uses contrastive learning to search for task vectors and finetune the representation on preference data, thus directly driving policy learning using preferences instead of a reward model.
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models
Juncheng Wu (Amazon), Yuyin Zhou (UC Santa Cruz)
Data SynthesisRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningVision Language ModelImageTextMultimodalityChain-of-Thought
🎯 What it does: The paper proposes a phased post-training framework to enhance the perception and reasoning capabilities of vision-language models by decomposing visual perception, textual reasoning, and visual reasoning into three separate stages.
From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning
Jike Zhong (University of Southern California), Shao-Yuan Lo (National Taiwan University)
Explainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningTextMultimodalityRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This paper studies how to enhance the Theory of Mind (ToM) capability of large language models through post-training. For the first time, it systematically audits and removes shortcuts from the dataset, and then introduces Thinking-RFT (Thinking-based Reinforcement Fine-Tuning) to improve the model's reasoning and generalization performance.
From Similarity to Vulnerability: Key Collision Attack on LLM Semantic Caching
Zhixiang Zhang (Hong Kong University of Science and Technology), Dongdong She (Hong Kong University of Science and Technology)
Safty and PrivacyAdversarial AttackTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation
🎯 What it does: The paper proposes the CacheAttack framework targeting semantic caching in large language models (LLMs), enabling key collision attacks in multi-tenant environments to hijack LLM responses or proxy tool calls.
From Static Constraints to Dynamic Adaptation: Sample-Level Constraint Relaxation for Offline-to-Online Reinforcement Learning
Lipeng Zu (Florida State University), Xiaonan Zhang (Florida State University)
TransformerReinforcement LearningDiffusion modelTabularTime SeriesBenchmark
🎯 What it does: Propose the DARE framework, which relaxes constraints at the sample level through behavioral consistency in offline-to-online reinforcement learning to achieve more robust online fine-tuning.
From Statics to Dynamics: Physics-Aware Image Editing with Latent Transition Priors
Liangbing Zhao (King Abdullah University of Science and Technology), Mohamed Elhoseiny (King Abdullah University of Science and Technology)
GenerationExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringVision Language ModelDiffusion modelAuto EncoderContrastive LearningImageVideoTextMultimodalityBenchmarkPhysics Related
🎯 What it does: Proposed a framework called PhysicEdit that treats image editing as a physical state transition, and constructed a video-based PhysicTran38K dataset
From Talking to Singing: A New Challenge for Audio-Visual Deepfake Detection
Ke Liu (University of Electronic Science and Technology of China), Yang Yang (University of Electronic Science and Technology of China)
Anomaly DetectionConvolutional Neural NetworkTransformerVision Language ModelDiffusion modelAuto EncoderContrastive LearningVideoTextMultimodalityAudio
🎯 What it does: Constructed the first deepfake dataset for singing scenarios, SHDF, and proposed an unsupervised framework called T-AVFD, which achieves cross-scenario (speaking and singing) deepfake detection using text-guided facial authenticity patterns and multimodal differential weighted learning.
From Teacher Pathways to Invariant Manifolds: Consensus Subspace Distillation for TSFMs
Zexing Zhang (National University of Defense Technology), Yang Kewei (National University of Defense Technology)
CompressionAnomaly DetectionKnowledge DistillationRepresentation LearningTransformerLarge Language ModelMixture of ExpertsDiffusion modelScore-based ModelAuto EncoderContrastive LearningTime SeriesSequential
🎯 What it does: Propose a knowledge distillation method based on consensus subspace, which utilizes the spontaneous convergence of high-level embeddings from Transformers of different scales in a low-rank subspace, guiding the student model to learn without relying on the specific path of the teacher.
From Text to Forecasts: Bridging Modality Gap with Temporal Evolution Semantic Space
Lehui Li (Shandong University), Yongshun Gong (Shandong University)
Representation LearningTransformerLarge Language ModelPrompt EngineeringTextMultimodalityTime SeriesFinance Related
🎯 What it does: Proposes the TESS framework, which extracts quantifiable temporal evolution semantic space (four primitives: distribution shift, shape, fluctuation, lag) from text information through large language models and injects it as a condition into numerical time series predictors.
From Token to Token Pair: Efficient Prompt Compression for Large Language Models in Clinical Prediction
Mingcheng Zhu (University of Oxford), Tingting Zhu (University of Oxford)
CompressionComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextBiomedical DataElectronic Health Records
🎯 What it does: Propose MedTPE, a tokenization method that achieves lossless compression by merging subword pairs of medical terminology, aimed at reducing the input length for LLMs in clinical prediction.
From Volume to Value: Preference-Aligned Memory Construction for On-Device RAG
Changmin Lee (Ulsan National Institute of Science and Technology), Taesik Gong (Ulsan National Institute of Science and Technology)
RetrievalComputational EfficiencyData-Centric LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextRetrieval-Augmented Generation
🎯 What it does: Propose EPIC, a framework for building preference-aligned retrieval memory on devices, which can store only content relevant to user preferences under extremely low memory budgets and support real-time retrieval.
From Weak Cues to Real Identities: Evaluating Inference-Driven De-Anonymization in LLM Agents
Myeongseob Ko (Virginia Tech), Ruoxi Jia (Virginia Tech)
Safty and PrivacyTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Investigate the inference-driven identity reconstruction risks in de-anonymization using large language model agents, propose and evaluate the InferLink benchmark, and review historical cases and real-time human-computer interaction logs.
From Welfare to Utility: Generalized Objectives in Budget-Feasible Procurement
Alon Eden (Hebrew University), Thodoris Tsilivis (Boston University)
OptimizationFederated LearningReinforcement Learning from Human FeedbackPrompt EngineeringReview/Survey Paper
🎯 What it does: In the procurement problem with budget constraints, a mechanism that achieves DSIC-IR is designed, providing constant approximations for welfare, utility, and their generalized objectives under both prior-independent and Bayesian settings;
From Winning to Understanding: A Diagnostic Long-Horizon RTS Benchmark for LLMs
Jiacheng Li (University of Chinese Academy of Sciences), Chenghao Li (Tsinghua University)
TransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextTime SeriesSequentialBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Built a long-term, adversarial real-time strategy game benchmark, using LLM as the decision module, providing three-track evaluation (against rule-based AI, LLM adversaries, and human instruction following)
From Zero to Hero: Advancing Zero-Shot Foundation Models for Tabular Outlier Detection
Xueying Ding (Carnegie Mellon University), Leman Akoglu (Carnegie Mellon University)
Anomaly DetectionTransformerMixture of ExpertsGenerative Adversarial NetworkContrastive LearningTabularBenchmark
🎯 What it does: Proposed a novel zero-shot tabular anomaly detection foundation model called OUTFORMER, which can directly perform forward inference on new tasks in an unsupervised environment.
Front-Loaded Robust Conformal Prediction: Heavy Calibration, Minimal Test-Time Cost
Soroush H. Zargarbashi (CISPA Helmholtz Center for Information Security), Aleksandar Bojchevski (University of Cologne)
ClassificationRecognitionImage TranslationAnomaly DetectionComputational EfficiencyConvolutional Neural NetworkTransformerDiffusion modelImage
🎯 What it does: Propose the Front-Loaded Robust Conformal Prediction (FLoR) framework, which performs large-scale Monte Carlo sampling during the calibration phase and uses only a single noisy sampling during the testing phase, thereby significantly reducing the size of the prediction set while maintaining robust coverage guarantees; simultaneously extend to FLoR(k) through majority voting to further reduce the randomness of the prediction set;
Frontier Models Can Take Actions at Low Probabilities
Alex Serrano (ML Alignment & Theory Scholars), Erik Jenner (Google DeepMind)
Explainability and InterpretabilityComputational EfficiencyAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This paper systematically evaluates the ability of state-of-the-art large language models to perform specified actions in multi-task scenarios (coding, business emails, a modified rock-paper-scissors game) with extremely low probability (as low as 0.001%), and measures their probability calibration performance.
FrontierCS: Evolving Challenges for Evolving Intelligence
Qiuyang Mang (University of California Berkeley), Alvin Cheung (University of California Berkeley)
OptimizationTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose the FrontierCS benchmark, which includes 240 open-ended, verifiable, and scoreable computer science problems covering two major categories: algorithms and research. It achieves high-quality, reproducible evaluation through expert review;
FS-I2P: A Hierarchical Focus–Sweep Registration Network with Dynamically Allocated Depth
Zhixin Cheng (University of Science and Technology of China), Tianzhu Zhang (University of Science and Technology of China)
Pose EstimationDepth EstimationConvolutional Neural NetworkTransformerReinforcement LearningScore-based ModelContrastive LearningOptical FlowImageMultimodalityPoint Cloud
🎯 What it does: This paper proposes an image-point cloud registration framework named FS-I2P, which effectively alleviates matching errors caused by scale ambiguity by utilizing a hierarchical Focus-Sweep interaction (first global-scale calibration and then local refinement) and a Mamba-based multi-layer interaction module.
FT-Dojo: Towards Autonomous LLM Fine-Tuning with Language Agents
Qizheng Li (Microsoft Research Asia), Jiang Bian (Microsoft Research Asia)
OptimizationFederated LearningData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningAgentic AIPrompt EngineeringTextTabularSequentialBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose the FT-Dojo interactive benchmark environment and the FT-Agent automated LLM fine-tuning framework
Full-Batch Gradient Descent Outperforms One-Pass SGD: Sample Complexity Separation in Single-Index Learning
Filip Kovačević (Institute of Science and Technology Austria), Marco Mondelli (Institute of Science and Technology Austria)
OptimizationTabular
🎯 What it does: This paper investigates the difference in sample complexity between full-batch gradient descent (GD) and one-shot stochastic gradient descent (SGD) in high-dimensional single-index models (Gaussian single-index models), with a particular focus on quadratic activation functions and truncated quadratic activation functions;
Full-Spectrum Graph Neural Networks: Expressive and Scalable
Xiaohan Wang (Nanyang Technological University), Kelin Xia (Nanyang Technological University)
ClassificationComputational EfficiencyRepresentation LearningGraph Neural NetworkGraphBenchmark
🎯 What it does: This paper proposes a Full-Spectrum Pairwise Graph Neural Network (FSPECGNN), which elevates traditional single-variable spectral filtering to the node-pair domain and employs bivariate spectral filters to achieve second-order graph signal processing.
FullStack-Agent: Enhancing Agentic Full-Stack Web Coding via Development-Oriented Testing and Repository Back-Translation
Zimu Lu (Chinese University of Hong Kong), Hongsheng Li (Chinese University of Hong Kong)
AI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningAgentic AIPrompt EngineeringTextTabularBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposes a unified FullStack-Agent system, which includes a multi-agent development framework, self-improvement methods achieved through warehouse reverse translation and incremental expansion, and a benchmark that comprehensively evaluates frontend, backend, and database functions.
Fully Dynamic Coreset Spectral Clustering
Ben Jourdan (University of Edinburgh), Gregory Schwartzman (JAIST)
OptimizationComputational EfficiencyData-Centric LearningGraph Neural NetworkScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningGraph
🎯 What it does: A fully dynamic data structure was constructed to support the insertion, deletion of edges/nodes, and clustering membership queries on any subset when the graph evolves over time, with the goal of providing an approximate solution to the Normalised Cut problem.
Fully Zero-Shot Image Dehazing
Shuocheng Wang (Fudan University), Yibo Fan (Fudan University)
RestorationTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningImage
🎯 What it does: Proposes a fully zero-shot image dehazing framework that is trained only using clean images, utilizing color, structural, and illumination invariant representations derived from physical models, and achieving dehazing through diffusion models.
FunCQNet: A Functional Censored Quantile Neural Network for Predicting Long-Term Post-Transplant Kidney Survival
Jiaqi Men (Shanghai University of Finance and Economics), Jiguo Cao (Simon Fraser University)
Explainability and InterpretabilityComputational EfficiencyDrug DiscoveryRecurrent Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningTabularTime SeriesBiomedical DataElectronic Health RecordsStochastic Differential Equation
🎯 What it does: Propose the FunCQNet framework, which utilizes deep neural networks and truncated quantile loss to estimate the time-varying coefficients of interactions between functional biomarkers and scalar covariates, thereby predicting long-term survival after kidney transplantation;
Function-Valued Causal Influence in Nonlinear Time Series
Valentina V. Kuskova (University of Notre Dame), Michael Coppedge (University of Notre Dame)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningRecurrent Neural NetworkTime SeriesBenchmark
🎯 What it does: This paper views the causality of nonlinear time series as state-dependent functional objects and proposes a practical estimation framework based on ICE (Individual Conditional Expectation);
Functional Adjoint Sampler: Scalable Sampling on Infinite Dimensional Spaces
Byoungwoo Park (KAIST), Guan-Horng Liu (FAIR at Meta)
OptimizationDrug DiscoveryProtein Structure PredictionDiffusion modelScore-based ModelBiomedical DataBenchmarkStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Propose Functional Adjoint Sampler (FAS), a diffusion sampler that can sample Gibbs distributions in infinite-dimensional Hilbert spaces, capable of directly sampling in the trajectory space and achieving endpoint constraints.
Functional Attention: From Pairwise Affinities to Functional Correspondences
Jiefang Xiao (Technical University of Munich), Daniel Cremers (Technical University of Munich)
OptimizationComputational EfficiencyRepresentation LearningTransformerDiffusion modelScore-based ModelFlow-based ModelRectified FlowAuto EncoderContrastive LearningPoint CloudMeshGraphTabularTime SeriesSequentialPhysics Related
🎯 What it does: A novel Functional Attention framework is proposed, which replaces the traditional point-to-point attention with a linear mapping in the function space, thus achieving efficient and analytically tractable mapping of continuous fields.
Functional building blocks of neural networks: from network motifs to collective dynamics
Jian Zhang (Chinese Academy of Sciences), Tielin Zhang (Chinese Academy of Sciences)
ClassificationRecurrent Neural NetworkReinforcement LearningImageTime SeriesAudio
🎯 What it does: Analyze 13 types of three-node network subgraphs (network motifs) and embed them as adjustable structural regularizers into continuous-time RNNs, investigating their impact on network stability and flexibility.
Functional Cache Grafting: Robust and Rapid Code-Policy Synthesis for Embodied Agents
Saehun Chun (Sungkyunkwan University), Honguk Woo (Sungkyunkwan University)
Robotic IntelligenceAI Code AssistantTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextSequentialRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This study proposes a framework called FCGRAFT, which rapidly and robustly synthesizes code strategies generated by CodeLLM through function-level KV caching, cache stitching, and patching techniques.
Functional Decomposition and Shapley Interactions for Interpreting Survival Models
Sophie Hanna Langbein (Leibniz Institute for Prevention Research and Epidemiology BIPS), Julia Herbinger (Leibniz Institute for Prevention Research and Epidemiology BIPS)
Explainability and InterpretabilityTabularTime SeriesBiomedical Data
🎯 What it does: This paper proposes two methods, SurvFD (Survival Functional Decomposition) and SurvSHAP-IQ, for decomposing and interpreting time-dependent interactions in machine learning survival models.
Functional Equivalence in Attention: A Comprehensive Study with Applications to Linear Mode Connectivity
Viet-Hoang Tran (National University Of Singapore), Tan Minh Nguyen
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerAuto EncoderContrastive LearningImageText
🎯 What it does: This study formally investigates functional equivalence in Transformer models with positional encoding, particularly focusing on two widely used encoding methods: sine encoding and rotational position encoding (RoPE).
FunPhase: A Periodic Functional Autoencoder for Motion Generation via Phase Manifolds
Marco Pegoraro (Institute of Science and Technology Austria), Arianna Rampini (Autodesk Research)
GenerationData SynthesisPose EstimationTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningVideoSequential
🎯 What it does: Propose FunPhase, a functional periodic autoencoder that can map human or character motion into continuous spatiotemporal functions and learn the latent space of phase structures, enabling motion reconstruction, prediction, and generation without skeletons and at arbitrary temporal resolutions; meanwhile, train latent diffusion models on this phase space to achieve conditional motion generation.
Furina: Fragmented Uncertainty-Driven Refusal Instability Attack
Tongxi Wu (Nanjing University), Yang Gao (Nanjing University)
Safty and PrivacyAdversarial AttackTransformerLarge Language ModelPrompt EngineeringDiffusion modelScore-based ModelGenerative Adversarial NetworkImageTextMultimodalityBenchmarkRetrieval-Augmented Generation
🎯 What it does: This study investigates the non-binary instability of safety alignment and proposes Furina, which induces uncertainty in the model's refusal decisions by fragmenting and scenario-anchored questioning, thereby breaking through the safety defenses of large language models.
FUSE: Ensembling Verifiers with Zero Labeled Data
Joonhyuk Lee (Stanford University), Emmanuel Candes
Explainability and InterpretabilityComputational EfficiencyData-Centric LearningLarge Language ModelScore-based ModelContrastive LearningTextBenchmark
🎯 What it does: Proposes a fully unsupervised score ensemble method called FUSE, which aims to improve the validation quality of LLM outputs without any labeled data.
FUSE: FK-Steered Multi-Modal Flow Matching for Efficient Simulation-Based Posterior Estimation
Weichen Qin (ShanghaiTech University), Jiakai Zhang (ShanghaiTech University)
Data SynthesisOptimizationComputational EfficiencyRepresentation LearningTransformerDiffusion modelScore-based ModelFlow-based ModelRectified FlowImageMultimodalityTabularTime SeriesBenchmarkPhysics RelatedStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Propose the FUSE framework, combining dual-track multi-modal flow matching and Feynman-Kac steering to achieve efficient posterior estimation.
FUSE: Frequency-domain Unification and Spectral Energy Alignment for Multi-modal Object Re-Identification
Xuanhao Qi (Xi'an Jiaotong University), Lei Tan (National University of Singapore)
RecognitionRetrievalTransformerContrastive LearningImageVideoMultimodality
🎯 What it does: To address the issues of incomplete frequency domain information and unstable cross-modal alignment in multi-modal ReID, the FUSE framework is proposed, splitting the task into two stages: spectral decomposition and energy alignment.
FUSE: Full‑spectrum Unlearnable Examples via Spectral Equalization
Jiale Cai (Western University), Boyu Wang (Western University)
ClassificationAuto EncoderContrastive LearningImage
🎯 What it does: Designed a full-spectrum learnable example (UE) generation method called FUSE to protect data from being learned by models.
FUSE: Quantifying Uncertainty in Vision-Language Models by Bayesian Fusing Epistemic and Aleatoric Uncertainty
Harry Zhang (Massachusetts Institute of Technology), Luca Carlone (Massachusetts Institute of Technology)
Explainability and InterpretabilityRepresentation LearningTransformerVision Language ModelContrastive LearningImageTextMultimodality
🎯 What it does: This paper proposes the FUSE framework, which leverages data uncertainty and model uncertainty in vision-language models to generate a single interpretable confidence score;
FuseFSS: Efficient Secure LLM Inference with Function Secret Sharing
Yuhan Ma (Huawei Technologies Dusseldorf), Stefan Schmid (Technische Universitaet Berlin)
Safty and PrivacyComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringText
🎯 What it does: Propose the FuseFSS compiler, which unifies fixed-point nonlinearities and scaling operations in LLM inference into a two-step FSS evaluation, replacing manual protocols for each operator, significantly reducing implementation costs.
FusionCell: Cross-Attentive Fusion of Layout Geometry and Netlist Topology for Standard-Cell Performance Prediction
Haoyi Zhang (Peking University), Runsheng Wang (Peking University)
OptimizationComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerContrastive LearningImageGraphTabular
🎯 What it does: Propose FusionCell, a dual-modal standard cell performance prediction framework that jointly models layout geometry and netlist topology;
Future Dynamic 3D Reconstruction: A 3D World Model with Disentangled Ego-Motion
Nils Morbitzer (Technical University of Munich), Stefano Gasperini (Technical University of Munich)
Autonomous DrivingKnowledge DistillationRepresentation LearningTransformerAuto EncoderContrastive LearningWorld ModelOptical FlowImageVideo
🎯 What it does: Given a monocular image sequence, predict dynamic 3D reconstruction and camera trajectory for the next 2 seconds.
Future-Gain Guided Test-Time Learning for Large Language Models
LangYu Bian, Mingkui Tan (South China University of Technology)
Domain AdaptationOptimizationComputational EfficiencyRepresentation LearningTransformerLarge Language ModelTextBenchmark
🎯 What it does: Propose FG-TTL, a test-time learning method based on Future-Gain
FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs
Qian Chen (Fudan University), Xipeng Qiu (Fudan University)
Representation LearningAdversarial AttackData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVideoTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-ThoughtAudio
🎯 What it does: Construct the FutureOmni benchmark to evaluate the future prediction capability of multi-modal LLMs in audio-visual environments.
G-RANS: Generalizable Residual-Aware Neural Solvers for Sparse Systems
Weixin Liao (Shanghai Jiao Tong University), Junchi Yan (Shanghai Jiao Tong University)
OptimizationGraph Neural NetworkReinforcement LearningAuto EncoderContrastive LearningGraphTabularPhysics Related
🎯 What it does: Proposes G-RANS, a neuralized iterative solver that utilizes residual-aware subspace correction to accelerate the solution of sparse linear systems.
G$^2$RPO: Geometric GRPO; Escaping LLM's Reasoning Rut to Break Accuracy--Entropy Trade-off
Ali Rad (Cognichip AI), Ehsan Kamalinejad (Cognichip AI)
OptimizationRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextChain-of-ThoughtOrdinary Differential Equation
🎯 What it does: This paper proposes a RLVR method called G RPO for post-training of LLMs, which suppresses diversity collapse caused by GRPO by adding the reciprocal gain based on mode probability to the advantage function, and maintains the accuracy learning channel through a "neutralization" correction.
G$^2$TAM: Geometry Grounded Track Anything Model
Chenming Zhu (University of Hong Kong), Xihui Liu (University of Hong Kong)
Object TrackingSegmentationPose EstimationDepth EstimationConvolutional Neural NetworkTransformerPrompt EngineeringVision Language ModelNeural Radiance FieldContrastive LearningGaussian SplattingImageVideoTextPoint CloudMeshBenchmark
🎯 What it does: Propose a unified Geometry Grounded Tracking Anything Model (G2TAM), enabling promptable 3D instance tracking, 3D reconstruction, and video object segmentation using only RGB images/videos without pose information.
GAAVI: Global Asymptotic Anytime Valid Inference for the Conditional Mean Function
Brian M Cho (Cornell Tech), Nathan Kallus (Cornell Tech)
OptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyData-Centric LearningDiffusion modelScore-based ModelFlow-based ModelRectified FlowNeural Radiance FieldAuto EncoderContrastive LearningGaussian SplattingTabularTime SeriesSequentialReview/Survey PaperBenchmarkFinance Related
🎯 What it does: Propose a GAAVI method for global hypothesis testing of the conditional mean function (CMF) and CATE under continuous monitoring, which can provide high-confidence decisions at any time after sufficient samples are collected.
GADA: Geometry-Aware Deformable Aggregation for Image-Based Gaussian Splatting
Siwoo Lim (Korea Advanced Institute of Science and Technology), Chang D. Yoo (Korea Advanced Institute of Science and Technology)
RestorationSuper ResolutionComputational EfficiencyConvolutional Neural NetworkTransformerDiffusion modelNeural Radiance FieldGaussian SplattingOptical FlowImagePoint CloudBenchmark
🎯 What it does: Correct geometric errors in warping-based Gaussian Splatting in image rendering by using geometry-aware deformable aggregation to recover high-frequency details.
GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert
Mingyu Liu (Zhejiang University), Chunhua Shen (Zhejiang University)
Robotic IntelligenceReinforcement Learning from Human FeedbackTransformerSupervised Fine-TuningMixture of ExpertsVision Language ModelVision-Language-Action ModelDiffusion modelImageVideoTextPoint Cloud
🎯 What it does: Propose a task-agnostic general action expert (GAE), which converts the high-level intent of a vision-language model (VLM) into robotic continuous actions through a sparse 3D pose trajectory interface.
GAM-RAG: Gain-Adaptive Memory for Evolving Retrieval in Retrieval-Augmented Generation
Yifan Wang (Fudan University), Hongfeng Chai (Fudan University)
RetrievalComputational EfficiencyData-Centric LearningGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextGraphRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Construct a lightweight hierarchical graph index, and introduce two memory vectors, task and time, with the sentence as the minimum unit during retrieval. In the reasoning process, adaptively update the memory in real-time based on the judgment results of the LLM, thus achieving dynamic retrieval improvement without training.
Game-Theoretic Co-Evolution for LLM-Based Heuristic Discovery
Xinyi Ke (Institute of Automation, Chinese Academy of Sciences), Jian Cheng (Institute of Automation, Chinese Academy of Sciences)
OptimizationMeta LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningTabularSequentialBenchmark
🎯 What it does: Proposed a game-theoretic co-evolutionary framework called ASRO, which views LLM-driven heuristic discovery as a zero-sum game between a solver and an instance generator.
GameDevBench: Evaluating Agentic Capabilities Through Game Development
Wayne Chi (Carnegie Mellon University), Chris Donahue (Carnegie Mellon University)
AI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringImageVideoTextMultimodalityBenchmarkRetrieval-Augmented GenerationAudio
🎯 What it does: Develop the GameDevBench benchmark to evaluate the capabilities of LLM agents in Godot game development tasks;
GameVerse: Can Vision-Language Models Learn from Video-based Reflection?
Kuan Zhang (Tsinghua University), Yiming Li (Tsinghua University)
Reinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringVision Language ModelVideoTextMultimodalityBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose the GameVerse benchmark, combining cognitive hierarchical classification and video reflection loops, enabling VLMs to learn by watching failure and tutorial videos across 15 games.
Gaming Consensus: Coordinated Manipulation in Crowdsourced Fact-Checking
Nikil Roashan Selvam (Stanford University), Sanmi Koyejo (Stanford University)
Anomaly DetectionOptimizationFederated LearningExplainability and InterpretabilityAdversarial AttackData-Centric LearningText
🎯 What it does: Studied the vulnerability of bridge algorithms based on matrix factorization in community note systems to collective manipulation, and proposed and evaluated a two-stage 'synthetic consensus' attack.
GASS: Geometry-Aware Spherical Sampling for Disentangled Diversity Enhancement in Text-to-Image Generation
Ye Zhu (Ecole Polytechnique), Olga Russakovsky (Princeton University)
GenerationTransformerPrompt EngineeringVision Language ModelDiffusion modelContrastive LearningImageTextMultimodality
🎯 What it does: By decomposing the CLIP space geometrically, the diversity of text-to-image generation is divided into prompt-related and prompt-agnostic categories, and during sampling, Geometry-Aware Spherical Sampling (GASS) is used to guide the generation of more diversely distributed images.