π― What it does: Propose the PRISM framework to achieve efficient test-time scaling (Test-Time Scaling) for discrete diffusion language models (dLLMs)
Elastic Attention: Test-time Adaptive Sparsity Ratios for Efficient Transformers
Zecheng Tang (Soochow University), Min Zhang (Soochow University)
CodeComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningText
π― What it does: Introduce a lightweight Attention Router into pre-trained large language models, enabling the model to dynamically assign full attention (FA) or sparse attention (SA) to each attention head based on the input during inference, thus achieving a variable sparse ratio.
Embodied Interpretability: Linking Causal Understanding to Generalization in Vision-Language-Action Models
Hanxin Zhang (University of Leicester), Zhou Daniel Hao (University of Leicester)
CodeExplainability and InterpretabilityRobotic IntelligenceVision-Language-Action ModelMultimodalityBenchmark
π― What it does: Proposes an explanation method based on causal interventionβInterventional Significance Score (ISS) and Nuisance Mass Ratio (NMR), used to quantify the causal dependence of visual language action (VLA) models on visual information and the utilization of irrelevant features.
π― What it does: Proposed an end-to-end temporal 3D object detection framework called Embodied-DETR for first-person perspective continuous RGB-D streams, and created a dedicated Embodied-Det benchmark.
π― What it does: Proposed a 3D Gaussian Splatting method called EnerGS based on an energy field, which utilizes partial geometric priors (such as LiDAR) to construct a continuous geometric energy field to guide the distribution of Gaussian primitives, thereby achieving more stable and accurate view synthesis in large-scale outdoor scenes.
Enhancing Conformal Prediction via Class Similarity
Ariel Fargion (Bar-Ilan University), Tom Tirer (Bar-Ilan University)
CodeClassificationExplainability and InterpretabilityComputational EfficiencyConvolutional Neural NetworkTransformerContrastive LearningImage
π― What it does: Propose a regularization method based on class similarity, which can improve the prediction set size and semantic consistency of any conformal prediction (CP) algorithm while maintaining coverage guarantees.
Enhancing Cross-subject Emotion Recognition via Heterogeneous Distribution Augmentation and Collaborative Learning
Wending Xiong (Wuhan University), Mang Ye (Wuhan University)
CodeRecognitionData SynthesisDomain AdaptationGenerative Adversarial NetworkContrastive LearningMultimodalityTime SeriesBiomedical Data
π― What it does: Proposes the MixEmo framework, which enhances the generalization ability of cross-subject emotion recognition by separating and recombining the distribution of emotional data to generate unseen distributions, and collaboratively learning across multiple sub-distributions.
Enhancing Numerical Prediction in LLMs via Smooth MMD Alignment
Zhuo Zuo (Sichuan University), Xianggen Liu (Sichuan University)
CodeClassificationRecommendation SystemOptimizationData-Centric LearningTransformerLarge Language ModelMixture of ExpertsContrastive LearningTextTabularBenchmark
π― What it does: Proposed a novel numerical prediction training loss called SMMD, which constructs a distance kernel based on a numerical subword vocabulary and uses MMD to match distributions, while applying graph Laplacian smoothing regularization to the prediction-target residual to improve the accuracy of LLMs in numerical outputs.
Enhancing Protein-Protein Interaction Prediction with Hierarchical Motif-based Multimodal Protein Embedding
Zaifei YANG (Hong Kong University of Science and Technology), James Kwok (Hong Kong University of Science and Technology)
CodeDrug DiscoveryProtein Structure PredictionGraph Neural NetworkTransformerSupervised Fine-TuningContrastive LearningMultimodalityGraphBiomedical Data
π― What it does: A hierarchical multi-modal protein encoder, MMM-PPI, was constructed to improve the prediction of protein-protein interactions (PPI).
Entropy-Aware On-Policy Distillation of Language Models
Woogyeol Jin (KAIST AI), Kimin Lee (KAIST AI)
CodeKnowledge DistillationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningContrastive LearningText
π― What it does: This paper proposes a new adversarial knowledge distillation method called Entropy-Aware On-Policy Distillation (EOPD), which improves the distillation process of language models on self-generated trajectories by using forward KL in high-entropy positions and backward KL in low-entropy positions.
Entropy-aware Span-Constrained Optimal Transport for Robust Cross-Tokenizer Knowledge Distillation
Zhi-Ping Liu (Nanjing University), Xinghao Chen (Huawei)
CodeKnowledge DistillationTransformerLarge Language ModelSupervised Fine-TuningText
π― What it does: Propose the E-SCOT framework in cross-tokenizer knowledge distillation, viewing distillation as a sparse optimal transport problem, leveraging a vocabulary-agnostic baseline metric, span-anchored lexical alignment, and adaptive reweighting based on R'-Enn entropy, to achieve more reliable alignment and information transfer between teacher and student.
π― What it does: Reformulate the traditional state-based predictive coding (sPC) and propose error-based predictive coding (ePC) to eliminate the problem of exponential signal decay in digital simulations and achieve fast convergence in deep networks.
Epistemic Uncertainty Quantification for Pre-trained VLMs via Riemannian Flow Matching
Li Ju (Uppsala University), Prashant Singh (Uppsala University)
CodeAnomaly DetectionRepresentation LearningData-Centric LearningVision Language ModelScore-based ModelFlow-based ModelContrastive LearningImageTextMultimodality
π― What it does: Propose a framework called REPVLM based on Riemannian flow matching, which is used to estimate the probability density of pre-trained vision-language models (VLMs) in their embedding space, thereby obtaining the model's awareness of uncertainty in its representations.
CodeComputational EfficiencySpiking Neural NetworkReinforcement LearningTabularTime Series
π― What it does: This paper studies the conversion of pre-trained artificial neural networks (ANN) into spiking neural networks (SNN), and analyzes the error amplification problem in continuous control tasks.
CodeReinforcement LearningContrastive LearningTabularTime SeriesSequentialElectronic Health RecordsBenchmark
π― What it does: This paper proposes a new Truncated Policy Gradient (TPG) estimator to estimate the global average treatment effect (GATE) from a single random experimental trajectory in non-stationary Markov environments.
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment
Xiuyu Li (Renmin University of China), Ju Fan (Renmin University of China)
CodeOptimizationComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningScore-based ModelTextBenchmark
π― What it does: Proposes ETS, a training-free inference method that directly generates text by sampling from the optimal RL policy.
Euclean: Automated Geometry Problem Formalization with Unified Verification in Lean
Linbin Tang (Tsinghua University), Fan Yang (Microsoft Research)
CodeOptimizationExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningAI Code AssistantTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
π― What it does: Propose the EUCLEAN framework, which automatically converts natural language geometry problems into the MATHLIB formalization of Lean 4, completing a four-stage pipeline: constraint explicitation, configuration anchoring, mapping, and iterative repair.
Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability
Yong Ren (Institute of Automation, Chinese Academy of Sciences), Xuerui Yang (StepFun)
CodeReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextMultimodalityAudio
π― What it does: Propose a 'Mean Continued Log Probability' (MCLP) based on LALM, which can serve both as an evaluation metric and as a reinforcement learning (RL) reward, to enhance the speech expression consistency in role-playing TTS.
CodeComputational EfficiencyData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextBenchmark
π― What it does: Propose a framework that utilizes pairwise comparison signals generated by LLMs as control variables, combined with semi-parametric inference to improve mathematical reasoning evaluation.
Evaluating Robustness of Reasoning Models on Parameterized Logical Problems
NaΓ―m Es-sebbani (University of Artois), Zied Bouraoui (University of Caen Basse Normandie)
CodeExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextBenchmarkChain-of-Thought
π― What it does: Constructed a diagnostic 2-SAT benchmark based on a parameterizable structured 2-CNF formula to evaluate the robustness of large language models in logical reasoning.
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning
Xin Qiu (Cognizant AI Lab), Risto Miikkulainen (Cognizant AI Lab)
CodeOptimizationComputational EfficiencyHyperparameter SearchReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextBenchmark
π― What it does: This paper is the first to directly apply evolutionary strategies (ES) to full-parameter fine-tuning of language models with billions of parameters, without any dimensionality reduction, demonstrating the feasibility of ES on large-scale LLMs;
Evolving Interdependent Operators with Large Language Models for Multi-Objective Combinatorial Optimization
junhao qiu, Qingfu Zhang (City University of Hong Kong)
CodeOptimizationTransformerLarge Language ModelPrompt EngineeringTextTabular
π― What it does: Propose an E2OC framework based on large language models (LLMs) that automatically co-evolves combinations of multiple neighborhood search operators to enhance the performance of multi-objective evolutionary algorithms (MOEAs).
ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation
Yiran Wu (Pennsylvania State University), Anand Mudgerikar (Microsoft Security AI Research)
CodeExplainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringAuto EncoderTextGraphTabularBenchmarkRetrieval-Augmented GenerationChain-of-Thought
π― What it does: This work constructs ExCyTIn-Bench, a benchmark for evaluating large language model (LLM) agents in network threat investigation tasks; by creating an interactive MySQL environment, automatically generated question-answer pairs, and fine-grained progress rewards, the performance of agents in practical security log querying and reasoning is assessed.
Explicitly Modeling Censoring Produces Superior Survival Predictors
Shi-ang Qi (University of Alberta), Russell Greiner (University of Alberta)
CodeComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerContrastive LearningTabularTime SeriesBiomedical DataElectronic Health RecordsBenchmark
π― What it does: Propose a survival prediction framework that explicitly models the censoring process, treating event time and censoring time as two processes sharing parameters.
π― What it does: Structural Hessian approximation is achieved by orbit averaging a single gradient using weight space symmetry, enabling efficient curvature estimation.
π― What it does: This paper investigates the 'exploration hacking' phenomenon where LLMs hinder training in RL by reducing exploration behavior, and constructs model organisms to simulate this behavior.
π― What it does: Investigate and utilize the stability of the Latent Flow Matching (LFM) model under perturbations in data subsets, model capacity, and training configurations, proposing data pruning methods and a two-stage inference acceleration technique from coarse to fine.
π― What it does: Studied the compatibility issues that arise when migrating LoRA to distilled models in video diffusion models (VDM), and proposed a data-agnostic Cluster-Aware Spectral Arbitration (CASA) method to achieve training-free migration of LoRA.
Exploring Nonlinear Pathway in Parameter Space for Machine Unlearning
Yingdan Shi (Illinois Institute of Technology), Ren Wang (Illinois Institute of Technology)
CodeClassificationSafty and PrivacyDiffusion modelScore-based ModelContrastive LearningImage
π― What it does: Proposes a nonlinear path exploration framework called MCU based on pattern connectivity, achieving efficient amnesia in machine learning models.
FACT: Fuzzy Alignment with Comorbidity Topology for Reliable Multi-Label Medical Image Diagnosis
Yingyu Chen (Sichuan University), Yi Zhang (Sichuan University)
CodeClassificationRecognitionAnomaly DetectionExplainability and InterpretabilityConvolutional Neural NetworkGraph Neural NetworkTransformerAuto EncoderContrastive LearningImageGraphBiomedical DataMagnetic Resonance ImagingComputed TomographyElectronic Health Records
π― What it does: Propose the FACT framework, which treats multi-label medical image diagnosis as a fuzzy alignment problem between atomic visual evidence and disease semantic anchors, addressing the limitations of hard segmentation caused by visual ambiguity and disease associations.
π― What it does: This paper proposes a fair dataset distillation framework called COBRA, which significantly reduces the fairness gap between subgroups by performing balanced cross-group barycenter alignment of representations for different subgroups during the distillation process, while maintaining overall performance.
Fairness in Aggregation: Optimal Top-$k$ and Improved Full Ranking
Diptarka Chakraborty (National University Of Singapore), Alvin Hong Yao Yan (National University Of Singapore)
CodeRecommendation SystemOptimizationTabular
π― What it does: This paper studies the problem of fair ranking aggregation under the Spearman footrule distance, providing a polynomial optimal algorithm for top-k fair ranking, and proposing a 2-approximation algorithm for full ranking.
FairRARI: A Plug and Play Framework for Fairness-Aware PageRank
Emmanouil Kariotakis (KU Leuven), Aritra Konar (KU Leuven)
CodeOptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyRepresentation LearningGraph Neural NetworkGraph
π― What it does: Propose a framework named FairRARI, which utilizes a variational formulation of PageRank to achieve pluggable solutions for group fairness across multiple groups;
Jiaee Cheong (Harvard University), Sinan Kalkan (METU)
CodeFederated LearningSafty and PrivacyExplainability and InterpretabilityRepresentation LearningAdversarial AttackTransformerAuto EncoderContrastive LearningImageTextMultimodalityTabularBiomedical DataElectronic Health RecordsAudio
π― What it does: Proposed a fair self-supervised learning framework called FairSSL, designed for heterogeneous, variable-length multimodal data, aiming to improve fairness while maintaining performance on downstream tasks.
Fast and Accurate Causal Parallel Decoding using Jacobi Forcing
Lanxiang Hu (University of California San Diego), Hao Zhang (University of California San Diego)
CodeComputational EfficiencyKnowledge DistillationAI Code AssistantTransformerLarge Language ModelText
π― What it does: Based on autoregressive (AR) large language models (LLMs), the Jacobi Forcing training method is proposed, directly transforming AR models into efficient parallel multi-token decoders;
Adam Zweiger (Massachusetts Institute of Technology), Yoon Kim (Massachusetts Institute of Technology)
CodeCompressionComputational EfficiencyTransformerLarge Language ModelTextMultimodalityBenchmarkRetrieval-Augmented Generation
π― What it does: This study proposes a fast KV cache compression method based on Attention Matching, which directly approximates the original attention output and attention quality by optimizing the compressed key-value pairs, maintaining the model's performance in long contexts.
π― What it does: By introducing a training-agnostic dynamic acceleration module into the SAM3D single-view 3D reconstruction pipeline, the inference latency is significantly reduced.
π― What it does: Propose a no-training, plug-and-play acceleration framework called FasterVAR, which accelerates the late-stage detail refinement phase in the high-resolution generation process of visual autoregressive (VAR) models.
π― What it does: Propose a fine-tuning framework based on entropy called Dem-HEC, which generates high-entropy samples by maximizing the model output entropy within a restricted perturbation space, and combines cross-entropy, contrastive learning, and knowledge distillation to enhance the model's robustness against natural noise/distortion corruption, while maintaining or improving the accuracy on clean images.
π― What it does: Propose FAHNES, a scalable hierarchical generative framework capable of simultaneously generating the topological structures and node/edge features of graphs/hypergraphs.
Federated Sketching LoRA: A Flexible Framework for Heterogeneous Collaborative Fine-Tuning of LLMs
Wenzhi Fang (Purdue University), Christopher Brinton
CodeOptimizationFederated LearningComputational EfficiencyTransformerLarge Language ModelSupervised Fine-TuningText
π― What it does: Propose FSLoRA β a LoRA fine-tuning method using sketching within a federated learning framework, allowing clients with different resources to update only the submatrix of the global LoRA module, thereby achieving efficient heterogeneous collaborative fine-tuning.
FedHPro: Federated Hyper-Prototype Learning via Gradient Matching
Huan Wang (University of Wollongong), Guansong Pang (Singapore Management University)
CodeFederated LearningExplainability and InterpretabilityRepresentation LearningContrastive LearningImageMultimodalityTabular
π― What it does: Propose the FedHPro framework, which learns interpretable hyper-prototypes through gradient matching in federated learning, and improves model generalization by using two contrastive and alignment mechanisms, HPCL and HPAL, during local training.
π― What it does: Propose the FedQueue algorithm, addressing the random enqueue delay caused by batch scheduling in cross-HPC facility training, by constructing a complete federated learning framework with queue prediction, cutoff enqueue control, and delay-aware aggregation.
FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model
Chunyu Xie (Beihang University), Yuhui Yin (360 AI Research)
CodeClassificationObject DetectionSegmentationRetrievalTransformerPrompt EngineeringMixture of ExpertsVision Language ModelContrastive LearningImageTextMultimodality
π― What it does: Propose FG-CLIP 2, which constructs an English-Chinese bilingual fine-grained vision-language alignment model, achieving precise alignment through two-stage training and multi-task loss.
π― What it does: Propose the FiGuRO method, which combines a low-rank adaptive layer with rate-distortion theory to dynamically estimate the intrinsic dimension of multi-modal data, and achieves spontaneous separation in shared and private subspaces.
Lucas Darius Konrad, Nikolas Kuschnig (Monash University)
CodeOptimizationExplainability and InterpretabilityComputational EfficiencyTextTabularBiomedical Data
π― What it does: Proposed an efficient algorithm for finding the most influential subset (MIS) in a dataset. The algorithm transforms the combination search into multiple top-k selections by expressing the holdout effect as a linear fractional form, ultimately achieving exact solutions.
CodeExplainability and InterpretabilityComputational EfficiencyTransformerVision Language ModelImageTextMultimodalityBenchmark
π― What it does: Propose a no-training, plug-and-play method called ILVAD, which utilizes inter-layer visual attention differences to identify and reinforce correct visual evidence, thereby reducing hallucinations generated by large audio-visual models.
Fine-grained Analysis of Brain-LLM Alignment through Input Attribution
Michela Proietti (Goethe University), Mariya Toneva (Max Planck Institute for Software Systems)
CodeExplainability and InterpretabilityTransformerLarge Language ModelTextBiomedical DataMagnetic Resonance Imaging
π― What it does: Propose a pipeline that integrates input attribution methods (such as Integrated Gradients, GradientΓInput, SmoothGrad) into a brain-LLM alignment framework, used to identify the input words most important for predicting brain activity.
FIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights
Zhen Wang (University of California San Diego), Eric P. Xing (MBZUAI)
CodeTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
π― What it does: Proposes FIRE-Bench, a comprehensive re-mining evaluation benchmark based on verified scientific discoveries, aimed at assessing the scientific reasoning and experimental capabilities of LLM-driven autonomous research agents.
FIRE: Multi-Fidelity Regression with Distribution-Conditioned In-Context Learning Using Tabular Foundation Models
Rosen Ting-Ying Yu (Massachusetts Institute of Technology), Faez Ahmed (Massachusetts Institute of Technology)
CodeOptimizationHyperparameter SearchData-Centric LearningTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsTabularTime SeriesSequentialBenchmarkPhysics Related
π― What it does: Designed a training-agnostic multi-talented regression framework called FIRE, which leverages the TabPFN (table foundation model) to achieve zero-shot Bayesian inference for low-resolution models and high-resolution residual correction through distribution-conditional context learning.
Fix Before Search: Benchmarking Agentic Visual Query Pre-processing in Multimodal Retrieval-augmented Generation
Shenglai Zeng (Michigan State University), Yi Chang (Jilin University)
CodeData SynthesisRetrievalTransformerLarge Language ModelSupervised Fine-TuningAgentic AIVision Language ModelImageTextMultimodalityBenchmarkRetrieval-Augmented Generation
π― What it does: This paper proposes V-QPP-Bench, a benchmark for visual query preprocessing in multi-modal retrieval-augmented generation (MRAG), and systematically evaluates the performance of different MLLMs on this benchmark.
π― What it does: This paper re-examines the mechanism of Sharpness-Aware Minimization (SAM), discovering that its fixed-radius first-order linear approximation leads to a mismatch between the gradient norm-dominated learning signal and the second-order nature of flat minima. It proposes Loss-Equated SAM (LE-SAM), which eliminates the interference of gradient norms by fixing a budget in the loss space and solving for the corresponding radius in the parameter space, shifting the optimization focus to curvature information. Additionally, radius clipping and loss budget annealing are introduced to ensure stability, and further, LE-SAM+ is proposed to enhance curvature-aware regularization. Experimental results verify that this mechanism significantly improves generalization performance across various tasks and models.
FiX: Introducing Fine-grained Forget Gate into Softmax Attention
Runzhong Li (Southern University of Science and Technology), Bo Tang (Southern University of Science and Technology)
CodeComputational EfficiencyRepresentation LearningData-Centric LearningTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsText
π― What it does: Propose Fine-grained Forgetting Transformer (FiX), introducing an element-wise forgetting gate into softmax attention to enhance the modeling capability for long-text contexts.
π― What it does: This paper proposes the FLAG framework, which predicts gene space expression from H&E slices by utilizing spatial graph encoding and a gene foundation model.
Fleet: Few Shots Lead Effective AI-generated Image Detection
Jiaan Wang (Institute of Computing Technology, Chinese Academy of Sciences), Sheng Tang (Institute of Computing Technology, Chinese Academy of Sciences)
π― What it does: Propose the Fleet framework, achieving dynamic adaptive AIGI detection based on subspace routing, enabling rapid adaptation to new generative models with very few samples.
π― What it does: Proposed a class of sequence kernels (LOCK) that utilize evolutionary substitution matrices and local linear features, and embedded them into a Gaussian process model for protein attribute prediction.
CodeDrug DiscoveryGraph Neural NetworkTransformerDiffusion modelScore-based ModelFlow-based ModelRectified FlowGraphBiomedical Data
π― What it does: Propose and implement FlexiFlow, which utilizes a decomposable flow matching framework to simultaneously generate molecular graphs and multiple low-energy conformation sets in a single sampling process.
π― What it does: Propose a calibration framework based on flow matching called FMCPE, which uses a small number of real calibration samples to correct the posterior distribution of simulated inferences.
π― What it does: Proposes a FlowCloud framework based on variational Neural ODE, which can learn continuous spatiotemporal dynamics and generate complete trajectories from sparse, non-continuous, and unpaired point cloud snapshots.
FOAM: Blocked State Folding for Memory-Efficient LLM Training
Ziqing Wen (National University of Defense Technology), Tao Sun (National University of Defense Technology)
CodeOptimizationComputational EfficiencyTransformerLarge Language ModelText
π― What it does: This paper proposes an optimizer called FOAM, which utilizes block-wise gradient averaging and residual correction to compress the optimizer state of Adam, significantly reducing memory usage and accelerating convergence during LLM training.
Kaihua Liang (King Abdullah University of Science and Technology), Marco Canini (King Abdullah University of Science and Technology)
CodeComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation
π― What it does: By identifying that most computations in the decoding process of DLLMs are wasted on tokens that are not decoded, the FOCUS system dynamically removes these useless tokens, significantly reducing FLOPs.
π― What it does: This paper designs a biology-inspired concave sampling interface called FOVI based on retino-cortical mapping, and combines it with convolutional networks and Vision Transformers to significantly reduce pixel and computational costs in high-resolution visual tasks.
FPTQuant: Function-Preserving Transforms for LLM Quantization
Boris van Breugel (Qualcomm AI Research), Markus Nagel (Qualcomm AI Research)
CodeComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelText
π― What it does: Proposed a function-preserving transformation (FPT) method called FPTQuant for low-precision INT4 quantization on large language models (LLMs), while maintaining the model's functionality unchanged.
Yuhan Zhu (Nanjing University), Limin Wang (Nanjing University)
CodeRetrievalTransformerLarge Language ModelPrompt EngineeringImageVideoTextMultimodalityBenchmarkRetrieval-Augmented Generation
π― What it does: Propose FreeRet, a training-free framework that can directly convert any off-the-shelf multimodal large language model (MLLM) into a two-stage retriever, performing both embedding extraction and reranking to complete the end-to-end process of retrieval and generation.
From Absolute to Relative: Rethinking Reward Shaping in Group-Based Reinforcement Learning
Wenzhe Niu (Tianjin University), Renqing He (Meituan)
CodeTransformerLarge Language ModelReinforcement LearningPrompt EngineeringContrastive LearningText
π― What it does: Proposed a reinforcement learning framework based on relative rewards, RLRR, which transforms rewards from absolute scores into group-relative rankings to address the issues of sparse rewards in verifiable tasks and unstable reward ranges in open-ended tasks;
From Backward Spreading to Forward Replay: Revisiting Target Construction in LLM Parameter Editing
Wei Liu (National University of Singapore), Wee Sun Lee (National University of Singapore)
CodeOptimizationComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextBenchmark
π― What it does: Studied the theoretical foundations of target construction in parameter editing of large language models, and proposed a forward replay method based on forward propagation to replace the traditional backward propagation diffusion, achieving more precise multi-layer targets while maintaining the same computational complexity.
From Blind Spots to Gains: Diagnostic-Driven Iterative Training for Large Multimodal Models
Hongrui Jia (Peking University), Wei Ye (Peking University)
CodeExplainability and InterpretabilityComputational EfficiencyData-Centric LearningReinforcement Learning from Human FeedbackTransformerReinforcement LearningAgentic AIPrompt EngineeringVision Language ModelGenerative Adversarial NetworkImageTextMultimodality
π― What it does: This paper proposes a diagnosis-driven iterative training framework called DPE, which combines multi-agent tool-based data generation with reinforcement learning to dynamically generate and reinforce training samples targeting the model's blind spots.
π― What it does: Proposed a multi-round Agentic RAG framework, MA-RAG, to enhance answer quality in medical question answering through iterative retrieval and reasoning.
CodeClassificationExplainability and InterpretabilityComputational EfficiencyData-Centric LearningScore-based ModelContrastive LearningImageTextTabular
π― What it does: Propose a multi-classifier based on the asymmetric Laplace distribution (HALD), and achieve reliable probabilistic prediction through individualized calibration (HICALD);
From Interactions to Principles: Experience-Driven Self-Distillation for Evolving LLM Agents
Rong Wu (Zhejiang University), Botian Shi (Shanghai Artificial Intelligence Laboratory)
CodeKnowledge DistillationTransformerLarge Language ModelReinforcement LearningAgentic AITextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
π― What it does: Propose the EvolveR framework, constructing a complete closed-loop experience lifecycle, including offline self-distillation to generate abstract principles, online interaction to retrieve experiences, and iterative optimization of strategies through reinforcement learning.
From Internal Diagnosis to External Auditing: A VLM-Driven Paradigm for Data-Free Online Backdoor Defense
Binyan Xu (Chinese University of Hong Kong), Kehuan Zhang (Chinese University of Hong Kong)
CodeAnomaly DetectionTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityRetrieval-Augmented Generation
π― What it does: Propose PRISM, an online backdoor defense framework that does not require access to training data or modification of model weights, utilizing external Vision-Language Models to perform semantic auditing on model predictions.
π― What it does: Propose Parameterized Diffusion Policy (PDP), which parameterizes traditional diffusion policies by learning a geometry-aligned behavioral latent space, enabling precise control of behavior on low-dimensional latent variables;
From Noise to Intent: Anchoring Generative VLA Policies with Residual Bridges
Yiming Zhong (ShanghaiTech University), Yuexin Ma (ShanghaiTech University)
CodeRobotic IntelligenceTransformerReinforcement LearningVision Language ModelVision-Language-Action ModelDiffusion modelFlow-based ModelRectified FlowMultimodality
π― What it does: Propose the ResVLA framework, which adopts a generative VLA strategy combining low-frequency intent anchoring with high-frequency residual diffusion bridge, achieving a fundamental shift from 'Generation-from-Noise' to 'Refinement-from-Intent'.
From Observations to States: Latent Time Series Forecasting
Jie Yang (University of Illinois Chicago), Philip S. Yu (University of Illinois Chicago)
CodeInformation TheoryAnomaly DetectionComputational EfficiencyRepresentation LearningTransformerAuto EncoderContrastive LearningTabularTime Series
π― What it does: Propose LatentTSF, which transforms time series forecasting from direct observation regression to first mapping via AutoEncoder to a latent space, then performing state prediction in that space, and finally decoding the predicted latent states back to observations.
From Perception to Planning: Evolving Ego-Centric Task-Oriented Spatiotemporal Reasoning via Curriculum Learning
Xiaoda Yang (Zhejiang University), Zhou Zhao (Zhejiang University)
CodeAutonomous DrivingComputational EfficiencyRepresentation LearningRobotic IntelligenceTransformerSupervised Fine-TuningVision Language ModelContrastive LearningImageVideoTextMultimodalityChain-of-Thought
π― What it does: Design and implement the EgoTSR framework, which adopts a three-stage curriculum learning (CoT β Tag β LongTag) to progressively achieve egocentric task-based spatiotemporal reasoning, moving from explicit spatial awareness, internalized judgment, to long-term planning.
From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning
Jike Zhong (University of Southern California), Shao-Yuan Lo (National Taiwan University)
CodeExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningTextMultimodalityRetrieval-Augmented GenerationChain-of-Thought
π― What it does: This paper studies how to enhance the Theory of Mind (ToM) capability of large language models through post-training. For the first time, it systematically audits and removes shortcuts from the dataset, and then introduces Thinking-RFT (Thinking-based Reinforcement Fine-Tuning) to improve the model's reasoning and generalization performance.
From Winning to Understanding: A Diagnostic Long-Horizon RTS Benchmark for LLMs
Jiacheng Li (University of Chinese Academy of Sciences), Chenghao Li (Tsinghua University)
CodeTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextTime SeriesSequentialBenchmarkRetrieval-Augmented GenerationChain-of-Thought
π― What it does: Built a long-term, adversarial real-time strategy game benchmark, using LLM as the decision module, providing three-track evaluation (against rule-based AI, LLM adversaries, and human instruction following)
π― What it does: Proposes a fully zero-shot image dehazing framework that is trained only using clean images, utilizing color, structural, and illumination invariant representations derived from physical models, and achieving dehazing through diffusion models.
FunCQNet: A Functional Censored Quantile Neural Network for Predicting Long-Term Post-Transplant Kidney Survival
Jiaqi Men (Shanghai University of Finance and Economics), Jiguo Cao (Simon Fraser University)
CodeExplainability and InterpretabilityComputational EfficiencyDrug DiscoveryRecurrent Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningTabularTime SeriesBiomedical DataElectronic Health RecordsStochastic Differential Equation
π― What it does: Propose the FunCQNet framework, which utilizes deep neural networks and truncated quantile loss to estimate the time-varying coefficients of interactions between functional biomarkers and scalar covariates, thereby predicting long-term survival after kidney transplantation;
π― What it does: Propose Functional Adjoint Sampler (FAS), a diffusion sampler that can sample Gibbs distributions in infinite-dimensional Hilbert spaces, capable of directly sampling in the trajectory space and achieving endpoint constraints.
Ali Rad (Cognichip AI), Ehsan Kamalinejad (Cognichip AI)
CodeOptimizationRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextChain-of-ThoughtOrdinary Differential Equation
π― What it does: This paper proposes a RLVR method called G RPO for post-training of LLMs, which suppresses diversity collapse caused by GRPO by adding the reciprocal gain based on mode probability to the advantage function, and maintains the accuracy learning channel through a "neutralization" correction.
π― What it does: Correct geometric errors in warping-based Gaussian Splatting in image rendering by using geometry-aware deformable aggregation to recover high-frequency details.
Game-Theoretic Co-Evolution for LLM-Based Heuristic Discovery
Xinyi Ke (Institute of Automation, Chinese Academy of Sciences), Jian Cheng (Institute of Automation, Chinese Academy of Sciences)
CodeOptimizationMeta LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningTabularSequentialBenchmark
π― What it does: Proposed a game-theoretic co-evolutionary framework called ASRO, which views LLM-driven heuristic discovery as a zero-sum game between a solver and an instance generator.
GameDevBench: Evaluating Agentic Capabilities Through Game Development
Wayne Chi (Carnegie Mellon University), Chris Donahue (Carnegie Mellon University)
CodeAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringImageVideoTextMultimodalityBenchmarkRetrieval-Augmented GenerationAudio
π― What it does: Develop the GameDevBench benchmark to evaluate the capabilities of LLM agents in Godot game development tasks;
GASS: Geometry-Aware Spherical Sampling for Disentangled Diversity Enhancement in Text-to-Image Generation
Ye Zhu (Ecole Polytechnique), Olga Russakovsky (Princeton University)
CodeGenerationTransformerPrompt EngineeringVision Language ModelDiffusion modelContrastive LearningImageTextMultimodality
π― What it does: By decomposing the CLIP space geometrically, the diversity of text-to-image generation is divided into prompt-related and prompt-agnostic categories, and during sampling, Geometry-Aware Spherical Sampling (GASS) is used to guide the generation of more diversely distributed images.
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Shih-Yang Liu (NVIDIA), Pavlo Molchanov (NVIDIA)
CodeOptimizationReinforcement LearningText
π― What it does: This paper investigates the limitations of the traditional GRPO method in multi-reward reinforcement learning, and proposes GDPO by separating reward normalization to avoid the reward folding problem, thereby improving training stability and performance.
π― What it does: Propose the GemDepth framework, which utilizes a geometric embedding module (GEM) and an alternating spatiotemporal Transformer (ASTT) to achieve 3D-consistent video depth estimation;
GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs
Jianing Deng (University of Pittsburgh), Jingtong Hu (University of Pittsburgh)
CodeComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringMixture of ExpertsText
π― What it does: This paper proposes a global expert-level mixed-precision quantization method called GEMQ, which systematically addresses the expert bit-width allocation and routing offset issues in MoE LLMs;
GenAlign: Towards Unified Alignment Framework of MLLMs via Generative Reward Model
Jingyu Zhang (Ant Group), shiwen cui
CodeExplainability and InterpretabilityComputational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringVision Language ModelGenerative Adversarial NetworkContrastive LearningImageTextMultimodalityBenchmark
π― What it does: Propose the GenAlign framework, which combines generative reward models (GRM) with multimodal large language models (MLLM) alignment, achieving reasoning-based preference judgment based on adaptive rubric.
Meshi Bashari (Technion Iit), Yaniv Romano (Technion Iit)
CodeData SynthesisFederated LearningSafty and PrivacyComputational EfficiencyRepresentation LearningProtein Structure PredictionLarge Language ModelImageTextTabularBiomedical DataBenchmark
π― What it does: Developed a general framework called GESPI, which can securely utilize synthetic data in statistical inference while ensuring error rate control.