ICML 2026 Papers — Page 21
International Conference on Machine Learning · 6554 papers
Fast Mixture of Curvature-Aware Experts for Diverse and Dynamic Graph Topologies
Jiayi Yang (Tongji University), Wei Ye (Tongji University)
OptimizationComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerMixture of ExpertsContrastive LearningGraphTime SeriesSequential
🎯 What it does: Proposes DyGMoCE — a Transformer that uses a mixture of curvature experts in dynamic graphs, to adaptively embed nodes into geometric spaces at different time points.
Fast Non-Episodic Finite-Horizon RL with K-Step Lookahead Thresholding
Jiamin Xu (Cornell Tech), Kyra Gan (Cornell Tech)
Reinforcement LearningTabular
🎯 What it does: Proposed a K-step lookahead threshold strategy and designed the LGKT algorithm to achieve fast sample efficiency in non-periodic finite-horizon MDPs;
Fast Reconstruction of Mixtures of Bernoulli Product Distributions
Sanyam Agarwal (Universitaet des Saarlandes), Markus Bläser (Universitaet des Saarlandes)
OptimizationComputational EfficiencyRepresentation Learning
🎯 What it does: The authors propose an algorithm to exactly recover the parameters of a Bernoulli product mixture distribution from probabilistic generating polynomial (PGP) black-box queries, with a query complexity of O(n·r²);
Fast Spectrally Sparse Signal Reconstruction via Jacobi-Preconditioned Gradient Descent
Jian-Feng Cai (Hong Kong University of Science and Technology), Jiaxi Ying (Hong Kong University of Science and Technology)
Optimization
🎯 What it does: Propose a Jacobi preconditioned gradient descent (JHGD) method for reconstructing spectrally sparse signals through low-rank Hankel matrix completion.
FAST-AR: Fast Autoregressive Video Diffusion and World Models with Temporal Cache Compression and Sparse Attention
Dvir Samuel (OriginAI), Rami Ben-Ari (OriginAI)
GenerationComputational EfficiencyTransformerDiffusion modelWorld ModelVideo
🎯 What it does: This paper proposes a training-agnostic FAST-AR framework to accelerate the generation of autoregressive video diffusion and world models;
Fast-SAM3D: 3Dfy Anything in Images but Faster
Weilun Feng (Chinese Academy of Sciences), Zhulin An (Chinese Academy of Sciences)
GenerationDepth EstimationComputational EfficiencyTransformerDiffusion modelAuto EncoderContrastive LearningGaussian SplattingImagePoint CloudMeshBenchmark
🎯 What it does: By introducing a training-agnostic dynamic acceleration module into the SAM3D single-view 3D reconstruction pipeline, the inference latency is significantly reduced.
Faster Activation Functions at the Edge for Post-Training Speedups
Anton Lydike (University of Edinburgh), Jackson Woodruff (University of Edinburgh)
Anomaly DetectionComputational EfficiencyConvolutional Neural NetworkTransformerSupervised Fine-TuningAuto EncoderImageText
🎯 What it does: Propose the FFCC compiler, which automatically generates precise and efficient activation function approximations using floating-point bit operations, enabling efficient inference on edge devices without hardware acceleration.
Faster Query-Key Learning Sharpens Attention in Self-Attention Models
Rahul Vashisht (IIT Madras), Harish G. Ramaswamy (IIT Madras)
OptimizationExplainability and InterpretabilityTransformerPrompt EngineeringContrastive LearningText
🎯 What it does: Investigate the impact of the learning rate of query-key (QK) compared to the learning rate of output-value (OV) in single-layer Transformer self-attention on attention sharpening, and verify this relationship through theoretical closed-form analysis and experiments.
Faster Than Flash: Exploiting Attention Sparsity for Efficient Long-Context Decoding
Zhigeng Liu (Fudan University), Xipeng Qiu (Fudan University)
Computational EfficiencyTransformerLarge Language ModelText
🎯 What it does: Propose Faster Flash Decoding (FFD), achieving efficient acceleration of sparse attention during the long context decoding phase through hardware-software co-design.
FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive Models
Senmao Li (Nankai University), Yaxing Wang (Jilin University)
GenerationComputational EfficiencyTransformerPrompt EngineeringDiffusion modelImage
🎯 What it does: Propose a no-training, plug-and-play acceleration framework called FasterVAR, which accelerates the late-stage detail refinement phase in the high-resolution generation process of visual autoregressive (VAR) models.
FastSESR: Fast Scene-level Explicit Surface Reconstruction
Jueqi Liu (Beijing Jiaotong University), Congyan Lang (Beijing Jiaotong University)
GenerationComputational EfficiencyGraph Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningGaussian SplattingPoint CloudMesh
🎯 What it does: Proposes a two-stage explicit surface reconstruction framework called FastSESR, which utilizes a Triangle Candidate Network to predict local triangle probabilities and learns point displacements through an Offset Optimization Network in a single forward inference, thereby directly generating high-quality triangular meshes.
Faults in Our Formal Benchmarking: Dataset Defects and Evaluation Failures in Lean Theorem Proving
Pawan Sasanka Ammanamanchi (Eleuther AI), Stella Biderman (Eleuther AI)
Explainability and InterpretabilityData-Centric LearningLarge Language ModelSupervised Fine-TuningTextBenchmark
🎯 What it does: Audited and exposed dataset defects and evaluation failures in the Lean formal theorem proving benchmark;
Feasible Fusion: Constrained Joint Estimation under Structural Non-Overlap
Yuxi Du (Shanghai University of Finance and Economics), Jiecheng Guo (Didi Chuxing)
Domain AdaptationOptimizationFederated LearningRepresentation LearningAuto EncoderContrastive LearningTabularTime SeriesBenchmark
🎯 What it does: Propose a constrained joint estimation framework that combines the predictive accuracy of observational data with causal identification constraints from randomized experiments under the condition of structured non-overlap (Conditional Structural Non-Overlap).
Feature Bagging Provides Stability
Yuheng Ma (East China Normal University), Qiang Sun (University of Toronto)
Explainability and InterpretabilityComputational EfficiencyData-Centric LearningImageTextTabularBenchmark
🎯 What it does: Studied the performance of feature bagging in terms of algorithm stability, proposed feature instability (FI) as a measure of instability along the feature axis, and analyzed the effects of feature bagging in parametric linear models and model-free settings.
Feature Collapse Under Corruption: An Entropy Perspective on Robust Neural Networks
Vishesh Kumar (Indian Institute of Science Education and Research Bhopal), Akshay Agarwal (Indian Institute of Science Education and Research Bhopal)
ClassificationKnowledge DistillationRepresentation LearningAdversarial AttackConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningImage
🎯 What it does: Propose a fine-tuning framework based on entropy called Dem-HEC, which generates high-entropy samples by maximizing the model output entropy within a restricted perturbation space, and combines cross-entropy, contrastive learning, and knowledge distillation to enhance the model's robustness against natural noise/distortion corruption, while maintaining or improving the accuracy on clean images.
Feature Resemblance: Towards a Theoretical Understanding of Analogical Reasoning in Transformers
Ruichen Xu (Chinese University of Hong Kong), Ying-Jun Angela Zhang (Chinese University of Hong Kong)
Explainability and InterpretabilityRepresentation LearningTransformerContrastive LearningText
🎯 What it does: Investigate the formation mechanism of analogical reasoning in Transformer models
Feature-Aware (Hyper)graph Generation via Next-Scale Prediction
Dorian Gailhard (Télécom Paris), Jhony H. Giraldo (Télécom Paris)
GenerationData SynthesisGraph Neural NetworkTransformerDiffusion modelScore-based ModelFlow-based ModelAuto EncoderGenerative Adversarial NetworkPoint CloudMeshGraph
🎯 What it does: Propose FAHNES, a scalable hierarchical generative framework capable of simultaneously generating the topological structures and node/edge features of graphs/hypergraphs.
FedARC: Anchor-Guided Residual Compensation for Data and Model Heterogeneous Federated Learning
Chentao Lu (Changchun University), Liehuang Zhu (Beijing Institute of Technology)
ClassificationFederated LearningConvolutional Neural NetworkTransformerContrastive LearningImage
🎯 What it does: Propose the FedARC framework to address the feature mismatch problem caused by data and model heterogeneity in federated learning.
FedCDWA: Decoupled Federated Prototype Distillation with Hierarchical Wasserstein Aggregation
Zhenshen Liu (Xidian University), Yintang Yang (Xidian University)
ClassificationFederated LearningKnowledge DistillationConvolutional Neural NetworkContrastive LearningImage
🎯 What it does: Propose the FedCDWA framework, which achieves prototype-level knowledge sharing in non-IID federated learning by decoupling client-side personalization distillation from server-side mutual distillation.
FedEBA+: Towards Fair and Effective Federated Learning via Entropy-Based Model
Zhichao Wang (Chinese University of Hong Kong), Xiaoying Tang (Chinese University of Hong Kong)
OptimizationFederated LearningComputational EfficiencyContrastive LearningImage
🎯 What it does: Propose FedEBA+, a method that simultaneously improves fairness and the accuracy of the global model in federated learning.
FedEMoE: Improving Personalization on Heterogeneous Federated Learning via Elastic Mixture of Experts Architecture
Haizhou Du (Shanghai University Of Electric Power), Huan Huo (University Of Technology Sydney)
ClassificationFederated LearningKnowledge DistillationMixture of ExpertsContrastive LearningImage
🎯 What it does: This paper proposes FedEMoE, a flexible expert mixture architecture that decouples personalization and generalization in heterogeneous federated learning.
Federated Bilevel Performative Prediction
Liangxin Qian (Nanyang Technological University), Kwok-Yan Lam (Nanyang Technological University)
OptimizationFederated LearningConvolutional Neural NetworkContrastive LearningImageTextTabular
🎯 What it does: Proposed a federated bilevel executable feasibility prediction framework, and defined the federated bilevel feasibility stable point (FBPS), providing theoretical analysis on its existence, uniqueness, and convergence properties.
Federated Causal Inference on Multi-Site Observational Data via Propensity Score Aggregation
Rémi Khellaf (Inria, Inserm, Idesp, Université de Montpellier), Julie Josse (Inria, Inserm, Idesp, Université de Montpellier)
Federated LearningExplainability and InterpretabilityComputational EfficiencyTabularBiomedical DataElectronic Health Records
🎯 What it does: Estimate the average treatment effect (ATE) on multi-site distributed observational data using federated learning, and propose a propensity score aggregation method based on member weights.
Federated Data and Feature Selection by Generalized CUR Decomposition
Ying-Peng Tang (Nanyang Technological University), Han Yu (Nanyang Technological University)
CompressionOptimizationFederated LearningSafty and PrivacyAuto EncoderContrastive LearningImageTabular
🎯 What it does: Proposes the FedGCUR framework for unified data and feature selection in federated learning, based on the general CUR decomposition to achieve compression and reconstruction.
Federated Distillation for Whole Slide Image via Gaussian-Mixture Feature Alignment and Curriculum Integration
Luru Jing (Peking University), Yongzhi Cao (Peking University)
ClassificationFederated LearningSafty and PrivacyExplainability and InterpretabilityKnowledge DistillationData-Centric LearningTransformerDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBiomedical DataComputed Tomography
🎯 What it does: Propose FedHD, a federated distillation-based whole slide image (WSI) learning framework, which generates synthetic features by aligning local Gaussian Mixture features, and achieves collaborative training under heterogeneous models and without data sharing through curriculum fusion.
Federated Graph Learning via Structure-Aware Fusion Using a Kalman Framework with Learnable Dynamics
Bisheng Tang (Shaoyang University), Xiaojun Chen (Chinese Academy of Sciences)
ClassificationFederated LearningGraph Neural NetworkContrastive LearningGraph
🎯 What it does: Proposes a federated graph learning framework called Fed-Kalter, which first separates structural information from node features, uses a Kalman-style Kalter-Conv to filter feature noise, and then aggregates only the structural module parameters on the federated server, achieving cross-client structural knowledge sharing.
Federated Learning with Unlabeled Clients: Personalization Can Happen in Low Dimensions
Hossein Zakerinia (Institute of Science and Technology Austria), Christoph H. Lampert (Institute of Science and Technology Austria)
ClassificationFederated LearningTransformerAuto EncoderContrastive LearningImageTabular
🎯 What it does: Propose the FLowDUP method, which can generate a personalized model by performing only one forward pass using unlabeled data from clients.
Federated Manifold Learning (FML): Tackling Domain Heterogeneity with Structural Knowledge Transfer
Xutong Mu (Xidian University), Yulong Shen (Xidian University)
CompressionDomain AdaptationFederated LearningContrastive LearningImage
🎯 What it does: Propose a structured knowledge transfer method in federated learning that utilizes compressed sensing subspace learning to address domain heterogeneity issues.
Federated Multi-view Clustering for Remote Sensing Data
Renxiang Guan (National University of Defense Technology), Yuhua Tang (Harbin University of Science and Technology)
Federated LearningSafty and PrivacyComputational EfficiencyRepresentation LearningHyperparameter SearchGraph Neural NetworkAuto EncoderContrastive LearningImageMultimodalityPoint CloudGraph
🎯 What it does: A federated deep clustering framework called FedRSMVC for remote sensing multi-view data is proposed, which avoids transmitting raw data and only shares noisy clustering prototypes and soft labels, achieving privacy protection and efficient communication.
Federated Sketching LoRA: A Flexible Framework for Heterogeneous Collaborative Fine-Tuning of LLMs
Wenzhi Fang (Purdue University), Christopher Brinton
OptimizationFederated LearningComputational EfficiencyTransformerLarge Language ModelSupervised Fine-TuningText
🎯 What it does: Propose FSLoRA — a LoRA fine-tuning method using sketching within a federated learning framework, allowing clients with different resources to update only the submatrix of the global LoRA module, thereby achieving efficient heterogeneous collaborative fine-tuning.
Federated Variational Preference Alignment with Gumbel-Softmax Prior for Personalized User Preferences
Jabin Koo (POSTECH), Jungseul Ok (POSTECH)
Recommendation SystemFederated LearningReinforcement Learning from Human FeedbackTransformerContrastive LearningText
🎯 What it does: The study proposes a federated variational preference alignment framework, FedVPA-GP, which utilizes Gumbel-Softmax priors to achieve personalized user preference alignment.
FedFit: Federated Dynamic Sparse Training via Fisher Information scoring
Meng Bi (Mohamed bin Zayed University of Artificial Intelligence), Xue Liu (Mohamed bin Zayed University of Artificial Intelligence)
OptimizationFederated LearningComputational EfficiencyTransformerContrastive LearningImageText
🎯 What it does: Address the insufficiency of structural adjustment in federated dynamic sparse training, proposing the FedFit framework.
FedGain: Toward Negative-Gain-Free Client Collaboration in Federated Learning
Yuqing Zhang (Huaqiao University), Hui Tian (Huaqiao University)
OptimizationFederated LearningData-Centric LearningContrastive LearningImageTabular
🎯 What it does: Propose the FedGain framework, which eliminates the negative gain problem in federated learning by optimizing client clustering.
FedHera: Towards Drift-Resilient Federated Fine-tuning with Heterogeneous Resources
Ke Xiao (University of Glasgow), Wenhao Li (University of Glasgow)
Federated LearningTransformerLarge Language ModelSupervised Fine-TuningText
🎯 What it does: This paper studies federated fine-tuning on resource-heterogeneous edge devices and proposes the FEDHERA framework to achieve scalable federated adaptation.
FedHPro: Federated Hyper-Prototype Learning via Gradient Matching
Huan Wang (University of Wollongong), Guansong Pang (Singapore Management University)
Federated LearningExplainability and InterpretabilityRepresentation LearningContrastive LearningImageMultimodalityTabular
🎯 What it does: Propose the FedHPro framework, which learns interpretable hyper-prototypes through gradient matching in federated learning, and improves model generalization by using two contrastive and alignment mechanisms, HPCL and HPAL, during local training.
FedPAT: Federated Test-Time Adaptation via Prototype Affinity Topology
Shunxin Guo (Southeast University), Xin Geng (Southeast University)
ClassificationDomain AdaptationFederated LearningConvolutional Neural NetworkTransformerReinforcement LearningPrompt EngineeringDiffusion modelScore-based ModelAuto EncoderContrastive LearningOptical FlowImageBenchmark
🎯 What it does: Propose FedPAT, which implements federated test-time adaptation (Federated Test-Time Adaptation) based on prototype affinity topology (PAT). It constructs a global PAT by aggregating category prototypes from source clients, and achieves unsupervised online adaptation at the target client through topological diffusion and parameterized classifier fusion.
FedPDG: Prediction Discrepancy–Guided Data Generation for Heterogeneous Federated Learning
Yuqi Wang (Beihang University), Xin Hao (Beihang University)
GenerationData SynthesisFederated LearningTransformerDiffusion modelScore-based ModelImage
🎯 What it does: The paper proposes the FedPDG framework, which utilizes prediction differences to guide data generation in federated learning, in order to bridge the gap between the local data distribution on clients and the globally imbalanced distribution.
FedPissa: Towards Federated Personalized Adaptation of Foundation Models via LoRA Subspace Mapping
Wenwen He (Wuhan University), Mang Ye (Wuhan University)
Federated LearningComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningImageText
🎯 What it does: Propose the FedPissa framework, which utilizes a single LoRA model to achieve federated personalization by selectively aggregating the B matrix and constructing a decorrelated subspace projection.
FedQueue: Queue-Aware Federated Learning for Cross-Facility HPC Training
Yijiang Li (Argonne National Laboratory), Kibaek Kim (Argonne National Laboratory)
OptimizationFederated LearningComputational EfficiencyTransformerSupervised Fine-TuningAgentic AIContrastive LearningTextTabular
🎯 What it does: Propose the FedQueue algorithm, addressing the random enqueue delay caused by batch scheduling in cross-HPC facility training, by constructing a complete federated learning framework with queue prediction, cutoff enqueue control, and delay-aware aggregation.
FedReLa: Imbalanced Federated Learning via Re-Labeling
Guangzheng Hu (University of Melbourne), Liuhua Peng (University of Melbourne)
ClassificationFederated LearningSupervised Fine-TuningContrastive LearningImage
🎯 What it does: A data re-labeling based scheme called FedReLa is proposed in federated learning to simultaneously address data heterogeneity and global class imbalance issues.
FedRGL: Robust Federated Graph Learning under Label Noise
De Li (Guangxi Normal University), Xianxian LI (Guangxi Normal University)
ClassificationFederated LearningGraph Neural NetworkAuto EncoderContrastive LearningGraph
🎯 What it does: Propose the FedRGL method to address the label noise problem in federated subgraph node classification.
FedRot-LoRA: Mitigating Rotational Misalignment in Federated LoRA
Haoran Zhang (University of Texas at Austin), Haris Vikalo (University of Texas at Austin)
Federated LearningComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningText
🎯 What it does: Parameter-efficient fine-tuning of large language models in federated learning, with an analysis of aggregation errors in the low-rank factors of the LoRA method and the proposal of a rotation alignment mechanism;
FedScar: Correcting Geometric Bias for Flatness-Consistent Federated Learning
Jianfeng Lu (Wuhan University of Science and Technology), Guanghui Wen (Southeast University)
OptimizationFederated LearningComputational EfficiencyContrastive LearningImage
🎯 What it does: Propose the FedScar framework, which explicitly corrects geometric bias caused by data heterogeneity in federated learning, thereby improving the generalization performance of the global model.
FedSDR: Federated Self-Distillation with Rectification
Ziheng Ren (Beijing University of Aeronautics and Astronautics), You Song (Beijing University of Aeronautics and Astronautics)
Federated LearningKnowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextRetrieval-Augmented Generation
🎯 What it does: Propose Federated Self-Distillation (FedSD) and its improved version FedSDR, which utilize self-distillation to project raw data from each client into a unified model understanding space, thereby alleviating model drift caused by statistical heterogeneity in federated learning. Additionally, it addresses factual bias and redundancy caused by the 'Rewrite Paradox' through a dual-stream LoRA architecture and selective aggregation.
FedSSM: State Space Model-based Proactive Inference for Heterogeneous Multimodal Federated Learning
Hengyi Ren (Nanjing Forestry University), Lijuan Sun (Nanjing University of Posts and Telecommunications)
OptimizationFederated LearningComputational EfficiencyRepresentation LearningRecurrent Neural NetworkTransformerAgentic AIMixture of ExpertsContrastive LearningWorld ModelImageVideoTextMultimodality
🎯 What it does: Propose the FedSSM framework, which transforms client selection into active inference, utilizing decision-aware state space models to predict training dynamics and perform counterfactual selection, combined with surprise-driven adaptive budgeting and trust-weighted aggregation.
FedTreeLoRA: Reconciling Statistical and Functional Heterogeneity in Federated LoRA Fine-Tuning
Jieming Bian (University of Florida), Jie Xu (University of Florida)
Federated LearningComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningAgentic AIPrompt EngineeringTextBenchmark
🎯 What it does: Perform low-rank adaptation (LoRA) fine-tuning of large language models in federated learning, and propose the FedTreeLoRA framework, which achieves aggregation through a tree structure;
FedUSD: Unbiased Synthetic Data for Federated Learning
Weiying Xie (Xidian University), Yunsong Li (Xidian University)
Data SynthesisOptimizationFederated LearningConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelFlow-based ModelRectified FlowAuto EncoderGenerative Adversarial NetworkContrastive LearningImage
🎯 What it does: Proposes FedUSD, a free aggregation federated learning method that achieves unbiased synthetic data optimization by matching high-energy orthogonal bases (HOB) and variance, eliminating performance degradation caused by client data heterogeneity.
FedVeer: Self-Adaptive Skew Estimation for Robust Federated Learning
Yun Xin (Wuhan University of Science and Technology), Guanghui Wen (Southeast University)
Federated LearningContrastive LearningImageTabular
🎯 What it does: Proposes FedVeer, a federated learning framework based on adaptive kernel density estimation and Kalman filtering, aimed at mitigating the impact of data skew and noise.
Feedback Control for Multi-Objective Graph Self-Supervision
Karish Grover (Carnegie Mellon University), Christos Faloutsos (Carnegie Mellon University)
OptimizationRepresentation LearningReinforcement Learning from Human FeedbackGraph Neural NetworkReinforcement LearningContrastive LearningGraph
🎯 What it does: Propose the ControlG framework, which treats multi-objective graph self-supervised learning as a closed-loop scheduling problem, allocating the optimization budget to each pre-training task over time.
FeRA: Frequency-Energy Constrained Routing for Effective Diffusion Adaptation Fine-Tuning
Bo Yin (National University Of Singapore), Shuicheng YAN
GenerationData SynthesisExplainability and InterpretabilityComputational EfficiencyTransformerMixture of ExpertsDiffusion modelScore-based ModelContrastive LearningImage
🎯 What it does: Designed and verified a parameter-efficient fine-tuning framework called FeRA based on frequency energy, which can dynamically adapt the diffusion process at different stages while preserving the original model's prior knowledge.
Few-Shot Design Optimization by Exploiting Auxiliary Information
Arjun Mani (Columbia University), Richard Zemel (Columbia University)
OptimizationHyperparameter SearchRobotic IntelligenceMeta LearningTransformerPoint CloudMeshGraphTabularTime Series
🎯 What it does: Proposes a method for few-shot design optimization in black-box optimization that leverages high-dimensional auxiliary information, enabling rapid prediction of target performance and accelerating the search process with limited experiments;
FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model
Chunyu Xie (Beihang University), Yuhui Yin (360 AI Research)
ClassificationObject DetectionSegmentationRetrievalTransformerPrompt EngineeringMixture of ExpertsVision Language ModelContrastive LearningImageTextMultimodality
🎯 What it does: Propose FG-CLIP 2, which constructs an English-Chinese bilingual fine-grained vision-language alignment model, achieving precise alignment through two-stage training and multi-task loss.
FIBER: A Differentially Private Optimizer with Filter-Aware Innovation Bias Correction
Duc Dm (Korea Advanced Institute of Science and Technology), Huy Nguyen
OptimizationSafty and PrivacyContrastive LearningGaussian SplattingImageText
🎯 What it does: This paper proposes a differential privacy optimizer named FIBER, which significantly improves model performance during long-term training by utilizing innovative spatial filtering and filter-aware second-moment correction.
FIDIA: Function-Informed Sequence Design via Inference-Aligned Policy Optimization
Minghan Li (Beihang University), Yue Deng (Beihang University)
Drug DiscoveryReinforcement Learning from Human FeedbackProtein Structure PredictionTransformerReinforcement LearningPrompt EngineeringScore-based ModelSequentialBiomedical Data
🎯 What it does: Developed a reinforcement learning framework called FIDIA, which aligns the Best-of-N reasoning strategy with training objectives in protein sequence design, enabling function-driven sequence generation.
FiGuRO: Intrinsic Dimension Estimation for Multi-Modal Data
Viktoria Schuster (Massachusetts Institute of Technology), Caroline Uhler (Massachusetts Institute of Technology)
OptimizationRepresentation LearningDiffusion modelScore-based ModelFlow-based ModelRectified FlowAuto EncoderContrastive LearningImageMultimodalityPoint CloudAudio
🎯 What it does: Propose the FiGuRO method, which combines a low-rank adaptive layer with rate-distortion theory to dynamically estimate the intrinsic dimension of multi-modal data, and achieves spontaneous separation in shared and private subspaces.
Find, Fix, Reason: Context Repair for Video Reasoning
Haojian Huang (Hong Kong University of Science and Technology), Ying-Cong Chen
Computational EfficiencyKnowledge DistillationTransformerReinforcement LearningPrompt EngineeringVision Language ModelVideoTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This paper proposes a video reasoning training framework named Find, Fix, Reason (FFR), which utilizes frozen tool-integrated teachers to perform minimal context repair on failed reasoning trajectories. Subsequently, students re-reason with supplemented evidence, thereby improving exploration efficiency and accuracy.
Finding Differentially Private Second Order Stationary Points in Stochastic Minimax Optimization
Difei Xu (King Abdullah University of Science and Technology), Di Wang (King Abdullah University of Science and Technology)
OptimizationSafty and PrivacyTabular
🎯 What it does: This paper proposes an algorithm based on differential privacy that can find approximate second-order stationary points (SOSP) in non-convex-strongly concave stochastic adversarial optimization problems.
Finding DoRI: Discovery of Retained Images in Diffusion Models
Antoni Kowalczuk (CISPA Helmholtz Center for Information Security), Franziska Boenisch (CISPA Helmholtz Center for Information Security)
GenerationData SynthesisSafty and PrivacyAdversarial AttackTransformerPrompt EngineeringDiffusion modelContrastive LearningImageTextMultimodality
🎯 What it does: The study finds that traditional weight pruning methods can only hide the memory mapping from text to images, but cannot truly eliminate the memory images, and proposes a technique to discover and remove retained images by adversarial text embedding.
Finding Most Influential Sets
Lucas Darius Konrad, Nikolas Kuschnig (Monash University)
OptimizationExplainability and InterpretabilityComputational EfficiencyTextTabularBiomedical Data
🎯 What it does: Proposed an efficient algorithm for finding the most influential subset (MIS) in a dataset. The algorithm transforms the combination search into multiple top-k selections by expressing the holdout effect as a linear fractional form, ultimately achieving exact solutions.
Finding Stationary Points by Comparisons
Helin Wang (Peking University), Tongyang Li (Peking University)
Optimization
🎯 What it does: Propose an algorithm that can access an ε-stationary point after approximately ˜O(n / ε^{5/2}) comparison queries in an environment where only the size information of function values can be obtained through comparison oracles; also provide its quantum version, which only requires ˜O(n / ε^{5/2}) queries using quantum comparison oracles.
Finding the Correct Visual Evidence Without Forgetting: Mitigating Hallucination in LVLMs via Inter-Layer Visual Attention Discrepancy
Yutong Xie (Southeast University), Yuheng Jia (Southeast University)
Explainability and InterpretabilityComputational EfficiencyTransformerVision Language ModelImageTextMultimodalityBenchmark
🎯 What it does: Propose a no-training, plug-and-play method called ILVAD, which utilizes inter-layer visual attention differences to identify and reinforce correct visual evidence, thereby reducing hallucinations generated by large audio-visual models.
Finding the Minimal Parameter Budget for Implicit Reasoning: A Data Complexity Driven Scaling Law for Language Models
Xinyi Wang (Princeton Language and Intelligence Lab), Yikang Shen (MIT-IBM Watson AI Lab)
Computational EfficiencyRepresentation LearningData-Centric LearningGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringTextGraph
🎯 What it does: This paper studies the minimum parameter budget required for language models to perform implicit reasoning during the pre-training phase.
Fine-grained Analysis of Brain-LLM Alignment through Input Attribution
Michela Proietti (Goethe University), Mariya Toneva (Max Planck Institute for Software Systems)
Explainability and InterpretabilityTransformerLarge Language ModelTextBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Propose a pipeline that integrates input attribution methods (such as Integrated Gradients, Gradient×Input, SmoothGrad) into a brain-LLM alignment framework, used to identify the input words most important for predicting brain activity.
Fine-to-Coarse Fairness-Informed Multi-View Clustering
Shengju Yu (Hong Kong Baptist University), Yiu-ming Cheung (Hong Kong Baptist University)
OptimizationFederated LearningComputational EfficiencyRepresentation LearningContrastive LearningMultimodalityBenchmark
🎯 What it does: Propose a fairness-guided multi-view clustering method called FCFMVC, which improves anchor point allocation to match the sizes and tightness of different clusters.
Fine-Tune Once, Reuse Across Models: Bayesian Task-Update Factors and Approximations
Siyang Guo (Sun Yat-sen University), Zibin Zheng (Sun Yat-sen University)
ClassificationRecommendation SystemOptimizationKnowledge DistillationMeta LearningTransformerSupervised Fine-TuningContrastive LearningImageTextMultimodality
🎯 What it does: This paper proposes a task update factor that can be reused across different base models, enabling the achievement of the same task performance on new models with just a single fine-tuning step;
Fine-Tuning Masked Diffusion for Provable Self-Correction
Jaeyeon Kim (Harvard University), Sitan Chen (Harvard University)
GenerationData SynthesisAI Code AssistantTransformerSupervised Fine-TuningDiffusion modelTextSequential
🎯 What it does: Lightweight fine-tuning is performed on the pre-trained Masked Diffusion Model (MDM), incorporating self-correction capabilities, leading to the proposal of the PRISM method.
Fine-Tuning of Transformer models with Frames
Harshavardhan Adepu (University of Wisconsin-Madison), Vikas Singh (University of Wisconsin-Madison)
OptimizationComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningAuto EncoderContrastive LearningImageText
🎯 What it does: A parameter-efficient fine-tuning framework called FrameFT based on fusion frames is proposed. It reparameterizes the Transformer weight updates using sparse subspace coefficients, achieving task adaptation with extremely few learnable parameters while keeping the pre-trained weights frozen.
Fine-Tuning Without Forgetting In-Context Learning: A Theoretical Analysis of Linear Attention Models
Chungpa Lee (Yonsei University), Kangwook Lee (University of WisconsinMadison)
Computational EfficiencyKnowledge DistillationRepresentation LearningMeta LearningTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningTextBenchmark
🎯 What it does: This paper investigates the impact of fine-tuning on the original in-context learning ability in large language models when using linear attention models, through theoretical derivation and experimental validation.
Fingerprinting Pre-trained Encoders under Arbitrary Downstream Fine-Tuning via Adversarial Shifting
Tianlong Xu (Huazhong University of Science and Technology), Xiaoyi Fan (Jiangxing Intelligence Technology Inc)
Federated LearningSafty and PrivacyRepresentation LearningAdversarial AttackConvolutional Neural NetworkTransformerContrastive LearningImageText
🎯 What it does: This paper proposes a fingerprinting method for pre-trained encoders, utilizing adversarial shifting to construct stable fingerprint clusters in the encoder's feature space, and achieving downstream task-agnostic black-box ownership verification through clustering voting.
Finite and Corruption-Robust Regret Bounds in Online Inverse Linear Optimization under M-Convex Action Sets
Taihei Oki (Institute for Chemical Reaction Design and Discovery, Hokkaido University), Shinsaku Sakaue (CyberAgent)
OptimizationReinforcement Learning
🎯 What it does: Study the finite and robust return upper bounds of online inverse linear optimization under M-convex (including base graph) action sets.
Finite-time Convergence Analysis of Actor-Critic with Evolving Reward
Rui Hu (Tsinghua University), Longbo Huang (Tsinghua University)
Reinforcement Learning
🎯 What it does: This paper provides the first finite-time convergence analysis of single-time-scale actor-critic algorithms in scenarios where the reward evolves over time and Markov sampling is used;
Finite-Width Neural Tangent Kernels from Feynman Diagrams
Max Guillen (Chalmers University of Technology and University of Gothenburg), Jan E Gerken (Chalmers University of Technology and University of Gothenburg)
Symbolic ComputationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningConvolutional Neural NetworkRecurrent Neural NetworkGaussian SplattingTabularTime SeriesPhysics Related
🎯 What it does: A framework based on Feynman diagrams is proposed to compute the statistics of the neural tangent kernel (NTK) and its higher-order derivatives under finite width, along with deriving the corresponding layer recurrence relations.
FIPN: Forward Self-Organizing Interpretable Polynomial Networks for Time Series Forecasting
YiZhen Wang, Zunwei Fu (University of Suwon)
Explainability and InterpretabilityComputational EfficiencyTime SeriesBenchmark
🎯 What it does: Studied a forward self-organizing interpretable polynomial network FIPN without backpropagation, used for long-term time series prediction.
FIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights
Zhen Wang (University of California San Diego), Eric P. Xing (MBZUAI)
TransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposes FIRE-Bench, a comprehensive re-mining evaluation benchmark based on verified scientific discoveries, aimed at assessing the scientific reasoning and experimental capabilities of LLM-driven autonomous research agents.
FiRE: Fine-grained Ranking Evaluation for Machine Translation
Wenyang Gao (Zhejiang University), Yue Zhang (Westlake University)
Explainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringTextBenchmark
🎯 What it does: Propose the FiRE method, which uses LLM to perform fine-grained pairwise ranking evaluation of MT outputs without reference, and constructs the first human-annotated fine-grained benchmark.
FIRE: Learning to Navigate and Act on Real-World Files via Stateful Reinforcement Learning
Jingyuan Ma (Peking University), Zhifang Sui (Peking University)
AI Code AssistantTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextMultimodalityTabularBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This paper proposes the File Reasoning task, placing LLM directly within an executable sandbox, allowing the model to perform interactive reasoning on unprocessed XLSX, PDF, DOCX, and PPTX files;
FIRE: Multi-Fidelity Regression with Distribution-Conditioned In-Context Learning Using Tabular Foundation Models
Rosen Ting-Ying Yu (Massachusetts Institute of Technology), Faez Ahmed (Massachusetts Institute of Technology)
OptimizationHyperparameter SearchData-Centric LearningTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsTabularTime SeriesSequentialBenchmarkPhysics Related
🎯 What it does: Designed a training-agnostic multi-talented regression framework called FIRE, which leverages the TabPFN (table foundation model) to achieve zero-shot Bayesian inference for low-resolution models and high-resolution residual correction through distribution-conditional context learning.
FiSeR: Fine-Grained Source Representations for Cross-Domain AI Image Detection
Shan Zhang (Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences), Lei Ma (Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences)
Domain AdaptationAnomaly DetectionTransformerContrastive LearningImage
🎯 What it does: Designed a hierarchical supervised contrastive learning framework called FiSeR, aiming to enhance the generalization ability of AI-generated image detection in cross-domain scenarios.
Fisher-Preserving Guidance: Training-Free Manifold Constraints for Safe Diffusion Control
Hao Ren (Sun Yat-sen University), Hui Cheng (Sun Yat-sen University)
Autonomous DrivingSafty and PrivacyComputational EfficiencyRobotic IntelligenceTransformerDiffusion modelScore-based ModelContrastive LearningImageVideoPoint Cloud
🎯 What it does: Proposed a training-free, no additional training cost Fisher preserving guidance (FPG-OPS), which keeps the sampling trajectory on the training data manifold by projecting onto the Fisher iso-surface during the diffusion reverse sampling process, while combining task guidance;
Fix Before Search: Benchmarking Agentic Visual Query Pre-processing in Multimodal Retrieval-augmented Generation
Shenglai Zeng (Michigan State University), Yi Chang (Jilin University)
Data SynthesisRetrievalTransformerLarge Language ModelSupervised Fine-TuningAgentic AIVision Language ModelImageTextMultimodalityBenchmarkRetrieval-Augmented Generation
🎯 What it does: This paper proposes V-QPP-Bench, a benchmark for visual query preprocessing in multi-modal retrieval-augmented generation (MRAG), and systematically evaluates the performance of different MLLMs on this benchmark.
Fix the Loss, Not the Radius: Rethinking the Adversarial Perturbation of Sharpness-Aware Minimization
Jinping Wang (Wenzhou-Kean University), Zhiqiang Gao (Wenzhou-Kean University)
ClassificationDomain AdaptationOptimizationAdversarial AttackConvolutional Neural NetworkTransformerSupervised Fine-TuningDiffusion modelContrastive LearningImageMultimodalityTabular
🎯 What it does: This paper re-examines the mechanism of Sharpness-Aware Minimization (SAM), discovering that its fixed-radius first-order linear approximation leads to a mismatch between the gradient norm-dominated learning signal and the second-order nature of flat minima. It proposes Loss-Equated SAM (LE-SAM), which eliminates the interference of gradient norms by fixing a budget in the loss space and solving for the corresponding radius in the parameter space, shifting the optimization focus to curvature information. Additionally, radius clipping and loss budget annealing are introduced to ensure stability, and further, LE-SAM+ is proposed to enhance curvature-aware regularization. Experimental results verify that this mechanism significantly improves generalization performance across various tasks and models.
Fix the Mind, Not the Move: Interpretable AI Assistance via Knowledge-Gap Localization
Ayano Hiranaka (University of Southern California), Daniel Seita (University of Southern California)
Explainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringTextSequentialRetrieval-Augmented Generation
🎯 What it does: Proposes the SENSEI framework, which can identify and correct users' cognitive biases in long-term decision-making tasks based on expert knowledge and user behavior, thereby improving subsequent task performance;
FiX: Introducing Fine-grained Forget Gate into Softmax Attention
Runzhong Li (Southern University of Science and Technology), Bo Tang (Southern University of Science and Technology)
Computational EfficiencyRepresentation LearningData-Centric LearningTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsText
🎯 What it does: Propose Fine-grained Forgetting Transformer (FiX), introducing an element-wise forgetting gate into softmax attention to enhance the modeling capability for long-text contexts.
Fixed Aggregation Features Can Rival GNNs
Celia Rubio-Madrigal (CISPA Helmholtz Center for Information Security), Rebekka Burkholz (CISPA Helmholtz Center for Information Security)
ClassificationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerAuto EncoderContrastive LearningGraphTabular
🎯 What it does: Propose Fixed Aggregation Features (FAFs), which transform graph node neighborhood information into tabular features through fixed aggregation functions (such as mean, sum, max, min, std, etc.), followed by using MLP for node classification.
Fixed Budget is No Harder Than Fixed Confidence in Best-Arm Identification up to Logarithmic Factors
Kapilan Balagopalan (University of Arizona), Kwang-Sung Jun (POSTECH)
OptimizationMeta LearningReinforcement Learning from Human FeedbackTabularStochastic Differential Equation
🎯 What it does: Propose a meta-algorithm called FC2FB that converts a fixed-confidence algorithm into a fixed-budget algorithm, and prove that in terms of sample complexity, FB is no harder than FC (only differing by a logarithmic factor).
FLAC: Maximum Entropy RL via Kinetic Energy Regularized Bridge Matching
Lei Lv (Shanghai Research Institute for Intelligent Autonomous Systems), Xiao Ma (ByteDance Seed)
Reinforcement LearningScore-based ModelTabularTime SeriesSequentialStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Propose the FLAC framework, which transforms maximum entropy reinforcement learning into a Generalized Schrödinger Bridge, achieving maximum entropy RL without likelihood by using kinetic regularization in an iterative policy generation strategy.
FLAG: Foundation model representation with Latent diffusion Alignment via Graph for spatial gene expression prediction
Qi Si (Shanghai Academy Of Artificial Intelligence For Science), Yuan Cheng (Fudan University)
Representation LearningData-Centric LearningGraph Neural NetworkTransformerDiffusion modelContrastive LearningImageBiomedical DataBenchmark
🎯 What it does: This paper proposes the FLAG framework, which predicts gene space expression from H&E slices by utilizing spatial graph encoding and a gene foundation model.
FLARE-AI: Flaw Reporting for AI
Shayne Longpre (Massachusetts Institute of Technology), Alex Pentland (Massachusetts Institute of Technology)
Federated LearningSafty and PrivacyExplainability and InterpretabilityTextBenchmark
🎯 What it does: Proposed and implemented an open-source AI incident reporting system called FLARE-AI, helping researchers submit AI incidents, vulnerabilities, and events in a unified and standardized manner.
Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization
Xiaoxuan He (Zhejiang University), Bohan Zhuang (Zhejiang University)
GenerationOptimizationComputational EfficiencyTransformerReinforcement LearningPrompt EngineeringDiffusion modelScore-based ModelVideoTextStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Propose Flash-GRPO, a single-step training framework for video diffusion model alignment, significantly reducing training costs while maintaining performance equivalent to full trajectory training;
Flash-VAED: Plug-and-Play VAE Decoders for Efficient Video Generation
Lunjie Zhu (Hong Kong University of Science and Technology), Jun Zhang (Hong Kong University of Science and Technology)
GenerationComputational EfficiencyKnowledge DistillationDiffusion modelAuto EncoderContrastive LearningVideoText
🎯 What it does: This study proposes Flash-VAED, an acceleration framework that can directly replace the existing VAE decoder, significantly improving the decoding speed of video generation.
FlashBlock: Attention Caching for Efficient Long-Context Block Diffusion
Zhuokun Chen (Monash University), Bohan Zhuang (Zhejiang University)
GenerationComputational EfficiencyKnowledge DistillationTransformerDiffusion modelScore-based ModelVideoText
🎯 What it does: Investigate the efficiency bottleneck of long context block diffusion models, proposing the FlashBlock mechanism, which reduces computation and KV cache access by caching and reusing cross-block attention.
FlashOptim: Optimizers for Memory-Efficient Training
Jose Javier Gonzalez Ortiz (Databricks AI Research), Davis Blalock (Databricks AI Research)
OptimizationComputational EfficiencyImageText
🎯 What it does: Propose FlashOptim, a memory compression technique tailored for commonly used optimizers (SGD, AdamW, Lion), significantly reducing the GPU memory consumption related to parameters;
FlashSinkhorn: IO-Aware Entropic Optimal Transport on GPU
Felix X.-F. Ye (University at Albany), Davis Wertheimer (IBM T. J. Watson Research Center)
OptimizationDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImagePoint CloudTabular
🎯 What it does: Implemented an IO-aware GPU implementation of Entropic Optimal Transport (EOT) solver FlashSinkhorn, leveraging FlashAttention-style streaming computation to transform Sinkhorn iterations into log-sum-exp (LogSumExp) plus bias dot product, enabling forward, backward, and Hessian-vector product computation without large memory consumption.
FlashSketch: Sketch-Kernel Co-Design for Fast Sparse Sketching on GPUs
Rajat Vadiraj Dwaraknath (Stanford University), Mert Pilanci (Stanford University)
OptimizationComputational EfficiencyImageTextPoint CloudGraphTabular
🎯 What it does: This paper designs a block-level structured sparse Johnson-Lindenstrauss transform (BLOCKPERM-SJLT), and based on it, implements an efficient GPU kernel (FLASHSKETCH), achieving fast execution of sparse Sketch;
Flat Minima and Generalization: Insights from Stochastic Convex Optimization
Matan Schliserman (Tel Aviv University), Tomer Koren (Google Research)
OptimizationContrastive LearningStochastic Differential Equation
🎯 What it does: This paper conducts theoretical analysis of three flat minima methods (SA-ERM, SA-GD, SAM) under the framework of convex smooth stochastic convex optimization (SCO), exploring the relationship between flatness and generalization performance.
FlatLab: A Unified Methodology Framework and Simulation-Based Benchmark for Robotic Manipulation of Flat Objects
Xingyu Zhu (Jilin University), Yixing Gao (Jilin University)
Depth EstimationRepresentation LearningRobotic IntelligenceConvolutional Neural NetworkTransformerReinforcement LearningVision-Language-Action ModelDiffusion modelContrastive LearningSimultaneous Localization and MappingImageMultimodalityPoint CloudBenchmark
🎯 What it does: Propose a unified robotic planar object grasping framework, divided into a policy generator and an action execution module, and build a simulation benchmark called FlatLab based on Isaac Sim;
FlatLand: Personalized Graph Federated Learning via Tailored Lorentz Space
Jiahong Liu (Chinese University of Hong Kong), Irwin King (Chinese University of Hong Kong)
Federated LearningRepresentation LearningGraph Neural NetworkContrastive LearningGraph
🎯 What it does: Proposes the FlatLand framework, which personalizes graph data learning by assigning adaptive Lorentz spaces (hyperbolic curvature) to each federated learning client, and designs a parameter separation strategy for temporal and spatial dimensions within this space;
Flatland: The Adventures of Gradient Descent with Large Step Sizes
Leonardo Galli (LMU Munich), Holger Rauhut (LMU Munich)
OptimizationData-Centric LearningConvolutional Neural NetworkRecurrent Neural NetworkTransformerContrastive LearningImage
🎯 What it does: This paper proposes a formal definition of 'maximum convergent step size' and designs a first-order algorithm based on equivalent line search (including backtracking and extrapolation), proving that these step sizes can enable gradient descent to enter the edge-of-stability (EoS) state early in training;
Flatness-Aware Stochastic Gradient Langevin Dynamics
Stefano Bruno (Ulsan National Institute of Science and Technology), Dongyoung Lim
OptimizationImageStochastic Differential Equation
🎯 What it does: Propose a novel first-order optimization algorithm called fSGLD, which introduces stochastic perturbations coupled with temperature in Langevin dynamics, maintaining the computational efficiency of SGD/SGLD while adaptively guiding the training process toward flat minima.
Fleet: Few Shots Lead Effective AI-generated Image Detection
Jiaan Wang (Institute of Computing Technology, Chinese Academy of Sciences), Sheng Tang (Institute of Computing Technology, Chinese Academy of Sciences)
Image TranslationGenerationDomain AdaptationAnomaly DetectionKnowledge DistillationRepresentation LearningMeta LearningTransformerMixture of ExpertsDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageBenchmark
🎯 What it does: Propose the Fleet framework, achieving dynamic adaptive AIGI detection based on subspace routing, enabling rapid adaptation to new generative models with very few samples.