ICML 2026 Papers — Page 22
International Conference on Machine Learning · 6554 papers
Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model
Xinyin Ma (National University Of Singapore), Xinchao Wang (National University Of Singapore)
GenerationData SynthesisKnowledge DistillationTransformerPrompt EngineeringDiffusion modelScore-based ModelVideoText
🎯 What it does: Propose Flex-Forcing, a unified training and inference framework that allows a single video diffusion model to flexibly switch between autoregressive and bidirectional modes during inference, and achieve arbitrary inference strategies through a variable blocking mechanism.
Flexibility-Aware Geometric Latent Diffusion for Full-Atom Peptide Design
Dongjiang Niu (Qingdao University), Zhen Li (Qingdao University)
Drug DiscoveryProtein Structure PredictionGraph Neural NetworkTransformerDiffusion modelAuto EncoderGenerative Adversarial NetworkGraphBiomedical Data
🎯 What it does: Propose a receptor-oriented, flexibility-aware geometric potential diffusion framework called PepFGLD that can capture peptide chain flexibility and generate peptide molecules at the all-atom level.
Flexible Kernels for Protein Property Prediction
Martin Jankowiak (Generate Biomedicines), Gevorg Grigoryan (Generate Biomedicines)
Drug DiscoveryProtein Structure PredictionConvolutional Neural NetworkGraph Neural NetworkContrastive LearningGaussian SplattingSequentialBiomedical Data
🎯 What it does: Proposed a class of sequence kernels (LOCK) that utilize evolutionary substitution matrices and local linear features, and embedded them into a Gaussian process model for protein attribute prediction.
FlexiFlow: decomposable flow matching for generation of flexible molecular ensemble
Riccardo Tedoldi (AstraZeneca), Alessandro Tibo (AstraZeneca)
Drug DiscoveryGraph Neural NetworkTransformerDiffusion modelScore-based ModelFlow-based ModelRectified FlowGraphBiomedical Data
🎯 What it does: Propose and implement FlexiFlow, which utilizes a decomposable flow matching framework to simultaneously generate molecular graphs and multiple low-energy conformation sets in a single sampling process.
FlexRank: Nested Low-Rank Knowledge Decomposition for Adaptive Model Deployment
Riccardo Zaccone (Polytechnic of Turin), Samuel Horváth (Mohamed bin Zayed University of Artificial Intelligence)
Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelDiffusion modelScore-based ModelContrastive LearningImageTextBenchmark
🎯 What it does: Propose the FlexRank method, which utilizes data-aware low-rank decomposition and dynamic programming to extract nested, adjustable submodels from pre-trained models, achieving one-time training for multiple deployments.
FLIP2: Expanding Protein Fitness Landscape Benchmarks for Real-World Machine Learning Applications
Kieran Didi (NVIDIA), Kevin K Yang
Drug DiscoveryTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningBiomedical DataBenchmark
🎯 What it does: This work proposes the FLIP2 benchmark, expanding the protein fitness prediction evaluation of FLIP by adding seven new multifunctional datasets and designing various generalization splits.
FLIPS: Instance-Fingerprinting for LLMs via Pseudo-random Sequences
Richardeau Gurvan (PEReN), Gilles Tredan (LAAS)
ClassificationFederated LearningSafty and PrivacyExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringContrastive LearningText
🎯 What it does: Propose instance-level fingerprinting and design the FLIPS method to distinguish different configuration instances of the same LLM in a black-box setting.
Float8@2bits: Entropy Coding Enables Data-Free Model Compression
Patrick Putzky (Merantix Momentum), Stefan Dietzel (Merantix Momentum)
CompressionComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelText
🎯 What it does: Proposes a post-training quantization framework called EntQuant, which achieves data-free model compression with extremely low bitrates (≈2bit/parameter) by utilizing entropy coding, while maintaining the functionality of large language models without the need for calibration data.
Floating-Point Networks with Automatic Differentiation Can Represent Almost All Floating-Point Functions and Their Gradients
Sejun Park (Korea University), Geonho Hwang (Gwangju Institute of Science and Technology)
OptimizationExplainability and InterpretabilityComputational EfficiencyAuto Encoder
🎯 What it does: This paper proves that under floating-point arithmetic and automatic differentiation, neural networks can approximate any floating-point function and its gradient.
FloorplanQA: A Benchmark for Spatial Reasoning in LLMs using Structured Representations
Fedor Rodionov (King Abdullah University of Science and Technology), Peter Wonka (King Abdullah University of Science and Technology)
TransformerLarge Language ModelPrompt EngineeringImageTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Constructed the FloorplanQA benchmark, using symbolized 2D indoor floorplans (JSON/XML) to evaluate the spatial reasoning capabilities of LLMs.
Flow Equivariant World Models: Structured Memory for Dynamic Environments
Hansen Lillemark, T. Anderson Keller (Harvard University)
Convolutional Neural NetworkRecurrent Neural NetworkTransformerDiffusion modelFlow-based ModelAuto EncoderWorld ModelImageVideo
🎯 What it does: Proposes the Flow Equivariant World Model (FloWM), which maintains the structure of dynamic environments in implicit memory by leveraging temporal symmetry, thereby achieving long-term stable dynamic prediction in partially observable environments.
Flow for Future: Geometric SE(3)-Equivariant Flow Matching for 3D Trajectory Prediction
Junwei Wu (Shandong University), Jian Sun (Xi'an Jiaotong University)
GenerationDrug DiscoveryProtein Structure PredictionGraph Neural NetworkTransformerDiffusion modelScore-based ModelFlow-based ModelRectified FlowPoint CloudGraphTime SeriesSequentialBenchmark
🎯 What it does: Proposes a geometric SE(3)-equivariant framework called GSE-Flow based on conditional flow matching for 3D trajectory prediction.
Flow Matching Calibration for Simulation-Based Inference under Model Misspecification
Pierre-Louis Ruhlmann (University Grenoble Alpes), Pedro L. C. Rodrigues (University Grenoble Alpes)
OptimizationData-Centric LearningTransformerFlow-based ModelTabularTime SeriesPhysics RelatedOrdinary Differential Equation
🎯 What it does: Propose a calibration framework based on flow matching called FMCPE, which uses a small number of real calibration samples to correct the posterior distribution of simulated inferences.
Flow Sampling : Learning to Sample from Unnormalized Densities via Denoising Conditional Processes
Aaron J Havens, Neta Shaul (Weizmann Institute of Science)
GenerationDrug DiscoveryTransformerDiffusion modelScore-based ModelFlow-based ModelImagePoint CloudGraphTabularBiomedical DataBenchmarkStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: This paper proposes a framework called Flow Sampling, which utilizes diffusion models and flow matching techniques to learn an approximate sampler from unnormalized densities (energy functions).
Flow-Based Density Ratio Estimation for Intractable Distributions with Applications in Genomics
Egor Antipov (Helmholtz Munich), Fabian J Theis
Drug DiscoveryScore-based ModelFlow-based ModelTabularBiomedical DataOrdinary Differential Equation
🎯 What it does: Propose a dynamic density ratio estimation method called scRatio based on conditional continuous normalizing flows (CNF), which directly tracks the logarithmic ratio on a single trajectory.
FlowBot: Inducing LLM Workflows with Bilevel Optimization and Textual Gradients
Hongyeon Yu (Naver Search US), Yoon Kim (Massachusetts Institute of Technology)
OptimizationAI Code AssistantReinforcement Learning from Human FeedbackNeural Architecture SearchTransformerLarge Language ModelPrompt EngineeringDiffusion modelFlow-based ModelTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose an automated method that uses bi-level optimization and text gradients to simultaneously learn the structure of LLM workflows and the prompts for each invocation.
FlowCloud: Learning Continuous Spatiotemporal Dynamics from Unpaired Sparse Point Cloud Snapshots
Yinbo Liu (Wuhan University), Tian Tian (Wuhan University)
GenerationData SynthesisTransformerOptical FlowPoint CloudTime SeriesBiomedical DataStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Proposes a FlowCloud framework based on variational Neural ODE, which can learn continuous spatiotemporal dynamics and generate complete trajectories from sparse, non-continuous, and unpaired point cloud snapshots.
Flowers: A Warp Drive for Neural PDE Solvers
Till Muser (University of Basel), Ivan Dokmanić (University of Basel)
OptimizationComputational EfficiencyConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderOptical FlowImageVideoTabularTime SeriesPhysics RelatedStochastic Differential Equation
🎯 What it does: Proposes FLOWERS, a neural PDE solver based on a multi-head warp mechanism, which utilizes point-to-point predicted displacements for global non-local mixing;
FlowMAP: Flow Matching for Generalizable Agent Planning
Jiarun Fu (Beijing Institute of Technology), Guoren Wang (Beijing Institute of Technology)
Reinforcement LearningFlow-based ModelTabularTime SeriesSequentialBenchmarkStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Propose FlowMAP, which learns planning dynamics of meta-state distributions through continuous-time flow matching, to enhance generalization in dynamic heterogeneous environments.
FlowNar: Scalable Streaming Narration for Long-Form Videos
Zeyun Zhong (Karlsruhe Institute of Technology), Jürgen Beyerer
Computational EfficiencyRepresentation LearningTransformerLarge Language ModelVision Language ModelContrastive LearningVideoTextRetrieval-Augmented Generation
🎯 What it does: Proposed the FLOWNAR framework, utilizing Dynamic Context Management (DCM) and Cross Linear Attentive Memory (CLAM) to achieve scalable streaming narration for long-duration videos;
FlowPET: Physics-Informed Symplectic Flow Matching for Low-Count PET Reconstruction
Zheng Zhang (Hong Kong Polytechnic University), Jing Qin (Hong Kong Polytechnic University)
RestorationConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelFlow-based ModelBiomedical DataPositron Emission TomographyStochastic Differential Equation
🎯 What it does: Propose FlowPET, a physics-informed symplectic flow matching framework for low-count PET reconstruction.
FlowSeg: Dynamic Semantic Guidance for LLM-Conditioned Segmentation
Zekang Zhang (Beijing Institute of Technology), Ting Liu (Meitu.inc)
SegmentationTransformerLarge Language ModelVision Language ModelDiffusion modelContrastive LearningImageTextMultimodality
🎯 What it does: Propose FlowSeg, which utilizes bidirectional semantic flow to dynamically guide segmentation under the condition of LLM, addressing the issue of semantic mismatch that occurs during the query-first generation-and-selection process.
FlowState: Sampling-Rate‑Equivariant Time‑Series Forecasting
Lars Graf (IBM Research Europe), Angeliki Pantazi (IBM Research Europe)
Computational EfficiencyRepresentation LearningRecurrent Neural NetworkTransformerFlow-based ModelAuto EncoderContrastive LearningTime SeriesBenchmarkStochastic Differential Equation
🎯 What it does: Propose the FlowState model, which utilizes an SSM encoder with variable sampling rates and a functional basis decoder to achieve continuous time and arbitrary sampling rate time series forecasting;
FluxNet: Learning Capacity-Constrained Local Transport Operators for Conservative and Bounded PDE Surrogates
Zishuo Lan (Northwestern Polytechnical University), Jincheng Wang (Northwestern Polytechnical University)
Convolutional Neural NetworkGraph Neural NetworkDiffusion modelAuto EncoderContrastive LearningTabularTime SeriesBenchmarkPhysics Related
🎯 What it does: Propose FluxNet, which achieves autoregressive approximation for conservative PDEs by learning capacity-limited cumulative transport, and strictly satisfies physical boundaries and constraint limits through modular L, U, D transport heads and virtual lattice technology;
FOAM: Blocked State Folding for Memory-Efficient LLM Training
Ziqing Wen (National University of Defense Technology), Tao Sun (National University of Defense Technology)
OptimizationComputational EfficiencyTransformerLarge Language ModelText
🎯 What it does: This paper proposes an optimizer called FOAM, which utilizes block-wise gradient averaging and residual correction to compress the optimizer state of Adam, significantly reducing memory usage and accelerating convergence during LLM training.
FOAM: Frequency and Operator-Error Based Adaptive Damping Method for Reducing Staleness-Oriented Error for Shampoo
Kyunghun Nam (KENTECH), Sumyeong Ahn (KENTECH)
OptimizationTransformerImageTextAudio
🎯 What it does: This paper proposes an adaptive damping method called FOAM based on frequency and operator error, which addresses the computational bottleneck and numerical instability caused by using outdated preconditioners in the Shampoo optimizer.
FOCA: Future-Oriented Conditioning for Data-Efficient Vision-Language-Action Adaptation
Minh Duc Nguyen (Vinrobotics), Vien Anh Ngo (Vinrobotics)
Representation LearningData-Centric LearningRobotic IntelligenceTransformerVision-Language-Action ModelDiffusion modelScore-based ModelFlow-based ModelContrastive LearningWorld ModelImageVideoTextMultimodality
🎯 What it does: This paper proposes a FOCA framework that utilizes future observations to perform explicit prediction and implicit alignment in the latent space, thereby enhancing the adaptation capability of vision-language-action models under limited demonstration data.
FocalPolicy: Frequency-Optimized Chunking and Locally Anchored Flow Matching for Coherent Visuomotor Policy
Qian He (Shenyang Institute of Automation, Chinese Academy of Sciences), Jiandong Tian (Shenyang Institute of Automation, Chinese Academy of Sciences)
Robotic IntelligenceTransformerReinforcement LearningDiffusion modelScore-based ModelFlow-based ModelContrastive LearningVideoSequential
🎯 What it does: Propose a prospective visual motion policy called FocalPolicy, which achieves smooth generation of long-horizon actions through frequency-optimized blocking and local anchored flow matching.
FOCUS & RePAIR: Mitigating Text Degeneration via Token-Level Guidance For Pruned Large Language Models
Junyoung Lee (POSTECH), Yeseong Kim (POSTECH)
Computational EfficiencyKnowledge DistillationTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringContrastive LearningText
🎯 What it does: This paper addresses the text degradation problem that occurs in large language models after pruning, proposing a token-level guided method that significantly reduces repetition cycles and improves generation quality.
Focus and Dilution: The Multi-stage Learning Process of Attention
Zheng-An Chen (Shanghai Jiao Tong University), Tao Luo (Shanghai Jiao Tong University)
Explainability and InterpretabilityRepresentation LearningTransformerAuto EncoderContrastive LearningText
🎯 What it does: Studied the multi-stage learning of the attention mechanism during Transformer training, discovering that attention first focuses on high-frequency words and gradually spreads out (focus-dilution cycle), and provided a stage-wise dynamical explanation.
Focus-Then-Contact: Speeding Up Robotic Contact-Rich Task Learning with Affordance-Guided Real-World Residual Reinforcement Learning
Guanren Qiao (Chinese University of Hong Kong (Shenzhen)), Guiliang Liu (Chinese University of Hong Kong (Shenzhen))
Robotic IntelligenceReinforcement Learning from Human FeedbackTransformerSupervised Fine-TuningReinforcement LearningContrastive LearningImageVideoSequential
🎯 What it does: Propose a human-robot collaborative real-time reinforcement learning framework named Focus-Then-Contact (FTC), which utilizes residual RL for fine-tuning on existing affine policies and accelerates the learning of contact-based fine-grained robotic manipulation tasks through affordance-guided dense rewards based on keyframes.
Focus, Align, and Sustain: Counteracting Gradient Dilution in Incremental Object Detection
Aoting Zhang (Chinese Academy of Sciences), Yu ZHOU
Object DetectionKnowledge DistillationRepresentation LearningTransformerContrastive LearningImage
🎯 What it does: This study proposes the FAS (Focus-Align-Sustain) framework to address the gradient dilution problem in incremental object detection based on DETR.
FOCUS: DLLMs Know How to Tame Their Compute Bound
Kaihua Liang (King Abdullah University of Science and Technology), Marco Canini (King Abdullah University of Science and Technology)
Computational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation
🎯 What it does: By identifying that most computations in the decoding process of DLLMs are wasted on tokens that are not decoded, the FOCUS system dynamically removes these useless tokens, significantly reducing FLOPs.
FOCUS: Forcing In-Context Object Localization through Visual Support Constraints and Policy Optimization
Mohammed Asad Karim (Amazon), Vinay Kumar Verma (Amazon)
Object DetectionTransformerReinforcement LearningVision Language ModelContrastive LearningImageVideo
🎯 What it does: Propose a two-stage, category-free visual context attention and reinforcement learning framework called FOCUS, aimed at achieving instance-level object localization based on support examples.
Focusing Where Vision Matters: Selective Training for Large Vision Language Models via Visual Information Gain
Seulbi Lee (Seoul National University of Science and Technology), Sangheum Hwang (Seoul National University of Science and Technology)
Explainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerSupervised Fine-TuningVision Language ModelContrastive LearningImageTextMultimodalityBenchmark
🎯 What it does: Propose the Visual Information Gain (VIG) metric, and use this metric to selectively train large-scale vision-language models (LVLMs) on training samples and tokens, aiming to enhance the model's visual inductive capability and reduce language bias.
Focusing: View-Consistent Sparse Voxels for Efficient 3D VAE Training
Xuhui Chen (Key Laboratory of System Software Chinese Academy of Sciences Institute of Software Chinese Academy of Sciences), Ying He (Nanyang Technological University)
GenerationComputational EfficiencyRepresentation LearningTransformerDiffusion modelScore-based ModelNeural Radiance FieldAuto EncoderContrastive LearningGaussian SplattingPoint CloudMesh
🎯 What it does: Trained a 3D VAE based on sparse voxels, utilizing perspective-consistent depth-driven voxel carving and adaptive scaling to significantly reduce memory and computational costs, enabling training at 1024³ resolution;
FoeGlass: Simple In-Context Learning Is Enough for Red Teaming Audio Deepfake Detectors
Sepehr Dehdashtian (Michigan State University), Gaurav Bharaj (Reality Defender)
Adversarial AttackTransformerLarge Language ModelPrompt EngineeringTextChain-of-ThoughtAudio
🎯 What it does: Leverage the adaptive context learning of large language models (LLM) to automatically search for natural adversarial examples in the input space of text-to-speech (TTS) models that can deceive audio deepfake detectors (ADD);
Follow-the-Perturbed-Leader for Decoupled Bandits: Best-of-Both-Worlds and Practicality
Chaiwon Kim (Seoul National University), Min-hwan Oh (Seoul National University)
OptimizationReinforcement Learning
🎯 What it does: Proposed a Follow-the-Perturbed-Leader (FTPL) algorithm based on Pareto perturbation for decoupling multi-armed bandit problems, achieving best-of-both-worlds (BOBW) performance in both robust and stochastic environments.
ForceForget: Reinforcement Concept Removal for Enhancing Safety in Text-to-Image Models
Dong Han (Huawei Technologies), Yong Li (Friedrich Schiller University Jena)
GenerationSafty and PrivacyTransformerSupervised Fine-TuningReinforcement LearningPrompt EngineeringDiffusion modelImageTextMultimodality
🎯 What it does: Optimize Concept Elimination Reward (CER) using reinforcement learning and introduce a safety adapter to regulate partial text embeddings in cross-attention, thereby eliminating unsafe concepts in text-to-image models without compromising generation quality.
Foreground-Aware Token Routing Vision Transformer for Real-Time Satellite Video Tracking
Jiahao Wang (Xidian University), Xu Liu (Xidian University)
Object TrackingTransformerAuto EncoderContrastive LearningVideo
🎯 What it does: Proposed the FATrack framework, achieving real-time satellite video tracking through two core modules: FA-ViT and ASM.
Forensic Prompting with Dual-Action Policy Optimization for Vision-Language Forgery Detection and Localization
Ye Zhu (Hebei University of Technology), Jinwei Wang (Nankai University)
Anomaly DetectionOptimizationReinforcement Learning from Human FeedbackTransformerReinforcement LearningPrompt EngineeringVision Language ModelDiffusion modelContrastive LearningImageTextMultimodality
🎯 What it does: Combine visual-language prompts and dual-action strategies for joint detection and localization of image forgery.
ForensicConcept: Transferable Forensic Concepts for AIGI Detection
Menyanshu Zhou (Xiamen University), Rongrong Ji (Xiamen University)
Anomaly DetectionExplainability and InterpretabilityTransformerDiffusion modelAuto EncoderContrastive LearningImage
🎯 What it does: Propose the FORENSICCONCEPT framework, which uses Transformer attribution to locate key image patches, clusters to obtain interpretable forensic concept codebooks, and employs concept-aligned projection in the detector to provide auditable evidence; meanwhile, it utilizes CleanDIFT's diffusion features as a reference for generative trajectories, quantifies the neighborhood structure consistency between the detector and diffusion features (CKNNA), and achieves knowledge transfer across backbone networks through concept codebook injection.
Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding
Zelin Zheng (University of Chinese Academy of Sciences), Laiyun Qing (University of Chinese Academy of Sciences)
RecognitionExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelVideoTextRetrieval-Augmented Generation
🎯 What it does: Propose a video temporal localization framework named Foresee-to-Ground (F2G), which redefines the video temporal localization task as a verifiable Identify-then-Measure process.
ForesightKV: Optimizing KV Cache Eviction for Reasoning Models by Learning Long-Term Contribution
Zican Dong (Renmin University of China), Xin Zhao
OptimizationComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningScore-based ModelTextBenchmark
🎯 What it does: This paper proposes ForesightKV, which trains a lightweight scoring model to dynamically evaluate the future contribution of KV pairs and efficiently eliminates KV pairs during the generation process according to the budget;
Forget by Uncertainty: Orthogonal Entropy Unlearning for Quantized Neural Networks
Tian Zhang (Wuhan University of Technology), Jingling Yuan (Wuhan University of Technology)
ClassificationFederated LearningSafty and PrivacyComputational EfficiencyConvolutional Neural NetworkSupervised Fine-TuningContrastive LearningImage
🎯 What it does: Propose a machine forgetting framework called OEU for quantized neural networks, which can eliminate the influence of specified data without complete retraining.
Forget to Know, Remember to Use: Context-Aware Unlearning for Large Language Models
Yuefeng Peng (University of Massachusetts Amherst), Dezhi Hong (Amazon)
Federated LearningSafty and PrivacyExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningAdversarial AttackData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Studied the ability of large language models (LLMs) to use previously forgotten knowledge in context after unlearning, and proposed a context-aware unlearning method that maintains contextual utility.
Forget-It-All: Multi-Concept Machine Unlearning via Concept-Aware Neuron Masking
Kaiyuan Deng (University of Arizona), Xiaolong Ma (University of Arizona)
GenerationSafty and PrivacyTransformerDiffusion modelAuto EncoderContrastive LearningImageTextMultimodality
🎯 What it does: Propose a training-free multi-concept machine forgetting framework called FIA, which achieves multi-concept forgetting in text-to-image diffusion models such as Stable Diffusion by utilizing concept-sensitive neuron masks.
Forgetting Whenever You Want: A Decentralized Continual Learning Framework with On-Demand Unlearning
Xiao Zhang (Shandong University), Dongxiao Yu (Shandong University)
Data SynthesisFederated LearningKnowledge DistillationRepresentation LearningConvolutional Neural NetworkTransformerDiffusion modelContrastive LearningImage
🎯 What it does: Proposed the DCU framework, achieving continuous learning and on-demand forgetting in decentralized environments;
FormAct: Agentic Source Editing for Rich-Format Document Generation
Eugene J. Yu (Peking University), Sujian Li (Peking University)
GenerationData SynthesisAI Code AssistantTransformerLarge Language ModelAgentic AIPrompt EngineeringVision Language ModelTextMultimodalityBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose FORMACT, an agent-based source editing framework based on LLM, capable of generating professional-level rich-format (HTML) documents from scratch and performing multi-round iterative optimization through rendering feedback.
Formal Concept Lattices are Good Semantic Scaffolds for Concept-Based Learning
Deepika Vemuri, Vineeth N. Balasubramanian
Explainability and InterpretabilityRepresentation LearningTransformerContrastive LearningImage
🎯 What it does: Constructing a concept lattice using formal concept analysis, performing hierarchical learning on the concept hierarchy, migrating concept learning from single layers to various depths of the network, thereby enhancing interpretability and model performance.
Formalizing and Falsifying Causal Pathways of Rare Events
Anahita Haghighat (Independent), Dominik Janzing (Amazon Research)
Explainability and InterpretabilityLarge Language ModelScore-based ModelContrastive LearningTextReview/Survey Paper
🎯 What it does: This paper proposes a formal definition of the causal pathway for rare events and introduces metrics such as the explanation score to evaluate causal explanations from root causes to target events. The authors theoretically elaborate on pathway abstraction using tools such as structural equation models, interventions (do-operations), and KL divergence, and provide several theorems and examples for illustration.
Formalizing Learning from Language Feedback with Provable Guarantees
Wanqiao Xu (Stanford University), Ching-An Cheng (Google Research)
Reinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningTextBenchmarkChain-of-Thought
🎯 What it does: Proposes the theoretical framework of Learning from Language Feedback (LLF), and provides a proof that regret-free learning can be achieved under certain assumptions;
Formalizing the Binding Problem
Lianghuan Huang (University of Pennsylvania), Konrad Kording (University of Pennsylvania)
Explainability and InterpretabilityRepresentation LearningTransformerContrastive LearningImageMultimodality
🎯 What it does: Formalize the binding problem within an information theory framework, and design a trainable detector to quantify binding information in model representations, subsequently evaluating binding capability on Vision Transformers.
FormalJudge: A Neuro-Symbolic Paradigm for Agentic Oversight
Jiayi Zhou (Peking University), Jie Fu (Shanghai AI Lab)
Safty and PrivacyExplainability and InterpretabilityAI Code AssistantTransformerLarge Language ModelAgentic AITextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposes the FormalJudge neuro-symbolic framework, using LLM as a specification compiler to generate Dafny specifications, which are then deterministically verified by the Z3 SMT solver.
Formally Exploring Visual Anomaly Detection Evaluation Metrics
Nasar Iqbal (University of Udine), Marius Kloft (University of KaiserslauternLandau)
Anomaly DetectionImageBenchmark
🎯 What it does: This paper formally analyzes the evaluation metrics for visual anomaly detection (VAD), proposes eight verifiable attributes, and evaluates whether existing metrics satisfy these attributes through case studies and experiments; subsequently, it designs and proposes a new metric called SAAM-ALARM, which theoretically satisfies all attributes and demonstrates its superiority on real data.
FormalRx: Rectify and eXamine Semantic Failures in Autoformalization
Haocheng Wang (Hong Kong University of Science and Technology), Zhijiang Guo (Hong Kong University of Science and Technology)
Explainability and InterpretabilityData-Centric LearningAI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose the FORMALRX framework, which provides fine-grained automated formula evaluation and diagnosis;
FormulaCode: Evaluating Agentic Optimization on Large Codebases
Atharva Sehgal (University of Texas at Austin), Yisong Yue (California Institute of Technology)
OptimizationAI Code AssistantTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Designed and released the FORMULACODE benchmark to evaluate the performance optimization capabilities of large LLM coding agents at the real-world repository level across multiple workloads.
Forward-Chaining Temporal Point Process
Chao Yang (Chinese University of Hong Kong), Shuang Li (Chinese University of Hong Kong)
GenerationData SynthesisRecurrent Neural NetworkTransformerScore-based ModelGenerative Adversarial NetworkContrastive LearningTabularTime SeriesBiomedical DataElectronic Health RecordsChain-of-Thought
🎯 What it does: This paper proposes a Forward Chain Temporal Point Process (FC-TPP), which can generate controllable and constraint-satisfying event sequences in continuous time based on hidden symbolic states.
Forward-KL Convergence of Time-Inhomogeneous Langevin Diffusions
Andreas Habring (Graz University of Technology), Martin Zach (Ecole Polytechnique Fédérale de Lausanne)
OptimizationTabularTime SeriesPhysics RelatedStochastic Differential Equation
🎯 What it does: Studies the convergence of non-homogeneous Langevin diffusion in time under the forward KL divergence and provides non-asymptotic error upper bounds for the discretization.
Foundation Inference Models for Ordinary Differential Equations
Johannes R. Hübers, Ramses J Sanchez (Lamarr Institute For Machine Learning and Artificial Intelligence)
Data SynthesisOptimizationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerSupervised Fine-TuningContrastive LearningTabularTime SeriesSequentialBenchmarkStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: This paper proposes and pretrains FIM-ODE, a fundamental inference model that can directly infer ODE vector fields from noisy trajectories in a single forward pass.
Foundation VAE for CT Reconstruction, Augmentation, and Generation
Qi Chen (Johns Hopkins University), Jingjing Fu (Microsoft Research)
RestorationGenerationData SynthesisTransformerVision Language ModelDiffusion modelAuto EncoderContrastive LearningImageTextBiomedical DataComputed Tomography
🎯 What it does: Propose using a Foundation VAE, which has been pre-trained on a large scale on natural images and videos, as a zero-shot representation interface for 3D CT, achieving CT reconstruction, data augmentation, and generation.
Foundations of Equivariant Deep Learning: Unifying Graph and Sheaf Neural Networks
Yoshihiro Maruyama (Kyoto University)
Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningGraph Neural NetworkReinforcement LearningContrastive LearningGraphReview/Survey Paper
🎯 What it does: Proposed an order-equivariant neural network (OENN) based on posets, unifying graph neural networks, Sheaf networks, and higher-order message passing models, and providing a complete construction from equivariant bundles to nonlinear layers.
FoundObj: Self-supervised Foundation Models as Rewards for Label-free 3D Object Segmentation
Zihui Zhang (Hong Kong Polytechnic University), Bo Yang (Hong Kong Polytechnic University)
SegmentationTransformerReinforcement LearningAuto EncoderContrastive LearningPoint Cloud
🎯 What it does: Propose the FoundObj framework, which utilizes self-supervised 2D/3D foundation models as rewards to train RL agents to discover multi-class objects in 3D point clouds without annotations.
Fourier Features Let Agents Learn High Precision Policies with Imitation Learning
Balázs Gyenes (Karlsruhe Institute of Technology), Gerhard Neumann (Karlsruhe Institute of Technology)
Robotic IntelligenceReinforcement Learning from Human FeedbackTransformerDiffusion modelScore-based ModelNeural Radiance FieldPoint Cloud
🎯 What it does: Propose the use of Fourier feature mapping in point cloud inputs to alleviate spectral bias in neural networks, thereby improving the performance of differential imitation learning in high-precision robotic manipulation tasks.
FourTune: Towards Fully 4-Bit Efficient Post-Training for Diffusion Models
Bowen Xue (Stanford University), Muyang Li (Nunchux AI)
GenerationComputational EfficiencyKnowledge DistillationTransformerDiffusion modelImage
🎯 What it does: Proposes FourTune, a fully 4-bit (W4A4G4) efficient framework for post-training of diffusion models.
FOVI: A biologically-inspired foveated interface for deep vision models
Nicholas Blauch, Talia Konkle (Harvard University)
ClassificationComputational EfficiencyRepresentation LearningConvolutional Neural NetworkTransformerContrastive LearningImage
🎯 What it does: This paper designs a biology-inspired concave sampling interface called FOVI based on retino-cortical mapping, and combines it with convolutional networks and Vision Transformers to significantly reduce pixel and computational costs in high-resolution visual tasks.
Fox in the Henhouse: Supply-Chain Backdoor Attacks Against Reinforcement Learning
Shijie Liu (University of Melbourne), Benjamin I. P. Rubinstein
Adversarial AttackConvolutional Neural NetworkRecurrent Neural NetworkReinforcement LearningPrompt EngineeringImageVideoSequential
🎯 What it does: Investigated the feasibility of using pre-trained strategies from supply chains to launch backdoor attacks in reinforcement learning, and implemented a black-box attack method called SCAB.
FPTQuant: Function-Preserving Transforms for LLM Quantization
Boris van Breugel (Qualcomm AI Research), Markus Nagel (Qualcomm AI Research)
Computational EfficiencyKnowledge DistillationTransformerLarge Language ModelText
🎯 What it does: Proposed a function-preserving transformation (FPT) method called FPTQuant for low-precision INT4 quantization on large language models (LLMs), while maintaining the model's functionality unchanged.
FRACTAL: State Space Model with Fractional Recurrent Architecture for Computational Temporal Analysis of Long Sequences
Mengqi Li (Northwestern Polytechnical University), Lixin Li (Northwestern Polytechnical University)
Computational EfficiencyRepresentation LearningRecurrent Neural NetworkTransformerContrastive LearningImageTextTime SeriesSequentialBenchmarkStochastic Differential Equation
🎯 What it does: Proposed a fractional-order measure-based state space model called FRACTAL for efficiently processing long-sequence temporal features.
Fractional is Better: Learnable Derivative Orders in Neural Operator Learning
Fares B. Mehouachi (New York University Abu Dhabi), Saif Jabari
Convolutional Neural NetworkGraph Neural NetworkTransformerDiffusion modelScore-based ModelContrastive LearningTabularTime SeriesBenchmarkPhysics RelatedStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: This paper proposes a neural operator framework with learnable fractional derivative features (∂-NO), which alleviates the network's burden on implicit differentiation by introducing adjustable-order fractional derivatives into the input;
FrameOracle: Learning What to See and How Much to See in Videos
Chaoyu Li (Arizona State University), Pooyan Fazli (NewsBreak)
CompressionComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelContrastive LearningVideoTextMultimodality
🎯 What it does: Designed and implemented FrameOracle, a lightweight and pluggable frame selection module that can predict which frames are most important and how many frames need to be retained based on the query.
FreeRet: MLLMs as Training-Free Retrievers
Yuhan Zhu (Nanjing University), Limin Wang (Nanjing University)
RetrievalTransformerLarge Language ModelPrompt EngineeringImageVideoTextMultimodalityBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose FreeRet, a training-free framework that can directly convert any off-the-shelf multimodal large language model (MLLM) into a two-stage retriever, performing both embedding extraction and reranking to complete the end-to-end process of retrieval and generation.
FreeText: Training-Free Text Rendering via Attention Localization and Spectral Glyph Injection
RuiQiang Zhang (University of Science and Technology of China), Weiming Zhang (University of Science and Technology of China)
RestorationGenerationTransformerPrompt EngineeringDiffusion modelImageTextMultimodalityBenchmark
🎯 What it does: Propose a training-free, plug-and-play text rendering enhancement framework called FreeText, which utilizes the intrinsic attention of diffusion Transformers to locate text positions and injects glyph priors through frequency domain modulation;
Frequency Matching in Spiking Neural Networks for mmWave Sensing
Zhenyu Liao (Zhejiang University), Shuiguang Deng (Zhejiang University)
ClassificationOptimizationComputational EfficiencyConvolutional Neural NetworkRecurrent Neural NetworkSpiking Neural NetworkTransformerSupervised Fine-TuningPoint CloudPhysics Related
🎯 What it does: This paper investigates the effectiveness of spiking neural networks in millimeter-wave sensing through frequency domain analysis.
Frequency-Aware Perceptual Optimization for Low-Complexity Implicit Image Compression
Haotian Wu (Zhejiang University), Deniz Gunduz (Imperial College London)
CompressionConvolutional Neural NetworkDiffusion modelScore-based ModelAuto EncoderContrastive LearningImage
🎯 What it does: Proposed a low-complexity implicit image compression framework called Re IC, which combines significance-driven region partitioning with Wavelet-Wasserstein loss to achieve frequency-aware perceptual optimization.
Frequentist Consistency of Prior-Data Fitted Networks for Causal Inference
Valentyn Melnychuk (LMU Munich), Rahul G Krishnan
Explainability and InterpretabilityComputational EfficiencyData-Centric LearningScore-based ModelContrastive LearningTabularTime SeriesBiomedical DataReview/Survey PaperBenchmark
🎯 What it does: This paper studies the frequentist consistency of prior data fitting networks (PFN) in causal inference and proposes a 'first-order posterior correction' (MP-OSPC) based on the martingale posterior to correct prior-induced confounding bias.
Frictional Q-Learning
Hyunwoo Kim (Korea University), Hyo Kyung Lee (Korea University)
TransformerReinforcement LearningAuto EncoderContrastive LearningTabularTime Series
🎯 What it does: By analogy between the support-related questions of batch data and the phenomenon of static friction, an offline reinforcement learning algorithm called Frictional Q-Learning is designed to separate the tangential (supported) and normal (unsupported) directions in the action space.
FRIGID: Scaling Diffusion-Based Molecular Generation from Mass Spectra at Training and Inference Time
Montgomery Bohde (Massachusetts Institute of Technology), Connor W. Coley (Massachusetts Institute of Technology)
Drug DiscoveryTransformerLarge Language ModelPrompt EngineeringDiffusion modelContrastive LearningGraphTabularBiomedical DataBenchmarkRetrieval-Augmented Generation
🎯 What it does: This paper proposes FRIGID, a molecular generation framework based on diffusion language models, used to reverse infer the chemical structures of unknown small molecules from tandem mass spectrometry (MS/MS) data. The framework can be pre-trained on billions of unlabeled molecules at a large scale, and during the inference stage, it utilizes a forward fragmentation model (ICEBERG) to detect and correct structurally inconsistent substructures, thereby achieving scalable inference acceleration.
FRISM: Fine-Grained Reasoning Injection via Subspace-Level Model Merging for Vision–Language Models
Chenyu Huang (Fudan University), Tao Chen (Fudan University)
Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelMixture of ExpertsVision Language ModelScore-based ModelContrastive LearningImageTextMultimodality
🎯 What it does: Proposed the FRISM framework, which injects fine-grained reasoning capabilities into vision-language models at the subspace level by performing SVD decomposition on the task vectors of large language reasoning models.
From 2D Grids to 1D Tokens: Reforming Shared Representations for Multimodal Image Fusion
Yuchen Xian (Zhejiang University), Yi Yang (Zhejiang University)
Image HarmonizationRestorationObject DetectionSegmentationTransformerAuto EncoderContrastive LearningImageMultimodalityBiomedical DataBenchmark
🎯 What it does: This paper proposes a hybrid multi-modal image fusion framework based on a frozen one-dimensional image tokenizer, utilizing one-dimensional tokens to carry global appearance information while retaining two-dimensional spatial channels to recover local details.
From Absolute to Relative: Rethinking Reward Shaping in Group-Based Reinforcement Learning
Wenzhe Niu (Tianjin University), Renqing He (Meituan)
TransformerLarge Language ModelReinforcement LearningPrompt EngineeringContrastive LearningText
🎯 What it does: Proposed a reinforcement learning framework based on relative rewards, RLRR, which transforms rewards from absolute scores into group-relative rankings to address the issues of sparse rewards in verifiable tasks and unstable reward ranges in open-ended tasks;
From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model
Bing Hu (Harbin Institute of Technology), Liqiang Nie (Harbin Institute of Technology)
Representation LearningRobotic IntelligenceReinforcement Learning from Human FeedbackTransformerMixture of ExpertsVision-Language-Action ModelDiffusion modelFlow-based ModelContrastive LearningImageVideoTextMultimodality
🎯 What it does: Propose the BehaviorVLA framework, which learns cross-domain temporal coherent behavior representations, achieving a visual-language-action model from abstraction to instantiation.
From Associations to Activations: Comparing Behavioral and Hidden-State Semantic Geometry in LLMs
Louis Schiekiera (Humboldt-Universitaet zu Berlin), Fritz Günther (Humboldt-Universitaet zu Berlin)
Explainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelText
🎯 What it does: Studied how the geometric structure of LLM hidden states can be recovered from behavioral data in psycholinguistic experiments (forced choice and free association), and compared the consistency between behavioral semantic geometry and hierarchical hidden state geometry.
From Backward Spreading to Forward Replay: Revisiting Target Construction in LLM Parameter Editing
Wei Liu (National University of Singapore), Wee Sun Lee (National University of Singapore)
OptimizationComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextBenchmark
🎯 What it does: Studied the theoretical foundations of target construction in parameter editing of large language models, and proposed a forward replay method based on forward propagation to replace the traditional backward propagation diffusion, achieving more precise multi-layer targets while maintaining the same computational complexity.
From Basis to Basis: Gaussian Particle Representation for Interpretable PDE Operators
Zhihao Li (Hong Kong University of Science and Technology), Wei Wang (Hong Kong University of Science and Technology)
Explainability and InterpretabilityComputational EfficiencyTransformerContrastive LearningGaussian SplattingPoint CloudTabularTime SeriesPhysics Related
🎯 What it does: Propose a Gaussian Particle Operator (GPO) neural operator, using a learnable Gaussian particle basis and Petrov–Galerkin attention to represent and predict fluid PDEs.
From Bits to Rounds: Parallel Decoding with Exploration for Diffusion Language Models
Hengyu Fu (University of California, Berkeley), Jiantao Jiao (University of California, Berkeley)
Computational EfficiencyAI Code AssistantTransformerLarge Language ModelDiffusion modelTextBenchmark
🎯 What it does: Proposed a theoretical lower bound for parallel decoding under information theory and designed an exploration-exploitation algorithm ETE, significantly improving the parallel decoding efficiency of diffusion language models.
From Blind Spots to Gains: Diagnostic-Driven Iterative Training for Large Multimodal Models
Hongrui Jia (Peking University), Wei Ye (Peking University)
Explainability and InterpretabilityComputational EfficiencyData-Centric LearningReinforcement Learning from Human FeedbackTransformerReinforcement LearningAgentic AIPrompt EngineeringVision Language ModelGenerative Adversarial NetworkImageTextMultimodality
🎯 What it does: This paper proposes a diagnosis-driven iterative training framework called DPE, which combines multi-agent tool-based data generation with reinforcement learning to dynamically generate and reinforce training samples targeting the model's blind spots.
From Coarse to Fine: Deep Prototype Refinement Network for Few-Shot Point Cloud Semantic Segmentation
Changshuo Wang (University College London), Dimitrios Kanoulas (University College London)
SegmentationGraph Neural NetworkTransformerMixture of ExpertsAuto EncoderContrastive LearningPoint Cloud
🎯 What it does: The paper proposes a deep prototype refinement network called DPR-Net, which can perform semantic segmentation on point clouds with a very small number of labeled samples, achieving gradual adaptation through a multi-stage prototype evolution trajectory from coarse to fine;
From Conflict to Consensus: Boosting Medical Reasoning via Multi-Round Agentic RAG
Wenhao Wu (Nanjing University), Zhi Wang (Nanjing University)
RetrievalExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningDrug DiscoveryTransformerAgentic AIPrompt EngineeringTextBiomedical DataBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposed a multi-round Agentic RAG framework, MA-RAG, to enhance answer quality in medical question answering through iterative retrieval and reasoning.
From Content to Knowledge: Lightning Fast Long-Video Understanding with Neural Knowledge Representations
Yuchen Guan (Tsinghua University), Yan Lu (Microsoft Research Asia)
Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningAgentic AIVideoTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposes compressing the semantic content of long videos into pluggable network weights through Neural Knowledge Representation (NKR), enabling instant querying without reprocessing the video.
From Correspondence to Actions: Human-Like Multi-Image Spatial Reasoning in Multi-modal Large Language Models
Masanari Oi (Institute of Science Tokyo), Naoaki Okazaki (Institute of Science Tokyo)
Image TranslationAutonomous DrivingOptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningVision Language ModelVision-Language-Action ModelContrastive LearningImageMultimodalityBenchmark
🎯 What it does: Studies multi-view image spatial reasoning, proposing the HATCH framework, which explicitly supervises cross-view correspondence and progressive view transformation
From Denoising to De-Channeling: Integrating Physical Channel Priors into Diffusion Models for Radio Signal Understanding
Yaoqi Liu (Pengcheng Laboratory), Chuan Shi (Beijing University of Posts and Telecommunications)
ClassificationAnomaly DetectionTransformerDiffusion modelContrastive LearningTime SeriesPhysics RelatedAudio
🎯 What it does: Propose a diffusion model named PWC-Diff, which utilizes physical wireless channel priors to achieve de-channeling from received signals to transmitted signals, and performs wireless signal understanding tasks based on this.
From Diagrams to Code: Multilingual Programming with Visual Design
Linzheng Chai (Beihang University), Zhoujun Li (Beihang University)
GenerationData SynthesisAI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmark
🎯 What it does: Developed a multilingual multimodal instruction tuning dataset M2 C-INSTRUCT, a multimodal code generation model M2-CODER trained on this dataset, and a benchmark for evaluating code generation from diagrams, M2 EVAL.
From Directions to Regions: Decomposing Activations in Language Models via Local Geometry
Or David Shafran (Tel Aviv University), Mor Geva (Tel Aviv University)
Explainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelMixture of ExpertsAuto EncoderText
🎯 What it does: This paper proposes an activation decomposition framework based on the Mixture of Factor Analyzers (MFA), which partitions the activation space of a language model into multiple local low-dimensional Gaussian regions, capturing the regional affiliation and local variations of each activation through the region centroid and local subspace;
From Distribution to Geometry: Stable Graph Generalization via Invariant Barycenters
Hangyuan Du (Shanxi University), Wenjian Wang (Shanxi University)
Domain AdaptationRepresentation LearningGraph Neural NetworkContrastive LearningGraph
🎯 What it does: Proposes the DIGL method, which learns graph representations generalizable across different distributions through optimal transport barycenter, enhancing the OOD generalization ability of GNNs.
From Drift to Coherence: Stabilizing Beliefs in LLMs
SongEun Kim (Seoul National University), Juho Lee (Korea Advanced Institute of Science & Technology)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextBenchmark
🎯 What it does: Propose a method based on prompted predictive resampling (PPR) and seed-answer induced self-consistency loss (SC loss), which makes the beliefs of LLMs stabilize over time when answering the same question multiple times, eliminating early drift.
From Evaluation to Design: Using Potential Energy Surface Smoothness Metrics to Guide ML Interatomic Potential Architectures
Ryan Liu (California Institute of Technology), Aditi S. Krishnapriyan (University of California, Berkeley)
Drug DiscoveryGraph Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningGraphTabularBenchmarkPhysics Related
🎯 What it does: Proposed and implemented the Bond Smoothness Characterization Test (BSCT), which evaluates the smoothness of the potential energy surface of machine learning interatomic potentials (MLIP) through one-dimensional bond stretching, and used this metric to make iterative improvements during the model design process.
From Extraction to Deduction: Resolving Functional Misalignment in RAG via a Collaborative Critic-Reasoner Framework
Yufei Chen (Xidian University), Tao Gu (Macquarie University)
RetrievalExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose a collaborative CRITIC-REASONER framework, which disassembles the retrieval context through fine-grained masking technology, obscuring erroneous answer paragraphs while preserving useful background information, thereby shifting the generation model from extractive copying to reasoning-based answering;
From Extrinsic to Intrinsic: Geodesic-Guided Representation Learning for 3D Geometric Data
Yuming Zhao (City University of Hong Kong), Ying He (Nanyang Technological University)
Representation LearningTransformerAuto EncoderContrastive LearningPoint CloudMesh
🎯 What it does: Proposes the PRISM framework, which performs self-supervised pre-training by predicting the intrinsic geodesic distances of 3D surfaces, learning isometric embeddings to obtain fine-grained geometric features.
From Feasible to Practical: Pareto-Optimal Synthesis Planning
Friedrich Hastedt (Imperial College London), Antonio Del rio chanona
OptimizationDrug DiscoveryGraphTabular
🎯 What it does: Proposed a multi-objective retrosynthetic planning method called MORetro*, which generates Pareto front synthetic routes through weight scalarization and Bayesian optimization sampling