ICML 2026 Papers — Page 33
International Conference on Machine Learning · 6554 papers
LithoGRPO: Fast Inverse Lithography via GRPO Reinforced Flow Matching
Yao Lai (University of Hong Kong), Ping Luo (University of Hong Kong)
OptimizationTransformerReinforcement LearningDiffusion modelScore-based ModelFlow-based ModelRectified FlowImageBenchmarkStochastic Differential Equation
🎯 What it does: Developed an integrated framework for inverse lithography called LithoGRPO, combining flow matching with GRPO reinforcement learning to achieve efficient, multi-objective optimizable mask generation.
LitReview Arena: Evaluating Literature Review Agents with Battle-style Peer Review Platform
Ruotong Zhao (Tsinghua University), Yong Li (Tsinghua University)
TransformerLarge Language ModelAgentic AIPrompt EngineeringMixture of ExpertsTextReview/Survey PaperBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Establish the LitReview Arena platform and the LitReviewBench evaluation framework, collect expert preferences for literature review generation models, and provide an offline evaluation tool called LitJudge.
Little By Little: Continual Learning via Incremental Mixture of Rank-1 Associative Memory Experts
Haodong Lu (University of New South Wales), Dong Gong (University of New South Wales)
Computational EfficiencyKnowledge DistillationRepresentation LearningMeta LearningTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsContrastive LearningImageTextMultimodalityBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposes a continuous learning framework called MoRAM based on a fine-grained rank-1 associative memory expert, which can incrementally expand parameters step by step and dynamically schedule memory through content-addressable retrieval during inference.
LittleBit-2: Maximizing the Spectral Energy Gain in Sub-1-Bit LLMs via Latent Geometry Alignment
Banseok Lee (Samsung Research), Youngmin Kim (Samsung Research)
Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelText
🎯 What it does: This paper proposes the LittleBit‑2 framework to address the performance bottleneck of large language models during sub-1-bit quantization, maximizing spectral energy gains.
LIVE: Long-horizon Interactive Video World Modeling
Junchao Huang (Chinese University of Hong Kong), Li Jiang (Chinese University of Hong Kong)
GenerationData SynthesisTransformerDiffusion modelScore-based ModelContrastive LearningWorld ModelVideo
🎯 What it does: Proposes LIVE, a method for long-term interactive video world modeling through cyclic consistency constraints, which can explicitly limit error accumulation without the need for a teacher model.
LIVEditor-14B: Lightning Unified Video Editing via In-Context Sparse Attention
Shitong Shao (Hong Kong University of Science and Technology (GZ)), Zeke Xie (Hong Kong University of Science and Technology (GZ))
GenerationComputational EfficiencyTransformerPrompt EngineeringVision Language ModelDiffusion modelContrastive LearningImageVideoText
🎯 What it does: Proposed the LIVEditor-14B video editing model and designed In-Context Sparse Attention (ISA) to significantly reduce the computational cost of attention in video editing.
LiveFigure: Generating Editable Scientific Illustration with VLM Agents
Chenyang Shao (Tsinghua University), Yong Li (Tsinghua University)
GenerationData SynthesisTransformerAgentic AIPrompt EngineeringVision Language ModelImageTextMultimodalityRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose the LiveFigure framework, which uses a Visual Language Model Agent (VLM Agent) to simulate the expert drawing process. It first generates a graphical blueprint through visual priors, then creates an executable PowerPoint script using a standardized skill library, and finally achieves automated generation of editable vector graphics through a visual diagnostic loop.
LiveNewsBench: Evaluating Web Search Agents with Freshly Curated News
Yunfan Zhang (Columbia University), Smaranda Muresan (Columbia University)
TransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Designed and released a regularly updated benchmark (LIVENEWSBENCH), which automatically scrapes recent news and generates question-answer pairs to evaluate the agentic web search capability of large language models (LLMs).
LiveOIBench: Can Large Language Models Outperform Human Contestants in Informatics Olympiads?
Kaijian Zou (University of Michigan), Lu Wang (University of Michigan)
AI Code AssistantTransformerLarge Language ModelPrompt EngineeringTextBenchmarkChain-of-Thought
🎯 What it does: Designed and released LiveOIBench, based on 403 expert-designed problems from the 2023-2025 Informatics Olympiad competitions, equipped with fine-grained subtask scoring and a complete offline evaluation system, to assess the programming and reasoning capabilities of LLMs.
LK Losses: Direct Acceptance Rate Optimization for Speculative Decoding
Alexander Samarin (Nebius), Alexander Golubev (Nebius)
OptimizationAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsTextBenchmark
🎯 What it does: This paper proposes the LK loss, which directly trains and optimizes the acceptance rate in speculative decoding, improving the traditional method of training draft models based on KL divergence.
LLawCo: Learning Laws of Cooperation for Modeling Embodied Multi-Agent Behavior
Qinhong Zhou (University of Massachusetts Amherst), Anoop Cherian (Mitsubishi Electric Research Laboratories)
Autonomous DrivingOptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyRobotic IntelligenceReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextSequentialBenchmark
🎯 What it does: Proposes the LLawCo framework, enabling embedded multi-agent systems to achieve adaptive alignment with partners and task goals by automatically learning 'cooperation rules';
LLM Priors for ERM over Programs
Shivam Singhal (Texas A&M University), Tomer Galanti (Texas A&M University)
Computational EfficiencyData-Centric LearningAI Code AssistantTransformerLarge Language ModelPrompt EngineeringTextTabularSequentialRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This paper proposes a propose-and-verify method based on pre-trained large language models (LLM-PV), which generates candidate executable programs using LLM and verifies them on a validation set, thereby achieving program learning with a small number of samples.
LLM Self-Recognition: Steering and Retrieving Activation Signatures
Thibaud Ardoin (Freie Universitaet Berlin), Gerhard Wunder (Freie Universitaet Berlin)
RecognitionExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelPrompt EngineeringAuto EncoderContrastive LearningText
🎯 What it does: Utilize the self-identifiable features of internal activation flows within large language models, and during inference, introduce sparse vectors to intermediate activations to create a retrievable model fingerprint, thereby achieving text attribution and watermarking.
LLM Watermark Evasion via Bias Inversion
Jeongyeon Hwang (Pohang University of Science and Technology), Jungseul Ok (Pohang University of Science and Technology)
Adversarial AttackTransformerLarge Language ModelPrompt EngineeringText
🎯 What it does: Propose a rewrite attack based on negative logit bias (BIRA), which can effectively eliminate LLM watermarks and maintain semantics under black-box conditions;
LLM-based Embeddings: Attention Values Encode Sentence Semantics Better Than Hidden States
Yeqin Zhang (Nanjing University), Cam-Tu Nguyen (Nanjing University)
ClassificationRetrievalRepresentation LearningTransformerLarge Language ModelAgentic AIPrompt EngineeringContrastive LearningTextBenchmark
🎯 What it does: Propose VA and AlignedWVA for generating sentence embeddings by aggregating the value vectors of LLM instead of hidden states.
LLM-Guided Communication for Cooperative Multi-Agent Reinforcement Learning
Sangjun Bae (UNIST), Seungyul Han (UNIST)
TransformerLarge Language ModelReinforcement LearningPrompt EngineeringTabularSequentialChain-of-Thought
🎯 What it does: Propose a multi-agent communication framework called LMAC based on large language models (LLMs), which leverages the reasoning ability of LLMs to design and iteratively improve communication protocols, enabling agents to more accurately and consistently reconstruct the global state;
LLM-Guided Diagnostic Evidence Alignment for Medical Vision–Language Pretraining under Limited Pairing
Huimin Yan (Shanxi University), Long Chen (Hong Kong University of Science and Technology)
ClassificationRetrievalRepresentation LearningData-Centric LearningGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringContrastive LearningImageTextBiomedical DataMagnetic Resonance ImagingComputed TomographyElectronic Health RecordsReview/Survey Paper
🎯 What it does: Proposed a diagnostic evidence alignment pre-training method based on LLM called LGDEA, which constructs a diagnostic evidence space and achieves cross-modal alignment using limited paired medical vision-language data.
LLM-Guided Loop Bound Generation for Program Termination Verification
Zan Gong (Tsinghua University), Fei He (Tsinghua University)
OptimizationExplainability and InterpretabilityComputational EfficiencyAI Code AssistantTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposes the LIFT framework, which utilizes large language models to generate loop bounds and combines them with a formal verifier to prove program termination.
LLM-MatLogic: Executable Exchange Contracts for Knowledge-Graph Query Answering with Scoped Negation
Dezhuang Miao (Beihang University), Yirui Qi (Beihang University)
Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextGraphBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposed an Executable Exchange Contract (EEC) and the corresponding matrix executor MATLOGIC for precise handling of scoped negation and exclusion in knowledge graph question answering.
LLM4Branch: Large Language Model for Discovering Efficient Branching Policies of Integer Programs
Zhinan Hou (Tsinghua University), Keyou You (Tsinghua University)
OptimizationTransformerLarge Language ModelPrompt EngineeringText
🎯 What it does: Leverage large language models to automatically generate and optimize branching strategies in order to improve the efficiency of solving mixed integer linear programming problems
LLM4Cov: Execution-Aware Agentic Learning for High-coverage Testbench Generation
Hejia Zhang (University of California San Diego), Jishen Zhao (University of California San Diego)
OptimizationFederated LearningAI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AIPrompt EngineeringTextTabularSequentialBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose an offline execution-aware LLM agent learning framework called LLM4Cov, used for generating testbench for high coverage hardware verification.
LLMInertia: Adaptive Counter-Inertial Reasoning to Improve Evidence Faithfulness in Large Language Models
Xinxin You (Tsinghua University), Ji Wu (Tsinghua University)
Explainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextBiomedical DataBenchmarkChain-of-Thought
🎯 What it does: Through systematic experiments, this paper finds that large language models develop 'cognitive inertia' during the pre-training phase towards high-frequency co-occurrence relationships, leading the models to persist in their original associations even when faced with conflicting input evidence, thus generating hallucinations; subsequently, the paper proposes the LLMInertia framework, which actively detects the model's internal co-occurrence preferences during inference, automatically generates targeted 'anti-inertia reminders', and injects them into the prompts to guide the model to follow the input evidence, significantly reducing hallucination rates and improving accuracy.
LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws
Xu Ouyang (University of Virginia), Yiyuan Ma (ByteDance Seed)
Explainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningTextBenchmark
🎯 What it does: View the training of large language models as a noisy channel, and propose the Shannon Scaling Law based on the Shannon-Hartley theorem, uniformly describing the scaling patterns of both monotonic and U-shaped non-monotonic behaviors.
LLMs Lean on Priors, Not Programming Language Semantics
Aditya Thimmaiah (The University of Texas at Austin), Milos Gligoric (The University of Texas at Austin)
AI Code AssistantTransformerLarge Language ModelPrompt EngineeringTextBenchmarkChain-of-Thought
🎯 What it does: This paper investigates whether large language models (LLMs) can perform program execution reasoning under given formal semantic rules, and proposes the PLSEMANTICSBENCH benchmark to systematically evaluate the model's rule conditioning capability under different semantic transformations and program complexities.
LMCleaner: Efficient and Certified Online Unlearning via Influence Propagation Truncation
Jie Xu (City University of Hong Kong), Xiaohua Jia (City University of Hong Kong)
Federated LearningSafty and PrivacyComputational EfficiencyTransformerLarge Language ModelContrastive LearningGaussian SplattingTextBenchmark
🎯 What it does: Propose an online no-learning framework named LMCLEANER, which can instantly respond to deletion requests during the training of large language models, and achieve efficient and certified no-learning through influence propagation truncation and subspace noise.
LMM4-IC4K: A Large Multimodal Model Powered Integrated Circuit Footprint Geometry Understanding
Yida Wang (Shanghai Jiaotong University), Mahanth Gowda (Pennsylvania State University)
RecognitionSegmentationData SynthesisComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningVision Language ModelImageTextMultimodalityBenchmarkChain-of-Thought
🎯 What it does: Built a framework specifically for mechanical graphics understanding of IC footprints, and proposed the LMM4-IC4K model;
LoBCD-GW: A Fast and Data-Dependent Algorithm for Computing Gromov-Wasserstein Distance via Localized Block Coordinate Descent
Jingni Song (University of Science and Technology of China), Hu Ding (University of Science and Technology of China)
OptimizationComputational EfficiencyData-Centric LearningContrastive LearningGraphTabularTime Series
🎯 What it does: Propose a method utilizing Local Block Coordinate Descent (LoBCD-GW) to efficiently compute the Gromov-Wasserstein (GW) distance, significantly reducing the per-iteration complexity from O(n³) to O(r³) (r ≪ n)
LOCA-bench: Benchmarking Language Agents Under Controllable and Extreme Context Growth
Weihao Zeng (Hong Kong University of Science and Technology), Junxian He (Hong Kong University of Science and Technology)
TransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose the LOCA-bench benchmark, which uses adjustable environmental description lengths to evaluate the performance of language agents under extreme long contexts, and provides a pluggable context engineering tool.
Local Constrained Bayesian Optimization
Jingzhe Jing (Chinese Academy of Sciences), Qingpei Hu (Chinese Academy of Sciences)
OptimizationReinforcement LearningTabular
🎯 What it does: Propose a local constraint Bayesian optimization framework named LCBO, which utilizes a penalty function and a local exploration-exploitation strategy to address high-dimensional constrained optimization problems.
Local Covariate Selection for Average Causal Effect Estimation without Pretreatment and Causal Sufficiency Assumptions
Zeyu Liu (Beijing Technology and Business University), Kun Zhang (Carnegie Mellon University)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningGraph Neural NetworkTabular
🎯 What it does: Studies the use of local learning methods for covariate selection under the assumptions of no pretreatment and causal sufficiency, in order to achieve unbiased estimation of the average causal effect.
Local Hessian Spectral Filtering for Robust Intrinsic Dimension Estimation
GENKI OSADA
Anomaly DetectionOptimizationComputational EfficiencyRepresentation LearningDiffusion modelScore-based ModelImagePoint CloudTabularBenchmark
🎯 What it does: Propose Local Hessian Spectral Dimension (LHSD), estimating the local intrinsic dimension by performing spectral filtering on the logarithmic density Hessian, and achieving linear scalability through Stochastic Lanczos Quadrature.
Local Intrinsic Dimension of Representations Predicts Alignment and Generalization in AI Models and Human Brain
Junjie Yu (Southern University of Science and Technology), Quanying Liu (Southern University of Science and Technology)
Explainability and InterpretabilityRepresentation LearningConvolutional Neural NetworkTransformerContrastive LearningImageBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Studied the alignment between AI models and human brain representations, finding that models with better generalization performance exhibit more similar representations both between models and with the human brain, and revealed the geometric origin of this phenomenon through Local Intrinsic Dimension (LID).
Local Linearity of LLMs Enables Activation Steering via Model-Based Linear Optimal Control
Julian Skifstad (Georgia Institute of Technology), Glen Chou (Georgia Institute of Technology)
OptimizationExplainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningTextOrdinary Differential Equation
🎯 What it does: This paper proposes a model-based linear optimal control framework (A-LQR) for real-time adjustment of the activation layer in LLMs, and demonstrates that Transformer layers can be approximated as locally linear near reachable activations, enabling closed-loop feedback control;
Local MAP Sampling for Diffusion Models
Shaorong Zhang (University of California, Riverside), Greg Ver Steeg (University of California, Riverside)
RestorationOptimizationDiffusion modelScore-based ModelImagePoint CloudMagnetic Resonance ImagingComputed TomographyBenchmark
🎯 What it does: Propose and implement Local MAP Sampling (LMAPS), a method that solves inverse problems during the inference process of diffusion models by iteratively solving local MAP subproblems.
Local Mechanisms of Compositional Generalization in Conditional Diffusion
Arwen Bradley (Apple)
GenerationData SynthesisConvolutional Neural NetworkDiffusion modelScore-based ModelContrastive LearningImage
🎯 What it does: Theoretical and experimental study on the compositional generation ability of conditional diffusion models, proposing the concept of local conditional scores and verifying the relationship between length generalization and locality.
Local Minima in Quadratic-Penalty Relaxations of Binary Linear Programs
Cheng-Han Huang (Michigan State University), Rongrong Wang (Michigan State University)
Optimization
🎯 What it does: Investigated whether the local optima obtained by gradient projection methods in quadratic penalty relaxation of binary linear programming are necessarily feasible binary solutions, and provided sufficient structural conditions for feasibility.
Local Policies for Graph-Structured Markov Decision Processes
Fathima Zarin Faizal (Massachusetts Institute of Technology), Martin J Wainwright
Graph Neural NetworkReinforcement LearningGraph
🎯 What it does: Studied approximating the global optimal policy using local policies that rely only on m-hop neighborhood information in graph-structured Markov Decision Processes (MDPs), and proved that the approximation error decays exponentially with m when the spectral radius of the influence matrix meets certain conditions.
Local Redundancy: An Information-Theoretic Measure of Plasticity from Synthetic Memorization
Jiaxuan Cheng (Massachusetts Institute of Technology)
ClassificationCompressionOptimizationFederated LearningExplainability and InterpretabilityRepresentation LearningSupervised Fine-TuningDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageTabularTime Series
🎯 What it does: Proposed local redundancy, an information-theoretic measure, to quantify the plasticity of neural networks at the current parameter state;
Local-Minima-Preserving Polynomial Relaxation of Ising Problems
Debraj Banerjee (Indian Institute of Science), Kunal N. Chaudhury (Indian Institute of Science)
OptimizationPhysics Related
🎯 What it does: This paper proposes MiP-CRIM, a continuous optimization framework based on polynomial attractor relaxation, which can maintain the first-order local optimality of Ising problems in the continuous domain.
Localize and Neutralize: Gradient-Guided Token Suppression Against Visual Prompt Injection Attack
Dongpeng Zhang (University of Chinese Academy of Sciences), Qingming Huang (University of Chinese Academy of Sciences)
Safty and PrivacyExplainability and InterpretabilityComputational EfficiencyAdversarial AttackTransformerPrompt EngineeringVision Language ModelImageTextMultimodality
🎯 What it does: Proposes a lightweight inference-time defense mechanism called Gradient Token Masking (GTM) to counteract perturbation-based visual prompt injection attacks.
Localized, High-resolution Geographic Representations with Slepian Functions
Arjun Rao (University of Colorado Boulder), Marc Rußwurm (University of Bonn)
ClassificationImage TranslationImage HarmonizationRestorationObject DetectionObject TrackingSegmentationGenerationData SynthesisPose EstimationDepth EstimationSuper ResolutionRetrievalCompressionDomain AdaptationRecommendation SystemAnomaly DetectionAutonomous DrivingOptimizationFederated LearningSafty and PrivacyExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningAdversarial AttackHyperparameter SearchData-Centric LearningRobotic IntelligenceMeta LearningDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackNeural Architecture SearchProtein Structure PredictionDiffusion modelScore-based ModelFlow-based ModelRectified FlowNeural Radiance FieldAuto EncoderGenerative Adversarial NetworkContrastive LearningImageMultimodalityPoint CloudGraphTabularTime SeriesPhysics Related
🎯 What it does: To address the learning representation of geographic coordinates, this paper proposes a localized high-resolution position encoder based on Slepian functions, and presents an architecture that combines it with global spherical harmonic bases (SH). This approach significantly improves the performance of local geographic prediction tasks while ensuring pole safety and maintaining spherical distances.
Localizing Memorized Regions in Diffusion Models via Coordinate-Wise Curvature Differences
Gwangho Kim (Hanyang University), Sungyoon Lee (Hanyang University)
GenerationExplainability and InterpretabilityDiffusion modelScore-based ModelImage
🎯 What it does: Study how to locate local memorized regions in images generated by diffusion models, and provide a geometrically interpretable method to detect and locate these regions.
Locally Coherent Parallel Decoding in Diffusion Language Models
Michael Hersche (IBM Research - Zurich), Abbas Rahimi (IBM Research - Zurich)
Computational EfficiencyAI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningDiffusion modelText
🎯 What it does: Propose the CoDiLA framework, which combines discrete diffusion language models (DLM) with local autoregressive models (AR) to achieve parallel decoding while maintaining local consistency;
Locate then Correct: Debiasing Attention Heads in CLIP
Wei Jie Yeo (Nanyang Technological University), Ranjan Satapathy (Institute of High Performance Computing A STAR)
Domain AdaptationExplainability and InterpretabilityKnowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningImageTextMultimodality
🎯 What it does: Propose the LOCATE-THEN-CORRECT (LTC) framework, which identifies and corrects attention heads carrying bias in the Vision Transformer of CLIP, to enhance the model's robustness across different subpopulations.
LoCoT2V-Bench: Benchmarking Long-Form and Complex Text-to-Video Generation
Xiangqing Zheng (Harbin Institute of Technology), Min Zhang (Harbin Institute of Technology)
GenerationData SynthesisTransformerLarge Language ModelPrompt EngineeringVision Language ModelDiffusion modelVideoTextMultimodalityBenchmark
🎯 What it does: Proposes LoCoT2V-Bench, a benchmark for evaluating complex text-to-video models in long video generation.
Log-Normal Multiplicative Dynamics for Stable Low-Precision Deep Learning
Keigo Nishida (RIKEN Center for Biosystems Dynamics Research), Thomas Möllenhoff
OptimizationComputational EfficiencyRepresentation LearningTransformerLarge Language ModelContrastive LearningImageText
🎯 What it does: Proposed a new optimizer called Log-Normal Multiplicative Dynamics (LMD) for training large-scale models from scratch (ViT and GPT-2), achieving stable convergence under low-precision forward multiplication (MXFP6).
Logarithmic Switching Regret for Online Convex Optimization
Wenhao Yang (Nanjing University), Lijun Zhang (Nanjing University)
OptimizationMeta LearningTabular
🎯 What it does: Propose a new meta-algorithm called IRESET for online convex optimization in non-stationary environments, specifically optimized for switching regret.
Logical Guidance for the Exact Composition of Diffusion Models
Francesco Alesiani (Nec Laboratories Europe), Mathias Niepert (Nec Laboratories Europe)
GenerationDrug DiscoveryDiffusion modelScore-based ModelImageBiomedical DataStochastic Differential Equation
🎯 What it does: Designed and implemented a logic-guided framework called LOGDIFF, which can precisely synthesize complex Boolean logic expressions during the inference of diffusion models, supporting operations such as AND, OR, and NOT.
LogicSAGE: Neuro-Symbolic Reasoning with Socratic-Guided Enhancement
Jinlong Tian (National University of Defense Technology), Shixuan Liu (National University of Defense Technology)
Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Constructed a dual-process neuro-symbolic reasoning framework called LogicSAGE, which utilizes the Socratic Agent to achieve dynamic error correction and active verification, thereby improving the accuracy and explainability of LLM logical reasoning.
Logit Distance Bounds Representational Similarity
Beatrix Miranda Ginn Nielsen (IT University of Copenhagen), Simon Buchholz (Max Planck Institute for Intelligent Systems)
Explainability and InterpretabilityKnowledge DistillationRepresentation LearningContrastive LearningImageText
🎯 What it does: Studied using logit distance to measure distribution similarity within a general discriminative model family, and proved that this distance can quantitatively bound the linear similarity of internal representations; simultaneously proposed a representation dissimilarity (d_rep) based on linear identifiability and provided its upper bound with logit distance; based on this theory, proposed a distillation method targeting logit distance, and verified through experiments that it better preserves the teacher's linear representations and interpretable concepts compared to traditional KL distillation.
Logit-Attention Divergence: Mitigating Position Bias in Multi-Image Retrieval via Attention-Guided Calibration
Mingtao Xian (Shanghai Jiao Tong University), Nanyang Ye (Shanghai Jiao Tong University)
RetrievalTransformerContrastive LearningImageMultimodality
🎯 What it does: Proposes a training-free, attention-based calibration framework to eliminate position bias in multi-image retrieval.
LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining
Qiuwu Chen (AIGCode), Mingkui Tan (South China University Of Technology)
Computational EfficiencyRepresentation LearningTransformerLarge Language ModelMixture of ExpertsText
🎯 What it does: Designed a novel large language model architecture called LoKiFormer, which combines Local Fusion Attention (LFA) and Knowledge Memory Module (KMM) to improve pre-training efficiency.
Long Grounded Thoughts: Synthesizing Visual Problems and Reasoning Chains at Scale
David Acuna (NVIDIA), Yejin Choi (NVIDIA)
GenerationData SynthesisExplainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringVision Language ModelImageTextMultimodalityRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposes a two-stage visual question and reasoning chain generation framework that can synthesize over 1 million visual-centric questions containing reasoning trajectories, preference data, and instruction prompts on a large scale.
Long Live The Balance: Information Bottleneck Driven Tree-based Policy Optimization
Hao Jiang (Alibaba Cloud Computing), Minying Zhang (Alibaba Cloud Computing)
OptimizationTransformerLarge Language ModelReinforcement LearningTextChain-of-Thought
🎯 What it does: Proposes the IB-Score metric based on the information bottleneck theory to measure the exploration-exploitation balance in large models during online reinforcement learning, and designs the IB-TPO framework with this goal, combining IBTree tree sampling to achieve more efficient thinking path search and reward estimation.
Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference
Siheng Xiong (Georgia Institute of Technology), Yae Jee Cho (Google)
Computational EfficiencyTransformerLarge Language ModelText
🎯 What it does: Proposed a dynamic sparse attention mechanism called Dynamic Hierarchical Sparse Attention (DHSA), for achieving long-context LLM inference in environments with limited GPU memory.
Long-Horizon Model-Based Offline Reinforcement Learning Without Explicit Conservatism
Tianwei Ni (Mila -Quebec AI Institute), Pierre-Luc Bacon (Mila -Quebec AI Institute)
Recurrent Neural NetworkTransformerReinforcement LearningMixture of ExpertsWorld ModelTabularTime SeriesSequentialBenchmark
🎯 What it does: Propose a Bayesian offline reinforcement learning framework called NEUBAY without explicit conservatism constraints, which achieves test-time adaptability by utilizing history-dependent policies and model posterior inference.
Long-term Fairness with Selective Labels
Giovani Valdrighi (Universidade Estadual de Campinas), Marcos M. Raimundo (Universidade Estadual de Campinas)
Anomaly DetectionAutonomous DrivingOptimizationFederated LearningExplainability and InterpretabilityData-Centric LearningSupervised Fine-TuningReinforcement LearningDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningTabularTime SeriesSequentialFinance Related
🎯 What it does: This paper studies the long-term fairness problem under the scenario where labels are only observable when they are accepted, proposing a framework based on label predictors and designing a reinforcement learning algorithm based on PPO (SELLF) to achieve long-term fair decisions.
LongCoT: Benchmarking Long-Horizon Chain-of-Thought Reasoning
Sumeet Ramesh Motwani (University of Oxford), Christian Schroeder de Witt (University of Oxford)
TransformerLarge Language ModelTextBenchmarkChain-of-Thought
🎯 What it does: Propose the LongCoT benchmark to evaluate the ability of language models in long-horizon chain-of-thought reasoning.
Look on Demand: A Cognitive Scheduling Framework for Visual Evidence Acquisition in Multimodal Reasoning
Yang Zhang (Xiamen University), Rongrong Ji (Xiamen University)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmarkChain-of-Thought
🎯 What it does: Proposes a cognitive scheduling framework called CSMR, which uses a language model to dynamically schedule visual perception for multi-modal reasoning.
Lookahead Path Likelihood Optimization for Diffusion LLMs
Xuejie Liu (Peking University), Anji Liu (National University Of Singapore)
OptimizationComputational EfficiencyAI Code AssistantTransformerLarge Language ModelDiffusion modelText
🎯 What it does: Proposes a decoding path optimization framework based on the path log-likelihood (Path LL), and designs an approximate estimator of the expected future path log-likelihood called POKE. It then embeds POKE into the Sequential Monte Carlo (SMC) search to form POKE-SMC, dynamically selecting the optimal step-by-step decoding path during the inference process of Diffusion LLM.
Lookahead Sample Reward Guidance for Test-Time Scaling of Diffusion Models
Yeongmin Kim (Korea Advanced Institute of Science and Technology), Il-chul Moon
GenerationComputational EfficiencyTransformerReinforcement LearningPrompt EngineeringDiffusion modelScore-based ModelImageTextMultimodalityStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: This paper proposes the LiDAR (Lookahead Sample Reward Guidance) sampling method, which achieves sampling in high-reward regions during the test phase of diffusion models through lookahead sampling and reward guidance, thereby improving the consistency between generated samples and human intent.
Lookahead Unmasking Elicits Reliable Decoding in Diffusion Language Models
Sanghyun Lee (KAIST), Dongmin Park (KRAFTON)
GenerationComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringDiffusion modelScore-based ModelText
🎯 What it does: Propose the Lookahead Unmasking (LookUM) framework, which utilizes an internal uncertainty validator to guide multi-path reasoning during the denoising phase of diffusion language models, thereby significantly reducing local errors.
Lookahead-GCG: Improving Universal Multi-Model Optimization-Based Jailbreaking Attacks via Stochastic Nesterov Optimization
Rong Feng (Jilin University), Runsheng Yu (Jilin University)
OptimizationAdversarial AttackTransformerLarge Language ModelPrompt EngineeringTextStochastic Differential Equation
🎯 What it does: Propose a general multi-model jailbreak attack framework based on Lookahead-GCG, which improves the cross-model transfer performance of the attack by utilizing stochastic Nesterov accelerated gradient (SNAG) in the discrete token space, embedding space momentum accumulation, and maximum distance initialization.
Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation
Manuel Traub (University of Tübingen), Martin V. Butz (University of Tübingen)
SegmentationComputational EfficiencyTransformerContrastive LearningGaussian SplattingImage
🎯 What it does: Proposed the FLIP model, achieving efficient single-object segmentation through multi-resolution, fovea-style input sampling based on the target object
LoPhyDA: Low-Rank Tensor and Physics Gradient Guided Diffusion for Atmospheric Data Assimilation
Danyang Peng (Nanjing University), Xiaotong Yuan
OptimizationTransformerDiffusion modelScore-based ModelAuto EncoderTabularTime SeriesPhysics Related
🎯 What it does: Propose a dual-guided diffusion-based atmospheric data assimilation method called LoPhyDA, combining low-rank tensor observation reconstruction and physical gradient guidance;
LoRA-DA: Data-Aware Initialization for Low-Rank Adaptation via Asymptotic Analysis
Qingyue Zhang (Tsinghua University), Shao-Lun Huang (Tsinghua University)
Domain AdaptationFederated LearningRepresentation LearningData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningText
🎯 What it does: This paper proposes a LoRA data-aware initialization method called LoRA‑DA, based on asymptotic analysis, aiming to provide a better low-rank subspace initialization for LoRA (including LoRA‑FA) using only a small number of target domain samples.
LORD-GoF: A Robust Online Detection Approach for LLM Watermarks in Sparse and Mixed Streams
Jiade Xu (Lanzhou University), Zhouping Li (Lanzhou University)
Anomaly DetectionData-Centric LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningText
🎯 What it does: Proposes the LORD‑GoF framework, combining the Goodness‑of‑Fit (GoF) statistic with LORD online false discovery rate (FDR) control, to achieve watermark detection in sparse and mixed human-machine text streams.
LoRDO: Distributed Low-Rank Optimization with Infrequent Communication
Andrej Jovanovic, Nicholas D. Lane (University of Cambridge)
OptimizationFederated LearningComputational EfficiencyLarge Language ModelAuto EncoderContrastive LearningText
🎯 What it does: Propose LoRDO, a framework that uses low-rank optimization with low communication frequency in distributed training;
LoRe: Adaptive Interaction-Evaluation Routing with Per-step Interaction Budgets for Iterative Graph Solvers
Jintao Li (Beijing Academy of Quantum Information Sciences), Heng Fan (Institute of Physics, Chinese Academy of Sciences)
OptimizationComputational EfficiencyGraph Neural NetworkTransformerDiffusion modelScore-based ModelContrastive LearningGraphTabularTime SeriesBenchmarkPhysics Related
🎯 What it does: Introduce a training-agnostic, inference-time 'LoRe' package for iterative graph solvers, which forces each step to evaluate only a fixed proportion of high-conflict or high-uncertainty interactions, thereby achieving significant computational and memory compression without modifying the original model.
LoSA: Locality Aware Sparse Attention in Diffusion Language Models
Haocheng Xi (University of California Berkeley), Amir Gholami (University of California Berkeley)
Computational EfficiencyTransformerLarge Language ModelDiffusion modelText
🎯 What it does: This paper proposes a Locality-Aware Sparse Attention (LoSA), which performs sparse attention computation only on tokens with significant changes during the block-level decoding process of diffusion language models, while reusing the prefix attention results from the previous step for stable tokens.
Loss-Aware Distributionally Robust Optimization via Trainable Optimal Transport Ambiguity Sets
Jonas Ohnemus (ETH Zürich), John Lygeros (ETH Zürich)
OptimizationFederated LearningData-Centric LearningTabularFinance Related
🎯 What it does: Designed and implemented an end-to-end OT-DRO framework that automatically constructs decision-oriented fuzzy sets by learning transportation cost parameters, achieving a less conservative robust optimization.
Lost in Context: Adressing Context Anxiety in Large Language Models
Ifueko Igbinedion (Massachusetts Institute of Technology), Eric So (Massachusetts Institute of Technology)
Explainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelSupervised Fine-TuningText
🎯 What it does: Studied the 'contextual anxiety' of large language models — models prematurely giving up on solvable problems due to overestimated token requirements, and proposed methods to detect, measure, and alleviate this phenomenon.
Lottery Prior: Randomized Neural Compression for Zero-Shot Inverse Problems
Haotian Wu (Zhejiang University), Deniz Gunduz (Imperial College London)
RestorationSuper ResolutionCompressionDiffusion modelScore-based ModelAuto EncoderImageMagnetic Resonance Imaging
🎯 What it does: Propose Lottery Prior, a zero-shot inverse problem solver based on random networks and compression constraints.
LOTTERY: Learning from Reference-Only Samples in Two-Sample Testing under Size Asymmetry
Xunye Tian (University of Melbourne), Feng Liu (University of Melbourne)
Anomaly DetectionRepresentation LearningData-Centric LearningContrastive LearningImageMultimodalityTabularPhysics Related
🎯 What it does: This paper proposes the LOTTERY framework, which learns multiple reference-dependent representations (RDR) using only reference samples, and performs unbiased two-sample testing through pooled permutation testing without querying data splitting, particularly suitable for extremely imbalanced sample and few querying sample scenarios.
LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation
Jiarui Wang (Shanghai Jiao Tong University), Xiongkuo Min (Shanghai Jiao Tong University)
GenerationData SynthesisExplainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision Language ModelVideoTextMultimodalityBenchmark
🎯 What it does: Constructed the largest AI-generated video evaluation dataset, AIGVE-60K, and proposed the LMM-driven LOVE evaluation framework and LOVE-Reward reward strategy to multidimensionally evaluate and improve text-to-video generation models;
Low Kruskal-Rank Adaptation
Yixing Xu (Advanced Micro Devices, Inc.), Emad Barsoum (Advanced Micro Devices, Inc.)
OptimizationComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningText
🎯 What it does: Propose two novel parameter-efficient fine-tuning algorithms, LoKRA and LoKRA+, which use low Kruskal rank optimization on the LoRA update matrix, significantly improving the performance of LLM downstream tasks.
Low-Compute Watermark Removal via Dual-Domain Natural Projection
Pragati Shuddhodhan Meshram (University of Illinois Urbana-Champaign), Varun Chandrasekaran (University of Illinois Urbana-Champaign)
RestorationConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderImage
🎯 What it does: Propose a lightweight, training-free single-image watermark removal method called DAWN, which suppresses watermarks and restores image quality simultaneously by utilizing dual-domain natural projection.
Low-cost Full Fine-tuning: Learning What to Update for LLMs
Jin Li (Hong Kong University of Science and Technology), Hui Xiong (Hong Kong University of Science and Technology)
Computational EfficiencyKnowledge DistillationTransformerLarge Language ModelSupervised Fine-TuningText
🎯 What it does: This paper proposes a low-cost full-parameter fine-tuning (LC-FT) framework that utilizes the Think-Touch thinking-touch paradigm to update only a limited number of parameter groups at each step, thereby significantly reducing memory and computational overhead while maintaining full expressiveness.
Low-dimensional topology of deep neural networks
Junyu Ren (University of Chicago), Lek-heng Lim
Explainability and InterpretabilityRepresentation LearningConvolutional Neural NetworkRecurrent Neural NetworkTransformerAuto EncoderContrastive LearningImagePoint CloudGraph
🎯 What it does: The study investigates how low-dimensional topological invariants (such as the number of chains) change across network layers in depth networks with a fixed width of 3 (including feedforward, ResNet, and Transformer architectures), revealing the expressive power of different architectures in handling topologically coupled data.
Low-Rank and Sparsity Are All You Need: Exploring Robust Hierarchical Latent Subspaces for Transferable Adversarial Attack
Shuangshuang Pu (Chongqing University of Technology), Di Ming (Chongqing University of Technology)
Representation LearningAdversarial AttackConvolutional Neural NetworkTransformerMixture of ExpertsImage
🎯 What it does: Designed and implemented an adversarial attack method based on low-rank and sparse decomposition (LRS-Attack), enhancing the transferability of black-box attacks by extracting global semantic subspaces and local discriminative patterns in the intermediate feature space.
Lower Bounds for Frank-Wolfe on Strongly Convex Sets
Jannis Halbey (Zuse Institute Berlin), Sebastian Pokutta (Zuse Institute Berlin)
Optimization
🎯 What it does: Minimize a strongly convex quadratic function on the Euclidean unit ball, construct the worst-case Frank-Wolfe trajectory, and prove that the lower bound of FW on strongly convex sets is Ω(1/√ε).
Lower Complexity Bounds for Nonconvex-Strongly-Convex Bilevel Optimization with First-Order Oracles
Kaiyi Ji (University at Buffalo)
OptimizationDiffusion modelScore-based ModelFlow-based ModelRectified FlowContrastive LearningStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: In the paper, the authors construct a series of worst-case instances to prove that when solving smooth non-convex-strongly convex bilevel optimization problems using standard first-order oracles (deterministic and stochastic), any first-order zero-responsiveness algorithm requires at least Ω(κ^{3/2}/ε^2) calls (deterministic) and Ω(κ^{5/2}/ε^4) calls (stochastic), thus providing new lower bounds;
LOZO+: Provably Efficient Zeroth-Order Fine-Tuning via Greedy Low-Rank Subspace Selection
Jinjie Fang (Jilin University), Bin Gu (Jilin University)
OptimizationFederated LearningComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningText
🎯 What it does: Propose LOZO+, a zeroth-order optimization method based on greedy low-rank subspace selection, for memory-efficient fine-tuning of large language models.
LRAgent: Efficient KV Cache Sharing for Multi-LoRA LLM Agents
Hyesung Jeon (Seoul National University), Jae-Joon Kim (Seoul National University)
Computational EfficiencyTransformerLarge Language ModelAgentic AIText
🎯 What it does: This paper proposes the LRAgent framework, which achieves efficient sharing of KV cache among multiple LoRA role agents, significantly reducing memory and computational overhead.
LS$^{2}$MC-GDA: A Smoothed Algorithm for Federated Stochastic Multi-Level Compositional Minimax Optimization
Xinwen Zhang (Temple University), Hongchang Gao (Temple University)
OptimizationFederated LearningConvolutional Neural NetworkContrastive LearningImageStochastic Differential Equation
🎯 What it does: Proposed the LS-MC-GDA algorithm for solving stochastic multi-layer compositional min-max optimization problems in federated learning scenarios.
LSGQuant: Layer-Sensitivity Guided Quantization for One-Step Diffusion Real-World Video Super-Resolution
Tianxing Wu (Shanghai Jiao Tong University), Yulun Zhang (Shanghai Jiao Tong University)
Super ResolutionCompressionTransformerDiffusion modelScore-based ModelVideo
🎯 What it does: Propose the LSGQuant framework, which applies first-order diffusion models to video super-resolution and achieves low-bit quantization compression of the model;
LUCID: Attention with Preconditioned Representations
Sai Surya Duvvuri (University of Texas at Austin), Inderjit S Dhillon
RetrievalComputational EfficiencyRepresentation LearningTransformerLarge Language ModelContrastive LearningTextRetrieval-Augmented Generation
🎯 What it does: Propose LUCID Attention, which eliminates key correlations in RKHS by preconditioning key-key similarity, achieving precise and efficient long-sequence attention.
LUGS: Latent-aware Guidance for Efficient Unmasking in Diffusion Large Language Models
Nuanqiao Shan (Zhejiang University), Kun Kuang (Zhejiang University)
Computational EfficiencyRepresentation LearningAI Code AssistantTransformerLarge Language ModelDiffusion modelScore-based ModelTextRetrieval-Augmented Generation
🎯 What it does: Propose the LUGS framework, which utilizes the hidden states of Diffusion LLM for latent-aware guidance to improve the order of token unmasking.
LUVE : Latent-Cascaded Ultra-High-Resolution Video Generation with Dual Frequency Experts
Chen Zhao (Nanjing University), Ying Tai (Nanjing University)
GenerationSuper ResolutionTransformerPrompt EngineeringMixture of ExpertsDiffusion modelScore-based ModelAuto EncoderVideo
🎯 What it does: Propose a three-stage Latent-Cascaded framework, first generating low-resolution motion, then up-sampling in the latent space, and finally using dual-frequency experts to enhance semantic consistency and detail quality, directly generating ultra-high-resolution videos.
M-IDoL: Information Decomposition for Modality-Specific and Diverse Representation Learning in Medical Foundation Model
Yihang Liu (Tongji University), Heng Tao Shen (Tongji University)
Representation LearningData-Centric LearningTransformerMixture of ExpertsContrastive LearningImageMultimodalityBiomedical DataComputed Tomography
🎯 What it does: Proposes M-IDoL, a self-supervised medical foundation model that achieves modality-specific and diverse representation learning through information decomposition.
M-Star: Markovian Projection of Star-Shaped Diffusion for Exponential Family Distributions
François Bertholom (Institut Polytechnique de Paris), Khalid Oublal (Institut Polytechnique de Paris)
GenerationData SynthesisKnowledge DistillationDiffusion modelScore-based ModelContrastive LearningGaussian SplattingImageVideoTabularTime SeriesBenchmark
🎯 What it does: Propose a diffusion framework named M-Star, which utilizes Markov projection to convert the star-shaped forward process into a labelable Markov process, and restores temporal coherence during the reverse sampling;
M+Adam: Low-Precision Training via Additive–Multiplicative Optimization
Xiaoyuan Liang (California Institute of Technology), Anima Anandkumar (California Institute of Technology)
OptimizationComputational EfficiencyHyperparameter SearchTransformerLarge Language ModelMixture of ExpertsContrastive LearningText
🎯 What it does: Propose an optimizer called M+Adam, which combines the additive updates of Adam and the multiplicative updates of Madam, for training with low-precision weights. It avoids two failure modes: additive updates being lost due to rounding in low precision and multiplicative updates failing near zero.
MA$^3$S: Model-Agnostic Active Annotation Strategy for Crowdsourcing
Wenjun Zhang (Zhongnan University of Economics and Law), Shanshan Si (China University of Geosciences)
Federated LearningData-Centric LearningContrastive LearningImageTextTabular
🎯 What it does: Propose a model-agnostic active repeated annotation strategy, MA S, aiming to reduce label redundancy and instance redundancy in crowdsourcing annotation, and improve annotation efficiency through online updates.
MAC-NeRF: Motion-Aware Curriculum Learning for Dynamic LiDAR NeRFs
Shangshu Yu (Northeastern University), Cheng Wang (Xiamen University)
Autonomous DrivingRectified FlowNeural Radiance FieldAuto EncoderGenerative Adversarial NetworkContrastive LearningOptical FlowVideoPoint Cloud
🎯 What it does: This paper proposes MAC-NeRF, a NeRF framework for dynamic LiDAR scenes, achieving high-fidelity view synthesis through motion-aware curriculum learning.
MACD: Model-Aware Contrastive Decoding via Counterfactual Data for Video-LLMs
Qixin Xiao (University of Michigan), Kun Zhou (University of California San Diego)
Explainability and InterpretabilityComputational EfficiencyAdversarial AttackData-Centric LearningTransformerPrompt EngineeringVision Language ModelGenerative Adversarial NetworkContrastive LearningVideoText
🎯 What it does: During the inference of video large language models, optimize object- and frame-level occlusion masks using the model's own loss signal to generate targeted adversarial videos, and use them together with the original video for contrastive decoding, reducing hallucinations and improving task accuracy.
Machine Learning Hamiltonians are Accurate Energy-Force Predictors
Seongsu Kim (Korea Advanced Institute of Science and Technology), Sungsoo Ahn (Korea Advanced Institute of Science and Technology)
Drug DiscoveryGraph Neural NetworkTransformerFlow-based ModelGraphTabularBenchmarkPhysics Related
🎯 What it does: Proposed a machine learning Hamiltonian model QHFlow2 that can be directly used for computing energy and forces, and constructed a unified benchmark for direct evaluation.
MACKO: Sparse matrix-vector multiplication for low sparsity
Vladimír Macko (Comenius University), Vladimír Boža (Comenius University)
Computational EfficiencyLarge Language ModelTabularBenchmark
🎯 What it does: Proposed the MACKO format and the corresponding GPU SpMV implementation, improving memory and speed for LLM sparse multiplication at low sparsity rates of 30-90%.
MacroGuide: Topological Guidance for Macrocycle Generation
Alicja Maksymiuk (University of Oxford), Ismail Ilkan Ceylan (TU Wien)
Drug DiscoveryDiffusion modelScore-based ModelContrastive LearningBiomedical Data
🎯 What it does: Proposed a topology-guided diffusion model called MACROGUIDE, which generates arbitrary macrocyclic molecules and supports unconditional and protein-conditioned generation.
MAD: Manifold Attracted Diffusion
Dennis Elbrächter (University of Vienna), Matteo Santacesaria (University of Genoa)
GenerationDiffusion modelScore-based ModelImageBiomedical DataOrdinary Differential Equation
🎯 What it does: In the scenario where training data is a noisy version, a reasoning algorithm called Manifold Attracted Diffusion (MAD) is proposed, which modifies the score model during the inference phase. It utilizes extended scores to suppress noise, generating samples that are nearly noise-free.
MADA-Attack: Transferable Multi-modal Attention Distraction Adversarial Attack against Vision Language Models
Zhihan Qin (Chongqing University), Shouling Ji (Zhejiang University)
RetrievalRepresentation LearningAdversarial AttackTransformerPrompt EngineeringVision Language ModelAuto EncoderContrastive LearningImageTextMultimodality
🎯 What it does: A multi-modal attention-diverted target-free universal adversarial perturbation attack framework called MADA-Attack was developed, which can achieve high transferability attacks across tasks and models on visual language models.