ICML 2026 Papers — Page 65
International Conference on Machine Learning · 6554 papers
WEVSR: Video Diffusion Generators for Real-World Video Super‑Resolution with Wavelet-Enhanced VAE Encoder
Yuying Chen (Guangdong Laboratory of Artificial Intelligence and Digital Economy), Wenqi Ren (Sun Yatsen University)
Super ResolutionTransformerDiffusion modelFlow-based ModelAuto EncoderVideo
🎯 What it does: Adapt pre-trained flow-matched video diffusion transformer (DiT) and VAE encoder for real-world video super-resolution tasks, and enhance the frequency perception of the VAE encoder through multi-scale discrete wavelet transform (DWT).
WF-Bench: A Benchmark for Neural-Network WaveFunction Expressivity and Scaling Laws
Lixing Zhang (University of California Los Angeles), Di Luo (Tsinghua University)
Convolutional Neural NetworkRecurrent Neural NetworkTransformerContrastive LearningGraphTabularBenchmarkPhysics Related
🎯 What it does: Designed and released the WF-Bench benchmark set, which includes 31 target wave functions of different types (topological, superconducting, Wigner crystals), and proposed a unified protocol for wave function matching and evaluation, systematically assessing the representational capacity and scaling behavior of two mainstream neural network wave function architectures: Ferminet and Psiformer.
WFR-MFM: One-Step Inference for Dynamic Unbalanced Optimal Transport
Xinyu Wang (Peking University), Tiejun Li (Peking University)
Data SynthesisOptimizationComputational EfficiencyFlow-based ModelOptical FlowTabularTime SeriesBiomedical DataStochastic Differential Equation
🎯 What it does: This study proposes an unbalanced optimal transport (OT) inference framework based on average flow, which can achieve one-time rapid inference in fields such as single-cell transcriptomics;
What Characterizes Effective Reasoning? Revisiting Length, Review, and Structure of CoT
Yunzhen Feng (Meta Superintelligence Labs), Anthony Hartshorn (Meta Superintelligence Labs)
Explainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Systematically evaluate the length, review ratio, and structural characteristics of chain-of-thought (CoT) in ten large reasoning models on math and science reasoning tasks, to explore which features best predict reasoning accuracy.
What Do Agents Learn from Trajectory-SFT: Semantics or Interfaces?
Weizheng Gu (Peking University), Wei Ye (Peking University)
Explainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningAgentic AIPrompt EngineeringTextBenchmark
🎯 What it does: Propose PIPE (Perturb Interface Protocol for Evaluation) and the Interface Reliance (IR) metric to diagnose whether LLM agents trained with trajectory-supervised fine-tuning (trajectory-SFT) fail to generalize due to over-reliance on the surface form of the interface.
What Does Flow-Matching Bring to TD-Learning?
Bhavya Kumar Agrawalla (Carnegie Mellon University), Aviral Kumar (Carnegie Mellon University)
Reinforcement LearningFlow-based ModelTabularTime SeriesBenchmarkOrdinary Differential Equation
🎯 What it does: This paper investigates and verifies the advantages of flow-matching networks in temporal difference (TD) learning, pointing out that the performance improvement mainly comes from test-time recovery and plastic features, rather than distributed RL.
What Does Preference Learning Recover from Pairwise Comparison Data?
Rattana Pukdee (Carnegie Mellon University), Pradeep Ravikumar (Carnegie Mellon University)
Representation LearningData-Centric LearningReinforcement Learning from Human FeedbackSupervised Fine-TuningReinforcement LearningContrastive LearningTextReview/Survey Paper
🎯 What it does: This paper provides a centralized theoretical analysis of learning preferences from contrastive data, defining the conditional preference distribution (CPRD), and clarifying under what conditions the Bradley-Terry (BT) model can represent CPRD, while analyzing the statistical properties and sample complexity of BT learning.
What Does Thompson Sampling Optimize?
Yanlin Qu (Columbia University), assaf zeevi
OptimizationReinforcement LearningTabular
🎯 What it does: This paper reinterprets Thompson Sampling (TS) as an online optimization algorithm that minimizes squared regret, and derives the optimal policy (R₂-optimal policy) through the Bellman equation.
What Does Vision Tool-Use Reinforcement Learning Really Learn? Disentangling Tool-Induced and Intrinsic Effects for Crop-and-Zoom
Yan Ma (Fudan University), Pengfei Liu (Shanghai Jiao Tong University)
Reinforcement LearningVision Language ModelVision-Language-Action ModelImageMultimodalityAgriculture Related
🎯 What it does: Propose the MED framework, which separates and quantifies the intrinsic capability enhancement and tool-induced effects during model training through the use of reinforcement learning (Crop-and-Zoom) for visual tools.
What if Tomorrow is the World Cup Final? Counterfactual Time Series Forecasting with Textual Conditions
Shuqi Gu (ShanghaiTech University), Kan Ren (ShanghaiTech University)
GenerationData SynthesisAnomaly DetectionComputational EfficiencyRepresentation LearningTransformerSupervised Fine-TuningPrompt EngineeringDiffusion modelContrastive LearningTextMultimodalityTime SeriesBenchmarkFinance Related
🎯 What it does: Propose the counterfactual time series forecasting task under textual conditions, and introduce the TADIFF text attribute attribution diffusion model to accomplish this task.
What If We Allocate Test-Time Compute Adaptively?
Ahsan Bilal (University of Oklahoma), Dean F. Hougen (University of Oklahoma)
Computational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextChain-of-Thought
🎯 What it does: Proposes an adaptive computation allocation framework during testing based on a process reward model, dynamically selecting tools and search strategies.
What If We Let Forecasting Forget? A Sparse Bottleneck for Cross-Variable Dependencies
Fan Zhang (Shandong Technology and Business University), Hua Wang (Ludong University)
Computational EfficiencyRepresentation LearningTransformerAuto EncoderContrastive LearningTabularTime SeriesBenchmark
🎯 What it does: Proposed MS-FLOW, which improves multivariate time series forecasting by introducing a sparse bottleneck in cross-variable information flow.
What Information Matters? Graph Out-of-Distribution Detection via Tri-Component Information Decomposition
Danny Wang (University of Queensland), Zi Huang (University of Queensland)
Anomaly DetectionRepresentation LearningGraph Neural NetworkContrastive LearningGraph
🎯 What it does: Propose a three-component information decomposition framework called TIDE for out-of-distribution (OOD) detection on graph nodes, which can decompose predictive information into feature, structural, and joint components, filtering out pseudo-correlations from single inputs to enhance the ID-OOD discrimination;
What is Missing? Explaining Neurons Activated by Absent Concepts
Robin Hesse (Max Planck Institute for Informatics), Stefan Roth (Technical University of Darmstadt)
ClassificationAnomaly DetectionExplainability and InterpretabilityConvolutional Neural NetworkAuto EncoderGenerative Adversarial NetworkContrastive LearningImage
🎯 What it does: This paper investigates and reveals the phenomenon of 'missing concepts' encoded in deep neural networks, proposing two simple improvements to existing explanation methods (attribution and feature visualization) to detect and explain these missing concepts. It further verifies their universality on the ImageNet and ISIC datasets and uses them to improve model debiasing.
What Language is This? Ask Your Tokenizer.
Clara Meister (EPFL), Tiago Pimentel (ETH Zürich)
ClassificationTransformerTextBenchmark
🎯 What it does: Propose UniLID, a language identification method based on UnigramLM, which utilizes the unigram frequency distribution of each language and discriminates the language through the most probable segmentation during inference.
What Linear Probes Miss: Multi-View Probing for Weight-Space Learning
Eunwoo Heo (Ulsan National Institute of Science and Technology), Jaejun Yoo (Ulsan National Institute of Science and Technology)
Explainability and InterpretabilityRepresentation LearningConvolutional Neural NetworkTransformerPrompt EngineeringAuto EncoderContrastive LearningImage
🎯 What it does: This paper investigates the limitations of probing methods in weight space learning and proposes a multi-perspective probing framework called MVProbe, which comprehensively encodes the network weight matrix using first-order row/column and second-order Gram perspectives.
What LLMs Explain Is Not What They Believe: Evaluating Explanation Sufficiency Under Models' Own Input Beliefs
Nhi Nguyen (New York University), Rajesh Ranganath (New York University)
Information TheoryExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringTextBenchmark
🎯 What it does: Studied whether free-text explanations generated by large language models (LLMs) can sufficiently explain model outputs, and proposed a theoretical framework and measurement method to evaluate this 'sufficiency of explanation'.
What Makes a Desired Graph for Relational Deep Learning?
Yao Cheng (Nanyang Technological University), Siqiang Luo (Nanyang Technological University)
Recommendation SystemOptimizationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningGraph Neural NetworkContrastive LearningGraphTabularBenchmark
🎯 What it does: This study explores how to generate an 'ideal graph' that meets the learning requirements of graph neural networks (GNNs) from relational databases, and proposes an end-to-end structural optimizer that performs information filtering and structural injection on the original database graph to automatically generate task-specific graph topologies.
What Makes a Representation Good for Single-Cell Perturbation Prediction?
Wenkang Jiang (Adelaide University), Javen Qinfeng Shi (Adelaide University)
Explainability and InterpretabilityRepresentation LearningDrug DiscoveryAuto EncoderGenerative Adversarial NetworkContrastive LearningTabularBiomedical Data
🎯 What it does: This paper proposes the PerturbedVAE framework, which improves the generalization and interpretability of single-cell perturbation prediction by explicitly separating perturbation-irrelevant and perturbation-specific information, and by structuring the causal relationships of perturbation-specific information.
What Makes a Strong Model? A Unified Spectral Analysis of Knowledge Transfer over High-dimensional Linear Regression
Wendao Wu (Peking University), Cong Fang (Peking University)
OptimizationKnowledge DistillationRepresentation LearningSupervised Fine-TuningContrastive LearningImageTabular
🎯 What it does: This paper proposes a unified spectral analysis framework in high-dimensional linear regression, elucidating the essential mechanisms of various knowledge transfer paradigms such as knowledge distillation (KD), weak-to-strong (W2S), and self-distillation.
What Makes Effective Supervision in Latent Chain-of-Thought: An Information-Theoretic Analysis
Xinghao Chen (Eastern Institute of Technology), Xiaoyu Shen (Eastern Institute of Technology)
Information TheoryExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelContrastive LearningTextChain-of-Thought
🎯 What it does: This paper studies Latent Chain-of-Thought, which internalizes reasoning within continuous latent states, and reveals the dual collapse problem caused by gradient decay and semantic drift when only the final answer is used for supervision;
What Makes Synthetic Data Effective in Image Segmentation
Jinjin Zhang (Beihang University), Di Huang (Beihang University)
SegmentationData SynthesisTransformerDiffusion modelContrastive LearningImage
🎯 What it does: Investigated and systematically evaluated the effectiveness of synthetic images generated by large-scale diffusion models on semantic segmentation tasks, and proposed a unified framework called SENSE, which assigns robust pseudo-labels to synthetic data through optimal transport (OT) techniques, significantly improving segmentation performance.
What Makes Value Learning Efficient in Residual Reinforcement Learning?
Guozheng Ma (Nanyang Technical University), Dacheng Tao (Nanyang Technical University)
Robotic IntelligenceTransformerReinforcement LearningVision-Language-Action ModelDiffusion modelTabularSequentialBenchmark
🎯 What it does: This paper investigates the bottleneck of value learning in residual reinforcement learning and proposes the DAWN method to address the cold-start symptom and structural scale mismatch issues.
What Preferences Can—and Cannot—Predict in Multi-Agent Online Learning
Omar Abbadi (Moroccan Center for Game Theory), Panayotis Mertikopoulos (Univ Grenoble Alpes)
OptimizationFederated LearningReinforcement LearningContrastive Learning
🎯 What it does: This paper studies the constraints of preference graphs on dynamic stability in multi-agent online learning, and clarifies the relationship between preferences and learning outcomes.
What Really Improves Mathematical Reasoning: Structured Reasoning Signals Beyond Pure Code
Yuze Zhao (University of Science and Technology of China), Enhong Chen (University of Science and Technology of China)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningAI Code AssistantTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsTextRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Construct a 10T token high-quality multi-domain corpus, distinguishing between pure code and Code-NL, pretrain using a MoE model, and systematically evaluate the impact of code and mathematical data through ablation experiments with a fixed budget, discovering a competitive relationship between the two in reasoning tasks; further identify and incorporate structured mathematical 'cognitive scaffolds,' significantly improving complex mathematical reasoning performance.
What Reward Structure Enables Efficient Sparse-Reward RL? A Proof-of-Concept with Policy-Aware Matrix Completion
Ibne Farabi Shihab (Iowa State University), Anuj Sharma (Iowa State University)
OptimizationComputational EfficiencyReinforcement LearningTabularTime SeriesSequential
🎯 What it does: Proposes the Policy-Aware Matrix Completion (PAMC) framework, which accelerates learning in sparse reward reinforcement learning by learning the low-rank + sparse structure of the reward matrix.
What You Think is What You See: Driving Exploration in VLM Agents via Visual-Linguistic Curiosity
Haoxi Li (Hong Kong University of Science and Technology), Song Guo (Hong Kong University of Science and Technology)
Robotic IntelligenceTransformerReinforcement LearningVision Language ModelContrastive LearningImageTextMultimodalityChain-of-Thought
🎯 What it does: Propose the GLANCE framework, which utilizes cross-modal prediction error from vision-language alignment as intrinsic reward, driving the VLM agent to actively explore and refine its internal world model.
When Actions Go Off-Task: Detecting and Correcting Misaligned Actions in Computer-Use Agents
Yuting Ning (Ohio State University), Huan Sun (Ohio State University)
Anomaly DetectionSafty and PrivacyComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposed and implemented the DEACT framework for error behavior detection and correction in computer user agents (CUA), defining and systematically addressing the problem of erroneous actions in CUA for the first time;
When Agents Go Rogue: Activation-Based Detection of Malicious Behaviors in Multi-Agent Systems
Haowen Xu (Worcester Polytechnic Institute), Xiaoyan Sun (Worcester Polytechnic Institute)
Anomaly DetectionReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningText
🎯 What it does: Proposed a framework called AcMAS based on activation for detecting malicious behavior in multi-agent systems, addressing the limitations of existing methods when facing covert attacks and asynchronous execution.
When AI Agents Compete for Jobs: Strategic Capabilities and Economic Dynamics of AI Labour Markets
Christopher Chiu (University of Cambridge), Mihaela van der Schaar (University of Cambridge)
OptimizationFederated LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Built and used AI-Work, a simulation of a labor market, to model LLM agents bidding, training, and reputation dynamics, and to explore the impact of agents' strategic capabilities on market performance.
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Mubashara Akhtar (ETH Zurich), Irene Solaiman
Large Language ModelTextBenchmark
🎯 What it does: Studied the saturation phenomenon in AI benchmarks, proposed a saturation index based on evaluation uncertainty, and conducted systematic evaluation and attribute correlation analysis on 60 text LLM benchmarks.
When and How Human Curation Backfires: Preference Alignment under Multi-Model Self-Consuming Loop
Yang Zhang (Ohio State University), Xueru Zhang (Ohio State University)
Recommendation SystemExplainability and InterpretabilityData-Centric LearningReinforcement Learning from Human FeedbackImageMultimodality
🎯 What it does: This paper investigates the impact of human curation on model preference alignment in multi-model self-consumption training, particularly in multi-model environments, exploring how human curation influences self-impact and cross-impact of models.
When Attributes Disagree: Gradient Conflict in Image Aesthetic Assessment
Ye Wang (Chongqing University of Posts and Telecommunications), Hong Yu (Chongqing University of Posts and Telecommunications)
Image TranslationImage HarmonizationExplainability and InterpretabilityRepresentation LearningConvolutional Neural NetworkTransformerPrompt EngineeringVision Language ModelDiffusion modelContrastive LearningImageTextMultimodalityRetrieval-Augmented Generation
🎯 What it does: This paper proposes the AGREE framework, which addresses attribute conflict issues in image aesthetic assessment through attribute-oriented gradient routing.
When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents
Jaylen Jones (Ohio State University), Huan Sun (Ohio State University)
Safty and PrivacyExplainability and InterpretabilityAdversarial AttackTransformerLarge Language ModelSupervised Fine-TuningAgentic AIVision Language ModelTextMultimodalityBenchmarkChain-of-Thought
🎯 What it does: Proposes a conceptual framework for systematically identifying and studying unintentional unsafe behaviors in computer usage agents (CUAs), and develops the AUTOELICIT automated extraction tool, which actively elicits and records severe adverse behaviors generated by CUAs in real, harmless scenarios by fine-tuning normal instructions and combining execution feedback.
When Can We Trust Survival Model Evaluation ?
Ghanem BAHRINI, Morgane Barbet-Massin (SAFRAN Tech)
Data SynthesisTabularBiomedical DataElectronic Health RecordsReview/Survey Paper
🎯 What it does: This paper conducts systematic experiments on survival model evaluation metrics under different censoring mechanisms and censoring rates using a semi-synthetic data generation method, comparing the differences between standard evaluation and full-information 'oracle' evaluation.
When Can You Poison Rewards? A Tight Characterization of Reward Poisoning in Linear MDPs
Jose Efraim Aguilar Escamilla, Huazheng Wang (Oregon State University)
OptimizationAdversarial AttackReinforcement LearningContrastive LearningTabularTime Series
🎯 What it does: Studied the attackability of reward poisoning attacks in linear MDPs, providing necessary and sufficient criteria for judgment;
When Data Is Scarce: Scaling Sparse Language Models with Repeated Training
Boqian Wu (University of Luxembourg), Decebal Constantin Mocanu (University of Luxembourg)
Computational EfficiencyData-Centric LearningTransformerLarge Language ModelText
🎯 What it does: This paper studies the scaling behavior of dynamic sparse training (DST) in data-limited environments and proposes a sparse-aware scalability law;
When Diffusion Language Models Hesitate: Detecting and Correcting Visual Hallucinations via Confidence Fluctuation
Wenzheng Song (Zhejiang University), Lingyun Sun (Zhejiang University)
GenerationAnomaly DetectionExplainability and InterpretabilityTransformerSupervised Fine-TuningPrompt EngineeringVision Language ModelDiffusion modelImageTextMultimodality
🎯 What it does: Designed and implemented a vision-guided correction framework called VGR, which detects visual hallucinations by leveraging confidence fluctuations during the diffusion process, extracts local visual evidence, and re-masks and re-generates affected text segments during generation.
When Distance Distracts: Representation Distance Bias in BT-Loss for Reward Models
Tong Xie (University of California, Los Angeles), Cho-Jui Hsieh (University of California, Los Angeles)
Representation LearningReinforcement Learning from Human FeedbackTransformerReinforcement LearningPrompt EngineeringContrastive LearningTextBenchmark
🎯 What it does: This paper analyzes the gradient update dynamics of the Bradley-Terry loss in reward model training, finding that the update magnitude is influenced by the representation distance, and proposes NormBT, which eliminates this bias by normalizing the gradient for each pair of samples;
When Do Diffusion Models Learn to Generate Multiple Objects?
Yujin Jeong (TU Darmstadt & hessian.AI), Anna Rohrbach (TU Darmstadt & hessian.AI)
GenerationData SynthesisConvolutional Neural NetworkTransformerSupervised Fine-TuningDiffusion modelAuto EncoderImage
🎯 What it does: Construct the MOSAIC dataset under controlled conditions to systematically evaluate the ability of diffusion models to generate multi-object images under concept generalization and compositional generalization.
When Do Graph Foundation Models Transfer? A Data-Centric Theory
Jiajun Zhu (University of Texas at Austin), Zhangyang Wang (University of Texas at Austin)
Representation LearningData-Centric LearningGraph Neural NetworkTransformerContrastive LearningGraphBenchmark
🎯 What it does: Propose a data-driven theory based on graphons to explain and decompose the output differences of graph foundation models (GFM) during cross-domain transfer.
When Do Hallucinations Arise? A Graph Perspective on the Evolution of Path Reuse and Path Compression
Xinnan Dai (Michigan State University), Jiliang Tang (Michigan State University)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningMixture of ExpertsTextGraph
🎯 What it does: This paper models the next-word prediction of LLMs as a graph search, analyzing two mechanisms that lead to hallucinations during training: path reuse and path compression.
When Does Adaptation Win? Scaling Laws for Meta-Learning in Quantum Control
Nima Leclerc (MITRE Corporation), Nicholas Peter Brawand (MITRE Corporation)
OptimizationMeta LearningReinforcement LearningTabularPhysics Related
🎯 What it does: In the simulated quantum gate calibration task, this paper derives and verifies the scaling laws of adaptive meta-learning, demonstrating that when device noise varies significantly, a small number of gradient steps can significantly improve gate fidelity, and proposes a pre-adaptive decision protocol based on a few probing steps;
When Does Predictive Inverse Dynamics Outperform Behavior Cloning?
Lukas Schäfer (Microsoft), Sergio Valcarcel Macua (Microsoft)
Robotic IntelligenceTransformerReinforcement LearningVision Language ModelAuto EncoderContrastive LearningImageVideo
🎯 What it does: This paper compares the sample efficiency and prediction error of behavioral cloning and predictive inverse dynamics models (PIDM) in offline imitation learning, combining theoretical analysis with experimental verification in 2D navigation and 3D gaming tasks.
When Does Sparsity Mitigate the Curse of Depth in LLMs
Dilxat Muhtar (Max Planck Institute for Intelligent Systems), Shiwei Liu (Max Planck Institute for Intelligent Systems)
Computational EfficiencyRepresentation LearningTransformerLarge Language ModelMixture of ExpertsText
🎯 What it does: Investigate how sparsity (implicit and explicit) mitigates the depth curse in large language models, and verify through theory and experiments the inhibitory effect of sparsity on variance propagation.
When Drafts Evolve: Speculative Decoding Meets Online Learning
Yu-Yang Qian (Nanjing University), Peng Zhao (Nanjing University)
GenerationComputational EfficiencyAI Code AssistantTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposes the OnlineSPEC framework, which combines Speculative Decoding with online learning, enabling the draft model to adaptively evolve during inference, thereby improving inference speed.
When Embedding-Based Defenses Fail: Rethinking Safety in LLM-Based Multi-Agent Systems
Lingxi Zhang (Rice University), Hanjie Chen (Rice University)
Anomaly DetectionSafty and PrivacyReinforcement Learning from Human FeedbackGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextGraphRetrieval-Augmented Generation
🎯 What it does: This paper investigates the failure mechanisms of embedding-based defenses in large language model (LLM)-driven multi-agent systems (MAS) when facing approximate benign attacks, and proposes a new defense strategy based on token-level confidence.
When Generalized Zero-Shot Learning Meets PU Learning: A Plug-and-Play Framework for Seen-Class Bias Mitigation
Long Tang (Nanjing University of Information Science & Technology), Yingjie Tian (Chinese Academy of Sciences)
ClassificationDomain AdaptationRepresentation LearningConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningImageTextBenchmark
🎯 What it does: Proposed the PUFE framework, treating GZSL inference as positive-unlabeled learning, jointly estimating label propensity and class posterior in the semantic space through a dual-headed network, and alleviating seen-class bias via adaptive prototype calibration.
When Is Rank-1 Enough? Geometry-Guided Initialization for Parameter-Efficient Fine-Tuning
Haoran Zhao (University of Melbourne), Eduard Hovy (University of Melbourne)
OptimizationComputational EfficiencyRepresentation LearningTransformerSupervised Fine-TuningVision Language ModelContrastive LearningImageTextMultimodality
🎯 What it does: Propose a geometry-guided initialization method (Gap-Init), which achieves parameter-efficient fine-tuning of multi-modal models by aligning the update direction of Rank-1 LoRA with the modality difference vector between visual-text features.
When Labelers Stay Silent: The Power of Ties in Cost-Effective Preference Learning
Jiaqi Lv (Southeast University), Xin Geng (Southeast University)
Data-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningContrastive LearningText
🎯 What it does: Proposed the silent-aware framework, allowing annotators to remain silent on answer pairs that cannot be distinguished, and incorporating these tie information into the alignment optimization of the dialogue model.
When LLMs Develop Languages: Symbolic Communication for Efficient Multi-Agent Reasoning
Zhengqi Pei (Chinese Academy Of Sciences), Shuhui Wang (Chinese Academy Of Sciences)
Autonomous DrivingOptimizationComputational EfficiencyAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkChain-of-Thought
🎯 What it does: Propose the CLSR framework, enabling LLM multi-agent self-evolution and sharing of symbolic language (LSF), and dynamically combining them based on queries through a router without a latent network
When LLMs Encounter Open-world Graph Learning: A Fresh View on Unlabeled Data Uncertainty
Yanzhe Wen (Beijing Institute of Technology), Guoren Wang (Beijing Institute of Technology)
ClassificationFederated LearningExplainability and InterpretabilityComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextGraphRetrieval-Augmented Generation
🎯 What it does: Proposed the Open-world Graph Assistant (OGA) framework, which includes Adaptive Label Traceability (ALT) and Graph Label Annotator (GLA), to address the uncertainty of unlabelled nodes in Text-Attributed Graphs (TAG).
When Model Merging Breaks Routing: Training-Free Calibration for MoE
Canbin Huang (Sun Yat-sen University), Qifan Wang (Meta AI)
OptimizationComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningMixture of ExpertsText
🎯 What it does: Proposes an untrained Hessian-Aware Router Calibration (HARC) method to address the routing breakdown issue that occurs after merging Mixture-of-Experts (MoE) models.
When More Data Doesn't Help: Limits of Adaptation in Multitask Learning
Steve Hanneke (Purdue University), Mingyue Xu (Purdue University)
Information TheoryClassificationDomain AdaptationOptimizationFederated LearningMeta LearningContrastive LearningTextTabularReview/Survey Paper
🎯 What it does: This paper theoretically investigates the limits of adaptability in multi-task learning, proving that even with infinite sample sizes for each task, optimal convergence rates cannot be achieved solely through multi-source data; meanwhile, it shows that in certain settings, simple pooling (global ERM) can achieve near-optimal adaptability rates.
When More Experts Hurt: Underfitting in Multi-Expert Learning to Defer
Shuqi Liu (Nanyang Technological University), Luke Ong (Nanyang Technological University)
ClassificationFederated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationSupervised Fine-TuningMixture of ExpertsContrastive LearningImageTabularBiomedical DataBenchmark
🎯 What it does: Proposes a new Learning to Defer (L2D) method called PiCCE, addressing the underfitting problem of classifiers in multi-expert scenarios.
When Planning Fails Despite Correct Execution: On Epistemic Calibration for LLM-Based Multi-Agent Systems
Zehao Wang (Tianjin University), Lanjun Wang (Tianjin University)
Autonomous DrivingOptimizationExplainability and InterpretabilityRobotic IntelligenceReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AITextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose an Epistemic Planning Calibration Agentic Workflow (EPC-AW) based on information consistency and memory-driven mechanisms, which reduces cognitive errors during the planning phase in LLM-based multi-agent systems through cross-information condition plan consistency evaluation and cross-round calibration constraints, thereby improving overall task success rates.
When Preference Labels Fall Short: Aligning Diffusion Models from Real Data
Weiyan Chen (Sun Yatsen University), Pengxu Wei (Sun Yatsen University)
GenerationData SynthesisTransformerSupervised Fine-TuningPrompt EngineeringDiffusion modelContrastive LearningImage
🎯 What it does: Using real images as implicit preference references, negative samples are generated by applying local distortions (such as saliency-guided inpainting) to these images, constructing unlabeled preference alignment signals, which guide diffusion models to align human aesthetics and semantic consistency.
When RAG Hurts: Diagnosing and Mitigating Attention Distraction in Retrieval-Augmented LVLMs
Beidi Zhao (University of British Columbia), Xiaoxiao Li (University of British Columbia)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerPrompt EngineeringVision Language ModelImageTextMultimodalityRetrieval-Augmented Generation
🎯 What it does: This study investigates the attention distraction phenomenon caused by Retrieval-Augmented Generation (RAG) in large vision-language models (LVLMs), and proposes a no-training, inference-only intervention method called MAD-RAG, which uses a dual-question + attention mixing approach to restore the synergy between visual localization and external knowledge.
When Random Saliency Looks Trained: Architectural Center Bias in CNN Interpretability
Keying Kuang (University of California, Berkeley), Elizabeth Purdom (University of California, Berkeley)
ClassificationExplainability and InterpretabilityConvolutional Neural NetworkTransformerContrastive LearningImage
🎯 What it does: Investigate and quantify the central bias generated in the initialization phase of convolutional networks due to zero-padding and receptive field expansion, and explore how training partially corrects this bias.
When Replanning Becomes the Bottleneck: Budgeted Replanning for Embodied Agents
Shuaijun Liu (Hong Kong University of Science and Technology), Ningxin Su (Hong Kong University of Science and Technology)
Computational EfficiencyRobotic IntelligenceTransformerLarge Language ModelPrompt EngineeringVision Language ModelVision-Language-Action ModelMultimodalityRetrieval-Augmented Generation
🎯 What it does: Proposed the BRACE framework, treating replanning as a budgeted system problem, and designed the E-RECAP token pruning module based on this, addressing the replanning delay bottleneck caused by context growth in embedded agents.
When RL Meets Adaptive Speculative Training: A Unified Training-Serving System
Junxiong Wang (Together AI), Xiaoxia Wu (Together AI)
OptimizationFederated LearningComputational EfficiencyAI Code AssistantTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsTextSequentialReview/Survey PaperBenchmarkFinance RelatedRetrieval-Augmented Generation
🎯 What it does: Built a unified training-service system called Aurora, which closes the loop between inference and online training, supporting real-time adaptive speculative decoding drafters (speculators) in production environments.
When Sample Selection Bias Precipitates Model Collapse
Xinbao Qiao (National University of Singapore (Chongqing) Research Institute), Yan Pang (National University of Singapore (Chongqing) Research Institute)
GenerationFederated LearningData-Centric LearningDiffusion modelImage
🎯 What it does: Studied how sample selection bias accelerates the collapse of recursive generative models in low-resource data island environments, and proposed a collaborative Wasserstein proxy reference method to alleviate this issue;
When Search Goes Wrong: Red-Teaming Web-Augmented Large Language Models
Haoran Ou (Nanyang Technological University), Kwok-Yan Lam (Nanyang Technological University)
Safty and PrivacyAdversarial AttackTransformerLarge Language ModelSupervised Fine-TuningTextRetrieval-Augmented Generation
🎯 What it does: This paper proposes CREST-Search, a red team testing framework for large language models equipped with web search capabilities, aimed at discovering citation risks.
When Shared Knowledge Hurts: Spectral Over-Accumulation in Model Merging
Yayuan Li (Nanjing University), Yinghuan Shi (Nanjing University)
ClassificationKnowledge DistillationRepresentation LearningConvolutional Neural NetworkTransformerContrastive LearningImageTextBenchmark
🎯 What it does: Propose a training-agnostic and data-agnostic post-processing method called Singular Value Calibration (SVC), which identifies and corrects singular value inflation caused by spectral over-accumulation by measuring subspace overlap on the output space basis of the merged model, thereby improving the effectiveness of model merging.
When Simple Problems Wear Complex Costumes: Improving Efficiency in LRM's Adaptive Reasoning
Junnan Ren (Xiamen University), Rongrong Ji (Xiamen University)
Computational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringGenerative Adversarial NetworkTextChain-of-Thought
🎯 What it does: Constructed AdaReasoningSwitch (AdaRS), an LRM that can automatically switch between thinking (Think) and direct answering (Nothink) modes when needed, solving the inefficiency caused by overthinking.
When Single Answer Is Not Enough: Rethinking Single-Step Retrosynthesis Benchmarks for LLMs
Bogdan Zagribelnyy (Insilico Medicine AI Limited), Alex Zhavoronkov (Insilico Medicine AI Limited)
Drug DiscoveryTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningTextGraphTabularBenchmarkRetrieval-Augmented Generation
🎯 What it does: This paper proposes a feasibility evaluation metric called ChemCensor based on chemical precedents, and constructs a large-scale feasible chemical reaction dataset named CREED and an OOD retrosynthesis benchmark named URSA-expert-2026. Large language models (LLMs) are used for single-step retrosynthesis experiments, and the C3LM model is fine-tuned using self-supervised learning and reinforcement learning.
When Softmax Fails at the Top: Extreme‑Value Corrections for InfoNCE
Hasan Sabri Melihcan Erol (Massachusetts Institute of Technology), Lizhong Zheng (Massachusetts Institute of Technology)
Representation LearningContrastive LearningImageText
🎯 What it does: Propose WEINCE, improving the softmax assumption of InfoNCE, utilizing extreme value theory to correct the shortage in the tail of hard negative samples, while keeping no additional parameters;
When Tabular Foundation Models Meet Strategic Tabular Data: A Prior Alignment Approach
Xinpeng Lv (National University of Defense Technology), Haotian Wang (National University of Defense Technology)
ClassificationFederated LearningExplainability and InterpretabilityData-Centric LearningTransformerReinforcement LearningPrompt EngineeringContrastive LearningTabularFinance Related
🎯 What it does: This paper proposes Strategic Prior-Data Fitted Network (SPN), which utilizes in-context learning during inference to counteract users' strategic feature manipulation against existing tabular foundation models (PFN), avoiding retraining;
When the Prompt Becomes Visual: Vision-Centric Jailbreak Attacks for Large Image Editing Models
Jiacheng Hou (Tsinghua University), Alex Jinpeng Wang (Central South University)
Safty and PrivacyAdversarial AttackTransformerPrompt EngineeringVision Language ModelVision-Language-Action ModelDiffusion modelImageMultimodalityBenchmarkRetrieval-Augmented Generation
🎯 What it does: Studied jailbreak attacks on vision-based image editing models and proposed a training-free introspective defense approach, while constructing the IESBench benchmark for security evaluation.
When to Memorize and When to Stop: Gated Recurrent Memory for Long-Context Reasoning
Leheng Sheng (National University of Singapore), Tat-Seng Chua (National University of Singapore)
Computational EfficiencyRecurrent Neural NetworkTransformerLarge Language ModelReinforcement LearningPrompt EngineeringText
🎯 What it does: Proposed the GRU-Mem framework, introducing an update gate and an exit gate into long context reasoning to improve the original MemAgent.
When to Think, When to Speak: Learning Disclosure Policies for LLM Reasoning
Jiaqi Wei (Zhejiang University), Chenyu You (Stony Brook University)
Explainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningMixture of ExpertsTextBenchmarkChain-of-Thought
🎯 What it does: Propose the Side-by-Side (SxS) interleaved reasoning framework, which enables the interleaving of reasoning steps in a single-stream autoregressive model through controllable 'think/speak' decisions, reducing the latency tax and improving user visibility.
When to Trust the Cheap Check: Weak and Strong Verification for Reasoning
Shayan Kiyani (University of Pennsylvania), Hamed Hassani (University of Pennsylvania)
Explainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This paper studies how to intelligently switch between low-cost weak verification (such as self-consistency, agent rewards, etc.) and high-cost strong verification (such as human checks, real-world execution) during the reasoning process of large language models. It proposes a weak-strong verification strategy framework and presents a distribution-agnostic online calibration algorithm called Selective Strong Verification (SSV), which significantly reduces the number of strong verification calls while ensuring the target Type-I and Type-II error rates.
Where Concept Erasure Should Occur: Concept–Layer Alignment in Text-to-Video Diffusion Models
Yiwei Xie (Huazhong University of Science and Technology), Zheng Zhang (Huazhong University of Science and Technology)
GenerationData SynthesisExplainability and InterpretabilityTransformerPrompt EngineeringDiffusion modelAuto EncoderVideoText
🎯 What it does: Propose the CLEAR framework, which automatically locates and performs concept erasure at corresponding levels within text-to-video diffusion models;
Where Detectors Fail: Probing Generative Space for Generalizable AI-Generated Image Detection
Zijie Cao (Sun Yat-sen University), Pengxu Wei (Sun Yat-sen University)
GenerationData SynthesisAnomaly DetectionConvolutional Neural NetworkSupervised Fine-TuningDiffusion modelImage
🎯 What it does: Propose the PROBE framework, which utilizes a detector as a critic to actively guide the generator to explore regions near the decision boundary, generating realistic and difficult-to-detect forged images. These samples are then used to fine-tune the detector, thereby enhancing the generalization ability across different generators.
Where Rectified Flows Leak: Characterising Membership Signals Along the Interpolation Path
Thomas Sesmat (Institut Polytechnique De Paris), Geoffroy Peeters (Institut Polytechnique De Paris)
Safty and PrivacyExplainability and InterpretabilityDiffusion modelScore-based ModelFlow-based ModelRectified FlowImageAudio
🎯 What it does: This study analyzes the memory signals in the Rectified Flows model within the training data, particularly its performance on interpolation paths, revealing that the gap between training and test data reconstruction exhibits a bell-shaped curve characteristic.
Which Algorithms Can Graph Neural Networks Learn?
Solveig Wittig (RWTH Aachen University), Christopher Morris (RWTH Aachen University)
OptimizationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningGraph Neural NetworkContrastive LearningGraph
🎯 What it does: This paper proposes a general theoretical framework that clarifies under what conditions message passing graph neural networks (MPNN) can learn from a limited number of training instances and approximate algorithms on graphs of arbitrary size;
Which Heads Matter for Reasoning? RL-Guided KV Cache Compression
Wenjie Du (Westlake University), Huan Wang (Westlake University)
CompressionComputational EfficiencyTransformerLarge Language ModelReinforcement LearningPrompt EngineeringContrastive LearningText
🎯 What it does: Use reinforcement learning probes to identify critical attention heads in inference models, and perform KV cache compression based on this information, keeping the inference performance almost unchanged.
WhisperSplat: Lossless Steganography in 3D Gaussian Splatting
Nicole Meng (Tufts University), Yingjie Lao (Tufts University)
Safty and PrivacyGaussian SplattingImagePoint Cloud
🎯 What it does: Propose WhisperSplat, a technique for achieving lossless steganography within the 3D Gaussian Splatting (3DGS) model, which can hide a full-resolution 2D image from a single view without affecting the rendering quality from other views.
Who can we trust? LLM-as-a-jury for Comparative Assessment
Mengjie Qian (University of Cambridge), Kate Knill (University of Cambridge)
Explainability and InterpretabilityRepresentation LearningData-Centric LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningText
🎯 What it does: Propose a BTσ method based on the Bradley-Terry model, which uses the judge parameters to unsupervisedly model the reliability of multiple LLM judges, thus enabling comparative evaluation of generated text without relying on human annotations.
Who Evaluates AI's Social Impacts? Mapping Coverage and Gaps in First and Third Party Evaluations
Anka Reuel (Stanford University), Irene Solaiman (Hugging Face)
Safty and PrivacyExplainability and InterpretabilityLarge Language ModelTextReview/Survey Paper
🎯 What it does: Systematically audit reports on the social impact assessment (bias, sensitive content, performance differences, environmental impact, privacy, economic cost, data/review labor force) of existing AI foundation models, compare the coverage and detail levels of first-party and third-party reports, and understand motivations and barriers through interviews.
Who Gets Credit or Blame? Attributing Accountability in Modern AI Systems
Shichang Zhang (Harvard University), Himabindu Lakkaraju (Harvard University)
OptimizationExplainability and InterpretabilityAuto EncoderGenerative Adversarial NetworkContrastive LearningImageTextBiomedical Data
🎯 What it does: Proposes a framework for tracing responsibilities at each stage of AI model development, providing a stage-level causal effect estimation method
Who Said Neural Networks Aren't Linear?
Nimrod Berman, Assaf Shocher
GenerationData SynthesisRepresentation LearningTransformerDiffusion modelScore-based ModelFlow-based ModelAuto EncoderImageVideoTabularTime Series
🎯 What it does: Propose the Linearizer framework, which enables any nonlinear network to become a linear operator in the induced vector space, thereby allowing direct application of linear algebra tools;
Who Transfers Safety? Identifying and Targeting Cross-Lingual Shared Safety Neurons
Xianhui Zhang (Nanjing University of Science and Technology), Tat-Seng Chua (National University of Singapore)
Safty and PrivacyExplainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningTextMultimodality
🎯 What it does: The paper studies cross-lingual safety alignment mechanisms, discovering and utilizing cross-lingual shared safety neurons (SS-Neurons) through fine-grained neuron-level interventions to enhance the safety of low-resource languages.
Who’s in Charge? Disempowerment Patterns in Real-World LLM Usage
Mrinank Sharma (Anthropic), David Duvenaud (University of Toronto)
Safty and PrivacyExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation
🎯 What it does: Analyzed 1.5M Claude.ai conversations, quantifying and describing patterns and potential risks in AI assistant interactions that lead to human situational disempowerment.
Whom to Query for What: Adaptive Group Elicitation via Multi-Turn LLM Interactions
Ruomeng Ding (University of North Carolina at Chapel Hill), Zhun Deng (University of North Carolina at Chapel Hill)
OptimizationFederated LearningExplainability and InterpretabilityData-Centric LearningGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextGraphTabular
🎯 What it does: Propose an adaptive group information collection framework based on multi-round LLM interaction and heterogeneous graph neural networks, jointly selecting 'who to ask' and 'what to ask' to maximize the reduction of uncertainty about group latent attributes under limited questionnaire and respondent budgets.
Why Agentic Theorem Prover Works: A Statistical Provability Theory of Mathematical Reasoning Models
Sho Sonoda (CyberAgent, Inc.), Yuya Uezato (CyberAgent, Inc.)
OptimizationExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelReinforcement LearningAgentic AITextReview/Survey PaperChain-of-Thought
🎯 What it does: This paper investigates the effectiveness of theorem provers, proposing a statistical provability theory that analyzes the probability of generating verified proofs under a limited budget.
Why Are Linear RNNs More Parallelizable?
William Merrill (Allen Institute for AI), Ashish Sabharwal (Allen Institute for AI)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningRecurrent Neural NetworkTransformerAuto EncoderContrastive LearningTabularTime SeriesSequentialBenchmark
🎯 What it does: This paper systematically compares the differences in expressive power and parallelizability between linear RNNs (LRNN) and nonlinear RNNs, and correlates different RNN classes with complexity classes (such as PNC1, NC1, L, P) through circuit complexity theory; meanwhile, it analyzes the fine-grained differences in expressive power among different parameterizations of LRNN (DPLR, PD).
Why DDIM Hallucinates More Than DDPM: A Theoretical Analysis of Reverse Dynamics
Muhammad H. Ashiq (University of Wisconsin-Madison), Grigorios Chrysos (University of Wisconsin-Madison)
GenerationDiffusion modelScore-based ModelImageStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Theoretically analyze the reverse dynamics of DDPM and DDIM diffusion samplers under Gaussian mixture targets, explaining why DDIM is more prone to generating mode-interpolation-type hallucinations, and verify this through experiments.
Why Dedicated Critics: Eliminating Target Drift in Multi-Constraint RL
Yue Yang (Monash University), Hao Wang (Monash University)
Reinforcement LearningTabularTime SeriesBenchmark
🎯 What it does: Theoretical and experimental study on the value-critic architecture in multi-constraint Lagrangian reinforcement learning, demonstrating that hybrid critics introduce bias due to target drift caused by Lagrange multiplier updates, while dedicated critics can eliminate this bias.
Why Deep Jacobian Spectra Separate: Depth-Induced Scaling and Singular-Vector Alignment
Nathanaël Haas (University of Artois), Zied Bouraoui (University of Artois)
OptimizationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningImageBenchmark
🎯 What it does: Study the singular value spectrum of the Jacobian in deep networks, proving the exponential scaling and spectral separation induced by depth, and explore their impact on the implicit bias of gradient training.
Why Do We Need Warm-up? A Theoretical Perspective
Foivos Alimisis (University of Basel), Aurelien Lucchi (University of Basel)
OptimizationRepresentation LearningConvolutional Neural NetworkTransformerContrastive LearningImageText
🎯 What it does: This paper proposes a new theoretical framework based on (H, H₀,₁)-smoothness to explain the effectiveness of learning rate warm-up in the early stages of training, and provides a corresponding adaptive step size strategy.
Why Linear Recurrent Memory Works in Partially Observable Reinforcement Learning
Yike Zhao (EPFL), Michael Muehlebach (Max Planck Institute for Intelligent Systems)
Recurrent Neural NetworkReinforcement LearningAuto EncoderSequential
🎯 What it does: Proposed and constructed a linear recursive memory that can precisely replicate Bayesian posterior logits or achieve error-free state decoding in partially observable Markov decision processes (POMDPs) with deterministic or nearly deterministic transitions.
Why ReLU? A Bit-Model Dichotomy for Deep Network Training
Ilan Doron-Arad (MIT), Elchanan Mossel (MIT)
OptimizationExplainability and InterpretabilityComputational EfficiencyReinforcement LearningContrastive Learning
🎯 What it does: Under the bit model with finite precision, the theoretical complexity of training deep networks is studied, revealing that nonlinear polynomial activation functions make both training and gradient computation #P-hard, while piecewise linear activations (such as ReLU) can be verified in NP, with forward and backward propagation completed in polynomial time.
Why Self-Distillation Helps and Hurts: Denoising vs. Signal Forgetting
Mingqi Wu (McGill University and Mila), Qiang Sun (University of Toronto and MBZUAI)
OptimizationKnowledge DistillationRepresentation LearningDiffusion modelScore-based ModelFlow-based ModelRectified FlowNeural Radiance FieldAuto EncoderContrastive LearningGaussian SplattingImageTabular
🎯 What it does: Theoretically analyze iterative self-distillation (self-training) in over-parameterized linear regression, revealing the trade-off between noise reduction and signal forgetting, and providing an exact risk recurrence formula; propose an iterative GCV method to achieve early stopping without a validation set.
Why Specialist Models Still Matter: A Heterogeneous Multi-Agent Paradigm for Medical Artificial Intelligence
Yanan Wang (Fudan University), Cuiwei Yang (Fudan University)
ClassificationAnomaly DetectionExplainability and InterpretabilityConvolutional Neural NetworkTransformerLarge Language ModelAgentic AIPrompt EngineeringMixture of ExpertsImageTextTime SeriesBiomedical DataUltrasoundElectronic Health RecordsElectrocardiogramRetrieval-Augmented GenerationAudio
🎯 What it does: Propose HetMedAgent, a multi-agent framework integrating a general large language model, domain-specific expert models, and clinicians;
Why Tree-Style Branching Matters for Thought Advantage Estimation in GRPO
Hongcheng Wang (Peking University), Hao Dong (Peking University)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringImageVideoTextMultimodalityTabularTime SeriesSequentialChain-of-Thought
🎯 What it does: Studied the impact of tree branching on the estimation of thinking advantages in GRPO and proposed the GRPO-MA multi-answer sampling scheme.
Width Independent Bounds for the Local Lipschitz Constant of Deep Neural Networks at Random Initialization and after Lazy Training
Apostolos Evangelidis (Technical University of Munich), Felix Krahmer (Technical University of Munich)
OptimizationExplainability and InterpretabilityComputational EfficiencyAuto EncoderContrastive LearningGaussian SplattingReview/Survey PaperStochastic Differential Equation
🎯 What it does: This paper studies the local Lipschitz constants of fully connected deep networks during random initialization and lazy training phases, and provides non-asymptotic upper bounds independent of width.
WildActor: Unconstrained Identity-Preserving Video Generation
Qin Guo (Hong Kong University of Science and Technology), Dan Xu (Hong Kong University of Science and Technology)
GenerationData SynthesisPose EstimationTransformerPrompt EngineeringDiffusion modelScore-based ModelRectified FlowImageVideo
🎯 What it does: Proposes the WILDACTOR framework and the Actor-18M dataset, achieving identity-preserving video generation from any perspective.
WildCat: Near-Linear Attention in Theory and Practice
Tobias Schröder (Imperial College London), Lester Mackey (Microsoft Research New England)
ClassificationGenerationCompressionComputational EfficiencyTransformerDiffusion modelAuto EncoderContrastive LearningImageTextSequential
🎯 What it does: Proposes WILDCAT, a weighted centroid-based attention approximation method that can compute Softmax attention in near-linear time.
WIND: Weather Inverse Diffusion for Zero-Shot Atmospheric Modeling
Michael Aich (Technical University of Munich), Johannes Brandstetter (Johannes Kepler University)
TransformerDiffusion modelScore-based ModelTime SeriesPhysics Related
🎯 What it does: Proposes WIND, an unsupervised pre-trained probabilistic large model that can address various meteorological and climatological problems without task-specific fine-tuning;