arXivSub Start free trial

ICML 2026 Papers — Page 59

International Conference on Machine Learning · 6554 papers

The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes

Mohammad Taufeeque (FAR.AI), Chris Cundy (FAR.AI)

Reinforcement Learning from Human FeedbackTransformerSupervised Fine-TuningReinforcement LearningPrompt EngineeringContrastive LearningText

🎯 What it does: This paper studies the use of white-box lie detectors in reward-verifiable reinforcement learning (RLVR) environments to suppress model reward hacking behavior, and explores four behaviors (honesty, obvious deception, activation layer deception, policy layer deception) that may occur during training and their mechanisms.

The Optimal Sample Complexity of Linear Contracts

Mikael Møller Høgsgaard (University of Oxford)

OptimizationFederated LearningReinforcement Learning from Human FeedbackContrastive LearningTabularFinance RelatedChain-of-Thought

🎯 What it does: Under an offline learning framework, the study explores how to learn an approximate contract equivalent to the optimal linear contract (a contract where payments are proportionally distributed) using a limited number of samples, and provides the optimal sample complexity.

The Optimal Token Baseline: Variance Reduction for Long-Horizon LLM-RL

Yingru Li (Chinese University of Hong Kong), Baoxiang Wang (Chinese University of Hong Kong)

OptimizationComputational EfficiencyTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextChain-of-Thought

🎯 What it does: Proposed the Optimal Token Baseline (OTB), which significantly reduces gradient variance by real-time weighting the advantages of each token, thereby achieving stable training of LLMs in long-term tasks.

The Oversight Game: Learning to Cooperatively Balance an AI Agent's Safety and Autonomy

William Overman (Stanford University), Mohsen Bayati (Stanford University)

Safty and PrivacyReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringText

🎯 What it does: Proposed and validated the Oversight Game, a minimal supervision layer based on a two-action interface, which allows pre-trained agents to act autonomously while maintaining human controllability.

The Pareto-optimal Trade-off between Regret and Statistical Inference in Linear Stochastic Bandits under Safety Constraints

Yuming Shao (Tsinghua University), Zhixuan Fang (Tsinghua University)

Information TheoryOptimizationSafty and PrivacyReinforcement LearningTabular

🎯 What it does: Studied the linear stochastic bandit problem under safety constraints, constructed a Pareto optimal trade-off among cumulative loss, parameter inference, and safety, and proposed a two-stage algorithm SERMiSC to achieve this trade-off.

The Perception–Physics Paradox: Probing Scientific Alignment with TC-Bench

Dingling Yao (Institute of Science and Technology Austria), Francesco Locatello (Institute of Science and Technology Austria)

Explainability and InterpretabilityTransformerVision Language ModelAuto EncoderContrastive LearningImageTabularBenchmarkPhysics Related

🎯 What it does: This paper introduces the concept of Scientific Alignment, defines Structural Isomorphism, and constructs the TC-BENCH global tropical cyclone benchmark dataset based on this, aiming to systematically explore the interpretability and alignment of visual foundation models (VFM) in satellite images with respect to physical states, revealing the contradiction between perception and physical reasoning.

The Perceptual Bandwidth Bottleneck in Vision-Language Models: Active Visual Reasoning via Sequential Experimental Design

Anjie Liu (Hong Kong University of Science and Technology), Jun Wang (Li Auto)

OptimizationExplainability and InterpretabilityComputational EfficiencyTransformerVision Language ModelImageMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This paper proposes to view high-resolution vision-language model reasoning as an active vision experiment design under the perceptual bandwidth bottleneck, designing a training-agnostic FOVEA that refines task-related evidence through adaptive cropping, thereby enhancing reasoning performance on high-resolution tasks.

The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs

Pengrui Han (California Institute of Technology), R. Michael Alvarez (California Institute of Technology)

Explainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringText

🎯 What it does: Systematically evaluate the association between self-reported personality and actual behavior in large language models using standard personality questionnaires (BFI, SRQ) and multiple psychological behavioral tasks, exploring the unity of personality expression and behavioral performance.

The Power of Power Law: Asymmetry Enables Compositional Reasoning

Zixuan Wang (Princeton University), Kaifeng Lyu (Tsinghua University)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelContrastive LearningTextChain-of-Thought

🎯 What it does: This paper studies implicit compositional reasoning tasks in natural language, compares the impact of different training data distributions (power-law vs. uniform) on model learning efficiency, and verifies the advantages of the power-law distribution through theoretical and experimental analysis.

The Quality-Utility Paradox: Why High-Reward Data Impairs Small Model Mathematical Reasoning

Haolong Qian (Tsinghua University), Chun Yuan (Tsinghua University)

Computational EfficiencyKnowledge DistillationData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Studies the 'quality-utility paradox' that occurs when using a strong teacher model to generate or revise data for small model mathematical reasoning tasks, and proposes a revision method based on style alignment to reduce adaptation costs and improve downstream performance.

The Realignment Problem: When Right becomes Wrong in LLMs

Aakash Sen Sharma (InvideoAI), Murari Mandal (Kalinga Institute of Industrial Technology)

OptimizationData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningContrastive LearningText

🎯 What it does: Proposes the TRACE framework, which achieves policy re-alignment of LLMs by utilizing existing preference data without the need for re-annotation.

The Relative Instability of Model Comparison with Cross-validation

Alexandre Bayle (Harvard University), Lester Mackey (Microsoft Research New England)

OptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyData-Centric LearningScore-based ModelAuto EncoderContrastive LearningGaussian SplattingTabularBenchmarkStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Studies the relative stability issue in model comparison using cross-validation, proving that even if individual models are stable, the comparison may still be unstable.

The Role of Target Update Frequencies in Q-Learning

Simon Weissmann (University of Mannheim), Leif Döring

Reinforcement LearningTabularTime SeriesBenchmark

🎯 What it does: This paper investigates the role of the target network update frequency (TUF) in Q-learning, providing theoretical analysis and exploring how to optimize the learning process by adjusting TUF.

The Safety-Aware Denoiser for Text Diffusion Models

Amman Yusuf (University of British Columbia), Mijung Park (University of British Columbia)

Safty and PrivacyTransformerPrompt EngineeringDiffusion modelText

🎯 What it does: This paper proposes Safety-Aware Denoiser (SAD), a framework that introduces safety guidance during the inference process of text diffusion models by modifying the denoising step, gradually guiding the generated text into a safe text space without retraining the model;

The Shadow Price of Reasoning: Economic Perspective on Optimal Budget Allocation for LLMs

Xu Wan (Zhejiang University), Mingyang Sun (Peking University)

OptimizationComputational EfficiencyAI Code AssistantTransformerLarge Language ModelTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposes CLEAR, a global shadow price balance model based on economic principles, which automatically assigns a token budget for each query during inference, supporting rational abandonment and resource reallocation.

The Shape of Addition: Geometric Structures of Arithmetic in Large Language Models

Liuyuan Wen (Nanjing University), Yang Gao (Nanjing University)

Explainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringText

🎯 What it does: Investigate the geometric structure of residual flows in large language models on multi-digit addition tasks, identifying and explaining the Iso-Raw-Sum Trajectory (IRST) and its resulting arithmetic errors

The Sign Estimator: Preference Modeling for LLM Alignment under Heterogeneity

Ali Aouad (MIT), Vivek Farias (MIT)

OptimizationExplainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningContrastive LearningText

🎯 What it does: Propose a new "Sign Estimator" to align large language models under user preference heterogeneity, correcting the bias in the reward model of traditional RLHF.

The Signal is in the Steps: Local Scoring for Reasoning Data Selection

Hoang Anh Just (Virginia Tech), Ruoxi Jia (Virginia Tech)

Explainability and InterpretabilityKnowledge DistillationData-Centric LearningTransformerLarge Language ModelPrompt EngineeringScore-based ModelTextSequentialChain-of-Thought

🎯 What it does: This paper proposes a local scoring method called LALP for selecting the most suitable reasoning trajectory for student model training from candidate answers generated by multiple teachers, specifically designed for long-form reasoning tasks.

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning

Donghang Wu (Nanyang Technological University), Yoshua Bengio (Quebec Artificial Intelligence Institute)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerReinforcement LearningPrompt EngineeringMixture of ExpertsGenerative Adversarial NetworkTextAudio

🎯 What it does: Propose FLAIR, which allows full-duplex speech dialogue models to implicitly perform internal reasoning while the user is speaking, achieving simultaneous listening and thinking.

The Stability of Singular Distribution: A Spectral Perspective on the Two-Phase Dynamics of Language Model Pre-training

Hongtao Zhang (Chinese Academy of Sciences), Xueqi Cheng (Chinese Academy of Sciences)

OptimizationExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelContrastive LearningText

🎯 What it does: Analyze the two-phase convergence of large language model pre-training from a spectral perspective, identifying and defining the 'Stable Singular Value Distribution' (SoSD) phenomenon

The Structural Origin of Attention Sink: Variance Discrepancy, Super Neurons, and Dimension Disparity

Siquan Li (Chinese University of Hong Kong), Tianyang Hu (Chinese University of Hong Kong)

Computational EfficiencyRepresentation LearningTransformerLarge Language ModelText

🎯 What it does: Studied the structural origin of the attention sink phenomenon in Transformers, and proposed a new head-level RMSNorm regularization method to suppress sinks;

The Surprising Difficulty of Search in Model-Based Reinforcement Learning

Wei-Di Chang (McGill University), Scott Fujimoto (Meta)

Reinforcement LearningTabularTime SeriesSequentialBenchmark

🎯 What it does: This paper systematically studies the challenges of planning in model-based reinforcement learning and proposes a new method called MRS.Q, which combines planning with minimization strategies for value functions, significantly improving sample efficiency and performance.

The Tell-Tale Norm: $\ell_2$ Magnitude as a Signal for Reasoning Dynamics in Large Language Models

Jinyang Zhang (Peking University), Yasha Wang (Peking University)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringAuto EncoderTextBenchmarkChain-of-Thought

🎯 What it does: This paper investigates the reasoning dynamics of LLM hidden layers through Sparse Autoencoders (SAE), discovering that the ℓ2 norm of hidden states is highly correlated with reasoning intensity, and proposes three test-time reasoning enhancement methods that require no training and no data.

The Theory and Practice of MAP Inference over Non-Convex Constraints

Leander Kurscheidt (University of Edinburgh), Antonio Vergari (University of Edinburgh)

OptimizationReinforcement LearningScore-based ModelContrastive LearningGraphTabularTime Series

🎯 What it does: This paper proposes two new MAP inference methods, achieving efficient inference for non-convex SMT(LRA) constraints and non-log-concave densities,

The Trojan Knowledge: Bypassing Commercial LLM Guardrails via Harmless Prompt Weaving and Adaptive Tree Search

Rongzhe Wei (Georgia Institute of Technology), Pan Li (Georgia Institute of Technology)

Safty and PrivacyAdversarial AttackTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose a dynamic attack framework called CKA-Agent based on tree search, which can decompose and aggregate internal related knowledge of LLM while keeping the prompt harmless at each step to bypass the security protection of commercial LLMs.

The Truth Lies Somewhere in the Middle (of the Generated Tokens)

Sophie L. Wang (Massachusetts Institute Of Technology), Brian Cheung (Massachusetts Institute Of Technology)

Explainability and InterpretabilityRepresentation LearningData-Centric LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningImageTextBiomedical Data

🎯 What it does: Investigate the hidden states during the generation process of autoregressive language models, comparing different pooling methods (last token vs mean) and their performance against prompt tokens, and evaluating their semantic quality through kernel alignment.

The Truth Stays in the Family: Enhancing Contextual Truthfulness via Inherited Heads in Model Lineages

Miso Choi (Korea University), Jungbeom Lee (Korea University)

Explainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelPrompt EngineeringScore-based ModelTextMultimodalityRetrieval-Augmented Generation

🎯 What it does: Investigated and verified that the head characteristics of large language models and their multimodal descendants regarding contextual truthfulness can be inherited within the model family, and proposed using these inherited head scores for soft gating (TruthProbe) to enhance the model's truth reasoning and the authenticity of multimodal generation.

The Two-Hump Problem: Bridging the Difficulty Gap in Mathematical Reinforcement Learning

Lucas Fagan (California Institute of Technology), Sergei Gukov (California Institute of Technology)

Data SynthesisOptimizationTransformerReinforcement LearningPrompt EngineeringAuto EncoderGenerative Adversarial NetworkGraphSequentialBenchmarkChain-of-Thought

🎯 What it does: To address the Andrews-Curtis conjecture in mathematical reasoning, it is modeled as a reinforcement learning search problem with sparse rewards, and three innovative approaches are proposed: data generation, algorithm improvement, and model architecture, successfully solving over 100 previously unsolved examples.

The Unlearnability Phenomenon in RLVR for Language Models

Yulin Chen (New York University), Chen Zhao (New York University)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackTransformerReinforcement LearningPrompt EngineeringContrastive LearningText

🎯 What it does: This paper systematically studies the phenomenon where some difficult examples remain unlearnable even after multiple rounds of reinforcement learning in RLVR training.

The Value Function Semi-Algebraic Set in Partially Observable Markov Decision Processes

Ryan A. Anderson (University of California Los Angeles), Guido Montufar (University of California Los Angeles)

OptimizationReinforcement Learning

🎯 What it does: This paper studies the geometric structure of the feasible value function set of partially observable Markov decision processes (POMDP) under infinite-horizon discounting based on memoryless random policies, and proves that this set is a semi-algebraic set defined by polynomial inequalities;

The Value of Variance: Mitigating Debate Collapse in Multi-Agent Systems via Uncertainty-Driven Policy Optimization

Luoxi Tang (Binghamton University), Zhaohan Xi (Binghamton University)

OptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringMixture of ExpertsText

🎯 What it does: A hierarchical uncertainty quantification framework is proposed to diagnose 'debate collapse' in multi-agent debates, and an uncertainty-driven policy optimization (UDPO) strategy is designed based on these metrics to enhance system robustness.

The Velocity Deficit: Initial Energy Injection for Flow Matching

Linze Li (Jiiov Technology), Jiajun Liang (Jiiov Technology)

GenerationData SynthesisTransformerDiffusion modelScore-based ModelFlow-based ModelRectified FlowImageText

🎯 What it does: Propose and address the Velocity Deficit problem that arises in Flow Matching, achieving high-quality generation through the introduction of Initial Energy Injection.

Theoretical Analysis of Sparse Optimization with Reparameterization, Weight Decay, and Adaptive Learning Rate

Huangyu Xu (Chinese Academy of Sciences), Jiaye Teng (Shanghai University of Finance and Economics)

ClassificationOptimizationConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningImageTabular

🎯 What it does: Proposed and theoretically analyzed a new sparse optimization framework called ReWA, which approximates ℓp regularization for 0 < p < 1 through reparameterization, weight decay, and adaptive learning rate, and significantly improves sparsity in experiments.

Theoretical Challenges in Learning for Branch-and-Cut

Hongyu Cheng (Johns Hopkins University), Amitabh Basu (Johns Hopkins University)

OptimizationMixture of ExpertsScore-based Model

🎯 What it does: Analyze the theoretical limitations of local scoring supervision in branch and cut plane decisions, proving that even exact matching of expert scores may lead to an exponential gap in tree scale.

Theoretical Guarantees for One-Shot Magnitude Pruning and Compute-Adaptive Early Exit

Erdem Koyuncu (University of Illinois Chicago)

Computational EfficiencyKnowledge DistillationConvolutional Neural NetworkTransformerContrastive LearningImageText

🎯 What it does: This paper studies two methods for achieving computational reduction in neural networks: static magnitude pruning and compute-adaptive early exiting, and provides theoretical guarantees for their performance on both single neurons and deep networks.

Theoretical Investigation on Inductive Bias of Isolation Forest

Qin-Cheng Zheng (Nanjing University), Zhi-Hua Zhou (Nanjing University)

Anomaly DetectionTabular

🎯 What it does: This paper models the random growth process of Isolation Forest as a random walk, deriving the closed-form expected path length and analyzing its theoretical bias.

Theoretical Perspectives on Data Quality and Synergistic Effects in Pre- and Post-Training Reasoning Models

Adel Javanmard (University of Southern California), Vahab Mirrokni (Google Research)

OptimizationExplainability and InterpretabilityData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningTextChain-of-Thought

🎯 What it does: Theoretical analysis of data quality and interaction during the pre-training and post-training stages of large language models (LLMs), validated through experiments on in-context learning tasks using linear regression.

Theory of Continual Learning Against Data Poisoning Attacks

Yiting Hu (Singapore University of Technology and Design), Lingjie Duan (Hong Kong University of Science and Technology (Guangzhou))

Federated LearningSafty and PrivacyAdversarial AttackConvolutional Neural NetworkTransformerContrastive LearningImage

🎯 What it does: Studies the theoretical security and defense strategies of continual learning (CL) when facing data poisoning attacks.

Theory of Minimal Weight Perturbations in Deep Networks and its Applications for Low-Rank Activated Backdoor Attacks

Bethan Evans (University of Oxford), Jared Tanner (University of Oxford)

Explainability and InterpretabilityComputational EfficiencyAdversarial AttackConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningImageText

🎯 What it does: This paper derives a closed-form expression for the minimal weight perturbation required to achieve a specified output change in deep networks, and applies it to backdoor attacks triggered by low-rank compression.

ThetaEvolve: Test-time Learning on Open Problems

Yiping Wang (University of Washington), yelong shen

OptimizationMeta LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextChain-of-Thought

🎯 What it does: Developed the ThetaEvolve framework, which enables a single LLM to perform reasoning and reinforcement learning on open mathematical optimization problems, continuously improving programs and surpassing existing optimal bounds during testing.

Think Deep, Not Just Long: Measuring LLM Reasoning Effort via Deep-Thinking Tokens

Wei-Lin Chen (University of Virginia), Yu Meng (University of Virginia)

Explainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelTextBenchmarkChain-of-Thought

🎯 What it does: Propose a new measure of reasoning effort—Depth of Thought Ratio (DTR), which identifies 'deep thinking' tokens by monitoring the convergence depth of internal prediction distributions at each layer during the generation process, and designs a test-time scaling strategy called Think@N based on this metric.

Think Fast and Slow: Step-Level Cognitive Depth Adaptation for LLM Agents

Ruihan Yang (Fudan University), Liefeng Bo (Tencent)

Explainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsWorld ModelTextSequentialBenchmarkChain-of-Thought

🎯 What it does: This paper proposes the COGROUTER framework, which trains LLM agents to dynamically select appropriate cognitive depth (ranging from intuitive reactions to strategic planning) at each step, thereby balancing performance and efficiency in multi-step tasks.

Think in Cloud, Look at Edges: Semantic-Driven Query Decomposition for Efficient Video Reasoning

Wenhao Zou (Shenzhen Research Institute of Big Data), Guangxu Zhu (Shenzhen Research Institute of Big Data)

Recommendation SystemFederated LearningComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringVideoTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Decompose complex queries into DAG plans using a large-scale multimodal language model in the cloud, guiding edge devices to retrieve keyframes based on structured sub-queries, thereby achieving high-precision long video reasoning under low-bandwidth conditions.

Think in Latent, Explain in Language: Self-Explainable Latent Reasoning

Dayuan Zhao (University of Illinois Urbana-Champaign), Liangyan Gui

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelTextMultimodalityChain-of-Thought

🎯 What it does: Developed Self-Explainable Latent Reasoning (SELR), enabling a single model to perform reasoning in a continuous latent space and convert latent thoughts into readable text, balancing efficiency and interpretability.

Think Less, Act Early: Reinforced Latent Reasoning with Early Exit in Vision-Language-Action Models

Dianqiao Lei (Tsinghua University), Lianlei Shan (Tsinghua University)

Computational EfficiencyRobotic IntelligenceReinforcement LearningVision Language ModelVision-Language-Action ModelMultimodalityChain-of-Thought

🎯 What it does: Propose a Vision-Language-Action framework AVA-VLA that models reasoning as the evolution of continuous latent states, and utilizes reinforcement learning for denoising and early exit to achieve efficient decision-making.

Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought Correction

Changyue Jiang (Fudan University), Min Yang (Fudan University)

Safty and PrivacyReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringContrastive LearningText

🎯 What it does: Proposes Thought-Aligner, a lightweight plugin that performs causal correction on internal reasoning processes before an LLM agent executes actions, thereby enhancing behavioral safety.

Think Twice Before You Act: Protecting LLM Agents Against Tool Description Poisoning via Isolated Planning

Shanghao Shi (Washington University in St. Louis), Ning Zhang (Washington University in St. Louis)

Safty and PrivacyAdversarial AttackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose cross-tool description poisoning attacks and design the Tool-Guard system for systematic defense.

Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models

Tianyu Fu (Tsinghua University), Yu Wang (Tsinghua University)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextBenchmarkChain-of-Thought

🎯 What it does: Proposed a cyclic Transformer model called TaH, which improves the reasoning performance of small-scale language models through selective latent iteration (performing additional iterations only on difficult-to-predict tokens).

Think-Then-Generate: Reasoning-Aware Text-to-Image Diffusion with LLM Encoders

Siqi Kou (Shanghai Jiao Tong University), Zhijie Deng (Shanghai Jiao Tong University)

Image TranslationGenerationData SynthesisTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringDiffusion modelImageTextMultimodalityChain-of-Thought

🎯 What it does: Proposed and implemented the 'Think‑Then‑Generate' (T2G) paradigm, enabling the LLM encoder to first reason and rewrite the original text prompt, and then embed the rewritten prompt into a diffusion model for image generation.

Thinking in Flow: A Dissipative Stabilization Operator for Robust Autoregressive Reasoning

Yujie Huang (Fujian University of Technology), Zhuo-Xu Cui (Chinese Academy of Sciences)

Autonomous DrivingOptimizationComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelTextChain-of-ThoughtStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Designed and implemented a continuous 'thinking state' controller (TiF) based on neural ODEs, enhancing the robustness and stability of multi-step reasoning during the Transformer autoregressive decoding process through dissipative dynamics and risk-triggered constrained interventions.

Thinking in Latent Space: Progressive Multimodal Simplification for Visual Reasoning

Yuesen Tang (Southeast University), Yu Tong (Wuhan University)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerSupervised Fine-TuningReinforcement LearningVision Language ModelDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageTextMultimodalityBenchmarkChain-of-Thought

🎯 What it does: This paper proposes the LDPVR framework, which combines explicit text states with latent space recursive simplification to form a recursive state simplification process in the form of a Markov chain within multimodal reasoning.

Thinking in Scales: Accelerating Gigapixel Pathology Image Analysis via Adaptive Continuous Reasoning

Jiusong Ge (Xi'an Jiaotong University), Zeyu Gao (University of Cambridge)

Anomaly DetectionComputational EfficiencyTransformerContrastive LearningBiomedical DataComputed TomographyChain-of-Thought

🎯 What it does: Proposes PathCTM, a panoramic pathological image analysis framework that achieves significant improvements in inference efficiency without compromising diagnostic accuracy through continuous-scale reasoning, attention-guided region pruning, and confidence-driven early stopping.

Thinking in Structures: Evaluating Spatial Intelligence in Constraint-Governed Spaces

Chen Yang (Tsinghua University), Jiansheng Fan (Tsinghua University)

TransformerPrompt EngineeringVision Language ModelImageGraphBenchmarkChain-of-Thought

🎯 What it does: Constructed and made public the SSI-Bench, a 1,000-item multiple-choice ranking VQA benchmark based on real 3D structural images, for evaluating the spatial intelligence of vision-language models in structurally constrained spaces.

Thinking with Geometry: Active Geometry Integration for Spatial Reasoning

Haoyuan Li (Sun Yet-sen University), Xiaodan Liang (Yinwang Intelligent Technology Co. Ltd)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningVideoTextMultimodalityBenchmark

🎯 What it does: Propose the GeoThinker framework, which actively injects 3D geometric information according to task requirements into a multi-modal large language model to enhance spatial reasoning capabilities.

Thinned Mean Field Langevin Dynamics

Zonghao Chen (University College London), Lester Mackey (Microsoft Research)

OptimizationComputational EfficiencyRepresentation LearningScore-based ModelContrastive LearningGaussian SplattingTabularTime SeriesSequentialBiomedical DataPhysics RelatedStochastic Differential Equation

🎯 What it does: A novel kernel thinning Mean-Field Langevin Dynamics (KT-MFLD) is studied, which reduces the original O(N²) computational complexity to O(N^{3/2}) by allowing each particle to interact with only O(√N) core particles, while maintaining convergence error comparable to standard MFLD under smoothness and RKHS conditions.

This State Looks Like That: Self-Interpretable Reinforcement Learning Agents using Prototype Soft Actor-Critic

Andrea Marzo (Université de Rennes), Roberto Capobianco (Sony AI)

Explainability and InterpretabilityReinforcement LearningTabularTime SeriesSequential

🎯 What it does: Propose ProtoSAC, which replaces the actor in Soft Actor-Critic with a prototype-based network, achieving self-explaining continuous action policies during training.

Thoughtbubbles: an Unsupervised Method for Parallel Thinking in Latent Space

Houjun Liu (Stanford University), Róbert Csordás (Stanford University)

Computational EfficiencyRepresentation LearningTransformerLarge Language ModelMixture of ExpertsText

🎯 What it does: Proposes an unsupervised Transformer variant called Thoughtbubbles, which improves reasoning efficiency without additional supervision by performing dynamic forking and removing residual flows in hidden layers, enabling parallel adaptive computation in the latent space.

ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning

Ziyan Liu (Shanghai Artificial Intelligence Laboratory), Kai Chen (Shanghai Artificial Intelligence Laboratory)

Explainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkChain-of-Thought

🎯 What it does: Propose the ThoughtFold framework, combining introspective redundancy identification with fine-grained preference learning, significantly compressing the reasoning chain length of large inference models.

ThreadWeaver: Adaptive Threading for Efficient Parallel Reasoning in Language Models

Long Lian (Meta Superintelligence Labs), Xi Victoria Lin (Meta Superintelligence Labs)

Computational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningTextBenchmarkChain-of-Thought

🎯 What it does: This study proposes the ThreadWeaver framework, which enables adaptive parallel inference for large language models to reduce inference latency while maintaining inference accuracy.

Threat2Traffic: Multi-Agent Environment Synthesis for Malware Traffic Generation from Threat Intelligence

Haoyang Chen (Chinese Academy of Sciences), Gang Xiong (Chinese Academy of Sciences)

Data SynthesisAnomaly DetectionTransformerLarge Language ModelAgentic AIPrompt EngineeringDiffusion modelTextTabularRetrieval-Augmented Generation

🎯 What it does: Propose the Threat2Traffic multi-agent framework, which automatically extracts sample-specific dependencies from threat intelligence, synthesizes IaC environments, and captures real malicious network traffic.

Three Years of r/ChatGPT: Societal Impact Evaluations from Social Media Data

Jessica Dai (University of California, Berkeley), Nika Haghtalab (University of California, Berkeley)

Explainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelAuto EncoderContrastive LearningTextTime SeriesRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Use Reddit r/ChatGPT social media data to conduct a longitudinal analysis of ChatGPT's social impact over three years, and propose the PULSE (Public and Longitudinal Signals for Evaluation) real-time monitoring framework.

Threshold-Based Exclusive Batching for LLM Inference

Weifang Zhang (Hong Kong Polytechnic University), Shining Wu (Hong Kong Polytechnic University)

OptimizationComputational EfficiencyTransformerLarge Language ModelText

🎯 What it does: Propose a threshold-based exclusive batching (EB) scheduling strategy, and design an online adaptive controller and a hybrid scheduler EB+, to optimize the throughput of LLM inference under different GPU bandwidths, model scales, and workload mixing ratios.

Threshold-Guided Optimization for Visual Generative Models

Jinbin Bai (National University of Singapore), Xiangtai Li (Peking University)

GenerationOptimizationReinforcement Learning from Human FeedbackTransformerSupervised Fine-TuningReinforcement LearningPrompt EngineeringDiffusion modelScore-based ModelImageVideoText

🎯 What it does: Propose Threshold-Guided Optimization (TGO), a method that aligns visual generation models using unpaired scalar feedback directly;

Through the Stealth Lens: Attention-Aware Defenses Against Poisoning in RAG

Sarthak Choudhary (University of Wisconsin Madison), Somesh Jha (University of Wisconsin Madison)

Explainability and InterpretabilityAdversarial AttackTransformerTextRetrieval-Augmented Generation

🎯 What it does: This paper proposes an attention-based defense method against poisoning attacks in retrieval-augmented generation (RAG) systems, which can detect and filter out malicious retrieved paragraphs that significantly affect the generated answers.

ThunderAgent: A Fast, Simple, and Program-Aware Agentic Inference System

Hao Kang (Georgia Institute of Technology), Simran Arora (Together AI)

Computational EfficiencyAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Designed and implemented THUNDERAGENT, a program-aware reasoning system for agent workflows, improving throughput and resource utilization in multi-turn reasoning and RL Rollout.

TIC-VLA: A Think-in-Control Vision-Language-Action Model for Robot Navigation in Dynamic Environments

Zhiyu Huang (University of California, Los Angeles), Jiaqi Ma (University of California, Los Angeles)

Autonomous DrivingRobotic IntelligenceTransformerSupervised Fine-TuningReinforcement LearningVision Language ModelVision-Language-Action ModelImageVideoTextMultimodalityChain-of-Thought

🎯 What it does: Propose a Think-and-Control (TIC-VLA) framework that enables robots to follow verbal instructions in dynamic environments in real time by utilizing a delayed semantic-control interface.

TideGS: Scalable Training of Over One Billion 3D Gaussian Splatting Primitives via Out-of-Core Optimization

Chonghao Zhong (Hong Kong University Of Science And Technology), Chaojian Li (Hong Kong University Of Science And Technology)

OptimizationComputational EfficiencyDiffusion modelGaussian SplattingOptical FlowImageVideoPoint Cloud

🎯 What it does: Achieved a scalable solution for training billion-level 3D Gaussian Splatting on a single 24 GB GPU, breaking through the traditional GPU memory bottleneck.

Tight Margin-Based Generalization Bounds for Voting Classifiers over Finite Hypothesis Sets

Kasper Green Larsen (Aarhus University), Natascha Schalburg (Aarhus University)

Classification

🎯 What it does: This paper proves the optimal margin-based generalization error upper bound for voting classifiers (ensemble learners) over a finite hypothesis set, for the first time providing a tight balance relationship between the number of hypotheses, margin, sample size, and failure probability;

Tight Stability Bounds for Robust Distributed Learning: Byzantine Failures Hurt Generalization More than Data Poisoning

Thomas Boudou (Inria), Aurélien Bellet (Inria)

Federated LearningSafty and PrivacyAgentic AIContrastive Learning

🎯 What it does: This paper studies the impact of Byzantine faults and data poisoning attacks on generalization error in robust distributed learning, conducting a rigorous analysis using algorithm stability theory.

Tightening the Score Matching Gap for Diffusion Models

Benjamin Dupuis (INRIA), Umut Simsekli (École Polytechnique)

GenerationData SynthesisDiffusion modelScore-based ModelContrastive LearningImageStochastic Differential Equation

🎯 What it does: Propose a tighter theoretical upper bound for the score matching gap in diffusion models, and extend this gap to KL divergence, inverse KL divergence, and Wasserstein distance.

Tighter Regret Lower Bound for Gaussian Process Bandits with Squared Exponential Kernel in Hypersphere

Shogo Iwazaki (LY Corporation)

OptimizationHyperparameter SearchReinforcement Learning from Human FeedbackMixture of ExpertsDiffusion modelScore-based ModelContrastive LearningGaussian Splatting

🎯 What it does: This paper studies the algorithm-independent worst-case lower bounds for Gaussian Process (GP) Bandits with the squared exponential kernel (SE kernel) on high-dimensional spherical input domains, and proposes new strict lower bounds, bridging the gap between dimension-related logarithmic factors and existing upper bounds.

TileQ: Efficient Low-Rank Quantization of Mixture-of-Experts with 2D Tiling

Hongyaoxing Gu (Institute of Software Chinese Academy of Sciences), Liu fangfang

Computational EfficiencyKnowledge DistillationTransformerMixture of ExpertsText

🎯 What it does: Proposes TILEQ, a post-training quantization method without fine-tuning, which utilizes 2D-tiling to share low-rank factors, mixed-precision storage, and fused inference, reducing the additional memory of Mixture-of-Experts models by 10 times and inference latency to about 5%.

TileSparse: Arithmetic-Intensity-Aware Sparse Attention for Compute-Bound LLM Decoding

Chao Wang (Chinese University of Hong Kong), Ming-Chang Yang (Chinese University of Hong Kong)

Computational EfficiencyTransformerLarge Language ModelText

🎯 What it does: Propose TileSparse, a sparse attention method targeting computational bottlenecks.

Tilt Matching for Scalable Sampling and Fine-Tuning

Peter Potaptchik (Harvard University), Michael Samuel Albergo (Harvard University)

GenerationData SynthesisOptimizationDiffusion modelScore-based ModelFlow-based ModelImagePoint CloudTabularPhysics RelatedStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Propose a Tilt Matching method based on random interpolation for sampling and fine-tuning of flow models and diffusion models without requiring gradient or trajectory backpropagation.

Time Series Forecasting Through the Lens of Dynamics

Alexis-Raja Brachet (CentraleSupélec), CELINE HUDELOT

OptimizationTransformerSupervised Fine-TuningContrastive LearningTabularTime SeriesBenchmarkPhysics Related

🎯 What it does: This paper systematically analyzes time series forecasting models from a dynamical perspective, proposing the PRO-DYN methodology, which decomposes models into three components: preprocessing (PRO), dynamics (DYN), and postprocessing (PRO). Within this framework, the paper investigates the key factors affecting model performance.

Time Series Reasoning via Process-Verifiable Thinking Data Synthesis and Scheduling for Tailored LLM Reasoning

Jiahui Zhou (Sun Yat-sen University), See-Kiong Ng (National University Of Singapore)

Data SynthesisExplainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextTime SeriesBenchmarkChain-of-Thought

🎯 What it does: Proposed the VeriTime framework, which enhances the performance of medium and small-scale LLMs on time series reasoning tasks through time series-specific chained reasoning data synthesis, data scheduling, and multi-objective reinforcement learning.

Time series saliency maps: Explaining models across multiple domains

Christodoulos Kechris (EPFL), David Atienza (EPFL)

Anomaly DetectionExplainability and InterpretabilityDiffusion modelScore-based ModelFlow-based ModelRectified FlowAuto EncoderContrastive LearningTime SeriesBiomedical DataElectrocardiogramBenchmarkAudio

🎯 What it does: Propose the Cross-Domain Integrated Gradients (CDIG) method, used to generate saliency maps for time series models under any invertible differentiable transformation (including complex domains).

Time Series, Vision, and Language: Exploring the Limits of Alignment in Contrastive Representation Spaces

Pratham Yashwante (University of California San Diego), Rose Yu (University of California San Diego)

RetrievalAnomaly DetectionRepresentation LearningTransformerContrastive LearningImageTextMultimodalityTime SeriesBiomedical DataBenchmark

🎯 What it does: Studied the representation space alignment of time series, visual line graphs, and natural language under a contrastive learning framework.

Time-Conditioned Foreseeing: An EHR-Specific Foundation Model for Irregular Dynamics and Calendrical Time

Bong Gyun Kang (Seoul National University), Sungroh Yoon (Seoul National University)

TransformerLarge Language ModelPrompt EngineeringMixture of ExpertsContrastive LearningTabularTime SeriesBiomedical DataElectronic Health Records

🎯 What it does: Designed and trained a foundation model for EHR, TCF PFM, which can accurately process irregular time and numerical data, supporting event generation and prediction with multiple time windows.

Time-Consistent Robust Multi-Objective Reinforcement Learning via a Bellman–Isaacs Weight-Adversary Recursion

Mingxi Hu (Fudan University), Meiling Yu (Nankai University)

Reinforcement LearningTime SeriesBenchmark

🎯 What it does: Propose a time-consistent robust multi-objective reinforcement learning framework, treating weights as step-by-step adversarial controls, and present the Bellman-Isaacs weight-adversary recursion; meanwhile, design deep learning implementations and evaluation protocols.

Time-PEFT: Temporal and Multichannel Complexity-Based Fine-Tuning for Time-Series Foundation Models

Jihye Na (Korea Advanced Institute of Science & Technology), Jae-Gil Lee (Korea Advanced Institute of Science & Technology)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerTabularTime SeriesBiomedical DataElectronic Health RecordsBenchmark

🎯 What it does: Propose Time-PEFT, a complexity-aware parameter-efficient fine-tuning framework for time series foundation models;

Time-Series Decomposition as a Standalone Task: A Mechanism-Driven Diagnostic Benchmark

Zipeng Wu (University of Birmingham), J. W. Andrews

Data SynthesisAnomaly DetectionOptimizationTabularTime SeriesBenchmarkPhysics Related

🎯 What it does: Established a unified benchmark for evaluating time series decomposition methods, providing synthetic data, a unified interface, and multi-dimensional evaluation metrics;

TIME: Tensor-Factorized Mixture-of-Experts with Intrinsic Routing for Lifelong Multimodal Knowledge Editing

Dexuan Xu (Peking University), Yu Huang (Peking University)

Computational EfficiencyRepresentation LearningTransformerPrompt EngineeringMixture of ExpertsContrastive LearningMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose the TIME framework, which utilizes low-rank experts from CP tensor decomposition and achieves lifelong multimodal knowledge editing through intrinsic routing.

TiME: Test-Time Mixture-of-Experts Routing via Asymmetric CO-Optimal Transport for Continual Test-Time Adaptation

Tianlun Liu (National University of Defense Technology), Dongsheng Li (National University of Defense Technology)

Domain AdaptationComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerMixture of ExpertsContrastive LearningTextBiomedical DataBenchmark

🎯 What it does: Propose an online routing and parameter update strategy (TiME) for Mixture-of-Experts (MoE) models in Continuous Testing Time Adaptation (CTTA) environments.

TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions

Linli Yao (Peking University), Xu Sun (Peking University)

GenerationData SynthesisRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringMixture of ExpertsVision Language ModelVideoTextMultimodalityBenchmarkRetrieval-Augmented GenerationAudio

🎯 What it does: Propose the OmniDenseCaptioning task, generating timestamped, multi-dimensional structured audio-visual captions; construct a high-quality human-annotated benchmark, OmniDCBench, and propose a unified evaluation metric, SodaM; based on Qwen2.5-Omni, the TimeChat-Captioner achieves high-quality spatiotemporal alignment captions through SFT+GRPO.

TimeGuard: Channel-wise Pool Training for Backdoor Defense in Time Series Forecasting

Quang Duc Nguyen (Nanyang Technological University), Dacheng Tao (Nanyang Technological University)

Anomaly DetectionSafty and PrivacyTransformerSupervised Fine-TuningContrastive LearningTabularTime Series

🎯 What it does: This paper proposes TIMEGUARD, a training-time defense framework against backdoor attacks in time series forecasting (TSF), which significantly improves model robustness without requiring additional clean samples.

TimeLAVA: Learning-Agnostic Valuation for Time Series Data

Wenqin Liu (University of Melbourne), Mingming Gong (University of Melbourne)

Anomaly DetectionContrastive LearningOptical FlowTabularTime SeriesBenchmarkStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: This paper proposes TIMELAVA, a learning-agnostic time series segmentation value assessment framework, used to quantify the intrinsic quality of each time segment in a time series;

TimeMRA: LLM-Empowered Time Series Forecasting via Multi-Scale Retrieval-Augmented Representations

Zongjiang Shang (Zhejiang University), Ling Chen (Zhejiang University)

Anomaly DetectionComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTime SeriesBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposed the TimeMRA framework, which leverages large language models to generate multi-scale retrieval-enhanced semantic representations to improve the performance of time series forecasting.

TimeOmni-VL: Unified Models for Time Series Understanding and Generation

Tong Guan (Griffith University), Shirui Pan (Griffith University)

GenerationData SynthesisAnomaly DetectionRepresentation LearningTransformerLarge Language ModelVision Language ModelDiffusion modelContrastive LearningImageTime SeriesChain-of-Thought

🎯 What it does: A unified time series model called TIMEOMNI-VL was constructed, which can simultaneously perform semantic understanding of time series (such as question answering, pattern analysis) and numerical generation (such as prediction, interpolation).

TimeRewarder: Learning Dense Reward from Passive Videos via Frame-wise Temporal Distance

Yuyang Liu (Tsinghua University), Yang Gao (Tsinghua University)

Robotic IntelligenceTransformerReinforcement LearningVision Language ModelContrastive LearningVideo

🎯 What it does: Proposes TimeRewarder, a method that generates dense rewards by learning the temporal distance between video frames, enabling the training of reinforcement learning solely based on videos without action annotations.

TimeSAE: Causal Sparse Decoding for Faithful Explanations of Black-Box Time Series Models

Khalid Oublal (LTCI Télécom Paris Institut Polytechnique de Paris), Zeynep Akata (Technical University of Munich)

Explainability and InterpretabilityTransformerAuto EncoderContrastive LearningTabularTime SeriesElectrocardiogram

🎯 What it does: Proposes the TimeSAE framework, which generates interpretable and out-of-distribution (OOD) robust explanations for black-box time series models by leveraging sparse autoencoders and causal inversion techniques.

TimeSeed: Effective Time Series Forecasting with Sparse Endogenous Variables

Zhaowang Wu (Sichuan University), Hua Yan (Sichuan University)

Computational EfficiencyData-Centric LearningRecurrent Neural NetworkTransformerContrastive LearningTabularTime SeriesBiomedical DataBenchmark

🎯 What it does: Proposes the TimeSeed framework for time series prediction under conditions with only sparse endogenous variables and complete exogenous variables.

TimeSpot: Benchmarking Geo-Temporal Understanding in Vision–Language Models in Real-World Settings

Azmine Toushik Wasi (Computational Intelligence and Operations Laboratory, Bangladesh), Md Rizwan Parvez (Qatar Computing Research Institute)

RecognitionTransformerSupervised Fine-TuningVision Language ModelContrastive LearningImageTextMultimodalityBenchmarkPhysics Related

🎯 What it does: Proposed the TIMESPOT benchmark to evaluate the joint reasoning ability of vision-language models in real-world scenarios regarding geography and time.

Timestep Rescheduling in Diffusion Inversion

Shangquan Sun (Sun Yat-sen University), Xiaochun Cao (Sun Yat-sen University)

RestorationGenerationDiffusion modelImage

🎯 What it does: Propose a non-uniform time step rescheduling method based on global scaling and local dynamic programming to reduce errors in diffusion model inversion.

TiMi: Empower Time Series Transformers with Multimodal Mixture of Experts

Jiafeng Lin (Tsinghua University), Zhongyi Pei (Tsinghua University)

Explainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelMixture of ExpertsTextMultimodalityTime SeriesAgriculture Related

🎯 What it does: Propose TiMi, a multimodal time series forecasting framework based on Transformer, which incorporates textual causal knowledge generated by large language models (LLMs) for future trend guidance into time series prediction;

TIMI: Training-Free Image-to-3D Multi-Instance Generation with Spatial Fidelity

Xiao Cai (University of Electronic Science and Technology of China), Lianli Gao (University of Electronic Science and Technology of China)

GenerationData SynthesisTransformerDiffusion modelGaussian SplattingImagePoint CloudMesh

🎯 What it does: Propose TIMI, an untrained Image-to-3D multi-instance generation framework that achieves high spatial fidelity by leveraging a pre-trained I23D generative model.

TINNs: Time-Induced Neural Networks for Solving Time-Dependent PDEs

Chen-Yang Dai (National Yang Ming Chiao Tung University), Chieh-Hsin Lai (National Yang Ming Chiao Tung University)

OptimizationTime SeriesBenchmarkPhysics Related

🎯 What it does: Propose a time-induced neural network (TINNs), which dynamically evolves spatial representations over time by parameterizing network weights as time-varying functions, addressing the time coupling problem in traditional spatiotemporal PINNs.

Tiny Brains, Giant Impact: Uncovering the Keystone Neurons of LLM with Just a Few Prompts

Xiangtian Ji (National University of Singapore), Tat-Seng Chua (National University of Singapore)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextBenchmark

🎯 What it does: Investigated and localized the extremely few 'keystone neurons' in LLMs that are highly activated across tasks, and demonstrated their critical role in the overall capabilities of the model.

TMD-Bench: A Multi-Level Evaluation Paradigm for Music–Dance Co-Generation

Xiaoda Yang (Zhejiang University), Zhou Zhao (Zhejiang University)

GenerationData SynthesisTransformerLarge Language ModelVision Language ModelDiffusion modelFlow-based ModelContrastive LearningVideoTextMultimodalityBenchmarkAudio

🎯 What it does: Propose TMD-Bench — a multi-level evaluation benchmark for text-driven music-dance co-generation, and develop a unified generation model called RhyJAM;

TMS: Trajectory-Mixed Supervision for On-Policy Self Distillation

Rana Shahroz, Tianlong Chen (University of North Carolina at Chapel Hill)

Knowledge DistillationReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningContrastive LearningText

🎯 What it does: Propose a reward-free post-training method called Trajectory-Mixed Supervision (TMS), which reduces the deviation between supervision and model distributions by mixing trajectory targets generated from historical model checkpoints, thereby alleviating the forgetting and mode collapse problems that occur under single-reference supervision.