ICML 2026 Papers — Page 61
International Conference on Machine Learning · 6554 papers
Towards Universal Gene Regulatory Network Inference: Unlocking Generalizable Regulatory Knowledge in Single-cell Foundation Models
Jiaxin Qi (Computer Network Information Center, Chinese Academy Of Sciences), Jianqiang Huang (Computer Network Information Center, Chinese Academy Of Sciences)
Explainability and InterpretabilityRepresentation LearningDrug DiscoveryTransformerSupervised Fine-TuningContrastive LearningGraphTabularBiomedical DataBenchmark
🎯 What it does: Developed a universal gene regulatory network inference framework (UGRN), and proposed two methods: virtual value perturbation (VVP) and gradient trajectory (GDT), to extract generalizable inter-gene features from frozen single-cell base models, for gene regulatory prediction across datasets and species.
Towards Whole-corpus Reconstruction of Heterogeneous RAG Knowledge Bases
Peiru Yang (Beijing University of Posts and Telecommunications), Tao Qi (Beijing University of Posts and Telecommunications)
RetrievalRepresentation LearningAdversarial AttackTransformerLarge Language ModelContrastive LearningGaussian SplattingTextMultimodalityBenchmarkFinance RelatedRetrieval-Augmented Generation
🎯 What it does: Propose the GeoEx framework, which utilizes a proxy encoder to plan queries in the embedding space and generates natural language queries through embedding inversion, achieving a full reconstruction of multi-source RAG knowledge bases without prior knowledge.
TPGDiff : Hierarchical Triple-Prior Guided Diffusion for Image Restoration
Yanjie Tu (Northwestern Polytechnical University), Jiacong Tang (Northwestern Polytechnical University)
RestorationKnowledge DistillationConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelContrastive LearningImageStochastic Differential Equation
🎯 What it does: Proposed a hierarchical diffusion framework called TPGDiff based on triple priors (semantic, structural, degradation) for unified handling of various image degradations.
TPV: Parameter Perturbations Through the Lens of Test Prediction Variance
Devansh Arpit (Modelable AI)
OptimizationExplainability and InterpretabilityData-Centric LearningConvolutional Neural NetworkTransformerContrastive LearningImageTabular
🎯 What it does: This paper proposes and verifies a new robustness evaluation metric called Test Prediction Variance (TPV), which measures the fluctuation of model predictions on the test set under small perturbations of the training model parameters.
TQL: Scaling Q-Functions with Transformers by Preventing Attention Collapse
Perry Dong (Stanford University), Chelsea Finn
TransformerReinforcement LearningContrastive LearningTabularSequentialBenchmark
🎯 What it does: This paper studies how to extend the Transformer architecture as a Q-value function and proposes Transformer Q-Learning (TQL), which achieves stable training of large-scale value functions by controlling attention entropy to prevent attention collapse.
TRACE: Toulmin-based Reasoning Assessment through Constructive Elements for LLM CoT Evaluation
Yundong Kim (Korea Institute of Science and Technology Information (KISTI)), Heyoung Yang (Korea Institute of Science and Technology Information (KISTI))
Explainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextChain-of-Thought
🎯 What it does: The paper proposes the TRACE evaluation framework for structured, reference-free assessment of the Chain-of-Thought reasoning process in LLMs.
TRACE: Trajectory Recovery for Continuous Mechanism Evolution in Causal Representation Learning
Shicheng Fan (University of Illinois at Chicago), Lu Cheng (University of Illinois at Chicago)
Representation LearningReinforcement Learning from Human FeedbackMixture of ExpertsAuto EncoderContrastive LearningTabularTime SeriesSequential
🎯 What it does: Proposed the TRACE framework for recovering the trajectory of continuous mechanism evolution in causal representation learning.
TRACER: Persistent Regularization for Robust Multimodal Finetuning
Hesam Asadollahzadeh (University of Melbourne), Sarah Monazam Erfani (University of Melbourne)
Domain AdaptationKnowledge DistillationRepresentation LearningTransformerSupervised Fine-TuningContrastive LearningImageTextMultimodality
🎯 What it does: Propose the TRACER method, which uses WMA teacher dynamic self-distillation combined with contrastive learning to achieve robust fine-tuning of multi-modal models, addressing the generalization degradation caused by the collapse of traditional EMA teachers.
TRACER: Trajectory Risk Aggregation for Critical Episodes in Agentic Reasoning
Sina Tayebati (University of Illinois Chicago), Amit Ranjan Trivedi (University of Illinois Chicago)
Anomaly DetectionExplainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextSequentialBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Design and implement a trajectory-level uncertainty assessment method called TRACER, specifically for conversational LLM agents used in multi-turn tool usage.
TraceRouter: Robust Safety for Large Foundation Models via Path-Level Intervention
Chuancheng Shi (University of Sydney), Tat-Seng Chua (National University of Singapore)
Safty and PrivacyExplainability and InterpretabilityTransformerPrompt EngineeringVision Language ModelDiffusion modelAuto EncoderImageTextMultimodality
🎯 What it does: This paper proposes TraceRouter, a security intervention framework that identifies, isolates, and cuts off harmful semantic propagation paths.
TraCeS: Learning Per-Timestep Constraint-Violation Credit from Sparse Trajectory-Level Labels
Siow Meng Low (Singapore Management University), Akshat Kumar (Singapore Management University)
OptimizationRecurrent Neural NetworkReinforcement LearningScore-based ModelContrastive LearningTime SeriesSequential
🎯 What it does: In the absence of cost function and threshold information, learn the violation credit for each time step using sparse trajectory-level accept/reject labels, thereby achieving safe reinforcement learning.
Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection
Xulin Hu (Peking University), Zhong Chen (Peking University)
Anomaly DetectionSafty and PrivacyExplainability and InterpretabilityConvolutional Neural NetworkTransformerLarge Language ModelPrompt EngineeringAuto EncoderText
🎯 What it does: Identify and quantify the dynamic trajectories of refusal within large language models using causal tracing methods, and then design a white-box detector called SALO based on sparse activation localization, which can real-time detect jailbreaks across multiple models and attack scenarios.
Tracing the Persona Circuit: How Large Language Models Encode and Express Character Traits
Guanzheng Qin (University of Science and Technology of China), Xinmei Tian (University of Science and Technology of China)
Explainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBenchmark
🎯 What it does: Investigates how large language models internally encode and express personality traits, and achieves causal tracing of the personality generation process through a differentiable 'latent personality vector'.
Tracking Drift: Variation-Aware Entropy Scheduling for Non-Stationary Reinforcement Learning
Tongxi Wang (School of Future Technology, Southeast University), Shan Liu (School of Automation, Southeast University)
Reinforcement LearningTabularTime Series
🎯 What it does: This paper proposes AES (Variation-Aware Entropy Scheduling) — a entropy weight scheduling method based on environment drift awareness, used to adaptively control exploration intensity in non-stationary reinforcement learning scenarios.
Tractable Expected Information Gains for Exponential Family Posteriors
Rik Knowles (University of Oxford), Tom Rainforth (University of Oxford)
Information TheoryOptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyTabularTime SeriesBenchmark
🎯 What it does: Studied how, in Bayesian experimental design, when the posterior belongs to the exponential family and depends on the experimental design only through the natural parameters, the expected information gain (EIG) can be reduced from doubly intractable to singly intractable, and provided the corresponding unbiased estimator and its gradient.
Trading Complexity for Expressivity Through Structured Generalized Linear Token Mixing
Erwan Fagnou (Universite Paris Dauphine PSL), Alexandre Allauzen (Universite Paris Dauphine PSL)
Computational EfficiencyRepresentation LearningTransformerLarge Language ModelMixture of ExpertsText
🎯 What it does: This paper explores balancing expressive power and complexity through structured generalized linear token mixing, proposing a unified framework that separates the direct impact of inputs on outputs from the recursive propagation of information.
Train for Truth, Keep the Skills: Binary Retrieval-Augmented Reward Mitigates Hallucinations
Tong Chen (University of Washington), Faeze Brahman (Allen Institute for AI)
Federated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningContrastive LearningTextRetrieval-Augmented Generation
🎯 What it does: Proposed and implemented a binary reward (Binary RAR) based on retrieval augmentation, embedding it into online reinforcement learning to reduce hallucination generation in large language models.
Train Once, Reuse Everywhere: Generalizable Implicit ICL by Routing Attention
Jiaqian Li (Brown University), Wenya Wang (Nanyang Technological University)
ClassificationDomain AdaptationRecommendation SystemMeta LearningTransformerPrompt EngineeringMixture of ExpertsAuto EncoderContrastive LearningText
🎯 What it does: A new implicit context learning method called In-Context Routing (ICR) is proposed without increasing model parameters, which simulates the effect of multi-task examples through routing attention, enabling direct inference across multiple domains after a single training.
Trainable Nonexpansive Denoisers for Contractive Image Reconstruction
Arghya Sinha (Indian Institute of Science), Kunal N. Chaudhury (Indian Institute of Science)
RestorationSuper ResolutionConvolutional Neural NetworkAuto EncoderImage
🎯 What it does: A weighted aggregation structure based on permutation sets is proposed, and a trainable denoiser with a global Lipschitz constant no greater than 1 is designed and embedded into the PnP-HQS framework, theoretically ensuring convergence;
Training AI Co-Scientists Using Rubric Rewards
Shashwat Goel (ELLIS Institute Tübingen), Chenxi Whitehouse (Meta Superintelligence Labs)
Meta LearningDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackNeural Architecture SearchTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposed an scalable training process that utilizes automatically extracted research objectives and corresponding scoring rubrics, combined with reinforcement learning to enable language models to generate superior research plans.
Training Data Efficiency in Multimodal Process Reward Models
Jinyuan Li (Washington University in St. Louis), Jiaxin Huang (Washington University in St. Louis)
Computational EfficiencyData-Centric LearningReinforcement Learning from Human FeedbackTransformerReinforcement LearningVision Language ModelContrastive LearningMultimodalityBenchmark
🎯 What it does: This paper investigates the data efficiency of the Multimodal Process Reward Model (MPRM), proposing to select the most informative rollouts from the annotated Monte Carlo (MC) samples using the Balanced-Information Score (BIS), in order to reduce the amount of training data and computational costs.
Training Deep Spiking Neural Networks without Normalization
Xinyu Shi (Peking University), Zhaofei Yu (Peking University)
ClassificationRecognitionOptimizationComputational EfficiencyConvolutional Neural NetworkSpiking Neural NetworkTransformerImageText
🎯 What it does: This paper proposes a non-normalization deep spiking neural network training initialization framework called SpikeInit, addressing the dependencies and limitations of traditional BN in SNN training.
Training Diffusion Language Models for Black-Box Optimization
Zipeng Sun (McGill University), Xue Liu (McGill University)
OptimizationTransformerSupervised Fine-TuningReinforcement LearningPrompt EngineeringDiffusion modelTextTabular
🎯 What it does: Adapt diffusion language models (diffusion LLM) to offline black-box optimization (BBO), achieving high-quality sampling of offline design data through unified prompt-response corpora, delimiter tokens, domain adaptation, supervised fine-tuning, and reinforcement learning.
Training Language Model Agents to Find Vulnerabilities with CTF-Dojo
Terry Yue Zhuo (Monash University), Zijian Wang (Meta Superintelligence Labs)
AI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Built CTF-DOJO — a large-scale executable Dockerized CTF environment, and automatically generated 658 challenges via CTF-FORGE; collected LLM interaction trajectories and trained using writing prompts and runtime augmentation.
Training LLM Agents to Empower Humans
Evan Ellis (University of California, Berkeley), Benjamin Eysenbach (Princeton University)
AI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AIPrompt EngineeringWorld ModelTextRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose a self-supervised method called Empower, which trains LLM assistants by maximizing human empowerment, improving code generation and interaction without requiring human feedback.
Training Prompt Matters: State-Adaptive Optimization for Robust Fine-Tuning
Wenhang Shi (Renmin University of China), Xiaoyong Du (Renmin University of China)
OptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningAdversarial AttackHyperparameter SearchData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringText
🎯 What it does: In the process of fine-tuning large language models, the impact of prompt selection during the training phase on model performance was systematically studied, and a dynamic prompt optimization method based on pre-updated loss (SAPO) was proposed. By generating multiple semantically equivalent prompts and selecting the one with the lowest pre-updated loss as the training prompt, this method significantly reduces forgetting and enhances generalization.
Training with Honeypots: Reshaping How LLMs Fail Under Adversarial Attacks
Samuel Simko (ETH Zürich), Bernhard Schölkopf (MPI for Intelligent Systems)
Safty and PrivacyAdversarial AttackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringContrastive LearningText
🎯 What it does: By introducing honeypot (bait) responses during training, reshape the failure patterns of large models when subjected to adversarial attacks, making the model tend to generate low-risk responses that are not actionable even if the defense is bypassed.
Training-Free Adaptation of Diffusion Models via Doob's $h$-Transform
Qijie Zhu (Northwestern University), Minshuo Chen (Northwestern University)
Domain AdaptationReinforcement Learning from Human FeedbackReinforcement LearningDiffusion modelScore-based ModelImageTextStochastic Differential Equation
🎯 What it does: Propose the DOIT method based on Doob's h-Transform to achieve reward-oriented adaptation of pre-trained diffusion models without requiring additional training.
Training-Free Adversarial Robustness in Computational MRI
Mahdi Saberi (University of Minnesota), Mehmet Akcakaya (University of Minnesota)
RestorationAdversarial AttackConvolutional Neural NetworkRecurrent Neural NetworkDiffusion modelContrastive LearningBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Proposes a training-free adversarial robustness method that does not require retraining, reducing the sensitivity of deep learning MRI reconstruction models to small adversarial perturbations through cyclic measurement consistency.
Training-Free Bayesian Filtering with Generative Emulators
Thomas Savary (University of Liège), Gilles Louppe (University of Liège)
OptimizationFederated LearningComputational EfficiencyRepresentation LearningGraph Neural NetworkSpiking Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningImageVideoTabularTime SeriesSequentialPhysics RelatedStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Proposes a particle filtering method that requires no additional training, utilizing a diffusion model generator to directly achieve the optimal proposal distribution, addressing the integration of Kalman filtering and particle filtering in high-dimensional systems.
Training-Free Coverless Multi-Image Steganography with Access Control
Minyeol Bae (Korea Advanced Institute of Science and Technology), Si-Hyeon Lee (Korea Advanced Institute of Science and Technology)
Safty and PrivacyTransformerDiffusion modelAuto EncoderImage
🎯 What it does: Propose MIDAS, a training-agnostic image-coverless multi-image steganography framework that supports access control based on private keys;
Training-Free Guided Diffusion for Planning: A Unified Framework via Doob’s h-Transform with Safety Guarantees
Kenta Hoshino (Institute of Science Tokyo), Yorie Nakahira (Carnegie Mellon University)
Autonomous DrivingOptimizationSafty and PrivacyRobotic IntelligenceDiffusion modelScore-based ModelImageTextTabularStochastic Differential Equation
🎯 What it does: Study the guiding mechanism of continuous-time fractional basis diffusion models, propose a theoretical framework based on Doob's h-transform, provide error upper bounds and safety guarantees, and achieve training-agnostic guidance through stochastic optimal control.
Training-Free Hashing-Based Attention via Binary Principal Components
Daohai Yu (Xiamen University), Rongrong Ji (Xiamen University)
Computational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringAuto EncoderTextBenchmark
🎯 What it does: BinaryPC proposes a training-agnostic, data-aware hashing sparse attention mechanism, achieving efficient retrieval and decoding by compressing the KV cache using binary principal component analysis;
Training-Free Multimodal Large Language Model Orchestration
Tianyu Xie (Xiamen University), Rongrong Ji (Xiamen University)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsImageVideoTextMultimodalityBenchmarkRetrieval-Augmented GenerationAudio
🎯 What it does: Designed and implemented a training-agnostic multimodal LLM scheduling framework, utilizing an LLM controller to issue closed token routing to experts and combining cross-modal memory to realize real-time full-duplex multimodal assistant.
Training-Free Rate-Distortion-Perception Traversal With Diffusion
Yuhan Wang (Chinese University of Hong Kong), Ying-Jun Angela Zhang (Chinese University of Hong Kong)
CompressionTransformerDiffusion modelScore-based ModelFlow-based ModelRectified FlowImageStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Propose a training-free framework that utilizes pre-trained diffusion models to traverse the complete rate-distortion-perception (RDP) surface through inverse channel coding (RCC) and score-scaled probability flow ODE.
Training-Free Vector Quantization via Gaussian VAEs
Tongda Xu (Tsinghua University), Jie Tang (Tsinghua University)
GenerationCompressionRepresentation LearningTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningGaussian SplattingImage
🎯 What it does: This paper proposes an untrained vector quantization method called Gaussian Quant (GQ), which first trains a constrained Gaussian Variational Autoencoder (Gaussian VAE), and then matches its posterior mean with a set of pre-generated Gaussian noise codebook, achieving a complete process of directly converting Gaussian VAE into VQ-VAE.
Training-Trajectory-Aware Token Selection
Zhanming Shen (Zhejiang University), Junbo Zhao (Zhejiang University)
Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningDiffusion modelText
🎯 What it does: This paper studies the 'imitation shock' phenomenon that occurs in continuous reasoning distillation, and proposes the Training Trajectory-aware Token Selection (T3S) method. It identifies and masks the easily learned 'Anchor' tokens at bottleneck points during the training process, thereby freeing up the optimization path for subsequent difficult-to-learn tokens.
Training–Inference Consistent Segmented Execution for Long-Context LLMs
Xianpeng Shang (Inner Mongolia University), Xiangdong Su (Inner Mongolia University)
Computational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation
🎯 What it does: Proposes a training-inference consistent segmented execution framework, allowing training and inference to use the same segmented forward execution and state passing.
Trajectory Consistency for One-Step Generation on Euler Mean Flows
Zhiqi Li (Georgia Institute of Technology), Bo Zhu (Georgia Institute of Technology)
GenerationData SynthesisComputational EfficiencyDiffusion modelScore-based ModelFlow-based ModelImagePoint CloudMeshTabular
🎯 What it does: Proposes the Euler Mean Flows (EMF) framework, which uses local linear approximation to convert trajectory consistency constraints into directly data-supervised objectives, achieving one-shot/ few-step generation while enabling training without JVP (Jacobian-Vector Product);
Trajectory Seriation via Spectral Tangent Alignment and Global Embedding
Zhixin Zhou (Alpha Benito Research), Arash A. Amini (University of California)
Computational EfficiencyRepresentation LearningAuto EncoderContrastive LearningOptical FlowPoint CloudTime SeriesBiomedical Data
🎯 What it does: Developed the STAGE algorithm to recover the one-dimensional curve order from noisy point clouds.
Trajectory Stitching for Solving Inverse Problems with Flow-Based Models
Alexander Denker (Helmholtz Imaging, Deutsches Elektronen-Synchrotron Desy), Moshe Eliasof (University Of Cambridge)
Image TranslationRestorationSuper ResolutionOptimizationComputational EfficiencyConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelFlow-based ModelImageComputed TomographyStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Proposes the Multi-Shot Flow (MS-Flow) framework, which solves inverse problems by inserting intermediate latent states along the flow model trajectory, avoiding backpropagation over the full ODE trajectory.
Trajectory-Aware Certified Decentralized Unlearning via SGD Stability
Hengliang Wu (Shandong University), Youming Tao (TU Berlin)
Federated LearningSafty and PrivacyConvolutional Neural NetworkTransformerContrastive LearningImageText
🎯 What it does: Proposes a general trajectory-aware verified decentralized federated learning framework, TRACE-DU, which can efficiently and verifiably remove the influence of specified clients in a collaborative learning environment without a central server;
Trajectory-Aware Heuristic Learning for Combinatorial Search
Mustafa Seddiqi (Concordia University), Tiberiu Popa (Concordia University)
OptimizationComputational EfficiencyReinforcement Learning from Human FeedbackRecurrent Neural NetworkTransformerAuto EncoderContrastive LearningSequentialBenchmark
🎯 What it does: Propose a trajectory-aware probabilistic framework that treats the depth of the Rubik's Cube as a trajectory-level latent variable, using HMM combined with sparse precise anchors and forward-backward inference to achieve label refinement and training of value heuristics without search;
Trajectory-Aware Spiking DiTs Conversion via Membrane Potential Error-Feedback
Haoran Fang (Jilin University), Bin Gu (Jilin University)
GenerationSpiking Neural NetworkTransformerDiffusion modelImage
🎯 What it does: Proposes a training-agnostic framework for converting Diffusion Transformer (DiT) into Spiking Neural Networks (SNN)
Trajectory-Level Data Augmentation for Offline Reinforcement Learning
Tobias Schmähling (University of Applied Sciences Kempten), Tobias Windisch (University of Applied Sciences Kempten)
Reinforcement LearningDiffusion modelContrastive LearningImageTabularTime SeriesSequential
🎯 What it does: Propose a trajectory-based shortcut data augmentation method (LIFT), which improves data quality in offline reinforcement learning by inserting high-value actions during the logging policy collection process.
Trajectory-Level Speculative Decoding for Diffusion Language Models
Tianxiang Pan (Li Auto Inc.), Kaiwen Long (Li Auto Inc.)
GenerationComputational EfficiencyAI Code AssistantTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsDiffusion modelText
🎯 What it does: Proposes a trajectory-level speculative decoding framework that constructs draft denoised trajectories using confidence-tiered tree search, and accelerates diffusion language model inference through block-level parallel verification and cross-block speculation.
Trajectory-Stabilized Inference for Diffusion-Based Video Inpainting
Zhanhe Zhang (Xidian University), Cheng Deng (Xidian University)
RestorationGenerationConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelOptical FlowVideo
🎯 What it does: Propose a trajectory stabilization framework during inference, which monitors motion alignment drift in the video inpainting process of diffusion models and performs soft correction when risk accumulates, to enhance spatiotemporal consistency in long sequences.
Transfer Learning in High-dimensional Ising Models
Joonho Kim (KAIST), Seyoung Park (Yonsei University)
Domain AdaptationRepresentation LearningContrastive LearningGraphTabularBiomedical Data
🎯 What it does: Designed a transfer learning method called Trans-Ising for high-dimensional Ising models, combining pseudo-likelihood-based source selection with a two-step estimation process involving dual penalties;
Transfer Learning in Nonparametric Regression with Deep ReLU Networks
Junpeng Ren (University of California, Los Angeles), OSCAR HERNAN MADRID PADILLA
ImageTabularBenchmark
🎯 What it does: This paper proposes a two-stage offset learning framework for multi-group nonparametric regression, first estimating the shared mean function across all groups, then estimating the offset term for each group, ultimately obtaining group-specific conditional mean estimates, and providing a general upper bound on the error;
Transferable Reinforcement Learning via Probabilistic Latent Embeddings and Dynamic Policy Adaptation for Sim-to-Real Deployment
Gengyue Han (Purdue University), Yiheng Feng (Purdue University)
Domain AdaptationAutonomous DrivingTransformerReinforcement LearningDiffusion modelAuto EncoderContrastive LearningTabularTime SeriesSequential
🎯 What it does: Construct a transferable control framework based on probabilistic latent embeddings and distributed reinforcement learning, designed to dynamically adjust risk and ensure safety during deployment from simulation to real-world environments (Sim-to-Real).
Transform Trained Transformer for Accelerating Native 4K Video Generation
Jiangning Zhang (APRIL Lab, Zhejiang University), Yong Liu (APRIL Lab, Zhejiang University)
GenerationComputational EfficiencyTransformerSupervised Fine-TuningDiffusion modelAuto EncoderVideo
🎯 What it does: This paper achieves a reconstruction of the full attention mechanism of a pre-trained Transformer through a Transformer retrofit method (T3-Video), significantly reducing the computational cost of 4K video generation without altering the core architecture of the model, thus enabling the efficient generation of 81-frame 4K videos.
Transformed Latent Variable Multi-Output Gaussian Processes
Xiaoyu Jiang (University of Manchester), Mauricio A Álvarez
Computational EfficiencyRepresentation LearningData-Centric LearningTransformerDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningGaussian SplattingTabularTime SeriesBiomedical DataBenchmarkStochastic Differential Equation
🎯 What it does: Proposes the Transformed Latent Variable Multi-Output Gaussian Processes (T-LVMOGP) framework, aimed at maintaining scalability and capturing rich output correlations in high-dimensional output spaces.
Transformer Circuits Can Realize Clustering Algorithms
Kenneth L. Clarkson (IBM Research), Parikshit Ram (IBM Research)
OptimizationExplainability and InterpretabilityComputational EfficiencyTransformerPrompt EngineeringAuto EncoderContrastive LearningTabular
🎯 What it does: Proposed a Transformer architecture that precisely implements the Lloyd algorithm for k-means clustering, and further improves clustering quality through training.
Transformers Can Learn Posterior Predictive Distributions In-Context
Gyeonghun Kang (Duke University), Xiang Cheng (Duke University)
OptimizationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerTabularTime Series
🎯 What it does: This paper constructs a proof demonstrating that Transformer-based Prior-Data-Fitted Networks (PFNs) can approximate the posterior predictive distribution (PPD) in Gaussian Process (GP) regression problems, and provides an error upper bound;
Transformers Efficiently Perform In-Context Logistic Regression via Normalized Gradient Descent
Chenyang Zhang (University of Hong Kong), Yuan Cao (University of Hong Kong)
OptimizationComputational EfficiencyMeta LearningTransformerTabular
🎯 What it does: Investigate the implementation of normalized gradient descent (NGD) in logistic regression for in-context learning tasks using soft-max attention Transformers, and obtain an efficient weight prediction model by retraining a single-layer Transformer.
Transformers learn factored representations
Adam Shai (Simplex, Astera Institute), Paul M. Riechers (Beyond Institute for Theoretical Science)
Representation LearningTransformerAuto EncoderGenerative Adversarial NetworkContrastive LearningSequential
🎯 What it does: Investigate whether Transformer models decompose the world into independent factors and represent these factors as orthogonal subspaces in the activation space during next-token pre-training.
Transformers Learn the Optimal DDPM Denoiser for Multi-Token GMMs
Hongkang Li (University of Pennsylvania), Rene Vidal
GenerationData SynthesisOptimizationRepresentation LearningTransformerDiffusion modelScore-based ModelImage
🎯 What it does: The paper provides a theoretical analysis of diffusion models (DDPM) with the Transformer architecture, proving that under a single-head Transformer with one layer and multi-token Gaussian mixture data (MTGM), gradient descent can globally converge to the Bayes optimal denoiser, and further achieve approximate score matching.
Transformers Provably Learn Algorithmic Solutions for Graph Connectivity, But Only with the Right Data
Qilin Ye (Duke University), Vatsal Sharan (University of Southern California)
OptimizationExplainability and InterpretabilityRepresentation LearningTransformerSupervised Fine-TuningContrastive LearningGraph
🎯 What it does: This paper studies how Transformer learns algorithmic solutions rather than fragile heuristics in graph connectivity tasks, and provides theoretical and experimental validations;
Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently
Bochen Lyu (University of Southampton), Zhanxing Zhu (University of Southampton)
Computational EfficiencyRepresentation LearningTransformerSupervised Fine-TuningReinforcement LearningTabularChain-of-Thought
🎯 What it does: Studied how to achieve chain-of-thought (CoT) on Transformers through reinforcement learning or supervised fine-tuning, and proved that they can learn k-sparse Boolean functions, but the learning behaviors of the two methods are different.
Transforming Weather Data from Pixel to Latent Space
Sijie Zhao (Nanjing University), LEI BAI
CompressionRepresentation LearningTransformerAuto EncoderGenerative Adversarial NetworkContrastive LearningImageTime Series
🎯 What it does: Propose Weather Latent Autoencoder (WLA), which maps weather data from the pixel space to a low-dimensional latent space, achieving compression, denoising, and disentanglement, and supporting unified processing of multi-pressure-variable subsets;
Transitive Representation Learning Enhances Histopathology Annotation
Moritz Schaefer (Stanford University), Zinaida Good (Stanford University)
RetrievalRepresentation LearningDrug DiscoveryTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityTabularBiomedical DataBenchmark
🎯 What it does: By constructing a tri-modal (image-gene expression-text) contrastive learning framework named SPATIALWHISPERER, gene expression bridges images and text, thereby achieving zero-shot cell type annotation of H&E images.
Transitivity Meets Cyclicity: Explicit Preference Decomposition for Dynamic Large Language Model Alignment
Yucong Huang (Harbin Institute of Technology), Jing Li (Harbin Institute of Technology)
OptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringScore-based ModelTextBenchmarkChain-of-Thought
🎯 What it does: Proposed the Hybrid Reward-Cyclic (HRC) model, which explicitly decomposes human preferences into a transitive scalar reward and a cyclic vector component, and combines Dynamic Self-Play Preference Optimization (DSPPO) for time-varying game optimization.
Translation Heads: Disentangling meaning from language in LLM-based machine translation
Théo Lasnier (Inria), Benoît Sagot (Inria)
Explainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextMultimodality
🎯 What it does: This paper analyzes the internal workings of LLMs in sentence-level machine translation from the perspective of mechanism interpretability, using activation patching to perform fine-grained localization of attention heads, and decomposing the translation task into two subtasks: target language identification and sentence equivalence.
TransLight: Image-Guided Customized Lighting Control with Generative Decoupling
Zongming Li (Huazhong University of Science and Technology), Xinggang Wang (Huazhong University of Science and Technology)
Image TranslationRestorationGenerationTransformerSupervised Fine-TuningDiffusion modelContrastive LearningImage
🎯 What it does: This paper proposes the TransLight method, which accurately transfers geometric light effects (such as Tyndall beams, glows, and lens flares) from reference images to target portrait images;
TransNormal: Dense Visual Semantics for Diffusion-based Transparent Object Normal Estimation
Mingwei Li (Zhejiang University), Yi Yang
Pose EstimationDepth EstimationConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningImage
🎯 What it does: Propose the TransNormal framework, which utilizes dense visual semantics from Stable Diffusion and DINOv3 for single-step transparent object normal estimation.
Transolver-3: Scaling Up Transformer Solvers to Industrial-Scale Geometries
Hang Zhou (Tsinghua University), Mingsheng Long (Tsinghua University)
OptimizationComputational EfficiencyGraph Neural NetworkTransformerPrompt EngineeringMixture of ExpertsDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningMeshGraphPhysics Related
🎯 What it does: Propose Transolver-3, a Transformer PDE solver designed for industrial-scale geometry (10⁸+ elements), capable of being trained on a single GPU and achieving high-fidelity field prediction during inference.
Transport and Merge: Cross-Architecture Merging for Large Language Models
Chenhang Cui (National University of Singapore), Tat-Seng Chua (University of Science and Technology of China)
Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningTextBenchmarkFinance Related
🎯 What it does: This paper proposes a cross-architecture model merging method based on optimal transport (OT), which can directly transfer knowledge from large high-resource language models to small low-resource target models, achieving parameter fusion without gradient updates.
Transport Clustering: Solving Low-Rank Optimal Transport via Clustering
Henri Schmidt (Princeton University), Benjamin Raphael
OptimizationComputational EfficiencyRepresentation LearningMixture of ExpertsScore-based ModelAuto EncoderContrastive LearningGaussian SplattingImagePoint CloudTabularBiomedical DataBenchmark
🎯 What it does: Propose a new Transport Clustering (TC) algorithm, which converts the low-rank optimal transport (LR-OT) problem into a generalized K-means clustering problem. After registering with a known full-rank optimal transport plan, the low-rank transport plan can be obtained by performing clustering only once.
Transport or Discard: Robust Unbalanced Optimal Transport for Cross-Domain Policy Adaptation
Wenyu Chen (North University of China), Jianchao Zeng (North University of China)
Domain AdaptationOptimizationReinforcement LearningContrastive LearningTabularTime SeriesSequentialBenchmark
🎯 What it does: Proposed a robust unbalanced optimal transport framework called ROOT, which automatically evaluates cross-domain usability by leveraging the geometric similarity of source data, and either transports or discards incompatible samples.
Transporting Task Vectors across Different Architectures without Training
Filippo Rinaldi (University of Modena and Reggio Emilia), Simone Calderara (University of Modena and Reggio Emilia)
Domain AdaptationKnowledge DistillationRepresentation LearningTransformerSupervised Fine-TuningAuto EncoderContrastive LearningImageText
🎯 What it does: Proposes a training-free task vector transfer framework called THESEUS, which can transfer task-specific parameter updates obtained during fine-tuning between Transformers of different widths (even different pre-training distributions or depths).
TranX-Adapter: Bridging Artifacts and Semantics within MLLMs for Robust AI-generated Image Detection
Wenbin Wang (Wuhan University), Yong Luo (Wuhan University)
Anomaly DetectionTransformerLarge Language ModelVision Language ModelContrastive LearningImageMultimodality
🎯 What it does: Designed and implemented a lightweight fusion adapter, TranX-Adapter, to bridge texture-level artifact features with semantic features in multimodal large language models (MLLM), thereby enhancing the ability to detect AI-generated images (AIGI).
TRAP: Hijacking VLA CoT-Reasoning via Adversarial Patches
Zhengxian Huang (Zhejiang University), Wenyuan Xu (Zhejiang University)
Adversarial AttackRobotic IntelligenceTransformerPrompt EngineeringVision-Language-Action ModelDiffusion modelContrastive LearningImageTextMultimodalityChain-of-Thought
🎯 What it does: Proposes a targeted behavior hijacking attack called TRAP against vision-language-action (VLA) models with chain-of-thought (CoT), by placing adversarial patches in the scene to manipulate the model's intermediate CoT reasoning, thereby making the robot perform malicious actions specified by the attacker without altering the user's instruction.
Treatment Responder Classification with Abstention
Haoxiang Wang (Peking University), Mingming Gong (University of Melbourne)
ClassificationOptimizationFederated LearningExplainability and InterpretabilityDrug DiscoveryReinforcement LearningContrastive LearningTabularBiomedical DataElectronic Health Records
🎯 What it does: Studies how to learn treatment responder classification rules under the allowance of abstention.
Tree-Structured Orthonormal Decomposition of the Aitchison Simplex
Daisuke Yamada (University of Wisconsin Madison), Vikas Singh (University of Wisconsin Madison)
Explainability and InterpretabilityRepresentation LearningContrastive LearningBiomedical Data
🎯 What it does: Propose an orthogonal reference system called PolyILR, which maps the Aitchison simplex space to a coordinate system corresponding to any given tree structure;
TreeCUA: Efficiently Scaling GUI Automation with Tree-Structured Verifiable Evolution
Deyang Jiang (Meituan), Zhixiong Zeng (Meituan)
Autonomous DrivingOptimizationComputational EfficiencyRobotic IntelligenceReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsVision-Language-Action ModelDiffusion modelWorld ModelTextSequentialRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: TreeCUA proposes a tree-structured, verifiable evolutionary multi-agent framework, achieving efficient synthesis of large-scale GUI automation trajectories.
TreePO: Enhancing Policy Efficacy and Inference Efficiency with Tree Modeling
Yizhi LI, Wenhao Huang (Seed ByteDance)
Computational EfficiencyTransformerLarge Language ModelReinforcement LearningPrompt EngineeringText
🎯 What it does: Proposes TreePO, a framework that transforms the reinforcement learning sampling of LLMs into a tree-structured segmented search, significantly reducing computational costs while maintaining or enhancing the quality of reasoning.
Trees to Flows and Back: Unifying Decision Trees and Diffusion Models
Sai Niranjan Ramachandran (Technical University of Munich), Suvrit Sra (Technical University of Munich)
GenerationData SynthesisKnowledge DistillationDiffusion modelScore-based ModelFlow-based ModelRectified FlowTabularOrdinary Differential Equation
🎯 What it does: Propose a strict correspondence between decision trees and diffusion models, and based on this, demonstrate the optimality of gradient boosting in discrete domains through a unified 'Global Trajectory Score Matching (GTSM)' framework;
Tri-Scale Neural ODEs for Continuous Multi-Omics Disease Modeling
Shohaib Shaffiey (University of Nebraska-Lincoln), Massimiliano Pierobon (University of Nebraska-Lincoln)
Drug DiscoveryRecurrent Neural NetworkAuto EncoderContrastive LearningTabularTime SeriesSequentialBiomedical DataStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Design and verify a three-scale neural ODE framework to model disease progression over continuous time using multi-omics data (transcriptomics, proteomics, metabolomics), and perform drug repurposing screening based on this.
Triadic Dynamics Aware Diffusion Posterior Sampling for Inverse Problems: Optimizing Guidance and Stochasticity Schedules
Junseo Bang (Seoul National University), Se Young Chun (Seoul National University)
RestorationSuper ResolutionTransformerReinforcement LearningDiffusion modelScore-based ModelImage
🎯 What it does: Propose the TriPS framework, which improves diffusion posterior sampling for solving inverse problems by dynamically scheduling three components over time: data consistency (DC), classifier-free guidance (CFG), and randomness.
TriAttention: Efficient Long Reasoning with Trigonometric KV Compression
Weian Mao (Massachusetts Institute of Technology), Yukang Chen (Nvidia)
Computational EfficiencyKnowledge DistillationTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation
🎯 What it does: This paper proposes a KV cache compression method based on trigonometric functions called TriAttention, aimed at improving the efficiency of large language models during long-text reasoning.
TriForces: Augmenting Atomistic GNNs for Transferable Representations
Ali Ramlaoui (Entalpic), Joseph Musielewicz (Entalpic)
Computational EfficiencyKnowledge DistillationRepresentation LearningDrug DiscoveryGraph Neural NetworkTransformerSupervised Fine-TuningAuto EncoderContrastive LearningGraphTabularBenchmark
🎯 What it does: Propose the TriForces three-stream framework and self-supervised pre-training to enhance the transferability and retrieval performance of MLIP.
Trifuse: Enhancing Attention-Based GUI Grounding via Multimodal Fusion
Longhui Ma (National University of Defense Technology), Miao Wang (Academy of Military Sciences)
RecognitionImage TranslationRestorationObject DetectionRetrievalComputational EfficiencyTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposes the Trifuse framework, achieving GUI grounding by fusing heatmaps from three modalities—attention, OCR text information, and icon captions—without task-specific fine-tuning.
TRIM: Token-wise Attention-Derived Saliency for Data-Efficient Instruction Tuning
Manish Nagaraj (Purdue University), Kaushik Roy (Purdue University)
Explainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerSupervised Fine-TuningPrompt EngineeringContrastive LearningText
🎯 What it does: Propose TRIM (Token-wise Attention-Derived Saliency for Instruction Tuning), a forward-only, attention-based token-level coreset selection method for instruction tuning.
TRIP-Bench: A Benchmark for Long-Horizon Interactive Agents in Real-World Scenarios
Yuanzhe Shen (Meituan), Ke Zeng (Meituan)
TransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextSequentialBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: A long-sequence interactive evaluation benchmark called TRIP-Bench based on real travel planning scenarios was constructed, and a reinforcement learning method called GTPO tailored for multi-round interactions was proposed.
TritonGym: A Benchmark for Agentic LLM Workflows in Triton GPU Code Generation
Yue Guan (University of California), Adnan Aziz (Meta)
OptimizationLarge Language ModelAgentic AITextBenchmark
🎯 What it does: Proposes the TritonGym benchmark, providing a standardized function call interface for evaluating agentic workflows in GPU kernel generation.
Trojan-Speak: Bypassing Constitutional Classifiers with No Jailbreak Tax via Adversarial Finetuning
Bilgehan Sel (Anthropic), Jerry Wei (Anthropic)
Safty and PrivacyAdversarial AttackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextBenchmark
🎯 What it does: Studied a对抗微调 method called Trojan-Speak, which can bypass Anthropic's constitutional classifier without significantly compromising the model's capabilities.
Trust Functions: Near Lossless Weak-to-Strong Generalization by Learning to Trust the Weak Teacher
Arda Uzunoglu (Johns Hopkins University), Daniel Khashabi
Explainability and InterpretabilityKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningTextBenchmark
🎯 What it does: Proposed and implemented a trust function based on neural networks to evaluate the reliability of weak labels from weak teachers, thereby achieving generalization from weak to strong.
Trust It or Not: Evidential Uncertainty for Feed-Forward 3D Reconstruction with Trust3R
Zihao Zhu (Texas A&M University), Zhiwen Fan (Texas A&M University)
Pose EstimationDepth EstimationExplainability and InterpretabilityComputational EfficiencyTransformerAuto EncoderContrastive LearningGaussian SplattingPoint Cloud
🎯 What it does: Propose the Trust3R framework, incorporating interpretable evidence uncertainty estimation into Feed-Forward 3D reconstruction, generating a multivariate Student-t predictive distribution for point clouds.
Trust Region Inverse Reinforcement Learning: Explicit Dual Ascent using Local Policy Updates
Anish Abhijit Diwan (Technical University of Darmstadt), Oleg Arenz (Technical University of Darmstadt)
OptimizationReinforcement LearningContrastive LearningTabularTime SeriesSequential
🎯 What it does: Propose a Trust Region Inverse Reinforcement Learning (TRIRL) algorithm that utilizes trust region policy updates and reward correction to achieve monotonic improvement in non-adversarial IRL, learning a reward function that is globally optimal and robust to changes in system dynamics.
Trust Region Masking for Long-Horizon LLM Reinforcement Learning
Yingru Li, Baoxiang Wang (The Chinese University Of Hong Kong Shenzhen)
TransformerLarge Language ModelReinforcement LearningPrompt EngineeringText
🎯 What it does: Studied and proposed Trust Region Masking (TRM) at the sequence level to address the training collapse problem in long-sequence LLM RL caused by implementation bias.
Trust-Region Diffusion Policies for Massively Parallel On-Policy RL
Huy Le (Bosch), Gerhard Neumann (Karlsruhe Institute Of Technology)
Reinforcement LearningDiffusion modelStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Proposes Trust-Region Diffusion Policies (TruDi), a training framework that uses diffusion models in on-policy RL within large-scale parallel simulation environments.
Trustworthy Federated Label Distribution Learning under Annotation Quality Disparity
Junxiang Wu (Southeast University), Xin Geng (Southeast University)
ClassificationImage TranslationImage HarmonizationRestorationRecommendation SystemFederated LearningAuto EncoderGenerative Adversarial NetworkContrastive LearningImageVideoTextMultimodalityBenchmarkAudio
🎯 What it does: Propose the FedQual framework to address the credibility issue caused by annotation quality differences in federated label distribution learning.
TrustworthyQENN: A Quantum Evidential Neural Network Based on Complex-Valued Contrastive Learning for Uncertainty Pattern Classification
Xiaolong Chen (Chongqing University), Chin-teng Lin
ClassificationAnomaly DetectionConvolutional Neural NetworkContrastive LearningImage
🎯 What it does: Propose a Trustworthy Quantum Evidence Neural Network (TrustworthyQENN), which extracts amplitude-phase collaborative features through supervised complex-valued contrastive learning, formally defines OOD states as quantum empty sets, and utilizes GQET for quantum evidence generation and fusion, achieving reliable detection of OOD.
Truthfulness Does Not Scale Like Reasoning: Why Polling Fails as a Proxy Verifier
Yegor Denisov-Blanch (Stanford Computer Science), Sanmi Koyejo (Stanford Computer Science)
ClassificationTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmark
🎯 What it does: Investigate whether aggregation methods (voting, confidence-weighted, surprise-based, etc.) can improve the honesty of language models in scenarios without external verifiers.
TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning
Zhepei Wei (University of Virginia), Xin Luna Dong (Meta Reality Labs)
Data-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose the TruthRL framework, which directly optimizes the honesty of LLMs through ternary rewards (correct answer reward +1, unknown/rejection reward 0, hallucination penalty -1) using reinforcement learning, encouraging the model to reasonably refuse to answer when uncertain.
TSFAdv: Frequency-Guided Black-Box Adversarial Attacks on Time Series Forecasting
Qizhuo Han (Nankai University), Zheli Liu (Nankai University)
Adversarial AttackTime Series
🎯 What it does: Developed TSFAdv, a black-box attack framework based on frequency-domain sensitivity analysis and natural evolution strategies, capable of inducing significant errors in long-sequence prediction models within a limited number of queries.
TsLLM: Augmenting LLMs for General Time Series Understanding and Prediction
Felix Parker (Johns Hopkins University), Kimia Ghobadi (Johns Hopkins University)
ClassificationAnomaly DetectionOptimizationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringAuto EncoderContrastive LearningTextMultimodalityTime SeriesBenchmarkFinance Related
🎯 What it does: A unified autoregressive model that can alternately process text and time series within the same sequence was achieved by adding a specialized time series encoder-decoder (patch-based β-VAE) to a pre-trained large language model (LLM), and embedding it into the LLM's embedding space through cross-attention adapters. This model supports multiple tasks such as time series forecasting, classification, anomaly detection, and question answering.
TSMGen: Target-Specific Molecule Generation via Higher-Order Structural Dependencies and Context-Aware Bidirectional Fusion
Yaoyu Chen (Wuhan University of Science and Technology), Jun Pang (Wuhan University of Science and Technology)
Drug DiscoveryGraph Neural NetworkTransformerDiffusion modelContrastive LearningGraphBiomedical Data
🎯 What it does: Proposed the TSMGen framework, which uses a hypergraph to capture high-order structural dependencies in protein pockets, and employs a context-aware bidirectional fusion module to achieve deep complementarity between molecular and pocket features, generating high-affinity molecules specific to particular targets.
TSP with Predictions: Heatmap to Tour with Provable Guarantees
Marek Elias, Eleonora Vercesi (Università della Svizzera italiana)
OptimizationGraphTabularBenchmark
🎯 What it does: Propose a learning-augmented algorithm that converts the heatmap prediction of TSP into a feasible route;
TSRBench: A Comprehensive Multi-task Multi-modal Time Series Reasoning Benchmark for Generalist Models
Fangxu Yu (University of Maryland), Tianyi Zhou (Mohamed bin Zayed University of Artificial Intelligence)
Autonomous DrivingTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextMultimodalityTime SeriesBenchmarkFinance RelatedRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposed TSRBENCH, a comprehensive multitask multimodal time series reasoning benchmark for evaluating the capabilities of general models in four dimensions: perception, reasoning, prediction, and decision-making;