ICML 2026 Papers — Page 62
International Conference on Machine Learning · 6554 papers
TT-Sparse: Learning Sparse Rule Models with Differentiable Truth Tables
Hans Farrell Soegeng (Nanyang Technological University), Thomas Peyrin (Nanyang Technological University)
ClassificationOptimizationExplainability and InterpretabilityAuto EncoderContrastive LearningTabular
🎯 What it does: Designed and implemented TT-SPARSE, a sparse rule learning module based on differentiable truth tables, to generate globally interpretable and precise rule sets within a single network.
Tucker Attention: A generalization of approximate attention mechanisms
Timon Klein (Otto von Guericke University Magdeburg), Steffen Schotthöfer (Oak Ridge National Laboratory)
Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelAuto EncoderContrastive LearningImageText
🎯 What it does: This paper proposes a general Tucker Attention mechanism, treating the pre-softmax and post-softmax weights of self-attention as third-order tensors and performing Tucker decomposition on them, unifying approximate attention methods such as MHA, GQA, and MLA, achieving higher parameter and KV cache compression.
TuneAhead: Predicting Fine-tuning Performance Before Training Begins
Yuxiang Luo (Hong Kong University of Science and Technology (Guangzhou)), Nan Tang (Hong Kong University of Science and Technology (Guangzhou))
Explainability and InterpretabilityComputational EfficiencyHyperparameter SearchData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringContrastive LearningText
🎯 What it does: Propose the TUNEAHEAD framework, which provides fine-grained performance prediction and can diagnose failure reasons by combining a 100-step probe with static dataset descriptors before the full fine-tuning begins.
Tuning the Implicit Regularizer of Masked Diffusion Language Models: Enhancing Generalization via Insights from $k$-Parity
Jianhao Huang (University of California, Los Angeles), Baharan Mirzasoleiman (University of California, Los Angeles)
GenerationRepresentation LearningData-Centric LearningTransformerSupervised Fine-TuningDiffusion modelText
🎯 What it does: Analyzed and improved the training process of masked diffusion language models (MDLM) on the k-parity task, demonstrating that their objective inherently includes signal-noise decomposition, with the noise part acting as implicit regularization to eliminate 'grokking'; based on this, proposed a signal-rich masked sampling strategy;
Tuning-Free One-Class Discriminant Learning for Tabular Anomaly Detection
Xuan-Ha Nguyen (University College Dublin), Nhien-An Le-Khac (University College Dublin)
Anomaly DetectionAuto EncoderContrastive LearningTabular
🎯 What it does: Propose a no-parameter tuning one-class discriminative learning method called DVM-AD for anomaly detection on tabular data.
TUR-DPO: Topology- and Uncertainty-Aware Direct Preference Optimization
Abdulhady abas, Mourad Oussalah (University of Oulu)
OptimizationComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringContrastive LearningTextGraphBenchmark
🎯 What it does: Proposes TUR-DPO, an RL-free alignment method that incorporates reasoning topology and uncertainty weighting into direct preference optimization.
Turbo Connection: Reasoning as Information Flow from Higher to Lower Layers
Tang Mohan (UCLA), Sidi Lu (UCLA)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningTextChain-of-Thought
🎯 What it does: Introduce a 'Turbo Connection' mechanism into the Transformer decoder, which adds downward residual connections between different layers and between different tokens, allowing the model's effective inference depth to grow linearly with the sequence length.
Turbo4DGen: Ultra-Fast Acceleration for 4D Generation
Yuanbin Man (University of Texas at Arlington), Miao Yin (University of Texas at Arlington)
GenerationComputational EfficiencyTransformerDiffusion modelScore-based ModelImageVideo
🎯 What it does: Propose Turbo4DGen, a high-speed acceleration framework for 4D generation, significantly reducing computational and memory consumption in multi-view 4D content generation.
TurboGS: Accelerating 3D Gaussian Splatting via Error-Guided Sparse Pixel Sampling and Optimization
Zheng Dong (Zhejiang Sci-Tech University), Weiwei Xu (Zhejiang University)
OptimizationComputational EfficiencyNeural Radiance FieldGaussian SplattingImagePoint Cloud
🎯 What it does: The TurboGS framework is constructed through error-guided sparse pixel sampling and sparse supervision, significantly reducing unnecessary pixel computations during the training process of 3D Gaussian Splatting, and achieving fast optimization while maintaining or even improving rendering quality.
Turning Adaptation into Assets: Cross-Domain Bridging for Online Vision-Language Navigation
Zixuan Hu (Peking University), LINGYU DUAN
Domain AdaptationOptimizationComputational EfficiencyReinforcement LearningPrompt EngineeringVision-Language-Action ModelContrastive LearningImageTextMultimodality
🎯 What it does: Propose the IDEA framework, which transforms online adaptation in the VLN task into a sustainable asset library for accumulation and reuse, and constructs dynamic assets through soft prompts and domain coordinates; utilize Fisher-guided weighted optimization for soft prompts, and rapidly construct cross-domain bridges without training via convex hull projection;
Turning Back Without Forgetting: Selective Backward Refinement for Parameter-Efficient Continual Learning
Anushka Tiwari (University at Buffalo), Kaiyi Ji (University at Buffalo)
Computational EfficiencyRepresentation LearningTransformerPrompt EngineeringTextBenchmark
🎯 What it does: Proposed a replay-free SABER framework to achieve controllable forward-backward knowledge transfer in prompt-based continual learning;
Turning Bias into Bugs: Bandit-Guided Style Manipulation Attacks on LLM Judges
XIANGLIN YANG, Jin Song Dong (National University of Singapore)
Explainability and InterpretabilityAdversarial AttackReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: This work explores black-box attacks on the style preferences of answer text using large language model (LLM) judges, proposing the BITE (Bandit-guided Style Manipulation Attack) framework. It can artificially improve the evaluation scores through semantically preserved style editing without accessing internal parameters.
Turning Drift into Constraint: Robust Reasoning Alignment in Non-Stationary Multi-Stream Environments
Xiaoyu Yang (University of Technology Sydney), Jie Lu (University of Technology Sydney)
Domain AdaptationAutonomous DrivingOptimizationTransformerSupervised Fine-TuningMixture of ExpertsContrastive LearningImageBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission TomographyUltrasoundElectronic Health RecordsRetrieval-Augmented Generation
🎯 What it does: Proposes the Autonomous Preference Optimization (APO) framework, which achieves robust inference alignment in non-stationary multi-stream environments by treating drift from multi-source models as negative constraints.
Turning Stale Gradients into Stable Gradients: Coherent Coordinate Descent with Implicit Landscape Smoothing for Lightweight Zeroth-Order Optimization
Chen Liang (Yale University), Daniel Rakita (Yale University)
OptimizationImageTabular
🎯 What it does: Proposed a new zeroth-order optimization method called Coherent Coordinate Descent (CoCD), which achieves efficient deterministic coordinate descent by utilizing old gradients and implicit smoothing.
Tvcache: A Tool-Value Cache for Post-Training LLM Agents
Abhishek Vijaya Kumar (Cornell University), Rachee Singh (Cornell University)
Computational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextTabularSequentialBenchmarkRetrieval-Augmented Generation
🎯 What it does: Designed a state-aware tool call caching system named TVCACHE to accelerate post-training of LLM agents.
TVDRNet: Text-driven Viewpoint Optimization via Differentiable Rendering for 3D Reasoning Segmentation
Zaiyang Yu (Chinese Academy of Sciences), Weijun Li (University of Chinese Academy of Sciences)
SegmentationTransformerSupervised Fine-TuningPrompt EngineeringVision Language ModelDiffusion modelNeural Radiance FieldContrastive LearningImageTextMultimodalityPoint Cloud
🎯 What it does: Propose TVDRNet, which achieves 3D reasoning segmentation through text-driven differentiable rendering, first learning text-guided view parameters and then fusing point cloud and multi-view features to generate segmentation masks.
TVI-CoT: Text-Visual Interleaved Chain-of-Thought Reasoning for Multimodal Understanding
Lianyu Hu (Nanyang Technological University), Yang Liu
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningVision Language ModelImageTextMultimodalityChain-of-Thought
🎯 What it does: Propose the TVI-CoT framework, which achieves interleaved access between text reasoning and visual features through learnable control tokens, addressing the vision-blind reasoning issue in multi-modal LLMs.
Twice Sequential Monte Carlo for Tree Search
Yaniv Oren (Delft University of Technology), Wendelin Boehmer
Reinforcement Learning
🎯 What it does: Developed and evaluated a tree search algorithm called TSMCTS based on sequential Monte Carlo (SMC) for policy improvement in reinforcement learning.
TwinQuant: Learnable Subspace Decomposition for 4-Bit LLM Quantization
Haodong WANG, Xu Chen (Sun Yat-sen University)
Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelContrastive LearningTextStochastic Differential Equation
🎯 What it does: This paper proposes the TwinQuant framework, which uses a learnable subspace decomposition to divide the weights into low-rank and residual components, and quantizes both parts using 4-bit quantization.
Twins: Learn to Predict Unified Representations with Focal Loss
Kaixiong Gong (Chinese University of Hong Kong), Xiangyu Yue (Chinese University of Hong Kong)
RestorationGenerationRepresentation LearningTransformerDiffusion modelFlow-based ModelAuto EncoderContrastive LearningImageMultimodality
🎯 What it does: Propose Twins to unify the continuous visual token space by concatenating ViT (SigLIP2) features with VAE (Flux.2) latent vectors along the channel dimension on the same token grid, using a Diffusion Transformer for flow matching training, and introducing Focal Loss to reweight the VAE dimension, addressing the optimization imbalance caused by overfitting to ViT.
TwinWeaver: An LLM-Based Foundation Model Framework for Pan-Cancer Digital Twins
Nikita Makarov (Roche), Michael Patrick Menden (Helmholtz Munich)
Explainability and InterpretabilityRepresentation LearningData-Centric LearningDrug DiscoveryReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningTextTabularTime SeriesBiomedical DataElectronic Health RecordsRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose the TwinWeaver framework, which leverages LLM to serialize multimodal medical records into text, training the Genie Digital Twin (GDT) to achieve joint time series prediction and clinical event prediction across 20 types of cancer;
TWLA: Achieving Ternary Weights and Low-Bit Activations for LLMs via Post-Training Quantization
Zhixiong Zhao (Houmo AI), Dawei Yang (Houmo AI)
Computational EfficiencyKnowledge DistillationTransformerLarge Language ModelAuto EncoderContrastive LearningText
🎯 What it does: Propose a post-training quantization framework called TWLA, which can compress the weights of large language models to 1.58 bits and activations to 4 bits while maintaining high precision;
Two Calm Ends and the Wild Middle: A Geometric Picture of Memorization in Diffusion Models
Nick Dodson (University of Missouri), Zhengchao Wan (University of California San Diego)
GenerationData SynthesisDiffusion modelScore-based ModelGaussian SplattingImage
🎯 What it does: Investigate the memorization mechanism of diffusion models under different noise levels, and propose a geometric framework based on posterior weight aggregation and Gaussian shell coverage, identifying the intermediate noise region as a high-risk memorization 'danger zone', and suppressing memorization by skipping or reducing the noise intensity in this region during training.
Two Modalities Are Better Than One: Efficient Adversarial Purification via Multimodal Diffusion Models
Mingyuan Bai (RIKEN), Qibin Zhao (RIKEN)
RestorationComputational EfficiencyAdversarial AttackTransformerPrompt EngineeringDiffusion modelScore-based ModelImageTextMultimodality
🎯 What it does: MultiDAP proposes a framework for efficient adversarial purification using multimodal diffusion models (text-image).
Two-dimensional quantization for geometry-aware audio coding
Tal Shuster (Ben Gurion University Of Negev), Eliya Nachmani (Ben Gurion University Of Negev)
CompressionAuto EncoderAudio
🎯 What it does: This paper proposes a two-dimensional geometric quantization (Q2D2), which maps features to hexagonal, rectangular, or rhombic grids for quantization;
Two-Layer Linear Auto-Regressive Models Estimate Latent States
Yahya Sattar (Cornell University), Sarah Dean (Cornell University)
OptimizationRepresentation LearningReinforcement Learning from Human FeedbackRecurrent Neural NetworkAuto EncoderTabularTime SeriesBenchmark
🎯 What it does: Train a two-layer linear autoregressive model, using empirical risk minimization to learn hidden state estimation for partially observed linear dynamical systems from a single trajectory, demonstrating that the model naturally approximates the hidden layer representation of the Kalman filter.
Two-Parameter Flows for Learning Population Dynamics of Physical Systems
Paul Schwerdtner (New York University), Benjamin Peherstorfer (New York University)
TransformerScore-based ModelFlow-based ModelRectified FlowAuto EncoderTabularTime SeriesPhysics RelatedStochastic Differential Equation
🎯 What it does: Learning the dynamics of high-dimensional probability density evolution over time, using only unlabeled temporal marginal samples rather than trajectories.
Two-Stage Unit Tying for Simplifying Differentiable Logic Gate Networks
Seungheon Lee (Yonsei University), Jaeyong Chung (Yonsei University)
ClassificationComputational EfficiencyKnowledge DistillationContrastive LearningGaussian SplattingImage
🎯 What it does: Proposes a unit tying method for post-training simplified differentiable logic gate networks, which can compress the model's logical operations into the deployable range of FPGAs.
U-Cast: A Surprisingly Simple and Efficient Frontier Probabilistic AI Weather Forecaster
Salva Rühling Cachay (UC San Diego), Rose Yu (UC San Diego)
OptimizationComputational EfficiencyRepresentation LearningAI Code AssistantConvolutional Neural NetworkTransformerMixture of ExpertsAuto EncoderContrastive LearningTabularTime Series
🎯 What it does: Developed a probabilistic weather forecasting model named U-Cast, which utilizes the standard U-Net architecture and achieves high-precision 15-day medium-range forecasts through a two-stage training process (pre-training with MAE, followed by fine-tuning with CRPS).
U$^3$CF: Unbiased, Unconfounding, and Unified Causal Framework for Multi-Target Domain Adaptation
Wenxu Wang (Ocean University of China), Wenbo Gong (Microsoft Research)
Domain AdaptationGraph Neural NetworkTransformerContrastive LearningImage
🎯 What it does: Propose the U CF³ framework to achieve multi-target domain adaptation, realizing unbiased and non-confusing generalization through prototype-driven progressive alignment, feature disentanglement, and causal classification.
UAV$^2$: A Unified and Adaptive Scheduling Framework for UAV Autopilot System with Reinforcement Learning
Zeying Li (Sun Yat-sen University), Kai Huang (Sun Yat-sen University)
Autonomous DrivingOptimizationRecurrent Neural NetworkTransformerReinforcement LearningSimultaneous Localization and MappingOptical FlowTime SeriesSequential
🎯 What it does: Propose a unified adaptive scheduling framework based on reinforcement learning, UAV 2, which integrates navigation and flight control into a single platform and learns task frequency scheduling through RL.
UB-SMoE: Universally Balanced Sparse Mixture-of-Experts for Resource-adaptive Federated Fine-tuning of Foundation Models
Van-Tuan Tran (Trinity College Dublin), Merim Dzaferagic (Trinity College Dublin)
OptimizationFederated LearningSafty and PrivacyComputational EfficiencyTransformerLarge Language ModelSupervised Fine-TuningMixture of ExpertsContrastive LearningTextTabular
🎯 What it does: To address the computational capability differences among devices in federated learning, the UB-SMoE method is proposed, which utilizes a sparse expert network (SMoE) under the LoRA adaptation framework to achieve resource-adaptive model fine-tuning.
Ubiquity of Emergent Hebbian Dynamics in Regularized Learning
David Aaron Koplow (Massachusetts Institute of Technology), Liu Ziyin (Massachusetts Institute of Technology)
OptimizationRepresentation LearningConvolutional Neural NetworkTransformerContrastive LearningImageTabularSequential
🎯 What it does: The study investigates how regularization (L2 weight decay) and noise lead to Hebbian or anti-Hebbian characteristics in learning updates during gradient descent training, and experiments are conducted across various networks, tasks, and optimizers to verify these findings.
UCPO: Uncertainty-Aware Policy Optimization
Xianzhou Zeng (Ant Group), Xingzhong Xu (Ant Group)
OptimizationTransformerLarge Language ModelReinforcement LearningPrompt EngineeringText
🎯 What it does: By introducing ternary rewards into the reinforcement learning framework and proposing two techniques, Ternary Advantage Decoupling and Dynamic Uncertainty Reward Adjustment, large language models can actively express uncertainty during reasoning, thereby reducing errors caused by overconfidence;
UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models
Jiaqi Wang (Beijing University of Posts and Telecommunications), Xinlong Wang (Beijing Academy of Artificial Intelligence)
GenerationOptimizationReinforcement Learning from Human FeedbackTransformerReinforcement LearningPrompt EngineeringDiffusion modelScore-based ModelImageTextMultimodality
🎯 What it does: Proposed and implemented UDM-GRPO, which integrates the Uniform Discrete Diffusion model with Group Relative Policy Optimization (GRPO), for text-to-image generation tasks.
UFO: Chain-of-Evaluation for Omni-Condition Alignment in Multi-Modal Image Generation
Danning Zhang (University of Science and Technology of China), Zhendong Mao (Institute of Artificial Intelligence Hefei Comprehensive National Science Center)
GenerationData SynthesisTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBenchmarkChain-of-Thought
🎯 What it does: Propose the UFO framework, which atomically chain-evaluates full conditional alignment in multi-modal image generation, and constructs the UFO-Bench benchmark.
UGround: Towards Unified Visual Grounding with Unrolled Transformers
Rui Qian (Fudan University), Dejing Dou (Fudan University)
SegmentationReinforcement Learning from Human FeedbackTransformerReinforcement LearningPrompt EngineeringVision Language ModelContrastive LearningImageText
🎯 What it does: Proposes UGround, a unified visual orientation segmentation framework that dynamically selects intermediate layers in the unrolled Transformer as 'mask as prompt' and directly connects with SAM through skip-connection.
UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing
Yunkai Dang (Nanjing University), Yang Gao (Nanjing University)
CompressionComputational EfficiencyTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodality
🎯 What it does: Propose a budget-aware token compression framework called UHR-BAT for ultra-high-resolution remote sensing images, which can maintain query-related details and global context under strict visual token budgets;
UI2Code^N: UI-to-Code Generation as Interactive Visual Optimization
Zhen Yang (Tsinghua University), Jie Tang (Tsinghua University)
GenerationOptimizationAI Code AssistantTransformerSupervised Fine-TuningReinforcement LearningVision Language ModelImageTextMultimodalityChain-of-Thought
🎯 What it does: Convert UI screenshots into executable frontend code and propose an interactive visual optimization framework that allows the code to iteratively improve through executable rendering feedback.
Ultrafast On-Chip Online Learning via Spline Locality in Kolmogorov–Arnold Networks
Duc Hoang (MIT), Philip Harris (MIT)
OptimizationComputational EfficiencyReinforcement LearningAuto EncoderContrastive LearningTabularTime SeriesSequentialBenchmark
🎯 What it does: Implement and evaluate Kolmogorov-Arnold network (KAN) for ultra-fast online learning on FPGA, utilizing B-spline locality to achieve sparse gradient updates and performing full-chip training under fixed precision;
UltraHorizon: Benchmarking LLM-Agent Capabilities in Ultra Long-Horizon Scenarios
Haotian Luo (Didichuxing Co. Ltd), Li Shen (Sun Yatsen University)
Autonomous DrivingOptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyRobotic IntelligenceMeta LearningDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackNeural Architecture SearchProtein Structure PredictionTransformerLarge Language ModelAgentic AIPrompt EngineeringTextSequentialBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose the UltraHorizon benchmark, which evaluates the ability of LLM agents in ultra-long time and partially observable environments through exploration tasks.
UltraLIF: Fully Differentiable Spiking Neural Networks via Ultradiscretization and Max-Plus Algebra
Jose Marie Antonio Miñoza (Center for AI Research Philippines)
ClassificationComputational EfficiencyRepresentation LearningSpiking Neural NetworkSupervised Fine-TuningContrastive LearningImageVideoTabularBenchmarkAudio
🎯 What it does: Propose two fully differentiable spiking neural network models, UltraLIF/UltraDLIF, by replacing the empirical surrogate gradient with ultradiscretization.
UMEM: Unified Memory Extraction and Management Framework for Generalizable Memory
Yongshi Ye (Institute of Artificial Intelligence, Xiamen University), Weihua Luo (Alibaba Group)
OptimizationRepresentation LearningTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringContrastive LearningTextRetrieval-Augmented Generation
🎯 What it does: Proposed a self-evolving agent framework called UMEM, which enhances the long-term memory capability of large models through joint optimization of memory extraction and management.
Unbiased Alignment for Large Language Models with Noisy Preferences
Jialiang Wang (Harbin Institute of Technology), Haoliang Li (City University of Hong Kong)
OptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringContrastive LearningText
🎯 What it does: This paper proposes a theoretical framework that can achieve unbiased alignment of large language models using only noisy preference data.
Unbiased and Second-Order-Free Training for High-Dimensional PDEs
Jaemin Seo (Chung-Ang University), Jae Yong Lee (Chung-Ang University)
OptimizationComputational EfficiencyScore-based ModelTabularTime SeriesBenchmarkPhysics RelatedStochastic Differential Equation
🎯 What it does: Proposes an unbiased BSDE training framework (Un-EM-BSDE) that does not require second-order derivatives, eliminating the bias caused by Euler-Maruyama discretization through random independent sampling.
Unbiased Dynamic Pruning for Efficient Group-Based Policy Optimization
Haodong Zhu (Beihang University), Baochang Zhang (Beihang University)
OptimizationComputational EfficiencyTransformerReinforcement LearningPrompt EngineeringTextChain-of-Thought
🎯 What it does: Propose a dynamic unbiased pruning framework DPPO, which combines importance sampling to compensate for distribution shift, significantly accelerating the GRPO algorithm;
Unbiased Principles, Robust Rewards
Qingnan Ren (University of Science and Technology of China), Feng Zhao (University of Science and Technology of China)
Reinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringText
🎯 What it does: Proposes a generative reward model generated under the independent principle (IP-GRM) and introduces a principle caching mechanism to eliminate principle drift and reward hijacking, thereby improving the quality and robustness of rewards in open-ended tasks.
Unbiased Reward Modeling from Implicit Feedback for LLM Alignment
Hao Wang (Xiaohongshu Inc.), Zhouchen Lin (Peking University)
OptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningContrastive LearningText
🎯 What it does: Proposed the ImplicitRM framework, which learns an unbiased reward model from implicit user feedback (such as clicks, copies, skips) to improve the alignment of large language models.
Uncertainty-Aware Clarification in LLM Agents with Information Gain
Mengyi DENG, Wei Wang (Hong Kong University of Science and Technology)
OptimizationExplainability and InterpretabilityTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringTextBenchmark
🎯 What it does: Propose a clarification framework based on information gain rewards, enabling LLM agents to proactively raise targeted clarification questions before tool calls to reduce uncertainty.
Uncertainty-Constrained Trustworthiness for Graph Learning
Chunhui Zhang (Beijing Institute of Technology), Guoren Wang
Graph Neural NetworkGraph
🎯 What it does: This paper proposes the DICT framework, which utilizes Wasserstein distribution uncertainty to perturb graph structures, features, and labels, thereby achieving dual improvements in robustness and fairness of graph learning models.
Uncertainty-Guided Exploration and Stable Planning for Sparse-Reward Manipulation from Limited Demonstrations
Haowen Sun (Xi'an Jiaotong University), Xuguang Lan (Xi'an Jiaotong University)
Robotic IntelligenceReinforcement LearningMixture of ExpertsContrastive LearningWorld ModelImageVideoTabular
🎯 What it does: Propose a model-based RL framework called QUEST, which achieves multi-stage robotic manipulation under sparse rewards and limited demonstrations through uncertainty-guided exploration and stable planning, and realizes zero-shot simulation-to-reality transfer.
Uncovering Bias Mechanisms in Observational Studies
Ilker Demirel (Massachusetts Institute of Technology), David Sontag (Massachusetts Institute of Technology)
Federated LearningExplainability and InterpretabilityData-Centric LearningSupervised Fine-TuningContrastive LearningOptical FlowTabularBiomedical DataElectronic Health RecordsReview/Survey PaperStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: A diagnostic framework is constructed by linking the magnitude of bias in observational studies with the predictive performance of nuisance functions, aiming to reveal the mechanisms behind the bias.
Uncovering Competency Gaps in Large Language Models and Their Benchmarks
Maty Bohacek (Stanford University), Stephanie C.Y. Chan (Google DeepMind)
Explainability and InterpretabilityTransformerLarge Language ModelAuto EncoderTextBenchmark
🎯 What it does: This paper proposes an automated method based on the concept activation of sparse autoencoders (SAE), called Competency Gaps (CG), to finely identify the weaknesses of large language models (model gaps) and the shortcomings of benchmark evaluations (benchmark gaps).
Uncovering Grounding IDs: How External Cues Shape Multi-Modal Binding
Amirmohammad Izadi (Sharif University of Technology), Mahdieh Soleymani Baghshah (Sharif University of Technology)
Explainability and InterpretabilityRepresentation LearningTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodality
🎯 What it does: Investigated how external visual cues (such as symbols, grid lines) generate potential 'Grounding IDs' in large vision-language models, thereby enhancing the binding and reasoning capabilities between visual and textual information.
Uncovering Hidden Triggers: Backdoor Attribution in Language Models
Miao Yu (University of Science and Technology of China), Qingsong Wen (Squirrel Ai Learning)
Explainability and InterpretabilityRepresentation LearningAdversarial AttackTransformerLarge Language ModelPrompt EngineeringText
🎯 What it does: A three-stage interpretable framework called BkdAttr is proposed and studied for detecting, locating, and controlling backdoor attacks in LLMs.
Uncovering Latent Communication Patterns in Brain Networks via Adaptive Flow Routing
Tianhao Huang (University of Virginia), Chen Chen (University of Virginia)
ClassificationExplainability and InterpretabilityComputational EfficiencyGraph Neural NetworkTransformerFlow-based ModelContrastive LearningGraphBiomedical DataAlzheimer's DiseaseBenchmark
🎯 What it does: By constructing a physics-informed flow pathway network, integrating brain structural connectivity (SC) with functional connectivity (FC), AFR-Net is proposed to reveal potential neural communication pathways.
Uncovering the Gradient Geometry of Long CoT: A Spectral-guided Approach to Reasoning Distillation
Sinan Fan (Zhejiang University), Jieping Ye (Tongyi Lab, Alibaba Group)
Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerSupervised Fine-TuningPrompt EngineeringTextBenchmarkChain-of-Thought
🎯 What it does: Identify low-rank consensus subspaces in long Chain-of-Thought (CoT) by gradient spectral decomposition, dynamically select steps, and perform distillation to student models to enhance reasoning capabilities.
Uncovering the Latent Potential of Deep Intermediate Representations
Arnesh Batra (Indraprastha Institute of Information Technology), Anubha Gupta (Indraprastha Institute of Information Technology)
Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningSupervised Fine-TuningContrastive LearningImageTextMultimodalityAudio
🎯 What it does: Proposes a hierarchical optimal embedding selection method (LOES) and geometric regularization (GeoReg), which identify and fuse the most task-discriminative intermediate layer representations from pre-trained models through spectral decomposition, orthogonalization, and isotropic variance constraints, thereby achieving more efficient transfer learning and model interpretability.
Understand and Accelerate Memory Processing Pipeline for Large Language Model Inference
Zifan He (University of California Los Angeles), Jason Cong (University of California Los Angeles)
OptimizationComputational EfficiencyTransformerLarge Language ModelTextRetrieval-Augmented Generation
🎯 What it does: This paper unifies memory processing in LLM inference into a four-stage pipeline (Prepare Memory, Compute Relevance, Retrieve, Apply to Inference), revealing that it accounts for 22%-97% of inference latency, and proposes methods to accelerate different stages.
Understanding and Mitigating Token-Pruning-Induced Vulnerabilities in VLMs
Shuailong Wang (University of Electronic Science and Technology of China), Lianli Gao (University of Electronic Science and Technology of China)
Safty and PrivacyComputational EfficiencyTransformerVision Language ModelMultimodalityBenchmark
🎯 What it does: Propose a mechanism called 'Safe Awareness Pruning' (SAP) in vision-language models to mitigate cross-modal jailbreak risks caused by Token-Pruning without sacrificing inference acceleration.
Understanding Behavior Cloning with Action Quantization
Haoqun Cao (University of Wisconsin-Madison), Tengyang Xie (University of Wisconsin-Madison)
Explainability and InterpretabilityRepresentation LearningReinforcement LearningContrastive Learning
🎯 What it does: Studied the impact of action quantization on learning performance in behavioral cloning within continuous action spaces, establishing a theoretical analysis framework.
Understanding Catastrophic Forgetting In LoRA via Mean-Field Attention Dynamics
Hugo Koubbi (Université Paris Dauphine), Matthieu Boussard (Craft AI)
Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningAuto EncoderContrastive LearningTextStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Studies the catastrophic forgetting phenomenon caused by LoRA in fine-tuning Transformers, proposes and analyzes a computable mean-field self-attention toy model, treating forgetting as a geometric drift in representation space.
Understanding Data Temporality Impact on Large Language Models Pre-training
Romain Fabre (Kyutai), Edouard Grave (Kyutai)
Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextTime SeriesSequentialBenchmarkRetrieval-Augmented Generation
🎯 What it does: The study investigates the impact of data order during pre-training on the acquisition of temporal knowledge in large language models, constructing a time-sensitive question-answering benchmark named KairosQA with over 7,000 items, and comparing pre-trained models of 6B parameters trained with ordered and shuffled data.
Understanding Dynamic Compute Allocation in Recurrent Transformers
Ibraheem Muhammad Moosa (Pennsylvania State University), Wenpeng Yin (Pennsylvania State University)
Computational EfficiencyRepresentation LearningTransformerLarge Language ModelMixture of ExpertsTextSequentialBenchmark
🎯 What it does: Studied methods for dynamically allocating computational resources per token in autoregressive Transformers, and proposed the ANIRA framework, which supports two decision mechanisms (early depth prediction and online stopping) to achieve variable computational depth at the token level.
Understanding Dynamics of Adam in Zero-Sum Games: An ODE Approach
Yi Feng (Aarhus University), Xiao Wang (Shanghai University of Finance and Economics)
OptimizationReinforcement Learning from Human FeedbackConvolutional Neural NetworkGenerative Adversarial NetworkImageOrdinary Differential Equation
🎯 What it does: Studied the continuous dynamics of Adam-DA in zero-sum games, derived the corresponding ODE, and analyzed local convergence and implicit gradient regularization.
Understanding Generalization and Forgetting in In-Context Continual Learning
Guangyu Li (Mohamed bin Zayed University of Artificial Intelligence), Lijie Hu (Mohamed bin Zayed University of Artificial Intelligence)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextSequential
🎯 What it does: The study investigates continuous learning through the attention mechanism in language models during inference without updating parameters.
Understanding Generalization from Embedding Dimension and Distributional Convergence
Junjie Yu (Southern University of Science and Technology), Quanying Liu (Southern University of Science and Technology)
Explainability and InterpretabilityRepresentation LearningContrastive LearningImageText
🎯 What it does: Proposed and proved an upper bound on the posterior generalization error for trained models, linking the intrinsic dimension of the embedding space with the Wasserstein convergence rate and the Lipschitz sensitivity of downstream mappings, thereby explaining the impact of embedding geometry on generalization.
Understanding LoRA as Knowledge Memory: An Empirical Analysis
Seungju Back (KAIST), Sungjin Ahn (KAIST)
Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Systematically evaluate the scalability, efficiency, and capability of LoRA for multi-module combination as a knowledge memory, and compare it with traditional methods such as ICL and RAG.
Understanding MARS: When Scaling Momentum Provably Helps
Egor Shulgin (King Abdullah University of Science and Technology), Peter Richtárik (King Abdullah University of Science and Technology)
OptimizationTransformerLarge Language ModelText
🎯 What it does: Studied the theoretical advantages of the MARS optimizer in training large-scale language models, proposed the γ-similarity metric, and proved its convergence in non-convex optimization.
Understanding Multimodal Learning: A Loss Landscape Smoothness Perspective
Jae-Jun Lee (Ulsan National Institute of Science and Technology), Sung Whan Yoon (Ulsan National Institute of Science and Technology)
ClassificationRetrievalRepresentation LearningDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageVideoTextMultimodalityAudio
🎯 What it does: This paper proposes that multi-modal learning can make the loss landscape flatter through the convolutional smoothing effect, thereby improving robustness and generalization, as verified by theoretical analysis and experiments; meanwhile, it designs a distributed multi-modal training (DML) method based on random pairing to further enhance this smoothing effect.
Understanding Performance Collapse in Layer-Pruned Large Language Models via Decision Representation Transitions
boyu shi, Xin Geng (Southeast University)
Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelText
🎯 What it does: Investigate the mechanism by which layer pruning causes performance degradation in LLMs, proposing an analysis from the perspective of decision representation;
Understanding Private Learning From Feature Perspective
Meng Ding (University at Buffalo), Jinhui Xu (University of Science and Technology of China)
Safty and PrivacyRepresentation LearningConvolutional Neural NetworkDiffusion modelScore-based ModelNeural Radiance FieldContrastive LearningGaussian SplattingImage
🎯 What it does: Analyze the feature learning and noise memory process in differentially private learning through a theoretical framework, and propose a two-layer CNN feature decomposition method under noisy gradient descent, providing a quantitative relationship between SNR and privacy budget on learning performance.
Understanding SAM through Minimax Perspective
Ying Chen (University of California, Berkeley), Javad Lavaei (University of California, Berkeley)
OptimizationExplainability and InterpretabilityComputational EfficiencyImageTabularStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Reinterpret SAM from a min-max perspective, derive a two-scale ODE, and propose the MSAM algorithm with multi-step inner loops.
Understanding Self-Supervised Learning via Latent Distribution Matching
Fabian A Mikulasch (Friedrich Miescher Institute for Biomedical Research), Friedemann Zenke (Friedrich Miescher Institute for Biomedical Research)
Representation LearningAuto EncoderContrastive LearningImageVideoBiomedical Data
🎯 What it does: Proposed and validated the potential distribution matching (LDM) framework for self-supervised learning, unifying contrastive, non-contrastive, and predictive methods, and deriving a new model based on Kalman filtering.
Understanding the Ability of LLMs to Handle Character-Level Perturbation
Anyuan Zhuo (Shanghai University of Finance and Economics), Pinyan Lu (Shanghai University of Finance and Economics)
Explainability and InterpretabilityComputational EfficiencyAdversarial AttackTransformerLarge Language ModelPrompt EngineeringTextChain-of-Thought
🎯 What it does: This study investigates the robustness of large models under character-level perturbations, designs and evaluates three perturbation methods (Typo, Shuffle, UCC-Inj), and analyzes their impact on model performance and internal mechanisms.
Understanding the Gaps in Satisficing Bandits
Chloé Rouyer (Universität Potsdam), Peter Auer (Montanuniversität Leoben)
Reinforcement Learning
🎯 What it does: Provides new lower bound analysis and a novel Uncertain-UCB algorithm for threshold-satisficing multi-armed bandits.
Understanding the Parameter Space Geometry of Transformers Encoding Boolean Functions
Blanka Kövér, Ryan Cotterell (ETH Zürich)
Explainability and InterpretabilityRepresentation LearningTransformerText
🎯 What it does: This paper studies the geometric structure of the parameter space of Transformers when encoding Boolean functions, and reveals their preference for low-sensitivity functions.
Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO
Ruizhe Shi (University of Washington), Simon Shaolei Du
Reinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningText
🎯 What it does: This paper systematically analyzes the performance gap between two-stage RLHF and direct preference optimization (DPO) under different model misspecification scenarios from both theoretical and experimental perspectives.
Understanding Transfer Learning of RNA Foundation Models on Downstream Tasks
Yuan Li (University of Exeter), Ke Li (University of Exeter)
Domain AdaptationComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerSupervised Fine-TuningPrompt EngineeringContrastive LearningBiomedical DataBenchmark
🎯 What it does: A systematic evaluation and analysis of the transfer learning mechanism of RNA foundation models in downstream tasks related to structure and function.
Understanding Truncated Positional Encodings for Graph Neural Networks
James Flora (Oregon State University), Amir Nayyeri (Oregon State University)
Computational EfficiencyRepresentation LearningGraph Neural NetworkTransformerContrastive LearningGraphTabularBenchmark
🎯 What it does: This paper studies the expressive power of truncated positional encodings (spectral, walk, k-harmonic) in graph neural networks and proves that these encodings are no longer equivalent in expressiveness after truncation;
Unfolded Laplacian Spectral Embedding: A Theoretically Grounded Approach to Dynamic Network Representation
Haruka Ezoe (University of Tokyo), Ryohei Hisano (University of Tokyo)
Representation LearningGraph Neural NetworkContrastive LearningGraphTime Series
🎯 What it does: Proposed a dynamic network embedding method called ULSE based on normalized Laplacian, and provided theoretical stability proof and experimental verification.
Unfolding Generative Flows with Koopman Operators: Trajectory-Preserving Linearization
Erkan Turan (Ecole Polytechnique), Maks Ovsjanikov (Ecole Polytechnique)
GenerationExplainability and InterpretabilityComputational EfficiencyDiffusion modelFlow-based ModelImageStochastic Differential Equation
🎯 What it does: A global linearization framework based on the Koopman operator is constructed, performing a first-order linear approximation on the pre-trained Conditional Flow Matching (CFM) generative model, enabling its generation process to be analytically solved and sampled in one go.
UnHype: CLIP-Guided Hypernetworks for Dynamic LoRA Unlearning
Piotr Wójcik (University of Cologne), Maciej Zieba (Tooploox)
GenerationSafty and PrivacyTransformerDiffusion modelContrastive LearningImageText
🎯 What it does: Propose a CLIP-guided hypernetwork (UnHype) that dynamically generates LoRA weights to achieve adaptive forgetting for both single and multiple concepts, supporting zero-shot and context-aware concept erasure;
Uni-DocRobust: Universal Plug-and-Play Robustness Enhancement for Multi-modal LLMs via Feature Restoration
Yuxuan Zhou (Wangxuan Institute of Computer Technology Peking University), Zhi Tang
RestorationData SynthesisComputational EfficiencyKnowledge DistillationRepresentation LearningAdversarial AttackTransformerLarge Language ModelPrompt EngineeringVision Language ModelDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageTextMultimodalityBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose the DocRobust-VQA large-scale clean/deggraded document pair + QA dataset, and design the Uni-DocRobust framework, achieving feature-level recovery core and lightweight adapter, enhancing the robustness of multimodal large language models on low-quality documents.
UniCode: Augmenting Evaluation for Code Reasoning
Xinyue Zheng (Peking University), Yitao Liang (Peking University)
OptimizationData-Centric LearningAI Code AssistantTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose the UniCode framework, which uses a generative approach to perform multi-dimensional enhanced evaluation of code reasoning and builds an expandable test generation pipeline;
UniDrag: Unified Multi-Field Prediction and Robust Shape Optimization for Vehicle Aerodynamics
Ye Liu (Eastern Institute of Technology), Yuntian Chen (Eastern Institute of Technology)
Autonomous DrivingOptimizationTransformerMixture of ExpertsDiffusion modelAuto EncoderContrastive LearningMeshTabularPhysics Related
🎯 What it does: Propose UniDrag, a unified multi-field prediction and robust shape optimization framework for vehicle aerodynamic design.
UniFast-HGR: Scalable and Efficient Maximal Correlation for Multimodal Models
Hongkang Zhang (Tsinghua University), Ercan Engin KURUOGLU (Tsinghua University)
ClassificationRecognitionSegmentationRetrievalComputational EfficiencyRepresentation LearningAuto EncoderContrastive LearningImageVideoTextMultimodality
🎯 What it does: Propose UniFast-HGR, a scalable and efficient method for approximating maximum correlation, replacing traditional covariance whitening. It enhances dependency learning in multi-modal models by utilizing centralized and ℓ2-normalized cosine alignment and Gram space regularization.
Unified Multimodal Autoregressive Modeling with Shared Context—Visual Tokenizer is Key to Unification
Wujian Peng (Fudan University), Shuai Bai (Qwen Team, Alibaba Group)
Image TranslationRestorationSegmentationGenerationCompressionRepresentation LearningReinforcement Learning from Human FeedbackTransformerSupervised Fine-TuningReinforcement LearningVision Language ModelDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageTextMultimodality
🎯 What it does: Proposed a unified autoregressive framework, UNIAR, which uses a single discrete visual tokenizer to simultaneously support visual understanding, generation, and editing;
Unified Multimodal Visual Tracking with Dual Mixture-of-Experts
Lingyi Hong (Fudan University), Wenqiang Zhang (Fudan University)
Object TrackingTransformerMixture of ExpertsContrastive LearningImageVideoMultimodality
🎯 What it does: Propose a unified multi-modal visual tracking framework, OneTrackerV2, which supports end-to-end training and inference for RGB and RGB+X (Depth, Thermal, Event, Language) modalities, avoiding the need for separate multi-task training and multi-stage fine-tuning.
Unified Safe In-context Image Generation in Multimodal Diffusion Transformers via Restricting Unsafe Information Flows
Xiang Yang (Fudan University), Min Yang (Fudan University)
GenerationSafty and PrivacyTransformerPrompt EngineeringDiffusion modelImageTextMultimodality
🎯 What it does: Designed an untrained safe generation framework UVR to suppress the generation of unsafe content in multimodal diffusion transformers.
Unified Time Series Explanations via Amortized Optimization and Instance-level Multi-Expert Knowledge Distillation
Viet-Hung Tran (Queen's University Belfast), Son Thai Mai (Queen's University Belfast)
ClassificationExplainability and InterpretabilityKnowledge DistillationConvolutional Neural NetworkRecurrent Neural NetworkTransformerMixture of ExpertsTime SeriesElectrocardiogram
🎯 What it does: Propose the XMA framework, achieving high-quality interpretability in temporal classification models through instance-level multi-expert knowledge distillation and target-regularized holistic optimization
UniFLoW: Universal Multi-Modal Federated LoRA Fine-Tuning Framework with Analytical Aggregation
Haoyuan Liang (Sun Yat-sen University), Juepeng Zheng (Sun Yat-sen University)
Federated LearningComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningAgentic AIContrastive LearningImageTextMultimodalityAudio
🎯 What it does: Proposes the UniFLoW framework to achieve multi-modal federated LoRA fine-tuning, supporting multiple modalities and addressing issues of heterogeneity and communication costs.
Unifying Adversarial Robustness and Training Across Text Scoring Models
Manveer Singh Tamber (University of Waterloo), Jimmy Lin (University of Waterloo)
RetrievalAdversarial AttackReinforcement Learning from Human FeedbackTransformerReinforcement LearningPrompt EngineeringScore-based ModelGenerative Adversarial NetworkContrastive LearningText
🎯 What it does: Unify the study of adversarial robustness and training of text scoring models (retriever, re-ranker, reward model), propose multiple adversarial training methods, and evaluate their robustness across attacks and models.
Unifying and Optimizing Data Values for Selection via Sequential Decision-Making
Hongliang Chi (Rensselaer Polytechnic Institute), Yao Ma (Rensselaer Polytechnic Institute)
OptimizationExplainability and InterpretabilityData-Centric LearningTextTabularTime SeriesSequential
🎯 What it does: Propose to treat data selection as a sequential decision-making problem, using dynamic programming to obtain the optimal selection sequence, and generalize existing data value methods as greedy strategies in approximate dynamic programming; propose a bilateral graph coverage approximation that preserves submodularity, enabling scalable greedy selection;
Unifying Dataset Pruning and Distillation for Efficient Large-scale Compression
Lingao Xiao (Agency for Science, Technology and Research), Xinchao Wang (National University of Singapore)
ClassificationObject DetectionCompressionKnowledge DistillationData-Centric LearningConvolutional Neural NetworkContrastive LearningImage
🎯 What it does: This study explores the relationship between dataset pruning (DP) and data distillation (DD) when compressing large-scale datasets, proposes a unified evaluation benchmark, and designs a Prune, Combine, Augment (PCA) framework that uses only hard labels;
Unifying Deep Stochastic Processes for Image Enhancement
Wojciech Maciej Kozłowski (Wrocław University of Science and Technology), Maciej Zieba (Wrocław University of Science and Technology)
RestorationSuper ResolutionDiffusion modelScore-based ModelImageStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Unify multiple image enhancement methods into a continuous-time stochastic differential equation (SDE) framework, decoupling the process definition from implementation details such as scheduling, sampling, and discretization.
Unifying Heterogeneous Degradations: Uncertainty-Aware Diffusion Bridge Model for All-in-One Image Restoration
Luwei Tu (Sun Yat-sen University), Zhi Jin (Sun Yat-sen University)
RestorationTransformerDiffusion modelScore-based ModelOptical FlowImageStochastic Differential Equation
🎯 What it does: Proposes an uncertainty-guided diffusion bridge model (UDBM) that solves the all-in-one image restoration (AiOIR) problem in a single-step reasoning manner.
Unifying Heterogeneous Multi-Modal Remote Sensing Detection Via Language-Pivoted Pretraining
Yuxuan Li (Nankai University), jian Yang
Object DetectionTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodality
🎯 What it does: Proposed the BabelRS framework, achieving unified heterogeneous multi-modal remote sensing object detection through language-centric pre-training, decoupling modal alignment from task learning;
Unifying Low Dimensional Spectra in Deep Learning
Connall Garrod (University of Oxford), Jonathan P. Keating (University of Oxford)
ClassificationRepresentation LearningConvolutional Neural NetworkContrastive LearningImage
🎯 What it does: This paper analyzes the low-dimensional spectral structures of the Hessian, gradient, and weight matrix in over-parameterized classification networks through the deep unconstrained feature model (UFM) and the theory of deep neural collapse (DNC), proving that DNC is the root cause of these spectral phenomena;
Unifying Masked Diffusion Models with Various Generation Orders and Beyond
Chunsan Hong (KAIST), Jong Chul Ye (KAIST)
GenerationData SynthesisTransformerLarge Language ModelDiffusion modelScore-based ModelAuto EncoderText
🎯 What it does: Proposed a unified mask diffusion model (OeMDM) capable of handling multiple generation orders, and on this basis designed a learnable order mask diffusion model (LoMDM) that can learn the generation order and diffusion network from scratch.