arXivSub Start free trial

ICML 2026 Papers — Page 6

International Conference on Machine Learning · 6554 papers

Automatic Unsupervised Ensemble Outlier Model Selection

Hong-Phuc Phan (FPT University), Christian S. Jensen (Aalborg University)

Anomaly DetectionMeta LearningMixture of ExpertsContrastive LearningImageTextTabular

🎯 What it does: Propose MetaEns, a framework for building an adaptive anomaly detection model ensemble without labels by leveraging meta-learning to predict gains and constructing the ensemble through redundancy discounting and family risk regularization.

Automatically Finding Reward Model Biases

Zifan Wang (MATS), Arthur Conmy

Reinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextMultimodality

🎯 What it does: This paper proposes a black-box auditing framework based on evolutionary algorithms and LLM auto-generation with iterative refinement, used to automatically discover natural language biases in reward models.

AutoMoT: A Unified Vision-Language-Action Model with Asynchronous Mixture -of-Transformers for End-to-End Autonomous Driving

Wenhui Huang (Nanyang Technological University), Chen Lv (Nanyang Technological University)

Autonomous DrivingTransformerMixture of ExpertsVision Language ModelVision-Language-Action ModelDiffusion modelImageVideoTextMultimodality

🎯 What it does: Propose AutoMoT, an end-to-end autonomous driving framework that unifies the three tasks of vision-language-action (VLA) and achieves asynchronous inference;

AutoMS: Multi-Agent Evolutionary Search for Cross-Physics Inverse Microstructure Design

Zhenyuan Zhao (Shandong University), Lin Lu (Shandong University)

OptimizationTransformerLarge Language ModelReinforcement LearningAgentic AIGraphTabularPhysics Related

🎯 What it does: Developed the AutoMS multi-agent evolutionary search framework, which utilizes LLM and simulation feedback closed-loop iteration to achieve cross-physical field inverse microstructure design.

AutoNumerics-Zero: Automated Discovery of State-of-the-Art Mathematical Functions

Esteban Real (Google DeepMind), David H. Park (Google)

OptimizationComputational EfficiencyAI Code AssistantTabularPhysics Related

🎯 What it does: This paper uses a zero-knowledge evolutionary symbolic regression method to automatically discover high-precision, low-computation programs for computing transcendental functions such as exponential, trigonometric, Bessel, Airy, and error functions.

AutoQRA: Joint Optimization of Mixed-Precision Quantization and Low-rank Adapters for Efficient LLM Fine-Tuning

Changhai Zhou (Fudan University), WEIZHONG ZHANG

OptimizationComputational EfficiencyHyperparameter SearchTransformerLarge Language ModelSupervised Fine-TuningText

🎯 What it does: Propose the AutoQRA framework, which jointly performs discrete search for quantization bit-width and LoRA rank for each layer of the LLM, to achieve efficient, low-memory fine-tuning.

AutoRAS: Learning Robust Agentic Systems with Primitive Representations

Yang Yue (Beijing University of Posts and Telecommunications), Jingfeng Zhang (Fudan University)

OptimizationRepresentation LearningAdversarial AttackRobotic IntelligenceReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringTextBenchmark

🎯 What it does: AutoRAS achieves the construction of automated robust agent systems by transforming the design of agent systems into the generation of sequences of symbolic primitives and continuously optimizing them within an execution feedback loop.

Autoregression with Self-Token Prediction

Dengsheng Chen (Key Laboratory of System Software (Chinese Academy of Sciences)), Enhua Wu (Key Laboratory of System Software (Chinese Academy of Sciences))

GenerationComputational EfficiencyTransformerDiffusion modelScore-based ModelFlow-based ModelAuto EncoderGenerative Adversarial NetworkImageVideo

🎯 What it does: Propose a self-token prediction framework, enabling the prediction of multiple tokens in a single step, and build SAR (Spatially Autoregressive Image Generator) based on this, significantly accelerating inference speed while maintaining strict spatial causality;

Autoregressive Boltzmann Generators

Danyal Rehman (Mila Quebec AI Institute), Alexander Tong (Aithyra)

GenerationDrug DiscoveryProtein Structure PredictionTransformerMixture of ExpertsDiffusion modelScore-based ModelFlow-based ModelBiomedical Data

🎯 What it does: Proposed a Boltzmann generator based on autoregressive models (ARBG), which constructs molecular conformation distributions by conditioning dimension by dimension, breaking the reversibility constraints of traditional flow models and improving sampling efficiency and expressive power.

Autoregressive Direct Preference Optimization

Masanari Oi (Institute of Science Tokyo), Nakamasa Inoue (Institute of Science Tokyo)

OptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningText

🎯 What it does: Proposes Autoregressive Direct Preference Optimization (ADPO), explicitly introducing the autoregressive assumption into the DPO framework and constructing a prefix-closed Bradley-Terry model;

Autoregressive Image Generation with Masked Bit Modeling

Qihang Yu (Amazon FAR), Xi Chen (Amazon FAR)

GenerationTransformerMixture of ExpertsDiffusion modelAuto EncoderContrastive LearningImage

🎯 What it does: Studied the gap between discrete and continuous image generation, and proposed the Masked Bit Autoregressive Modeling (BAR) framework, which replaces traditional large vocabulary classification with a bit-by-bit prediction approach, enabling efficient generation of discrete tokenizers under sufficient bit budgets;

Autoregressive Language Models are Secretly Energy-Based Models: Insights into the Lookahead Capabilities of Next-Token Prediction

Mathieu Blondel (Google DeepMind), Vincent Roulet (Google DeepMind)

Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationReinforcement Learning from Human FeedbackTransformerLarge Language ModelScore-based ModelContrastive LearningTextChain-of-Thought

🎯 What it does: This paper proves the bijective relationship between autoregressive language models (ARM) and energy-based models (EBM) in function space through the chain rule, and uses this mapping to clarify the equivalence of the two types of models in learning and inference, further deriving the KL upper bound and training error bounds from EBM to ARM.

Autoregressive, Yet Revisable: In Decoding Revision for Secure Code Generation

Chengran Yang (Singapore Management University), David Lo (Singapore Management University)

Safty and PrivacyAI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: Proposes the 'Stream of Revision' framework, allowing large language models to dynamically trigger, locate, and fix vulnerabilities in code during a single autoregressive decoding process, thereby achieving immediate self-correction.

AutoRPA: Efficient GUI Automation through LLM-Driven Code Synthesis from Interactions

Minghao Chen (Hangzhou Dianzi University), Yufei Yin (Hangzhou Dianzi University)

Computational EfficiencyAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: Automatically convert the interaction trajectory of LLM agents based on ReAct into reusable RPA scripts, thereby achieving efficient GUI automation.

AutoSizer: Automatic Sizing of Analog and Mixed-Signal Circuits via Large Language Model (LLM) Agents

Xi Yu (Brookhaven National Laboratory), Yihui Ren (Brookhaven National Laboratory)

OptimizationTransformerLarge Language ModelAgentic AIPrompt EngineeringMixture of ExpertsTextTabularBenchmark

🎯 What it does: Proposed a dual-loop self-adaptive optimization framework called AUTOSIZER based on large language models, for automatically optimizing transistor sizing in analog/mixed-signal circuits.

AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning

Jiaru Zou (Princeton University), Mengdi Wang (Princeton University)

OptimizationExplainability and InterpretabilityComputational EfficiencyData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AIPrompt EngineeringTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose the AutoTool framework, enabling large language models to dynamically select and integrate tools during the reasoning process, and supporting the evolution of the toolset over time.

AutoVSR: Automatic Visual-to-Symbolic Reasoning for Symbolic Expression Generation from Circuit Schematic

Zhe Xiao (Hunan University), Mingyu Liu (Huazhong University of Science and Technology)

Image TranslationGenerationOptimizationComputational EfficiencyAI Code AssistantTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextGraphBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose the AutoVSR framework, which can automatically convert circuit schematic images into executable intermediate representations (Executable IR). Subsequently, it utilizes a visual language model and a symbolic tool library to perform multi-step symbolic reasoning, generating accurate circuit symbolic expressions.

AutoWebWorld: Synthesizing Infinite Verifiable Web Environments via Finite State Machines

Yifan Wu (Hong Kong University of Science and Technology), Yuyu Luo (Hong Kong University of Science and Technology)

Data SynthesisReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringGenerative Adversarial NetworkWorld ModelTextSequential

🎯 What it does: Propose the AutoWebWorld framework, which generates verifiable synthetic web environments using finite state machines (FSM) and enumerates and verifies interaction trajectories through BFS search.

AvAtar: Learning to Align via Active Optimal Transport

Qi Yu (University of Illinois Urbana Champaign), Hanghang Tong (University of Illinois Urbana Champaign)

RetrievalOptimizationRepresentation LearningData-Centric LearningContrastive LearningImageTextMultimodalityGraph

🎯 What it does: Propose an active alignment framework called AVATAR based on optimal transport, which evaluates the impact of candidate nodes on the global alignment results through gradient propagation, thus actively selecting the most informative samples for querying.

AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation

Ziwei Zhou (Fudan University), Chong Luo (Microsoft Research Asia)

GenerationData SynthesisTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsDiffusion modelVideoTextMultimodalityBenchmarkAudio

🎯 What it does: Propose AVGen-Bench, a task-driven, multi-granularity evaluation benchmark for text-to-audio-visual generation.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs

Yaoting Wang (Fudan University), Yunxin Liu (Tsinghua University)

TransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningVideoTextMultimodalityBenchmarkAudio

🎯 What it does: Proposes the AVI-Bench cross-modal cognition benchmark for Omni-MLLM, and extends it to AVI-Bench-PriSe, to evaluate the capabilities in four stages: perception, understanding, reasoning, and low-semantic primitive perception.

Avoid What You Know: Divergent Trajectory Balance for GFlowNets

Pedro Dall'Antonia (Getulio Vargas Foundation), Diego Mesquita (MBZUAI)

OptimizationReinforcement LearningFlow-based ModelTextTabularSequentialBenchmark

🎯 What it does: Proposed a new algorithm called Adaptive Complementary Exploration (ACE), which is used to effectively explore under-explored high-reward regions when training generative flow networks (GFlowNets).

AVTrack: Audio-Visual Tracking in Human-centric Complex Scenes

Yaoting Wang (Fudan University), Henghui Ding (Fudan University)

Object TrackingSegmentationTransformerPrompt EngineeringVision Language ModelContrastive LearningOptical FlowVideoMultimodalityBenchmarkRetrieval-Augmented GenerationAudio

🎯 What it does: A specialized audio-visual instance segmentation and tracking benchmark (AVTrack) for human-centric complex scenes was constructed, and a scalable three-stage baseline method, AVTracker, was proposed on this benchmark.

Awakening Visual Reasoning: Mitigating Post-Training Failure in Vision-Text Compression

Xing Xi (South China University of Technology), Ronghua Luo (South China University of Technology)

CompressionKnowledge DistillationRepresentation LearningReinforcement Learning from Human FeedbackTransformerSupervised Fine-TuningReinforcement LearningVision Language ModelContrastive LearningImageTextMultimodalityTabularRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposes the CoRe (Coordinated Reasoning) framework to address the issue of performance degradation caused by visual prompts in the Visual Text Compression (VTC) scenario, achieving a significant improvement in visual reasoning capabilities.

Axiomatic Atlas: A Prescriptive Framework for Neural Architecture Design

Minghao Guo (MIT), Wojciech Matusik (MIT)

Neural Architecture SearchGraph Neural NetworkTransformerMixture of ExpertsScore-based ModelFlow-based ModelAuto EncoderContrastive LearningTextGraph

🎯 What it does: Propose Axiomatic Atlas, a framework that presets, diagnoses, and repairs neural network architectures through composable axioms.

B-Spar: Bayesian Sparse-Reward Modeling for RL-based Image Editing

Shusong Xu (vivo Mobile Communication Co Ltd), Bo Li (vivo Mobile Communication Co Ltd)

Image HarmonizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringDiffusion modelImageMultimodality

🎯 What it does: Proposed a RL framework called B-Spar based on Bayesian sparse reward modeling, used to train image editing agents to achieve perceptually consistent image refinement under sparse human feedback.

BabyVision: Visual Reasoning Beyond Language

Liang Chen (UniPat AI), Kuan Li (UniPat AI)

TransformerLarge Language ModelReinforcement LearningPrompt EngineeringVision Language ModelDiffusion modelImageMultimodalityBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose the BabyVision benchmark to evaluate the early visual reasoning capabilities of multimodal large language models without linguistic intervention, and extend it to BabyVision-GEN for generative evaluation;

Backjump-on-Graph: Empowering Large Language Models with Reinforced Retrospective Exploration for Agentic Knowledge Graph Reasoning

Yunqi Zhang (Zhongguancun Laboratory), Yubo Chen (Zhongguancun Laboratory)

OptimizationExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AITextGraphBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This paper proposes an agent framework based on large language models (LLMs) to address the dead-end problem in knowledge graph question answering caused by mismatches between queries and graph structures. It introduces a 'Backjump' mechanism, allowing LLMs to retreat to historical nodes and re-explore alternative paths during reasoning.

Backward Oversmoothing: why is it hard to train deep Graph Neural Networks?

Nicolas Keriven (French National Center for Scientific Research)

OptimizationGraph Neural NetworkContrastive LearningGraph

🎯 What it does: This paper theoretically analyzes the backward oversmoothing in graph neural networks (GNNs) from an optimization perspective, and reveals the fundamental mechanism that leads to a large number of easily reachable spurious stationary points in deep GNNs.

Backward SDE–Based Diffusion for Physics-Constrained Generation

Zihao WANG (University of Tennessee)

RestorationGenerationDiffusion modelScore-based ModelImageTime SeriesBiomedical DataComputed TomographyStochastic Differential Equation

🎯 What it does: Based on a pre-trained score-based SDE prior, this paper proposes to achieve terminal condition-based physical constraint generation and inverse problem solving through associated BSDE, resulting in an unconditional prior, terminal-consistent, and trainable inverse mapping.

Bad Seeing or Bad Thinking? Rewarding Perception for Multimodal Reasoning

Haozhe Wang (Hong Kong University of Science and Technology), Fangzhen Lin (Hong Kong University of Science and Technology)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBenchmark

🎯 What it does: Propose the MoCA framework, which splits the VLM generation process into two stages: perception and reasoning, and uses reinforcement learning to provide rewards for perception, thereby improving multimodal reasoning performance.

Baguan-TS: dual in-context learning model for time series forecasting with covariates

Linxiao Yang (DAMO Academy Alibaba Group), Liang Sun (DAMO Academy Alibaba Group)

Recommendation SystemAnomaly DetectionAutonomous DrivingOptimizationFederated LearningComputational EfficiencyData-Centric LearningMeta LearningTransformerPrompt EngineeringMixture of ExpertsContrastive LearningTabularTime SeriesBenchmarkFinance RelatedPhysics RelatedRetrieval-Augmented Generation

🎯 What it does: Propose Baguan-TS, a scenario learning model based on 3D Transformer, which can achieve multi-variable time series prediction without manual features, and enhance robustness through local calibration and context overfitting within the original sequence.

Balanced LoRA: Removing Parameter Invariance to Accelerate Convergence

Valérie Castin (École Normale Supérieure PSL), Gabriel Peyré (École Normale Supérieure PSL)

OptimizationComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsAuto EncoderContrastive LearningTextBenchmark

🎯 What it does: This paper studies the problem of different optimal point condition numbers caused by LoRA over-parameterization, and proposes Balanced LoRA (BaLoRA) to accelerate convergence by projecting to the balanced manifold after each iteration.

Balancing Fidelity and Diversity in Diffusion Models via Symmetric Attention Decomposition: Hopfield Perspective

Hyunmin Cho (Korea University), Kyong Hwan Jin (Korea University)

GenerationTransformerDiffusion modelImageText

🎯 What it does: The authors view the self-attention matrix QKᵀ in diffusion models as an associative memory matrix, decomposing it into symmetric energy and skew-symmetric cyclic parts, and propose an adjustable cyclic control mechanism to balance the fidelity and diversity of generation.

Balancing Learning Rates Across Layers: Exact Two-Step Dynamics and Optimal Scaling in Linear Neural Networks

Tianyu Pang (Dartmouth College), Yaoqing Yang (Dartmouth College)

OptimizationRepresentation LearningHyperparameter SearchAuto EncoderContrastive LearningStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Analyze the impact of learning rate allocation on early training dynamics and generalization in one- to two-step gradient descent for two-layer and three-layer linear neural networks, deriving closed-form expressions for gradients and test errors.

Balancing Plasticity and Stability with Fast and Slow Successor Features

Raymond Chua (McGill University), Blake Aaron Richards (McGill University)

TransformerReinforcement LearningTime SeriesBenchmarkStochastic Differential Equation

🎯 What it does: Constructed a natural continuous non-stationary reinforcement learning benchmark, and studied the trade-off between stability and plasticity, proposing to combine multi-timescale synaptic consolidation (SC) with successor features (SF) to enhance continuous adaptation performance.

Balancing Understanding and Generation in Discrete Diffusion Models

Yue Liu (University Of Chinese Academy Of Sciences), Yunfan Liu (University Of Chinese Academy Of Sciences)

GenerationData SynthesisComputational EfficiencyRepresentation LearningTransformerDiffusion modelScore-based ModelImageText

🎯 What it does: Propose a hybrid noise diffusion model XDLM, which combines the advantages of Mask and Uniform noise;

BALLAST: Bayesian Active Learning with Look-ahead Amendment for Sea-drifter Trajectories under Spatio-Temporal Vector Fields

Rui-Yang Zhang (Lancaster University), Henry Moss (Lancaster University)

Autonomous DrivingOptimizationComputational EfficiencyReinforcement Learning from Human FeedbackDiffusion modelContrastive LearningGaussian SplattingTabularTime SeriesSequentialPhysics RelatedStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: This paper proposes a Bayesian active learning framework named BALLAST, which deploys observation points along the trajectories of Lagrangian drifters in ocean current fields, thereby more efficiently inferring spatiotemporal vector fields.

Bandit Social Leaning Dynamics with Exploration Episodes

Kiarash Banihashem (University of Maryland), Aleksandrs Slivkins (Microsoft Research)

Reinforcement Learning

🎯 What it does: The study investigates the scenario where an agent controls a multi-armed bandit with limited rounds in a social learning context, exploring the issue of learning failure even when the agent itself performs exploration;

BandPO: Bridging Trust Regions and Ratio Clipping via Probability-Aware Bounds for LLM Reinforcement Learning

Yuan Li (Fudan University), Xipeng Qiu (Fudan University)

Reinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningText

🎯 What it does: Propose the BandPO framework, which maps the f-divergence trust region to a probability-aware dynamic clipping interval, addressing the issue in LLM RL where fixed clipping suppresses low-probability high-advantage actions.

BARRED: Synthetic Training of Custom Policy Guardrails via Asymmetric Debate

Arnon Mazza (Plurai Inc), Elad Levi (Plurai Inc)

ClassificationSafty and PrivacyExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelAgentic AIPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: By leveraging task descriptions and a few unannotated samples, high-quality synthetic training data is automatically generated using dimension decomposition and multi-agent debate to train custom guardian models.

Barriers to Counterfactual Credit Attribution for Autoregressive Models

Aloni Cohen (University of Chicago), Chenhao Zhang (Northwestern University)

Information TheoryExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringText

🎯 What it does: This paper investigates the challenges of achieving counterfactual credit assignment (CCA) in autoregressive language models and presents two main theoretical obstacles;

BAS: Bridging Adam and SignSGD for Memory-Efficient LLM Training

Yijie Zhou (Chinese University of Hong Kong Shenzhen), Shi Pu (Chinese University of Hong Kong Shenzhen)

OptimizationComputational EfficiencyTransformerLarge Language ModelSupervised Fine-TuningText

🎯 What it does: Propose Block Adaptive Signum (BAS) as a low-memory optimizer, combining the adaptiveness of Adam with the efficiency of SignSGD, using block-level scaling to update parameters.

Base Models Know How to Reason, Thinking Models Learn When

Constantin Venhoff (University of Oxford), Neel Nanda

Explainability and InterpretabilityKnowledge DistillationTransformerSupervised Fine-TuningReinforcement LearningAuto EncoderTextChain-of-Thought

🎯 What it does: This paper proposes an unsupervised method that uses sparse autoencoders to extract the reasoning mechanisms of thinking language models, and explains the content learned under different training paradigms (RL vs. SFT-distillation) by constructing model differences (category vectors + heuristic of when they are activated).

BASIL: Scalable Bayesian Semi-supervised Clustering with Feature Selection and Adaptive Constraint Weighting

Luwei Wang (University of Edinburgh), Sohan Seth (University of Edinburgh)

Explainability and InterpretabilityComputational EfficiencyData-Centric LearningImageTabularBiomedical DataElectronic Health Records

🎯 What it does: Proposes an scalable Bayesian semi-supervised clustering framework called BASIL, which can jointly learn clustering assignments, feature importance, and adaptive constraint weights.

BAT: Better Audio Transformer Guided by Convex Gated Probing

Houtan Ghaffari (Ghent University), Paul Devos (Ghent University)

RecognitionTransformerSupervised Fine-TuningAuto EncoderContrastive LearningAudio

🎯 What it does: Propose Convex Gated Probing (CGP) as an efficient frozen feature evaluation method, and improve the audio self-supervised learning model based on CGP, introducing Better Audio Transformer (BAT), which achieves new state-of-the-art (SOTA) results on multiple audio and speech benchmarks.

Batch Normalization for Neural Networks on Complex Domains

Xuan Son Nguyen (ETIS, UMR 8051, CY Cergy Paris University, ENSEA, CNRS), Nistor Grozavu (ETIS, UMR 8051, CY Cergy Paris University, ENSEA, CNRS)

ClassificationRecurrent Neural NetworkGraph Neural NetworkAuto EncoderContrastive LearningGraphTabularTime SeriesSequentialBenchmark

🎯 What it does: This paper proposes a batch normalization (BN) layer tailored for complex domains (such as Siegel disks and the complex unit ball), utilizing automorphisms, Fréchet mean, and Kobayashi pseudo-distance to achieve geometric centering and bias, and provides a closed-form formula that can be directly used in deep networks;

Batched Contextual Reinforcement

Bangji Yang (University of Illinois at Urbana-Champaign), Ge Liu (Tsinghua University)

Computational EfficiencyTransformerLarge Language ModelReinforcement LearningPrompt EngineeringText

🎯 What it does: Propose Batched Contextual Reinforcement (BCR), a single-stage training paradigm that stimulates efficient reasoning by enabling LLMs to simultaneously solve multiple problems within a shared context window.

Batched First-Order Methods for Parallel LP Solving in MIP

Nicolas Blin (NVIDIA), Bartolomeo Stellato (Princeton University)

OptimizationTabularBenchmark

🎯 What it does: A batch first-order method for solving multiple linear programming problems in parallel on a GPU is proposed, directly extending the PDHG algorithm and utilizing matrix-matrix multiplication to achieve efficient parallelism, targeting strong branching and bound tightening in mixed integer programming.

Bayes-inspired Integration of Pretrained Priors and Few-Shot Evidence for Few-Shot Classification

Mingyang Zhou (Shenzhen University), Rui Mao (Shenzhen University)

ClassificationMeta LearningTransformerVision Language ModelContrastive LearningImageText

🎯 What it does: Propose a Bayesian-inspired dual-path fusion framework called BOIF, which integrates pre-trained models as priors and few samples as likelihoods for optimal integration in few-shot classification.

Bayesian Gated Non-Negative Contrastive Learning

Peng Cui (Mohamed bin Zayed University of Artificial Intelligence), Lijie Hu (Mohamed bin Zayed University of Artificial Intelligence)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningContrastive LearningImage

🎯 What it does: Propose Bayesian Gated Non-negative Contrastive Learning (BayesNCL), which dynamically suppresses shared background features through a variational Bayesian gating mechanism, addressing optimization conflicts in traditional contrastive learning, thus achieving interpretable sparse representations.

Bayesian Meta-Learning with Expert Feedback for Task-Shift Adaptation through Causal Embeddings

Lotta Mäkinen (Aalto University), Samuel Kaski (Aalto University)

Domain AdaptationRepresentation LearningMeta LearningContrastive LearningTabularBiomedical DataElectronic Health Records

🎯 What it does: A Bayesian meta-learning framework combining causal embeddings with expert feedback to enhance adaptation performance under distribution shift across tasks.

Bayesian model selection and misspecification testing in imaging inverse problems only from noisy and partial measurements

Tom Sprunck (Université Paris-Saclay), Tobías I. Liaudat (Université Paris-Saclay)

RestorationAnomaly DetectionDiffusion modelScore-based ModelImageBiomedical DataMagnetic Resonance Imaging

🎯 What it does: This paper proposes an unsupervised model selection and mismatch detection framework based on data fission and Bayesian cross-validation, which can evaluate and diagnose Bayesian imaging models using only a single noisy measurement.

Bayesian Rain Field Reconstruction using Commercial Microwave Links and Diffusion Model Priors

Badr MOUFAD, Eric Moulines (MBZUAI)

RestorationTransformerDiffusion modelScore-based ModelTabularTime Series

🎯 What it does: This paper proposes using commercial microwave links (CML) for rainfall field reconstruction, employing a Bayesian inverse problem framework combined with a diffusion model prior.

Bayesian Tensor Decomposition with Diffusion Model Prior

Zerui Tao (RIKEN Center for Advanced Intelligence Project), Qibin Zhao (RIKEN Center for Advanced Intelligence Project)

RestorationDiffusion modelImage

🎯 What it does: Propose a Bayesian CP tensor decomposition framework fused with pre-trained diffusion models (DiffBCP), for achieving high-quality restoration on missing or noisy images.

Bayesian-LoRA: Probabilistic Low-Rank Adaptation of Large Language Models

Moule Lin (Trinity College Dublin), Goetz Botterweck (Trinity College Dublin)

Explainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelSupervised Fine-TuningFlow-based ModelText

🎯 What it does: In the parameter-efficient fine-tuning of large-scale language models, the deterministic low-rank updates of LoRA are transformed into a probabilistic low-rank adaptation framework. By utilizing variational inference with sparse Gaussian processes, uncertainty is injected into the low-rank subspace of LoRA, and self-consistent calibrated learning is achieved through regularization and flow transformations.

BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference

Janghyeon Kim (Hanyang University), Jungwook Choi (Hanyang University)

Computational EfficiencyKnowledge DistillationTransformerLarge Language ModelContrastive LearningTextBenchmark

🎯 What it does: Propose BeaconKV, an untrained KV cache compression method that utilizes 'Beacon Query' to capture global queries that revisit old contexts during long inference processes;

BEAR: Dissecting Embodied Abilities in Multimodal Language Models through Skill-level Evaluation and Diagnosis

Yu Qi (Northeastern University), Lawson L.S. Wong (Northeastern University)

Robotic IntelligenceTransformerAgentic AIPrompt EngineeringVision-Language-Action ModelImageVideoTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose the BEAR benchmark, which decomposes and diagnoses the capability bottlenecks of multimodal language models in embodied tasks through skill-level evaluation;

BEAT: Tokenizing and Generating Symbolic Music by Uniform Temporal Steps

Lekai Qian (South China University Of Technology), Ziyu Wang (Mohamed Bin Zayed University Of Artificial Intelligence)

GenerationTransformerLarge Language ModelPrompt EngineeringSequentialAudio

🎯 What it does: Proposed a symbolic music segmentation method called BEAT based on uniform beats, implemented autoregressive generation on Transformer.

BEDTime: A Unified Benchmark for Automatically Describing Time Series

Medhasweta Sen (University of Virginia), Thomas Hartvigsen (University of Virginia)

RecognitionGenerationAnomaly DetectionTransformerLarge Language ModelPrompt EngineeringVision Language ModelTextMultimodalityTime SeriesBenchmarkChain-of-Thought

🎯 What it does: Proposed the BEDTime benchmark to evaluate models on tasks involving identification, discrimination, and generation of structural descriptions for univariate time series, and conducted systematic comparisons using a unified multi-task, cross-modal evaluation framework.

Behavior-Invariant Task Representation Learning with Transformer-based World Models for Offline Meta-Reinforcement Learning

Fuyuan Qian (Southern University of Science and Technology), Quanying Liu (Southern University of Science and Technology)

Meta LearningTransformerReinforcement LearningAuto EncoderContrastive LearningWorld ModelTabularTime SeriesSequential

🎯 What it does: In the offline meta reinforcement learning framework, a random world model based on the Transformer structure is proposed to learn behavior-invariant task representations, and robust transfer from offline to online is achieved by combining context-aware imagination with conservative policy optimization.

Behavioral Mode Discovery for Fine-tuning Multimodal Generative Policies

Alberta Longhini (Stanford University), Seungsu Kim (Naver Labs Europe)

Robotic IntelligenceTransformerReinforcement LearningDiffusion modelMultimodalitySequential

🎯 What it does: Studies how to preserve the multimodality of the action distribution when fine-tuning pre-trained multimodal generation policies in reinforcement learning, and proposes a fine-tuning framework based on pattern discovery.

Being More Lightweight and Practical: Mini-sized Contrastive Learning Pre-trained Models for Fine-grained Traffic Task

Shuhao Li (Fudan University), Fan Zhang (Guangzhou University)

Autonomous DrivingComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerContrastive LearningGraphTime Series

🎯 What it does: Proposes MiniTraffic, a lightweight pre-training framework that pre-trains using road-level data and transfers to lane-level fine-grained traffic prediction tasks;

Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering

Eric Bigelow (Goodfire AI), Ekdeep Singh Lubana

Explainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelPrompt EngineeringText

🎯 What it does: Propose a unified Bayesian belief dynamics framework to explain and predict the effects of in-context learning (ICL) and activation steering on the behavior of large language models.

Belief Propagation Converges to Gaussian Distributions in Sparsely-Connected Factor Graphs

Tom Yates (Imperial College London), Andrew J. Davison (Imperial College London)

OptimizationContrastive LearningGaussian SplattingGraph

🎯 What it does: It is proved that in sparse and low-degree factor graphs, the variable beliefs of Belief Propagation (BP) converge to a Gaussian distribution after multiple iterations, with theoretical guarantees provided.

Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks

Zexue He (Stanford University), Alex Pentland (Stanford University)

TransformerLarge Language ModelTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose the MEMORYARENA benchmark to evaluate the memory and action coupling capabilities of LLM agents in multi-session, interdependent tasks;

Benchmarking and Enhancing VLM for Compressed Image Understanding

Zifu Zhang (Tsinghua University), Yan Wang (Tsinghua University)

CompressionComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerPrompt EngineeringVision Language ModelImageMultimodalityBenchmark

🎯 What it does: Constructed a VLM benchmark covering over 1M compressed images, 11 types of bitstreams, and 7 evaluation metrics, and proposed a lightweight adapter to enhance the understanding ability of VLMs for compressed images.

Benchmarking and Evolving Reason-Reflect-Rectify for Reflective Visual Generation

Junjie Wang (Harbin Institute of Technology), Liqiang Nie (Harbin Institute of Technology)

GenerationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringVision Language ModelDiffusion modelImageTextMultimodalityBenchmark

🎯 What it does: Propose R3-Bench benchmark and R3-Refiner framework to evaluate and improve the Reason-Reflect-Rectify reflection-correction loop in visual generation

Benchmarking and Improving Fine-Grained Text-to-Image Alignment via Paired Reinforcement Learning

Kaihang Pan (Zhejiang University), Siliang Tang (Zhejiang University)

GenerationData SynthesisTransformerReinforcement LearningPrompt EngineeringVision Language ModelDiffusion modelImageTextMultimodalityBenchmark

🎯 What it does: Propose DeltaBench benchmark and FocusDiff framework, leveraging paired reinforcement learning to enhance fine-grained semantic alignment in text-to-image generation.

Benchmarking at the Edge of Comprehension

Samuele Marro (University of Oxford), Philip Torr (University of Oxford)

Large Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose the Critique-Resilient Benchmarking framework, which replaces full-answer evaluation with adversarial critiques and verifiable local evidence.

Benchmarking Dense and Indiscernible Object Counting with Blueberries

Weihao Bo (Nanjing University of Science and Technology), Zechao Li (Nanjing University of Science and Technology)

TransformerVision Language ModelContrastive LearningImageMultimodalityBenchmarkAgriculture Related

🎯 What it does: Proposed the DIOCblueberry dataset and the MaskCount method to address the counting problem of blueberries in dense and hard-to-distinguish scenarios.

Benchmarking LLM-Assisted Blue Teaming via Standardized Threat Hunting

Yuqiao Meng (State University of New York at Binghamton), Zhaohan Xi (State University of New York at Binghamton)

Anomaly DetectionExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Constructed the CYBERTEAM benchmark to evaluate the performance and process of large language models in blue team threat hunting.

Benchmarking Physics-Informed Time-Series Models for Operational Global Station Weather Forecasting

Tao Han (Hong Kong University of Science and Technology), LEI BAI

Graph Neural NetworkTransformerContrastive LearningTime SeriesBenchmarkPhysics RelatedOrdinary Differential Equation

🎯 What it does: Proposes the global observational weather dataset WEATHER-5K and designs the PhysicsFormer physics-informed temporal forecasting model for full-station weather prediction.

Benchmarking Reward Hack Detection in Code Environments via Contrastive Analysis

Darshan Girish Deshpande (Patronus AI), Rebecca Qian (Patronus AI)

Anomaly DetectionReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextSequentialBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Created and made public the TRACE benchmark, which contains 517 manually verified multi-label reward cheating trajectories, covering 54 fine-grained categories, and evaluated the ability of LLMs to detect reward cheating through comparative anomaly detection.

Benchmarking the Limits of In-Context Reinforcement Learning for Ad-Hoc Teamwork

Yuheng Jing (Chinese Academy of Sciences), Jian Cheng (Chinese Academy of Sciences)

TransformerReinforcement LearningSequentialBenchmark

🎯 What it does: Constructed a large-scale reproducible ICRL4AHT benchmark to evaluate the adaptability of Transformer-based ICRL in Ad-Hoc team collaboration.

Benchmarking the Scientific Mind: A Pathology-Derived Biomedical VQA Benchmark for Complex Scientific Reasoning

Ziyu Zhao (University of Chinese Academy of Sciences), Haixin Wang (Chinese Academy of Sciences)

TransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextMultimodalityBiomedical DataBenchmarkChain-of-Thought

🎯 What it does: Constructed a biomedical VQA benchmark called SORBE based on pathological multi-graphs, requiring models to perform evidence-driven multi-step scientific reasoning across multiple images

Benchmarking World-Model Learning with Environment-Level Queries

Archana Warrier (Basis Research Institute), Zenna Tavares (Basis Research Institute)

Representation LearningReinforcement Learning from Human FeedbackReinforcement LearningPrompt EngineeringWorld ModelImageVideoTextTabularBenchmark

🎯 What it does: Propose the WorldTest framework, which evaluates world models using a two-stage reward-free interaction and environment-level query testing approach, and based on this, implements the AutumnBench benchmark.

Bend the Basics: Degradation-Aware Deformable Tokenization for All-in-One Image Restoration

Zihao He (Shanghai Jiao Tong University), Songhua Liu (Shanghai Jiao Tong University)

RestorationTransformerDiffusion modelScore-based ModelFlow-based ModelRectified FlowNeural Radiance FieldAuto EncoderGenerative Adversarial NetworkContrastive LearningImage

🎯 What it does: Propose a unified image restoration Transformer (FIT) with full-process deformable block processing, achieving spatially adaptive image denoising, de-raining, de-fogging, de-blurring, and low-light enhancement through degradation-aware blocking and reconstruction.

Benign Overfitting in Adversarial Training for Vision Transformers

Jiaming Zhang (King Abdullah University of Science and Technology), Di Wang (King Abdullah University of Science and Technology)

ClassificationAdversarial AttackTransformerContrastive LearningImage

🎯 What it does: This paper, for the first time in theory and experiment, explains that under appropriate signal-to-noise ratio and perturbation range, visual Transformers can exhibit benign overfitting during adversarial training, achieving zero robust training error while maintaining low robust test error.

BESplit: Bias-Compensated Split Federated Learning with Evidential Aggregation

Yuhan Xie (Shanghai University of Finance and Economics), Jingrong Huang (Shanghai University of Finance and Economics)

OptimizationFederated LearningKnowledge DistillationContrastive LearningImageTabular

🎯 What it does: Propose the BESplit framework to address the optimization bias and convergence instability caused by non-IID data in Split Federated Learning.

BESPOKE: Benchmark for Search-Augmented Large Language Model Personalization via Diagnostic Feedback

Hyunseo Kim (Yonsei University), Dongha Lee (Yonsei University)

Recommendation SystemExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Constructed the BESPOKE benchmark, collected real user chat and search history, and designed query-response pairs with fine-grained preference scoring and diagnostic feedback to evaluate the personalization capabilities of search-enhanced large language models.

Best of Both Worlds: Multimodal Reasoning and Generation via Unified Discrete Flow Matching

Onkar Kishor Susladkar (University of Illinois Urbana Champaign), Ismini Lourentzou (University of Illinois Urbana Champaign)

RecognitionImage TranslationRestorationSegmentationGenerationData SynthesisRepresentation LearningTransformerPrompt EngineeringMixture of ExpertsVision Language ModelDiffusion modelFlow-based ModelContrastive LearningImageTextMultimodalityBenchmark

🎯 What it does: Propose UniDFlow, a unified discrete flow matching framework that integrates multimodal understanding, text-to-image generation, and instruction-driven editing.

Best-of-Both-Worlds for Heavy-Tailed Markov Decision Processes

Yu Chen (Tsinghua University), Longbo Huang (Tsinghua University)

OptimizationReinforcement Learning

🎯 What it does: This paper proposes two algorithms for Markov decision processes with heavy-tailed losses (HTMDP), achieving Best-of-Both-Worlds (BoBW) performance in both scenarios where the transition probabilities are known and unknown;

BEST: Benchmarking Efficiency in Space and Time for LLM-Generated Code

Aocheng Shen (Huazhong University of Science and Technology), Xianjun Deng (Huazhong University of Science and Technology)

Computational EfficiencyAI Code AssistantTransformerLarge Language ModelTextBenchmark

🎯 What it does: This paper proposes the BEST benchmark to evaluate the time and space efficiency of code generated by LLMs.

Better, Faster: Harnessing Self-Improvement in Large Reasoning Models

Qihuang Zhong (Nanyang Technological University), Dacheng Tao (Nanyang Technological University)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextBiomedical DataBenchmarkChain-of-Thought

🎯 What it does: Through self-improving training, large reasoning models (LRM) generate high-quality reasoning trajectories and perform iterative self-training without external supervision.

Beyond Accuracy and Complexity: The Effective Information Criterion for Structurally Stable Symbolic Regression

Zihan Yu (Tsinghua University), Yong Li (Tsinghua University)

OptimizationExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerReinforcement LearningPrompt EngineeringGenerative Adversarial NetworkContrastive LearningTextTabularBenchmarkPhysics Related

🎯 What it does: Propose Effective Information Criterion (EIC) to quantify the structural stability of symbolic regression formulas, and apply it to both search-based and generative SR;

Beyond Accuracy: Latent Perturbations for Cognitive-Aware Diagnosis

Yuting Yan (Chinese University of Hong Kong), Shuang Li (Chinese University of Hong Kong)

Anomaly DetectionExplainability and InterpretabilityTransformerAuto EncoderGenerative Adversarial NetworkContrastive LearningTabularBiomedical DataElectronic Health Records

🎯 What it does: This paper proposes a cognition-aware adversarial diagnostic framework based on Denoising Masked AutoEncoder (DMAE), which generates adversarial 'hypotheses' through perturbations in the latent space to correct anchoring bias in clinical diagnosis and improve the detection accuracy of rare diseases.

Beyond Additive Decompositions: Interpretability Through Separability

Jinyang Liu (University of Copenhagen), Munir Hiabu (University of Copenhagen)

Explainability and InterpretabilityTabular

🎯 What it does: Proposed and implemented Tensor Separation Learning (TSL), an interpretable regression model based on separable product differentiation.

Beyond Attention Imbalance: Mitigating Hallucinations via Spectral Surgery

Siqi Lu (National University of Defense Technology), PENG WANG

Explainability and InterpretabilityComputational EfficiencyTransformerVision Language ModelContrastive LearningImageTextMultimodality

🎯 What it does: Designed FLASH, a training-agnostic and contrastive decoding-free spectral modulation framework, which automatically detects visual attention heads and performs dual-stream spectral modulation on attention scores and value matrices, thereby alleviating hallucination problems in large vision-language models.

Beyond Benchmarks: Toward Causally Faithful Evaluation of Large Language Models

Zhengshuyuan Tian (Institute of Computing Technology, Chinese Academy of Sciences), Jianfeng Zhan (Institute of Computing Technology, Chinese Academy of Sciences)

Explainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelTextBenchmark

🎯 What it does: Proposes the LLM Evaluatology framework, which performs causal decomposition and experimental design on LLM evaluation, systematically examining the impact of each component of the evaluation system on performance.

Beyond Binary: Continuous State Optimization with Graph-Structured Objectives

Corinna Cortes (Google Research), Mehryar Mohri (Google Research)

OptimizationGraph Neural NetworkReinforcement LearningGraphTabular

🎯 What it does: This paper extends competitive target optimization from binary states to continuous state spaces, and captures local dependencies between targets through graph-structured linear approximation. It proposes a lazy Graph-LinUCB algorithm to balance exploration and stability, and further provides three structural utilization schemes: asynchronous updates, graph structure learning, and joint estimation.

Beyond Blind Noising: Disentangled Visual Rectification for Hallucination Mitigation in MLLMs

Yujia Chen (University of Science and Technology of China), Wenzhang SUN (Tsinghua University)

RestorationAnomaly DetectionExplainability and InterpretabilityRepresentation LearningTransformerPrompt EngineeringVision Language ModelAuto EncoderContrastive LearningImageTextMultimodality

🎯 What it does: To address the image hallucination problem in multi-modal large language models (MLLM), the authors propose a training-free, plug-and-play dual-stream framework called Disentangled Visual Rectification (DVR). The framework utilizes the response differences of LIP (e.g., CLIP, SigLIP) and SSL (e.g., DINOv3) encoders under noise perturbations. It first adaptively suppresses or enhances the original features in the visual encoding layer, and then further weakens the residual hallucination-inducing components through a contrastive mechanism in the decoding layer.

Beyond Buffer Limits: Energy-Based Data Reassembly for Continual Learning

Zhenyi Wang (University of Central Florida), Heng Huang (University of Maryland)

ClassificationComputational EfficiencyRepresentation LearningData-Centric LearningMeta LearningConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningImageBiomedical Data

🎯 What it does: Propose an energy-based data reorganization (EBDR) method, which generates composite samples with higher information density for memory replay in continual learning by splitting the original image into patches and reorganizing them under an energy framework.

Beyond Confidence: Adaptive and Coherent Decoding for Diffusion Language Models

Kecheng Chen (City University of Hong Kong), Haoliang Li (City University of Hong Kong)

GenerationAI Code AssistantTransformerLarge Language ModelPrompt EngineeringDiffusion modelContrastive LearningText

🎯 What it does: Propose a sampling framework called CCD based on historical context consistency to improve the decoding process of diffusion language models.

Beyond Continuity: Simulation-free Reconstruction of Discrete Branching Dynamics from Single-cell Snapshots

Junda Ying (Peking University), Lei Zhang (Peking University)

Data SynthesisComputational EfficiencyRepresentation LearningTransformerDiffusion modelScore-based ModelFlow-based ModelTabularTime SeriesBiomedical DataBenchmarkStochastic Differential Equation

🎯 What it does: This paper proposes the Unbalanced Schrödinger Bridge (USB) framework, which utilizes the branching Schrödinger bridge theory and unbalanced score matching to simultaneously infer trajectories of stochasticity and imbalanced mass changes in single-cell snapshot data, and supports discrete cell birth/apoptosis simulations.

Beyond Correctness: Distance-Based Social Dynamics of Multi-Agent Debate

Seungwoong Ha (Santa Fe Institute), Melanie Mitchell (Santa Fe Institute)

Explainability and InterpretabilityRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringTextTabularBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Studied the microdynamics of answers in multi-agent debate systems, analyzing distance-aware revisions and convergence behaviors in social interactions using the ConceptARC two-dimensional grid task.

Beyond Description: Federated Adaptation via Semantic-Visual Prototype Alignment

Jiarong Yang (South China University of Technology), Yuan Liu (South China University of Technology)

Domain AdaptationFederated LearningRepresentation LearningTransformerVision Language ModelAuto EncoderContrastive LearningImageTextMultimodality

🎯 What it does: In the federated learning scenario, the FedSPA scheme is proposed, which utilizes the pre-trained Vision-Language model (CLIP) to achieve lightweight alternating optimization of visual prototypes and global semantic prototypes;

Beyond Detection: A Structure-Aware Framework for Scene Text Tracking

Chenmin Yu (Nankai University), Yu ZHOU (Nankai University)

Object TrackingTransformerContrastive LearningOptical FlowVideoTextBenchmark

🎯 What it does: This paper proposes a detection-free framework called SymTrack for scene text tracking (STT), and constructs three STT benchmarks based on video text detection data, systematically evaluating and significantly improving tracking performance.

Beyond Distribution Estimation: Simplex Anchored Structural Inference Towards Universal Semi-Supervised Learning

Yaxin Hou (Southeast University), Yuheng Jia (Southeast University)

ClassificationDomain AdaptationRepresentation LearningGraph Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningImage

🎯 What it does: Proposed a general semi-supervised learning framework called SAGE, which is designed for unknown, arbitrarily distributed, and extremely few labeled samples. The core idea is to replace distribution estimation with structural reasoning;

Beyond Drift: Stabilizing Subjective LLM Evaluation with Information-Theoretic Rubrics

Wang Xu (HeFei University of Technology), Qian Wan (Central China Normal University)

Explainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextBenchmarkChain-of-Thought

🎯 What it does: Construct a questionnaire-based evaluation framework based on Expected Information Gain (EIG) to address the dimension drift problem in subjective evaluation of LLMs.