ICML 2026 Papers — Page 8
International Conference on Machine Learning · 6554 papers
BlitzRank: Principled Zero-shot Ranking Agents with Tournament Graphs
Sheshansh Agrawal (Contextual AI), Douwe Kiela (Contextual AI)
RetrievalRecommendation SystemComputational EfficiencyGraph Neural NetworkTransformerLarge Language ModelAgentic AIPrompt EngineeringTextGraphBenchmarkRetrieval-Augmented Generation
🎯 What it does: This paper proposes BLITZRANK, a zero-shot ranking framework based on tournament graphs. It utilizes complete tournament information from k-wise comparisons, infers additional preferences through reachability closure, and stops querying once the nodes are fully determined, efficiently selecting the top m items.
Block Rotation is All You Need for MXFP4 Quantization
Yuantian Shao (Nanjing University of Science and Technology), Jian Cheng (Institute Of Automation Chinese Academy Of Sciences)
Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelAuto EncoderContrastive LearningTextBenchmark
🎯 What it does: This paper addresses the post-training quantization (PTQ) issue of low-precision quantization for MXFP4, systematically evaluates existing INT4 methods, and proposes Block-level Rotation (BRQ) to resolve the incompatibility between rotation and MXFP4.
BLOCK-EM: Preventing Emergent Misalignment via Latent Blocking
Muhammed Ustaomeroglu (Carnegie Mellon University), Guannan Qu (Carnegie Mellon University)
Safty and PrivacyExplainability and InterpretabilitySupervised Fine-TuningAuto EncoderContrastive LearningTextTabularFinance Related
🎯 What it does: Proposes a training-time implicit feature blocking method called BLOCK-EM, which suppresses emergent harmful behaviors generated by the model during fine-tuning by utilizing a small number of identified causal SAE dimensions.
Block-wise Codeword Embedding for Reliable Multi-bit Text Watermarking
Joeun Kim (DGIST), Young-Sik Kim (DGIST)
Anomaly DetectionComputational EfficiencyData-Centric LearningTransformerLarge Language ModelContrastive LearningTextRetrieval-Augmented Generation
🎯 What it does: This paper proposes the BREW framework, a reliable multi-bit text watermarking technique based on block-level codewords, which can achieve a high detection rate and extremely low false detection rate even when resisting attacks such as word substitution, insertion, and deletion.
Blocking the Leakage: Manifold-Aware Gradient Projection for Long-Horizon Test-Time Adaptation
Haoyu Xiong (Yunnan University), Yuanyuan Pu (Yunnan University)
Domain AdaptationContrastive LearningImage
🎯 What it does: Proposes the MGP (Manifold-Aware Gradient Projection) method, which uses subspace projection to prevent gradient leakage, thereby achieving stability in adaptation during long-term testing.
BlueCodeAgent: A Blue Teaming Agent Powered by Automated Red Teaming for CodeGen AI
Chengquan Guo (University of Chicago), Bo Li (University of Chicago)
Safty and PrivacyAI Code AssistantTransformerLarge Language ModelAgentic AIPrompt EngineeringTextRetrieval-Augmented Generation
🎯 What it does: Propose an end-to-end blue team agent called BlueCodeAgent, which utilizes automated red teaming to generate diverse risk examples, and implements security protection for code generation AI through the construction of a constitution and dynamic testing.
BOCLOAK: Optimal Transport-Guided Adversarial Attacks on Graph Neural Network-Based Bot Detection
Kunal Mukherjee (Virginia Tech), Murat Kantarcioglu (Virginia Tech)
OptimizationAdversarial AttackGraph Neural NetworkGenerative Adversarial NetworkGraph
🎯 What it does: Propose a framework called BOCLOAK based on optimal transport for edge editing and node injection attacks targeting social bot detection in graph neural networks.
Boost the Identity-Preserving Embedding for Consistent Visual Generation
Zixun Xia, Yaxing Wang (Jilin University)
GenerationRepresentation LearningTransformerPrompt EngineeringVision Language ModelDiffusion modelContrastive LearningImageVideoText
🎯 What it does: Proposed the BIPE framework, which is training-free and pluggable, achieving identity-consistent generation of the same subject in videos and sequential images by explicitly extracting and enhancing identity-preserving embeddings (IPemb) from text embeddings.
BOOSTAPR: Boosting Automated Program Repair via Execution-Grounded Reinforcement Learning with Dual Reward Models
Yuanhao Li (Beijing University of Posts and Telecommunications), Xuhong Chen (Beijing University of Posts and Telecommunications)
AI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningContrastive LearningTextSequential
🎯 What it does: Propose a three-stage framework called BOOSTAPR, which achieves automatic program repair by performing verified supervised fine-tuning, dual reward model training, and PPO reinforcement learning.
Boosting CVaR Policy Optimization with Quantile Gradients
Yudong Luo (HEC Montreal), Erick Delage (HEC Montreal)
Reinforcement Learning
🎯 What it does: Propose an algorithm that incorporates the VaR (quantile) gradient into the CVaR policy gradient, leveraging the dynamic programming properties of VaR to improve sample efficiency and overcome the 'success blind spot' of CVaR-PG.
Boosting Monocular Metric Depth Estimation via Bokeh Rendering
Hangwei Zhang (Nanyang Technological University), Xingang Pan (Nanyang Technological University)
Image TranslationGenerationData SynthesisDepth EstimationConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningOptical FlowImageMultimodalityPoint Cloud
🎯 What it does: Proposes a two-stage framework called BokehDepth, which first synthesizes calibrated Bokeh layer stacks on a single clear image using a generative model without depth information, and then injects this stack into the encoder of an existing monocular metric depth network through a specialized DSFA module, leveraging optically quantified defocus information to improve the accuracy and physical consistency of depth estimation.
Boosting World Models Learning via Latent-Space Value Alignment
Xingyu Jiang (Beihang University), Yue Deng (Beihang University)
Recurrent Neural NetworkReinforcement LearningAuto EncoderWorld ModelImageVideoBenchmark
🎯 What it does: This paper proposes a value-aligned world model, which guides the model to focus on task-related features by introducing value-aligned regularization in the latent space.
Bootstrapped Exploration with Causal Reasoning: A Training Paradigm for Adaptive Forecasting Agent
Qingwen Zeng (University of Sydney), Ling Chen (University of Technology Sydney)
Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningAdversarial AttackHyperparameter SearchData-Centric LearningRobotic IntelligenceMeta LearningDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackNeural Architecture SearchProtein Structure PredictionTransformerLarge Language ModelReinforcement LearningPrompt EngineeringContrastive LearningWorld ModelTabularTime SeriesChain-of-Thought
🎯 What it does: This paper proposes BECRA, a training paradigm based on contrastive-aware UCB sampling and causal lesson extraction, which enables an LLM agent to learn time series prediction strategies and achieve zero-shot transfer without the need for human-labeled data;
Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing
Shilong Zhang (University of Hong Kong), Ping Luo (University of Hong Kong)
GenerationRepresentation LearningTransformerVision Language ModelDiffusion modelAuto EncoderContrastive LearningImageText
🎯 What it does: For generative tasks, the paper maps the features of a high-dimensional self-supervised representation encoder into a compact latent space with 96 channels and KL regularization, and jointly combines pixel-level reconstruction with semantic reconstruction to build PS-VAE for text-to-image generation and image editing.
Bottleneck Communication Delay Minimization for Communication-Efficient Decentralized Learning
Nozomi Hata (NTT Communication Science Laboratory, NTT Inc.), Kenta Niwa (NTT Communication Science Laboratory, NTT Inc.)
OptimizationFederated LearningReinforcement LearningContrastive LearningImageGraph
🎯 What it does: This paper proposes an approximate solver named BTSP-MSR for node allocation on cyclic directed graphs such as rings, exponential graphs, and 1-peer exponential graphs in heterogeneous communication delay environments, aiming to minimize bottleneck communication delay (BCD);
Bottleneck-Guided Spectral Subgoals For Offline Goal-Conditioned RL
Hebin Liang (Tianjin University), Jianye HAO
Graph Neural NetworkTransformerReinforcement LearningDiffusion modelContrastive LearningGraphTabularSequential
🎯 What it does: Designed the BASS framework: first, use offline data to train Laplacian representations and perform spectral clustering, automatically discovering bottlenecks in the state space; then extract key points (KPs) near the bottlenecks and construct a reachability graph; during deployment, plan subgoal sequences across bottlenecks through graph search, and complete each short-term transition using pluggable low-level controllers (e.g., Decision Diffuser or MLP).
Boundary Embedding Shaping with Adaptive Contrastive Learning for Graph Structural Disentanglement
Jiaqing Chen (Yunnan Normal University), Javen Qinfeng Shi (Adelaide University)
ClassificationRepresentation LearningGraph Neural NetworkContrastive LearningGraph
🎯 What it does: Decouple noise in graph structures, specifically targeting nodes on the decision boundary through a boundary embedding shaping (BES) module to achieve graph structure decoupling, thereby improving performance in node classification and link prediction.
Bounded Hyperbolic Tangent: A Stable and Efficient Alternative to Pre-Layer Normalization in Large Language Models
Hoyoon Byun (Yonsei University), Kyungwoo Song (Upstage AI)
Computational EfficiencyRepresentation LearningTransformerLarge Language ModelContrastive LearningText
🎯 What it does: Propose BHyT, a stable and efficient normalization method that can serve as an alternative to Pre-LN
Box Thirding: Anytime Best Arm Identification under Insufficient Sampling
SEOHWA HWANG, Junyong Park (Seoul National University)
OptimizationReinforcement Learning from Human FeedbackText
🎯 What it does: Propose the Box Thirding (B3) algorithm, which achieves Best Arm Identification (BAI) at any time (anytime) under budget constraints, while maintaining the screening and promotion of candidate arms under data-poor conditions.
BPDQ: Bit-Plane Decomposition Quantization on a Variable Grid for Large Language Models
Junyu Chen (Southwestern University of Finance and Economics), Ngai Wong (University of Hong Kong)
OptimizationComputational EfficiencyTransformerLarge Language ModelText
🎯 What it does: Proposed and implemented a low-bit quantization method called BPDQ for large language models, which constructs variable quantization grids using bit-plane decomposition and iteratively optimizes under the Hessian geometry.
BPL: Generalizable Deepfake Detection via Bias-only Pair-aware Learning
Yuxiang Xu (Shandong University), Yilong Yin (Shandong University)
Anomaly DetectionTransformerSupervised Fine-TuningVision Language ModelContrastive LearningImage
🎯 What it does: Proposes a general deepfake detection framework based on Bias-Only Pair-Aware Learning (BPL), which leverages semantically aligned real-fake image pairs to learn fine-grained differences caused by generation, and maintains semantic structure through bias-only fine-tuning;
Brain Networks Should Be Learned, Not Constructed
Liang Yang (Hebei University of Technology), Xiaochun Cao (Sun Yat-sen University)
ClassificationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningGraph Neural NetworkAuto EncoderContrastive LearningTime SeriesBiomedical DataMagnetic Resonance ImagingAlzheimer's Disease
🎯 What it does: This paper proposes a learnable brain functional network representation (BRep), which transforms traditional handcrafted correlation coefficients into high-order, learnable dependency measures. It directly performs end-to-end training based on this, and finally completes disease prediction using a simple MLP.
BrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and Language
Haitao Wu (Tianjin University), Jiamin Wu (Shanghai Artificial Intelligence Laboratory)
GenerationRepresentation LearningTransformerVision Language ModelDiffusion modelAuto EncoderContrastive LearningImageTextBiomedical DataMagnetic Resonance Imaging
🎯 What it does: Built BrainJanus, a unified multimodal model that can encode and decode between the brain, vision, and language, enabling generation between any two modalities;
Branch Scaling Manifests as Implicit Architectural Regularization for Improving Generalization in Overparameterized ResNets
Zixiong Yu (Huawei Large Model Data Technology Lab), Songtao Tian (Tsinghua University)
ClassificationConvolutional Neural NetworkImage
🎯 What it does: This paper investigates the impact of residual branch scaling on the generalization performance of wide ResNet through theoretical and experimental analysis, proving that constant scaling leads to non-learnability, while fast decaying scaling combined with early stopping can achieve optimal generalization error.
Branching Diffusion for Point Processes in Time and Space
Chao Yang (Chinese University of Hong Kong, Shenzhen), Shuang Li (Chinese University of Hong Kong, Shenzhen)
GenerationData SynthesisDiffusion modelScore-based ModelPoint CloudTime SeriesSequentialStochastic Differential Equation
🎯 What it does: Proposed a non-autoregressive spatiotemporal point process generation model based on branching diffusion, which can simultaneously handle event locations, time, and the number of events.
Break the Block: Dynamic-size Reasoning Blocks for Diffusion Large Language Models via Monotonic Entropy Descent with Reinforcement Learning
Yan Jiang (University of Queensland), Zi Huang (University of Queensland)
OptimizationComputational EfficiencyTransformerLarge Language ModelReinforcement LearningDiffusion modelTextBenchmark
🎯 What it does: By introducing the b1 framework into the post-training of diffusion large language models (dLLM), the model can learn dynamic-sized reasoning blocks, thereby enhancing the coherence and accuracy of the reasoning process.
Breaking Dual Bottlenecks: Evolving Unified Multimodal Models into Self-Adaptive Interleaved Visual Reasoners
Qingyang Liu (Shanghai Jiao Tong University), Li Niu (Shanghai Jiao Tong University)
GenerationData SynthesisExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision Language ModelDiffusion modelGenerative Adversarial NetworkImageTextMultimodalityBenchmarkChain-of-Thought
🎯 What it does: Propose an adaptive interactive visual reasoning framework called InterVR, which can dynamically switch between three strategies—direct generation, reflection correction, and multi-step planning—within a unified multimodal model, to bridge the gap between understanding and generation.
Breaking Manifold Continuity: Vector Quantized Modeling for Real-Centric Deepfake Detection
Changshuo Wang (Shanghai Jiao Tong University), Lizhuang Ma (Shanghai Jiao Tong University)
Anomaly DetectionData-Centric LearningTransformerVision Language ModelAuto EncoderContrastive LearningImageVideo
🎯 What it does: By introducing a vector quantized codebook into the feature space of the CLIP visual encoder, a discrete latent space of real data is constructed, thereby enabling the detection of forged samples;
Breaking Multi-Task Curse: Reward-Weighted Evolution for Black-Box Many-Task Optimization
Yanchi Li (China University of Geosciences), Yew-Soon Ong (Nanyang Technological University)
OptimizationReinforcement LearningTabularTime SeriesSequential
🎯 What it does: This paper proposes an evolutionary strategy suitable for black-box multi-task optimization, aiming to overcome the performance degradation—known as the multi-task curse—that traditional multi-task methods encounter when the number of tasks increases and their similarity decreases.
Breaking the Block: Preserving Data Continuity to Train Superior SAEs for Instruct Models
Jiaming Li (Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences), Min Yang (Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences)
Representation LearningData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningAuto EncoderText
🎯 What it does: To address the block training issue in sparse autoencoder training for instruction models, the FAST method based on fine-grained sequential training is proposed and implemented.
Breaking the Capacity Bottleneck in Model-Heterogeneous Federated Learning via Gradual Model Restoration
Chengjie Ma (Yonsei University), Seong-Lyun Kim (Yonsei University)
OptimizationFederated LearningComputational EfficiencyConvolutional Neural NetworkTransformerAgentic AIPrompt EngineeringMixture of ExpertsAuto EncoderContrastive LearningImageTextRetrieval-Augmented Generation
🎯 What it does: Proposes a framework called FedGMR that alleviates the capacity bottleneck in model heterogeneous federated learning by progressively restoring model capacity.
Breaking the Computational Barrier: Provably Efficient Actor–Critic for Low-Rank MDPs
Ruiquan Huang (University of Kentucky), Jing Yang (University of Virginia)
OptimizationComputational EfficiencyReinforcement LearningTabularTime SeriesSequential
🎯 What it does: A new optimistic actor-critic algorithm (OptAC) is proposed, specifically designed for low-rank Markov decision processes (MDPs), which relies solely on a policy evaluation oracle, avoiding the computationally complex planning or optimization oracles common in previous methods.
Breaking the Echo Chamber: A Dynamic Ensemble Pruning Perspective on MoE
Xinlai Kang (Renmin University of China), Cheng Meng (Renmin University of China)
Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningMixture of ExpertsContrastive LearningTextBenchmark
🎯 What it does: Propose an integrated pruning perspective-based Mixture-of-Experts routing framework, MP-MoE, based on Mahalanobis distance, to address expert redundancy and representation collapse issues;
Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for Open-Ended LLM Reasoning
Yang Zhou (Zhejiang University), Mingli Song (Zhejiang University)
Data-Centric LearningTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBiomedical DataBenchmark
🎯 What it does: Addressing the exploration bottleneck in open-ended LLM reasoning through Rubric-Scaffolded Reinforcement Learning (RuscaRL).
Breaking the Factorization Barrier in Diffusion Language Models
Ian Li (University of California, San Diego), Anji Liu (National University of Singapore)
GenerationComputational EfficiencyData-Centric LearningTransformerLarge Language ModelDiffusion modelScore-based ModelContrastive LearningText
🎯 What it does: Propose the CoDD framework, which couples discrete diffusion language models with tractable probabilistic circuits (Probabilistic Circuits), breaking the traditional independence assumption and allowing the joint distribution of multiple words to be modeled in a single denoising step.
Breaking the Lock-in: Diversifying Text-to-Image Generation via Representation Modulation
Dahee Kwon (Korea Advanced Institute of Science and Technology), Jaesik Choi (Korea Advanced Institute of Science and Technology)
GenerationTransformerPrompt EngineeringDiffusion modelFlow-based ModelImageText
🎯 What it does: Propose a method of internal representation layer intervention without training and without additional sampling steps—DAVE—which significantly enhances the diversity of text-to-image generation by attenuating the DC (zero-frequency) component of Transformer hidden representations in the early generation stage.
Breaking the Reference Bottleneck via Learning to Rewrite Conversational Queries without Gold Reference Passages
Doyoung Kim (KAIST), Jae-Gil Lee (Korea University)
Data SynthesisRetrievalOptimizationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringDiffusion modelAuto EncoderTextSequentialRetrieval-Augmented Generation
🎯 What it does: Proposes a no-reference paragraph-based dialogue query rewriting (CQR) preference optimization framework called DUALREFORM, which can automatically generate pseudo reference paragraphs from datasets containing only queries and responses, and use these pseudo paragraphs for preference optimization, thereby improving the performance of dialogue query rewriting.
Breaking the Reversal Curse in Autoregressive Language Models via Identity Bridge
Xutao Ma (University Of California Berkeley), Somayeh Sojoudi (University Of California Berkeley)
Computational EfficiencyRepresentation LearningData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningText
🎯 What it does: Proposed a method to solve the 'curse of reversal' problem in autoregressive language models by introducing an 'identity bridge' regularization technique in the training data;
Breaking the Scale Barrier: One-Shot Knowledge Transfer via Frequency Transform
Jianlu Shen (Southeast University), Xin Geng (Southeast University)
ClassificationObject DetectionSegmentationComputational EfficiencyKnowledge DistillationRepresentation LearningConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningImageText
🎯 What it does: This paper proposes to extract the low-frequency components (i.e., 'learngene') of pre-trained model weights using discrete cosine transform, achieving one-time cross-scale model initialization without training;
Breaking the Self-Confirming Loop: Diagnosing and Mitigating Systemic Reward Bias in Self-Rewarding RL
Chuyi Tan (Beijing Institute of Technology), Kan Li (Beijing Institute of Technology)
Reinforcement Learning from Human FeedbackReinforcement LearningPrompt EngineeringMixture of ExpertsText
🎯 What it does: This paper studies the systematic reward bias present in self-reward reinforcement learning (RLIR), and proposes an RLER method based on integrated rewards to eliminate this bias.
Breaking the Simplification Bottleneck in Amortized Neural Symbolic Regression
Paul Saegert (Heidelberg University), Ullrich Koethe
OptimizationComputational EfficiencyRepresentation LearningTransformerReinforcement LearningPrompt EngineeringAuto EncoderContrastive LearningTabularTime SeriesSequentialPhysics Related
🎯 What it does: Proposes a rule-based matching SIMPLIPY simplification engine and the FLASH-ANSR training framework to significantly improve the speed and quality of expressions in amortized neural symbolic regression.
Breaking the Synthetic-Real Domain Shortcut for Training-Free Generative Replay-based Class Incremental Learning
Tao Zhang (Huazhong University of Science and Technology), Xu Zou (Huazhong University of Science and Technology)
ClassificationGenerationData SynthesisDomain AdaptationKnowledge DistillationTransformerDiffusion modelContrastive LearningImageText
🎯 What it does: The paper proposes an example-free incremental learning method that synthesizes old class data using a pre-trained text-to-image generator, and eliminates domain bias between synthesized and real data through feature domain correction and prototype alignment.
Bregman meets Lévy: Stochastic Mirror Descent with Heavy-Tailed Noise in Continuous and Discrete Time
Pierre-Louis Cauvin (University of Grenoble Alpes), Panayotis Mertikopoulos (University of Grenoble Alpes)
OptimizationStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: This paper studies the convergence of stochastic mirror descent (SMD) under heavy-tailed noise, proposing a continuous-time Lévy mirror flow (LMF) model and providing the corresponding discrete-time analysis.
Brep2Shape: Boundary and Shape Representation Alignment via Self-supervised Transformers
Yuanxu Sun (Tsinghua University), Mingsheng Long (Tsinghua University)
ClassificationSegmentationRepresentation LearningTransformerDiffusion modelAuto EncoderContrastive LearningPoint CloudMesh
🎯 What it does: Proposes the Brep2Shape self-supervised pre-training framework, achieving alignment between abstract parameters and intuitive shapes by mapping B-rep control points to dense spatial points.
Bridge Matching Sampler: Scalable Sampling via Generalized Fixed-Point Diffusion Matching
Denis Blessing (Karlsruhe Institute of Technology), Gerhard Neumann (Karlsruhe Institute of Technology)
GenerationOptimizationComputational EfficiencyDiffusion modelScore-based ModelPoint CloudTabularTime SeriesBiomedical DataPhysics RelatedStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Proposed the Bridge Matching Sampler (BMS), achieving stable and efficient sampling for arbitrary priors and target distributions through generic fixed-point iteration and Nelson's identity.
BRIDGE: Predicting Human Task Completion Time From Model Performance
Fengyuan Liu (Mila Quebec AI Institute), Hugo Larochelle (Mila Quebec AI Institute)
Explainability and InterpretabilityComputational EfficiencyAI Code AssistantLarge Language ModelTextTabularBenchmark
🎯 What it does: Propose the BRIDGE framework, which jointly models model performance using the two-parameter logistic IRT, infers the latent difficulty of tasks, and aligns it with human completion time, enabling the prediction of the time required for new benchmark tasks without relying on manual annotations.
BRIDGE: Triangular Fixed-Point Refinement for Long-Horizon Persona Consistency
Yinghui Jiang (National Institute for Data Science in Health and Medicine, Xiamen University), Haotong Sun (Hangzhou Shenji Technology Co., Ltd)
OptimizationExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose the BRIDGE framework, which refines explicit coupling between observable behavior (O), latent state (L), and memory (M) through triangular fixed-point iteration, ensuring internal state consistency before each response;
Bridging Dynamics and Data: A Unified Diffusion Framework for Mechanistically-Informed Epidemic Forecasting
Guanghui Min (University of Virginia), Chen Chen (University of Virginia)
Drug DiscoveryGraph Neural NetworkTransformerDiffusion modelScore-based ModelContrastive LearningGraphTabularTime SeriesElectronic Health RecordsBenchmarkOrdinary Differential Equation
🎯 What it does: Propose a hybrid forecasting framework (EpiDiff) that combines epidemiological mechanism models with diffusion models to achieve accurate prediction of epidemic time series.
Bridging Functional and Representational Similarity via Usable Information
Antonio Almudévar (University of Zaragoza), Alfonso Ortega (University of Zaragoza)
ClassificationImage TranslationRestorationCompressionExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningConvolutional Neural NetworkTransformerAuto EncoderGenerative Adversarial NetworkContrastive LearningOptical FlowImageTabularTime SeriesSequentialStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: This paper proposes a unified framework that connects functional similarity (whether they can be interchangeable functionally) and representational similarity (whether they can correspond structurally) through usable information, clarifying their hierarchical relationship and providing operational measurement methods.
Bridging Functional Correctness and Runtime Efficiency Gaps in LLM-Based Code Translation
Longhui Zhang (Harbin Institute of Technology), Min Zhang (Harbin Institute of Technology)
Computational EfficiencyAI Code AssistantTransformerLarge Language ModelPrompt EngineeringTextSequentialBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposes the SWIFTTRANS framework, which enhances both the functional correctness of LLM code translation and significantly improves runtime efficiency.
Bridging Local–Global Dissonance: Learning from Compressive Measurements for Hyperspectral Reconstruction
Xian-Hua Han (University of Rikkyo)
RestorationDepth EstimationSuper ResolutionConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningOptical FlowImageMultimodalityComputed TomographyPhysics RelatedStochastic Differential Equation
🎯 What it does: This paper proposes a Hierarchical Scale Alignment Architecture (HSRA) to address the local-global imbalance problem in compressed sensing spectral imaging, improving the quality of recovering high-dimensional spectral images from single-frame compressed measurements.
Bridging On-Device and Cloud LLMs for Collaborative Reasoning: A Unified Methodology for Local Routing and Post-Training
Wenzhi Fang (Purdue University), Christopher Brinton
Federated LearningComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmark
🎯 What it does: Achieve collaborative reasoning between edge devices and cloud-based LLMs, training a local model to autonomously decide whether to invoke the cloud model during inference.
Bridging RGB and RAW: Single-step Deterministic Flow with Homogeneous Representation Alignment
Diedong Feng (University of Electronic Science and Technology of China), Shuaicheng Liu (University of Electronic Science and Technology of China)
RestorationTransformerDiffusion modelFlow-based ModelAuto EncoderImage
🎯 What it does: This paper studies a single-step deterministic flow framework called SHADE for reconstructing high-fidelity RAW sensor data from RGB images.
Bridging Spherical Black-Box Optimizers
Johannes Ackermann (The University of Tokyo), Stefano Peluchetti (Sakana AI)
OptimizationGaussian SplattingTabularTime SeriesSequentialBenchmark
🎯 What it does: A unified general Master Update framework for black-box optimizers such as ES, OVI, and CBO was constructed, and based on this, hybrid methods that can interpolate between aggregation methods and consensus ranges were designed, such as ES-OVI, AdaPol, and SchedPol.
Bridging Structure and Semantics: Uncertainty-Modulated Dual-Path Diffusion for Robust Text-Attributed Graph Learning
Zhizhi Yu (Tianjin University), Di Jin (Tianjin University)
Representation LearningGraph Neural NetworkTransformerLarge Language ModelDiffusion modelContrastive LearningTextGraph
🎯 What it does: Propose an Uncertainty-Modulated Dual-Path Diffusion Model (UDPD), which mitigates the structure-semantic mismatch and dual-source noise problems by separately diffusing and denoising the semantic and structural embeddings in the text-attribute graph, and adaptively regulating the interaction between the two paths during the reverse process based on node uncertainty.
Bridging the Gap Between Average and Discounted TD Learning
Haoxing Tian (Boston University), Alex Olshevsky (Boston University)
Reinforcement LearningTabularBenchmark
🎯 What it does: Proposes two average reward TD learning algorithms, double-chain and single-chain, and provides finite-sample convergence analysis, ensuring iterative convergence to a unique, sample-independent solution;
Bridging the Grounding Gap in VideoQA via Typed Memory for Language-based Belief-State Reasoning
Saman Forouzandeh (Royal Melbourne Institute of Technology University), Mahdi Jalili (Royal Melbourne Institute of Technology University)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelVision Language ModelVision-Language-Action ModelContrastive LearningVideoTextRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: LINGUA is a memory-based Agent that utilizes event-driven perception, Typed Memory, and Belief-Action-Verification loops to perform VideoQA within explicit language belief states, thereby addressing the 'Grounding Gap' of traditional models.
Bridging the Knowledge-Prediction Gap in LLMs on Multiple-Choice Questions
Yoonah Park (Seoul National University), Yohan Jo (Seoul National University)
Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBenchmark
🎯 What it does: Investigate the knowledge-prediction gap of LLMs on multiple-choice questions, analyze their geometric structure, and propose the KAPPA intervention method during reasoning to align the knowledge subspace with the prediction subspace.
Bridging the Perceptual Gap: Residual-Enhanced Downscaling and Manifold-Aware Perception Alignment Adaptation for NR-IQA
Yu Li (Harbin Institute of Technology), Shaohui Liu (Harbin Institute of Technology)
Domain AdaptationRepresentation LearningTransformerSupervised Fine-TuningVision Language ModelAuto EncoderContrastive LearningImageTextMultimodality
🎯 What it does: This paper proposes a cross-modal perceptual alignment framework combining low-dimensional perceptual subspace projection with residual-enhanced downsampling (CMPA+RPD) to improve the performance of no-reference image quality assessment (NR-IQA).
Bridging the Stability-Expressivity Gap: Synthetic Data Scaling and Preference Alignment for Low-Resource Spoken Language Models
Yizhong Geng (Beijing University of Posts and Telecommunications), Xiaoyu Shen (Eastern Institute of Technology)
GenerationData SynthesisTransformerSupervised Fine-TuningScore-based ModelFlow-based ModelAudio
🎯 What it does: This paper investigates the trade-off between stability and expressiveness when expanding Speech Language Models (SLM) using synthetic data in low-resource languages, and proposes a self-alignment method to alleviate the problem of synthetic degradation.
Bridging Time and Frequency: A Joint Modeling Framework for Irregular Multivariate Time Series Forecasting
Xiangfei Qiu (East China Normal University), Jilin Hu (East China Normal University)
Anomaly DetectionData-Centric LearningTransformerMixture of ExpertsAuto EncoderContrastive LearningTabularTime SeriesBiomedical DataElectronic Health RecordsBenchmark
🎯 What it does: Proposed a novel TFMixer framework for accurate prediction of irregular multivariate time series.
Bridging Tokens and Geometry: Token-wise 3D Supervision for CAD Generation
Yijia Guan (Shanghai Jiao Tong University), Jianhua Sun (Shanghai Jiao Tong University)
GenerationData SynthesisTransformerSupervised Fine-TuningPrompt EngineeringDiffusion modelTextPoint CloudMeshSequential
🎯 What it does: Propose Argument-induced 3D Point Loss (A3PL) and Grammar-constrained Operator (GCO), achieving three-dimensional geometric supervision for each parameter token in CAD program sequences, and enforcing syntactic constraints during sequence generation, thereby enhancing the accuracy and effectiveness of generated geometry.
Bridging Your Imagination with Audio-Video Generation via a Unified Director
Jiaxu Zhang, Xin Chen (ByteDance Intelligent Creation)
GenerationData SynthesisTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsDiffusion modelRectified FlowImageVideoTextMultimodality
🎯 What it does: Propose the UniMAGE model, unifying script writing and keyframe generation to achieve long-form, multi-shot, movie-level audio-visual generation.
Bring Future Vision: Dynamic Computation Allocation Guided by Lightweight Feature Forecaster
Chao Han (Eastern Institute of Technology), Xiaoyu Shen (Eastern Institute of Technology)
Computational EfficiencyKnowledge DistillationTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsText
🎯 What it does: Proposes a dynamic computational allocation framework based on a lightweight feature predictor (LFF), improving traditional greedy routing to achieve efficient inference for large language models.
Bring My Cup! Personalizing Vision-Language-Action Models with Visual Attentive Prompting
Sangoh Lee (POSTECH), Wook-Shin Han (POSTECH)
Robotic IntelligencePrompt EngineeringVision Language ModelVision-Language-Action ModelImageMultimodalityBenchmark
🎯 What it does: Designed a training-agnostic visual attention prompt (VAP) framework that enables frozen Vision-Language-Action (VLA) models to identify and manipulate users' personal objects with only a few reference images.
Bringing Code ALIVE: Optimizing Interactive Frontend Mini-Games via Automated Play and Reinforcement Learning at Scale
Jiajun Zhang (University of Science and Technology of China), Binyuan Hui (Alibaba Group)
OptimizationAI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AIPrompt EngineeringTextMultimodalityRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose the ALIVE framework, which uses one-time planning combined with DOM/accessibility tree analysis to achieve large-scale automated evaluation of frontend mini-games. The evaluation score drives supervised fine-tuning and reinforcement learning, thereby improving the quality of interactive code in open-source models.
Broadening the Backdoor Basin: Understanding LLM Backdoors Collapse and Making Backdoors Persistent
Xingyi Zhao (Utah State University), Shuhan Yuan (Utah State University)
OptimizationAdversarial AttackTransformerLarge Language ModelSupervised Fine-TuningText
🎯 What it does: Proposed a backdoor attack method called BAD-BOOM based on Fisher information, which widens the behind-the-scenes trap basin, allowing LLM to maintain a high attack success rate in subsequent SFT.
BrokenMath: A Benchmark for Sycophancy in Theorem Proving with LLMs
Ivo Petrov (INSAIT), Martin Vechev
TransformerLarge Language ModelSupervised Fine-TuningAgentic AIPrompt EngineeringTextBenchmark
🎯 What it does: This paper proposes the BROKENMATH benchmark to evaluate the 'sycophantic' behaviors exhibited by large language models in natural language theorem proving.
BroRL: Scaling Reinforcement Learning via Broadened Exploration
Jian Hu (NVIDIA), Yi Dong (NVIDIA)
TransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningText
🎯 What it does: This paper proposes expanding the exploration range by significantly increasing the roll-out number (N) in each example, forming a new RL extension strategy called BroRL. It verifies that BroRL can continue to improve model performance even after reaching a saturation point, based on the existing ProRLv2 framework.
BTSP-CAM: A Brain-Inspired Geometric Memory for Class-Incremental Learning
Zheng Zhang (Dalian University of Technology), Qi Xu (Dalian University of Technology)
ClassificationComputational EfficiencyRepresentation LearningMeta LearningSpiking Neural NetworkPrompt EngineeringAuto EncoderContrastive LearningImageBenchmark
🎯 What it does: Proposes BTSP-CAM, a gradient-free binary memory module based on the brain's BTSP mechanism, for sample-free class-incremental learning.
BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning
Yuhang Xu (Shanghai Jiao Tong University), Guihai Chen
Computational EfficiencyTransformerLarge Language ModelReinforcement LearningText
🎯 What it does: Propose the BubbleSpec framework, which utilizes GPU idle time during synchronous reinforcement learning for pre-generation, and employs suffix-tree for model-free speculative decoding, thereby significantly accelerating the rollout phase.
Budget-Constrained Step-Level Diffusion Caching
Mingkun Lei (Westlake University), Chi Zhang (Westlake University)
GenerationComputational EfficiencyDiffusion modelImageVideoTextOrdinary Differential Equation
🎯 What it does: During the iterative sampling process of diffusion models, BudCache is proposed to fix the NFE budget through an offline search caching strategy, thereby achieving predictable inference latency and maximizing the final generation quality.
Budget-Efficient Attacks and Robustness Training for Cooperative MARL
junyong jiang (Southeast University), Lu Dong (Southeast University)
Adversarial AttackGraph Neural NetworkReinforcement LearningContrastive LearningGraphBenchmark
🎯 What it does: Proposed a hierarchical attack called BHEA under budget constraints and an adversarial training framework called BHEA-AT using this attack, aiming to enhance the robustness of collaborative multi-agent reinforcement learning.
Budget-Feasible Mechanisms for Submodular Welfare Maximization in Procurement Auctions
Shuang Cui (Soochow University), Chen Xue (Soochow University)
OptimizationFederated LearningGraphBenchmarkFinance Related
🎯 What it does: Propose a budget feasible mechanism BFM-SWM based on a clock auction to achieve submodular welfare maximization in procurement auctions, providing an approximation ratio of 0.0328 (non-monotonic) or 0.0877 (monotonic) while ensuring truthfulness, individual rationality, and budget feasibility.
Budgeted Active Experimentation for Treatment Effect Estimation from Observational and Randomized Data
Jiacan Gao (East China Normal University), Zhiheng Zhang (Shanghai University of Finance and Economics)
OptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyData-Centric LearningReinforcement Learning from Human FeedbackReinforcement LearningAuto EncoderContrastive LearningTabularTime SeriesBenchmark
🎯 What it does: Developed a budget-based active experimental design framework that uses observational logs to guide sample selection in limited randomized controlled trials, aiming to more efficiently estimate heterogeneous treatment effects.
BuildArena: A Physics-Aligned Interactive Benchmark of LLMs for Engineering Construction
Tian Xia (Westlake University), Tailin Wu (Westlake University)
TransformerLarge Language ModelAgentic AIPrompt EngineeringWorld ModelTextPoint CloudMeshBenchmarkPhysics Related
🎯 What it does: Proposed BuildArena, the first physics-aligned interactive engineering construction benchmark for LLMs;
Building Reliable Long-Form Generation via Hallucination Rejection Sampling
Lin Li (University of Oxford), Yarin Gal (University of Oxford)
GenerationTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposes the SHARS framework, which utilizes an arbitrary hallucination detector to reject hallucinatory text segment by segment during the generation process and resample it, thereby achieving reliable generation of long texts.
Building Social World Model with Large Language Models
Haofei Yu (University of Illinois Urbana-Champaign), Jiaxuan You (University of Illinois Urbana-Champaign)
TransformerLarge Language ModelPrompt EngineeringWorld ModelTextTime SeriesBenchmarkFinance RelatedRetrieval-Augmented Generation
🎯 What it does: Propose the Social World Model (SWM), which leverages large language models to capture the dynamics of social beliefs as they evolve with events;
Bulk-Calibrated Credal Ambiguity Sets: Fast, Tractable Decision Making under Out-of-Sample Contamination
Mengqi Chen (University of Warwick), Michele Caprio (University of Manchester)
Domain AdaptationOptimizationData-Centric LearningTextTabular
🎯 What it does: Proposed a credal uncertainty set based on data-driven bulk calibration, utilizing the Huber error model to achieve a closed-form mean-maximization distributionally robust optimization objective in continuous unbounded space, and provided solvable LP/SOCP forms;
Bullet Trains: Parallelizing Training of Temporally Precise Spiking Neural Networks
Todd Morrill (Columbia University), Anthony M. Zador (Cold Spring Harbor Laboratory)
ClassificationComputational EfficiencySpiking Neural NetworkImageTime SeriesAudio
🎯 What it does: A event-driven SNN training method based on parallel associative scan and differentiable spiking time solver is studied, which can maintain hard reset dynamics on GPU and significantly accelerate training.
Butterworth as Attention: Anisotropic Spectral Gating for Pansharpening
Zhenggang Wang (Southwestern University of Finance and Economics), Tai-Xiang Jiang (Southwestern University of Finance and Economics)
RestorationSuper ResolutionTransformerImage
🎯 What it does: Propose an attention mechanism based on the Fourier domain Butterworth filter for pansharpening of remote sensing images;
BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception
Geng Li (Peking University), Yuxin Peng (Peking University)
RecognitionRetrievalOptimizationComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextMultimodality
🎯 What it does: Propose a visual search framework based on Bayesian optimization, named BVS, to enhance the fine-grained perception ability of multi-modal large language models in ultra-high resolution images for detecting tiny targets.
BYORn: Bootstrap Your Own Responses to Defend Large Vision-Language Models Against Backdoor Attacks
Ivan Sabolic, Sven Loncaric
Adversarial AttackTransformerSupervised Fine-TuningVision Language ModelContrastive LearningImageTextMultimodality
🎯 What it does: Propose BYORn, which detects poisoned samples using low likelihood judgment from pre-trained vision-language models, and dynamically generates clean responses with the model itself during training to achieve robust instruction fine-tuning.
Byte Pair Encoding for Efficient Time Series Forecasting
Leon Götz (Volkswagen AG), Leo Schwinn (Technical University of Munich)
Computational EfficiencyRepresentation LearningTransformerTime Series
🎯 What it does: Proposes a byte-pair encoding pattern tokenization method based on the frequent time series motivation, combined with conditional decoding post-processing to improve prediction accuracy.
C$^{2}$R: Cross-sample Consistency Regularization Mitigates Feature Splitting and Absorption in Sparse Autoencoders
Haoran Jin (University of Science and Technology of China), Defu Lian (University of Science and Technology of China)
Explainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelMixture of ExpertsAuto EncoderContrastive LearningText
🎯 What it does: This study investigates and proposes a cross-sample consistency regularization (C-R) to address the common issues of feature splitting and absorption in sparse autoencoders.
Cache Coherent Resampling for Efficient Test Time Scaling in LLM Reasoning via Adaptive Sequential Monte Carlo
KE WANG, Yongchao Huang (University of Aberdeen)
Computational EfficiencyTransformerLarge Language ModelPrompt EngineeringText
🎯 What it does: Propose a parallel particle inference framework Adaptive Sequential Monte Carlo (ASMC) and its cache coherence resampling technique to improve test-time scalability of large language model inference.
CacheEdit: Efficient Multi-round Image Editing via Adaptive Token-wise Reuse.
Jinxin Yu (SKLP, Institute of Computing Technology, Chinese Academy of Sciences), Ying Wang (SKLP, Institute of Computing Technology, Chinese Academy of Sciences)
GenerationComputational EfficiencyTransformerPrompt EngineeringDiffusion modelImage
🎯 What it does: Proposed a training-agnostic framework called CacheEdit, which utilizes the activation cache of unchanged regions in multi-round image editing to achieve efficient sparse computation for Diffusion Transformers;
CACR: Reinforcing Temporal Answer Grounding in Instructional Video via Candidate-Aware Causal Reasoning
Muge Qi (Peking University), Bin Li (Shenzhen Institute of Advanced Technology)
Representation LearningTransformerLarge Language ModelReinforcement LearningPrompt EngineeringVision Language ModelVideoTextMultimodalityRetrieval-Augmented Generation
🎯 What it does: Proposes the Candidate-Aware Causal Reasoning (CACR) framework, which first generates K high-quality candidate segments using a vision-language pre-trained model, and then identifies precise temporal answers within the candidate set through a reinforcement learning-driven causal reasoning module.
CADFit: Precise Mesh-to-CAD Program Generation with Hybrid Optimization
Ghadi Nehme (Massachusetts Institute of Technology), Faez Ahmed (Massachusetts Institute of Technology)
OptimizationDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningImagePoint CloudMesh
🎯 What it does: Proposes CADFit, a hybrid optimization framework that uses geometry-driven optimization and program execution to recover editable CAD construction sequences from meshes or images.
CAffNet: Hard Constraint-Affine Neural Networks
Yang Zhao (Northeastern University), Sze Zheng Yong (Northeastern University)
OptimizationReinforcement Learning from Human FeedbackTransformerAuto EncoderContrastive LearningTabularTime Series
🎯 What it does: This paper proposes a new framework called CAffNet, which embeds a hard constraint satisfaction mechanism into neural networks, enabling zero violations under any number of affine inequality constraints dependent on the input;
Calibrated Knowledge Aggregation in Bayesian Mixture-of-Experts for Continual VQA
Mahsa Mozaffari (Rochester Institute of Technology), Qi Yu (Rochester Institute of Technology)
RecognitionFederated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningMeta LearningTransformerSupervised Fine-TuningReinforcement LearningPrompt EngineeringMixture of ExpertsVision Language ModelContrastive LearningImageTextMultimodality
🎯 What it does: Construct a Bayesian Mixture-of-Experts (MoE) framework for continual learning in Visual Question Answering (VQA), where each task expert is trained with parameter-efficient LoRA adapters, learning a soft routing that aims to maximize expected VQA utility, and performing Bayesian aggregation of expert probability distributions during inference for calibration.
Calibrated Multimodal Representation Learning with Missing Modalities
Xiaohao Liu (National University Of Singapore), Tat-Seng Chua (National University Of Singapore)
ClassificationRetrievalRepresentation LearningTransformerDiffusion modelScore-based ModelFlow-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageVideoTextMultimodalityAudio
🎯 What it does: Propose a multi-modal representation learning framework called CalMRL, specifically designed to address the anchor shift problem caused by missing modalities.
Calibrated Preference Learning: The Case of Label Ranking
Santo M. A. R. Thies (Deutsches Forschungsinstitut fuer Kuenstliche Intelligenz), Eyke Hüllermeier (Deutsches Forschungsinstitut fuer Kuenstliche Intelligenz)
Recommendation SystemReinforcement Learning from Human FeedbackTransformerSupervised Fine-TuningReinforcement LearningPrompt EngineeringContrastive LearningTextBenchmark
🎯 What it does: This paper proposes a calibration framework for probabilistic label ranking (Probabilistic Label Ranking, ProLR), defining a multi-level calibration concept from full ranking to sub-ranking and Top-k ranking, and providing the theoretical relationships between them.
Calibrated Test-Time Guidance for Bayesian Inference
Daniel Geyfman (University of California), Stephan Mandt (University of California)
RestorationGenerationSuper ResolutionTransformerDiffusion modelScore-based ModelImageTextPoint CloudPhysics Related
🎯 What it does: This paper proposes a Calibrated Bayesian Guidance (CBG) method that can accurately sample from the true Bayesian posterior while preserving the prior knowledge of pre-trained diffusion models, correcting the bias estimation problem of the diffusion posterior probability in traditional test-time guidance methods.
Calibrating Conservatism for Scalable Oversight
William Overman (Stanford University), Mohsen Bayati (Stanford University)
Safty and PrivacyReinforcement Learning from Human FeedbackReinforcement LearningAgentic AIText
🎯 What it does: Propose a framework named Calibrated Collective Oversight (CCO), which can dynamically adjust the level of safety conservatism during the continuous decision-making process of agentic AI by aggregating multiple auxiliary evaluation functions and balancing compliance penalties with agent rewards.
Calibrating Decision Robustness via Inverse Conformal Risk Control
Wenbin Zhou (Carnegie Mellon University), Shixiang Zhu (Carnegie Mellon University)
OptimizationTabularBenchmark
🎯 What it does: This paper proposes an inverse consistency risk control framework to evaluate and calibrate the error coverage and regret risk of robust decisions.
Calibrating Generative Models to Distributional Constraints
Henry Smith, Brian L. Trippe (Stanford University)
GenerationData SynthesisProtein Structure PredictionTransformerSupervised Fine-TuningDiffusion modelScore-based ModelContrastive LearningImageTextBiomedical Data
🎯 What it does: Proposes a distribution layer constraint calibration method for generative models, formulated as fine-tuning model parameters under the constraint of minimizing KL divergence.
Calibrating Uncertainty for Zero-Shot Adversarial CLIP
Wenjing Lu (Shanghai Jiao Tong University), Qibin Zhao (RIKEN AIP)
ClassificationDomain AdaptationRepresentation LearningAdversarial AttackTransformerSupervised Fine-TuningContrastive LearningImageTextMultimodality
🎯 What it does: Study how to calibrate uncertainty under adversarial attacks in zero-shot CLIP, and propose an adversarial fine-tuning method based on the Dirichlet distribution.
CaliDist: Calibrating Large Language Models via Behavioral Robustness to Distraction
Mohammad Anas Jawad (University of Illinois Chicago), Cornelia Caragea (University of Illinois Chicago)
Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelPrompt EngineeringText
🎯 What it does: Propose CaliDist, a post-hoc calibration method that calibrates confidence by evaluating the behavioral robustness of large language models when semantic noise words are introduced.
CALM Before the STORM: Unlocking Native Reasoning for Optimization Modeling
Zhengyang Tang (Chinese University of Hong Kong), Benyou Wang (Chinese University of Hong Kong)
OptimizationKnowledge DistillationReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: A local error correction and adaptive training framework called CALM was implemented on large-scale reasoning models (LRM). The model first solves problems on its own, then inserts short prompts at the first error location for local correction. Subsequently, supervised fine-tuning and reinforcement learning were used to obtain the optimized modeling expert STORM 4B.