ICML 2026 Papers — Page 7
International Conference on Machine Learning · 6554 papers
Beyond Euclidean Clipping: Overcoming Exploration Collapse in LLM RL via Riemannian Isometric Policy Optimization
Zhicheng Cai (Tsinghua University), Hao Zhou (Tsinghua University)
OptimizationTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextChain-of-Thought
🎯 What it does: Propose a RL algorithm called RIPO based on Riemannian isometric shearing to address the exploration collapse problem caused by PPO-Clip in LLM training.
Beyond Euclidean Summaries: Online Change Point Detection for Distribution-Valued Data
Yingyan Zeng (University of Cincinnati), Xiaoyu Chen (University of Buffalo)
Anomaly DetectionOptimizationComputational EfficiencyTabularTime SeriesBiomedical DataBenchmark
🎯 What it does: Propose a distribution-valued online change point detection framework (IDD) based on the 2-Wasserstein space, which explicitly monitors distribution deformation by mapping the empirical distribution of each batch to the tangent space of the reference Fréchet barycenter.
Beyond Explicit Edges: Robust Reasoning over Noisy and Sparse Knowledge Graphs
Hang Gao (Rutgers University), Dimitris N. Metaxas (Rutgers University)
Recommendation SystemOptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningAdversarial AttackHyperparameter SearchData-Centric LearningRobotic IntelligenceMeta LearningDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackNeural Architecture SearchProtein Structure PredictionGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringTextGraphBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose the INSES framework, which utilizes LLM-guided navigation and embedding similarity expansion to dynamically overcome the limitations of missing or noisy edges in knowledge graphs, achieving multi-hop reasoning;
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring
Guanxu Chen (Shanghai Artificial Intelligence Laboratory), Dongrui Liu (Shanghai Artificial Intelligence Laboratory)
Safty and PrivacyExplainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningText
🎯 What it does: By performing contrastive learning and reorganization on the internal representations of LLMs, model transparency is improved, allowing monitors to directly identify inappropriate behaviors on internal features.
Beyond Extrapolation: Knowledge Utilization Paradigm with Bidirectional Inspiration for Time Series Forecasting
Liu Chong (Sichuan University), Ce Zhu (University of Electronic Science and Technology of China)
OptimizationData-Centric LearningTransformerContrastive LearningTabularTime SeriesBenchmarkFinance RelatedRetrieval-Augmented Generation
🎯 What it does: Propose a knowledge utilization paradigm called KUP-BI based on the historical-goal-subsequent chain of the training set, introducing an 'auxiliary stream' for subsequent evolution into the traditional unidirectional prediction framework, and enhancing long-term time series prediction through lightweight gated fusion.
Beyond First-order Asymptotics in Sequential Mean Testing
VIKAS DEEP, Shubhada Agrawal (Indian Institute of Science)
OptimizationTabularTime SeriesSequentialAgriculture Related
🎯 What it does: Study sequential mean testing for nonparametric bounded distributions, propose an optimal test based on KL inf and prove that its stopping time satisfies the central limit theorem, providing second-order approximation and confidence intervals.
Beyond Fixed Biases: Decoding the Role of Reasoning Uncertainty in MLLM Modality Conflicts
Zhuoran Zhang (Peking University), Lijie Hu (Mohamed bin Zayed University of Artificial Intelligence)
Autonomous DrivingExplainability and InterpretabilityTransformerSupervised Fine-TuningVision Language ModelContrastive LearningImageTextMultimodalityBenchmark
🎯 What it does: Propose a framework that decomposes the conflict resolution behavior of multimodal models into 'relative preference uncertainty' and 'inherent preference,' and measures uncertainty using entropy;
Beyond Gemini-3-Pro: Revisiting LLM Routing and Aggregation at Scale
Shengji Tang (Shanghai AI Lab), Wanli Ouyang (Shanghai AI Lab)
Federated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelMixture of ExpertsTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposed JiSi, a training-agnostic and scalable multi-LLM collaboration framework, which enables collaboration among different open-source LLMs through routing and aggregation, aiming to surpass closed-source models such as Gemini-3-Pro.
Beyond Generative Priors: Minority Sampling with JEPA-Guided Diffusion
Sol Park (Agency for Defense Development), Soobin Um (Kookmin University)
GenerationData SynthesisTransformerDiffusion modelScore-based ModelContrastive LearningImageText
🎯 What it does: Propose a minority sample sampling method based on world prior, JEPA-Guided Diffusion, which uses JEPA encoder to guide diffusion models to generate instances of low density and real-world scarcity.
Beyond Global Alignment: Fine-Grained Motion-Language Retrieval via Pyramidal Shapley-Taylor Learning
Hanmo Chen (Hangzhou Institute of Technology), Cheng Deng (Xidian University)
RetrievalKnowledge DistillationRepresentation LearningTransformerVision Language ModelContrastive LearningVideoTextMultimodality
🎯 What it does: Proposes a pyramid-based Shapley-Taylor learning framework (PST), achieving fine-grained motion-text retrieval from the joint level to the segment level and finally to the overall level.
Beyond Hamming: Query-Aware Decoding of Binary Cosine Sketches
DaeHun Nyang (Ewha Womans University)
RetrievalRepresentation LearningTransformerScore-based ModelContrastive LearningTextBenchmark
🎯 What it does: Propose a query-aware decoder named QA-COS, which is used to more accurately estimate the cosine similarity between a query and database vectors, even when only binary cosine projection bits are stored.
Beyond Heuristic Tuning: Power-Calibrated LLM Watermarking
Xiaopu Wang (Pennsylvania State University), Runze Li (Pennsylvania State University)
OptimizationExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelText
🎯 What it does: Proposes a statistical framework based on power analysis, quantifying the relationship between parameters (γ, δ) of Logit-based watermark and detection power and semantic distortion, and provides an optimization scheme targeting power/distortion.
Beyond Heuristics: Learnable Density Control for 3D Gaussian Splatting
Zhenhua Ning (Pengcheng Laboratory), Wenjie Pei (Harbin Institute of Technology)
OptimizationReinforcement LearningGaussian SplattingImagePoint Cloud
🎯 What it does: Proposes a learnable density control framework called LeGS based on reinforcement learning for 3D Gaussian Splatting scene reconstruction.
Beyond Independence: Learning Correlated Views for Variational Incomplete Multi-View Clustering
Zheming Xu (Beijing Jiaotong University), Michael Kampffmeyer (UiT Arctic University of Norway)
OptimizationRepresentation LearningData-Centric LearningTransformerMixture of ExpertsAuto EncoderContrastive LearningMultimodality
🎯 What it does: This paper proposes a novel variational framework called ACOVA for incomplete multi-view clustering, which can enhance clustering performance by learning cross-view correlations even in the absence of missing view information.
Beyond Independent Genes: Learning Module-Inductive Representations for Single-Cell Gene Perturbation Prediction
Jiafa Ruan (Zhejiang University), Yi Yang (Zhejiang University)
Explainability and InterpretabilityRepresentation LearningDrug DiscoveryGraph Neural NetworkTransformerFlow-based ModelContrastive LearningTabularBiomedical Data
🎯 What it does: Propose the scBIG framework, which utilizes modular representations to predict transcriptional responses to single-cell gene perturbations.
Beyond Instance-Level Self-Supervision in 3D Multi-Modal Medical Imaging
Tan Pan (Fudan University), Mahsa Baktashmotlagh (University of Queensland)
ClassificationSegmentationRepresentation LearningTransformerAuto EncoderContrastive LearningMultimodalityBiomedical DataMagnetic Resonance ImagingComputed TomographyPositron Emission Tomography
🎯 What it does: Propose the TACO framework, which utilizes cross-individual and cross-modal topological consistency for self-supervised pre-training on 3D multi-modal medical imaging.
Beyond Language Modeling: An Exploration of Multimodal Pretraining
Shengbang Tong (FAIR, Meta), Saining Xie (New York University)
GenerationData SynthesisRepresentation LearningTransformerMixture of ExpertsVision Language ModelDiffusion modelAuto EncoderWorld ModelImageVideoTextMultimodality
🎯 What it does: This paper starts from scratch to unify multi-modal pre-training, using the Transfusion framework to simultaneously learn the next-token prediction for text and the diffusion prediction for vision, and systematically evaluates different visual representations, data combinations, architectural designs, and scaling strategies.
Beyond Literal Translation: Evaluating Cultural Effectiveness in Social Media UGC
Linjuan Wu (Zhejiang University), Weiming Lu (Zhejiang University)
Data-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningTextBenchmark
🎯 What it does: Constructed the Chinese-to-English UGC translation benchmark CULTURE-MT, and proposed a cultural effectiveness evaluation metric and an automatic evaluator called JUDGER
Beyond Log Likelihood: Probability-Based Objectives for Supervised Fine-Tuning across the Model Capability Continuum
Gaotang Li (University of Illinois Urbana Champaign), Hanghang Tong (University of Illinois Urbana Champaign)
OptimizationComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringContrastive LearningTextMultimodalityBenchmark
🎯 What it does: Studied probability-based objectives beyond negative log-likelihood (NLL) in the post-training (SFT) of large language models, and systematically evaluated their effectiveness across different model capability ranges.
Beyond Logits: Coherent Hallucination Mitigation via Attention Contrastive Decoding
Yujia Chen (University of Science and Technology of China), Tianzhu Zhang (University of Science and Technology of China)
GenerationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodality
🎯 What it does: This paper studies the hallucination problem that occurs during the generation process of large-scale vision-language models (LVLMs), and proposes Attention Contrastive Decoding (ACD) as a training-free plug-in. It transfers the operations of traditional contrastive decoding from the logit layer to the attention layer, and further introduces the Adaptive Subtraction Strategy (ASS) to achieve position-adaptive suppression.
Beyond Logits: Metastable Latent Dynamics for Sample-Efficient Best-of-N Selection in LLMs
Xinrong Li (Tsinghua University), Shangqi Guo (Tsinghua University)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose an unsupervised Latent Velocity Entropy (LVE) metric, which ranks candidate inference trajectories based on the aggregation degree of hidden state updates during the generation process, thereby improving the effectiveness of Best-of-N selection.
Beyond Looking Up, Try Looking Around: Harmonizing Global Structure and Local Consistency in Optimal Transport for Short Text Clustering
Zhihao Yao (Harbin Engineering University), Bo Li (Harbin Engineering University)
OptimizationRepresentation LearningData-Centric LearningTransformerContrastive LearningTextMultimodality
🎯 What it does: Propose an end-to-end short text clustering framework that utilizes consistency-aware adaptive optimal transport (CAOT) to generate reliable pseudo labels, and constructs semantic similarity between samples through an instance-level attention network, achieving the unification of global structure and local consistency.
Beyond Magnitude: Scale-Invariant Evidential Fusion for Multi-View Classification
Wei Liu (Tongji University), Xiaodong Yue (Shanghai University)
ClassificationDiffusion modelScore-based ModelContrastive LearningImageBiomedical DataBenchmarkStochastic Differential Equation
🎯 What it does: Propose Scale-Invariant Evidential Fusion (SAEF) to address the issue of credibility distortion caused by scale mismatch in multi-view fusion.
Beyond Majority Voting: LLM Aggregation by Leveraging Higher-Order Information
Rui Ai (Massachusetts Institute of Technology), Haifeng Xu (University of Chicago)
OptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningData-Centric LearningTransformerLarge Language ModelAgentic AIMixture of ExpertsContrastive LearningTextRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Studies the answer aggregation problem in multi-model LLMs, proposing aggregation algorithms Optimal Weight (OW) and Inverse Surprisingly Popular (ISP) that utilize first-order (accuracy) and second-order (answer relevance) information, along with theoretical optimality proofs and empirical evaluations.
Beyond Majority Voting: Self-Reflective Test-Time Reinforcement Learning for LLM Reasoning
Sitong Wu (Chinese University of Hong Kong), Jiaya Jia (Chinese University of Hong Kong)
Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Designed and implemented a self-reflective test-time reinforcement learning framework (SR-TTRL), which enhances the reasoning performance of large language models (LLMs) in an unsupervised manner by generating high-quality pseudo-labels through sampling multiple reasoning trajectories, constructing a candidate pool, abstract compression, and model self-checking.
Beyond Mode Collapse: Distribution Matching for Diverse Reasoning
Xiaozhe Li (Tongji University), Kai Chen (Shanghai AI Lab)
OptimizationReinforcement LearningPrompt EngineeringDiffusion modelScore-based ModelContrastive LearningTextMultimodalityBenchmark
🎯 What it does: Propose Distribution-Matching Policy Optimization (DMPO), which addresses mode collapse in RL by approximating the forward KL divergence at the group level; construct a multi-modal NP-Bench evaluation framework; and verify the superiority of DMPO in multi-modal, text, and mathematical reasoning tasks.
Beyond Model Base Retrieval: Weaving Knowledge to Master Fine-grained Neural Network Design
Jialiang Wang (Hong Kong University of Science and Technology), Xiaofang Zhou (Hong Kong University of Science and Technology)
Neural Architecture SearchGraph Neural NetworkImageGraphTabularRetrieval-Augmented Generation
🎯 What it does: Designed a retrieval-enhanced model improvement framework called M-DESIGN, which can quickly approach the optimal neural network architecture through fine-grained structural modifications under a limited evaluation budget.
Beyond Model Ranking: Predictability-Aligned Evaluation for Time Series Forecasting
Wanjin Feng (Tsinghua University), Yong Li (Tsinghua University)
Anomaly DetectionExplainability and InterpretabilityComputational EfficiencyRecurrent Neural NetworkTransformerSupervised Fine-TuningAuto EncoderContrastive LearningTabularTime SeriesBenchmark
🎯 What it does: Propose a predictability assessment framework based on spectral coherence called Spectral Coherence Predictability (SCP), and introduce Linear Utilization Ratio (LUR) to diagnose how models utilize linearly predictable information across different frequency intervals.
Beyond Next-Token Alignment: Distilling Multimodal Large Language Models via Token Interactions
Lin Chen (MAIS, Institute of Automation, Chinese Academy of Sciences), Shiming Xiang (MAIS, Institute of Automation, Chinese Academy of Sciences)
Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningVision Language ModelContrastive LearningImageTextMultimodalityBenchmark
🎯 What it does: Proposes an Align‑TI framework based on token interaction for knowledge distillation, aiming to compress large-scale multi-modal language models (MLLM) into parameter-efficient small models.
Beyond Normalization: Rethinking the Partition Function as a Difficulty Scheduler for RLVR
Dohyung Kim (Seoul National University), Kyomin Jung (Seoul National University)
AI Code AssistantTransformerLarge Language ModelReinforcement LearningPrompt EngineeringFlow-based ModelGenerative Adversarial NetworkText
🎯 What it does: This paper proposes a post-training framework for LLMs called PACED-RL, which reinterprets the partition function of GFlowNet as an online accuracy estimate for each prompt question, thereby achieving adaptive prompt selection based on difficulty and accuracy error-priority replay.
Beyond Perplexity: UTF-8 Validity in Byte-aware Language Models
Sangwhan Moon (Google LLC), Naoaki Okazaki (Institute of Science Tokyo)
GenerationData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextSequentialRetrieval-Augmented Generation
🎯 What it does: Trained a 355M-parameter GPT-2-style byte-level language model, and systematically measured its UTF-8 generation reliability through a custom Level0 and Level1 evaluation protocol.
Beyond Pixel Histories: World Models with Persistent 3D State
Samuel Garcin (University of Edinburgh), Jiang Bian (Microsoft Research)
GenerationData SynthesisRobotic IntelligenceTransformerDiffusion modelScore-based ModelFlow-based ModelRectified FlowAuto EncoderWorld ModelVideoPoint Cloud
🎯 What it does: Propose the PERSIST framework, achieving continuous 3D latent world states for interactive world model generation
Beyond Pixels: Mining Compressed Domain Artifacts for Efficient AI-Generated Video Detection
Anran Zhu (Xi'an Jiaotong University), Chao Shen (Xi'an Jiaotong University)
Image TranslationRestorationAnomaly DetectionComputational EfficiencyConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningOptical FlowVideo
🎯 What it does: Propose a AI-generated video detection framework called STREAM based on compressed domain features, which directly utilizes I-frames, motion vectors, and residuals for detection without requiring full decoding;
Beyond Point Predictions: Manifold Expansion and Dual Alignment for Robust Time Series Distillation
Junyao Hong (Southeast University), Youyong Kong (Southeast University)
Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerSupervised Fine-TuningTime Series
🎯 What it does: Propose the Dynamic Structural Distillation (DSD) framework and design LMP-Net to achieve lightweight long-term sequence prediction by expanding through a high-dimensional manifold.
Beyond Point-wise Neural Collapse: A Topology-Aware Hierarchical Classifier for Class-Incremental Learning
HuiYu Yi, Furao Shen (Nanjing University)
ClassificationFederated LearningComputational EfficiencyRepresentation LearningMeta LearningGraph Neural NetworkAuto EncoderContrastive LearningGaussian SplattingOptical FlowImageStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Proposed HC-SOINN, a topology-aware hierarchical classifier, combined with the STAR module to achieve point-level adaptation for nonlinear feature drift;
Beyond Policy Training: Recursive Solution Search from Unannotated Videos
Lipeng Wan (Xi'an Jiaotong University), Xuguang Lan (Xi'an Jiaotong University)
Recurrent Neural NetworkTransformerReinforcement LearningDiffusion modelScore-based ModelAuto EncoderContrastive LearningVideoChain-of-Thought
🎯 What it does: Proposed a policy-free recursive search framework called PFR-Search, which directly recovers feasible solutions from unannotated videos.
Beyond Prediction: Tail-Aware Scheduling for LLM Inference
Yueying Li (Cornell University), Udit Gupta (Cornell University)
OptimizationComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation
🎯 What it does: Propose a prediction-agnostic, soft-priority-based LLM inference scheduling framework called UNIBOOST, which combines KV cache-aware preemption strategies.
Beyond Problem Solving: UOJ-Bench for Evaluating Code Generation, Hacking, and Repair in Competitive Programming
Tingqiang Xu (Tsinghua University), Kaifeng Lyu (Tsinghua University)
AI Code AssistantTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Constructed the UOJ-Bench benchmark to evaluate the code generation, code hacking (generating test cases that make opponents' code fail), and code repair (providing minimal patches) capabilities of large language models (LLMs) in competitive programming, and validated it directly using the native judging interface of UOJ.
Beyond Procedure: Substantive Fairness in Conformal Prediction
Pengqi Liu (McGill University), Jesse C. Cresswell (Layer 6 AI)
Federated LearningExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelDiffusion modelScore-based ModelFlow-based ModelRectified FlowNeural Radiance FieldAuto EncoderGenerative Adversarial NetworkContrastive LearningImageTextMultimodalityTabularAudio
🎯 What it does: This paper investigates the fairness of conformal prediction (CP) throughout the entire decision-making process, proposing an evaluation framework that transforms procedural fairness into substantive fairness, verifying and recommending CP methods that can simultaneously achieve efficiency and fairness.
Beyond Rational Illusion: Behaviorally Realistic Strategic Classification
Xinpeng Lv (National University of Defense Technology), Haotian Wang (National University of Defense Technology)
ClassificationOptimizationExplainability and InterpretabilityReinforcement LearningContrastive LearningTabularFinance Related
🎯 What it does: Studied strategic classification under behavioral realism, proposing the Pro-SF framework to improve traditional rationality assumptions.
Beyond Reactivity: Proactive Adaptive Conformal Inference for Online LLM Factuality
Xinyu Liu (Michigan State University), Jun Wu (Michigan State University)
Domain AdaptationFederated LearningExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelTextTabular
🎯 What it does: Propose an online LLM factuality assurance framework called PACE based on active adaptive consistency reasoning.
Beyond ReLU: Bifurcation, Oversmoothing, and Topological Priors
Erkan Turan (Ecole Polytechnique), Maks Ovsjanikov (Ecole Polytechnique)
ClassificationGraph Neural NetworkContrastive LearningGraph
🎯 What it does: This paper reinterprets the oversmoothing problem in GNNs through dynamic bifurcation theory, and demonstrates that by using activation functions with stable cubic nonlinearities (such as Sine, tanh) and bifurcation-aware initialization, feature homogenization can be avoided in deep GNNs (up to 64 layers), leading to the emergence of stable non-uniform patterns. Subsequently, polynomial spectral filters are used to further control the topological patterns selected by the model, and the effectiveness of this method is validated on multiple node classification benchmarks.
Beyond Rewards in RL for Cyber Defence
Elizabeth Bates (Alan Turing Institute), Vasilios Mavroudis (Alan Turing Institute)
Convolutional Neural NetworkRecurrent Neural NetworkTransformerReinforcement LearningScore-based ModelContrastive LearningGraphTabular
🎯 What it does: Studied the impact of sparse and dense rewards on autonomous network defense agents, and proposed a Ground Truth scoring method based on real network states.
Beyond Sample-Level Forgetting: Improving Reliability in Multimodal Unlearning
Jianzhou Wang (Hohai University), Jun Liu (Lancaster University)
Safty and PrivacyRepresentation LearningTransformerContrastive LearningImageTextMultimodality
🎯 What it does: This study proposes a multi-modal model forgetting method based on a causal perspective, achieving fine-grained control over forgetting through knowledge decoupling and contrastive semantic editing.
Beyond Scalars: Evaluating and Understanding LLM Reasoning via Geometric Progress and Stability
Xinyan Jiang (Mohamed bin Zayed University of Artificial Intelligence), Lijie Hu (Mohamed bin Zayed University of Artificial Intelligence)
Explainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelReinforcement LearningContrastive LearningTextBenchmarkChain-of-Thought
🎯 What it does: This paper proposes the TRACED framework, which decomposes the displacement and curvature of LLM reasoning trajectories through geometric kinematics to evaluate reasoning quality.
Beyond Single Embedding: Modeling User Preferences as Distribution in Federated Recommendation
Chunxu Zhang (Jilin University), Bo Yang (Jilin University)
Recommendation SystemFederated LearningTransformerDiffusion modelTabularSequential
🎯 What it does: In the federated recommendation scenario, FedDistRec is proposed, which uses diffusion generative models to model user preferences as distributions rather than single-point representations, and obtains the final user embeddings through diverse sample generation and aggregation, thereby improving ranking stability and robustness.
Beyond Single-View Indexing: Structure-Aware Multi-View Retrieval for Knowledge-Based VQA
Hao Wang (HKUST(GZ)), Lei Chen (HKUST(GZ))
RetrievalGraph Neural NetworkTransformerContrastive LearningImageTextMultimodalityRetrieval-Augmented Generation
🎯 What it does: Propose a no-training, structure-aware multi-view retrieval framework called SCAR, used for the coarse-grained retrieval phase in knowledge-driven visual question answering
Beyond Softmax: A Natural Parameterization for Categorical Random Variables
Alessandro Manenti (Università della Svizzera italiana), Cesare Alippi (Università della Svizzera italiana)
OptimizationReinforcement LearningContrastive LearningImageGraphTabular
🎯 What it does: Propose a new method for parameterizing category distributions (catnat), which replaces the traditional softmax with a hierarchical binary decision to improve gradient descent optimization.
Beyond Static Allocation: Dynamic Sensitivity-Aware Fine-Tuning for Vision Transformers
Yuanyang Cao (Xi'an Jiaotong University), Jianji Wang (Xi'an Jiaotong University)
ClassificationObject DetectionSegmentationComputational EfficiencyRepresentation LearningConvolutional Neural NetworkTransformerSupervised Fine-TuningAuto EncoderContrastive LearningImage
🎯 What it does: This paper proposes a dynamic parameter-efficient fine-tuning framework called Dynamic Adaptive Fine-Tuning (DAF), which periodically perceives, decides, and reconstructs the trainable structure, allowing Vision Transformers to adaptively allocate limited parameter budgets on downstream tasks.
Beyond Static Endpoints: Tool Programs as an Interface for Flexible Agentic Web Services
Mugeng Liu (Peking University), Yun Ma (Peking University)
OptimizationFederated LearningComputational EfficiencyTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose TOOLPRO, a system that replaces traditional static endpoint calls with executable tool programs with effect types, allowing LLM agents to uniformly execute multi-step workflows on the server side.
Beyond Static Pipelines: Learning Dynamic Workflows for Text-to-SQL
Yihan Wang (Renmin University of China), Wei Xu (Renmin University of China)
TransformerSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextTabularBenchmark
🎯 What it does: This paper proposes a dynamic workflow construction framework called SquRL based on reinforcement learning, for the text-to-SQL (Text-to-SQL) task;
Beyond Structural Symmetries: Linear Mode Connectivity via Neuron Identifiability
Vincent Bürgin, Stefanie Jegelka (Technical University Of Munich)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningImageTabular
🎯 What it does: This paper studies the symmetry breaking and linear mode connectivity achieved through neuron identifiability in deep networks.
Beyond Sunk Costs: Boosting LLM Pre-training Efficiency via Orthogonal Growth of Mixture-of-Experts
Ruizhe Wang (University of Science and Technology of China), Yeyun Gong (Microsoft Research Asia)
Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsText
🎯 What it does: Propose a 'recycling' strategy that utilizes a pre-trained Mixture-of-Experts (MoE) model, expanding the converged checkpoints in depth (layer replication) and width (expert replication) through 'orthogonal growth', while preserving the knowledge already learned by the model, thereby achieving a larger and stronger LLM under a limited additional computational budget.
Beyond Temperature: Hyperfitting as a Late-Stage Geometric Expansion
Meimingwei Li (LMU Munich), Christian Heumann (LMU Munich)
GenerationData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringText
🎯 What it does: This paper reveals that the 'overfitting' process, where LLMs are fine-tuned to extremely low training error, significantly enhances the diversity of generated text and reduces repetition rates, which is attributed to an adaptive vocabulary ranking reordering mechanism;
Beyond Test-Time Memory: State-Space Optimal Control for LLM Reasoning
Peihao Wang (University of Texas at Austin), Rene Vidal
OptimizationComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningTextChain-of-Thought
🎯 What it does: Introduce a Test-Time Control (TTC) layer into language models, treating reasoning as a linear quadratic regulation (LQR) problem, and planning over internal representations during inference, thereby achieving System-2 style long-term reasoning.
Beyond Text-to-SQL: Can LLMs Really Debug Enterprise ETL SQL?
Jing Ye (Independent Researcher), Xing Chen (Bytedance Inc.)
AI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningAgentic AIPrompt EngineeringTextTabularBenchmark
🎯 What it does: Proposed the Squirrel Benchmark, a benchmark for enterprise-level SQL debugging, and evaluated the performance of multiple LLMs.
Beyond the Bellman Recursion: A Pontryagin-Guided Framework for Non-Exponential Discounting
Hojin Ko (Sungkyunkwan University), Jeonggyu Huh (Sungkyunkwan University)
OptimizationReinforcement LearningTabularTime SeriesBenchmarkStochastic Differential Equation
🎯 What it does: Proposed a direct policy optimization framework called PG-DPO based on the Pontryagin maximum principle, for handling continuous time control problems under non-exponential discounting.
Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents
Wonjoong Kim (Korea Advanced Institute of Science and Technology), Chanyoung Park (Korea Advanced Institute of Science and Technology)
Explainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposed a no-reference trajectory evaluation framework called TRACE, designed to comprehensively evaluate the reasoning trajectories of tool-enhanced LLM agents.
Beyond the Proxy: Trajectory-Distilled Guidance for Offline GFlowNet Training
Ruishuo Chen (Tsinghua University), Longbo Huang (Tsinghua University)
OptimizationDrug DiscoveryGraph Neural NetworkReinforcement LearningDiffusion modelGenerative Adversarial NetworkContrastive LearningGraphTabularSequential
🎯 What it does: Propose a model-free offline GFlowNet training framework called TD-GFN, which uses inverse reinforcement learning to extract edge rewards from trajectories for DAG pruning and priority reverse sampling;
Beyond Theorem Proving: Formulation, Framework and Benchmark for Formal Problem-Solving
Qi Liu (Shanghai Jiao Tong University), Junchi Yan (Shanghai Jiao Tong University)
AI Code AssistantTransformerLarge Language ModelPrompt EngineeringDiffusion modelTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This paper proposes the Formal Problem-Solving (FPS) framework, unifying problem-solving and proof processes as dependent pairs in Lean 4, and further designs Deductive FPS (D-FPS) to decompose solving into forward deduction and backward verification. It also introduces the Restricted Propositional Equivalence (RPE) evaluation metric and constructs three benchmarks: FormalMath500, MiniF2F-Solving, and PutnamBench-Solving.
Beyond Token-level Supervision: Unlocking the Potential of Decoding-based Regression via Reinforcement Learning
Ming Chen (Nanjing University), Chao Qian (Nanjing University)
OptimizationKnowledge DistillationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringScore-based ModelContrastive LearningTextTabularSequentialBenchmark
🎯 What it does: This paper proposes a framework called GenRe 2, which combines decoding regression with reinforcement learning, to address the mismatch between traditional token-level supervision based on cross-entropy and continuous numerical targets.
Beyond Tokens: Enhancing RTL Quality Estimation via Structural Graph Learning
Yi Liu (Chinese University of Hong Kong), Qiang Xu (Chinese University of Hong Kong)
OptimizationKnowledge DistillationRepresentation LearningGraph Neural NetworkTransformerLarge Language ModelContrastive LearningTextGraph
🎯 What it does: Proposes a structured self-supervised learning framework called StructRTL based on control data flow graphs (CDFG), aimed at improving RTL design quality estimation;
Beyond Trajectory-Level Attribution: Graph-Based Credit Assignment for Agentic Reinforcement Learning
Xin Cheng (Nanyang Technological University), Bo An (Nanyang Technological University)
OptimizationExplainability and InterpretabilityGraph Neural NetworkTransformerReinforcement LearningAgentic AIPrompt EngineeringVision Language ModelImageTextGraphBenchmark
🎯 What it does: Construct a unified state transition graph, and achieve fine-grained step-level reward allocation based on the graph's advantage estimation, thereby improving the reinforcement learning training of LLM/VLM agents.
Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning
Liang Chen (Chinese University of Hong Kong), Kam-Fai Wong (Chinese University of Hong Kong)
OptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningContrastive LearningTextBenchmark
🎯 What it does: Proposes a new training framework called BRIDGE, which leverages the idea of second-order optimization to enable supervised fine-tuning (SFT) to actively supervise reinforcement learning (RL) in order to enhance the reasoning capabilities of large language models (LLMs).
Beyond Unidirectional Bias: Reciprocal Perspective Calibration in Scene Graph Generation
Haifeng Zhao (Anhui University), Dengdi Sun (Anhui University)
GenerationKnowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningImageTextMultimodality
🎯 What it does: This paper addresses the unidirectional bias problem in scene graph generation by proposing the Mutual-Perspective Inverse Relations (MPIR) principle, and achieves logical consistency from a bidirectional perspective through the Reciprocal Perspective Calibration (RPC) framework.
Beyond VLM-Based Rewards: Diffusion-Native Latent Reward Modeling
Gongye Liu (Hong Kong University of Science and Technology), Wenhan Luo (Hong Kong University of Science and Technology)
GenerationOptimizationComputational EfficiencyReinforcement Learning from Human FeedbackTransformerReinforcement LearningPrompt EngineeringMixture of ExpertsVision Language ModelDiffusion modelScore-based ModelContrastive LearningImageMultimodality
🎯 What it does: Proposed DiNa-LRM, a diffusion model-based latent space reward model that performs preference learning directly in noisy states using noise-calibrated Thurstone likelihood;
BFCL Audio: An Audio Function Calling Evaluation for Large Language Models
Huanzhi Mao (University of California, Berkeley), Joseph E. Gonzalez (University of California, Berkeley)
TransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-ThoughtAudio
🎯 What it does: Proposed the BFCL Audio benchmark to evaluate the ability of large language models to perform tool calls under audio input.
BFTS: Thompson Sampling with Bayesian Additive Regression Trees
Ruizhe Deng (National University of Singapore), Yan Shuo Tan (National University of Singapore)
OptimizationReinforcement LearningTabularTime Series
🎯 What it does: Proposes Bayesian Forest Thompson Sampling (BFTS), which utilizes the full posterior of Bayesian Additive Regression Trees (BART) for Thompson Sampling in contextual bandits, achieving exploration and exploitation in nonlinear reward environments;
Bi-Anchor Interpolation Solver for Accelerating Generative Modeling
Hongxu chen, Long Chen (Hong Kong University of Science and Technology)
GenerationTransformerDiffusion modelFlow-based ModelImageOrdinary Differential Equation
🎯 What it does: Propose the Bi-Anchor Interpolation Solver (BA-solver), which significantly reduces the number of function evaluations (NFE) in flow matching models and improves sampling speed through a lightweight SideNet and dual-anchor interpolation.
Bias in Zeroth-Order Normal Estimation for Decision-Based Attacks
Feiyang Wang (Beijing University of Posts and Telecommunications), Ivor Tsang (Agency for Science, Technology and Research)
Adversarial AttackConvolutional Neural NetworkTransformerScore-based ModelContrastive LearningImageStochastic Differential Equation
🎯 What it does: This paper studies the bias caused by input sensitivity non-uniformity in zeroth-order decision boundary normal estimation in decision-based black-box attacks, and proposes a sensitivity-aware rescaling (SAR) algorithm based on this.
Bias-Spectrum Neural Processes for Parametric PDEs: Architecture Priors Meet PDE Constraints
Hui Li (Beijing Jiaotong University), Liping Jing (Beijing Jiaotong University)
Convolutional Neural NetworkContrastive LearningTabularTime SeriesPhysics RelatedStochastic Differential Equation
🎯 What it does: Proposed a Bias‑Spectrum Neural Processes (BSNP) framework for building fast and reliable surrogate models for parameterized PDEs under sparse and irregular observations.
Biased Generalization in Diffusion Models
Jerome Garnier-Brun (Bocconi University), Luca Saglietti (Bocconi University)
GenerationData SynthesisConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelImageSequential
🎯 What it does: Studied the 'bias generalization' phenomenon in diffusion models during the training process, where the model generates biased samples from the training set even when the test loss is still decreasing.
Biases in the Blind Spot: Detecting What LLMs Fail to Mention
Iván Arcuschin (Poseidon Research), Oana-Maria Camburu (Imperial College London)
Explainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposed a fully automated, black-box pipeline for detecting implicit biases in large language models that are not verbalized during chain-of-thought reasoning.
BiCrossNet with Decoupled Dual Generators: A Parameter‑Efficient and Generalizable Few‑Shot Custom Gesture Recognition Framework
Chunyang Yu (OPPO Research Institute), Shenyue Wang (OPPO Research Institute)
RecognitionConvolutional Neural NetworkTransformerDiffusion modelContrastive LearningTime Series
🎯 What it does: Proposed a BiCrossNet framework based on IMU and a decoupled dual generator data augmentation scheme to achieve few-shot custom gesture recognition;
Bilevel Optimization over Saddle Points of Zero-Sum Markov Games
Zihao Zheng (Chinese University of Hong Kong), Songtao Lu (Chinese University of Hong Kong)
OptimizationReinforcement LearningScore-based ModelContrastive Learning
🎯 What it does: Proposes a two-layer reinforcement learning framework with regularized zero-sum Markov games at the lower level, and designs the PANDA algorithm.
Bilinear Bandits with Partially Observable Features
Wooseong Cho (Seoul National University), Min-hwan Oh (Seoul National University)
Recommendation SystemReinforcement LearningContrastive LearningTabular
🎯 What it does: Propose the BiRoLF algorithm to address the bandit problem with a bilinear reward model where features on both sides are only partially observable, directly utilizing the bilinear structure without performing Kronecker linearization.
Bimodal masked language modeling for bulk RNA-seq and DNA methylation representation learning
Maxence Gélard (InstaDeep), Paul-Henry Cournède (InstaDeep)
Representation LearningConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningTabularBiomedical Data
🎯 What it does: Designed and trained the MOJO model, using a dual-modal masked language model to learn joint representations of bulk RNA-seq and DNA methylation, followed by fine-tuning on tasks of cancer type classification and survival prediction.
Bio-Inspired Self-Supervised Learning for Wrist-worn Accelerometer Data
Prithviraj Tarale (University of Massachusetts), Sunghoon Lee
RecognitionConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningTime SeriesBiomedical DataBenchmark
🎯 What it does: Propose a tokenization of motion segments (zero-crossing segments) based on the theory of sub-motions, using a Transformer Encoder for self-supervised pre-training through masked reconstruction, and transfer the pre-trained model to multi-task activity recognition on wrist-worn accelerometers.
Bio-Vision-Inspired Spiking Neural Networks for Object Detection with Event Cameras
Dongyang Ma (Peking University), Yonghong Tian (Peking University)
Object DetectionSpiking Neural NetworkSupervised Fine-TuningContrastive LearningImageVideo
🎯 What it does: A biovision-inspired event camera object detection framework is proposed, including STATNF denoising neurons, Events-to-Spikes representation, and bidirectional multi-scale Spiking networks, achieving end-to-end processing from event streams to object detection.
Bioacoustic Geolocation: Species Sounds as Geographic Signals
Mustafa Chasmai (University of Massachusetts Amherst), Grant Van Horn
RetrievalRepresentation LearningData-Centric LearningConvolutional Neural NetworkTransformerMixture of ExpertsContrastive LearningMultimodalityBenchmarkRetrieval-Augmented GenerationAudio
🎯 What it does: Proposed a global-scale audio geolocation method, utilizing bioacoustic (species calls) and location embedding retrieval techniques, developed the AG-CLIP model, and constructed a multi-species audio benchmark XCDC along with spatiotemporal aggregation strategies.
BioAgent Bench: An AI Agent Evaluation Suite for Bioinformatics
Dionizije Fa (Entropic), Mateo Čupić (Entropic)
Drug DiscoveryTransformerLarge Language ModelAgentic AIPrompt EngineeringTabularBiomedical DataBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Built the BioAgent Bench — a benchmarking suite containing 10 handpicked end-to-end bioinformatics tasks (such as RNA-seq, variant detection, metagenomics, etc.), and implemented automated evaluation through an LLM judge.
BIOARC: Discovering Optimal Neural Architectures for Biological Foundation Models
Yi Fang (Virginia Tech), Xuan Wang (Virginia Tech)
Neural Architecture SearchProtein Structure PredictionConvolutional Neural NetworkTransformerAgentic AIContrastive LearningBiomedical Data
🎯 What it does: Propose the BIOARC framework, which automatically discovers efficient biological foundation model architectures through NAS in a heterogeneous search space.
BioDynaSpec: Harmonic-Guided Spatio-Spectral Autoregressive Diffusion for Protein Dynamics Generation
Mujie Lin (Peking University), Jie Chen (Peking University)
GenerationProtein Structure PredictionTransformerDiffusion modelAuto EncoderTime SeriesBiomedical Data
🎯 What it does: This study proposes a generative framework based on frequency-domain autoregressive diffusion, capable of generating high-quality protein dynamics trajectories over long time scales;
BioFormer: Rethinking Cross-Subject Generalization via Spectral Structural Alignment in Biomedical Time-Series
Guikang Du (Nankai University), Jin Zhang (Nankai University)
Domain AdaptationAnomaly DetectionConvolutional Neural NetworkTransformerContrastive LearningTime SeriesBiomedical DataElectrocardiogramReview/Survey PaperBenchmark
🎯 What it does: This paper proposes a new cross-subject generalization framework called BioFormer, which focuses on suppressing subject-specific variations in medical time series through spectral structure alignment;
Biologically plausible heavy-tailed connectivity enhances generalizations on cognitive tasks in recurrent neural networks
Zhe Jiao (Northwestern Polytechnical University), Shanglin Zhou (Fudan University)
OptimizationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackRecurrent Neural NetworkSequential
🎯 What it does: A new DSco-SGD optimization algorithm based on optimal transport has been developed, which can train RNNs to perform various cognitive tasks while maintaining the Dale principle and heavy-tailed connection statistics.
BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science
Yuyang Liu (Peking University), Yonghong Tian (Peking University)
Autonomous DrivingDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringTextBiomedical DataBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Built BioProBench, a corpus containing 27,000 professional biological experimental protocols and its derived 556,000+ multi-task instances, and proposed a multi-dimensional evaluation metric for procedural reasoning;
BioToken and BioFM – Biologically-Informed Tokenization Enables Accurate and Efficient Genomic Foundation Models
Aleksandr Medvedev (M42), Shadab Khan (M42)
Computational EfficiencyRepresentation LearningDrug DiscoveryProtein Structure PredictionTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningBiomedical DataBenchmark
🎯 What it does: Proposed a tokenization method based on biological annotations called BioToken, and built a lightweight genomic foundation model called BioFM based on this method;
Bipartite Graph Attention-based Clustering for Large-scale scRNA-seq Data
Zhuomin Liang (Shanxi University), Xian Yang (University of Manchester)
Computational EfficiencyRepresentation LearningGraph Neural NetworkTransformerAuto EncoderContrastive LearningBiomedical Data
🎯 What it does: Proposed a Transformer model called BGFormer based on bilateral graph attention for clustering large-scale single-cell RNA sequencing data.
BiRQA: Bidirectional Robust Quality Assessment for Images
Aleksandr Gushchin (Trusted AI Research Center RAS), Anastasia Antsiferova (Trusted AI Research Center RAS)
Computational EfficiencyAdversarial AttackConvolutional Neural NetworkContrastive LearningImage
🎯 What it does: Propose a lightweight full-reference image quality assessment model called BiRQA, and design Anchored Adversarial Training (AAT) on this basis to enhance robustness.
BIT-LLM: Brain Instruction Tuned LLM with persistent Cross-Attention for fMRI-to-Text Decoding
Sunghwan Lee (Korea University), Jong-Hwan Lee (Korea University)
GenerationRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningContrastive LearningTextMultimodalityBiomedical DataMagnetic Resonance ImagingRetrieval-Augmented Generation
🎯 What it does: Designed and implemented BIT-LLM, a self-regressive text generation model that treats fMRI as the first modality through persistent cross-attention, and achieved cross-subject fMRI-to-text decoding using a three-stage training pipeline (multi-modal contrastive pre-training, supervised fine-tuning, and reinforcement learning fine-tuning).
BiTrajDiff: Bidirectional Trajectory Generation with Diffusion Models for Offline Reinforcement Learning
Yunpeng Qing (Zhejiang University), Changqing Zou (Zhejiang Lab)
TransformerReinforcement LearningDiffusion modelTabularTime SeriesSequentialBenchmark
🎯 What it does: Proposed a bidirectional trajectory augmentation framework called BiTrajDiff, which generates forward and backward trajectories on shared anchor states using a bidirectional diffusion model, and concatenates them to produce global trajectories across behavior patterns.
Bits That Count: Quantifying and Predicting Capabilities of Language Models
Elizabeth Donoway (University of California, Berkeley), Jan Leike (Anthropic)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringContrastive LearningText
🎯 What it does: Investigated and quantified whether language models improve performance during fine-tuning by activating existing capabilities (elicitation) or learning new capabilities (teaching).
BizFinBench.v2: Towards Reliable LLMs in Finance via Real-User Data and Offline/Online Bilingual Evaluation
Xin Guo (HiThink Research), Liwen Zhang (Shanghai University of Finance and Economics)
TransformerLarge Language ModelTextBenchmarkFinance RelatedRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Built BizFinBench.v2, which includes 28,860 real business Q&A pairs covering Chinese and U.S. stock markets, providing an offline and online dual-track evaluation framework.
Black-Box Assisted Regression: Phase Transitions and Minimax Optimality
Yan Zhou (Changsha University of Science and Technology)
OptimizationData-Centric LearningImageText
🎯 What it does: Studied black-box assisted non-parametric regression under the scenario where only a fixed black-box predictor can be queried and labeled data is limited, and provided corresponding risk minimization and phase transition.
Black-Box Combinatorial Optimization with Order-Invariant Reinforcement Learning
Olivier Goudet (Universite d'Angers), Sylvain Lamprier (Universite d'Angers)
OptimizationNeural Architecture SearchRecurrent Neural NetworkTransformerReinforcement LearningPrompt EngineeringDiffusion modelContrastive LearningGraphTabularBenchmark
🎯 What it does: Proposed an unordered quantization reinforcement learning framework to solve black-box combinatorial optimization problems using neural network generators;
Black-Box Detection of LLM-Generated Text Using Generalized Jensen Shannon Divergence
Shuangyi Chen (University of Toronto), Ashish J Khisti (University of Toronto)
Anomaly DetectionExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringDiffusion modelContrastive LearningText
🎯 What it does: Proposed a black-box text generation detection method called SurpMark, which calculates the surprisal of tokens using a proxy language model, discretizes it into interpretable states, constructs a first-order Markov transition matrix, and then computes the generalized Jensen-Shannon (GJS) distance between transition matrices of two reference corpora (human text and machine text), making judgments based on their differences.
Blending Neural Control Density Functions for Stabilization and Safety
Sahil Chaudhary (Indian Institute of Science), Chiranjib Bhattacharyya (Indian Institute of Science)
OptimizationSafty and PrivacyReinforcement LearningTabularTime SeriesSequential
🎯 What it does: This paper proposes to use neural networks to learn density functions (NCDFs) to provide almost everywhere stability guarantees for nonlinear controllers, and to achieve smooth fusion of controllers through density functions, thereby expanding the system's region of convergence.
Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling
Zeyu Huang (University of Edinburgh), Ivan Titov (University of Amsterdam)
Reinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringText
🎯 What it does: Propose Prefix‑RFT, a post-training method that integrates supervised fine-tuning (SFT) with reinforcement learning fine-tuning (RFT), leveraging offline demonstration prefixes to guide policy generation.
BLIPs: Bayesian Learned Interatomic Potentials
Dario Coscia (SISSA), Max Welling (University of Amsterdam)
Drug DiscoveryGraph Neural NetworkContrastive LearningGraphPhysics Related
🎯 What it does: This paper proposes a scalable, architecture-agnostic variational Bayesian framework called BLIP for precise confidence quantification in machine learning potential energy models for molecules and materials;
BLISS: A Lightweight Bilevel Influence Scoring Method for Data Selection in Language Model Pretraining
Jie Hao (George Mason University), Mingrui Liu (George Mason University)
OptimizationKnowledge DistillationData-Centric LearningTransformerLarge Language ModelContrastive LearningText
🎯 What it does: Propose a lightweight two-level optimization data selection method called BLISS, which starts from scratch and does not require any external pre-trained models, for pre-training large language models.