arXivSub Start free trial

ICML 2026 Papers with AI Summaries

International Conference on Machine Learning · 6554 papers

“Do Diffusion Models Dream of Electric Planes?” Discrete and Continuous Simulation-Based Inference for Aircraft Design

Aurelien Ghiglino (SRI), Adam D. Cobb (SRI)

Autonomous DrivingOptimizationFederated LearningComputational EfficiencyRobotic IntelligenceDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackTransformerDiffusion modelScore-based ModelSimultaneous Localization and MappingWorld ModelOptical FlowTabularTime SeriesSequentialPhysics RelatedStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: This paper proposes a hybrid diffusion model called MixeDiT-MaskeDiT, which uses the SBI method to learn the complete posterior distribution of the conceptual design space for electric vertical take-off and landing (eVTOL) aircraft;

“very likely” Means “uncertain”? How LLMs Diverge from Humans in Linguistic Uncertainty Quantification

Jinhao Duan (UNC-Chapel Hill), Tianlong Chen (UNC-Chapel Hill)

OptimizationExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelPrompt EngineeringAuto EncoderContrastive LearningTextBenchmark

🎯 What it does: This paper constructs a lookup table of human uncertainty vocabulary (UM-Lookup) and proposes an optimized VOCAL algorithm to align the language expressions of LLMs with numerical uncertainty, thereby investigating the differences between LLMs and humans in verbal expressions of uncertainty.

(1D) Ordered Tokens Enable Efficient Test-Time Search

Zhitong Gao (Swiss Federal Institute of Technology Lausanne), Oğuzhan Fatih Kar (Swiss Federal Institute of Technology Lausanne)

GenerationComputational EfficiencyTransformerPrompt EngineeringDiffusion modelImage

🎯 What it does: This paper investigates the impact of 1D ordered (coarse-to-fine) token structures on autoregressive image generation models during test-time search, demonstrating significant advantages in search efficiency and generation quality;

(Be Cautious!) Bio-Foundation Models Are Not Yet Robust to Biologically Plausible Perturbations and ML Transformations

Jinhao Duan (University Of North Carolina Chapel Hill), Tianlong Chen (University Of North Carolina Chapel Hill)

Anomaly DetectionExplainability and InterpretabilityComputational EfficiencyAdversarial AttackDrug DiscoveryProtein Structure PredictionGraph Neural NetworkTransformerPrompt EngineeringDiffusion modelScore-based ModelAuto EncoderContrastive LearningImageMultimodalityGraphBiomedical DataReview/Survey PaperBenchmark

🎯 What it does: This paper conducts 2128 experiments on 7 downstream tasks (sequence, structure, imaging, etc.) using 11 mainstream bio-foundation models (Bio-FM), systematically evaluating the robustness of two types of perturbations—biologically feasible perturbations (e.g., coordinate noise, label renaming) and ML-side transformations (e.g., graph construction parameters, tokenization), and constructs a unified evaluation framework;

(Doubly) Exponential Lower Bounds for Follow the Regularized Leader in Potential Games

Ioannis Anagnostides (Carnegie Mellon University), Tuomas Sandholm (Carnegie Mellon University)

OptimizationReinforcement LearningContrastive LearningReview/Survey Paper

🎯 What it does: In two-player potential games, it is proven that the convergence time of the Follow the Regularized Leader (FTRL) algorithm can be exponential; in multi-player potential games, a double-exponential convergence lower bound for FTRL is provided; and an upper bound is given for a 'lazy alternating' no-regret dynamic that converges to an ϵ-Nash equilibrium in exponential time, showing that the upper bound matches the lower bound tightly at the exponential scale.

(Sparse) Attention to the Details: Preserving Spectral Fidelity in ML-based Weather Forecasting Models

Maksim Zhdanov (AMLab), Jan-Willem van de Meent (AMLab)

Computational EfficiencyRepresentation LearningTransformerTabularTime SeriesPhysics Related

🎯 What it does: Proposed and implemented a probabilistic weather forecasting model called MOSAIC, which addresses three major spectral distortion issues—spectral attenuation, frequency aliasing, and residual high-frequency leakage—by using block sparse attention, functional perturbation, and direct prediction of the next state at the original resolution.

[CLS] is Not Enough: Multi-Label Recognition via Patch-Level Inference and Adaptive Aggregation

Akang Wang (Yunnan Normal University), Huafeng Li (Kunming University of Science and Technology)

RecognitionComputational EfficiencyRepresentation LearningTransformerVision Language ModelAuto EncoderContrastive LearningImageMultimodality

🎯 What it does: Leveraging multi-label image recognition with vision-language models such as CLIP, this paper proposes a training-free, single-forward-pass framework called PIAA, which decomposes the recognition task into local patch-level reasoning and adaptive aggregation.

*MemPot*: Defend Against Memory Extraction Attack with Optimized Honeypots

Yuhao Wang (National University of Singapore), Jiaheng Zhang (National University of Singapore)

Anomaly DetectionOptimizationSafty and PrivacyTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextElectronic Health RecordsRetrieval-Augmented Generation

🎯 What it does: Proposed a honeypot defense framework called MemPot with zero online cost, aimed at preventing memory extraction attacks by LLM agents.

``Someone Hid It!'': Query-Agnostic Black-Box Attacks on LLM-Based Retrieval

Jiate Li (University of Southern California), Yue Zhao (University of Southern California)

RetrievalAdversarial AttackTransformerLarge Language ModelPrompt EngineeringContrastive LearningText

🎯 What it does: Propose a completely black-box, query-agnostic attack method, which injects transferable tokens at the end of the target document, making large language model retrievers (LLMR) unable to retrieve the document without knowing.

$\alpha$-PFN: Fast Entropy Search via In-Context Learning

Herilalaina Rakotoarison (University of Helsinki), Eytan Bakshy (Meta)

OptimizationHyperparameter SearchTransformerContrastive LearningGaussian SplattingTabularTime SeriesBenchmark

🎯 What it does: This paper proposes a two-stage Prior-Data Fitted Network (PFN) framework, which uses a single forward pass to approximate the acquisition function of Entropy Search, thereby achieving fast information gain evaluation.

$\mathbb{R}^{2k}$ is Theoretically Large Enough for Embedding-based Top-$k$ Retrieval

Zihao Wang (TSY Capital), Simon See (NVIDIA AI Technology Center)

RetrievalOptimizationRepresentation LearningContrastive LearningTabular

🎯 What it does: This paper studies the theoretical limits of the minimum embeddable dimension (MED) and its robust version (RMED) in embedded retrieval, proving that MED is Θ(k) under inner product, Euclidean distance, and cosine similarity, while RMED is constrained by m and ε and is O(k·log m) within the feasible range;

$\mathcal{O}(\log N)$ Latent Dimension Suffices for Universal Approximation of Permutation-invariant Function

Min ZHOU, Minghua Chen (Chinese University of Hong Kong)

Computational EfficiencyRepresentation LearningGraph Neural NetworkImageTabularBenchmark

🎯 What it does: Studied permutation-invariant functions under the Wasserstein stability assumption, and proved that there exists a DeepSets architecture with width O(log N) that can achieve global approximation

$\mu$pscaling small models: Principled warm starts and hyperparameter transfer

Yuxin Ma (Johns Hopkins University), Soledad Villar (Johns Hopkins University)

Computational EfficiencyKnowledge DistillationHyperparameter SearchConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningImageTextTabular

🎯 What it does: Proposes a width upscaling method based on µP, which can seamlessly transfer a trained small model into a large model while maintaining dynamic equivalence during training, achieved through weight copying and perturbation.

$\phi$-Balancing for Mixture-of-Experts Training

Lizhang Chen (University Of Texas At Austin), qiang liu

OptimizationFederated LearningComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningMixture of ExpertsContrastive LearningTextBenchmark

🎯 What it does: Proposes the φ-balance framework, which directly balances the utilization of experts across the overall data distribution using a strictly convex symmetric differentiable potential function, and derives an online update algorithm based on exponential moving average (EMA) through convex duality and mirror descent.

$\sigma$: Sigmoid Modulation for Ultra High Resolution Diffusion

Bingxuan Zhao (Northwestern Polytechnical University), Qi Wang (Northwestern Polytechnical University)

GenerationSuper ResolutionTransformerDiffusion modelImage

🎯 What it does: Propose SigMa, a training-free Sigmoid modulation framework that extends pre-trained Diffusion Transformers to ultra-high resolutions (up to 16MP) without retraining.

$\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Victor Barres (Sierra), Karthik R Narasimhan (Sierra)

Reinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AITextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose the τ 2 -bench dual-control environment evaluation framework, covering Dec-POMDP tasks where both agents and users can interact with tools.

$\tau$-Knowledge: Evaluating Conversational Agents over Unstructured Knowledge

Quan Shi (Sierra), Victor Barres (Princeton University)

TransformerLarge Language ModelPrompt EngineeringTextBenchmarkFinance RelatedRetrieval-Augmented Generation

🎯 What it does: Constructed the τ-Knowledge benchmark, added the τ-Banking domain, and evaluated conversational agents' ability to complete complex tasks by retrieving and reasoning over unstructured knowledge in a financial customer service scenario.

$\tau$-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains

Soham Ray (Sierra.ai), Karthik R Narasimhan (Princeton University)

TransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationAudio

🎯 What it does: Proposed the τ-Voice benchmark for evaluating the performance of full-duplex voice agents in real-world multi-turn tasks, integrating task completion, real-time dialogue management, and real audio environments;

$\text{DT}^\text{2}$: Decision-Targeted Digital Twins

Harry Amad (University of Cambridge), Mihaela van der Schaar (University of Cambridge)

OptimizationTransformerSupervised Fine-TuningReinforcement LearningContrastive LearningTabularTime SeriesSequentialBiomedical DataBenchmark

🎯 What it does: Studied a digital twin training method called DT2 for decision support, improving upon traditional training methods that only minimize one-step transition errors.

$\textit{S}$-SPPO: Semantic-Calibrated Self-Play Preference Optimization

Xiwen Chen (Morgan Stanley), Yuriy Nevmyvaka (Morgan Stanley)

OptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningContrastive LearningTextChain-of-Thought

🎯 What it does: To address the strategy degradation problem in self-play preference optimization (SPPO), we propose S-SPPO, which stabilizes training through a dual-space mechanism of supervised calibration and representation calibration.

$\texttt{FHAIM}$: Fully Homomorphic AIM for Private Tabular Synthetic Data Generation

Mayank Kumar (University of Central Florida), Sikha Pentyala (University of Washington Tacoma)

Data SynthesisSafty and PrivacyComputational EfficiencyDiffusion modelGenerative Adversarial NetworkContrastive LearningTabularElectronic Health Records

🎯 What it does: Developed FHAIM, a privacy-preserving table data synthesis framework that utilizes fully homomorphic encryption (FHE), enabling service providers to train the AIM synthesizer without decrypting any original data;

$\texttt{FlashSchNet}$: Fast and Accurate Coarse-Grained Neural Network Molecular Dynamics

Pingzhi Li (University of North Carolina at Chapel Hill), Tianlong Chen (University of North Carolina at Chapel Hill)

Computational EfficiencyDrug DiscoveryGraph Neural NetworkDiffusion modelScore-based ModelFlow-based ModelRectified FlowNeural Radiance FieldAuto EncoderContrastive LearningGaussian SplattingSimultaneous Localization and MappingWorld ModelOptical FlowGraphBiomedical DataBenchmarkStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: A IO-friendly SchNet-style graph neural network molecular dynamics framework called FlashSchNet was constructed to significantly improve the computational speed and memory usage of coarse-grained molecular dynamics.

$\texttt{MetaDistill}$: Unlocking the Performance Ceiling for Pretrained Optimizers

Muqi Han (Guangzhou Institute of Technology), Zilong Wang (Xidian University)

OptimizationKnowledge DistillationMeta LearningSupervised Fine-TuningReinforcement LearningTabularBenchmark

🎯 What it does: MetaDistill improves the performance upper bound of learnable black-box optimizers through pre-training and multi-teacher distillation, and performs self-supervised fine-tuning during testing.

$\texttt{Multi}^2$: Hierarchical Multi-Agent Decision-Making with LLM-Based Agents in Interactive Environments

Sangeun Park (Sungkyunkwan University), Minhae Kwon (Sungkyunkwan University)

Autonomous DrivingOptimizationRobotic IntelligenceTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AIPrompt EngineeringWorld ModelTextSequentialBenchmark

🎯 What it does: Designed and implemented Multi 2 — a hierarchical multi-agent decision-making framework that uses a high-level LLM (System 1) to generate context-aware subgoals, and a low-level LLM (System 2) to execute atomic actions, while enhancing execution robustness through offline-to-online reinforcement learning training.

$\texttt{PRISM}$:A 3D Probabilistic Neural Representation for Interpretable Shape Modeling

Yining Jiao (University of North Carolina at Chapel Hill), Marc Niethammer (University of California San Diego)

Anomaly DetectionExplainability and InterpretabilityRepresentation LearningDiffusion modelScore-based ModelNeural Radiance FieldAuto EncoderGenerative Adversarial NetworkContrastive LearningGaussian SplattingPoint CloudMeshBiomedical DataMagnetic Resonance ImagingComputed Tomography

🎯 What it does: Built PRISM, a probabilistic shape model based on implicit neural representations, which can generate conditional distributions of shapes given continuous covariates (e.g., age) and estimate confidence intervals for individuals' intrinsic developmental time and spatial variation.

$\texttt{ShaplEIG}$: Bayesian Experimental Design for Shapley Value Estimation

David Rundel (LMU Munich), Matthias Feurer (TU Dortmund University)

OptimizationExplainability and InterpretabilityComputational EfficiencyContrastive LearningTabularTime SeriesSequentialBenchmark

🎯 What it does: Proposes ShaplEIG, an adaptive method based on Bayesian experimental design, which uses a Gaussian process surrogate to approximate the value function of Shapley values, and selects the most valuable subset for evaluation based on information gain.

$A_2$DEPT: Large Language Model–Driven Automated Algorithm Design via Evolutionary Program Trees

Bin Chen (University of Electronic Science and Technology of China), Zhengqiu Zhu (National University of Defense Technology)

OptimizationTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmark

🎯 What it does: A framework named A DEPT for automatic algorithm design targeting complete solvers is proposed by combining large language models (LLMs) with evolutionary computation. It utilizes tree-structured evolutionary search and program maintenance loops, enabling LLMs to automatically generate, modify, and execute complete combinatorial optimization solvers without template constraints.

$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses

Di Wu (University of Virginia), Cong Shen (University of Virginia)

Reinforcement Learning from Human FeedbackReinforcement LearningContrastive LearningTabularSequential

🎯 What it does: This paper studies online algorithms using general f-divergence regularization in reinforcement learning with self-feedback (RLHF);

$f$-Divergence Self-Play for Tabular Anomaly Detection via Large Language Models

Hoang Tran Vuong (Hanoi University of Science and Technology), Trung Le (Monash University)

Anomaly DetectionTransformerLarge Language ModelReinforcement LearningContrastive LearningTabular

🎯 What it does: Proposed a self-play based f-Divergence refinement framework called DiSPaT, which uses large language models for unsupervised anomaly detection on mixed-type tabular data.

$f$-Trajectory Balance: A Loss Family for Tuning GFlowNets, Generative Models, and LLMs with Off- and On-Policy Data

Jake Fawkes (University College London), Jason Hartford (Valence Labs)

GenerationData SynthesisOptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningDiffusion modelScore-based ModelFlow-based ModelImageTextTabular

🎯 What it does: This paper proposes a family of loss functions L_f based on f-divergence, extending the known KL-square (trajectory balance) loss. When on-policy, the gradient is equivalent to the corresponding f-divergence, and when off-policy, it still maintains the same global optimal solution. This method is applied to GFlowNets, diffusion models, and the fine-tuning of large language models (LLMs).

$G^2$-Reader: Dual Evolving Graphs for Multimodal Document QA

Yaxin Du (Shanghai Jiao Tong University), Siheng Chen (Shanghai Jiao Tong University)

RetrievalExplainability and InterpretabilityComputational EfficiencyGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextMultimodalityGraphRetrieval-Augmented Generation

🎯 What it does: Propose a dual-graph structure multi-modal document question answering system (G2-Reader), which retains the native structure and cross-modal semantics of the document through a content graph, and uses a planning graph as a sub-problem planning of a directed acyclic graph to progressively retrieve and combine evidence, ultimately generating an answer.

$L^3$: Large Lookup Layers

Albert Tseng (Cornell University), Christopher De Sa (Cornell University)

Computational EfficiencyRepresentation LearningTransformerLarge Language ModelMixture of ExpertsContrastive LearningText

🎯 What it does: Propose the Large Lookup Layer (L3) to expand the token embedding table, using context-aware static routing in the decoding layer to improve the performance of sparse models

$R^3$DAO: Reactive Recovery and Reconstruction for Long-horizon Data Agent Orchestration

Quanxin Liu (Huazhong University of Science and Technology), Yijun Mo (Huazhong University of Science and Technology)

Computational EfficiencyKnowledge DistillationData-Centric LearningRobotic IntelligenceReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringImageTextMultimodalityTabularBenchmarkAudio

🎯 What it does: Propose a long-term data science workflow automation framework RDAO 3 based on reactive recovery and reconfiguration

$V_0$: A Generalist Value Model for Any Policy at State Zero

Yi-Kai Zhang (Nanjing University), Han-Jia Ye (Nanjing University)

Computational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsContrastive LearningTextTabular

🎯 What it does: Proposes V0, a general value model capable of estimating value on any policy, which can predict the success rate of LLMs in initial states (state zero) without parameter updates.

1-Bit Wonder: Improving QAT Performance in the Low-Bit Regime through K-Means Quantization

Sohir Maskey (Aleph Alpha Research), Douglas Orr (Graphcore Research)

Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningTextMultimodalityBenchmark

🎯 What it does: The study investigates low-bit quantization in QAT, comparing the effects of k-means nonlinear quantization and uniform integer quantization on LLMs.

3D MeanFlow: One-Step Point Cloud Completion and Generation via Average-Velocity Transport

Haowen Zhong (Tongji University), Shangce Gao (University of Toyama)

Object DetectionGenerationAutonomous DrivingGraph Neural NetworkDiffusion modelScore-based ModelFlow-based ModelOptical FlowPoint Cloud

🎯 What it does: Propose a single-step point cloud completion and generation model called 3DMF, based on a teacher-free training framework with average velocity transport, and introduce the PointPlug plugin to enhance 3D object detection performance.

3D Scene Assertion Verification

Jun Lin (China University of Geosciences), Wenqian Wang (University of the Chinese Academy of Sciences)

ClassificationRecognitionAnomaly DetectionExplainability and InterpretabilityConvolutional Neural NetworkRecurrent Neural NetworkTransformerPrompt EngineeringMixture of ExpertsVision Language ModelContrastive LearningTextMultimodalityPoint CloudBenchmarkRetrieval-Augmented Generation

🎯 What it does: This paper proposes a new 3D scene assertion verification task, requiring strict binary classification of truth values for given natural language assertions in 3D scenes;

3D-DLP: Self-supervised 3D Object-centric Scene Representation Learning

Ellina Zhang (Carnegie Mellon University), Tal Daniel (Carnegie Mellon University)

Representation LearningRobotic IntelligenceConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningGaussian SplattingImagePoint CloudMeshTabular

🎯 What it does: Proposed the 3D-DLP model, achieving self-supervised 3D object-centric scene representation learning, which can decompose RGB-D or voxel observations into 3D latent particles and support scene reconstruction and editing.

3D-RFT: Reinforcement Fine-Tuning for Video-based 3D Scene Understanding

Xiongkun Linghu (State Key Laboratory of General Artificial Intelligence), Siyuan Huang (State Key Laboratory of General Artificial Intelligence)

OptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision Language ModelVideoTextPoint CloudChain-of-Thought

🎯 What it does: Proposed the 3D-RFT framework, which directly optimizes video-based 3D scene understanding tasks using reinforcement learning with verifiable rewards (such as 3D IoU, F1-Score), replacing the traditional token-level cross-entropy supervision.

3DGS-HPC: Distractor-free 3D Gaussian Splatting with Hybrid Patch-wise Classification

Jiahao Chen (Sun Yat-sen University), Guanbin Li (Sun Yat-sen University)

RestorationContrastive LearningGaussian SplattingImagePoint Cloud

🎯 What it does: This paper proposes a Hybrid Patch-wise Classification (HPC) framework to suppress the impact of transient artifacts (such as pedestrians and shadows) on reconstruction quality during the training process of 3D Gaussian Splatting (3DGS);

3DGS$^2$-TR: A Scalable Second-Order Trust-Region Method for 3D Gaussian Splatting

Roger Hsiao (University of California Berkeley), Sophia Shao (University of California Berkeley)

GenerationData SynthesisOptimizationComputational EfficiencyScore-based ModelGaussian SplattingPoint CloudMesh

🎯 What it does: Designed and implemented a second-order optimizer called 3DGS-TR based on the diagonal estimation of the Hessian matrix, combined with a parameter-level trust region using the squared Hellinger distance, significantly improving the training efficiency and reconstruction quality of 3D Gaussian Splatting.

3DMedAgent: Unified Perception-to-Understanding for 3D Medical Analysis

Ziyue Wang (National University of Singapore), Yueming Jin (National University of Singapore)

ClassificationRecognitionSegmentationAnomaly DetectionExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageMultimodalityBiomedical DataComputed TomographyBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose 3DMedAgent, a unified agent based on 2D multimodal large language models, which can progressively advance from low-level perception (measurement, localization) to high-level clinical understanding (pathological reasoning) of three-dimensional CT scans by invoking various visual and textual tools and utilizing long-term structured memory. Meanwhile, the DeepChestVQA benchmark is constructed to evaluate the three-dimensional visual question answering capability of chest CT.

3DPoV: Improving 3D understanding via Patch Ordering on Videos

Ioana Simion (University of Amsterdam), Yuki M Asano (University of Technology Nuremberg)

Pose EstimationDepth EstimationRepresentation LearningTransformerSupervised Fine-TuningContrastive LearningOptical FlowVideoPoint Cloud

🎯 What it does: Lightweight self-supervised fine-tuning of visual foundation models is performed using point tracking, differentiable sorting, and a teacher-student framework in videos, thereby enhancing their 3D spatial consistency and geometric understanding.

3ViewSense: Spatial and Mental Perspective Reasoning from Orthographic Views in Vision-Language Models

Shaoxiong Zhan (Tsinghua University), Hai-Tao Zheng (Tsinghua University)

Representation LearningReinforcement Learning from Human FeedbackReinforcement LearningPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBenchmark

🎯 What it does: Propose the 3ViewSense framework, introducing simulated and reasoned orthogonal projection perspectives into Vision-Language Models (VLM) to address the 'spatial intelligence gap' in spatial reasoning.

4DPC$^2$hat: Towards Dynamic Point Cloud Understanding with Failure-Aware Bootstrapping

Xindan Zhang (Jilin University), Hehe Fan (Zhejiang University)

RecognitionData SynthesisRepresentation LearningTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningVideoTextMultimodalityPoint CloudRetrieval-Augmented Generation

🎯 What it does: Proposes a large-scale multimodal language model, 4DPC hat, specifically designed for understanding dynamic 4D point cloud sequences, and constructs a corresponding cross-modal dataset, 4DPC hat-200K.

4RC: 4D Reconstruction via Conditional Querying Anytime and Anywhere

Yihang Luo (Nanyang Technological University), Chen Change Loy (University of Oxford)

Object TrackingData SynthesisPose EstimationDepth EstimationAutonomous DrivingTransformerNeural Radiance FieldAuto EncoderContrastive LearningOptical FlowVideoPoint Cloud

🎯 What it does: Propose a unified feed-forward Transformer framework, 4RC, which can encode an entire monocular video into a 4D representation in one go and support querying geometry and motion from any source frame to any target time point at any time and place.

A Bayesian Approach to Quantify the Uncertainty of Human Ratings in a Single-Instance Multimodal Framework

Zijian Chen (Boston University), Archana Venkataraman (Boston University)

Explainability and InterpretabilityRepresentation LearningGraph Neural NetworkTransformerAuto EncoderMultimodalityBiomedical DataAlzheimer's Disease

🎯 What it does: Propose a Bayesian graph model that estimates instance-level uncertainty in human ratings using auxiliary objective data and learns interpretable feature selection masks;

A Behavioural and Representational Evaluation of Goal-Directedness in Language Model Agents

Raghu Arghal (University of Pennsylvania), Mario Giulianelli (University College London)

Explainability and InterpretabilityRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringWorld ModelTextSequentialChain-of-Thought

🎯 What it does: Propose a framework that combines behavioral evaluation and internal representation probing to assess the goal-directedness of large language model agents, using 2D grid world navigation as a case study.

A Benchmark and Framework for Evaluating Next Action Predictions in Spreadsheets

Tejas Agrawal (Microsoft), Gust Verbruggen (Microsoft)

Data-Centric LearningAI Code AssistantRecurrent Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTabularSequentialBenchmark

🎯 What it does: A benchmark dataset and an online evaluation framework for assessing the prediction of the next operation in spreadsheets are proposed, aiming to promote the development of call-free auto-completion technologies.

A Bi-metric Framework for Efficient Nearest Neighbor Search

Haike Xu (Massachusetts Institute of Technology), Piotr Indyk (Massachusetts Institute of Technology)

RetrievalComputational EfficiencyContrastive LearningTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposed and implemented a bi-metric framework that uses a low-cost proxy metric to build an index, and employs a high-cost true metric during queries to achieve precise nearest neighbors;

A Call to Lagrangian Action: Learning Population Mechanics from Temporal Snapshots

Vincent Guan (University of British Columbia), Kirill Neklyudov (Mila -Quebec AI Institute)

OptimizationDiffusion modelNeural Radiance FieldAuto EncoderGenerative Adversarial NetworkContrastive LearningOptical FlowPoint CloudGraphTabularTime SeriesBenchmarkPhysics RelatedStochastic Differential Equation

🎯 What it does: Propose Wasserstein Lagrangian Mechanics (WLM) and design an algorithm for learning second-order collective dynamics from time snapshots;

A Capacity-Based Rationale for Multi-Head Attention

Micah Adler (MIT)

RecognitionExplainability and InterpretabilityComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerLarge Language ModelMixture of ExpertsContrastive LearningTextGraphReview/Survey Paper

🎯 What it does: This paper views the 'finding relations' in self-attention as a graph edge recovery task (Relational Graph Recognition), and investigates how many different token-token relations a single-layer attention can reliably represent and recover under a fixed Q-K dimension budget, proposing a capacity theory with upper and lower bounds.

A Cartesian-3j Framework for Machine Learning Interatomic Potentials

Zemin Xu (ShanghaiTech University), Peijun Hu (ShanghaiTech University)

Representation LearningGraph Neural NetworkContrastive LearningGraphBenchmarkPhysics Related

🎯 What it does: This paper constructs a machine learning atomic potential framework based on irreducible Cartesian tensors (ICT), and implements ICT multiplication and contraction on e3nn.

A Causal Decomposition Approach for Fair Contextual Multi-Armed Bandits

Jiajun Chen (Iowa State University), Christopher John Quinn (Iowa State University)

OptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyReinforcement LearningContrastive LearningTabular

🎯 What it does: This paper proposes a fair contextual multi-armed bandit learning framework that utilizes causal decomposition (direct, indirect, confounding effects) as constraints, and provides online algorithms under both nonparametric and logistic regression settings, achieving sublinear regret degradation and upper bounds on constraint violation accumulation.

A Close Look at Negative Label Guided Out-of-distribution Detection in Pre-trained Vision-Language Models

Bo Peng (University of Technology Sydney), Guangquan Zhang (University of Technology Sydney)

Anomaly DetectionRepresentation LearningTransformerVision Language ModelContrastive LearningImageTextMultimodality

🎯 What it does: This paper studies post-processing OOD detection for pre-trained vision-language models based on negative labels, and proposes a new energy-based framework aimed at improving the effectiveness of the model in open-world scenarios.

A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness

Leo Schwinn (Helmholtz AI), Stephan Günnemann

Safty and PrivacyAdversarial AttackTransformerLarge Language ModelPrompt EngineeringTextBenchmark

🎯 What it does: Audit the reliability of LLM-as-a-Judge in adversarial security evaluation, and find that three types of distribution shifts—attack, model, and data—lead to an increase in the misjudgment rate of the judge, thereby affecting the assessment of attack success rates.

A Computational Framework for Evaluating Human-likeness in LLMs' Open-ended Human Behaviors

Yuxuan Lei (University of Science and Technology of China), Xing Xie (Microsoft Research Asia)

Recommendation SystemData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextSequentialRetrieval-Augmented Generation

🎯 What it does: A framework based on distributed evaluation was constructed to measure the realism and credibility of LLMs in simulating human behavior using large-scale network behavioral data.

A Conflict-aware Evidential Framework for Reliable Sleep Stage Classification

Yunzhi Tian (Northwest University), Jun Feng (Northwest University)

ClassificationRecurrent Neural NetworkGraph Neural NetworkContrastive LearningTime SeriesSequentialBiomedical Data

🎯 What it does: Proposed ConfSleepNet, a conflict-aware evidence framework for multi-view sleep staging;

A Consensus Anchor-guided Hypergraph Framework for Incomplete Multi-view Clustering

Yipin Hu (Harbin Institute of Technology), Guoqing Chao (Harbin Institute of Technology)

OptimizationRepresentation LearningData-Centric LearningGraph Neural NetworkAuto EncoderContrastive LearningImageGraphTabular

🎯 What it does: Proposed the HA-IMVC framework for large-scale incomplete multi-view clustering.

A Constrained Optimization Perspective of Unrolled Transformers

Javier Porras-Valenzuela (University of Pennsylvania), Alejandro Ribeiro (University of Pennsylvania)

ClassificationAnomaly DetectionOptimizationTransformerLarge Language ModelSupervised Fine-TuningVideoText

🎯 What it does: Proposes a framework that introduces hierarchical descent constraints into Transformer training, ensuring that the model achieves a decreasing expected loss at each layer;

A Control-Theoretic View of Mamba on Stability and Robustness

Liang Cao (University of British Columbia), Yan Qin (Chongqing University)

OptimizationExplainability and InterpretabilityComputational EfficiencyRobotic IntelligenceRecurrent Neural NetworkAuto EncoderContrastive LearningTime SeriesSequentialBenchmarkStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: This paper constructs a control theory framework to systematically analyze the stability and robustness of selective state space models such as Mamba, and provides certificates that can be directly verified during the inference phase.

A Critical Look at Targeted Instruction Selection: Disentangling What Matters (and What Doesn’t)

Nihal V. Nayak (Harvard University), David Alvarez-Melis (Harvard University)

Computational EfficiencyRepresentation LearningData-Centric LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: This paper proposes a framework that decomposes the selection of target instructions into two core elements: data representation and selection algorithm, and systematically evaluates its performance across different models and tasks.

A Decision-Theoretic View of Test-Time Training: When, How Far, and Which Directions to Adapt

Tomoya Wakayama (RIKEN)

Domain AdaptationOptimizationRepresentation LearningMeta LearningTransformerPrompt EngineeringImageTextBiomedical DataReview/Survey PaperBenchmark

🎯 What it does: Propose to view test-time training as implicit Bayesian inference, analyze its risks, step selection, and subspace selection, and provide theoretical guidance.

A Deep Learning Model of Mental Rotation Informed by Interactive VR Experiments

Raymond Khazoum (Aalto University), Stephane Deny

Explainability and InterpretabilityRepresentation LearningReinforcement Learning from Human FeedbackConvolutional Neural NetworkTransformerVision-Language-Action ModelDiffusion modelContrastive LearningImagePoint CloudSequentialChain-of-Thought

🎯 What it does: This study designs and implements a deep learning-based mental rotation model, integrating equivariant neural encoders, neuro-symbolic modules, and neural decision agents, and verifies its behavior similar to humans through interactive VR experiments.

A Diagnostic Study of Multi-Agent LLMs for Real-World Debates

Priya Pitre (Virginia Tech), Xuan Wang (Virginia Tech)

Explainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Propose a diagnostic evaluation framework for multi-agent LLM debates, which evaluates both debate outcomes and the debate process, revealing process attributes such as engagement, responsiveness, influence asymmetry, and balance.

A Diffusive Classification Loss for Learning Energy-based Generative Models

RuiKang OuYang (University of Cambridge), José Miguel Hernández-Lobato (University of Cambridge)

GenerationData SynthesisOptimizationRepresentation LearningDiffusion modelScore-based ModelBiomedical Data

🎯 What it does: Propose Diffusive Classification Loss (DiffCLF), a new objective function for training energy-based generative models, which balances energy learning with resolution consistency;

A Dirac-Frenkel-Onsager Principle: Instantaneous Residual Minimization with Gauge Momentum for Nonlinear Parametrizations of PDE Solutions

Matteo Raviola (EPFL), Benjamin Peherstorfer (New York University)

OptimizationBenchmarkPhysics RelatedStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Propose a PDE parameterized dynamics framework (Dirac-Frenkel-Onsager) based on the Dirac-Frenkel variational principle and the Onsager minimum dissipation principle, achieving unambiguous and smooth parameter evolution by injecting memory momentum into the Jacobian null space;

A Direct Approach for Handling Contextual Bandits with Latent State Dynamics

Zhen LI, Gilles Stoltz (Université Paris-Saclay)

Reinforcement LearningTabularFinance Related

🎯 What it does: A direct method is proposed, combining phased LinUCB with belief estimation from a hidden Markov model to handle linear contextual bandits with latent state dynamics. Upper bounds on pseudo-loss with high probability are provided: T^{3/4} for simplified models and T^{7/8} for complex models.

A Direct Second-Order Method for Solving Two-Player Zero-Sum Games

David Yang (Columbia University), Christian Kroer (Columbia University)

OptimizationReinforcement LearningTabular

🎯 What it does: Proposed the first direct second-order method (based on Douglas-Rachford splitting and semi-smooth Newton) to compute Nash equilibria in two-player zero-sum games, and designed a hybrid algorithm that smoothly transitions from the first-order PRM+ to the second-order SSN.

A Distributional View for Visual Mechanistic Interpretability: KL-Minimal Soft-Constraint Principle

Guancheng Zhou (Xi'an Jiaotong University), Xipeng Qiu (Shanghai Innovation Institute)

Explainability and InterpretabilityTransformerDiffusion modelAuto EncoderContrastive LearningImage

🎯 What it does: A distribution perspective for visual mechanism interpretation was constructed, and the KL minimum soft constraint principle and energy-guided diffusion posterior sampling method were proposed;

A Factorized Low-Rank RNN Framework for Uncovering Independent Neural Latent Dynamics and Connectivity

Chengrui Li (Georgia Institute of Technology), Anqi Wu (Georgia Institute of Technology)

Explainability and InterpretabilityRepresentation LearningRecurrent Neural NetworkAuto EncoderContrastive LearningTime SeriesSequentialBiomedical DataElectrocardiogramReview/Survey Paper

🎯 What it does: Propose FacRNN, a low-rank RNN structure, to simultaneously learn the low-dimensional dynamics and functional connectivity of neural networks, and achieve group-level independence in the latent space.

A Fine-Grained Understanding of Uniform Convergence for Halfspaces

Aryeh Kontorovich (Ben Gurion University), Kasper Green Larsen (Aarhus University)

OptimizationExplainability and InterpretabilityComputational EfficiencyRepresentation LearningContrastive Learning

🎯 What it does: This paper conducts a fine-grained analysis of the uniform convergence of halfspaces, revealing the significant impact of different dimensions and whether they are homogeneous on the convergence behavior.

A Flat Vocabulary or a Rich Hierarchy? Re-introducing Intrinsic Structure Transforms the Autoregressive Image Generation

Landis He, Shikang Zheng (Shanghai Jiao Tong University)

GenerationTransformerDiffusion modelScore-based ModelContrastive LearningImage

🎯 What it does: Proposes the MASC framework, which constructs a hierarchical semantic tree using the manifold geometry of the code itself to improve autoregressive image generation.

A Formal Comparison Between Chain of Thought and Latent Thought

Kevin Xu (University of Tokyo), Issei Sato (University of Tokyo)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkChain-of-Thought

🎯 What it does: This paper formally compares the expressive power, parallelism, and approximate counting/sampling capabilities of chain-of-thought (CoT) and latent thought reasoning. It verifies the strengths and weaknesses of the two modes of thinking in different tasks through limited experiments.

A Fourier perspective on the learning dynamics of neural networks: from sample complexities to mechanistic insights

Fabiola Ricci (International School of Advanced Studies), Sebastian Goldt (International School of Advanced Studies)

ClassificationImage TranslationImage HarmonizationRestorationSuper ResolutionExplainability and InterpretabilityComputational EfficiencyRepresentation LearningData-Centric LearningDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningOptical FlowImageVideoTextMultimodalityPoint CloudMeshGraphTabularTime SeriesSequentialBiomedical DataReview/Survey PaperBenchmarkPhysics RelatedStochastic Differential EquationOrdinary Differential EquationAudio

🎯 What it does: This paper models and analyzes the learning dynamics of neural networks trained with gradient descent from a Fourier perspective, proposing a synthetic data model with controllable amplitude and phase, and theoretically and experimentally verifying the phenomenon that learning is difficult under flat covariance with phase information, but can significantly accelerate under the power-law spectrum of natural images.

A Fully First-Order Layer for Differentiable Optimization

Zihao Zhao (Georgia Institute of Technology), Kai Wang (Georgia Institute of Technology)

OptimizationReinforcement LearningDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive Learning

🎯 What it does: Proposed a fully first-order based differentiable optimization layer (FFOLayer), which approximates the supergradient through bi-level optimization and active set Lagrangian approximation.

A Game-Theoretic Analysis of Attacks on Large Language Models via Compositional Skills

Xinbo Wu (University of Illinois Urbana-Champaign), Lav R. Varshney (Stony Brook University)

Safty and PrivacyAdversarial AttackTransformerLarge Language ModelPrompt EngineeringText

🎯 What it does: Propose a game-theoretic framework to model the interaction between attacks and defenses in large language models (LLMs), where attackers hide malicious intent by combining skills, and design optimal response attacks and optimal defense strategies within this framework;

A Game-Theoretic Framework for Measuring and Explaining Metric Compatibility in Fair Machine Learning

Lingfeng Zhang (East China Normal University), Zhang Qing

Explainability and InterpretabilityTabular

🎯 What it does: Propose a framework based on game theory, using Harsanyi interaction to decompose fairness and efficacy metrics into attribute interaction vectors, and using cosine similarity to measure compatibility between metrics, thereby achieving interpretable attribution of the fairness-efficacy conflict mechanism.

A General Framework for Dynamic Consistent Submodular Maximization

Paul Duetting (Google Research), Morteza Zadimoghaddam (Google Research)

OptimizationReinforcement LearningMixture of ExpertsContrastive Learning

🎯 What it does: Propose a general framework for achieving consistency submodular maximization algorithms in fully dynamic environments, providing approximate algorithms for cardinality constraints and matroid constraints respectively.

A General Framework for Fair and Robust Regression

Wenhai Cui (Hong Kong Polytechnic University), Xingqiu Zhao (Hong Kong Polytechnic University)

OptimizationFederated LearningExplainability and InterpretabilityData-Centric LearningContrastive LearningTabularBenchmark

🎯 What it does: Propose a unified post-processing framework that performs robust regression with multiple robust loss functions (e.g., Cauchy, Huber, LAD, quantile, Tukey) while satisfying population fairness (demographic balance) constraints.

A General Neural Backbone for Mixed-Integer Linear Optimization via Dual Attention

Peixin Huang (Shandong University), Wen Song (Shandong University)

OptimizationGraph Neural NetworkTransformerAuto EncoderContrastive LearningGraphTabularBenchmark

🎯 What it does: Proposed an attention-based MILP representation learning architecture that adopts an element-centric perspective, treating variables and constraints as two types of elements, and using dual-channel self-attention and cross-attention for parallel fusion.

A Generalist Pair-wise Progress Critic Model for Vision-Language-Action Robots

Qi Zhang (Shanghai AI Lab), Ming Zhou (Shanghai AI Lab)

Robotic IntelligenceReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringVision Language ModelVision-Language-Action ModelContrastive LearningImageVideoTextMultimodality

🎯 What it does: Proposes VLAC, a Vision-Language-Action-Critic model that unifies action generation with task progress evaluation within a single autoregressive framework.

A Geometric Analysis of Small-sized Language Model Hallucinations

Emanuele Ricco (King Abdullah University of Science and Technology), Roberto Di Pietro (King Abdullah University of Science and Technology)

Anomaly DetectionExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextRetrieval-Augmented Generation

🎯 What it does: This study investigates the hallucination phenomenon produced by small-scale language models when generating responses to repeated prompts. It proposes viewing hallucination as a geometric problem of retrieval instability, and analyzes and detects it through structural differences in the sentence embedding space.

A Geometric Lens on Physics-Aligned Data Compression

Aleix Segui Ugalde, Wesley Armour (University of Oxford)

CompressionAuto EncoderPoint CloudTabularTime SeriesBiomedical DataBenchmarkPhysics Related

🎯 What it does: Proposes the theory of physically aligned compressed local geometry, explaining the trade-off between physical quantities and standard reconstruction error under fixed code rate, and provides corresponding alignment diagnostic methods.

A geometric relation of the error introduced by sampling a language model's output distribution to its internal state

Albert F. Modenbach (King's College London)

Explainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelDiffusion modelScore-based ModelContrastive LearningWorld ModelTextStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: This paper investigates the geometric relationship between token sampling ambiguities (top-two word distribution overlaps) during language model generation and internal hidden states. It proposes an SO(n) 1-form and its curvature based on token embedding geometry, and demonstrates the coupling between this geometric uncertainty measure and the world model learned by the model through constructing parallel transport and holonomy operations. Furthermore, it verifies that the rotation direction matches piece importance using three-dimensional PCA and Spearman correlation.

A Geometry-Aware Efficient Algorithm for Compositional Entropic Risk Minimization

Xiyuan Wei (Texas Aandm University), Tianbao Yang (Texas Aandm University)

ClassificationOptimizationContrastive LearningImageTabularStochastic Differential Equation

🎯 What it does: Proposed a geometry-aware stochastic optimization algorithm called SCENT for Combinatorial Entropy Risk Minimization (CERM), which effectively updates dual variables by adopting Stochastic Proximal Mirror Descent (SPMD) on the dual min-min formulation, addressing issues such as non-convergence, numerical instability, and slow convergence in traditional methods.

A Geometry-Based View of Mahalanobis OOD Detection

Denis Janiak (Wrocław University of Science and Technology), Tomasz Jan Kajdanowicz

Anomaly DetectionRepresentation LearningTransformerContrastive LearningImage

🎯 What it does: Propose analyzing the performance of Mahalanobis OOD detection from a geometric perspective, and discover that it is highly sensitive to the geometric structure of the feature space;

A Graph Foundation Model with Cross-Modal Alignment and Modality-Aware Expert Fusion for Multi-Modal Graphs

Dongxiao He (Tianjin University), Di Jin (Tianjin University)

Domain AdaptationRecommendation SystemAnomaly DetectionFederated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningMeta LearningReinforcement Learning from Human FeedbackNeural Architecture SearchGraph Neural NetworkTransformerSupervised Fine-TuningPrompt EngineeringMixture of ExpertsVision Language ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningImageTextMultimodalityGraphRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: A multi-modal graph foundation model named CAME was constructed, pre-trained on cross-modal alignment and modality-aware expert fusion, learning unified representations of structure and multi-source features, and evaluated on multi-task and few-shot transfer tasks.

A Graphop Analysis of Graph Neural Networks on Sparse Graphs: Generalization and Universal Approximation

Ofek Amran (Technion Israel Institute of Technology), Ron Levie (Technion Israel Institute of Technology)

Federated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningGraph Neural NetworkDiffusion modelScore-based ModelFlow-based ModelRectified FlowAuto EncoderGenerative Adversarial NetworkContrastive LearningGraph

🎯 What it does: A unified metric space was constructed, utilizing graphops theory and bounded fiber operators (bofops) to describe full-scale inputs of both sparse and dense graphs, enabling message passing graph neural networks (MPNN) to perform generalization and approximation analysis within the same framework.

A hitchhiker's guide to Poisson gradient estimation

Michael Ibrahim (UC Berkeley), Hadi Vafaii (UC Berkeley)

OptimizationRepresentation LearningReinforcement Learning from Human FeedbackScore-based ModelAuto EncoderContrastive LearningImageTabularBiomedical DataReview/Survey Paper

🎯 What it does: Systematically compare gradient estimation methods in Poisson latent variable models, and improve the EAT method to make its mean unbiased and its variance more consistent with the original distribution.

A Hypertoroidal Covering for Perfect Color Equivariance

Yulong Yang (Princeton University), Christine Allen-Blanchette (Princeton University)

ClassificationImage TranslationRestorationConvolutional Neural NetworkDiffusion modelContrastive LearningImageBiomedical Data

🎯 What it does: This paper proposes a color equivariant network called T3CEN based on dual covering, achieving full equivariance with respect to hue, saturation, and brightness.

A Judge-Aware Ranking Framework for Evaluating Large Language Models without Ground Truth

Mingyuan Xu (National University of Singapore), Doudou Zhou (National University of Singapore)

Recommendation SystemExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelText

🎯 What it does: Propose a benchmark-free label-free evaluation framework for large language models (LLMs), which can directly infer model quality through binary comparison results from LLM judges;

A Kinetic Energy Perspective of Flow Matching

Ziyun Li (Kth Royal Institute Of Technology), Henrik Boström

GenerationData SynthesisScore-based ModelFlow-based ModelRectified FlowImagePhysics RelatedOrdinary Differential Equation

🎯 What it does: Proposes a trajectory-level kinetic energy measure called KPE in flow matching, and investigates its relationship with the semantic quality of generated samples and local sparsity. Subsequently, a training-agnostic two-phase energy shaping strategy called KTS is designed based on KPE to improve generation quality and reduce memorization.

A KL-regularization framework for learning to plan with adaptive priors

Alvaro Serra-Gomez (Leiden University), Thomas M. Moerland (Leiden University)

OptimizationRobotic IntelligenceReinforcement LearningWorld ModelTabularTime SeriesSequential

🎯 What it does: Proposes the PO-MPC framework, combining MPPI planning with KL regularized RL, learning through sampling strategies guided by adapted priors.

A Language-Guided Bayesian Optimization for Efficient LoRA Hyperparameter Search

Baek Seong-Eun (POSTECH), Tae-Hyun Oh (KAIST)

OptimizationHyperparameter SearchData-Centric LearningTransformerLarge Language ModelPrompt EngineeringText

🎯 What it does: Propose a framework that combines large language models (LLMs) with Bayesian optimization (BO) for efficiently searching hyperparameters of LoRA, significantly improving the performance of LoRA fine-tuning.

A Linearly Convergent Proximal Subgradient Algorithm for Sparse Portfolio Optimization with Transaction Cost

Xiaoting Yao (Sun Yat-sen University), Na Zhang (South China Agricultural University)

OptimizationDiffusion modelScore-based ModelFlow-based ModelRectified FlowContrastive LearningOptical FlowTabularTime SeriesFinance RelatedStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Proposes a K-sparse TCO model that simultaneously considers transaction costs and portfolio sparsity, and designs a linearly convergent proximal subgradient algorithm (PSGA) combined with ADMM to solve its difference-of-convex (DC) reformulation problem.

A Machine-Learned Comorbidity Index

Suleman Baloch (University of Iowa), Bijaya Adhikari (University of Iowa)

Representation LearningElectronic Health Records

🎯 What it does: Propose and implement a single severity score model based on diagnostic codes, MLCI, which utilizes multi-endpoint normalized HSIC maximization to capture statistical dependencies between diagnosis and multiple clinical outcomes.

A Mechanistic Analysis of Sim-and-Real Co-Training in Generative Robot Policies

Yu Lei (University of Texas at Austin), Yuke Zhu (University of Texas at Austin)

Domain AdaptationRepresentation LearningRobotic IntelligenceConvolutional Neural NetworkTransformerDiffusion modelImageVideoTabular

🎯 What it does: Investigated and validated the mechanism of sim-and-real co-training for diffusion-based generative robot policies, identifying and quantifying two key intrinsic effects: structured representation alignment (balancing cross-domain alignment and domain discriminability) and importance weighting; empirically verified their impact through controlled experiments, sim-sim and sim-real robotic grasping tasks; based on this, proposed the CFG-ADDA combined method, significantly improving the success rate of real-world tasks.

A Minimal Agent for Automated Theorem Proving

Borja Requena (Axiomatic AI), Leopoldo Sarra (Axiomatic AI)

Explainability and InterpretabilityComputational EfficiencyAI Code AssistantTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose a minimalist automated theorem proving agent that supports iterative proof refinement, memory management, and tool calling, facilitating systematic comparisons across different AI reasoners.

A Minimax Approach for Optimal Intervention Policy Learning with Two-Stage Outcomes

Chenyang Li (Renmin University of China), Yue Liu (Renmin University of China)

Recommendation SystemOptimizationReinforcement LearningContrastive LearningTabularSequential

🎯 What it does: This paper proposes a minimax intervention strategy learning framework under two-stage outcomes, utilizing partial identifiability of the principal strata to optimize intervention allocation.