ICML 2026 Papers — Page 42
International Conference on Machine Learning · 6554 papers
Physics from Video: Identifiability of Time-Invariant Second-Order ODEs under Minimal Trajectory Conditions
Yuanyuan Wang (Mohamed bin Zayed University of Artificial Intelligence), Mingming Gong (Mohamed bin Zayed University of Artificial Intelligence)
Convolutional Neural NetworkAuto EncoderContrastive LearningOptical FlowVideoPhysics RelatedOrdinary Differential Equation
🎯 What it does: Studied how to identify physical parameters that satisfy a homogeneous second-order linear ordinary differential equation from raw video under a framework that uses only an encoder and no decoder, and provided geometric conditions for identifiability and upper bounds on finite-sample error.
Physics in 2-Steps: Locking Motion Priors Before Visual Refinement Erases Them
Woojung Han (Yonsei University), Seong Jae Hwang (Yonsei University)
GenerationDiffusion modelVideoPhysics Related
🎯 What it does: Propose a training-agnostic framework called PhaseLock, which captures motion priors through two-step reasoning and maintains physical consistency in subsequent high-order reasoning using Latent Delta Guidance,
Physics-Guided Motion Loss for Video Generation Model
Bowen Xue (University of Manchester), Zahra Montazeri (University of Manchester)
GenerationData SynthesisSupervised Fine-TuningDiffusion modelOptical FlowVideoTextPhysics Related
🎯 What it does: This paper proposes a physics-guided motion loss based on the frequency domain to improve the physical motion performance in video diffusion models, avoiding issues such as rubber sheet stretching, periodic flickering, and inconsistent scaling and rotation.
Physics-informed coarsening for multigrid graph neural networks surrogates
Amir Bazzi (CEMEF Mines Paris PSL), Elie Hachem (CEMEF Mines Paris PSL)
OptimizationGraph Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningMeshGraphPhysics Related
🎯 What it does: Proposed a physics-informed multigrid graph neural network for high-precision, long-term solid mechanics simulation as an alternative to traditional finite element solvers.
Physics-Informed Diffusion Models in Spectral Space
Davide Gallon (ETH Zurich), Arnulf Jentzen (Chinese University of Hong Kong)
GenerationOptimizationConvolutional Neural NetworkTransformerDiffusion modelScore-based ModelContrastive LearningImageTabularPhysics RelatedStochastic Differential Equation
🎯 What it does: Developed a physics-informed diffusion model called PISD based on a spectral domain latent space, which can generate solutions and parameters of PDEs during inference by guiding Adam to satisfy PDE constraints and observational conditions.
Physics-Informed Distillation of Diffusion Models for PDE-Constrained Generation
Yi Zhang (University of Hong Kong), Difan Zou (University of Hong Kong)
GenerationData SynthesisOptimizationKnowledge DistillationDiffusion modelScore-based ModelTabularTime SeriesSequentialBenchmarkPhysics Related
🎯 What it does: This paper proposes Physics-Informed Distillation of Diffusion Models (PIDDM), a diffusion model framework that applies PDE constraints directly to the final generated samples through post-training distillation;
Physics-informed Neural Operator Learning for Nonlinear Grad-Shafranov Equation
Siqi Ding (ENN Science and Technology Development Co Ltd), Tianyuan Liu (ENN Science and Technology Development Co Ltd)
TransformerDiffusion modelContrastive LearningTabularPhysics Related
🎯 What it does: A physics-anchored neural operator framework was constructed for predicting plasma equilibrium governed by the nonlinear Grad-Shafranov equation (GSE), achieving real-time inference at the millisecond level.
Physics-Informed Pre-training on Efficient Electron-Density Images for Organic Material Property Prediction
Zhixiang Cheng (Hunan University), xiangxiang Zeng
Computational EfficiencyRepresentation LearningDrug DiscoveryTransformerSupervised Fine-TuningContrastive LearningImageTabularPhysics Related
🎯 What it does: Proposed a pre-training framework called VisionED based on electronic density images for predicting the properties of organic materials.
Physics-Informed Residual Flows
Jephte Abijuru (RPTU University Kaiserslautern-Landau), Sophie Fellenz (RPTU University Kaiserslautern-Landau)
OptimizationConvolutional Neural NetworkTransformerFlow-based ModelTabularTime SeriesPhysics RelatedStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Reformulate physics-informed neural networks (PINN) as residual flows (Residual Flow), proposing architectures such as ResPINN, which solve PDEs through step-by-step fine-tuning of predictions.
Physiology as Language: Translating Respiration to EEG during Sleep
Kaiwen Zha (MIT), Dina Katabi (MIT)
GenerationData SynthesisRepresentation LearningTransformerDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningTime SeriesBiomedical DataElectrocardiogramAudio
🎯 What it does: Generate sleep EEG using respiratory signals and achieve non-contact neural monitoring
Physiology-Aware Masked Cross-Modal Reconstruction for Biosignal Representation Learning
Hao Zhou (Pennsylvania State University), Sharanya Arcot Desai (Samsung Research America)
Anomaly DetectionExplainability and InterpretabilityRepresentation LearningTransformerAuto EncoderContrastive LearningTime SeriesBiomedical DataElectrocardiogram
🎯 What it does: By pretraining PPG representations through masked cross-reconstruction of synchronized ECG and PPG, the generalization performance of single PPG signals in health tasks is enhanced.
Pianist Transformer: Towards Expressive Piano Performance Rendering via Scalable Self-Supervised Pre-Training
Hong-Jie You (Nanjing University), Yu-Feng Li (Nanjing University)
GenerationData SynthesisTransformerLarge Language ModelContrastive LearningSequentialAudio
🎯 What it does: Propose Pianist Transformer, a large-scale self-supervised pre-trained piano performance rendering model that can convert sheet music into human-style performances;
PICACO: Pluralistic In-Context Value Alignment via Total Correlation Optimization
Han Jiang (Johns Hopkins University), Xing Xie (Microsoft Research Asia)
OptimizationRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBenchmark
🎯 What it does: Propose a no-training multi-value alignment method PICACO that optimizes meta-instructions by maximizing total correlation;
PINE: Pruning Boosted Tree Ensembles with Conformal In-Distribution Prediction Equivalence
Haruki Yajima (University of Tokyo), Yusuke Matsui (University of Tokyo)
CompressionExplainability and InterpretabilityComputational EfficiencyMixture of ExpertsContrastive LearningOptical FlowTabularStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Proposed a new tree ensemble pruning method called PINE, which can guarantee prediction equivalence within distributional regions defined after conformal calibration, while achieving higher compression rates.
PINNfluence: Interpreting PINNs through Influence Functions
Aleksander Krasowski (Fraunhofer HHI), René Pascal Klausen (Fraunhofer HHI)
Explainability and InterpretabilityTabularTime SeriesPhysics Related
🎯 What it does: This paper proposes PINNFLUENCE, a training data attribution framework based on influence functions, used to explain the learning and prediction behavior of Physics-Informed Neural Networks (PINNs) with respect to physical constraints;
PinTok: Tokenizers Deserve Dedicated Pinned CPU-Compute and Memory
Sean Choi (Santa Clara University), Ernest K. Ryu (University of California, Los Angeles)
Computational EfficiencyLarge Language ModelText
🎯 What it does: System-level acceleration for LLM tokenizers, achieving PinTok.
PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding
Yunhe Han (Zhejiang University), Yanfeng Zhang (Northeastern University)
Federated LearningComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Designed and implemented the PipeSD framework, which accelerates LLM inference in cloud-edge collaborative inference by utilizing token batch pipeline scheduling and dual-threshold NAV triggering.
PISA: Privacy-Preserving Split Adaptation with Model IP Protection
Haocheng Yang (Beijing University of Posts and Telecommunications), Sen Su (Beijing University of Posts and Telecommunications)
Federated LearningSafty and PrivacyTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningText
🎯 What it does: Designed the PISA framework, achieving a split learning scheme that simultaneously satisfies data privacy (LDP) and model IP protection during LLM fine-tuning.
PISCES: Annotation-free Text-to-Video Post-Training via Optimal Transport-Aligned Rewards
Minh-Quan Le (Microsoft), Mei Chen (Microsoft)
GenerationData SynthesisRepresentation LearningReinforcement Learning from Human FeedbackTransformerReinforcement LearningVision Language ModelDiffusion modelContrastive LearningVideoTextMultimodality
🎯 What it does: This paper proposes an annotation-free text-to-video (T2V) post-training framework called PISCES. After aligning text and video embeddings using OT, it provides denoiser with distributed quality rewards and discrete token-level semantic rewards, significantly improving the visual quality and semantic consistency of the generated videos.
Pix2Key: Controllable Open-Vocabulary Retrieval with Semantic Decomposition and Self-Supervised Visual Dictionary Learning
Guoyizhe Wei (Johns Hopkins University), Yan Gao (Amazon.com)
RetrievalTransformerPrompt EngineeringVision Language ModelAuto EncoderContrastive LearningImageTextMultimodality
🎯 What it does: Propose a zero-training composed image retrieval (CIR) framework called Pix2Key, which utilizes an open-vocabulary visual dictionary to split reference images and natural language edits into positive and negative attributes and open anchors. It unifies them in the text embedding space for nearest neighbor retrieval and incorporates diversity re-ranking. Additionally, it provides a self-supervised visual dictionary autoencoder, V-Dict-AE, to enhance dictionary quality.
PixCLIP: Towards Fine-grained Vision-Language Understanding via Any-granularity Pixel-Text Alignment
YiCheng Xiao, Jinqiao Wang (Institute of Automation Chinese Academy of Sciences)
ClassificationSegmentationRetrievalRepresentation LearningConvolutional Neural NetworkTransformerLarge Language ModelVision Language ModelContrastive LearningImageTextMultimodality
🎯 What it does: Develop the PixCLIP framework to achieve alignment between arbitrary-shaped pixel-level regions and arbitrary-length text, and construct a large-scale long-text mask-aligned dataset called LongGRIT.
PLaID++: A Preference Aligned Language Model for Targeted Inorganic Materials Design
Andy Xu (Harvey Mudd College), Gabriel Hope (Swarthmore College)
GenerationOptimizationDrug DiscoveryReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringDiffusion modelContrastive LearningTextGraphPhysics Related
🎯 What it does: Developed a post-training framework for LLMs called PLaID++, based on Wyckoff text encoding, which utilizes RLIP (Reward-based Direct Preference Optimization with MLIP) to generate thermodynamically stable, unique, and novel inorganic crystal structures.
Plain Transformers are Surprisingly Powerful Link Predictors
Quang Truong (Michigan State University), Jiliang Tang (Michigan State University)
Computational EfficiencyRepresentation LearningGraph Neural NetworkTransformerContrastive LearningGraph
🎯 What it does: Propose PENCIL, a link prediction model based purely on Transformers, which encodes sampled local subgraphs and does not rely at all on manually designed structural heuristics, node ID embeddings, or pre-trained global structural encodings;
Plan for Speed: Dilated Scheduling for Masked Diffusion Language Models
Omer Luxembourg (Ben-Gurion University of the Negev), Eliya Nachmani (Ben-Gurion University of the Negev)
Computational EfficiencyAI Code AssistantTransformerLarge Language ModelPrompt EngineeringDiffusion modelText
🎯 What it does: Propose a sparse decoding scheduling method called Dilated Unmasking Scheduler (DUS), which is only used during inference. By dividing the positions within a block into sparse hierarchical groups and revealing them in parallel, the number of denoising calls per block is reduced from O(B) to O(log B), achieving significant acceleration and quality improvement.
Plan in Sandbox, Navigate in Open Worlds: Learning Physics-Grounded Abstracted Experience for Embodied Navigation
Zhixuan Shen (Southwest Jiaotong University), Tianrui Li (Southwest Jiaotong University)
Robotic IntelligenceTransformerReinforcement LearningPrompt EngineeringVision Language ModelWorld ModelImageTextMultimodalityRetrieval-Augmented Generation
🎯 What it does: Propose the SAGE framework, which utilizes a physically constrained semantic sandbox for self-evolving experience generation and reinforcement learning, and maps the learned abstract strategies to real-world navigation.
Plan Then Action: High-Level Planning Guidance Reinforcement Learning for LLM Reasoning
Zhihao Dou (Case Western Reserve University), Sumon Biswas (Case Western Reserve University)
OptimizationExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextMultimodalityRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose a two-stage Plan-Then-Action framework (PTA-GRPO), where the LLM first generates a concise high-level plan, and then uses this plan to guide fine-grained Chain-of-Thought reasoning;
Plan, Decouple, Assimilate: Physics-Aware Object Insertion in Remote Sensing Imagery
Yingyan Hou (Aerospace Information Research Institute, Chinese Academy of Sciences), Xian Sun (Aerospace Information Research Institute, Chinese Academy of Sciences)
Image TranslationImage HarmonizationGenerationData SynthesisTransformerPrompt EngineeringDiffusion modelAuto EncoderContrastive LearningImagePhysics Related
🎯 What it does: Proposed a physics-aware object insertion framework for remote sensing images called PDA (Plan, Decouple, Assimilate), which achieves high-fidelity synthesis through three stages: planning, decoupling, and assimilation.
Planar Symmetric Pattern Generation
Ning Lin (Renmin University of China), Hao Sun (Renmin University of China)
GenerationData SynthesisOptimizationTransformerPrompt EngineeringDiffusion modelScore-based ModelImageTextPoint CloudBenchmark
🎯 What it does: Propose a framework that transforms any two-dimensional continuous representation into a strictly planar group symmetric and continuous representation.
PLANTAIN: Plan-Answer Interleaved Reasoning
Anthony Liang (Google DeepMind), Jacob Eisenstein (Google DeepMind)
AI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextChain-of-Thought
🎯 What it does: Propose the PLANTAIN framework, which enables post-training language models to perform plan-answer interleaved reasoning, i.e., first generating explicit plans and then continuing the reasoning process, with early user intervention incorporated during the reasoning process;
PLASH: Provably Linear-Time Attention with Selective Higher-Order Feature Sketching
Yuwen Huang (Hong Kong University of Science and Technology), Xiang Pan (Lingnan University)
Computational EfficiencyTransformerLarge Language ModelMixture of ExpertsContrastive LearningTextTime Series
🎯 What it does: Propose the PLASH attention mechanism, which compresses keys and values to a shorter length M, and then enhances them with random high-order features through TensorSketch, while maintaining the standard softmax interface.
Plasticity Activation via Polar Operator: A Plug-in Method for Balancing Stability and Plasticity
Guodong Zheng (Huazhong University of Science and Technology), Li Shen (Sun Yat-sen University)
ClassificationOptimizationComputational EfficiencyMeta LearningTransformerSupervised Fine-TuningContrastive LearningImageTextBenchmark
🎯 What it does: Propose a plugin called PAPO based on the Polar Operator to activate suppressed gradient directions, improving the stability-plasticity balance in continual learning.
PLATE: Plasticity-Tunable Efficient Adapters for Geometry-Aware Continual Learning
Romain Cosentino (Salesforce AI Research)
Computational EfficiencyKnowledge DistillationRepresentation LearningMeta LearningTransformerLarge Language ModelSupervised Fine-TuningImageTextMultimodality
🎯 What it does: Propose a continual learning method called PLATE that does not require access to old task data. It leverages the redundancy of pre-trained networks to construct an input low-energy subspace Q and an output redundant channel selection B, achieving a low-rank adapter ΔW = BAQᵀ.
Platonic Transformers: A Solid Choice For Equivariance
Mohammad Mohaiminul Islam (QurAI, University of Amsterdam), Erik J Bekkers (University of Amsterdam)
GenerationOptimizationDrug DiscoveryProtein Structure PredictionTransformerVision-Language-Action ModelDiffusion modelAuto EncoderContrastive LearningImagePoint CloudGraphTabular
🎯 What it does: Designed a framework that can achieve spatial equivariance within the Transformer architecture
PLoRA: Efficient Concurrent LoRA Training for Large Language Models
Minghao Yan (University of Wisconsin-Madison), Yida Wang (Amazon Web Services)
OptimizationComputational EfficiencyHyperparameter SearchTransformerLarge Language ModelSupervised Fine-TuningText
🎯 What it does: Built a system called PLoRA, which achieves parallel fine-tuning of multiple LoRA adapters on the same GPU, significantly improving hardware utilization and shortening overall training time.
PlotCraft: Pushing the Limits of LLMs for Complex and Interactive Data Visualization
Jiajun Zhang (University of Science and Technology of China), Junyang Lin (Alibaba Group)
AI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningAgentic AIVision Language ModelTextMultimodalityTabularBenchmark
🎯 What it does: Constructed a new complex visualization benchmark PlotCraft, and performed SFT on Qwen3-Coder-30B based on the SynthVis-30K dataset with 30k scale, proposing a lightweight code generation model called PlotCraftor; meanwhile, conducted unified evaluations of 23 LLMs on PlotCraft, VisEval, and PandasPlotBench.
Plug-and-Play Benchmarking of Reinforcement Learning Algorithms for Large-Scale Flow Control
Jannis Becktepe (TU Dortmund University), Sebastian Peitz (TU Dortmund University)
OptimizationReinforcement LearningDiffusion modelAuto EncoderPoint CloudMeshTabularTime SeriesBenchmarkPhysics Related
🎯 What it does: Propose FluidGym, a fully differentiable, CFD-independent reinforcement learning benchmark for active flow control;
Plug-and-Play Diffusion Meets ADMM: Dual-Variable Coupling for Robust Medical Image Reconstruction
Chenhe Du (ShanghaiTech University), Yuyao Zhang (ShanghaiTech University)
RestorationConvolutional Neural NetworkTransformerDiffusion modelBiomedical DataMagnetic Resonance ImagingComputed Tomography
🎯 What it does: Propose Dual-Coupled PnP Diffusion Prior (DC-PnPDP), fully recover dual variables in the traditional PnP-ADMM framework, and introduce the Spectral Homogenization module to eliminate steady-state bias and artifact generation.
Plug-and-Play Guidance for Discrete Diffusion Models via Gradient-Informed Logit Correction
Hongkun Dou (Beihang University), Yue Deng (Beihang University)
GenerationDrug DiscoveryProtein Structure PredictionTransformerReinforcement LearningDiffusion modelScore-based ModelImageTextPoint CloudGraphBiomedical Data
🎯 What it does: Propose a pluggable guidance framework named GILC, which corrects logits using a pre-trained discrete diffusion model and gradient information, thereby achieving controllable generation without requiring additional training.
Plug-and-Play Label Map Diffusion for Universal Goal-Oriented Navigation
Zhixuan Shen (Southwest Jiaotong University), Haonan Luo (Southwest Jiaotong University)
Autonomous DrivingOptimizationReinforcement Learning from Human FeedbackConvolutional Neural NetworkTransformerSupervised Fine-TuningDiffusion modelScore-based ModelContrastive LearningImagePoint CloudTabular
🎯 What it does: Proposed a label map completion method called PLMD based on diffusion models, used to complete obstacles and semantic labels in partially observed environments, thus assisting goal-oriented navigation.
Plug-and-Play Spiking Operators: Breaking the Nonlinearity Bottleneck in Spiking Transformers
Xinzhe Yuan (Harbin Institute of Technology), Huan Xiong (Harbin Institute of Technology)
Computational EfficiencyKnowledge DistillationSpiking Neural NetworkTransformerLarge Language ModelRectified FlowContrastive LearningText
🎯 What it does: This paper proposes a pluggable implementation of nonlinear operations based on LIF neuron groups, supporting key nonlinear operations in Transformers such as Softmax, SiLU, and RMSNorm, achieving ANN-to-SNN conversion without additional training.
PlugGuard: A Streaming Safeguard for Large Models via Latent Dynamics-Guided Risk Detection
Xiaodan Li (Alibaba Group), Hui Xue (Alibaba Group)
Safty and PrivacyComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextMultimodalityBenchmark
🎯 What it does: This paper proposes a streaming risk detection plugin called PlugGuard, which can assess and intercept harmful content in real-time during large model generation.
PlugMem: A Task-Agnostic Plugin Memory Module for LLM Agents
Ke Yang (University of Illinois Urbana Champaign), ChengXiang Zhai (University of Illinois Urbana Champaign)
Explainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelTextGraphBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This paper proposes PLUGMEM, a pluggable and task-agnostic memory module that can abstract the long-term experience of large language model agents into propositional and procedural knowledge, organizing it into a knowledge graph.
Pluralistic Leaderboards
Nika Haghtalab (University of California Berkeley), Kunhe Yang (University of California Berkeley)
Recommendation SystemReinforcement Learning from Human FeedbackLarge Language ModelTextBenchmark
🎯 What it does: Design a leaderboard mechanism that takes into account the diverse preferences of users, called pluralistic leaderboard;
PluRel: Synthetic Data unlocks Scaling Laws for Relational Foundation Models
Vignesh Kothapalli (Stanford University), Jure Leskovec (Stanford University)
Data SynthesisGraph Neural NetworkTransformerDiffusion modelScore-based ModelGenerative Adversarial NetworkContrastive LearningTabular
🎯 What it does: Propose the PLUREL framework to synthesize multi-table relational databases from scratch for training relational foundation models.
PMSPO: Progressive Matching and Semantic-Aware Policy Optimization for Camouflaged Object Detection
Maosheng Su (Huazhong Agricultural University), Jun Luo (Huazhong Agricultural University)
Object DetectionTransformerLarge Language ModelReinforcement LearningVision Language ModelContrastive LearningImageMultimodality
🎯 What it does: A multi-modal large language model based on reinforcement learning is proposed, named PMSPO framework, to achieve covert target detection.
PnP-Corrector: A Universal Correction Framework for Coupled Spatiotemporal Forecasting
Hao Wu (Tsinghua University), Yuan Gao (Tsinghua University)
Convolutional Neural NetworkRecurrent Neural NetworkGraph Neural NetworkDiffusion modelAuto EncoderContrastive LearningOptical FlowImageVideoTabularTime SeriesPhysics RelatedStochastic Differential Equation
🎯 What it does: Proposed the PnP-Corrector framework and the DSLCast model to address the problem of recursive error amplification in coupled spatiotemporal prediction.
PODiff: Latent Diffusion in Proper Orthogonal Decomposition Space for Scientific Super-Resolution
Onkar Jadhav (University of Western Australia), Nicole L. Jones (University of Western Australia)
Super ResolutionConvolutional Neural NetworkDiffusion modelScore-based ModelAuto EncoderContrastive LearningImageTime SeriesPhysics Related
🎯 What it does: A probabilistic super-resolution framework called PODDiff is proposed, which performs conditional diffusion in the orthogonal low-rank POD coefficient space, aiming to efficiently generate high-resolution scientific fields and their uncertainties.
POET-X: Memory-efficient LLM Training by Scaling Orthogonal Transformation
Zeju Qiu, Weiyang Liu (Chinese University of Hong Kong)
Computational EfficiencyTransformerLarge Language ModelText
🎯 What it does: Proposed a memory-efficient training framework called POET-X, which utilizes scalable orthogonal equivalent transformations to train large-scale language models.
PointCHR: Point Cloud Analysis via Curvature-Aware Hyperbolic Rectification
Xinxing Yu (Macau University of Science and Technology), Yanyan Liang (Macau University of Science and Technology)
ClassificationSegmentationGraph Neural NetworkAuto EncoderContrastive LearningPoint Cloud
🎯 What it does: Propose a curvature-aware hyperbolic remapping (PointCHR) that alleviates the representation congestion problem in Euclidean space by expanding the embedding space at points with high curvature.
PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation
Haofei Xu (Google), Michael Niemeyer (Google)
Depth EstimationTransformerDiffusion modelScore-based ModelFlow-based ModelImagePoint Cloud
🎯 What it does: This paper proposes PointDiT, a point cloud diffusion model based on the pixel space, which directly infers 3D point clouds from monocular RGB images using ViT, completely eliminating the need for VAE and complex hybrid networks;
Poison with Style: A Practical Poisoning Attack on Code Large Language Models
Khang Tran (New Jersey Institute of Technology), Md Rizwan Parvez (Qatar Computing Research Institute)
Adversarial AttackAI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringText
🎯 What it does: Proposed a passive backdoor attack based on code style, leveraging code style as a covert trigger to make large code language models generate code containing vulnerabilities when triggered.
PolarDepth: Monocular Transparent Object Depth from Polar-Physics Priors
Wen Dong (Dalian University of Technology), Xin Yang (Dalian University of Technology)
SegmentationDepth EstimationTransformerVision-Language-Action ModelContrastive LearningImagePoint CloudPhysics Related
🎯 What it does: A monocular depth estimation framework called PolarDepth was developed, which utilizes polarization information (DoLP, AoLP) to extract physical priors (refractive index, polar angle, azimuthal angle) from transparent objects and injects them into geometric representations. Combined with transparent object segmentation, it achieves depth recovery and refinement.
Polaris: Coupled Orbital Polar Embeddings for Hierarchical Concept Learning
Sahil Mishra (Indian Institute of Technology Delhi), Tanmoy Chakraborty (Indian Institute of Technology Delhi)
RetrievalRepresentation LearningGraph Neural NetworkContrastive LearningImageTextMultimodalityGraph
🎯 What it does: Propose Polaris, a polar coordinate embedding framework that separates semantic directions and hierarchical radii on a unit high-dimensional sphere, for hierarchical structure learning and expansion under weak supervision.
POLCA: Stochastic Generative Optimization with LLM
Xuanfei Ren (University of Wisconsin Madison), Ching-An Cheng (Google Research)
OptimizationHyperparameter SearchAI Code AssistantTransformerLarge Language ModelTextBenchmarkRetrieval-Augmented GenerationStochastic Differential Equation
🎯 What it does: Propose an scalable LLM-based random generation optimization framework called POLCA, which automatically searches for optimal parameters under uncertain feedback using a priority queue and an ε-grid filter.
POLIA: Policy Optimization with Visual-Object-Level Intrinsic Advantage for Multimodal Reasoning
Yiran Zeng (South China University of Technology), Mengchen Zhao (South China University of Technology)
OptimizationTransformerReinforcement LearningVision Language ModelMultimodalityBenchmarkChain-of-Thought
🎯 What it does: This paper proposes POLIA, a group-based reinforcement learning method that introduces visual object-level internal advantages in multi-modal reasoning, achieving finer-grained credit assignment.
Policies Permitting LLM Use for Polishing Peer Reviews Are Currently Not Enforceable
Rounak Saha (Indian Institute of Science), Danish Pruthi (Indian Institute of Science)
ClassificationAnomaly DetectionTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextReview/Survey PaperRetrieval-Augmented Generation
🎯 What it does: This paper constructs a peer review text dataset containing five levels of AI collaboration, evaluates the identification performance of various existing AI text detectors (including commercial and open-source ones) on these reviews, and explores the feasibility of improving detection by leveraging signals such as paper context, AI-generated reference text similarity, and writing style features.
Policy Search via Bayesian Optimization with Temporal Difference Gaussian Processes
Armin Lederer (National University of Singapore), Andreas Krause (ETH Zurich)
OptimizationReinforcement LearningTabularTime Series
🎯 What it does: A Bayesian optimization framework utilizing Gaussian processes for temporal difference learning is proposed, aimed at low-dimensional policy search.
Policy-Driven World Model Adaptation for Robust Offline Model-based Reinforcement Learning
Jiayu Chen (University of Hong Kong), Jeff Schneider (Carnegie Mellon University)
OptimizationRobotic IntelligenceReinforcement LearningWorld ModelTabularTime SeriesSequentialBenchmark
🎯 What it does: Proposes an offline model-based reinforcement learning framework (ROMBRL) with online dynamic adaptation of the world model, aiming to maximize the worst-case return of the policy under an uncertain model, thereby enhancing deployment robustness.
PolicyGuard: Towards Test-time and Step-level Adversary Defense for Reinforcement Learning Agent
Junfeng Guo (University of Maryland College Park), Heng Huang (University of Maryland College Park)
Anomaly DetectionAdversarial AttackConvolutional Neural NetworkRecurrent Neural NetworkReinforcement LearningAuto EncoderContrastive LearningSequential
🎯 What it does: Proposes a test-time, step-level defense method called PolicyGuard, which uses the posterior variance of Gaussian Processes (GP) to detect backdoor behaviors triggered by reinforcement learning agents during deployment.
PolyFlow: Safe and Efficient Polytope-Constrained Flow Matching with Constraint Embedding and Projection-free Update
jianming ma, Yue Gao (Shanghai Jiao Tong University)
OptimizationRobotic IntelligenceTransformerDiffusion modelFlow-based ModelTabularTime SeriesSequentialOrdinary Differential Equation
🎯 What it does: Propose PolyFlow, a discrete-time, projection-free flow matching framework for generating safe distributions under polyhedral constraints;
Polyphonia: Zero-Shot Timbre Transfer in Polyphonic Music with Acoustic-Informed Attention Calibration
Haowen Li (South China University of Technology), Qi Liu (South China University of Technology)
GenerationData SynthesisTransformerPrompt EngineeringDiffusion modelContrastive LearningMultimodalityAudio
🎯 What it does: Propose a zero-shot polyphonic music timbre transfer framework called Polyphonia based on unsupervised attention calibration, which can accurately change the timbre of the target track to a specified instrument while keeping the accompaniment unchanged;
PolySAE: Modeling Feature Interactions in Sparse Autoencoders via Polynomial Decoding
Panagiotis Koromilas (Cyprus Institute), Mihalis Nicolaou (University of Cyprus)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerAuto EncoderContrastive LearningText
🎯 What it does: This paper proposes PolySAE, a model that introduces a polynomial decoder into sparse autoencoders (SAE) to capture interactions between features.
PoMtVRS: Preference-Optimized Multi-Task Vehicle Routing Solver with Preference Gating
Dian Meng (Dalian University of Technology), Zhiguang Cao (Singapore Management University)
Autonomous DrivingOptimizationReinforcement Learning from Human FeedbackTransformerReinforcement LearningMixture of ExpertsContrastive LearningGraph
🎯 What it does: Proposed a pluggable Preference-Optimized Multi-Task Vehicle Routing Solver (PoMtVRS), which enhances the representation capability and exploration efficiency of multi-task VRP solvers by inserting a Preference-Gated Block (PGB) into the Transformer decoder and adopting Preference Optimization (PO).
PonderLM-2: Pretraining LLM with Latent Thoughts in Continuous Space
Boyi Zeng (Shanghai Jiao Tong University), Zhouhan Lin (Shanghai Jiao Tong University)
Computational EfficiencyRepresentation LearningData-Centric LearningTransformerLarge Language ModelTextChain-of-Thought
🎯 What it does: Propose that during the pre-training phase, the LLM first generates intermediate latent thinking states, and then uses these states to predict the next word, thereby performing multi-step reasoning in a continuous space.
Population-Aware Imitation Learning in Mean-field Games with Common Noise
Grégoire Lambrecht (New York University), Mathieu Lauriere
OptimizationAdversarial AttackReinforcement LearningContrastive LearningTabularSequentialStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: This paper studies the method of imitation learning in mean field games with common noise, and proposes theoretical analysis and algorithmic frameworks.
Population-Free Pareto Tracking for Sample-Efficient Multi-Policy MORL
Zeyu Zhao (Shenzhen University), Junmei Yao (Shenzhen University)
Reinforcement LearningTabularTime Series
🎯 What it does: Propose a multi-strategy non-population Pareto front tracking framework (MPFT), which starts from a single-objective extreme point strategy, uses Pareto reverse and improvement directions to achieve local tracking, and finally generates a high-quality set of strategies that approximate the complete Pareto front.
PortraitRL: Reinforcement Learning for Personalized Portrait Pose Transfer with Multi-Objective Reward Modeling
Jiahui Wu (Renmin University of China), Zhiwu Lu (Renmin University of China)
Image TranslationPose EstimationReinforcement Learning from Human FeedbackTransformerReinforcement LearningPrompt EngineeringVision Language ModelScore-based ModelFlow-based ModelImageMultimodality
🎯 What it does: Propose a post-training framework called PortraitRL based on reinforcement learning, which utilizes multi-objective rewards and negative preference scores to improve detail preservation and instruction following in portrait pose transfer.
Pose-ICL: 3D-Aware In-Context Learning for Pose-Controllable Subject Customization
Xuan Han (Tongji University), Mingyu You (Tongji University)
GenerationData SynthesisPose EstimationTransformerSupervised Fine-TuningPrompt EngineeringDiffusion modelAuto EncoderContrastive LearningImagePoint CloudMesh
🎯 What it does: Proposes Pose-ICL, a parameter-free, 3D perception-based In-Context Learning framework that achieves pose-controllable subject customization with the help of multi-pose reference images through Surface-Anchored Position Embedding (SAPE).
Position Is All You Need: A Free Lunch Token Compression Strategy for MLLM-based Referring Expression Segmentation
Yuhan Liu (Huazhong University of Science and Technology), Ruixuan Li (Huazhong University of Science and Technology)
SegmentationCompressionComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextMultimodality
🎯 What it does: This paper proposes a training-free compression strategy called PAYN, which relies solely on position information to address the excessive computational cost of multimodal large language models in the task of referring expression segmentation.
Position: 'AI Alignment' Encompasses Competing Technical Priorities
Tushita Jha (Mimir Center for Long Term Futures Research), Mateusz Bagiński (AFFINE)
Explainability and InterpretabilityReview/Survey Paper
🎯 What it does: The authors clarify the concept of AI alignment, pointing out that it has multiple meanings, and propose three core ideals (Task Reliability, Social Judiciousness, Takeover Avoidance). They explore the conflicts and trade-offs among these ideals in terms of their goal attributes and objects.
Position: *Beyond Text* The Text-Centric Bias in Foundation Models Must Be Revisited for a Speech-First Future
Deepak Piskala
RecognitionRepresentation LearningData-Centric LearningTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextMultimodalityReview/Survey PaperRetrieval-Augmented GenerationAudio
🎯 What it does: Discusses and advocates for a shift from text-dominated models to audio-centric 'speech-first' foundational models, elaborating on technical feasibility, user habit barriers, and future research directions.
Position: A Dynamical Systems Perspective is Needed to Advance Time Series Modeling
Daniel Durstewitz (Central Institute of Mental Health), Lukas Eisenmann (Central Institute of Mental Health)
Anomaly DetectionExplainability and InterpretabilityComputational EfficiencyRepresentation LearningRecurrent Neural NetworkTransformerSupervised Fine-TuningContrastive LearningTabularTime SeriesBiomedical DataReview/Survey PaperPhysics RelatedStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: This paper proposes introducing a dynamical systems (DS) perspective into time series modeling and provides a review of DS reconstruction (DSR) methods, exploring how they can enhance short-term forecasting, recovery of long-term statistical features, and prediction of tipping points.
Position: Academic Conferences are Potentially Facing Denominator Gaming Caused by Fully Automated Scientific Agents
Rong Shan (Shanghai Jiao Tong University), Jianghao Lin (Shanghai Jiao Tong University)
Recommendation SystemFederated LearningExplainability and InterpretabilityComputational EfficiencyData-Centric LearningAI Code AssistantReinforcement Learning from Human FeedbackNeural Architecture SearchTransformerLarge Language ModelAgentic AIPrompt EngineeringDiffusion modelScore-based ModelFlow-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningTextReview/Survey PaperBenchmarkPhysics RelatedRetrieval-Augmented GenerationStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: This paper describes and analyzes the threat of 'denominator manipulation' in academic conference submission systems, proposing that malicious parties can use fully automated scientific agents to generate and submit a large number of low-quality papers, artificially increasing the number of submissions to a conference, thereby increasing the acceptance probability of target papers while maintaining a fixed acceptance rate.
Position: Accountable Deployment of Agentic AI Demands Layered, System-Level Interpretability
Judy Zhu (Vector Institute for Artificial Intelligence), Shaina Raza (Vector Institute for Artificial Intelligence)
Explainability and InterpretabilityAgentic AITextTabularTime SeriesReview/Survey Paper
🎯 What it does: Propose the ATLIS framework, constructing a system-level explainability stack for Agentic AI, covering all stages of the lifecycle and multi-level explainability needs
Position: Adopting AI in Practice Does Not Guarantee the Productivity Boost
Won Ik Cho (Samsung Electronics), Geunhye Kim (Hankuk University of Foreign Studies)
OptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackReview/Survey PaperFinance Related
🎯 What it does: This paper proposes that AI in organizational practices may not necessarily enhance productivity, identifies five major moderating factors related to human and environmental aspects, and extends the partial equilibrium model of Gries & Naudé;
Position: Adversarial ML for LLMs Is Not Making Any Progress
Javier Rando (Anthropic), Florian Tramèr (ETH Zurich)
Explainability and InterpretabilityAdversarial AttackTransformerLarge Language ModelPrompt EngineeringTextReview/Survey PaperRetrieval-Augmented Generation
🎯 What it does: This paper systematically reviews recent adversarial machine learning (Adversarial ML) research conducted in large language models (LLM) in the form of a position paper, pointing out three core challenges faced by existing research: unclear definitions, difficult solving, and strict evaluation. It elaborates on how these issues hinder technological progress through several cases (such as jailbreak, prompt injection, Membership Inference, Unlearning, etc.).
Position: Age Estimation Models Do Not Process Biometric Data
Nikita Marshalkin (Sumsub GmbH)
RecognitionSafty and PrivacyConvolutional Neural NetworkTransformerContrastive LearningImageBenchmark
🎯 What it does: This paper experimentally verifies that age estimation models are almost unable to distinguish individuals at any layer through hierarchical embedding evaluation of 14 age estimation models on three facial verification benchmarks, indicating that they lack identification capabilities, thus answering whether they process biometric data.
Position: Agent Evaluation Should Be Agentified for Openness, Standardization, and Reproducibility
Xiaoyuan Liu (UC Berkeley), Dawn Song (UC Berkeley)
Autonomous DrivingFederated LearningSafty and PrivacyExplainability and InterpretabilityComputational EfficiencyRobotic IntelligenceReinforcement LearningAgentic AIPrompt EngineeringTextTabularTime SeriesSequentialBenchmark
🎯 What it does: Proposes the Agentified Agent Assessment (AAA) framework, treating the evaluation process itself as an agent, enabling the evaluator and the evaluated agent to communicate through standardized A2A (Agent-to-Agent) and MCP (Model Context Protocol) protocols, completely separating the evaluator from the evaluated agent, reducing the original integration cost from N×M to N+M, supporting multi-agent evaluation, openness, and reproducibility;
Position: Agent Security Needs Redefinition through a Holistic Framework
Vincent Siu (University of California Santa Cruz), Dawn Song (University of California Berkeley)
Safty and PrivacyExplainability and InterpretabilityTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: The paper proposes shifting the focus of AI agent safety from solely concentrating on the content of instructions to context-based safety evaluation. It constructs four core attributes—Source Authorization, Task Alignment, Action Alignment, and Data Isolation—and emphasizes the necessity of continuously checking these attributes throughout the agent's lifecycle.
Position: Agentic AI Orchestration Should Be Bayes-Consistent
Theodore Papamarkou (Polyshape), Alexey Zaytsev (Bimsa)
Explainability and InterpretabilityAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AITextReview/Survey Paper
🎯 What it does: Discusses embedding Bayesian decision theory into the control layer of agent-based AI systems driven by large language models to achieve normalized selection under uncertainty and cost.
Position: Agentic AI System Is a Foreseeable Pathway to AGI
Junwei Liao (Shanghai Jiao Tong University), Weinan Zhang (Shanghai Jiao Tong University)
Computational EfficiencyRepresentation LearningMeta LearningReinforcement Learning from Human FeedbackGraph Neural NetworkTransformerAgentic AIContrastive Learning
🎯 What it does: Propose and theoretically prove that Agentic AI (multi-agent collaborative system) has higher sample and parameter efficiency than a single large model when achieving AGI, and formalize it as a DAG structure;
Position: Agentic Safety is an Epistemic Property, Not a Behavioral One
Charles L. Wang (Columbia University), Peter Jin (Columbia University)
Safty and PrivacyReinforcement LearningAgentic AITabularReview/Survey Paper
🎯 What it does: Propose defining the safety of self-modifying agents as epistemic property, emphasizing future correctability;
Position: Agentic Systems Should be General
Elron Bandel (IBM Research), Leshem Choshen (IBM Research)
Meta LearningAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextMultimodalityBenchmark
🎯 What it does: This paper calls for the construction of a general intelligent agent system that can flexibly adapt to new environments and operate across terminals, networks, physical environments, and other diverse settings;
Position: Agents Should Invoke External Tools ONLY When Epistemically Necessary
Hongru WANG, Kam-Fai Wong (Chinese University of Hong Kong)
Explainability and InterpretabilityRobotic IntelligenceMeta LearningReinforcement LearningAgentic AIReview/Survey Paper
🎯 What it does: Propose a theoretical framework called Theory of Agent (ToA), which argues that agents only invoke external tools when cognitively necessary, clarifying the roles and boundaries of internal reasoning and external interaction in knowledge acquisition.
Position: AGI Requires a Coordination Layer on Top of Pattern Repositories
Edward Y Chang
Autonomous DrivingFederated LearningSafty and PrivacyExplainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringDiffusion modelScore-based ModelContrastive LearningTextTabularTime SeriesSequentialBenchmarkRetrieval-Augmented GenerationChain-of-ThoughtStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: The paper proposes constructing a system-2 coordination layer on top of large language models (LLMs), arguing that a pattern repository is a necessary subsystem, while the missing system-2 coordination layer is a bottleneck for achieving artificial general intelligence (AGI).
Position: AI Capabilities May Not Increase Exponentially
Haosen Ge (University of Pennsylvania), Osbert Bastani (University of Pennsylvania)
Explainability and InterpretabilityComputational EfficiencyData-Centric LearningLarge Language ModelTabularTime SeriesBenchmark
🎯 What it does: This paper questions whether the growth of AI capabilities is exponential, and proposes a stacked model that separates the capabilities of foundational models from reasoning capabilities, aiming to explain the exponential growth observed in the recent METR report;
Position: AI Evaluation Should Work With Humans
Jan Kulveit (Charles University), Raymond Douglas (Charles University)
Explainability and InterpretabilityReinforcement Learning from Human FeedbackReinforcement LearningPrompt EngineeringTextReview/Survey PaperChain-of-Thought
🎯 What it does: Propose shifting AI evaluation from merely replacing humans to assessing the collaborative capabilities of human-AI teams, explain the limitations of current evaluation paradigms, and provide a theoretical framework, metrics, and practical recommendations for future collaborative evaluation.
Position: AI Evaluations Should be Grounded on a Theory of Capability
Nathanael Jo (Massachusetts Institute of Technology), Ashia C. Wilson (Massachusetts Institute of Technology)
Explainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This paper proposes to view AI evaluation as an inference task based on explicit ability theory, and demonstrates through experiments that different theoretical assumptions significantly affect evaluation results.
Position: AI for Science Should Treat Measurement-to-Dataset Pipelines as Inference Components
Ling Zhan (Southwest University), Tao Jia (Southwest University)
Explainability and InterpretabilityComputational EfficiencyData-Centric LearningDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackNeural Architecture SearchProtein Structure PredictionBiomedical DataAlzheimer's DiseaseElectrocardiogramReview/Survey PaperAudio
🎯 What it does: This paper proposes that in the AI for Science (AI4Science) workflow, the pipeline for measuring data should be regarded as an inference component, rather than a fixed data interface, and this view is validated through a large-scale audit in the field of EEG functional connectomics.
Position: AI Governance Needs ISO-like Interoperability Protocols, Not Just Laws
Azmine Toushik Wasi (Computational Intelligence and Operations Laboratory (CIOL), Bangladesh), Dong-Kyu Chae (Hanyang University)
Safty and PrivacyExplainability and InterpretabilityTextReview/Survey PaperRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This paper proposes a framework for AI risk statements (AI nutritional labels) based on ISO specifications, aiming to provide interoperable, machine-readable risk communication and compliance credentials for cross-border AI regulation.
Position: AI Leaderboards Are Underserving the Global South: A Case Study from India
Sourav Banerjee (Indian Institute of Technology Kharagpur), Saikat Saha (Nasscom)
TextReview/Survey PaperBenchmarkAgriculture RelatedAudio
🎯 What it does: The paper points out that AI leaderboards are structurally unable to effectively serve the Global South, and uses India as a case to propose the need for region-specific leaderboards with independent governance, conflict-of-interest policies, and mechanisms for indicator evolution.
Position: AI Lock-In Is in Progress, and We Must Be Prepared
Jaeho Kim (Korea University), Changhee Lee (Korea University)
Review/Survey Paper
🎯 What it does: Proposes the concept of AI Lock‑In, explains its risks at the individual, organizational, and national levels, and calls for preventive and mitigation measures.
Position: AI Must Become Planet-Centered, Not Just Human-Centered
Maria Perez-Ortiz (University College London)
OptimizationFederated LearningSafty and PrivacyExplainability and InterpretabilityComputational EfficiencyDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackNeural Architecture SearchProtein Structure PredictionImageTextTabularTime SeriesReview/Survey PaperBenchmarkAgriculture RelatedFinance RelatedPhysics Related
🎯 What it does: Proposes a novel AI design philosophy called Planet-Centered AI (PCAI), advocating for the integration of AI research with the complexity, systemic risks, and long-term impacts of Earth systems;
Position: AI Researchers Must Help Lead Arms Control to Mitigate Military AI Risks
Ted Fujimoto (Pacific Northwest National Laboratory), Jacob Benz (Pacific Northwest National Laboratory)
Explainability and InterpretabilityReview/Survey Paper
🎯 What it does: The authors propose that AI researchers should take the lead in controlling military AI to prevent risks associated with military AI.
Position: AI Should Facilitate Democratic Deliberation at Scale
José Ramón Enríquez (Stanford), Alex Pentland (Stanford)
Recommendation SystemAutonomous DrivingOptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyData-Centric LearningRobotic IntelligenceMeta LearningDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackNeural Architecture SearchProtein Structure PredictionTransformerLarge Language ModelPrompt EngineeringTextMultimodalityReview/Survey PaperRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: The paper proposes using AI to assist in democratic discussions, aiming to reduce cognitive and social friction and promote large-scale democratic deliberation.
Position: AI Usage Policies Should Be Aligned with International Human Rights Law
Jordi Calvet-Bademunt (Vanderbilt University)
Safty and PrivacyExplainability and InterpretabilityLarge Language ModelTextReview/Survey Paper
🎯 What it does: Propose a framework for aligning AI usage policies with international human rights law (ICCPR Article 19), and use this framework to evaluate the usage policies of eight major AI service providers.
Position: AI Welfare Is Bullshit
Yunze Xiao (Carnegie Mellon University), Mona T. Diab (Carnegie Mellon University)
Recommendation SystemSafty and PrivacyExplainability and InterpretabilityAI Code AssistantReinforcement Learning from Human FeedbackReview/Survey Paper
🎯 What it does: Analyze and criticize current AI welfare evaluation methods, pointing out their structural defects such as lack of co-design with systems and external validation, and argue that AI welfare should not be used as a governance threshold.
Position: AI/ML Deepfake Research is Misaligned with AI Generated Non-Consensual Intimate Imagery (AIG-NCII)
Li Qiwei (University of Michigan), Eric Gilbert (University of Michigan)
Safty and PrivacyExplainability and InterpretabilityDiffusion modelGenerative Adversarial NetworkImageTextReview/Survey Paper
🎯 What it does: Conduct a systematic literature review of deepfake-related papers published in top conferences from 2020 to 2025, evaluating the attention paid to AI-generated non-consensual intimate imagery (AIG-NCII), and proposing a critical analysis of the mismatch between technical interventions and harm to victims' dignity.
Position: Anthropomorphic Misalignment Research Needs Stronger Evidence
Vansh Gupta (ETH Zurich), Anna Hedström (ETH Zurich)
Explainability and InterpretabilityData-Centric LearningReview/Survey Paper
🎯 What it does: This paper evaluates methodological flaws in humanized misalignment (AMR) research and proposes improvement strategies.
Position: Artificial Intelligence Needs Meta Intelligence - the Case for Metacognitive AI
Sergei Chuprov (University of Texas Rio Grande Valley), Dmitrii Korobeinikov (Rochester Institute of Technology)
OptimizationFederated LearningSafty and PrivacyComputational EfficiencyHyperparameter SearchAgentic AIPrompt EngineeringImageBiomedical Data
🎯 What it does: This paper proposes using metacognition (self-monitoring and resource allocation) as a general design principle for AI, and implements metacognitive mechanisms in the federated learning (FL) scenario through a framework called InteFL, demonstrating how metacognitive control can enhance learning efficiency, robustness, and security.