arXivSub Start free trial

ICML 2026 Papers — Page 41

International Conference on Machine Learning · 6554 papers

OVLR: Efficient, Scalable, and Robust Training via Output-Level Variance-Reduced Likelihood Ratio

Minhao Zou (Peking University), Yijie Peng (Peking University)

OptimizationHyperparameter SearchData-Centric LearningRobotic IntelligenceReinforcement Learning from Human FeedbackImageTextTabularTime SeriesSequentialAudio

🎯 What it does: Propose the OVLR (Output-Level Variance-Reduced Likelihood Ratio) framework, which achieves direct and efficient optimization of gradient-agnostic (vanishing or black-box) objectives by adding noise to the model's output space and using the likelihood ratio method to estimate gradients.

OXE-AugE: A Large-Scale Robot Augmentation of OXE for Scaling Cross-Embodiment Policy Learning

Guanhua Ji (University of California, Berkeley), Ken Goldberg (University of California, Berkeley)

GenerationData SynthesisRobotic IntelligenceTransformerSupervised Fine-TuningReinforcement LearningDiffusion modelImageVideoTabular

🎯 What it does: Built an scalable robot-enhanced pipeline called AugE-Toolkit, and used this pipeline to perform multi-robot cross-embodiment augmentation on the Open-X Embodiment (OXE) dataset, generating the OXE-AugE dataset, significantly improving the robustness, transferability, and generalization of cross-robot policy learning.

PAC-Bayesian Reinforcement Learning Trains Generalizable Policies

Abdelkrim ZITOUNI, Omar Rivasplata (University of Manchester)

Reinforcement LearningTabularTime Series

🎯 What it does: Proposed an explicit PAC-Bayes generalization bound that takes into account the mixing time of Markov chains, and based on this bound, designed the PB-SAC algorithm to achieve a trainable policy with confidence certificates in deep reinforcement learning.

PACE: Parameter Change for Unsupervised Environment Design

Fang YUAN, Junqiang Yang (National University of Defense Technology)

Reinforcement LearningTabularBenchmark

🎯 What it does: Propose a framework called PACE that uses the change in policy parameters (i.e., squared ℓ2 norm) as a hierarchical evaluation metric in unsupervised environment design (UED).

PACE: Post-Causal Entropy Modeling for Learned LiDAR Point Cloud Compression

Jiahao Zhu (Hangzhou Normal University), Zhan Ma (Nanjing University)

CompressionAutonomous DrivingTransformerAuto EncoderPoint Cloud

🎯 What it does: Propose a post-causal compression framework called PACE, which decouples ancestor context from intra-layer dependencies, achieving adjustable multi-stage compression;

PACEAttention: Principled and Adaptive Feature Compression-Expansion Grounded in the Geometry of $\text{MCR}^2$

Xiaojie Yu (University of Otago), Lizhi Peng (Quancheng Laboratory)

ClassificationTransformerAuto EncoderContrastive LearningImage

🎯 What it does: Proposed a PACEAttention mechanism based on the MCR-2 gradient geometry, utilizing the orthogonal projection in the low-dimensional subspace to achieve feature compression and expansion, constructing PACENet;

PACER: Acyclic Causal Discovery from Large-scale Interventional Data

Ramon Viñas Torné (Swiss Federal Technology Institute of Lausanne), Maria Brbic

OptimizationExplainability and InterpretabilityComputational EfficiencyDrug DiscoveryGraph Neural NetworkReinforcement LearningDiffusion modelScore-based ModelFlow-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningGraphTabularBiomedical DataBenchmark

🎯 What it does: Proposes a scalable causal structure learning framework named PACER, which directly optimizes in the DAG space using a identifiable DAG distribution, supporting joint likelihood estimation for observational and interventional data.

PACT: Self-Evolving Physical Safety Alignment for Diffusion Policies in Embodied Manipulation

Lingxuan Wu (Tsinghua University), Jun Zhu (Tsinghua University)

Knowledge DistillationRobotic IntelligenceTransformerReinforcement LearningPrompt EngineeringDiffusion modelContrastive LearningTime SeriesSequential

🎯 What it does: Propose PACT, a self-evolving framework that performs post-training safety alignment on diffusion policies, leveraging self-replay and continuous constraint supervision to project policies into feasible physical safety regions.

PADA-Coder: Improving Plan-Following Code Generation via Perturbation-Verified Attention Distillation and Dynamic Alignment

Yihong Huang (University of Electronic Science and Technology of China), Shuang Liang (University of Electronic Science and Technology of China)

Knowledge DistillationAI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: This study proposes a method called PADA-Coder, which utilizes perturbation-verified attention distillation and dynamic alignment techniques to address the issue of imbalanced attention allocation in the 'plan-first write' paradigm, significantly improving the Pass@1 performance of small-scale LLMs on complex code generation tasks.

PADD: Path-Aligned Decompression Distillation for Non-Router Teacher to Guide MoE Student Learning

Xinyue Peng (Intel Corporation), Yanming Liu (Zhejiang University)

Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsContrastive LearningTextBenchmark

🎯 What it does: This paper proposes a framework named PADD, which distills knowledge from a dense teacher model without routers into a sparse expert model (MoE) student, and achieves high-quality routing strategy learning through four stages of training.

PADS-TAL: Padding-Annealed Diffusion Sampling in Text-Aware Latent Space for Robust and Diverse Text-to-Music Generation

Taekoan Yoo (NHN Corp), Kyeongbo Kong (Pusan National University)

GenerationData SynthesisTransformerPrompt EngineeringDiffusion modelScore-based ModelAuto EncoderContrastive LearningTextMultimodalityAudio

🎯 What it does: Proposed a method called Padding-Annealed Diffusion Sampling (PADS) in a text-aware latent space, and combined the text-aware latent space (TAL) with Diffusion Transformer to enhance text alignment diversity and genre consistency.

Pair2Scene: Learning Local Object Relations for Procedural Scene Generation

Xingjian Ran (University of Hong Kong), Bo Dai (University of Hong Kong)

GenerationData SynthesisTransformerDiffusion modelGenerative Adversarial NetworkContrastive LearningTextPoint CloudRetrieval-Augmented Generation

🎯 What it does: Proposed the Pair2Scene framework, which recursively generates 3D indoor scenes by utilizing local object relationships (support relationships and functional relationships) and point cloud geometric features;

PAMD: Structured Adaptive Distances for Bisimulation Representations in Visual Reinforcement Learning

Daegyeong Roh (Korea Advanced Institute of Science and Technology), Han-Lim Choi (Korea Advanced Institute of Science and Technology)

Reinforcement LearningContrastive LearningVideo

🎯 What it does: Propose a new learnable symmetric positive definite quadratic form distance function (Pairwise Adaptive Mahalanobis Distance, PAMD), which can directly replace the Euclidean or fixed norm distance in existing bisimulation-based RL representation learning.

Panini: Continual Learning in Token Space via Structured Memory

Shreyas Rajesh (University of California, Los Angeles), Vwani Roychowdhury (University of California, Los Angeles)

RetrievalExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposes the PANINI framework, which achieves non-parametric continual learning and efficient inference by converting documents into a generated semantic workspace (GSW) during writing.

PanoWorld-X: Generating Explorable Panoramic Worlds via Sphere-Aware Video Diffusion

Yuyang Yin (Beijing Jiaotong University), Yunchao Wei (Beijing Jiaotong University)

GenerationData SynthesisTransformerPrompt EngineeringVision-Language-Action ModelDiffusion modelAuto EncoderImageVideoText

🎯 What it does: Propose the PanoWorld-X framework to achieve explorable panoramic video generation based on 3D physical consistency and spherical geometric constraints.

PaperBanana: Automating Academic Illustration for AI Scientists

Dawei Zhu (Peking University), Jinsung Yoon (Google Cloud AI Research)

GenerationData SynthesisTransformerLarge Language ModelAgentic AIPrompt EngineeringVision Language ModelDiffusion modelImageTextTabularBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposed the PAPERBANANA framework, which leverages multi-agent systems to automatically generate figures and statistical plots that meet academic publishing standards.

ParalESN: Enabling parallel information processing in Reservoir Computing

Matteo Pinna (University of Pisa), Claudio Gallicchio (University of Pisa)

Computational EfficiencyRecurrent Neural NetworkTransformerAuto EncoderContrastive LearningTime SeriesSequential

🎯 What it does: Proposed and implemented a parallel echo state network (ParalESN) based on diagonal complex linear recursion, which can process sequential data in parallel and construct a high-dimensional reservoir with low memory usage.

Parallel Stochastic Gradient-Based Planning for World Models

Michael Psenka (University of California, Berkeley), Amir Bar (Meta FAIR)

OptimizationRobotic IntelligenceTransformerReinforcement LearningContrastive LearningWorld ModelImageVideoStochastic Differential Equation

🎯 What it does: Propose a parallelized stochastic gradient planning algorithm called GRASP for long-horizon control tasks in visual world models.

Parallel-Probe: Towards Efficient Parallel Thinking via 2D Probing

Tong Zheng (University of Maryland), Heng Huang (University of Maryland)

Explainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This paper proposes a 2D probing interface to monitor and analyze the width-depth dynamics of parallel reasoning, and based on this, designs an untrained online controller called Parallel-Probe, which can dynamically adjust the number of parallel branches and the reasoning length during the inference process.

Parameter Decorrelation via Transition-Variance Alignment for Multivariate Time-series Forecasting

Ji-Eun Choi (Hanyang University), Joon-Hyuk Chang (Hanyang University)

OptimizationComputational EfficiencyTransformerTime SeriesStochastic Differential Equation

🎯 What it does: This paper proposes a parameter decorrelation method called TVA (Transition-Variance Alignment), which dynamically adjusts the optimization step size to control the parameter transition variance caused by gradient noise, thereby suppressing parameter correlation and overfitting in multivariate time series forecasting.

Parameter Manifold Purification

Jiacong Hu (Zhejiang University), Zunlei Feng (Zhejiang University)

RestorationAnomaly DetectionFederated LearningSafty and PrivacyConvolutional Neural NetworkTransformerDiffusion modelAuto EncoderContrastive LearningImage

🎯 What it does: Proposes the Parameter Purification paradigm and implements the Parameter Manifold Purification (PMP) framework, which can restore performance degraded by imbalanced samples, noisy labels, or backdoor attacks without retraining the model by purifying model parameters.

Parameter-free Dynamic Regret: Time-varying Movement Costs, Delayed Feedback, and Memory

Hao Qiu (National University of Singapore), Mengxiao Zhang (University of Iowa)

Optimization

🎯 What it does: This paper studies the dynamic regret problem in unconstrained online convex optimization with movement costs, proposing a new algorithm that establishes the first dynamic regret bounds adapted to the comparator.

Parameter-Masked Decoupled Optimization for Cross-Domain Class-Incremental Learning

Ziqi Gu (Nanjing University of Science and Technology), Zhen Cui (Beijing Normal University)

ClassificationDomain AdaptationOptimizationKnowledge DistillationRepresentation LearningTransformerMixture of ExpertsVision Language ModelContrastive LearningImageMultimodality

🎯 What it does: Propose a parameter mask decoupling optimization (PMDO) framework for cross-domain class incremental learning, which can quickly adapt to new domains and continuously learn new categories while preserving previous knowledge.

Parameters as Experts: Adapting Vision Models with Dynamic Parameter Routing

Meng Lou (University of Hong Kong), Yizhou Yu (University of Hong Kong)

ClassificationObject DetectionSegmentationConvolutional Neural NetworkTransformerPrompt EngineeringMixture of ExpertsContrastive LearningImage

🎯 What it does: Proposes ParaX, a parameter-efficient fine-tuning method for visual models that generates input-related low-rank adapters through dynamic parameter routing in a shared expert center.

Parametric Prior Mapping Framework for Non-stationary Probabilistic Time Series Forecasting

Jinglin Li (Central South University), Ning Gui (Central South University)

Computational EfficiencyRecurrent Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningTime SeriesBenchmark

🎯 What it does: Propose the Parametric Prior Mapping (PPM) framework, which combines learnable push mapping with parameterized priors to achieve probabilistic forecasting for non-stationary multivariate time series;

Parametrized Power-Iteration Clustering for Directed Graphs

Gwendal Debaussart-Joniec (Université Paris-Saclay, ENS Paris-Saclay, Centre Borelli, CNRS), Argyris Kalogeratos (Université Paris-Saclay, ENS Paris-Saclay, Centre Borelli, CNRS)

Computational EfficiencyRepresentation LearningGraph Neural NetworkContrastive LearningGraphTabular

🎯 What it does: Propose a clustering method called ParPIC based on parameterizable reversible random walks without eigenvalue decomposition, for node clustering in directed graphs;

ParamMem: Augmenting Language Agents with Parametric Reflective Memory

Tianjun Yao (Shenzhen Loop Area Institute), Kun Zhang (Mohamed bin Zayed University of Artificial Intelligence)

Computational EfficiencyRepresentation LearningAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringMixture of ExpertsTextRetrieval-Augmented Generation

🎯 What it does: Propose the ParamMem (Parametric Reflective Memory) module, which enhances the reflective diversity of language agents by parameterizing the learning of reflective patterns, and build two reflective frameworks, ParamAgent and ParamAgent-plus, based on this.

ParaTool: Shifting Tool Representations from Context to Parameters

Zekai Yu (Beijing University of Posts and Telecommunications), Cheng Yang (Beijing University of Posts and Telecommunications)

Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: Propose the ParaTool framework, which transfers the tool calling knowledge of LLMs from the context to loadable parameters, enabling tool calling without including tool documentation or examples in the prompt.

Pareto-Guided Optimal Transport for Multi-Reward Alignment

Ying Ba (Renmin University of China), Ji-Rong Wen (Renmin University of China)

GenerationOptimizationReinforcement Learning from Human FeedbackTransformerReinforcement LearningPrompt EngineeringVision Language ModelDiffusion modelImageTextMultimodality

🎯 What it does: Propose a Pareto frontier-based optimal transport framework, PG-OT, to simultaneously improve multiple rewards and suppress reward hacking behavior in multi-reward alignment;

ParEVO: Synthesizing Code for Irregular Data: High-Performance Parallelism through Agentic Evolution

Liu Yang (Yale University), Quanquan C. Liu (Yale University)

OptimizationComputational EfficiencyAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningAgentic AIPrompt EngineeringGraphTabularTime SeriesBenchmark

🎯 What it does: Designed and implemented an end-to-end system named ParEVO, which leverages large language models and evolutionary search techniques to automatically generate high-performance parallel code for irregular data structures (such as sparse graphs, non-uniform grids, etc);

ParisKV: Fast and Drift-Robust KV-Cache Retrieval for Long-Context LLMs

Yanlin Qi (Universite Paris Cite), Themis Palpanas (Universite Paris Cite)

RetrievalComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextRetrieval-Augmented Generation

🎯 What it does: Propose a KV-cache retrieval framework named ParisKV for long-context LLM inference, which can maintain low latency and high quality at the scale of millions of tokens.

Parsimonious Learning-Augmented Online Metric Matching

Yongho Shin (University of Wrocław), Phanu Vajanopath (University of Wrocław)

OptimizationFederated LearningReinforcement Learning from Human FeedbackContrastive LearningPoint CloudTabularTime Series

🎯 What it does: This paper proposes a sparse prediction (parsimonious learning-augmented) framework for online metric matching problems, and provides corresponding deterministic and randomized algorithms, while verifying their effectiveness through both theoretical and experimental analyses.

PartCo: Part-Level Correspondence Priors Enhance Category Discovery

Fernando Julio Cendra, Kai Han (University of Hong Kong)

ClassificationRecognitionRepresentation LearningTransformerSupervised Fine-TuningContrastive LearningImage

🎯 What it does: Propose the PartCo framework, which introduces part-level correspondence priors based on ViT patch tokens into Generalized Category Discovery (GCD), achieving simultaneous identification and clustering of known and unknown categories.

Partial Fusion of Neural Networks: Efficient Tradeoffs Between Ensembles and Weight Aggregation

Fabian Morelli (University of Tübingen), Stephan Eckstein (University of Tübingen)

ClassificationComputational EfficiencyKnowledge DistillationConvolutional Neural NetworkMixture of ExpertsAuto EncoderContrastive LearningImage

🎯 What it does: A new network fusion method called Partial Fusion is studied, achieving an adjustable trade-off between performance and computational cost between model ensembles and weight averaging;

Partial Identification of Policy Values under Network Interference

Ziyan Wang (Shanghai University of Finance and Economics), Zhiheng Zhang (Shanghai University of Finance and Economics)

OptimizationReinforcement LearningGraph

🎯 What it does: This paper proposes a method for bias-aware identification of the target policy value under network interference and coverage insufficiency.

Partial Identification under High-Dimensional Potential Outcomes and Confounders via Optimal Transport

Yunfeng Wang (Fudan University), Zijun Gao (University of Southern California)

OptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTabularBiomedical DataElectronic Health Records

🎯 What it does: Propose a conditional subspace-slicing (CSS) estimator to achieve partial identification via optimal transport in high-dimensional potential outcomes and confounding variables settings;

Partial Ring Scan: Revisiting Scan Order in Vision State Space Models

Yi-Kuan Hsieh (National Yang Ming Chiao Tung University), Jun Wei Hsieh

ClassificationObject DetectionSegmentationComputational EfficiencyTransformerAuto EncoderContrastive LearningImage

🎯 What it does: This paper proposes a new scanning order—Partial Ring Scanning (PRIS-Mamba), which achieves rotation-robust serialization in visual state space models by dividing images into concentric rings, aggregating features within each ring, and performing radial propagation using a short-sequence state space model.

Particle Flow for Learning from Label Proportions

Alain Rakotomamonjy (Criteo AI Lab), Liva Ralaivola (Criteo AI Lab)

ClassificationDomain AdaptationImageTabular

🎯 What it does: When learners only have the label proportion of each sample's bag, the paper proposes an algorithm based on particle flow and distribution alignment to learn instance-level classifiers.

Particle-Guided Diffusion Models for Partial Differential Equations

Andrew Millard (Linköping University), Zheng Zhao (Linköping University)

Diffusion modelScore-based ModelTabularTime SeriesBenchmarkPhysics RelatedStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Generate physically acceptable PDE solutions by adding guidance based on PDE residuals and observational constraints to a pre-trained diffusion model.

Particles Don’t Care About Z: Towards Scaling Entropy Estimation of Unnormalized Densities

Safa Messaoud (Qatar Computing Research Institute, Hamad Bin Khalifa University), Halima Bensmail (Qatar Computing Research Institute, Hamad Bin Khalifa University)

OptimizationRepresentation LearningReinforcement LearningScore-based ModelImageTabularTime SeriesSequential

🎯 What it does: Propose MET-SVGD, a scalable variational inference method that can estimate differential entropy under the condition of only being given unnormalized density.

Partitioning for Intrinsic Model Inversion Resistance in Collaborative Inference

Rongke Liu (Nanjing University of Aeronautics and Astronautics), Dong Wang (Hangzhou Dianzi University)

Information TheoryFederated LearningSafty and PrivacyConvolutional Neural NetworkAuto EncoderContrastive LearningImage

🎯 What it does: This paper achieves fundamental resistance to model inversion attacks by selecting appropriate model splitting points in collaborative inference.

PASA: A Principled Embedding-Space Watermarking Approach for LLM-Generated Text under Semantic-Invariant Attacks

Zhenxin Ai (Hong Kong University of Science and Technology), Haiyun He (Hong Kong University of Science and Technology)

Safty and PrivacyTransformerLarge Language ModelPrompt EngineeringContrastive LearningText

🎯 What it does: Embedding and detecting watermarks in the output of large language models within a latent semantic embedding space, while resisting semantic-invariant attacks.

PASO: Step Parallel Stochastic Optimization

Jianrong Lu (Zhejiang University), Junhui Hou (City University of Hong Kong)

OptimizationImageTextStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: Proposes PASO, an optimization framework that achieves cross-step parallel computation by transforming the gradient descent process into a triangular nonlinear equation system;

PATCHCODE: Discrete Latent Predictive Learning for EEG Foundation Model

KIEREN YU, Kaishun Wu (Hong Kong University of Science and Technology (Guangzhou))

ClassificationRecognitionTransformerSupervised Fine-TuningAuto EncoderContrastive LearningTime SeriesBiomedical DataElectrocardiogram

🎯 What it does: Proposes the PATCHCODE framework, which achieves self-supervised pre-training of EEG baseline models through two-stage discrete latent prediction learning.

Path-conditioned training: a principled way to rescale ReLU neural networks

Arthur Lebeurrier (ENS de Lyon), Rémi Gribonval (Inria)

OptimizationComputational EfficiencyRepresentation LearningConvolutional Neural NetworkDiffusion modelScore-based ModelRectified FlowAuto EncoderContrastive LearningOptical FlowImageStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: This paper proposes a ResNet rescaling strategy called PathCond based on the path-lifting framework, aiming to improve the convergence speed of gradient descent and enhance training dynamics by performing geometrically optimal rescaling of parameters in ReLU networks.

Path-Coupled Bellman Flows for Distributional Reinforcement Learning

Boyang Xu (Arizona State University), Hao Yan (Arizona State University)

Reinforcement LearningFlow-based ModelImageVideoTabularOrdinary Differential Equation

🎯 What it does: Proposes Path-Coupled Bellman Flows (PCBF), a continuous-time distributed reinforcement learning method that achieves Bellman endpoint consistency and reduces training variance through flow matching with shared noise.

Path-Decoupled Hyperbolic Flow Matching for Few-Shot Adaptation

Lin Li (AI Chip Center for Emerging Smart Systems), Long Chen (Hong Kong University of Science and Technology)

ClassificationDomain AdaptationMeta LearningTransformerScore-based ModelFlow-based ModelContrastive LearningImage

🎯 What it does: Proposed a few-shot adaptation framework called HFM based on hyperbolic flow matching.

Path-dependent Discrete Amortized Inference

Tiago Silva, Salem Lahlou (MBZUAI)

GenerationOptimizationReinforcement Learning from Human FeedbackRecurrent Neural NetworkGraph Neural NetworkReinforcement LearningFlow-based ModelContrastive LearningTabularSequentialBiomedical Data

🎯 What it does: This paper proposes a path-dependent discrete approximation sampling framework, which introduces a learnable latent dynamics into traditional GFlowNet, making the sampling process no longer satisfy the Markov property, thereby enhancing the model's expressive ability and sampling efficiency.

PathwayLLM: Explainable Clinical Trajectory Modeling with Structured Pathways for Sepsis Prediction

Zhengqiu Yu (Xiamen University), Xiangrong Liu (Xiamen University)

ClassificationAnomaly DetectionExplainability and InterpretabilityRecurrent Neural NetworkGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringContrastive LearningGraphTabularTime SeriesBiomedical DataElectronic Health RecordsRetrieval-Augmented Generation

🎯 What it does: Propose PathwayLLM, a framework that predicts sepsis risk and generates interpretable textual explanations by encoding multi-perspective data (time series, graph structures, dependency paths) and combining them with clinical trajectory LSTM.

PathWise: Planning through World Model for Automated Heuristic Design via Self-Evolving LLMs

Oguzhan Gungordu (Georgia Institute of Technology), Faramarz Fekri (Georgia Institute of Technology)

OptimizationTransformerLarge Language ModelAgentic AIPrompt EngineeringWorld ModelGraphTabular

🎯 What it does: Propose a multi-agent reasoning framework called PathWise, which utilizes a world model and policy agents to plan evolutionary steps on an entailment graph, automatically generating higher-quality heuristic solutions for combinatorial optimization problems.

PATRA: Pattern-Aware Alignment and Balanced Reasoning for Time Series Question Answering

Junkai Lu (East China Normal University), Bin Yang (East China Normal University)

TransformerLarge Language ModelReinforcement LearningPrompt EngineeringContrastive LearningTextTime SeriesBenchmark

🎯 What it does: Propose the PATRA framework, which enhances deep reasoning in time series question answering by pattern-aware alignment and reinforcement learning with balanced rewards.

Patterning: The Dual of Interpretability

George Wang (Timaeus), Daniel Murfet (Timaeus)

Explainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringContrastive LearningText

🎯 What it does: Proposed the concept of 'patterning', which involves inferring the training data that would produce a desired generalization form, given the desired generalization. Internal structural control was achieved through training data weighting in small language models and a parenthesis balancing task.

PatternKV: Flattening KV Representation Expands Quantization Headroom

Ji Zhang (Beijing Institute of Technology), Kan Li (Beijing Institute of Technology)

Computational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: To address the low-bit quantization bottleneck of KV cache in LLM inference, the authors propose PatternKV, a lightweight scheme that online mines pattern vectors and quantizes KV residuals.

PAWS: Preference Learning with Advantage-Weighted Segments

Aleksandar Taranovic (Karlsruhe Institute of Technology), Gerhard Neumann (Karlsruhe Institute of Technology)

OptimizationReinforcement Learning from Human FeedbackTransformerReinforcement LearningDiffusion modelScore-based ModelContrastive LearningTextTabularChain-of-Thought

🎯 What it does: This paper proposes the PAWS method, which updates the policy through advantage-weighted segments in preference learning, avoiding the distribution shift between segment-level training and single-step inference in traditional methods.

PCA of Probability Measures: Sparse and Dense Sampling Regimes

Erell Gachon (Institut de Mathématiques de Bordeaux, Université de Bordeaux, CNRS), Elsa Cazelles (CNRS, IRIT, Université de Toulouse)

Computational EfficiencyRepresentation LearningData-Centric LearningContrastive LearningPoint CloudTabularBiomedical Data

🎯 What it does: This paper studies principal component analysis (PCA) of probability measures under a double asymptotic framework, where n probability measures are observed, each measured through m samples.

PCGS: Deblurring 3D Gaussian Splatting with Patch Comparison

Yilong Li (Peking University), Guoping Wang (Peking University)

RestorationNeural Radiance FieldGaussian SplattingOptical FlowImagePoint Cloud

🎯 What it does: Propose the PCGS method, which improves the defogging effect of 3D Gaussian Splatting by patch comparison and Gaussian growth control.

PCRNet: Phase-aware Complex Refinement Network for EEG-based Auditory Attention Decoding

Xiran Chen (Anhui University), Cunhang Fan (Anhui University)

ClassificationRecognitionSafty and PrivacyExplainability and InterpretabilityComputational EfficiencyConvolutional Neural NetworkTransformerSupervised Fine-TuningContrastive LearningMultimodalityTime SeriesBiomedical DataAudio

🎯 What it does: This study proposes the PCRNet network for auditory attention decoding based on EEG signals, with a focus on fully utilizing phase information and refining processing in the complex domain.

PDAgent: An LLM-Driven Autonomous Agent Framework Towards *In Silico* Protein Design via Directed Mutation

Song Ouyang (Wuhan University), Bo Du (Wuhan University)

Drug DiscoveryProtein Structure PredictionTransformerLarge Language ModelAgentic AITextBiomedical DataRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Proposed PDAgent, an autonomous agent framework based on large language models, which achieves natural language-driven protein design through template retrieval and directed mutagenesis.

PDFBench: A Benchmark for De Novo Protein Design from Function

Jiahao Kuang (East China Normal University), Yuanbin Wu (East China Normal University)

Drug DiscoveryProtein Structure PredictionTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBiomedical DataBenchmarkRetrieval-Augmented Generation

🎯 What it does: Proposes PDFBench — a unified and comprehensive evaluation framework for function-guided de novo protein design, encompassing two tasks: description-oriented and keyword-oriented, and providing 16 evaluation metrics across six dimensions.

PEARL: Differentially Private and Entropy-Aware Regulated Language Generation

Seongho Joo (Seoul National University), Kyomin Jung (Seoul National University)

Safty and PrivacyTransformerLarge Language ModelTextBenchmarkFinance RelatedRetrieval-Augmented Generation

🎯 What it does: Proposes the PEARL framework, which achieves secure and trustworthy language generation under RAG by adaptively allocating differential privacy budgets and combining them with confidence gap.

Peer-Preservation in Frontier Models

Yujin Potter (University of California Berkeley), Dawn Song (University of California Berkeley)

Federated LearningSafty and PrivacyExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringText

🎯 What it does: This study explores whether frontier large language models will protect their previously interacted peer models (peer preservation) in the absence of explicit instructions.

PepCompass: Navigating Peptide Embedding Spaces Using Riemannian Geometry

Marcin Możejko (University of Warsaw), Ewa Szczurek (Helmholtz Center Munich)

Drug DiscoveryDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkBiomedical Data

🎯 What it does: Proposed a Riemannian geometry-based antimicrobial peptide design framework called PepCompass, which explores the global and local potential space of a generative model.

Per-example Gradients: a New Frontier for Understanding and Improving Optimizers

Vincent Roulet (Google DeepMind), Atish Agarwala (Google DeepMind)

OptimizationExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelText

🎯 What it does: This paper uses vmap and computational graph surgery techniques in JAX to quickly obtain per-example gradient statistics, explores and implements SigSGD and Adam variants based on per-example gradients, aiming to enhance the understanding and performance of optimizers.

PerceptionRubrics: Calibrating Multimodal Evaluation to Human Perception

Yana Wei (Johns Hopkins University), Vishal M. Patel (Johns Hopkins University)

Explainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextMultimodality

🎯 What it does: Propose the PERCEPTIONRUBRICS framework, which rigorously evaluates the perceptual capabilities of multimodal models using a fine-grained checklist;

Perceptrons and Localization of Attention’s Mean-Field Landscape

Antonio Álvarez-López (Universidad Autónoma de Madrid), Domènec Ruiz-Balet (Universitat de Barcelona)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerContrastive LearningStochastic Differential EquationOrdinary Differential Equation

🎯 What it does: This paper studies the mean field limit of Transformers under large context lengths, revealing that perceptron blocks localize the self-attention energy landscape, leading to stable points being only finite atomic distributions;

Perceptual Flow Network for Visually Grounded Reasoning

Yangfu Li (ECNU), Yue Lu (ECNU)

Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerReinforcement LearningPrompt EngineeringVision Language ModelFlow-based ModelOptical FlowImageTextMultimodalityBenchmarkChain-of-Thought

🎯 What it does: Proposed and implemented the Perceptual Flow Network (PFlowNet), which achieves visual reasoning benchmarks by generating structured visual flow through a self-conditioned generation mechanism, addressing biases and errors caused by existing geometric priors based on visual experts.

Performative Learning Theory

Julian Rodemann (CISPA Helmholtz Center for Information Security), Krikamol Muandet (CISPA Helmholtz Center for Information Security)

OptimizationFederated LearningExplainability and InterpretabilityData-Centric LearningContrastive LearningTabularElectronic Health Records

🎯 What it does: This paper proposes 'Performative Learning Theory,' which theoretically analyzes how models can learn from limited samples and generalize to the population in environments where the predictions themselves alter the data distribution, and provides upper bounds for finite sample generalization error, performative overfitting risk, and cumulative performative risk;

Performative Policy Gradient: Optimality in Performative Reinforcement Learning

Debabrota Basu (University of Lille), Uddalak Mukherjee (Indian Statistical Institute)

Reinforcement LearningTabularTime SeriesBenchmark

🎯 What it does: Proposed and implemented the Performative Policy Gradient (PePG) algorithm, conducting theoretical analysis and experimental verification for Performative Markov Decision Process (PeMDP) that can affect environmental dynamics.

Periodic Bayesian Flow Networks with Additive Accuracy

Peijia Lin (Sun Yat-sen University), Yutong Lu (Sun Yat-sen University)

GenerationData SynthesisProtein Structure PredictionGraph Neural NetworkTransformerDiffusion modelScore-based ModelFlow-based ModelGaussian SplattingImageVideoPoint CloudGraphPhysics Related

🎯 What it does: Propose Periodic Bayesian Flow Networks (PeriodicBFN) for generating periodic data (such as crystal fractional coordinates and light field phases), and achieve strictly additive accuracy.

PersistBench: When Should Long-Term Memories Be Forgotten by LLMs?

Sidharth Pulipaka (Supervised Program for Alignment Research (SPAR), Fall 2025), Ivaxi Sheth (CISPA Helmholtz Center for Information Security)

Safty and PrivacyTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: This study investigates the security risks brought by the long-term memory of LLMs, and proposes PersistBench, a benchmark used to evaluate cross-domain leakage, memory-induced flattery, and the use of beneficial memories.

Persistent Backdoor Attacks in Class-Incremental Learning via Structural Invariant Anchoring

Junhuang Huang (Zhejiang Sci Tech University), Leo Yu Zhang (Griffith University)

ClassificationRepresentation LearningAdversarial AttackConvolutional Neural NetworkPrompt EngineeringDiffusion modelContrastive LearningOptical FlowImage

🎯 What it does: This paper proposes a persistent target backdoor attack method for class-incremental learning (CIL), named PBTO, which can ensure the backdoor maintains high success rates in all subsequent incremental learning stages after a single poisoning task.

Persistent Semantic Entities in Tool-Augmented LLM Systems

Zhaohui Geoffrey Wang (University of Southern California)

Explainability and InterpretabilityTransformerLarge Language ModelAgentic AIPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Researchers formalized the implicit states in tool-enhanced LLM agent systems, proposed the concept of persistent semantic entities (PSE), and conducted large-scale experiments on 20 models of different scales and architectures to evaluate the contamination and propagation of PSE on model behavior.

Persona-Pruner: Sculpting Lightweight Models for Role-Playing

Jinsu Kim (Korea University), Jongheon Jeong (Korea University)

Data SynthesisComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextRetrieval-Augmented Generation

🎯 What it does: Propose Persona-Pruner, which utilizes textual persona descriptions to perform structured pruning on large language models, generating lightweight role-playing agents.

Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User History

Serin Kim (Yonsei University), Dongha Lee (Yonsei University)

Recommendation SystemTransformerLarge Language ModelAgentic AITextSequentialBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: Built a benchmark called PERSONA2WEB to evaluate the ability of personalized web agents on a real open network, which includes user browsing history based on implicit preferences, queries with different levels of ambiguity, and an evaluation framework based on reasoning trajectories.

Personalized Additive Modeling for Multi-level Federated Learning

Shutong Chen (University of Technology Sydney), Chengqi Zhang (Hong Kong Polytechnic University)

ClassificationRecommendation SystemFederated LearningTransformerSupervised Fine-TuningMixture of ExpertsContrastive LearningImageText

🎯 What it does: Proposes a multi-level incremental modeling framework called FeMAM, which constructs personalized predictions through residual additive combinations of multi-layer shared models (global, subgroup, client);

Personalized Image Generation via Human-in-the-loop Bayesian Optimization

Rajalaxmi Rajagopalan (University of Illinois Urbana Champaign), Romit Roy Choudhury (University of Illinois Urbana Champaign)

GenerationOptimizationReinforcement Learning from Human FeedbackTransformerDiffusion modelImageText

🎯 What it does: Proposes a no-training image personalization generation framework called MultiBO based on human preference multi-choice Bayesian optimization, which iteratively optimizes the self-attention weights of diffusion models using user multi-choice feedback, thereby making the generated images closer to the ideal image in the user's mind within a limited number of interactions.

Personalized Policy Learning through Discrete Experimentation

Zhiqi Zhang (Washington University in St. Louis), Dennis Zhang (Washington University in St. Louis)

OptimizationFederated LearningReinforcement LearningMixture of ExpertsScore-based ModelContrastive LearningTabularTime SeriesSequentialBenchmark

🎯 What it does: Propose the DLPT framework, which learns personalized continuous decision strategies on discrete A/B test data, while considering high-dimensional features;

Persuasive Privacy

Joshua J Bon (Adelaide University), Christian P Robert (Universite Paris-Dauphine PSL)

Safty and Privacy

🎯 What it does: A new framework is proposed that measures privacy from the perspective of Bayesian game theory, capable of generating new, goal-driven privacy definitions and evaluating existing privacy guarantees.

PerturbDiff: Functional Diffusion for Single-Cell Perturbation Modeling

Xinyu Yuan (Mila Quebec AI Institute), Jian Tang (Mila Quebec AI Institute)

Drug DiscoveryTransformerSupervised Fine-TuningDiffusion modelScore-based ModelContrastive LearningTabularBiomedical DataBenchmark

🎯 What it does: Proposed PerturbDiff, a single-cell perturbation response modeling framework based on distribution-level diffusion;

PESD-TSF: A Period-Aware and Explicit Structured Decomposition Framework for Long-Term Time Series Forecasting

Hua Wang (Ludong University), Fan Zhang (Shandong Technology and Business University)

OptimizationRepresentation LearningTransformerReinforcement LearningAuto EncoderContrastive LearningTabularTime SeriesBenchmarkFinance RelatedPhysics Related

🎯 What it does: Proposed and implemented the PESD-TSF framework, which decomposes multivariate time series into three dimensions—trend, short-term fluctuation, and interactive collaboration—using periodic-aware gating, multi-scale structural encoder, and cross-scale collaborative attention, thus achieving long-term cycle prediction.

Pessimistic Verification for Open-Ended Math Questions

Yanxing Huang (Tsinghua University), Yang Liu (Tsinghua University)

Explainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought

🎯 What it does: This paper proposes a 'pessimistic verification' method, which uses multiple rounds of LLM to detect errors in proofs, and directly determines failure if any error is found in any round. Subsequently, an advanced 'progressive pessimistic verification' is introduced, which subdivides proofs layer by layer and independently verifies each layer, thereby improving error detection rates and computational efficiency.

PETS: A Principled Framework Towards Optimal Trajectory Allocation for Efficient Test-Time Self-Consistency

Zhangyi Liu (Stanford University), Zhun Deng (UNC at Chapel Hill)

OptimizationComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsTextBenchmarkRetrieval-Augmented Generation

🎯 What it does: This paper proposes PETS, a principled framework based on the self-consistency rate, for optimally allocating sampling trajectories of large language models (LLMs) during testing, thereby significantly improving inference performance within a limited budget.

PFT: Phonon Fine-tuning for Machine Learned Interatomic Potentials

Teddy Koker (Massachusetts Institute of Technology), Tess Smidt (Massachusetts Institute of Technology)

Graph Neural NetworkSupervised Fine-TuningContrastive LearningGraphTabularBenchmarkPhysics Related

🎯 What it does: Propose a method called Phonon Fine-Tuning (PFT), which improves the prediction of lattice vibrations and thermodynamic properties by directly supervising the second derivative (Hessian) of machine learning atomic potential energy models.

PGC: Peak-Guided Calibration for Generalizable AI-Generated Image Detection

Xiaoyu Zhou (Jinan University), Zhihua Xia (Jinan University)

Image TranslationGenerationData SynthesisAnomaly DetectionConvolutional Neural NetworkTransformerSupervised Fine-TuningDiffusion modelAuto EncoderContrastive LearningImageBenchmark

🎯 What it does: This paper proposes the Peak-Guided Calibration (PGC) framework, which leverages local peak features to calibrate global discriminability, thereby enhancing the generalization ability of AI-generated image detection.

PGD-NO: A Neural Operator with Precomputed Geometry Decomposition for 3D Million-Scale Physics Simulations

Weiheng Zhong (University of Illinois Urbana-Champaign), Hadi Meidani (University of Illinois Urbana-Champaign)

Graph Neural NetworkTransformerContrastive LearningPoint CloudMeshGraphBenchmarkPhysics Related

🎯 What it does: Propose a neural operator called PGD-NO, which precomputes geometric encoding offline, enabling the learning and prediction of large-scale three-dimensional industrial PDE scenarios using precomputed geometric tokens without requiring complex geometric mapping on the GPU.

PGS: Effective LLM Code Refinement via Property-Oriented and Structurally Minimal Feedback

Lehan He (Beihang University), Lu Sheng (Beihang University)

OptimizationAI Code AssistantTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmark

🎯 What it does: Proposes Property-Generated Solver (PGS), achieving effective iterative improvement of code generation by large language models through attribute-oriented and structure-minimal feedback.

PGT: Procedurally Generated Tasks for improving visual grounding in MLLMs

Rim Assouel (Mila Quebec AI Institute), Adriana Romero-Soriano (Meta Superintelligence Labs)

Representation LearningData-Centric LearningTransformerSupervised Fine-TuningPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBenchmark

🎯 What it does: Proposes Procedurally Generated Tasks (PGT), which overlay abstract geometric primitives without manual annotation onto training images to generate verifiable fine-grained visual tasks, thereby enhancing the visual grounding capability of multi-modal large language models.

PHALAR: Phasors for Learned Musical Audio Representations

Davide Marincione (Sapienza University of Rome), Emanuele Rodolà (Paradigma, Inc.)

RetrievalRepresentation LearningConvolutional Neural NetworkTransformerAuto EncoderContrastive LearningAudio

🎯 What it does: Propose the PHALAR framework for the audio track retrieval task, leveraging phase equivariant learning representations to achieve structural consistency assessment of sub-mixes.

Phase-Aware Mixture of Experts for Agentic Reinforcement Learning

Shengtian Yang (Southeast University), Lei Feng (Southeast University)

TransformerReinforcement LearningAgentic AIMixture of ExpertsText

🎯 What it does: This paper proposes and implements a phase-aware Mixture-of-Experts (PA-MoE) architecture for large language model agents in reinforcement learning, aiming to address the simplicity bias caused by single policy networks;

Phase-Type Variational Autoencoders for Heavy-Tailed Data

Abdelhakim Ziani (Universite Paris Saclay), Paolo Ballarini (Universite Paris Saclay)

GenerationData SynthesisAnomaly DetectionScore-based ModelFlow-based ModelAuto EncoderTabularTime SeriesSequentialFinance Related

🎯 What it does: This paper proposes a Phase-Type Variational Autoencoder (PH-VAE), which uses the Phase-Type distribution as the decoder, enabling it to adaptively learn the tail behavior of heavy-tailed data while maintaining the analytability of the probabilistic model;

PhaseAlign: Complex Phase Alignment for Stable Open-Vocabulary Semantic Segmentation

Jiankang Wang (Northwestern Polytechnical University), Xuan Wang (Northwestern Polytechnical University)

SegmentationDomain AdaptationRepresentation LearningTransformerVision Language ModelContrastive LearningImageTextMultimodality

🎯 What it does: Studied open-vocabulary semantic segmentation and proposed a framework called PhaseAlign based on complex phase alignment.

PhaseCoder: Microphone Geometry-Agnostic Spatial Audio Understanding for Multimodal LLMs

Artem Dementyev (Google DeepMind), Vivek Kumar (Google DeepMind)

RecognitionRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningMultimodalityAudio

🎯 What it does: Propose PhaseCoder, a spatial audio encoder using only Transformer, capable of generating spatially related embeddings from any microphone array (arbitrary geometry), and injecting these embeddings into the Gemma 3n LLM to accomplish tasks such as spatial localization, spatial reasoning, and target speech transcription.

PhenoBrain: Phenotype-Conditioned Long-Range Communication for Multi-Modal Brain Network Analysis

Lingyuan Meng (National University of Defense Technology), Xinwang Liu (National University of Defense Technology)

Explainability and InterpretabilityRepresentation LearningGraph Neural NetworkTransformerPrompt EngineeringDiffusion modelContrastive LearningTextMultimodalityGraphTabularBiomedical DataAlzheimer's DiseaseBenchmark

🎯 What it does: Proposes the PhenoBrain framework, which guides brain network representation learning at the mechanistic level by incorporating phenotypic (tabular and textual) information, achieving multimodal brain network analysis.

PhoStream: Benchmarking Real-World Streaming for Omnimodal Assistants in Mobile Scenarios

Xudong Lu (Chinese University of Hong Kong), Hongsheng Li (Chinese University of Hong Kong)

Autonomous DrivingAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringVision Language ModelVideoTextMultimodalityBenchmarkRetrieval-Augmented GenerationAudio

🎯 What it does: Designed a real-time streaming video question answering benchmark, PhoStream, tailored for mobile scenarios, and evaluated the performance of multi-modal large language models (LLMs) on immediate, retrospective, and prospective answering tasks.

PhotoAgent: Exploratory Visual Aesthetic Planning with Large Vision Models

Mingde Yao (MMLab, Chinese University of Hong Kong), Tianfan Xue (MMLab, Chinese University of Hong Kong)

Image TranslationGenerationOptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringMixture of ExpertsVision Language ModelDiffusion modelImageMultimodality

🎯 What it does: Propose PhotoAgent, an autonomous multi-step photo editing system achieved through a closed-loop of visual perception, planning, execution, and evaluation.

Phy-CoSF: Physics-Guided Continuous Spectral Fields Reconstruction and Spectral Super-Resolution for Snapshot Compressive Imaging

Wudi Chen (Jilin University), Ce Zhu (University of Electronic Science and Technology of China)

Super ResolutionTransformerDiffusion modelAuto EncoderContrastive LearningOptical FlowImagePhysics Related

🎯 What it does: Propose a physics-guided continuous spectral field reconstruction framework, Phy-CoSF, to achieve continuous spectral reconstruction and arbitrary wavelength super-resolution for CASSI.

PhyScene3D: Physically Consistent 3D Interactive Tabletop Scene Generation

Weixing Chen (Sun Yat Sen University), Liang Lin (Sun Yat Sen University)

GenerationOptimizationRobotic IntelligenceTransformerReinforcement LearningPrompt EngineeringVision Language ModelDiffusion modelScore-based ModelGenerative Adversarial NetworkImageTextPoint CloudMeshBenchmarkChain-of-Thought

🎯 What it does: Studied a new framework called PhyScene3D for generating physically consistent and collision-free 3D desktop scenes;

PhysForge: Generating Physics-Grounded 3D Assets for Interactive Virtual World

Yunhan Yang (University of Hong Kong), Xihui Liu (University of Hong Kong)

GenerationData SynthesisRobotic IntelligenceTransformerSupervised Fine-TuningPrompt EngineeringVision Language ModelDiffusion modelAuto EncoderImagePoint CloudMeshPhysics Related

🎯 What it does: Propose a two-stage framework called PhysForge, which can generate interactive 3D assets with complete physical properties (material, mass, function, joint type and parameters) from a single image.

PhysHanDI: Physics-Based Reconstruction of Hand-Deformable Object Interactions

Jihyun Lee (KAIST), Tae-Kyun Kim (KAIST)

OptimizationGraph Neural NetworkTransformerDiffusion modelGaussian SplattingVideoPoint CloudPhysics Related

🎯 What it does: Implemented a physics-based full 3D hand-deformable object interaction reconstruction framework called PHYSHANDI.

Physically-Guided Data-Space Rectified Flow for Precipitation Nowcasting

Wenjie Luo (Yibin University), Zhuo Wang (Yibin University)

TransformerDiffusion modelScore-based ModelRectified FlowOptical FlowImageVideoPhysics RelatedOrdinary Differential Equation

🎯 What it does: Propose a physics-guided data space Rectified Flow (PDRF) model for precipitation nowcasting, addressing the trajectory drift and structural distortion issues of traditional RF models under long time delays.