ICML 2026 Papers — Page 10
International Conference on Machine Learning · 6554 papers
CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models
Junming Huang (Zhejiang University), Weiwei Xu (Zhejiang University)
GenerationData SynthesisTransformerLarge Language ModelMixture of ExpertsVision Language ModelAuto EncoderImageTextMultimodalityPoint CloudMesh
🎯 What it does: Propose a multi-modal large language model, CG-MLLM, which can generate 3D content and perform 3D description (captioning) within the same framework.
CGRiC: Compositional Risk Certification for Structured LLM Outputs
Ibne Farabi Shihab (Iowa State University), Anuj Sharma (Iowa State University)
Safty and PrivacyExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextGraphRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Constructed the Claim Graph Risk Control (CGRiC) framework, which decomposes structured outputs generated by large language models (such as citation-based summaries, chain-of-thought reasoning, and tool-enhanced answers) into verifiable assertion dependency graphs. It calibrates the uncertainty of each assertion through information lift and ultimately provides a formal upper bound on the global failure probability; when the risk exceeds a threshold, it performs localized repairs rather than rejecting the entire output.
CGSVD: Cascaded Granular Singular Value Decomposition for Large Language Model Compression
Yuli Chen (Beijing Information Science and Technology University), Xiulei Liu (Beijing Information Science and Technology University)
CompressionRepresentation LearningTransformerLarge Language ModelAuto EncoderText
🎯 What it does: Propose a training-agnostic low-rank compression framework called CGSVD, which compresses large language models through hierarchical non-uniform rank allocation and iterative residual filling.
Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding
Zhixuan Wu (Beijing University of Posts and Telecommunications), Soujanya Poria (Nanyang Technological University)
Explainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision Language ModelVideoTextBenchmarkChain-of-Thought
🎯 What it does: Studied the video understanding method Chain‑of‑Glimpse based on search-guided progressive object reasoning through information chains.
Chain-of-Goals Hierarchical Policy for Long-Horizon Offline Goal-Conditioned RL
Jinwoo Choi (Seoul National University), Seung-Woo Seo (Seoul National University)
Reinforcement LearningTabularSequentialBenchmark
🎯 What it does: Propose a Chain-of-Goals Hierarchical Policy (CoGHP), which transforms hierarchical decision-making in offline goal-oriented reinforcement learning into a unified framework for autoregressive subgoal sequence generation;
Chain-of-Thought Gradient Descent
Hong-Yu Chen (Northwestern University), Han Liu (Northwestern University)
OptimizationComputational EfficiencyTransformerPrompt EngineeringChain-of-Thought
🎯 What it does: Demonstrate that Chain-of-Thought (CoT) can simulate gradient descent in deep networks with arbitrary steps and capacities through dynamic masking in fixed-depth Transformers.
Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
Iván Arcuschin (Poseidon Research), Arthur Conmy
Explainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelPrompt EngineeringWorld ModelTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This paper systematically evaluates the reliability of large language models in generating Chain-of-Thought (CoT) through experiments on natural language question answering and mathematical problems, finding that models often produce unfaithful reasoning chains without explicit bias prompts, leading to inconsistencies between the final answer and the reasoning process.
Chamaileon: Cross-Context Binder Design with Contextualized Modeling and Mixed Sampling
Hengyuan Cao (Zhejiang University), Min Zhang (Zhejiang University)
Drug DiscoveryProtein Structure PredictionTransformerDiffusion modelScore-based ModelFlow-based ModelBiomedical DataBenchmarkStochastic Differential Equation
🎯 What it does: This study proposes the Chamaileon framework, achieving cross-binding design of a single protein sequence across multiple states (different conformations of the same target) and multiple targets (same epitope across different target proteins);
Channel Adapter for Time Series Foundation Models in Zero-Shot Multivariate Forecasting
Dongyuan Li (University of Tokyo), Jiang Bian (Microsoft Research Asia)
Domain AdaptationAutonomous DrivingOptimizationRepresentation LearningTransformerContrastive LearningTabularTime SeriesBenchmarkFinance RelatedPhysics RelatedStochastic Differential Equation
🎯 What it does: A lightweight and pluggable channel adapter, ChaTSFM, is proposed to enable pre-trained time series foundation models (TSFM) to capture spatial dependencies among multivariate variables in a zero-shot setting, thereby improving the accuracy of multivariate time series forecasting.
ChaosNexus: A Foundation Model for ODE-based Chaotic System Forecasting with Hierarchical Multi-scale Awareness
Chang Liu (Tsinghua University), Yong Li (Tsinghua University)
OptimizationComputational EfficiencyRepresentation LearningRecurrent Neural NetworkTransformerMixture of ExpertsAuto EncoderContrastive LearningTabularTime SeriesSequentialPhysics RelatedStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Designed and trained a foundation model called ChaosNexus based on ScaleFormer for long-term prediction of ODE-based chaotic systems, achieving zero-shot/ few-shot inference through multi-scale hierarchy, Mixture-of-Experts, and frequency fingerprints.
Characterization of Gaussian Universality Breakdown in High-Dimensional Empirical Risk Minimization
Mohamed Chiheb Yaakoubi (Chinese University of Hong Kong), Zhenyu Liao (Huazhong University of Science and Technology)
OptimizationContrastive LearningGaussian SplattingImageTabular
🎯 What it does: This paper studies the limiting behavior under non-Gaussian design within the framework of high-dimensional empirical risk minimization, and provides asymptotic analytical expressions for the mean and covariance of the estimator.
Characterizing the Effect of Noise in Language Generation in the Limit
Aaron Li (Harvard University), Ian Zhang (Duke University)
GenerationTextReview/Survey Paper
🎯 What it does: Studied the generative capacity of language generation in the limit model after introducing noise, provided the equivalence between noise level 1 and any finite noise level in generation, and proved the existence of sets that can be generated without noise but cannot be generated with noise level 1, thus answering a previous open question.
Characterizing the Predictive Impact of Modalities with Supervised Latent-Variable Modeling
Divyam Madaan (New York University), Kyunghyun Cho (New York University)
Explainability and InterpretabilityRepresentation LearningAuto EncoderContrastive LearningImageMultimodalityTabularBiomedical DataElectronic Health RecordsAudio
🎯 What it does: Proposes PRIMO, a supervised latent variable model for quantifying the impact of missing modalities on predictions in multimodal learning, and trains and infers under both complete and missing modalities.
Characterizing Vision-Language-Action Models across XPUs: Constraints and Acceleration for On-Robot Deployment
Kaijun Zhou (School of Computer Science, Shanghai Jiao Tong University), Jinyu Gu (School of Computer Science, Shanghai Jiao Tong University)
Computational EfficiencyRobotic IntelligenceTransformerVision Language ModelVision-Language-Action ModelAuto EncoderContrastive LearningImageTextMultimodality
🎯 What it does: This paper systematically evaluates Vision-Language-Action (VLA) models on multiple XPUs for robot deployment, revealing two-stage computational bottlenecks, and proposes two training-free acceleration methods: DP-Cache and V-AEFusion;
Characterizing, Evaluating, and Optimizing Complex Reasoning
Haoran Zhang (Shanghai Jiao Tong University), Yu Cheng (Chinese University of Hong Kong)
OptimizationExplainability and InterpretabilityComputational EfficiencyReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringMixture of ExpertsTextGraphChain-of-Thought
🎯 What it does: Propose a unified ME2 principle, using DAG abstraction to convert any reasoning trajectory into a comparable structure, and train a Thinking Reward Model (TRM) to achieve large-scale reasoning quality evaluation and optimization.
ChartE$^{3}$: A Comprehensive Benchmark for End-to-End Chart Editing
Shuo Li (Fudan University), Xuanjing Huang (Fudan University)
Image TranslationTransformerLarge Language ModelVision Language ModelDiffusion modelContrastive LearningImageTextMultimodalityBenchmark
🎯 What it does: Proposed the ChartE 3 benchmark to evaluate the end-to-end chart editing capabilities of multimodal models without the need for intermediate code;
Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models
Mickel Liu (University of Washington), Natasha Jaques (University of Washington)
Safty and PrivacyReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningTextChain-of-Thought
🎯 What it does: Proposed a fully online self-play multi-agent reinforcement learning framework called SELF-REDTEAM, used for continuously co-evolving attackers and defenders in language models.
CHB: A Diagnostic Toolkit for Hardness-Aware Clustering Evaluation
Walid Durani (LMU Munich), Christian Böhm (University of Vienna)
Explainability and InterpretabilityRepresentation LearningAuto EncoderContrastive LearningImagePoint CloudTabularBenchmark
🎯 What it does: Proposes the Clustering Hardness Benchmark (CHB), an external evaluation-based diagnostic toolkit that describes datasets using interpretable hardness fingerprints (separability, cohesion, topological structure), and achieves mechanism-level diagnosis of clustering methods' performance and auditing of representation learning through partitioning in the hardness space.
Cheap2Rich: A Multi-Fidelity Framework for Data Assimilation and System Identification of Multiscale Physics - Rotating Detonation Engines
Yuxuan Bao (University of Washington), J. Nathan Kutz (University of Washington)
OptimizationExplainability and InterpretabilityComputational EfficiencyTransformerDiffusion modelScore-based ModelAuto EncoderGenerative Adversarial NetworkContrastive LearningTabularTime SeriesPhysics Related
🎯 What it does: Propose the Cheap2Rich multi-scale data assimilation framework, which reconstructs the full state of high-fidelity RDE (rotating detonation engine) using low-fidelity simulations and sparse sensor historical data, and identifies and compensates for missing physics through low-frequency/high-frequency decomposition.
Chebyshev Policies and the Mountain Car Problem: Reinforcement Learning for Low-Dimensional Control Tasks
Stefan Huber (University of Applied Sciences Salzburg), Jakob Rehrl (University of Applied Sciences Salzburg)
OptimizationComputational EfficiencyRepresentation LearningReinforcement LearningTabularTime Series
🎯 What it does: Analytically solve the Mountain Car problem to obtain the optimal control strategy, and introduce a multivariate Chebyshev polynomial as a general RL policy representation, replacing the traditional MLP approximation.
CHESS: Chebyshev Spectral Synthesis for Trajectory Condensation
Ruituo Wu (University of Electronic Science and Technology of China), Bing Li (University of Electronic Science and Technology of China)
CompressionRepresentation LearningAuto EncoderContrastive LearningTime Series
🎯 What it does: Propose CHESS, a function-priority trajectory compression framework based on low-rank spatial modeling and piecewise Chebyshev polynomial time parameterization;
Chiral Symmetry Breaking in Transformers: A Group-Equivariant Framework for Addressing the Reversal Curse via Adjoint Manifold Mappings
Hanji Du (Shanghai University of International Business and Economics)
RetrievalTransformerContrastive LearningTextRetrieval-Augmented Generation
🎯 What it does: Propose the Chiral Transformer, which addresses the 'reversal curse' in autoregressive models by learning inverse transformation mappings to achieve inverse relationship retrieval.
Chunk-Guided Q-Learning
Gwanwoo Song (Yonsei University), Youngwoon Lee (Yonsei University)
Reinforcement LearningFlow-based ModelTabularTime SeriesSequential
🎯 What it does: Propose an offline reinforcement learning method called Chunk-Guided Q-Learning (CGQ), which combines single-step TD learning with action-block-based TD learning, using block-level Q learning as a regularization signal to guide the single-step Q learning, thereby reducing long-term cumulative errors.
CINOC: Cardinality-Invariant Neural Operator Policies for Scalable PDE Control
Pietro Zanotta (Johns Hopkins University), Jan Drgona (Johns Hopkins University)
OptimizationGraph Neural NetworkTransformerReinforcement LearningDiffusion modelContrastive LearningTabularTime SeriesPhysics RelatedStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: The CINOC framework is proposed by modeling control policies as neural operators and end-to-end training with a differentiable PDE solver, achieving PDE control for multi-agent systems with variable sensor/actuator configurations.
CIRBench: Evaluating Large Language Models as LLVM IR Optimizers
Zi Yang, Haojie Zhou (Jiangnan University)
OptimizationTransformerLarge Language ModelPrompt EngineeringTextBenchmark
🎯 What it does: By constructing CIRBench, an LLVM IR benchmark that includes four levels: analysis, repair, restructuring, and transformation, the paper evaluates the understanding and optimization capabilities of large language models in compilers for IR.
Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Vision-Language Models
Chengcheng Wang (University of Sydney), Kai Han (Huawei Noah's Ark Lab)
Computational EfficiencyRepresentation LearningTransformerSupervised Fine-TuningPrompt EngineeringVision Language ModelContrastive LearningImageTextMultimodalityBenchmark
🎯 What it does: Proposes Circle-RoPE, a method based on conical decoupled rotational position encoding, which can eliminate cross-modal relative position information bias caused by the unified indexing of RoPE in vision-language models. It quantifies and ensures PTD = 0 through the Per-Token Distance (PTD) metric, thereby maintaining the spatial structure within images while achieving isometric decoupling between text and images; meanwhile, it introduces Alternating Geometry Encoding (AGE), which alternates between Circle-RoPE and traditional M-RoPE across different Transformer layers to balance cross-modal alignment and image local feature extraction.
CircuitPrint: Mechanistic Circuit Fingerprints for Large Language Models
Zhenxiong Yan (Hunan University), Wenqiang Jin (Hunan University)
Safty and PrivacyExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringText
🎯 What it does: CircuitPrint proposes a non-intrusive intellectual property fingerprinting technique that verifies the source of large language models (LLMs) by using specially synthesized trigger prefixes in standard API queries.
CiteGuard: Conformal False-Discovery Control for Faithful Retrieval-Augmented Generation
Xiangyu Jiang (Southwestern University of Finance and Economics)
GenerationRetrievalData-Centric LearningTransformerPrompt EngineeringContrastive LearningTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose CiteGuard to achieve controllable citation credibility in retrieval-augmented generation through multiple testing and compliant p-values.
CL-GCL: Comprehensive and Lightweight Graph Contrastive Learning
Jianqing Liang (Southeast University), Zhiqiang Li (Shanxi University)
Computational EfficiencyRepresentation LearningGraph Neural NetworkContrastive LearningGraph
🎯 What it does: Propose a comprehensive lightweight graph contrastive learning framework named CL-GCL, which constructs community-level positive and negative samples through graph coarsening and manifold learning, and trains a hybrid contrastive loss on the node encoder.
CLAA: Cross-Layer Attention Aggregation for Accelerating LLM Prefill
Bradley McDanel (Franklin and Marshall College), Harshit Khaitan (Meta Reality Labs)
Computational EfficiencyTransformerLarge Language ModelTextBenchmark
🎯 What it does: Studied an answer-informed reference framework to evaluate the importance of tokens during the prefilling phase, and proposed a Cross-Layer Attention Aggregation (CLAA) method to accelerate the prefilling of LLMs.
Clarify Before You Draw: Proactive Agents for Robust Text-to-CAD Generation
Bo Yuan (Georgia Institute of Technology), Yongxin Chen (Georgia Institute of Technology)
GenerationAI Code AssistantTransformerLarge Language ModelSupervised Fine-TuningAgentic AIPrompt EngineeringTextMesh
🎯 What it does: A dual-agent system called ProCAD is proposed. First, a clarification agent proactively identifies and asks questions about missing or conflicting geometric information. Then, an encoding agent generates executable CadQuery code based on the clarified specifications, achieving robust text-to-CAD generation.
CLARITree: Cholesky and Lookahead Accelerations for Regression with Interpretable Piecewise Linear Trees
Yixiao Wang (Duke University), Cynthia Rudin (Duke University)
Explainability and InterpretabilityComputational EfficiencyTabularBenchmark
🎯 What it does: Proposed and implemented a near-optimal sparse piecewise linear regression tree algorithm called CLARITree, which combines lookahead-style split search with rank-one Cholesky updates to efficiently evaluate continuous feature thresholds.
CLASP: Online learning algorithms for Convex Losses And Squared Penalties
Ricardo N. Ferreira (NOVA School of Science and Technology), Claudia Soares
Optimization
🎯 What it does: Proposes the CLASP (Convex Losses And Squared Penalties) framework, aiming to minimize the accumulated loss and squared constraint violations, addressing online convex optimization problems with dynamic constraints.
Class-Conditional Distribution Balancing for Group Robust Classification
Miaoyun Zhao (Dalian University), Qiang Zhang (Dalian University of Technology)
ClassificationDomain AdaptationFederated LearningExplainability and InterpretabilityData-Centric LearningContrastive LearningGaussian SplattingImageTextMultimodality
🎯 What it does: Propose a sample weighting method based on category-conditioned distribution balancing (CCDB), which eliminates pseudo-relevance in training data without biased annotations or pre-trained models, thereby enhancing model robustness in out-of-distribution and imbalanced scenarios.
Class-Grouped Normalized Momentum and Faster Hyperparameter Exploration to Tackle Class Imbalance in Federated Learning
Haemin Park (Northwestern University), Balakrishnan Ananthanarayanan (Intel Corporation)
OptimizationFederated LearningHyperparameter SearchContrastive LearningImageTabularBiomedical DataBenchmark
🎯 What it does: Proposed FedCGNM client optimizer and FedHOO hyperparameter search method to address class imbalance issues in federated learning.
Class-Prior Perturbation-Robust Regularization for Imbalanced Unreliable Partial Label Learning
Congyu Qiao (Southeast University), Ning Xu (Southeast University)
ClassificationOptimizationData-Centric LearningContrastive LearningImage
🎯 What it does: Propose a novel CLAPOR framework to address the Imbalanced Unreliable Partial Label Learning (I-UPLL) problem where both class imbalance and label unreliability coexist; achieve robust regularization against prior uncertainty by perturbing class priors during training, avoiding reliance on inaccurate prior estimates.
CLEAR: Context-Aware Learning with End-to-End Mask-Free Inference for Adaptive Video Subtitle Removal
Qingdong He (University of Electronic Science and Technology of China), Xiaobin Hu (National University of Singapore)
RestorationTransformerDiffusion modelGenerative Adversarial NetworkContrastive LearningVideo
🎯 What it does: Proposes CLEAR, an end-to-end video caption removal framework based on self-supervised pre-training and LoRA adaptation, without requiring masks.
ClimateAR: Multi-Scale Autoregressive Generative Modeling for Climate Forecasting
Yue Yu (Zhejiang University), Ling Chen (Zhejiang University)
GenerationData SynthesisTransformerSupervised Fine-TuningPrompt EngineeringDiffusion modelAuto EncoderTabularTime SeriesPhysics Related
🎯 What it does: Proposed ClimateAR, a self-attention generative model for probabilistic climate forecasting, combining an aligned tokenizer and multi-scale conditional mechanisms;
CLIMB: Taming the LoRA Residency Cliff in Multi-LoRA Serving
Haoran Zhang (Harbin Institute of Technology), Hongzhi Wang (Harbin Institute of Technology)
OptimizationComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringText
🎯 What it does: Proposed and implemented a minimal entry controller named CLIMB, which migrates the adapter loading bottleneck to a controllable queue through pre-check feasibility in multi-tenant multi-LoRA inference, thereby eliminating the 'LoRA residency cliff' caused by continuous batching.
CLINIC : Evaluating Multilingual Trustworthiness in Language Models for Healthcare
Akash Ghosh (Indian Institute of Technology Patna), Chirag Agarwal (University of Virginia)
Safty and PrivacyExplainability and InterpretabilityAdversarial AttackDrug DiscoveryTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBiomedical DataElectronic Health RecordsBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposed the CLINIC multilingual reliability benchmark for systematically evaluating the reliability of medical language models in multilingual environments.
CLINIC: Towards High-quality Graph Out-Of-Distribution Detection
Yifan Wang (University of International Business and Economics), Xiao Luo (University of Wisconsin-Madison)
Anomaly DetectionMeta LearningGraph Neural NetworkContrastive LearningGraph
🎯 What it does: Proposed the CLINIC method for detecting out-of-distribution (OOD) samples in graph data.
ClinTutor-R1: Advancing Scalable and Robust One-to-Many Alignment in Clinical Socratic Education
Zhitao He (Hong Kong University of Science and Technology), Yi R. Fung (Hong Kong University of Science and Technology)
Autonomous DrivingFederated LearningSafty and PrivacyExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningAdversarial AttackHyperparameter SearchData-Centric LearningRobotic IntelligenceMeta LearningDrug DiscoveryReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AIVision Language ModelImageTextBiomedical DataElectronic Health RecordsChain-of-Thought
🎯 What it does: Proposed ClinEdu multi-agent simulator, ClinTeach large-scale Socratic teaching dialogue dataset, and the first vision-language agent for clinical one-to-many teaching, ClinTutor-R1, which uses an internal thinking mechanism to achieve a balance between individual and group guidance.
CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large Vision-Language Models
Sangin Lee (Sejong University), Yukyung Choi (Sejong University)
SegmentationComputational EfficiencyTransformerPrompt EngineeringVision Language ModelContrastive LearningImageVideoText
🎯 What it does: Developed a training-agnostic, text-guided visual token pruning method called LiteLVLM for efficient pixel-level localization inference.
Clipped Q-Learning: Your Value Clipping Is Secretly A Robust Operator
Zhishuai Liu (Duke University), Pan Xu (Duke University)
OptimizationReinforcement LearningContrastive LearningTabular
🎯 What it does: Proposes 'Clipped Q-Learning', which truncates Bellman backups in Q-learning with a threshold λ, and proves its equivalence to robust Bellman iterations in TV distance regularized robust MDPs (RRMDPs), thereby learning policies robust to specific dynamic perturbations; further provides UCB-Hoeffding and UCB-Bernstein variants for both discrete and deep value-based algorithms, and presents polynomial regret upper bounds; finally, experiments validate that the truncation mechanism significantly improves robustness to target environment perturbations without sacrificing performance in the original environment, both in simple MDPs and the Inverted Pendulum task in OpenAI Gym.
Clipping Bottleneck: Stabilizing RLVR via Stochastic Recovery of Near-Boundary Signals
Shuo Yang (Peking University), Jingren Zhou (Alibaba)
OptimizationReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningMixture of ExpertsText
🎯 What it does: Propose the Near-boundary Stochastic Rescue (NSR) method, which restores the learning signals that were clipped during RLVR training by randomly retaining tokens near the boundary that exceed the hard clipping threshold, thereby improving the model's training stability and convergence performance.
Clipping Low-Probability Tokens in SFT Yields a Generalizable Initialization for RL
Tian-Shuo Liu (Nanjing University), Yang Yu (Nanjing University)
Data-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringText
🎯 What it does: This paper proposes trimming low-probability (off-policy) tokens during the supervised fine-tuning (SFT) phase to reduce forgetting of prior knowledge, thereby obtaining a more general and exploratory RL initialization; and validates the effectiveness of this method through theoretical and experimental analysis.
Clipping Makes Distributed and Federated Asynchronous SGD Robust to Stragglers
Samuel Erickson (KTH Royal Institute of Technology), Mikael Johansson (KTH Royal Institute of Technology)
OptimizationFederated LearningReinforcement LearningContrastive LearningImageText
🎯 What it does: This paper studies the robustness of gradient clipping in asynchronous stochastic gradient descent (ASGD), provides expected convergence and high probability convergence theories in homogeneous and heterogeneous environments, and proves that clipping can eliminate the impact of maximum delay on convergence speed.
Closing the Expression Gap in LLM Instructions via Socratic Questioning
Jianwen Sun (Nankai University), Kaipeng Zhang (Shanghai Innovation Institute)
Explainability and InterpretabilityData-Centric LearningAI Code AssistantTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringDiffusion modelScore-based ModelFlow-based ModelRectified FlowNeural Radiance FieldAuto EncoderGenerative Adversarial NetworkContrastive LearningGaussian SplattingImageTextTabularStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Built a Socratic questioning agent called Nous that actively asks users questions to bridge the intention expression gap.
Closing the Loop: Universal Repository Representation with RPG-Encoder
Jane Luo (Microsoft), Scarlett Li
Representation LearningAI Code AssistantGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringDiffusion modelTextGraphRetrieval-Augmented Generation
🎯 What it does: By converting the Repository Planning Graph (RPG) from a unidirectional code generation blueprint into a bidirectional, high-fidelity repository representation through the RPG-Encoder, the paper completes a closed-loop reasoning process for repository understanding and generation;
Closing the Sim-to-Real Gap in Non-Markovian Spreading Processes via GPU-Accelerated Distributional RL
Heman Shakeri (University of Virginia)
Domain AdaptationOptimizationFederated LearningComputational EfficiencyRecurrent Neural NetworkGraph Neural NetworkTransformerReinforcement LearningDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningGraphTabularTime SeriesSequentialBenchmarkPhysics RelatedStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: This study addresses the 'simulation-to-reality' gap in the network propagation process by proposing a Stratified Mean-Field Observer and a GPU-accelerated non-Markovian simulator called FLASHSPREAD, and combines them with distributed truncated quantile RL (Truncated Quantile Critics) to achieve zero-copy, efficient control of large-scale networks.
CLoVE: Personalized Federated Learning through Clustering of Loss Vector Embeddings
Randeep Bhatia (Nokia Bell Labs), Sayak Chakrabarty (Northwestern University)
ClassificationFederated LearningContrastive LearningImageText
🎯 What it does: Proposed the CLoVE algorithm, which achieves clustering federated learning by embedding based on model loss vectors, enabling simultaneous client clustering and model personalization.
Cluster-Aware Causal Mixer for Online Anomaly Detection in Multivariate Time Series
Md Mahmuddun Nabi Murad (University of South Florida), Yasin Yilmaz (University of South Florida)
Anomaly DetectionRecurrent Neural NetworkTransformerDiffusion modelScore-based ModelAuto EncoderContrastive LearningGaussian SplattingTabularTime Series
🎯 What it does: Propose a cluster-aware causal mixer (CCM-TAD) for online anomaly detection in multivariate time series data, combining clustering embedding, causal mixing layers, and serialized anomaly scoring;
Clustered Influence Functions
Miklós Máté Badó (Eotvos Lorand University), Kristian Fenech (Eotvos Lorand University)
Explainability and InterpretabilityComputational EfficiencyConvolutional Neural NetworkTransformerImage
🎯 What it does: Propose Clustered Influence Functions (CiF), which achieve reusable subset influence functions and low query cost by clustering gradients and caching the inverse curvature response for each cluster.
Clustering as Reasoning: A $k$-Means Interpretation of Chain-of-Thought Graph Learning
Xuanting Xie (University of Electronic Science and Technology of China), Yuan Fang (Singapore Management University)
Explainability and InterpretabilityRepresentation LearningGraph Neural NetworkTransformerPrompt EngineeringContrastive LearningTextGraphChain-of-Thought
🎯 What it does: Propose a unified KCOT framework that combines chain-of-thought (CoT) with graph neural networks (GNN), and theoretically equates Transformer blocks to k-means clustering, thereby interpreting CoT as an iterative process of 'clustering + updating'.
Clustering in Deep Stochastic Transformers
Lev Fedorov (New York University), Mathieu Lauriere (New York University)
Computational EfficiencyRepresentation LearningTransformerImageStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: This paper analyzes the self-attention dynamics of deep Transformers under random initialization, derives the spherical SDE continuous-time limit, and proves that noise can prevent tokens from clustering to a single point.
CMI-RewardBench: Evaluating Music Reward Models with Compositional Multimodal Instruction
Yinghao Ma (Queen Mary University of London), Emmanouil Benetos (Queen Mary University of London)
GenerationData SynthesisReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringDiffusion modelContrastive LearningTextMultimodalityBenchmarkAudio
🎯 What it does: Constructed CMI-RewardBench — a unified benchmark for evaluating the music quality and instruction following capability of music generation models under multimodal instructions such as text, lyrics, and reference audio. Based on this, we released two large-scale datasets: CMI-Pref-Pseudo (pseudo-label pairs) and CMI-Pref (human-annotated pairs). Subsequently, we trained and launched the parameter-efficient CMI-RM reward model that can handle multimodal inputs.
cMoLLM at Scale: Horizontal Scaling Laws for Convolutionally-Gated Mixture-of-LLMs
Xin Yang (Zhejiang University), Ryan Dong
Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelMixture of ExpertsText
🎯 What it does: Propose a convolutional gated Mixture-of-LLMs (cMoLLM), achieving horizontal scaling of capacity at the Transformer pipeline level through dynamic convolution;
Co-Evolving Latent Action World Models
Yucen Wang (Nanjing University), Jiang Bian (Microsoft Research Asia)
GenerationRobotic IntelligenceTransformerVision-Language-Action ModelDiffusion modelFlow-based ModelAuto EncoderWorld ModelVideo
🎯 What it does: Propose a framework called CoLA-World that jointly trains a pre-trained video generation model with a latent action model (LAM), achieving co-evolution between the world model and latent actions by first freezing the pre-trained model for warm-up, then performing end-to-end joint training.
Co-Generative De Novo Functional Protein Design
Xinrui Chen (Tsinghua University), Zaiqing Nie (Tsinghua University)
GenerationProtein Structure PredictionTransformerDiffusion modelBiomedical DataRetrieval-Augmented Generation
🎯 What it does: Propose a co-generative model called CodeFP that simultaneously generates protein sequences and structures for de novo functional protein design without templates;
Co-RedTeam: Orchestrated Security Discovery and Exploitation with LLM Agents
Pengfei He (Google Cloud AI Research), Long Le
AI Code AssistantTransformerLarge Language ModelAgentic AITextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This paper proposes CO-REDTEAM, a multi-agent framework designed for red team operations, aimed at automating the discovery and exploitation of software vulnerabilities.
Coarse-Grained Boltzmann Generators
Weilong Chen (Technical University of Munich), Julija Zavadlav (Technical University of Munich)
Drug DiscoveryProtein Structure PredictionScore-based ModelFlow-based ModelBiomedical Data
🎯 What it does: Proposed Coarse-Grained Boltzmann Generators (CG-BGs), which use flow-based generative models to generate samples in the coarse-grained coordinate space and achieve scalable sampling of the equilibrium distribution of molecular systems through importance reweighting with learned PMF.
COBRA: Contribution-Based Bayesian Rank Allocation for Parameter-Efficient Fine-Tuning
Hongcheng Ding (Zhejiang University of Finance and Economics), Deshinta Arrova Dewi (INTI International University)
OptimizationExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelSupervised Fine-TuningTextFinance Related
🎯 What it does: To address the issue of uneven hierarchical importance in parameter-efficient fine-tuning of large language models, the COBRA framework is proposed. It first obtains the contribution of each layer to prediction using hierarchical conductance, then combines the adaptation requirements of gradient amplitude to generate a dual-factor importance distribution. Optimal heterogeneous low-rank matrix decomposition (LoRA) hierarchical rank configurations are achieved through Bayesian optimization.
CoCoEdit: Content-Consistent Image Editing via Region Regularized Reinforcement Learning
Yuhui Wu (Hong Kong Polytechnic University), Lei Zhang (Hong Kong Polytechnic University)
Image TranslationRestorationGenerationTransformerReinforcement LearningPrompt EngineeringVision Language ModelDiffusion modelContrastive LearningImageMultimodalityBenchmark
🎯 What it does: A post-training framework named CoCoEdit was developed to enhance image editing models by achieving more refined editing results while maintaining content consistency with the original image.
CoCoEmo: Composable and Controllable Human-Like Emotional TTS via Activation Steering
Siyi Wang (University of Melbourne), Ting Dang (University of Melbourne)
GenerationData SynthesisTransformerMixture of ExpertsTextAudio
🎯 What it does: Proposes a hybrid emotion control framework based on activation steering, which can achieve quantitative interpolation of emotions and synthesis of text-emotion mismatch without retraining the model.
CoCoQuant: Breaking the Bandwidth Wall via Co-Optimized Communication and Computation Quantization
Haojie Duanmu (Shanghai Jiao Tong University), Dahua Lin (Chinese University of Hong Kong)
OptimizationComputational EfficiencyTransformerLarge Language ModelText
🎯 What it does: For multi-GPU distributed LLM inference, joint optimization of communication and computation quantization to break the bandwidth wall
CoCoReviewBench: A Completeness- and Correctness-Oriented Benchmark for AI Reviewers
Hexuan Deng (Harbin Institute of Technology), Min Zhang (Harbin Institute of Technology)
Explainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This paper constructs a benchmark for AI review systems—CoCoReviewBench—by enhancing the completeness and correctness of human reviews, providing a fine-grained category-level evaluation framework.
CocoRNA: Collective RNA Design with Cooperative Multi-agent Reinforcement Learning
Tianmeng Hu (University of Exeter), Ke Li (University of Exeter)
OptimizationDrug DiscoveryProtein Structure PredictionTransformerReinforcement LearningPrompt EngineeringMixture of ExpertsGenerative Adversarial NetworkBiomedical Data
🎯 What it does: Proposed and implemented a collaborative multi-agent reinforcement learning framework called COCORNA for RNA inverse design, which generates RNA sequences that conform to a given secondary structure.
CoD-Lite: Real-Time Diffusion-Based Generative Image Compression
Zhaoyang Jia (Microsoft Research Asia), Yan Lu (Microsoft Research Asia)
CompressionKnowledge DistillationConvolutional Neural NetworkDiffusion modelGenerative Adversarial NetworkContrastive LearningImage
🎯 What it does: Proposed a real-time, lightweight one-step convolutional diffusion model for image compression
CoDA-Bench: Can Code Agents Handle Data-Intensive Tasks?
Yuxin Zhang (Renmin University of China), Xiaoyong Du (Renmin University of China)
Data-Centric LearningAI Code AssistantGraph Neural NetworkTransformerLarge Language ModelAgentic AIPrompt EngineeringImageTextGraphTabularBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose and implement CODA-BENCH, a data-intensive Linux sandbox benchmark built within the Kaggle ecosystem that simultaneously evaluates code generation and data discovery capabilities, containing 1,009 tasks;
Code2Video: A Code-centric Paradigm for Educational Video Creation
Yanzhe Chen (National University of Singapore), Mike Zheng Shou (National University of Singapore)
GenerationExplainability and InterpretabilityData-Centric LearningAI Code AssistantTransformerLarge Language ModelAgentic AIPrompt EngineeringVision Language ModelDiffusion modelVideoTextMultimodalityBenchmark
🎯 What it does: Propose Code2Video, a multi-agent framework based on executable Manim code, for generating structured and interpretable educational videos.
Code2Worlds: Empowering Coding LLMs for 4D World Generation
Yi Zhang (Peking University), Hao Tang (Peking University)
GenerationData SynthesisRobotic IntelligenceAI Code AssistantTransformerLarge Language ModelVision Language ModelDiffusion modelImageVideoTextMultimodalityMeshBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Through a dual-stream architecture and a closed-loop self-reflective mechanism, programming LLMs are utilized to generate executable 4D physical simulation scenarios, achieving the transformation from textual instructions to dynamically runnable holographic worlds.
CodeChemist: Test-Time Scaling for Low-Resource Code Generation via Functional Knowledge Transfer
KaiXin Wang, Bin Shi (Xi'an Jiaotong University)
AI Code AssistantTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose the CodeChemist framework, which improves code generation for low-resource programming languages during testing by using multi-temperature hedged sampling and synthetic test cases.
CodeClash: Benchmarking Goal-Oriented Software Engineering
John Yang (Stanford University), Diyi Yang (Stanford University)
AI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringTextBenchmarkChain-of-Thought
🎯 What it does: Proposed the CodeClash benchmark, which uses a multi-round competition format to let language models (LMs) repeatedly edit a codebase and compete in different arenas (such as BattleSnake, Poker, RoboCode, etc.) without explicit task instructions, in order to achieve high-level goals.
CodeMamba: Shifting from Target Semantics to Self-Supervised Background Manifold Learning for Singularity Detection in Infrared Sequences
Jingwen Ma (Tianjin University of Technology), Shengyong Chen (Tianjin University of Technology)
Object DetectionAnomaly DetectionTransformerContrastive LearningImageVideo
🎯 What it does: Propose the CodeMamba dual-stream framework, which uses self-supervised background manifold learning and motion singularity capture to detect tiny targets in infrared sequences.
CodeTaste: Can LLMs Generate Human-Level Code Refactorings?
Alex Thillen (ETH Zurich), Martin Vechev (ETH Zurich)
AI Code AssistantTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Constructed and released the CODETASTE benchmark to evaluate the performance of large language models on real multi-file code refactoring tasks and compare their consistency with human refactoring choices.
CODiff: One-Step Diffusion Model for Camouflaged Object Detection
Xiaotong Fu (Zhejiang University), Shibo He (Zhejiang University)
Object DetectionConvolutional Neural NetworkTransformerDiffusion modelContrastive LearningImageStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Proposed and implemented a single-step diffusion model called CODiff for camouflaged object detection.
CoEvol-NO: State and Coordinate Co-Evolution with an Error-Driven Predictor-Corrector Paradigm for Neural Operator Transformer
Jianqiao Zeng (Fudan University), Junchi Yan (Shanghai Jiao Tong University)
OptimizationComputational EfficiencyTransformerPoint CloudMeshGraphBenchmarkPhysics Related
🎯 What it does: Designed CoEvol-NO, a linear complexity neural operator that achieves co-evolution of potential states and grid coordinate sequences through a predictor-corrector framework.
Coevolutionary Continuous Discrete Diffusion: Make Your Diffusion Language Model a Latent Reasoner
Cai Zhou (Massachusetts Institute of Technology), Dinghuai Zhang (Microsoft Research)
GenerationComputational EfficiencyRepresentation LearningData-Centric LearningTransformerLarge Language ModelMixture of ExpertsDiffusion modelScore-based ModelTextSequentialStochastic Differential EquationOrdinary Differential Equation
🎯 What it does: Proposed and implemented the Coevolutionary Continuous Discrete Diffusion (CCDD) language model, which jointly integrates continuous diffusion and discrete diffusion within the same network, achieving implicit reasoning and stronger expressiveness.
CoF-T2I: Video Models as Pure Visual Reasoners for Text-to-Image Generation
Chengzhuo Tong (Peking University), Wentao Zhang (Peking University)
GenerationTransformerDiffusion modelRectified FlowAuto EncoderImageVideoText
🎯 What it does: Reusing pre-trained video models for text-to-image generation by generating high-quality images through chained frames (CoF) step-by-step visual reasoning.
CofactGVR: Counterfactual Intervention for Grounded Visual Reasoning
Yan Zhang (Tsinghua University), Jungong Han (Tsinghua University)
Explainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningVision Language ModelContrastive LearningImageTextMultimodality
🎯 What it does: Propose CofactGVR, which explicitly encourages the model to truly rely on visual evidence during the answering process by introducing counting counterfactual interventions into multimodal large models.
CoFrGeNet: Continued Fraction Architectures for Language Generation
Amit Dhurandhar (IBM Research), Rahul Nair (IBM Research)
GenerationComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelSupervised Fine-TuningAuto EncoderTextMultimodality
🎯 What it does: Designed and implemented CoFrGeNet, a Transformer alternative component based on continued fractions, to replace multi-head attention and feed-forward networks, significantly reducing the number of parameters and accelerating training and inference;
COFT: Counterfactual–Conformal Decoding for Fair Chain‑of‑Thought Reasoning in Large Language Models
Arya Fayyazi (University of Southern California), Massoud Pedram (University of Southern California)
Federated LearningSafty and PrivacyExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This paper proposes COFT, a training-agnostic decoding framework that is executed only during inference, aiming to ensure fairness in large language models during chain-of-thought reasoning;
CoGe-GCD: Reframing Generalized Category Discovery with Compositional Generalization
Luyao Tang (University of Hong Kong), Cheng Chen (University of Hong Kong)
ClassificationRecognitionTransformerVision Language ModelContrastive LearningImage
🎯 What it does: Propose the CoGe-GCD framework for Generalized Category Discovery, which consists of two stages: Compositional Perception and Generalizing Induction, performing primitive abstraction and geometric calibration on patch tokens.
CoGenCast: A Coupled Autoregressive–Flow Generative Framework for Time Series Forecasting
Mingyue Cheng (University of Science and Technology of China), Qi Liu (University of Science and Technology of China)
GenerationComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringDiffusion modelScore-based ModelFlow-based ModelMultimodalityTabularTime SeriesBenchmark
🎯 What it does: Proposes CoGenCast, a temporal prediction framework that couples pre-trained LLMs with a flow-matching generation mechanism.
CoGeoAD: Hierarchical Color-Geometric Fusion with Multi-View Attention for Zero-Shot 3D Anomaly Detection
Ke Xu (Anhui University), Jianfeng Qiu (Anhui University)
Anomaly DetectionTransformerPrompt EngineeringVision Language ModelContrastive LearningImageTextPoint Cloud
🎯 What it does: Proposed a zero-shot 3D anomaly detection framework called CoGeoAD based on CLIP, which can integrate color and geometric information and achieve high-precision localization.
Cognitive Fatigue in Autoregressive Transformers: Formalization and Measurement
Riju Marwah (Guru Gobind Singh Indraprastha University), Amit Sheth (University of South Carolina)
Explainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelText
🎯 What it does: Propose the concept of cognitive fatigue and construct the Fatigue Index (FI) through three internal signals—attention decay, representation drift, and entropy imbalance—during the autoregressive generation process, to monitor the model's generation quality in real time.
COGNOS: Universal Enhancement for Time Series Anomaly Detection via Constrained Gaussian-Noise Optimization and Smoothing
Wenlong Shang (Beijing University of Technology), Peng Chang (Beijing University of Technology)
Anomaly DetectionOptimizationAuto EncoderContrastive LearningGaussian SplattingTabularTime Series
🎯 What it does: Propose the COGNOS framework, which enhances the statistical quality of reconstructed residuals and the stability of anomaly scores in TSAD through Gaussian white noise regularization training and adaptive Kalman filter post-processing.
CoIRL-AD: Collaborative-Competitive Imitation-Reinforcement Learning in Latent World Models for Autonomous Driving
Xiaoji Zheng (Tsinghua University), Jiangtao Gong (Shanghai Jiao Tong University)
Autonomous DrivingTransformerReinforcement LearningAuto EncoderContrastive LearningWorld ModelVideoPoint Cloud
🎯 What it does: Propose the CoIRL-AD framework, which integrates imitation learning and reinforcement learning in offline end-to-end autonomous driving through competitive dual strategies, and utilizes a latent world model to achieve future trajectory prediction and reward estimation;
CoLA: Cross-Modal Low-rank Adaptation for Multimodal Downstream Tasks
Wish Suharitdamrong (University of Surrey), Sara Atito (University of Surrey)
Federated LearningComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringMixture of ExpertsVision Language ModelContrastive LearningImageTextMultimodalityRetrieval-Augmented GenerationAudio
🎯 What it does: Designed a cross-modal low-rank adaptation framework called CoLA, for parameter-efficient fine-tuning of vision, language, audio, and other modalities' base models within a dual-stream architecture.
Cold-Start Personalization via Bayesian Adaptive Questioning
Avinandan Bose (Meta Superintelligence Labs), Asli Celikyilmaz (Meta Superintelligence Labs)
Recommendation SystemReinforcement Learning from Human FeedbackTransformerReinforcement LearningPrompt EngineeringWorld ModelText
🎯 What it does: Propose CAPEn, a cold start personalization framework, which achieves efficient question selection through offline learning of preference associations followed by online Bayesian inference.
CollabBench: Benchmarking and Unleashing Collaborative Ability of LLMs with Diverse Players via Proactive Engagement
Hong Qian (East China Normal University), Aimin Zhou (East China Normal University)
Reinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningAgentic AIPrompt EngineeringMixture of ExpertsTextSequentialBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Constructed a collaborative game benchmark named CollabBench, training and evaluating the collaborative capabilities of LLM agents in diverse player environments through multiple player simulations, unified agentic rollout, and hybrid reward mechanisms.
Collaborative and Efficient Fine-tuning: Leveraging Task Similarity
Gagik Magakyan (Massachusetts Institute of Technology), Asuman E. Ozdaglar (Massachusetts Institute of Technology)
Federated LearningComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningText
🎯 What it does: Propose the CoLoRA framework, which utilizes task similarity to perform parameter-efficient personalized fine-tuning in distributed/federated environments.
Collaborative Disagreement Resolution for Scalable Oversight
Yuyang Jiang (University of Chicago), Chenhao Tan (University of Chicago)
Federated LearningExplainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringTextBenchmarkChain-of-Thought
🎯 What it does: Propose a collaborative 'Disagreement Resolution' protocol to replace traditional adversarial debates, helping AI models self-correct and approach truth in the absence of strong judges.
Collaborative Learning for Semi-Supervised LiDAR Semantic Segmentation
Bin Yang (Robert Bosch GmbH), Alexandru Paul Condurache (Robert Bosch GmbH)
SegmentationAutonomous DrivingFederated LearningKnowledge DistillationRepresentation LearningDiffusion modelAuto EncoderGenerative Adversarial NetworkContrastive LearningPoint Cloud
🎯 What it does: Propose the CoLLiS framework, which enables multiple LiDAR representations (range-view, BEV, voxel, etc.) to collaboratively learn from each other within the same training step, thereby improving the performance of semi-supervised semantic segmentation.
Collaborative Threshold Watermarking
Tameem Bakr (MBZUAI), Nils Lukas (MBZUAI)
Federated LearningSafty and PrivacyImageText
🎯 What it does: A threshold watermark mechanism was designed in federated learning, enabling model ownership verification only when at least t clients collaborate.
Collapsed Effective Operators for Higher-order Structures
Maximilian Krahn (Imperial College London), Tolga Birdal (Imperial College London)
ClassificationOptimizationRepresentation LearningGraph Neural NetworkContrastive LearningGraphTabularTime SeriesSequentialBiomedical Data
🎯 What it does: This paper proposes an effective operator (Collapsed Effective Operator) that efficiently compresses high-order topological structures to the vertex level, achieving a unified representation of high-order information at the node level.
COLLIE: Guiding Skill Discovery in Semantically Coherent Latent Space
Yao Luan (Tsinghua University), Qing-Shan Jia (Tsinghua University)
Robotic IntelligenceReinforcement LearningContrastive LearningImageTabularSequential
🎯 What it does: Proposes a training-free skill discovery framework called COLLIE, which constructs a semantically consistent latent space using unsupervised data and generates training-agnostic guidance signals through sparse human feedback to guide skill learning.
Colorful Pinball: Density-Weighted Quantile Regression for Conditional Guarantee of Conformal Prediction
Qianyi Chen (Tsinghua University), Bo Li (Tsinghua University)
OptimizationExplainability and InterpretabilityComputational EfficiencyData-Centric LearningRecurrent Neural NetworkTransformerContrastive LearningTabularTime SeriesSequential
🎯 What it does: This paper addresses the conditional coverage problem in split conformal prediction by proposing Colorful Pinball Conformal Prediction (CPCP), which significantly improves conditional coverage performance through density-weighted quantile regression.
Coloring the Noise: Adversarial Sobolev Alignment for Faithful Image Super Resolution
Hongbo Wang (Chinese Academy of Sciences), Ran He (Chinese Academy of Sciences)
RestorationSuper ResolutionTransformerDiffusion modelScore-based ModelFlow-based ModelRectified FlowGenerative Adversarial NetworkContrastive LearningImage
🎯 What it does: Propose a generative super-resolution method ASASR that combines Sobolev space for noise inpainting and adversarial Sobolev alignment, addressing the distortion of high-frequency details in traditional generative priors.
Column Thresholding for Sparse Spiked Wigner Models: Improved Signal Strength Requirements
Jian-Feng Cai (Hong Kong University of Science and Technology), Jiaxi Ying (Hong Kong University of Science and Technology)
OptimizationComputational EfficiencyRepresentation LearningContrastive LearningTabularTime SeriesSequentialBenchmarkPhysics Related
🎯 What it does: Study the sparse spike Wigner model, propose a two-stage algorithm combining column thresholding and truncated power iteration for recovering sparse vectors from noisy matrices.