CodeAdversarial AttackTransformerReinforcement LearningPrompt EngineeringVision Language ModelDiffusion modelContrastive LearningImageTextMultimodalityRetrieval-Augmented Generation
π― What it does: Propose the VERA-V framework, which transforms the Jailbreak task of vision-language models into joint posterior inference of text-image pairs, thereby enabling cross-modal red team attacks.
VeriSimpl: Robust Optimization Modeling from Natural Language using Simplification-based Verification
Sumaya Abdul Rahman (Hamad Bin Khalifa University), Mohammad Raza (Qatar Computing Research Institute)
CodeOptimizationExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented Generation
π― What it does: Propose a robust verification framework called VERISIMPL that utilizes a solver to generate simplified diagnostic queries and combines a large language model (LLM) for natural language to optimization model verification.
Varun Ursekar (Scale AI), Samuel Marc Denton (Scale AI)
CodeOptimizationAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
π― What it does: This paper proposes the VERO external harness for automatically improving LLM agents through code modifications, and constructs the VERO-BENCH evaluation suite to conduct systematic experiments on the agent optimization process.
Guangzhi Sun (Tsinghua University), Chao Zhang (Tsinghua University)
CodeFederated LearningSafty and PrivacyExplainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelContrastive LearningVideoTextMultimodalityBenchmarkRetrieval-Augmented GenerationAudio
π― What it does: Propose video-SALMONN, a streaming audio-visual large language model capable of processing videos longer than 3 hours at 1 FPS and 360p resolution; introduce test-time training (TTT) to enhance memory, and design TTTMEM layer, two-stage training, and modality-aware memory retrieval mechanism; also propose the ELViM benchmark to evaluate the model's learning and reuse capabilities in long videos.
VideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding
Jiapeng Shi (Fudan University), Zuxuan Wu (Fudan University)
CodeRecognitionSegmentationRetrievalTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelContrastive LearningVideoTextMultimodalityBenchmark
π― What it does: Proposed a unified Video LLM (VideoLoom) that can simultaneously achieve spatiotemporal understanding of videos, supporting temporal localization and spatial segmentation.
π― What it does: Designed and trained VideoMAETok, a video tokenizer based on ViT, which simulates the denoising task of diffusion models during training using a high proportion of masking and interpolated Gaussian noise to generate representations suitable for latent diffusion models.
VideoVeritas: AI-Generated Video Detection via Perception Pretext Reinforcement Learning
Hao Tan (University of Chinese Academy of Sciences), Zhen Lei (University of Chinese Academy of Sciences)
CodeAnomaly DetectionReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringContrastive LearningVideoMultimodality
π― What it does: This paper proposes the VIDEOVERITAS framework for detecting AI-generated videos;
π― What it does: Propose an architecture (VIAR) that replaces the deep explicit intermediate stack in visual autoregressive models (VAR) with a single implicit balanced layer, and achieves computational control at each scale through adjustable iteration counts.
CodeRobotic IntelligenceReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningVision Language ModelVision-Language-Action ModelContrastive LearningMultimodalityTime Series
π― What it does: This paper conducts an in-depth analysis of the Vision-Language-Action (VLA) design space through a unified framework and systematic evaluation, proposing a practical construction route from basic components, perceptual inputs to action modeling, and achieving a state-of-the-art VLANeXt model.
VT-Bench: A Unified Benchmark for Visual-Tabular Multi-Modal Learning
Ziyi Jia, Lan-Zhe Guo (Nanjing University)
CodeData-Centric LearningTransformerPrompt EngineeringVision Language ModelContrastive LearningImageMultimodalityTabularElectronic Health RecordsBenchmarkRetrieval-Augmented GenerationChain-of-Thought
π― What it does: Proposes VT-Bench, a unified visual-table multimodal learning benchmark covering two paradigms: discriminative prediction and generative reasoning;
WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving
Chiheng Lou (Peking University), Xin Jin (Peking University)
CodeOptimizationComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringMixture of ExpertsTextRetrieval-Augmented Generation
π― What it does: This paper proposes WarmServe, a multi-model GPU preheating system based on workload prediction, which can pre-load multi-model parameters before request peaks, significantly reducing inference time and improving throughput.
Weak-to-Strong Generalization via Bregman BiasβVariance Decomposition
Gengze Xu (Renmin University of China), Yong Liu (Renmin University of China)
CodeClassificationKnowledge DistillationRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningContrastive LearningTextReview/Survey Paper
π― What it does: This paper theoretically analyzes the weak-to-strong generalization (W2SG) phenomenon through Bregman bias-variance decomposition, and provides error inequalities that do not rely on convexity and realizability assumptions. It further explores the sufficient conditions for student models approaching the posterior mean teacher, as well as their impacts on student capacity, cross-entropy, and inverse cross-entropy training objectives.
Weakly Supervised Cross-Modal Learning for 4D Radar Scene Flow Estimation
Jingyun Fu (Zhejiang University), Na Zhao (Singapore University of Technology and Design)
CodeAutonomous DrivingOptimizationComputational EfficiencyRepresentation LearningConvolutional Neural NetworkRecurrent Neural NetworkVision Language ModelDiffusion modelScore-based ModelContrastive LearningSimultaneous Localization and MappingOptical FlowImageVideoPoint Cloud
π― What it does: Propose a weakly supervised cross-modal learning framework called IterFlow, which provides auxiliary supervision for 4D radar scene flow estimation during training using only RGB images and odometry;
π― What it does: Proposed MS-FLOW, which improves multivariate time series forecasting by introducing a sparse bottleneck in cross-variable information flow.
Clara Meister (EPFL), Tiago Pimentel (ETH ZΓΌrich)
CodeClassificationTransformerTextBenchmark
π― What it does: Propose UniLID, a language identification method based on UnigramLM, which utilizes the unigram frequency distribution of each language and discriminates the language through the most probable segmentation during inference.
What Linear Probes Miss: Multi-View Probing for Weight-Space Learning
Eunwoo Heo (Ulsan National Institute of Science and Technology), Jaejun Yoo (Ulsan National Institute of Science and Technology)
CodeExplainability and InterpretabilityRepresentation LearningConvolutional Neural NetworkTransformerPrompt EngineeringAuto EncoderContrastive LearningImage
π― What it does: This paper investigates the limitations of probing methods in weight space learning and proposes a multi-perspective probing framework called MVProbe, which comprehensively encodes the network weight matrix using first-order row/column and second-order Gram perspectives.
π― What it does: Investigated and systematically evaluated the effectiveness of synthetic images generated by large-scale diffusion models on semantic segmentation tasks, and proposed a unified framework called SENSE, which assigns robust pseudo-labels to synthetic data through optimal transport (OT) techniques, significantly improving segmentation performance.
CodeData SynthesisTabularBiomedical DataElectronic Health RecordsReview/Survey Paper
π― What it does: This paper conducts systematic experiments on survival model evaluation metrics under different censoring mechanisms and censoring rates using a semi-synthetic data generation method, comparing the differences between standard evaluation and full-information 'oracle' evaluation.
When Is Rank-1 Enough? Geometry-Guided Initialization for Parameter-Efficient Fine-Tuning
Haoran Zhao (University of Melbourne), Eduard Hovy (University of Melbourne)
CodeOptimizationComputational EfficiencyRepresentation LearningTransformerSupervised Fine-TuningVision Language ModelContrastive LearningImageTextMultimodality
π― What it does: Propose a geometry-guided initialization method (Gap-Init), which achieves parameter-efficient fine-tuning of multi-modal models by aligning the update direction of Rank-1 LoRA with the modality difference vector between visual-text features.
When LLMs Develop Languages: Symbolic Communication for Efficient Multi-Agent Reasoning
Zhengqi Pei (Chinese Academy Of Sciences), Shuhui Wang (Chinese Academy Of Sciences)
CodeAutonomous DrivingOptimizationComputational EfficiencyAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkChain-of-Thought
π― What it does: Propose the CLSR framework, enabling LLM multi-agent self-evolution and sharing of symbolic language (LSF), and dynamically combining them based on queries through a router without a latent network
When Planning Fails Despite Correct Execution: On Epistemic Calibration for LLM-Based Multi-Agent Systems
Zehao Wang (Tianjin University), Lanjun Wang (Tianjin University)
CodeAutonomous DrivingOptimizationExplainability and InterpretabilityRobotic IntelligenceReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AITextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought
π― What it does: Propose an Epistemic Planning Calibration Agentic Workflow (EPC-AW) based on information consistency and memory-driven mechanisms, which reduces cognitive errors during the planning phase in LLM-based multi-agent systems through cross-information condition plan consistency evaluation and cross-round calibration constraints, thereby improving overall task success rates.
π― What it does: Propose a training-agnostic and data-agnostic post-processing method called Singular Value Calibration (SVC), which identifies and corrects singular value inflation caused by spectral over-accumulation by measuring subspace overlap on the output space basis of the merged model, thereby improving the effectiveness of model merging.
π― What it does: Propose WEINCE, improving the softmax assumption of InfoNCE, utilizing extreme value theory to correct the shortage in the tail of hard negative samples, while keeping no additional parameters;
WinDeskGround: A Benchmark for Robust GUI Grounding in Complex Multi-Window Desktop Environments
Haoren Zhao (Hangzhou Dianzi University), Zhen Wang (Hangzhou Dianzi University)
CodeData SynthesisTransformerLarge Language ModelPrompt EngineeringVision Language ModelDiffusion modelImageTextMultimodalityBenchmark
π― What it does: This paper proposes the WinDeskGround benchmark, which utilizes a parameterized multi-window synthesis framework to generate complex desktop scenes across dimensions such as multi-window, occlusion, and semantic similarity, in order to evaluate the GUI localization robustness of multimodal large language models (MLLMs).
WinQ: Accelerating Quantization-Aware Training of Language Models Around Saddle Points
Dongyue Li (Northeastern University), Steven Li (Meta AI)
CodeOptimizationComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelContrastive LearningTextStochastic Differential Equation
π― What it does: Propose the WINQ method to accelerate low-precision language model quantization-aware training through weight linear interpolation reset and noise injection;
CodeGenerationTransformerLarge Language ModelPrompt EngineeringDiffusion modelImageTextMultimodalityBenchmarkChain-of-Thought
π― What it does: This paper proposes a benchmark called WISE to evaluate the ability of text-to-image models in world knowledge reasoning and complex semantic understanding.
π― What it does: This paper proposes WorldCache, a training-agnostic acceleration framework that significantly improves the inference speed of diffusion world models (such as HunyuanVoyager and Aether) through heterogeneous token caching.
π― What it does: Proposed the WorldComp2D framework, which maps local observations to a structured spatial semantic latent space and achieves object localization through a local locator.
XRPO: Pushing the Limits of GRPO with Targeted Exploration and Exploitation
Udbhav Bamba (Unaffiliated), Fan Lai (University Of Illinois Urbana Champaign)
CodeAI Code AssistantTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextSequentialChain-of-Thought
π― What it does: Propose the XRPO framework, improving rollout allocation based on GRPO, introducing ICL seeds and advantage sharpening techniques to enhance the reasoning and programming performance of large language models.
You Don't Protect if You Don't Expect: Breaking the Key Assumption behind CLIP's Test-Time Defenses
Ruize Zhang (Institute of Computing Technology, Chinese Academy of Sciences), Sheng Tang (Institute of Computing Technology, Chinese Academy of Sciences)
CodeAdversarial AttackTransformerVision Language ModelContrastive LearningImageTextMultimodality
π― What it does: This paper investigates the key assumptions of the CLIP model in test-time defense and proposes the CLIP-MAD attack strategy, revealing the vulnerability of existing defense methods when facing adaptive attacks.
π― What it does: Based on pre-trained models, we propose a zeroth-order forward-only trained spiking neural network method (SZO) to achieve online learning.