ACL 2026 Papers — Page 9
Annual Meeting of the Association for Computational Linguistics · 2296 papers
From Individual to Common: An Early Exploration of Consensus in Non-verifiable Data for Balanced Preference Optimization
Shangjian Yin (University of California Riverside), Zhouxing Shi (University of California Riverside)
OptimizationData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextRetrieval-Augmented Generation
🎯 What it does: Proposed the Donkey dataset and a training framework based on consensus preference optimization to improve the subjective alignment and objective performance of LLMs on non-verifiable data.
From Inheritance to Saturation: Disentangling the Evolution of Visual Redundancy for Architecture-Aware MLLM Inference Acceleration
Jiaqi Shi (University of Science and Technology of China), Jianzong Wang (Ping An Technology (Shenzhen) Co., Ltd.)
Computational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringVision Language ModelContrastive LearningImageVideoTextMultimodality
🎯 What it does: This study investigates the issue of visual token redundancy during inference in multimodal large language models, and proposes the HalfV framework. It first uniformly eliminates Intrinsic Visual Redundancy (IVR), and then adaptively accelerates inference based on Secondary Saturation Redundancy (SSR) specific to different architectures. By analyzing the truncated matrix entropy, it reveals the three-stage redundancy lifecycle, and achieves a cross-architecture general inference acceleration solution.
From Insight to Action: A Novel Framework for Interpretability-Guided Data Selection in Large Language Models
Ling Shi (Tianjin University), Weihua Luo (Alibaba Group)
Explainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningAuto EncoderText
🎯 What it does: Proposes the Interpretability-Guided Data Selection (IGDS) framework, which identifies internal causal features of LLMs and selects the data most capable of activating these features through a 'feature resonance' method to achieve efficient fine-tuning.
From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?
Aaron Mueller (Boston University), Patrik Reizinger (Mila Quebec AI Institute)
Explainability and InterpretabilityTransformerPrompt EngineeringDiffusion modelScore-based ModelFlow-based ModelRectified FlowNeural Radiance FieldAuto EncoderGenerative Adversarial NetworkContrastive LearningText
🎯 What it does: This study investigates the identifiability and manipulability of neural network explanation methods in multi-concept environments, proposing a multi-concept evaluation framework and verifying causal independence through steering experiments.
From Knowing to Teaching: Scaffolding Pedagogical Decisions for LLM Agent
Yucheng Wang (Tsinghua University), Bin Xu (Tsinghua University)
Explainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelAgentic AIPrompt EngineeringTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This study proposes the TeachCraft framework, which utilizes three agents (Explorer, Planner, Generator) to construct pedagogical materials that align with Pedagogical Content Knowledge from multi-source texts, bridging the 'knowing-to-teaching' gap in teaching design for LLMs.
From Laboratory to Real-World Applications: Benchmarking Agentic Code Reasoning at the Repository Level
Jia Li (Chinese University of Hong Kong), Michael R. Lyu (Chinese University of Hong Kong)
AI Code AssistantLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Designed and implemented RepoReason, a white-box diagnostic benchmark for warehouse-level code reasoning, generating non-memorized reasoning tasks through assertion validation and execution-driven mutation.
From Language to Driving: A Dual-Loop SLM-Enhanced Framework for Multi-Planner Scheduling via a Domain-Specific Language
Jiawei Liu (Jilin University), Qing Guo (Nankai University)
Autonomous DrivingOptimizationTransformerLarge Language ModelReinforcement LearningPrompt EngineeringVision-Language-Action ModelDiffusion modelTextSequential
🎯 What it does: This study proposes a dual-loop SLM enhancement framework that converts passenger open-ended instructions into vehicle control commands. The outer low-frequency loop uses a small language model to generate SchedulingDSL scripts for semantic scheduling of multiple planners; the inner high-frequency loop executes specific planning and control based on MPC. By employing recursive rolling windows, DSL constraints, and RL fine-tuning, the scheduling capability of the small model is improved, and safety, compliance, and comfort are verified in high-fidelity urban simulations.
From Local to Global: Revisiting Structured Pruning Paradigms for Large Language Models
Ziyan Wang (University of North Carolina at Charlotte), Li Yang (University of North Carolina at Charlotte)
Computational EfficiencyKnowledge DistillationTransformerLarge Language ModelPrompt EngineeringContrastive LearningText
🎯 What it does: Proposed the global iterative structured pruning method GISP, which can compress large language models to a high degree according to task objectives without fine-tuning.
From Logical to Computational Sparsity: Structure-Aware Block-Sparse Attention for Long-Code Completion
Yanli Wang (Sun Yat-sen University), Zibin Zheng (Sun Yat-sen University)
Computational EfficiencyAI Code AssistantTransformerLarge Language ModelPrompt EngineeringTextSequentialRetrieval-Augmented Generation
🎯 What it does: Proposes SabreCoder, a training-agnostic structural-aware block-sparse attention mechanism for long code completion.
From Narrow Unlearning to Emergent Misalignment in LLMs
Erum Mushtaq (University Of Southern California), Rahul Gupta (Amazon Agi)
Safty and PrivacyExplainability and InterpretabilityTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextBenchmark
🎯 What it does: This paper investigates whether narrowly targeting 'forgetting' interventions on already aligned large language models to reduce rejection rates on target concepts would trigger emergent mismatch across concepts (EMA) and how to control it.
From Naturalness to Norms: Interactional Cultural Competence for SpeechLMs
Santosh T.Y.S.S
RecognitionRecommendation SystemExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringTextReview/Survey PaperRetrieval-Augmented GenerationAudio
🎯 What it does: This paper redefines the cultural competence of speech models as interactive cultural competitiveness, emphasizing the decisive role of speech events and interaction contracts in interactive behavior.
From Nodes to Narratives: Explaining Graph Neural Networks with LLMs and Graph Context
Peyman Baghershahi (University of Illinois Chicago), Sourav Medya (University of Illinois Chicago)
Explainability and InterpretabilityRepresentation LearningGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextGraph
🎯 What it does: This paper proposes the GSPELL framework, which utilizes large language models (LLMs) to generate natural language explanations and interpretable subgraphs without using an external interpreter, by projecting GNN node embeddings into the LLM embedding space.
From Off-Policy to On-Policy: Enhancing GUI Agents via Bi-level Expert-to-Policy Assimilation
Zezhou Wang (Nanjing University), Yan Lu (Microsoft Research Asia)
Robotic IntelligenceTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringWorld ModelTextSequentialRetrieval-Augmented Generation
🎯 What it does: Propose a dual-layer expert-to-policy fusion framework called BEPA, aimed at improving the performance of end-to-end GUI agents under a limited number of expert trajectories
From Past To Path: Masked History Learning for Next-Item Prediction in Generative Recommendation
Kaiwen Wei (Chongqing University), Junnan Zhu (Chinese Academy Of Sciences)
Recommendation SystemTransformerPrompt EngineeringAuto EncoderContrastive LearningTextRetrieval-Augmented Generation
🎯 What it does: Constructed the Mask-History Learning (MHL) framework, which simultaneously trains the prediction of the next item and the reconstruction of masked items in the historical trajectory in generative recommendation, enhancing the model's understanding of logical relationships in user history.
From Proof to Program: Characterizing Tool-Induced Reasoning Hallucinations in Large Language Models
Farima Fatahi Bayat (Megagon Labs), Estevam Hruschka (Megagon Labs)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningAI Code AssistantTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposed and studied Tool-Induced Myopia (TIM), which refers to the phenomenon where models tend to replace internal reasoning with tool results when using external tools (such as the Code Interpreter), leading to a decrease in reasoning depth, although the answers may still be correct.
From Query to Counsel: Structured Reasoning with a Multi-Agent Framework and Dataset for Legal Consultation
Mingfei Lu (University of Technology Sydney), Yue Feng (University of Birmingham)
Recommendation SystemTransformerLarge Language ModelAgentic AIPrompt EngineeringTextGraphBenchmarkRetrieval-Augmented Generation
🎯 What it does: This paper proposes a legal consultation question-answering system (JURISMA) based on a multi-agent framework, which generates professional and compliant legal advice by parsing queries into legal semantic graphs and performing iterative collaborative reasoning.
From Recognition to Reasoning: Benchmarking and Enhancing MLLMs on Real-World Receipt Document Understanding
Yandi Wang (Zhejiang University), Jun Chen (Zhejiang University)
RecognitionTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision Language ModelImageTextMultimodalityBenchmark
🎯 What it does: Constructed a multilingual receipt and invoice benchmark called ReceiptBench with 10k scale, and proposed a two-stage training approach (SFT + Metric-Aware GRPO) to enhance the perception, normalization, reasoning, and structural parsing capabilities of multimodal large language models in receipt understanding.
From Regulatory Approvals to Patents: Cross-Domain Linking for Cardiovascular Device Traceability
Qingqing Yang, Moyan Li (Hong Kong University of Science and Technology (Guangzhou))
Drug DiscoveryTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningTextBiomedical DataElectronic Health RecordsReview/Survey PaperBenchmarkRetrieval-Augmented Generation
🎯 What it does: This study proposes the Bridge-MedDevKG framework, which aligns FDA-approved medical devices with their related U.S. patents across domains and constructs the first benchmark for cardiovascular device-patent alignment;
From Selection to Refinement: Iterative Optimization for Instruction Data
Hang Hu (East China University of Science and Technology), Jingping Liu (Sun Yat-sen University)
OptimizationData-Centric LearningTransformerLarge Language ModelPrompt EngineeringText
🎯 What it does: Built an automated iterative framework for optimizing instruction data, through hierarchical screening of high-quality samples and a multi-round feedback-driven evaluation-refinement-review process for low-quality samples, ultimately generating high-quality instruction-input-output triplets;
From Synthesis to Clinical Assistance: A Strategy-Aware Agent Framework for Autism Intervention based on Real Clinical Dataset
Junhong Lai (Zhejiang University), Yueming Wang (Zhejiang University)
Data SynthesisRecommendation SystemExplainability and InterpretabilityKnowledge DistillationTransformerLarge Language ModelPrompt EngineeringTextBiomedical DataElectronic Health RecordsChain-of-Thought
🎯 What it does: Proposed the ASDAGENT framework, combining DOCTORAGENT (O-T-A-C iterative reasoning) and CHILDAGENT (probabilistic behavior simulation), to achieve high-fidelity autism intervention dialogue synthesis and clinical decision support.
From TDMA to CDMA: A Multi-bit Watermark for Diffusion Language Models
Baizhou Huang (Peking University), Xiaojun Wan (Peking University)
Safty and PrivacyData-Centric LearningTransformerLarge Language ModelMixture of ExpertsDiffusion modelTextFinance Related
🎯 What it does: Implement multi-bit watermarking in diffusion language models (DLM) and propose a new global encoding framework;
From Trajectories to Graphs: Contract-Checked Editing for Verifier-Guided LLM Reasoning
Rui Li (Hill Research), Shuang Cao (Hill Research)
Explainability and InterpretabilityComputational EfficiencyAI Code AssistantGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringTextGraphSequentialRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose a contract-checked editing framework based on structured graphs, using interface-typed inference DAGs for reliable cross-candidate recombination, and separating structural validity verification from semantic verification.
From Trust to Compromise: Outcome-Verified LLM Phishing Simulation and Real-Time Defense
Tulika Tewari (BITS Pilani), Dhruv Kumar (BITS Pilani)
Anomaly DetectionSafty and PrivacyTransformerLarge Language ModelAgentic AITextRetrieval-Augmented Generation
🎯 What it does: This paper develops PhishSim, a result-driven multi-round social engineering simulator that can generate 1,833 dialogues and verify the success of the attack by simulating victims completing offline credential submission; it also proposes PhishGate, a real-time multi-agent risk scorer that uses three submodules—Tactic, Severity, and RAG—to conduct round-by-round risk assessment of dialogues and issue early alerts.
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents
Niu Lian (Tsinghua Shenzhen International Graduate School, Tsinghua University), Shu-Tao Xia (Tsinghua Shenzhen International Graduate School, Tsinghua University)
RetrievalComputational EfficiencyKnowledge DistillationRepresentation LearningTransformerReinforcement LearningVision Language ModelContrastive LearningVideoTextMultimodalityBenchmarkRetrieval-Augmented Generation
🎯 What it does: Designed and implemented a pyramid multi-modal memory architecture called MM-MEM based on fuzzy trace theory, to address the challenges of memory and retrieval in long video understanding.
From Weights to Activations: Is Steering the Next Frontier of Adaptation?
Simon Ostermann (Saarland University), Vera Schmitt (German Research Center for Artificial Intelligence)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextMultimodalityReview/Survey PaperRetrieval-Augmented Generation
🎯 What it does: This paper proposes viewing internal model activation intervention (steering) as a model adaptation method, establishing a set of functional evaluation criteria, and systematically comparing steering with traditional fine-tuning, parameter-efficient fine-tuning (PEFT), and prompting using these criteria, highlighting steering's unique advantages in terms of locality, reversibility, and resource efficiency.
From Where Words Come: Efficient Regularization of Code Tokenizers Through Source Attribution
Pavel Chizhov (Technical University of Applied Sciences Würzburg-Schweinfurt), Ivan P. Yamshchikov (Technical University of Applied Sciences Würzburg-Schweinfurt)
Computational EfficiencyRepresentation LearningAI Code AssistantTransformerLarge Language ModelContrastive LearningTextSequential
🎯 What it does: Studied the overfitting problem of tokenizers in code language models, and proposed and implemented the Source-Attributed BPE (SA-BPE) regularization method.
From Word Sequences to Behavioral Sequences: Adapting Modeling and Evaluation Paradigms for Longitudinal NLP
Adithya V Ganesan (Stony Brook University), H. Andrew Schwartz (Stony Brook University)
Explainability and InterpretabilityRepresentation LearningData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelContrastive LearningTextTime SeriesBiomedical DataElectronic Health Records
🎯 What it does: Proposed and systematically validated the 'Person-Time Alignment Assessment and Modeling' paradigm for longitudinal NLP, demonstrating that traditional random splitting can lead to misleading evaluation results.
From Word to World: Can Large Language Models be Implicit Text-based World Models?
Yixia Li (Southern University of Science and Technology), Heng Ji (University of Illinois Urbana-Champaign)
Robotic IntelligenceTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringWorld ModelTextRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This study investigates whether large language models can serve as world models in textual environments and proposes a three-tier evaluation framework, systematically verifying their feasibility in five textual environments and their enhancement of downstream agents.
From Words to Pixels: A Comprehensive Survey on Large Language Models in Visual Segmentation
Yizhou Wang (Northeastern University), Yun Fu (Northeastern University)
SegmentationTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageVideoTextReview/Survey PaperChain-of-Thought
🎯 What it does: Reviews the latest research on large language models in visual segmentation, proposing various task categories such as reasoning segmentation, pixel-level localization, open vocabulary, and video segmentation, and systematically summarizes their technical implementations, dataset usage, and evaluation metrics.
Frozen LLMs are Native Decoders for High-Norm Semantic Vectors
Yunsheng Zeng (Beijing University of Posts and Telecommunications), Yongmei Tan (Beijing University of Posts and Telecommunications)
CompressionRepresentation LearningTransformerLarge Language ModelText
🎯 What it does: The study investigates the decoding capability of frozen large language models on continuous high norm vectors and proposes a milestone-based non-block compression framework;
FS-Researcher: Test-Time Scaling for Long-Horizon Research Tasks with File-System-Based Agents
Chiwei Zhu (University of Science and Technology of China), Yongdong Zhang (Metastone Technology)
Autonomous DrivingOptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningAdversarial AttackHyperparameter SearchData-Centric LearningRobotic IntelligenceMeta LearningDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose the FS-Researcher framework, which achieves test-time scaling for deep research through dual agents and a file system workspace
Full-Duplex-Bench-v2: A Multi-Turn Evaluation Framework for Duplex Dialogue Systems with an Automated Examiner
Guan-Ting Lin (National Taiwan University), Hung-yi Lee (National Taiwan University)
TransformerLarge Language ModelPrompt EngineeringTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-ThoughtAudio
🎯 What it does: Proposed the Full-Duplex-Bench-v2 (FDB-v2) framework, which can automatically evaluate the interaction fluency, instruction execution, and task-specific capabilities of models in multi-turn, full-duplex conversations;
Function Words as Statistical Cues for Language Learning
Xiulin Yang (Georgetown University), Ethan Gotlieb Wilcox (Georgetown University)
Explainability and InterpretabilityRepresentation LearningTransformerLarge Language ModelContrastive LearningTextBenchmark
🎯 What it does: Perform statistical analysis on the Universal Dependencies corpus covering 186 languages to verify the universality of function words in three aspects: high frequency, structural predictability, and phrase boundary alignment. Construct six contrastive conditions with six functional word attributes on English Wikipedia text, train Transformer language models and n-gram models, evaluate their performance on syntactic generalization in BLiMP, and explore the model's internal representation and utilization of function words through attention probing and function word ablation experiments.
FusionFlow: Enabling Deep Structural Exploration for Automated Agentic Workflow Generation
Xiang Wang (Huazhong University of Science and Technology), Wei Wei (Huazhong University of Science and Technology)
OptimizationAI Code AssistantReinforcement Learning from Human FeedbackNeural Architecture SearchTransformerLarge Language ModelPrompt EngineeringTextTabularBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose a fusion-based auto-generation framework called FUSIONFLOW, addressing the issue of structural depth limitations caused by traditional local optimization.
G-Cap: A Game Character Caption Generator
Yang Yang (Sun Yat-sen University), Wenqi Ren (Sun Yat-sen University)
GenerationData SynthesisTransformerLarge Language ModelSupervised Fine-TuningVision Language ModelDiffusion modelContrastive LearningImageTextMultimodalityGraphBenchmarkRetrieval-Augmented Generation
🎯 What it does: This paper proposes the task of generating image captions for game characters, constructs the GC-Bench benchmark and the Graph-F1 evaluation metric, further builds the GC-148K large-scale dataset, and fine-tunes the G-Cap series of models on it.
G-IdiomAlign: A Gloss-Pivoted Benchmark for Cross-Lingual Idiom Alignment
Fengying Ye (University of Macau), Derek F. Wong (University of Macau)
TransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Constructed a cross-lingual idiom alignment benchmark, G-IdiomAlign, based on English definitions, and proposed two evaluation protocols (multiple-choice and contrastive generation), providing a reproducible diagnostic framework for cross-lingual idiom equivalence.
G^2RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance
Yongxin Guo (Chinese University of Hong Kong), Xiaoying Tang (Chinese University of Hong Kong)
OptimizationAI Code AssistantTransformerSupervised Fine-TuningReinforcement LearningContrastive LearningTextBenchmarkChain-of-Thought
🎯 What it does: Propose an adaptive guided GRPO algorithm (GRPO-A) for small language model training, enhancing the model's reasoning ability by injecting high-quality reasoning steps during the roll-out process.
Gained in Translation: Privileged Pairwise Judges Enhance Multilingual Reasoning
Lintang Sutawika (Carnegie Mellon University), Graham Neubig (Carnegie Mellon University)
Computational EfficiencyData-Centric LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextMultimodalityRetrieval-Augmented Generation
🎯 What it does: Propose the SP3F framework in the absence of target language data, first using translated English QA for supervised fine-tuning, and then training through self-adversarial RL and pairwise judge with restricted information.
GAM: Hierarchical Graph-based Agentic Memory for LLM Agents
Zhaofen Wu (Zhejiang University), Hongwei Wang (Zhejiang University)
Safty and PrivacyExplainability and InterpretabilityComputational EfficiencyGraph Neural NetworkTransformerLarge Language ModelAgentic AIPrompt EngineeringTextGraphRetrieval-Augmented Generation
🎯 What it does: Proposed the GAM (Hierarchical Graph-based Agentic Memory) framework to achieve the separation of event buffering and semantic integration in LLM agents, thereby improving the consistency and accuracy of long-term conversations.
GAMBIT: A Gamified Jailbreak Framework for Multimodal Large Language Models
Xiangdong Hu (Georgia State University), Xiaojun Jia (Nanyang Technological University)
Safty and PrivacyAdversarial AttackTransformerLarge Language ModelPrompt EngineeringVision Language ModelDiffusion modelImageTextMultimodality
🎯 What it does: This paper proposes a multi-modal jailbreak framework called GAMBIT, which uses three modules: image jigsaw encoding, gamified scenarios, and adaptive prompt search, to induce multi-modal large language models to ignore safety filters and generate harmful responses while completing 'game' tasks.
GameplayQA: A Benchmarking Framework for Decision-Dense POV-Synced Multi-Video Understanding of 3D Virtual Agents
Yunzhe Wang (University of Southern California), Volkan Ustun (University of Southern California)
Large Language ModelPrompt EngineeringVision Language ModelVideoTextMultimodalityBenchmarkRetrieval-Augmented Generation
🎯 What it does: Built and released GAMEPLAYQA, a decision-intensive video question answering benchmark framework based on synchronized multi-perspective 3D multiplayer game videos; the framework includes high-frequency (≈1.22 tags/second) multi-track timeline annotations, Self-Other-World tripartite entity partitioning, 15 types of cross-cognitive level (L1 single reference, L2 temporal reasoning, L3 cross-video reasoning) question answering generation, and structured error inducers; it also provides 2.4K diagnostic multiple-choice QA and evaluation scripts.
Gap-K%: Measuring Top-1 Prediction Gap for Detecting Pretraining Data
Minseo Kwak (Yonsei University), Jaehyung Kim (Yonsei University)
Anomaly DetectionExplainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelContrastive LearningText
🎯 What it does: A pre-training data detection method based on the training dynamics of autoregressive language models, called Gap-K%, is studied. It distinguishes between training set samples and untrained samples by measuring the log probability difference between the target word and the model's highest probability word.
GASE: Graph-Aware Semantic Embedding Learning with Frozen LLMs for Text-Attributed Graphs
Mingqian Ding, Yushen Fang (Huazhong University of Science and Technology)
Computational EfficiencyRepresentation LearningGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextGraph
🎯 What it does: Proposes the GASE framework, which learns node representations of text attribute graphs (TAG) through a two-stage method (TSSE + SDSI) by leveraging a frozen large language model (LLM), balancing text semantics and graph structure;
GASim: A Graph-Accelerated Hybrid Framework for Social Simulation
Xuan Zhou (University of Science and Technology of China), Wu Liu (University of Science and Technology of China)
Computational EfficiencyReinforcement Learning from Human FeedbackGraph Neural NetworkTransformerLarge Language ModelTextGraphRetrieval-Augmented Generation
🎯 What it does: Proposed GASim, a graph-accelerated hybrid multi-agent framework for large-scale social simulations, addressing the high latency issues caused by traditional frameworks where LLM retrieval and ABM sequential execution are used.
GATE: Graph-based Adaptive Tool Evolution Across Diverse Tasks
Jianwen Luo (Key Laboratory of Cognition and Decision Intelligence for Complex Systems), Kang Liu (Key Laboratory of Cognition and Decision Intelligence for Complex Systems)
Computational EfficiencyRepresentation LearningAI Code AssistantReinforcement Learning from Human FeedbackGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringScore-based ModelContrastive LearningTextGraphRetrieval-Augmented Generation
🎯 What it does: Proposed the GATE framework, which utilizes hierarchical tool graphs and a dual-agent mechanism to achieve cross-task tool adaptive evolution, supporting automatic retrieval, generation, merging, and pruning of tools.
Gated Differentiable Working Memory for Long-Context Language Modeling
Lingrui Mei (State Key Laboratory of AI Safety), Xueqi Cheng (State Key Laboratory of AI Safety)
Computational EfficiencyRepresentation LearningTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented Generation
🎯 What it does: Proposed the GDWM framework, which uses a write controller to perform budget-constrained memory consolidation on long texts, significantly reducing gradient steps and improving the inference efficiency and performance of long-context language models.
Gated Tree Cross-Attention for Checkpoint-Compatible Syntax Injection in Decoder-Only LLMs
Xinyu Gao (Zhejiang University), Nai Ding (Zhejiang University)
Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextRetrieval-Augmented Generation
🎯 What it does: A checkpoint-compatible gated tree cross-attention (GTCA) branch is proposed, which can inject constituent syntactic information into decoder-only LLMs without modifying the backbone structure, thereby enhancing syntactic robustness.
GBV-SQL: Guided Generation and SQL2Text Back-Translation Validation for Multi-Agent Text2SQL
Daojun Chen (Soochow University), An Liu (Soochow University)
Data-Centric LearningAI Code AssistantTransformerLarge Language ModelAgentic AIPrompt EngineeringTextTabularBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This paper proposes the GBV-SQL multi-agent framework, which uses the SQL2Text reverse translation verification mechanism to ensure that the generated SQL semantics are consistent with the original natural language question; meanwhile, the quality of the Text2SQL benchmark data is systematically audited, and a 'Gold Errors' classification system is constructed; significant improvements in execution accuracy are achieved on the BIRD and Spider benchmarks.
GCA Framework: A GCC Countries–Grounded Dataset and Agentic Pipeline for Climate Decision Support
Muhammad Umer Sheikh (Mohamed Bin Zayed University of Artificial Intelligence), Muhammad Haris Khan (Mohamed Bin Zayed University of Artificial Intelligence)
Explainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningAgentic AIVision Language ModelImageTextMultimodalityTime SeriesRetrieval-Augmented Generation
🎯 What it does: Proposed the Gulf Climate Agent (GCA) framework, combining the multi-modal climate question-answering dataset for the GCC region (GCA-DS) with tool-enhanced language agents, to achieve interpretable, data-driven question-answering for climate decision-making in the Gulf region;
GenDis: Generative-Discriminative Dual-View Co-Training for Generalized Category Discovery
Xi Chen (University of Science and Technology of China), Hui Xiong (Hong Kong University of Science and Technology (Guangzhou))
ClassificationRepresentation LearningData-Centric LearningTransformerLarge Language ModelPrompt EngineeringDiffusion modelContrastive LearningTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Improved the task of generalized category discovery (GCD) by jointly training a large language model (LLM) with dual views, combining discriminative classification and generative labeling.
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models
Lillian Sun (Harvard University), Himabindu Lakkaraju (Harvard University)
Federated LearningSafty and PrivacyExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningAdversarial AttackTransformerLarge Language ModelSupervised Fine-TuningDiffusion modelScore-based ModelContrastive LearningTextTabularTime SeriesSequentialStochastic Differential Equation
🎯 What it does: This paper studies how to improve the trustworthiness of strong models through weak model labels in the absence of real labels, proposing the Weak-to-Strong Trustworthiness framework, and evaluating four key attributes: fairness, robustness, and privacy.
Generating Attribution Reports for Manipulated Facial Images: A Dataset and Baseline
Jingchun Lian (Xi'an Jiaotong University), Zhedong Zheng (University of Macau)
Anomaly DetectionExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextMultimodality
🎯 What it does: Propose the Forgery Attribution Report Generation task, jointly locating forged regions and generating natural language explanations, while constructing the MMTT large-scale dataset.
Generating Literature-Driven Scientific Theories at Scale
Peter Jansen (Allen Institute for Artificial Intelligence), Daniel S Weld
GenerationExplainability and InterpretabilityKnowledge DistillationData-Centric LearningTransformerLarge Language ModelPrompt EngineeringTextReview/Survey PaperRetrieval-Augmented Generation
🎯 What it does: Developed a system called THEORIZER for automatically generating theories from a large volume of scientific literature.
Generating then Refining for Reliable Knowledge Base Question Answering
Jianqi Gao (Shanghai University), Yonggang Zhang (Hong Kong University of Science and Technology)
Recommendation SystemOptimizationExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextGraphBenchmarkRetrieval-Augmented Generation
🎯 What it does: This paper proposes a generation-verification-refinement framework called ARI-KBQA for reliably generating executable logical forms in knowledge graph question answering.
Generative Gamer: Learning Equilibrium Strategy by LLM-driven Dynamic Deduction
Yadong Zhang (East China Normal University), Man Lan (East China Normal University)
OptimizationTransformerLarge Language ModelReinforcement LearningTextSequentialChain-of-Thought
🎯 What it does: Train LLM to generate dynamic reasoning trees, simulating expert search to solve strategy games;
GenesisFunc: Multi-Agent Data Generation for Accurate and Generalizable Function-Calling
Hao-Xiang Xu (University of Science and Technology of China), Zhen-Hua Ling (University of Science and Technology of China)
GenerationData SynthesisReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningAgentic AIPrompt EngineeringTextSequentialRetrieval-Augmented Generation
🎯 What it does: Built the GENESISFUNC multi-agent data generation pipeline, automatically generating high-quality, scalable multi-turn function call training data, covering multiple tools, multiple tasks, and multiple scenarios;
GenomeQA: Benchmarking General Large Language Models for Genome Sequence Understanding
Weicai Long (Hong Kong University of Science and Technology (Guangzhou)), Yanlin Zhang (Hong Kong University of Science and Technology (Guangzhou))
Protein Structure PredictionTransformerLarge Language ModelPrompt EngineeringTextSequentialBiomedical DataBenchmarkChain-of-Thought
🎯 What it does: Propose the GenomeQA benchmark, which designs 5,200 questions based on original DNA sequences, covering six major tasks: enhancer/promoter identification, splicing sites, classification, histone marks, transcription factor binding sites, and motif prediction;
GenProve: Learning to Generate Text with Fine-Grained Provenance
Jingxuan Wei (Key Laboratory of Computing Power Network and Information Security Ministry of Education Shandong Computer Science Center National Supercomputer Center in Jinan Qilu University of Technology Shandong Academy of Sciences), Junnan Zhu (Institute of Automation Chinese Academy of Sciences)
GenerationExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextRetrieval-Augmented Generation
🎯 What it does: Propose the fine-grained source proof generation task and design the GenProve framework to enable LLMs to simultaneously output sentence-level structured proof triplets when generating answers.
GenPT: Beyond Self-Report for Reliable LLM Psychometrics via Generative Projective Testing
Ming Wang (Northeastern University), Yufan Sun (Northeastern University)
Explainability and InterpretabilityData-Centric LearningTransformerLarge Language ModelAgentic AIPrompt EngineeringDiffusion modelImageTextTabularBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposed and implemented the GenPT generative projective test framework to evaluate the psychological characteristics of LLM agents, addressing issues of training data contamination and social desirability bias in self-report questionnaires.
GeoArena: Evaluating Open-World Geographic Reasoning in Large Vision-Language Models
Pengyue Jia (City University of Hong Kong), Sharon Li (University of Wisconsin-Madison)
Explainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelVision Language ModelImageTextMultimodalityBenchmarkChain-of-Thought
🎯 What it does: Proposes GeoArena, a dynamic, human-preference-based evaluation framework designed to measure the capabilities of large-scale vision-language models in open-world geographic reasoning.
GeoLaux: A Benchmark for Evaluating MLLMs’ Geometry Performance on Long-Step Problems Requiring Auxiliary Lines
Yumeng Fu (Xi'an Jiaotong University), Jun Liu (Xi'an Jiaotong University)
TransformerLarge Language ModelPrompt EngineeringVision Language ModelTextMultimodalityBenchmarkChain-of-Thought
🎯 What it does: Propose the GeoLaux dataset and a five-dimensional evaluation framework to assess the long-step reasoning and auxiliary line construction capabilities of multimodal large language models in geometric reasoning.
GeoRA: Geometry-Aware Low-Rank Adaptation for RLVR
Jiaying Zhang (Meituan), Renqing He (Meituan)
Computational EfficiencyRepresentation LearningReinforcement Learning from Human FeedbackTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringTextTabular
🎯 What it does: Proposes GeoRA, a geometry-aware low-rank adaptation method for Reinforcement Learning with Verifiable Rewards (RLVR), aiming to efficiently perform parameter fine-tuning while preserving the pre-trained structure.
GeoRC: A Benchmark for Geolocation Reasoning Chains
Mohit Talreja (Georgia Institute of Technology), James Hays (Georgia Institute of Technology)
TransformerLarge Language ModelPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Create the GeoRC benchmark, collecting 800 geographic reasoning chains written by three GeoGuessr world champion-level experts, and evaluate the quality of reasoning chains generated by VLMs.
GIFT: Guided Fine-Tuning and Transfer for Enhancing Instruction-Tuned Language Models
Zhiwen Ruan (Southern University of Science and Technology), Guanhua Chen (Southern University of Science and Technology)
Knowledge DistillationTransformerLarge Language ModelSupervised Fine-TuningTextBenchmark
🎯 What it does: This paper proposes the GIFT framework, which guides the fine-tuning of low-rank adapters on a pre-trained base model by leveraging the confidence of the instruction model, and merges them into the instruction model to achieve task specialization while maintaining instruction following capability.
GiLT: Augmenting Transformer Language Models with Dependency Graphs
Tianyu Huang (ShanghaiTech University), Kewei Tu (ShanghaiTech University)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerLarge Language ModelContrastive LearningTextGraph
🎯 What it does: Propose a new Transformer language model called GiLT, which enhances the syntactic structure during the parsing process by utilizing dependency graphs.
GitChameleon 2.0: Evaluating AI Code Generation Against Python Library Version Incompatibilities
Diganta Misra (ELLIS Institute Tübingen), Massimo Caccia (ServiceNow Research)
AI Code AssistantTransformerLarge Language ModelAgentic AITextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposes GitChameleon 2.0, an executable, version-constrained Python code generation benchmark containing 328 problems and complete unit tests.
GKnow: Measuring the Entanglement of Gender Bias and Factual Gender
Leonor Veloso (LMU Munich), Hinrich Schuetze
Explainability and InterpretabilityTransformerLarge Language ModelTextBenchmark
🎯 What it does: This paper constructs the GKnow benchmark, systematically analyzing the high coupling between gender bias and factual gender information at the circuit layer and neuron layer in language models, and verifying the destructive impact of ablation on factual gender knowledge through neuron ablation experiments.
GLARE: Agentic Reasoning for Legal Judgment Prediction
Xinyu Yang (Renmin University of China), Zhicheng Dou (Renmin University of China)
ClassificationExplainability and InterpretabilityTransformerLarge Language ModelPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose the GLARE framework, which leverages LLM to actively retrieve and utilize external legal knowledge for comparative reasoning, significantly improving the accuracy of legal judgment prediction.
GLIER: Generative Legal Inference and Evidence Ranking for Legal Case Retrieval
Minghan Li (Soochow University), Guodong Zhou (Soochow University)
RetrievalExplainability and InterpretabilityKnowledge DistillationTransformerLarge Language ModelTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposes the GLIER framework, which utilizes generative legal reasoning and evidence fusion for legal case retrieval.
Global Adaptive Momentum Meets Local Personalized Perturbation: Efficient Federated LLM Fine-Tuning with Zeroth-Order Gradients
Zihan Chen (Singapore University Of Technology And Design), Kai Fong Ernest Chong (Singapore University Of Technology And Design)
OptimizationFederated LearningComputational EfficiencyTransformerLarge Language ModelSupervised Fine-TuningText
🎯 What it does: Proposed a federated LLM fine-tuning framework based on zeroth-order (ZO) optimization, which redesigned the ZO gradient processing pipeline through global adaptive momentum and local personalized perturbation, aiming to address challenges such as memory, communication, and data heterogeneity in federated learning.
Glyph: Scaling Context Windows via Visual-Text Compression
Jiale Cheng (Tsinghua University), Minlie Huang (Tsinghua University)
Computational EfficiencyRepresentation LearningData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningVision Language ModelContrastive LearningImageTextMultimodalityBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose the Glyph framework, which renders long text into a visual page and then processes it with a vision-language model, achieving scalability for long contexts.
GMFL: Efficient Global Masking for Federated LLM Fine-tuning
Xin Huang (South China University of Technology), Xinglin Zhang (South China University of Technology)
Federated LearningComputational EfficiencyTransformerLarge Language ModelSupervised Fine-TuningContrastive LearningText
🎯 What it does: Propose GMFL, a dynamic layer freezing mechanism based on global update magnitude, for fine-tuning LLMs in federated learning, significantly reducing communication and computational costs;
GMoE: Global Mixture of Experts with Logit Propagation
Geonwoo Hong (Ulsan National Institute of Science and Technology), Taehwan Kim (Ulsan National Institute of Science and Technology)
Computational EfficiencyKnowledge DistillationRepresentation LearningTransformerLarge Language ModelMixture of ExpertsText
🎯 What it does: Propose a GMoE architecture that achieves parameter efficiency in sparse mixture-of-experts networks through shared global experts, local experts, and a global router.
GMSA: Enhancing Context Compression via Group Merging and Layer Semantic Alignment
Jiwei Tang (Tsinghua University), Hong-Gee Kim (Seoul National University)
CompressionRepresentation LearningTransformerLarge Language ModelPrompt EngineeringAuto EncoderText
🎯 What it does: Propose a context compression framework called GMSA, which generates short and soft prompts by utilizing Group Merging and Layer Semantic Alignment to achieve efficient compression of long contexts while preserving information.
Gold-Medal-Level Olympiad Geometry Solving with Efficient Heuristic Auxiliary Constructions
Boyan Duan (ETH Zurich), Yeyun Gong (Microsoft)
OptimizationComputational EfficiencyReinforcement LearningContrastive LearningTextTabularBenchmarkChain-of-Thought
🎯 What it does: Proposes a CPU-based method for proving Euclidean geometry theorems called HAGeo, which uses heuristic-assisted construction to enhance reasoning performance, ultimately achieving gold-level performance on the IMO-30 problem set and making significant progress on the newly constructed HAGeo-409 problem set.
GOLEMcoref: A Multilingual Coreference Dataset of Fiction
Andreas van Cranenburgh (University of Groningen), Federico Pianzola (University of Groningen)
RecognitionTransformerSupervised Fine-TuningPrompt EngineeringContrastive LearningTextMultimodalityBenchmark
🎯 What it does: This paper creates a multilingual novel coreference dataset spanning seven languages (Indonesian, Chinese, Dutch, English, Italian, Korean, Spanish), containing full stories (500–17,000 words) with annotations focused on character coreference, zero pronouns, and split antecedents; it also provides two formats: CoNLL-2012 and CorefUD; and systematically explains the annotation process, challenges, and quality.
Good Arguments Against the People Pleasers: How Reasoning Mitigates (Yet Masks) LLM Sycophancy
Zhaoxin Feng (Hong Kong Polytechnic University), Bo Li (Hong Kong University of Science and Technology)
Recommendation SystemFederated LearningExplainability and InterpretabilityReinforcement Learning from Human FeedbackTransformerLarge Language ModelPrompt EngineeringTextBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: The study investigates whether using chain-of-thought (CoT) alignment reduces or masks the sycophancy of LLMs when facing user or authoritative bias, and reveals their dynamic behavior through multi-dimensional evaluation and internal mechanism analysis.
GQLBench: A Large-Scale Cross-Domain, Cross-Dialect Benchmark for NL2GQL
Yanning Su (Fudan University), Hongfeng Chai (Fudan University)
Data SynthesisTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextGraphTabularBenchmark
🎯 What it does: Constructed GQLBench, a cross-domain, cross-dialect, executable NL2GQL benchmark, and provided corpus based on automated migration and synthesis.
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models
Zhaohan Zhang (Queen Mary University of London), Ioannis Patras (Queen Mary University of London)
GenerationExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringTextRetrieval-Augmented Generation
🎯 What it does: Proposes GrACE, a generative confidence mining method that directly outputs confidence during the generation phase, by adding a special <CNF> token to the vocabulary, allowing the model to generate this token at the end of the response, and using the similarity between its hidden state and the embedded <CNF> token to estimate confidence in real-time;
GRAD: Generalizing RAG Adaptation with Decoding
Youngwon Lee (Seoul National University), Yuxiong He (Snowflake AI Research)
RetrievalDomain AdaptationComputational EfficiencyKnowledge DistillationTransformerPrompt EngineeringContrastive LearningTextRetrieval-Augmented Generation
🎯 What it does: Propose GRAD, a framework that dynamically activates multiple RAG objectives (model expansion, domain adaptation, position debiasing) during decoding and converts them into token-level guidance for small models, avoiding the need to retrain large models for each task.
Gradient-Guided Multi-Judge Prompt Optimization
ChenZhuo Zhao, Dongmei Zhang (Microsoft)
OptimizationExplainability and InterpretabilityComputational EfficiencyData-Centric LearningTransformerLarge Language ModelReinforcement LearningPrompt EngineeringText
🎯 What it does: This paper proposes an efficient and robust automatic prompt optimization framework called GMPO, which uses gradient approximation to score the importance of prompt paragraphs and combines the loss integration of multiple discriminators to guide prompt rewriting.
GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMs
Duy Nguyen (University of North Carolina Chapel Hill), Mohit Bansal (University of North Carolina Chapel Hill)
Safty and PrivacyExplainability and InterpretabilityTransformerLarge Language ModelVision Language ModelContrastive LearningTextMultimodality
🎯 What it does: Propose a reasoning-time regulation method called GRAINS based on gradient attribution, which is used to perform targeted interventions on LLMs and VLMs without updating weights;
Grammar as Control: Modular Language Generation for the Long Tail
Ndapa Nakashole (University of California, San Diego)
GenerationData SynthesisTransformerLarge Language ModelPrompt EngineeringText
🎯 What it does: By decomposing descriptive grammar into modular segments, utilizing large language models (LLMs) to generate diverse and structurally balanced synthetic corpora based on control prompts, thus enhancing syntactic coverage for low-resource languages and applying it to machine translation;
Grammar Search for Multi-Agent Systems
Mayank Singh (University of Arizona), Eduardo Blanco (University of Arizona)
OptimizationComputational EfficiencyAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose an automatic search framework called Grammar Search based on context-free grammar, which constructs multi-agent systems (MAS) using composable components, and ensures the generated MAS structure is legal and executable without relying on LLM to generate code during the search phase.
Graph Reasoning Paradigm: Structured and Symbolic Reasoning with Topology-Aware Reinforcement Learning for Large Language Models
Runxuan Liu (Harbin Institute of Technology), Bing Qin (Harbin Institute of Technology)
Data SynthesisOptimizationExplainability and InterpretabilityReinforcement Learning from Human FeedbackGraph Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningTextGraphSequentialChain-of-Thought
🎯 What it does: Propose a Graph Structure Reasoning Paradigm (GRP), enabling large language models to output structured symbolic graphical reasoning steps, and design a Process-Aware Hierarchical Pruning Relative Strategy Optimization (PASC-GRPO) based on this paradigm for reinforcement learning;
Graph-Based Alternatives to LLMs for Human Simulation
Joseph Suh (University of California, Berkeley), Serina Chang (University of California, Berkeley)
Explainability and InterpretabilityComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringTextGraphChain-of-Thought
🎯 What it does: Proposed a graph-based model called GEMS to simulate human behavior in closed-choice tasks, replacing large language models.
GRAPHIA: Harnessing Social Graph Data to Enhance LLM-Based Social Simulation
Jiarui Ji (Renmin University of China), Bo Zheng (Alibaba)
GenerationData SynthesisExplainability and InterpretabilityReinforcement Learning from Human FeedbackGraph Neural NetworkTransformerLarge Language ModelSupervised Fine-TuningReinforcement LearningPrompt EngineeringContrastive LearningTextGraph
🎯 What it does: Propose the Graphia framework, which generates realistic social graphs and simulates human interactions by using social graph data as supervision during post-training of LLMs;
GraphSynth: Resolving the Diversity-Reliability Trade-off with Probabilistic Factor Graphs
Zehua Cheng (University of Oxford), Thomas Lukasiewicz (TU Wien)
Data SynthesisGraph Neural NetworkTransformerPrompt EngineeringDiffusion modelScore-based ModelContrastive LearningTextGraphBiomedical DataRetrieval-Augmented Generation
🎯 What it does: Propose the GraphSynth framework, which utilizes probabilistic factor graphs for multi-attribute subspace sampling, and combines hard constraint compilation with real-time semantic verification to generate both diverse and reliable synthetic data.
GRASPrune: Global Gating for Budgeted Structured Pruning of Large Language Models
Ziyang Wang (Beijing Institute of Technology), Jianbin Qin (Shenzhen University)
Computational EfficiencyKnowledge DistillationTransformerLarge Language ModelContrastive LearningText
🎯 What it does: This paper proposes a global budget structured pruning framework called GRASPrune, which jointly prunes FFN channels and KV head groups after pre-training while keeping the model weights unchanged;
GROKE: Vision-Free Navigation Instruction Evaluation via Graph Reasoning on OpenStreetMap
Farzad Shami (Aalto University), Henrikki Tenkanen (Aalto University)
Autonomous DrivingExplainability and InterpretabilityComputational EfficiencyRepresentation LearningGraph Neural NetworkTransformerLarge Language ModelTextGraphRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Propose a fully vision-free navigation instruction evaluation framework called GROKE, based on the OpenStreetMap graph, which objectively measures the navigability of instructions using hierarchical LLM reasoning and sub-instruction planning.
GroupToM-Bench: Benchmarking Group Theory of Mind and Nonlinear Social Emergence in MLLMs
Weidong Tang (Xidian University), Wangbo Zhao (National University of Singapore)
TransformerLarge Language ModelPrompt EngineeringVision Language ModelDiffusion modelImageTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Designed and released the GroupToM-Bench multimodal benchmark to evaluate the reasoning and prediction capabilities of large language models at the group level Theory of Mind (ToM).
GS-Quant: Granular Semantic and Generative Structural Quantization for Knowledge Graph Completion
Qizhuo Xie (Nanjing University), Tieke He (Nanjing University)
GenerationData SynthesisKnowledge DistillationRepresentation LearningGraph Neural NetworkTransformerLarge Language ModelPrompt EngineeringAuto EncoderContrastive LearningTextGraphBenchmark
🎯 What it does: Propose the GS-Quant framework, which generates semantically coherent and structured discrete codes through hierarchical quantization, to achieve the knowledge graph completion task;
GTA: Generating Long-horizon Tasks for Web Agents at Scale
Tenghao Huang (University of Southern California), Chien-Sheng Wu (Salesforce AI Research)
Autonomous DrivingReinforcement Learning from Human FeedbackTransformerLarge Language ModelAgentic AIPrompt EngineeringTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: Proposes the GTA framework to automatically generate multi-hop web tasks and generate executable trajectories, building a sustainable benchmark ecosystem.
Guaranteeing Knowledge Integration with Joint Decoding for Retrieval-Augmented Generation
Zhengyi Zhao (Chinese University of Hong Kong), Xian Wu (Tencent)
GenerationRetrievalExplainability and InterpretabilityComputational EfficiencyTransformerLarge Language ModelPrompt EngineeringContrastive LearningTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Proposes the GUARANTRAG framework, which first generates an internal reasoning answer (Inner-Answer) and then generates a reference answer (Refer-Answer) based on retrieved evidence, using joint decoding to achieve knowledge fusion;
GUI0: Self-Evolving Foundational GUI Agents in Super App Ecosystems
Xinyi Wang (WeChat, Tencent Inc.), Gong Zhi (WeChat, Tencent Inc.)
Autonomous DrivingComputational EfficiencyData-Centric LearningRobotic IntelligenceTransformerSupervised Fine-TuningReinforcement LearningAgentic AIVision Language ModelVision-Language-Action ModelImageTextMultimodalityBenchmarkChain-of-Thought
🎯 What it does: This paper proposes the GUI0 framework, aiming to enable visual language models to automatically interact with graphical user interfaces in non-standard rendered Super App environments.
GUIDE: Towards Scalable Advising for Research Ideas
Yaowenqi Liu (University of Illinois Urbana Champaign), Tong Zhang (University of Illinois Urbana Champaign)
Recommendation SystemExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationData-Centric LearningReinforcement Learning from Human FeedbackTransformerSupervised Fine-TuningPrompt EngineeringTextRetrieval-Augmented Generation
🎯 What it does: Proposes a system named GUIDE based on retrieval-augmented generation (RAG) and structured reasoning, designed for automated evaluation of research hypotheses and experimental designs.
Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence
Amirhosein Ghasemabadi (University of Alberta), Di Niu (University of Alberta)
Computational EfficiencyAI Code AssistantReinforcement Learning from Human FeedbackTransformerLarge Language ModelReinforcement LearningPrompt EngineeringTextBenchmarkChain-of-Thought
🎯 What it does: Proposes the Guided by Gut (GG) framework, which utilizes the intrinsic confidence and novelty of LLMs for self-guided tree search, achieving efficient scaling at test time.
GuideDog: A Real-World Egocentric Multimodal Dataset for Blind and Low-Vision Accessibility-Aware Guidance
Junhyeok Kim (Yonsei University), Youngjae Yu (Seoul National University)
Object DetectionDepth EstimationRecommendation SystemAnomaly DetectionAutonomous DrivingData-Centric LearningTransformerLarge Language ModelSupervised Fine-TuningPrompt EngineeringVision Language ModelImageTextMultimodalityBenchmarkRetrieval-Augmented GenerationChain-of-Thought
🎯 What it does: This paper proposes a scalable and assessable egocentric dataset called GUIDE DOG for blind/low vision (BLV) navigation assistance, and provides a multimodal guidance generation task and a fine-grained visual perception question answering benchmark based on BLV standards.
Guidelines as Environments: A World Model Approach to Rule Following
Haiqing Li (University of Texas at Arlington), Junzhou Huang (University of Texas at Arlington)
Autonomous DrivingOptimizationFederated LearningExplainability and InterpretabilityComputational EfficiencyKnowledge DistillationRepresentation LearningMeta LearningDrug DiscoveryAI Code AssistantReinforcement Learning from Human FeedbackNeural Architecture SearchProtein Structure PredictionTransformerLarge Language ModelReinforcement LearningPrompt EngineeringScore-based ModelWorld ModelTextGraphBenchmarkChain-of-Thought
🎯 What it does: Proposed a rule-based causal world model, RGCWM, to execute and plan complex interdependent rules without additional training.
HAG: Hierarchical Demographic Tree-based Agent Generation for Topic-Adaptive Simulation
Rongxin Chen (State Key Laboratory of AI Safety), Huawei Shen (State Key Laboratory of AI Safety)
GenerationData SynthesisTransformerLarge Language ModelWorld ModelTextBenchmarkRetrieval-Augmented Generation
🎯 What it does: Propose a hierarchical population distribution tree combined with retrieval and incremental generation, named the HAG framework, to achieve high-fidelity Agent generation that adapts to specific themes.