Tutorial Materials
Beyond Relevance: Utility-Centric Retrieval in the LLM Era — SIGIR 2026
Slides
Reading List
Tutorial Proposal: "Beyond Relevance: Utility-Centric Retrieval in the LLM Era"
Paper Repository: awesome-papers-on-Utility-focused-RAG (GitHub)
Section 1: Introduction and Foundations
1.1 RAG Foundations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. (Lewis et al. NeurIPS 2020.)
- REPLUG: Retrieval-Augmented Black-Box Language Models. (Shi et al. NAACL 2024.)
- Atlas: Few-shot Learning with Retrieval Augmented Language Models. (Izacard et al. JMLR 2023.)
1.2 Classical Relevance and Utility
- A Definition of Relevance for Information Retrieval. (Cooper. Information Storage and Retrieval 1971.)
- Relevance: A Review of and a Framework for the Thinking on the Notion in Information Science. (Saracevic. JASIS 1975.)
- Relevance Reconsidered. (Saracevic. CoLIS 1996.)
- A Study of Information Seeking and Retrieving. I. Background and Methodology. (Saracevic et al. JASIS 1988.)
1.3 User-Centric Utility in Web Search
- Click Data as Implicit Relevance Feedback in Web Search. (Jung et al. IP&M 2007.)
- Display Time as Implicit Feedback: Understanding Task Effects. (Kelly. SIGIR 2004.)
- Segment-Level Display Time as Implicit Feedback: A Comparison to Eye Tracking. (Buscher et al. SIGIR 2009.)
- Does Document Relevance Affect the Searcher's Perception of Time? (Luo et al. WSDM 2017.)
- More than Relevance: High Utility Query Recommendation by Mining Users' Search Behaviors. (Zhu et al. CIKM 2012.)
- Beyond Ranking: Optimizing Whole-Page Presentation. (Wang et al. WSDM 2016.)
1.5 Related Perspective: Denoising-First IR
- LLM-Oriented Information Retrieval: A Denoising-First Perspective. (Dai et al. SIGIR 2026.)
Section 2: What Is LLM-Centric Utility?
2.2 Task-Specific Definitions of Utility
Factoid QA
- Are Large Language Models Good at Utility Judgments? (Zhang et al. SIGIR 2024.)
- Evaluating Retrieval Quality in Retrieval-Augmented Generation. (Salemi et al. SIGIR 2024.)
Non-Factoid and Complex QA
- WebGLM: Towards an Efficient Web-Enhanced Question Answering System with Human Preferences. (Liu et al. KDD 2023.)
- DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research. (Shao et al. ICML 2026.)
Fact Verification and Classification
- From Relevance to Utility: Evidence Retrieval with Feedback for Fact Verification. (Zhang et al. EMNLP 2023.)
- Read It Twice: Towards Faithfully Interpretable Fact Verification by Revisiting Evidence. (Hu et al. SIGIR 2023.)
Code Generation
- Preference-Guided Refactored Tuning for Retrieval Augmented Code Generation. (Gao et al. ASE 2024.)
- SelfRACG: Enabling LLMs to Self-Express and Retrieve for Code Generation. (Dong et al. EMNLP 2025.)
Tool and Skill Retrieval
- Tool-to-Agent Retrieval: Bridging Tools and Agents for Scalable LLM Multi-Agent Systems. (Lumer et al. 2025.)
- Skill Retrieval Augmentation for Agentic AI. (Su et al. 2026.)
- Multi-Field Tool Retrieval. (Tang et al. 2026.)
- SkillResolve-Bench: Measuring and Resolving Same-Capability Ambiguity in Agent Skill Retrieval. (Ding et al. 2026.)
2.3 LLM-Agnostic vs. LLM-Specific Utility
LLM-Agnostic Utility
- Are Large Language Models Good at Utility Judgments? (Zhang et al. SIGIR 2024.)
- An Iterative Utility Judgment Framework via LLMs Inspired by Relevance in Philosophy. (Zhang et al. ACL 2026.)
LLM-Specific Utility
- LLM-Specific Utility: A New Perspective for Retrieval-Augmented Generation. (Zhang et al. 2025.)
- SePer: Measure Retrieval Utility Through the Lens of Semantic Perplexity Reduction. (Dai et al. ICLR 2025.)
- Uplift-RAG: Uplift-Driven Knowledge Preference Alignment for Retrieval-Augmented Generation. (Qu et al. EMNLP 2025.)
2.5 Pointwise and Set-Level Utility
Pointwise Utility
- REPLUG: Retrieval-Augmented Black-Box Language Models. (Shi et al. NAACL 2024.)
- Atlas: Few-shot Learning with Retrieval Augmented Language Models. (Izacard et al. JMLR 2023.)
- SePer: Measure Retrieval Utility Through the Lens of Semantic Perplexity Reduction. (Dai et al. ICLR 2025.)
- Uplift-RAG: Uplift-Driven Knowledge Preference Alignment for Retrieval-Augmented Generation. (Qu et al. EMNLP 2025.)
Set-Level Utility
- DIVERGE: Diversity-Enhanced Retrieval-Augmented Generation for Open-Ended Information Seeking. (Hu et al. 2026.)
- DF-RAG: Query-Aware Diversity for Retrieval-Augmented Generation. (Khan et al. ACL 2026.)
- Modeling Contextual Passage Utility for Multihop Question Answering. (Jain et al. IJCNLP 2025.)
- Distilling a Small Utility-Based Passage Selector to Enhance Retrieval-Augmented Generation. (Zhang et al. SIGIR-AP 2025.)
- Shifting from Ranking to Set Selection for Retrieval Augmented Generation. (Lee et al. ACL 2025.)
- OptiSet: Unified Optimizing Set Selection and Ranking for Retrieval-Augmented Generation. (Jiang et al. 2026.)
- BGM: Bridging the Preference Gap between Retrievers and LLMs. (Ke et al. ACL 2024.)
2.6 Evidence Format and Presentation to LLMs
- Relevance to Utility: Process-Supervised Rewrite for RAG. (Kim et al. ACL 2026.)
- ECoRAG: Evidentiality-Guided Compression for Long Context RAG. (Jeong et al. ACL 2025.)
- Parametric Retrieval-Augmented Generation. (Su et al. SIGIR 2025.)
- Chain-of-Note: Enhancing Robustness in Retrieval-Augmented Language Models. (Yu et al. EMNLP 2024.)
- Injecting External Knowledge into the Reasoning Process Enhances Retrieval-Augmented Generation. (Tang et al. SIGIR-AP 2025.)
- The Last Human-Written Paper: Agent-Native Research Artifacts. (Liu et al. 2026.)
Section 3: Utility Modeling and Optimization Methods
3.1 Utility Signals
Signal A: Verbalized LLM Utility Judgment (no ground-truth answer required)
- Are Large Language Models Good at Utility Judgments? (Zhang et al. SIGIR 2024.)
- An Iterative Utility Judgment Framework via LLMs Inspired by Relevance in Philosophy. (Zhang et al. ACL 2026.)
- Bridging Relevance and Reasoning: Rationale Distillation in Retrieval-Augmented Generation. (Jia et al. ACL 2025.)
Signal B: Attention / Reader-Derived Signals (no ground-truth answer required)
- FiD: Distilling Knowledge from Reader to Retriever for Question Answering. (Izacard et al. ICLR 2021.)
- Atlas: Few-shot Learning with Retrieval Augmented Language Models. (Izacard et al. JMLR 2023.)
- LLM-Specific Utility: A New Perspective for Retrieval-Augmented Generation. (Zhang et al. 2025.)
Signal C: Likelihood and Perplexity (ground-truth answer required)
- REPLUG: Retrieval-Augmented Black-Box Language Models. (Shi et al. NAACL 2024.)
- Aligning Dense Retrievers with LLM Utility via Distillation. (Sandhu et al. 2026.)
- Atlas: Few-shot Learning with Retrieval Augmented Language Models. (Izacard et al. JMLR 2023.)
- Predicting Retrieval Utility and Answer Quality in Retrieval-Augmented Generation. (Tian et al. ECIR 2026.)
- SEER: Self-Aligned Evidence Extraction for Retrieval-Augmented Generation. (Zhao et al. EMNLP 2024.)
- Utility-Oriented Visual Evidence Selection for Multimodal Retrieval-Augmented Generation. (Luo et al. ACL 2026.)
- Training a Utility-based Retriever Through Shared Context Attribution for Retrieval-Augmented Language Models. (Xu et al. EMNLP 2025.)
Signal D: Downstream Performance (ground-truth answer required)
3.2 Efficiency Ladder: From Small Candidate Sets to Corpus-Scale
Level 1: Small Candidate Sets (10–20) — High-Fidelity Judgments
- Are Large Language Models Good at Utility Judgments? (Zhang et al. SIGIR 2024.)
- An Iterative Utility Judgment Framework via LLMs Inspired by Relevance in Philosophy. (Zhang et al. ACL 2026.)
- LLM-Specific Utility: A New Perspective for Retrieval-Augmented Generation. (Zhang et al. 2025.)
Level 2: Medium Candidate Sets (dozens to hundreds) — Selectors and Rerankers
- Distilling a Small Utility-Based Passage Selector to Enhance Retrieval-Augmented Generation. (Zhang et al. SIGIR-AP 2025.)
- Predicting Retrieval Utility and Answer Quality in Retrieval-Augmented Generation. (Tian et al. ECIR 2026.)
- From Relevance to Utility: Evidence Retrieval with Feedback for Fact Verification. (Zhang et al. EMNLP 2023.)
- Read It Twice: Towards Faithfully Interpretable Fact Verification by Revisiting Evidence. (Hu et al. SIGIR 2023.)
- Modeling Contextual Passage Utility for Multihop Question Answering. (Jain et al. IJCNLP 2025.)
- Bridging Relevance and Reasoning: Rationale Distillation in Retrieval-Augmented Generation. (Jia et al. ACL 2025.)
- RAG-DDR: Optimizing Retrieval-Augmented Generation Using Differentiable Data Rewards. (Li et al. ICLR 2025.)
- Retrieve What You Need: A Mutual Learning Framework for Open-Domain Question Answering. (Wang et al. TACL 2024.)
- Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning. (Wu et al. ACL 2026.)
Level 3: Corpus-Scale — Utility-Oriented Retrievers
- Utility-Focused LLM Annotation for Retrieval and Retrieval-Augmented Generation. (Zhang et al. EMNLP 2025.)
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. (Lewis et al. NeurIPS 2020.)
- Stochastic RAG: End-to-End Retrieval-Augmented Generation through Expected Utility Maximization. (Zamani et al. SIGIR 2024.)
- REPLUG: Retrieval-Augmented Black-Box Language Models. (Shi et al. NAACL 2024.)
- Aligning Dense Retrievers with LLM Utility via Distillation. (Sandhu et al. 2026.)
- GripRank: Bridging the Gap between Retrieval and Generation via Generative Knowledge Improved Passage Ranking. (Bai et al. CIKM 2023.)
- FiD: Distilling Knowledge from Reader to Retriever for Question Answering. (Izacard et al. ICLR 2021.)
- Atlas: Few-shot Learning with Retrieval Augmented Language Models. (Izacard et al. JMLR 2023.)
- Training a Utility-based Retriever Through Shared Context Attribution for Retrieval-Augmented Language Models. (Xu et al. EMNLP 2025.)
- RRAML: Reinforced Retrieval Augmented Machine Learning. (Bacciu et al. 2023.)
3.3 Evaluation: What Should We Measure?
- eRAG: Evaluating Retrieval Quality in Retrieval-Augmented Generation. (Salemi et al. SIGIR 2024.)
- RARE: Redundancy-Aware Retrieval Evaluation Framework for High-Similarity Corpora. (Cho et al. ACL 2026.)
- UsefulBench: Towards Decision-Useful Information as a Target for Information Retrieval. (Schimanski et al. 2026.)
Section 4: Utility in Agentic RAG
4.2 Query-side Utility
- Active Retrieval Augmented Generation. (Jiang et al. EMNLP 2023.)
- DRAGIN: Dynamic Retrieval Augmented Generation Based on the Real-Time Information Needs of Large Language Models. (Su et al. ACL 2024.)
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning. (Jin et al. COLM 2025.)
- AgentIR: Reasoning-Aware Retrieval for Deep Research Agents. (Chen et al. 2026.)
- SubSearch: Intermediate Rewards for Unsupervised Guided Reasoning in Complex Retrieval. (Petcu et al. 2026.)
4.3 Document-side Utility
- Revisiting Text Ranking in Deep Research. (Meng et al. SIGIR 2026.)
- Agentic-R: Learning to Retrieve for Agentic Search. (Liu et al. ACL 2026.)
- Learning to Retrieve from Agent Trajectories. (Zhou et al. SIGIR 2026.)
- Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems. (Zhao et al. ACL 2026.)
4.4 Action-side Utility
- StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization. (Zheng et al. EMNLP 2025.)
- Information Gain-Based Policy Optimization: A Simple and Effective Approach for Multi-Turn Search Agents. (Wang et al. ICLR 2026.)
- Adaptive-RAG: Learning to Adapt Retrieval-Augmented Large Language Models Through Question Complexity. (Jeong et al. NAACL 2024.)
- Beyond Monolithic Architectures: A Multi-Agent Search and Knowledge Optimization Framework for Agentic Search. (Chen et al. 2026.)
- Beyond Semantic Similarity: Rethinking Retrieval for Agentic Search via Direct Corpus Interaction. (Li et al. 2026.)
Citation
@inproceedings{zhang2026beyond,
title={Beyond Relevance: Utility-Centric Retrieval in the LLM Era},
author={Zhang, Hengran and Tang, Minghao and Bi, Keping and Guo, Jiafeng},
booktitle={Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval},
pages={5357--5361},
year={2026}
}