Deep Research is Still All About Retrieval

Abstract: Deep research applications that involve retrieving information and synthesizing comprehensive responses have become increasingly popular. While improvements in LLMs and AI agents have made deep research systems possible, the retrieval component in such systems is still a critical piece that has received less attention. Instead of retrieving relevant but repetitive information from different sources, finding unique pieces of relevant information is a key factor in improving the utility of the final responses. In this talk, I introduce a reproducible evaluation method for information coverage, discuss how information coverage in responses is bottlenecked by upstream retrieval, and present several optimization directions for coverage.
Bio: Eugene Yang is a senior research scientist at the Human Language Technology Center of Excellence (HLTCOE) at Johns Hopkins University, where he works on multilingual retrieval and retrieval for deep research. His recent work centers on evaluation and optimization methods for RAG applications and their upstream retrieval engines. He co-organizes the TREC RAGTIME track, which builds shared evaluation for RAG systems, and also its predecessor, TREC NeuCLIR, which focused on cross-language and multilingual retrieval.