WebGPT: Browser-assisted question-answering with human feedback
Reiichiro NakanoJacob HiltonSuchir BalajiJeff WuLong OuyangChristina KimChristopher HesseShantanu JainVineet KosarajuWilliam Saunders
Demonstrates that fine-tuning language models to search the web and cite supporting sources using human feedback produces long-form answers preferred by human evaluators over human-written responses.
Large language models often struggle to provide accurate, comprehensive answers to complex questions, frequently generating plausible-sounding falsehoods or outdated information. This article addresses these challenges by evaluating how equipping a language model with web-browsing capabilities and optimizing its outputs using human feedback can enhance factual accuracy and overall response quality in long-form question answering.
The researchers developed a text-based web environment that allows the GPT-3 model to search the web using a search engine, navigate links, and collect cited quotes before drafting answers. To train the system, known as WebGPT, they collected thousands of human browsing demonstrations to teach the model how to use the interface, along with pairwise human preference comparisons across model-generated outputs. The primary training pipeline involved supervised fine-tuning on human demonstrations, training a reward model on human preferences, and applying rejection sampling—selecting the highest-scoring response from multiple sampled candidates.
The findings show that WebGPT significantly improves answer quality over standard generative models and competitive baselines. On the open-ended ELI5 benchmark, human evaluators preferred answers from the top 175-billion-parameter WebGPT model 56% of the time over those written by human demonstrators and 69% of the time over the highest-voted answers on Reddit. On TruthfulQA, a benchmark designed to trigger human misconceptions, the model answered truthfully and informatively 54% of the time, substantially outperforming standard GPT-3. Furthermore, rejection sampling proved more effective and compute-efficient than reinforcement learning, especially as inference-time compute was scaled.
These results demonstrate that enabling models to retrieve and cite supporting evidence substantially improves transparency and fact-checking efficiency, mitigating common factual hallucinations. However, the article highlights important risks, including user overreliance due to automation bias and the authoritative appearance of cited answers. The system also remains vulnerable to accepting false premises in user queries, perpetuating cultural reference-point biases, and quoting from unreliable sources when handling unfamiliar or adversarial questions.
Organizations considering similar retrieval-augmented systems should implement robust governance, including clear documentation of system boundaries and rigorous source-verification protocols. Further research and development are recommended to improve debiasing methods, explore multi-agent debate to prevent models from cherry-picking evidence, and establish stricter safeguards before deploying autonomous web-enabled AI systems.
- Paper: Language Models are Few-Shot Learners, T. B. Brown et al. (2020). Provides the foundational GPT-3 architecture and few-shot capabilities that WebGPT directly fine-tunes to interact with a web-browsing environment.
- Paper: Learning to summarize from human feedback, Nisan Stiennon et al. (2020). Establishes the core reinforcement learning from human feedback (RLHF) framework and reward modeling methodology adapted by WebGPT to optimize answer quality.
- Paper: Fine-Tuning Language Models from Human Preferences, Daniel M. Ziegler et al. (2019). Introduces the foundational pipeline of fine-tuning pretrained language models on human preference data using reward models and reinforcement learning.
- Paper: Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering, Gautier Izacard et al. (2021). Pioneers the open-domain question answering paradigm of conditioning sequence-to-sequence generative models on retrieved textual evidence.
- Paper: Reading Wikipedia to Answer Open-Domain Questions, Danqi Chen et al. (2017). Introduces the foundational two-step retrieve-and-read paradigm for open-domain question answering that web-based language models build upon.
- Paper: Toolformer: Language Models Can Teach Themselves to Use Tools, Timo Schick et al. (2023). Extends the concept of grounding language models with external sources by enabling them to autonomously teach themselves how and when to use tools such as search engines.
- Paper: Active Retrieval Augmented Generation, Zhengbao Jiang et al. (2023). Builds upon interactive retrieval for long-form generation by actively querying external text only when the model predicts low confidence in its generated tokens.
- Paper: Training language models to follow instructions with human feedback, Long Ouyang et al. (2022). Scales the human feedback alignment and reward modeling techniques used in WebGPT to general-purpose instruction following.
- Paper: Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback, Yuntao Bai et al. (2022). Develops preference modeling and reinforcement learning from human feedback further to simultaneously target helpfulness and harmlessness in interactive assistants.
- Paper: TruthfulQA: Measuring How Models Mimic Human Falsehoods, Stephanie C. Lin et al. (2022). Provides a targeted benchmark to evaluate whether aligned or retrieval-augmented models successfully avoid mimicking human misconceptions and falsehoods.
- Paper: REPLUG: Retrieval-Augmented Black-Box Language Models, Weijia Shi et al. (2024). Applies retrieval augmentation to black-box large language models without requiring fine-tuning of the base generator.
- Paper: GPQA: A Graduate-Level Google-Proof Q&A Benchmark, David Rein et al. (2023). Introduces a challenging benchmark designed to test expert-level reasoning and question answering that resists simple web retrieval lookups.
