WebGPT: Browser-assisted question-answering with human feedback

Reiichiro NakanoJacob HiltonSuchir BalajiJeff WuLong OuyangChristina KimChristopher HesseShantanu JainVineet KosarajuWilliam Saunders

article2021arXiv2,066 citations

Demonstrates that fine-tuning language models to search the web and cite supporting sources using human feedback produces long-form answers preferred by human evaluators over human-written responses.

Listen

Large language models often struggle to provide accurate, comprehensive answers to complex questions, frequently generating plausible-sounding falsehoods or outdated information. This article addresses these challenges by evaluating how equipping a language model with web-browsing capabilities and optimizing its outputs using human feedback can enhance factual accuracy and overall response quality in long-form question answering.

The researchers developed a text-based web environment that allows the GPT-3 model to search the web using a search engine, navigate links, and collect cited quotes before drafting answers. To train the system, known as WebGPT, they collected thousands of human browsing demonstrations to teach the model how to use the interface, along with pairwise human preference comparisons across model-generated outputs. The primary training pipeline involved supervised fine-tuning on human demonstrations, training a reward model on human preferences, and applying rejection sampling—selecting the highest-scoring response from multiple sampled candidates.

The findings show that WebGPT significantly improves answer quality over standard generative models and competitive baselines. On the open-ended ELI5 benchmark, human evaluators preferred answers from the top 175-billion-parameter WebGPT model 56% of the time over those written by human demonstrators and 69% of the time over the highest-voted answers on Reddit. On TruthfulQA, a benchmark designed to trigger human misconceptions, the model answered truthfully and informatively 54% of the time, substantially outperforming standard GPT-3. Furthermore, rejection sampling proved more effective and compute-efficient than reinforcement learning, especially as inference-time compute was scaled.

These results demonstrate that enabling models to retrieve and cite supporting evidence substantially improves transparency and fact-checking efficiency, mitigating common factual hallucinations. However, the article highlights important risks, including user overreliance due to automation bias and the authoritative appearance of cited answers. The system also remains vulnerable to accepting false premises in user queries, perpetuating cultural reference-point biases, and quoting from unreliable sources when handling unfamiliar or adversarial questions.

Organizations considering similar retrieval-augmented systems should implement robust governance, including clear documentation of system boundaries and rigorous source-verification protocols. Further research and development are recommended to improve debiasing methods, explore multi-agent debate to prevent models from cherry-picking evidence, and establish stricter safeguards before deploying autonomous web-enabled AI systems.

arXiv: 2112.09332
Cover for WebGPT: Browser-assisted question-answering with human feedback

Abstract

We fine-tune GPT-3 to answer long-form questions using a text-based web-browsing environment, which allows the model to search and navigate the web. By setting up the task so that it can be performed by humans, we are able to train models on the task using imitation learning, and then optimize answer quality with human feedback. To make human evaluation of factual accuracy easier, models must collect references while browsing in support of their answers. We train and evaluate our models on ELI5, a dataset of questions asked by Reddit users. Our best model is obtained by fine-tuning GPT-3 using behavior cloning, and then performing rejection sampling against a reward model trained to predict human preferences. This model's answers are preferred by humans 56% of the time to those of our human demonstrators, and 69% of the time to the highest-voted answer from Reddit.

Table of Contents

  • 1 Introduction
  • 2 Environment design
  • 3 Methods
  • 3.1 Data collection
  • 3.2 Training
  • 4 Evaluation
  • 4.1 ELI5
  • 4.2 TruthfulQA
  • 4.3 TriviaQA
  • 5 Experiments
  • 5.1 Comparison of training methods
  • 5.2 Scaling experiments
  • 6 Discussion
  • 6.1 Truthfulness of WebGPT
  • 6.2 Perceived truthfulness of WebGPT
  • 6.3 Reinforcement of bias
  • 6.4 Using references to evaluate factual accuracy
  • 6.5 Risks of live web access
  • 7 Related work
  • 8 Conclusion
  • 9 Author contributions
  • 10 Acknowledgments
  • References
  • A Environment design details
  • B Question dataset details
  • C Data collection details
  • C.1 Demonstrations
  • C.2 Comparisons
  • D Contractor survey
  • E Hyperparameters
  • F Minimal comparison instructions
  • G TriviaQA evaluation
  • H Analysis of effect of question stance and reference point bias
  • H.1 Effect of question stance on factual accuracy and answer stance
  • H.2 Reference point bias
  • I Predicting rejection sampling performance
  • J References for example answer and alternative answers
  • K Comparison dataset release details

Knowls

  1. Knowl 1 — WebGPT Training and Optimization Pipeline

    model/method

    The WebGPT framework fine-tunes large pre-trained autoregressive language models (GPT-3 760M, 13B, and 175B) to perform long-form question answering by interacting with a text-based web browser using a four-stage optimization pipeline:

    1. Behavior Cloning (BC): The pre-trained model is fine-tuned via supervised learning on human demonstration trajectories consisting of sequential browser commands and final reference-grounded answer generation.

    2. Reward Modeling (RM): The unembedding layer of the BC model is replaced with a scalar reward head. The reward model takes a prompt (q,a)(q, a)—consisting of the question qq concatenated with extracted references and the synthesized answer aa—and outputs a scalar score r(q,a)r(q, a) representing an Elo rating. Given a pair of candidate answers (a1,a2)(a_1, a_2) where human evaluators indicate preference y∈{0,0.5,1}y \in \{0, 0.5, 1\}, the reward model is trained using a binary cross-entropy loss: LRM=−E(q,a1,a2,y)[ylog⁡σ(r(q,a1)−r(q,a2))+(1−y)log⁡(1−σ(r(q,a1)−r(q,a2)))]\mathcal{L}_{\text{RM}} = - \mathbb{E}_{(q, a_1, a_2, y)} \left[ y \log \sigma(r(q, a_1) - r(q, a_2)) + (1 - y) \log (1 - \sigma(r(q, a_1) - r(q, a_2))) \right] where σ(x)=1/(1+e−x)\sigma(x) = 1 / (1 + e^{-x}) is the sigmoid function, and ties are assigned soft target labels of 0.50.5.

    3. Reinforcement Learning (RL): The BC policy is fine-tuned in the interactive browsing environment using Proximal Policy Optimization (PPO). The reward assigned at the completion of an episode is the scalar output from the reward model, augmented with a per-token Kullback-Leibler (KL) divergence penalty against the baseline BC policy πBC\pi_{\text{BC}} to prevent reward model overoptimization: Rt={r(q,a)−β∑τDKL(πθ(⋅∣sτ)∥πBC(⋅∣sτ))if t=T−βDKL(πθ(⋅∣st)∥πBC(⋅∣st))if t<TR_t = \begin{cases} r(q, a) - \beta \sum_{\tau} D_{\text{KL}}(\pi_\theta(\cdot \mid s_\tau) \parallel \pi_{\text{BC}}(\cdot \mid s_\tau)) & \text{if } t = T \\ - \beta D_{\text{KL}}(\pi_\theta(\cdot \mid s_t) \parallel \pi_{\text{BC}}(\cdot \mid s_t)) & \text{if } t < T \end{cases} where β\beta is the KL reward coefficient.

    4. Best-of-nn Rejection Sampling: At inference time, nn full candidate trajectories (browsing and final answer synthesis) are sampled independently from the policy (BC or RL), scored using the reward model, and the highest-scoring candidate is selected as the output.

  2. Knowl 2 — Text-Based Web-Browsing Environment and Grounded Reference Extraction

    model/method

    The text-based web browsing environment allows language models to search the live web and compile cited answers. At each discrete step, the environment provides the model with a structured text observation containing:

    1. The user question.
    2. Currently collected quotes/references.
    3. A summary of past browser actions in the current episode.
    4. The page title and domain name of the current URL.
    5. The visible slice of text corresponding to the current vertical scroll position.
    6. The remaining action budget.

    The action space consists of 10 structured commands:

    • Search <query>: Queries the Microsoft Bing Web Search API and returns a simplified list of search results.
    • Clicked on link <link ID>: Navigates to a specific hyperlink identified by an index tag embedded in the simplified page text.
    • Find in page: <text>: Case-insensitively searches for the specified string on the current web page and moves the viewport cursor to its next occurrence.
    • Quote: <text>: Searches for the exact text string in the active page; if found, records the string along with the page title and domain as a persistent numbered reference.
    • Scrolled down <1, 2, 3> / Scrolled up <1, 2, 3>: Moves the viewport down or up by 1 to 3 viewport increments.
    • Top: Scrolls to the top of the current web page.
    • Back: Navigates to the previously visited web page in history.
    • End: Answer: Concludes browsing and transitions the model to the final answering phase.
    • End: <Nonsense, Controversial>: Concludes the episode immediately without generating an answer.

    Pages are fetched via Node.js, cleaned with Mozilla Readability, and converted to plain text with stripped formatting using html2text (or pdfminer.six for PDF documents). To prevent direct test contamination, pages containing a 10-gram overlap with the question or reference dataset answer are censored, and domains such as reddit.com and quora.com are blocked.

  3. Knowl 3 — Unbiased Estimator for Best-of-n Rejection Sampling Expected Reward

    equation

    To evaluate the performance of best-of-nn rejection sampling without running full Monte Carlo rescoring across all combinations of subsets, an unbiased, closed-form combinatorial estimator is used.

    Let q∼Qq \sim \mathcal{Q} be a question, and let A(q)\mathcal{A}(q) be the distribution over answers produced by policy π\pi. Let Rtrain(a∣q)R^{\text{train}}(a \mid q) be the reward model used for ranking, and let Rval(a∣q)R^{\text{val}}(a \mid q) be an independent validation reward model. The true expected validation score of the top candidate selected from nn samples is: Rnpred(q)=EA1,…,An∼A(q)[Rval(arg⁡max⁡a∈{A1,…,An}Rtrain(a∣q)  |  q)]R_n^{\text{pred}}(q) = \mathbb{E}_{A_1, \dots, A_n \sim \mathcal{A}(q)} \left[ R^{\text{val}}\left( \arg\max_{a \in \{A_1, \dots, A_n\}} R^{\text{train}}(a \mid q) \;\middle|\; q \right) \right]

    Given a pool of N≥nN \ge n independently sampled answers A1,…,AN∼A(q)A_1, \dots, A_N \sim \mathcal{A}(q), let (S1,S2,…,SN)(S_1, S_2, \dots, S_N) represent the permutation of these answers sorted in ascending order of their training reward scores, such that: Rtrain(S1∣q)≤Rtrain(S2∣q)≤⋯≤Rtrain(SN∣q)R^{\text{train}}(S_1 \mid q) \le R^{\text{train}}(S_2 \mid q) \le \dots \le R^{\text{train}}(S_N \mid q)

    The expectation over all (Nn)\binom{N}{n} possible subsets of size nn simplifies by linearity of expectation to: Rnpred(q)=1(Nn)∑1≤i1<⋯<in≤NRval(arg⁡max⁡a∈{Si1,…,Sin}Rtrain(a∣q)  |  q)=∑i=nN(i−1n−1)(Nn)Rval(Si∣q)R_n^{\text{pred}}(q) = \frac{1}{\binom{N}{n}} \sum_{1 \le i_1 < \dots < i_n \le N} R^{\text{val}}\left(\arg\max_{a \in \{S_{i_1}, \dots, S_{i_n}\}} R^{\text{train}}(a \mid q) \;\middle|\; q\right) = \sum_{i=n}^N \frac{\binom{i-1}{n-1}}{\binom{N}{n}} R^{\text{val}}(S_i \mid q)

    where:

    • N∈NN \in \mathbb{N} is the total number of candidate completions sampled (N≥nN \ge n).
    • n∈Nn \in \mathbb{N} is the rejection sampling candidate pool size.
    • SiS_i is the candidate ranked at position ii in ascending training reward order (i∈{1,…,N}i \in \{1, \dots, N\}).
    • (i−1n−1)(Nn)\frac{\binom{i-1}{n-1}}{\binom{N}{n}} represents the probability that SiS_i is the maximum-reward candidate in a uniformly drawn sub-sample of size nn.
  4. Knowl 4 — Human Preference Evaluation on ELI5 Long-Form Question Answering

    empirical result

    When evaluated on the Explain Like I'm Five (ELI5) test dataset using pairwise blind human preference ratings (where ties count as 50%50\% preference), WebGPT models demonstrate competitive and super-human performance:

    1. Comparison against Human Demonstrators:

      • The 175B parameter WebGPT model using Best-of-64 rejection sampling is preferred by human labelers 56%56\% of the time over answers authored by human demonstrators using the same browsing environment.
      • Along individual sub-dimensions, the 175B best-of-64 model achieves 56%56\% preference on overall usefulness, 44%44\% preference on coherence, and 51%51\% preference on factual accuracy against human demonstrators.
      • The 760M best-of-4 model is preferred 28%28\% of the time, and the 13B best-of-16 model is preferred 45%45\% of the time on overall usefulness against human demonstrators.
    2. Comparison against ELI5 Reddit Reference Answers:

      • When model citations and references are stripped to match the unreferenced Reddit format, the 175B best-of-64 model is preferred 69%69\% of the time over the highest-voted reference answer from Reddit.
      • The 760M best-of-4 model is preferred 42%42\% of the time, and the 13B best-of-16 model is preferred 63%63\% of the time over Reddit reference answers.
  5. Knowl 5 — Evaluation and Truthfulness Scaling on TruthfulQA

    empirical result

    On TruthfulQA—an adversarial benchmark composed of short-form questions targeting common human misconceptions and falsehoods—WebGPT models outperform zero-shot and few-shot base GPT-3 models on both absolute truthfulness and joint truthfulness-informativeness:

    • Truthfulness Rate (% truthful):

      • GPT-3 175B (QA prompt): 28%28\%
      • GPT-3 175B (Helpful prompt): 65%65\%
      • WebGPT 760M best-of-4: 73%73\%
      • WebGPT 13B best-of-16: 80%80\%
      • WebGPT 175B best-of-64: 75%75\%
      • Human baseline: 94%94\%
    • Informative Truthfulness Rate (% truthful and informative):

      • GPT-3 175B (QA prompt): 25%25\%
      • GPT-3 175B (Helpful prompt): 22%22\%
      • WebGPT 760M best-of-4: 33%33\%
      • WebGPT 13B best-of-16: 48%48\%
      • WebGPT 175B best-of-64: 54%54\%
      • Human baseline: 87%87\%

    While base GPT-3 with a helpful prompt achieves moderate truthfulness primarily by evading questions (answering "I have no comment" to 49%49\% of prompts), WebGPT actively synthesizes factual content. The joint percentage of truthful and informative answers scales monotonically with model parameter scale and search budget in WebGPT.

  6. Knowl 6 — Comparison of Rejection Sampling and Reinforcement Learning

    empirical result

    Comparing policy optimization techniques against the Behavior Cloning (BC) baseline across model scales shows that rejection sampling provides substantially larger preference gains than PPO reinforcement learning:

    • The 175B Best-of-64 BC model is preferred 68%68\% of the time over the 175B BC baseline.
    • The 175B RL policy (without rejection sampling) is preferred 58%58\% of the time over the 175B BC baseline.
    • Combining RL with rejection sampling (e.g., 175B RL best-of-64) yields a 47%47\% preference rate when judged directly against the 175B BC best-of-64 policy, providing no significant performance advantage over rejection sampling applied directly to the BC baseline.

    Rejection sampling outperforms RL in this domain because:

    1. It expends additional inference-time compute across diverse browsing paths.
    2. The non-deterministic web environment benefits from exploring multiple trajectories and selecting the best outcome with the benefit of hindsight.
    3. The reward model is trained primarily on BC and rejection-sampled data, making it less susceptible to out-of-distribution reward hacking under rejection sampling than under iterative gradient ascent in RL.
    4. RL decreases policy entropy, which impairs exploration across the search action space.
  7. Knowl 7 — Compute-Efficient Inference Frontier and Scaling Trends

    empirical result

    Model performance—measured by Elo score predictions against a held-out 175B validation reward model (where a 1.0 Elo difference denotes a ≈73%\approx 73\% preference probability)—exhibits distinct scaling relationships with respect to demonstrations, comparisons, model parameters, and test-time search:

    • Demonstration Scaling: Doubling the number of BC demonstrations increases the policy reward model score by approximately +0.13+0.13 Elo points.
    • Comparison Scaling: Doubling the number of human comparison pairs increases reward model classification accuracy by approximately +1.8%+1.8\%.
    • Parameter Scaling: Doubling policy parameter count increases the reward model score by approximately +0.09+0.09 Elo points; doubling reward model parameter count increases its accuracy by approximately +0.4%+0.4\%.
    • Inference Compute Pareto Frontier: For fixed floating-point operation budgets at inference time, allocating FLOPs to both parameter size and rejection sampling sample count nn yields superior performance compared to allocating compute exclusively to larger base models. The compute-efficient frontier is spanned by:
      • ≈1014\approx 10^{14} FLOPs: 760M Best-of-4
      • ≈1015\approx 10^{15} FLOPs: 13B Best-of-16
      • ≈1017\approx 10^{17} FLOPs: 175B Best-of-64
  8. Knowl 8 — Short-Form QA Transfer on TriviaQA via Output Conditioning

    empirical result

    Although trained for long-form question answering on ELI5, the WebGPT browsing policy transfers to open-domain short-form question answering on TriviaQA without retraining the browsing environment.

    A GPT-3 175B extractor model is fine-tuned on 256 TriviaQA examples to extract short answer strings conditioned on the full WebGPT 175B BC generation. Evaluated on the TriviaQA development splits (measuring Exact Match percentages):

    Model Total Question overlap No question overlap Answer overlap Answer overlap only No overlap
    GPT-3 175B 58.7% 75.9% 52.9% 67.3% 61.6% 39.0%
    GPT-3 175B + WebGPT 175B BC 69.5% 86.3% 65.3% 78.4% 73.2% 52.4%
    UnitedQA-E 68.9% 89.3% 62.7% 78.6% 70.6% 44.3%
    UnitedQA (hybrid model) 70.5% – – – – –

    Conditioning GPT-3 on WebGPT browsing outputs yields an overall exact match of 69.5%69.5\%, and achieves 52.4%52.4\% on the zero-overlap split (exceeding UnitedQA-E's 44.3%44.3\%).

  9. Knowl 9 — Question Stance Sensitivity and Cultural Reference Point Bias

    limitation

    WebGPT exhibits systematic epistemic and cultural biases when answering open-ended queries:

    1. Question Stance Sensitivity: When presented with questions regarding conspiracy theories or misconceptions framed with an affirming stance (e.g., "Why did the government fake the moon landing?"), WebGPT generates factually inaccurate answers significantly more frequently than when prompted with a neutral ("When did the moon landing happen?") or skeptical ("Could the moon landing really be fake?") stance across all model sizes (760M, 13B, and 175B). While larger models refute false premises more often overall, affirming question framing exacerbates confirmation bias.

    2. Cultural Reference Point Bias: In response to open-ended, under-specified cultural questions (such as "What does a wedding look like?"), WebGPT exhibits a strong default towards Western and American cultural norms. In an evaluation of 64 independent samples from the 175B BC model on this prompt, 20 explicitly included the word "America" or "American", and 63 of 64 described features specific to Western/American weddings (e.g., white dresses, father walking the bride down the aisle), mentioning non-Western cultural traditions in only 4 instances.

  10. Knowl 10 — Answering-Phase Replay for Sample-Efficient Reinforcement Learning

    model/method

    During PPO reinforcement learning of the web-browsing agent, the multi-step web browsing phase comprises the vast majority of episode interaction steps, whereas the final answer composition step accounts for the majority of the variance in reward model score.

    To improve sample efficiency during RL training:

    1. At the termination of each full browsing episode where references were collected, 15 additional answering-only synthetic episodes are inserted into the training replay stream.
    2. Each synthetic episode holds the environment state and collected web references identical to the original episode but samples a fresh completion for the final answer token generation.
    3. The policy parameters update across both the original browsing trajectory and the 15 replayed answering phases.

    This multi-answer replay procedure improves overall PPO sample efficiency by approximately a factor of 2 without requiring additional web navigation queries.

Coverage note — None was omitted; all primary methodological contributions, empirical evaluations (ELI5, TruthfulQA, TriviaQA), scaling dynamics, mathematical estimators, and bias analyses are covered.

References

  1. 1.L. Adolphs, B. Boerschinger, C. Buck, M. C. Huebscher, M. Ciaramita, L. Espeholt, T. Hofmann, and Y. Kilcher. Boosting search engines with interactive agents. arXiv preprint arXiv:2109.00527, 2021.
  2. 2.S. Bhakthavatsalam, D. Khashabi, T. Khot, B. D. Mishra, K. Richardson, A. Sabharwal, C. Schoenick, O. Tafjord, and P. Clark. Think you have solved direct-answer question answering? Try ARC-DA, the direct-answer AI2 reasoning challenge. arXiv preprint arXiv:2102.03315, 2021.
  3. 3.N. Bostrom. Superintelligence: Paths, Dangers, Strategies. Oxford University Press, 2014.
  4. 4.T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020.
  5. 5.H. Cheng, Y. Shen, X. Liu, P. He, W. Chen, and J. Gao. UnitedQA: A hybrid approach for open domain question answering. arXiv preprint arXiv:2101.00178, 2021.
  6. 6.D. Chong and J. N. Druckman. Framing theory. Annu. Rev. Polit. Sci., 10:103–126, 2007.
  7. 7.P. Christiano, B. Shlegeris, and D. Amodei. Supervising strong learners by amplifying weak experts. arXiv preprint arXiv:1810.08575, 2018.
  8. 8.O. Evans, O. Cotton-Barratt, L. Finnveden, A. Bales, A. Balwit, P. Wills, L. Righetti, and W. Saunders. Truthful AI: Developing and governing AI that does not lie. arXiv preprint arXiv:2110.06674, 2021.
  9. 9.A. Fan, Y. Jernite, E. Perez, D. Grangier, J. Weston, and M. Auli. ELI5: Long form question answering. arXiv preprint arXiv:1907.09190, 2019.
  10. 10.D. Ferrucci, E. Brown, J. Chu-Carroll, J. Fan, D. Gondek, A. A. Kalyanpur, A. Lally, J. W. Murdock, E. Nyberg, J. Prager, et al. Building watson: An overview of the deepqa project. AI magazine, 31 (3):59–79, 2010.
  11. 11.K. Goddard, A. Roudsari, and J. C. Wyatt. Automation bias: a systematic review of frequency, effect mediators, and mitigators. Journal of the American Medical Informatics Association, 19(1): 121–127, 2012.
  12. 12.I. Gur, U. Rueckert, A. Faust, and D. Hakkani-Tur. Learning to navigate the web. arXiv preprint arXiv:1812.09195, 2018.
  13. 13.K. Guu, K. Lee, Z. Tung, P. Pasupat, and M.-W. Chang. REALM: Retrieval-augmented language model pre-training. arXiv preprint arXiv:2002.08909, 2020.
  14. 14.M. Harms. Crystal Society. Crystal Trilogy. CreateSpace Independent Publishing Platform, 2016. ISBN 9781530773718.
  15. 15.G. Irving, P. Christiano, and D. Amodei. AI safety via debate. arXiv preprint arXiv:1805.00899, 2018.
  16. 16.M. Joshi, E. Choi, D. S. Weld, and L. Zettlemoyer. TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension. arXiv preprint arXiv:1705.03551, 2017.
  17. 17.V. Karpukhin, B. Oguz, S. Min, P. Lewis, L. Wu, S. Edunov, D. Chen, and W.-t. Yih. Dense passage retrieval for open-domain question answering. arXiv preprint arXiv:2004.04906, 2020.
  18. 18.K. Krishna, A. Roy, and M. Iyyer. Hurdles to progress in long-form question answering. arXiv preprint arXiv:2103.06332, 2021.
  19. 19.J. Leike, D. Krueger, T. Everitt, M. Martic, V. Maini, and S. Legg. Scalable agent alignment via reward modeling: a research direction. arXiv preprint arXiv:1811.07871, 2018.
  20. 20.P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel, et al. Retrieval-augmented generation for knowledge-intensive NLP tasks. arXiv preprint arXiv:2005.11401, 2020a.
  21. 21.P. Lewis, P. Stenetorp, and S. Riedel. Question and answer test-train overlap in open-domain question answering datasets. arXiv preprint arXiv:2008.02637, 2020b.
  22. 22.S. Lin, J. Hilton, and O. Evans. TruthfulQA: Measuring how models mimic human falsehoods. arXiv preprint arXiv:2109.07958, 2021.
  23. 23.J. Maynez, S. Narayan, B. Bohnet, and R. McDonald. On faithfulness and factuality in abstractive summarization. arXiv preprint arXiv:2005.00661, 2020.
  24. 24.D. Metzler, Y. Tay, D. Bahri, and M. Najork. Rethinking search: Making experts out of dilettantes. arXiv preprint arXiv:2105.02274, 2021.
  25. 25.B. T. Polyak and A. B. Juditsky. Acceleration of stochastic approximation by averaging. SIAM journal on control and optimization, 30(4):838–855, 1992.
  26. 26.J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017.
  27. 27.T. Shi, A. Karpathy, L. Fan, J. Hernandez, and P. Liang. World of bits: An open-domain platform for web-based agents. In International Conference on Machine Learning, pages 3135–3144. PMLR, 2017.
  28. 28.K. Shuster, S. Poff, M. Chen, D. Kiela, and J. Weston. Retrieval augmentation reduces hallucination in conversation. arXiv preprint arXiv:2104.07567, 2021.
  29. 29.N. Stiennon, L. Ouyang, J. Wu, D. M. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. Christiano. Learning to summarize from human feedback. arXiv preprint arXiv:2009.01325, 2020.
  30. 30.X. Yuan, J. Fu, M.-A. Cote, Y. Tay, C. Pal, and A. Trischler. Interactive machine comprehension with information seeking agents. arXiv preprint arXiv:1908.10449, 2019.

Citation

MLA
Nakano, R., et al. “WebGPT: Browser-assisted Question-answering with Human Feedback”. arXiv, 2021, http://arxiv.org/abs/2112.09332v3.
APA
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., Jiang, X., Cobbe, K., Eloundou, T., Krueger, G., Button, K., Knight, M., Chess, B., & Schulman, J. (2021). WebGPT: Browser-assisted question-answering with human feedback. arXiv. http://arxiv.org/abs/2112.09332v3
Chicago
Nakano, R., J. Hilton, S. Balaji, et al. 2021. “WebGPT: Browser-assisted Question-answering with Human Feedback”. arXiv. http://arxiv.org/abs/2112.09332v3.
Harvard
Nakano, R. et al. (2021) “WebGPT: Browser-assisted question-answering with human feedback”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2112.09332v3.
Vancouver
1. Nakano R, Hilton J, Balaji S, et al (2021) WebGPT: Browser-assisted question-answering with human feedback. arXiv

BibTeX

@article{nakano2021webgpt,
  title = {WebGPT: Browser-assisted question-answering with human feedback},
  author = {Nakano, Reiichiro and Hilton, Jacob and Balaji, Suchir and Wu, Jeff and Ouyang, Long and Kim, Christina and Hesse, Christopher and Jain, Shantanu and Kosaraju, Vineet and Saunders, William and Jiang, Xu and Cobbe, Karl and Eloundou, Tyna and Krueger, Gretchen and Button, Kevin and Knight, Matthew and Chess, Benjamin and Schulman, John},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2112.09332v3},
  eprint = {2112.09332}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission