Entailer: Answering Questions with Faithful and Truthful Chains of Reasoning

Oyvind TafjordBhavana Dalvi MishraPeter Clark

article2022EMNLP62 citations

Proposes a question-answering framework that couples backward-chaining entailment generation with self-query verification to construct multi-step reasoning trees directly reflecting a language model's internal beliefs.

Listen

Large pretrained language models excel at answering complex questions, yet they typically function as black boxes. When these models provide explanations or rationale chains, those explanations are frequently unfaithful, meaning the final answer does not strictly follow from the reasoning, or untruthful, meaning the model generates statements it does not internally verify as true. This lack of transparency limits organizational trust, complicates compliance and safety audits, and makes it difficult to diagnose why a model failed. The article addresses this challenge by designing an architecture that answers questions through systematic, verifiable reasoning grounded in the model's own internal beliefs.

The main objective of the article is to demonstrate Entailer, a question-answering framework that produces faithful, multistep entailment proofs by combining backward-chaining generation with self-verification. The approach utilizes an 11-billion parameter language model trained across three core tasks: generating premises that imply a given answer hypothesis, scoring whether the model believes those individual premises are true, and validating whether the logical entailment step is sound. To find the strongest answer, the system searches backward from candidate answers, generates potential premises, and filters them out if self-querying shows low belief scores or flawed logic. The model was trained on the EntailmentBank science dataset augmented by thousands of crowdsourced positive and negative examples, and was subsequently evaluated zero-shot without task-specific fine-tuning on multiple-choice reasoning benchmarks.

The findings show that Entailer generates structured reasoning proofs while preserving strong baseline accuracy. On benchmark evaluations, Entailer achieved a question-answering accuracy of roughly 75%, matching the accuracy of direct, unreasoned answers. In human evaluations, human judges determined that over 70% of Entailer's generated reasoning chains clearly proved the conclusion from the premises, compared to only 34% for explanations from a leading alternative question-answering model. Additionally, annotators confirmed that approximately 90% of the self-verified premises generated by Entailer were factually correct, preferring its structured explanations over the baseline by a margin of 57% to 23%. An analysis of errors revealed that 47% stemmed from reasoning flaws such as near-tautologies, 33% from incorrect factual beliefs, and 20% from dataset ambiguities.

These results demonstrate that language models can expose their latent knowledge as logical proof trees without incurring an accuracy penalty. For operational workflows, this capability enhances transparency and enables targeted debugging: when an answer is wrong, operators can inspect the reasoning tree to identify precisely which belief or deduction failed. This structured visibility lays the foundation for interactive and teachable artificial intelligence systems, where human feedback can correct an isolated erroneous premise rather than requiring opaque, expensive model retraining.

Stakeholders should treat Entailer as an effective framework for high-stakes domains that demand explainable reasoning, such as technical diagnostics or compliance support. Next steps should focus on implementing user-feedback loops where human corrections are stored in a retrieval memory to dynamically override flawed beliefs. However, leaders should note key operational limitations: Entailer is computationally intensive, requiring up to 360 seconds per question due to recursive multi-step search, and it can occasionally produce circular reasoning or verify contradictory statements. Initial deployments should therefore target asynchronous diagnostic tasks rather than real-time consumer applications until faster search mechanisms are implemented.

  • Paper: Language Models (Mostly) Know What They Know, Saurav Kadavath et al. (2022). Read this account of how language models assess the correctness of their own answers first; it grounds Entailer’s use of self-reported beliefs to screen premises.
Cover for Entailer: Answering Questions with Faithful and Truthful Chains of Reasoning

Abstract

Our goal is for a system to be able to answer questions and justify its answers with a faithful and truthful chain of reasoning. We present Entailer, a system that produces chains of reasoning by decomposing the question into subquestions and then answering each subquestion using a textual entailment model. The model is trained on a new dataset of 10K examples of multistep reasoning, and uses a novel method to generate the training data automatically from a large corpus. Entailer achieves state-of-the-art results on two multistep reasoning datasets, and its reasoning chains are more faithful and truthful than those of a large language model, as measured by human evaluation and a new automatic metric. We also show that Entailer can be used to improve the performance of a large language model by providing it with a faithful derivation of the answer.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Approach
  • 3.1 Hypothesis Generation
  • 3.2 Generating Entailment Trees
  • 3.2.2 Backward Chaining
  • 4 Model Training
  • 4.1 Data Sources
  • 4.1.1 EntailmentBank
  • 4.1.2 Crowdsourced Data
  • 4.1.3 Optional Fields
  • 4.2 Model Details
  • 5 Evaluation
  • 5.1 QA Accuracy
  • 5.2 Results
  • 5.3 Human Judgements
  • 5.4 Analysis
  • 5.4.2 When do proofs do better?
  • 6 Towards Teachable Systems
  • 7 Conclusion
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • Appendix A. Entailer's Backward Chaining Algorithm
  • A.1 Generating One Backward-Chaining Step
  • A.2 Backward Chaining
  • Appendix B. Crowdsourcing Instructions for Verifying Entailments (Section 4.1.2)
  • Appendix C: Model Training
  • C.1 Dataset Preparation
  • C.2 Model Details
  • Appendix D. Examples of Macaw Explanations and Entailer Proofs

Knowls

  1. Knowl 1 — Entailer makes its internal beliefs explicit as reasoning trees

    model/method

    Entailer answers a question by converting candidate answers into declarative hypotheses and seeking a tree of natural-language entailments for each hypothesis. Each internal node states a conclusion and the premises that jointly support it; the leaves are statements that require no further proof. The system operationalizes a model “believing” a statement as the model answering yes when queried about its truth, and it separately checks whether each premise is believed and whether each inference is valid. A chain is intended to be truthful in the operational sense that it reflects the model’s checked beliefs, and faithful in the sense that its entailment steps support the answer. These checks do not guarantee that the beliefs are true in the world. The resulting tree exposes some of the model’s latent beliefs and how they support a selected answer.

  2. Knowl 2 — A single-step entailment model generates and verifies premises

    model/method

    Entailer uses three model behaviors for a hypothesis HH, an English statement to prove, and a set of premises P={p1,…,pm}P=\{p_1,\ldots,p_m\} that may jointly entail it. First, H→PH\rightarrow P generates candidate premises. Second, H→SdH\rightarrow S_d assigns a direct score SdS_d in [0,1][0,1] to whether a statement is true according to the model. Third, PH→SePH\rightarrow S_e scores whether the inference P⊢HP\vdash H is valid, independently of whether its premises are true. The question, answer choice, and optional context may also be supplied as inputs; the reported test-time experiments did not use context.

    For one backward-chaining step, the system samples kk candidate premise sets, checks every premise with the direct-scoring behavior and each inference with the entailment-scoring behavior, and rejects candidates with a premise or inference score below 0.50.5. It ranks the remaining candidates by

    sr-1deep(H)=(∏i=1msd(pi))se(P⊢H),s_{r\text{-}1deep}(H)=\left(\prod_{i=1}^{m}s_d(p_i)\right)s_e(P\vdash H),

    where sd(pi)s_d(p_i) is the direct truth score for premise pip_i and se(P⊢H)s_e(P\vdash H) is the inference-validity score. The highest-scoring candidate is returned as the step supporting HH.

  3. Knowl 3 — Recursive backward chaining selects a proof or a direct answer

    algorithm

    For each answer hypothesis, Entailer recursively expands premises to search for a higher-confidence proof, stopping at a maximum depth or when expansion no longer improves the node’s confidence. The system requires at least a one-step proof for the top-level hypothesis. For multiple-choice questions it creates a declarative hypothesis for each answer option; open-ended candidate answers can first be generated by Macaw and converted into hypotheses.

    At a node NN, let sd(N)∈[0,1]s_d(N)\in[0,1] be the model’s direct score that the statement is true. Its direct confidence in its predicted truth value is cd(N)=max⁡(sd(N),1−sd(N))c_d(N)=\max(s_d(N),1-s_d(N)). For a candidate entailment P⊢NP\vdash N, the recursive proof score is

    sr(N)=(∏i=1ms(pi))se(P⊢N),s_r(N)=\left(\prod_{i=1}^{m}s(p_i)\right)s_e(P\vdash N),

    where pip_i are the premises, s(pi)s(p_i) is each premise’s overall selected score (direct or proof-derived), and se(P⊢N)s_e(P\vdash N) is the entailment-validity score. The proof confidence is cr(N)=sr(N)c_r(N)=s_r(N). The system retains a proof when its proof confidence exceeds the direct confidence; otherwise it uses the direct answer and treats the node as a leaf. The top-level hypothesis is expanded even if it would not otherwise pass this comparison, to ensure a nontrivial proof. The answer returned is the hypothesis with the highest overall score across the candidate answers. In evaluation, greedy search used the first generated step (k=1k=1); sampled top-level search selected the best of k=6k=6 root steps.

  4. Knowl 4 — Crowdsourced annotations add negative entailments to training data

    data/table

    To add negative examples to the positive entailment steps in EntailmentBank, the authors used an earlier Entailer model to generate one-step proofs for incorrect answer options in its 1,313 training questions. This produced 3,939 candidate proofs for false hypotheses. After excluding examples with more than two premises, crowdworkers annotated 3,673 proofs, judging each premise’s truth and the validity of the inference separately with true, false, or unsure labels. Each proof received three annotations, with three additional workers consulted when the initial annotations had no majority. After removing items without a final majority verdict, the annotations contributed 7,013 labeled premises for direct truth scoring and 3,391 labeled inferences for entailment scoring. The paper describes this new crowdsourced multi-premise entailment data as doubling the amount of data available in EntailmentBank and adding negative examples, which EntailmentBank did not contain.

  5. Knowl 5 — Entailer is trained once as a multi-task T5-11B model

    experimental setup

    The authors trained Entailer as a T5-11B multi-task model for premise generation, direct truth scoring, and entailment-validity scoring, then froze it and applied it zero-shot to new question-answering datasets. Its training data included 4,175 one-step positive entailments obtained by breaking apart EntailmentBank trees, about 9,000 true premises, and the crowdsourced premise and inference labels for negative examples. Training examples included variants with optional question-answer inputs and contexts of up to five sentences. The contexts mixed gold proof facts with less relevant facts, including sentences retrieved using BM25 from a science-filtered Wikipedia corpus of about 1.5 million sentences; no context was used in the paper’s test-time experiments.

    The model was fine-tuned for 20,000 steps with batch size 8, using the T5 library’s default hyperparameters, including Adafactor, and the checkpoint with the best validation score. For sampled generation, the paper used nucleus sampling with temperature 2.0 and top-p=0.95p=0.95.

  6. Knowl 6 — Zero-shot proof-based QA is competitive with direct QA

    empirical result

    The authors evaluated zero-shot multiple-choice accuracy on 500 OBQA questions and 557 QuaRTz questions. Direct QA used Entailer’s direct-scoring behavior; proof-based QA generated and scored entailment trees. The chart reports the following accuracies (percent):

    • OBQA: direct QA 75.2; Entailer greedy, maximum depth 1/3: 72.8/69.0; Entailer sampled top-level (k=6k=6), depth 1/3: 76.8/74.8; sampled top-level Entailer+Direct, depth 1/3: 74.6/75.2.
    • QuaRTz: direct QA 74.1; Entailer greedy, maximum depth 1/3: 73.6/68.8; Entailer sampled top-level (k=6k=6), depth 1/3: 74.3/73.8; sampled top-level Entailer+Direct, depth 1/3: 74.7/75.4.

    Sampling candidate root proofs outperformed greedy proof selection in the reported comparisons. The authors report that deeper proofs did not significantly change accuracy. The Entailer+Direct combination selected the more confident answer from the two methods, but it could select an answer other than the one with the highest-scoring proof and thus did not always preserve faithfulness; this occurred for 16.8% of OBQA questions and 14.2% of QuaRTz questions. The combination had no significant performance gain in these experiments.

  7. Knowl 7 — Human judges rated Entailer’s proof explanations more favorably

    empirical result

    Six annotators compared one-step Entailer proofs with Macaw explanations for 100 correct answers to OBQA development questions. They rated whether the conclusion clearly followed from the premises, whether the premises were correct and relevant, whether the explanation added something useful, and which explanation they preferred. The percentage of “yes” judgments was 71% for Entailer versus 34% for Macaw on whether the conclusion clearly followed; 91% versus 86% on premise correctness; 81% versus 62% on premise relevance; and 64% versus 57% on whether the explanation introduced something new and useful. Annotators preferred Entailer in 57% of pairs and Macaw in 23% (the remaining 20% were rated similar). For Entailer, the authors report that nearly all of the remaining premise-correctness judgments were “unsure”; only one Entailer fact was judged incorrect.

  8. Knowl 8 — Wrong answers arise from belief, reasoning, and dataset errors

    empirical result

    The authors manually classified the 51 of 500 OBQA cases in which Entailer chose an incorrect answer while direct QA chose the correct one. They attributed 33% to belief errors, where a false premise supported the answer; 20% to dataset errors, such as ambiguity or multiple defensible answer options; and 47% to reasoning errors. The reasoning errors comprised near-tautological proofs (20% of the 51 cases), invalid basic entailments (10%), incorrect abductive inferences (9%), and quantification or exception problems (8%). The categories distinguish false model beliefs from invalid inference and from questions whose answer labels or wording may themselves be problematic.

  9. Knowl 9 — The system remains limited by proof quality, belief consistency, and search cost

    limitation

    Entailer does not always generate coherent proofs: some entailments are invalid and some chains are nearly tautological. Its reasoning operation relies on natural-language entailment, whose validity is imprecisely defined and can introduce noise into training and evaluation. The method also assumes that a model is generally consistent about its beliefs, but it does not resolve cases where the model verifies contradictory statements. Recursive search is computationally expensive: for a four-option multiple-choice question with sample size k=6k=6, the reported average runtime on a 48-GB GPU was about 80 seconds for depth-1 proofs and about 360 seconds for proofs up to depth 3. The paper’s proposal that users could correct faulty beliefs through supplied context and persistent memory is presented as a future possibility, not as a capability evaluated in these experiments.

Coverage note — The proposed persistent-memory interaction for users to correct model beliefs is not included as a separate knowl because the paper presents it as a conjectural future direction rather than an implemented or evaluated contribution.

References

  1. 1.Antoine Bosselut, Hannah Rashkin, Maarten Sap, Chaitanya Malaviya, Asli Celikyilmaz, and Yejin Choi. 2019. COMET: Commonsense transformers for automatic knowledge graph construction. In ACL.
  2. 2.Kaj Bostrom, Zayne Sprague, Swarat Chaudhuri, and Greg Durrett. 2022. Natural language deduction through search over statement compositions. ArXiv, abs/2201.06028.
  3. 3.Kaj Bostrom, Xinyu Zhao, Swarat Chaudhuri, and Greg Durrett. 2021. Flexible generation of natural language deductions. In EMNLP.
  4. 4.Tom B. Brown et al. 2020. Language models are few-shot learners. In NeurIPS.
  5. 5.Peter Clark, Oyvind Tafjord, and Kyle Richardson. 2020. Transformers as soft reasoners over language. In IJCAI’20.
  6. 6.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training verifiers to solve math word problems. ArXiv, abs/2110.14168.
  7. 7.Antonia Creswell, Murray Shanahan, and Irina Higgins. 2022. Selection-inference: Exploiting large language models for interpretable logical reasoning. ArXiv, abs/2205.09712.
  8. 8.Ido Dagan, Dan Roth, Mark Sammons, and Fabio Massimo Zanzotto. 2013. Recognizing Textual Entailment: Models and Applications. Morgan and Claypool.
  9. 9.Bhavana Dalvi, Peter A. Jansen, Oyvind Tafjord, Zhengnan Xie, Hannah Smith, Leighanna Pipatanangkura, and Peter Clark. 2021. Explaining answers with entailment trees. In EMNLP.
  10. 10.Dorottya Demszky, Kelvin Guu, and Percy Liang. 2018. Transforming question answering datasets into natural language inference datasets. ArXiv, abs/1809.02922.
  11. 11.Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, E. Hovy, H. Schutze, and Yoav Goldberg. 2021. Measuring and improving consistency in pretrained language models. ArXiv, abs/2102.01017.
  12. 12.Saadia Gabriel, Chandra Bhagavatula, Vered Shwartz, Ronan Le Bras, Maxwell Forbes, and Yejin Choi. 2021. Paragraph-level commonsense transformers with recurrent memory. In AAAI.
  13. 13.Ari Holtzman, Jan Buys, Maxwell Forbes, and Yejin Choi. 2019. The curious case of neural text degeneration. In ICLR.
  14. 14.Ruixin Hong, Hongming Zhang, Xintong Yu, and Changshui Zhang. 2022. Metgen: A module-based entailment tree generation framework for answer explanation. ArXiv, abs/2205.02593.
  15. 15.Aditya Kalyanpur, Tom Breloff, and David A. Ferrucci. 2020. Braid: Weaving symbolic and neural knowledge into coherent logical explanations. arXiv: Computation and Language.
  16. 16.Nora Kassner and H. Schütze. 2020. Negated and misprimed probes for pretrained language models: Birds can talk, but cannot fly. In ACL.
  17. 17.Nora Kassner, Oyvind Tafjord, Hinrich Schutze, and Peter Clark. 2021. BeliefBank: Adding memory to a pre-trained language model for a systematic notion of belief. In EMNLP.
  18. 18.Tao Li, Vivek Gupta, Maitrey Mehta, and Vivek Srikumar. 2019. A logic-driven framework for consistency of neural models. In EMNLP.
  19. 19.Yujia Li, David H. Choi, et al. 2022. Competition-level code generation with alphacode. ArXiv, abs/2203.07814.
  20. 20.Christopher D. Manning and Bill MacCartney. 2009. Natural language inference. Stanford University.
  21. 21.Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018. Can a suit of armor conduct electricity? A new dataset for open book question answering. In EMNLP.
  22. 22.Bhavana Dalvi Mishra, Oyvind Tafjord, and Peter Clark. 2022. Towards teachable reasoning systems: Using a dynamic memory of user feedback for continual system improvement. In EMNLP.
  23. 23.Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, and Christopher D. Manning. 2021. Fast model editing at scale. ArXiv, abs/2110.11309.
  24. 24.Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, Charles Sutton, and Augustus Odena. 2021. Show your work: Scratchpads for intermediate computation with language models. ArXiv, abs/2112.00114.
  25. 25.F. Petroni, Tim Rocktäschel, Patrick Lewis, A. Bakhtin, Y. Wu, Alexander H. Miller, and S. Riedel. 2019. Language models as knowledge bases? In EMNLP.
  26. 26.Danilo Neves Ribeiro, Shen Wang, Xiaofei Ma, Rui Dong, Xiaokai Wei, Henry Zhu, Xinchi Chen, Zhiheng Huang, Peng Xu, Andrew O. Arnold, and Dan Roth. 2022. Entailment tree explanations via iterative retrieval-generation reasoner. ArXiv, abs/2205.09224.
  27. 27.Marco Tulio Ribeiro, Carlos Guestrin, and Sameer Singh. 2019. Are red roses red? Evaluating consistency of question-answering models. In ACL.
  28. 28.Eric Schwitzgebel. 2019. Belief. Stanford Encyclopedia of Philosophy. Https://plato.stanford.edu/entries/belief/.
  29. 29.Vered Shwartz, Peter West, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2020. Unsupervised commonsense question answering with self-talk. In EMNLP, pages 4615–4629.
  30. 30.Oyvind Tafjord and Peter Clark. 2021. General-purpose question-answering with Macaw. ArXiv, abs/2109.02593.
  31. 31.Oyvind Tafjord, Matt Gardner, Kevin Lin, and Peter Clark. 2019. QuaRTz: An open-domain dataset of qualitative relationship questions. ArXiv, abs/1909.03553.
  32. 32.Oyvind Tafjord, B. D. Mishra, and P. Clark. 2020. ProofWriter: Generating implications, proofs, and abductive statements over natural language. ArXiv, abs/2012.13048.
  33. 33.Alon Talmor, Oyvind Tafjord, P. Clark, Y. Goldberg, and Jonathan Berant. 2020. LeapOfThought: Teaching pre-trained models to systematically reason over implicit knowledge. In NeurIPS.
  34. 34.Niket Tandon, Aman Madaan, Peter Clark, and Yiming Yang. 2022. Memory-assisted prompt editing to improve GPT-3 after deployment. In ACL Workshop on Commonsense Representation and Reasoning (CSRR’22). (also arxiv:2201.06009).
  35. 35.Stefano Teso and Kristian Kersting. 2019. Explanatory interactive machine learning. Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society.
  36. 36.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, and Denny Zhou. 2022. Rationale-augmented ensembles in language models. ArXiv, abs/2207.00747.
  37. 37.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. ArXiv, abs/2201.11903.
  38. 38.Sarah Wiegreffe and Ana Marasovic. 2021. Teach me to explain: A review of datasets for explainable NLP. ArXiv, abs/2102.12060.
  39. 39.Zhengnan Xie, Sebastian Thiem, Jaycie Martin, Elizabeth Wainwright, Steven Marmorstein, and Peter Jansen. 2020. WorldTree V2: A corpus of science-domain structured explanations and inference patterns supporting multi-hop inference. In LREC.
  40. 40.Kaiyu Yang, Jia Deng, and Danqi Chen. 2022. Generating natural language proofs with verifier-guided search. ArXiv, abs/2205.12443.

Citation

MLA
Tafjord, O., et al. “Entailer: Answering Questions with Faithful and Truthful Chains of Reasoning”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 2078–93, https://doi.org/10.18653/V1/2022.EMNLP-MAIN.134.
APA
Tafjord, O., Dalvi Mishra, B., & Clark, P. (2022). Entailer: Answering Questions with Faithful and Truthful Chains of Reasoning. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2078–2093. https://doi.org/10.18653/V1/2022.EMNLP-MAIN.134
Chicago
Tafjord, O., B. Dalvi Mishra, and P. Clark. 2022. “Entailer: Answering Questions with Faithful and Truthful Chains of Reasoning”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2078–93. https://doi.org/10.18653/V1/2022.EMNLP-MAIN.134.
Harvard
Tafjord, O., Dalvi Mishra, B. and Clark, P. (2022) “Entailer: Answering Questions with Faithful and Truthful Chains of Reasoning”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 2078–2093. Available at: https://doi.org/10.18653/V1/2022.EMNLP-MAIN.134.
Vancouver
1. Tafjord O, Dalvi Mishra B, Clark P (2022) Entailer: Answering Questions with Faithful and Truthful Chains of Reasoning. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 2078–2093

BibTeX

@inproceedings{Tafjord_2022, title={Entailer: Answering Questions with Faithful and Truthful Chains of Reasoning}, url={http://dx.doi.org/10.18653/V1/2022.EMNLP-MAIN.134}, DOI={10.18653/v1/2022.emnlp-main.134}, booktitle={Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing}, publisher={Association for Computational Linguistics}, author={Tafjord, Oyvind and Dalvi Mishra, Bhavana and Clark, Peter}, year={2022}, pages={2078–2093} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/