An Information-theoretic Approach to Prompt Engineering Without Ground Truth Labels

Taylor SorensenJoshua RobinsonChristopher Michael RyttingAlexander Glenn ShawKyle Jeffrey RogersAlexia Pauline DeloreyMahmoud KhalilNancy FuldaDavid Wingate

article2022ACL159 citations

Proposes an unsupervised, black-box prompt selection method that optimizes mutual information between inputs and language model outputs to identify top-performing prompt templates without requiring labeled data or access to model weights.

Listen

Large language models store vast linguistic and factual knowledge, but their performance depends heavily on the exact wording of the natural language prompt templates used to query them. Minor prompt rephrasings can cause severe swings in accuracy. Currently, identifying effective prompt templates requires either large quantities of expensive labeled data for validation or direct access to internal model weights for gradient-based optimization, which is often impossible when using proprietary or commercial application programming interfaces.

The article demonstrates a method for selecting high-performing prompt templates without requiring any labeled ground-truth data or access to model parameters. Specifically, the researchers evaluate whether an information-theoretic metric called mutual information can serve as a dependable surrogate for task accuracy when choosing among candidate prompt templates.

The researchers developed a five-step framework centered on maximizing the mutual information between the model's inputs and outputs across candidate templates. In high-level terms, this metric favors prompt designs where the model exhibits high output confidence on individual inputs while maintaining an unbiased, balanced distribution of predictions overall. The team evaluated 20 candidate templates per dataset across eight benchmark datasets covering seven standard language processing tasks, including question-answering, sentiment analysis, and reading comprehension. The evaluation spanned eight causal language models ranging in scale from 124 million to 175 billion parameters.

The analysis yielded several key findings. First, mutual information strongly correlates with actual test accuracy, particularly in larger language models. Second, on the largest 175-billion-parameter model, prompt selection via mutual information closed approximately 90% of the gap between the average candidate prompt accuracy and the optimal prompt accuracy, achieving near-perfect selection performance without labels. Third, mutual information proved effective even in extremely low-data settings, selecting better prompts using as few as two unlabeled examples than standard methods could select using two labeled examples. Fourth, creating a combined ensemble of the top five prompt templates ranked by mutual information matched or outperformed an ensemble of all twenty candidate prompts while reducing computational costs by about 75%. Finally, prompt templates chosen via mutual information demonstrated strong cross-model transferability, showing that prompts selected on large models generalize well to others.

These findings indicate that organizations can effectively optimize language model performance without incurring the timeline delays and high labor costs associated with manual data labeling. The approach also enables high-quality prompt selection when interacting with closed-source, black-box systems where model parameters are hidden. Furthermore, prompt ensembling using top mutual information scores offers a cost-effective hedge against occasionally deceptive or erratic prompt templates.

Practitioners should adopt mutual information-based selection when deploying models on tasks lacking labeled validation data, starting with a diverse set of candidate templates. For mission-critical applications requiring risk mitigation, teams should ensemble predictions across the top-ranked templates. However, decision-makers should note that this method relies on the underlying model already possessing the capability to perform the given task; mutual information cannot compensate if all candidate prompts are poor or if the task exceeds the model's baseline reasoning capabilities. Confidence remains highest when applying the technique to large, well-calibrated models.

arXiv: 2203.11364
Cover for An Information-theoretic Approach to Prompt Engineering Without Ground Truth Labels

Abstract

Pre-trained language models derive substantial linguistic and factual knowledge from the massive corpora on which they are trained, and prompt engineering seeks to align these models to specific tasks. Unfortunately, existing prompt engineering methods require significant amounts of labeled data, access to model parameters, or both. We introduce a new method for selecting prompt templates without labeled examples and without direct access to the model. Specifically, over a set of candidate templates, we choose the template that maximizes the mutual information between the input and the corresponding model output. Across 8 datasets representing 7 distinct NLP tasks, we show that when a template has high mutual information, it also has high accuracy on the task. On the largest model, selecting prompts with our method gets 90% of the way from the average prompt accuracy to the best prompt accuracy and requires no ground truth labels.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methods
  • 3.1 Task Definition
  • 3.2 Mutual Information
  • 4 Experimental Setup
  • 4.1 Datasets
  • 4.2 Models
  • 5 Results
  • 5.1 Template Selection Performance
  • 5.2 Correlation between Template Mutual Information and Accuracy
  • 5.3 Compared to Few Labeled Examples
  • 5.4 Method Robustness and Ensembling
  • 5.5 Transferability across Models
  • 6 Conclusion
  • 7 Ethics
  • Acknowledgements
  • References
  • A Prompt Engineering Process
  • B Additional Figures
  • B.1 Mutual Information vs. Accuracy
  • B.2 Per Dataset Transfer Heatmaps
  • C Template Examples
  • C.1 SQuAD

Knowls

  1. Knowl 1 — Mutual information as a label-free prompt-selection criterion

    model/method

    For a fixed language model and a set of candidate prompt templates, select the template whose prompt representation has the greatest mutual information with the task answer. The criterion uses model-predicted answer distributions on unlabeled inputs rather than comparisons to ground-truth answers, and requires no direct access to model parameters. The paper motivates this as a surrogate: high mutual information favors templates whose predictions vary across inputs while remaining confident for individual inputs. The authors hypothesize that this tends to align model outputs with task answers; it does not guarantee high accuracy.

  2. Knowl 2 — Estimating prompt mutual information from model predictions

    equation

    For a candidate template fθf_\theta, draw NN inputs xix_i independently from the task input distribution XX. Let Y\mathcal{Y} be the answer space and let pi(y)=P(Y=y∣fθ(xi))p_i(y)=P(Y=y\mid f_\theta(x_i)) be the language model's collapsed probability for answer yy on input xix_i. Estimate mutual information in nats as the entropy of the average predicted distribution minus the average entropy of the individual predictions:

    I^θ=H(pˉθ)−1N∑i=1NH(pi),pˉθ(y)=1N∑i=1Npi(y),H(q)=−∑y∈Yq(y)ln⁡q(y).\widehat I_\theta=H(\bar p_\theta)-\frac{1}{N}\sum_{i=1}^{N}H(p_i), \qquad \bar p_\theta(y)=\frac{1}{N}\sum_{i=1}^{N}p_i(y), \qquad H(q)=-\sum_{y\in\mathcal{Y}}q(y)\ln q(y).

    Here θ\theta identifies the candidate template, NN is the number of sampled inputs, pip_i and pˉθ\bar p_\theta are probability distributions over answers, and HH is categorical entropy. The selected template maximizes I^θ\widehat I_\theta.

  3. Knowl 3 — One-token Response task representation and answer collapsing

    model/method

    The One-token Response (OTR) framework turns a task into prediction over answers that each begin with a distinct model token. A template maps an input to natural-language context; the language model supplies a probability distribution over next tokens; and a collapsing function groups token probabilities by answer, sums probabilities for tokens that begin an answer (including lexical variants such as capitalization and surrounding whitespace), ignores tokens that match no answer, and normalizes the answer scores to obtain P(Y∣x)P(Y\mid x). The predicted answer is the one with greatest collapsed probability. OTR is convenient for classification and other tasks with distinct answer-initial tokens, but is not directly suited to tasks such as translation or summarization. For open-ended answers, a wrong completion can share the correct answer's first token; on ROCStories, an exact-match temperature-zero check changed accuracy by only −0.03-0.03 on average and increased the MI–accuracy correlation from 0.680.68 to 0.720.72.

  4. Knowl 4 — Label-free prompt-selection procedure

    algorithm

    Given a task, an inference-capable language model, candidate input instances, and a set of candidate answer mappings:

    1. Construct KK diverse natural-language templates. For each template, define a collapsing function that maps model next-token probabilities to a normalized distribution over task answers. The paper's templates were written by hand and varied in framing, including question-and-answer, dialogue, quiz, and code-like formats.
    2. Test each template on a few inputs without consulting their labels. Keep or revise templates according to whether the model assigns probability to tokens that can sensibly be mapped to the task answers.
    3. Sample NN unlabeled inputs. For every template and input, obtain the model's next-token probabilities and collapse them to an answer distribution.
    4. Estimate mutual information for each template using the entropy of the mean answer distribution minus the mean entropy of the per-input answer distributions. Select the template with the highest estimate; with additional inference budget, an ensemble of the highest-MI templates can be used.
    5. Apply the selected template or ensemble to inputs for inference. The model used at inference may differ from the selection model, including being smaller when inference cost is a concern.

    The procedure needs neither ground-truth labels nor parameter access; it does require model probability outputs, candidate templates, and a task-to-token answer mapping.

  5. Knowl 5 — Experimental tasks, models, and sampling

    experimental setup

    The evaluation used eight datasets spanning seven task types: SQuAD (open-book question answering), LAMBADA and ROCStories (cloze prediction), CoQA (five-choice closed-book question answering), IMDB (sentiment classification), BoolQ (reading comprehension), COPA (causal alternative selection), and WiC (word-in-context classification). The authors sampled N=500N=500 instances per dataset. For ROCStories, they randomly masked a word in each five-sentence story. To fit OTR, SQuAD questions without one-word answers were removed, and CoQA questions whose choices shared a first word were removed. The eight causal language models ranged from GPT-2 124M and 1.5B, through GPT-Neo 2.7B, GPT-J 6B, and GPT-3 2.7B, 6.7B, and 13B, to GPT-3 Davinci 175B. Twenty candidate templates were evaluated for each dataset/model pair. Data were sampled from the train split for SQuAD and CoQA, train plus validation for WiC, COPA, and BoolQ, the full datasets for ROCStories and IMDB, and the test split for LAMBADA.

  6. Knowl 6 — Prompt-selection performance against random and oracle choices

    empirical result

    On GPT-3 Davinci (175B), the maximum-mutual-information template outperformed both the mean and median accuracy among the 20 candidates on all eight datasets, and it selected the most accurate candidate on three datasets. Averaged across datasets, the selected template achieved 90% of the distance from mean prompt accuracy to best-prompt accuracy, without labels. Across model sizes, MI selection was above the average prompt on several datasets for every model; excluding COPA and WiC, it selected an above-average template in 83% of model/dataset cases and in 100% of cases for the two largest models. The paper reports weaker selection where candidate templates were low-signal, including COPA and WiC.

  7. Knowl 7 — Correlation between template mutual information and accuracy

    empirical result

    For GPT-3 Davinci (175B), Pearson correlations across the 20 candidate templates between mean mutual information and mean accuracy were: SQuAD 0.920.92, LAMBADA 0.860.86, ROCStories 0.680.68, CoQA 0.560.56, IMDB 0.710.71, BoolQ 0.700.70, COPA 0.620.62, and WiC 0.330.33. The paper found the relationship more consistently strong for larger language models; it was less reliable for smaller models on some datasets. These results support MI as a label-free indicator of relative prompt quality, rather than establishing that high MI always entails high accuracy.

  8. Knowl 8 — Comparison with prompt selection using a few labeled examples

    empirical result

    For GPT-3 Davinci (175B), the authors compared MI-based selection with choosing a template by labeled training accuracy. They repeatedly partitioned the 500 instances into training sets of N=2,4,8,…,256N=2,4,8,\ldots,256 and held-out test sets, using 100 random partitions at each size. MI was estimated using only the NN input instances, without their labels. At low sample sizes, including N=2N=2, MI selection chose a substantially better-than-average template and, on average, beat selection by labeled training accuracy across all eight datasets. Labeled selection often improved at larger NN, but required labels.

  9. Knowl 9 — Ensembling the highest-MI templates

    empirical result

    To reduce sensitivity to an anomalous high-MI, low-accuracy template, the authors evaluated ensembles of five templates from each set of 20. They compared the top-five-MI ensemble with all (205){20\choose5} five-template ensembles and with the ensemble of all 20 templates. The top-five-MI ensemble achieved at least the accuracy of the all-20 ensemble on seven of eight datasets, while requiring one quarter as many templates at inference; the all-20 ensemble won in the remaining case. Across six datasets, the highest-MI templates were generally strong, whereas COPA and WiC were more brittle, making diverse candidate generation and caution about outliers important.

  10. Knowl 10 — Prompt transfer across selection and inference models

    empirical result

    The authors selected a template using one model and evaluated it using another, averaging normalized accuracy across datasets. The normalization maps average candidate-template accuracy to 00 and the best candidate-template accuracy to 11. MI-based prompt transfer was most effective when GPT-3 Davinci (175B) served as both selection and inference model, reaching an average normalized score of 0.900.90. Performance was most consistently high when a largest model was used for selection or inference. Across 64 selection/inference model pairings, only one had a negative average gain over the average-template baseline, indicating that transfer was often beneficial but not uniformly so.

Coverage note — The paper's individual prompt-template examples and per-dataset transfer heatmaps are omitted because they illustrate particular candidates or repeat the aggregate method and results rather than adding a separate, load-bearing contribution.

References

  1. 1.Asaf Amrami and Yoav Goldberg. 2018. Word Sense Induction with Neural biLM and Symmetric Patterns. pages 4860–4867.
  2. 2.Sid Black, Leo Gao, Phil Wang, Connor Leahy, and Stella Biderman. 2021. GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow.
  3. 3.Antoine Bosselut, Hannah Rashkin, Maarten Sap, Chaitanya Malaviya, Asli Celikyilmaz, and Yejin Choi. COMET : Commonsense Transformers for Automatic Knowledge Graph Construction.
  4. 4.Zied Bouraoui, Jose Camacho-collados, and Steven Schockaert. Inducing Relational Knowledge from BERT.
  5. 5.Peter F. Brown, Vincent J. Della Pietra, Peter V. deSouza, Jenifer C. Lai, and Robert L. Mercer. 1992. Class-based n-gram models of natural language. Computational Linguistics, 18(4):467–480.
  6. 6.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. arXiv.
  7. 7.Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019. Boolq: Exploring the surprising difficulty of natural yes/no questions. CoRR, abs/1905.10044.
  8. 8.Thomas M. Cover and Joy A. Thomas. 2006. Elements of Information Theory 2nd Edition (Wiley Series in Telecommunications and Signal Processing). Wiley-Interscience.
  9. 9.Jacob Devlin, Ming Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. NAACL HLT 2019 - 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies - Proceedings of the Conference, 1(Mlm):4171–4186.
  10. 10.Andrew Gordon, Zornitsa Kozareva, and Melissa Roemmele. 2012. SemEval-2012 task 7: Choice of plausible alternatives: An evaluation of commonsense causal reasoning. In *SEM 2012: The First Joint Conference on Lexical and Computational Semantics – Volume 1: Proceedings of the main conference and the shared task, and Volume 2: Proceedings of the Sixth International Workshop on Semantic Evaluation (SemEval 2012), pages 394–398, Montréal, Canada. Association for Computational Linguistics.
  11. 11.Yuxian Gu, Xu Han, Zhiyuan Liu, and Minlie Huang. 2021. PPT: Pre-trained Prompt Tuning for Few-shot Learning.
  12. 12.Ziwei Ji, Justin D. Li, and Matus Telgarsky. 2021. Early-stopped neural networks are consistent.
  13. 13.Lingpeng Kong, Cyprien de Masson d’Autume, Wang Ling, Lei Yu, Zihang Dai, and Dani Yogatama. 2019. A mutual information maximization perspective of language representation learning.
  14. 14.Brian Lester, Rami Al-Rfou, and Noah Constant. 2021. The Power of Scale for Parameter-Efficient Prompt Tuning.
  15. 15.Xiang Lisa Li and Percy Liang. 2021. Prefix-Tuning: Optimizing Continuous Prompts for Generation. pages 4582–4597.
  16. 16.Nelson F. Liu, Matt Gardner, Yonatan Belinkov, Matthew E. Peters, and Noah A. Smith. 2019. Linguistic knowledge and transferability of contextual representations. NAACL HLT 2019 - 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies - Proceedings of the Conference, 1:1073–1094.
  17. 17.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2021a. Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing. pages 1–46.
  18. 18.Xiao Liu, Yanan Zheng, Zhengxiao Du, Ming Ding, Yujie Qian, Zhilin Yang, and Jie Tang. 2021b. GPT Understands, Too.
  19. 19.Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. 2021. Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity.
  20. 20.Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011. Learning word vectors for sentiment analysis. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, pages 142–150, Portland, Oregon, USA. Association for Computational Linguistics.
  21. 21.Nasrin Mostafazadeh, Nathanael Chambers, Xiaodong He, Devi Parikh, Dhruv Batra, Lucy Vanderwende, Pushmeet Kohli, and James F. Allen. 2016. A corpus and evaluation framework for deeper understanding of commonsense stories. CoRR, abs/1604.01696.
  22. 22.Preetum Nakkiran and Yamini Bansal. 2020. Distributional generalization: A new kind of generalization. CoRR, abs/2009.08092.
  23. 23.Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Quan Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández. 2016. The LAMBADA dataset: Word prediction requiring a broad discourse context. CoRR, abs/1606.06031.
  24. 24.Ethan Perez, Douwe Kiela, and Kyunghyun Cho. 2021. True Few-Shot Learning with Language Models. (Cv):1–21.
  25. 25.Fabio Petroni, Tim Rocktäschel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H. Miller, and Sebastian Riedel. 2020. Language models as knowledge bases? EMNLP-IJCNLP 2019 - 2019 Conference on Empirical Methods in Natural Language Processing and 9th International Joint Conference on Natural Language Processing, Proceedings of the Conference, pages 2463–2473.
  26. 26.Mohammad Taher Pilehvar and José Camacho-Collados. 2018. Wic: 10, 000 example pairs for evaluating context-sensitive representations. CoRR, abs/1808.09121.
  27. 27.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners.
  28. 28.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. pages 1–53.
  29. 29.Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018. Know what you don’t know: Unanswerable questions for squad. CoRR, abs/1806.03822.
  30. 30.Laria Reynolds and Kyle McDonell. 2021. Prompt programming for large language models: Beyond the few-shot paradigm.
  31. 31.Karl Stratos. 2019. Mutual information maximization for simple and accurate part-of-speech induction.
  32. 32.Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2018. Commonsenseqa: A question answering challenge targeting commonsense knowledge. CoRR, abs/1811.00937.
  33. 33.Elena Voita and Ivan Titov. 2020. Information-Theoretic Probing with Minimum Description Length. pages 183–196.
  34. 34.Ben Wang and Aran Komatsuzaki. 2021. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/mesh-transformer-jax.
  35. 35.Ningyu Zhang, Luoqiu Li, Xiang Chen, Shumin Deng, Zhen Bi, Chuanqi Tan, Fei Huang, and Huajun Chen. 2021. Differentiable Prompt Makes Pre-trained Language Models Better Few-shot Learners. pages 1–18.
  36. 36.Tony Z. Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. Calibrate Before Use: Improving Few-Shot Performance of Language Models.
  37. 37.Yukun Zuo, Quan Fang, Shengsheng Qian, Xiaorui Zhang, and Changsheng Xu. 2018. Representation Learning of Knowledge Graphs with Entity Attributes and Multimedia Descriptions. 2018 IEEE 4th International Conference on Multimedia Big Data, BigMM 2018, pages 2659–2665.

Citation

MLA
Sorensen, T., et al. “An Information-theoretic Approach to Prompt Engineering Without Ground Truth Labels”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 819–62, https://doi.org/10.18653/v1/2022.acl-long.60.
APA
Sorensen, T., Robinson, J., Rytting, C., Shaw, A., Rogers, K., Delorey, A., Khalil, M., Fulda, N., & Wingate, D. (2022). An Information-theoretic Approach to Prompt Engineering Without Ground Truth Labels. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 819–862. https://doi.org/10.18653/v1/2022.acl-long.60
Chicago
Sorensen, T., J. Robinson, C. Rytting, et al. 2022. “An Information-theoretic Approach to Prompt Engineering Without Ground Truth Labels”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 819–62. https://doi.org/10.18653/v1/2022.acl-long.60.
Harvard
Sorensen, T. et al. (2022) “An Information-theoretic Approach to Prompt Engineering Without Ground Truth Labels”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 819–862. Available at: https://doi.org/10.18653/v1/2022.acl-long.60.
Vancouver
1. Sorensen T, Robinson J, Rytting C, Shaw A, Rogers K, Delorey A, Khalil M, Fulda N, Wingate D (2022) An Information-theoretic Approach to Prompt Engineering Without Ground Truth Labels. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 819–862

BibTeX

@inproceedings{sorensen-etal-2022-information,
    title = "An Information-theoretic Approach to Prompt Engineering Without Ground Truth Labels",
    author = "Sorensen, Taylor  and
      Robinson, Joshua  and
      Rytting, Christopher  and
      Shaw, Alexander  and
      Rogers, Kyle  and
      Delorey, Alexia  and
      Khalil, Mahmoud  and
      Fulda, Nancy  and
      Wingate, David",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.60/",
    doi = "10.18653/v1/2022.acl-long.60",
    pages = "819--862"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/