MatSci-NLP: Evaluating Scientific Language Models on Materials Science Language Tasks Using Text-to-Schema Modeling

Yu SongSantiago MiretBang Liu

article2023ACL70 citations

Presents a benchmark spanning seven materials science NLP tasks and introduces a unified text-to-schema modeling approach that improves low-resource multitask performance across domain-specific language models.

Listen

Accelerating the discovery, synthesis, and manufacturing of advanced materials—such as those needed for clean energy, electronics, and sustainable manufacturing—requires extracting valuable scientific insights from vast volumes of unstructured text across research papers and technical reports. However, applying natural language processing (NLP) to materials science is hindered by fragmented research, complex specialized terminology, and an acute scarcity of high-quality annotated training data.

The main objective of the article is to establish MatSci-NLP, a standardized natural language benchmark for materials science, and demonstrate how domain-specific pretraining combined with a unified "text-to-schema" fine-tuning framework improves model performance in data-scarce environments.

To evaluate this, the researchers unified several publicly available materials science datasets into a seven-task benchmark spanning conventional tasks (such as named entity recognition, relation extraction, and event argument extraction) and domain-specific tasks (such as synthesis action retrieval and experimental slot filling). They tested several specialized language models (including MatBERT, MatSciBERT, SciBERT, BatteryBERT, BioBERT, and ScholarBERT) alongside a general-domain baseline (BERT). To replicate realistic data-scarce conditions, the models were fine-tuned on just 1% of the training data and tested on the remaining 99%, comparing single-task learning, multitask Mixture-of-Experts, and several question-answering-inspired input schemas.

The findings show that domain-specific pretraining significantly boosts downstream performance, with MatBERT achieving the strongest overall performance (overall micro-F1 of 0.722 vs. 0.658 for general BERT). Pretraining on broad scientific literature (such as SciBERT at 0.685 micro-F1) also substantially outperformed general-purpose BERT, indicating that scientific text across disciplines shares vocabulary distributions that standard language models struggle to capture. Furthermore, the proposed structured Task-Schema method consistently surpassed standard fine-tuning approaches across all evaluated models, lifting average micro-F1 from 0.493 in single-task learning to 0.688 in the unified schema format.

These results demonstrate that organizations can deploy higher-performing materials intelligence tools without the prohibitive expense of creating massive labeled training datasets. Adopting structured input schemas and domain-adapted base models directly reduces data annotation overhead and model training complexity. When selecting or developing models, the quality and curation of pretraining data matter greatly; not all scientific models perform equally well, as evidenced by general ScholarBERT lagging behind general BERT.

Organizations developing NLP tools for scientific discovery should leverage specialized pretraining models like MatBERT or SciBERT rather than standard off-the-shelf general language models. Teams should also adopt unified, structured text-to-schema input formatting for fine-tuning rather than separate single-task pipelines. Because the underlying benchmark datasets contain significant class imbalance—which skews basic accuracy metrics—practitioners should implement weighted loss functions (such as focal loss) or balanced data samplers during model deployment.

Decision-makers should note that the study evaluated BERT-scale encoder models on low-resource splits and did not analyze massive modern autoregressive large language models. While the text-to-schema methodology is modular and transferable to adjacent scientific domains like chemistry and biology, additional validation on broader datasets and larger model families is recommended before large-scale production deployment.

No sufficiently relevant recommendations were found.

Cover for MatSci-NLP: Evaluating Scientific Language Models on Materials Science Language Tasks Using Text-to-Schema Modeling

Abstract

We present MatSci-NLP, a natural language benchmark for evaluating the performance of natural language processing (NLP) models on materials science text. We construct the benchmark from publicly available materials science text data to encompass seven different NLP tasks, including conventional NLP tasks like named entity recognition and relation classification, as well as NLP tasks specific to materials science, such as synthesis action retrieval which relates to creating synthesis procedures for materials. We study various BERT-based models pretrained on different scientific text corpora on MatSci-NLP to understand the impact of pre-training strategies on understanding materials science text. Given the scarcity of high-quality annotated data in the materials science domain, we perform our fine-tuning experiments with limited training data to encourage the generalize across MatSci-NLP tasks. Our experiments in this low-resource training setting show that language models pretrained on scientific text outperform BERT trained on general text. Mat-BERT, a model pretrained specifically on materials science journals, generally performs best for most tasks. Moreover, we propose a unified text-to-schema for multitask learning on MatSci-NLP and compare its performance with traditional fine-tuning methods. In our analysis of different training methods, we find that our proposed text-to-schema methods inspired by question-answering consistently outperform single and multitask NLP fine-tuning methods. The code and datasets are publicly available¹.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 2.1 Scientific Language Models
  • 2.2 NLP in Materials Science
  • 3 MatSci-NLP Benchmark
  • 4 Unified Text-to-Schema Language Modeling
  • 4.1 Language Model Formulation
  • 4.2 Text-To-Schema Modeling
  • 4.3 Language Decoding & Evaluation
  • 5 Evaluation and Results
  • 5.1 How does in-domain pretraining of language models affect the downstream performance on MatSci-NLP tasks? (Q1)
  • 5.2 How do in-context data schema and multitasking affect the learning efficiency in low-resource training settings? (Q2)
  • 6 Conclusion and Future Works
  • Limitations
  • Broader Impacts and Ethics Statement
  • Acknowlegments
  • References
  • Appendix
  • A Experimental Details
  • B Additional Text-to-Schema Experiments

Knowls

  1. Knowl 1 — MatSci-NLP benchmark composition

    data/table

    MatSci-NLP combines publicly available annotated materials-science text datasets into a seven-task benchmark. It contains 169,197 samples in total; the component count records how many source datasets were combined for each task. The benchmark data uses a common JSON-based format containing text, task definitions, and annotations, which can be reformatted for different model inputs.

    TaskSamplesSource components
    Named Entity Recognition112,1914
    Relation Classification25,6743
    Event Argument Extraction6,5662
    Paragraph Classification1,5001
    Synthesis Action Retrieval5,5471
    Sentence Classification9,4661
    Slot Filling8,2531

    The benchmark spans material categories and applications including fuel cells, glasses, inorganic materials, superconductors, and synthesis procedures. The authors made the benchmark datasets and code publicly available.

  2. Knowl 2 — Operational definitions of the seven benchmark tasks

    definition

    MatSci-NLP evaluates seven tasks on materials-science text. Named Entity Recognition assigns entity types to text spans, with a null label for non-entities; entity types include materials, descriptors, properties, and applications. Relation Classification predicts a relation type for a pair of spans. Event Argument Extraction identifies arguments and their roles for a specified event trigger, allowing for multiple events in one text. Paragraph Classification determines whether a paragraph concerns glass science. Synthesis Action Retrieval classifies tokens into eight predefined synthesis-action categories. Sentence Classification identifies sentences describing relevant experimental facts. Slot Filling extracts predefined semantic slot values from a sentence describing an experimental frame.

  3. Knowl 3 — Unified text-to-schema input and output representation

    model/method

    The unified text-to-schema method reformulates heterogeneous materials-science NLP tasks as structured input-output prediction. Each model input can combine four elements: the source text; a description naming the task and its arguments; task instructions, potentially including answer choices or an example; and a predefined output schema. The output schema is task-specific but drawn from a common structured format: for example, an entity is represented by an entity name and type, a relation by its type and two entities, and an event by its trigger and role-labelled arguments. Classification outputs such as paragraph labels and synthesis actions also use predefined schemas. Because one text can support labels for several tasks, it can be supplied with different task descriptions and schemas for multitask training. The method is designed to let tasks share a unified modeling format while keeping outputs structured and evaluable.

  4. Knowl 4 — Modular encoder-decoder model for structured multitask prediction

    model/method

    The text-to-schema model pairs a domain-specific BERT encoder, such as MatBERT, MatSciBERT, or SciBERT, with a general transformer decoder. The encoder can be exchanged independently of the decoder. The decoder uses masked self-attention so that output positions cannot attend to future output tokens, and encoder-decoder attention to condition generation on the input text. It combines the self-attention and cross-attention representations to predict the next output token. The same architecture is fine-tuned across MatSci-NLP tasks, with the task schema structuring each prediction.

  5. Knowl 5 — Schema-constrained decoding and classification loss

    algorithm

    To evaluate and fine-tune text-to-schema predictions, the method first filters generated outputs against the predefined answer schema. It then matches each remaining prediction to the most similar valid class among the task's annotated labels; the matched class is treated as the final prediction. For example, an NER output that combines a span with the word “materials” can be matched to the valid entity-type class “material.” Cross-entropy loss is computed using the matched class. This procedure turns structured generation into classification over the task's valid labels without requiring the generated text to exactly reproduce a label.

  6. Knowl 6 — Low-resource evaluation and fine-tuning protocol

    experimental setup

    The experiments evaluate BERT-based encoders under a low-resource split: 1% of each MatSci-NLP dataset is used for training and the remaining 99% for testing. The evaluated encoders had not previously been exposed to the fine-tuning data. The authors compare micro-F1 and macro-F1 and report means over five runs, with uncertainty shown as two standard deviations. Fine-tuning used Adam, a learning rate of 2×10−52\times10^{-5}, a maximum of 20 epochs, and early stopping; encoder hidden size was 768 except for ScholarBERT, which used 1024. Experiments were run on a single GPU.

  7. Knowl 7 — Scientific and materials-specific pretraining performance

    data/table

    Under the unified Task-Schema setting, the table compares the aggregate performance of encoders with different pretraining corpora across all seven tasks. Each score is the mean over five runs, reported as mean ± two standard deviations. MatBERT, pretrained on materials-science journals, has the highest overall micro-F1 and macro-F1. SciBERT, pretrained on scientific text, also exceeds general-language BERT on both aggregate metrics. The results support the paper's finding that scientific-text pretraining generally helps on this benchmark, while showing that the pretraining corpus matters: MatSciBERT and BatteryBERT do not outperform MatBERT, and ScholarBERT has the lowest aggregate scores.

    EncoderOverall micro-F1Overall macro-F1
    MatSciBERT0.671 ± 0.0600.456 ± 0.042
    MatBERT0.722 ± 0.0230.517 ± 0.041
    BatteryBERT0.663 ± 0.0380.456 ± 0.048
    SciBERT0.685 ± 0.0560.460 ± 0.044
    ScholarBERT0.468 ± 0.0280.276 ± 0.024
    BioBERT0.670 ± 0.0610.442 ± 0.057
    BERT0.658 ± 0.0300.439 ± 0.021

    MatBERT is not the top encoder on every individual task: SciBERT has the highest relation-classification micro-F1 (0.819 ± 0.067), ScholarBERT has the highest event-argument-extraction micro-F1 (0.489 ± 0.083), and BioBERT has the highest sentence-classification micro-F1 (0.915 ± 0.021).

  8. Knowl 8 — Aggregate comparison of training schemas

    data/table

    This comparison averages performance across the seven MatSci-NLP tasks and the evaluated encoders. The four question-answering-inspired text-to-schema settings provide task descriptions with different amounts or forms of guidance: No Explanations supplies the description only; Potential Choices adds candidate labels; Examples adds an input-output example; and Task-Schema supplies the predefined answer structure. The conventional baselines are separate single-task training, single-task prompting, and MMOE multitask learning. Task-Schema obtains the highest aggregate micro-F1, while all four question-answering-inspired settings have higher aggregate micro- and macro-F1 than the three conventional settings.

    Training settingMicro-F1Macro-F1
    Single Task0.493 ± 0.0640.288 ± 0.063
    Single Task Prompt0.486 ± 0.0620.246 ± 0.032
    MMOE0.439 ± 0.0030.212 ± 0.022
    No Explanations0.644 ± 0.0340.430 ± 0.049
    Potential Choices0.622 ± 0.0350.402 ± 0.049
    Examples0.639 ± 0.0440.410 ± 0.043
    Task-Schema0.688 ± 0.0460.435 ± 0.039

    Scores are means over five runs, with ± two standard deviations. The Task-Schema aggregate exceeds the strongest conventional baseline, Single Task, by 0.195 micro-F1 and 0.147 macro-F1.

  9. Knowl 9 — Class imbalance affects benchmark interpretation

    empirical result

    The authors report that micro-F1 is substantially higher than macro-F1 across MatSci-NLP tasks, indicating class imbalance and showing that micro-F1 alone can overstate performance on less frequent classes. In paragraph classification, 492 of 1,500 samples are positive examples. In sentence classification, one label accounts for 876 of 9,466 samples, approximately 10%. The paper cautions that high imbalance can make evaluation misleading even when it reflects some practical information-extraction settings. As possible mitigations, the authors suggest weighting losses (including focal loss), using class-balanced samplers, and adjusting model architecture or regularization to emphasize minority classes.

  10. Knowl 10 — Scope and limitations of the study

    limitation

    The study's conclusions are limited by the small quantity and scope of available annotated materials-science data; the low-resource experiments use only 1% of datasets that are already limited in size. Small datasets may allow models to memorize answers rather than learn broader patterns. The evaluation covers materials science only, so transfer to related fields such as chemistry or physics was not tested. The model comparison is restricted to BERT-based encoders and does not evaluate autoregressive large language models. The authors also note that pretraining customized models has financial and environmental costs, and recommend building on existing large models where possible.

Coverage note — Detailed per-task and per-schema score tables are omitted because they are extensive; the aggregate comparisons and the clearest task-specific encoder exceptions preserve the main reported findings.

References

  1. 1.Georgios Balikas, Anastasia Krithara, Ioannis Partalas, and George Paliouras. 2015. Bioasq: A challenge on large-scale biomedical semantic indexing and question answering. In International Workshop on Multimodal Retrieval in the Medical Domain, pages 26–39. Springer.
  2. 2.Iz Beltagy, Kyle Lo, and Arman Cohan. 2019. Scibert: A pretrained language model for scientific text. arXiv preprint arXiv:1903.10676.
  3. 3.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  4. 4.Kamal Choudhary, Brian DeCost, Chi Chen, Anubhav Jain, Francesca Tavazza, Ryan Cohn, Cheol Woo Park, Alok Choudhary, Ankit Agrawal, Simon JL Billinge, et al. 2022. Recent advances and applications of deep learning methods in materials science. npj Computational Materials, 8(1):1–26.
  5. 5.Pradeep Dasigi, Kyle Lo, Iz Beltagy, Arman Cohan, Noah A Smith, and Matt Gardner. 2021. A dataset of information-seeking questions and answers anchored in research papers. arXiv preprint arXiv:2105.03011.
  6. 6.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805.
  7. 7.Annemarie Friedrich, Heike Adel, Federico Tomazic, Johannes Hingerl, Renou Benteau, Anika Maruscyk, and Lukas Lange. 2020. The sofc-exp corpus and neural approaches to information extraction in the materials science domain. arXiv preprint arXiv:2006.03039.
  8. 8.Alexandru B Georgescu, Peiwen Ren, Aubrey R Toland, Shengtong Zhang, Kyle D Miller, Daniel W Apley, Elsa A Olivetti, Nicholas Wagner, and James M Rondinelli. 2021. Database, features, and machine learning model to identify thermally driven metal–insulator transition compounds. Chemistry of Materials, 33(14):5591–5605.
  9. 9.Yu Gu, Robert Tinn, Hao Cheng, Michael Lucas, Naoto Usuyama, Xiaodong Liu, Tristan Naumann, Jianfeng Gao, and Hoifung Poon. 2021. Domain-specific language model pretraining for biomedical natural language processing. ACM Transactions on Computing for Healthcare (HEALTH), 3(1):1–23.
  10. 10.Tanishq Gupta, Mohd Zaki, NM Krishnan, et al. 2022. Matscibert: A materials domain language model for text mining and information extraction. npj Computational Materials, 8(1):1–11.
  11. 11.Kai Hakala and Sampo Pyysalo. 2019. Biomedical named entity recognition with multilingual bert. In Proceedings of the 5th workshop on BioNLP open shared tasks, pages 56–61.
  12. 12.Zhi Hong, Aswathy Ajith, Gregory Pauloski, Eamon Duede, Carl Malamud, Roger Magoulas, Kyle Chard, and Ian Foster. 2022. Scholarbert: Bigger is not always better. arXiv preprint arXiv:2205.11342.
  13. 13.Shu Huang and Jacqueline M Cole. 2022. Batterybert: A pretrained language model for battery database enhancement. Journal of Chemical Information and Modeling.
  14. 14.Zach Jensen, Soonhyoung Kwon, Daniel Schwalbe-Koda, Cecilia Paris, Rafael Gómez-Bombarelli, Yuriy Román-Leshkov, Avelino Corma, Manuel Moliner, and Elsa A Olivetti. 2021. Discovering relationships between osdas and zeolites through data mining and generative neural networks. ACS central science, 7(5):858–867.
  15. 15.Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William W Cohen, and Xinghua Lu. 2019. Pubmedqa: A dataset for biomedical research question answering. arXiv preprint arXiv:1909.06146.
  16. 16.Christopher Karpovich, Zach Jensen, Vineeth Venugopal, and Elsa Olivetti. 2021. Inorganic synthesis reaction condition prediction with generative machine learning. arXiv preprint arXiv:2112.09612.
  17. 17.Edward Kim, Zach Jensen, Alexander van Grootel, Kevin Huang, Matthew Staib, Sheshera Mysore, Haw-Shiuan Chang, Emma Strubell, Andrew McCallum, Stefanie Jegelka, et al. 2020. Inorganic materials synthesis planning with literature-trained neural networks. Journal of chemical information and modeling, 60(3):1194–1201.
  18. 18.Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980.
  19. 19.Olga Kononova, Tanjin He, Haoyan Huo, Amalie Trewartha, Elsa A Olivetti, and Gerbrand Ceder. 2021. Opportunities and challenges of text mining in materials research. Iscience, 24(3):102155.
  20. 20.Fusataka Kuniyoshi, Kohei Makino, Jun Ozawa, and Makoto Miwa. 2020. Annotating and extracting synthesis process of all-solid-state batteries from scientific literature. arXiv preprint arXiv:2002.07339.
  21. 21.Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. 2020. Biobert: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36(4):1234–1240.
  22. 22.Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár. 2017. Focal loss for dense object detection. In Proceedings of the IEEE international conference on computer vision, pages 2980–2988.
  23. 23.Yaojie Lu, Hongyu Lin, Jin Xu, Xianpei Han, Jialong Tang, Annan Li, Le Sun, Meng Liao, and Shaoyi Chen. 2021. Text2event: Controllable sequence-to-structure generation for end-to-end event extraction. arXiv preprint arXiv:2106.09232.
  24. 24.Minh-Thang Luong, Quoc V Le, Ilya Sutskever, Oriol Vinyals, and Lukasz Kaiser. 2015. Multi-task sequence to sequence learning. arXiv preprint arXiv:1511.06114.
  25. 25.Jiaqi Ma, Zhe Zhao, Xinyang Yi, Jilin Chen, Lichan Hong, and Ed H Chi. 2018. Modeling task relationships in multi-task learning with multi-gate mixture-of-experts. In Proceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining, pages 1930–1939.
  26. 26.Rubayyat Mahbub, Kevin Huang, Zach Jensen, Zachary D Hood, Jennifer LM Rupp, and Elsa A Olivetti. 2020. Text mining for processing conditions of solid-state battery electrolytes. Electrochemistry Communications, 121:106860.
  27. 27.MatSciRE. 2022. Material science relation extraction (matscire).
  28. 28.Santiago Miret, Marta Skreta, Benjamin Sanchez-Lengelin, Shyue Ping Ong, Zamyla Morgan-Chan, and Alan Aspuru-Guzik. Ai4mat - neurips 2022.
  29. 29.Sheshera Mysore, Zach Jensen, Edward Kim, Kevin Huang, Haw-Shiuan Chang, Emma Strubell, Jeffrey Flanigan, Andrew McCallum, and Elsa Olivetti. 2019. The materials science procedural text corpus: Annotating materials synthesis procedures with shallow semantic structures. arXiv preprint arXiv:1905.06939.
  30. 30.Elsa A Olivetti, Jacqueline M Cole, Edward Kim, Olga Kononova, Gerbrand Ceder, Thomas Yong-Jin Han, and Anna M Hiszpanski. 2020. Data-driven materials research enabled by natural language processing and information extraction. Applied Physics Reviews, 7(4):041317.
  31. 31.Yifan Peng, Shankai Yan, and Zhiyong Lu. 2019. Transfer learning in biomedical natural language processing: An evaluation of bert and elmo on ten benchmarking datasets. In Proceedings of the 2019 Workshop on Biomedical Natural Language Processing (BioNLP 2019).
  32. 32.Long N Phan, James T Anibal, Hieu Tran, Shaurya Chanana, Erol Bahadroglu, Alec Peltekian, and Grégoire Altan-Bonnet. 2021. Scifive: a text-to-text transformer model for biomedical literature. arXiv preprint arXiv:2106.03598.
  33. 33.Ghanshyam Pilania. 2021. Machine learning in materials science: From explainable predictions to autonomous design. Computational Materials Science, 193:110360.
  34. 34.Chen Qu, Liu Yang, Minghui Qiu, W Bruce Croft, Yongfeng Zhang, and Mohit Iyyer. 2019. Bert with history answer embedding for conversational question answering. In Proceedings of the 42nd international ACM SIGIR conference on research and development in information retrieval, pages 1133–1136.
  35. 35.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21(140):1–67.
  36. 36.Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilic, Daniel Hesslow, Roman ´ Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2022. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100.
  37. 37.Hoo-Chang Shin, Yang Zhang, Evelina Bakhturina, Raul Puri, Mostofa Patwary, Mohammad Shoeybi, and Raghav Mani. 2020. Biomegatron: Larger biomedical domain language model. arXiv preprint arXiv:2010.06060.
  38. 38.Ross Taylor, Marcin Kardas, Guillem Cucurull, Thomas Scialom, Anthony Hartshorn, Elvis Saravia, Andrew Poulton, Viktor Kerkez, and Robert Stojnic. 2022. Galactica: A large language model for science. arXiv preprint arXiv:2211.09085.
  39. 39.Minh Van Nguyen, Bonan Min, Franck Dernoncourt, and Thien Nguyen. 2022. Joint extraction of entities, relations, and events via modeling inter-instance and inter-label dependencies. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4363–4374.
  40. 40.Vineeth Venugopal, Sourav Sahoo, Mohd Zaki, Manish Agarwal, Nitya Nand Gosvami, and NM Anoop Krishnan. 2021. Looking through glass: Knowledge discovery from materials science literature using natural language processing. Patterns, 2(7):100290.
  41. 41.Shoya Wada, Toshihiro Takeda, Shiro Manabe, Shozo Konishi, Jun Kamohara, and Yasushi Matsumura. 2020. A pre-training technique to localize medical bert and enhance biobert.
  42. 42.Nicholas Walker, Amalie Trewartha, Haoyan Huo, Sanghoon Lee, Kevin Cruse, John Dagdelen, Alexander Dunn, Kristin Persson, Gerbrand Ceder, and Anubhav Jain. 2021. The impact of domain-specific pre-training on named entity recognition tasks in materials science. Available at SSRN 3950755.
  43. 43.Zheren Wang, Kevin Cruse, Yuxing Fei, Ann Chia, Yan Zeng, Haoyan Huo, Tanjin He, Bowen Deng, Olga Kononova, and Gerbrand Ceder. 2022a. Ulsa: Unified language of synthesis actions for the representation of inorganic synthesis protocols. Digital Discovery.
  44. 44.Zheren Wang, Olga Kononova, Kevin Cruse, Tanjin He, Haoyan Huo, Yuxing Fei, Yan Zeng, Yingzhi Sun, Zijian Cai, Wenhao Sun, et al. 2022b. Dataset of solution-based inorganic materials synthesis procedures extracted from the scientific literature. Scientific Data, 9(1):1–11.
  45. 45.Leigh Weston, Vahe Tshitoyan, John Dagdelen, Olga Kononova, Amalie Trewartha, Kristin A Persson, Gerbrand Ceder, and Anubhav Jain. 2019. Named entity recognition and normalization applied to large-scale information extraction from the materials science literature. Journal of chemical information and modeling, 59(9):3692–3702.
  46. 46.Shanchan Wu and Yifan He. 2019. Enriching pre-trained language model with entity information for relation classification. In Proceedings of the 28th ACM international conference on information and knowledge management, pages 2361–2364.
  47. 47.Kyosuke Yamaguchi, Ryoji Asahi, and Yutaka Sasaki. 2020. Sc-comics: a superconductivity corpus for materials informatics. In Proceedings of The 12th Language Resources and Evaluation Conference, pages 6753–6760.

Citation

MLA
Song, Y., et al. “MatSci-NLP: Evaluating Scientific Language Models on Materials Science Language Tasks Using Text-to-Schema Modeling”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 3621–39, https://doi.org/10.18653/v1/2023.acl-long.201.
APA
Song, Y., Miret, S., & Liu, B. (2023). MatSci-NLP: Evaluating Scientific Language Models on Materials Science Language Tasks Using Text-to-Schema Modeling. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3621–3639. https://doi.org/10.18653/v1/2023.acl-long.201
Chicago
Song, Y., S. Miret, and B. Liu. 2023. “MatSci-NLP: Evaluating Scientific Language Models on Materials Science Language Tasks Using Text-to-Schema Modeling”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3621–39. https://doi.org/10.18653/v1/2023.acl-long.201.
Harvard
Song, Y., Miret, S. and Liu, B. (2023) “MatSci-NLP: Evaluating Scientific Language Models on Materials Science Language Tasks Using Text-to-Schema Modeling”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 3621–3639. Available at: https://doi.org/10.18653/v1/2023.acl-long.201.
Vancouver
1. Song Y, Miret S, Liu B (2023) MatSci-NLP: Evaluating Scientific Language Models on Materials Science Language Tasks Using Text-to-Schema Modeling. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 3621–3639

BibTeX

@inproceedings{song-etal-2023-matsci,
    title = "{M}at{S}ci-{NLP}: Evaluating Scientific Language Models on Materials Science Language Tasks Using Text-to-Schema Modeling",
    author = "Song, Yu  and
      Miret, Santiago  and
      Liu, Bang",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.201/",
    doi = "10.18653/v1/2023.acl-long.201",
    pages = "3621--3639"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/