OceanGPT: A Large Language Model for Ocean Science Tasks

Zhen BiNingyu ZhangYida XueYixin OuDaxiong JiGuozhou ZhengHuajun Chen

article2024ACL75 citations

Presents OceanGPT, the first domain-specific large language model for oceanography, trained using a multi-agent instruction generation framework and evaluated on a new dedicated benchmark to handle specialized scientific tasks and ocean robotics commands.

Listen

Oceans cover over 70% of the Earth's surface and play a critical role in global climate regulation, biodiversity, and economic activity. While general-purpose large language models have advanced scientific research across various disciplines, they frequently fail to meet the complex demands of oceanography. This shortfall stems from the vast, intricate nature of marine data and the need for deep, specialized knowledge across diverse oceanic subfields.

To address this gap, the article develops and evaluates OceanGPT, the first dedicated large language model for ocean science tasks. The initiative aimed to create a domain-specific model capable of answering complex oceanographic queries, executing scientific language tasks, and generating actionable commands for marine engineering systems.

To build the system, researchers constructed a specialized training corpus from 67,633 open-access ocean science documents. They then developed DoInstruct, an automated multi-agent framework that generated more than 150,000 instruction tuning samples across five major oceanic topics: science and research, resources and development, ecology and environment, technology and engineering, and culture. The framework utilized specialized agents acting as evolving generators, literature extractors, and inspectors with rule-based quality controls. To evaluate performance, the researchers introduced OceanBench, an oceanography benchmark spanning 15 distinct tasks, utilizing both automated evaluations calibrated by GPT-4 and human evaluations by marine science experts.

Key findings show that OceanGPT significantly outperformed leading general-purpose open-source models, including LLaMA-2-7B-Chat, Vicuna-1.5-7B, and ChatGLM2-6B, winning the majority of tested sub-tasks in both automated and human evaluations. In task-level human evaluations, OceanGPT won 12 to 14 out of 15 tasks against the baseline models. The model demonstrated superior domain expertise in complex scientific workflows, such as radioactive nuclide research, providing practical experimental and risk assessment steps where baseline models offered only generic responses. Furthermore, OceanGPT demonstrated preliminary embodied intelligence for ocean engineering, successfully generating control code and console commands to direct underwater robots in physical simulation environments.

These findings indicate that domain-adapted language models can substantially improve analytical accuracy, technical problem-solving, and operational planning in maritime industries and scientific research. By integrating domain-specific literature and multi-agent synthetic data generation, organizations can reduce the high costs and labor bottlenecks typically required to train specialized artificial intelligence systems.

Decision-makers and research teams should consider adopting specialized multi-agent instruction generation frameworks to scale domain-specific models cost-effectively. For marine technology deployment, organizations should conduct further pilot testing of OceanGPT’s robotic code generation within real-world maritime systems before relying on it for high-stakes operational planning.

Confidence in the model's domain expertise is supported by strong agreement among expert evaluators (an inter-annotator agreement score of 0.82) and consistent benchmark performance. However, users must remain cautious regarding standard language model limitations present in the system, including potential data distribution biases from public literature and occasional factual hallucinations.

arXiv: 2310.02031
Cover for OceanGPT: A Large Language Model for Ocean Science Tasks

Abstract

Ocean science, which delves into the oceans that are reservoirs of life and biodiversity, is of great significance given that oceans cover over 70% of our planet's surface. Recently, advances in Large Language Models (LLMs) have transformed the paradigm in science. Despite the success in other domains, current LLMs often fall short in catering to the needs of domain experts like oceanographers, and the potential of LLMs for ocean science is under-explored. The intrinsic reasons are the immense and intricate nature of ocean data as well as the necessity for higher granularity and richness in knowledge. To alleviate these issues, we introduce OCEANGPT, the first-ever large language model in the ocean domain, which is expert in various ocean science tasks. We also propose DOINSTRUCT, a novel framework to automatically obtain a large volume of ocean domain instruction data, which generates instructions based on multi-agent collaboration. Additionally, we construct the first oceanography benchmark, OCEANBENCH, to evaluate the capabilities of LLMs in the ocean domain. Though comprehensive experiments, OCEANGPT not only shows a higher level of knowledge expertise for oceans science tasks but also gains preliminary embodied intelligence capabilities in ocean technology.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 OCEANGPT
  • 3.1 Pre-training Stage
  • 3.2 Domain Instruction Data Generation
  • 4 Benchmarking Ocean Science Tasks
  • 4.1 Implementation Details and Baselines
  • 5 Results
  • 5.1 Insights from Performance Results
  • 5.2 Exploring the Potential of OceanGPT
  • 5.2.1 OceanGPT for Ocean Science
  • 5.2.2 OceanGPT for Ocean Engineering
  • 6 Conclusion
  • Limitations
  • Acknowledgments
  • References
  • A Appendix
  • The Cost for Fine-tuning GPT-3.5-Turbo
  • Comparison between Our Fine-tuning Method and the Prefix Prompts
  • The Similarity Calculating Method in the Deduplication Procedure
  • Prompt for "Fine-Tuned Agent as the Literature Extractor":
  • Prompt for "Evolving Agent as the Generator":
  • Prompt for 'Agent as the Inspector with Rule Constraints':
  • Prompt for automatic evaluation using GPT4:

Knowls

  1. Knowl 1 — DoInstruct Multi-Agent Instruction Generation Algorithm

    algorithm

    The DoInstruct framework automates the generation of domain-specific instruction datasets for oceanography tasks by coordinating multiple large language model agents acting as specialized domain experts.

    Input: Seed instruction dataset S={(inst,output)}S = \{(\text{inst}, \text{output})\}, Ocean literature corpus OO, Pre-defined rule constraints RR
    Output: Filtered high-quality instruction dataset HH
    Initialize empty sets: Step1Data←∅\text{Step1Data} \leftarrow \emptyset, Step2Data←∅\text{Step2Data} \leftarrow \emptyset, H←∅H \leftarrow \emptyset
    # Stage 1: Seed Evolution via Generator Agent
    for each sample (inst,output)∈S(\text{inst}, \text{output}) \in S do
        enriched_sample ←Enrich(inst,output)\leftarrow \text{Enrich}(\text{inst}, \text{output})
        refined_sample ←Refine(inst,output)\leftarrow \text{Refine}(\text{inst}, \text{output})
        \text{Step1Data} \leftarrow \text{Step1Data} \cup \{\text{enriched_sample}, \text{refined_sample}\}
    end for
    # Stage 2: Literature Extraction via Fine-Tuned Agent
    RetrievedTexts←BM25_Retrieve(O)\text{RetrievedTexts} \leftarrow \text{BM25\_Retrieve}(O)
    M←FineTune(gpt-3.5-turbo,Sreverse)M \leftarrow \text{FineTune}(\text{gpt-3.5-turbo}, S_{\text{reverse}})
    for each document ∈RetrievedTexts\in \text{RetrievedTexts} do
        output←document.content\text{output} \leftarrow \text{document.content}
        inst←M(output)\text{inst} \leftarrow M(\text{output})
        Step2Data←Step2Data∪{(inst,output)}\text{Step2Data} \leftarrow \text{Step2Data} \cup \{(\text{inst}, \text{output})\}
    end for
    # Stage 3: Constraint Filtering via Inspector Agent
    MergedData←Inspector(Step1Data∪Step2Data,R)\text{MergedData} \leftarrow \text{Inspector}(\text{Step1Data} \cup \text{Step2Data}, R)
    # Stage 4: Multi-Agent Debate Quality Control
    for each sample (inst,output)∈MergedData(\text{inst}, \text{output}) \in \text{MergedData} do
        debate_result ←Debate(inst,output)\leftarrow \text{Debate}(\text{inst}, \text{output})
        if debate_result is high-quality then
            H←H∪{(inst,output)}H \leftarrow H \cup \{(\text{inst}, \text{output})\}
        end if
    end for
    return HH

    In Stage 1, two gpt-3.5-turbo agents perform iterative evolution by adding background knowledge and analyzing entities in depth. In Stage 2, gpt-3.5-turbo is fine-tuned on inverted seed pairs Sreverse={(output,inst)}S_{\text{reverse}} = \{(\text{output}, \text{inst})\} to generate questions given retrieved unstructured text chunks. In Stage 3, an inspector agent filters pairs against natural language constraints RR (publication year, domain keywords, literature type, geographic regions). In Stage 4, two additional agents debate each instance to finalize inclusion in HH, yielding over 150,000 synthetic ocean instructions.

  2. Knowl 2 — OceanGPT Model Architecture and Training Pipeline

    model/method

    OceanGPT is an ocean-domain large language model built on the LLaMA-2 base architecture via a two-stage training procedure:

    1. Ocean Pre-training Stage: The foundational model is trained on a raw domain corpus of 67,633 open-access scientific publications across physical oceanography, marine chemistry, marine biology, marine geology, and hydrology. Text is extracted from PDF files using pdfminer and cleaned with regular expressions to strip headers, footers, page numbers, URLs, references, tables, and garbled symbols. Deduplication is performed using hash-based matching to prevent memorization and over-fitting.

    2. Instruction-Tuning Stage: OceanGPT is fine-tuned on instructions generated by the DoInstruct framework using Low-Rank Adaptation (LoRA).

    Key hyperparameters for the LoRA instruction tuning are:

    • Base Model: LLaMA-2 (7B parameters; with releases also adapted to MiniCPM-2B and Meta-Llama-3-8B-Instruct)
    • LoRA Rank (rr): 8
    • LoRA Scaling Factor (α\alpha): 16
    • LoRA Dropout: 0.05
    • Batch Size: 512
    • Learning Rate: 1×10−41\times 10^{-4}
    • Training Epochs: 10
    • Compute Hardware: 6 ×\times NVIDIA A800 GPUs
  3. Knowl 3 — Ocean Science Topic Categorization

    definition

    The domain instruction data and benchmark for ocean science tasks are structured around five major, mutually exclusive top-level topics comprising over 500 fine-grained sub-categories:

    1. Science and research: Fundamental oceanographic theories and phenomena, including ocean currents, sea surface temperature, hydrology, and marine biodiversity.
    2. Resources and development: Rational utilization and exploration of marine resources, including fisheries, seabed minerals, offshore oil and gas, and renewable seawater resources.
    3. Ecology and environment: Environmental protection, ecological sustainability, ocean acidification, marine pollution, ecological degradation, and ocean-climate interactions.
    4. Technology and engineering: Marine observational instrumentation, ship engineering, offshore structures, deep-sea exploration technologies, and ocean energy generation systems.
    5. Life, culture and others: Historical and sociocultural interactions with the ocean, including maritime history, marine policy, coastal tourism, and maritime leisure.
  4. Knowl 4 — OceanBench Benchmark Dataset

    data/table

    OceanBench is a domain-specific evaluation benchmark containing 10,446 expert-verified test instances distributed across 15 distinct ocean science task categories.

    Task Num Task Num
    Analysis 674 Classification 895
    Judgment 655 Letter Writing 359
    Open-ended Generation 930 Extraction 1,078
    Recommendation 1,089 Description 1,246
    Summary 149 Editing 1,075
    Identification 464 Transformation 401
    Question Answering 1,230 Others 157
    Commonsense Reasoning 1,024
    Total 10,446

    Samples are generated from the validated seed dataset, filtered to remove bad cases via expert manual inspection, and deduplicated against the training dataset using keyword-hash comparison to prevent test set contamination.

  5. Knowl 5 — Pairwise Position-Balanced Automatic and Human Evaluation Protocol

    experimental setup

    To evaluate models on OceanBench while controlling for positional bias in LLM evaluators, a bidirectional pairwise evaluation protocol is used:

    • Automatic Evaluation: GPT-4 serves as the automated judge. For each test question and candidate pair (M1,M2)(M_1, M_2), GPT-4 is queried twice with swapped presentation orders: Prompt 1 evaluates (M1,M2)(M_1, M_2) and Prompt 2 evaluates (M2,M1)(M_2, M_1). The final win, tie, or loss count between M1M_1 and M2M_2 is the direct sum of the outcomes across both ordered evaluations.
    • Human Evaluation: Five annotators with academic backgrounds in ocean science independently rank model outputs on randomly sampled subsets of 200 instances per evaluation condition.
    • Task-level metric: A model wins a task if it outperforms the competitor on the majority of instances within that task category.
    • Instance-level metric: Overall win, tie, and loss percentages aggregated across all individual test instances without task weighting.
  6. Knowl 6 — OceanBench Performance of OceanGPT vs Baseline Open-Source LLMs

    empirical result

    Across 15 ocean science sub-tasks on OceanBench, OceanGPT (7B) achieves superior win rates compared to general-purpose open-source baselines under both automated (GPT-4) and expert human evaluations.

    Comparison Task-Level Wins (out of 15) Instance-Level Rates (%)
    Win Tie Loss Win (%) Tie (%) Loss (%)
    Automatic Evaluation (GPT-4)
    OceanGPT vs ChatGLM2-6B 15 0 0 84.0 2.0 14.0
    OceanGPT vs LLaMA-2-chat-7B 14 1 0 65.0 32.0 3.0
    OceanGPT vs Vicuna-1.5-7B 12 3 0 60.0 35.0 5.0
    Human Evaluation
    OceanGPT vs ChatGLM2-6B 14 1 0 - - -
    OceanGPT vs LLaMA-2-chat-7B 13 2 0 - - -
    OceanGPT vs Vicuna-1.5-7B 12 2 1 - - -

    OceanGPT outperforms baseline models across all individual task categories, showing especially large margins in domain-specific tasks requiring deep technical grounding, such as editing and technical analysis.

  7. Knowl 7 — Functional Role Specialization of DoInstruct Agent Types

    empirical result

    Human manual assessment scoring data on a 1-to-5 scale across Quality, Expertise, and Diversity reveals distinct contributions for each agent role in the DoInstruct multi-agent pipeline:

    • Evolving Generator Agent: Achieves the highest Diversity score (~4.7 / 5.0), with moderate Expertise (~3.0) and lower raw Quality (~2.1), serving primarily to expand data breadth and coverage.
    • Fine-Tuned Literature Extractor Agent: Produces the highest Expertise score (~4.5 / 5.0), with moderate Diversity (~3.0) and lower Quality (~2.2), ensuring technical precision derived directly from oceanographic literature.
    • Inspector Agent: Achieves the highest Quality score (~4.6 / 5.0) while exhibiting low independent Diversity (~1.8) and Expertise (~1.2), acting as a quality filter enforcing formatting, syntactic, and factual domain rules.
  8. Knowl 8 — Embodied Control Capabilities for Underwater Robotics

    model/method

    OceanGPT acquires preliminary embodied intelligence for ocean engineering by training on synthesized machine code and robotics command instructions. When prompted with natural language mission descriptions, OceanGPT outputs executable Robot Operating System (ROS) launch scripts and console command parameters for underwater vehicles simulated in Gazebo using the uuv_simulator environment.

    For example, given high-level navigation instructions, the model generates parameter-complete launch commands such as:

    roslaunch uuv_control_utils start_helical_trajectory.launch uuv_name:=rexrov n_turns:=2
    

    This enables the simulated underwater vehicle (RexROV) to execute multi-turn 3D helical path planning and trajectory tracking in an aquatic environment directly from text instructions.

  9. Knowl 9 — Inherent Biases and Hallucinations in Ocean Science LLMs

    limitation

    OceanGPT and related domain-specific large language models exhibit two primary operational limitations:

    1. Data Distribution Bias: The pre-training corpus and human-authored instruction seeds inherit geographic, linguistic, and socioeconomic biases from open-access scientific publications and contributors. This can cause the model to favor specific regional marine ecosystems (such as heavily studied coastal areas) and underrepresent others.
    2. Domain Hallucinations: OceanGPT can generate plausible-sounding but factually inaccurate or contextually contradictory statements regarding complex oceanographic dynamics, chemical reactions, and taxonomic identifications.

Coverage note — None was omitted; all key contributions including the pre-training pipeline, the DoInstruct framework and algorithm, the topic taxonomy, OceanBench specifications, empirical evaluation results, agent role analysis, embodied robotics extension, and limitations are fully covered.

References

  1. 1.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  2. 2.Zhuo Chen, Wen Zhang, Yufeng Huang, Mingyang Chen, Yuxia Geng, Hongtao Yu, Zhen Bi, Yichi Zhang, Zhen Yao, Wenting Song, Xinliang Wu, Yi Yang, Mingyi Chen, Zhaoyang Lian, Yingying Li, Lei Cheng, and Huajun Chen. 2023. Tele-knowledge pre-training for fault analysis.
  3. 3.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.
  4. 4.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, and et al. 2022. Palm: Scaling language modeling with pathways. CoRR, abs/2204.02311.
  5. 5.Cheng Deng, Tianhang Zhang, Zhongmou He, Qiyuan Chen, Yuanyuan Shi, Le Zhou, Luoyi Fu, Weinan Zhang, Xinbing Wang, Chenghu Zhou, Zhouhan Lin, and Junxian He. 2023. Learning A foundation language model for geoscience knowledge understanding and utilization. CoRR, abs/2306.05064.
  6. 6.Wayne E Esaias, Mark R Abbott, Ian Barton, Otis B Brown, Janet W Campbell, Kendall L Carder, Dennis K Clark, Robert H Evans, Frank E Hoge, Howard R Gordon, et al. 1998. An overview of modis capabilities for ocean science observations. IEEE Transactions on Geoscience and Remote Sensing, 36(4):1250–1265.
  7. 7.Paul Falkowski. 2012. Ocean science: the power of plankton. Nature, 483(7387):S17–S20.
  8. 8.Yin Fang, Qiang Zhang, Ningyu Zhang, Zhuo Chen, Xiang Zhuang, Xin Shao, Xiaohui Fan, and Huajun Chen. 2023. Knowledge graph-enhanced molecular contrastive learning with functional prompt. Nature Machine Intelligence, 5:1–12.
  9. 9.Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Qianyu Guo, Meng Wang, and Haofen Wang. 2023. Retrieval-augmented generation for large language models: A survey. CoRR, abs/2312.10997.
  10. 10.Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. Lora: Low-rank adaptation of large language models.
  11. 11.Jinhao Jiang, Kun Zhou, Zican Dong, Keming Ye, Wayne Xin Zhao, and Ji-Rong Wen. 2023. Structgpt: A general framework for large language model to reason over structured data. CoRR, abs/2305.09645.
  12. 12.Xuchen Jin, Xianqiang He, Difeng Wang, Jianyun Ying, Fang Gong, Qiankun Zhu, Chenghu Zhou, and Delu Pan. 2023. Impact of rain effects on l-band passive microwave satellite observations over the ocean. IEEE Trans. Geosci. Remote. Sens., 61:1–16.
  13. 13.Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2020. Generalization through memorization: Nearest neighbor language models. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  14. 14.Zeljko Kraljevic, Anthony Shek, Daniel Bean, Rebecca Bendayan, James T. Teo, and Richard J. B. Dobson. 2021. Medgpt: Medical concept prediction from clinical narratives. CoRR, abs/2107.03134.
  15. 15.Patrick S. H. Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-augmented generation for knowledge-intensive NLP tasks. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  16. 16.Chengshu Li, Jacky Liang, Andy Zeng, Xinyun Chen, Karol Hausman, Dorsa Sadigh, Sergey Levine, Li Fei-Fei, Fei Xia, and Brian Ichter. 2023. Chain of code: Reasoning with a language model-augmented code emulator.
  17. 17.Zeming Lin, Halil Akin, Roshan Rao, Brian Hie, Zhongkai Zhu, Wenting Lu, Nikita Smetanin, Robert Verkuil, Ori Kabeli, Yaniv Shmueli, Allan dos Santos Costa, Maryam Fazel-Zarandi, Tom Sercu, Salvatore Candido, and Alexander Rives. 2023. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science, 379(6637):1123–1130.
  18. 18.Renqian Luo, Liai Sun, Yingce Xia, Tao Qin, Sheng Zhang, Hoifung Poon, and Tie-Yan Liu. 2022. Biogpt: generative pre-trained transformer for biomedical text generation and mining. Briefings Bioinform., 23(6).
  19. 19.Musa Morena Marcusso Manhães, Sebastian A. Scherer, Martin Voss, Luiz Ricardo Douat, and Thomas Rauschenbach. 2016. UUV simulator: A gazebo-based package for underwater intervention and multi-robot simulation. In OCEANS 2016 MTS/IEEE Monterey. IEEE.
  20. 20.Michael Moor, Oishi Banerjee, Zahra Shakeri, Harlan Krumholz, Jure Leskovec, Eric Topol, and Pranav Rajpurkar. 2023. Foundation models for generalist medical artificial intelligence. Nature, 616:259–265.
  21. 21.OpenAI. 2023. Gpt-4 technical report.
  22. 22.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In NeurIPS.
  23. 23.Shuofei Qiao, Honghao Gui, Huajun Chen, and Ningyu Zhang. 2023a. Making language models better tool learners with execution feedback. CoRR, abs/2305.13068.
  24. 24.Shuofei Qiao, Yixin Ou, Ningyu Zhang, Xiang Chen, Yunzhi Yao, Shumin Deng, Chuanqi Tan, Fei Huang, and Huajun Chen. 2023b. Reasoning with language model prompting: A survey. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, pages 5368–5393. Association for Computational Linguistics.
  25. 25.Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, H. Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks, Maribeth Rauh, Po-Sen Huang, Amelia Glaese, Johannes Welbl, Sumanth Dathathri, Saffron Huang, Jonathan Uesato, John Mellor, Irina Higgins, Antonia Creswell, Nat McAleese, Amy Wu, Erich Elsen, Siddhant M. Jayakumar, Elena Buchatskaya, David Budden, Esme Sutherland, Karen Simonyan, Michela Paganini, Laurent Sifre, Lena Martens, Xiang Lorraine Li, Adhiguna Kuncoro, Aida Nematzadeh, Elena Gribovskaya, Domenic Donato, Angeliki Lazaridou, Arthur Mensch, Jean-Baptiste Lespiau, Maria Tsimpoukelli, Nikolai Grigorev, Doug Fritz, Thibault Sottiaux, Mantas Pajarskas, Toby Pohlen, Zhitao Gong, Daniel Toyama, Cyprien de Masson d’Autume, Yujia Li, Tayfun Terzi, Vladimir Mikulik, Igor Babuschkin, Aidan Clark, Diego de Las Casas, Aurelia Guy, Chris Jones, James Bradbury, Matthew J. Johnson, Blake A. Hechtman, Laura Weidinger, Iason Gabriel, William Isaac, Edward Lockhart, Simon Osindero, Laura Rimell, Chris Dyer, Oriol Vinyals, Kareem Ayoub, Jeff Stanway, Lorrayne Bennett, Demis Hassabis, Koray Kavukcuoglu, and Geoffrey Irving. 2021. Scaling language models: Methods, analysis & insights from training gopher. CoRR, abs/2112.11446.
  26. 26.Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilic, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, Jonathan Tow, Alexander M. Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanchi, Thomas Wang, Benoît Sagot, Niklas Muennighoff, Albert Villanova del Moral, Olatunji Ruwase, Rachel Bawden, Stas Bekman, Angelina McMillan-Major, Iz Beltagy, Huu Nguyen, Lucile Saulnier, Samson Tan, Pedro Ortiz Suarez, Victor Sanh, Hugo Laurençon, Yacine Jernite, Julien Launay, Margaret Mitchell, Colin Raffel, Aaron Gokaslan, Adi Simhi, Aitor Soroa, Alham Fikri Aji, Amit Alfassy, Anna Rogers, Ariel Kreisberg Nitzav, Canwen Xu, Chenghao Mou, Chris Emezue, Christopher Klamm, Colin Leong, Daniel van Strien, David Ifeoluwa Adelani, and et al. 2022. BLOOM: A 176b-parameter open-access multilingual language model. CoRR, abs/2211.05100.
  27. 27.Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023. Toolformer: Language models can teach themselves to use tools. CoRR, abs/2302.04761.
  28. 28.Karan Singhal, Shekoofeh Azizi, Tao Tu, S. Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Kumar Tanwani, Heather Cole-Lewis, Stephen Pfohl, Perry Payne, Martin Seneviratne, Paul Gamble, Chris Kelly, Nathaneal Schärli, Aakanksha Chowdhery, Philip Andrew Mansfield, Blaise Agüera y Arcas, Dale R. Webster, Gregory S. Corrado, Yossi Matias, Katherine Chou, Juraj Gottweis, Nenad Tomasev, Yun Liu, Alvin Rajkomar, Joelle K. Barral, Christopher Semturs, Alan Karthikesalingam, and Vivek Natarajan. 2022. Large language models encode clinical knowledge. CoRR, abs/2212.13138.
  29. 29.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca.
  30. 30.Christina Theodoris, Ling Xiao, Anant Chopra, Mark Chaffin, Zeina Sayed, Matthew Hill, Helene Mantineo, Elizabeth Brydon, Zexian Zeng, Shirley Liu, and Patrick Ellinor. 2023. Transfer learning enables predictions in network biology. Nature, 618:1–9.
  31. 31.Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, YaGuang Li, Hongrae Lee, Huaixiu Steven Zheng, Amin Ghafouri, Marcelo Menegali, Yanping Huang, Maxim Krikun, Dmitry Lepikhin, James Qin, Dehao Chen, Yuanzhong Xu, Zhifeng Chen, Adam Roberts, Maarten Bosma, Yanqi Zhou, Chung-Ching Chang, Igor Krivokon, Will Rusch, Marc Pickett, Kathleen S. Meier-Hellstern, Meredith Ringel Morris, Tulsee Doshi, Renelito Delos Santos, Toju Duke, Johnny Soraker, Ben Zevenbergen, Vinodkumar Prabhakaran, Mark Diaz, Ben Hutchinson, Kristen Olson, Alejandra Molina, Erin Hoffman-John, Josh Lee, Lora Aroyo, Ravi Rajakumar, Alena Butryna, Matthew Lamm, Viktoriya Kuzmina, Joe Fenton, Aaron Cohen, Rachel Bernstein, Ray Kurzweil, Blaise Agüera y Arcas, Claire Cui, Marian Croak, Ed H. Chi, and Quoc Le. 2022. Lamda: Language models for dialog applications. CoRR, abs/2201.08239.
  32. 32.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023a. Llama: Open and efficient foundation language models. CoRR, abs/2302.13971.
  33. 33.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton-Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, and et al. 2023b. Llama 2: Open foundation and fine-tuned chat models. CoRR, abs/2307.09288.
  34. 34.Martin Visbeck. 2018. Ocean science research is key for a sustainable future. Nature communications, 9(1):690.
  35. 35.Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, Wayne Xin Zhao, Zhewei Wei, and Ji-Rong Wen. 2023a. A survey on large language model based autonomous agents. CoRR, abs/2308.11432.
  36. 36.Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Qi Liu, Tianyu Liu, and Zhifang Sui. 2023b. Large language models are not fair evaluators.
  37. 37.Xingyao Wang, Zihan Wang, Jiateng Liu, Yangyi Chen, Lifan Yuan, Hao Peng, and Heng Ji. 2023c. Mint: Evaluating llms in multi-turn interaction with tools and language feedback.
  38. 38.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2023d. Self-instruct: Aligning language models with self-generated instructions. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, pages 13484–13508. Association for Computational Linguistics.
  39. 39.Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. 2022. Finetuned language models are zero-shot learners. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net.
  40. 40.Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yiwen Ding, Boyang Hong, Ming Zhang, Junzhe Wang, Senjie Jin, Enyu Zhou, Rui Zheng, Xiaoran Fan, Xiao Wang, Limao Xiong, Yuhao Zhou, Weiran Wang, Changhao Jiang, Yicheng Zou, Xiangyang Liu, Zhangyue Yin, Shihan Dou, Rongxiang Weng, Wensen Cheng, Qi Zhang, Wenjuan Qin, Yongyan Zheng, Xipeng Qiu, Xuanjing Huan, and Tao Gui. 2023. The rise and potential of large language model based agents: A survey. CoRR, abs/2309.07864.
  41. 41.Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, and Daxin Jiang. 2023. Wizardlm: Empowering large language models to follow complex instructions. CoRR, abs/2304.12244.
  42. 42.Shukang Yin, Chaoyou Fu, Sirui Zhao, Ke Li, Xing Sun, Tong Xu, and Enhong Chen. 2023. A survey on multimodal large language models. CoRR, abs/2306.13549.
  43. 43.Aohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, Weng Lam Tam, Zixuan Ma, Yufei Xue, Jidong Zhai, Wenguang Chen, Zhiyuan Liu, Peng Zhang, Yuxiao Dong, and Jie Tang. 2023. GLM-130B: an open bilingual pre-trained model.
  44. 44.Liangyu Zha, Junlin Zhou, Liyao Li, Rui Wang, Qingyi Huang, Saisai Yang, Jing Yuan, Changbao Su, Xiang Li, Aofeng Su, Tao Zhang, Chen Zhou, Kaizhe Shou, Miao Wang, Wufang Zhu, Guoshan Lu, Chao Ye, Yali Ye, Wentao Ye, Yiming Zhang, Xinglong Deng, Jie Xu, Haobo Wang, Gang Chen, and Junbo Zhao. 2023. Tablegpt: Towards unifying tables, nature language and commands into one GPT. CoRR, abs/2307.08674.
  45. 45.Ningyu Zhang, Jintian Zhang, Xiaohan Wang, Honghao Gui, Kangwei Liu, Yinuo Jiang, Xiang Chen, Shengyu Mao, Shuofei Qiao, Yuqi Zhu, Zhen Bi, Jing Chen, Xiaozhuan Liang, Yixin Ou, Runnan Fang, Zekun Xi, Xin Xu, Lei Li, Peng Wang, Mengru Wang, Yunzhi Yao, Bozhong Tian, Yin Fang, Guozhou Zheng, and Huajun Chen. 2023a. Knowlm technical report.
  46. 46.Shengyu Zhang, Linfeng Dong, Xiaoya Li, Sen Zhang, Xiaofei Sun, Shuhe Wang, Jiwei Li, Runyi Hu, Tianwei Zhang, Fei Wu, and Guoyin Wang. 2023b. Instruction tuning for large language models: A survey. CoRR, abs/2308.10792.
  47. 47.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona T. Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2022. OPT: open pre-trained transformer language models. CoRR, abs/2205.01068.
  48. 48.Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie, and Ji-Rong Wen. 2023. A survey of large language models. CoRR, abs/2303.18223.

Citation

MLA
Bi, Z., et al. “OceanGPT: A Large Language Model for Ocean Science Tasks”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 3357–72, https://doi.org/10.18653/v1/2024.acl-long.184.
APA
Bi, Z., Zhang, N., Xue, Y., Ou, Y., Ji, D., Zheng, G., & Chen, H. (2024). OceanGPT: A Large Language Model for Ocean Science Tasks. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3357–3372. https://doi.org/10.18653/v1/2024.acl-long.184
Chicago
Bi, Z., N. Zhang, Y. Xue, et al. 2024. “OceanGPT: A Large Language Model for Ocean Science Tasks”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3357–72. https://doi.org/10.18653/v1/2024.acl-long.184.
Harvard
Bi, Z. et al. (2024) “OceanGPT: A Large Language Model for Ocean Science Tasks”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 3357–3372. Available at: https://doi.org/10.18653/v1/2024.acl-long.184.
Vancouver
1. Bi Z, Zhang N, Xue Y, Ou Y, Ji D, Zheng G, Chen H (2024) OceanGPT: A Large Language Model for Ocean Science Tasks. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 3357–3372

BibTeX

@inproceedings{bi-etal-2024-oceangpt,
    title = "{O}cean{GPT}: A Large Language Model for Ocean Science Tasks",
    author = "Bi, Zhen  and
      Zhang, Ningyu  and
      Xue, Yida  and
      Ou, Yixin  and
      Ji, Daxiong  and
      Zheng, Guozhou  and
      Chen, Huajun",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.184/",
    doi = "10.18653/v1/2024.acl-long.184",
    pages = "3357--3372"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/