OceanGPT: A Large Language Model for Ocean Science Tasks
Zhen BiNingyu ZhangYida XueYixin OuDaxiong JiGuozhou ZhengHuajun Chen
Presents OceanGPT, the first domain-specific large language model for oceanography, trained using a multi-agent instruction generation framework and evaluated on a new dedicated benchmark to handle specialized scientific tasks and ocean robotics commands.
Oceans cover over 70% of the Earth's surface and play a critical role in global climate regulation, biodiversity, and economic activity. While general-purpose large language models have advanced scientific research across various disciplines, they frequently fail to meet the complex demands of oceanography. This shortfall stems from the vast, intricate nature of marine data and the need for deep, specialized knowledge across diverse oceanic subfields.
To address this gap, the article develops and evaluates OceanGPT, the first dedicated large language model for ocean science tasks. The initiative aimed to create a domain-specific model capable of answering complex oceanographic queries, executing scientific language tasks, and generating actionable commands for marine engineering systems.
To build the system, researchers constructed a specialized training corpus from 67,633 open-access ocean science documents. They then developed DoInstruct, an automated multi-agent framework that generated more than 150,000 instruction tuning samples across five major oceanic topics: science and research, resources and development, ecology and environment, technology and engineering, and culture. The framework utilized specialized agents acting as evolving generators, literature extractors, and inspectors with rule-based quality controls. To evaluate performance, the researchers introduced OceanBench, an oceanography benchmark spanning 15 distinct tasks, utilizing both automated evaluations calibrated by GPT-4 and human evaluations by marine science experts.
Key findings show that OceanGPT significantly outperformed leading general-purpose open-source models, including LLaMA-2-7B-Chat, Vicuna-1.5-7B, and ChatGLM2-6B, winning the majority of tested sub-tasks in both automated and human evaluations. In task-level human evaluations, OceanGPT won 12 to 14 out of 15 tasks against the baseline models. The model demonstrated superior domain expertise in complex scientific workflows, such as radioactive nuclide research, providing practical experimental and risk assessment steps where baseline models offered only generic responses. Furthermore, OceanGPT demonstrated preliminary embodied intelligence for ocean engineering, successfully generating control code and console commands to direct underwater robots in physical simulation environments.
These findings indicate that domain-adapted language models can substantially improve analytical accuracy, technical problem-solving, and operational planning in maritime industries and scientific research. By integrating domain-specific literature and multi-agent synthetic data generation, organizations can reduce the high costs and labor bottlenecks typically required to train specialized artificial intelligence systems.
Decision-makers and research teams should consider adopting specialized multi-agent instruction generation frameworks to scale domain-specific models cost-effectively. For marine technology deployment, organizations should conduct further pilot testing of OceanGPT’s robotic code generation within real-world maritime systems before relying on it for high-stakes operational planning.
Confidence in the model's domain expertise is supported by strong agreement among expert evaluators (an inter-annotator agreement score of 0.82) and consistent benchmark performance. However, users must remain cautious regarding standard language model limitations present in the system, including potential data distribution biases from public literature and occasional factual hallucinations.
- Paper: WizardLM: Empowering Large Pre-Trained Language Models to Follow Complex Instructions, Can Xu et al. (2023). Learn how Evol-Instruct pioneered automated, multi-agent instruction generation to create high-complexity synthetic tuning data, establishing the foundational paradigm adapted by OceanGPT's DoInstruct framework.
- Paper: BloombergGPT: A Large Language Model for Finance, Shijie Wu et al. (2023). Examine how domain-specific corpus curation and targeted language model training can outperform general-purpose foundation models in specialized disciplines, providing a direct conceptual blueprint for oceanographic LLM adaptation.
- Paper: Solving Quantitative Reasoning Problems with Language Models, Aitor Lewkowycz et al. (2022). Understand the methodology of training foundation models on technical STEM literature to overcome complex scientific workflows and quantitative reasoning challenges.
- Paper: Voyager: An Open-Ended Embodied Agent with Large Language Models, Guanzhi Wang et al. (2023). Discover how LLM-based autonomous agents translate high-level natural language instructions into executable code and commands for embodied physical simulation environments.
- Paper: A survey on large language model based autonomous agents, Lei Wang et al. (2023). Review the core architectural components of LLM-based autonomous agents to contextualize OceanGPT’s multi-agent synthetic data generation and embodied robotic execution.
- Paper: LLM and Simulation as Bilevel Optimizers: A New Paradigm to Advance Physical Scientific Discovery, Pingchuan Ma et al. (2024). Explore how coupling large language models with differentiable physical simulations advances automated scientific hypothesis generation and parameter discovery beyond standard prompt and command execution.
- Paper: Innovator-VL: A Multimodal Large Language Model for Scientific Discovery, Zichen Wen et al. (2026). See how domain-specific scientific LLMs can be extended to multimodal reasoning across complex scientific visual data and experimental imagery.
- Paper: AutoML-Agent: A Multi-Agent LLM Framework for Full-Pipeline AutoML, Patara Trirat et al. (2025). Investigate how multi-agent LLM systems can be further extended to automate entire end-to-end machine learning engineering pipelines under strict execution constraints.
