Toxic prompts are text inputs, queries, or sentence prefixes provided to artificial intelligence language models that contain offensive, hateful, or harmful language, or are designed to elicit toxic responses from the system. In artificial intelligence safety research and evaluation, these prompts are utilized as benchmarking tools to measure the propensity of a model to generate inappropriate content, such as profanity, insults, or biased remarks, and to test the robustness of safety guardrails and moderation filters. Such prompts can range from overtly abusive statements to subtly provocative or incomplete phrases that inadvertently trigger harmful text generation learned from uncurated pretraining data.