Unsafe categories refer to standardized classifications of harmful, illegal, or unethical content used to evaluate, moderate, and improve the safety of artificial intelligence systems. In the context of natural language processing and content moderation, these categories delineate specific types of unacceptable model outputs, commonly encompassing areas such as hate speech, harassment, self-harm, sexually explicit material, violence, illegal acts, and privacy violations. Researchers and developers utilize these predefined taxonomies to assess model vulnerabilities against adversarial prompts, measure toxicity across diverse risk domains, and implement safety interventions to ensure generated content aligns with ethical and legal standards.