Toxicity classification is a natural language processing task that uses machine learning algorithms to automatically identify and categorize harmful, offensive, or abusive content within text. This process involves analyzing digital text to distinguish benign communication from various forms of toxic language, which can include explicit hate speech, harassment, profanity, and threats, as well as subtle, context-dependent implicit toxicity such as microaggressions or veiled insults. Widely applied in content moderation platforms, online discussion forums, and artificial intelligence safety auditing, toxicity classification helps filter inappropriate material, enforce community standards, and monitor outputs from generative language models to maintain safe and respectful communication environments.