A jailbreak loss function is a mathematical objective function used in optimization-based adversarial attacks on large language models to evaluate how close the model is to generating a prohibited or safety-violating response. Typically formulated as a cross-entropy or negative log-likelihood loss over target output tokens, such as affirmative introductory phrases or compliant completions, it quantifies the probability of the model following a restricted instruction. Adversarial optimization algorithms use gradients or heuristic token substitutions guided by this loss to iteratively alter the input prompt or append adversarial suffixes, effectively lowering the loss to bypass safety alignment mechanisms and compel the model into fulfilling forbidden queries.