JudgeBench is an evaluation benchmark designed to assess the accuracy and reliability of large language models when they are used as automated judges to evaluate other artificial intelligence outputs. Unlike traditional evaluation suites that primarily measure how well an automated judge aligns with subjective human preferences, JudgeBench focuses on objective factual and logical correctness across complex, technical domains such as mathematics, computer programming, reasoning, and advanced knowledge. It tests these evaluating systems by presenting them with challenging pairs of model-generated responses labeled with verified ground truth, enabling researchers to determine whether automated judges can accurately distinguish correct technical solutions and arguments from plausible yet flawed alternatives.