Pairwise evaluation is an assessment methodology in machine learning and natural language processing where an evaluator directly compares two candidate outputs generated for the same prompt to determine which response is superior according to defined criteria such as correctness, quality, or helpfulness. Rather than assigning absolute numerical scores to individual outputs in isolation, this side-by-side approach relies on relative comparison, enabling human annotators or automated model judges to choose the preferred output or declare a tie. By focusing on relative preference, pairwise evaluation reduces calibration bias and subjective scale variance inherent in independent scoring systems, making it a foundational method for benchmarking generative models, calculating comparative performance rankings, and training reward models.