Generative evaluation is an assessment methodology in machine learning and artificial intelligence in which benchmark instances and diagnostic test cases are programmatically generated rather than selected from fixed, static datasets. By dynamically producing diverse inputs, controlled perturbations, or task variations through algorithmic or generative techniques, this approach systematically probes model capabilities across wide ranges of conditions, edge cases, and difficulty levels. It helps overcome the limitations of static testing, such as data contamination, benchmark memorization, and dataset-specific artifacts, while enabling automated, scalable, and continuous analysis of model behavior, robustness, and failure modes.