Language model-based evaluation is an assessment method in artificial intelligence where a language model is employed to analyze, score, or compare the quality of generated text or the performance of other models. Unlike traditional reference-matching metrics that rely strictly on word overlap, this approach leverages the semantic understanding and instruction-following abilities of language models to evaluate complex criteria such as coherence, factual accuracy, helpfulness, and style. Common implementations include direct scoring against predefined rubrics, pairwise ranking of competing outputs, and generating explanatory critiques. While it offers a scalable, automated alternative or supplement to human evaluation, practitioners often monitor and calibrate evaluator models to address potential biases, inconsistencies, and discrepancies with human judgment.