Fine-tuned judges are large language models that have undergone specialized training on curated evaluation datasets to assess, score, or compare the outputs generated by other artificial intelligence models. Unlike prompted judges that rely purely on in-context instructions supplied to general-purpose foundation models, fine-tuned judges update their underlying model parameters through supervised training or preference optimization using evaluation rubrics, critique feedback, and ranked response pairs. This targeted adaptation enables the models to learn domain-specific assessment standards, generate consistent evaluation rationales, and mitigate common evaluator tendencies such as position or length bias, providing a scalable and specialized approach to automated quality assessment and model benchmarking.