Self-play fine-tuning is an iterative machine learning training method in which a language model improves its capabilities by generating its own training data and competing against earlier versions of itself. In this approach, which typically builds upon an initial supervised baseline, the current iteration of the model generates candidate responses while the training process optimizes the model to discern and favor high-quality target demonstration data over its self-generated outputs. By progressively refining its policy across multiple rounds so that its generated responses become indistinguishable from the target data distribution, the model enhances its overall performance and alignment without requiring additional human-annotated labels, external reward models, or feedback from stronger teacher systems.