keyword
instruction-following difficulty score
An instruction-following difficulty score is a quantitative metric used in machine learning to measure how challenging it is for a language model to generate a target response when prompted with a specific instruction. Rather than merely evaluating the intrinsic complexity of generating the output text on its own, the score assesses the discrepancy between the model probability of generating the response conditioned on the instruction and the probability of generating that response unconditionally or directly. By comparing these two values, typically through a ratio of cross-entropy loss or perplexity, the score determines whether an instruction effectively guides the model or presents a significant alignment hurdle. This metric is primarily applied during instruction tuning to identify, filter, and select the most informative, high-impact training samples, enabling more efficient model optimization using smaller, higher-quality datasets.
1 item

