Length-normalized scoring is a sequence evaluation method in natural language processing and probabilistic modeling that calculates the overall likelihood or confidence of a generated sequence by dividing its cumulative log-probability by the total number of tokens in that sequence. Because joint probabilities naturally diminish as more tokens are generated due to repeated multiplication of fractional values, unadjusted sequence scores inherently penalize longer outputs regardless of their quality. Length-normalized scoring mitigates this length bias by measuring the average per-token score, enabling balanced comparisons across candidate outputs of differing lengths during sequence decoding, response ranking, and uncertainty estimation.