Speaker similarity refers to the degree of perceptual and acoustic resemblance between two voice samples, measuring how closely a synthesized, converted, or recorded speech output matches the unique vocal identity and timbre of a target speaker. It serves as a fundamental evaluation metric in text-to-speech synthesis, voice conversion, and speaker verification systems to determine whether generated or processed audio faithfully replicates an intended individual personal voice characteristics, including pitch, resonance, and articulation patterns. This attribute is evaluated either through subjective listening tests, where human raters score the likeness between voice pairs, or through objective computational methods, such as calculating the cosine similarity or distance between speaker embedding vectors extracted by deep neural networks.