Evaluation of Phonetic Encoding Algorithms on Transcription Datasets
This paper proposes a new evaluation scheme based on the H\"ullermeier-Rifqi Index, used to measure the degree of compliance of speech encoding algorithms in IPA transcription. This scheme calculates the absolute difference in similarity between the true transcription obtained through normalized edit distance and the corresponding speech encoding, to obtain a score of incoherence, and corrects it based on the results of a random string generator using the same alphabet. The study evaluated the recall capabilities of various speech encoders and their collision-rate-based methods on multilingual transcription datasets. Additionally, the applicability of this scheme was verified; it can be used to measure the transparency of orthography when considering writing systems as inherent speech representations.