Generating Feature Vectors from Phonetic Transcriptions in Cross-Linguistic Data Formats

Arne Rubehn; Jessica Nieder; Robert Forkel; Johann-Mattis List

doi:10.7275/scil.2144

Options

Paper

Generating Feature Vectors from Phonetic Transcriptions in Cross-Linguistic Data Formats

Authors

Arne Rubehn (University of Passau)
Jessica Nieder (University of Passau)
Robert Forkel (DLCE, MPI-EVA)
Johann-Mattis List (University of Passau)

Abstract

When comparing speech sounds across languages, scholars often make use of feature representations of individual sounds in order to determine fine-grained sound similarities. Although binary feature systems for large numbers of speech sounds have been proposed, large-scale computational applications often face the challenges that the proposed feature systems -- even if they list features for several thousand sounds -- often only cover a smaller part of the numerous speech sounds reflected in actual cross-linguistic data. In order to address the problem of missing data for attested speech sounds, we propose a new approach that can create binary feature vectors dynamically for all sounds that can be represented in the the standardized version of the International Phonetic Alphabet proposed by the Cross-Linguistic Transcription Systems (CLTS) reference catalog. Since CLTS is actively used in large data collections, covering more than 2,000 distinct language varieties, our procedure for the generation of binary feature vectors provides immediate access to a very large collection of multilingual wordlists. Testing our feature system in different ways on different datasets proves that the system is not only useful to provide a straightforward means to compare the similarity of speech sounds, but also illustrates its potential to be used in future cross-linguistic machine learning applications.

Keywords: phonological features, vector representation, computer-assisted language comparison, cross-linguistically linked data

How to Cite:

Rubehn, A., Nieder, J., Forkel, R. & List, J., (2024) “Generating Feature Vectors from Phonetic Transcriptions in Cross-Linguistic Data Formats”, Society for Computation in Linguistics 7(1), 205–216. doi: https://doi.org/10.7275/scil.2144

Downloads:
Download PDF

1029 Views

721 Downloads

Published on
2024-06-24

Peer Reviewed

License

Creative Commons Attribution 4.0

Authors

Arne Rubehn (University of Passau)
Jessica Nieder (University of Passau)
Robert Forkel (DLCE, MPI-EVA)
Johann-Mattis List (University of Passau)

Publication details

Pages: 205–216
Submitted on: 2024-06-11
Accepted on: 2024-06-17

File Checksums (MD5)

PDF: No checksum could be calculated.

Generating Feature Vectors from Phonetic Transcriptions in Cross-Linguistic Data Formats

Abstract

Harvard-Style Citation

Vancouver-Style Citation

APA-Style Citation

Non Specialist Summary