Skip to main content
Abstract

Mind the Gap: Computational Quality Assurance of Crowd-Sourced Linguistic Knowledge on Latin and Italian Morphological Gaps

Authors
  • Jonathan Sakunkoo
  • Annabella Sakunkoo
  • Annabella Sakunkoo orcid logo (Stanford University)

Abstract

Morphological defectivity is an intriguing and understudied phenomenon in linguistics. Addressing defectivity, where expected inflectional forms are absent, is essential for improving the accuracy of NLP tools in morphologically rich languages. However, traditional linguistic resources often lack coverage of morphological gaps as such knowledge requires significant human expertise and effort to document and verify. For scarce linguistic phenomena in under-explored languages, Wikipedia and Wiktionary often serve as among the few accessible resources. Despite their extensive reach, their reliability has been a subject of controversy. This study customizes a novel neural morphological analyzer to annotate Latin and Italian corpora. Using the massive annotated data, crowd-sourced lists of defective verbs compiled from Wiktionary are validated computationally. Our results indicate that while Wiktionary provides a highly reliable account of  Italian morphological gaps, 7% of Latin lemmata listed as defective show strong corpus evidence of being non-defective. This discrepancy highlights  potential limitations of crowd-sourced wikis as definitive sources of linguistic knowledge, particularly for less-studied phenomena and languages, despite their value as resources for rare linguistic features. By providing scalable tools and methods for quality assurance of crowd-sourced data, this work advances computational morphology and expands linguistic knowledge of defectivity in non-English, morphologically rich languages.

Keywords: Morphology, NLP, Morphological Gaps, Defectivity, Defectiveness, Defective Words, Inflectional Gaps, UDTube, UD, Linguistics, Latin, Italian

How to Cite:

Sakunkoo, J., Sakunkoo, A. & Sakunkoo, A., (2025) “Mind the Gap: Computational Quality Assurance of Crowd-Sourced Linguistic Knowledge on Latin and Italian Morphological Gaps”, Society for Computation in Linguistics 8(1): 47. doi: https://doi.org/10.7275/scil.3186

Downloads:
Download PDF

24 Views

6 Downloads

Published on
2025-06-14

Peer Reviewed