Sinitic Languages Diversity and Digital Humanities – DLS-HN
The development of practices in digital humanities (DH) and the increasing availability of corpora open new fields of application and new challenges for Natural Language Processing. The NLP of Sinitic languages is no exception; yet it presents some specific issues.
This project aims at describing and addressing some of these challenges by approaching these issues from the perspective of variation.
We will distinguish three axes of variation: temporal (diachronic), geographical (diatopical/dialectal) and grapholinguistic (language-script relationship). We will question the formal representations (incl. normalization and vectorization of data) and the choices of corpora at the basis of any processing of Sinitic languages.
We will study different cases of variation and different applications of NLP to DH and heritage languages.
Our contribution will be twofold. On the one hand, it will focus on the evaluation and design of NLP methods on data located at different positions along these variation axes, and on the other hand, on the dissemination of these methods and their applications. Our work will address both text processing and speech processing issues.
The temporal axis will be explored mainly through the corpus of the Shun-Pao, the first daily newspaper printed in sinograms between 1872 and 1949. This corpus allows us to address both linguistic and historical issues, and will be worked on in collaboration with the historians involved in the ENP-China project. It will be complemented by a comparison to more ancient materials.
The geographical axis will be studied with the cases of Taiwanese Hokkien and Teochew (with a focus on the variant spoken in France). These are two languages of the same family, relatively close to each other and distant from Mandarin. However, they are in quite different in terms of sociolinguistic and NLP situations. They will allow us to explore NLP transfer methods. This work will be conducted in collaboration with colleagues in Taiwan and Wikimedia France so it can benefit speaker communities.
Project coordination
Pierre MAGISTRY (EQUIPE DE RECHERCHE : TEXTES, INFORMATIQUE, MULTILINGUISME)
The author of this summary is the project coordinator, who is responsible for the content of this summary. The ANR declines any responsibility as for its contents.
Partnership
ERTIM EQUIPE DE RECHERCHE : TEXTES, INFORMATIQUE, MULTILINGUISME
Graduate Institute of Linguistics, National Taiwan University
Help of the ANR 333,633 euros
Beginning and duration of the scientific project:
November 2023
- 42 Months