Enriching bilingual reading and interaction with cross-lingual alignments – TransRead
Owing to the globalization of economies and to the advent of the Internet as a truly universal media, a growing number of interactions take place between individuals and/or companies speaking different languages. As a result, the demand for translation services is more important than ever, speeding up the development of numerous effective Machine Translation (MT) solutions, making MT one of the key technologies of a multilingual Internet.
In recent years, research and development in MT have focused on systems that attempt to attain human performance while being used as black boxes. If the quality of the resulting translations is commonly far below that of a human professional translator, a state of affair which will likely persist in a foreseeable future, MT outputs are nowadays sufficiently good to deliver valuable services for the general public. Not less importantly, MT is now considered a standard tool by the translation industry.
In this context, the objective of TransRead is to study new multilingual text processing applications, aimed at facilitating the reading of multilingual documents for readers with intermediate knowledge of a foreign language. Contrarily to black-box approaches, which target users without any knowledge of the original language of some text, TransRead is primarily concerned with the visualization of bilingual texts and of their cross-lingual links.
Subtitling is a commonplace technology that exploits such bilingual alignments to enable viewers to read in a familiar language what is said in some unknown language. By analogy, one of the objectives of TransRead is to explore the subtitling of books, and to devise ways in which cross-lingual alignments, computed at the sentential and sub-sentential levels, or cross-lingual dictionary access, will help and enrich the reading of texts in their original language. To this end, we intend to take advantage of the opportunities created by the availability of new mobile terminals (touchpad tablets, electronic readers) and by the recent advances in information visualization technologies.
Between the complete ignorance of a language (where translation is the only option) and bilingualism (where translation is useless), there exist a variety of contexts of partial bilingualism, where such bilingual electronic readers would prove highly useful: for second language learners, for workers in international environments, for migrant settling in a new country, for inhabitants of multilingual states, for speakers of related languages, or, in more industrial contexts, for editors in the publishing industry or for professional translators.
For this latter type of professional uses, we propose to study a second kind of application related to the crucial issue of quality control of human translation and of translation memories. Bodies of document translations, so called parallel corpora, constitute key resources in modern computer assisted translation environments. They are used to derive bilingual lexicons or terminologies and to feed translation memories, not to mention their use for MT applications. Yet, the quality of these resources and its impact on their reusability have, surprisingly, not received much attention to date. A second objective of TransRead is therefore to study and develop innovative visualization strategies for such bilingual texts, with a view on assessing and improving their quality. This raises difficult scientific issues, such as the design of numerical confidence measures evaluating the validity of alignments or the conception of effective visualization techniques for large bilingual corpora.
Project coordination
François YVON (Laboratoire d'Informatique pour la Mécanique et les Sciences de l'Ingénieur)
The author of this summary is the project coordinator, who is responsible for the content of this summary. The ANR declines any responsibility as for its contents.
Partnership
CEDRIC Centre d'Etudes et de Recherche en Informatique et Communications
Reverso Softissimo Softissimo
LIMSI-CNRS Laboratoire d'Informatique pour la Mécanique et les Sciences de l'Ingénieur
Help of the ANR 593,637 euros
Beginning and duration of the scientific project:
September 2012
- 36 Months