CE23 - Intelligence Artificielle 2021

Coreference resolution into machine translation – CREMA

Submission summary

In this project we would like to make a step forward in the domain of document-level neural machine translation by dynamically choosing the contextual information that the models uses to generate its translations.
This is opposed to current works where the contextual information is fixed chosen a-priori. The latter solution does not prove very effective even when using a relatively short context. This is due to the fact that, in most cases, the model can translate correctly a sentence without using any context. Words needing a context for their correct translation are relatively rare, thus learning specific model contextual parameters for taking them into account is difficult, as the training signal is sparse.
In this project we would like to study models with a more compact context. The choice of such context is leaded however by an additional module which is able to detect the most ambiguous words a translation model can face: discourse phenomena, and in particular anaphora and coreferences.
Another aspect we would like to study in the project is the specificity of the evaluation of Document-Level Neural Machine Translation models.
Indeed for such models, the BLEU evaluation metric is not adapted. The words needing a context for their correct translation are relatively rare, their impact on an automatic evaluation metric like BLEU is thus limited. Their correct translation however, and more in general contextualized translation, has a non negligible effect on the translation quality as perceived by a reader.
For a better evaluation of contextual models, contrastive test suites have been designed. We find that such kind of evaluations can be improved by using more realistic sentences. Current test suites contain indeed mostly artificial sentences choosen ad-hoc.
The main objectives of the CREMA project (Coreference REsolution into MAchine translation) are: 1) designing new models for coreference resolution; 2) integrating a coreference module into Document-Level NMT models so that to allow a dynamic context choice, based on ambigous discourse phenomena detected by such module; 3) designing a new test suite, more effective for Document-Level NMT models evaluation.

Project coordination

Marco Dinarelli (Laboratoire d'Informatique de Grenoble)

The author of this summary is the project coordinator, who is responsible for the content of this summary. The ANR declines any responsibility as for its contents.

Partnership

LIG Laboratoire d'Informatique de Grenoble

Help of the ANR 253,055 euros
Beginning and duration of the scientific project: December 2021 - 48 Months

Useful links

Explorez notre base de projets financés

 

 

ANR makes available its datasets on funded projects, click here to find more.

Sign up for the latest news:
Subscribe to our newsletter