CE45 - Mathématiques et sciences du numérique pour la biologie et la santé 2020

Statistics for characterizing interactions between plant and its environment – Stat4Plant

Statistical methods for characterizing interactions between plants and their environment: accounting for heterogeneity in plant populations and environmental variations

Developing new statistical methods and algorithmic tools for modelling and analyzing genetic variability and the interactions between plants and their environment in the context of climate change.

Predicting crop performance in changing environments and improving variety selection within heterogeneous populations using numerical methods

The agricultural sector currently faces major challenges linked to the need to increase global food production despite the increasing scarcity of soil and water resources, and the intensifying impacts of climate change, such as heatwaves, drought and pest infestations. These constraints have a significant impact on plant development and yields, as illustrated by the stagnation of arable crops in Europe. Genotype-environment interactions play a key role: a given variety may develop differently depending on conditions, and different varieties may react differently within the same environment. These interactions offer the opportunity to select better-adapted and more resilient varieties. To exploit this potential, new strategies must incorporate the uncertainty of environmental conditions, genetic variability and the dynamics of plant development. Predictive approaches combining ecophysiology, genetics and mathematical modelling are promising, allowing numerous environmental scenarios to be explored without testing all combinations experimentally. Non-linear dynamic models enable growth processes and genotype-environment interactions to be accurately represented, but pose significant statistical and computational challenges. The following factors must be taken into account: (i) the interaction of complex non-linear dynamics (crop, climate, pests), (ii) the influence of heterogeneous and high-dimensional covariates, and (iii) different levels of variability, particularly genetic variability. These models involve numerous parameters to be estimated from complex data, requiring suitable tools for estimation, variable selection and model comparison in order to identify the main sources of variability. The Stat4Plant project aims to develop these original methods and analyse dynamic crop data from multiple environments and conditions. It aims to provide innovative predictive tools to improve crop performance in the face of climate change. Its four main objectives are: to identify genotype-dependent parameters governing plant development; to jointly model flowering and certain dynamic phenotypic traits; to identify influential covariates from a large dataset; and to propose variety selection criteria that incorporate environmental uncertainty. These methodologies will be generic and applicable to other dynamic models, with potential impact in other applied fields, such as medicine.

Most of our studies combine mixed-effects statistical models with mechanistic mathematical models to better quantify the variability within a heterogeneous population of varieties. More specifically, we consider non-linear mixed-effects models which are widely used in agronomy and pharmacology to account for different levels of variability in repeated measures data, in particular intra-individual and inter-individual variability. These models combine fixed effects, shared by all individuals, and random effects, specific to each individual. Accurate modelling of these effects is essential for parameter estimation and prediction. In the context of plant growth models, intra-individual variability is captured by the non-linear mechanistic model, whilst inter-individual variability is represented by individual parameters treated as random effects, the latter being latent variables. For genotypes, this inter-individual variability corresponds to inter-genotype variability. One of the objectives is to quantify and characterise this inter-individual variability. The aim is therefore to assess whether or not individual parameters vary within the population. We also seek to characterise this variation using descriptive covariates of the varieties, potentially in a high-dimensional context. We also consider joint models for longitudinal data on phenotypic traits of interest and survival time data. These models combine mixed-effects statistical models, mechanistic mathematical models and survival analysis models. The approaches developed for inference are based on methods of estimation, variable selection and prediction. They also make use of recent numerical optimisation tools. We are also developing new multi-criteria approaches to simultaneously optimise several traits of interest within a population of genotypes by incorporating the uncertainties of environmental conditions.

The Stat4Plant project has led to the development of a body of recent research aimed at improving the analysis of complex data, by combining advances in modelling, statistics and artificial intelligence.

 

We have thus developed a new statistical testing method, which is more reliable than traditional approaches, even in highly complex models, for identifying variability within a mixed-effects model. In particular, this method allows for the inclusion of nuisance parameters that were previously difficult to account for. It is based on a technique known as ‘bootstrap’, which improves the robustness of the results. We have also proposed a new efficient algorithm for the numerical calculation of mathematical quantities that are very difficult to evaluate within the framework of this procedure. This algorithm is more accurate and stable than existing methods, even in situations where the data is heterogeneous. Its theoretical properties have been studied, guaranteeing its long-term reliability.

 

In parallel, a new machine learning method has been developed to estimate parameters in complex latent variable models. It not only enables accurate estimates to be obtained, but also allows their uncertainty to be assessed. This approach is flexible, efficient and applicable to many types of models.

 

The first two developments have been used to analyze genotypic variability of Arabidopsis Thaliana based on the biological plant model, known as ARNICA, which has first been improved and refined to better understand internal exchanges (carbon, nitrogen) and differences between plants.

 

We have also developed joint models capable of linking time-varying data to events of interest. These models can incorporate a large number of variables. We have developed methods to identify those that actually have an impact. These methods utilize modern optimization techniques to manage the complexity of the data. These tools have been applied to real-world problems in agriculture. They have enabled us to carry out a first quantitative study of the impact of pests on maize development. They have also been used to identify biological factors linked to plant resistance.

 

New methods have been developed to select the most important variables when there are a very large number of them, as in genetics. These approaches, which are both Bayesian and frequentist, allow us to better target the factors that have an impact. Their effectiveness has been demonstrated on both simulated and real data. Robust theoretical results have also been obtained, showing that these methods remain reliable even in highly complex and high-dimensional contexts.

 

Finally, multi-criteria optimization methods have been proposed to facilitate the selection of the best plant varieties. These methods make it possible to consider multiple objectives simultaneously, such as yield and quality. They also account for uncertainties related to environmental conditions.

The work carried out in the Stat4Plant project opens up several promising research perspectives.

 

Variable selection in high-dimensional settings remains a major challenge. During the evaluation of the performance of the algorithmic tools developed to integrate and select high-dimensional covariates in models with random effects, it appeared that potential correlations between covariates in real data, such as genetic markers, strongly degrade selection performance. New hybrid approaches combining statistics and deep learning could be explored.

 

The bootstrap-based methods developed to test variance components in mixed-effects models, which are well suited to small sample sizes, require long computation times in practice. This can be a limitation, and further work is needed to explore possibilities for acceleration through more efficient computational approaches or new numerical approximations.

 

Joint models could be further enriched to capture more detailed interactions between biological and environmental processes, in particular by incorporating a spatial component.

 

In multi-criteria optimization, it would be useful to integrate future climate scenarios in order to anticipate the effects of climate change.

 

Another direction would be to adapt the developed methods to even larger datasets, particularly those produced by high-throughput phenotyping technologies. Improving algorithmic efficiency will therefore be essential to reduce computation times. It would also be relevant to integrate more heterogeneous data (omic, environmental, imaging) into the models.

 

Finally, transferring these methods into operational applications represents a key challenge. The development of user-friendly decision-support tools is crucial for innovation and for enhancing the impact of research. This will require the development of accessible software and strong collaborations with stakeholders in the agricultural sector.

Agriculture has currently to tackle new challenges, largely due to the need to increase global food supply under the declining availability of soil and water resources and increasing threats from climate change. It has to face main changes and to adapt to new conditions, in particular environmental ones. To better handle this adaptation, it is necessary to better understand several key notions such as genetic variability and interactions between the plant and its environment. In this context, predictive approaches relying on ecophysiology and genetic knowledge, as well as mathematical modeling are very promising.

The Stat4Plant project aims at developing new statistical methodologies and new algorithmic tools for modeling and analyzing genotypic variability and interaction between plant and its environment in a context of climate change. The project consortium gathers scientists in modeling and applied statistics with large experience in interdisciplinary collaborations in plant sciences and biologists with strong expertise in phenotype-genotype relations. The project is structured in four main research axes, supported by strong collaborations between statisticians and biologists and motivated by practical questions linked with biological dataset.

The first axis aims at developing new methods for identifying key biological processes driving plant development lying behind the observed genotypic variability. These works combine mechanistic ecophysiologic modeling highly nonlinear of plant development, statistical mixed effects modeling for genotypic variability and statistical testing procedures, in particular adapted to small data samples, to identify genotype-dependent parameters. These approaches will allow to better understand genotype by environment interactions and to identify new tools for varietal selection.

The second axis is dedicated to joint modeling of a time of interest such as flowering time or harvest time and of a phenotypic dynamical trait depending on time such as biomass or pest presence. The considered joint models combine survival models with random effects and covariates of high dimension and nonlinear mixed effects models for the dynamical trait. The objective is to identify the relevant covariates, to estimate the parameters and to predict the time of interest. These methods will allow for example to better predict flowering time or optimal harvest time.

The third axis aims at developing new methods for identifying among a high number of covariates those who are the most influent for a phenotypic trait of interest, solely or jointly with a time of interest. Nonlinear mixed effects models combining mechanistic models of plant development and genetic models integrating a high number of genetic covariates will be used to model genotypic variability of the trait of interest. New covariates selection methods adapted to the nonlinear context will be developed. These methods will allow to identify the main genetic factors influencing the trait.

Finally, the last axis aims at building new criteria for varietal selection, integrating randomness of environmental conditions and targeting simultaneously several objectives, such as maximizing yield and minimizing yield variability. New methodologies for optimizing these criteria will be developed. Such criteria will be new tools for decision support system in agriculture.

Project coordination

ESTELLE KUHN (Mathématiques et Informatique Appliquée du Génome à l'Environnement)

The author of this summary is the project coordinator, who is responsible for the content of this summary. The ANR declines any responsibility as for its contents.

Partnership

MaIAGE Mathématiques et Informatique Appliquée du Génome à l'Environnement
MIA Mathématiques et Informatique Appliquées
MICS Laboratoire de Mathématiques et Informatique pour la Complexité et les Systèmes
HEUDIASYC Heuristique et diagnostic des systèmes complexes
GQE Génétique quantitative et Evolution - Le Moulon
IJPB Institut Jean-Pierre BOURGIN

Help of the ANR 495,249 euros
Beginning and duration of the scientific project: January 2021 - 48 Months

Useful links

Explorez notre base de projets financés

 

 

ANR makes available its datasets on funded projects, click here to find more.

Sign up for the latest news:
Subscribe to our newsletter