CE45 - Mathématiques et sciences du numérique pour la biologie et la santé 2020

Predictive Boolean Network Ensembles – BNeDiction

The BNeDiction project primarily relies on answer set programming (ASP) to synthesize sets of Boolean networks (BNs) that are compatible with a given network architecture and predefined dynamic properties. With this approach, the set of constraints is represented by a single logic program, such that each solution corresponds to a distinct Boolean network that satisfies the desired constraints. Our approach aims to be applied to networks comprising hundreds or thousands of nodes, depending on the specific constraints and the network architecture.

This logical approach will be combined with statistical methods that bridge the gap between quantitative data (experimental measurements of cellular state) and the qualitative state of the models.

The BNeDiction project has focused on the automatic construction of sets of Boolean networks based on specifications regarding their structure (permitted gene regulations) and their dynamics (expected behaviors). Our method relies on logic programming and automated symbolic reasoning to discover Boolean networks that satisfy these specifications.

 

The BoNesis software (https://bnediction.github.io/bonesis) implements this approach, offering a high-level language to model the inference problem, and algorithms to explore and evaluate the set of solutions.

 

In addition to significant progress on the richness of behaviors that can be specified (e.g., doi.org/10.1007/978-3-031-42697-1_11), the project has focused on methods to better sample and analyze sets of Boolean networks that satisfy the specification (https://doi.org/10.1038/s41540-025-00569-z, theses.hal.science/tel-05493556v1).

 

Finally, various theoretical results were obtained regarding the computational complexity of certain inference tasks (https://doi.org/10.1137/23M1553248) and the prediction of molecular control targets (https://doi.org/10.24072/pcjournal.255).

 

Throughout the project, the methodologies and tools developed were tested in real-world case studies in biology and ecology through interdisciplinary collaborations (e.g., doi.org/10.1038/s41540-025-00569-z, doi.org/10.1111/2041-210x.70090)

 

A key challenge is therefore to bridge the gap between quantitative experimental data and qualitative logical models.

We have thus developed scBoolSeq (https://github.com/bnediction/scBoolSeq), which implements statistical methods to classify gene activity based on scRNA-seq data. scBoolSeq also allows for the reverse process: given a Boolean network and a sequence of possible states, our method can generate fictitious quantitative datasets that closely resemble real ones. The ability to generate such data is invaluable for thoroughly validating RNA-Seq data analysis methods, as the ground truth is then known.

The BNeDiction project has resulted in a set of methods and tools that enable the automatic construction of sets of mathematical models that replicate observed biological behaviors. These sets of models thus provide insights into the underlying biological mechanisms and serve as a basis for predicting molecular targets to control cellular behavior.

 

The tools and their applications to concrete cases in biology demonstrate the feasibility and scalability of this approach.

 

These results open up various avenues:

 

- theoretical, with the development of new algorithms to predict molecular targets that modify cellular behavior based on sets of Boolean networks, and the study of additional constraints for predicting the possible behaviors of a Boolean network

 

- methodological, with the creation of synthetic datasets whose ground truth is known in order to thoroughly evaluate Boolean network inference methods based on quantitative data

 

- software-related, involving the addition of new features to BoNesis to accommodate even richer dynamic properties and facilitate its use

 

- application-related, involving the use of the developed methodology and software to address new modeling case studies in biology, health, and ecology.

Bridging the gap between dynamical systems and their partial observation, computational models of biological processes aim at uncovering key mechanisms driving cellular dynamics, and ultimately predict their behaviour under unobserved conditions.
Computational models of molecular interaction networks are usually built from data related to the structure of the biological system, such as
known interactions; and data related to its dynamics, such as measurements of gene expressions or proteins activity at different time and/or conditions.
However, despite huge advances in experimental technologies, observations of biological processes stay very scarce, either in terms of temporal resolution, number of observed entities, synchronisation between measure points, or variety of experimental conditions. Combined with complex structures for molecular interactions, the model engineering problem in this case appears to be largely under-specified, leading to (too) many potential candidate models.

Boolean Networks (BNs), and logical models in general, are widely adopted for the modelling of signalling pathways and gene and transcription factors networks as they are not demanding for the exact knowledge of the quantitative parameters of the selected molecular interactions. With BNs, the activity of components is caricatured to “off” and “on”, and their evolution is computed according to logical rules
However, in practice, biological data still let open a multitude of candidate BNs. Thus, arbitrary modelling choices have to be made, e.g., by prioritizing certain logics between regulators or by preferring smallest/largest models, which may introduce biases in subsequent model predictions.

The BNeDiction project aims at providing a general methodology for making predictions from data on systems structure and dynamics by the means of ensembles of Boolean networks (BNs), an unexplored direction.

By focusing on ensembles of models, we aim at capturing the diversity of admissible models and reduce biases due to the selection of a single “best” model from arbitrary criteria. Similarly to random forest approaches, we will constitute ensembles of BNs representative of the whole multitude of admissible models, and then compute predictions from the ensemble.
Based on recent advances on the symbolic and implicit formal characterization of the compatible models using logic programming, the key challenges relate to the sampling of ensemble of diverse models, and the evaluation and maximization of its predictive power, with a thorough benchmarking of the pipeline.

Ensemble modelling has the potential of improving the robustness of predictions by accounting for potential model variability and uncertainty, We plan to demonstrate the BNeDiction pipeline for the ensemble modelling of mouse hematopoietic system from single-cell RNA-seq differentiation data available with different mutant conditions. These type of data brings strong dynamical constraints for the model inference, and the different mutant conditions can be split into training and testing data to evaluate the predictive power of ensemble modeling.

Overall, the project aspires at delivering a convincing methodology for assessing the adequacy of automated logical modelling from experimental data, a key and recurring question at the intersection of artificial intelligence and life sciences

Project coordination

Loïc Paulevé (Laboratoire Bordelais de Recherche en Informatique)

The author of this summary is the project coordinator, who is responsible for the content of this summary. The ANR declines any responsibility as for its contents.

Partnership

LaBRI Laboratoire Bordelais de Recherche en Informatique

Help of the ANR 250,992 euros
Beginning and duration of the scientific project: February 2021 - 48 Months

Useful links

Explorez notre base de projets financés

 

 

ANR makes available its datasets on funded projects, click here to find more.

Sign up for the latest news:
Subscribe to our newsletter