Runtime Self-adaptable Approximate Hardware Accelerators for Deep Learning – REAxION
Deep learning (DL) has transformed many aspects of human society, influencing industries, daily life, and the socio-economic landscape. Indeed, DL-based systems are now pervasive. Their exceptional accuracy stems from the extremely costly training using massive input datasets. Also at runtime – once the trained DL system is deployed – the computational demands are increasing constantly. Interestingly, data complexity varies within a dataset, and DL models with very different computational demands can perform correct predictions on most of the data.
As we shift the gears towards carbon-aware sustainable computing, we must develop new techniques to lower the computing demands while achieving comparable accuracy to state-of-the-art DL models. Specialized hardware accelerators lead to order-of-magnitude efficiency gains.
Approximate Computing (AxC) has been studied for the last 10-15 years as an alternative computing paradigm that can significantly improve the efficiency of computing systems by carefully trading off some accuracy. DL systems often are robust to minor errors, enabling the exploration of AxC methods for DL hardware accelerators, leading to order-of-magnitude efficiency gains even on large DL models.
Since the error entailed by AxC depends on the application and data distribution, dynamically tuning approximation degrees allows for improving the final output accuracy.
Therefore, (i) DL models do not require the same computational demands for all the inputs, and (ii) approximation tuning during computations leads to high efficiency gains. Hence, depending on the actual system inputs, selecting the correct approximation degree of DL accelerators would significantly improve the overall efficiency while still providing high accuracy. REAxION proposes a co-design methodology to perform a Design Space Exploration (DSE) of runtime self-adaptable approximate hardware accelerators for DL inference tasks. The runtime adaptation will depend on the input complexity.
Project coordination
Marcello TRAIOLA (INSTITUT NATIONAL DE LA RECHERCHE EN INFORMATIQUE ET AUTOMATIQUE)
The author of this summary is the project coordinator, who is responsible for the content of this summary. The ANR declines any responsibility as for its contents.
Partnership
INSTITUT NATIONAL DE LA RECHERCHE EN INFORMATIQUE ET AUTOMATIQUE
Help of the ANR 319,956 euros
Beginning and duration of the scientific project:
- 42 Months