Extraction of information from semi-structured handwritten tables in the Fichier Domiciliaire for a history of the population of Strasbourg (1871-1939) – POPSTRAS
POPSTRAS
Extraction of information from semi-structured handwritten tables in the Fichier Domiciliaire for a history of the population of Strasbourg (1871-1939)
Construction of a database from an unprecedented data in France using methods based on deep learning.
Historical demography has focused mainly on small populations (villages) and has done little analysis of the population of towns and cities because the date collection time was long. As a result , the urban populations remain less studied, even through that cities underwent major transformations (industrialisation, urbanisation). Recent advances in Deep Learning have made it possible to overcome these difficulties and exploit new sources of data. <br />The aim of this project is to build a large-scale database using a data source unprecedented in France: the Fichier Domiciliaire of the city of Strasbourg (1871-1939). The database makes it possible to track households over time and space, and to reconstruct their family and residential trajectories. Its richness will make it possible to carry out innovative analyses on a wide range of subjects. For that, we will be using innovative computer methods for automatic recognition of handwritten characters based on Deep Learning.<br />While the source used in this project, provides a unique opportunity to gain a better understanding of the urban population of this period, it also represents specific and original challenges for IT, particularly in terms of the automatic digitisation of historical sources. This corpus is complex to process, as there are specific difficulties linked to handwriting and the diversity of semi-tabular layouts (handwritten nature of the text, different styles of cursive writing in Latin and German, variation in the spatial layout of lines and information fields, sentences sometimes cramped or overflowing into neighboring fields). This project will be carried out in collaboration with a team of computer scientists specializing in Deep Learning for vision and language processing for automatic reading of documents, and a multidisciplinary team of researchers in the humanities and social sciences (demographers, historians, geographers).
The Fichier Domiciliaire de la ville de Strasbourg (FDS), maintained by the police, was opened at the beginning of the annexation and remained after Alsace-Moselle was returned to France. It contains 1.2 million household cards, written in German (Gothic script, Kurrentschrift) from 1871 to 1919 and in French after 1919. Each card contains a wealth of socio-demographic information on all household members, as well as information on the household's residential history in and outside the city.
Recent progress in Deep Learning make it possible to recognise the information contained in written data sources and record it in a tabular database. Therefore, our project requires new research to address the scientific and technical challenges posed by the nature of the FDS. Reading the FSD cards faces the challenge of understanding semi-structured documents (irregular tables) handwritten in the two languages. Most of the information to be extracted are named entities (first names and family names, places and dates) that pertain to very large lexicons which make the handwriting recognition all the more difficult. To overcome these difficulties, we plan to develop a Semi-structured Document Understanding system, the output of which will be disambiguated by exploiting the printed population yearbooks of Strasbourg. The project plans to develop, based on the DANIEL architecture, a tabular attention network (TAN) for extracting named entities from bilingual (French-German) handwritten tables. The system will first be pre-trained on synthetic data automatically generated from texts taken from Wikipedia and the Strasbourg city directories, reproducing the layout and handwritten style of the FDS files, and then refined on at least 5,000 real annotated files. A first version (TAN V1) will target simple files. A second version (TAN V2) will use Visual Question Answering technology to handle difficult cases (overlaps, annotations, etc.). This approach aims to extract specific elements (named entities) without having to read the entire irregular table, by querying the system via targeted ‘questions’ or requests.
The database created using deep learning methods will be enriched with geographical data (geocoding of addresses). A historical GIS will be developed, which will contain the geolocation of the city's street numbers over time. It will take into account changes in the built environment, street names and house numbers that the city has undergone over time.
Once the database has been finalised, we will initially use the analytical potential of the FDS to make original contributions to the demographic transition of urban populations. To this end, we have chosen to analyse the mechanisms behind changes in fertility and mortality in Strasbourg over the period 1871-1939 in terms of six cross-cutting dimensions (immigration, religion, socio-economic structure, intergenerational transfer, mobility).
On the computer science side, to date no study has been devoted to named entity recognition in semi structured handwritten tables. More generally, understanding semi-structured tables (either printed or handwritten) requires more effort from the research community. POPSTRAS will contribute to this ambition by the development of a generic VQA-based interactive extraction system for Table Understanding that will be made available open-source to the research community. The production of a new specific but challenging FDS annotated dataset (METS/ALTO, and JSON exports), will be another significant result of POPSTRAS that will be made publicly available to the research community (images + training annotations). As a side effect of the project, upon completion of this extraction process, a new Strasbourg’s Population Yearbook Database will be produced and made available to the SHS community (Relational database or plain csv exports will be made available). This database will serve as a secondary source to automatically control, correct and validate the extraction results of the system.
On the humanities and social sciences side, the POPSTRAS database will bring together in a single relational database a wealth of information that usually has to be extracted from different sources: civil status (births, marriages, deaths, divorces), censuses (characteristics of individuals and households) and population registers (migration and residential mobility). It will be longitudinal, allowing individuals to be followed in time and space over a long period (1871-1939), when Europe moved from a rural to a dominantly urban society. Its analytical potential will make Strasbourg an exceptional case study for the analysis of urban populations. The database will allows cross-sectional and longitudinal analyses of demographic phenomena (fertility, marriage and family dynamics, mortality, migration and residential mobility). Its completeness will allow detailed analyses and the study of sub-populations that are usually difficult to analyse statistically. Thanks to the precise location of individuals' residences, analyses will be possible at different spatial scales (building, street, neighbourhood, etc.). Researchers will be able to study the spatial differentiation of demographic phenomena in the city, their evolution and diffusion, and to measure «neighbourhood effects« on individual behaviours. The study of the individual life courses, of the interaction between life domains, like family, residential and occupational trajectories, will be possible, as well as intergenerational researches on sedentary families. With a data source documenting gender, marital status, place of birth, nationality, occupation and religion, the potential for differential analytical perspectives is enormous.
The POPSTRAS database will open up new research opportunities for historical demography and population studies, as well as for a wide range of humanities and social sciences. A wide range of topics in social, family and migration history can be addressed, such as social stratification, individual and intragenerational social mobility, social and religious homogamy, family ties and proximity in and outside the city, integration and assimilation of immigrants, etc.
The AI-oriented technology of POPSTRAS also opens the door to a wealth of new perspectives. Indeed, the collaboration with computer scientists and their recent progress in optical character and manuscript reading beckon new collaborations for the creation of other large databases, allowing us to link even more databases together and bring out new research questions. This will be an amazing breakthrough for work in quantitative history as it will greatly reduce data collection time while enabling the construction of much larger databases, allowing for much more refined statistical processing. The ability of the system to analyze semi-structured documents and to incorporate external sources of knowledge inside the information extraction process itself will open access to historical sources that had never been exploitable before.
Scientific dissemination: During the project, the team will publish articles on the methodological and technical innovations of the project. Members of the IT and H&SS teams will communicate at international conferences and jointly publish articles to inform both scientific communities of the progress made. We will pay particular attention to publishing in open access journals to ensure that the results of the project are widely disseminated. An end-of-project conference will be organised to disseminate the main results.
Data dissemination: we have drawn up a Data Management Plan (DMP). For the duration of the project, all images and data will be stored on the IR* Huma-Num servers. At the end of the project, the high-resolution FDS images will be made available to the archives of the city (Archives de la Ville et de l’Eurométropole de Strasbourg), which will distribute them free of charge on their website. This will meet a strong demand, particularly from the genealogy community. Currently, it is necessary to visit the archives in person to consult the FDS.
All the data (complete database, codes used for correction and recoding, historical GIS, etc.) and metadata from the project will be archived at IR* PROGEDO, which will ensure their dissemination to the scientific community. This process will make it possible to obtain a DOI for the data, guaranteeing greater visibility for the project.
The aim of this project is to build a large-scale database using a data source unprecedented in France, the Fichier Domiciliaire of the city of Strasbourg (1871-1939). The database makes it possible to track households over time and space, and to reconstruct their family and residential trajectories. Its richness will make it possible to carry out innovative analyses on a wide range of subjects. For that, we will be using innovative computer methods for automatic recognition of handwritten characters based on Deep Learning.
The urban populations of this period remain little studied, despite the fact that cities underwent major transformations (industrialisation, urbanisation). Historical demography has focused mainly on small populations (villages), and has done little analysis of the population of towns and cities because the data collection time was long. Recent advances in Deep Learning have made it possible to overcome these difficulties and exploit new sources of data.
While the source used in this project, largely unexploited, provides a unique opportunity to gain a better understanding of the urban population of this period, it also represents specific and original challenges for IT, particularly in terms of the automatic digitisation of historical sources. This corpus is complex to process, as there are specific difficulties linked to handwriting and the diversity of semi-tabular layouts (handwritten nature of the text, different styles of cursive writing in Latin and German, variations in the spatial layout of lines and information fields, sentences sometimes cramped or overflowing into neighbouring fields).
This project will be carried out in collaboration with a team of computer scientists specialising in Deep Learning for vision and language processing for automatic reading of documents, and a multidisciplinary team of researchers in the humanities and social sciences (demographers, historians, geographers).
Project coordination
Bénédicte GERARD (UNIVERSITÉ STRASBOURG)
The author of this summary is the project coordinator, who is responsible for the content of this summary. The ANR declines any responsibility as for its contents.
Partnership
SAGE UNIVERSITÉ STRASBOURG
LITIS UNIVERSITÉ ROUEN
IDEES IDENTITE ET DIFFERENCIATION DE L'ESPACE, DE L'ENVIRONNEMENT ET DES SOCIETES
Help of the ANR 597,914 euros
Beginning and duration of the scientific project:
November 2025
- 54 Months