Contract type : Fixed-term contract
Level of qualifications required : Graduate degree or equivalent
Other valued qualifications : PhD
Fonction : Temporary scientific engineer
Level of experience : Recently graduated
This position is part of the ANR project DECIDON (Débats parlementaires et Espace médiatique (1870-1940) : comprendre la CIrculation du discours politique grâce à des méthodes à forte intensité de DONnées), which is a collaboration between multiple French team, led at Epita Paris by Marie Puren, and including the Inria ALMAnaCH (Inria Paris center) project team as a work package lead for natural language processing techniques. The objective of the project is to enable a better understanding of how the parliamentary debates are being set up in the Third Republic and how thematics can appear and disappear across the newspaper and the parliamentary debate proceedings. The ALMAnaCH team role specifically focus on:
The employee will be supervised by Thibault Clérice, permanent researcher at Inria, and will work in close collaboration with other members of the team and project, namely Florian Cafiero and Marie Puren (EPITA)
They will work at Inria, within the ALMAnaCH project-team. Within the team, they will find researchers connected to the topic outside of the project itself, including Cecilia Graiff, a PhD student in the team, supervised by Benoît Sagot and Chloé Clavel, working on multilingual and cross-cultural automatic analysis of argumentation structures in political debates.
Participation to national meetings and national/international conference are to be expected.
The post-doc / NLP engineer will design and train models to support historians in constructing focused sub-corpora from large, noisy, OCR'd historical text collections. The work involves:
The research component focuses on efficient, low-resource sequence classification for historical/noisy text — including handling class imbalance, domain-shift across a ~70-year corpus, and figurative/contextual language. Publishable outputs (methods, benchmarks, and the resulting tool) are expected as part of the project's open, FAIR-data pipeline.
Bibliography:
El Assadi, A., Muennighoff, N., & Lee, J. (2026). The embedder's dilemma: LLMs are better, but at what cost? arXiv. https://doi.org/10.48550/arXiv.2608.12875
Martin, L., Muller, B., Ortiz Suárez, P. J., Dupont, Y., Romary, L., Villemonte de la Clergerie, É., Seddah, D., & Sagot, B. (2020). CamemBERT: A tasty French language model. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.645
van Strien, D., Beelen, K., Coll Ardanuy, M., Hosseini, K., McGillivray, B., & Colavizza, G. (2020). Assessing the impact of OCR quality on downstream NLP tasks. In A. Rocha, L. Steels, & J. van den Herik (Eds.), Proceedings of the 12th International Conference on Agents and Artificial Intelligence: ICAART 2020 (Vol. 1, pp. 484–496). SCITEPRESS. https://doi.org/10.5220/0009169004840496
Luo, X., Shinnick, Z., Griesshaber, N., Wang, Y., Yu, J., Shi, F., Torr, P., & Lu, Y. (2026). Pretraining language models on historical text. arXiv. https://doi.org/10.48550/arXiv.2606.02991
Bian, D., Puren, M., & Cafiero, F. (2026). How to efficiently explore noisy historical data? Leveraging corpus pre-targeting to enhance graph-based RAG. In D. Alves, Y. Bizzoni, S. Degaetano-Ortlieb, A. Kazantseva, J. Pagel, & S. Szpakowicz (Eds.), Proceedings of the 10th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature 2026 (pp. 241–250). Association for Computational Linguistics. https://doi.org/10.18653/v1/2026.latechclfl-1.23
Clerice, T. (2024). Detecting sexual content at the sentence level in first millennium Latin texts. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) (pp. 4772–4783). ELRA and ICCL. https://aclanthology.org/2024.lrec-main.427/
The main activities of the applicant will include:
Technical skills and level required :
Languages :
Relational skills :
Additional skills considered an asset:
Warning : you must enter your e-mail address in order to save your application to Inria. Applications must be submitted online on the Inria website. Processing of applications sent from other channels is not guaranteed.
Defence Security :
This position is likely to be situated in a restricted area (ZRR), as defined in Decree No. 2011-1425 relating to the protection of national scientific and technical potential (PPST).Authorisation to enter an area is granted by the director of the unit, following a favourable Ministerial decision, as defined in the decree of 3 July 2012 relating to the PPST. An unfavourable Ministerial decision in respect of a position situated in a ZRR would result in the cancellation of the appointment.
Recruitment Policy :
As part of its diversity policy, all Inria positions are accessible to people with disabilities.
Inria, the French national institute for research in digital science and technology, supports the French government in national research and innovation strategies in the digital field, acting as Digital Programs Agency. Inria leads over 300 research and innovation projects with its 3,500 scientists, engineers, and support staff, in partnership with universities and the digital ecosystem (businesses, entrepreneurs, and public stakeholders). Together, we explore strategic fields such as artificial intelligence, cybersecurity, quantum computing, cloud technologies, digital transformation in healthcare, digital twins, and digital technologies for defence. We develop practical solutions such as software, tech startups, partnerships with national companies, and cutting-edge training programmes. Our goal is to drive scientific, technological, and industrial excellence to ensure France’s digital sovereignty.
Job details are sourced from the employer's original posting.
Open job postingAbout the company
INRIA is the French national research institute for digital science and technology.