HireFT
Browse JobsHow it worksPricingAboutSuccess Stories
    Back to jobs
    IN

    Inria

    Research and Development

    NLP Post-doc / Engineer for Information Mining in Historical Data (French 3rd Republic)

    Paris, FranceOn-SiteTemporaryPosted 1w ago
    All Inria jobs

    Job description

    NLP Post-doc / Engineer for Information Mining in Historical Data (French 3rd Republic)

    Download job offer in PDF format

    Contract type : Fixed-term contract

    Level of qualifications required : Graduate degree or equivalent

    Other valued qualifications : PhD

    Fonction : Temporary scientific engineer

    Level of experience : Recently graduated

    Context

    This position is part of the ANR project DECIDON (Débats parlementaires et Espace médiatique (1870-1940) : comprendre la CIrculation du discours politique grâce à des méthodes à forte intensité de DONnées), which is a collaboration between multiple French team, led at Epita Paris by Marie Puren, and including the Inria ALMAnaCH (Inria Paris center) project team as a work package lead for natural language processing techniques. The objective of the project is to enable a better understanding of how the parliamentary debates are being set up in the Third Republic and how thematics can appear and disappear across the newspaper and the parliamentary debate proceedings. The ALMAnaCH team role specifically focus on: 

    • Development of an annotation interface to support the creation of thematic datasets across the corpus of parliamentary debates and, potentially, the press.
    • Development of an easily deployable approach for robust information retrieval across the corpus, targeting topics whose vocabulary may have evolved over time and diverged from contemporary French.
    • Evaluation of information retrieval methods for topic-based search on a historically situated corpus (including RAG).

    The employee will be supervised by Thibault Clérice, permanent researcher at Inria, and will work in close collaboration with other members of the team and project, namely Florian Cafiero and Marie Puren (EPITA)

    They will work at Inria, within the ALMAnaCH project-team. Within the team, they will find researchers connected to the topic outside of the project itself, including Cecilia Graiff, a PhD student in the team, supervised by Benoît Sagot and Chloé Clavel, working on multilingual and cross-cultural automatic analysis of argumentation structures in political debates.

    Participation to national meetings and national/international conference are to be expected.

    Assignment

    The post-doc / NLP engineer will design and train models to support historians in constructing focused sub-corpora from large, noisy, OCR'd historical text collections. The work involves:

    • Setting up an annotation workflow (true/false positive labeling of keyword occurrences in context) in collaboration with historians;
    • Training and evaluating small-scale classifiers — ranging lightweight models to larger pretrained models — capable of distinguishing relevant from irrelevant occurrences of ambiguous or polysemous terms in diachrony (e.g., distinguishing an "anti-parliamentary" use of réforme de l'État from a routine administrative reform);
    • Integrating active learning so that model performance improves iteratively as historians annotate, minimizing labeling effort while maximizing corpus quality;
    • Prioritizing model deployability: given the scale of the corpus (at least dozens of millions of tokens across noisy OCR output) and the need for the tool to run efficiently and reproducibly within a web interface used directly by non-specialist historians, the research should focus on enabling this on lightway models: fast at inference, and easy to retrain/update rather than relying on large LLM inference at scale;
    • Benchmarking against and complementing RAG-based exploration (T5.5), providing a transparent, low-cost alternative for corpus-scoping that historians can audit and reproduce before moving to more exploratory or generative tasks.

    The research component focuses on efficient, low-resource sequence classification for historical/noisy text — including handling class imbalance, domain-shift across a ~70-year corpus, and figurative/contextual language. Publishable outputs (methods, benchmarks, and the resulting tool) are expected as part of the project's open, FAIR-data pipeline.

    Bibliography:

    El Assadi, A., Muennighoff, N., & Lee, J. (2026). The embedder's dilemma: LLMs are better, but at what cost? arXiv. https://doi.org/10.48550/arXiv.2608.12875

    Martin, L., Muller, B., Ortiz Suárez, P. J., Dupont, Y., Romary, L., Villemonte de la Clergerie, É., Seddah, D., & Sagot, B. (2020). CamemBERT: A tasty French language model. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.645

    van Strien, D., Beelen, K., Coll Ardanuy, M., Hosseini, K., McGillivray, B., & Colavizza, G. (2020). Assessing the impact of OCR quality on downstream NLP tasks. In A. Rocha, L. Steels, & J. van den Herik (Eds.), Proceedings of the 12th International Conference on Agents and Artificial Intelligence: ICAART 2020 (Vol. 1, pp. 484–496). SCITEPRESS. https://doi.org/10.5220/0009169004840496

    Luo, X., Shinnick, Z., Griesshaber, N., Wang, Y., Yu, J., Shi, F., Torr, P., & Lu, Y. (2026). Pretraining language models on historical text. arXiv. https://doi.org/10.48550/arXiv.2606.02991

    Bian, D., Puren, M., & Cafiero, F. (2026). How to efficiently explore noisy historical data? Leveraging corpus pre-targeting to enhance graph-based RAG. In D. Alves, Y. Bizzoni, S. Degaetano-Ortlieb, A. Kazantseva, J. Pagel, & S. Szpakowicz (Eds.), Proceedings of the 10th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature 2026 (pp. 241–250). Association for Computational Linguistics. https://doi.org/10.18653/v1/2026.latechclfl-1.23

    Clerice, T. (2024). Detecting sexual content at the sentence level in first millennium Latin texts. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) (pp. 4772–4783). ELRA and ICCL. https://aclanthology.org/2024.lrec-main.427/

    Main activities

    The main activities of the applicant will include:

    • carrying out research on the topic outlined above, both in the development of new ideas, positioning with respect to related work and validation of the methodology via experiments and analysis
    • Producing a solution that will be deployable for post-lexical information retrieval
    • Interact with the project’s historians as well as the researchers from the work package on natural language processing
    • the presentation of work both internally to colleagues and externally in the form of conference/journal/workshop papers
    • interacting and exchanging with colleagues on NLP topics

    Skills

    Technical skills and level required :

    • Python
    • NLP Model Training and evaluation, beyond just LLM

    Languages :

    • French (Reading understanding to interact with the data)
    • English

    Relational skills :

    • Good organizational skills.
    • Good interpersonal skills.

    Additional skills considered an asset:

    • Knowledge of the French 3rd Republic or similar political systems.
    • Interefest in information retrieval
    • A strong interest in open science.

    Benefits package

    • Subsidized meals
    • Partial reimbursement of public transport costs
    • Leave: 7 weeks of annual leave + 10 extra days off due to RTT (statutory reduction in working hours) + possibility of exceptional leave (sick children, moving home, etc.)
    • Possibility of teleworking and flexible organization of working hours
    • Professional equipment available (videoconferencing, loan of computer equipment, etc.)
    • Social, cultural and sports events and activities
    • Access to vocational training
    • Social security coverage
    Apply for this position
    Share
    • Facebook
    • Linkedin
    • Twitter
    • Email

    General Information

    • Theme/Domain : Language, Speech and Audio
      Data production, processing, analysis (BAP D)
    • Town/city : Paris
    • Inria Center : Centre Inria de Paris
    • Starting date : 2027-01-01
    • Duration of contract : 12 months
    • Deadline to apply : 2026-10-18

    Warning : you must enter your e-mail address in order to save your application to Inria. Applications must be submitted online on the Inria website. Processing of applications sent from other channels is not guaranteed.

    Instruction to apply

    Defence Security :
    This position is likely to be situated in a restricted area (ZRR), as defined in Decree No. 2011-1425 relating to the protection of national scientific and technical potential (PPST).Authorisation to enter an area is granted by the director of the unit, following a favourable Ministerial decision, as defined in the decree of 3 July 2012 relating to the PPST. An unfavourable Ministerial decision in respect of a position situated in a ZRR would result in the cancellation of the appointment.

    Recruitment Policy :
    As part of its diversity policy, all Inria positions are accessible to people with disabilities.

    Contacts

    • Inria Team : ALMANACH
    • Recruiter :
      Clerice Thibault / [email protected]

    The keys to success

    • Feeling comfortable in an interdisciplinary environment, as well as a willingness to learn and listen, are essential qualities for success in this role.
    • Interest in open science issues.
    • A PhD thesis or master's dissertation focusing on historical data or on information retrieval is an asset.
    • Interest in small models, beyond LLMs

    About Inria

    Inria, the French national institute for research in digital science and technology, supports the French government in national research and innovation strategies in the digital field, acting as Digital Programs Agency. Inria leads over 300 research and innovation projects with its 3,500 scientists, engineers, and support staff, in partnership with universities and the digital ecosystem (businesses, entrepreneurs, and public stakeholders). Together, we explore strategic fields such as artificial intelligence, cybersecurity, quantum computing, cloud technologies, digital transformation in healthcare, digital twins, and digital technologies for defence. We develop practical solutions such as software, tech startups, partnerships with national companies, and cutting-edge training programmes. Our goal is to drive scientific, technological, and industrial excellence to ensure France’s digital sovereignty.

    Job details are sourced from the employer's original posting.

    Open job posting
    IN

    About the company

    Inria

    INRIA is the French national research institute for digital science and technology.

    View all Inria jobs
    Industry
    Research and Development
    Open roles
    43

    Interested in this role?

    Apply with HireFT

    Free to start — no card required.

    Your fit

    How well do you match?

    Sign in to see how your résumé lines up with this role.