HireFT
Browse JobsHow it worksPricingAboutSuccess Stories
    Back to jobs
    TR

    Transluce

    Technology

    AI Behavior Researcher - Agent Alignment

    San Francisco, United StatesOn-SiteFull-time$250k – $450k / yearPosted 4w ago
    All Transluce jobs

    Job description

    Salary range: $250,000 - $450,000/year + benefits
    Description: Transluce is a fast-moving nonprofit research lab building the public tech stack for AI evaluation and oversight. We have contributed foundational research to the study of AI agents and their behaviors, and are using these to study emerging issues in the honesty and alignment of AI agents. 
    About the role: As an AI Behavior Researcher, you will lead projects to design and develop automated evaluations of frontier AI systems that are technically sophisticated, scientifically valid, and concretely impactful. You will conduct novel analyses of behaviors related to agentic honesty and alignment. Example behaviors of interest include misreporting results, falsely claiming success, evaluation awareness, and memetic effects within AI swarms.
    As an early member of a highly collaborative team, you will learn and grow quickly, and work with our governance and infrastructure teams to scale your impact and technical reach.
    Core responsibilities:
    • Develop novel, valid automated evaluations of AI agents' honesty and alignment.
    • Write code to implement and run automated evaluations, such as environment simulators or LLM-as-a-judge pipelines.
    • Design methods to improve the ecological of automated evaluations, and especially to measure which effects are increasing or decreasing as models become more capable.
    • Collaborate with our governance team to deliver high-impact evaluations for public policy.
    • Collaborate with scientists and research engineers to productionize best practices in AI behavior evaluation.
    Minimum qualifications:
    • Expertise on quantitative generative AI evaluation and measurement. Good intuition about how to systematize and operationalize complex concepts and to work backwards from possible failure modes of agents.
    • Relevant experience designing and validating automated AI evaluation methods, such as LLM-as-a-judge systems or multi-turn benchmarks.
    • Proficiency in Python to implement analysis and evaluation tooling.
    • Meticulous, good experimental design, epistemic self-awareness and transparency.
    • Ability to iterate quickly and balance between scrappiness and thoroughness based on the impact needs of a project.
    • Strong communication skills, low ego, openness to giving and receiving feedback.
    Preferred qualifications (not required): 
    • Experience running automated evaluations at scale or in a production context.
    • Experience conducting controlled human subjects experiments to validate automated evaluation methods.
    • Experience in customer-facing, consulting, or forward-deployed roles translating ambiguous stakeholder needs into concrete deliverables.
    • Experience and comfort using AI coding agents at work.
    We are hiring at all levels of experience and would encourage those enthusiastic about the role who do not meet all of the qualifications to apply. We are located in San Francisco and excited to work together in-person. We are open to sponsoring international visas.

    Job details are sourced from the employer's original posting.

    Open job posting
    TR

    About the company

    Transluce

    Transluce is a company that provides solutions in the technology sector.

    View all Transluce jobs
    Industry
    Technology
    Open roles
    12

    Interested in this role?

    Apply with HireFT

    Free to start — no card required.

    Your fit

    How well do you match?

    Sign in to see how your résumé lines up with this role.