AI Behavior Researcher - Human Impacts

Transluce
  • San Francisco, California
    30+ days ago

    Job Description

    Salary range: $250,000 - $450,000/year + benefits

    Description: Transluce is a fast-moving nonprofit research lab building the public tech stack for AI evaluation and oversight. We are pioneering research into the behaviors of AI chatbots and their effect on user wellbeing, and we’re improving outcomes for millions of sensitive AI interactions with vulnerable users. 

    About the role: As an AI Behavior Researcher, you will lead projects to design and develop automated evaluations of frontier AI systems that are technically sophisticated, scientifically valid, and concretely impactful. This includes expanding on our existing evaluation pipelines to conduct novel analyses of AI behaviors that affect the autonomy and wellbeing of specific user groups (e.g., children or users located in countries beyond the United States).

    As an early member of a highly collaborative team, you will learn and grow quickly, and work directly with frontier labs to improve AI evaluations design, with regulators to improve independent oversight of AI, and with domain experts and affected populations to enhance the realism and relevance of our evaluations. 

    Core responsibilities: 
    • Develop novel, valid automated evaluations of AI’s impacts on users, including their mental health and decision making.
    • Write code to implement and run automated evaluations, such as user simulators or LLM-as-a-judge pipelines.
    • Design methods to improve the ecological validity and realism of automated evaluations for specific populations, such as customizing existing user simulation methods to capture the vocabulary used by children.
    • Write and revise judge rubrics to evaluate model behaviors related to user wellbeing and decision making, systematizing abstract, socially situated concepts into clear measurement criteria.
    • Collaborate with scientists and research engineers to productionize best practices in AI behavioral evaluation.

    Minimum qualifications:
    • Expertise on quantitative generative AI evaluation and measurement. Good intuition about how to systematize and operationalize complex social concepts.
    • Relevant experience designing and validating automated AI evaluation methods, such as LLM-as-a-judge systems or multi-turn benchmarks.
    • Proficiency in Python to implement analysis and evaluation tooling.
    • Meticulous, good experimental design, epistemic self-awareness and transparency.
    • Ability to balance between the needs of AI researchers and domain experts, as well as between researchers and senior decision makers.
    • Strong communication skills, low ego, openness to giving and receiving feedback.

    Preferred qualifications (not required): 
    • Experience running automated evaluations at scale or in a production context.
    • Experience conducting controlled human subjects experiments to validate automated evaluation methods.
    • Experience in customer-facing, consulting, or forward-deployed roles translating ambiguous stakeholder needs into concrete deliverables.
    • Experience or training in human-centered design or HCI research methods, including working with domain experts or impacted communities.
    • Experience or demonstrated interest in studying AI’s psychological or social impacts, such as for crisis support, manipulation or sycophancy, political persuasion, or displacing human relationships.
    • Experience designing multilingual generative AI evaluations.
    • Experience and comfort using AI coding agents at work.

    We are hiring at all levels of experience and would encourage those enthusiastic about the role who do not meet all of the qualifications to apply. We are located in San Francisco and excited to work together in-person. We are open to sponsoring international visas.


    Numbers & Facts

    LocationSan Francisco, California

    Skills

    • Analysis Skillsunmatched
    • Artificial Intelligence (AI)unmatched
    • Artificial Intelligence (AI) Agentsunmatched
    • Benchmarkingunmatched
    • Best Practicesunmatched
    • Communication Skillsunmatched
    • Computer Programmingunmatched
    • Concreteunmatched
    • Consultingunmatched
    • Conversation Engineunmatched
    • Customer Relationsunmatched
    • Customer/Consumer Behaviorunmatched
    • Experiment Designunmatched
    • Fundingunmatched
    • Human-Computer Interactionunmatched
    • Machine Toolunmatched
    • Multilingualunmatched
    • Nonprofitunmatched
    • Persuasion Skillsunmatched
    • Project Designunmatched
    • Psychiatry and Mental Healthunmatched
    • Psychologyunmatched
    • Python Programming/Scripting Languageunmatched
    • Quantitative Analysisunmatched
    • Research Laboratoryunmatched
    • Research Skillsunmatched
    • Scientific Researchunmatched
    • Simulationunmatched
    • Team Playerunmatched
    • User Groupsunmatched

    Be found by employers

    5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

    Level up your application

    Professional resume templates

    Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

    Free resume templates

    Free resume builder

    Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

    Free resume builder