Data Scientist — Agent Evaluations & Quality

Clera
  • Palo Alto, California
    Today

    Job Description

    About the Role

    This is an applied data science role focused on measuring, understanding, and improving the quality of AI agents that handle real-world tasks — email, calendar, browser, and business software. You'll sit at the intersection of evaluation design, statistics, and production engineering, building the feedback loops that directly guide how the product and engineering teams make decisions. Getting agent quality measurement right is core to how this product improves.

    What You'll Do

    • Architect and maintain automated evaluation pipelines that measure agent quality across capabilities and product surfaces.

    • Translate agent capabilities into explicit success criteria, including pass, partial-pass, and failure definitions for complex multi-step tasks.

    • Build representative gold datasets and regression suites covering common workflows, edge cases, long-tail behavior, and adversarial scenarios.

    • Define and track metrics such as task success, tool-selection accuracy, instruction adherence, factual consistency, latency, cost, and reliability.

    • Design deterministic and model-based graders; calibrate LLM-as-a-judge systems and measure false positives, false negatives, variance, and grader agreement.

    • Analyze traces, tool calls, model outputs, and production outcomes to identify root causes and build a useful failure taxonomy.

    • Compare models, prompts, tools, and capability implementations using rigorous offline experiments and production evidence.

    • Build dashboards and release-quality signals that make evaluation results actionable for engineering, product, and leadership.

    • Partner with capability engineers to verify that fixes improve quality without unacceptable regressions in cost, latency, or reliability.

    What We're Looking For

    • 5+ years in data science, machine learning, or analytics roles delivering evaluation systems, metrics frameworks, or quality measurement for production systems.

    • Demonstrated experience designing and implementing evaluation frameworks and grading systems for ML or AI systems in production.

    • Production-quality Python and SQL; ability to build automated data pipelines and analysis code at scale.

    • Strong evaluation methodology skills: success criteria definition, dataset construction, metric selection, and identifying misleading benchmarks.

    • Statistical and experimental design knowledge including sampling, variance, uncertainty quantification, bias detection, and significance testing for non-deterministic systems.

    • Experience with ground-truth data development: labeling guidelines, annotation quality control, ambiguity resolution, and dataset maintenance.

    • Working knowledge of LLM behavior, tool use, retrieval, multi-step execution, and practical failure modes of language model systems.

    • Ability to connect quantitative patterns to individual system traces and identify failure origins across model, prompt, context, tools, and application logic.

    • Experience building dashboards and communicating evaluation results, methodology, and trade-offs to both technical and non-technical stakeholders.

    • Familiarity with LLM-as-a-judge systems, agentic pipelines, or benchmarking platforms for AI is a strong plus.

    Location

    On-site in Palo Alto, CA. Visa sponsorship is not available for this role.

    Numbers & Facts

    LocationPalo Alto, California

    Skills

    • Analysis Skillsunmatched
    • Artificial Intelligence (AI)unmatched
    • Artificial Intelligence (AI) Agentsunmatched
    • Benchmarkingunmatched
    • Business Solutionsunmatched
    • Constructionunmatched
    • Data Analysisunmatched
    • Data Managementunmatched
    • Data Scienceunmatched
    • Data Setsunmatched
    • Experiment Designunmatched
    • Graderunmatched
    • Leadershipunmatched
    • Machine Learningunmatched
    • Metricsunmatched
    • Modeling Languagesunmatched
    • Product Documentationunmatched
    • Product Engineeringunmatched
    • Production Systemsunmatched
    • Python Programming/Scripting Languageunmatched
    • Quality Controlunmatched
    • Quality Managementunmatched
    • Quality Metricsunmatched
    • Quantitative Analysisunmatched
    • Reporting Dashboardsunmatched
    • Root Cause Analysisunmatched
    • SQL (Structured Query Language)unmatched
    • Statisticsunmatched
    • Taxonomiesunmatched
    • Web Browsersunmatched

    Be found by employers

    5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

    Level up your application

    Professional resume templates

    Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

    Free resume templates

    Free resume builder

    Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

    Free resume builder