Job Description AI Evaluation Engineer
Location: USA RemoteEmployment Type: Contract
Position Summary
We are seeking an AI Evaluation Engineer to establish and maintain the quality standards used to determine whether production AI and LLM systems are ready to launch.
This role will build centralized evaluation frameworks, create golden datasets, establish measurable quality thresholds, audit AI evaluation results, and implement continuous production monitoring. The successful candidate will bring strong technical evaluation expertise along with the independence and analytical rigor required to objectively determine whether AI systems meet production quality and safety standards.
Key Responsibilities
Design and build centralized LLM/AI evaluation frameworks and reusable evaluation templates .
Establish evaluation standards that can be adopted consistently across multiple AI engineering teams.
Create and maintain golden datasets in partnership with business, domain, and subject-matter experts.
Develop evaluation suites covering functional quality, accuracy, safety, reliability, and domain-specific requirements.
Implement and assess LLM-as-a-judge evaluation approaches while accounting for their limitations and failure modes.
Define measurable pass/fail thresholds and production-readiness criteria for AI applications.
Independently review and audit evaluation suites developed by individual engineering teams.
Build regression testing approaches for LLM applications, agents, prompts, retrieval systems, and model changes.
Establish production monitoring for AI quality degradation, drift, regressions, and incidents.
Analyze and report evaluation pass rates, quality trends, regressions, and production incidents.
Partner with AI engineers, platform engineers, product teams, and domain experts to continuously improve AI quality.
Required Qualifications
4+ years of experience in ML/LLM evaluation, AI quality engineering, applied research engineering, or related areas .
Hands-on experience designing and implementing AI/LLM evaluation frameworks.
Experience developing golden/reference datasets and measurable evaluation criteria.
Strong understanding of LLM-as-a-judge methodologies and associated failure modes .
Experience evaluating LLM applications, agents, RAG systems, prompts, or other probabilistic AI systems.
Strong statistical and analytical skills, including evaluation methodologies for relatively small sample sizes.
Experience establishing quality thresholds and regression criteria.
Strong engineering skills with the ability to build reusable evaluation tooling and automation.
Ability to independently assess system quality and challenge release decisions when evaluation evidence does not meet established standards.
Preferred Qualifications
Healthcare, clinical, or other safety-critical AI evaluation experience.
Experience designing medical-accuracy or domain-specific evaluation suites.
AI red-teaming or adversarial testing experience.
Experience with production AI monitoring, drift detection, and continuous evaluation.
Show more Numbers & Facts Location NULL, NJ (Remote ) Salary $70–$75
Skills
Accountingunmatched
Analysis Skillsunmatched
Artificial Intelligence (AI)unmatched
Automationunmatched
Continuous Improvementunmatched
Data Setsunmatched
Engineeringunmatched
Healthcareunmatched
Machine Toolunmatched
Production Controlunmatched
Quality Engineeringunmatched
Quality Managementunmatched
Quality Metricsunmatched
Quality Monitoringunmatched
Safety Standardsunmatched
Safety/Work Safetyunmatched
Systems Analysisunmatched
Technical Analysisunmatched
Trend Analysisunmatched
Show more Level up your application Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.
Free resume templates
Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.
Free resume builder