Open to remote or Hybrid. Local to HDQ
Will serve as the primary hands-on contributors responsible for building the AI testing framework, developing automated AI evaluations, validating AI-enabled products, and providing continuous quality insights that improve the reliability and effectiveness of AI solutions across our DevEx portfolio.
Description:
We are seeking a highly motivated AI Test Engineer to help establish and scale our AI testing and evaluation capability. This role will work closely with AI engineers, platform teams, and product stakeholders to build automated testing frameworks, evaluate AI-driven solutions, and ensure the quality, reliability, and performance of AI-enabled products.
The AI Test Engineer will be hands-on contributors responsible for creating automated tests, executing evaluations against AI systems, analyzing model outputs, and helping define best practices for AI quality assurance. This role requires a strong combination of software testing expertise, automation engineering, data analysis, and an understanding of modern AI and generative AI technologies.
Key Responsibilities:
- Help design, develop, and maintain AI testing and evaluation frameworks for generative AI, predictive AI, and AI-powered developer experience (DevEx) solutions.
- Create and execute automated test suites that validate AI model behavior, output quality, accuracy, reliability, and performance.
- Develop automated evaluation pipelines that integrate with CI/CD workflows and support continuous AI quality validation.
- Perform functional, regression, performance, and scenario-based testing of AI-enabled applications and services.
- Analyze AI-generated outputs to identify defects, inconsistencies, hallucinations, bias, performance issues, and other quality concerns.
- Build and maintain test datasets, benchmark scenarios, prompt libraries, and evaluation artifacts used to assess AI systems.
- Collaborate with software engineers, AI engineers, and platform teams to improve model quality and testing coverage.
- Define, measure, and report AI quality metrics such as accuracy, relevance, latency, consistency, robustness, and user satisfaction.
- Support root-cause analysis, troubleshooting, and validation activities for AI-related defects and incidents.
- Document test strategies, test results, evaluation methodologies, and recommendations for technical and non-technical stakeholders.
- Contribute to AI governance and responsible AI efforts by supporting evaluation activities related to reliability, fairness, transparency, and safety.
- Support delivery of strategic DevEx initiatives and AI-enabled platform capabilities across the software development lifecycle.
Required Qualifications:
- Bachelor's degree in Computer Science, Information Systems, Software Engineering, Data Science, or a related technical field.
- 3+ years of experience in software testing, quality engineering, test automation, or AI/ML testing.
- Experience designing and executing automated test strategies for complex software applications.
- Strong programming and scripting skills in Python.
- Experience building automated test frameworks and test pipelines.
- Experience working with APIs, automated integration testing, and CI/CD environments.
- Understanding of machine learning, generative AI, large language models (LLMs), or AI-powered applications.
- Experience performing data validation, test result analysis, and defect investigation.
- Familiarity with modern testing methodologies, quality engineering practices, and software delivery lifecycles.
- Knowledge of statistical analysis and evaluation techniques used to validate system performance.
- Experience using source control systems and collaborative development practices.
- Strong analytical, problem-solving, and troubleshooting skills.
- Excellent written and verbal communication skills.
- Ability to work independently in a fast-paced, agile environment while collaborating effectively across teams.
Preferred Qualifications:
- Experience testing generative AI applications, LLM-based solutions, or Retrieval-Augmented Generation (RAG) systems.
- Experience with AI evaluation frameworks such as DeepEval, Ragas, LangSmith, or similar tools.
- Knowledge of LLM observability, monitoring, and traceability practices.
- Experience working with Azure AI, GitHub Copilot, OpenAI, Hugging Face, LangChain, or comparable AI platforms.
- Familiarity with prompt engineering, benchmark dataset creation, and AI output validation techniques.
- Understanding of Responsible AI principles and AI governance frameworks.
- Experience with performance testing, scalability testing, and reliability engineering.
- Experience supporting developer platforms, DevEx initiatives, or software engineering productivity tools.
- Experience working within enterprise or regulated technology environments.
Ideal Candidate Profile:
The ideal candidate is a hands-on quality engineer who enjoys building testing capabilities rather than simply executing predefined test cases. They are curious about how AI systems behave, comfortable analyzing large volumes of AI-generated outputs, and passionate about improving reliability through automation.
They bring a strong software testing foundation, practical automation experience, and a desire to help create scalable AI quality processes from the ground up. They thrive in collaborative environments, communicate findings clearly, and can translate evaluation results into actionable improvements for engineering teams.