Working knowledge of generative AI models, AI agents, and related concepts such as retrieval augmented generation (RAG), prompt engineering, context engineering, explainability, traceability, observability, guard rails, reasoning, specificity, etc. Validates AI models and agents for accuracy, safety, bias, and performance through structured testing, benchmarking, and continuous evaluation pipelines and will be responsible for following: Build and maintain AI evaluation pipelines to test, measure, and evaluate the behavior and performance of AI systems.