8-11 years in IT (Strong SDET) experience, including a minimum of 3 years dedicated to evaluating AI, LLM, and Agentic systems, preferably in a product-led organization or AI research lab. Define evaluation models for LLM and Agentic workflow safety and reliability, combining traditional QA foundations with next-gen AI testing paradigms to lead niche skills groups. Build automated evaluation frameworks for LLMs, RAG pipelines, and Agentic workflows. Implement "LLM-as-a-judge" benchmarks, HITL, and tools like LangSmith, TruLens, DeepEval, Phoenix, and Ragas into CI/CD. Deploy agentic frameworks (Perception/Reasoning/Action) in AWS, GCP, and Azure using containerization, service governance, and cloud orchestration. Validate functional output quality (hallucinations, relevance, tool-calling precision) and non-functional vectors (red-teaming, prompt injection, jailbreaking, toxicity, latency, cost). Mentor QA/SDETs into niche AI testing specialists, support presales, and drive market authority through articles and open-source benchmarks. Requires expert-level Java, Python, and Node.JS scripting alongside advanced hands-on automation with Selenium or Playwright, RestAssured, mobile testing, and strong manual testing foundations. AI/ML expertise must cover Transformer architectures, Vector DBs (Pinecone, Milvus), fine-tuning, and frameworks like LangChain, LlamaIndex, or AutoGen. Proven experience with BLEU, ROUGE, G-Eval, BERTScore, Context Precision, Faithfulness, MLflow, Weights & Biases, and cloud infrastructure-as-code. Requires a product mindset for testing non-deterministic systems, consultative exec-level communication, and management of onshore/offshore agile teams.
| Location | Jersey City, NJ |
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.
Free resume templatesImprove your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.
Free resume builder