Practical experience with AI/LLM evaluation frameworks (e.g., Ragas, DeepEval, LangSmith, Promptfoo, OpenAI Evals, TruLens) - building eval suites, scoring rubrics, and golden datasets. Working knowledge of eval metrics for generative AI: hallucination rate, faithfulness/groundedness, relevance, answer correctness, toxicity/bias scoring, BLEU/ROUGE/semantic similarity where applicable.