This role sits at the intersection of large language models, agentic systems, and legal workflows, with a focus on building rigorous evaluation frameworks that measure, benchmark, and advance AI performance across complex legal tasks. Design, own, and evolve evaluation frameworks for AI agents operating in legal domains, including benchmark suites, scoring methodologies, quality rubrics, and research-grade evaluation protocols.