Ability to build measurable frameworks assessing response quality, agent behavior, tool-selection accuracy, and regression risk - golden datasets, scenario suites, model-as-judge scoring with human calibration, continuous evaluation pipelines, and drift detection - plus adversarial, jailbreak, and grounding testing. Build experiences that hold context across turns, hand off cleanly between automated and human agents, and behave consistently across channels - accounting for what voice imposes: latency budgets, barge-in, speech recognition error, disambiguation, and confirmation before consequential actions.