Experience with observability and monitoring stacks (e.g., OpenTelemetry, Datadog, Grafana/Prometheus) and, ideally, LLM-specific observability tools (e.g., LangSmith, Langfuse, Arize Phoenix, Braintrust, W&B Weave). Demonstrated experience building evaluation systems for ML or LLM applications: test harnesses, benchmark datasets, automated scoring (including LLM-as-judge approaches), and regression detection.