You will design the observability platform that ingests signals from building and electrical systems, server and network fabrics, Kubernetes, and GPU/accelerator clusters - then apply AI/ML models on top of that telemetry to optimize utilization, predict failures, reduce energy cost, and surface insights operators can act on. This role spans facilities/OT telemetry (cooling, power) and IT/AI infrastructure observability (compute, network, accelerators), unified by a single goal: complete, real-time, predictive visibility into how AI infrastructure consumes power, generates heat, moves data, and delivers compute.