This role requires strong technical judgment across the ML lifecycle (data training online inference monitoring), a strong understanding of how to enable AI applications to operate safely at scale, and a proven ability to build and lead a high performance team that operates production ML/AI systems with the same rigor as core infrastructure. Employ a diverse set of tools and platforms, including Python, AWS, Databricks, Docker, Kubernetes, Terraform, Snowflake, and GitHub, to guide your team in developing, deploying, and maintaining scalable and robust machine learning systems.