Special Skill Requirements: (1) designing, building, and iterating large-scale scalable software systems; (2) designing and building ML recommender systems at high scale; (3) Delivering large and complex systems with significant business impact; (4) Cross-functional collaboration on large-scale projects with complex dependencies across teams; (5) Object-oriented programming (Python, Golang, or Java); (6) Building ML models with PyTorch or TensorFlow; (7) API design and integration with REST, HTTP, Thrift, or gRPC; (8) Working with largescale key-value and NoSQL storage and caching systems (Redis, Cassandra, DynamoDB); (9) Working with large-scale messaging and event-driven systems (Apache Kafka, AWS SQS, SNS, SES); (10) Building workflows using orchestration systems (Kubeflow, Ray, Apache Airflow, AWS Step Functions); (11) Working with large-scale analytics and data warehousing platforms (Google BigQuery, Amazon Redshift); (12) Experience with observability, logging, tracing and monitoring tools such as Prometheus, Grafana, AWS Cloudwatch; (13) Designing, running, and analyzing A/B experiments to measure and optimize system performance. Architect and implement end-to-end ML pipelines for notifications, including data ingestion, feature computation, model training, evaluation, deployment, and online serving in production environments.