Build batch and streaming ingestion/ETL/ELT pipelines using AWS Glue (jobs, crawlers) Implement data cataloging, governance, and access control, in collaboration with governance/security stakeholders, using the Glue Data Catalog and Lake Formation (finegrained, cross-team permissions), and optionally DataZone for data product publishing/discovery across science teams Enable self-service analytics and data access for scientists across the Science Office via Athena including query patterns, cost controls, and onboarding documentation. AWS Lambda and Step Functions for orchestration and event-driven workflows S3 as the primary data/artifact layer, with lifecycle and access policies CloudWatch (and related tooling) for logging, metrics, and alerting IAM & VPC design following least-privilege and network security best practices Infrastructure as Code via AWS CDK or CloudFormation Implement CI/CD pipelines for ML (data, model, and code), including automated testing, packaging, and promotion of models across dev/staging/production environments.