2MCR - 5828945
Required skills:
Assess cloud (AWS/Azure/GCP), on-prem, and vendor-hosted infrastructure against 8 NFR resilience domains with a focus on platform-layer gaps
Platform Evaluation: Evaluate networking, compute, storage, and platform configurations for resilience: load balancers, auto-scaling, multi-AZ/region, failover routing
DR / BCP Validation: Validate DR/BCP documentation, RTO/RPO targets, failover test evidence, and chaos engineering results for each assigned application
Observability Review: Assess observability maturity: alerting coverage, distributed tracing, log aggregation, SLO/SLA definitions, incident detection latency
Container & IaC: Evaluate container orchestration and IaC practices: Kubernetes health, Terraform state management, Helm releases, cluster-level resilience
BGC required
Nice to have skills:
10 15 years total IT / infrastructure / platform engineering
7+ years cloud infrastructure architecture (AWS, Azure, or GCP)
5+ years DR/BCP design, testing, and validation on enterprise systems
3+ years banking / financial services regulated environment
5+ years Kubernetes, Docker, container-native architecture
3+ years infrastructure-as-code (Terraform, Helm, Ansible)
3+ years observability tooling (Dynatrace, Splunk, Prometheus, Grafana)