Title: Systems Architect, Disaster RecoveryLocation: RemoteDuration: 12+ MonthsPay Rate: $100-125/hr. on w2.Job Description:
Architectural Leadership:
- Define end to end DR and high availability (HA) architectures for enterprise-wide workloads, incorporating multi region cloud, hybrid, and on prem solutions.
- Develop architectural blueprints, reference designs, and pattern libraries that align with security, compliance, and cost optimization policies.
Solution Design & Implementation:
- Design and implement automated fail over, replication, and fail back mechanisms (e.g., Site Recovery Manager, Kubernetes based HA, database mirroring, storage level replication).
- Evaluate and integrate emerging technologies (e.g., Immutable Infrastructure, Chaos Engineering, Serverless DR) to improve resiliency and reduce mean time to recover (MTTR).
Governance & Compliance
- Ensure all DR solutions meet corporate policies CRX 301, CRX 302, and relevant regulatory requirements (e.g., NIST?800 34, ISO?22301, FedRAMP).
- Create and maintain DR documentation, run books, and test plans; conduct periodic reviews and updates.
Testing & Validation
- Lead full scale DR exercise planning, execution, and post mortem analysis for multi site, multi cloud environments.
- Define success criteria, metrics, and KPIs; report findings to senior leadership and stakeholders.
Stakeholder Collaboration
- Partner with IT Infrastructure, Cloud Engineering, Application Development, Security, and Governance teams to embed DR/HA considerations early in the SDLC.
- Serve as the technical authority for DR during design reviews (SRR, PDR, CDR, TRR) and program risk assessments.
Continuous Improvement
- Conduct risk assessments, threat modeling, and capacity planning to anticipate emerging resiliency challenges.
- Drive adoption of Model Based Systems Engineering (MBSE) and automated documentation tools to keep architecture artefacts current.
DR Plan Modernization & Compliance
- Review existing DR plan architectures across the enterprise, assessing their alignment with current resilience standards, best practices, and organizational Recovery Objectives.
- Collaborate with internal teams (Application Owners, IT Service Managers, Engineering) to update and refine DR plans, ensuring that all applications and IT services meet the latest RTO/RPO targets.
- Develop and implement remediation plans to bring legacy systems and applications up to date with modern resilience standards, ensuring compliance with corporate policies (CRX 301, CRX 302) and regulatory requirements.
- Track progress and report status to senior leadership, providing insights into plan modernization efforts and risk mitigation strategies.
Basic Qualifications :
- 5+?years of experience designing and implementing DR/HA solutions for enterprise scale workloads in cloud, hybrid, and on prem environments.
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related technical discipline (or equivalent experience).
- Proven hands on experience with cloud platforms (AWS, Azure, GCP) and related services (e.g., Disaster Recovery, Site Recovery Manager, cross region replication, networking, IAM).
- Strong understanding of networking, storage, virtualization, container orchestration (Kubernetes), and database technologies as they relate to resiliency.
- Excellent written and verbal communication skills; ability to translate complex technical concepts for both technical and non technical audiences.
Desired skills :
Familiarity with automated disaster recovery (DR) solutions, including but not limited to:
Amazon Web Services (AWS) Disaster Recovery Service (DRS): Experience with configuring and managing replication, fail over, and fail back processes for AWS workloads.
Microsoft Azure Site Recovery (ASR): Knowledge of setting up and managing site recovery between on premises environments, Azure, and other clouds.
Zerto: Hands on experience with continuous data protection (CDP) and near zero RPO replication across VMware, Hyper V, and cloud environments.
Veeam Backup & Replication: Experience with agent less backup, replication, and automated fail over testing for virtual, physical, and cloud workloads.
IBM Resiliency Services (formerly IBM Disaster Recovery as a Service): Familiarity with managed DR services for hybrid cloud environments, including integration with IBM Cloud and on premises infrastructure.
Experience with DR automation, including:
Scripting and integration with IaC tools (Terraform, CloudFormation) and CI/CD pipelines (Jenkins, GitLab) to automate DR workflows.
DR exercise planning and execution, including defining success criteria, metrics, and KPIs for recovery processes.
Strong analytical skills: ability to perform risk assessments, impact analysis, and cost benefit modeling for DR solutions.
Hybrid/multi-cloud deployments, with ability to manage DR across multiple cloud providers and on premises environments.
Advanced certifications (e.g., AWS Certified Solutions Architect – Professional, Azure Solutions Architect Expert, VMware VCAP DCV).