Overview:
This is a hybrid role - 2 days remote and 3 days in the Malvern, PA office.
We are seeking a highly experienced Senior Site Reliability & Cloud Systems Engineer to architect, build, automate, and operate scalable, secure, resilient, and highly available cloud platforms in AWS. This role combines hands-on reliability engineering with cloud architecture and automation expertise, with a strong emphasis on building immutable infrastructure and improving system resilience.
You will play a critical role in evolving our AWS ecosystem into a fully automated, self-service, “push-button” platform, minimizing manual operational intervention while improving reliability, security, performance, scalability, and engineering velocity. You will establish and champion engineering standards for immutable infrastructure, automation, observability, resilience, and operational excellence.
This role is well suited for a senior engineer who thrives at the intersection of SRE, cloud architecture, platform engineering, DevOps, and systems engineering, and who can independently drive complex technical initiatives from architecture and design through implementation and production operation.
Who we are:
At CubeSmart, we’re intentional about culture. You can experience it everywhere from our mission statement of “genuine care” to our “It’s What’s Inside That Counts” tagline to calling each other “teammates” rather than employees. This spirit fosters a fun and collaborative environment that has resulted in our rapid growth and being recognized amongst the top in our industry.
CubeSmart’s award-winning team is made up of people who genuinely care. Teammates care about our customers and the life events and/or business needs they are facing. Teammates are passionate, responsible and understanding. The CubeSmart team is made up of people who have a can-do attitude, are committed to their own success and the success of the company, and lead by example.
If this sounds like a team and culture that matches your personal values and motivations, we want to hear from you.
Responsibilities:
Reliability, Performance & Production Operations
- Own the reliability, availability, scalability, performance, and operational health of AWS-hosted, Linux-based production platforms and associated lower environments.
- Define and drive SRE practices, reliability standards, SLOs, SLIs, error budgets, and operational maturity across critical systems.
- Architect and implement comprehensive observability using platforms such as Datadog, Amazon CloudWatch, Prometheus/Grafana, and PagerDuty.
- Establish proactive monitoring, alerting, capacity planning, performance engineering, and predictive reliability practices.
- Lead complex production incident response, including technical triage, mitigation, service restoration, root cause analysis, and executive/stakeholder communication.
- Drive blameless post-incident reviews and ensure corrective actions are translated into measurable reliability improvements.
- Participate in and provide leadership during 24/7 on-call operations, including escalation management for high-severity production incidents.
- Identify systemic reliability risks and proactively eliminate single points of failure, operational bottlenecks, and sources of technical debt.
- Develop and implement disaster recovery, business continuity, backup, restoration, and resilience strategies.
- Perform capacity and performance analysis for distributed applications and infrastructure at scale.
Cloud Architecture & Automation
- Architect and implement highly automated, ephemeral, immutable, and reproducible AWS environments across production and non-production workloads.
- Lead the design of scalable, fault-tolerant, secure distributed systems using AWS Well-Architected principles.
- Establish infrastructure patterns that enable engineering teams to provision and manage environments through self-service and “push-button” automation.
- Eliminate manual infrastructure operations through Infrastructure as Code using Terraform, Ansible, Packer, and related technologies.
- Design and maintain reusable infrastructure modules, automation frameworks, and engineering standards.
- Build and evolve CI/CD and GitOps workflows using technologies such as Jenkins, GitHub Actions, GitLab CI, ArgoCD, and Flux.
- Develop sophisticated automation and operational tooling using Python and Bash.
- Identify opportunities to reduce operational toil through automation, platform engineering, and intelligent operational workflows.
- Establish engineering patterns for immutable infrastructure, automated provisioning, blue/green deployments, canary releases, and automated rollback.
Infrastructure & Platform Engineering
- Architect, deploy, and operate AWS services including:
- EKS, ECS, Fargate, Lambda
- RDS and Aurora PostgreSQL
- OpenSearch
- Redis and ElastiCache
- Load balancing, networking, storage, compute, and supporting AWS services
- Design and manage enterprise AWS networking architectures, including Transit Gateways, VPCs, routing, security groups, network segmentation, load balancers, and service meshes.
- Architect Kubernetes and container platforms for highly available production workloads.
- Establish standardized platform capabilities that enable development teams to deploy applications safely and independently.
- Design and implement caching, asynchronous processing, microservices, service discovery, and other distributed-system patterns.
- Evaluate emerging AWS and cloud-native technologies and make recommendations based on reliability, security, scalability, operational complexity, and cost.
- Provide technical leadership for major infrastructure modernization and cloud transformation initiatives.
Security, Risk & Governance
- Design and implement zero-trust cloud security architectures across AWS environments.
- Establish secure identity and access patterns using IAM, Organizations, SCPs, OIDC, KMS, Secrets Manager, and related AWS security services.
- Implement least-privilege access models and automated security controls across infrastructure and deployment pipelines.
- Embed security into CI/CD and Infrastructure as Code workflows using SAST, DAST, SCA, and vulnerability-management tools, including technologies such as Snyk.
- Implement automated compliance auditing, configuration validation, security scanning, backup verification, and governance controls.
- Partner with security and compliance teams to establish cloud security standards and ensure infrastructure meets organizational and regulatory requirements.
- Secure secrets, credentials, certificates, and sensitive configuration using technologies such as HashiCorp Vault and Ansible Vault.
- Continuously assess infrastructure for security vulnerabilities and proactively remediate risks.
Platform Engineering, Leadership & Strategy
- Serve as a senior technical authority for cloud infrastructure, reliability engineering, automation, and platform architecture.
- Mentor engineers and establish best practices for cloud engineering, SRE, DevOps, Infrastructure as Code, observability, and operational excellence.
- Partner closely with development, security, architecture, and operations teams to design reliable and scalable application platforms.
- Influence architectural decisions and provide technical guidance on complex infrastructure and distributed-system challenges.
- Establish and maintain engineering standards, reference architectures, reusable patterns, and operational guidelines.
- Create and maintain comprehensive architecture documentation, runbooks, troubleshooting guides, disaster recovery procedures, and operational standards.
- Drive FinOps initiatives, including cloud cost optimization, resource utilization, capacity forecasting, rightsizing, and infrastructure efficiency.
- Integrate infrastructure and operational changes into ITIL-compliant change-management and service-management processes, including platforms such as Freshservice.
- Identify and prioritize opportunities to improve reliability, automation, security, developer experience, and operational efficiency.
- Lead cross-functional technical initiatives and drive complex projects from requirements and architecture through implementation and production adoption.
Qualifications:
- 10+ years of professional experience in Site Reliability Engineering, DevOps, Cloud Engineering, Infrastructure Engineering, Platform Engineering, or a closely related discipline.
- Demonstrated experience designing, implementing, and operating enterprise-scale AWS cloud environments.
- Deep hands-on expertise with AWS architecture, services, networking, security, and operational best practices.
- Extensive Linux systems engineering experience, preferably with Ubuntu.
- Advanced experience with Infrastructure as Code, particularly Terraform and Ansible.
- Proven ability to design and implement highly automated, immutable, and reproducible infrastructure.
- Extensive experience designing and maintaining enterprise CI/CD pipelines and deployment automation.
- Strong programming and automation capabilities using Python and Bash.
- Extensive experience with production observability, monitoring, logging, alerting, and incident-management platforms.
- Strong understanding of distributed systems, scalability, fault tolerance, high availability, and performance engineering.
- Demonstrated experience leading production incident response and conducting root cause analysis.
- Strong understanding of cloud security principles, including IAM, KMS, secrets management, least privilege, OIDC, and zero-trust architectures.
- Experience implementing security controls and automated governance within cloud infrastructure and CI/CD pipelines.
- Demonstrated ability to independently lead complex technical initiatives and make sound architectural decisions.
- Strong written and verbal communication skills, including the ability to communicate technical risks and decisions to both engineering and non-engineering stakeholders.
- Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related field, or equivalent practical experience.
- Candidates must be authorized to work in the U.S. without the need for current or future sponsorship.
Preferred Qualifications
- Extensive experience with Docker, Kubernetes, Amazon EKS, ECS, and Fargate.
- Experience designing and operating Kubernetes platforms at enterprise scale.
- Advanced experience with GitOps, including ArgoCD or Flux.
- Experience implementing SAST, DAST, SCA, and secure SDLC practices.
- Strong knowledge of distributed systems, microservices, caching, messaging, and asynchronous architectures.
- Experience designing highly available and geographically resilient cloud architectures.
- Experience with service meshes and cloud-native networking.
- Experience implementing automated disaster recovery and multi-region resilience strategies.
- Experience with FinOps, cloud cost optimization, and capacity forecasting.
- Experience with ITIL processes, change management, and enterprise service-management platforms.
- Experience mentoring engineers and establishing technical standards across engineering teams.
- AWS professional-level certification or equivalent demonstrated expertise.
- Experience leading large-scale cloud migrations, infrastructure modernization, or platform transformation initiatives.
#LI-MT1