Director - Infra Engineering - Platform Security and Lifecycle Management

American Express Co
  • Phoenix, AZ
    30+ days ago

    Job Description

    American Express Enterprise Cloud team is looking for innovators to help us build world-class applications, Cloud platforms and infrastructure supported by integrated CICD, Observability and security capabilities.

    The Director Infrastructure Engineering - Platform Security and Lifecycle Management is responsible for leading the strategy, governance, and execution of Platform as a Service and Middleware services lifecycle management across Private Cloud environments. This role leverages Gen AI/Agentic AI to drive fully automated platform upgrades, security posture management, capacity management, and operational resilience to ensure a secure, scalable, highly available, and compliant Private Cloud platform supporting mission-critical business applications.

    The Director partners closely with Platform Engineering, Information Security, Architecture, Infrastructure, SRE, DevOps, and Application Development teams to drive platform modernization, automate operations, reduce technology risk, and enable cloud-native adoption at scale.

    At American Express, our culture is built on a 175-year history of innovation, shared values and Leadership Behaviors, and an unwavering commitment to back our customers, communities, and colleagues. From delivering differentiated products to providing world-class customer service, we operate with a strong risk mindset, ensuring we continue to uphold our brand promise of trust, security, and service.

    As part of Team Amex, you'll experience our powerful backing with comprehensive support for your holistic well-being and many opportunities to learn new skills, develop as a leader, and grow your career. Here, your voice and ideas matter, your work makes an impact, and together, you will help us define the future of American Express.

    • Bachelor's degree in Computer Science, Engineering, or related field (Master's preferred).

    • 8+ years of experience in Platform Engineering & Operations, Site Reliability Engineering (SRE), Platform lifecycle management with a proven track record of leading teams in managing large-scale cloud infrastructure with a focus on automation, reliability and resilience.

    • Deep hands-on experience with any Kubernetes platform(multi-cloud preferred).

    • Experience building end to end platform upgrade and fleet management automation leveraging Gen AI / Agentic AI

    • Strong experience with:

    • Infrastructure as Code (Terraform, CloudFormation, ARM)

    • Infrastructure automation tools like Ansible

    • Container platforms (OpenShift/Kubernetes)

    • Monitoring tools (Prometheus, OTEL, LOKI)

    • CI/CD pipelines (Jenkins, GitHub Actions)

    • Open source based messaging, caching, and database technologies like Kafka, Redis, Elastic

    • Strong understanding of cloud networking, security, and architecture.

    • Experience managing large-scale, mission-critical production environments.

    • Relevant certifications preferred

    • Experience with DevOps practices and methodologies, including CI/CD pipelines, configuration management, and infrastructure as code.

    • Experience with observability tools such as Prometheus, Splunk, ELK, Dynatrace.

    • Strong analytical and problem-solving skills, with the ability to troubleshoot complex issues and drive resolution in a fast-paced environment.

    • Excellent communication and leadership skills, with the ability to effectively collaborate with cross-functional teams and influence decision-making at all levels of the organization.

    Employment eligibility to work with American Express in the United States is required as the company will not pursue visa sponsorship for these positions.

    Private Cloud Platform Upgrade Strategy & Modernization

    • Define and execute the enterprise strategy and roadmap for Enterprise PaaS platform using Redhat Openshift and Data Middleware platform upgrades and lifecycle management.
    • Lead major version upgrades, cluster modernization, and infrastructure refresh initiatives across development, test, and production environments.
    • Establish standards, reference architectures, and best practices for platform lifecycle and security management.
    • Leverage Gen AI/Agentic AI to drive a fully automated pipeline for platform provisioning, upgrades, patching, and configuration management

    Upgrades & Release Management

    • Own end-to-end upgrade planning, governance, risk assessment, and execution.
    • Establish upgrade readiness processes, validation frameworks, rollback strategies, and post-upgrade monitoring.
    • Coordinate with application teams to ensure platform compatibility and minimize business disruption during upgrades.
    • Manage lifecycle risks associated with OpenShift, Kubernetes, operating systems, middleware, and supporting infrastructure.
    • Track and report platform currency and technology lifecycle compliance metrics.

    Security & Compliance

    • Establish and maintain a strong security posture across clusters and supporting infrastructure.
    • Lead vulnerability management, container image security, platform hardening, patch management, and remediation efforts.
    • Ensure compliance with enterprise security policies, regulatory requirements, and industry standards.
    • Partner with Information Security and Risk teams to manage security assessments, audits, and remediation activities.
    • Manage security exceptions and risk acceptance processes while driving long-term remediation strategies.

    Capacity & Performance Management

    • Own capacity planning and forecasting for Platform, including compute, memory, storage, and network resources.
    • Develop predictive capacity models to support business growth and application onboarding.
    • Establish monitoring and reporting processes for cluster utilization, performance, and scalability.
    • Optimize infrastructure consumption and platform efficiency while maintaining service-level objectives.
    • Lead resource optimization initiatives to improve workload density and reduce infrastructure costs.

    Leadership & Team Management

    • Build and lead high-performing cloud native DevOps engineers, Kubernetes administrators, security specialists, and capacity planners.
    • Foster a culture of operational excellence, continuous learning, innovation, and accountability.
    • Mentor leaders and technical experts within the organization.
    • Drive workforce planning, succession planning, and talent development initiatives.

    Stakeholder & Vendor Management

    • Serve as the senior technology leader for Private Cloud platform services.
    • Partner with application development, architecture, security, infrastructure, and business leaders to align platform capabilities with organizational priorities.
    • Manage relationships with key vendors, managed service providers, and strategic partners.
    • Present platform health, upgrade status, security posture, risks, and capacity forecasts to executive leadership.

    Private Cloud Platform Upgrade Strategy & Modernization

    • Define and execute the enterprise strategy and roadmap for Enterprise PaaS platform using Redhat Openshift and Data Middleware platform upgrades and lifecycle management.
    • Lead major version upgrades, cluster modernization, and infrastructure refresh initiatives across development, test, and production environments.
    • Establish standards, reference architectures, and best practices for platform lifecycle and security management.
    • Leverage Gen AI/Agentic AI to drive a fully automated pipeline for platform provisioning, upgrades, patching, and configuration management

    Upgrades & Release Management

    • Own end-to-end upgrade planning, governance, risk assessment, and execution.
    • Establish upgrade readiness processes, validation frameworks, rollback strategies, and post-upgrade monitoring.
    • Coordinate with application teams to ensure platform compatibility and minimize business disruption during upgrades.
    • Manage lifecycle risks associated with OpenShift, Kubernetes, operating systems, middleware, and supporting infrastructure.
    • Track and report platform currency and technology lifecycle compliance metrics.

    Security & Compliance

    • Establish and maintain a strong security posture across clusters and supporting infrastructure.
    • Lead vulnerability management, container image security, platform hardening, patch management, and remediation efforts.
    • Ensure compliance with enterprise security policies, regulatory requirements, and industry standards.
    • Partner with Information Security and Risk teams to manage security assessments, audits, and remediation activities.
    • Manage security exceptions and risk acceptance processes while driving long-term remediation strategies.

    Capacity & Performance Management

    • Own capacity planning and forecasting for Platform, including compute, memory, storage, and network resources.
    • Develop predictive capacity models to support business growth and application onboarding.
    • Establish monitoring and reporting processes for cluster utilization, performance, and scalability.
    • Optimize infrastructure consumption and platform efficiency while maintaining service-level objectives.
    • Lead resource optimization initiatives to improve workload density and reduce infrastructure costs.

    Leadership & Team Management

    • Build and lead high-performing cloud native DevOps engineers, Kubernetes administrators, security specialists, and capacity planners.
    • Foster a culture of operational excellence, continuous learning, innovation, and accountability.
    • Mentor leaders and technical experts within the organization.
    • Drive workforce planning, succession planning, and talent development initiatives.

    Stakeholder & Vendor Management

    • Serve as the senior technology leader for Private Cloud platform services.
    • Partner with application development, architecture, security, infrastructure, and business leaders to align platform capabilities with organizational priorities.
    • Manage relationships with key vendors, managed service providers, and strategic partners.
    • Present platform health, upgrade status, security posture, risks, and capacity forecasts to executive leadership.

    Numbers & Facts

    LocationPhoenix, AZ

    Skills

    • Analysis Skillsunmatched
    • Ansibleunmatched
    • Artificial Intelligence (AI)unmatched
    • Automationunmatched
    • Best Practicesunmatched
    • Business Growthunmatched
    • Business Solutionsunmatched
    • Business Supportunmatched
    • Capacity Managementunmatched
    • Capacity and Performance Managementunmatched
    • Cloud Applicationsunmatched
    • Cloud Computingunmatched
    • Communication Skillsunmatched
    • Computer Scienceunmatched
    • Computer Securityunmatched
    • Configuration Managementunmatched
    • Continuous Deployment/Deliveryunmatched
    • Continuous Integrationunmatched
    • Cost Controlunmatched
    • Cross-Functionalunmatched
    • Customer Support/Serviceunmatched
    • DevOpsunmatched
    • Enterprise Protectionunmatched
    • Fleet Managementunmatched
    • Forecastingunmatched
    • High Availabilityunmatched
    • Identify Issuesunmatched
    • Industry Standardsunmatched
    • Information Architectureunmatched
    • Information/Data Security (InfoSec)unmatched
    • Leadershipunmatched
    • Maintain Complianceunmatched
    • Memory Hardwareunmatched
    • Mentoringunmatched
    • Metricsunmatched
    • Middlewareunmatched
    • Network Securityunmatched
    • Onboardingunmatched
    • Operating Systemsunmatched
    • Platform as a Service (PaaS)unmatched
    • Predictive Modelingunmatched
    • Private Cloudunmatched
    • Problem Solving Skillsunmatched
    • Process Validationunmatched
    • Product Lifecycleunmatched
    • Production Systemsunmatched
    • Productivity Modelunmatched
    • Red Hat Linux Operating Systemunmatched
    • Regulatory Requirementsunmatched
    • Release Management/Engineeringunmatched
    • Reliability Engineeringunmatched
    • Riskunmatched
    • Risk Analysisunmatched
    • Risk Managementunmatched
    • Sales Pipelineunmatched
    • Security Analysisunmatched
    • Security Architectureunmatched
    • Security Auditingunmatched
    • Security Complianceunmatched
    • Security Infrastructureunmatched
    • Security Monitoringunmatched
    • Software Developmentunmatched
    • Software Engineeringunmatched
    • Software Patchesunmatched
    • Splunkunmatched
    • Succession Planningunmatched
    • Supplier Relationship Management (SRM)unmatched
    • Systems Administration/Managementunmatched
    • Talent Managementunmatched
    • Team Lead/Managerunmatched
    • Team Playerunmatched
    • Technical Leadershipunmatched
    • Technical Operationsunmatched
    • Test Plan/Scheduleunmatched
    • Vendor/Supplier Managementunmatched
    • Vendor/Supplier Relationsunmatched
    • Workforce Planningunmatched

    Be found by employers

    5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

    Level up your application

    Professional resume templates

    Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

    Free resume templates

    Free resume builder

    Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

    Free resume builder