Infrastructure Engineer III

American Express Co
  • Phoenix, AZ
    17 days ago

    Job Description

    Joining Amex Tech means discovering and shaping your contribution to something big. Here, you can work alongside talented tech teams and build a unique career with the Powerful Backing of American Express. With a range of opportunities to work with the latest technologies, and a commitment to back the broader engineering community through open source, our mission is to power your success. Because Amex Tech is powered by our technology, our culture, and our colleagues.

    The Technology organization enables and accelerates the company's growth strategies, delivering global capabilities and services in support of Amex's customers and colleagues, while maintaining 24/7 servicing and availability to ensure an uninterrupted, high-quality customer experience. Technology provides the foundation for everything we do in the company while driving differentiation through building and leveraging innovative technology and data insights.

    At American Express, our culture is built on a 175-year history of innovation, shared values and Leadership Behaviors, and an unwavering commitment to back our customers, communities, and colleagues. From delivering differentiated products to providing world-class customer service, we operate with a strong risk mindset, ensuring we continue to uphold our brand promise of trust, security, and service.

    As part of Team Amex, you'll experience our powerful backing with comprehensive support for your holistic well-being and many opportunities to learn new skills, develop as a leader, and grow your career. Here, your voice and ideas matter, your work makes an impact, and together, you will help us define the future of American Express.

    • Overall 5+ years of IT experience, with 3+ years of hands-on experience administering Splunk, Grafana, and/or Dynatrace in large-scale enterprise environments.
    • Strong expertise in observability concepts, including logs, metrics, traces, APM, infrastructure monitoring, and synthetic monitoring.
    • Proven experience with Splunk indexers, forwarders, search heads, clustering, and performance optimization.
    • Hands-on experience with Grafana dashboard development and data sources such as Prometheus, Loki, Elasticsearch, and cloud-native monitoring tools.
    • Strong Dynatrace experience, including OneAgent deployment, service monitoring, custom metrics, and AI-driven root cause analysis.
    • Experience integrating observability platforms with cloud infrastructure, Kubernetes, and microservices-based architectures.
    • Proficiency in scripting, automation, and API integrations (Python, Shell, REST APIs, Terraform, Ansible).
    • Knowledge of security, compliance, and governance practices in monitoring platforms (SOC2, ISO, NIST, audit requirements).
    • Relevant certifications (Splunk Certified Admin/Architect, Dynatrace Associate/Professional, Grafana certifications) preferred.
    • Strong communication skills with the ability to collaborate across engineering, SRE, and leadership teams.

    Depending on factors such as business unit requirements, the nature of the position, cost and applicable laws, American Express may provide visa sponsorship for certain positions

    • Administer, configure, upgrade, and optimize Splunk, Grafana, and Dynatrace platforms, ensuring high availability, security, scalability, and performance.
    • Manage end-to-end observability solutions, including log aggregation (Splunk), metrics and visualization (Grafana), and application performance monitoring (Dynatrace).
    • Design, implement, and maintain monitoring, alerting, and dashboarding standards across infrastructure, applications, and cloud platforms.
    • Perform platform upgrades, patching, and lifecycle management with minimal downtime and adherence to enterprise change management processes.
    • Integrate Splunk, Grafana, and Dynatrace with cloud services (AWS, Azure, GCP), container platforms (Kubernetes, OpenShift), and CI/CD pipelines.
    • Configure and manage data ingestion pipelines, indexes, retention policies, parsing rules, and performance tuning for large-scale log and metric volumes.
    • Implement Dynatrace OneAgent deployments, SmartScape topology mapping, service flow monitoring, and root cause analysis (Davis AI).
    • Develop and maintain Grafana dashboards using Prometheus, Loki, InfluxDB, and other supported data sources.
    • Enforce security, access controls, and compliance requirements (SSO/SAML, RBAC, audit logging, data masking, retention policies).
    • Automate operational tasks using scripting and APIs (Python, Shell, REST APIs, Terraform, Ansible).
    • Troubleshoot platform issues, performance bottlenecks, ingestion delays, alert noise, and data gaps across observability tools.
    • Collaborate with application, infrastructure, and SRE teams to improve monitoring coverage, incident response, and operational resilience.
    • Document platform standards, best practices, and provide guidance or training to engineering and operations team.
    • Administer, configure, upgrade, and optimize Splunk, Grafana, and Dynatrace platforms, ensuring high availability, security, scalability, and performance.
    • Manage end-to-end observability solutions, including log aggregation (Splunk), metrics and visualization (Grafana), and application performance monitoring (Dynatrace).
    • Design, implement, and maintain monitoring, alerting, and dashboarding standards across infrastructure, applications, and cloud platforms.
    • Perform platform upgrades, patching, and lifecycle management with minimal downtime and adherence to enterprise change management processes.
    • Integrate Splunk, Grafana, and Dynatrace with cloud services (AWS, Azure, GCP), container platforms (Kubernetes, OpenShift), and CI/CD pipelines.
    • Configure and manage data ingestion pipelines, indexes, retention policies, parsing rules, and performance tuning for large-scale log and metric volumes.
    • Implement Dynatrace OneAgent deployments, SmartScape topology mapping, service flow monitoring, and root cause analysis (Davis AI).
    • Develop and maintain Grafana dashboards using Prometheus, Loki, InfluxDB, and other supported data sources.
    • Enforce security, access controls, and compliance requirements (SSO/SAML, RBAC, audit logging, data masking, retention policies).
    • Automate operational tasks using scripting and APIs (Python, Shell, REST APIs, Terraform, Ansible).
    • Troubleshoot platform issues, performance bottlenecks, ingestion delays, alert noise, and data gaps across observability tools.
    • Collaborate with application, infrastructure, and SRE teams to improve monitoring coverage, incident response, and operational resilience.
    • Document platform standards, best practices, and provide guidance or training to engineering and operations team.

    Numbers & Facts

    LocationPhoenix, AZ

    Skills

    • Access Controlunmatched
    • Amazon Web Services (AWS)unmatched
    • Ansibleunmatched
    • Application Programming Interface (API)unmatched
    • Artificial Intelligence (AI)unmatched
    • Best Practicesunmatched
    • Change Managementunmatched
    • Cloud Applicationsunmatched
    • Cloud Computingunmatched
    • Communication Skillsunmatched
    • Continuous Deployment/Deliveryunmatched
    • Continuous Integrationunmatched
    • Customer Experienceunmatched
    • Customer Support/Serviceunmatched
    • Data Managementunmatched
    • Documentation Standardsunmatched
    • Elasticsearchunmatched
    • Emerging Technologyunmatched
    • Forwarderunmatched
    • GCP (Good Clinical Practices)unmatched
    • High Availabilityunmatched
    • ISO (International Organization for Standardization)unmatched
    • Identify Issuesunmatched
    • Incident Responseunmatched
    • Information Technology & Information Systemsunmatched
    • Leadershipunmatched
    • Metricsunmatched
    • Microservicesunmatched
    • Microsoft Windows Azureunmatched
    • Open Sourceunmatched
    • People Managementunmatched
    • Performance Analysisunmatched
    • Performance Tuning/Optimizationunmatched
    • Python Programming/Scripting Languageunmatched
    • REST (Representational State Transfer)unmatched
    • Regulatory Complianceunmatched
    • Reporting Dashboardsunmatched
    • Riskunmatched
    • Root Cause Analysisunmatched
    • Sales Pipelineunmatched
    • Scripting (Scripting Languages)unmatched
    • Security Assertion Markup Language (SAML)unmatched
    • Single Sign-On (SSO)unmatched
    • Software Patchesunmatched
    • Splunkunmatched
    • Systems Administration/Managementunmatched
    • Team Playerunmatched
    • Topologyunmatched
    • Training/Teachingunmatched
    • U.S. National Institute of Standards and Technology (NIST)unmatched
    • Unix Shell Programmingunmatched

    Be found by employers

    5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

    Level up your application

    Professional resume templates

    Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

    Free resume templates

    Free resume builder

    Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

    Free resume builder