Observability Lead-Sr. Infrastructure Engineer

Truist Bank
  • Charlotte, North Carolina
    2 days ago

    Job Description

    The position is described below. If you want to apply, click the Apply Now button at the top or bottom of this page. After you click Apply Now and complete your application, you'll be invited to create a profile, which will let you see your application status and any communications. If you already have a profile with us, you can log in to check status.

    Need Help?

    If you have a disability and need assistance with the application, you can request a reasonable accommodation. Send an email to

    careers@truist.com?subject=Accommodation%20request

    (accommodation requests only; other inquiries won't receive a response).

    Regular or Temporary:

    Regular

    Language Fluency:  English (Required)

    Work Shift:

    1st shift (United States of America)

    Please review the following job description:

    We are seeking a highly experienced and strategic Observability Lead-Sr. Infrastructure Engineer with a strong Forward Deployed Engineering (FDE), software development, event streaming, and technical leadership focus to define, implement, and scale enterprise observability capabilities. This role combines technical leadership in observability strategy with hands-on, customer-facing engineering, application development expertise, and Kafka/event-driven integration knowledge to deliver high-impact solutions across complex enterprise environments.
    As an Observability Lead, you will establish the vision and direction for metrics, logs, traces, telemetry pipelines, event streaming, and developer-enabled observability patterns using modern standards such as OpenTelemetry and a combination of open-source and commercial tooling. You will also operate as a forward deployed partner, working directly with engineering, SRE, platform teams, software development teams, and business stakeholders to solve real-world problems and deliver tailored observability implementations in production.
    In this role, you will drive a transition from reactive monitoring toward proactive, intelligence-driven observability. You will influence architectural decisions, embed observability into the software development lifecycle, guide code-level instrumentation and performance engineering practices, and ensure solutions are scalable and adaptable to diverse use cases, including event-driven and Kafka-based architectures. The role requires the ability to read, troubleshoot, and guide improvements to application code; design automation, APIs, integrations, collectors, dashboards, Kafka-aware telemetry flows, and reusable tooling; and help engineering teams adopt observability as part of day-to-day software delivery.
    Success in this position means improving system reliability, reducing mean-time-to-detect and resolve (MTTD/MTTR), enabling faster root-cause analysis, increasing engineering velocity, and creating a consistent, high-fidelity observability experience. You will also translate field learnings into reusable implementation patterns, code assets, automation, Kafka/event-streaming practices, developer standards, and enterprise platform capabilities that inform platform strategy and observability standards.

    For this opportunity, Truist will not sponsor an applicant for work visa status or employment authorization, nor will we offer any immigration-related support for this position (including, but not limited to H-1B, F-1 OPT, F-1 STEM OPT, F-1 CPT, J-1, TN-1 or TN-2, E-3, O-1, or future sponsorship for U.S. lawful permanent residence status.)

    This position is office-centric 5 days a week in one of our Truist hub locations.

    ESSENTIAL DUTIES AND RESPONSIBILITIES
    Following is a summary of the essential functions for this job. Other duties may be performed, both major and minor, which are not mentioned below. Specific activities may change from time to time.
    1. Works to achieve operational targets with major impact on the infrastructure engineering department or job area results and contributes to the development of goals for area of responsibility.
    2. Designs, builds, manages, and implements enterprise infrastructure technology platforms and systems across cloud, network, database, storage, platform, computing, or middleware domains.
    3. Develops and applies automation, monitoring, and optimization techniques to ensure high availability and performance of infrastructure.
    4. Manages large infrastructure projects and programs aligned with organizational strategy and regulatory requirements.
    5. Collaborates with cross-functional teams and external partners to integrate new technologies and continuously improve infrastructure standards and processes.
    6. Troubleshoots and resolves complex technical issues impacting infrastructure performance and reliability.
    7. Ensures compliance with technology strategies, standards, and governance to mitigate risks and ensure regulatory adherence.
    8. Reviews infrastructure designs, configurations, and procedures to support operational consistency and knowledge sharing.
    9. Provides strong technical guidance, training, and direction to infrastructure teams and lower-level technical professionals.
    10. Demonstrates innovative influence with stakeholders in supporting business objectives and technical strategic objectives.

    Qualifications
    Required Qualifications
    The requirements listed below are representative of the knowledge, skill and/or ability required. Reasonable accommodations may be made to enable individuals with disabilities to perform the essential functions.
    1. Bachelor’s degree in Computer Science, Engineering, Information Systems, or related field.
    2. Minimum of 7 years of professional experience in infrastructure engineering.
    3. Advanced knowledge of enterprise infrastructure technologies including cloud, network, database, storage, platform, computing, and middleware.

    Preferred Qualifications:

    - Bachelor’s degree and ten or more years of experience, or an equivalent combination of education and work experience

    - Hands-on experience with OpenTelemetry, including application instrumentation, collector configuration, semantic conventions, and telemetry pipeline design

    - Experience with observability tools such as Prometheus, Grafana, Jaeger, Elastic, Splunk, Dynatrace, Datadog, or similar platforms

    - Strong software development experience in one or more languages such as Python, Go, Java, JavaScript/TypeScript, .NET/C#, Bash, or PowerShell

    - Ability to read, debug, profile, and guide improvements to application code to identify performance bottlenecks, inefficient queries, memory pressure, dependency latency, threading issues, or error-handling gaps

    - Experience designing and developing APIs, automation frameworks, command-line utilities, integrations, collectors, dashboards, or reusable internal tools that support observability and operational workflows

    - Strong Kafka or event streaming platform experience, including producers, consumers, topics, partitions, consumer groups, Kafka Connect, event-driven integration patterns, consumer lag analysis, throughput troubleshooting, and monitoring of Kafka-based services

    - Strong understanding of software development lifecycle practices including Git-based source control, code review, CI/CD pipelines, automated testing, release practices, secure coding, and developer enablement

    - Strong background in Kubernetes, cloud platforms, containers, microservices, service-oriented architectures, event-driven architectures, and cloud-native application patterns

    - Experience with infrastructure as code or configuration automation tools such as Terraform, Ansible, Helm, or similar technologies

    - Experience defining observability strategies, technical standards, developer enablement patterns, event-streaming observability patterns, and engineering roadmaps across multiple application or platform teams

    - Experience in forward deployed engineering, solutions engineering, developer advocacy, or customer-facing technical leadership roles

    - Proven ability to translate customer-specific implementations, code assets, event-streaming patterns, and field learnings into reusable platform capabilities, enterprise standards, and strategic technology recommendations

    General Description of Available Benefits for Eligible Employees of Truist Financial Corporation: All regular teammates (not temporary or contingent workers) working 20 hours or more per week are eligible for benefits, though eligibility for specific benefits may be determined by the division of Truist offering the position. Truist offers medical, dental, vision, life insurance, disability, accidental death and dismemberment, tax-preferred savings accounts, and a 401k plan to teammates. Teammates also receive no less than 10 days of vacation (prorated based on date of hire and by full-time or part-time status) during their first year of employment, along with 10 sick days (also prorated), and paid holidays. For more details on Truist’s generous benefit plans, please visit our Benefits site. Depending on the position and division, this job may also be eligible for Truist’s defined benefit pension plan, restricted stock units, and/or a deferred compensation plan. As you advance through the hiring process, you will also learn more about the specific benefits available for any non-temporary position for which you apply, based on full-time or part-time status, position, and division of work.

    Truist is an Equal Opportunity Employer that does not discriminate on the basis of race, gender, color, religion, citizenship or national origin, age, sexual orientation, gender identity, disability, veteran status, or other classification protected by law. Truist is a Drug Free Workplace.

    EEO is the Law    E-VerifyIER Right to Work

    Numbers & Facts

    LocationCharlotte, North Carolina
    Websitebenefits.truist.com

    Skills

    • Accidental Death and Dismemberment (AD&D)unmatched
    • Analysis Skillsunmatched
    • Ansibleunmatched
    • Application Programming Interface (API)unmatched
    • Architectural Servicesunmatched
    • Automationunmatched
    • Bash Scriptingunmatched
    • Business Strategyunmatched
    • Business Supportunmatched
    • Cloud Applicationsunmatched
    • Cloud Architectureunmatched
    • Cloud Computingunmatched
    • Code Reviewsunmatched
    • Command Lineunmatched
    • Compensation and Benefitsunmatched
    • Computer Scienceunmatched
    • Continuous Deployment/Deliveryunmatched
    • Continuous Improvementunmatched
    • Continuous Integrationunmatched
    • Cross-Functionalunmatched
    • Customer Relationsunmatched
    • Debugging Skillsunmatched
    • Disability Insuranceunmatched
    • English Languageunmatched
    • Error Handlingunmatched
    • Gitunmatched
    • Go Programming Language (Golang)unmatched
    • High Availabilityunmatched
    • Identify Issuesunmatched
    • Information Technology & Information Systemsunmatched
    • Instrumentationunmatched
    • Instrumentation Engineeringunmatched
    • Javaunmatched
    • JavaScriptunmatched
    • Knowledge Managementunmatched
    • Life Insuranceunmatched
    • Machine Toolunmatched
    • Maintain Complianceunmatched
    • Memory Hardwareunmatched
    • Metricsunmatched
    • Microservicesunmatched
    • Microsoft .NETunmatched
    • Microsoft C# (C Sharp)unmatched
    • Middlewareunmatched
    • Multiplatform/Cross-Platformunmatched
    • Open Sourceunmatched
    • Operational Supportunmatched
    • Operations Processesunmatched
    • Performance Engineeringunmatched
    • Performance Managementunmatched
    • Problem Solving Skillsunmatched
    • Python Programming/Scripting Languageunmatched
    • Quality Assurance Methodologyunmatched
    • Regulatory Requirementsunmatched
    • Reliability Engineeringunmatched
    • Reporting Dashboardsunmatched
    • Risk Managementunmatched
    • Root Cause Analysisunmatched
    • Secure Codingunmatched
    • Service-Oriented Architecture (fka Distributed Object Architecture)unmatched
    • Set Goalsunmatched
    • Software Developmentunmatched
    • Software Development Lifecycle (SDLC)unmatched
    • Source Code/Configuration Management (SCM)unmatched
    • Splunkunmatched
    • Standards Developmentunmatched
    • Standards Strategyunmatched
    • Streaming Technologyunmatched
    • Systems Reliabilityunmatched
    • Technical Leadershipunmatched
    • Technical Strategyunmatched
    • Telemetryunmatched
    • Test Automationunmatched
    • Time Managementunmatched
    • Use Casesunmatched
    • User Documentationunmatched
    • Vision Planunmatched
    • Windows PowerShellunmatched

    Be found by employers

    5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

    Level up your application

    Professional resume templates

    Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

    Free resume templates

    Free resume builder

    Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

    Free resume builder