Want to know if you’re a fit? Upload your resume and let our AI show you.
Skills
Application Programming Interface (API)unmatched
Atlassian JIRAunmatched
Automationunmatched
CISSP - Certified Information Systems Security Professionalunmatched
Cloud Computingunmatched
Communication Skillsunmatched
CompTIA Security+unmatched
DevOpsunmatched
Documentationunmatched
Enterprise Applicationsunmatched
High Availabilityunmatched
Identify Issuesunmatched
Incident Managementunmatched
Incident Responseunmatched
Linux Administrationunmatched
Linux Operating Systemunmatched
On Callunmatched
Operating Systemsunmatched
Operational Strategyunmatched
Operational Supportunmatched
Operations Planningunmatched
Performance Metricsunmatched
Reliability Engineeringunmatched
Risk Managementunmatched
Sensitive Compartmented Information (SCI)unmatched
Software Patchesunmatched
Standard Operating Procedures (SOP)unmatched
Systems Administration/Managementunmatched
Systems Analysisunmatched
Systems Engineeringunmatched
Technical Leadershipunmatched
Top Secret Clearanceunmatched
United States Citizenunmatched
Virtualizationunmatched
Description
Job Title: Tier III Operations Engineer Location: Onsite – Dallas, TX Area Clearance: Active TS/SCI with CI Poly Citizenship: U.S. Citizen
Position Summary
The Tier III Operations Engineer is responsible for maintaining the availability, performance, and reliability of mission-critical systems. This role serves as the highest level of operational support, providing advanced technical troubleshooting and incident resolution for complex infrastructure and application issues. The engineer collaborates with Operations, Engineering, and external vendors to ensure system stability and continuous service availability.
Key Responsibilities
Lead diagnosis and resolution of complex incidents that cannot be resolved by Tier I or Tier II support teams
Perform advanced troubleshooting across servers, operating systems, networking, databases, storage, virtualization, and enterprise applications
Participate in incident response during critical outages and coordinate technical recovery efforts
Monitor system health, performance, and capacity to proactively identify potential issues
Develop and maintain operational documentation, troubleshooting guides, and standard operating procedures
Support planned maintenance, system upgrades, patching, and change‑management activities
Collaborate with engineering teams to improve system resiliency, automation, and operational efficiency
Identify recurring issues and recommend long-term improvements to increase reliability and reduce operational risk
Required Qualifications
Security+ or CISSP certification
Strong communication skills with the ability to explain complex technical topics
Strong understanding of Linux operating systems
Availability to work extended hours and provide after-hours and on-call support
Experience supporting and troubleshooting APIs and security appliances such as API gateways
Ability to analyze system logs, performance metrics, and diagnostic data to quickly isolate technical issues
Ability to work independently and effectively under pressure during critical production incidents
Preferred Qualifications
Bachelor’s degree in a STEM field and 5+ years of related experience
Experience supporting high‑availability or mission‑critical environments
Experience with Grafana or similar monitoring tools
Familiarity with incident management and change‑management processes
Broad knowledge of servers, networking, storage, and virtualization
Experience with cloud monitoring and logging tools such as Grafana, Prometheus, Promtail, and Loki
Familiarity with DevOps collaboration tools such as Jira and Confluence