Site Reliability Engineer (SRE) – Production Services

TechDigital Corporation

Pittsburgh, PA

JOB DETAILS
SKILLS
Apache Kafka, Artificial Intelligence (AI), Automation, Budgeting, Continuous Deployment/Delivery, Continuous Integration, Metrics, Operational Improvement, Problem Solving Skills, Quality Management, Reliability Engineering, Reporting Dashboards, Spring Framework
LOCATION
Pittsburgh, PA
POSTED
4 days ago
Mandatory Skills:
1. Java Spring boot 2. Apache Kafka 3. Dev Ops 4. CI/CD automation

Years of experience required: 8-10

Job Description:
Automation & Efficiency
· Automate the top 5 high-volume support and request types
· Build self-service and agent-driven solutions to reduce manual work
· Harden operational workflows for consistency, auditability, and resilience
· Implement auto-retry and backoff for recurring failure patterns

Reliability Engineering
· Define and manage Service Level Objectives (SLOs) for critical services and batch processes
· Apply error budget concepts to guide reliability and release decisions
· Improve batch reliability through standardized recovery patterns and monitoring

Observability & Metrics
· Build reliability dashboards tracking incidents, repeat issues, failure rates, and automation coverage
· Improve operational reporting and visibility across incidents, problems, and changes

Runbooks & Self-Service
· Develop and expand runbooks for key production scenarios
· Convert runbooks into automated remediation workflows
· Enable self-service for repeat operational requests
· Drive conversion of repeat incidents into permanent fixes and known problems

Self-Healing & Intelligent Operations
· Implement self-healing capabilities to minimize manual intervention
· Optimize alerting systems (e.g., Moogsoft) to reduce noise and improve signal quality
· Leverage automation and AI to resolve recurring issues with minimal human involvement

About the Company

T

TechDigital Corporation

COMPANY SIZE
100 to 499 employees
INDUSTRY
Other/Not Classified