Site Reliability Engineer (SRE) Production Services

TechDigital Corporation
  • Pittsburgh, PA
  • Quick Apply
25 days ago

Job Description

Mandatory Skills:
1. Java Spring boot 2. Apache Kafka 3. Dev Ops 4. CI/CD automation

Years of experience required: 8-10

Job Description:
Automation & Efficiency
· Automate the top 5 high-volume support and request types
· Build self-service and agent-driven solutions to reduce manual work
· Harden operational workflows for consistency, auditability, and resilience
· Implement auto-retry and backoff for recurring failure patterns

Reliability Engineering
· Define and manage Service Level Objectives (SLOs) for critical services and batch processes
· Apply error budget concepts to guide reliability and release decisions
· Improve batch reliability through standardized recovery patterns and monitoring

Observability & Metrics
· Build reliability dashboards tracking incidents, repeat issues, failure rates, and automation coverage
· Improve operational reporting and visibility across incidents, problems, and changes

Runbooks & Self-Service
· Develop and expand runbooks for key production scenarios
· Convert runbooks into automated remediation workflows
· Enable self-service for repeat operational requests
· Drive conversion of repeat incidents into permanent fixes and known problems

Self-Healing & Intelligent Operations
· Implement self-healing capabilities to minimize manual intervention
· Optimize alerting systems (e.g., Moogsoft) to reduce noise and improve signal quality
· Leverage automation and AI to resolve recurring issues with minimal human involvement

Numbers & Facts

LocationPittsburgh, PA
IndustryOther/Not Classified
Company Size100 to 499 employees

More jobs like this

See more jobs

Be found by employers

5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.