Site Reliability Engineer II Data & Intelligence

Danta Technologies
  • Atlanta, GA
  • $40–$45 Per Hour
  • Quick Apply
1 day ago

Job Description

Employment Eligibility Statement:
Due to specific project and client requirements, this position is open to U.S. Citizens and U.S. Lawful Permanent Residents (Green Card holders). Sponsorship is not available at this time.
Danta Technologies evaluates all candidates in compliance with the Immigration and Nationality Act (INA) and EEOC guidelines. All hiring decisions are made without regard to race, color, religion, sex, gender identity, sexual orientation, national origin, age, disability, veteran status, or any other protected characteristic.

Note:
This position is based in Dallas, TX / Overland Park, KS / Atlanta, GA / Bellevue, WA, and requires 5 days per week working from the office.
Only candidates who are currently located within the United States will be considered.
Applications from candidates residing outside the USA will not be accepted.

Job Summary

The Site Reliability Engineer II (SRE II) is responsible for ensuring the reliability, scalability, performance, security, and operational excellence of data platforms, analytics systems, AI/ML services, and business intelligence applications. This role combines software engineering, systems engineering, automation, and operational expertise to build resilient and highly available data services while driving continuous improvement through observability, automation, and reliability engineering practices.

The ideal candidate is passionate about large-scale distributed systems, cloud-native technologies, data platforms, and operational excellence. They partner closely with Data Engineering, Data Science, Analytics, Platform Engineering, and Product teams to maintain and improve critical business services.


Key Responsibilities:

Reliability & Operations:

  • Ensure the availability, performance, scalability, and reliability of data and intelligence platforms.
  • Manage production environments supporting data ingestion, processing, transformation, storage, analytics, and AI/ML workloads.
  • Participate in on-call rotations and incident response activities.
  • Lead troubleshooting efforts for complex production issues and drive root cause analysis (RCA).
  • Develop and implement service level indicators (SLIs), service level objectives (SLOs), and error budgets.

Automation & Engineering:

  • Design and develop automation to improve operational efficiency and system reliability.
  • Build self-healing solutions and automate routine operational tasks.
  • Create tools and scripts to monitor, deploy, and manage large-scale distributed systems.
  • Improve deployment processes through CI/CD pipelines and Infrastructure as Code (IaC).

Observability & Monitoring:

  • Design and maintain monitoring, logging, tracing, and alerting solutions.
  • Create dashboards and actionable alerts to proactively identify service degradation.
  • Analyze system performance metrics and recommend optimization opportunities.
  • Drive observability standards across data services and platforms.

Platform & Infrastructure Management:

  • Support cloud-based infrastructure and platform services across Azure, AWS, or GCP environments.
  • Optimize compute, storage, networking, and data platform resources.
  • Work with containerized and Kubernetes-based workloads.
  • Ensure high availability and disaster recovery capabilities are implemented and tested.

Data Platform Reliability:

  • Support modern data ecosystems including data lakes, warehouses, streaming platforms, and analytics environments.
  • Monitor ETL/ELT pipelines, batch processing, real-time streaming, and data orchestration services.
  • Partner with Data Engineers to improve pipeline reliability and data quality monitoring.
  • Ensure platform scalability for growing data volumes and user demands.

Security & Compliance:

  • Implement security best practices and operational controls.
  • Support compliance requirements related to data governance and privacy.
  • Collaborate with security teams to remediate vulnerabilities and improve platform security posture.

Continuous Improvement:

  • Conduct post-incident reviews and drive corrective and preventive actions.
  • Identify reliability risks and implement long-term improvements.
  • Promote a culture of operational excellence, resilience, and automation.
  • Contribute to engineering standards, runbooks, knowledge sharing, and best practices.

Required Qualifications:

Education:

  • Bachelor's degree in computer science, Information Technology, Engineering, or related field, or equivalent practical experience.

Experience:

  • 3+ years of experience in Site Reliability Engineering, DevOps, Systems Engineering, Cloud Engineering, or a related role.
  • Experience supporting production environments with high availability requirements.
  • Experience managing cloud infrastructure and distributed systems.

Technical Skills:

  • Strong knowledge of Linux systems administration.
  • Proficiency in one or more programming languages such as Python, Java, Go, C#, or JavaScript.
  • Experience with CI/CD tools and deployment automation.
  • Experience with Infrastructure as Code tools such as Terraform, ARM, or Bicep.
  • Experience with Kubernetes and container technologies.
  • Knowledge of monitoring and observability technologies such as Prometheus, Grafana, Datadog, Azure Monitor, Splunk, or OpenTelemetry.
  • Understanding of networking, DNS, load balancing, and distributed systems concepts.

Experience supporting data platforms such as:

  • Azure Data Lake
  • Azure Synapse Analytics
  • Databricks
  • Snowflake
  • Kafka
  • SQL/NoSQL databases
  • Data orchestration platforms

Preferred Qualifications:

  • Experience supporting AI/ML platforms and MLOps environments.
  • Experience with Azure cloud-native services.
  • Familiarity with data governance and data quality frameworks.
  • Knowledge of reliability engineering best practices and SRE methodologies.
  • Experience implementing SLOs, SLIs, and error budgets.
  • Experience supporting large-scale analytics and business intelligence environments.
  • Azure, AWS, Kubernetes, Terraform, or DevOps certifications.

Key Competencies:

  • Problem-solving and analytical thinking
  • Incident management and troubleshooting
  • Automation-first mindset
  • Collaboration and stakeholder management
  • Strong communication skills
  • Continuous learning and innovation
  • Customer-focused approach
  • Operational excellence and accountability Success Measures
  • Maintaining high availability and reliability targets for critical services.
  • Reducing operational toil through automation.
  • Improving platform observability and incident response effectiveness.
  • Meeting service-level objectives and performance goals.
  • Enhancing deployment reliability and operational efficiency.
  • Driving measurable improvements in system resilience, scalability, and customer experience.

Notes:- All qualified applicants will receive consideration for employment without regard to race, color, religion, religious creed, sex, national origin, ancestry, age, physical or mental disability, medical condition, genetic information, military and veteran status, marital status, pregnancy, gender, gender expression, gender identity, sexual orientation, or any other characteristic protected by local law, regulation, or ordinance.

Benefits: Danta offers a compensation package to all W2 employees that are competitive in the industry. It consists of competitive pay, the option to elect healthcare insurance (Dental, Medical, Vision), Major holidays and Paid sick leave as per state law.

The rate/ Salary range is dependent on numerous factors including Qualification, Experience and Location.

Numbers & Facts

LocationAtlanta, GA
Salary$40–$45 Per Hour

Skills

  • ARM (Advanced RISC Machine)unmatched
  • Amazon Web Services (AWS)unmatched
  • Analysis Skillsunmatched
  • Artificial Intelligence (AI)unmatched
  • Automationunmatched
  • Best Practicesunmatched
  • Budgetingunmatched
  • Business Intelligenceunmatched
  • Business Intelligence Softwareunmatched
  • Business Servicesunmatched
  • Cloud Computingunmatched
  • Communication Skillsunmatched
  • Computer Scienceunmatched
  • Computer Securityunmatched
  • Continuous Deployment/Deliveryunmatched
  • Continuous Improvementunmatched
  • Continuous Integrationunmatched
  • Corrective Actionunmatched
  • Customer Experienceunmatched
  • Customer Relationsunmatched
  • DNS (Domain Name System)unmatched
  • Data Analysisunmatched
  • Data Lakeunmatched
  • Data Qualityunmatched
  • Data Scienceunmatched
  • Data Warehousingunmatched
  • Database Extract Transform and Load (ETL)unmatched
  • DevOpsunmatched
  • Disaster Recoveryunmatched
  • Distributed Computingunmatched
  • Ecosystemsunmatched
  • GCP (Good Clinical Practices)unmatched
  • Go Programming Language (Golang)unmatched
  • High Availabilityunmatched
  • High Reliabilityunmatched
  • Identify Issuesunmatched
  • Incident Managementunmatched
  • Incident Responseunmatched
  • Information Technology & Information Systemsunmatched
  • Javaunmatched
  • JavaScriptunmatched
  • Large-Scale Systemsunmatched
  • Linux Administrationunmatched
  • Load Balancingunmatched
  • Microsoft C# (C Sharp)unmatched
  • Microsoft Windows Azureunmatched
  • NoSQLunmatched
  • On Callunmatched
  • Operational Improvementunmatched
  • Operational Strategyunmatched
  • Performance Analysisunmatched
  • Performance Metricsunmatched
  • Problem Solving Skillsunmatched
  • Process Improvementunmatched
  • Production Managementunmatched
  • Production Supportunmatched
  • Production Systemsunmatched
  • Programming Languagesunmatched
  • Python Programming/Scripting Languageunmatched
  • Quality Monitoringunmatched
  • Regulatory Complianceunmatched
  • Reliability Engineeringunmatched
  • Reporting Dashboardsunmatched
  • Risk Analysisunmatched
  • Root Cause Analysisunmatched
  • SQL Databasesunmatched
  • Scripting (Scripting Languages)unmatched
  • Software Engineeringunmatched
  • Splunkunmatched
  • Systems Analysisunmatched
  • Systems Engineeringunmatched
  • Systems Reliabilityunmatched
  • Systems Scalabilityunmatched
  • Technology Analysisunmatched

Be found by employers

5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

Level up your application

Professional resume templates

Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

Free resume templates

Free resume builder

Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

Free resume builder