We are seeking a highly skilled Senior Observability Engineer to join our Observability Engineering team.
This role is responsible for designing, implementing, administering, and automating enterprise observability solutions with a primary focus on Grafana, Open Telemetry, monitoring, alerting, and telemetry management.
The ideal candidate has strong experience building and operating observability platforms at scale, automating infrastructure through Terraform, and enabling application and infrastructure teams to adopt standardized observability practices.
This role will play a key part in modernizing observability capabilities and driving migration from legacy monitoring tools to Grafana-based solutions.
Key Responsibilities
Observability Platform Engineering
Administer and support Grafana Cloud and on-premises Grafana deployments.
Design and implement enterprise observability solutions for metrics, logs, traces, synthetic monitoring, and alerting.
Establish and maintain observability standards, best practices, and governance processes.
Configure and manage Grafana data sources, alerting, RBAC, folders, teams, and integrations.
Ensure platform scalability, reliability, resiliency, and operational excellence.
Automation & Infrastructure as Code
Develop and maintain Terraform modules for Grafana infrastructure and configuration management.
Automate onboarding of applications, infrastructure, dashboards, alerts, and data sources.
Build self-service capabilities that reduce manual operational effort and improve adoption.
Integrate observability capabilities into CI/CD and infrastructure provisioning workflows.
Monitoring, Alerting & Incident Management
Design meaningful monitoring and alerting strategies based on service health and business-critical workflows.
Implement and optimize alerting standards to reduce noise and improve signal quality.
Support incident response, troubleshooting, root cause analysis, and post-incident reviews.
Drive continuous improvement of operational visibility and platform health.
Open Telemetry & Telemetry Engineering
Implement and support Open Telemetry instrumentation across applications and infrastructure.
Establish standards for logs, metrics, traces, and telemetry collection.
Support telemetry pipelines, agent deployments, and data collection strategies.
Assist teams with instrumentation design and observability adoption.
Migration & Modernization
Support migration initiatives from legacy observability platforms to Grafana.
Analyze existing monitoring, alerting, logging, and tracing implementations and recommend modernization approaches.
Develop reusable migration patterns, automation, and engineering standards.
Partner with application teams to accelerate adoption of enterprise observability capabilities.
Collaboration & Leadership
Work closely with application development, infrastructure, cloud, and SRE teams.
Provide technical leadership and mentoring to engineers across the organization.
Contribute to observability architecture, strategy, and roadmap development.
Promote observability as a core engineering practice across the enterprise.
Required Qualifications
Bachelor's degree in Computer Science, Engineering, Information Systems, or related field.
5 years of experience in observability, monitoring, operations, or platform engineering.
Hands-on experience administering Grafana in large-scale enterprise environments.
Strong experience with Terraform and Infrastructure as Code practices.
Experience implementing monitoring, alerting, logging, and distributed tracing solutions.
Experience with Open Telemetry concepts, instrumentation, and telemetry pipelines.
Strong Linux and cloud platform administration skills.
Experience with scripting and automation using Python, PowerShell, Bash, or similar languages.
Knowledge of operational excellence, reliability engineering, and incident management practices.
Preferred Qualifications
Experience migrating from tools such as Splunk, Dynatrace, AppDynamics, New Relic, OpenText OBM, or similar platforms.
Experience with Grafana Alloy, Tempo, Loki, Mimir, or Prometheus.
Experience operating observability platforms in AWS environments.
Knowledge of Kubernetes, containers, and cloud-native observability.
Experience designing enterprise observability strategies and governance models.
Familiarity with CI/CD platforms and DevOps practices.
Desired Skills
Grafana Administration
Terraform
Open Telemetry (OTEL)
Monitoring & Alerting
Observability Engineering
Platform Engineering
Linux Administration
AWS Cloud Services
Automation & Scripting
Incident Management
Infrastructure as Code
Telemetry Pipelines
Reliability Engineering
Root Cause Analysis
Enterprise Monitoring Architecture
About US Tech Solutions: US Tech Solutions is a global staff augmentation firm providing a wide range of talent on-demand and total workforce solutions. To know more about US Tech Solutions, please visit www.ustechsolutions.com.
US Tech Solutions is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, colour, religion, sex, sexual orientation, gender identity, national origin, disability, or status as a protected veteran.
AI Statement: By applying, you acknowledge that AI-assisted tools may be used during hiring.
#LI-AS140
Numbers & Facts
Location
Coppell, TX
Skills
Administrative Skillsunmatched
Amazon Web Services (AWS)unmatched
Analysis Skillsunmatched
Automationunmatched
Bash Scriptingunmatched
Best Practicesunmatched
Cloud Computingunmatched
Computer Scienceunmatched
Configuration Managementunmatched
Continuous Deployment/Deliveryunmatched
Continuous Improvementunmatched
Continuous Integrationunmatched
Data Collectionunmatched
Data Managementunmatched
DevOpsunmatched
Enterprise Architectureunmatched
Identify Issuesunmatched
Incident Managementunmatched
Incident Responseunmatched
Information Technology & Information Systemsunmatched
Instrumentationunmatched
Leadershipunmatched
Linux Administrationunmatched
Linux Operating Systemunmatched
Mentoringunmatched
Metricsunmatched
Onboardingunmatched
Operational Improvementunmatched
Process Improvementunmatched
Python Programming/Scripting Languageunmatched
Quality Managementunmatched
Reliability Engineeringunmatched
Reporting Dashboardsunmatched
Root Cause Analysisunmatched
Scripting (Scripting Languages)unmatched
Software Developmentunmatched
Splunkunmatched
System Migrationunmatched
Systems Administration/Managementunmatched
Technical Leadershipunmatched
Telemetryunmatched
Windows PowerShellunmatched
🎯
Be found by employers
5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.
Level up your application
Professional resume templates
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.