Lead Monitoring and Observability Engineer

CoreTechs
  • Evansville, IN
  • Remote
  • $1–$86 Per Hour
  • Instant Apply
4 days ago

Job Description

Lead Monitoring and Observability Engineer
Remote, USA

Note: MUST be legally authorized to work in the United States. 

SUMMARY
  • The Company is the country's largest lending-exclusive financial company, proudly serving millions of customers with safe, affordable, and transparent installment loans. Our customers turn to us every day - online and at over 1,400 branches in 44 states - to help them take control and improve their financial lives. It's all about doing the right thing - a mission that hasn't changed for more than 100 years.
  • We are seeking a Lead Platform Engineer to; Able to collaboratively establish availability and performance objectives, measure progress towards those objectives, and implement necessary changes. Able to interpret changing technical & process needs of product teams and adjust platform and tooling meets those needs.
  • The Lead Monitoring and Observability Engineer serve as a senior technical contributor responsible for ensuring monitoring reliability, telemetry quality, automation maturity, and operability across OMF’s infrastructure and application ecosystem. This role acts as a technical mentor, monitoring and networking SME, and hands-on engineer who guides monitoring-platform evolution, improves service quality, and collaborates with product and engineering teams to deliver scalable, stable, and observable systems.

KEY RESPONSIBILITIES
Monitoring Reliability & Performance
  • Establish platform SLOs, availability goals, latency/error budgets, reliability metrics, and monitoring coverage expectations in partnership with teams.
  • Continuously measure service and platform health and implement changes to improve reliability, performance, alert quality, and operational stability.
  • Define and maintain standards for actionable alerting, dashboards, logs, metrics, traces, and service-health reporting.
Architecture & Technical Leadership
  • Lead observability design efforts for Elastic/ELK, telemetry pipelines, monitoring platforms, dashboards, alerting, distributed tracing, synthetic monitoring, microservices platforms, cloud infrastructure, or network monitoring, depending on assignment.
  • Provide “out-of-the-box” technical solutions that balance velocity, reliability, operational visibility, and cost.
  • Evaluate monitoring and observability tools, patterns, integrations, and emerging requirements.
  • Define reusable observability patterns, reference architectures, onboarding guidance, and validation practices for product and platform teams.
Hands-On Engineering & Automation
  • Perform advanced configuration, IaC development (Terraform/Ansible), monitoring-as-code development, CI/CD pipeline engineering, and cloud platform automation.
  • Build and maintain scalable, resilient monitoring and observability components that support product teams.
  • Implement and maintain monitoring, alerting, data visualization, logging, metrics, tracing, and telemetry collection capabilities as noted in the IC3 role reference.
  • Improve automation for monitoring onboarding, dashboard creation, alert configuration, telemetry collection, and operational workflows.
Operational Excellence
  • Reduce manual operations through automation and self-service observability tooling.
  • Review, optimize, and maintain observability capabilities including logs, metrics, traces, dashboards, alerting, and service-health reporting.
  • Improve alert signal quality, reduce alert noise, and ensure alerts have clear ownership, escalation paths, and actionable runbook guidance.
  • Participate in and lead high-severity incident response for monitoring and observability-owned domains.
  • Support root-cause analysis by correlating telemetry, infrastructure conditions, application behavior, and operational events.
Network Monitoring
  • Define and maintain network-monitoring standards for network availability, reachability, latency, packet loss, interface health, capacity, device health, routing, and dependency-related service impact.
  • Partner with Networking, Cloud, AppDev, Security, and SRE teams to establish monitoring coverage for critical network devices, services, and customer journeys.
  • Build and maintain network-monitoring dashboards, alerting patterns, service-health views, and operational workflows.
  • Establish validation practices for network-device onboarding, telemetry collection, alert quality, dashboard completeness, and operational readiness.
  • Support the migration of network monitoring from OpsRamp to the selected network-monitoring tool, including requirements definition, technical design, migration planning, testing, validation, and operational handoff.
  • Evaluate network-monitoring tools, integrations, automation opportunities, and emerging network observability requirements.
  • Connect network telemetry and alerts to application, infrastructure, and customer-impact signals to improve triage and incident response.
Cross-Team Collaboration
  • Partner with AppDev, Security, Observability, Networking, Cloud, and SRE teams to deliver cohesive monitoring services and shared solutions.
  • Adjust monitoring-platform design, observability standards, and processes based on evolving product team needs.
  • Collaborate with service owners to prioritize observability improvements for critical services.
Mentorship & Influence
  • Mentor IC1–IC2 engineers in engineering best practices, TDD, agile ceremonies, monitoring and observability fundamentals, and operational workflows.
  • Provide code reviews, design feedback, and internal technical enablement.
  • Guide partner teams in adopting agreed observability standards and reusable implementation patterns.
  • Provide technical leadership on monitoring architecture, telemetry quality, alerting practices, and operational readiness.

QUALIFICATIONS
Required
  • Demonstrated expertise in monitoring and observability engineering, including Elastic/ELK, dashboards, alerting, logging, metrics, tracing, telemetry pipelines, cloud infrastructure, networking, or equivalent.
  • Proven ability to define availability and performance objectives and implement monitoring-driven improvements.
  • Hands-on experience with automation, including Terraform, Ansible, scripts, monitoring-as-code, or cloud-platform automation.
  • Experience designing, implementing, and maintaining monitoring, alerting, dashboards, and service-health reporting.
  • Experience with observability practices across logs, metrics, traces, dashboards, alerting, and incident response.
  • Ability to evaluate alternatives, provide thought leadership, and guide technical direction.
  • Proficiency with TDD, agile development, and iterative delivery.
  • Strong communication and mentoring capability.
Network Monitoring Requirements
  • Experience with enterprise network monitoring, network operations, network observability, or a related infrastructure discipline.
  • Working knowledge of network technologies and protocols, including TCP/IP, DNS, HTTP/S, SNMP, ICMP, routing, switching, firewalls, load balancing, VPN, and WAN connectivity.
  • Experience monitoring network devices, interfaces, availability, latency, packet loss, bandwidth utilization, capacity, and reachability.
  • Ability to design network-focused dashboards, alerts, escalation workflows, and service-health views.
  • Experience collaborating with Networking and incident-response teams during complex service-impacting events.
  • Familiarity with network-monitoring migrations, platform consolidations, tool evaluations, or proof-of-concepts.
  • Preferred
  • Experience with distributed tracing, observability platforms, monitoring analytics, or OpenTelemetry-aligned telemetry practices.
  • Familiarity with DevOps/SRE practices, including SLOs, runbooks, deployments, incident response, and post-incident improvement.
  • Experience with multi-cloud environments or hybrid architectures.
  • Experience with OpsRamp, SolarWinds, Elastic/ELK, Grafana, ServiceNow, or related monitoring, alerting, and incident-management integrations.
  • Experience with synthetic monitoring, customer-journey monitoring, or service-level reporting.
  • Experience leading a network-monitoring platform migration or enterprise monitoring-tool evaluation.
We are an equal opportunity employer, and we are an organization that values diversity. We welcome applications from all qualified candidates, including minorities and persons with disabilities. 
 
Req45-1

Numbers & Facts

LocationEvansville, IN (
Remote
)
Salary$1–$86 Per Hour

Skills

  • Agile Programming Methodologiesunmatched
  • Ansibleunmatched
  • Automationunmatched
  • Best Practicesunmatched
  • Budgetingunmatched
  • Capacity Utilizationunmatched
  • Cloud Computingunmatched
  • Code Reviewsunmatched
  • Communication Skillsunmatched
  • Continuous Deployment/Deliveryunmatched
  • Continuous Integrationunmatched
  • DNS (Domain Name System)unmatched
  • Data Visualizationunmatched
  • DevOpsunmatched
  • Diversityunmatched
  • Ecosystemsunmatched
  • Establish Prioritiesunmatched
  • Financeunmatched
  • Firewallsunmatched
  • HTTP (HyperText Transport Protocol)unmatched
  • Hybrid Cloudunmatched
  • ICMPunmatched
  • Incident Managementunmatched
  • Incident Responseunmatched
  • Integrated Circuits (ICs)unmatched
  • Leadershipunmatched
  • Load Balancingunmatched
  • Loansunmatched
  • Machine Toolunmatched
  • Mentoringunmatched
  • Metricsunmatched
  • Microservicesunmatched
  • Network Administration/Managementunmatched
  • Network Connectivityunmatched
  • Network Designunmatched
  • Network Monitoringunmatched
  • Network Protocolsunmatched
  • Network Routingunmatched
  • Network Switchingunmatched
  • Onboardingunmatched
  • Performance Analysisunmatched
  • Performance Managementunmatched
  • Process Improvementunmatched
  • Product Engineeringunmatched
  • Product Supportunmatched
  • Product Testingunmatched
  • Proof of Conceptunmatched
  • Quality Managementunmatched
  • Reliability Engineeringunmatched
  • Reporting Dashboardsunmatched
  • Requirements Managementunmatched
  • Root Cause Analysisunmatched
  • SNMP (Simple Network Management Protocol)unmatched
  • Scalable System Developmentunmatched
  • Scripting (Scripting Languages)unmatched
  • Service Deliveryunmatched
  • ServiceNowunmatched
  • Software Engineeringunmatched
  • Standards Developmentunmatched
  • System Migrationunmatched
  • TCP/IP (Transmission Control Protocol/Internet Protocol)unmatched
  • Team Playerunmatched
  • Technical Leadershipunmatched
  • Technical/Engineering Designunmatched
  • Telemetryunmatched
  • Test Driven Development (TDD)unmatched
  • Thought Leadershipunmatched
  • VPN (Virtual Private Network)unmatched
  • Validation Testingunmatched
  • Wide Area Network (WAN)unmatched

Be found by employers

5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

Level up your application

Professional resume templates

Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

Free resume templates

Free resume builder

Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

Free resume builder