AI DevOps/Observability Engineer

PeopleNTech LLC
  • Alexandria, VA
  • $70 Per Hour
  • Instant Apply
7 days ago

Job Description

Role : AI DevOps/Observability Engineer

Location : Charlotte NC (Onsite)

Rate : $70

Indent : SF_OP_205828-13

We are seeking a highly skilled AI DevOps/Observability Engineer to join our production operations team. In this role, you will be responsible for the reliability, performance, and operational readiness of our priority Generative AI and LLM agent releases. You will bridge the gap between AI development and production operations by implementing robust telemetry, automated evaluation pipelines, and comprehensive monitoring systems. Your work will ensure our intelligent agents are stable, accurate, efficient, and scale seamlessly in production environments.

Key Responsibilities

  • Observability & Telemetry: Implement and maintain deep tracing, logging, and telemetry solutions specifically tailored for LLM application architectures and multi-agent workflows.
  • Monitoring & Insights: Design, build, and maintain production dashboards that track systemic health, infrastructure metrics, and specialized AI performance indicators.
  • Operational Readiness: Establish actionable alerting systems, define Service Level Objectives (SLOs), curate runtime runbooks, and provide concrete engineering evidence for production readiness.
  • Evaluation & Testing Pipelines: Develop automated continuous evaluation suites to assess LLM agent behavior, safety, and output quality prior to and during deployment.
  • Performance Analysis: Analyze prompt efficiency, token usage, latency, and overall model performance to optimize cost, speed, and accuracy.
  • Production Operations: Support the deployment pipeline, participate in incident management, and continually improve the resilience of our live AI services.

Required Skills and Qualifications

Core Technical Skills

  • LLM & Agent Evaluation: Experience with framework-based evaluation tools (e.g., Ragas, DeepEval, TruLens) to measure hallucination, faithfulness, and relevancy.
  • Tracing & Telemetry: Proficiency with LLM-specific tracing tools (e.g., LangSmith, LangFuse, Phoenix, Arize) and open standards like OpenTelemetry.
  • Metrics & Dashboards: Hands-on experience building monitoring views in platforms like Datadog, Prometheus/Grafana, New Relic, or cloud-native suites.
  • Production Alerting & SLOs: Proven ability to define meaningful Service Level Indicators (SLIs) and SLOs, minimizing alert fatigue while maximizing system reliability.
  • Test Automation: Strong background in integrating automated test frameworks into CI/CD pipelines for continuous integration of AI features.
  • Prompt & Model Analysis: Analytical mindset to benchmark prompt variants, track regression in model behavior, and profile latency across model providers.

Programming & Operations

  • Python Mastery: Advanced Python programming skills, including experience with async execution, API integration, and AI frameworks (e.g., LangChain, LlamaIndex).
  • Production Operations: Solid understanding of cloud infrastructure (AWS/GCP/Azure), containerization (Docker, Kubernetes), and DevOps best practices.

Preferred Qualifications

  • 3+ years of experience operationalizing LLMs or generative AI applications in production.
  • Experience managing vector databases (e.g., Pinecone, Milvus, Chroma) and tracking RAG pipeline performance.
  • Background in Site Reliability Engineering (SRE) or specialized MLOps roles.

Numbers & Facts

LocationAlexandria, VA

Skills

  • Amazon Web Services (AWS)unmatched
  • Analysis Skillsunmatched
  • Application Programming Interface (API)unmatched
  • Artificial Intelligence (AI)unmatched
  • Benchmarkingunmatched
  • Best Practicesunmatched
  • Cloud Computingunmatched
  • Computer Programmingunmatched
  • Concreteunmatched
  • Continuous Deployment/Deliveryunmatched
  • Continuous Integrationunmatched
  • Cost Controlunmatched
  • Database Administrationunmatched
  • DevOpsunmatched
  • Dockerunmatched
  • GCP (Good Clinical Practices)unmatched
  • Incident Managementunmatched
  • Metricsunmatched
  • Microsoft Windows Azureunmatched
  • Operational Supportunmatched
  • Performance Analysisunmatched
  • Performance Modelingunmatched
  • Performance Tuning/Optimizationunmatched
  • Production Controlunmatched
  • Production Supportunmatched
  • Production Systemsunmatched
  • Python Programming/Scripting Languageunmatched
  • Reliability Engineeringunmatched
  • Reporting Dashboardsunmatched
  • Safety/Work Safetyunmatched
  • Software Agentsunmatched
  • Systems Reliabilityunmatched
  • Telemetryunmatched
  • Test Automationunmatched
  • Test Harnessunmatched

Be found by employers

5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

Level up your application

Professional resume templates

Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

Free resume templates

Free resume builder

Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

Free resume builder