Observability & Evaluation Engineer

Kasmo Inc
  • Charlotte, NC
  • Quick Apply
11 days ago

Job Description



Description:
Charlotte, NC (Onsite)
Observability & Evaluation Engineer will build telemetry, tracing, dashboards, evaluation suites, alerts, service objectives, runbooks, and readiness evidence for Tachyon agent releases. This role ensures production AI systems can be monitored, evaluated, improved, and supported with clear operational visibility.
Key Responsibilities
Implement observability and telemetry for LLM-powered applications, agents, tools, and platform services.
Build evaluation suites for agent behavior, prompt quality, response quality, retrieval performance, latency, reliability, and safety signals.
Develop dashboards, alerts, traces, metrics, service objectives, and reporting for production readiness.
Work with platform engineers and Product Owners to define monitoring requirements and evaluation metrics.
Automate evidence collection for release readiness, operational reviews, and governance checkpoints.
Create runbooks and support documentation for priority agent releases.
Analyze production behavior and recommend improvements to reliability, performance, and quality.
Required Qualifications
7+ years of engineering experience with observability, monitoring, test automation, platform operations, or AI/ML systems.
5+ years of strong hands-on Python experience.
5+ years of Experience with dashboards, metrics, alerts, traces, logs, SLOs, and production monitoring.
5+ years of Understanding of LLM evaluation, prompt evaluation, RAG evaluation, or AI quality assessment approaches.
5+ years of Experience working in Agile engineering teams and production support environments.
Required Skills / Knowledge
Python, telemetry, tracing, monitoring, dashboards, alerting, SLOs, evaluation frameworks, test automation, and production operations.
Understanding of LLMs, agents, RAG, prompt performance, retrieval quality, latency, and reliability metrics.
Experience with observability tools and open telemetry concepts.
Preferred Qualifications
Experience with GenAI observability, AI evaluation tools, ML monitoring, or platform reliability engineering.
Experience in regulated environments with evidence and readiness documentation.
Kubernetes, cloud platforms, and CI/CD experience.
Expected Outcomes
Operational dashboards and evaluation suites for priority agent releases.
Clear readiness evidence, alerts, SLOs, and runbooks.
Improved quality, reliability, and trust in production Agentic AI systems.

Numbers & Facts

LocationCharlotte, NC

Skills

  • Agile Programming Methodologiesunmatched
  • Analysis Skillsunmatched
  • Artificial Intelligence (AI)unmatched
  • Engineeringunmatched
  • Metricsunmatched
  • Operational Auditunmatched
  • Operational Supportunmatched
  • Production Controlunmatched
  • Production Supportunmatched
  • Production Systemsunmatched
  • Python Programming/Scripting Languageunmatched
  • Quality Managementunmatched
  • Reliability Engineeringunmatched
  • Reporting Dashboardsunmatched
  • Requirements Managementunmatched
  • Support Documentationunmatched
  • Team Playerunmatched
  • Telemetryunmatched
  • Test Automationunmatched

Be found by employers

5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

Level up your application

Professional resume templates

Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

Free resume templates

Free resume builder

Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

Free resume builder