We respectfully request that 3rd parties refrain from contacting us regarding this posting.
Overview
LABUR is partnering with a client to hire a Senior AI Platform & Production Engineer for a hands-on engineering role focused on building, operating, and continuously improving shared Enterprise AI platform capabilities and the production AI systems that depend on them. The ideal candidate combines strong production software engineering fundamentals with practical experience operating LLM and AI applications at scale, and is comfortable diagnosing complex issues across platform infrastructure and application layers. This role is fully remote.
Responsibilities
Build, operate, and evolve the shared Enterprise AI platform alongside the production AI applications it supports
Develop and enhance the LLM API Gateway, developer tooling, onboarding workflows, and reusable platform services
Design and maintain secure MCP and tool integration infrastructure for agentic AI systems
Implement and harden authentication, authorization, policy enforcement, and audit controls across the platform
Establish and maintain production-grade observability, including tracing, monitoring, alerting, debugging, and operational dashboards
Lead production reliability efforts such as KTLO support, incident response, root-cause analysis, and remediation
Optimize AI systems for quality, reliability, latency, scalability, and cost while driving incremental enhancements to deployed applications
Collaborate with application engineers, platform teams, and security stakeholders to define standards and enable reliable AI development and operations
Qualifications
Strong production software engineering skills in Python and/or TypeScript/Node.js with experience designing backend APIs and distributed systems
Hands-on experience with AWS, Kubernetes/EKS, and CI/CD pipelines
Proven track record building, operating, and optimizing LLM or AI applications in production for reliability, latency, and cost
Deep understanding of authentication, authorization, security, policy enforcement, and audit controls
Experience with observability tooling (tracing, monitoring, alerting, debugging) and incident response, with a demonstrated history of evolving production systems post-launch
Familiarity with MCP, tool integrations, agentic-system architectures, or shared AI platform and gateway development (preferred)
Experience with React or full-stack application development is a plus; strong collaboration, documentation, and communication skills expected
Compensation
$80-$85/hr - Dependent on fit and experience
Numbers & Facts
Location
San Francisco, CA (Remote)
Salary
$80–$85 Per Hour
Skills
Access Authorizationunmatched
Amazon Web Services (AWS)unmatched
Application Programming Interface (API)unmatched
Artificial Intelligence (AI)unmatched
Authenticationunmatched
Communication Skillsunmatched
Continuous Deployment/Deliveryunmatched
Continuous Improvementunmatched
Continuous Integrationunmatched
Debugging Skillsunmatched
Distributed Computingunmatched
Documentationunmatched
Engineeringunmatched
Identify Issuesunmatched
Incident Responseunmatched
MCP - Microsoft Certified Professionalunmatched
Machine Toolunmatched
Multiplatform/Cross-Platformunmatched
Node.jsunmatched
Onboardingunmatched
Production Systemsunmatched
Programming Toolsunmatched
Python Programming/Scripting Languageunmatched
React.jsunmatched
Reporting Dashboardsunmatched
Root Cause Analysisunmatched
Security Policyunmatched
Software Developmentunmatched
Software Engineeringunmatched
Standards Developmentunmatched
System Architectureunmatched
🎯
Be found by employers
5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.
Level up your application
Professional resume templates
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.