Want to know if you’re a fit? Upload your resume and let our AI show you.
Skills
Amazon Web Services (AWS)unmatched
Artificial Intelligence (AI)unmatched
Automationunmatched
Best Practicesunmatched
Cloud Applicationsunmatched
Cloud Computingunmatched
Computer Securityunmatched
Computer Systemsunmatched
Continuous Deployment/Deliveryunmatched
Continuous Integrationunmatched
DevOpsunmatched
Distributed Computingunmatched
Dockerunmatched
Ecosystemsunmatched
GCP (Good Clinical Practices)unmatched
GPU (Graphics Processing Unit)unmatched
GitHubunmatched
High Availabilityunmatched
Identify Issuesunmatched
Infrastructure Softwareunmatched
Jenkinsunmatched
Linux Administrationunmatched
Machine Learningunmatched
Machine Toolunmatched
Microsoft Windows Azureunmatched
Performance Tuning/Optimizationunmatched
Python Programming/Scripting Languageunmatched
Reliability Engineeringunmatched
Root Cause Analysisunmatched
Scalable System Developmentunmatched
Scripting (Scripting Languages)unmatched
Security Infrastructureunmatched
Software Development Lifecycle (SDLC)unmatched
Splunkunmatched
Vulnerability Scannersunmatched
Description
Job Description:
We are seeking a highly skilled DevOps / Site Reliability Engineer (SRE) with experience supporting modern AI platforms and cloud-native infrastructure. This role will focus on building scalable, reliable infrastructure for AI workloads while partnering closely with security and engineering teams to operationalize findings from Mythos AI, an emerging AI-driven security platform used to identify code vulnerabilities and infrastructure risks.
While prior hands-on experience with Mythos AI is not expected, candidates should understand its purpose within the AI security ecosystem and be comfortable implementing the remediation work it identifies.
This position is ideal for an engineer who enjoys automating infrastructure, improving software delivery pipelines, and supporting the rapid adoption of AI technologies in enterprise environments.
Responsibilities:
Design, build, and maintain highly available infrastructure supporting AI and machine learning platforms.
Develop scalable platform engineering solutions that enable reliable deployment and operation of AI services.
Partner with development and security teams to remediate vulnerabilities and infrastructure issues identified by Mythos AI.
Improve platform reliability through automation, monitoring, observability, and proactive performance tuning.
Build and maintain robust CI/CD pipelines for application and infrastructure deployments.
Automate operational workflows using Python and Infrastructure-as-Code practices.
Implement DevSecOps best practices throughout the software development lifecycle.
Support containerized workloads and cloud-native applications.
Troubleshoot production issues, perform root cause analysis, and implement long-term reliability improvements.
Optimize deployment strategies, release automation, and infrastructure scalability.
Collaborate with AI engineering teams to ensure AI services are secure, resilient, and production-ready.
Required Qualifications:
5+ years of experience in DevOps, Site Reliability Engineering, or Platform Engineering
Strong experience designing and maintaining CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins, Azure DevOps, or similar)
Strong Python scripting and automation skills
Experience supporting cloud infrastructure (AWS, Azure, or GCP)
Experience with Infrastructure as Code (Terraform, CloudFormation, or Pulumi)
Hands-on experience with Docker and Kubernetes
Strong understanding of Linux systems administration