Platform Engineer

AMroute LLC
  • St. Louis, MO
  • Quick Apply
20 days ago

Job Description

Job Title : Platform Engineer
Note : US Citizen or Green Card candidates

Duration : 160 hours
Location: St. Louis, MO
100% Onsite


Role Overview
Highly skilled AI Infrastructure Engineer to design, build, and operate scalable GPU-enabled Kubernetes platforms for AI/ML workloads.

Must have skills: Kubernetes, Linux, Terraform, GPUs and NVIDIA stack

Required Qualifications
  • 3–8+ years' experience
  • Strong Kubernetes knowledge
  • Experience with GPUs and NVIDIA stack
  • Linux (Ubuntu) expertise
  • Experience with Terraform
Preferred Qualifications
  • Longhorn or Ceph experience
  • Canonical ecosystem (MAAS, Juju)
  • AI/ML tools like Kubeflow
  • Certifications (CKA, NVIDIA)
Soft Skills
  • Strong problem-solving and troubleshooting mindset
  • Ability to collaborate with cross-functional teams (ML engineers, data scientists)
  • Clear communication and documentation skills
  • Passion for automation and platform scalability

Key Responsibilities
  • Design and manage Kubernetes clusters
  • Build GPU-enabled infrastructure
  • Deploy Longhorn storage
  • Automate infrastructure using Terraform
  • Monitor systems using Prometheus and Grafana
  • Knowledge Transfer & Client Enablement
  • Provide structured knowledge transfer (KT) sessions to client teams on all core platform components, including:
    • Kubernetes architecture, operations, and troubleshooting
    • GPU infrastructure (NVIDIA stack, scheduling, resource optimization)
    • Longhorn storage management and performance tuning
    • Canonical ecosystem tools (MAAS, Juju, Charmed Kubernetes)
  • Develop and deliver technical documentation, runbooks, and training materials to support ongoing operations
  • Conduct hands-on workshops and guided sessions to enable client teams to independently manage and scale the platform
  • Act as a technical advisor, helping client stakeholders understand best practices in:
  • Cloud-native infrastructure
    o AI/ML platform operations
    o Reliability, performance, and cost optimization
  • Ensure smooth handoff of production systems with full operational readiness and support knowledge

Numbers & Facts

LocationSt. Louis, MO

Skills

  • Artificial Intelligence (AI)unmatched
  • Automationunmatched
  • Best Practicesunmatched
  • Cloud Computingunmatched
  • Communication Skillsunmatched
  • Cost Controlunmatched
  • Cross-Functionalunmatched
  • Data Scienceunmatched
  • Documentationunmatched
  • Ecosystemsunmatched
  • GPU (Graphics Processing Unit)unmatched
  • Identify Issuesunmatched
  • Knowledge Transferunmatched
  • Linux Operating Systemunmatched
  • Operational Supportunmatched
  • Performance Tuning/Optimizationunmatched
  • Problem Solving Skillsunmatched
  • Production Systemsunmatched
  • Scalable System Developmentunmatched
  • System Operationsunmatched
  • Team Playerunmatched
  • Technical Deliveryunmatched
  • Technical Supportunmatched
  • Technical Writingunmatched
  • Ubuntuunmatched
  • United States Citizenunmatched

Be found by employers

5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

Level up your application

Professional resume templates

Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

Free resume templates

Free resume builder

Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

Free resume builder