Production Support Engineer

ASM Tech Solutions LLC
  • Lake mary, FL
  • Quick Apply
1 day ago

Job Description

As of now regular shift during US hours. Hybrid work ( 3 days from Office is Mandatory and 2 days remote)
Location: Lake Mary or 240 NY Office

Key Responsibilities
  1. Monitor, troubleshoot, and resolve Level 2 production incidents across AI platforms, cloud infrastructure, data pipelines, model-serving environments, and associated services.
  2. Provide hands-on support for deployment, orchestration, and operational management of AI/ML workloads across cloud-native environments.
  3. Build and maintain monitoring, alerting, and observability capabilities for infrastructure, applications, data pipelines, model operations, and distributed compute workloads.
  4. Perform root cause analysis for production issues, implement permanent fixes, and drive problem-management activities to improve service reliability.
  5. Collaborate with DevOps, MLOps, data engineering, platform engineering, and application teams to maintain and enhance CI/CD and deployment automation.
  6. Develop and support automation, self-service tooling, recovery mechanisms, and self-healing controls to reduce manual operational effort.
  7. Monitor and troubleshoot batch processes, workflow orchestration, data ingestion, model training, and model deployment failures.
  8. Apply ITIL-based incident, problem, change, and release management processes to support stable production operations.
  9. Analyse support-ticket and incident trends; recommend and implement operational improvements, including AI-driven automation where appropriate.

Qualifications & Skills
Mandatory:
  1. 3–5 years of experience in Level 2 application, platform, DevOps, or production support roles.
  2. Strong hands-on experience with UNIX/Linux, SQL, and shell or Python scripting.
  3. Experience troubleshooting cloud-native applications, distributed systems, containers, and Kubernetes-based environments.
  4. Working knowledge of CI/CD pipelines, deployment automation, and source-control platforms such as GitLab.
  5. Experience with monitoring, logging, and observability tools such as Splunk, Grafana, AppDynamics, Prometheus, or similar tools.
  6. Understanding of incident management, root cause analysis, problem management, and ITIL support processes.
  7. Strong analytical and problem-solving skills, with a client-service mindset.
  8. Ability to troubleshoot data-pipeline, workflow, API, and production deployment issues.

Good-to-Have:
  1. Exposure to MLOps practices, including model deployment, model monitoring, feature/data pipelines, and AI workload orchestration.
  2. Experience with cloud platforms such as AWS, Azure, or GCP.
  3. Experience with infrastructure automation and configuration-management tools such as Ansible, Terraform, or Ansible Tower.
  4. Familiarity with workflow and job-scheduling tools, such as Contro-M or equivalent enterprise schedulers.
  5. Knowledge of Docker, Kubernetes, and distributed compute/data-processing technologies.
  6. Exposure to Kafka, MQ, or other messaging and event-streaming platforms.
  7. Experience implementing self-healing, auto-remediation, resilience, backup, or disaster-recovery mechanisms.
  8. Familiarity with AI-driven operational tooling, ticket-trend analysis, or automation agents.

Numbers & Facts

LocationLake mary, FL

Skills

  • Amazon Web Services (AWS)unmatched
  • Analysis Skillsunmatched
  • Ansibleunmatched
  • Apache Kafkaunmatched
  • Application Programming Interface (API)unmatched
  • Artificial Intelligence (AI)unmatched
  • Automationunmatched
  • Cloud Applicationsunmatched
  • Cloud Computingunmatched
  • Configuration Managementunmatched
  • Continuous Deployment/Deliveryunmatched
  • Continuous Integrationunmatched
  • Customer Support/Serviceunmatched
  • Data Managementunmatched
  • Data Modelingunmatched
  • DevOpsunmatched
  • Disaster Recoveryunmatched
  • Distributed Computingunmatched
  • Dockerunmatched
  • GCP (Good Clinical Practices)unmatched
  • ITIL (IT Infrastructure Library)unmatched
  • Identify Issuesunmatched
  • Incident Managementunmatched
  • Linux Operating Systemunmatched
  • Machine Toolunmatched
  • Messaging Middlewareunmatched
  • Microsoft Windows Azureunmatched
  • Multiplatform/Cross-Platformunmatched
  • Operational Improvementunmatched
  • Operational Supportunmatched
  • Operations Managementunmatched
  • Problem Solving Skillsunmatched
  • Production Supportunmatched
  • Python Programming/Scripting Languageunmatched
  • Release Management/Engineeringunmatched
  • Reliability Engineeringunmatched
  • Root Cause Analysisunmatched
  • SQL (Structured Query Language)unmatched
  • Splunkunmatched
  • Technical Supportunmatched
  • Trend Analysisunmatched
  • Unix Operating Systemsunmatched
  • Unix Shell Programmingunmatched

Be found by employers

5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

Level up your application

Professional resume templates

Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

Free resume templates

Free resume builder

Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

Free resume builder