Role: HPC (High-Performance Computing) Consultant
Location: Remote
Duration: Fulltime
Interview Type: Video
Must Have: Kubernetes + Slurm + NVIDIA GPU + AI/ML Infrastructure + Terraform + Python + AWS (EKS/FSx Lustre) + HPC Storage + Monitoring/SRE
We are looking for engineers with expertise across Kubernetes, cloud infrastructure, HPC platforms, GPU computing, Terraform, and automation. Depending on experience, candidates may be considered for Platform Engineer, Kubernetes Engineer, HPC Engineer, Cloud Infrastructure Engineer, DevOps Engineer, AI Infrastructure Engineer, or Site Reliability Engineer (SRE) roles.
Key Responsibilities
Design, deploy, operate, and support large-scale Kubernetes platforms across AWS, GCP, CoreWeave, OCI, and other cloud environments.
Manage Kubernetes cluster lifecycle activities including provisioning, scaling, node pool management, upgrades, troubleshooting, and performance optimization.
Support AI/ML and HPC workloads, including GPU-enabled compute infrastructure for training and inference environments.
Provision and automate cloud and infrastructure resources using Terraform and CI/CD pipelines.
Troubleshoot Kubernetes scheduling, networking, storage, and platform reliability issues.
Implement monitoring, observability, alerting, SLIs/SLOs, and incident response processes.
Collaborate with Networking, Security, Storage, AI/ML, Data Engineering, and Application teams.
Develop automation and operational tooling using Python and cloud-native technologies.
Participate in production support, root cause analysis, capacity planning, and platform optimization initiatives.
Required Skills
Kubernetes & Container Platforms
Kubernetes (EKS, GKE, AKS, OpenShift, CoreWeave CKS)
Cluster Lifecycle Management
Node Pool Management
Scheduler Troubleshooting
CNI Troubleshooting
Networking Policies
RBAC
Helm
Autoscaling
Rolling Upgrades
Cloud Infrastructure
AWS (EC2, S3, IAM, VPC, EKS, EFS, FSx for Lustre)
Google Cloud Platform (GCP)
OCI (Oracle Cloud Infrastructure)
Azure (Preferred)
Multi-Cloud Infrastructure
Infrastructure as Code & Automation
Terraform
Infrastructure as Code (IaC)
CI/CD Pipelines
GitHub Actions / Jenkins / GitLab CI
Ansible (Preferred)
Programming & Scripting
Python
Bash/Shell Scripting
Automation Development
REST API Integration
HPC & GPU Infrastructure (Preferred)
High-Performance Computing (HPC)
Slurm
NVIDIA GPU Platforms
GPU Scheduling
CUDA
Distributed Computing
AI/ML Infrastructure
Monitoring & Reliability
Prometheus
Grafana
Datadog
Splunk
Cloud Monitoring
SLI/SLO Management
Incident Response
Root Cause Analysis (RCA)
Preferred Experience
Experience with any of the following is highly desirable:
AI/ML Infrastructure Platforms
Kubeflow
KServe
Ray
MLflow
vLLM
Vector Databases
Distributed Training Platforms
CoreWeave
AWS ParallelCluster
FSx for Lustre
Lustre
WekaFS
InfiniBand
RDMA
AI Training & Inference Workloads
The pay range for this role is USD $180,000- $200,000 per annum including any bonuses or variable pay. Tech Mahindra also offers benefits like medical, vision, dental, life, disability insurance and paid time off (including holidays, parental leave, and sick leave, as required by law). Ask our recruiters for more details on our Benefits package. The exact offer terms will depend on the skill level, educational qualifications, experience and location of the candidate.
AI tools may assist in the recruitment process; however, all hiring decisions are made by the recruitment team based on a comprehensive evaluation of candidates.
“Tech Mahindra is an Equal Employment Opportunity employer. We promote and support a diverse workforce at all levels of the company. All qualified applicants will receive consideration for employment without regard to race, religion, color, sex, age, national origin, or disability. All applicants will be evaluated solely on the basis of their ability, competence, and performance of the essential functions of their positions with or without reasonable accommodations. Reasonable accommodations also are available in the hiring process for applicants with disabilities. Candidates can request a reasonable accommodation by contacting the company ADA Coordinator at ADA_Accomodations@TechMahindra.com.”
| Location | (Remote) |
| Job Type | Full-time |
| Salary | $180,000–$200,000 Per Year |
Requirements:
4+ years in infrastructure engineering, cloud platforms, or HPC.
Kubernetes is the core requirement. You should have hands-on experience operating clusters at meaningful scale: node pool sizing, scheduler debugging, CNI troubleshooting, and rolling upgrades across large fleets. Candidates whose K8s experience is limited to small or local environments are unlikely to be a fit.
Terraform proficiency. You'll write and review infrastructure-as-code daily.
Working knowledge of AWS (EC2, S3, EFS, FSx for Lustre).
Python for tooling and automation.
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.
Free resume templatesImprove your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.
Free resume builder