KEY RESPONSIBILITIES:
Manage platform deployment, availability, and stability for data center CPU/GPU systems
Lead system-level debugging efforts involving hardware, firmware, Linux, networking, power, and thermal behavior
Develop structured debug strategies, validation flows, and failure-analysis methodologies.
Installation, configuration of various OS Distros from console. Setup Network gear, perform and validate Network configurations.
Configure file systems using VG/LV/Partitions
Integrate automated testing in CI/CD environment (e.g. Jenkins, ansible)
Provide logs and statistics that will help in further debug of issues.
Own inventory database management and administration through a managed system.
Participate in the Agile method of planning, delivery, and collaboration with internal and scaled agile teams.
Work with a managed ticketing system and communicate clearly on activities and steps.
Track daily activities, prioritize issues, assign work, and monitor progress to resolution
Collaborate with silicon, firmware, validation, networking, and operations teams to assess risks and requirements
Execute hands-on laboratory validation to ensure systems operate as intended prior to and during deployment
Partner with OEMs, ODMs, and vendors to resolve issues and improve platform reliability
Drive continuous improvement in tools, processes, and platform readiness
EEO:
“Mindlance is an Equal Opportunity Employer and does not discriminate in employment on the basis of – Minority/Gender/Disability/Religion/LGBTQI/Age/Veterans.”