We are seeking a highly skilled Lead Site Reliability Engineer (AI & Cloud Operations) to drive the reliability, scalability, automation, and operational excellence of our cloud-native platforms and AI-powered solutions. This role will serve as a technical leader responsible for building resilient infrastructure, implementing modern DevOps and SRE practices, and enabling enterprise AI capabilities through automation, observability, and operational intelligence.
The ideal candidate combines deep expertise in AWS cloud technologies, Kubernetes, infrastructure automation, CI/CD, and incident management with hands-on experience supporting AI/ML and Generative AI platforms. This individual will partner closely with software engineering, data engineering, machine learning, security, and product teams to establish highly available systems, streamline deployments, optimize platform performance, and accelerate innovation through AI-driven operations.
Perficient is the global AI and technology consulting firm disrupting the traditional consulting model. Powered by our 7,000+ advisors, engineers, and designers, Perficient implements AI-first solutions that break conventions and deliver outcomes that matter. Proudly serving clients that represent the world's most innovative brands, and in collaboration with our powerful technology partner ecosystem, we bring deep industry expertise and data-driven design to redefine how businesses run and succeed. Perficient is different. For real. Learn more at perficient.com.
WHAT WE BELIEVE
At Perficient, we promise to challenge, champion, and celebrate our people. You will experience a unique and collaborative culture that values every voice. Join our team, and you'll become part of something truly special.
We believe in developing a workforce that is as diverse and inclusive as the clients we work with. We're committed to actively listening, learning, and acting to further advance our organization, our communities, and our future leaders… and we're not done yet.
Perficient, Inc. proudly provides equal employment opportunities (EEO) to all employees and applicants for employment without regard to race, color, religion, gender, sexual orientation, national origin, age, disability, genetic information, marital status, amnesty, or status as a protected veteran in accordance with applicable federal, state and local laws. Perficient, Inc. complies with applicable state and local laws governing non-discrimination in employment in every location in which the company has facilities. This policy applies to all terms and conditions of employment, including, but not limited to, hiring, placement, promotion, termination, layoff, recall, transfer, leaves of absence, compensation, and training. Perficient, Inc. expressly prohibits any form of unlawful employee harassment based on race, color, religion, gender, sexual orientation, national origin, age, genetic information, disability, or covered veterans. Improper interference with the ability of Perficient, Inc. employees to perform their expected job duties is absolutely not tolerated.
Disability Accommodations:
Perficient is committed to providing a barrier-free employment process with reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures. If you need assistance or accommodation due to a disability, please contact us.
Applications will be accepted until the position is filled or the posting is removed.
The salary range for this position takes into consideration a variety of factors, including but not limited to skill sets, level of experience, applicable office location, training, licensure and certifications, and other business and organizational needs. The new hire salary range displays the minimum and maximum salary targets for this position across all US locations, and the range has not been adjusted for any specific state differentials. It is not typical for a candidate to be hired at or near the top of the range for their role, and compensation decisions are dependent on the unique facts and circumstances regarding each candidate. A reasonable estimate of the current salary range for this position is $111,300 to $144,600. Please note that the salary range posted reflects the base salary only and does not include benefits or any potential variable compensation programs. Information regarding the benefits available for this position are in our benefits overview.
Disclaimer: The above statements are not intended to be a complete statement of job content, rather to act as a guide to the essential functions performed by the employee assigned to this classification. Management retains the discretion to add or change the duties of the position at any time.
#LI-RS1
Lead the design, implementation, and continuous improvement of Site Reliability Engineering (SRE) practices to ensure highly available, scalable, and resilient cloud platforms.
Architect, deploy, and support AWS-based infrastructure and services, including containerized and serverless environments.
Build, maintain, and optimize CI/CD pipelines and Infrastructure as Code (IaC) solutions to accelerate and standardize deployments.
Develop automation solutions, operational tooling, and self-healing capabilities using Python, Shell, and modern DevOps technologies.
Manage Kubernetes and container platforms, ensuring performance, scalability, and operational stability.
Establish and enhance observability through monitoring, logging, alerting, and performance management tools to proactively identify and resolve issues.
Lead incident response, root cause analysis, problem management, and service reliability improvement initiatives.
Partner with engineering, data, AI/ML, and security teams to support enterprise applications, analytics platforms, and cloud-native solutions.
Design and implement AI-driven operational capabilities, including intelligent monitoring, automated remediation, predictive analytics, and chatbot-enabled support workflows.
Support MLOps and AI platform operations, including model deployment, monitoring, governance, and lifecycle management.
Define and track reliability metrics, service-level objectives (SLOs), and operational KPIs to drive continuous improvement.
Mentor and provide technical leadership to engineering teams while promoting best practices in reliability, automation, cloud operations, and AI-enabled innovation.
Lead the design, implementation, and continuous improvement of Site Reliability Engineering (SRE) practices to ensure highly available, scalable, and resilient cloud platforms.
Architect, deploy, and support AWS-based infrastructure and services, including containerized and serverless environments.
Build, maintain, and optimize CI/CD pipelines and Infrastructure as Code (IaC) solutions to accelerate and standardize deployments.
Develop automation solutions, operational tooling, and self-healing capabilities using Python, Shell, and modern DevOps technologies.
Manage Kubernetes and container platforms, ensuring performance, scalability, and operational stability.
Establish and enhance observability through monitoring, logging, alerting, and performance management tools to proactively identify and resolve issues.
Lead incident response, root cause analysis, problem management, and service reliability improvement initiatives.
Partner with engineering, data, AI/ML, and security teams to support enterprise applications, analytics platforms, and cloud-native solutions.
Design and implement AI-driven operational capabilities, including intelligent monitoring, automated remediation, predictive analytics, and chatbot-enabled support workflows.
Support MLOps and AI platform operations, including model deployment, monitoring, governance, and lifecycle management.
Define and track reliability metrics, service-level objectives (SLOs), and operational KPIs to drive continuous improvement.
Mentor and provide technical leadership to engineering teams while promoting best practices in reliability, automation, cloud operations, and AI-enabled innovation.
| Location | NY |
| Industry | Management Consulting Services |
| Salary | $111,300–$144,600 Per Year |
| Company Size | 2,500 to 4,999 employees |
| Website | http://www.perficient.com/ |
Perficient is a leading information technology consulting firm serving Global 2000 and other enterprise customers throughout North America. We are experts in designing, building and delivering business-driven technology solutions. We help our clients gain competitive advantage by using Internet-based technologies to make their business more responsive to market opportunities and threats, strengthen relationships with customers, suppliers and partners, improve productivity and reduce information technology costs. We succeed by focusing on services and solutions that leverage best-of-breed technologies from our network of partners - companies like IBM, Microsoft, Oracle-Siebel, TIBCO, EMC Documentum and more.
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.
Free resume templatesImprove your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.
Free resume builder