Microsoft Corp logo

Service Engineer

Microsoft Corp
  • Redmond, WA
  • $102,100–$202,200 Per Year
30+ days ago

Job Description

Overview

Are you a customer-obsessed, AI-curious problem-solver who thrives in an inclusive, collaborative global team? Join Engineering Operations (EngOps) - the organization driving operational excellence across the Microsoft Cloud to strengthen quality, reliability, security, and customer trust. As part of EngOps, you'll design solutions that prevent issues before they happen, embed AI-powered automation, and turn signals into actions that deliver measurable customer impact. Our culture of empowerment, inclusion, and growth mindset defines how we work.

The Customer Reliability Engineering (CRE) team within Azure EngOps is a top-level pillar of Azure Engineering responsible for world-class live-site management, customer reliability engagements, modern customer-first experiences for scale, and drives deep customer insights and empathy into the broader Azure Engineering organization. Our "no dead-end's" philosophy ensures that every customer, regardless of size or scale, can realize their full potential through the Microsoft Cloud

We are seeking decisive and experienced Service Engineers for Live Site Issues, Problem Management and driving Customer reliability space. This role is accountable for enhancing the customer experience across Azure, including First Party Services. The ideal candidate will demonstrate strong breadth in managing complex, highly available services, paired with deep technical expertise in Azure Core Services and their inter dependencies. You will work closely with Customers, First Parties, Customer Support, Livesite, and Engineering teams to deliver critical, customer-facing features. Success in this role requires the ability to influence and collaborate across many Azure servicing teams to ensure customer needs are met.

In addition, this role includes on-call responsibilities for managing and resolving complex multi-service outages. It requires the ability to remain effective under pressure, apply broad technical and analytical skills, and coordinate seamlessly with internal service teams and stakeholders. Strong communication skills-both written and verbal-are essential. You will also lead the evolution of Azures Incident Management practice through Post-Incident Reviews, process development, and system automation. By leveraging telemetry and metrics, you will identify and drive platform-wide improvements with global impact. You'll be the single point of command and control during high-severity incidents, orchestrating cross-functional engineering, operations, and communications to minimize impact, restore services quickly, and protect the trust of our global customer base.

This role offers a unique opportunity to make immediate impact, improve systems at scale.

#engops, #CRE

Responsibilities

  • To be successful in this role, you must have a great track record of customer compassion, an engineering mindset, an innate aptitude for agility, and technical excellence in software engineering. Collaborate closely with Engineering/PM to ensure the availability, performance of Live Site and the satisfaction of our customers
  • Lead and manage high-severity incidents across Azure services, serving as the single point of accountability to ensure rapid detection, triage, resolution, and customer communication.
  • Act as the central authority during live site incidents, driving real-time decision-making and coordination across Engineering, Support, PM, Communications, and Field teams.
  • Contribute to the design of V. Next architecture for Cloud infrastructure services, based on Customer/ First party engagements.
  • Engage in major production triage efforts and work with different teams in the identification of root cause of highly impactful or complex issues as required and identify Product gaps and work with Product teams to bridge the gaps.
  • Partner closely with Software developers, Product Managers, architects, and Infrastructure teams to drive delivery of sustainable and reusable design solution patterns to ensure non-functional production support requirements are adopted early in the Migration /Deployment
  • Promote a customer-first culture by prioritizing availability, reliability, and platform trust in every response.
  • Participate in the on-call rotation.
  • Analyze customer-impacting signals from telemetry, support cases, and feedback to identify root causes, drive incident reviews (RCAs/PIRs), and implement preventative service improvements.
  • Drive continuous improvement of the Azure platform by incorporating learnings from live site events and customer feedback, ensuring improved reliability, observability, and supportability.
  • Collaborate closely with Engineering and Product teams to influence and implement service resiliency enhancements, auto-remediation tools, and customer-centric mitigation strategies.
  • Identify and advocate for customer self-service capabilities, improved documentation, and scalable solutions that empower customers to resolve common issues independently.
  • Design and drive adoption of incident response playbooks, mitigation levers, and operational frameworks aligned to real-world support scenarios and strategic customer needs.
  • Contribute to the design of next-generation architecture for cloud infrastructure services with a focus on reliability and strategic customer support outcomes.
  • Build and maintain cross-functional partnerships, ensuring alignment across engineering, business, and support organizations.
  • Be data-driven and results-focused, using metrics to evaluate incident response effectiveness and platform health.
  • Bring an engineering mindset to operational challenges, balancing agility, scalability, and technical excellence.
  • Exhibit strong cross-team collaboration, engineering mindset, and results-oriented execution under pressure

Qualifications

Required Qualifications:

  • Bachelors Degree in Computer Science, Information Technology, Mechanical Engineering, Electrical Engineering, Aerospace Engineering, Data Science, Cybersecurity, or related field AND 2+ years technical experience in software engineering, network engineering, service engineering, systems engineering, or industrial controls
  • OR equivalent experience

Other Requirements:

  • Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include but are not limited to the following specialized security screenings: Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud Background Check upon hire/transfer and every two years thereafter

Preferred Qualifications:

  • 2-4+ Yrs of experience in roles cloud operations, incident response, SRE or large-scale system engineering preferably in platforms like Azure, AWS, or GCP.
  • Must hav Service Engineering experience in a 24 x 7 x 365 enterprise environments.
  • Exceptional command-and-control communication skills-able to drive clarity and direction with customers - internal Microsoft stake holders and third-party vendors during ambiguity and chaos.
  • Deep understanding of cloud architecture patterns, microservices, and containerization.
  • Demonstrated ability to make decisions quickly, under pressure, and with limited data-without compromising long-term reliability.
  • Familiarity with monitoring and observability tools (e.g., Grafana, Prometheus, Datadog, Splunk, New Relic).
  • Fluency in one or more automation languages (PowerShell, Python, CLI etc.)
  • Understanding ITIL or other incident management frameworks is a must.
  • Understand High Availability, Disaster Recovery, Business Continuity, Performance Tuning.
  • Demonstrates strategic thinking, quantitative and analytical skills, team leadership, and collaboration.
  • Excellent problem resolution, judgment, negotiating and decision-making skills .
  • Desired Strong knowledge of Windows Platform or Linux, developer tools and ability to diagnose and debug user code.
  • Effectively manage and prioritize multiple tasks in accordance with high level objectives/projects.
  • 3+ Years of demonstrated experience as an Incident Commander or Crisis Manager for critical, high-severity incidents in high-availability, distributed environments.
  • Experience with SRE (Site Reliability Engineering) principles and practices.
  • Exposure to chaos engineering, fault injection, or high availability architecture.
  • AI/ML Experience: [Beginner to Intermediate]
  • Familiarity with how AI/ML models are integrated into cloud infrastructure and their potential failure modes.
  • Experience using AI-powered tools for incident analysis, log correlation, or predictive alerting.
  • An understanding of the challenges and risks associated with AI/ML systems in a production environment.
  • Excellent communication skill (written + verbal) in English, especially in high-pressure scenarios.
  • Ability to communicate with a variety of audiences; including high-profile customers, executive management, and engineering teams.
  • Desired BS/BA in Computer Science, Engineering, Math or equivalent experience.
  • Certifications: Relevant cloud certifications (e.g., AWS Certified DevOps Engineer, Azure Solutions Architect, GCP Professional Cloud Architect).
  • Certifications in ITIL, SRE, or other relevant frameworks.

Service Engineering IC3 - The typical base pay range for this role across the U.S. is USD $102,100 - $202,200 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $133,800 - $219,200 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:

https://careers.microsoft.com/us/en/us-corporate-pay

This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.

Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Numbers & Facts

LocationRedmond, WA
IndustryComputer Software
Salary$102,100–$202,200 Per Year
Company Size10,000 employees or more
Year Founded1975
Websitehttp://www.microsoft.com

About Company

DO WHAT YOU LOVE
Make your mark on the world’s most used technologies. Develop the next hit mobile application. Pioneer a startup that could be the next big thing. At Microsoft, you choose your path.

Headquartered in Redmond, Washington, Microsoft is a top innovator in both the consumer and enterprise technology industry. Just a few of the many things our products do are unleash creativity, connect businesses, and make learning more fun. But our continued success is based on one thing: our employees. We hire amazing, talented people and give them the opportunities—and the tools—to succeed.

WHY MICROSOFT?
As a Microsoft employee, you’re surrounded by a diverse group of the smartest people in your field. This fosters new ideas, better business results, and creates a dynamic work environment. In the office, you’re constantly challenged and supported by your colleagues. Every day holds something new and exciting.

We also offer unparalleled depth and breadth of career opportunities. As an industry leader in multiple fields, working for Microsoft means being able to do whatever you feel passionate about—and being able to make an impact in that field. From day one, we give our employees significant responsibility. This means that you’ll know that you directly contributed to something that has a positive impact on people worldwide. Whether you choose to work in management, dive deep into the newest technology, or explore multiple professions, you’ll find everything you need at Microsoft to drive your career—and to make a difference.

WE GET IT – YOU’RE MORE THAN YOUR JOB
Everyone works differently and is motivated by different things. We also understand that there’s more to you than your job. That’s why we offer competitive pay and a wide assortment of benefits-- to help you make the most of life at work and away from it.

GET THE BALL ROLLING

Skills

  • Aerospace Engineeringunmatched
  • Amazon Web Services (AWS)unmatched
  • Analysis Skillsunmatched
  • Artificial Intelligence (AI)unmatched
  • Automationunmatched
  • Automation System Developmentunmatched
  • Background Investigationunmatched
  • Business Supportunmatched
  • Cloud Architectureunmatched
  • Cloud Computingunmatched
  • Communication Skillsunmatched
  • Computer Scienceunmatched
  • Continuous Improvementunmatched
  • Cross-Functionalunmatched
  • Customer Experienceunmatched
  • Customer Relationsunmatched
  • Customer Satisfactionunmatched
  • Customer Support/Serviceunmatched
  • Customer/Client Researchunmatched
  • Data Scienceunmatched
  • Debugging Skillsunmatched
  • DevOpsunmatched
  • Disaster Recoveryunmatched
  • Documentationunmatched
  • Electrical Engineeringunmatched
  • Engineeringunmatched
  • English Languageunmatched
  • Establish Prioritiesunmatched
  • GCP (Good Clinical Practices)unmatched
  • High Availabilityunmatched
  • Hyperion Pillarunmatched
  • ITIL (IT Infrastructure Library)unmatched
  • Identify Issuesunmatched
  • Incident Managementunmatched
  • Incident Responseunmatched
  • Industrial Engineeringunmatched
  • Information Technology & Information Systemsunmatched
  • Infrastructure as a Service (IaaS)unmatched
  • Injectionsunmatched
  • Internet Securityunmatched
  • Large-Scale Systemsunmatched
  • Linux Operating Systemunmatched
  • Mathematicsunmatched
  • Mechanical Engineeringunmatched
  • Metricsunmatched
  • Microservicesunmatched
  • Microsoft Product Familyunmatched
  • Microsoft Windows Azureunmatched
  • Microsoft Windows Operating Systemunmatched
  • Microsoft Windows System Internals/Programmingunmatched
  • Multitaskingunmatched
  • Negotiation Skillsunmatched
  • Network Architecture/Engineeringunmatched
  • On Callunmatched
  • Operational Communicationsunmatched
  • Organizational Skillsunmatched
  • Performance Tuning/Optimizationunmatched
  • Philosophyunmatched
  • Presentation/Verbal Skillsunmatched
  • Problem Solving Skillsunmatched
  • Process Developmentunmatched
  • Process Improvementunmatched
  • Product Strategyunmatched
  • Production Supportunmatched
  • Production Systemsunmatched
  • Programming Toolsunmatched
  • Python Programming/Scripting Languageunmatched
  • Quantitative Analysisunmatched
  • Reliability Engineeringunmatched
  • Resolve Customer Issuesunmatched
  • Root Cause Analysisunmatched
  • Software Developmentunmatched
  • Software Engineeringunmatched
  • Splunkunmatched
  • Systems Engineeringunmatched
  • Team Lead/Managerunmatched
  • Team Playerunmatched
  • Telemetryunmatched
  • Telephone Skillsunmatched
  • Vehicle Drivingunmatched
  • Windows PowerShellunmatched
  • Writing Skillsunmatched

Be found by employers

5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

Level up your application

Professional resume templates

Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

Free resume templates

Free resume builder

Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

Free resume builder