Principal Software Engineer, Performance Tooling

Microsoft Corp

Mountain View, CA

JOB DETAILS
SALARY
$139,900–$274,800 Per Year
SKILLS
Algorithms, Application Programming Interface (API), Artificial Intelligence (AI), Background Investigation, Benchmarking, Business Strategy, C++ Programming Language, CUDA (Compute Unified Device Architecture), Capital Expenditure (CAPEX), Cellular Telephone, Cloud Computing, Computer Architecture, Computer Science, Computer Systems, Debugging Skills, Deep Learning, Desktop PC, Distributed Computing, Engineering, GPU (Graphics Processing Unit), Government Requirements, Hardware Architecture, Internet of Things, Kernel Programming, Leadership, Machine Tool, Mentoring, Microsoft Bing Search Engine, Microsoft Product Family, Microsoft SQL Server, Microsoft Windows Azure, Microsoft Windows Operating System, Modeling Languages, Performance Analysis, Performance Management, Performance Modeling, Performance Tuning/Optimization, Programming Languages, Python Programming/Scripting Language, Quality Management, Risk, Software Design, Software Development, Software Engineering, Systems Analysis, Systems/Internals Programming, Team Lead/Manager, Team Player, Technical Strategy, Time Management, Web Browsers
LOCATION
Mountain View, CA
POSTED
5 days ago

Overview

The Artificial Intelligence (AI) Frameworks team at Microsoft develops AI software that enables running AI models everywhere, from world's fastest AI supercomputers, to servers, desktops, mobile phones, internet of things (IoT) devices and internet browsers. We collaborate with our hardware teams and partners, both internal and external, and operate at the intersection of AI algorithmic innovation, purpose-built AI hardware, systems, and software. We are a team of highly capable and motivated people that pride themselves on a collaborative and inclusive culture. We own inference performance of OpenAI and other state of the art large language model (LLM) models and work directly with OpenAI on the models hosted on the Azure OpenAI service serving some of the largest workloads on the planet with trillions of inferences per day in major Microsoft products, including Office, Windows, Bing, SQL Server, and Dynamics.

As a Principal Software Engineer - Performance Tooling on the team, you will set technical direction across organizations, define the architecture and engineering strategy for AI performance validation and optimization, and drive execution across hardware, runtime, compiler, and model-serving partners. You will lead the design of benchmarking and performance tooling systems used to evaluate OpenAI and other frontier LLMs across GPUs and Microsoft hardware, identify systemic bottlenecks, and translate performance insights into production-ready improvements that reduce time-to-deploy, improve hardware efficiency, and directly support Microsoft Azures capex goals.

Microsoft's mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.

Responsibilities

Define and drive the technical strategy for AI performance tooling and validation across multiple layers of the AI software stack, including programming models, compilers, runtimes, libraries, and model-serving APIs.

Architect scalable benchmarking and regression-detection systems for OpenAI and other state-of-the-art LLMs across GPUs, Microsoft accelerators, and emerging AI hardware platforms.

Lead deep performance investigations across model architecture, kernels, runtimes, networking, scheduling, and hardware behavior, and guide teams toward durable optimizations for production-scale training and inference.

Establish performance quality bars, readiness signals, and operating mechanisms that enable faster model and hardware bring-up while reducing regressions, deployment risk, and total hardware footprint.

Influence and align senior engineers, researchers, product leaders, and hardware partners across Microsoft and OpenAI to deliver high-impact, production-ready AI performance improvements.

Mentor and raise the technical bar for engineers across teams by modeling engineering excellence, improving design quality, and creating reusable systems that scale beyond a single project or product line.

Qualifications

Required/Minimum Qualifications:

Bachelors Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to C++, Python OR equivalent experience.

Other Requirements:

Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. This includes passing the Microsoft Cloud background check upon hire/transfer and every two years thereafter.

Preferred/Additional Qualifications:

Masters Degree in Computer Science or related technical field AND 15+ years technical engineering experience with coding in languages including, but not limited to, C++, Python, or equivalent systems programming languages

OR Bachelors Degree in Computer Science or related technical field AND 18+ years technical engineering experience with coding in languages including, but not limited to, C++, Python, or equivalent systems programming languages

OR equivalent experience.

8+ years of practical experience building, debugging, and optimizing high-performance distributed systems, AI inference/training workloads, or accelerator-backed compute platforms.

Deep experience with DNN/LLM inference or training performance, including one or more deep learning frameworks such as PyTorch, TensorFlow, or ONNX Runtime and accelerator programming environments such as CUDA, ROCm, Triton, or equivalent.

Recognized technical depth in software engineering, distributed systems, computer architecture, GPU/accelerator architecture, and hardware/software co-design for AI workloads.

Demonstrated ability to lead end-to-end performance analysis and optimization for state-of-the-art LLMs, HPC applications, or large-scale production AI services, including expert-level use of profiling, tracing, and observability tools.

Proven ability to influence technical strategy across organizations, resolve ambiguity, and drive alignment among senior engineering, research, product, and hardware stakeholders.

Track record of independently leading multi-team or cross-organizational initiatives from strategy through execution, with measurable impact on performance, reliability, cost efficiency, or developer productivity.

#AIInfra

Software Engineering IC5 - The typical base pay range for this role across the U.S. is USD $139,900 - $274,800 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $188,000 - $304,200 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:

https://careers.microsoft.com/us/en/us-corporate-pay

This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.

Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

About the Company

M

Microsoft Corp

DO WHAT YOU LOVE
Make your mark on the world’s most used technologies. Develop the next hit mobile application. Pioneer a startup that could be the next big thing. At Microsoft, you choose your path.

Headquartered in Redmond, Washington, Microsoft is a top innovator in both the consumer and enterprise technology industry. Just a few of the many things our products do are unleash creativity, connect businesses, and make learning more fun. But our continued success is based on one thing: our employees. We hire amazing, talented people and give them the opportunities—and the tools—to succeed.

WHY MICROSOFT?
As a Microsoft employee, you’re surrounded by a diverse group of the smartest people in your field. This fosters new ideas, better business results, and creates a dynamic work environment. In the office, you’re constantly challenged and supported by your colleagues. Every day holds something new and exciting.

We also offer unparalleled depth and breadth of career opportunities. As an industry leader in multiple fields, working for Microsoft means being able to do whatever you feel passionate about—and being able to make an impact in that field. From day one, we give our employees significant responsibility. This means that you’ll know that you directly contributed to something that has a positive impact on people worldwide. Whether you choose to work in management, dive deep into the newest technology, or explore multiple professions, you’ll find everything you need at Microsoft to drive your career—and to make a difference.

WE GET IT – YOU’RE MORE THAN YOUR JOB
Everyone works differently and is motivated by different things. We also understand that there’s more to you than your job. That’s why we offer competitive pay and a wide assortment of benefits-- to help you make the most of life at work and away from it.

GET THE BALL ROLLING
COMPANY SIZE
10,000 employees or more
INDUSTRY
Computer Software
FOUNDED
1975
WEBSITE
http://www.microsoft.com