Overview
CADET (Customer and Analytics Driven Evals Team) is building a customer-grounded quality system for Copilot. Our mission is to rapidly identify the customer scenarios that matter most, represent them faithfully in evaluation and learning assets, run quality gates continuously, and turn every important failure into reusable product and model improvements. We bring together DSAT and other product signals, deep customer engagements to create representative eval sets. Operating in a fast-paced environment, we connect customer grounded quality issues with quality teams to advance Copilot quality and product innovation.
We are looking for a Senior Data Scientist to build the measurement and decision system that determines where CADET invests and whether Copilot is improving for the customers and intents that matter most. You will combine DSAT, usage, customer engagement, product feedback, evaluation, and other quality signals to create a representation- and coverage-aware view of customer quality. You will define the taxonomies, metrics, prioritization models, analyses, and reporting mechanisms that turn a fragmented signal landscape into clear decisions. You will synthesize quality opportunities, investigate loss patterns, shape evaluation portfolios, and ensure recurring quality gates reflect real customer experiences rather than static benchmarks.
Microsoft's mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.
Responsibilities
Build and operate a data-driven prioritization framework for CADET quality forums and partner teams that combines DSAT, customer impact, usage, severity, strategic importance, representation, and current evaluation coverage to guide quality investments.
Define and maintain intent, sub-intent, customer scenario, coverage, and loss-pattern taxonomies that can be used consistently across signal intake, triage, evaluation, and reporting.
Measure how well evaluation portfolios represent production traffic, customer segments, workflow complexity, locales, grounding paths, and material failure modes.
Identify underrepresented customers, intents, scenarios, and loss patterns, and translate those gaps into evaluation and data-collection priorities.
Design quality gates for top intents and top customers, including success thresholds, segmentation, run cadence, escalation criteria, and reporting.
Build recurring scorecards that connect offline evaluation movement with online measures such as DSAT, task completion, retries, abandonment, and escalation; detect meaningful quality changes; and alert accountable owners when action is required.
Analyze offline-online agreement, evaluation freshness, regression coverage, grader reliability, and quality movement over time.
Develop sampling, weighting, deduplication, clustering, and trend-detection approaches for noisy customer and product signals using resource- and performance-optimized data-analysis solutions that make effective use of CPU, GPU, and platform capacity.
Use causal and experimental methods where appropriate to distinguish correlation, attribution, and treatment impact, and translate the findings into concrete product, model, data, and evaluation investment decisions.
Partner to translate customer evidence into valid task distributions, datasets, metrics, and reward signals and to encode metrics, taxonomies, data-quality checks, and reporting into automated pipelines.
Leverage team signals from customer engagements to understand workflows, business impact, expected outcomes, and gaps hidden by aggregate metrics; produce clear recommendations for product, model, data, and evaluation investments; and communicate them to senior leaders.
Establish solid practices for data provenance, privacy, responsible use, reproducibility, and metric governance.
Qualifications
Required Qualifications:
Preferred Qualifications:
#cadets
Data Science IC4 - The typical base pay range for this role across the U.S. is USD $119,800 - $234,700 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $160,200 - $261,000 per year.
Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay
This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.
Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.
| Location | Redmond, WA |
| Industry | Computer Software |
| Salary | $119,800–$234,700 Per Year |
| Company Size | 10,000 employees or more |
| Year Founded | 1975 |
| Website | http://www.microsoft.com |
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.
Free resume templatesImprove your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.
Free resume builder