Apple Inc logo

Senior Applied Scientist, Multilingual AI Evaluation

Apple Inc

  • Seattle, WA
  • 5 days ago
    Want to know if you’re a fit?
    Upload your resume and let our AI show you.

    Skills

    • Appleunmatched
    • Artificial Intelligence (AI)unmatched
    • Benchmarkingunmatched
    • Communication Skillsunmatched
    • Computational Linguisticsunmatched
    • Computer Scienceunmatched
    • Cross-Functionalunmatched
    • Data Setsunmatched
    • English Languageunmatched
    • Human Interactionunmatched
    • Internationalizationunmatched
    • JAX (Java API for XML)unmatched
    • Linguisticsunmatched
    • Localizationunmatched
    • Machine Toolunmatched
    • Metricsunmatched
    • Modeling Languagesunmatched
    • Multilingualunmatched
    • Natural Language Processing (NLP)unmatched
    • Presentation/Verbal Skillsunmatched
    • Process Modelingunmatched
    • Publicationsunmatched
    • Python Programming/Scripting Languageunmatched
    • Scientific Researchunmatched
    • Scripting (Scripting Languages)unmatched
    • Software Engineeringunmatched
    • Test Designunmatched
    • Vehicle Drivingunmatched
    • Workflow Analysisunmatched
    • Writing Skillsunmatched

    Description

    AI systems are only as trustworthy as the methods used to evaluate them. At Apple, where AI powers experiences for billions of people around the world, getting evaluation right is not a support function-it is a foundational science. Our team, part of Apple Services Engineering, is building that scientific foundation: rigorous, scalable evaluation methodology for LLMs, agentic systems, and human-AI interaction.

    We're looking for a senior applied scientist to make the evaluation tooling we build work across every language and culture Apple serves. This is a role for someone who is fluent in both modern AI and the science of language, and who can set direction and drive initiatives independently, not just execute them. You'll do this on a deeply interdisciplinary team working alongside ML researchers, measurement scientists, and platform engineers.

    In this role, you'll help ensure Apple's AI features work well across languages and cultures. Your goal is to make our evaluation tooling multilingual from the start so that engineers building AI features can design, test, and ship across the world from day one. It's a broad applied science role: you'll shape how Apple evaluates AI wherever the hardest questions are, and you'll have the opportunity to publish novel work.

    The scientific challenge is real. How do we ensure we consistently evaluate AI features across different grammar, script, or cultural norms and how do we do this at scale? You'll bring linguistic judgment to questions like these and, working with measurement scientists and ML researchers, turn it into validated methodology that holds across dozens of languages.

    This is a hands-on role. You'll design and implement your own methods in Python, working closely with research and engineering partners, while staying focused on the science of getting evaluation right. Extend Apple's AI evaluation methodology and tooling to new languages and locales, so our AI experiences are equally capable, accurate, and culturally appropriate - not just translated English. Design and validate evaluation methods, benchmarks, and metrics that capture language- and culture-specific phenomena: grammar, morphology, script, register, dialect, code-switching, and cultural norms. Build and curate high-quality multilingual datasets and human evaluation protocols, partnering with linguists and native-speaker annotators. Investigate how LLMs and agentic systems behave across languages, identifying systematic failure modes, capability gaps, and quality disparities between high- and low-resource languages. Partner with engineers to productionize your methods so they run reliably and at scale, implementing your own work in Python. Communicate findings clearly - translating results into actionable guidance for model and product teams, and publishing novel work where appropriate.MS in Linguistics, Computational Linguistics, NLP, Computer Science, or a related field - or equivalent research/work experience. Deep expertise in linguistics, with working fluency in the structure of multiple languages beyond English. Strong proficiency in Python. Solid understanding of LLMs and AI evaluation fundamentals, including how language models process and generate across languages. Demonstrated experience shipping or evaluating features across multiple languages or locales. Experience designing benchmarks, datasets, or human evaluation protocols, with attention to statistical rigor and reproducibility. Ability to drive initiatives independently and collaborate across a cross-functional, interdisciplinary team. Strong written and verbal communication skills.PhD in Linguistics, Computational Linguistics, or NLP with a focus on multilingual or cross-lingual modeling. Publications in NLP, multilingual evaluation, or evaluation methodology. Hands-on experience with modern ML frameworks (PyTorch, JAX) and with fine-tuning or evaluating LLMs. Experience with low-resource languages, dialectal variation, or sociolinguistics. Familiarity with localization/internationalization workflows and quality assessment. Experience with LLM-as-judge approaches, rubric design, or bias and fairness evaluation across languages. Fluency or professional proficiency in one or more languages in addition to English.

    Numbers & Facts

    LocationSeattle, WA
    IndustryComputer/IT Services
    Company Size10,000 employees or more
    Year Founded1976
    Websitehttps://www.apple.com/jobs

    About Company

    We bring amazing people together to make amazing things happen.

    We’re a diverse collection of thinkers and doers, continually reimagining what’s possible to help us all do what we love in new ways. The people who work here have reinvented entire industries with the Mac, iPhone, iPad, and Apple Watch, as well as with services, including iTunes, the App Store, Apple Music, and Apple Pay. And the same passion for innovation that goes into our products also applies to our practices — strengthening our commitment to leave the world better than we found it.

    About Apple

    There’s a place here for every kind of brilliant. Everyone here is an innovator, or an innovator-to-be, no matter what your team or your role. So bring your passion, courage, and original thinking and get ready to share it, because every new product, service, or feature we invent is the result of people working together to make each others’ ideas stronger. Innovation at this level depends on people who represent the variety of the human experience and inspire us with their own fresh perspectives. Together, we’ll do amazing work that can make a difference in people’s lives. Including your own. Learn more about working at Apple.

    Similar Jobs