Role Summary
This role leads how client decides which AI models it can trust - the behavioral and model-risk evaluation that lets teams adopt new models fast and with evidence. You'll own model risk end to end: the initial clearance that gates a model into use, the ongoing evaluation that re-checks it on every version change and watches it in production, and the model cards that make the results durable - what was tested, what it's cleared for, and where its risks sit. Your scope is every model the enterprise makes available to use or build on: the models in our GenAI tools (Claude, OpenAI, Glean), the model gardens in Bedrock and other clouds, and models that come through SaaS platforms. You'll build and run the evaluation harness to the bar Safety Standards sets, alongside Cyber - who owns the security testing - and Responsible AI. Success looks like teams getting cleared access to new models quickly because the risk is measured, documented, and current.
If you want to help bring AI to client in a way that's fast, safe, and easy to use, we'd like to hear from you.
Job Responsibilities
- Own initial model-risk clearance and the approved-use rating that gates a model into use.
- Own ongoing evaluation - re-clear on version changes, evaluate against live traffic, and catch drift.
- Author and own the model cards as the durable record of what was tested and what each model is cleared for.
- Build and run the evaluation harness and test suites to the bar Safety Standards sets, with Cyber owning security red-teaming.
- Collaborate with cross-functional teams to ensure effective integration and functionality of AI systems
- Also responsible for other duties/projects as assigned by business management as needed
Education and Work Experience
- Bachelor's Degree plus 3 years of related work experience
- OR advanced degree with 1 year of related work experience
- OR a combination of education and experience deemed equivalent (Required)
- Acceptable areas of study include Computer Science, Engineering, Artificial Intelligence, Data Science, or a Related Field (Preferred)
- 2-4+ years Developing and deploying machine learning models, including fine-tuning and prompt engineering (Preferred)
- 2-4+ years Experience with AI tools and software development using programming languages such as Python or R (Preferred)
- 2-4+ years Collaborating with cross-functional teams to integrate AI solutions into business workflows (Preferred)
Required Knowledge, Skills and Abilities
- Cloud Computing
- Collaboration
- Customer-Focused
- Data Analysis
- Data Management
- Data Modeling
- Model Tuning
- Prompt Engineering
- Model Evaluation
- Model Cards & Documentation
- Model Risk & Clearance
Licenses and Certifications
- Certified Analytics Professional (CAP): Validates expertise in data analytics processes, including the ability to model and interpret complex data. (Preferred)
- Machine Learning Certification: Demonstrates proficiency in machine learning techniques and applications, relevant for developing AI models. (Preferred)
- Certified Data Scientist (CDS): Endorses skills in data science and its application in AI modeling. (Preferred)