Want to know if you’re a fit? Upload your resume and let our AI show you.
Skills
Application Programming Interface (API)unmatched
Benchmarkingunmatched
Customer/Client Researchunmatched
Data Setsunmatched
Public/Media/Press/Analyst Relationsunmatched
Scientific Researchunmatched
Supply Chainunmatched
Writing Skillsunmatched
Description
About the Team & Role
The core of this work is training foundational models for supply chain expertise, not wrapping a generalist frontier model in a better prompt. We think a specialist model, trained deep on the domain, beats a generalist model on the problems that actually matter here, and gets us to a level of inference speed, cost, and reliability that routing every decision through a frontier API simply cant reach.
What Makes You Succeed Here
Wed rather see your checkpoint get quantized, distilled, and forked into someone elses production stack than see it top a leaderboard for a week and disappear. A few things from The Auger Edge show up again and again in the people who do well here.
The instinct to Explore to Evolve looks like this in practice: you dont just call .fit() on a technique, you can derive why it works, and youll rebuild the pipeline from the tokenizer up when the domain demands it, whether thats continued pretraining into a knowledge-intensive vertical or an eval harness that measures something real instead of something convenient.
If youve built evaluation frameworks specifically to catch what standard benchmarks miss, youre already living Own the Fall, Rise Stronger: you treat a bad eval run as signal, not shame, and the loop from "heres where it breaks" to "heres the next checkpoint" is short.
The field dresses complexity up as sophistication constantly, which is exactly what it means to Crush Complexity here: we want the person who ships the clean dataset and the clean eval that a teammate can pick up cold, not the clever bespoke pipeline only its author can operate.
Tech-leading through v1, v2, v3, each release measurably stronger than the last, is what All In, All the Time looks like day to day, and its also why your job isnt done at a passing eval or a merged PR. Its done when youve watched the checkpoint run flawlessly in production, under real load, on real customer data. Ask anyone whos been here a while what that means in practice: the job is never actually done, theres always a v4.
What You Bring
Youve built foundational training data at scale, corpora and not just models, and understand that what goes into a model matters as much as its architecture.
Youve led a project across multiple release cycles, each one measurably better than the last.
Youve designed evaluation methodology that goes beyond standard benchmarks, built specifically to surface what those benchmarks miss.
Youve adapted general purpose models to specialized, knowledge intensive domains and understand what actually transfers versus what has to be rebuilt.
Youve created datasets that other researchers and practitioners now build on.
Youve taken research past the paper and into a real, end to end system that people other than researchers actually use.
Recognition, best paper or outstanding paper or otherwise, has followed the work, but wasnt the point of the work.
Were not hiring for a specific problem or a specific product. Were hiring for a pattern. If you read that list and thought "yes, and also," we want to talk to you.