About CaseGuild
CaseGuild builds the industry’s most advanced evidence reasoning platform for complex litigation, helping legal teams investigate massive document sets and surface critical facts with speed and precision.
We’re a fast growing, early-stage Seattle startup building systems at the intersection of information retrieval, machine learning, and large language models, operating at billions of tokens where accuracy and speed matter. You’ll join a small, senior team that ships production systems daily, measures everything, and iterates based on real-world performance.
The Role
This role is for a startup engineer who is data-first, evaluation-driven, and has built production systems.
You’ve spent years building ML or AI systems where success isn’t measured by demos, but by metrics, benchmarks, and real-world performance. You understand that modern LLM pipelines still require datasets, experiments, baselines, and failure analysis, and you enjoy owning that end-to-end.
You Might Thrive in This Role If You:
- Have 5+ years of experience building ML or applied AI systems where accuracy and evaluation mattered
- Have designed and owned experiment frameworks and evaluation pipelines in production
- Are fluent in metrics (precision, recall, F1) and know when each matters
- Have strong foundations in classic ML, NLP, and information retrieval, now applied to LLM-based systems
- Have experience working with multiple LLM providers and models, and don’t treat them as black boxes
- Enjoy end-to-end ownership and pragmatic tradeoffs in a startup environment
- Have a high sense of ownership and agency, with a bias toward getting 1% better every day
- Care deeply about correctness, rigor, and repeatability