What Youll DoCustomer Engagement & Technical DeliveryEmbed directly with strategic customers to understand their agent architectures, failure modes, and product goalsDesign and build custom RL environments, evaluation harnesses, and verifiers that capture what "good" looks like for each customers domainArchitect agent scaffolding - tool use, multi-step reasoning, memory, sandbox execution - tailored to customer workflowsConfigure and launch training runs on Lab, iterating on reward functions, rollout strategies, and evaluation criteriaServe as the technical lead for engagements end-to-end: from discovery through deployed, improved modelsPlatform Feedback & EcosystemIdentify repeatable patterns from customer engagements and codify them into reference implementations, templates, and documentationServe as the voice of the customer internally, shaping the roadmap for Lab, verifiers, the Environments Hub, and training infrastructureBuild high-quality examples and "recipes" that make it easy for new customers and open-source contributors to extend the stackContribute to technical content (blog posts, tutorials, case studies) that demonstrates real-world platform usageApplied Research & ExperimentationDevelop novel evaluation methodologies for agentic behavior - multi-step reasoning, tool use correctness, recovery from failure, long-horizon task completionPrototype and iterate on agent harnesses for real-world tasks: code generation, workflow automation, document processing, and moreExperiment with reward design, rubric construction, and environment shaping to improve training signal qualityStay current on the frontier of agentic AI, evals, and post-training methods, and bring that knowledge directly into customer workWhat Were Looking ForDeep hands-on experience building, evaluating, or deploying LLM-based agents in the past 1-2 years - youve seen what breaks in production and know what good evals look likeStrong intuition for evaluation design: you can look at a customers agent and quickly identify what to measure, how to construct a rubric, and where the reward signal is weakWorking understanding of RL and post-training concepts (GRPO, RLHF, reward modeling, SFT) - you dont need to have written a trainer from scratch, but you should understand what the knobs do and why they matterStrong Python skills and comfort with the modern AI stack (Hugging Face, inference engines, agent frameworks)Experience in a customer-facing or consulting-adjacent technical role, or as a technical founder - youre comfortable in a room with a customers engineering team figuring out what to buildExcellent written and verbal communication - you can write a clear environment spec, a compelling case study, and a useful Slack message to a frustrated customerHigh agency and comfort with ambiguity. We recently raised $15mm in funding (total of $20mm raised) led by Founders Fund, with participation from Menlo Ventures and prominent angels including Andrej Karpathy (Eureka AI, Tesla, OpenAI), Tri Dao (Chief Scientific Officer of Together AI), Dylan Patel (SemiAnalysis), Clem Delangue (Huggingface), Emad Mostaque (Stability AI) and many others.